opensources.devGet v8

V8 ADVANCED PROMPT / BENCHMARK

Measured, not claimed.
Every version. Every run.

Real coding agents work on real repositories twice: once with the developer's raw request, once with v8's brief. Hidden tests decide who solved it. Wins, ties and losses are all here.

Versions logged
—
Studies
—
Agent runs
—
Tasks with hidden tests
—

LATEST STUDY

— loading results

VERSION LOG

Each compiler version, and what it did to real agents.

Loading the version log…

TREND

v8 minus raw, study by study.

Points above the line: more tasks solved with v8's brief than with the raw request. Development-set studies are optimistic by design; held-out studies are the honest ones.

METHOD

How a number gets on this page.

  1. 01

    Independent tasks

    Small product repositories, each with a README or docs stating the rules a careful engineer would find. Held-out tasks are written by authors who never see the compiler.

  2. 02

    Hidden tests

    The agent sees the repository and the request, never the grader. A task counts as solved only when hidden tests pass, including the traps a partial fix falls into.

  3. 03

    Validated first

    Before any agent runs, every task is proven: the hidden tests fail on the untouched repository and pass with a reference solution.

  4. 04

    Real agents, isolated

    Claude Code (Sonnet and Haiku) and Codex run headless in a fresh copy of the repository, without the operator's plugins, hooks or settings.

  5. 05

    Raw against v8

    The same task twice: the developer's words as typed, and the brief the real v8 CLI writes inside that repository. Two repeats each, paired per task.

  6. 06

    Held-out discipline

    A set stays held out until its failures shape a rule. Then it is marked seen, and a new set written by new authors does the confirming.

JOURNEY

What the losses taught us.

  1. Briefs that read as the whole spec

    The first real-agent study went 88% raw against 75% with v8. Agents trusted the brief, explored a third less and missed rules written in the repository's own docs. Every brief now says it was written without reading the code.

  2. A 300 ms budget that silently failed

    On Windows the local content search rarely finished in time, so file ranking, docs and function names never reached the compiler. At 1200 ms the right file lands in the top eight 74% of the time instead of 38%.

  3. Gains that did not transfer

    97% against 88% on the development set became a tie on the first held-out set. We published it, kept the set, and wrote two more.

  4. Rules nobody asked for

    "Do not modify any files" inside a fix. "Limit changes to the listed files." A fence around the file that needed the change. Each one cost a task, and each wording now gets removed.

  5. Paraphrase drift

    "The info output" became "output info like yank". The request now travels with every brief, exactly as written, and wins wherever the two differ.

LIMITS

What these numbers do not say.

Every figure is computed from the study reports by a script, not typed by hand. Raw data · —