Loading the version log…
V8 ADVANCED PROMPT / BENCHMARK
Measured, not claimed.
Every version. Every run.
Real coding agents work on real repositories twice: once with the developer's raw request, once with v8's brief. Hidden tests decide who solved it. Wins, ties and losses are all here.
- Versions logged
- —
- Studies
- —
- Agent runs
- —
- Tasks with hidden tests
- —
LATEST STUDY
— loading results
VERSION LOG
Each compiler version, and what it did to real agents.
TREND
v8 minus raw, study by study.
Points above the line: more tasks solved with v8's brief than with the raw request. Development-set studies are optimistic by design; held-out studies are the honest ones.
METHOD
How a number gets on this page.
- 01
Independent tasks
Small product repositories, each with a README or docs stating the rules a careful engineer would find. Held-out tasks are written by authors who never see the compiler.
- 02
Hidden tests
The agent sees the repository and the request, never the grader. A task counts as solved only when hidden tests pass, including the traps a partial fix falls into.
- 03
Validated first
Before any agent runs, every task is proven: the hidden tests fail on the untouched repository and pass with a reference solution.
- 04
Real agents, isolated
Claude Code (Sonnet and Haiku) and Codex run headless in a fresh copy of the repository, without the operator's plugins, hooks or settings.
- 05
Raw against v8
The same task twice: the developer's words as typed, and the brief the real v8 CLI writes inside that repository. Two repeats each, paired per task.
- 06
Held-out discipline
A set stays held out until its failures shape a rule. Then it is marked seen, and a new set written by new authors does the confirming.
JOURNEY
What the losses taught us.
Briefs that read as the whole spec
The first real-agent study went 88% raw against 75% with v8. Agents trusted the brief, explored a third less and missed rules written in the repository's own docs. Every brief now says it was written without reading the code.
A 300 ms budget that silently failed
On Windows the local content search rarely finished in time, so file ranking, docs and function names never reached the compiler. At 1200 ms the right file lands in the top eight 74% of the time instead of 38%.
Gains that did not transfer
97% against 88% on the development set became a tie on the first held-out set. We published it, kept the set, and wrote two more.
Rules nobody asked for
"Do not modify any files" inside a fix. "Limit changes to the listed files." A fence around the file that needed the change. Each one cost a task, and each wording now gets removed.
Paraphrase drift
"The info output" became "output info like yank". The request now travels with every brief, exactly as written, and wins wherever the two differ.
LIMITS
What these numbers do not say.
- Small repositories. Most tasks live in 8 to 45 files with explicit docs. Large, messy codebases are where context should matter most, and are not yet covered.
- Strong agents are near the ceiling. Sonnet and Haiku solve 90–100% of held-out tasks from the raw request alone, so there is little room to add solves; what the brief changes most reliably is the number of turns.
- Two repeats. Agents are stochastic. A one-task difference on 24 tasks is noise until it repeats.
- More documentation edits. Agents given v8's brief edit README files more often than agents given the raw request; nearly all out-of-scope lines in held-out studies are documentation. A rule for the next version, to be measured.
- Hidden tests define solved. A valid but different reading of a request fails. That is intended for ambiguous tasks and a bias elsewhere.
Every figure is computed from the study reports by a script, not typed by hand. Raw data · —