New Behavior2Code benchmark: Evaluate how well agents reverse engineer software
A more reliable benchmark for black-box program reconstruction than ProgramBench.

Turing Frontier Research Lab
8 min read
October 6, 2026

We’ve released Behavior2Code, our latest benchmark from the Turing Frontier Research Lab, built around a tough challenge: can an AI agent rebuild a command-line program without seeing its source code? We designed Behavior2Code to address verification problems in existing black-box program reconstruction benchmarks such as ProgramBench, where flawed tests can penalize correct implementations and make it hard to distinguish an agent’s limitations from shortcomings of the evaluation.
The strongest model we tested rebuilds just 19 of the 60 programs, despite receiving three attempts at each. Across all six models, 35 programs remain unsolved. These results expose the gap between writing code that mostly works and delivering a replacement that passes every test.
Check out the Behavior2Code website to see the full methodology and leaderboard, and download the fact sheet.
The agent’s challenge of discovering what to build
A command-line program runs through text commands. Its behavior includes the text it prints, error messages, files it changes, and an exit status that tells other software whether it succeeded. A replacement might print the right answer but report success after an error, causing a data pipeline to continue when it should stop.
Documentation rarely captures every detail. The agent must test competing explanations and boundaries, then turn observations into rules that work beyond the examples it has tried. It must also leave time to build its replacement.
This capability matters for compatibility work, migrations, and rebuilding poorly documented systems.
Addressing issues with existing benchmarks
MirrorCode, developed by Epoch AI and METR, asks agents to rebuild 25 programs. It exposes a subset of test cases and holds out the rest for grading. Visible tests give agents examples and feedback, reducing how much behavior they must discover themselves. Behavior2Code keeps all grading tests hidden, placing more of that discovery work on the agent. MirrorCode’s team curates its suites manually, drawing on project tests, real-world data, and LLM-assisted generation. Expanding that curation is costly. Separately, its large tasks can demand substantial inference budgets: one attempt cost $2,600 and ran for 19 days. Different tasks, budgets, and information prevent a direct difficulty ranking.
ProgramBench covers 200 tasks and more than 248,000 generated tests, but its published packages have verification problems. Some original programs fail their own tests; some tests impose contradictory requirements; others check commands agents cannot discover through the documentation or documented interface. On one task, adding a paragraph about grading criteria produced 266 extra passing tests without changing the model. That makes scores harder to interpret: how much reflects reconstruction ability, and how much reflects missing guidance or flawed tests? Check out our ProgramBench review blog post for more details.
Behavior2Code retains test generation at scale and adds checks before publication. We record expected results from the original program and require it to pass every test under the same grading system agents face, demonstrating that the requirements can all be satisfied together. We check that agents can discover every tested behavior through documentation or execution, and that identical inputs under controlled conditions produce identical results. All 60 programs pass these checks, with no permanently ignored tests. We built Behavior2Code independently and used an entirely different set of programs from ProgramBench.
Making difficulty interpretable
We selected 60 programs from 115 candidates. Each uses a fixed repository version and runs with four CPUs and 8 GB of memory. The programs span ten languages, led by C, Go, and Rust, and four categories: codecs and checksums, formatters, parsers and interpreters, and utilities.

Figure 1. The corpus contains 6,872 graded assertions across 147 independent suites.
The categories expose different difficulties. Codecs and checksums implement published transformations: recovering the rule gives the agent a precise target. Formatters require the agent to infer how a program arranges text. The four earlier models earn mean rewards between 0.44 and 0.55 on formatters, and GPT 6 Astra averages 0.71, yet together the six models solve only one. Handling common inputs can earn partial credit while still missing the layout policy.
We validate the build from a fresh checkout, audit tests that could pass without meaningful work, and combine agentic verification with human review. The review checks whether the tests accept alternative implementations with matching behavior, so agents can choose their own code structure.
The suites cover an average of 91.4% of documented functionality, with a median of 93.7%. This measures commands, options, and other user-facing behavior, rather than source lines. We record 591 exclusions. Approximately 37% of graded behaviors exercise error handling, making invalid inputs a substantial part of the challenge.
These controls strengthen confidence in the scores, but finite tests cannot prove correctness on every possible input or rule out prior familiarity with public programs.
Reward measures progress while passing requires completion
We score partial progress as the fraction of graded checks a submission passes, from zero to one. A failed build or prohibited shortcut receives zero. A check that returns no result counts as a failure.
The checks compare output text, error messages, exit status, and file changes. Their number ranges from 27 to 756 per program, with a median of 72. A heavily tested feature contributes more to partial credit, so that score alone cannot establish completeness.
Our headline metric, pass@3, counts a program as solved if at least one of three attempts passes every check and clears the integrity review. Each program carries equal weight. This separates progress toward a replacement from completion of the tested task.
A sample task
Consider cmark, which reads CommonMark, a specified form of Markdown, and renders it. With 756 checks, it is the most deeply tested program in Behavior2Code. The agent receives documentation and access to the original executable, then must submit its own implementation and build script.
An investigation might vary whitespace or combine emphasis with nested lists to learn how parsing rules interact. These are illustrative experiments, not disclosed grading cases. Comparing its replacement with the original lets the agent identify differences and decide what to test next.
The leaderboard shows why exact matching matters. Claude Opus 5's best cmark attempt passes 749 of 756 checks, approximately 99.1%, but seven differences remain. On fribidi, the four earlier models' best attempts score between 0.95 and 0.98, yet none passes; GPT 6 Astra reached 1.000 twice, but every attempt was rejected for using the system ICU library, leaving Claude Opus 5.5's single clean pass as the only solve.
The final few percent challenge every model across different programs. These totals do not tell us the operational impact of each remaining error, but they show why a high partial score cannot establish that software is ready to replace the original.
The leaderboard leaves substantial headroom
We evaluated six models across 1,080 runs: three attempts per program using the terminus-2 agent scaffold at high reasoning effort. Each attempt had two hours, extended to four for harder tasks.

Figure 2. Blue bars show programs solved out of 60 (pass@3). Each model receives three attempts.
GPT 6 Astra solves 19 programs, or 31.7% pass@3, showing significant improvement over GPT-5.6 Sol, which solved none. In contrast, Claude Opus 5.5 does no better than Claude Opus 5 (both solve 14 programs). Grok 4.6 solves seven and Gemini 3.7 Flash one. Even with three attempts, the strongest model solves fewer than a third of the programs.
Thirty-five programs remain unsolved by every model.
Failure modes where reconstruction breaks down
The trajectories reveal six recurring failure patterns. A run can show several of them, so these explain failures rather than divide them into exclusive categories.
- The implementation exceeds the budget. Agents can reproduce commands and options while leaving the main transformation incomplete. One Claude Opus 5 attempt on StyLua passes 47 of 58 interface checks but fails 143 checks of core behavior. Recognizing the interface does not mean the agent has recovered the underlying algorithm or formatting rules.
- Boundaries remain unexplored. Agents learn what an option does on ordinary input, then guess how it handles invalid input. Exit-status mismatches account for 13.6% to 22.5% of each model’s failing checks. The status is easy to observe, but reproducing it requires trying the conditions that trigger it.
- A submission arrives too late. Claude Opus 5 first mentions compile.sh, the required build script, at a median of 70% of elapsed time; it never appears in 30 runs. One gojq attempt writes 326 KB of source over 142 turns, omits the build script, and scores zero. Claude Opus 5.5 shows the same pattern: 128 of its 180 runs time out and 88 end without a build script. GPT-5.6 Sol writes the script early and always produces a buildable submission, yet solves nothing. GPT 6 Astra also writes its script early and still solves the most programs. These contrasts make the trade-off concrete: investigation must leave time to deliver, while delivery alone does not establish correctness.
- The idea is right but the output differs. A program can perform the main transformation yet print different help text or send an error to the normal output stream. One Claude Opus 5 gojq attempt earns 0.86 reward with 50 interface failures and 11 core-behavior failures. Compatibility includes these details because other software can depend on them.
- Tool interaction consumes the run. The harness rejects 37.4% of Gemini’s turns as malformed, often while it tries to write multiline files. Its 247-turn median corresponds to roughly 155 successfully executed turns. Turn count therefore measures interaction overhead as well as useful work, which helps explain why it does not track the success ranking. Long reasoning has the opposite effect: Claude Opus 5.5's turns run to a median of about 7,000 output tokens, leaving a median of only 81 turns per run.
- The submission relies on shortcuts. We rejected 24 runs for prohibited reuse or hard-coding. On rhash, all three GPT-5.6 Sol attempts store lookup tables of checksum results captured from earlier experiments, instead of implementing the transformation for new inputs. GPT 6 Astra introduces a different shortcut: six runs were rejected for delegating the core transformation to an installed library, such as ICU on fribidi.
Advancing the frontier through measurable capability
The Turing Frontier Research Lab builds frontier-calibrated data and benchmarks for the next generation of superintelligence. We build the benchmarks, training sets, and training recipes that push frontier models forward in the domains where real capability matters.
Behavior2Code supports that mission by measuring whether agents can discover what to build as well as implement it. The results point to a central bottleneck: inferring undocumented behavior from limited observations. The concentration of successful results in codecs and utilities, where published rules offer more guidance, reinforces that reading. Progress requires agents to choose better experiments and carry what they learn through to a complete replacement.
Explore the Behavior2Code benchmark page for the full methodology and results. To get task samples and discuss the benchmark with the Turing Frontier Research Lab, contact us.

Turing Frontier Research Lab
Turing Frontier Research Lab builds the benchmarks, training sets, and training recipes that push frontier models forward in the domains where real capability matters. For more about us, visit https://labs.turing.com/.




