Behavior2Code: Challenging coding agents to reconstruct real software from behavior alone
Welcome to the Frontier AI Newsletter, Turing’s weekly recap of the most important AI & AGI developments...
Turing Staff
5 min read
October 9, 2026

Before an agent can write code, it often has to work out what the code is supposed to do. This week we look at that harder first step. We introduce Behavior2Code, a new benchmark from the Turing Frontier Research Lab that asks coding agents to rebuild 60 real command-line programs from behavior alone, with no source code, graded against 6,872 hidden differential tests. Even the strongest of six frontier configurations solved just 16, and 40 programs went unsolved by every model. We also celebrate our Co-founder and CEO Jonathan Siddharth's conversation on The Sourcery podcast. And we explore fresh research on how much of Anthropic's own AI R&D is now led by Claude and how it oversees roughly 30,000 internal agents, a benchmark that drops AI systems into "alien worlds" to test whether they can discover rules rather than recall them, and new tooling for catching reward hacking inside agent evaluation infrastructure. Finally, we're heading to San Francisco for COLM 2026.
What we're delivering
This week, we're spotlighting Behavior2Code, a new benchmark from the Turing Frontier Research Lab that challenges coding agents to reconstruct real software from behavior alone. Each target is an open-source command-line program handed to the agent as an execute-only binary with its user documentation and no source code. The agent has to probe it, infer how it behaves, and rebuild it from scratch.
Here's what we built:
- 60 real programs across four recovery categories: codecs and checksums, formatters and pretty-printers, parsers and interpreters, and utilities and validators, drawn from 10 source languages. Agents work in an offline cleanroom with no way to read, decompile, or debug the binary, and can write their replacement in any language.
- 6,872 hidden differential tests that run the submission and the original on the same input and compare standard output, standard error, exit status, and filesystem effects. About 37% of graded behaviors exercise error handling, which user documentation rarely describes. An instance counts as solved only when every assertion passes.
- Every test proven satisfiable. Targets must be deterministic, every graded behavior must be discoverable without source access, and the original program must score a perfect 1.0 under the same verifier. Only 60 of 115 candidate programs cleared all three gates, and the corpus carries no ignore lists. A separate integrity gate zeroes any submission that wraps, links to, vendors, or reuses the original.
💡 Across 1,080 runs from six frontier configurations, GPT 6 Astra led at 26.7% pass@3 (16 of 60), followed by Claude Opus 5 at 23.3% (14 of 60). 40 programs went unsolved by every model, and none was solved by all six. The last few percent proved hardest: on cmark, Claude Opus 5.5's best attempt passed 752 of 756 assertions and still failed, and exit-status mismatches made up 16.5% to 22.5% of every model's failing assertions, a sign that agents often find a feature without probing its edges.
Explore Turing OTS Datasets
🔬 Need better eval data this quarter?
Turing’s off-the-shelf (OTS) datasets are built for teams that need verifiable, high-signal data where frontier models still break — across multimodal STEM, HLE++ STEM, coding evaluation, and rubric-based reasoning.
Use them for reward modeling, RL post-training, outcome-supervised fine-tuning, frontier benchmarking, and failure-mode analysis.
What we're celebrating
🎉Jonathan Siddharth on The Sourcery
Our Co-founder and CEO Jonathan Siddharth joined The Sourcery podcast to talk about Frontier AI vs Sovereign AI, why enterprises are moving to open models, and what AI safety teams actually do. It's a look at how Turing works alongside frontier labs to advance model capabilities, and how that work carries into real-world AI systems for enterprises.
What we're reading
- Measurements for understanding the pace of AI development inside frontier labs Anthropic proposes three measurements to help the public track AI development inside frontier labs: how much AI R&D is performed by AI itself, how well AI agents' actions are overseen, and how compute is allocated. The headline number: Claude was leading 26% of Anthropic's measured AI R&D work as of August, up from less than 1% in February, while more than 90% involved AI at least as a collaborator.
On oversight, roughly 30,000 agents run at a time on Anthropic's most-used internal platform. A real-time monitor checked over a billion of their decisions in August and blocked 0.002%, about one in 47,000. Anthropic is also candid about the limits: a Claude judge assigned the automation ratings, and the company acknowledges that a judge model may repeat the errors of the model it is checking, so outside verification is needed. - ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds Researchers from Fudan, Tencent Hunyuan, and Tsinghua introduce verifiable "Alien Worlds". Their rules are executable, so every answer can be checked exactly, and they conflict with familiar knowledge, so recall alone cannot solve the tasks. The benchmark ships with AlienCode (31 rules, 70 tasks) and AlienLogic (24 rules, 70 tasks).
Across 10 AI systems, the strongest can acquire and apply unfamiliar rules, but performance varies substantially across trajectories, and continued exploration can stall or reverse earlier gains. In AlienLogic, simply being told the rules (93 to 97%) beats every system's own exploration, a clear signal of how far models remain from genuine discovery. - BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure Agent benchmarks increasingly function as interactive infrastructure: agents observe state, call tools, modify workspaces, and receive rewards from outcome procedures. That makes the benchmark itself an attack surface. To study it, the authors built a human-labeled corpus of 456 adjudicated trajectories from more than 31,000 public agent runs across three benchmarks.
Compared with an agentic hackability scanner baseline, BenchShield raises full-chain recall from 23 to 94% up to 77 to 100% and cuts per-task cost by up to 65%. It also detects reward hacking from infrastructure-side evidence with 96% accuracy, a practical step for any team whose eval scores feed into training decisions.
Where we’ll be
📍 San Francisco, CA United States | 🗓️ Oct 6-9
Turing is at COLM 2026 in San Francisco! Come connect with us and other researchers on the final day, stay tuned for specific Turing-hosted events.
Stay ahead with the Frontier AI Newsletter
Turing is leading the charge in bridging AI research with real-world applications. Subscribe to the Frontier AI Newsletter for weekly insights into breakthroughs, research, and industry shifts that matter.
Related resources
Ready to Optimize Your Model for Real-World Needs?
Partner with Turing to fine-tune, validate, and deploy models that learn continuously.



