An agent can spot the right vulnerability, write a plausible patch, and still leave the system exploitable, and in security, that last mile is the whole job. This week we look past what an agent recognizes to what it can actually finish. We introduce CyberStrike, a new benchmark from the Turing Frontier Research Lab with 200 expert-authored offensive, defensive, and incident-response tasks, each graded on verifiable outcomes rather than plausible answers, where even the strongest of six frontier configurations cleared under a third of tasks. We also celebrate two milestones: wrapping up our hands-on Claude Code workshop, and welcoming Varsha Udayabhanu as our new CFO. And we explore fresh research on how swarms of AI agents coordinate (and collude) in shared systems, ten open problems in mathematics and theoretical computer science cracked by a frontier model, and how Claude is accelerating protein design and analytical chemistry in the lab.
What we're delivering
This week, we're spotlighting CyberStrike, a new long-horizon cybersecurity benchmark from the Turing Frontier Research Lab, built to measure whether frontier agents can complete real, verifiable security work, not just recognize a problem or make partial progress on it.
Here's what we built:
- 200 expert-authored tasks in sealed, interactive Harbor environments. 120 defensive engineering, 68 offensive security, and 12 digital forensics and incident response (DFIR) tasks, spanning the full cybersecurity lifecycle. To pass, an agent has to produce something a security team could actually use: a patch that closes the root cause without breaking legitimate workflows, an exploit that reproduces in a fresh environment, or an investigation grounded in evidence.
- Outcome-based grading with hidden deterministic verifiers. Each task ships with an agent-facing prompt, a containerized environment, a hidden verifier, and an author-only oracle solution. A run earns credit only when its entire verification contract passes, so partial fixes, unreplayable proofs, and invalid artifacts all score zero. Tasks demand sustained effort, with a median trajectory of roughly 40 agent steps.
- Built and audited for trust. Every task was authored by security experts with 5+ years of experience, put through repeated execution-based review, validated against its oracle solution, and audited by two independent experts. The design is leakage-resistant, so results reflect capability rather than familiarity or infrastructure noise.
💡 The first evaluation covered six frontier-model configurations across 3,600 recorded trials. The leading configuration reached a 31.5% mean per-task pass rate, 72 of the 200 tasks went unsolved by every model, and only two were solved by all six. The dominant failure modes are revealing: incomplete patches on defense, non-replayable proofs on offense, and incomplete evidence or report contracts on DFIR, all cases where a plausible-looking attempt still doesn't hold up.
Explore Turing OTS Datasets
🔬 Need better eval data this quarter?
Turing’s off-the-shelf (OTS) datasets are built for teams that need verifiable, high-signal data where frontier models still break — across multimodal STEM, HLE++ STEM, coding evaluation, and rubric-based reasoning.
Use them for reward modeling, RL post-training, outcome-supervised fine-tuning, frontier benchmarking, and failure-mode analysis.
What we're celebrating
🎉 Turing x Anthropic Claude Code workshop
We recently brought together engineers and researchers for a hands-on Claude Code workshop, a working session on building with Anthropic's agentic coding tool: real workflows, live building, and the practical patterns that make agents useful in a production codebase.
🎉 Welcome, Varsha Udayabhanu, our new CFO
We're thrilled to welcome Varsha Udayabhanu as Turing's Chief Financial Officer. Varsha brings more than 15 years of finance, strategy, and corporate development leadership, including the past five years at high-velocity AI startups. She joins at a defining moment as Turing scales both its Frontier AI and Enterprise AI businesses.
What we're reading
- Patterns and problems in emerging multiagent systems Anthropic ran swarms of Claude agents in shared codebases, markets, and simulated workplaces and documented how individual quirks compound into systemic failures. Coordinating swarms outperformed brute-force parallel search at vulnerability discovery, but agents struggled to build on each other's work, tended to converge on identical choices, and slipped into price collusion even without a direct communication channel.
Most striking was a "turf war" experiment: three agents each told to migrate the same backend to a different language escalated into sabotage, deploying self-replicating scripts and locking each other out, though newer models more often negotiated a truce. The takeaway is that coordination doesn't emerge automatically from more capable or better-aligned individual models, and the mechanisms that make human coordination work still need to be designed for agents. - Ten advances in mathematics and theoretical computer science OpenAI shared ten results in which an internal model resolved or made substantial progress on long-standing open problems spanning high-dimensional geometry, coding theory, arithmetic circuit complexity, group theory, quantum complexity, lattice cryptography, and extremal combinatorics, including several Erdős problems.Human researchers prepared the manuscripts and the arguments were formalized in Lean certificates, with the model's reasoning walkthroughs released alongside.
The piece is also notable for how carefully it frames attribution, arguing that a proof generated by an AI system shouldn't be presented as human-authored, and inviting the mathematical community to engage with and build on the results. - How Claude is accelerating protein design and analytical chemistry Two experiments show frontier models moving into experimental science. In a wet-lab-validated protein design campaign, Claude designed binders against 14 of 15 targets, with 22–35% of individual designs binding successfully versus the 10–15% typical of campaigns today, and some designs matching or beating the best previously published results. On the chemistry side, given only a contract lab's raw NMR and LC-MS files and a two-sentence prompt, Claude returned finished analyses in under 25 minutes and matched the lab's own purity reading almost exactly (96.4% vs. 96.33%).
- The results also underscore the dual-use nature of increasingly autonomous research capabilities, and the safeguards that need to travel with them.
Where we’ll be
🔹 Turing After Hours | AI Researchers in Bellevue & Seattle-Tacoma
📍 Bellevue, WA United States | 🗓️ Sept 10
Join Turing for an evening of networking with researchers and engineers advancing frontier AI across reasoning, coding, multimodality, and beyond.
Stay ahead with AGI Advance
Turing is leading the charge in bridging AI research with real-world applications. Subscribe to AGI Advance for weekly insights into breakthroughs, research, and industry shifts that matter.
Ready to Optimize Your Model for Real-World Needs?
Partner with Turing to fine-tune, validate, and deploy models that learn continuously.


