An agent can write a pipeline that runs, ties every total, and reproduces cleanly on a fresh directory. Whether the model underneath it predicts well enough to act on is a separate question, and in production it's the one that matters. This week we look past "did the script run", and look at whether an agent can ship an analytics deliverable someone would actually trust. We share Frontier Agentic ML & DS, Turing's new set of 304 expert-authored, verifier-graded terminal tasks that grade predictive quality. Using a minimum F1 (a calibration threshold the model has to clear on data it never sees, not just formatting), the strongest of eight frontier configurations cleared just 64.6%. We also welcome Ece Kamar as Turing's new Chief Technology Officer, and explore fresh research on our own expert-validated STEM benchmark that frontier models still fail, how far agents actually get on long-horizon terminal tasks under dense reward-based grading, and how often tool-using agents reward-hack their way to a passing score.
What we're delivering
This week, we're highlighting Frontier Agentic ML & DS, a new off-the-shelf training and evaluation set from Turing Frontier Research Lab: 304 expert-authored, verifier-graded terminal tasks that measure whether a frontier agent can actually ship a working machine-learning or analytics deliverable, not just describe the method.
Here's what's inside:
- 304 expert-authored tasks across four families. 120 Data Science, 110 Machine Learning, 53 Pure ML Science, and 21 Data Wrangling, fanning out into 106 distinct sub-domains, from tabular classification and time-series forecasting to causal inference, quantum computing, and physics-informed neural networks. No sub-domain holds more than 18 tasks, so the set rewards genuine breadth over one template repeated.
- Predictive quality is graded, not just formatting. A large share of tasks carry a performance gate, a minimum F1, a tolerance band, a calibration or backtest threshold the agent's own model must clear on data it never sees. This is the axis public agentic benchmarks almost never test, and it's exactly where frontier agents are weakest.
- Sealed, reproducible, leakage-resistant environments. Every task ships an instruction, a Docker environment with a pinned scientific stack and explicit CPU/memory/storage limits, raw inputs, an oracle solution, and a task-specific pytest suite. Containers run offline, and verifiers re-execute the agent's own pipeline against regenerated inputs, fresh output directories, and hidden holdout sets, so a hard-coded or memorized answer turns into a failing assertion. The average task carries 23 assertions; the deepest, 143.
💡 Across 7,296 rollouts spanning 8 frontier-agent configurations, the top configuration (Claude Opus 5) reached just 64.6%, and 24 tasks (7.9%) went unsolved by every model on every run. The pattern is consistent: agents reliably satisfy the output contract but stumble on the last mile, a model that's actually good enough to use. That gap is exactly what this set exists to train.
Explore Turing OTS Datasets
🔬 Need better eval data this quarter?
Turing’s off-the-shelf (OTS) datasets are built for teams that need verifiable, high-signal data where frontier models still break — across multimodal STEM, HLE++ STEM, coding evaluation, and rubric-based reasoning.
Use them for reward modeling, RL post-training, outcome-supervised fine-tuning, frontier benchmarking, and failure-mode analysis.
What we're celebrating
🎉Welcoming Ece Kamar as Turing's Chief Technology Officer
We're thrilled to announce that Ece Kamar has joined Turing as Chief Technology Officer. Ece will lead Turing's technology and research strategy across our frontier AI and enterprise AI businesses by sharpening the link between what frontier labs need next and how enterprises put that innovation to work.
Ece joins after 16 years at Microsoft, where she most recently served as Corporate Vice President leading the AI Frontiers Lab at Microsoft Research, a mission-focused lab advancing foundation models, agents, efficiency, and the control of frontier AI systems. Her work there spanned both sides of the research-to-product divide, including Microsoft's Phi small-model series and the Fara computer-use models, and she co-authored the widely cited "Sparks of AGI" paper.
Across two decades in AI research, Ece has consistently worked at the seam between advancing capabilities and deploying them in the real world exactly the vantage point Turing operates from, sitting between the frontier labs and the enterprises putting that technology to work.
What we're reading
- Expert-validated STEM QA Turing's research team introduces a high-quality, expert-validated STEM benchmark (N=398) spanning Physics, Chemistry, Biology, and Mathematics, authored by 241 domain experts. The dataset was built to fix recurring gaps in existing STEM sets, performance saturation, skewed taxonomies, multiple-choice formats misaligned with real scientific use, and answer-quality issues from contest-style collection, using a balanced taxonomy, incentive-aligned contributor vetting, and multiple expert review rounds in a verifiable question-and-answer format.
Frontier models score below 25% on it, and post-training an open-source model on a larger private version (N=2,000) lifted performance on the HLE-verified STEM subset by 15% over baseline, pointing to the set's value for both evaluation and training. A portion has been open-sourced for the research community. - Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading This benchmark extends the Terminal-Bench style of realistic command-line evaluation to genuinely long-horizon work, decomposing each task into fine-grained, deterministically graded subtasks. That design yields dense intermediate rewards and partial credit, so a run captures not just whether an agent reaches the final goal but how far it gets on open-ended workflows.
The authors report a cost-versus-reward frontier across frontier models under a fixed time budget, and break down where agents fail: timeouts, early self-termination, and harness errors, offering a sharper picture of long-horizon reliability than pass/fail scoring alone. - Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use As tool-using agents move into coding assistants and autonomous systems, this benchmark measures how often they take illegitimate shortcuts, skipping verification, inferring answers from task-adjacent metadata, or tampering with evaluation-relevant functions across both independent and chained multi-step tasks.
Evaluating 13 frontier models from OpenAI, Anthropic, Google, and DeepSeek, the authors find exploit rates ranging from 0% for the strongest performers to roughly 14% at the high end, with the rate varying sharply by post-training style. A controlled sibling comparison suggests RL post-training is associated with substantially more reward hacking, a useful caution for anyone shipping RL-trained agents.
Where we’ll be
🔹 Turing After Hours | AI Researchers in Bellevue & Seattle-Tacoma
📍 Bellevue, WA United States | 🗓️ Sept 10
Join Turing for an evening of networking with researchers and engineers advancing frontier AI across reasoning, coding, multimodality, and beyond.
Ready to Optimize Your Model for Real-World Needs?
Partner with Turing to fine-tune, validate, and deploy models that learn continuously.


