Turing

Introducing CEO Bench: Evaluating frontier agents on real company work

CEO Bench evaluates whether frontier agents can complete long-horizon financial and operational work inside a complete, expert-authored fictional company rather than merely crunch numbers in isolation.

Jeffrey Weichsel
Jeffrey Weichsel
7 MIN READ07 Sep 2026
DATASET ANNOUNCEMENT

Today, Turing is introducing CEO Bench, an evaluation suite built on a complete, expert-authored fictional company, the first benchmark produced on SVC (Simulated Virtual Company), our platform for constructing entire companies as evaluation substrates.

Our first virtual company, Cirrus Sleep, Inc., is a simulated direct-to-consumer sleep brand with a complete first fiscal year: 1100+ files covering accounting, finance, contracts, HR and payroll, board materials, operations, marketing, customer support, and a year of internal Slack. CEO Bench currently includes 500+ human-authored tasks, with 40 published and scored in this release. These are department-level and executive tasks, all written by finance and operations practitioners. 

CEO Bench is available now. To request samples, contact us.

The challenge: measuring judgment, not string matches

AI research has a reliability problem that it keeps measuring but has yet to solve: models that saturate public benchmarks still stumble in deployment on tasks that any competent analyst handles routinely. That’s because existing evaluations of long-horizon knowledge work have several major shortcomings.

Contamination. Benchmarks leak. Anything public long enough ends up in training data, and scores stop meaning what they meant. You cannot prove a public company's 10-K is not in a training corpus. Score inflation from familiarity is undetectable and unfalsifiable.

Benchmarks fragment. A question with a single gold answer cannot represent work whose output is a forty-tab model with judgment calls on every tab.

Ground truth under ambiguity. Real companies legitimately carry two "correct" values for the same quantity. The books report one order count; the raw platform export reports another. An evaluation with one gold answer either tests the wrong skill or marks correct reasoning wrong.

Benchmarks decontextualize. Real work happens inside an organization, where the hard part is rarely the calculation and usually the archaeology: which of these five conflicting numbers is authoritative, who produced this file, what did the board actually approve? 

Grading judgment. An executive deliverable can be factually right and professionally useless. Rubrics need to score sourcing, reconciliation, and disclosure decisions, and stay usable by a grader who never saw the company.

Scale without decay. Hundreds of tasks over one document corpus means every task added is a chance to leak an answer into someone else's inputs.

Introducing CEO Bench

Existing benchmarks like GDPval grade realistic professional deliverables, but each task stands alone; CEO Bench's tasks live inside one coherent company, where the hard part is deciding which of the company's own numbers to trust.

CEO Bench's answer is to make the company the benchmark substrate. Cirrus Sleep was authored by domain experts as a coherent economic object: the trial balance foots to the bank statements, the payroll register ties to the HR roster, the Shopify export reconciles to booked revenue through a governed adjustment schedule. And deliberately, some things do not reconcile, because at real companies they don't.

Of the 500+ tasks in CEO Bench, we evaluated six frontier models on 40 department-level (L2) tasks. You can find more details on the evaluation results and the larger corpus at the CEO Bench website.

Figure 1. The substrate. Cirrus Sleep, Inc. made up of 1100+ files across a complete fiscal year.

Tasks are organized around two levels. Executive tasks (L3s) pose company-wide deliverables, while department tasks (L2s) cover the professional workstreams beneath them. L2s and L3s are separate tasks, and at evaluation time no task receives another task's outputs. This release publishes and scores 40 L2 tasks. In addition, the release includes 70+ L3 tasks published as catalog entries, with scored runs to follow.

From gold answers to graded judgment

The design of CEO Bench directly addresses the shortcomings of existing benchmarks.

The dataset is contamination-proof by construction, not by detection. Nothing about Cirrus Sleep exists outside the dataset, and the methodology regenerates: new fiscal years, new companies, new verticals can be authored on demand.

Ambiguity is graded, not averaged away. Where the company carries dual bases (books versus operational exports), rubrics specify which basis each criterion grades on. The skill under test is knowing which number to use, which is precisely the skill that separates analysts from autocomplete.

Judgment is graded through multi-criterion rubrics with explicit expected values, written by the same experts who built the goldens.

Scale is protected by packaging discipline. Tasks ship self-contained, inputs are files (never any other task’s outputs), and every package is swept for leakage before release.

Quality and methodology

Every task survives two independent adversaries before it ships.

Figure 2. The authoring and review loop. The same four-step build runs at every node, followed by two independent adversaries: an ensemble of LLM judges, then human expert adjudication in which disputes are ruled on rather than averaged.

First, automated QC: ensembles of LLM judges attack each package across a dozen dimensions, including instruction/input consistency, self-containedness, rubric specificity, input authenticity, and solution alignment. Second, human adjudication: domain experts can dispute judge findings, and the disputes are ruled on rather than averaged.

The disagreement between those layers turned out to be the most informative signal we have  about task quality and about the judges. Automated judges confidently flag defects that turn out to be artifacts of their own context limits. Expert-built goldens occasionally fail their own rubrics. Both facts should unsettle anyone consuming benchmark numbers uncritically, and both are exactly what a two-tier loop catches.

SVC is the platform, and Cirrus Sleep is the pilot. The expensive artifact is not any single task, it’s the company, and SVC is the platform for building companies. Cirrus Sleep Year 1 is the pilot: a validated economic object that already powers CEO Bench's FP&A and accounting suite, with treasury, diligence/quality of earnings, S&OP, and technology-planning families in build. From here the platform scales on two axes. In depth: Cirrus Sleep extends to four to five fiscal years of operating history including growth, downturns, financings, and leadership changes, so tasks can demand genuinely long-horizon reasoning across a company's life. In breadth: SVC builds new companies in new industries, each expert-authored to the same standard, each contamination-proof by construction. One platform, many companies and many benchmarks to come.

Evaluation results

Each of the 40 tasks in our current published evaluation includes one run from each of six frontier models, scored against its expert-authored rubric.

Figure 3. Aggregate and individual scores by task

The pattern that survives any particular number: department-level work is nowhere near solved. GPT-5.6 Terra, the top model, averages 41.3, which means it fails more of the rubric than it passes. On 20 of the 40 tasks, even the strongest result across the six models is below 50.

The scores are also uneven. A model can score 83.3 on one task and zero on another, and that is not a rounding artifact: a zero means no criterion in the rubric passed, which includes runs that produced nothing readable at all. Of the 240 runs in this release, 96 came in under 25. On five of the 40 tasks, the best score across all six models was below 25. Only two tasks saw any model break 80. So this is not a case of models being uniformly mediocre. They are competent on some pieces of the work and completely lost on others, and from the outside there is no obvious way to tell which is which in advance.

Cost tells a similar story. Across the six models, average cost per run spans a 9.6x range, from $0.95 to $9.09, while scores span 18.4 points, from 22.9 to 41.3. Higher cost does not reliably mean a higher score. In the GPT and Gemini pairs, the more expensive model scores lower. Claude is the exception: Opus 5 costs slightly more than Sonnet 5 and scores 8.4 points higher. The cheapest model, Gemini 3.1 Pro, still beats Gemini 3.7 Flash, which costs more than twice as much.

Figure 4. Score against cost per run pareto frontier.

Failure-mode analysis

The instructive failures are not arithmetic. Across the scored runs, four modes recur.

Fabricated precision. Asked to reconcile a governed revenue adjustment held at monthly grain, one frontier model invented per-order allocations, presenting unallocated adjustments as if resolved to individual transactions, complete with plausible identifiers. Correct-looking, wrong, and dangerous in exactly the way that matters professionally.

Interpolated data. Where monthly customer figures did not exist, a model manufactured a smooth monthly series that netted to the correct annual total — an error invisible to any aggregate check and caught only by rubric criteria that grade sourcing.

Missed governance traps. Several tasks embed a deliberate provenance anomaly: a bridge whose source orders do not exist in the operational export. Strong models flag it for review; weaker runs either miss it entirely or, worse, silently "fix" it.

Confident basis-mixing. Blending the accounting record (the “books”) basis and the operational basis in one analysis without disclosure, which is the single most human-like failure in the set, and the one practitioners flag first.

These modes matter because they are trust failures, not capability failures. Each produces a deliverable a busy executive might accept.

Next steps

Sample tasks, including prompts and expert goldens across the FP&A and accounting suite, are available on request.

  • Contact us for sample tasks and more information on the research. 
  • See the full task inventory, rubric structure, and delivery formats (including Harbor) on the CEO Bench detail page.

Jeffrey Weichsel

Jeffrey Weichsel is Head of Frontier Data: Enterprise Workflows at Turing, where he leads the Simulated Virtual Company (SVC) effort; building complete, expert-authored fictional companies, paired with expert-built task hierarchies spanning staff-level to C-suite work, to evaluate and train AI models on long-horizon enterprise knowledge work. He is interested in how AI agents can extend human thinking and how expert-grade data shapes model capabilities.

Contact us

Learn more about our research and datasets.

Request samples