Turing

CompanyBench: Evaluating Frontier AI Agents On Real Operational Knowledge Tasks

We identify systematic mistakes incurred when running agents on operational tasks in real enterprise settings.

Andrew Ma
3 MIN READ05 Aug 2026
AI/ML

Enterprises exist to coordinate and channel human efforts towards economically useful goals. Much like any human institution, enterprise data is messy, contextual, and scattered. Performing tasks correctly in such settings requires common-sense knowledge, correct operational depth, and deep technical expertise. These are often the implicit skills for which we hire human employees, and what we hope AI systems will be able to handle in the near future.

We introduce CompanyBench, a dataset that exists to address these issues. 

CompanyBench drops frontier agents into a real fintech company, sourced from historical operational data: a production-scale data warehouse, policy documents, Slack, Jira, email, and customer support systems. Agents solve tasks on a system that human employees have worked on for years, and which have impacted the finances of thousands of humans. Frontier agents attempt to do critical business tasks, but still come up short.

What makes CompanyBench different

Grounded in a real company

CompanyBench is built on five years of one fintech firm’s operational fabric: hundreds of examples of work that the firm’s employees have done over the course of its history while servicing customers - from quick the tasks of an analyst to strategic initiatives drawn out over the course of many months at the executive level. Every task is grounded in an actual historical work assignment at the company, like a Jira ticket or a Slack message. We require models to answer in a realistic manner: post a summary to Slack, file a Jira ticket, issue actionable decisions.

All data in CompanyBench is private and has never been made publicly available. All personally identifiable information is thoroughly and consistently scrubbed from the data before integration with our backend services. We obtained all data in the company’s lifecycle of five years, including tens of thousands of files, 148M database rows, and millions of Slack messages and Jira tickets.

Results: Frontier models still fall short

We release results on 64 hard tasks attempted ten times each. In this setting, the strongest model, GPT-5.5, clears every required check 44.4% of the time. Claude Opus 4.8 came in second at 35.9% followed by Gemini 3.1 Pro at 18.3%, and Gemini 3.5 Flash rarely performed all the tasks correctly, usually neglecting to perform required write actions to Slack or Jira.

Systematic failure patterns

We identified four common modes across model runs, suffered by Claude, GPT, and Gemini models.

Neglecting to read the docs. Instead of searching for and reading the correct documentation on Slack or Confluence, the model makes assumptions about data sources by guessing the relevant tables, columns, and analysis logic. 

Chasing non-existent documents. When it cannot immediately ground a value, the model over-invests its search budget and hunts for authoritative documents that do not exist, then terminates with an empty or guessed answer.

Unrequested search filters. The model "improves" a deliberately simple SQL request by adding plausible extra filters, such as transfer-shape constraints, time-window anchors, or status codes that silently exclude legitimate rows.

Dropped write actions. Models repeatedly treat the numeric reply as the finish line and skip the mandated Slack and Jira deliverables, even in runs where the underlying analysis was correct.

Conclusion

We believe that the future of work is agentic, and models will be increasingly trusted with high stakes workflows. However frontier agents clearly still have a long way to go before being reliable autonomous partners for complex distributed knowledge work tasks. Simple errors, such as common sense failure cases and a lack of thoroughness handicaps model results in these complex settings, where models have to be deeply aware of contextual nuance. We introduce CompanyBench, a benchmark based on real enterprises, that addresses these gaps.

Explore CompanyBench

The full leaderboard, a task explorer with complete reasoning trajectories, and an overview of the dataset are all available on the dashboard. CompanyBench is a living benchmark, and we will keep evaluating new models as they arrive.

Andrew Ma

Andrew Ma is Head of Research, Knowledge Work at Turing. Before Turing, Andrew worked on recommender systems and search backend at xAI and pre-training data at EssentialA

Kyle Waters

Kyle is an SPL at Turing, where he leads complex data projects for frontier AI labs, turning ambitious research goals into scalable operations and high-quality datasets.

Explore CompanyBench

Find more details about CompanyBench including the full leaderboard, tasks, and dataset profile.

Explore the benchmark

Ready for frontier model data?

Request dataset samples today and accelerate your research.

Request dataset samples