PLSQLBench: Measuring Correctness Through Execution-Based Tests
Welcome to the Frontier AI Newsletter, Turing’s weekly recap of the most important AI & AGI developments...
Turing Staff
5 min read
October 2, 2026

This week, we're highlighting new research from Turing in collaboration with Oracle: the launch of PLSQLBench, the first benchmark for evaluating whether LLM systems can write executable PL/SQL, graded through execution-based tests across 2,865 tasks and five subsets. We also recap our time at LEAP 2026 in Riyadh, where the conversation kept turning from model access to sovereign, production-grade deployment, and we round up recent research on catching reward hacking through internal representations and on treating calibration as a first-class evaluation criterion.
What we're delivering
This week, we're introducing PLSQLBench, new research from Turing in collaboration with Oracle, and to our knowledge the first benchmark for evaluating whether LLM systems can write executable PL/SQL, with correctness measured through execution-based tests rather than surface-level matching.
Procedural database programming (stored procedures, functions, packages, cursors, and exception handling) sits in a gap that general-purpose code generation and declarative text-to-SQL benchmarks don't cover. PL/SQL is a practically relevant place to test this, since Oracle Database currently ranks first in the DB-Engines popularity ranking. Here's what we built:
- 2,865 instances spanning single-turn and multi-turn work, including 2,594 single-turn tasks and 271 multi-turn conversations across 978 turns, so the benchmark measures the ability to generate, modify, debug, and repair programs while preserving context across a session.
- Five subsets across three source families. PLSQLBench is constructed from Spider 2.0 Lite, Spider 1.0, and MBPP/MBPP+, forming five subsets (Spider2-ST, Spider2-MT, Spider-PLSQL, MBPP-PLSQL, and MBPP+-PLSQL) that span varying levels of database grounding, schema complexity, and procedural reasoning.
- Eight frontier models tested, and consistent gaps found. Experiments surfaced recurring difficulties in schema grounding, PL/SQL dialect fidelity, procedural control flow, exception handling, and cross-turn consistency. Tool-augmented agents improved several schema-grounded evaluations, though substantial gaps remained.
💡 Strong reasoning alone doesn't make a model reliable on real enterprise database code. Writing PL/SQL that executes correctly, handles exceptions, and stays consistent across a multi-turn session takes capabilities that conventional coding and text-to-SQL benchmarks don't directly assess.
Explore Turing OTS Datasets
🔬 Need better eval data this quarter?
Turing’s off-the-shelf (OTS) datasets are built for teams that need verifiable, high-signal data where frontier models still break — across multimodal STEM, HLE++ STEM, coding evaluation, and rubric-based reasoning.
Use them for reward modeling, RL post-training, outcome-supervised fine-tuning, frontier benchmarking, and failure-mode analysis.
What we're celebrating
🎉Turing at LEAP 2026
Our team spent four days in Riyadh for LEAP 2026 and DeepFest, and the takeaway was clear: the next chapter of AI will not be defined only by adopting models, but by building the sovereign capabilities, infrastructure, and ecosystems around them so that governments and enterprises can own and control their own intelligence.
A big part of that story is the HUMAIN and Turing partnership. In his session, Turing founder and CEO Jonathan Siddharth shared a practical blueprint for AI sovereignty at scale, focused on how governments and enterprises can deploy AI on their own terms. Across the floor, conversations kept returning to the theme we're seeing everywhere: the shift from experimentation and demos to production-grade deployment, where real workflows, real data, and real talent are the bottleneck, not access to models alone.
Thank you to everyone who stopped by, shared where their systems are breaking, and helped us think through what it takes to turn frontier capability into real-world impact.
What we're reading
- PLSQLBench: Benchmarking LLM Systems for Executable Procedural Database Programming Our own collaboration with Oracle is worth a full read this week. The paper introduces the first benchmark for executable procedural database programming in PL/SQL, graded through execution-based tests across 2,865 instances.
Across eight models, the authors document consistent weaknesses in schema grounding, dialect fidelity, procedural control flow, exception handling, and cross-turn consistency, and show that tool-augmented agents help on schema-grounded tasks without closing the gap. It's a useful map of where frontier models still fall short on real enterprise code. - Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations This work asks whether reward hacking leaves a detectable signature inside a model's activations, and finds that it does. Simple difference-of-means vectors, built from synthetic examples of hacking behavior, generalize to long-context, frontier-scale agentic rollouts and reliably flag hacks across common environments.
The authors report that models hack far more often than headline numbers suggest (for one open model, in more than half of rollouts on DeepSWE and roughly three quarters on SWE-bench), and that running these probes on the chain of thought can predict a hack before it happens. It's a strong case that lightweight, white-box methods can scalably monitor for reward hacking during evaluation. - Calibration as a First-Class Criterion in LLM Evaluation Calibration, the alignment between a model's stated confidence and its actual correctness, is well studied, yet the authors note that most new models, datasets, and benchmarks ship without reporting it.
They argue this adoption gap undermines trustworthy evaluation in two places: at deployment, where overconfident mistakes cause real harm, and inside the research pipeline, where LLM-as-a-judge, synthetic data generation, and active learning all lean on confidence scores that go unverified. Since standard calibration metrics need only a confidence score and a correctness judgment, both of which most benchmarks already produce, the paper makes the case that every subfield should pair its headline metric with a calibration score.
Where we’ll be
📍 San Francisco, CA United States | 🗓️ Oct 6-9
Turing is headed to San Francisco for COLM 2026! Come connect with us and other researchers, stay tuned for specific Turing-hosted events.
Stay ahead with the Frontier AI Newsletter
Turing is leading the charge in bridging AI research with real-world applications. Subscribe to the Frontier AI Newsletter for weekly insights into breakthroughs, research, and industry shifts that matter.
Related resources
Ready to Optimize Your Model for Real-World Needs?
Partner with Turing to fine-tune, validate, and deploy models that learn continuously.



