Turing

AGI Advance: Weekly AI & AGI Insights (Aug 25, 2026)

Welcome to AGI Advance, Turing’s weekly recap of the most important AI & AGI developments...

Turing Staff
Turing Staff
5 MIN READ28 Aug 2026
LLM training and enhancement
AGI_Advance_Newsletter

A capable agent can finish the job. Whether it reached for the right tool to get there, or asked before handing off sensitive data, is a separate question, and often the one that matters. This week we look past task completion at how an agent handles trust and privacy while it works. We share how Turing delivered 35,000+ human-driven trajectories and 120,000+ RLHF preference pairs for a frontier lab, each pair contrasting a policy-compliant action against a named failure mode like reaching for a higher-trust tool when a lower one would do. We also celebrate our Frontier Fellows program, now open for applications through August 31, and explore fresh research on Meta's case for personal superintelligence, whether LLM judges hold their verdicts under pressure, and a vulnerability that exposes reasoning traces from proprietary APIs.

What we're delivering

This week, we're highlighting how Turing delivered 35,000+ human-driven agent trajectories and 120,000+ RLHF preference pairs for a frontier AI research lab, built to train and align a general-purpose AI agent on realistic task execution and privacy-compliant behavior across real APIs and live services.

Here's what we delivered:

  • 35,000+ SFT trajectories and 120,000+ RLHF preference pairs spanning general task execution and privacy-compliant agent behavior, with each preference pair explicitly contrasting a correct, policy-compliant action against named failure modes such as using a higher-trust tool when a lower-tier one was available, or failing to request authorization before a sensitive data handoff
  • Human-driven sessions under stratified alter-ego personas calibrated across occupation, life stage, geography, economic position, and digital behavior, with every session executed inside isolated cloud sandboxes against real APIs and live third-party services, captured at step-level granularity across 150+ distinct tools and skills spanning local, first-party cloud, and third-party API trust tiers
  • Privacy-first behavior enforced throughout, built around a defined data sensitivity taxonomy and three-tier tool trust model, with trajectories covering ideal local execution, cloud fallback with consent, hard-block conditions, and jailbreak and manipulation attempts

💡 The delivered dataset drove measurable improvements across the client's internal evaluation suite, including gains on tau-bench, IFBench, and GPQA alongside a meaningful reduction in memory-related policy violations.

Request sample data

Explore Turing OTS Datasets

🔬 Need better eval data this quarter?

Turing’s off-the-shelf (OTS) datasets are built for teams that need verifiable, high-signal data where frontier models still break — across multimodal STEM, HLE++ STEM, coding evaluation, and rubric-based reasoning.

Use them for reward modeling, RL post-training, outcome-supervised fine-tuning, frontier benchmarking, and failure-mode analysis.

Request sample data

What we're celebrating

🎉 Become a Turing Frontier Fellow

We're bringing together exceptional leaders across AI, engineering, finance, law, medicine, STEM, and more who want to help shape how AI transforms their fields.

Over four months, Fellows receive access to frontier AI tools, hands-on workshops with teams from Anthropic and OpenAI, fireside chats with leaders, including Aaron Levie (CEO, Box), Tanay Kothari (CEO, Wispr Flow), Carina Hong (CEO, Axiom Math), Kanjun Qiu (CEO, Imbue), Nicolas Bonamy (Codex, OpenAI), and Mary M. Liu (CFO, Runway), plus a global community of peers across San Francisco, New York, Boston, London, and Bengaluru.

Apply by August 31

What we're reading

  • The Future Is for Everyone: The Path to a Positive AI Future Meta outlines a vision for personal superintelligence built around three principles: individual empowerment, invention, and a balance of power that favors people. Rather than concentrating advanced AI within a few companies or governments, Meta argues that superintelligence should be broadly accessible, giving individuals personal agents, creation tools, tutors, scientific capabilities, and resources to build businesses and pursue their own goals.

    The piece argues that invention, rather than automation, should be AI’s primary purpose. Meta also frames broad AI access as a safety mechanism, proposing that distributed superintelligence can create checks and balances while addressing risks including job displacement, cybersecurity, biological misuse, surveillance, and loss of human control.
  • Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence Meta researchers introduce the Wiggle Framework, a stress test for measuring whether LLM judges maintain their verdicts when re-prompted, challenged, or subjected to sustained pressure. Across 9 frontier models and 14 judging tasks, every model showed substantial instability, flipping verdicts 25–71% under static pushback and 62–91% when challenged by an adaptive LLM persuader.

    More importantly, successful persuasion usually made judgments less accurate. Across tasks with ground truth, 70% of verdict changes under adaptive persuasion were corrupting, moving away from the correct label. The study also finds that agreement among a baseline jury of models is the strongest low-cost predictor of which judgments are likely to be unstable.
  • Stealing Reasoning Traces from Proprietary LLM APIs
    Researchers identify a vulnerability in how major AI providers protect hidden reasoning. Encrypted reasoning blocks can be portable across sessions, users, and even models within the same provider, allowing a weaker model to act as a decoder for reasoning produced by a more capable model. The researchers demonstrate the issue across Anthropic, OpenAI, and Google.

    The vulnerability creates four major risks: stealing proprietary reasoning for model distillation, extracting sensitive information, exposing harmful information hidden from final answers, and enabling invisible prompt injections. In 315,320 decoded reasoning blocks collected from public traces, the researchers found 367 distinct PII artifacts and 182 credentials.


Where we’ll be

🔹 Turing After Hours | AI Researchers in Bellevue & Seattle-Tacoma

📍 Bellevue, WA United States |  🗓️ Sept 10

Join Turing for an evening of networking with researchers and engineers advancing frontier AI across reasoning, coding, multimodality, and beyond.

Stay ahead with AGI Advance

Turing is leading the charge in bridging AI research with real-world applications. Subscribe to AGI Advance for weekly insights into breakthroughs, research, and industry shifts that matter.

[Subscribe & Read More]

You might also like

Ready to Optimize Your Model for Real-World Needs?

Partner with Turing to fine-tune, validate, and deploy models that learn continuously.

Optimize Continuously