Delivering 35,000+ agent trajectories to train and align a frontier AI assistant
Built human-driven agent trajectories to support Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF) for a frontier AI research lab's general-purpose AI agent. Turing produced structured trajectory data grounded entirely in real APIs and live services, with no simulated environments or mocked tool responses.
35,000+
SFT trajectories delivered spanning general task execution and privacy-compliant agent behavior.
120,000+
RLHF preference pairs produced, pairing optimal agent actions against suboptimal or policy-violating alternatives.
150+
tools and skills exercised, from local file operations to third-party API calls.

The challenge
The client needed training data that reflected how a general-purpose AI agent actually behaves when a human drives it through realistic, everyday tasks, not synthetic prompts or scripted demonstrations. This required:
- Human experts authentically representing diverse real-world users
- Every task executed against real APIs and live services, with no stubbed endpoints or simulated tool responses
- Coverage across a wide range of life domains, task structures, and escalating tool-use complexity (from structured CLI and API calls up to vision-based computer-use fallback)
- The agent to demonstrate careful, trustworthy handling of sensitive information throughout a task, recognizing when data required extra caution, using appropriately trusted tools and services, and responding appropriately when a task risked exposing or over-retaining that information
- Rigorous quality assurance to catch subtle failures, such as inefficient tool-use paths, unnatural interactions, and privacy-policy violations
The approach
1. Task design across life domains and task structures
Turing designed tasks against a taxonomy of real-world life domains (health, relationships, time management, etc.), each exercised through distinct task structures, including:
- standard multi-turn dialogue
- single-turn complex instructions
- long-context tasks
- cross-session memory recall
- skill discovery and use
- native tool capabilities like scheduling and sub-agent delegation
Tasks were designed across three tiers of increasing interface difficulty, from structured APIs and shell commands, through browser and accessibility-tree interaction, up to vision-based fallback for applications with no structured interface at all.
2. Human-driven sessions under stratified personas
Rather than operating as themselves, experts were assigned stratified human identities defined across various demographic and behavioral axes (occupation, life stage, family structure, geography, economic position, cultural context, digital behavior, and social network). These personas were calibrated against empirical user demographics for computer-use agents, ensuring the expert pool reflected genuinely representative users rather than convenience sampling. Each persona carried structured identity files that governed tone, vocabulary, and behavior consistently across a session.
3. Trajectory collection in isolated, live environments
Experts drove agent sessions inside isolated cloud sandboxes, using only real APIs and live third-party services. Every interaction was captured at step-level granularity: messages, tool calls, tool results, and reasoning traces, structured for direct use in fine-tuning or distillation.
4. Privacy-first engineering and behavior enforcement
Turing implemented prompt-level hooks for privacy rule enforcement, real-time PII redaction from conversation history, and leak-prevention logic during memory compression. Trajectories were built around a data sensitivity taxonomy (from public information up to critical data such as credentials and biometrics) and a three-tier tool trust model (local tools, first-party cloud services, and third-party APIs).
Experts generated trajectories across defined scenarios covering ideal local execution, cloud fallback with consent, third-party fallback and hard-block conditions, and jailbreak/manipulation attempts, ensuring the model saw correct behavior across the full spectrum of privacy-critical situations.
5. RLHF preference pair construction
Alongside SFT trajectories, Turing produced explicit RLHF preference pairs, each pairing a correct, policy-compliant agent action against one or more rejected alternatives representing specific, named failure modes (for example, using a higher-trust-tier tool when a lower-tier one was available, or failing to request authorization before a sensitive data handoff). This gave the client both a supervised signal for correct behavior and a contrastive signal for the specific mistakes to train away from.
6. Three-tier quality assurance
Every trajectory passed through a three-stage QA pipeline before entering a delivery batch:
- Pod lead review: full-coverage human review of persona fidelity, instruction quality, step-label accuracy, and task completion (extended to privacy-compliance checks in later phases)
- Automated assessment: full-coverage automated scoring against task completion, trace coherence, degenerate-loop detection, and schema validity, with an LLM-judge layer producing structured, per-dimension scores
- Human calibration: a manual spot-check sample from each delivery batch, where senior reviewers audited automated decisions and arbitrated edge cases
Each trajectory was scored across dimensions including correctness, completeness, efficiency, and naturality of the human-agent exchange.
Key results
- Delivered more than 35,000 SFT trajectories and 120,000 RLHF preference pairs spanning general task execution and privacy-compliant agent behavior
- Achieved ~90% pass rate on automated LLM-judge review, reinforced by a 20% manual QA spot-check
- Exercised more than 150 distinct tools and skills across local, first-party cloud, and third-party API trust tiers
The outcome
The delivered dataset drove measurable improvements across the majority of the client’s internal evaluation suite, showing gains after the data was integrated into the model's training mix. Reported improvements included substantial gains on public agentic benchmarks such as tau-bench, instruction-following (IFBench) and graduate-level reasoning (GPQA) evaluations, alongside a meaningful reduction in memory-related policy violations.
This gives the client a validated foundation to:
- Train agent models on realistic, human-authenticated task execution across a broad range of domains
- Reinforce correct privacy and data-handling behavior through both supervised examples and contrastive preference pairs
- Continue scaling trajectory collection using an established alter-ego methodology and three-tier QA pipeline
- Track measurable training impact against a broad internal evaluation suite
Need human-driven agent trajectories for SFT or RLHF?
Request a sample of step-annotated agent sessions, including reasoning traces, tool calls, and quality rubric scoring.
Request SampleFAQ
How were trajectories made authentic rather than scripted?
Experts operated under stratified "alter ego" personas calibrated to real user demographics across eight behavioral axes, rather than writing as themselves or following a rigid script.
Were any tool calls or API responses simulated?
No. Every trajectory was generated using real APIs and live services inside isolated cloud sandboxes; no mocked endpoints or simulated tool results were used at any stage.
How was privacy-compliant behavior verified?
Trajectories were built around a defined data-sensitivity taxonomy and tool-trust model. Agent behavior was checked against specific operational rules, including data elicitation, tool-tier selection, consent routing, and refusal behavior, with any violation flagged and gating trajectory acceptance.
What’s the NDA process?
A standard mutual NDA. Turing provides the countersigned agreement within one business day.
How fast can I get a sample?
Within three business days after NDA execution.
Building an agent that needs to reason and act safely across real tools?
Work with Turing to design trajectory collection and RLHF pipelines grounded in authentic human behavior and enforced policy compliance.
AGI Advance Newsletter
Weekly updates on frontier benchmarks, evals, fine-tuning, and agentic workflows read by top labs and AI practitioners.


