Reasoning datasets
Competitive programming tasks and rubric-aligned prompts that evaluate logic depth, planning, and correctness in code.
FEATURES AND BENEFITS
Real-world tasks to evaluate and improve model reasoning across the entire software development lifecycle.
Competitive programming tasks and rubric-aligned prompts that evaluate logic depth, planning, and correctness in code.
Stepwise code generation prompts with human-verified chain-of–thought traces, useful for reward modeling and SFT.
Applied coding problems with multimodal inputs including image, video, audio, and molecular structures.
MCP tools, computer-use, browser-use, and terminal environments for software engineering.