Model comparisons, agent setups, cost, and modalities, measured on real tasks. Read the methodology and results, or rerun a study on your own data.
Kimi K3 ran through Moonshot and Fireworks in 30 direct latency requests per provider and 60 planned Figma-to-HTML runs per provider. Paired intervals compare matched designs.
55% faster
Loop and Braintrust MCP + Codex each ran five workflows on identical 200-trace projects. Each tool was scored against a handwritten rubric measuring how well it surfaced insights about the traces and cited supporting evidence so traceability could be verified.
22 condition-neutral checks
39% faster
Four models answered 1,329 questions about specific events from LiveNewsBench. Each model ran without web search and with limited and wide You.com Search API access. GPT-5.6 Terra and Claude Sonnet 5 also used their built-in search tools.
Each skill packages best practices from leading AI teams and researchers into a plain SKILL.md file your coding agent can read. Use them to get started running evals on your data.
Ask it to do what the skill describes, like validating a scorer against human labels or finding hidden failure modes.