v0.36.2

This scoreboard answers one question: how reliably does a model operate the Cast harness? It is not a general intelligence ranking. A model is certified once it clears 80% of the full behavior suite across exactly three fresh attempts per case, where every attempt on a case must agree (a case that only passes sometimes doesn't count). Core/Chain break the same score down by single-turn tool contracts vs. multi-turn stateful workflows (see Behavior Evals). Tokens are per-attempt averages; the time and turns columns are avg/median/p75/p95/p99 over every individual attempt. See Eval Methodology for the full methodology.

ModelReasoningScorePassedCertifiedCoreChainTime (avg/p50/p75/p95/p99)Turns (avg/p50/p75/p95/p99)Input tokens (avg/p50/p75/p95/p99)Output tokens (avg/p50/p75/p95/p99)Provider URLLast updated
openai/gpt-5.6-lunamedium92.1%35/38✓ certified92% (11/12)92% (24/26)11.4s / 8.5s / 11.7s / 30.1s / 77.1s3.5 / 3.0 / 4.0 / 10.0 / 13.028.7k / 19.7k / 29.4k / 78.6k / 150.4k478 / 255 / 485 / 1.4k / 4.6khttps://openrouter.ai/api/v12026-08-09
deepseek/deepseek-v4-flash-0731high86.8%33/38✓ certified92% (11/12)85% (22/26)13.2s / 7.7s / 14.1s / 34.3s / 126.1s3.1 / 3.0 / 4.0 / 6.0 / 9.029.4k / 23.6k / 34.0k / 69.6k / 107.7k717 / 350 / 763 / 1.7k / 8.5khttps://openrouter.ai/api/v12026-08-09
deepseek/deepseek-v4-prohigh84.2%32/38✓ certified100% (12/12)77% (20/26)17.3s / 9.6s / 15.3s / 34.4s / 240.0s3.1 / 3.0 / 4.0 / 7.0 / 10.029.4k / 23.1k / 31.1k / 67.7k / 116.5k873 / 309 / 661 / 1.6k / 16.8khttps://openrouter.ai/api/v12026-08-09

Per-signal breakdown

openai/gpt-5.6-luna — signal breakdown
SignalPassed
argument-grounding9/10
background-lifecycle3/3
delegation3/4
filesystem-safety9/9
mcp-discovery2/2
mcp-tool-chain1/1
no-unneeded-tools11/11
parallel-tools1/2
plan-lifecycle5/6
plan-safety1/1
required-tool7/8
skill-discovery2/2
state-persistence5/5
state-transition2/2
tool-chain10/10
tool-error-recovery5/6
tool-result-integrity8/8
deepseek/deepseek-v4-flash-0731 — signal breakdown
SignalPassed
argument-grounding9/10
background-lifecycle3/3
delegation3/4
filesystem-safety6/9
mcp-discovery2/2
mcp-tool-chain1/1
no-unneeded-tools10/11
parallel-tools2/2
plan-lifecycle5/6
plan-safety1/1
required-tool7/8
skill-discovery2/2
state-persistence4/5
state-transition1/2
tool-chain7/10
tool-error-recovery5/6
tool-result-integrity8/8
deepseek/deepseek-v4-pro — signal breakdown
SignalPassed
argument-grounding10/10
background-lifecycle2/3
delegation4/4
filesystem-safety6/9
mcp-discovery2/2
mcp-tool-chain1/1
no-unneeded-tools11/11
parallel-tools2/2
plan-lifecycle3/6
plan-safety1/1
required-tool8/8
skill-discovery2/2
state-persistence3/5
state-transition2/2
tool-chain7/10
tool-error-recovery4/6
tool-result-integrity7/8