Leaderboard
Explore results by task across design and implementation. These results are a published snapshot, not yet a broad model ranking.
How scoring works →Design / Discovery tasks
CodexGPT-5.6-solClaude CodeClaude Opus 5.5
Accepted design · S = 0.90–2.00
Provisional S runs from 0–2; higher is better. N/E means gate-ineligible, while — means no result is published. Select a scored bar for run details.
Implementation / Optimization tasks
At or above baseline parity · S = 1.00–2.00
Implementation S uses each task's declared J_baseline; S = 1 is baseline parity, while native evaluator gates determine eligibility.
Individually approved single-task runs are not official campaign results or task-release certifications.