Published benchmark results

Leaderboard

Explore results by task across design and implementation. These results are a published snapshot, not yet a broad model ranking.

How scoring works →

Design / Discovery tasks

CodexGPT-5.6-solClaude CodeClaude Opus 5.5
Accepted design · S = 0.90–2.00

Provisional S runs from 0–2; higher is better. N/E means gate-ineligible, while — means no result is published. Select a scored bar for run details.

Implementation / Optimization tasks

At or above baseline parity · S = 1.00–2.00

Implementation S uses each task's declared J_baseline; S = 1 is baseline parity, while native evaluator gates determine eligibility.

Individually approved single-task runs are not official campaign results or task-release certifications.