← All blogs · Benchmark introduction

Public benchmark · photonic integrated circuits

PhySciBench

A physical-science design benchmark, beginning with photonic integrated circuits.

PhySciBench measures whether an autonomous agent can turn a scientific requirement into a device that survives evaluator-owned checks. The first release covers two complementary modes: open-ended design discovery and reference-grounded implementation optimization.

Conceptual silicon photonic chip with microring resonators, interferometer meshes, and a fiber array
Evaluator-owned checks
Benchmark domain 01PIC design & verificationConcept visualization · published evidence uses exact evaluated layouts
Why PhySciBench

A physical design is not a persuasive answer.

Scientific agents must turn an incomplete functional requirement into executable geometry, use tools under bounded resources, and deliver an artifact that survives checks they do not control. PhySciBench evaluates that complete chain.

01

Exact outputs

Small geometric, interface, or parameter errors can make an otherwise plausible design unusable.

02

Long-horizon work

Progress depends on choosing probes, reading evaluator feedback, and revising within a fixed budget.

03

Independent evidence

Success is determined by trusted evaluation of the submitted artifact, not by the agent’s explanation.

What it measures

From scientific intent to verified artifact

InterpretFunctional requirements
DesignExecutable device source
IterateBounded probe feedback
VerifyEvaluator-owned outcomes
16 tasks · Design / Discovery

Choose the approach from the requirement

Open-ended tasks ask the agent to select a device concept and workflow from functional goals rather than reproduce a fixed parameter set.

Explore design tasks →
10 tasks · Implementation / Optimization

Implement and improve a defined target

Reference-grounded tasks test faithful implementation and optimization against evaluator-owned performance and artifact checks.

Explore implementation tasks →
Evaluation protocol

From open requirement to verified artifact

The benchmark separates what the agent controls from what the evaluator owns. Progress is measured on the submitted device.

Agent-controlledInterpret → design → iterate
Frozen submissionSource + artifact digest
Evaluator-controlledExecute → gate → score
01

Task contract

Functional requirements, permitted resources, and measurable success criteria define the scientific problem.

02

Agent design loop

The agent authors the device, uses a bounded probe budget, and decides what to change from evaluator feedback.

03

Final artifact

One exact final submission is frozen with its run record and provenance digests for inspection.

04

Trusted verification

Private evaluator checks determine gate eligibility and the official objective J. They also determine the reporting score S when enabled.

Design / Discovery tasks

CodexGPT-5.6-solClaude CodeClaude Opus 5.5
Accepted design · S = 0.90–2.00

Provisional S runs from 0–2; higher is better. N/E means gate-ineligible, while — means no result is published. Select a scored bar for run details.

Implementation / Optimization tasks

At or above baseline parity · S = 1.00–2.00

Implementation S uses each task's declared J_baseline; S = 1 is baseline parity, while native evaluator gates determine eligibility.

Individually approved single-task runs are not official campaign results or task-release certifications.

What the current evidence shows

Early results, bounded conclusions

29/35

Finals are score-eligible

The remaining published runs did not clear every trusted gate, so J and S stay null.

1.5597

Highest reporting S

Microresonator-v1 leads this published snapshot; it is not yet a broad model ranking.

944

Evaluator probes recorded

Published trajectories expose iterative progress rather than only the final number.

35/35

Runs include a visual artifact

Final, checkpoint, or diagnostic visuals are explicitly labelled by their evidence status.

Inspect and reuse

Open evidence, explicit boundaries

Browse tasks and traces, inspect the public result catalog, read the scoring methodology, or cite the current benchmark release.