Exact outputs
Small geometric, interface, or parameter errors can make an otherwise plausible design unusable.
← All blogs · Benchmark introduction
A physical-science design benchmark, beginning with photonic integrated circuits.
PhySciBench measures whether an autonomous agent can turn a scientific requirement into a device that survives evaluator-owned checks. The first release covers two complementary modes: open-ended design discovery and reference-grounded implementation optimization.

Scientific agents must turn an incomplete functional requirement into executable geometry, use tools under bounded resources, and deliver an artifact that survives checks they do not control. PhySciBench evaluates that complete chain.
Small geometric, interface, or parameter errors can make an otherwise plausible design unusable.
Progress depends on choosing probes, reading evaluator feedback, and revising within a fixed budget.
Success is determined by trusted evaluation of the submitted artifact, not by the agent’s explanation.
Open-ended tasks ask the agent to select a device concept and workflow from functional goals rather than reproduce a fixed parameter set.
Explore design tasks →Reference-grounded tasks test faithful implementation and optimization against evaluator-owned performance and artifact checks.
Explore implementation tasks →The benchmark separates what the agent controls from what the evaluator owns. Progress is measured on the submitted device.
Functional requirements, permitted resources, and measurable success criteria define the scientific problem.
The agent authors the device, uses a bounded probe budget, and decides what to change from evaluator feedback.
One exact final submission is frozen with its run record and provenance digests for inspection.
Private evaluator checks determine gate eligibility and the official objective J. They also determine the reporting score S when enabled.
Provisional S runs from 0–2; higher is better. N/E means gate-ineligible, while — means no result is published. Select a scored bar for run details.
Implementation S uses each task's declared J_baseline; S = 1 is baseline parity, while native evaluator gates determine eligibility.
Individually approved single-task runs are not official campaign results or task-release certifications.
The remaining published runs did not clear every trusted gate, so J and S stay null.
Microresonator-v1 leads this published snapshot; it is not yet a broad model ranking.
Published trajectories expose iterative progress rather than only the final number.
Final, checkpoint, or diagnostic visuals are explicitly labelled by their evidence status.
Browse tasks and traces, inspect the public result catalog, read the scoring methodology, or cite the current benchmark release.