Methods · evaluator-owned verification

Methodology

PhySciBench tests whether an agent can produce a physical design artifact that satisfies a task contract.

What the benchmark measures

A complete design workflow

01

Requirement interpretation

Translate a functional scientific objective into a concrete device plan.

02

Tool-guided iteration

Use bounded evaluator feedback to decide what to inspect and change.

03

Artifact correctness

Produce executable source and the required physical layout or simulation artifact.

04

Verified performance

Clear trusted gates before evaluator-owned metrics can be reported.

Authority boundary

Agent control and evaluator control stay separate

Agent controls
  • Design source and parameter choices
  • Probe requests within the task budget
  • Iteration strategy and final submission
Evaluator controls
  • Private verification environment and scientific checks
  • Gate eligibility and official objective J
  • Canonical result and artifact provenance
Benchmark tracks

Two complementary forms of engineering work

Design / Discovery

Start from functional requirements

The agent chooses a device concept and design workflow without being asked to reproduce one paper’s parameters.

Implementation / Optimization

Start from a defined reference target

The agent implements and improves a specified structure while the evaluator checks fidelity and performance.

Scoring

Official objective first; reporting layer second

J

Official objective

J is task-specific and evaluator-owned. Lower is better. It remains null when the final submission does not clear trusted gates.

S = 2 / (1 + J / R)

Reporting score

S is a 0–2 reporting layer for cross-task reading. It does not replace J, change gate outcomes, or certify every specification. R is the provisional J_acc calibration for design tasks and the declared external J_baseline for implementation tasks.

Current evidence boundary

What this release can and cannot support

Supported

Inspection of released tasks, published run outcomes, trajectories, evaluator-approved visuals, and provenance digests.

Not yet supported

Broad model rankings, implementation-track conclusions, or claims about physical-science domains beyond the first PIC release.