Run
Run PhySciBench on your machine with Harbor. Start with one task, then inspect the agent’s work and evaluator results.
Install the benchmark
You’ll need Python 3.12 or newer, uv, Git, and a running Docker engine.
git clone https://github.com/MomeniAli/PICBench.git
cd PICBench
uv sync --extra viewer --extra devThis installs the repository’s patched Harbor version, together with the benchmark and viewer. Run the following commands from the repository root.
Configure your credentials
test -f .env || cp .env.example .envEdit .env and set OPENROUTER_API_KEY for the design run below. Add the simulation credentials required by your task, as described in its environment and task documentation.
Runs use your provider and simulation accounts. Keep .env private; it is ignored by Git.
Run a task
This example runs the four-channel electrothermal WDM channelizer. The repository runner selects the task’s container setup, evaluator, and run policy.
uv run python scripts/run_implementation_task.py \
wdm-channelizer-electrothermal-design-v1Design runs use the repository’s fixed policy: Codex through OpenRouter, medium reasoning, a 7,200-second agent deadline, and 40 evaluator probes. Task-level resource limits still apply.
Preview the run configuration first
Add --dry-run to print the planned command without launching the benchmark.
uv run python scripts/run_implementation_task.py \
wdm-channelizer-electrothermal-design-v1 \
--dry-runInspect the results
PICBench’s Harbor plugins normalize supported runs for the local viewer. Start it to inspect trajectories, evaluator outcomes, and visual artifacts.
uv run python scripts/run_viewer.py --directory logsOpen http://localhost:5001 in your browser. If you’re working on a remote VM, forward port 5001 to your computer.
A completed run is missing from the viewer
Normalize the latest completed trial, then refresh the viewer.
uv run python scripts/normalize_harbor_trial_to_logs.py --latest