From calibration to a 12-scenario corpus.
The corpus now spans calibration bugs, text processing, backend/API fixes, UI logic, migrations, monorepo upgrades, observability, document extraction, greenfield product work, and external-helper leverage.
Loading…
An evidence-first coding-agent benchmark. Pass/fail, wall-clock, diffs, and verifier outputs are backed by committed artifacts; token/cost figures are normalized public estimates with provenance caveats documented in the audit pack.
Corpus status
Loading…
Benchmark scenarios
All six agents attempted every scenario in isolated workspaces. Pass/fail is binary and Hermes-verified; outcome and efficiency compare scaffold+model configurations, not scaffold-only performance.
Decision lens
Lower cost and higher execution score is better. Agents without cost telemetry appear at the right edge at their actual execution score.
Methodology & interpretation
The corpus now spans calibration bugs, text processing, backend/API fixes, UI logic, migrations, monorepo upgrades, observability, document extraction, greenfield product work, and external-helper leverage.