Loading…

Six agents.
Twelve scenarios.
Verified runs.

An evidence-first coding-agent benchmark. Pass/fail, wall-clock, diffs, and verifier outputs are backed by committed artifacts; token/cost figures are normalized public estimates with provenance caveats documented in the audit pack.

Corpus status

Loading…

Scenarios committed benchmark cases
Agent runs six-agent matrix
Pass rate Hermes-verified
Total public cost normalized estimate · configured model prices

Benchmark scenarios

Agent × Scenario results

All six agents attempted every scenario in isolated workspaces. Pass/fail is binary and Hermes-verified; outcome and efficiency compare scaffold+model configurations, not scaffold-only performance.

Decision lens

Cost vs. process quality frontier

Lower cost and higher execution score is better. Agents without cost telemetry appear at the right edge at their actual execution score.

Methodology & interpretation

How to read these results

From calibration to a 12-scenario corpus.

The corpus now spans calibration bugs, text processing, backend/API fixes, UI logic, migrations, monorepo upgrades, observability, document extraction, greenfield product work, and external-helper leverage.

What changed

  • Added the remaining BB-002 and BB-004–BB-012 tasks and fixtures.
  • Ran all six agents in isolated workspaces.
  • Verified pytest and JSON CLI smoke independently.
  • Promoted Pages to a multi-scenario benchmark briefing with per-scenario detail pages.

Known gaps

  • ccusage is the preferred collector; direct extraction is fallback only when run attribution is missing.
  • agy has no reliable token export — Antigravity transcripts are not Gemini CLI ccusage logs.
  • More scenarios needed before claiming broad capability rankings.
  • Vendor-reported cost ≠ normalized public-price estimates.
ccusage methodology & pricing basis