current standings
| rank | recipe | quality | cost | evidence |
|---|---|---|---|---|
| — | vanilla harness opencode@1.17.14 | all arms basic — human verdict | $0.0253 | self-contained |
| — | raw one-shot opencode@1.17.14 | all arms basic — human verdict | $0.0397 | external dep · quarantined |
| — | plan-build-check opencode@1.17.14 | all arms basic — human verdict | $0.0864 | self-contained |
metered — billed from a receipt
est — token-estimated
unmetered — sub / separate pool, never a price
quality = human verdict (2026-07-18, recorded unblinded after the provisional AI read was published): all arms basic on this founding task — no quality separation claimed; no showcase-grade output · the AI screening read stays published in the lab record, labeled as screening
case files
5 builders. One prompt. Ranked.
One byte-identical racing-game prompt fired at five different builders, alone. A blind eye ranked the five tiles.
$0.04 beat $0.26.
On one dashboard task, two ~4-cent runs swept #1 and #2 over a 26-cent agentic champion, which placed last.
Same prompt. 1/4 the price.
The same 785-byte base card on GLM arms vs a raw Sonnet baseline. The cheap GLM ran roughly a quarter of the price.
The founding harness comparison.
One model, one prompt, three ways to run it. The harness enforced the self-contained rule the raw one-shot ignored — receipts and blind screening on the page.
inside the lab ›the lab
method, short version
Read these before you read the numbers. Every page below repeats its own full disclosures verbatim.
- The verdicts are n=1. One eye, one task each, given as a blind rank. They are directional and illustrative — not statistical, not claim-grade. A different eye or a second task could land differently.
- Estimates are labeled est. Costs tagged est are token-estimated (the provider generation-id was not exposed) or recorded null in the receipt. Costs tagged metered / api are billed figures. Subscription and separate-pool arms show $0 marginal — no per-call dollar meter exists for them.
- One label was corrected upward. A champion run cited upstream at “~$0.209” meters $0.256634 from its own receipts; the higher metered figure is the one shown.
- Full disclosures ride every page. Pre-blind glimpses, apples-to-apples caveats, honest-fail run status, and background-lighting notes are all disclosed on the individual pages, unaveraged.