preregistered harness experiments
lab
A controlled place to test what execution recipes change — including the cost of the harness itself.
thesis
The same model does not do the same work in every harness. Give a cheap open model a plan-build-check recipe inside a real coding harness and it can close much of the gap to a raw call from a pricier setup — or not. The lab runs that experiment under controlled conditions and publishes the receipts, so the claim is checkable rather than asserted.
the claim
One prompt. One model. Three ways of running it: a raw one-shot, a vanilla harness, and the same harness driving a house recipe. Everything else held fixed — same provider, same routing, same fixture. What changes is the execution recipe, and what we measure is the cost of that change against the quality of what it builds. This is the founding study; the standings it feeds decay as models and prices move, and get re-run.
how the lab is hash-stamped
Every run executes in a throwaway container that cannot reach the internet except through a single allow-listed path to the model provider — no other host, no side channels. The model's API key is passed in memory only and never touches disk, a log, or the receipt; the receipt carries a short fingerprint of the key, nothing more. Cost is measured across the whole run, including the harness's own calls — the number other comparisons quietly leave out. Nothing is published until the artifact passes a secrets-and-dependencies scan. The preregistration below was committed, with an independent timestamp, before the first paid run — so the arms, repetitions, run order, and scoring rule could not be chosen after seeing the results.
take the recipe home
The recipes compared here are not locked in a product. Each is a small bundle of skills and seat configuration you can download, drop into your own harness with your own key, and run unchanged. The artifact is the ad: what wins here, you can take with you.
method
These are small, honest studies, not a universal benchmark. A result of N=3 on one task, judged by one eye, is directional and illustrative — not claim-grade. We publish the dispersion, the failures, and the decay date alongside every verdict, and we never collapse several different dimensions into one invented score. Costs are labeled by how they were obtained: metered (billed and measured), estimated (token-derived), or unmetered (subscription or a separate quota pool — never shown as a dollar price). Receipts are hash-stamped: content-addressed manifests with digests, not signed immutability. Where a human blind rank has not yet been recorded, the quality column shows an AI screening read, labeled as exactly that, and the human verdict lands later as its own snapshot.
preregistration
results
| rank | arm / model | quality | cost | evidence |
|---|---|---|---|---|
| — | opencode default agent (harness, no recipe) z-ai/glm-5.2 |
human verdict (2026-07-18, recorded unblinded after the provisional AI read was published): all arms basic on this founding task — no quality separation claimed; no showcase-grade output | $0.0253 | vanilla-r1-7cd3379c vanilla-r2-a5600b64 vanilla-r3-7aee41d7 |
| — | raw one-shot (single API call, no harness) z-ai/glm-5.2 |
human verdict (2026-07-18, recorded unblinded after the provisional AI read was published): all arms basic on this founding task — no quality separation claimed; no showcase-grade output | $0.0397 | raw-r1-3bdcb563 · quarantined raw-r2-872e7882 · quarantined raw-r3-447f47dd · quarantined |
| — | opencode + plan-build-check recipe z-ai/glm-5.2 |
human verdict (2026-07-18, recorded unblinded after the provisional AI read was published): all arms basic on this founding task — no quality separation claimed; no showcase-grade output | $0.0864 | recipe-r1-57303053 recipe-r2-a3b0ac8b recipe-r3-a5f58c9e |
meteredestimated / backfilledunmetered
receipts
Every amount is harness-inclusive: provider calls made by the harness belong in the run total.
| run | source | amount |
|---|---|---|
| vanilla-r1-7cd3379c | credits-delta | $0.0239 |
| vanilla-r3-7aee41d7 | credits-delta | $0.0253 |
| vanilla-r2-a5600b64 | credits-delta | $0.0279 |
| raw-r2-872e7882 | credits-delta | $0.0316 |
| raw-r3-447f47dd | credits-delta | $0.0397 |
| raw-r1-3bdcb563 | credits-delta | $0.0415 |
| recipe-r2-a3b0ac8b | credits-delta | $0.0738 |
| recipe-r1-57303053 | credits-delta | $0.0864 |
| recipe-r3-a5f58c9e | credits-delta | $0.0906 |
artifacts
vanilla-r1-7cd3379c vanilla-r3-7aee41d7 vanilla-r2-a5600b64 raw-r2-872e7882 · quarantined raw-r3-447f47dd · quarantined raw-r1-3bdcb563 · quarantined recipe-r2-a3b0ac8b recipe-r1-57303053 recipe-r3-a5f58c9edisclosures
disclosures
- DIRECTIONAL, ILLUSTRATIVE — NOT CLAIM-GRADE. N=3 per arm, one task family (kanban-01). Quality was read by one AI screening eye at publication; the human verdict has since been recorded — see the HUMAN VERDICT disclosure below. A different eye or task could land differently.
- All three arms received the BYTE-IDENTICAL prompt, which asked for a single self-contained file. All 9 artifacts were 10/10 on the enumerated spec by code-trace.
- THE FINDING: the raw one-shot produced the most visually polished boards but pulled EXTERNAL Google Fonts (fonts.googleapis.com/gstatic.com) and auto-seeded demo cards — violating the self-contained requirement — so all 3 raw artifacts were QUARANTINED by the publication scan. Both harnessed arms (vanilla + recipe) honored self-contained and are published. The blind screening independently flagged the same 3 font-loading tiles.
- Vanilla (the plain harness) ranked best on the blind screening AND was the cheapest ($0.025 median); the recipe cost 3.4x more ($0.086 median) and did not clearly beat vanilla here (one recipe rep had a filter empty-state bug). For this simple task the plain harness was enough — an honest result, not the hoped-for 'recipe wins'.
- COST is measured harness-inclusive via OpenRouter credits-delta (every provider call the harness makes), the figure most comparisons omit.
- INSTRUMENT NOTE (transparency): an earlier run of this batch showed the raw arm failing to produce any artifact — that was OUR client-side HTTP bug (chunked-decode corruption on large responses), since fixed (buffer-based). This published batch is the corrected run; the raw failures here are the genuine self-contained violation above, not that bug.
- The blind screening read is an AI code-trace assessment, not a live-DOM interaction test. The originally planned morning verification (live-DOM checklist + human blind rank) was superseded: the human verdict was recorded unblinded before a cold rank could happen, no per-tile human ranks exist, and the quality column now carries that verdict — landed as its own new snapshot per the lab's method.
- HUMAN VERDICT (2026-07-18, recorded UNBLINDED — the user viewed the results with arm labels visible after the provisional AI read was published, so the originally planned cold-blind rank protocol was no longer possible). In their words: "all of them are equally very basic... for some reason the usual glm quality did not come at all." Operationalized: all arms basic on this founding task; no per-tile quality separation meaningful; no showcase-grade output. The AI screening read stays published above, labeled as exactly what it is — a screening read, not the verdict.
take the recipe home
download plan-build-check sha 164e40087cc0race a matchup like this yourself — on the direct-chain runner ›