vrinda.dev / evidence lab

preregistered harness experiments

lab

A controlled place to test what execution recipes change — including the cost of the harness itself.

thesis

The same model does not do the same work in every harness. Give a cheap open model a plan-build-check recipe inside a real coding harness and it can close much of the gap to a raw call from a pricier setup — or not. The lab runs that experiment under controlled conditions and publishes the receipts, so the claim is checkable rather than asserted.

the claim

One prompt. One model. Three ways of running it: a raw one-shot, a vanilla harness, and the same harness driving a house recipe. Everything else held fixed — same provider, same routing, same fixture. What changes is the execution recipe, and what we measure is the cost of that change against the quality of what it builds. This is the founding study; the standings it feeds decay as models and prices move, and get re-run.

how the lab is hash-stamped

Every run executes in a throwaway container that cannot reach the internet except through a single allow-listed path to the model provider — no other host, no side channels. The model's API key is passed in memory only and never touches disk, a log, or the receipt; the receipt carries a short fingerprint of the key, nothing more. Cost is measured across the whole run, including the harness's own calls — the number other comparisons quietly leave out. Nothing is published until the artifact passes a secrets-and-dependencies scan. The preregistration below was committed, with an independent timestamp, before the first paid run — so the arms, repetitions, run order, and scoring rule could not be chosen after seeing the results.

take the recipe home

The recipes compared here are not locked in a product. Each is a small bundle of skills and seat configuration you can download, drop into your own harness with your own key, and run unchanged. The artifact is the ad: what wins here, you can take with you.

method

These are small, honest studies, not a universal benchmark. A result of N=3 on one task, judged by one eye, is directional and illustrative — not claim-grade. We publish the dispersion, the failures, and the decay date alongside every verdict, and we never collapse several different dimensions into one invented score. Costs are labeled by how they were obtained: metered (billed and measured), estimated (token-derived), or unmetered (subscription or a separate quota pool — never shown as a dollar price). Receipts are hash-stamped: content-addressed manifests with digests, not signed immutability. Where a human blind rank has not yet been recorded, the quality column shows an AI screening read, labeled as exactly that, and the human verdict lands later as its own snapshot.

preregistration

sha
0da6084f7eeb8c831395057f13821dbfefc3823f57fe3b84f897b7edd0984c41
timing
committed before runs
apparatus
throwaway container, provider-only allow-list, in-memory key, secrets scan

results

rankarm / modelqualitycostevidence
opencode default agent (harness, no recipe)
z-ai/glm-5.2
human verdict (2026-07-18, recorded unblinded after the provisional AI read was published): all arms basic on this founding task — no quality separation claimed; no showcase-grade output $0.0253 vanilla-r1-7cd3379c vanilla-r2-a5600b64 vanilla-r3-7aee41d7
raw one-shot (single API call, no harness)
z-ai/glm-5.2
human verdict (2026-07-18, recorded unblinded after the provisional AI read was published): all arms basic on this founding task — no quality separation claimed; no showcase-grade output $0.0397 raw-r1-3bdcb563 · quarantined raw-r2-872e7882 · quarantined raw-r3-447f47dd · quarantined
opencode + plan-build-check recipe
z-ai/glm-5.2
human verdict (2026-07-18, recorded unblinded after the provisional AI read was published): all arms basic on this founding task — no quality separation claimed; no showcase-grade output $0.0864 recipe-r1-57303053 recipe-r2-a3b0ac8b recipe-r3-a5f58c9e

meteredestimated / backfilledunmetered

receipts

Every amount is harness-inclusive: provider calls made by the harness belong in the run total.

runsourceamount
vanilla-r1-7cd3379ccredits-delta$0.0239
vanilla-r3-7aee41d7credits-delta$0.0253
vanilla-r2-a5600b64credits-delta$0.0279
raw-r2-872e7882credits-delta$0.0316
raw-r3-447f47ddcredits-delta$0.0397
raw-r1-3bdcb563credits-delta$0.0415
recipe-r2-a3b0ac8bcredits-delta$0.0738
recipe-r1-57303053credits-delta$0.0864
recipe-r3-a5f58c9ecredits-delta$0.0906

artifacts

vanilla-r1-7cd3379c vanilla-r3-7aee41d7 vanilla-r2-a5600b64 raw-r2-872e7882 · quarantined raw-r3-447f47dd · quarantined raw-r1-3bdcb563 · quarantined recipe-r2-a3b0ac8b recipe-r1-57303053 recipe-r3-a5f58c9e

disclosures

disclosures
  1. DIRECTIONAL, ILLUSTRATIVE — NOT CLAIM-GRADE. N=3 per arm, one task family (kanban-01). Quality was read by one AI screening eye at publication; the human verdict has since been recorded — see the HUMAN VERDICT disclosure below. A different eye or task could land differently.
  2. All three arms received the BYTE-IDENTICAL prompt, which asked for a single self-contained file. All 9 artifacts were 10/10 on the enumerated spec by code-trace.
  3. THE FINDING: the raw one-shot produced the most visually polished boards but pulled EXTERNAL Google Fonts (fonts.googleapis.com/gstatic.com) and auto-seeded demo cards — violating the self-contained requirement — so all 3 raw artifacts were QUARANTINED by the publication scan. Both harnessed arms (vanilla + recipe) honored self-contained and are published. The blind screening independently flagged the same 3 font-loading tiles.
  4. Vanilla (the plain harness) ranked best on the blind screening AND was the cheapest ($0.025 median); the recipe cost 3.4x more ($0.086 median) and did not clearly beat vanilla here (one recipe rep had a filter empty-state bug). For this simple task the plain harness was enough — an honest result, not the hoped-for 'recipe wins'.
  5. COST is measured harness-inclusive via OpenRouter credits-delta (every provider call the harness makes), the figure most comparisons omit.
  6. INSTRUMENT NOTE (transparency): an earlier run of this batch showed the raw arm failing to produce any artifact — that was OUR client-side HTTP bug (chunked-decode corruption on large responses), since fixed (buffer-based). This published batch is the corrected run; the raw failures here are the genuine self-contained violation above, not that bug.
  7. The blind screening read is an AI code-trace assessment, not a live-DOM interaction test. The originally planned morning verification (live-DOM checklist + human blind rank) was superseded: the human verdict was recorded unblinded before a cold rank could happen, no per-tile human ranks exist, and the quality column now carries that verdict — landed as its own new snapshot per the lab's method.
  8. HUMAN VERDICT (2026-07-18, recorded UNBLINDED — the user viewed the results with arm labels visible after the provisional AI read was published, so the originally planned cold-blind rank protocol was no longer possible). In their words: "all of them are equally very basic... for some reason the usual glm quality did not come at all." Operationalized: all arms basic on this founding task; no per-tile quality separation meaningful; no showcase-grade output. The AI screening read stays published above, labeled as exactly what it is — a screening read, not the verdict.

take the recipe home

download plan-build-check sha 164e40087cc0

race a matchup like this yourself — on the direct-chain runner ›