vrinda.dev / evidence lab

immutable standings snapshot

the founding harness comparison — raw vs vanilla vs recipe, one model, one prompt

html-app-builds · results remain separate from the cost provenance attached to every run.

snapshot2026-07-18-a0b3a109core sha a0b3a109

# standings

rankarm / modelqualitycostevidence
opencode default agent (harness, no recipe)
z-ai/glm-5.2
human verdict (2026-07-18, recorded unblinded after the provisional AI read was published): all arms basic on this founding task — no quality separation claimed; no showcase-grade output $0.0253 vanilla-r1-7cd3379c vanilla-r2-a5600b64 vanilla-r3-7aee41d7
raw one-shot (single API call, no harness)
z-ai/glm-5.2
human verdict (2026-07-18, recorded unblinded after the provisional AI read was published): all arms basic on this founding task — no quality separation claimed; no showcase-grade output $0.0397 raw-r1-3bdcb563 · quarantined raw-r2-872e7882 · quarantined raw-r3-447f47dd · quarantined
opencode + plan-build-check recipe
z-ai/glm-5.2
human verdict (2026-07-18, recorded unblinded after the provisional AI read was published): all arms basic on this founding task — no quality separation claimed; no showcase-grade output $0.0864 recipe-r1-57303053 recipe-r2-a3b0ac8b recipe-r3-a5f58c9e

meteredestimated / backfilledunmetered

disclosures

disclosures
  1. DIRECTIONAL, ILLUSTRATIVE — NOT CLAIM-GRADE. N=3 per arm, one task family (kanban-01). Quality was read by one AI screening eye at publication; the human verdict has since been recorded — see the HUMAN VERDICT disclosure below. A different eye or task could land differently.
  2. All three arms received the BYTE-IDENTICAL prompt, which asked for a single self-contained file. All 9 artifacts were 10/10 on the enumerated spec by code-trace.
  3. THE FINDING: the raw one-shot produced the most visually polished boards but pulled EXTERNAL Google Fonts (fonts.googleapis.com/gstatic.com) and auto-seeded demo cards — violating the self-contained requirement — so all 3 raw artifacts were QUARANTINED by the publication scan. Both harnessed arms (vanilla + recipe) honored self-contained and are published. The blind screening independently flagged the same 3 font-loading tiles.
  4. Vanilla (the plain harness) ranked best on the blind screening AND was the cheapest ($0.025 median); the recipe cost 3.4x more ($0.086 median) and did not clearly beat vanilla here (one recipe rep had a filter empty-state bug). For this simple task the plain harness was enough — an honest result, not the hoped-for 'recipe wins'.
  5. COST is measured harness-inclusive via OpenRouter credits-delta (every provider call the harness makes), the figure most comparisons omit.
  6. INSTRUMENT NOTE (transparency): an earlier run of this batch showed the raw arm failing to produce any artifact — that was OUR client-side HTTP bug (chunked-decode corruption on large responses), since fixed (buffer-based). This published batch is the corrected run; the raw failures here are the genuine self-contained violation above, not that bug.
  7. The blind screening read is an AI code-trace assessment, not a live-DOM interaction test. The originally planned morning verification (live-DOM checklist + human blind rank) was superseded: the human verdict was recorded unblinded before a cold rank could happen, no per-tile human ranks exist, and the quality column now carries that verdict — landed as its own new snapshot per the lab's method.
  8. HUMAN VERDICT (2026-07-18, recorded UNBLINDED — the user viewed the results with arm labels visible after the provisional AI read was published, so the originally planned cold-blind rank protocol was no longer possible). In their words: "all of them are equally very basic... for some reason the usual glm quality did not come at all." Operationalized: all arms basic on this founding task; no per-tile quality separation meaningful; no showcase-grade output. The AI screening read stays published above, labeled as exactly what it is — a screening read, not the verdict.