vrinda.dev

the recipe referee · real runs · receipts

AI coding recipes, run under lab conditions. Ranked, with receipts.

A recipe is a bundle of skills and seat configs for Claude-Code-type coding harnesses. We run each one in a pinned harness, meter every call, blind-rank the artifacts, and publish all of it.

snapshot 2026-07-18-a0b3a109 · batch 79ba83e6 · immutable fresh
rankrecipequalitycostevidence
vanilla harness opencode@1.17.14 all arms basic — human verdict $0.0253 self-contained
raw one-shot opencode@1.17.14 all arms basic — human verdict $0.0397 external dep · quarantined
plan-build-check opencode@1.17.14 all arms basic — human verdict $0.0864 self-contained

metered — billed from a receipt est — token-estimated unmetered — sub / separate pool, never a price
quality = human verdict (2026-07-18, recorded unblinded after the provisional AI read was published): all arms basic on this founding task — no quality separation claimed; no showcase-grade output · the AI screening read stays published in the lab record, labeled as screening

full standings ›

The lab runs every recipe inside one pinned, sandboxed harness — the same model, the same prompt, raw against harnessed — and meters each call end to end.

The founding comparison is running now; its prereg was hash-stamped before the first paid call, so the claim is on the record before any result is.

inside the lab ›

method, short version

Read these before you read the numbers. Every page below repeats its own full disclosures verbatim.

  • The verdicts are n=1. One eye, one task each, given as a blind rank. They are directional and illustrative — not statistical, not claim-grade. A different eye or a second task could land differently.
  • Estimates are labeled est. Costs tagged est are token-estimated (the provider generation-id was not exposed) or recorded null in the receipt. Costs tagged metered / api are billed figures. Subscription and separate-pool arms show $0 marginal — no per-call dollar meter exists for them.
  • One label was corrected upward. A champion run cited upstream at “~$0.209” meters $0.256634 from its own receipts; the higher metered figure is the one shown.
  • Full disclosures ride every page. Pre-blind glimpses, apples-to-apples caveats, honest-fail run status, and background-lighting notes are all disclosed on the individual pages, unaveraged.
full method — /lab/ ›