vrinda.dev / inversion
‹ all comparisons

02 / inversion · canvas-dashboard-01

$0.04 beat $0.26.

On one dashboard task, on a blind board, two ~4-cent lean runs took #1 and #2 — over a full agentic ~26-cent champion pipeline, which the eye placed last. Same task, richer machinery, more money, lower rank.

read this first — mandatory framing

This page shows one striking counter-case, not a general win. The overall experiment it comes from (slice-f3) was REFUTED: the lean “slice” pipeline did not beat a bare cost-matched one-shot on a majority of tasks (2 of 3 — the racer and the forms task).

Do not read this as the cheap recipe broadly winning. It is canvas-dashboard-01 only, where the two cheap reps happened to sweep. n=3 tasks, one eye, 2 reps/arm — directional, not claim-grade.

# blind rank & reveal

ranktileroleartifactmetered cost
1A SLICE rep1
lean spec + harness + contract
16,756 B
sha 7d609e56…
$0.0431metered
2D SLICE rep2
lean spec + harness + contract
18,158 B
sha 7d5816ff…
$0.0399metered
3C floor — GLM one-shot
bare cost-matched baseline
37,765 B
sha 53aee12a…
$0.0437metered
4B CHAMPION — full agentic pipeline
multi-seat build · honest-fail / dead(s6)
32,486 B
sha 268981fd…
$0.256634metered

Champion cost, corrected honestly: the brief and rank doc cite the champion at “~$0.209”, but that figure is not reproducible from receipts. The champion run’s own manifest records cost_total = 0.256634 and its stage rows sum to the same — so the true metered cost is $0.256634, shown here. The inversion holds a fortiori: the champion was even pricier than advertised and still ranked last. It also came from an honest-fail / dead(s6) run — quarantined, but it renders and was critiqued on three concrete bugs (not a degenerate artifact).

blind rank source: agentic-v3/slice-f3/SPONSOR-RANK-SLICE-2026-07-14.md · sha map: slice-f3/board/final/out-private/mapping.json (all four tile shas recomputed and match)

# the blind read (verbatim)

When it comes to the first task, that is, the canvas dashboard, from a visual richness point of view:

- Tile B is good and well planned. There are a couple of minor bugs or minor issues: there is a static number shown for each product. The static number is not automatically changing on the date range click, but it is changing when I click back on the product again. That means this is trying to show the right numbers, but I think the sync between the number showcase and the date range click is not proper. That is one bug, I would say.

- Obviously, earlier I reported the same issue: when I click on the Remove Check buttons, the graph is not recalibrating or removing the bar. That's it. It is not a bad thing, but it doesn't feel the same way that others attempt the experience.

- The third issue is that the best day number is shown, but the day is, I think, repeated twice: 7 by 4, 7 by 4. I don't know what it is. I would also classify this as a mistake that I haven't reported before, but with visual richness, I cannot accept its mistakes, so I'm keeping it down.

Now, when it comes to tile D, design-wise, it is basic, but it is doing the job. It's professional-looking, but color-wise, it is not very polished. Most of it is all violet color, but it is working from a logic point of view. Here also, the best day is given as day 28. Usually, here I used to see the actual date. Here, it is showing day 28, which means it is thinking in a different way. That is about tile D and tile C. I've seen this before. There is a rendering mistake, even though the design is very good. The trend rendering did not happen well. That means it feels like a blank screen. The trend line is not shown. Otherwise, if the trend line is shown, this would have been one of the best visually rich designs. Again, at the bottom, it is showing implementation notes and stuff. Hence, this is also having some problems in terms of how it is planned.

Till now, we discussed tile B, tile C, and tile D. Everything has some minor problems. Nothing is perfected. When it comes to the remaining tile A, here also, structure-wise, it is working, and color-wise, it is also decent, but the best day is showing 7/14. Usually, I would expect to say, let's say, July 14 or something, but okay, at least this is not a mistake, in my opinion. From a completeness point of view, tile A actually is complete. Even though it is not rich, it is decent. From tile A, I will give the first rating, then I will give the next rating to tile D, then I'll come to tile C. Earlier, I thought tile B is good, but when I looked at all three mistakes, maybe if I discount only the trend line not appearing, I think tile C is looking good now. Then tile B, maybe I'll rank tile B last, even though it is good from a rich UI point of view. I'm unable to classify it as the best because it has three mistakes. I can see simple mistakes, but three is not there. My rating again is tile A first, then tile D, then tile C, then tile B.

verbatim, unaveraged: agentic-v3/slice-f3/SPONSOR-RANK-SLICE-2026-07-14.md (blind read, lines 9–18) · forced order: dashboard A > D > C > B

# open the artifacts

#1 · A
SLICE rep1
lean spec + harness + contract
$0.0431

artifact loads once published

sha 7d609e56…open full screen ↗
#2 · D
SLICE rep2
lean spec + harness + contract
$0.0399

artifact loads once published

sha 7d5816ff…open full screen ↗
#3 · C
Floor — GLM one-shot
bare cost-matched baseline
$0.0437

artifact loads once published

sha 53aee12a…open full screen ↗
#4 · B
Champion — full agentic
multi-seat · honest-fail / dead(s6)
$0.256634

artifact loads once published

sha 268981fd…open full screen ↗

Previews are sandboxed (sandbox="allow-scripts", no same-origin access) and load on click. Until the artifacts are published, each tile shows its placeholder and its open full screen link.

# receipts

tilerolemodelsmetered usdusd_sourcegen-id
Arep1 (slice) glm-5.2 ×3 + luna review $0.04309749 metered
usd == true_up.actual_usd
gen-…jTR7z4
blessing slice:3
Drep2 (slice) glm-5.2 ×3 + luna review $0.03988069 metered
usd == true_up.actual_usd
gen-…0VT00T
blessing slice:3
Cfloor (GLM one-shot) z-ai/glm-5.2 (fp8, Novita) $0.04369409 metered
openrouter == gen_stats
gen-…Ah5kKY
Bchampion (agentic) grok-4.5 + glm-5.2 + luna $0.256634 metered
cost_basis = metered
a3 rows carry no gen-id

Champion 11-row receipt, summarized — 3 setup rows (run_start, watchdog_envelope, vision_preflight) + 8 metered stages. The pipeline drafted, was refused by its own vision check, re-rolled, was accepted, seated, repaired, and re-seated:

#stagemodelcostcumulativenote
1plannerx-ai/grok-4.5$0.021814$0.0218142,191 out-tok
2draftz-ai/glm-5.2$0.057706$0.07952016,603 out-tok
3vlm_readgpt-5.6-luna$0.004128$0.083648verdict: refuse
4rerollz-ai/glm-5.2$0.033932$0.1175809,828 out-tok
5vlm_readgpt-5.6-luna$0.006288$0.123868verdict: accept
6seatgpt-5.6-luna$0.030104$0.153972
7repairz-ai/glm-5.2$0.072319$0.22629110,740 out-tok
8seatgpt-5.6-luna$0.030343$0.256633final

Eight metered stages sum to $0.256633 cumulative; the manifest records cost_total = 0.256634. That is ~6× the ~4-cent lean reps that outranked it. Cost was never the problem — the mean receipted $/run over all six slice-f3 runs was $0.059387, under the $0.080 bar.

show all verbatim receipt rows (reps A/D, floor C, champion B)
// TILE A — rep1 (slice) · slice-f3/money/receipts.jsonl lines 32,36,40,41
{"run_id":"f3-canvas-dashboard-01-rep1","call_id":"...:slice:1","arm":"slice","model":"z-ai/glm-5.2","provider":"StreamLake","prompt_tokens":1953,"completion_tokens":14,"usd":0.001845228,"status":"SUCCESS","generation_id":"gen-1784041629-tPmxaEv8AksK3MBsTKsJ","cost_basis":"max-provider-and-live-pin"}
{"run_id":"f3-canvas-dashboard-01-rep1","call_id":"...:slice:2","arm":"slice","model":"z-ai/glm-5.2","provider":"StreamLake","prompt_tokens":3424,"completion_tokens":5146,"usd":0.01810776,"status":"SUCCESS","generation_id":"gen-1784041630-bxo2e0Q8YTcYT5i0himD","artifact_sha256":"7d609e567029e1ba…"}
{"run_id":"f3-canvas-dashboard-01-rep1","call_id":"...:review:1","arm":"review","model":"openai/gpt-5.6-luna","provider":"OpenAI","prompt_tokens":9034,"completion_tokens":1627,"usd":0.02105375,"status":"SUCCESS","generation_id":"gen-1784041780-L7h2js5h0JAHw7hB1pew"}
{"run_id":"f3-canvas-dashboard-01-rep1","call_id":"...:slice:3","arm":"slice","model":"z-ai/glm-5.2","provider":"StreamLake","prompt_tokens":2203,"completion_tokens":19,"usd":0.002090748,"status":"SUCCESS","generation_id":"gen-1784041794-jTR7z4yindFsTVvTPXdi"}
// tile A total (blessing chain) = $0.04309749

// TILE D — rep2 (slice) · slice-f3/money/receipts.jsonl lines 47,48,52,53
{"run_id":"f3-canvas-dashboard-01-rep2","call_id":"...:slice:1","model":"z-ai/glm-5.2","provider":"Novita","completion_tokens":14,"usd":0.001845228,"generation_id":"gen-1784042150-bKVlm4usOxPIkjz3xAwv"}
{"run_id":"f3-canvas-dashboard-01-rep2","call_id":"...:slice:2","model":"z-ai/glm-5.2","provider":"StreamLake","completion_tokens":5547,"usd":0.019272264,"artifact_sha256":"7d5816ff099211d9…","generation_id":"gen-1784042163-mNXDpaocxo8gjUeMYyfW"}
{"run_id":"f3-canvas-dashboard-01-rep2","call_id":"...:review:1","model":"openai/gpt-5.6-luna","completion_tokens":749,"usd":0.01618075,"generation_id":"gen-1784042285-n6HN29FA1pm3NRknakxT"}
{"run_id":"f3-canvas-dashboard-01-rep2","call_id":"...:slice:3","model":"z-ai/glm-5.2","completion_tokens":174,"usd":0.002582448,"generation_id":"gen-1784042291-0VT00TXP7FqsaR0rnFdL"}
// tile D total = $0.03988069

// TILE C — floor (GLM one-shot) · p2-review/floor-ceiling/receipts.jsonl line 5
{"id":"canvas-dashboard-01","arm":"floor","model":"z-ai/glm-5.2","require_fp8":true,"artifact_bytes":37765,"prompt_tokens":170,"completion_tokens":14924,"reasoning_tokens":3463,"pinned_cost_usd":0.04369409,"openrouter_usage_cost_usd":0.0436940868,"served_provider":"Novita","served_quantization":"fp8","finish_reason":"stop","gen_id":"gen-1784014462-Ah5kKYbhMcQeWRfWKfny"}

// TILE B — champion (agentic) · a3-runs/a3qc-q-live-canvas-dashboard-01-…dead-s6/receipts.jsonl (8 metered stage rows)
{"stage":"planner","cost":0.021814,"cost_basis":"metered","cumulative":0.021814,"model":"x-ai/grok-4.5","finish_reason":"stop","completion_tokens":2191}
{"stage":"draft","cost":0.057706,"cost_basis":"metered","cumulative":0.07952,"model":"z-ai/glm-5.2","finish_reason":"stop","completion_tokens":16603}
{"stage":"vlm_read","cost":0.004128,"cost_basis":"metered","cumulative":0.083648,"model":"openai/gpt-5.6-luna","vlm_verdict":"refuse"}
{"stage":"reroll","cost":0.033932,"cost_basis":"metered","cumulative":0.11758,"model":"z-ai/glm-5.2","finish_reason":"stop","completion_tokens":9828}
{"stage":"vlm_read","cost":0.006288,"cost_basis":"metered","cumulative":0.123868,"model":"openai/gpt-5.6-luna","vlm_verdict":"accept"}
{"stage":"seat","cost":0.030104,"cost_basis":"metered","cumulative":0.153972,"model":"openai/gpt-5.6-luna","finish_reason":"stop"}
{"stage":"repair","cost":0.072319,"cost_basis":"metered","cumulative":0.226291,"model":"z-ai/glm-5.2","finish_reason":"stop","completion_tokens":10740}
{"stage":"seat","cost":0.030343,"cost_basis":"metered","cumulative":0.256633,"model":"openai/gpt-5.6-luna","finish_reason":"stop"}
// champion manifest cost_total = 0.256634

Rows are trimmed to the load-bearing fields for readability; usd/gen-id/tokens/cost_basis are copied verbatim from the source receipts. Full untrimmed rows live in the run receipts referenced above.

# disclosures

disclosures (6)
  1. mandatoryThe overall experiment was a refutation, not a win. The pre-registered floor-control test concluded the lean slice pipeline did not beat a bare cost-matched one-shot on a majority of tasks (2 of 3: racer + forms). Sealed verdict: REFUTED. This page is the one task where the two cheap reps swept #1 and #2 — a single counter-intuitive case, on the record.
  2. mandatoryVerdicts are n=1-eye illustrative. n=3 tasks, one eye, 2 reps/arm. Directional, not claim-grade / statistical.
  3. Champion cost is higher than its cited label. The brief and rank doc cite “~$0.209”; the run’s own manifest records cost_total = 0.256634 and its rows sum to the same. The “~$0.209” figure is not reproducible; the true metered cost is ~$0.2566. Direction unaffected — the champion was pricier than advertised and still last.
  4. The champion tile came from an honest-fail / dead run. Its source run is quarantined (dead_reason: s6). It was still placed as the champion anchor and blind-ranked #4 — it renders and was critiqued on three concrete bugs; all stage finish-reasons are “stop”. Not a degenerate/empty artifact.
  5. Second-viewing recognition occurred. The eye had previously seen the dashboard floor tile (“I’ve seen this before”). Critiques stayed concrete and quality-grounded; the caveat rides the verdict.
  6. This bundle compares pipeline architectures, not one prompt. Champion = full agentic multi-seat build; floor = bare GLM one-shot; reps = lean spec+harness+contract. Unlike the other two pages, no byte-identical-prompt claim is made here — the manifest prompt_sha256 is null by design.
source: agentic-v3/slice-f3/ — disclosures.md, SPONSOR-RANK-SLICE-2026-07-14.md, BLOCKERS.md (condensed verbatim)