read this first — mandatory framing
This page shows one striking counter-case, not a general win. The overall experiment it comes from (slice-f3) was REFUTED: the lean “slice” pipeline did not beat a bare cost-matched one-shot on a majority of tasks (2 of 3 — the racer and the forms task).
Do not read this as the cheap recipe broadly winning. It is canvas-dashboard-01 only, where the two cheap reps happened to sweep. n=3 tasks, one eye, 2 reps/arm — directional, not claim-grade.
# blind rank & reveal
| rank | tile | role | artifact | metered cost |
|---|---|---|---|---|
| 1 | A | SLICE rep1 lean spec + harness + contract |
16,756 B sha 7d609e56… |
$0.0431metered |
| 2 | D | SLICE rep2 lean spec + harness + contract |
18,158 B sha 7d5816ff… |
$0.0399metered |
| 3 | C | floor — GLM one-shot bare cost-matched baseline |
37,765 B sha 53aee12a… |
$0.0437metered |
| 4 | B | CHAMPION — full agentic pipeline † multi-seat build · honest-fail / dead(s6) |
32,486 B sha 268981fd… |
$0.256634metered |
† Champion cost, corrected honestly: the brief and rank doc cite the champion at “~$0.209”, but that figure is not reproducible from receipts. The champion run’s own manifest records cost_total = 0.256634 and its stage rows sum to the same — so the true metered cost is $0.256634, shown here. The inversion holds a fortiori: the champion was even pricier than advertised and still ranked last. It also came from an honest-fail / dead(s6) run — quarantined, but it renders and was critiqued on three concrete bugs (not a degenerate artifact).
blind rank source: agentic-v3/slice-f3/SPONSOR-RANK-SLICE-2026-07-14.md · sha map: slice-f3/board/final/out-private/mapping.json (all four tile shas recomputed and match)# the blind read (verbatim)
When it comes to the first task, that is, the canvas dashboard, from a visual richness point of view:
- Tile B is good and well planned. There are a couple of minor bugs or minor issues: there is a static number shown for each product. The static number is not automatically changing on the date range click, but it is changing when I click back on the product again. That means this is trying to show the right numbers, but I think the sync between the number showcase and the date range click is not proper. That is one bug, I would say.
- Obviously, earlier I reported the same issue: when I click on the Remove Check buttons, the graph is not recalibrating or removing the bar. That's it. It is not a bad thing, but it doesn't feel the same way that others attempt the experience.
- The third issue is that the best day number is shown, but the day is, I think, repeated twice: 7 by 4, 7 by 4. I don't know what it is. I would also classify this as a mistake that I haven't reported before, but with visual richness, I cannot accept its mistakes, so I'm keeping it down.
Now, when it comes to tile D, design-wise, it is basic, but it is doing the job. It's professional-looking, but color-wise, it is not very polished. Most of it is all violet color, but it is working from a logic point of view. Here also, the best day is given as day 28. Usually, here I used to see the actual date. Here, it is showing day 28, which means it is thinking in a different way. That is about tile D and tile C. I've seen this before. There is a rendering mistake, even though the design is very good. The trend rendering did not happen well. That means it feels like a blank screen. The trend line is not shown. Otherwise, if the trend line is shown, this would have been one of the best visually rich designs. Again, at the bottom, it is showing implementation notes and stuff. Hence, this is also having some problems in terms of how it is planned.
Till now, we discussed tile B, tile C, and tile D. Everything has some minor problems. Nothing is perfected. When it comes to the remaining tile A, here also, structure-wise, it is working, and color-wise, it is also decent, but the best day is showing 7/14. Usually, I would expect to say, let's say, July 14 or something, but okay, at least this is not a mistake, in my opinion. From a completeness point of view, tile A actually is complete. Even though it is not rich, it is decent. From tile A, I will give the first rating, then I will give the next rating to tile D, then I'll come to tile C. Earlier, I thought tile B is good, but when I looked at all three mistakes, maybe if I discount only the trend line not appearing, I think tile C is looking good now. Then tile B, maybe I'll rank tile B last, even though it is good from a rich UI point of view. I'm unable to classify it as the best because it has three mistakes. I can see simple mistakes, but three is not there. My rating again is tile A first, then tile D, then tile C, then tile B.
# open the artifacts
artifact loads once published
artifact loads once published
artifact loads once published
artifact loads once published
Previews are sandboxed (sandbox="allow-scripts", no same-origin access) and load on click. Until the artifacts are published, each tile shows its placeholder and its open full screen link.
# receipts
| tile | role | models | metered usd | usd_source | gen-id |
|---|---|---|---|---|---|
| A | rep1 (slice) | glm-5.2 ×3 + luna review | $0.04309749 | metered usd == true_up.actual_usd |
gen-…jTR7z4 blessing slice:3 |
| D | rep2 (slice) | glm-5.2 ×3 + luna review | $0.03988069 | metered usd == true_up.actual_usd |
gen-…0VT00T blessing slice:3 |
| C | floor (GLM one-shot) | z-ai/glm-5.2 (fp8, Novita) | $0.04369409 | metered openrouter == gen_stats |
gen-…Ah5kKY |
| B | champion (agentic) | grok-4.5 + glm-5.2 + luna | $0.256634 | metered cost_basis = metered |
— a3 rows carry no gen-id |
Champion 11-row receipt, summarized — 3 setup rows (run_start, watchdog_envelope, vision_preflight) + 8 metered stages. The pipeline drafted, was refused by its own vision check, re-rolled, was accepted, seated, repaired, and re-seated:
| # | stage | model | cost | cumulative | note |
|---|---|---|---|---|---|
| 1 | planner | x-ai/grok-4.5 | $0.021814 | $0.021814 | 2,191 out-tok |
| 2 | draft | z-ai/glm-5.2 | $0.057706 | $0.079520 | 16,603 out-tok |
| 3 | vlm_read | gpt-5.6-luna | $0.004128 | $0.083648 | verdict: refuse |
| 4 | reroll | z-ai/glm-5.2 | $0.033932 | $0.117580 | 9,828 out-tok |
| 5 | vlm_read | gpt-5.6-luna | $0.006288 | $0.123868 | verdict: accept |
| 6 | seat | gpt-5.6-luna | $0.030104 | $0.153972 | — |
| 7 | repair | z-ai/glm-5.2 | $0.072319 | $0.226291 | 10,740 out-tok |
| 8 | seat | gpt-5.6-luna | $0.030343 | $0.256633 | final |
Eight metered stages sum to $0.256633 cumulative; the manifest records cost_total = 0.256634. That is ~6× the ~4-cent lean reps that outranked it. Cost was never the problem — the mean receipted $/run over all six slice-f3 runs was $0.059387, under the $0.080 bar.
show all verbatim receipt rows (reps A/D, floor C, champion B)
// TILE A — rep1 (slice) · slice-f3/money/receipts.jsonl lines 32,36,40,41
{"run_id":"f3-canvas-dashboard-01-rep1","call_id":"...:slice:1","arm":"slice","model":"z-ai/glm-5.2","provider":"StreamLake","prompt_tokens":1953,"completion_tokens":14,"usd":0.001845228,"status":"SUCCESS","generation_id":"gen-1784041629-tPmxaEv8AksK3MBsTKsJ","cost_basis":"max-provider-and-live-pin"}
{"run_id":"f3-canvas-dashboard-01-rep1","call_id":"...:slice:2","arm":"slice","model":"z-ai/glm-5.2","provider":"StreamLake","prompt_tokens":3424,"completion_tokens":5146,"usd":0.01810776,"status":"SUCCESS","generation_id":"gen-1784041630-bxo2e0Q8YTcYT5i0himD","artifact_sha256":"7d609e567029e1ba…"}
{"run_id":"f3-canvas-dashboard-01-rep1","call_id":"...:review:1","arm":"review","model":"openai/gpt-5.6-luna","provider":"OpenAI","prompt_tokens":9034,"completion_tokens":1627,"usd":0.02105375,"status":"SUCCESS","generation_id":"gen-1784041780-L7h2js5h0JAHw7hB1pew"}
{"run_id":"f3-canvas-dashboard-01-rep1","call_id":"...:slice:3","arm":"slice","model":"z-ai/glm-5.2","provider":"StreamLake","prompt_tokens":2203,"completion_tokens":19,"usd":0.002090748,"status":"SUCCESS","generation_id":"gen-1784041794-jTR7z4yindFsTVvTPXdi"}
// tile A total (blessing chain) = $0.04309749
// TILE D — rep2 (slice) · slice-f3/money/receipts.jsonl lines 47,48,52,53
{"run_id":"f3-canvas-dashboard-01-rep2","call_id":"...:slice:1","model":"z-ai/glm-5.2","provider":"Novita","completion_tokens":14,"usd":0.001845228,"generation_id":"gen-1784042150-bKVlm4usOxPIkjz3xAwv"}
{"run_id":"f3-canvas-dashboard-01-rep2","call_id":"...:slice:2","model":"z-ai/glm-5.2","provider":"StreamLake","completion_tokens":5547,"usd":0.019272264,"artifact_sha256":"7d5816ff099211d9…","generation_id":"gen-1784042163-mNXDpaocxo8gjUeMYyfW"}
{"run_id":"f3-canvas-dashboard-01-rep2","call_id":"...:review:1","model":"openai/gpt-5.6-luna","completion_tokens":749,"usd":0.01618075,"generation_id":"gen-1784042285-n6HN29FA1pm3NRknakxT"}
{"run_id":"f3-canvas-dashboard-01-rep2","call_id":"...:slice:3","model":"z-ai/glm-5.2","completion_tokens":174,"usd":0.002582448,"generation_id":"gen-1784042291-0VT00TXP7FqsaR0rnFdL"}
// tile D total = $0.03988069
// TILE C — floor (GLM one-shot) · p2-review/floor-ceiling/receipts.jsonl line 5
{"id":"canvas-dashboard-01","arm":"floor","model":"z-ai/glm-5.2","require_fp8":true,"artifact_bytes":37765,"prompt_tokens":170,"completion_tokens":14924,"reasoning_tokens":3463,"pinned_cost_usd":0.04369409,"openrouter_usage_cost_usd":0.0436940868,"served_provider":"Novita","served_quantization":"fp8","finish_reason":"stop","gen_id":"gen-1784014462-Ah5kKYbhMcQeWRfWKfny"}
// TILE B — champion (agentic) · a3-runs/a3qc-q-live-canvas-dashboard-01-…dead-s6/receipts.jsonl (8 metered stage rows)
{"stage":"planner","cost":0.021814,"cost_basis":"metered","cumulative":0.021814,"model":"x-ai/grok-4.5","finish_reason":"stop","completion_tokens":2191}
{"stage":"draft","cost":0.057706,"cost_basis":"metered","cumulative":0.07952,"model":"z-ai/glm-5.2","finish_reason":"stop","completion_tokens":16603}
{"stage":"vlm_read","cost":0.004128,"cost_basis":"metered","cumulative":0.083648,"model":"openai/gpt-5.6-luna","vlm_verdict":"refuse"}
{"stage":"reroll","cost":0.033932,"cost_basis":"metered","cumulative":0.11758,"model":"z-ai/glm-5.2","finish_reason":"stop","completion_tokens":9828}
{"stage":"vlm_read","cost":0.006288,"cost_basis":"metered","cumulative":0.123868,"model":"openai/gpt-5.6-luna","vlm_verdict":"accept"}
{"stage":"seat","cost":0.030104,"cost_basis":"metered","cumulative":0.153972,"model":"openai/gpt-5.6-luna","finish_reason":"stop"}
{"stage":"repair","cost":0.072319,"cost_basis":"metered","cumulative":0.226291,"model":"z-ai/glm-5.2","finish_reason":"stop","completion_tokens":10740}
{"stage":"seat","cost":0.030343,"cost_basis":"metered","cumulative":0.256633,"model":"openai/gpt-5.6-luna","finish_reason":"stop"}
// champion manifest cost_total = 0.256634
Rows are trimmed to the load-bearing fields for readability; usd/gen-id/tokens/cost_basis are copied verbatim from the source receipts. Full untrimmed rows live in the run receipts referenced above.
# disclosures
disclosures (6)
- mandatoryThe overall experiment was a refutation, not a win. The pre-registered floor-control test concluded the lean slice pipeline did not beat a bare cost-matched one-shot on a majority of tasks (2 of 3: racer + forms). Sealed verdict: REFUTED. This page is the one task where the two cheap reps swept #1 and #2 — a single counter-intuitive case, on the record.
- mandatoryVerdicts are n=1-eye illustrative. n=3 tasks, one eye, 2 reps/arm. Directional, not claim-grade / statistical.
- Champion cost is higher than its cited label. The brief and rank doc cite “~$0.209”; the run’s own manifest records cost_total = 0.256634 and its rows sum to the same. The “~$0.209” figure is not reproducible; the true metered cost is ~$0.2566. Direction unaffected — the champion was pricier than advertised and still last.
- The champion tile came from an honest-fail / dead run. Its source run is quarantined (dead_reason: s6). It was still placed as the champion anchor and blind-ranked #4 — it renders and was critiqued on three concrete bugs; all stage finish-reasons are “stop”. Not a degenerate/empty artifact.
- Second-viewing recognition occurred. The eye had previously seen the dashboard floor tile (“I’ve seen this before”). Critiques stayed concrete and quality-grounded; the caveat rides the verdict.
- This bundle compares pipeline architectures, not one prompt. Champion = full agentic multi-seat build; floor = bare GLM one-shot; reps = lean spec+harness+contract. Unlike the other two pages, no byte-identical-prompt claim is made here — the manifest prompt_sha256 is null by design.