vrinda.dev / showdown
‹ all comparisons

01 / showdown · T9-racing-car

5 builders. One prompt. Ranked.

One byte-identical racing-game prompt, fired at five different builders — one independent invocation per builder (the Kimi harness completed in two turns). The tiles went onto a blind board with the labels hidden. The eye’s verbatim rank, best to worst: D, B, A, C, E.

# byte-identical prompt

task-card sha-256

2c0f286382c2580e2c88496582c31b618115d694648405f44b7baf91036b7763

The same 785-byte card, hashed to this sha-256, was fired byte-for-byte at all five builders — the hash is proof no arm quietly got an easier or richer prompt. (It also matches the prompt on the glm-vs-sonnet page — same base T9 card.)

show the exact prompt
Build a playable single-file HTML 3D car racing game using exactly ONE external script:
the Three.js CDN (no other dependencies, no import maps, no modules). Requirements:
a road with side barriers scrolling toward the camera; a player car controlled with
arrow keys / A-D (lane steering) that feels responsive; oncoming obstacle cars spawning
at increasing frequency; collision ends the run with a game-over panel and Restart
button (restart fully resets state); score = distance, increasing over time, shown in
a HUD; speed ramps up gradually; the FIRST SCREEN is the actual playable game already
running (no start overlay); mobile-responsive canvas; and a DOM-mirrored QA hook
<script id="qa-state" type="application/json"> updated every ~500ms with
{state, score, speed, crashes}.

785 bytes · sha-256 verified equal to the sealed card_sha256 in mapping-SEALED.json

# sealed blind rank

ranktilebuilderchannelartifactmetered cost
1 D Opus subagent, one-call self-contained HTML
arm-2-opus-sub
subscription 42,483 B
sha 54e378fd…
sub · unmetered
2 B moonshotai/kimi-k3 agentic in claude -p harness
arm-5-kimi-harness · 2 turns
OpenRouter 16,704 B
sha e91f7e06…
$0.283437est
3 A gpt-5.6-sol via codex exec, one-call
arm-3-sol-sub
Codex pool 22,807 B
sha 667331bb…
separate pool
4 C moonshotai/kimi-k3 raw OpenRouter one-call
arm-4-kimi-raw
OpenRouter 21,351 B
sha 365195ea…
~$0.12est
usd null · 7,015 out-tok
5 E Fable 5 (opus) orchestrator, one-call
arm-1-fable
subscription 15,164 B
sha 73452160…
sub · unmetered

Tile E was potentially glimpsed pre-blind (the browser pane auto-opened it during authoring); it placed last anyway. See disclosures. Sizes are exact bytes of the shipped files; costs are from manifest.json (est = token-estimated or null in receipt).

verbatim rank & reveal: kimi3-probe/showdown/VERDICT-REVEAL.md · sealed A–E map: showdown/board/mapping-SEALED.json

# open the artifacts

#1 · D
Opus subagent
one-call self-contained HTML · subscription
sub · unmetered

artifact loads once published

sha 54e378fd…open full screen ↗
#2 · B
Kimi-k3 agentic (harness)
claude -p, 2 turns · OpenRouter
$0.2834 est

artifact loads once published

sha e91f7e06…open full screen ↗
#3 · A
gpt-5.6-sol via codex
one-call · Codex pool
separate pool

artifact loads once published

sha 667331bb…open full screen ↗
#4 · C
Kimi-k3 raw
one-call · OpenRouter
~$0.12 est

artifact loads once published

sha 365195ea…open full screen ↗
#5 · E
Fable 5 orchestrator
one-call · subscription · glimpse-disclosed
sub · unmetered

artifact loads once published

sha 73452160…open full screen ↗

Previews are sandboxed (sandbox="allow-scripts", no same-origin access) and load on click. Until the artifacts are published, each tile shows its placeholder and its open full screen link.

# receipts

tile / armmodeltokens (in / out)durationusdusd_source
D
arm-2-opus-sub
Opus subagent $0 subscription ($0 API; no receipt row)
B
arm-5-kimi-harness
moonshotai/kimi-k3 33,129 / 12,270 468 s $0.283437est estimate(token*kimi-pin)
claude_reported_usd = 0.488713
A
arm-3-sol-sub
gpt-5.6-sol 90 s $0 separate-pool (codex/chatgpt-plan; no per-call USD)
C
arm-4-kimi-raw
moonshotai/kimi-k3 — / 7,015 162.4 s nullnull estimate (~$0.12 per reveal, 7,015 out-tok)
E
arm-1-fable
Fable 5 (opus) $0 subscription ($0 API; no receipt row)

Two arms (D, E) ran on subscription and have no metered receipt row — recorded as $0 API, flagged for auditability, not a data loss. Tile A ran in the Codex/ChatGPT separate quota pool with no per-call dollar figure. Tokens shown are the build-leg values from the receipts.

show all 7 verbatim receipt rows
{"arm": "kimi-raw", "model": "moonshotai/kimi-k3", "finish": "stop", "out_tok": 7015, "usd": null, "usd_source": "estimate", "provider": null, "dur_s": 162.4, "ts": "2026-07-16T22:13:55Z"}

{"wave":"showdown","arm":"arm-3-sol-sub","leg":"dry-probe","model":"gpt-5.6-sol","result":"PASS","note":"echo token matched via -o capture; separate quota pool ($0 metered)","ts":"2026-07-16T22:19:04Z"}

{"wave": "showdown", "arm": "arm-5-kimi-harness", "leg": "penny-probe", "model": "moonshotai/kimi-k3", "verdict": "BLOCKED", "rc": 1, "reply": "", "input_tokens": null, "output_tokens": null, "claude_reported_usd": null, "kimi_estimate_usd": null, "usd_source": "estimate(token*kimi-pin); OpenRouter gen_id not exposed by harness json", "err_tail": ["Error: Invalid MCP configuration:", "mcpServers: Invalid input: expected record, received undefined"], "ts": "2026-07-16T22:19:20Z"}

{"wave": "showdown", "arm": "arm-5-kimi-harness", "leg": "penny-probe", "model": "moonshotai/kimi-k3", "verdict": "BLOCKED", "rc": 0, "reply": "", "input_tokens": 20477, "output_tokens": 47, "claude_reported_usd": 0.10381599999999999, "kimi_estimate_usd": 0.062136, "usd_source": "estimate(token*kimi-pin); OpenRouter gen_id not exposed by harness json", "err_tail": [], "ts": "2026-07-16T22:20:17Z"}

{"wave": "showdown", "arm": "arm-5-kimi-harness", "leg": "penny-probe-AUTHORITATIVE", "model": "moonshotai/kimi-k3", "verdict": "WIRED", "served_model_confirmed": true, "served_model": "moonshotai/kimi-k3", "subtype": "success", "is_error": false, "stop_reason": "end_turn", "num_turns": 1, "terminal_reason": "completed", "input_tokens": 19762, "output_tokens": 42, "cache_read_input_tokens": 1280, "duration_ms": 115199, "claude_reported_usd": 0.10049999999999999, "usd_source": "api(total_cost_usd + modelUsage.costUSD from claude -p json)", "text_status": "empty-visible-text (42 out tok, reasoning-only; expected kimi-k3 behavior, NOT a wiring fault)", "ts": "2026-07-16T22:24:43Z"}

{"wave":"showdown","arm":"arm-3-sol-sub","leg":"build","model":"gpt-5.6-sol","rc":0,"duration_s":90,"artifact":"arm-A.html","usd_source":"separate-pool(codex/chatgpt-plan; no per-call USD)","ts":"2026-07-16T22:31:56Z"}

{"wave": "showdown", "arm": "arm-5-kimi-harness", "leg": "build", "model": "moonshotai/kimi-k3", "rc": 0, "duration_s": 468, "artifact": "arm-B.html", "bytes": 16704, "input_tokens": 33129, "output_tokens": 12270, "num_turns": 2, "claude_reported_usd": 0.4887129999999999, "kimi_estimate_usd": 0.283437, "usd_source": "estimate(token*kimi-pin)", "ts": "2026-07-16T22:38:17Z"}

source: showdown/receipts.jsonl (7 rows). Probe rows (dry-probe / penny-probe) are process rows, not build artifacts. Slots D & E have no rows ($0 subscription).

# disclosures

disclosures (5)
  1. mandatoryTile E pre-blind glimpse. Tile E (Fable orchestrator) was potentially glimpsed pre-blind — the browser pane auto-opened it during authoring. It placed last (#5) anyway; the source verdict argues the disclosure changes nothing favorable.
  2. mandatorySome USD values are estimates or null. A (gpt-5.6-sol): $0 metered, separate Codex pool, no per-call USD. B (kimi-k3 agentic): estimate token*kimi-pin = $0.283437 (claude_reported_usd = 0.488713); OpenRouter gen-id not exposed by the harness. C (kimi-k3 raw): usd = null, estimate ~$0.12, 7,015 out-tok. D & E: $0 API subscription, no receipt row.
  3. mandatoryn=1 task, one eye — directional, not claim-grade. Stated in the reveal. This board measured a ceiling on one task; it is not a statistical result.
  4. Unsealed duplicates excluded. A parallel non-blind fire wrote duplicate copies (arm-kimi / arm-fable / arm-opus.html) into the same directory with no recorded card sha. Those were not adopted into the sealed slots and are excluded; only sealed slots arm-A…arm-E ship here.
  5. Card integrity (positive). The shipped card’s sha-256 2c0f2863… equals the sealed card_sha256 in mapping-SEALED.json — the identical byte-for-byte card fired at all five arms.
source: kimi3-probe/showdown/ — disclosures.md, VERDICT-REVEAL.md, BLOCKERS.md (condensed verbatim)