Can gpt-5.3-codex-spark hold a judgment seat?
Could Spark hold a judgment seat under the same frozen-cell conditions as the comparator arms?
The method
frozen before flightThree frozen cells came from recently-merged units. Ground truth was authored and FROZEN per cell before any arm ran, so no arm could shape the answer key it was scored against.
lead-fill provenance + erasure
webhook auth + channel renewal
first inbound leg on a push-only connector
What it cost
quota ledgerThe number that decides it
Wave 1
quota log- 8
- seats
- 614,380
- tokens
- 0
- findings
Wave 2
quota log- 4
- seats
- 458,808
- tokens
- 0
- findings
An explicit token budget was included — and per-seat cost rose anyway.
12 seats · 0 findings
Weekly exhausted until 2026-09-09.
What we measured instead
all comparator arms freesonnet-swarm-4
- 11
- seats
- 0.125 / 0.250 / 0.500
- recall
- 0-of-2, 1-of-2, 1-of-3
- opus-critical
glm-swarm-8
- 24
- seats
- 0.458 / 0.250 / 0.409
- recall
- 1-of-2, 1-of-2, 1-of-3
- opus-critical
89 verified defects the ground truth missed; 11 arm claims verified false.
GLM rated 3 of the 4 opus-criticals it found “minor” — it finds the severe things and calls them small.
Verdict
Reject on capacity — the bench never measured capability, because Spark ran out of its weekly budget first.
- ·12 seats and 1,073,188 tokens produced 0 findings before the weekly quota hit 98% — the run never reached a capability signal.
- ·Implied ceiling is roughly 1.09M tokens per week — about 1.5 eight-seat swarms — too little to stand a judgment seat on.
- ·Per-seat cost ROSE under an explicit budget (114,702 vs 76,798 avg), so a budget instruction does not make Spark affordable here.
- ·The comparator arms, not Spark, produced this cycle's actionable findings; Spark stays unseated pending a capacity change.
What changed in the harness because of it
- 01Spark is not a candidate for any judgment seat this cycle — capacity, not capability.
- 02The incumbent sonnet swarm fails this bench's own circuit-breaker on citation discipline.
- 03An 8-seat orthogonal GLM swarm is non-inferior to the sonnet swarm on opus-critical coverage.
- 04A swarm without a verifying lead is not seatable: 4 within-arm contradictions vs 2 convergences.
- 05Ground truth is a floor on the defect set, never a ceiling — the arms found 89 defects it missed.
- 06Three defects were found in the bench's OWN ground truth, disclosed rather than repaired.
Raw data
generated 2026-09-02 by scripts/gen-spark.tsEvery number on this page is read from data/spark-seat.json, which the generator rewrites from these source files. A data update is a commit.
- /raw/spark/SPARK_QUOTA_LOG.mdThe quota ledger: per-wave token burn, 0% to 98%, and the implied weekly capacity — the evidence behind the capacity verdict.
- /raw/spark/scores.csvPer-arm aggregate scores: weighted recall, opus-critical coverage, precision, anti-hits, and severity fidelity across the three cells.
- /raw/spark/SCORING_RULES.mdThe scoring methodology actually applied: match / recall / precision definitions and the three-bucket citation policy. Cell-level defect detail and the ground-truth answer key are held private.