Study 05 · published 2026-09-02Rejected · capacity

Can gpt-5.3-codex-spark hold a judgment seat?

Could Spark hold a judgment seat under the same frozen-cell conditions as the comparator arms?

The method

frozen before flight

Three frozen cells came from recently-merged units. Ground truth was authored and FROZEN per cell before any arm ran, so no arm could shape the answer key it was scored against.

C19.5k-line diff
D15-REVISED

lead-fill provenance + erasure

C29.0k
INT7-RT-ZOHO

webhook auth + channel renewal

C313.8k
INT7-RT-NOTION

first inbound leg on a push-only connector

Ground truth36 musts31 anti-hitsAuthored and FROZEN per cell before any arm ran.

What it cost

quota ledger

The number that decides it

Wave 1

quota log
8
seats
614,380
tokens
0
findings
Spark weekly0% → 49%

Wave 2

quota log
4
seats
458,808
tokens
0
findings

An explicit token budget was included — and per-seat cost rose anyway.

Spark weekly49% → 98%
Total
1,073,188tokens

12 seats · 0 findings

Weekly exhausted until 2026-09-09.

Implied capacity
≈ 1.09M tokens / week — ≈ 1.5 eight-seat swarms per week
Per-seat cost
Per-seat cost ROSE in wave 2: 114,702 avg vs 76,798, despite the budget instruction.

What we measured instead

all comparator arms free

sonnet-swarm-4

incumbent
11
seats
0.125 / 0.250 / 0.500
recall
0-of-2, 1-of-2, 1-of-3
opus-critical
9 fabricated citations

glm-swarm-8

orthogonal
24
seats
0.458 / 0.250 / 0.409
recall
1-of-2, 1-of-2, 1-of-3
opus-critical
0 fabricated in 95 findings · 1 anti-hit
glm-floor-5.34 samples · 0 fabrications · precision 0.75–1.00
glm-4.7 probe2 fabricated citations in 6 findings — the existing ban holds
Across all arms

89 verified defects the ground truth missed; 11 arm claims verified false.

Severity calibration

GLM rated 3 of the 4 opus-criticals it found “minor” — it finds the severe things and calls them small.

Verdict

Rejected · 2026-09-02

Reject on capacity — the bench never measured capability, because Spark ran out of its weekly budget first.

  • ·12 seats and 1,073,188 tokens produced 0 findings before the weekly quota hit 98% — the run never reached a capability signal.
  • ·Implied ceiling is roughly 1.09M tokens per week — about 1.5 eight-seat swarms — too little to stand a judgment seat on.
  • ·Per-seat cost ROSE under an explicit budget (114,702 vs 76,798 avg), so a budget instruction does not make Spark affordable here.
  • ·The comparator arms, not Spark, produced this cycle's actionable findings; Spark stays unseated pending a capacity change.

What changed in the harness because of it

  1. 01Spark is not a candidate for any judgment seat this cycle — capacity, not capability.
  2. 02The incumbent sonnet swarm fails this bench's own circuit-breaker on citation discipline.
  3. 03An 8-seat orthogonal GLM swarm is non-inferior to the sonnet swarm on opus-critical coverage.
  4. 04A swarm without a verifying lead is not seatable: 4 within-arm contradictions vs 2 convergences.
  5. 05Ground truth is a floor on the defect set, never a ceiling — the arms found 89 defects it missed.
  6. 06Three defects were found in the bench's OWN ground truth, disclosed rather than repaired.

Raw data

generated 2026-09-02 by scripts/gen-spark.ts

Every number on this page is read from data/spark-seat.json, which the generator rewrites from these source files. A data update is a commit.

  • /raw/spark/SPARK_QUOTA_LOG.mdThe quota ledger: per-wave token burn, 0% to 98%, and the implied weekly capacity — the evidence behind the capacity verdict.
  • /raw/spark/scores.csvPer-arm aggregate scores: weighted recall, opus-critical coverage, precision, anti-hits, and severity fidelity across the three cells.
  • /raw/spark/SCORING_RULES.mdThe scoring methodology actually applied: match / recall / precision definitions and the three-bucket citation policy. Cell-level defect detail and the ground-truth answer key are held private.

I. Nocturne in E♭ major Op. 9 № 2 · Chopin

0:00 / 4:16