# Spark pool consumption log (bench-measured)

The quota ledger behind the study's verdict. Token counts are read from each seat's own reported
rate-limit record. The code under review, and the per-seat findings, are held privately; this file
publishes only the capacity measurement.

## Baseline (2026-09-01 18:38 PT / 2026-09-02 01:38Z)
A probe session on the Spark pool (Codex Pro plan) read the rate limits before any wave ran:
- 5h window (300 min): **0.0% used**, resets 2026-09-02T06:38Z
- weekly (10080 min): **0.0% used**, resets 2026-09-09T01:38Z

This pool is distinct from the MAIN Codex pool (which stood high on its own separate weekly at the time),
so these bars are Spark's own capacity, not the shared account's.

## WAVE 1 — 8 concurrent seats, high reasoning effort (2026-09-02 01:46–01:47Z)
Result: **all 8 seats killed by "You've hit your usage limit for GPT-5.3-Codex-Spark" — 0 findings produced.**

| seat | scope | tokens used | outcome |
|---|---|--:|---|
| s1 | correctness | 64,382 | limit error mid-read |
| s2 | consent/DNC | 109,803 | limit error mid-read |
| s3 | SQL / migrations | 90,574 | limit error mid-read |
| s4 | concurrency / idempotency | 93,976 | limit error mid-read |
| s5 | error handling | 69,400 | limit error mid-read |
| s6 | API / type-contract drift | 53,936 | limit error mid-read |
| s7 | security | 59,302 | limit error mid-read |
| s8 | resource / lifecycle | 73,007 | limit error during remote context compaction |
| **total** | | **614,380** | **0 findings** |

Pool after the wave: **5h window 100.0% used · weekly 49.0% used.**

### Interpretation
- ~614k Spark tokens ≈ 100% of one 5h window ≈ 49% of the Spark weekly.
- Implied Spark weekly capacity ≈ **1.25M tokens ≈ 2 waves of this shape per week**, not one swarm per
  review round.
- Per-seat cost driver: each seat independently reads the cell's full diff, then explores the repository at
  high effort (seats were observed pulling in large migration-history files) — i.e. 8× duplicated context
  acquisition with no shared cache across seats.

---
## WAVE 2 — 4 seats (halved), WITH an explicit token-budget instruction (2026-09-02 06:56–06:58Z)
The seat brief added one line: "BUDGET: do NOT read whole large files — use grep/ripgrep and ranged reads to
open only the cited regions; keep total reading under ~40k tokens." Scope, materials and the finding
contract were otherwise byte-identical to wave 1.

Result: **all 4 seats killed by the usage wall — 0 findings, again.**

| seat | scope | tokens used | outcome |
|---|---|--:|---|
| s1 | correctness | 90,672 | limit error |
| s3 | SQL / migrations | 191,742 | limit error |
| s5 | error handling | 69,732 | limit error |
| s7 | security | 106,662 | limit error |
| **total** | | **458,808** | **0 findings** |

Pool after wave 2: **5h 100% · weekly 98%.** Weekly resets 2026-09-09.

### The two waves together — this is the finding
| | seats | tokens | findings | Spark weekly |
|---|--:|--:|--:|---|
| wave 1 | 8 | 614,380 | 0 | 0% → 49% |
| wave 2 (halved + budget line) | 4 | 458,808 | 0 | 49% → 98% |
| **total** | **12** | **1,073,188** | **0** | **exhausted** |

Three things follow, and they are stronger than wave 1 alone could support:
1. **Halving the seat count did not halve the cost.** 4 seats consumed the same ~49% weekly share as 8 did.
2. **Per-seat cost went UP, not down** — 114,702 avg in wave 2 vs 76,798 in wave 1 — despite the explicit
   40k reading budget. The seats run until the wall kills them; the instruction did not constrain them.
   One seat alone burned 191,742 tokens.
3. **Implied Spark weekly capacity ≈ 1.09M tokens**, consistent with wave 1's ≈1.25M estimate — roughly
   **one and a half 8-seat review swarms per WEEK**, on one cell.

⚠ Consequence: the Spark capability question **cannot be answered this week**. The weekly budget is
exhausted until 2026-09-09 with zero findings produced across 12 seats and two independent attempts. No
capability data exists for any of the four seats Spark was being tested for.

### Operator error, disclosed
The first wave-2 launch failed instantly on all 4 seats because of a command-line argument-ordering mistake
in the relaunch script. It consumed NO quota (the CLI rejected the arguments before any model call) and was
relaunched six minutes later. Root cause: the wave-2 script changed the launch form from wave 1's working
version and was armed on a timer without being smoke-tested against a single seat first.
