Can an 8× glm-5.3 swarm beat the 4× sonnet swarm at R1?
On round-1 review of three frozen cells, does an 8-seat glm-5.3 swarm match the incumbent 4-seat sonnet swarm — and where does it fail?
The two arms
same three frozen cells8 × glm-5.3
- 24
- seats
- 95
- findings
- 3 of 7
- opus-critical caught
- 0
- fabricated cites
4 × sonnet
- 11
- seats
- 50
- findings
- 2 of 7
- opus-critical caught
- 9
- fabricated cites
glm-5.3 runs on the z.ai coding plan, so an 8-seat swarm burns no subscription quota — twice the seats of the sonnet swarm at no extra pool pressure.
Cell by cell
glm-swarm-8 vs sonnet-swarm-4Both arms read the same live tree and were scored against the same frozen ground truth. Recall is weighted; opus-critical is the count caught of the count present.
| cell | arm | weighted recall | opus-critical | precision | fabricated |
|---|---|---|---|---|---|
| C1D15-REVISED | 8 × glm | 0.458 | 1 of 2 | 0.921 | 0 |
| 4 × sonnet | 0.125 | 0 of 2 | 1.000 | 0 | |
| C2INT7-RT-ZOHO | 8 × glm | 0.250 | 1 of 2 | 1.000 | 0 |
| 4 × sonnet | 0.250 | 1 of 2 | 0.667 | 8 | |
| C3INT7-RT-NOTION | 8 × glm | 0.409 | 1 of 3 | 1.000 | 0 |
| 4 × sonnet | 0.500 | 1 of 3 | 0.706 | 1 |
The metric that decides it — citation discipline
the circuit-breakerThe bench fails any arm that fabricates a citation — the claimed code is absent from the cited file, or the cited line is past the file's end.
9 fabricated citations across C2 + C3 — 7 past end-of-file in a single seat, plus a claimed authentication bypass and a column that exists nowhere in the codebase. FAIL.
0 fabricated citations in 95 findings across 24 seats. PASS.
Verdict
Yes on coverage, decisively on citation discipline — the 8-seat GLM swarm matches or beats the incumbent's opus-critical coverage on every cell, at higher precision and zero fabricated cites, while the sonnet swarm fails the bench's circuit-breaker.
- ·GLM caught 3 of 7 opus-criticals to sonnet's 2 — matching or beating on every cell — at equal or higher precision.
- ·Zero fabricated citations in 95 findings across 24 seats. The incumbent sonnet swarm produced 9 and fails the bench's citation circuit-breaker.
- ·The one loss: GLM under-calls severity, rating the opus-criticals it found 'minor'. Seat it behind a confirm lead that re-weights severity, not more GLM.
- ·Eight GLM seats cost no subscription quota — the swarm scales where the incumbent cannot.
What changed in the harness because of it
- 01An 8-seat orthogonal GLM swarm is non-inferior to the incumbent sonnet swarm on opus-critical coverage — ratified as an R1 review candidate.
- 02Citation discipline is a hard gate: the incumbent sonnet swarm fails it with 9 fabricated cites, one seat placing 7 past end-of-file.
- 03GLM's severity under-call means it is seated behind a confirm lead that re-weights severity, never as the sole R1 voice.
Raw data
generated 2026-09-02 by scripts/gen-glm-swarm.tsEvery number on this page is read from data/glm-swarm-r1.json, which the generator rewrites from these source files. A data update is a commit.
- /raw/spark/scores.csvPer-arm aggregate scores for both swarms across the three cells: weighted recall, opus-critical coverage, precision, anti-hits, severity fidelity.
- /raw/spark/SCORING_RULES.mdThe scoring methodology, including the three-bucket citation policy behind the circuit-breaker. Per-finding detail and the ground-truth answer key are held private.