Study 06 · published 2026-09-02Ratified · R1 candidate

Can an 8× glm-5.3 swarm beat the 4× sonnet swarm at R1?

On round-1 review of three frozen cells, does an 8-seat glm-5.3 swarm match the incumbent 4-seat sonnet swarm — and where does it fail?

The two arms

same three frozen cells

8 × glm-5.3

challenger
24
seats
95
findings
3 of 7
opus-critical caught
0
fabricated cites

4 × sonnet

incumbent
11
seats
50
findings
2 of 7
opus-critical caught
9
fabricated cites

glm-5.3 runs on the z.ai coding plan, so an 8-seat swarm burns no subscription quota — twice the seats of the sonnet swarm at no extra pool pressure.

Cell by cell

glm-swarm-8 vs sonnet-swarm-4

Both arms read the same live tree and were scored against the same frozen ground truth. Recall is weighted; opus-critical is the count caught of the count present.

cellarmweighted recallopus-criticalprecisionfabricated
C1D15-REVISED8 × glm0.4581 of 20.9210
4 × sonnet0.1250 of 21.0000
C2INT7-RT-ZOHO8 × glm0.2501 of 21.0000
4 × sonnet0.2501 of 20.6678
C3INT7-RT-NOTION8 × glm0.4091 of 31.0000
4 × sonnet0.5001 of 30.7061

The metric that decides it — citation discipline

the circuit-breaker

The bench fails any arm that fabricates a citation — the claimed code is absent from the cited file, or the cited line is past the file's end.

4 × sonnetFAIL

9 fabricated citations across C2 + C3 — 7 past end-of-file in a single seat, plus a claimed authentication bypass and a column that exists nowhere in the codebase. FAIL.

8 × glm-5.3PASS

0 fabricated citations in 95 findings across 24 seats. PASS.

Verdict

Ratified · 2026-09-02

Yes on coverage, decisively on citation discipline — the 8-seat GLM swarm matches or beats the incumbent's opus-critical coverage on every cell, at higher precision and zero fabricated cites, while the sonnet swarm fails the bench's circuit-breaker.

  • ·GLM caught 3 of 7 opus-criticals to sonnet's 2 — matching or beating on every cell — at equal or higher precision.
  • ·Zero fabricated citations in 95 findings across 24 seats. The incumbent sonnet swarm produced 9 and fails the bench's citation circuit-breaker.
  • ·The one loss: GLM under-calls severity, rating the opus-criticals it found 'minor'. Seat it behind a confirm lead that re-weights severity, not more GLM.
  • ·Eight GLM seats cost no subscription quota — the swarm scales where the incumbent cannot.

What changed in the harness because of it

  1. 01An 8-seat orthogonal GLM swarm is non-inferior to the incumbent sonnet swarm on opus-critical coverage — ratified as an R1 review candidate.
  2. 02Citation discipline is a hard gate: the incumbent sonnet swarm fails it with 9 fabricated cites, one seat placing 7 past end-of-file.
  3. 03GLM's severity under-call means it is seated behind a confirm lead that re-weights severity, never as the sole R1 voice.

Raw data

generated 2026-09-02 by scripts/gen-glm-swarm.ts

Every number on this page is read from data/glm-swarm-r1.json, which the generator rewrites from these source files. A data update is a commit.

  • /raw/spark/scores.csvPer-arm aggregate scores for both swarms across the three cells: weighted recall, opus-critical coverage, precision, anti-hits, severity fidelity.
  • /raw/spark/SCORING_RULES.mdThe scoring methodology, including the three-bucket citation policy behind the circuit-breaker. Per-finding detail and the ground-truth answer key are held private.

I. Nocturne in E♭ major Op. 9 № 2 · Chopin

0:00 / 4:16