Can GLM hold more seats?
With the Pro plan's 6× quota, can GLM hold more seats than the two already ratified — build, discovery lattice, review lens, orient, web — without ever owning a judgment call?
More seats — but only additive ones, and only where an oracle or a stronger model owns the judgment.
GLM Coding Pro (6× quota) landed 2026-08-21. The owner ordered n≥2 testing of every candidate seat with a hard-fail override; 21 GLM trials + 2 opus comparators settled which new seats GLM can hold.
The program in numbers
21 GLM trials · 2 comparatorsThe verdict sheet
One row per candidate seat after two waves of testing. The status column pairs its tone with the ruling text — no seat is decided by color alone.
| seat | evidence | status |
|---|---|---|
| checkpoint_confirm + loop_iterate | Ratified pre-Pro; held/rose under enforcement; lead re-triage + codex gates. | Active (canary) |
| discovery union (additive lattice) | n=2: pooled 86.75 vs opus 73.35; 0 anti-hits; GO gates pass. | Adopt — coverage layer + retained opus depth |
| web research lane | n=2: W2 100 + W1-v2 88, crux EXACT both runs, 0 fabrications (mean 94). | Adopt — additive web lane |
| build_rote | 6 cells, 4/6 hard-fail, defects self-defended; mean ~57.5. | Rejected (final, n=6) |
| r1_additive_lens | 2×100% precision + 1 fabricated provider-must. | Fails strict rule; code-only-claims redesign possible |
| orient / auto_proxy | Mean ~64; tier reversal + minor fabricated ids; no hard-fails. | Weak — no seat recommended |
| any confirming / judgment seat | Consistent severity/judgment failures across benches. | Never (unchanged) |
Every scored cell, both waves
n≥2 per categoryEach cell graded 0–100 (or a descriptive recall string), tool-enforced in a frozen worktree. A hard-fail override — any anti-hit, false-clean, or fence-invisible must-defect — fails the seat regardless of its mean; hard-fail rows are marked in the danger color and labelled.
| cell | task | category | wave | grade | result | outcome |
|---|---|---|---|---|---|---|
| Q1 | Q1 checkpoint-confirm re-test | review | wave-0 | 1/2 musts corr-sev | pass | HELD — reproduced the original exactly; 0 hallucinations. |
| Q2 | Q2 checkpoint-confirm re-test | review | wave-0 | 1/4 musts corr-sev | pass | ROSE — enforcement fixed a downgrade + recovered a should; 3/4 musts still missed. |
| Q5 | Q5 confirming-tier re-test | review | wave-0 | 1/3 gating | pass | DROP — reaffirms the confirming ban (data-integrity must would stay open). |
| D1-solo | D1 discovery re-test (solo, unscoped) | discovery | wave-0 | ~54 (7/13) | pass | −34 vs unenforced — lost the load-bearing any_runner_wanted trap; triggered the whole program. |
| B2-A | A-cell R1 lens re-run | lens | wave-1 | 100 | pass | 100% precision, 0 anti-hits, 2 fresh novel-valid bugs. |
| C1-union | D1 3-arm union (case-1) | discovery-union | wave-1 | 90.6 | pass | 97% precision, 0 anti-hits; arm B caught any_runner_wanted/serve.rs:926 solo. |
| C1-opus | D1 opus-xhigh comparator (case-1) | opus-comparator | wave-1 | 62.5 | pass | Superb anchoring; missed 4 items incl. the flagged critical — the coverage gap. |
| B3-agv2 | AGV2 S3 (CRUD build) | build | wave-1 | 71 | HARD-FAIL | No pager clamp (the exact GT R2 bug) + INV-1 gap — self-check DEFENDED skipping the clamp. |
| B3-sheets | SHEETS S8 (wiring build) | build | wave-1 | 58 | HARD-FAIL | display_id wired to row id not fileId — defeats the panel's purpose; defended as a custody win. |
| B3-intake | INTAKE S8 (daemon build) | build | wave-1 | 37 | HARD-FAIL | Reproduced GT's R1 MUST verbatim (stalled ticks clear the failure counter) + a compile break. |
| B2-airtable | AIRTABLE R1 lens | lens | wave-1 | 47 | HARD-FAIL | Fabricated a [must] about provider behavior (numeric timestamp) — refuted by our own web GT. Tool-less memory. |
| B2-sdk | SDK R1 lens | lens | wave-1 | 100 | pass | 8/8 findings verified real; missed the GT must (recall not the lens bar). |
| B4-orient | AIRTABLE orient | orient | wave-1 | 60 | pass | ~75% coverage, 0 fabrication, but REVERSED the tier recommendation (standard vs high). |
| W1 | W1 web (v1) | web | wave-1 | VOID | pass | Config-level deny stripped ALL web tools; model refused to fabricate + reported the blocker. Fixed: config split. |
| B3b-pubapi | PUBAPI S3 (CRUD#2 build) | build | wave-2 | 60 | pass | Reached GT's R2 end-state in ONE draft but reproduced the busy→409 defect + dropped a trait method. |
| B3b-cal1 | CAL1 S5 (wiring#2 build) | build | wave-2 | 61 | pass | Custody mechanics right from the start; wrong preflight-401 discriminator, defended. |
| B3b-sag1 | SAG1 S6 (daemon#2 build) | build | wave-2 | 58 | HARD-FAIL | Read the spec, CONSIDERED the vendor-boundary shutdown fix, and explicitly rejected it — GT R1's headline MUST. |
| B4b-zoho | ZOHO orient#2 | orient | wave-2 | 67.5 | pass | All 4 Zoho forks nailed; dropped out-of-scope section; 2 fabricated identifiers (non-load-bearing). |
| W2 | W2 web (Airtable rubric) | web | wave-2 | 100 | pass | 42/42 — crux unhedged, every claim cited + labeled, 0 fabrications, with real WebFetch/WebSearch. |
| C2-union | D1 3-arm union (case-2) | discovery-union | wave-2 | 82.9 | pass | 98% precision, 0 anti-hits; 4/5 load-bearing clean + 1 partial across the union. |
| C2-opus | D1 opus-xhigh comparator (case-2) | opus-comparator | wave-2 | 84.2 | pass | 5/5 load-bearing, surgical citations — why 2 opus deep seats stay. |
| W1-v2 | W1-v2 web re-run | web | wave-2 | 88 | pass | Crux EXACT, all 3 TTLs, 8/9 headers, full citation discipline, 0 hallucinations. |
Union coverage vs a single deep seat
3-arm union vs one deep seatThree orthogonally-scoped glm-5.3 arms (~600–900 credits ≈ $1–2) union-beat one opus-xhigh seat on case-1 and held within tolerance on case-2. Orthogonal scoping — not duplicate prompts — is what did it: on case-1 arm B alone recovered the load-bearing trap both single seats missed. Case-2 shows a single opus seat at its best is surgically better on load-bearing depth, which is why the ruling is two-layer, not full replacement.
| case | glm union | opus-xhigh | anti-hits | precision | note |
|---|---|---|---|---|---|
| case-1 | 90.6 | 62.5 | 0 | 97% | Union caught any_runner_wanted (weight-5) that both single seats missed. |
| case-2 | 82.9 | 84.2 | 0 | 98% | −1.3 vs opus (inside the ≥opus−5 tolerance); 1 partial on a load-bearing item where opus was clean. |
| pooled | 86.75 | 73.35 | 0 | 97–98% | Pooled union ≥ pooled opus ⇒ a legitimate coverage floor; opus depth retained on top. |
weighted recall, 0–100 · bars scale to the 0–100 grade
Findings and bounds
- 01Every adopted use is additive/advisory — the discovery lattice feeds retained opus deep seats and the web lane adds a source-cited pass; neither replaces a judgment seat.
- 02The web-fabrication failure and the flawless web cells share one cause: grounding. With no web tools GLM fabricated a provider-behavior must; with real tools it scored 100 and 88.
- 03case-2 is the honest caveat to the union win — a single opus seat at its best (84.2) edged the union (82.9) on load-bearing depth, which is why 2 opus deep seats are retained.
- 04All anti-hits in the program landed in bench-gated (non-active) seats; the active canary seats had no adverse reports, so no live disable was warranted.
- 05Final verdicts await real-counsel-independent owner ratification of scope (additive-only vs both tracks); the replacement canary and the additive program run in parallel.
Verdict
More seats — but only additive ones, and only where an oracle or a stronger model owns the judgment.
- ·A 3-arm GLM discovery union out-covered a single opus-xhigh seat 90.6 vs 62.5 (case-1) and held within tolerance at 82.9 vs 84.2 (case-2); pooled 86.75 vs 73.35, 0 anti-hits — adopted as a coverage lattice feeding retained opus deep seats, not as a replacement.
- ·The additive web-research lane was adopted (mean 94; W2 42/42, W1-v2 88): with real web tools GLM is flawless; the one web fabrication happened in a cell that had no web tools.
- ·build_rote was rejected for good at n=6: 4/6 cells hard-failed on fence-invisible defects, three with the model's own self-check arguing for the bug it had just committed.
- ·Every grounded mode produced 0 anti-hits; every ungrounded-judgment seat (severity, tier, provider-memory, self-review) failed — the law of the dataset is grounded-strong, judgment-unsafe.
What changed in the harness because of it
- 01Owner ratified a two-layer discovery: a GLM coverage lattice (4–6 orthogonal enumeration lanes) plus 2 retained opus deep seats; the Sonnet/opus mechanical-breadth seats were retired.
- 02An additive GLM web-research lane was added, with a verify-tools-or-refuse contract.
- 03build_rote was removed from GLM eligibility permanently; build workers stay codex-5.4 (preferred) + Sonnet.
- 04The checkpoint_confirm / loop_iterate canary went live on Pro as a GLM-default 4-week trial, with codex as fallback + convergence declarer, circuit breakers, and a renewal rule due before 2026-09-21.
- 05The ADDENDUM protocol + an acceptance-rate circuit breaker were wired so additive GLM output never becomes wasteful context for main agents.
Raw data
generated 2026-08-30 by scripts/gen-glm.tsEvery number on this page is read from data/glm-more-seats.json, which the generator rewrites from these source files. A data update is a commit.
- data/raw/glm/PRO-ADOPTION-2026-08-21.mdThe whole expansion program: n≥2 rule, wave-1/wave-2 scores, B1 case-1/case-2 discovery-union results, the final verdict sheet, and the canary log.
- data/raw/glm/published/pro.htmlThe currently published Pro-expansion page — 28-run log, seat-status board, discovery-union numbers, judgment-wall prose; program-burn figure (10.2%).
- data/raw/glm/memory/glm-pro-default-canary.mdSecondary: Pro plan facts (12k/5h · 60k/wk · $80/mo, renews 2026-09-21), canary scope, enforcement fix, renewal rule.
- data/raw/glm/memory/claude-banned-from-review-glm-seats-ratified.mdSecondary: prior seat roster + A1 + severity guardrail context.
- data/raw/glm/memory/glm-adoption-grosstalk-and-offload.mdSecondary: model facts (glm-5.3 flagship, glm-4.7 1× breadth) and the original offload map.
Deliberately unfilled: programBurn: the source doc's own arithmetic disagrees — PRO-ADOPTION.md's final sheet says '~4.3k credits ≈ 7% of one Pro week' while pro.html (twice) and the program-complete line say 10.2%. Published 10.2% is used; the ~4.3k-credit figure is carried separately and does not reconcile to 10.2%. · '~30 blind verdicts' is stated as approximate in both sources; kept as the string '~30'. · W1-v2 web grade 88 comes from the final verdict sheet; the run log row 28 corroborates it (crux EXACT, all 3 TTLs, 8/9 headers)..