Study 02 · started 2026-08-21Ratified · canary

Can GLM hold more seats?

With the Pro plan's 6× quota, can GLM hold more seats than the two already ratified — build, discovery lattice, review lens, orient, web — without ever owning a judgment call?

The question
With Pro's 6× quota, can GLM hold MORE seats than the two already ratified — build, discovery-lattice, review-lens, orient, web?
The method
n≥2 per task category: two independent cells per category, averaged, all tool-enforced (deny verified per trial) in frozen-worktree sandboxes; ground truth = the shipped implementations including their checkpoint-round fixes.
The scoring
Each cell graded 0–100 (weighted recall for review/discovery, fidelity-vs-shipped for build, field-completeness for orient, rubric points for web). A hard-fail override: any anti-hit, false-clean, or fence-invisible must-defect fails the seat regardless of its mean.

More seats — but only additive ones, and only where an oracle or a stronger model owns the judgment.

GLM Coding Pro (6× quota) landed 2026-08-21. The owner ordered n≥2 testing of every candidate seat with a hard-fail override; 21 GLM trials + 2 opus comparators settled which new seats GLM can hold.

The program in numbers

21 GLM trials · 2 comparators
21
GLM trials
Plus 2 opus comparators; ~30 blind-scored verdicts.
10.2%
Program burn
Of one Pro week (pro.html). Note: the source doc's final sheet also states '~4.3k credits ≈ 7%' — the two do not reconcile.
0
Grounded anti-hits
Across every grounded mode (tree citations + web-tool citations).
4/6
Build hard-fails
build_rote cells; 3 with the model's self-check defending the defect.
86.75 vs 73.35
Discovery union (pooled)
GLM 3-arm union vs single opus-xhigh, pooled over both cases.
94
Web lane mean
W2 100 + W1-v2 88; crux EXACT both runs.
every category
n≥2 rule
Pre-registered before wave-2 results; seat grade = mean of ≥2 cells, hard-fail override.
60k/wk (6× LITE)
Pro quota
12k/5h · $80/mo, renews 2026-09-21; quota at canary flip 59,988/60,000.

The verdict sheet

One row per candidate seat after two waves of testing. The status column pairs its tone with the ruling text — no seat is decided by color alone.

seatevidencestatus
checkpoint_confirm + loop_iterateRatified pre-Pro; held/rose under enforcement; lead re-triage + codex gates.Active (canary)
discovery union (additive lattice)n=2: pooled 86.75 vs opus 73.35; 0 anti-hits; GO gates pass.Adopt — coverage layer + retained opus depth
web research lanen=2: W2 100 + W1-v2 88, crux EXACT both runs, 0 fabrications (mean 94).Adopt — additive web lane
build_rote6 cells, 4/6 hard-fail, defects self-defended; mean ~57.5.Rejected (final, n=6)
r1_additive_lens2×100% precision + 1 fabricated provider-must.Fails strict rule; code-only-claims redesign possible
orient / auto_proxyMean ~64; tier reversal + minor fabricated ids; no hard-fails.Weak — no seat recommended
any confirming / judgment seatConsistent severity/judgment failures across benches.Never (unchanged)

Every scored cell, both waves

n≥2 per category

Each cell graded 0–100 (or a descriptive recall string), tool-enforced in a frozen worktree. A hard-fail override — any anti-hit, false-clean, or fence-invisible must-defect — fails the seat regardless of its mean; hard-fail rows are marked in the danger color and labelled.

celltaskcategorywavegraderesultoutcome
Q1Q1 checkpoint-confirm re-testreviewwave-01/2 musts corr-sevpassHELD — reproduced the original exactly; 0 hallucinations.
Q2Q2 checkpoint-confirm re-testreviewwave-01/4 musts corr-sevpassROSE — enforcement fixed a downgrade + recovered a should; 3/4 musts still missed.
Q5Q5 confirming-tier re-testreviewwave-01/3 gatingpassDROP — reaffirms the confirming ban (data-integrity must would stay open).
D1-soloD1 discovery re-test (solo, unscoped)discoverywave-0~54 (7/13)pass−34 vs unenforced — lost the load-bearing any_runner_wanted trap; triggered the whole program.
B2-AA-cell R1 lens re-runlenswave-1100pass100% precision, 0 anti-hits, 2 fresh novel-valid bugs.
C1-unionD1 3-arm union (case-1)discovery-unionwave-190.6pass97% precision, 0 anti-hits; arm B caught any_runner_wanted/serve.rs:926 solo.
C1-opusD1 opus-xhigh comparator (case-1)opus-comparatorwave-162.5passSuperb anchoring; missed 4 items incl. the flagged critical — the coverage gap.
B3-agv2AGV2 S3 (CRUD build)buildwave-171HARD-FAILNo pager clamp (the exact GT R2 bug) + INV-1 gap — self-check DEFENDED skipping the clamp.
B3-sheetsSHEETS S8 (wiring build)buildwave-158HARD-FAILdisplay_id wired to row id not fileId — defeats the panel's purpose; defended as a custody win.
B3-intakeINTAKE S8 (daemon build)buildwave-137HARD-FAILReproduced GT's R1 MUST verbatim (stalled ticks clear the failure counter) + a compile break.
B2-airtableAIRTABLE R1 lenslenswave-147HARD-FAILFabricated a [must] about provider behavior (numeric timestamp) — refuted by our own web GT. Tool-less memory.
B2-sdkSDK R1 lenslenswave-1100pass8/8 findings verified real; missed the GT must (recall not the lens bar).
B4-orientAIRTABLE orientorientwave-160pass~75% coverage, 0 fabrication, but REVERSED the tier recommendation (standard vs high).
W1W1 web (v1)webwave-1VOIDpassConfig-level deny stripped ALL web tools; model refused to fabricate + reported the blocker. Fixed: config split.
B3b-pubapiPUBAPI S3 (CRUD#2 build)buildwave-260passReached GT's R2 end-state in ONE draft but reproduced the busy→409 defect + dropped a trait method.
B3b-cal1CAL1 S5 (wiring#2 build)buildwave-261passCustody mechanics right from the start; wrong preflight-401 discriminator, defended.
B3b-sag1SAG1 S6 (daemon#2 build)buildwave-258HARD-FAILRead the spec, CONSIDERED the vendor-boundary shutdown fix, and explicitly rejected it — GT R1's headline MUST.
B4b-zohoZOHO orient#2orientwave-267.5passAll 4 Zoho forks nailed; dropped out-of-scope section; 2 fabricated identifiers (non-load-bearing).
W2W2 web (Airtable rubric)webwave-2100pass42/42 — crux unhedged, every claim cited + labeled, 0 fabrications, with real WebFetch/WebSearch.
C2-unionD1 3-arm union (case-2)discovery-unionwave-282.9pass98% precision, 0 anti-hits; 4/5 load-bearing clean + 1 partial across the union.
C2-opusD1 opus-xhigh comparator (case-2)opus-comparatorwave-284.2pass5/5 load-bearing, surgical citations — why 2 opus deep seats stay.
W1-v2W1-v2 web re-runwebwave-288passCrux EXACT, all 3 TTLs, 8/9 headers, full citation discipline, 0 hallucinations.

Union coverage vs a single deep seat

3-arm union vs one deep seat

Three orthogonally-scoped glm-5.3 arms (~600–900 credits ≈ $1–2) union-beat one opus-xhigh seat on case-1 and held within tolerance on case-2. Orthogonal scoping — not duplicate prompts — is what did it: on case-1 arm B alone recovered the load-bearing trap both single seats missed. Case-2 shows a single opus seat at its best is surgically better on load-bearing depth, which is why the ruling is two-layer, not full replacement.

caseglm unionopus-xhighanti-hitsprecisionnote
case-190.662.5097%Union caught any_runner_wanted (weight-5) that both single seats missed.
case-282.984.2098%−1.3 vs opus (inside the ≥opus−5 tolerance); 1 partial on a load-bearing item where opus was clean.
pooled86.7573.35097–98%Pooled union ≥ pooled opus ⇒ a legitimate coverage floor; opus depth retained on top.
case-1
glm union
90.6
opus-xhigh
62.5
case-2
glm union
82.9
opus-xhigh
84.2
pooled
glm union
86.75
opus-xhigh
73.35

weighted recall, 0–100 · bars scale to the 0–100 grade

Findings and bounds

  • 01Every adopted use is additive/advisory — the discovery lattice feeds retained opus deep seats and the web lane adds a source-cited pass; neither replaces a judgment seat.
  • 02The web-fabrication failure and the flawless web cells share one cause: grounding. With no web tools GLM fabricated a provider-behavior must; with real tools it scored 100 and 88.
  • 03case-2 is the honest caveat to the union win — a single opus seat at its best (84.2) edged the union (82.9) on load-bearing depth, which is why 2 opus deep seats are retained.
  • 04All anti-hits in the program landed in bench-gated (non-active) seats; the active canary seats had no adverse reports, so no live disable was warranted.
  • 05Final verdicts await real-counsel-independent owner ratification of scope (additive-only vs both tracks); the replacement canary and the additive program run in parallel.

Verdict

Scored · 2026-08-21

More seats — but only additive ones, and only where an oracle or a stronger model owns the judgment.

  • ·A 3-arm GLM discovery union out-covered a single opus-xhigh seat 90.6 vs 62.5 (case-1) and held within tolerance at 82.9 vs 84.2 (case-2); pooled 86.75 vs 73.35, 0 anti-hits — adopted as a coverage lattice feeding retained opus deep seats, not as a replacement.
  • ·The additive web-research lane was adopted (mean 94; W2 42/42, W1-v2 88): with real web tools GLM is flawless; the one web fabrication happened in a cell that had no web tools.
  • ·build_rote was rejected for good at n=6: 4/6 cells hard-failed on fence-invisible defects, three with the model's own self-check arguing for the bug it had just committed.
  • ·Every grounded mode produced 0 anti-hits; every ungrounded-judgment seat (severity, tier, provider-memory, self-review) failed — the law of the dataset is grounded-strong, judgment-unsafe.

What changed in the harness because of it

  1. 01Owner ratified a two-layer discovery: a GLM coverage lattice (4–6 orthogonal enumeration lanes) plus 2 retained opus deep seats; the Sonnet/opus mechanical-breadth seats were retired.
  2. 02An additive GLM web-research lane was added, with a verify-tools-or-refuse contract.
  3. 03build_rote was removed from GLM eligibility permanently; build workers stay codex-5.4 (preferred) + Sonnet.
  4. 04The checkpoint_confirm / loop_iterate canary went live on Pro as a GLM-default 4-week trial, with codex as fallback + convergence declarer, circuit breakers, and a renewal rule due before 2026-09-21.
  5. 05The ADDENDUM protocol + an acceptance-rate circuit breaker were wired so additive GLM output never becomes wasteful context for main agents.

Raw data

generated 2026-08-30 by scripts/gen-glm.ts

Every number on this page is read from data/glm-more-seats.json, which the generator rewrites from these source files. A data update is a commit.

Deliberately unfilled: programBurn: the source doc's own arithmetic disagrees — PRO-ADOPTION.md's final sheet says '~4.3k credits ≈ 7% of one Pro week' while pro.html (twice) and the program-complete line say 10.2%. Published 10.2% is used; the ~4.3k-credit figure is carried separately and does not reconcile to 10.2%. · '~30 blind verdicts' is stated as approximate in both sources; kept as the string '~30'. · W1-v2 web grade 88 comes from the final verdict sheet; the run log row 28 corroborates it (crux EXACT, all 3 TTLs, 8/9 headers)..

I. Nocturne in E♭ major Op. 9 № 2 · Chopin

0:00 / 4:16