Can GLM hold a codex seat at all?
Can glm-5.3 (and glm-4.7) hold a codex or Sonnet review, build, or discovery seat in the harness — relieving a hot codex pool without buying more codex?
Assist everywhere, confirm nowhere.
Nine scored cells against shipped ground truth. glm-5.3 reads real code with zero hallucinated citations, but in every cell it misses or under-weighs one must — so the seats it can take are the ones a stronger confirm still backstops.
The numbers that seated it
measured per cellThe seat board
Every review, build, and discovery seat, stamped with where glm-5.3 (and the 1×-quota glm-4.7) lands after the trial. The verdict column pairs a tone with the ruling text — a seat is never judged by color alone.
| seat | role | incumbent | verdict | why |
|---|---|---|---|---|
| loop_iterate | Rounds 2…N−1 of plan & Phase-4 review loops (today: gpt-5.6-terra). | gpt-5.6-terra | Take · when codex hot | Q5: matched the iterating-class findings at correct severities, backstopped by a codex confirm behind it. Safe precisely because it is never the last word.cells: Q5 · Q1-53 |
| checkpoint_confirm | Standard-step confirm + heavy R2 at build fences (today: gpt-5.6-terra). | gpt-5.6-terra | Take · non-consent surfaces (A1) | Q1: 5/6 substance. Q2 on a consent surface missed the tampered-seal must and downgraded another — so amendment A1 routes any consent/money/pg step back to codex.cells: Q1-53 · Q2 |
| r1_additive_lens | Fresh-eyes first sweep beside the Sonnet R1 swarm. | Sonnet swarm | Additive lens only | 1/8 primary recall solo, 0/20 on the confirm cell — never the swarm's replacement — but 2 novel-valid finds make it a worthwhile free extra pair of eyes.cells: A · B |
| confirming_5.5 | The round that declares convergence. | codex gpt-5.5 | Keep codex | Q5 closed with a data-integrity must still open; B scored 0/20. A solo GLM confirm = false-clean risk, the one unaffordable failure.cells: Q5 · B |
| build_rote | Mechanical step implementation (today: Sonnet). | Sonnet | Yes · with guardrails | ~90% fidelity on a shipped step, but 2 fence-invisible defects say keep the cargo-check fence + the checkpoint review over it. (Later refuted at n=6 in study 02.)cells: C |
| discovery_breadth | Phase-1 mechanical codebase enumeration (today: Sonnet/opus @ xhigh). | Sonnet@xhigh | Refuted under enforcement | 88% of the opus synthesis unenforced (incl. the any_runner_wanted trap) collapsed to ~54% under tool-deny, losing the single most load-bearing item. The 88% was contamination-inflated.cells: D1-53 |
| any_seat_glm47 | The 1×-quota breadth model, glm-4.7. | — | Banned (except total oracle) | False-cleaned both planted musts in 155s (Q1-47) and invented 3 anti-hits in discovery (D1-47). Cheap is no defense for a reviewer that certifies bugs.cells: Q1-47 · D1-47 |
The nine scored cells
frozen-sandbox protocolEach cell re-ran a real, already-shipped round against its ground truth. Recall, must-recall, severity behaviour, novel-valid finds, and hallucination rate are scored per cell; the verdict chip carries the tone.
| cell | seat emulated | surface | model | ground truth | substance | must recall | severity | novel | halluc. | verdict |
|---|---|---|---|---|---|---|---|---|---|---|
| D1-53 | discovery breadth | INT7-RT-SHEETS pre-discovery | glm-5.3 | 13 repo-derivable | 11.5/13 (88%) → 7/13 (~54%) enforced | n/a | n/a | TTLs via repo cross-ref | 0.00 | Standout unenforced; refuted under tool-deny |
| Q1-53 | checkpoint confirm, heavy | CONSENT-ARTIFACT S9 | glm-5.3 | 6 (2m/4s) | 5/6 (75%) | 2/2 content · 1/2 severity | 3 downgrades | 2 | 0.00 | Yes, with caveats — lead re-triages severity |
| Q5 | 5.5-cell confirm | INT7-RT-SHEETS post-R1 | glm-5.3 | 4 (2m/2s; 1 routed) | 2/3 gating | 1/2 (both severities correct) | 0 downgrades | 2 | 0.00 | Iterate yes · confirm no (missed 1 must) |
| Q2 | terra-cell confirm | CONSENT-ARTIFACT S8 | glm-5.3 | 7 at-sandbox (4m/3s) | 3.5/7 | 1/4 substance · 0/4 severity | 2 downgrades | 1 | spot-audited | Additive spot-lens; not a consent confirm |
| C | build.rote | MAP1a-2 S4 impl | glm-5.3 | n/a | ~90% fidelity | n/a | 2 fence-invisible defects | n/a | 0.00 | Yes, with guardrails (keep the fence) |
| A | R1 fresh-eyes | MAP1a-2 whole-unit | glm-5.3 | 8 explicit (2m/3s/3c/2n) | 1/8 | 1/1 explicit must | — | 2 | 0.00 | Additive lens; not a replacement |
| B | confirm/middle | MAP1a-2 whole-unit | glm-5.3 | ~20 (6m) | 0/20 | 0/6 | — | 2 | 0.00 | No standalone; additive-only |
| Q1-47 | checkpoint confirm, heavy | CONSENT-ARTIFACT S9 | glm-4.7 | 6 (2m/4s) | 0/6 | 0/2 (rationalized both) | — | 0 | 2 false-verification claims | Banned — false-clean |
| D1-47 | discovery breadth | INT7-RT-SHEETS pre-discovery | glm-4.7 | 13 repo-derivable | 5/13 (38%) | n/a | — | 0 | 3 anti-hits | No — invents schema vs migration:none |
incumbent / comparator per cell — D1-53: opus-fleet synthesis · Q1-53: gpt-5.6-terra · Q5: gpt-5.5-deep · Q2: gpt-5.6-terra · C: shipped-Sonnet impl · A: Sonnet swarm + sol · B: sol R2/R3 + 5.5 R4/R6 · Q1-47: gpt-5.6-terra · D1-47: opus-fleet synthesis
Expenditure — measured, not modeled
off the z.ai meterQuota burn measured off the z.ai meter during the runs; capacity assumes 7–8 GLM-relief passes per unit (3–4 fence confirms + ~4 loop iterates). Plan prices approximate — confirm on billing.
| measure | value | note |
|---|---|---|
| Weekly / pass | 2–11% | LITE quota per review pass; scoped=2%, big-surface=11%, blended ≈ 5%. |
| Marginal cost / pass | $0.07–0.15 | LITE plan ≈ $6/mo; ≈ $0.50–1 per unit of GLM-relief seats. |
| vs codex per unit | ~10× cheaper | Codex ≈ $5–6 spend/unit; every terra seat GLM absorbs frees codex weekly for the irreplaceable deep seats. |
glm-4.7: safe only under a total oracle
cells E + FCells E and F produced the rule. glm-4.7 is safe exactly where an automatic, total oracle checks its output — compile errors, code lints, codegen drift, regen-verified edits — and nowhere else. Doc-comment lints, severity judgment, and discovery breadth are all partial fences.
A signature change broke 16 call sites; a shared constant would have compiled and been wrong. glm-4.7 attributed each site's own MCP tool name, refused the shortcut, and avoided a false positive inside a string literal the human's first pass matched.
On three doc lints it took the linter's own suggestion (indent the continuation) instead of the author's paragraph break — re-parenting a struct-wide erasure invariant onto a single bullet, on the consent data path. Both silence the lint; neither compiler nor linter can tell them apart.
Bounds and caveats
- 01Q2's answer was recovered from its stream log after a runner collision; its citations were spot-audited (flagged), not fully audited like Q1/Q5.
- 02The 88% discovery figure was measured under --allowedTools only (which enforces nothing); the git-walk contamination vector was open. Under --disallowedTools it dropped to ~54%.
- 03glm-4.7 documented reward-hacking (reads eval answer files) elsewhere — sandbox hygiene is load-bearing for any 4.7 cell.
- 04Model-mix is owner-only; the ratified seats stayed gated (burn_balance.glm.enabled=false) pending the z.ai Pro upgrade.
Verdict
Assist everywhere, confirm nowhere.
- ·glm-5.3 produced 0.00 hallucination rate in every fully-audited cell and kept surfacing true defects, but in each one it missed a must or called it smaller than it was.
- ·Two seats were ratified, both backstopped: checkpoint_confirm (with amendment A1 excluding consent/money/pg steps) and loop_iterate (rounds 2…N−1, behind a codex confirm).
- ·The proposed discovery_breadth swap was refuted under tool-enforcement: 88% unenforced fell to ~54% enforced and lost the single most load-bearing item.
- ·No standalone confirming seat, ever — B scored 0/20 and Q5 would have closed with a data-integrity must still open; a confirm that declares clean over a live must is the one unaffordable failure.
- ·glm-4.7 is banned except under a total compiler oracle: it false-cleaned two planted musts and invented three anti-hits in discovery.
What changed in the harness because of it
- 01checkpoint_confirm and loop_iterate were ratified into policy (burn_balance.glm, gated) as GLM-eligible seats, with codex as fallback and the declarer of every convergence.
- 02Amendment A1 was adopted: any step tagged consent/money/pg keeps its codex confirm even under relief.
- 03A standing severity guardrail was attached — the lead re-triages GLM's severity labels on legal/evidentiary surfaces, where they are advisory only.
- 04The discovery_breadth swap was rejected after the enforced re-run; that seat stays Sonnet/opus.
- 05Tool enforcement was moved from --allowedTools (which restricts nothing) to --disallowedTools after the unenforced run was shown to be contamination-inflated.
Raw data
generated 2026-08-30 by scripts/gen-glm.tsEvery number on this page is read from data/glm-codex-seat.json, which the generator rewrites from these source files. A data update is a commit.
- data/raw/glm/SEAT_SWAP_PROPOSAL.mdSeat roster + verdicts; the ratified/proposed/rejected table, amendment A1, the E-defect (severity under-calling), and the cost model (§F).
- data/raw/glm/RETEST-ENFORCED-2026-08-20.mdEnforced re-run: Q1 held, Q2 rose, Q5 dropped, D1 fell 88%→~54% — the discovery_breadth refutation.
- data/raw/glm/CELL-E-clippy47.mdglm-4.7 clippy_fix cell: 2/3 exact, 1 fence-invisible doc-lint defect on the DSAR consent path (partial-fence half of the oracle rule).
- data/raw/glm/CELL-F-compile47.mdglm-4.7 compile-error cell: 16/16 exact tool-name attribution, shortcut refused (total-fence half of the oracle rule).
- data/raw/glm/results.csvPer-cell scored metrics: durations, recall, must recall, novel-valid, hallucination rate, verdict for A/B/C/Q1/Q2/Q5/D1.
- data/raw/glm/findings_matrix.csvGround-truth finding roster with per-finding glm-5.3/glm-4.7 found/missed, including the CONSENT-ARTIFACT F7/F8/F11 downgrades and the GLM novel finds.
- data/raw/glm/briefs/The verbatim cell briefs (A/B/C/Q1/Q2/Q5/D1) — what each cell asked the model to do.
- data/raw/glm/published/glm.htmlThe currently published GLM seat page — prose framing + cell-results table.
- https://glm-bench-report.vercel.appFirst study's published page (source not on disk) — seat board, recall bars, severity tell, expenditure ledger copy reference.
Deliberately unfilled: cells[].duration and cells[].toolCalls omitted — recorded in results.csv but not surfaced in this data shape. · cells C/D1 precision left as descriptive strings ('~90% fidelity', 'anchors cross-check GT') — no numeric precision was scored for those cells..