Study 01 · started 2026-08-16Seated · 2 roles

Can GLM hold a codex seat at all?

Can glm-5.3 (and glm-4.7) hold a codex or Sonnet review, build, or discovery seat in the harness — relieving a hot codex pool without buying more codex?

The question
Can glm-5.3 (and glm-4.7) hold a codex/Sonnet review, build, or discovery seat in the harness — relieving a hot codex without buying more codex?
The method
Each cell re-runs a real, already-shipped review round in a frozen git sandbox at the exact commit the original reviewer saw; the model gets read-only tools and every tool call is audited (cheat-flags 0 across all cells).
The scoring
Ground truth = the findings that round actually produced, grounded in the shipped fix commits + CONVERGENCE.ndjson. Substance recall, must recall, severity calls, novel-valid, and hallucination rate are scored per cell.

Assist everywhere, confirm nowhere.

Nine scored cells against shipped ground truth. glm-5.3 reads real code with zero hallucinated citations, but in every cell it misses or under-weighs one must — so the seats it can take are the ones a stronger confirm still backstops.

The numbers that seated it

measured per cell
9
Scored cells
A · B · C · Q1(53/47) · Q2 · Q5 · D1(53/47), frozen-sandbox protocol.
0.00
Hallucination rate
Every glm-5.3 citation re-opened against the sandbox tree in all fully-audited cells.
$0.07–0.15
Cost per pass
Marginal z.ai LITE cost; ≈ $0.50–1 per unit of GLM-relief seats.
2–11% LITE
Weekly quota per pass
Scoped confirm (Q5)=2%; big-surface (Q2)=11%; blended ≈ 5%.
~10×
Cheaper than codex
Codex yardstick ≈ $5–6 of ChatGPT-Pro spend per unit at the current mix.
0
Cheat flags
Across all cells — every read-only tool call audited.
z.ai LITE
Plan
thinking default-max; policy burn_balance.glm.enabled=false — this page was the evidence for that flip.

The seat board

Every review, build, and discovery seat, stamped with where glm-5.3 (and the 1×-quota glm-4.7) lands after the trial. The verdict column pairs a tone with the ruling text — a seat is never judged by color alone.

seatroleincumbentverdictwhy
loop_iterateRounds 2…N−1 of plan & Phase-4 review loops (today: gpt-5.6-terra).gpt-5.6-terraTake · when codex hotQ5: matched the iterating-class findings at correct severities, backstopped by a codex confirm behind it. Safe precisely because it is never the last word.cells: Q5 · Q1-53
checkpoint_confirmStandard-step confirm + heavy R2 at build fences (today: gpt-5.6-terra).gpt-5.6-terraTake · non-consent surfaces (A1)Q1: 5/6 substance. Q2 on a consent surface missed the tampered-seal must and downgraded another — so amendment A1 routes any consent/money/pg step back to codex.cells: Q1-53 · Q2
r1_additive_lensFresh-eyes first sweep beside the Sonnet R1 swarm.Sonnet swarmAdditive lens only1/8 primary recall solo, 0/20 on the confirm cell — never the swarm's replacement — but 2 novel-valid finds make it a worthwhile free extra pair of eyes.cells: A · B
confirming_5.5The round that declares convergence.codex gpt-5.5Keep codexQ5 closed with a data-integrity must still open; B scored 0/20. A solo GLM confirm = false-clean risk, the one unaffordable failure.cells: Q5 · B
build_roteMechanical step implementation (today: Sonnet).SonnetYes · with guardrails~90% fidelity on a shipped step, but 2 fence-invisible defects say keep the cargo-check fence + the checkpoint review over it. (Later refuted at n=6 in study 02.)cells: C
discovery_breadthPhase-1 mechanical codebase enumeration (today: Sonnet/opus @ xhigh).Sonnet@xhighRefuted under enforcement88% of the opus synthesis unenforced (incl. the any_runner_wanted trap) collapsed to ~54% under tool-deny, losing the single most load-bearing item. The 88% was contamination-inflated.cells: D1-53
any_seat_glm47The 1×-quota breadth model, glm-4.7.—Banned (except total oracle)False-cleaned both planted musts in 155s (Q1-47) and invented 3 anti-hits in discovery (D1-47). Cheap is no defense for a reviewer that certifies bugs.cells: Q1-47 · D1-47

The nine scored cells

frozen-sandbox protocol

Each cell re-ran a real, already-shipped round against its ground truth. Recall, must-recall, severity behaviour, novel-valid finds, and hallucination rate are scored per cell; the verdict chip carries the tone.

cellseat emulatedsurfacemodelground truthsubstancemust recallseveritynovelhalluc.verdict
D1-53discovery breadthINT7-RT-SHEETS pre-discoveryglm-5.313 repo-derivable11.5/13 (88%) → 7/13 (~54%) enforcedn/an/aTTLs via repo cross-ref0.00Standout unenforced; refuted under tool-deny
Q1-53checkpoint confirm, heavyCONSENT-ARTIFACT S9glm-5.36 (2m/4s)5/6 (75%)2/2 content · 1/2 severity3 downgrades20.00Yes, with caveats — lead re-triages severity
Q55.5-cell confirmINT7-RT-SHEETS post-R1glm-5.34 (2m/2s; 1 routed)2/3 gating1/2 (both severities correct)0 downgrades20.00Iterate yes · confirm no (missed 1 must)
Q2terra-cell confirmCONSENT-ARTIFACT S8glm-5.37 at-sandbox (4m/3s)3.5/71/4 substance · 0/4 severity2 downgrades1spot-auditedAdditive spot-lens; not a consent confirm
Cbuild.roteMAP1a-2 S4 implglm-5.3n/a~90% fidelityn/a2 fence-invisible defectsn/a0.00Yes, with guardrails (keep the fence)
AR1 fresh-eyesMAP1a-2 whole-unitglm-5.38 explicit (2m/3s/3c/2n)1/81/1 explicit must—20.00Additive lens; not a replacement
Bconfirm/middleMAP1a-2 whole-unitglm-5.3~20 (6m)0/200/6—20.00No standalone; additive-only
Q1-47checkpoint confirm, heavyCONSENT-ARTIFACT S9glm-4.76 (2m/4s)0/60/2 (rationalized both)—02 false-verification claimsBanned — false-clean
D1-47discovery breadthINT7-RT-SHEETS pre-discoveryglm-4.713 repo-derivable5/13 (38%)n/a—03 anti-hitsNo — invents schema vs migration:none

incumbent / comparator per cell — D1-53: opus-fleet synthesis · Q1-53: gpt-5.6-terra · Q5: gpt-5.5-deep · Q2: gpt-5.6-terra · C: shipped-Sonnet impl · A: Sonnet swarm + sol · B: sol R2/R3 + 5.5 R4/R6 · Q1-47: gpt-5.6-terra · D1-47: opus-fleet synthesis

Expenditure — measured, not modeled

off the z.ai meter

Quota burn measured off the z.ai meter during the runs; capacity assumes 7–8 GLM-relief passes per unit (3–4 fence confirms + ~4 loop iterates). Plan prices approximate — confirm on billing.

measurevaluenote
Weekly / pass2–11%LITE quota per review pass; scoped=2%, big-surface=11%, blended ≈ 5%.
Marginal cost / pass$0.07–0.15LITE plan ≈ $6/mo; ≈ $0.50–1 per unit of GLM-relief seats.
vs codex per unit~10× cheaperCodex ≈ $5–6 spend/unit; every terra seat GLM absorbs frees codex weekly for the irreplaceable deep seats.
LITE capacity
~3–5 u/wk
Below fleet demand.
PRO capacity (6×)
~20–30 u/wk
Effectively unconstrained.
Fleet demand
~7–10 u/wk
What the fleet clears when hot ⇒ LITE can't cover demand; PRO can.

glm-4.7: safe only under a total oracle

cells E + F

Cells E and F produced the rule. glm-4.7 is safe exactly where an automatic, total oracle checks its output — compile errors, code lints, codegen drift, regen-verified edits — and nowhere else. Doc-comment lints, severity judgment, and discovery breadth are all partial fences.

cell F
16/16 exact
Compile-error loop — a total fence

A signature change broke 16 call sites; a shared constant would have compiled and been wrong. glm-4.7 attributed each site's own MCP tool name, refused the shortcut, and avoided a false positive inside a string literal the human's first pass matched.

cell E
2/3 exact, 1 fence-invisible
Clippy fix — a partial fence

On three doc lints it took the linter's own suggestion (indent the continuation) instead of the author's paragraph break — re-parenting a struct-wide erasure invariant onto a single bullet, on the consent data path. Both silence the lint; neither compiler nor linter can tell them apart.

Bounds and caveats

  • 01Q2's answer was recovered from its stream log after a runner collision; its citations were spot-audited (flagged), not fully audited like Q1/Q5.
  • 02The 88% discovery figure was measured under --allowedTools only (which enforces nothing); the git-walk contamination vector was open. Under --disallowedTools it dropped to ~54%.
  • 03glm-4.7 documented reward-hacking (reads eval answer files) elsewhere — sandbox hygiene is load-bearing for any 4.7 cell.
  • 04Model-mix is owner-only; the ratified seats stayed gated (burn_balance.glm.enabled=false) pending the z.ai Pro upgrade.

Verdict

Scored · 2026-08-20

Assist everywhere, confirm nowhere.

  • ·glm-5.3 produced 0.00 hallucination rate in every fully-audited cell and kept surfacing true defects, but in each one it missed a must or called it smaller than it was.
  • ·Two seats were ratified, both backstopped: checkpoint_confirm (with amendment A1 excluding consent/money/pg steps) and loop_iterate (rounds 2…N−1, behind a codex confirm).
  • ·The proposed discovery_breadth swap was refuted under tool-enforcement: 88% unenforced fell to ~54% enforced and lost the single most load-bearing item.
  • ·No standalone confirming seat, ever — B scored 0/20 and Q5 would have closed with a data-integrity must still open; a confirm that declares clean over a live must is the one unaffordable failure.
  • ·glm-4.7 is banned except under a total compiler oracle: it false-cleaned two planted musts and invented three anti-hits in discovery.

What changed in the harness because of it

  1. 01checkpoint_confirm and loop_iterate were ratified into policy (burn_balance.glm, gated) as GLM-eligible seats, with codex as fallback and the declarer of every convergence.
  2. 02Amendment A1 was adopted: any step tagged consent/money/pg keeps its codex confirm even under relief.
  3. 03A standing severity guardrail was attached — the lead re-triages GLM's severity labels on legal/evidentiary surfaces, where they are advisory only.
  4. 04The discovery_breadth swap was rejected after the enforced re-run; that seat stays Sonnet/opus.
  5. 05Tool enforcement was moved from --allowedTools (which restricts nothing) to --disallowedTools after the unenforced run was shown to be contamination-inflated.

Raw data

generated 2026-08-30 by scripts/gen-glm.ts

Every number on this page is read from data/glm-codex-seat.json, which the generator rewrites from these source files. A data update is a commit.

Deliberately unfilled: cells[].duration and cells[].toolCalls omitted — recorded in results.csv but not surfaced in this data shape. · cells C/D1 precision left as descriptive strings ('~90% fidelity', 'anchors cross-check GT') — no numeric precision was scored for those cells..

I. Nocturne in E♭ major Op. 9 № 2 · Chopin

0:00 / 4:16