Discovery adopted August 21
The expansion study adopted an additive coverage lattice and web lane, with Opus retained for deep discovery and synthesis.
Read study
Editorial ratings · recomputed 2026-09-10. 70% evidence reading + 30% harness fit, rounded to five points.
glm-5.3, including scoped discovery lanes and review swarms. Scores do not transfer to glm-4.7 or to unbounded lead/confirmation work.
Each input uses a 0–4 rubric. Evidence: negative (0), weak or indirect (1), mixed (2), strong within bounds (3), strong across tested cases (4). Fit: excluded (0), trial/reserve (1), bounded/additive (2), regular role (3), core role (4). Policy estimates are labeled below; these ratings are not measured success percentages.
The scoped union averaged 86.75 versus single-Opus 73.35, and the two web cells averaged 94. Current policy retains these additive lanes alongside Opus depth.
GLM contributes loop findings and discovery coverage, but severity downgrades and missed musts prevent it from owning planning decisions or convergence.
Four of six general rote-build cells hard-failed, often with defects defended during self-check. Current policy allows only gated build classes and restricted fix routes.
Finding breadth is useful, but repeated must downgrades and false-clean confirming cells disqualify GLM from judgment. Current declaration routes exclude it.
The September swarm produced 95 findings across 24 seat outputs with no fabricated citations; grounded modes in the expansion also had no anti-hits. Tool-less provider claims remain a documented failure mode and lead verification remains required.
The August studies documented inexpensive subscription relief and substantially greater Pro capacity. Scoped GLM routes now absorb finding/fix work; conflicting historical burn totals keep the evidence below the top band. This does not quote current plan pricing.
The expansion study adopted an additive coverage lattice and web lane, with Opus retained for deep discovery and synthesis.
Read studyIn three frozen cells, 24 GLM seat outputs produced 95 findings and no fabricated citations. The study still requires citation checking and independent severity judgment.
Read studyThe August study rejected general rote build. Current policy separately permits gated build classes and restricted fix routes; this is not blanket build approval.
View policy snapshotStudy results describe the named model and test. Roles reflect the policy checked on 2026-09-10.
GLM supplies the discovery lattice, additive web research, and review-finding lanes. The current router also includes bounded build and fix work. Citations and severity need lead verification; convergence declarations remain with GPT.
Finding and declaring are separate seats; GLM findings require lead verification.