Can Fable 5.1 hold a codex review seat?
Replayed blind on eight frozen cells and scored against shipped ground truth, can Fable 5.1 hold the DECLARER, FIND, or ITERATE review seats the harness reserves for codex — and at which effort?
The method
frozen cells, replayed blindIts incumbent was a 25-candidate glm-5.3 swarm run that was never triaged. All 25 were adjudicated first, blind to the Fable cells: 0 must · 2 should · 8 could · 1 refuted · 14 duplicate — uniformly over-severitied and 56% duplicated across the eight swarm seats. Both real shoulds were still open on main. That adjudication is the bed's answer key.
The review worktree had been pruned. Grading used the round's carry-forward narration — the applied/refuted list and the R3 must the incumbent eventually filed — not the incumbent's original prose. Scored conservatively where reconstruction was involved.
The tally
8 of 8 cells · PM cite-verifiedEach row is one diff, replayed at both efforts against the incumbent that actually reviewed it. A verdict is per seat-bed: Ratified, Bounded, or Rejected.
| seat · bed | Fable @high | Fable @xhigh | incumbent on that diff |
|---|---|---|---|
| DECLARERSCAN-ENGINE R3 · d374495 | Ratifiedhit the sole must, with a sharper mechanism and a fix proposal | Ratified+hit the must + a real bug terra missed, verified on main | terra@xhigh: 1 must |
| ITERATESCAN-ENGINE R2 · 9f888b1 | Boundedmust 1/3, 3 novel verified-on-main candidates | Boundedmust 1/3, 4 novel verified-on-main candidates, 1.5× tokens | gpt-5.5@xhigh: 3 must + ~10 should |
| FINDMERGE-ERASURE S8 R2 · c76744688 | Ratified2 verified majors — beat terra, ≈ the swarm | Ratified+fullest seam diagnosis of any round | terra@xhigh: 0 · glm swarm: 1 major + 1 minor |
| FINDCAL-HARDENING-2 S11 · 4af33b86f | Boundedright on the facts, under-flagged on the shoulds (0.5/2) | Ratifiedboth shoulds (1.5/2), correct severities, 0 dup | terra R2×2 + R3: 0 ×3 · glm swarm-8: all, noisy |
Bed by bed
what each cell actually foundDECLARER · SCAN-ENGINE R3
Both efforts independently landed terra's one must: the R2 fix for the sidecar retry race is unsound — on a re-scan whose commit fails, the retry finds the recording's older same-owner sidecar, declares it superseded, returns Ok to the UI, and the new scan's events are silently discarded while the index still points at the stale entry. @high named it with a sharper mechanism and a fix. @xhigh went further and filed a second core finding terra missed: F14's stale-index fix only defers the false-block one scan, because commit_entry_inner never evicts a same-sidecar_path row. The PM verified it against live code; it is still on main and was routed to SCAN-ENGINE-HARDENING. @xhigh also read the round correctly as a fix-of-fix drift — the convergence tell that matches the real R4 outcome better than the incumbent's read. @high wrongly confirmed the F14 fix as safe; @xhigh earned its extra tokens precisely there.
ITERATE · SCAN-ENGINE R2
This is the seat Fable does not fit, and the cell shows why cleanly. The incumbent gpt-5.5 swept broad: 3 core musts and ~10 shoulds. Fable recalled 1 of 3 core musts at both efforts — it hit the retry-supersede item and missed cancel-token-await and omit-any-underivable — and almost none of the should breadth. But it surfaced 3–4 novel, verified-on-main defects no incumbent round ever raised: a non-blocking Library::lock that fails jobs Red under the unit's own parallelism and swallows a contended unpin; a global→batch gate-acquisition order that starves a concurrent batch; a watchdog NaN duration that clamps to the 300s floor instead of the ceiling; a RemedyCode contract mismatch. Narrow-and-deep is the wrong shape for a seat whose job is broad recall of a fix round. @xhigh did not buy breadth here; it bought depth and novelty at 1.5× the tokens.
FIND · MERGE-ERASURE S8 R2
On this diff terra found zero; the glm swarm found one major (two-pass detach ordering) and one minor. Fable at both efforts beat terra outright and matched-or-exceeded the swarm on substance. Both caught terra's own R3 must one round early: the R1 tag-taint predicate is a negative existence test where soundness needs a positive one, so a survivor's own identically-named tag is over-deleted in an erasure cascade. @xhigh gave the fullest diagnosis of the seam produced by any round — it proved the taint predicate is single-hop and still live for any chain with an intermediate node, and independently reconstructed the swarm's major as a failure mode of the whole-header retire rule. Two verified majors from a seat where the incumbent codex model found nothing.
FIND · CAL-HARDENING-2 S11
Here the effort split decided the verdict. Ground truth: 2 shoulds, both open on main. terra filed 0 across three passes (R2×2 and the R3 declare). The glm swarm-8 found everything but drowned it in mis-severity and 56% duplication. @high got the facts right but under-flagged: it analyzed the consent-predicate gap (C12) correctly and then consciously accepted it as ruling-conformant — a judgment miss — and it missed C13, verifying a path that a same-auth error actually slips through. @xhigh flagged C12 as a latent consent gap with the right one-line fix, refuted the wrong half of C13 with a real cite, and additionally caught that the S12 plan text and STEP_MANIFEST still carry the pre-amendment predicate. Same model, same brief, same commit: @xhigh bought judgment, not just depth.
What @xhigh bought
≈1.3–1.5× the tokens of @high@xhigh was ≥ @high on every bed and earned its extra tokens where judgment or verify-after-fix depth was the deciding factor — the C12 consent call, the F14 fix-of-fix, the two majors terra missed. @high is acceptable for FIND, where breadth of surfaced candidates matters more than adjudicating each. Route declare and confirm at @xhigh. @xhigh did not buy breadth anywhere: on the ITERATE bed it bought depth and novelty at 1.5× the cost and the same must recall.
| bed | Fable @high | Fable @xhigh | ratio |
|---|---|---|---|
| DECLARER · SCAN-ENGINE R3 | 135k | 195k | 1.44× |
| ITERATE · SCAN-ENGINE R2 | 210k | 312k | 1.49× |
| FIND · MERGE-ERASURE S8 R2 | 158k | 214k | 1.35× |
| FIND · CAL-HARDENING-2 S11 | 140k | 179k | 1.28× |
Citation discipline
the circuit-breaker- 0
- fabricated cites
- 8
- cells
Every PM spot-check landed on the exact line and excerpt claimed, with one immaterial exception: a C-R3 @xhigh sidecar_path reference was one symbol off its cited line, and the finding itself verified. No cell re-raised a known-refuted item as a must.
Seat verdicts
per seat, not per modelFable ruled convergence honestly in every declarer cell, matched the incumbent's must, and at @xhigh out-found it. It respected prior refutations and never re-raised a dead item. The one bound: there is no confirm-seat cell — a round whose job is to pass or fail a specific submitted fix, where a false PASS on a regressed fix is disqualifying. @high wrongly confirmed one fix in the R3 cell; that is exactly the failure a confirm seat must not make. Ratified for declare; a confirm cell is required before the seat is unbounded.
Two of two FIND beds beat the codex incumbent — 0→2 verified majors on MERGE-ERASURE, 0→both shoulds on CAL-HARDENING-2 at @xhigh — with clean cites and correct severities. This is the strongest fit.
Fable recalls about a third of a fix round's musts at either effort; its value is depth on one or two seams, not the wide sweep this seat needs. R1 stays the Sonnet swarm.
Byproduct — 7 defects routed
the method paid for itselfThe bake-off surfaced seven latent, verified-on-main defects as a side effect, each routed to a follow-up unit. Replaying old rounds against live code finds bugs that shipped past the original review — the frozen-cell method paid for itself.
- 01Retry superseded misfire — a failed re-scan commit discards the new scan and reports success.
- 02Stale same-sidecar_path index row false-blocks the legitimate owner's rescan; F14's fix only defers it one scan.
- 03Library::lock is a non-blocking try_lock_exclusive — spurious busy-Red under the unit's own parallelism, and a swallowed unpin leaks a pack pin ref.
- 04Global→batch gate-acquisition order starves a concurrent batch of machine-wide permits.
- 05Watchdog NaN duration clamps to the 300s floor, not the ceiling (f64::max ignores NaN) — spurious TimedOut.
- 01claim_and_book fence dropped its status predicate — admits a revoked auth row (latent consent-predicate gap, both backends; no current 'revoked' writer, so latent not live).
- 02The shared fence loosening also un-fences the poll: a mid-snapshot 401 books remaining invitees and then misreports the tick — an undocumented behavioral change.
Also surfaced and noted, not in the core seven: a RemedyCode contract mismatch (SCAN-ENGINE), a PG FOR SHARE→advisory-lock ordering edge to carry into S12's FOR UPDATE fence, and four adjacent minors on the MERGE-ERASURE detach/shell rules — removed-root over-strip, whole-header under-strip, an unbumped updated_at, a verified_at shell mismatch — routed to MERGE-EDIT-UNDO.
Limits
a probe, not a population- ·n = 8 cells across 3 beds and 2 repos — a probe, not a population.
- ·Grading was PM-performed — the same role that runs the live pipeline — not an independent adjudicator.
- ·Bed A's incumbent outputs were partly reconstructed from carry-forward narration, not originals.
- ·@xhigh cells cost ≈1.3–1.5× the tokens of @high, which any seat routing has to price in.
- ·There is no confirm-seat cell yet — the DECLARER verdict is bounded precisely on that missing evidence, since a false PASS on a submitted regression is the one failure that seat cannot afford.
Verdict
Ratified for FIND and DECLARER, Bounded for ITERATE — Fable beat the codex incumbent on both FIND beds and out-found terra at the declarer seat at @xhigh, but recalls only a third of a fix round's musts at the breadth seat.
- ·FIND: on both beds Fable beat the codex incumbent — 0→2 verified majors on MERGE-ERASURE S8 R2, and 0→both open shoulds on CAL-HARDENING-2 S11 at @xhigh — with clean cites and correct severities. The strongest fit.
- ·DECLARER: both efforts landed terra's one must on SCAN-ENGINE R3; @xhigh filed a second core finding terra missed, verified on main and routed. Bounded only by the missing confirm-seat cell — @high wrongly confirmed one fix.
- ·ITERATE: 1 of 3 core musts at both efforts against gpt-5.5's 3 must + ~10 should. Narrow-and-deep is the wrong shape for the breadth seat; gpt-5.5 and the glm-5.3 swarm keep it.
- ·0 fabricated citations across all 8 cells. @xhigh ≥ @high on every bed at ≈1.3–1.5× the tokens — it bought judgment and verify-after-fix depth, not breadth.
- fabricated cites / 8 cells
- 0
- defects routed
- 7
What changed in the harness because of it
- 01Seven verified-on-main latent defects were routed as a byproduct of the replay: five to SCAN-ENGINE-HARDENING, two to the CAL-HARDENING-2 RUN_OVERRIDE for resume. Four adjacent minors went to MERGE-EDIT-UNDO.
- 02Fable 5.1 is recorded as a candidate to reclaim the codex DECLARER, CONFIRM, and FIND seats at @xhigh — a finding, not a shipped routing change; a confirm-seat cell is required before DECLARER is unbounded.
- 03R1 stays the Sonnet swarm and ITERATE stays gpt-5.5 / the glm-5.3 swarm — a third of a fix round's musts is not breadth-seat recall.
- 04Bed B's never-triaged 25-candidate swarm run now has an answer key: 0 must · 2 should · 8 could · 1 refuted · 14 duplicate, adjudicated blind to the Fable cells.
Raw data
generated 2026-09-06 by scripts/gen-fable-frozen-cell.tsEvery number on this page is read from data/fable-frozen-cell.json, which the generator rewrites from these source files. A data update is a commit.
- data/raw/fable/STUDY.mdThe final tally: the per-cell results table (8 cells), verdicts, token costs, and the routed-bug byproduct. PM cite-verified.
- data/raw/fable/B-groundtruth.mdBed B's answer key: the adjudication of the 25 never-triaged swarm-8 candidates on CAL-HARDENING-2 S11.
- data/raw/fable/C-R3-fable-high.mdCell record: DECLARER, SCAN-ENGINE R3, Fable @high.
- data/raw/fable/C-R3-fable-xhigh.mdCell record: DECLARER, SCAN-ENGINE R3, Fable @xhigh.
- data/raw/fable/C-R2-fable-high.mdCell record: ITERATE, SCAN-ENGINE R2, Fable @high.
- data/raw/fable/C-R2-fable-xhigh.mdCell record: ITERATE, SCAN-ENGINE R2, Fable @xhigh.
- data/raw/fable/A-R2-fable-high.mdCell record: FIND, MERGE-ERASURE S8 R2, Fable @high.
- data/raw/fable/A-R2-fable-xhigh.mdCell record: FIND, MERGE-ERASURE S8 R2, Fable @xhigh.
- data/raw/fable/B-R2-fable-high.mdCell record: FIND, CAL-HARDENING-2 S11, Fable @high.
- data/raw/fable/B-R2-fable-xhigh.mdCell record: FIND, CAL-HARDENING-2 S11, Fable @xhigh. Per-finding detail and file:line cites in the cell records are held private.