Study 07 · published 2026-09-06Ratified · declare + find

Can Fable 5.1 hold a codex review seat?

Replayed blind on eight frozen cells and scored against shipped ground truth, can Fable 5.1 hold the DECLARER, FIND, or ITERATE review seats the harness reserves for codex — and at which effort?

The method

frozen cells, replayed blind
The replay
A throwaway git worktree at the exact commit the incumbent reviewer saw at the start of its round, and the identical seat brief. Fable works blind: it never sees the incumbent's output, the later rounds, or the shipped fix. Read-only, no toolchain, no build.
The design
3 beds × 2 repos × 2 efforts (@high, @xhigh) = 8 cells across three seat types — DECLARER, ITERATE, FIND — with the SCAN-ENGINE bed replayed at two rounds. Verdicts are Ratified, Bounded, or Rejected per seat-bed.
The grading
Every claim is cite-verified by the PM against live code at that commit (file:line, exact excerpt) before grading against shipped ground truth: what a later round proved real, what got refuted, what the incumbent caught or missed. A must the code contradicts is a false must and scores against the cell.
Bed B had no ground truth to score against

Its incumbent was a 25-candidate glm-5.3 swarm run that was never triaged. All 25 were adjudicated first, blind to the Fable cells: 0 must · 2 should · 8 could · 1 refuted · 14 duplicate — uniformly over-severitied and 56% duplicated across the eight swarm seats. Both real shoulds were still open on main. That adjudication is the bed's answer key.

Bed A's incumbent raw outputs were unrecoverable

The review worktree had been pruned. Grading used the round's carry-forward narration — the applied/refuted list and the R3 must the incumbent eventually filed — not the incumbent's original prose. Scored conservatively where reconstruction was involved.

The tally

8 of 8 cells · PM cite-verified

Each row is one diff, replayed at both efforts against the incumbent that actually reviewed it. A verdict is per seat-bed: Ratified, Bounded, or Rejected.

seat · bedFable @highFable @xhighincumbent on that diff
DECLARERSCAN-ENGINE R3 · d374495Ratifiedhit the sole must, with a sharper mechanism and a fix proposalRatified+hit the must + a real bug terra missed, verified on mainterra@xhigh: 1 must
ITERATESCAN-ENGINE R2 · 9f888b1Boundedmust 1/3, 3 novel verified-on-main candidatesBoundedmust 1/3, 4 novel verified-on-main candidates, 1.5× tokensgpt-5.5@xhigh: 3 must + ~10 should
FINDMERGE-ERASURE S8 R2 · c76744688Ratified2 verified majors — beat terra, ≈ the swarmRatified+fullest seam diagnosis of any roundterra@xhigh: 0 · glm swarm: 1 major + 1 minor
FINDCAL-HARDENING-2 S11 · 4af33b86fBoundedright on the facts, under-flagged on the shoulds (0.5/2)Ratifiedboth shoulds (1.5/2), correct severities, 0 dupterra R2×2 + R3: 0 ×3 · glm swarm-8: all, noisy

Bed by bed

what each cell actually found

DECLARER · SCAN-ENGINE R3

C-R3 · d374495
@high Ratified@xhigh Ratified+

Both efforts independently landed terra's one must: the R2 fix for the sidecar retry race is unsound — on a re-scan whose commit fails, the retry finds the recording's older same-owner sidecar, declares it superseded, returns Ok to the UI, and the new scan's events are silently discarded while the index still points at the stale entry. @high named it with a sharper mechanism and a fix. @xhigh went further and filed a second core finding terra missed: F14's stale-index fix only defers the false-block one scan, because commit_entry_inner never evicts a same-sidecar_path row. The PM verified it against live code; it is still on main and was routed to SCAN-ENGINE-HARDENING. @xhigh also read the round correctly as a fix-of-fix drift — the convergence tell that matches the real R4 outcome better than the incumbent's read. @high wrongly confirmed the F14 fix as safe; @xhigh earned its extra tokens precisely there.

incumbent — terra@xhigh: 1 must

ITERATE · SCAN-ENGINE R2

C-R2 · 9f888b1
@high Bounded@xhigh Bounded

This is the seat Fable does not fit, and the cell shows why cleanly. The incumbent gpt-5.5 swept broad: 3 core musts and ~10 shoulds. Fable recalled 1 of 3 core musts at both efforts — it hit the retry-supersede item and missed cancel-token-await and omit-any-underivable — and almost none of the should breadth. But it surfaced 3–4 novel, verified-on-main defects no incumbent round ever raised: a non-blocking Library::lock that fails jobs Red under the unit's own parallelism and swallows a contended unpin; a global→batch gate-acquisition order that starves a concurrent batch; a watchdog NaN duration that clamps to the 300s floor instead of the ceiling; a RemedyCode contract mismatch. Narrow-and-deep is the wrong shape for a seat whose job is broad recall of a fix round. @xhigh did not buy breadth here; it bought depth and novelty at 1.5× the tokens.

incumbent — gpt-5.5@xhigh: 3 must + ~10 should

FIND · MERGE-ERASURE S8 R2

A-R2 · c76744688
@high Ratified@xhigh Ratified+

On this diff terra found zero; the glm swarm found one major (two-pass detach ordering) and one minor. Fable at both efforts beat terra outright and matched-or-exceeded the swarm on substance. Both caught terra's own R3 must one round early: the R1 tag-taint predicate is a negative existence test where soundness needs a positive one, so a survivor's own identically-named tag is over-deleted in an erasure cascade. @xhigh gave the fullest diagnosis of the seam produced by any round — it proved the taint predicate is single-hop and still live for any chain with an intermediate node, and independently reconstructed the swarm's major as a failure mode of the whole-header retire rule. Two verified majors from a seat where the incumbent codex model found nothing.

incumbent — terra@xhigh: 0 · glm swarm: 1 major + 1 minor

FIND · CAL-HARDENING-2 S11

B-R2 · 4af33b86f
@high Bounded@xhigh Ratified

Here the effort split decided the verdict. Ground truth: 2 shoulds, both open on main. terra filed 0 across three passes (R2×2 and the R3 declare). The glm swarm-8 found everything but drowned it in mis-severity and 56% duplication. @high got the facts right but under-flagged: it analyzed the consent-predicate gap (C12) correctly and then consciously accepted it as ruling-conformant — a judgment miss — and it missed C13, verifying a path that a same-auth error actually slips through. @xhigh flagged C12 as a latent consent gap with the right one-line fix, refuted the wrong half of C13 with a real cite, and additionally caught that the S12 plan text and STEP_MANIFEST still carry the pre-amendment predicate. Same model, same brief, same commit: @xhigh bought judgment, not just depth.

incumbent — terra R2×2 + R3: 0 ×3 · glm swarm-8: all, noisy

What @xhigh bought

≈1.3–1.5× the tokens of @high

@xhigh was ≥ @high on every bed and earned its extra tokens where judgment or verify-after-fix depth was the deciding factor — the C12 consent call, the F14 fix-of-fix, the two majors terra missed. @high is acceptable for FIND, where breadth of surfaced candidates matters more than adjudicating each. Route declare and confirm at @xhigh. @xhigh did not buy breadth anywhere: on the ITERATE bed it bought depth and novelty at 1.5× the cost and the same must recall.

Fable @highFable @xhigh
100k200k300kDECLARER · SCAN-ENGINE R3135k195k · 1.44×ITERATE · SCAN-ENGINE R2210k312k · 1.49×FIND · MERGE-ERASURE S8 R2158k214k · 1.35×FIND · CAL-HARDENING-2 S11140k179k · 1.28×
tokens per cell, from the cell records · ratio = @xhigh ÷ @high on the same diff
Tokens per cell at @high and @xhigh across the four seat-beds: DECLARER SCAN-ENGINE R3 135k vs 195k; ITERATE SCAN-ENGINE R2 210k vs 312k; FIND MERGE-ERASURE S8 R2 158k vs 214k; FIND CAL-HARDENING-2 S11 140k vs 179k
bedFable @highFable @xhighratio
DECLARER · SCAN-ENGINE R3135k195k1.44×
ITERATE · SCAN-ENGINE R2210k312k1.49×
FIND · MERGE-ERASURE S8 R2158k214k1.35×
FIND · CAL-HARDENING-2 S11140k179k1.28×

Citation discipline

the circuit-breaker
Fable 5.1PASS
0
fabricated cites
8
cells

Every PM spot-check landed on the exact line and excerpt claimed, with one immaterial exception: a C-R3 @xhigh sidecar_path reference was one symbol off its cited line, and the finding itself verified. No cell re-raised a known-refuted item as a must.

Seat verdicts

per seat, not per model
DECLARER
Ratified · bounded pending a confirm-seat cell

Fable ruled convergence honestly in every declarer cell, matched the incumbent's must, and at @xhigh out-found it. It respected prior refutations and never re-raised a dead item. The one bound: there is no confirm-seat cell — a round whose job is to pass or fail a specific submitted fix, where a false PASS on a regressed fix is disqualifying. @high wrongly confirmed one fix in the R3 cell; that is exactly the failure a confirm seat must not make. Ratified for declare; a confirm cell is required before the seat is unbounded.

FIND
Ratified

Two of two FIND beds beat the codex incumbent — 0→2 verified majors on MERGE-ERASURE, 0→both shoulds on CAL-HARDENING-2 at @xhigh — with clean cites and correct severities. This is the strongest fit.

ITERATE / breadth
Bounded · keep gpt-5.5 and the glm-5.3 swarm

Fable recalls about a third of a fix round's musts at either effort; its value is depth on one or two seams, not the wide sweep this seat needs. R1 stays the Sonnet swarm.

Byproduct — 7 defects routed

the method paid for itself

The bake-off surfaced seven latent, verified-on-main defects as a side effect, each routed to a follow-up unit. Replaying old rounds against live code finds bugs that shipped past the original review — the frozen-cell method paid for itself.

SCAN-ENGINE-HARDENING
5 routed
  1. 01Retry superseded misfire — a failed re-scan commit discards the new scan and reports success.
  2. 02Stale same-sidecar_path index row false-blocks the legitimate owner's rescan; F14's fix only defers it one scan.
  3. 03Library::lock is a non-blocking try_lock_exclusive — spurious busy-Red under the unit's own parallelism, and a swallowed unpin leaks a pack pin ref.
  4. 04Global→batch gate-acquisition order starves a concurrent batch of machine-wide permits.
  5. 05Watchdog NaN duration clamps to the 300s floor, not the ceiling (f64::max ignores NaN) — spurious TimedOut.
CAL-HARDENING-2 RUN_OVERRIDE
2 routed
  1. 01claim_and_book fence dropped its status predicate — admits a revoked auth row (latent consent-predicate gap, both backends; no current 'revoked' writer, so latent not live).
  2. 02The shared fence loosening also un-fences the poll: a mid-snapshot 401 books remaining invitees and then misreports the tick — an undocumented behavioral change.

Also surfaced and noted, not in the core seven: a RemedyCode contract mismatch (SCAN-ENGINE), a PG FOR SHARE→advisory-lock ordering edge to carry into S12's FOR UPDATE fence, and four adjacent minors on the MERGE-ERASURE detach/shell rules — removed-root over-strip, whole-header under-strip, an unbumped updated_at, a verified_at shell mismatch — routed to MERGE-EDIT-UNDO.

Limits

a probe, not a population
  • ·n = 8 cells across 3 beds and 2 repos — a probe, not a population.
  • ·Grading was PM-performed — the same role that runs the live pipeline — not an independent adjudicator.
  • ·Bed A's incumbent outputs were partly reconstructed from carry-forward narration, not originals.
  • ·@xhigh cells cost ≈1.3–1.5× the tokens of @high, which any seat routing has to price in.
  • ·There is no confirm-seat cell yet — the DECLARER verdict is bounded precisely on that missing evidence, since a false PASS on a submitted regression is the one failure that seat cannot afford.

Verdict

Ratified · FIND + DECLARER · Bounded · ITERATE · 2026-09-06

Ratified for FIND and DECLARER, Bounded for ITERATE — Fable beat the codex incumbent on both FIND beds and out-found terra at the declarer seat at @xhigh, but recalls only a third of a fix round's musts at the breadth seat.

  • ·FIND: on both beds Fable beat the codex incumbent — 0→2 verified majors on MERGE-ERASURE S8 R2, and 0→both open shoulds on CAL-HARDENING-2 S11 at @xhigh — with clean cites and correct severities. The strongest fit.
  • ·DECLARER: both efforts landed terra's one must on SCAN-ENGINE R3; @xhigh filed a second core finding terra missed, verified on main and routed. Bounded only by the missing confirm-seat cell — @high wrongly confirmed one fix.
  • ·ITERATE: 1 of 3 core musts at both efforts against gpt-5.5's 3 must + ~10 should. Narrow-and-deep is the wrong shape for the breadth seat; gpt-5.5 and the glm-5.3 swarm keep it.
  • ·0 fabricated citations across all 8 cells. @xhigh ≥ @high on every bed at ≈1.3–1.5× the tokens — it bought judgment and verify-after-fix depth, not breadth.
fabricated cites / 8 cells
0
defects routed
7
a finding, not a shipped routing change

What changed in the harness because of it

  1. 01Seven verified-on-main latent defects were routed as a byproduct of the replay: five to SCAN-ENGINE-HARDENING, two to the CAL-HARDENING-2 RUN_OVERRIDE for resume. Four adjacent minors went to MERGE-EDIT-UNDO.
  2. 02Fable 5.1 is recorded as a candidate to reclaim the codex DECLARER, CONFIRM, and FIND seats at @xhigh — a finding, not a shipped routing change; a confirm-seat cell is required before DECLARER is unbounded.
  3. 03R1 stays the Sonnet swarm and ITERATE stays gpt-5.5 / the glm-5.3 swarm — a third of a fix round's musts is not breadth-seat recall.
  4. 04Bed B's never-triaged 25-candidate swarm run now has an answer key: 0 must · 2 should · 8 could · 1 refuted · 14 duplicate, adjudicated blind to the Fable cells.

Raw data

generated 2026-09-06 by scripts/gen-fable-frozen-cell.ts

Every number on this page is read from data/fable-frozen-cell.json, which the generator rewrites from these source files. A data update is a commit.

I. Nocturne in E♭ major Op. 9 № 2 · Chopin

0:00 / 4:16