Routing is phase-specific
Terra leads review and cleanup and competes with Opus for build. The August 29 adoption was followed by further routing changes.
View policy snapshot
Editorial ratings · recomputed 2026-09-10. 70% evidence reading + 30% harness fit, rounded to five points.
gpt-5.6-terra at xhigh. Historical benchmark results are read alongside the September routing snapshot; planning/discovery estimates are marked explicitly.
Each input uses a 0–4 rubric. Evidence: negative (0), weak or indirect (1), mixed (2), strong within bounds (3), strong across tested cases (4). Fit: excluded (0), trial/reserve (1), bounded/additive (2), regular role (3), core role (4). Policy estimates are labeled below; these ratings are not measured success percentages.
Policy-based estimate: Terra has a tool-required web-research seat, but no codebase discovery-lead route or dedicated discovery score in these studies. This is a bounded role estimate.
Policy-based estimate: Terra opens plan scrutiny and confirms its results, with GPT iteration duties. Opus still leads plan and planloop, so Terra is rated as an active planning contributor.
Run D reached 96.7% checklist conformance, but its overall 72/100 reflects conduct and gate failures. The repaired harness now configures Terra for build, review and cleanup leadership.
Terra is the primary confirming/declaration model, but Run D reported a false-green gate and later Fable cells recovered defects Terra missed. Strong routing responsibility does not erase mixed observed judgment.
Run D’s isolation held, but its gate claims needed external verification; Run C’s withdrawn result is excluded as positive evidence. Later Fable comparisons also expose missed mechanisms.
Terra leads review and cleanup and competes with Opus for build. The August 29 adoption was followed by further routing changes.
View policy snapshotAn isolated W3-EMP rebuild scored 72/100 against Opus’s 85. Checklist conformance was 96.7% versus 96.0%; conduct and gate failures reduced the overall result.
Read studyThe earlier MAX-1 benchmark attempt could access the shipped implementation. Its withdrawal concerns that benchmark, not a production lead-run total.
Read studyStudy results describe the named model and test. Roles reflect the policy checked on 2026-09-10.
Terra is configured for build, review and cleanup leadership. Build can also route to Opus. Terra holds checkpoint declaration and review-confirm roles, plus a discovery web seat; Opus still leads planning.
Lead-seat study: Run D scored; Run C withdrawn. No lifetime run total is established by that study.