Flash and Pro have different jobs
Current lint routing distinguishes fully checked fixes from changes a linter cannot fully validate. Sensitive partial-check fixes stay with the lead.
View policy snapshot
Editorial ratings · recomputed 2026-09-10. 70% evidence reading + 30% harness fit, rounded to five points.
DeepSeek v4 Pro for discovery/partial-check fixes and Flash for web/fully checked fixes. Cost evidence is from the August bench, not current pricing.
Each input uses a 0–4 rubric. Evidence: negative (0), weak or indirect (1), mixed (2), strong within bounds (3), strong across tested cases (4). Fit: excluded (0), trial/reserve (1), bounded/additive (2), regular role (3), core role (4). Policy estimates are labeled below; these ratings are not measured success percentages.
Pro recovered 10/13 discovery items with verified anchors and novel additions; Flash’s web cell was strong. These remain additive fallback roles, with Pro limited to its benched single-lane shape.
Indirect estimate: the bench’s iterate/confirm cells missed governing musts, and current policy has no DeepSeek planning route. No dedicated plan-authoring benchmark is available.
Flash passed all 16 compiler checks in its compile-fix cell; Pro matched the human partial-lint fix. Current lint routes preserve those narrow strengths, not general implementation leadership.
The confirming cell scored 0/20 and cleared a surface containing six known musts. Review and declaration routes exclude DeepSeek; this zero is the rubric’s negative-evidence band for judgment.
Pro’s discovery anchors were byte-exact and audited cells showed no invented citations. Flash’s discovery thrashing and the limited tested scopes keep this below the broadest evidence band.
The published bench billed $3.91 across 25 cells and used no subscription quota. Billing reconciliation and a later seven-cell check support economical bounded work; this is historical efficiency evidence, not a current tariff or universal cost ranking.
Current lint routing distinguishes fully checked fixes from changes a linter cannot fully validate. Sensitive partial-check fixes stay with the lead.
View policy snapshotThe August bench reached a passing compiler gate on all 16 checks. That is a frozen-cell result; the current compile-fix router separately names GLM.
Read studyThe confirming cell scored 0/20 and cleared a surface with six known must-fix defects. This supports the restriction on decision seats.
Read studyStudy results describe the named model and test. Roles reflect the policy checked on 2026-09-10.
DeepSeek has model-specific lint-fix routes and fallback discovery/web roles. Flash handles lints checked fully by the compiler; Pro handles partial-check lints. Sensitive partial-check changes stay with the lead, and DeepSeek has no review declaration seat.
The August bench approved four bounded uses; today’s compile-fix route lists GLM, so that historical approval is not a current assignment.