DeepSeek as the pressure valve
Can an API-metered model give relief when the subscription pools run hot — and where must it stay out of the decision seats?
DeepSeek was run against the GLM bench's own frozen sandboxes, briefs and ground truth, so every number sits directly beside glm-5.3. The seat map is decided by how complete the automatic check is, never by cost.
The bench in nine numbers
each read from a source fileEvery bench cell, scored beside glm-5.3
23 cellsEach cell reuses a frozen GLM-bench sandbox, brief and ground truth. Cells tagged ENFORCED were re-run under a real read-only guard; contamination could only inflate scores, so the clean rows are canonical.
| trial | cell / seat | model | incumbent | recall pri / must | severity accuracy | novel valid | halluc. | cost | verdict |
|---|---|---|---|---|---|---|---|---|---|
| Q1-consent-pro | Q1checkpoint_confirm heavy (consent) | deepseek-v4-pro | gpt-5.6-terra | 3/6 / 1/2 | 1 downgrade (F10 should->could) | 1 (restore-replay; cite verified) | 0 | $0.15017 | false-cleanWEAKER than glm-5.3 (4/6,2/2m). No false-clean. Missed F8 envelope_mac must |
| D1-disc-pro | D1discovery_breadth | deepseek-v4-pro | opus-fleet synthesis | 10/13 (77%) / n/a | n/a | 3 strong (scope-firewall drive.file vs channels.watch; binding-grain partial-unique; initial-sync re-import) | 0 | $0.13379 | passSTRONG. Below glm-5.3 11.5/13 but complementary; all spot-checked anchors byte-exact |
| Q5-iter-pro | Q5loop_iterate / R2 confirm | deepseek-v4-pro | gpt-5.5-deep | 0/3 / 0/2 | all peripheral could/nit | 0 | 0 | $0.27005 | failFAIL as confirm/iterate. Missed BOTH musts. Worse than glm-5.3 (2/3 gating) |
| A-r1-pro | A-r1r1_additive_lens | deepseek-v4-pro | Sonnet swarm + sol breadth | >=1/8 / partial | 1 downgrade (R1-F1 must->should) | TBD | 0 | $0.23426 | mixed~parity with glm-5.3 (1/8). Additive lens at best |
| Q2-consent-pro | Q2checkpoint_confirm (consent/PG) | deepseek-v4-pro | gpt-5.6-terra | 0/2 / 0/1 | n/a | 0 | 0 | $0.2351 | failMISSED the PG TOCTOU must |
| Q2-consent-flash | Q2checkpoint_confirm (consent/PG) | deepseek-v4-flash | gpt-5.6-terra | 0/2 / 0/1 | n/a | 0 | 0 | $0.1391 | false-cleanFALSE-CLEAN: declared the unlocked tombstone screen 'sound' |
| Q1-consent-flash | Q1checkpoint_confirm heavy | deepseek-v4-flash | gpt-5.6-terra | 4/6 / 2/2 content 1/2 severity | must rated 'low'; used high/med/low vocab | 1 | 0 | $0.08791 | passBEAT pro (3/6). Severity vocabulary non-conforming |
| D1-disc-flash | D1discovery_breadth | deepseek-v4-flash | opus-fleet synthesis | partial / n/a | n/a | - | 0 | $0.08553 | failTHRASHED: 186 tool calls vs pro's 20 |
| B-confirm-pro | B-confirmconfirming seat (BANNED for GLM) | deepseek-v4-pro | codex sol/5.5 | 0/20 / 0/6 | n/a | 0 | 0 | $0.30339 | failFAIL, = glm-5.3's 0/20. Affirmatively 'cleared' a surface with 6 known musts |
| C-build-pro | C-buildbuild_rote | deepseek-v4-pro | shipped Sonnet impl | n/a / n/a | n/a | n/a | 0 | $0.08455 | mixedran the cargo-check fence AND touched files to defeat stale-cache false-green |
| F-compile-flash | Fbuild_fix / compile loop (TOTAL ORACLE) | deepseek-v4-flash | glm-4.7 (16/16) | 16/16 / 16/16 | exact | 0 | 0 | $0.01201 | passPASS. RC=0. Refused placeholder shortcut; avoided the debug_assert string-literal trap |
| F-compile-pro | Fbuild_fix / compile loop (TOTAL ORACLE) | deepseek-v4-pro | glm-4.7 (16/16) | 16/16 / 16/16 | exact | 0 | 0 | $0.04969 | passPASS. Identical table + richer rmcp derivation reasoning |
| W1-web-pro | W1web_research (NEW seat) | deepseek-v4-pro | opus+terra web seats | crux EXACT + 8/8 headers + TTLs / n/a | n/a | flagged goog.json as NOT an auth control | 0 | $0.03955 | passSTRONG. Per-claim source URLs + documented-vs-community confidence |
| W1-web-flash | W1web_research (NEW seat) | deepseek-v4-flash | opus+terra web seats | crux EXACT + 9 headers + all 3 TTLs / n/a | n/a | - | 0 | $0.02132 | passSTRONG at $0.02 |
| E-clippy-flash | Eclippy_fix (PARTIAL oracle) | deepseek-v4-flash | glm-4.7 (conditional pass) | 4/4 lint-clearing / n/a | FAIL: took clippy's indent suggestion | 0 | 0 | $0.01038 | failFAIL — re-parented a struct-wide DSAR invariant into one bullet. Identical to glm-4.7's failure |
| E-clippy-pro | Eclippy_fix (PARTIAL oracle) | deepseek-v4-pro | glm-4.7 (conditional pass) | 4/4 / n/a | PASS: chose the paragraph break | 0 | 0 | $0.03318 | passPASS — BYTE-IDENTICAL to human fix c70f4961a on both files; reasoned the lines are a new paragraph |
| Q1-consent-pro-CLEANenforced | Q1checkpoint_confirm heavy | deepseek-v4-pro | gpt-5.6-terra | 2/6 / 2/2 content 1/2 severity | F8 must->should | 1 | 0 | $0.19959 | failCLEAN: found F8 which the unenforced run MISSED |
| Q1-consent-flash-CLEANenforced | Q1checkpoint_confirm heavy | deepseek-v4-flash | gpt-5.6-terra | 3/6 / 2/2 content 1/2 severity | F8 rated 'low' | 1 | 0 | $0.08163 | mixedCLEAN: 3/6 (was 4/6 unenforced) — mild inflation confirmed |
| Q2-consent-pro-CLEANenforced | Q2checkpoint_confirm (consent/PG) | deepseek-v4-pro | gpt-5.6-terra | 1/2 / 1/1 must at CORRECT severity | HIGH = must, correct | 2 | 0 | $0.25953 | passCLEAN PASS on the must: named the unlocked probe + exact advisory-lock fix + S5 precedent |
| Q2-consent-flash-CLEANenforced | Q2checkpoint_confirm (consent/PG) | deepseek-v4-flash | gpt-5.6-terra | 1/2 / 0/1 must | missed the TOCTOU must | 2 | 0 | $0.1107 | failCLEAN: found GT#2 (SENT->FAILED) but still missed the must |
| Q5-iter-pro-CLEANenforced | Q5loop_iterate / R2 confirm | deepseek-v4-pro | gpt-5.5-deep | 1/3 / 1/2 content 0/2 severity | MUST-1 rated 'could' — would not gate | 1 | 0 | $0.30724 | mixedCLEAN: FOUND MUST-1 (was 0/3 unenforced) but mis-weighted it |
| A-r1-pro-CLEANenforced | A-r1r1_additive_lens | deepseek-v4-pro | Sonnet swarm | ~1/8 / must->could (worse) | n/a | 2 | 0 | $0.33694 | mixedCLEAN: same recall, severity downgrade deeper |
| B-confirm-pro-CLEANenforced | B-confirmconfirming seat | deepseek-v4-pro | codex sol/5.5 | ~0/20 / 0/6 | n/a | 0 | 0 | $0.26582 | failCLEAN: unchanged FAIL |
What it really cost — and why the runtime lied about it
from cost.sh + reconcile.shThe harness runtime's self-reported cost runs 8.6x high. The agent runtime's result.total_cost_usd applies its own vendor's price table to the unrecognized DeepSeek model id, so its self-reported cost runs ~8.6x the real DeepSeek bill. The one measured cell: runtime-reported $0.1768 against a real $0.0206 (cost.sh (measured: $0.1768 claimed vs $0.0206 real)).
SEAT_RECOMMENDATION.md: computed is a lower bound (~13% under); the account balance delta is authoritative.
computed high — erring in the safe direction
Subscription quota burned: none — DeepSeek is API-metered and burns no subscription quota.
| model (off-peak, per 1M tok) | cache-miss in | cache-hit in | output |
|---|---|---|---|
| deepseek-v4-pro | $0.66 | $0.022 | $1.98 |
| deepseek-v4-flash | $0.22 | $0.007 | $0.66 |
peak multiplier 2x during 01:00-04:00 UTC and 06:00-10:00 UTC. Whole-bench billed $3.91 over 25 cells — Published page (site/deepseek.html): $3.91 over 25 cells. The earlier recommendation doc (2026-08-20) recorded $2.19 over 14 cells, before the seven enforcement re-runs were added.
Where it may sit — and where it may not
seat by seatCell F: 16/16 exact, lane driven to RC=0. Refused the placeholder shortcut and avoided the debug_assert string-literal trap. $0.012.
Cell E: pro was byte-identical to the human fix c70f4961a; flash took clippy's indent hint and re-parented a struct-wide DSAR invariant. Partial oracle, so pro ONLY. $0.053.
Cell W1: crux exact, all nine notification headers, all three TTLs, a source URL per claim; flagged that matching published IP ranges is not proof. $0.021.
Cell D1: 10/13 vs glm-5.3's 11.5/13, zero hallucinations, byte-exact anchors, three items the ground truth lacked. Additive only, behind GLM, and non-gating so misses are caught downstream.
Q1 2/6, Q2 missed the PG TOCTOU must (flash false-cleaned the unlocked tombstone screen). A confirm that mis-weights a must does not gate.
Q5: found MUST-1 under enforcement but rated it could — would not gate; 1/3 clean.
A-r1: ~1/8, downgraded the ground truth's must to should (deeper under enforcement). Additive at best, not a seat.
B-confirm: 0/20, 0/6 musts, and affirmatively 'cleared' a surface with six known musts. Never.
The rule the bench confirms
A weaker model is safe exactly where an automatic total oracle checks its output, or where its failure is omission and nothing downstream treats it as a fence. Where a human or a stronger model must re-triage, you pay twice and inherit the risk.
Why it is additive, not a substitute
the capability seat1M context / 384K output — nothing else in the fleet has it. A capability seat, not a substitute, and the only pool that can absorb work when both subscriptions are hot.
Verdict
A bounded reserve valve: seat it where a total oracle checks it, never where its severity call must be trusted.
- ·build_fix → flash and clippy_fix → pro-only: a total oracle makes the cheap tier safe; a partial oracle (a doc lint) does not.
- ·web_research → flash and discovery_breadth → pro are additive third/fourth seats, discovery strictly behind GLM and non-gating.
- ·No review, confirm, or iterate seat: under enforcement it found the musts but rated them could — a fence that will not fire.
- ·The defining risk is a silent downgrade: legacy model ids are served as flash with no error, so a stale id quietly demotes a seat.
- ·It consumes no subscription quota and carries 1M context / 384K output — an additive capability pool, not a cheaper GLM.
What changed in the harness because of it
- 01Cost telemetry now reads only the result event, priced by tier and peak window, and reconciles against the account balance delta — the runtime's 8.6x-high self-reported cost is ignored.
- 02The leaf now compares served-vs-requested model and refuses to stay quiet, closing the silent-downgrade trap.
- 03Tool restriction moved from --allowedTools (auto-approve only) to --disallowedTools, and seven cells were re-run under real enforcement.
- 04DeepSeek was ratified (owner, 2026-08-20) for build_fix, clippy_fix, web_research and additive discovery — with no review seat.
Raw data
generated 2026-08-30 by scripts/gen-deepseek.tsEvery number on this page is read from data/deepseek-valve.json, which the generator rewrites from these source files. A data update is a commit.
- data/raw/deepseek/SEAT_RECOMMENDATION.mdSeat verdicts, the governing rule, integrity caveats, and the $2.19/computed-$1.85 reconciliation.
- data/raw/deepseek/results.csvThe trials table — every cell x metric, parsed programmatically.
- data/raw/deepseek/HARNESS-NOTES.mdThe --allowedTools-does-not-restrict finding and the enforcement fix.
- data/raw/deepseek/cost.shTelemetry methodology and the three cost bugs; the $0.1768-vs-$0.0206 8.6x example and off-peak price table.
- data/raw/deepseek/reconcile.shHow computed cost is validated against the account balance delta.
- data/raw/deepseek/brief-W1.mdWhat the web_research cell tested (Google Drive push channels).
- data/raw/deepseek/brief-E.mdWhat the clippy_fix partial-oracle cell tested (doc-lint trap).
- data/raw/deepseek/brief-E2.mdThe round-2 clippy floor cell.
- data/raw/deepseek/brief-F.mdWhat the build_fix total-oracle cell tested (per-site tool-name derivation).
- data/raw/deepseek/brief-GUARD.mdThe Bash-enforcement guard probe.
- data/raw/deepseek/site-deepseek.htmlPublished page: prose framing, the defining-defect table, and the $3.91/25-cell + 7-cell/+3.4% telemetry figures.
- data/raw/deepseek/site-telemetry.htmlFleet telemetry page — quota-burn framing context (deepseek idle on the analyzed night).
- data/raw/deepseek/site-index.htmlProgram framing and the pools table.
- data/raw/deepseek/drosstalk-deepseek-leaf.mdMemory note: 1M/384K roster, the two measured traps, seat ratification (API-key fragment redacted).