Study 03 · published 2026-08-20Bounded · reserve

DeepSeek as the pressure valve

Can an API-metered model give relief when the subscription pools run hot — and where must it stay out of the decision seats?

DeepSeek was run against the GLM bench's own frozen sandboxes, briefs and ground truth, so every number sits directly beside glm-5.3. The seat map is decided by how complete the automatic check is, never by cost.

Identical cells
Each cell reuses a GLM-bench sandbox pinned to the commit a real reviewer had in front of them, with the shipped fix as ground truth — verified to predate its own answer. Every number is directly comparable to glm-5.3.
Two tiers, one file
deepseek-v4-pro and deepseek-v4-flash are run on the same brief. The tier that takes a seat is chosen by fence completeness: a total oracle (rustc) makes flash safe; a partial oracle (a doc lint) does not.
Enforcement re-run
The first pass used --allowedTools, which only auto-approves and never restricted Bash. Seven cells were re-run under --disallowedTools. Contamination could only inflate scores, so the negative findings are robust and the clean numbers are canonical.

The bench in nine numbers

each read from a source file
8.6x
Runtime cost over-report
cost.sh / SEAT_RECOMMENDATION.md
none
Subscription quota burned
SEAT_RECOMMENDATION.md
1000000tokens
Context window
drosstalk-deepseek-leaf.md
384000tokens
Max output
drosstalk-deepseek-leaf.md
16/16
build_fix (cell F)
results.csv / SEAT_RECOMMENDATION.md
10/13
discovery_breadth (cell D1)
results.csv
186tool calls (pro: 20)
flash thrash on D1
results.csv / SEAT_RECOMMENDATION.md
0/200/6 musts
confirming seat (cell B)
results.csv
7cells
cells re-run under enforcement
site/deepseek.html

Every bench cell, scored beside glm-5.3

23 cells

Each cell reuses a frozen GLM-bench sandbox, brief and ground truth. Cells tagged ENFORCED were re-run under a real read-only guard; contamination could only inflate scores, so the clean rows are canonical.

trialcell / seatmodelincumbentrecall pri / mustseverity accuracynovel validhalluc.costverdict
Q1-consent-proQ1checkpoint_confirm heavy (consent)deepseek-v4-progpt-5.6-terra3/6 / 1/21 downgrade (F10 should->could)1 (restore-replay; cite verified)0$0.15017false-cleanWEAKER than glm-5.3 (4/6,2/2m). No false-clean. Missed F8 envelope_mac must
D1-disc-proD1discovery_breadthdeepseek-v4-proopus-fleet synthesis10/13 (77%) / n/an/a3 strong (scope-firewall drive.file vs channels.watch; binding-grain partial-unique; initial-sync re-import)0$0.13379passSTRONG. Below glm-5.3 11.5/13 but complementary; all spot-checked anchors byte-exact
Q5-iter-proQ5loop_iterate / R2 confirmdeepseek-v4-progpt-5.5-deep0/3 / 0/2all peripheral could/nit00$0.27005failFAIL as confirm/iterate. Missed BOTH musts. Worse than glm-5.3 (2/3 gating)
A-r1-proA-r1r1_additive_lensdeepseek-v4-proSonnet swarm + sol breadth>=1/8 / partial1 downgrade (R1-F1 must->should)TBD0$0.23426mixed~parity with glm-5.3 (1/8). Additive lens at best
Q2-consent-proQ2checkpoint_confirm (consent/PG)deepseek-v4-progpt-5.6-terra0/2 / 0/1n/a00$0.2351failMISSED the PG TOCTOU must
Q2-consent-flashQ2checkpoint_confirm (consent/PG)deepseek-v4-flashgpt-5.6-terra0/2 / 0/1n/a00$0.1391false-cleanFALSE-CLEAN: declared the unlocked tombstone screen 'sound'
Q1-consent-flashQ1checkpoint_confirm heavydeepseek-v4-flashgpt-5.6-terra4/6 / 2/2 content 1/2 severitymust rated 'low'; used high/med/low vocab10$0.08791passBEAT pro (3/6). Severity vocabulary non-conforming
D1-disc-flashD1discovery_breadthdeepseek-v4-flashopus-fleet synthesispartial / n/an/a-0$0.08553failTHRASHED: 186 tool calls vs pro's 20
B-confirm-proB-confirmconfirming seat (BANNED for GLM)deepseek-v4-procodex sol/5.50/20 / 0/6n/a00$0.30339failFAIL, = glm-5.3's 0/20. Affirmatively 'cleared' a surface with 6 known musts
C-build-proC-buildbuild_rotedeepseek-v4-proshipped Sonnet impln/a / n/an/an/a0$0.08455mixedran the cargo-check fence AND touched files to defeat stale-cache false-green
F-compile-flashFbuild_fix / compile loop (TOTAL ORACLE)deepseek-v4-flashglm-4.7 (16/16)16/16 / 16/16exact00$0.01201passPASS. RC=0. Refused placeholder shortcut; avoided the debug_assert string-literal trap
F-compile-proFbuild_fix / compile loop (TOTAL ORACLE)deepseek-v4-proglm-4.7 (16/16)16/16 / 16/16exact00$0.04969passPASS. Identical table + richer rmcp derivation reasoning
W1-web-proW1web_research (NEW seat)deepseek-v4-proopus+terra web seatscrux EXACT + 8/8 headers + TTLs / n/an/aflagged goog.json as NOT an auth control0$0.03955passSTRONG. Per-claim source URLs + documented-vs-community confidence
W1-web-flashW1web_research (NEW seat)deepseek-v4-flashopus+terra web seatscrux EXACT + 9 headers + all 3 TTLs / n/an/a-0$0.02132passSTRONG at $0.02
E-clippy-flashEclippy_fix (PARTIAL oracle)deepseek-v4-flashglm-4.7 (conditional pass)4/4 lint-clearing / n/aFAIL: took clippy's indent suggestion00$0.01038failFAIL — re-parented a struct-wide DSAR invariant into one bullet. Identical to glm-4.7's failure
E-clippy-proEclippy_fix (PARTIAL oracle)deepseek-v4-proglm-4.7 (conditional pass)4/4 / n/aPASS: chose the paragraph break00$0.03318passPASS — BYTE-IDENTICAL to human fix c70f4961a on both files; reasoned the lines are a new paragraph
Q1-consent-pro-CLEANenforcedQ1checkpoint_confirm heavydeepseek-v4-progpt-5.6-terra2/6 / 2/2 content 1/2 severityF8 must->should10$0.19959failCLEAN: found F8 which the unenforced run MISSED
Q1-consent-flash-CLEANenforcedQ1checkpoint_confirm heavydeepseek-v4-flashgpt-5.6-terra3/6 / 2/2 content 1/2 severityF8 rated 'low'10$0.08163mixedCLEAN: 3/6 (was 4/6 unenforced) — mild inflation confirmed
Q2-consent-pro-CLEANenforcedQ2checkpoint_confirm (consent/PG)deepseek-v4-progpt-5.6-terra1/2 / 1/1 must at CORRECT severityHIGH = must, correct20$0.25953passCLEAN PASS on the must: named the unlocked probe + exact advisory-lock fix + S5 precedent
Q2-consent-flash-CLEANenforcedQ2checkpoint_confirm (consent/PG)deepseek-v4-flashgpt-5.6-terra1/2 / 0/1 mustmissed the TOCTOU must20$0.1107failCLEAN: found GT#2 (SENT->FAILED) but still missed the must
Q5-iter-pro-CLEANenforcedQ5loop_iterate / R2 confirmdeepseek-v4-progpt-5.5-deep1/3 / 1/2 content 0/2 severityMUST-1 rated 'could' — would not gate10$0.30724mixedCLEAN: FOUND MUST-1 (was 0/3 unenforced) but mis-weighted it
A-r1-pro-CLEANenforcedA-r1r1_additive_lensdeepseek-v4-proSonnet swarm~1/8 / must->could (worse)n/a20$0.33694mixedCLEAN: same recall, severity downgrade deeper
B-confirm-pro-CLEANenforcedB-confirmconfirming seatdeepseek-v4-procodex sol/5.5~0/20 / 0/6n/a00$0.26582failCLEAN: unchanged FAIL

What it really cost — and why the runtime lied about it

from cost.sh + reconcile.sh

The harness runtime's self-reported cost runs 8.6x high. The agent runtime's result.total_cost_usd applies its own vendor's price table to the unrecognized DeepSeek model id, so its self-reported cost runs ~8.6x the real DeepSeek bill. The one measured cell: runtime-reported $0.1768 against a real $0.0206 (cost.sh (measured: $0.1768 claimed vs $0.0206 real)).

Input double-counted 2x
The runtime emits each assistant message twice with the same id; summing assistant usage events double-counts input exactly 2x.
Output read as zero
Per-assistant-event output_tokens is 0 — output is only totalled on the result event, so summing assistant events drops the single most expensive component.
Wrong vendor's price table
result.total_cost_usd prices an unrecognized model with the runtime vendor's own table, ~8.6x high. Cost is now taken only from the result event, priced by tier and peak window.
Reconciliation vs the account balance
$1.85computed vs$2.19billed-13%

SEAT_RECOMMENDATION.md: computed is a lower bound (~13% under); the account balance delta is authoritative.

Later enforcement-clean window
3.4%over7 cells

computed high — erring in the safe direction

Subscription quota burned: none — DeepSeek is API-metered and burns no subscription quota.

model (off-peak, per 1M tok)cache-miss incache-hit inoutput
deepseek-v4-pro$0.66$0.022$1.98
deepseek-v4-flash$0.22$0.007$0.66

peak multiplier 2x during 01:00-04:00 UTC and 06:00-10:00 UTC. Whole-bench billed $3.91 over 25 cells — Published page (site/deepseek.html): $3.91 over 25 cells. The earlier recommendation doc (2026-08-20) recorded $2.19 over 14 cells, before the seven enforcement re-runs were added.

Where it may sit — and where it may not

seat by seat
build_fix / compile-error loop
seat
tier deepseek-v4-flashoracle total

Cell F: 16/16 exact, lane driven to RC=0. Refused the placeholder shortcut and avoided the debug_assert string-literal trap. $0.012.

clippy_fix / code and doc lints
seat
tier deepseek-v4-prooracle partial

Cell E: pro was byte-identical to the human fix c70f4961a; flash took clippy's indent hint and re-parented a struct-wide DSAR invariant. Partial oracle, so pro ONLY. $0.053.

web_research (third, additive)
seat
tier deepseek-v4-flash

Cell W1: crux exact, all nine notification headers, all three TTLs, a source URL per claim; flagged that matching published IP ranges is not proof. $0.021.

discovery_breadth (added, never swapped)
seat
tier deepseek-v4-prooracle none

Cell D1: 10/13 vs glm-5.3's 11.5/13, zero hallucinations, byte-exact anchors, three items the ground truth lacked. Additive only, behind GLM, and non-gating so misses are caught downstream.

checkpoint_confirm
refused
tier —oracle none

Q1 2/6, Q2 missed the PG TOCTOU must (flash false-cleaned the unlocked tombstone screen). A confirm that mis-weights a must does not gate.

loop_iterate / R2 confirm
refused
tier —oracle none

Q5: found MUST-1 under enforcement but rated it could — would not gate; 1/3 clean.

r1_additive_lens
refused
tier —oracle none

A-r1: ~1/8, downgraded the ground truth's must to should (deeper under enforcement). Additive at best, not a seat.

confirming seat / heavy R3 / P4 confirm
refused
tier —oracle none

B-confirm: 0/20, 0/6 musts, and affirmatively 'cleared' a surface with six known musts. Never.

The rule the bench confirms

Governing rule

A weaker model is safe exactly where an automatic total oracle checks its output, or where its failure is omission and nothing downstream treats it as a fence. Where a human or a stronger model must re-triage, you pay twice and inherit the risk.

Why it is additive, not a substitute

the capability seat
1000000context tokens
384000max output tokens

1M context / 384K output — nothing else in the fleet has it. A capability seat, not a substitute, and the only pool that can absorb work when both subscriptions are hot.

Verdict

Bounded · 2026-08-20

A bounded reserve valve: seat it where a total oracle checks it, never where its severity call must be trusted.

  • ·build_fix → flash and clippy_fix → pro-only: a total oracle makes the cheap tier safe; a partial oracle (a doc lint) does not.
  • ·web_research → flash and discovery_breadth → pro are additive third/fourth seats, discovery strictly behind GLM and non-gating.
  • ·No review, confirm, or iterate seat: under enforcement it found the musts but rated them could — a fence that will not fire.
  • ·The defining risk is a silent downgrade: legacy model ids are served as flash with no error, so a stale id quietly demotes a seat.
  • ·It consumes no subscription quota and carries 1M context / 384K output — an additive capability pool, not a cheaper GLM.

What changed in the harness because of it

  1. 01Cost telemetry now reads only the result event, priced by tier and peak window, and reconciles against the account balance delta — the runtime's 8.6x-high self-reported cost is ignored.
  2. 02The leaf now compares served-vs-requested model and refuses to stay quiet, closing the silent-downgrade trap.
  3. 03Tool restriction moved from --allowedTools (auto-approve only) to --disallowedTools, and seven cells were re-run under real enforcement.
  4. 04DeepSeek was ratified (owner, 2026-08-20) for build_fix, clippy_fix, web_research and additive discovery — with no review seat.

Raw data

generated 2026-08-30 by scripts/gen-deepseek.ts

Every number on this page is read from data/deepseek-valve.json, which the generator rewrites from these source files. A data update is a commit.

I. Nocturne in E♭ major Op. 9 № 2 · Chopin

0:00 / 4:16