DeepSeek Harness Plugin Hub

Publish and manage complete Harness Profiles. Discover Plugins for your next setup.

Explore

PluginsPresetsDocsNews

Community

Publish a pluginContactReport an issue

Resources

Plugin Hub on GitHubDeepSeek HarnessSystem statusPrivacy notice
© 2026 DeepSeek Harness Plugin HubPowered byPaxTech

Independent and unofficial. Not affiliated with, authorized by, or endorsed by DeepSeek.

Plugin Save Token — DSH Plugin for DeepSeek Harness
← Plugins

dsh-plugin-save-token

Plugin Save Token

Cut token cost without cutting model intelligence — a DeepSeek Harness (dsh) bundle plugin: structure-aware compression + TOON-style lossless encoding first + reversible spill-to-disk + real-time dashboard.

The plugin will be installed here. Keep web if you are unsure.

npx -y @deepseek-ai/dsh plugin --profile web add dsh-plugin-save-token@2.4.1
READMECompatibilityVersions

Compatibility and provenance

Plugin Save Token is published as dsh-plugin-save-token and currently resolves to version 2.4.1. The Hub verifies its manifest and preserves the exact installation source for reproducible installs.

DSH compatibility
*
Runtime surfaces
web
Release source
npm
Registry updated
9/20/2026

Versions

2.4.1stable
9/4/2026
2.1.3stable
8/29/2026
2.1.2stable
8/27/2026
Show 2 more versionsCollapse versions
2.1.1stable
8/27/2026
2.1.0stable
8/27/2026

Related plugins

Loading related plugins…

Latest
2.4.1
DSH
*
HMR
Process restart
Tree shaking
Safe tree shaking not declared
Unpacked size
299.4 kB
Files
12
Surface
web
License
MIT
Source
npm
GitHub
★ 5
Weekly downloads
277
Security scan
✓ v2.4.1 scan passed
Last push
9/6/2026
View source ↗Project homepage ↗
README badge

Click the badge to copy Markdown for your README.

Do you maintain this Plugin?Claim benefit · Priority security scan

Verify the GitHub repository declared in package.json to manage this listing. After you claim it, Hub will prioritize a security scan of the current version and publish the result when it passes.

Claim this Plugin →
Report an issue
DeepSeek Harness Plugin Hub
ProfilesPluginsCategoriesNewsDocsSign inManage Profiles
ProfilesPluginsCategoriesNewsDocsSign in

Related plugins

More verified plugins in models-usage.

Usage@linxin666/dsh-usageUsage statistics plugin for the dsh web GUI: per-provider balance and coding-plan quota detection plus a live token usage ledger, with a dedicated pet bubble for the current providerWhale Widgetdsh-whale-widgetDeepSeek balance whale widget in the bottom-right corner of the DSH Web interface: balance/today’s usage/peak-off-peak pricing, customizable bubble click sequence (text/balance/today/peak-off-peak/image/random phrases and parallel weighted selection), per-line styles and fonts, floating quick editinUsage Stats@ychris12138/dsh-usage-statsToken usage heatmap, provider balances, and subscription quotas for the dsh web GUICodex Connectdsh-codex-connectChatGPT OAuth and Codex models for DeepSeek Harness.

README

dsh-plugin-save-token

English | 简体中文

In one sentence: a DeepSeek Harness (dsh) dynamic plugin that cuts token cost without cutting model intelligence.

It slims down oversized tool outputs at the entrance of every model request — reversibly and structure-aware. The full original text is always saved to disk; what the model sees is always a condensed version carrying a retrieval path. Every optimization obeys one red line: any replacement must be restorable in one step, and the estimated token count after compression must be strictly smaller than the original.


Why you need it

  • The bulk of agentic-session cost is tool output that keeps re-entering the context: large JSON from APIs, CLI tables, logs. The same 40KB price table may be billed again on every conversation turn.
  • But blunt truncation degrades intelligence: research shows that even with perfect retrieval, merely padding the context with irrelevant content drops accuracy by 13.9–85% (curated in "Context Length Alone Hurts"); conversely, over-compression fails too — in production randomized controlled trials, aggressive compression (keep ratio 0.2) actually made cost +1.8% worse, while moderate compression (0.5) delivered −27.9%.
  • Conclusion: the right way to save tokens is structure-preserving slimming, not content hacking. All strategies in this plugin are designed within that boundary.

Core optimizations

1. Structure-aware compression: no more blind table slicing

Detects pipe-delimited high-density table rows (≥70% of lines sharing the same separator profile). When matched, instead of a blind head/tail window it does: keep the first 60 rows verbatim + stride-sample the middle section every N rows with original line numbers annotated (L61: ...) + keep the last 40 rows verbatim. The model receives a "map with coordinates" — any segment can be fetched precisely by line number.

Reference: rtk's never-worse guard and error-line retention policy; the stride sampling is this plugin's improvement over rtk's head/tail window.

2. TOON-style lossless encoding first

Uniform JSON arrays (e.g. API responses with 300 homogeneous objects) first go through deterministic tabular re-encoding: prices[300]{model,input,output}: — one schema header line + CSV data rows. Keys are written once, zero information loss, and the notice explicitly says "zero information loss". Lossy paths are only used when the lossless route is unavailable.

Reference: TOON — Token-Oriented Object Notation; measured savings of 30–60% tokens on uniform arrays.

3. Compress means spill (CCR): never burn the bridge

Before every replacement, the original text is written to disk via the dsh spillStore; the replacement embeds two retrieval paths: the dynamic tool save_token_expand (one-call fetch by marker id) plus a file locator for the original (readable directly with read/grep). If spilling fails, compression is abandoned — reversibility is a hard precondition, not an option.

Reference: headroom's CCR (Compress-Cache-Retrieve) pattern.

4. Cross-turn dedup

Tool-call results byte-identical within a 90-second window (rerunning the same command, etc.) are replaced by one stub: "This output is identical to N seconds ago, refer to earlier context." Prevents the same large output from appearing twice in the context.

Reference: headroom's cross-turn dedup.

5. Never-worse double gate

A candidate compressed result is adopted only if it passes both gates:

  • Byte gate: compressed ≤ 72% of the original (keepRatioMax=0.72, more conservative than the RCT-validated 0.5), and absolute savings ≥500B;
  • Token gate: estimated tokens must strictly decrease (an llmtrim-style quality-gating idea). If either gate fails, the output passes through untouched.

Reference: RCT boundary data and quality-gating survey in awesome-llm-token-optimization.

6. Error-line protection

Within the omitted region of log-like output, up to 25 lines matching error/fatal/traceback/timeout... are kept (with line-number prefixes). Debugging evidence is never compressed away.

Reference: rtk's error-line keeps.

7. Compaction pressure coupling

At each reasoning-step boundary, check the session's most recent actual context size; above 120k tokens (10-minute cooldown), fire dsh's native compaction.compactIfNeeded() with 'pressure' and let the engine decide when to summarize history.

The threshold is a conservative water line (sized for 128k-class context windows); compaction itself is built into dsh — the plugin only hands over the trigger at the right moment.

8. Full-chain metering + dual-panel dashboard

Every llm/stream is intercepted: real billed tokens (input/cached/output/reasoning) and "tokens avoided from context" are accounted separately. Historical messages are scanned for [save-token #id] markers to total savings (including multi-turn replays). The Settings page hosts a full panel (KPIs, per-request stacked chart, top-tools leaderboard, activity feed), plus a persistent live strip under the input box.


Measured results

End-to-end A/B on real agent tasks (2026-08-29)

Randomized comparison on GAIA / Terminal-Bench / SWE-bench-Verified tasks — unique variable: the plugin's compress/dedupe switches; both arms under identical constraints that force large tool output to be printed directly into the conversation. Full data, scripts and per-episode records: bench/, report: bench/report_2026-08-29.md.

EpisodesSuccess rateTotal tokensCompressionsPer-episode median
24/24 (6 tasks × 2 arms × n=2)100% vs 100%5.25M vs 4.32M (−17.6%)16 events across 8/12 episodes−56%

Take-aways: the plugin pays off exactly when large tool output lands directly in the context (verbose test runs, raw log/JSON dumps); when agents go through the write-file-then-read pattern it never triggers — and costs nothing. Success rate was never hurt.

End-to-end A/B, round 2 — the optimized build (2026-09-03, v2.4.1)

Same harness re-run after the v2.4.1 optimizations (lossless-TOON-first pipeline, gated fallback, cache-aware layer), this time on a Linux host. 24/24 valid episodes again; unique variable unchanged. Report: bench/report_2026-09-03.md, records: bench/results/raw2/, summary: bench/results/summary-r2.md.

Metricbaselinetreatment (v2.4.1 on)
Success rate100% (12/12)100% (12/12) — 48/48 across both rounds, zero damage re-confirmed
Provider cache-hit rate90.0%90.7% — held at 91.1% even in the heaviest episode (1.86M tokens of repeated dump/expand cycles)
Tokens avoided (plugin estimate)—~560k — 11 compressions across 6/12 episodes (13.8% of their tokens)
SWE long-context tasks (sympy / django)—−19.7% / −18.0%, both repeats same direction

Take-aways: on the optimized build the value proposition sharpens — savings concentrate exactly where context is longest and output is largest (SWE-style agent runs), the cache-aware marker-replay layer keeps provider cache hits stable even under worst-case repeated re-reading, and success rate remains untouched. As with any n=2 study, one agent-side strategy outlier can outweigh per-output savings in the aggregate — see the report's paired per-task analysis for the breakdown.

Single-event compression strength

InputBeforeAfterStrategy
CLI price table (400-line pipe table)41,727 B15,191 B (−64%)Structure-aware: verbatim head/tail + stride-sampled middle with original line numbers
Model-price JSON registry (300-item uniform array)34,000 B19,935 B (−41%)TOON lossless route, zero information loss

Anti-pattern on record: an early version once applied blind head/tail windowing to a 35.5KB LiteLLM price registry; subagents couldn't locate middle rows and re-queried repeatedly — that failure is why the structure-aware strategy exists.

Single-event figures were measured in a development environment; session-level gains depend on how much of your workload is large tool output delivered directly to the model (write-file-then-read patterns bypass compression by design, at zero cost).


Installation & usage

Requirements: a running DeepSeek Harness (dsh) with its Web GUI, and pnpm on PATH. The web profile provides everything else the plugin needs (tools, webServer, React for the dashboard; spillStore is included in standard deployments — if it is ever missing, compression stays off by design).

Install

Run one of these commands — dsh plugin installs the package into the profile and activates its bundle layer automatically:

# from the npm registry
dsh plugin --profile web add dsh-plugin-save-token

# or straight from GitHub
dsh plugin --profile web add github:vibe-any/dsh-plugin-save-token

# or from a local checkout
dsh plugin --profile web add /absolute/path/to/dsh-plugin-save-token

That's the whole installation: no prompts to paste into the GUI, no dynamic-code authorization dialogs. Verify it's in the roster with dsh --profile web --dump-config | grep save-token, then restart the running dsh instance (ESM caches are per-process).

Removal: dsh plugin --profile web remove dsh-plugin-save-token.

Using it

Once installed there is nothing to operate: open Settings → Token Saver for the full panel, and look for the persistent live strip under the input box. Three toggles (Compress / Dedupe / Compact@120k) switch right on the panel. The panel and the strip follow dsh's language setting (Settings → General → Language: English / 简体中文).

Config defaults (the config: block of the save-token row in cordis.patch.yml; code fallbacks live in src/index.js)

ParameterDefaultMeaning
minBytes1400Minimum size for ordinary outputs to enter compression
errorMinBytes6000Higher threshold for error output (leave debugging scenes alone)
keepRatioMax0.72Byte-gate cap: compressed must not exceed 72% of original
maxLines / headLines / tailLines240/140/80Window shape for ordinary long outputs
tabularHeadRows / tabularTailRows / tabularStrideSamples60/40/50Retention and sampling density in table mode
longLineChars420Head/tail truncation threshold for single oversized lines
jsonlMinLines8Minimum uniform-object lines for the JSONL/NDJSON lossless route
noticeFullTrailerCount3First N adopted compressions carry the verbose retrieval notice; later ones use the compact trailer (same id + locator)
dedupeTtlMs600000Validity window for cross-turn dedup (fingerprints are byte-exact, so a longer window is information-safe)
dedupeTtlOverrides{}Per-tool TTL in ms; 0 disables dedupe for that tool (freshness-sensitive commands)
compactAssistEnabledfalseCompaction coupling master switch — off by default (see cache note below)
compactBudgetTokens / compactCooldownMs120000/600000Absolute fallback watermark and cooldown for the compaction assist
contextWindowTokens / compactWatermarkRatio0/0.85When the model window is known, the watermark is window × ratio instead of the absolute budget

v2.2.0 behavior notes:

  • Dedupe keys carry the owning session id and hash the FULL args/content strings (long shared prefixes can no longer produce false "byte-identical" stubs; two sessions sharing one process never see each other's stubs).
  • save_token_expand output is exempt from compression — unfolding a notice can never hand back the same elided preview again.
  • Lossless routes extended: JSONL/NDJSON logs, nested field groups (pos{x,y}), keyed maps, and a depth-2 dominant-array search ({data:{items:[...]}}). Lossless wins whenever it passes the never-worse gates; the lossy elision candidate is generated as the gate-checked fallback before the line compressor, and lossy notices disclose exactly what was omitted.
  • Compression counters increment on adopted candidates (previously on attempts that the gates could still reject).

v2.3.0 cache-aware layer (bench evidence: ~90% of measured input tokens were provider cache reads, billed at ~1/30 of the miss price on DeepSeek):

  • Compaction assist defaults to OFF and is repositioned as an anti-overflow measure, not a saver: summarizing 120k→40k tokens converts cheap cached replay into full-price input and breaks even only after ~60 further requests. Turn it on when sessions actually grow past the watermark; do not expect it to cut spend. The toggle and watermark are honest in the panel.
  • The watermark prefers the last real billed input for the session (main requests only) over the heuristic estimate, and scales with the model window (contextWindowTokens × compactWatermarkRatio) when configured.
  • Cache-hit sentinel KPI: cacheRead/cacheWrite are metered separately and the panel shows the cache-hit percentage. If a future change tanks that number, it is saving tokens while silently raising real cost.
  • Online calibration: a per-model EMA of billed/estimated tokens (learned from real usage each request) corrects the avoided-token accounting — no bundled tokenizer. The compression token gate needs no calibration (the ratio cancels in that comparison).

v2.4.0 closing items:

  • Dedupe TTL default 90s → 600s with per-tool overrides (dedupeTtlOverrides, 0 opts a tool out entirely).
  • Error-line protection in plain-text windows widened to ±1 context line (25 anchors, adjacent anchors merged): a bare assertion line rarely explains itself; the neighboring test name / stack header is what saves a re-run.
  • save_token_expand survives eviction and restarts: an id→locator side index outlives the text cache, so a miss returns the spill locator (with a transparent spillStore.readText attempt when the host offers one) instead of a dead end.
  • Top-level vs nested tool calls are counted on the panel — the nesting exemption currently skips compression for subagent-internal calls, and this counter finally quantifies that unexploited surface before anyone flips it.

How it works (30-second version)

tool returns ──► tools/post-execute (prepend)
             ├─ size ≤ threshold? ────────── pass through
             ├─ byte-identical within 90s? ─ spill original → replace with dedup stub
             ├─ JSON with uniform array? ─── TOON lossless re-encode (zero loss)
             ├─ JSONL of uniform objects? ── one TOON table for the whole log
             ├─ pipe/tab table shape? ────── stride-sampled window with line numbers
             └─ other long text ──────────── head/tail window + error-line protection
                      │  double gate per route: ≤72% bytes AND tokens strictly decrease
                      │  (lossless first; lossy elision is the gate-checked fallback)
                      ▼
             spill original to spillStore → inject [save-token #id] retrieval notice
                      ▼
every model request ◄── llm/stream metering (real billing + avoided tokens, per-model calibrated)
step boundaries ──► assist on AND billed > watermark? ──► compaction.compactIfNeeded('pressure')

All compression logic lives in src/compress.js (pure, side-effect free) and is pinned by the unit-test suite: npm test (node --test, zero extra dependencies).

Directory layout

dsh-plugin-save-token/
├── README.md             ← this file (English, default entry)
├── README.zh-CN.md       ← Chinese documentation
├── manifest.json         ← metadata + config defaults
├── package.json          ← npm manifest declaring dsh.bundle + ./client export
├── cordis.patch.yml      ← the bundle layer inserted into the profile roster
├── build.mjs             ← esbuild script producing lib/
├── src/
│   ├── index.js          ← Host half: waterfall hooks / orchestration / tool registration / API routes
│   ├── compress.js       ← pure compression brain (estimators, TOON tabular, gates, notices)
│   └── client/index.js   ← Client half: Dashboard panel + input-box live strip
├── test/                 ← node --test unit suite pinning the compression brain
└── lib/                  ← built artifacts (committed, so git installs need no build step)
    ├── index.js          ← bundled ESM host half (node)
    └── client.js         ← bundled client half wrapped in window.__ModuleLoader__.load({ id, factory })

Design red lines ("no dumbing down" promises)

  1. Reversible: failed disk write = abandon compression; read and save_token_expand outputs are never processed (by-design exemptions).
  2. Lossless first: lossless wins whenever it passes the never-worse gate; lossy elision runs only as the gate-checked fallback and its notices disclose what was omitted.
  3. Double gate: every replacement must prove itself "smaller in bytes AND cheaper in tokens," or it passes through.
  4. Error protection: high thresholds around failure scenes, mandatory retention of error lines (±1 context line).
  5. Cache-stable: compression happens once, at tool-result entry; history stays byte-stable afterwards, so the provider prompt cache keeps hitting (measured: ~90% of input tokens were cache reads at ~1/30 price, and 91.1% even in the heaviest dump-heavy episode). Replay-time rewriting of history is out of scope by design — it lowers the token meter while raising the real bill.