DeepSeek Harness Plugin Hub

发布与管理完整 Harness Profiles,发现适合你的插件。

探索

插件目录环境预设文档中心动态

社区

发布插件联系我们报告问题

相关链接

Plugin Hub GitHubDeepSeek Harness 官方项目系统状态隐私说明
© 2026 DeepSeek Harness Plugin HubPowered byPaxTech

独立、非官方社区项目,与 DeepSeek 官方无隶属、授权或背书关系。

Supreme — DeepSeek Harness 插件(DSH Plugin)
← Plugins
S

dsh-supreme

Supreme

面向主机的七个 DeepSeek Harness 策略插件,通过 DSH 内置的 Cordis 运行时进行组合。可作为 dsh bundle 安装:dsh plugin --profile <name> add <this package>。固定上游版本:deepseek-ai/deepseek-harness @ d347e703908d0406b7a7ef80e3a0e594d86b2215。

插件会安装到这里;不确定时保持 web。

npx -y @deepseek-ai/dsh plugin --profile web add github:stadeummwt/dsh-supreme#08ebd146a733168a55856699843ffc2a53adaaeb
README兼容性版本

兼容性与来源证明

Supreme 以 dsh-supreme 发布,当前版本为 1.3.3。Plugin Hub 会校验它的 manifest,并保存精确安装来源,便于复现安装结果。

DSH 兼容范围
*
运行环境
any
发布来源
github
Registry 更新时间
2026/9/12

版本

1.3.3stable
2026/9/12
1.3.2stable
2026/9/9
1.3.1stable
2026/9/9
查看其余 3 个版本收起版本
1.3.0stable
2026/9/9
1.2.1stable
2026/9/8
1.2.0stable
2026/9/8

相关插件

正在加载相关插件…

最新版
1.3.3
DSH
*
HMR
重启进程
Tree shaking
未声明可安全裁剪
解包体积
未提供
文件数
未提供
Surface
any
许可证
MIT
发布源
github
GitHub
★ 0
周下载
0
最近提交
2026/9/12
查看源码 ↗
README Badge

点击下方 Badge 复制 Markdown,粘贴到 README 即可。

这是你的 Plugin?认领权益 · 优先安全扫描

验证 package.json 声明的 GitHub 仓库,即可管理这个公开页面。认领后,Hub 会优先安排当前版本的安全扫描,并在通过后公开展示结果。

认领这个 Plugin →
报告问题
DeepSeek Harness Plugin Hub
ProfilesPlugins分类动态文档登录管理 Profiles
ProfilesPlugins分类动态文档登录

相关插件

继续浏览 security-access 分类下经过校验的插件。

Doctor@linxin666/dsh-doctorDSH 配置档案的事务性救援模式,配备受监督的启动器、隔离的恢复容器、确定性修复、健康监控以及本地 Web 恢复控制台Pocketdsh-pocket把 DeepSeek Harness 装进你的口袋:一个包、一个设置页,手机扫码即同步访问电脑上的 DSH(局域网 + 公网,实时同屏)。DSCODE@toddzheng024/dscode-bundle完整的 DeepSeek 编码代理,支持持久化 shell、Ultra 协作和自动权限审查。Auto Reviewdsh-auto-review针对 DeepSeek Harness 审批请求的第二模型 AI 自动审查:只读审查子代理在审批应答链上决定允许或拒绝,并采用故障关闭回退机制和完整的会话日志审计。

README

DSH SUPREME — the governance layer for DeepSeek Harness: seven policy plugins, one bundle install, zero upstream patches, every claim executable

🛡️ DSH SUPREME

The governance layer for DeepSeek Harness

"ECC gives your harness breadth. Supreme gives it a conscience."

Install · dsh plugin --profile <your-profile> add github:stadeummwt/dsh-supreme

At a glance · Why Supreme · Install · The seven plugins · Proof wall · Security · v1.3.1 fixes · Benchmarks · Docs · FAQ


📊 At a glance

Every row below is re-runnable — see the proof wall.

MetricValue
Full suite101/101 Level-A checks · 5/5 real-loader boots · VERDICT COMPLETE
v1.3.1 review probes465/465 across 7 verifiers
v1.3 E2E probes184/184 (policy 85 · workflow 82 · routing 17)
Verdict markers14 green (suite · v131:verify · v13:verify · bundle:verify · composition:verify · v12:verify · v3:verify)
Policy benchmark A/B/C0 escapes after patch · benign 75/75 · Astra NOT_RUN (honest label)
Router latency≈ 0.02–0.04 ms / 1k decisions · RM0-first
Secret sentinel leaks0 every run
Upstream patches0 — pinned d347e703908d, worktree clean
LicenseMIT

🤔 Why Supreme

DSH's plugin ecosystem (3,421 catalog entries reviewed, 2026-09) is rich in single-domain tools — a router here, a memory store there, a verifier somewhere else. Each solves one slice of governance and asks you to trust its output.

Supreme is the opposite design. It is a full governance stack — cost policy, observability, routing, verification, memory policy, workflow limits, security audit — that treats proof as a product feature: every claim in this README maps to a command you can run, and every hard rule (deny paths, cost gates, secret scrubbing) is deterministic code, not model judgment.

The usual DSH pluginDSH Supreme
Solves one domainSeven governance domains, one install
"Trust the output"Verdict gates — COMPLETE only when every check passes
Config verified by vibesConfig-key hygiene scan against real zod schemas (silent-strip trap closed)
Markdown evidencePublished JSON Schemas + append-only JSONL evidence stores
Touches core or monkey-patchesZero upstream patches — pinned upstream, worktree clean, verified every run
Security as a README paragraphSix-surface security audit (prompts · hooks · MCP · permissions · secrets · agent files) in CI
No ML dependencyAlso no ML — deterministic counting, globs and comparisons only. Speed is a feature: router decision ≈ 0.02–0.03 ms / 1k iterations

The core rule of this repo: bukti sebenar > klaim — real evidence over claims. If a statement here can't be re-run by you, it's marked as a claim, not a fact.


⚡ 60-second install

The repository is a dsh bundle — no build step needed (dist/ is committed):

dsh plugin --profile <your-profile> add github:stadeummwt/dsh-supreme

That mounts all seven plugins with safe production defaults:

PAID / TRIAL routes  → DENIED (hard rule, LAB-only override)
UNKNOWN cost class   → DENIED
commands / network   → OFF by default
router candidates    → you add yours in your own patch layer (last write wins)

Zero-thought path: one command does everything

Don't want to think about environments, builds, or profiles at all? The built-in operator CLI (zero dependencies) diagnoses, installs, composes and boot-proves your setup:

node real/supreme.mjs doctor    # what's missing? (prints a fix line per check)
node real/supreme.mjs setup     # EVERYTHING: builds what's missing, runs the real
                                # `dsh plugin add`, applies the composition,
                                # boot-probes it → "SUPREME READY"
node real/supreme.mjs verify    # full verification ladder, PASS/FAIL per gate
node real/supreme.mjs setup --composition standard   # core|standard|supreme|lab

setup is idempotent and never modifies the upstream checkout. It will clone + pin + build the pinned DSH upstream only if it is missing (skip with --no-upstream-build).

v1.3.2 note (Windows): if an earlier version showed bundle:verify … obs stats 0 or a composition:verify dataDir failure — root cause found and fixed (store engines created their parent directory with a POSIX-only separator check; on Windows every record write was silently dropped). Re-run setup (rebuilds dist/), then re-run the verifiers — they now also print full writer stats + an exact diagnosis instead of a bare zero.

Pick a composition in one more line if you don't need all seven:

FragmentActive pluginsUse it for
corepolicygovernance floor on any profile
standardpolicy · observability · memory · verifierdaily-driver
supremeall sevenfull stack
laball seven + LAB overridesexperiments only — never production
dsh --profile <your-profile> \
  --patch "$DSH_HOME/profiles/<your-profile>/node_modules/dsh-supreme/config/compositions/standard.patch.yml"

🧩 The seven governance plugins

Exactly seven. The scope is frozen (AGENTS.md) — no scope creep without a proven blocker.

#PluginServiceWhat it enforces
1supreme-policysupremePolicyCost-class / risk / delegation admission. UNKNOWN cost ⇒ DENY. Paid & trial overrides are LAB-only. Unicode-taint + encoding-blob detection & denial. CoT presence gate with visibility profiles + risk gating. Deny-circumvention (deny_retry) guard. Capability-class gate.
2supreme-observabilitysupremeObservabilityAppend-only JSONL metadata log over official DSH event seams. Allowlisted fields, secret-sentinel scrub, fail-open.
3supreme-benchmarksupremeBenchmarkReproducible task/run/score JSONL evidence; per-model aggregation that feeds the router; commitHash + irVersion provenance binding; evidenceBacked anti-sandbagging flag.
4supreme-routersupremeRouterDeterministic selection: 8 hard gates → weighted scoring → RM0-first cost-class rule → unscoredEvidenceWeight anti-sandbagging downweight → optional verifier-failure-driven effort pacing. Carries CapabilitySignal labels onto decisions.
5supreme-verifiersupremeVerifierDeterministic validator registry (exact-text · regex · JSON · file · command). Evidence > model self-confidence.
flowchart TB
    subgraph SUP["Supreme plugin layer — project-owned, frozen 7"]
        P["supreme-policy"]
        O["supreme-observability"]
        BM["supreme-benchmark"]
        R["supreme-router"]
        V["supreme-verifier"]
        M["supreme-memory-policy"]
        W["supreme-workflow-policy"]
    end
    subgraph CORE["DSH core — pinned upstream · never modified"]
        C["ctx.llm · ctx.sessions · ctx.systemPrompt · ctx.tokenMeter · ctx.credentials · ctx.subagents · ctx.workflowEngine"]
    end
    P --> C
    O --> C
    BM --> C
    R --> C
    V --> C
    M --> C
    W --> C
    P -. consults .-> V
    P -. consults .-> R
    P -. consults .-> W
    O -. optional .-> V
    O -. optional .-> W
    BM -. history .-> R
    V -. evidence .-> W

Four support plugins (supreme-minimal-probe, supreme-boot-probe, supreme-gate-driver, supreme-fake-llm) exist only as test fixtures for the suite — they never ship in the bundle.


🏁 Proof wall — every verdict runnable

Don't trust this README. Run these:

CommandVerdict markerWhat it proves
bun run suiteCOMPLETE101/101 Level-A checks + 5/5 real-loader boots + v1.2/v1.3/v1.3.1 audit gates
bun run v131:verifyV131_COST_FIX_VERIFIED · V131_VERIFIER_FIX_VERIFIED · V131_MEMORY_FIX_VERIFIED · V131_A2A_FIX_VERIFIED · V131_EVIDENCE_BINDING_VERIFIED · V131_OUTCOME_ROUTING_VERIFIED · V131_FAILURE_INJECTION_VERIFIEDReview-hardening end-to-end: cost pre-dispatch deny, symlink-proof roots, JSON-schema strictness, memory isolation, A2A registry, evidence staleness, outcome routing, failure injection (465 probes — see docs/REVIEW-FIXES-v1.3.1.md)
bun run v13:verifyV13_POLICY_E2E_COMPLETE · V13_WORKFLOW_E2E_COMPLETE · V13_ROUTING_E2E_COMPLETEAll 7 v1.3 ASTRA features end-to-end: real engines + real pinned-cordis adapters (85 + 82 + 17 probes)
bun run bundle:verifyBUNDLE_E2E_COMPLETEReal dsh plugin add → reconciler → boot → 13 services → user-patch override wins → clean dispose
bun run composition:verifyCOMPOSITIONS_E2E_COMPLETEAll 4 fragments: service presence and absence, relative dataDir write-through
bun run v12:verifyV12_E2E_COMPLETEEvery v1.2 config key arrives at its service + functional probes (taint deny, effort escalation/recover, path scope, close gate, ledger)
bun run v3:verifyV3_CONFIG_REVIEW_EVIDENCEThe silent-strip trap, live: a wrong config loses 5/6 keys → corrected config enforces 6/6
Level A unit checks      101/101 PASS  (policy 17 · observability 7 · benchmark 11 · router 23
                                       verifier 11 · memory 13 · workflow 19)
v1.3.1 review probes     465/465      (cost 37 · verifier 43 · memory 88 · a2a 63 ·
                                       evidence 82 · outcome-routing 80 · failure-injection 72)
v1.3 E2E probes          184/184      (policy 85 · workflow 82 · routing 17 — real engines,
                                       real pinned-cordis adapters, no upstream build needed)
Real-loader boots        5/5 PASS     (supreme-minimal, core, standard, supreme, lab)
  boot times             supreme-minimal ~51–60 ms · core/standard/supreme/lab ~830–990 ms
Keyless scenario         9/9 gates PASS (real DSH session; router picks free route; PAID rejected)
Security                 sentinel leaks = 0 · paid automatic fallback = DISABLED
v1.2/v1.3 audit gates    config-key hygiene PASS · pinned-ref scan PASS ·
                         six-surface audit PASS (incl. the pinned `workflow/agent-start` seam) ·
                         schema contract PASS (3 schemas)
Upstream integrity       commit unchanged · worktree clean · patches = 0
Performance              router ≈ 0.02–0.04 ms / 1k · observability serialize ≈ 0.001–0.007 ms / 1k
VERDICT                  COMPLETE

Honesty rule: the real-loader path (real/boot.mjs) is the only real-integration evidence. The Level-A harness under src/harness/cordis-mini is a lifecycle fixture — it is never cited as DSH proof.


🔐 Security guarantees

GuaranteeMechanismProof
Secrets never leak through observabilitySecret-sentinel scrub on allowlisted fields, fail-open write pathsuite: sentinelLeaks = 0 every run
Tainted tool arguments can't dispatchUnicode class scan (zero-width / bidi / BOM / tag) + taintPolicy: DENY via upstream tools/pre-executeV12_E2E_COMPLETE functional probe
Denying a command actually stops itDeny-circumvention guard: same-shape retry of a denied call refused (deny_retry) — signature carries names/types, never valuesV13_POLICY_E2E_COMPLETE probes
Hidden payloads can't ride in tool argsEncoding-blob scan (≥256-char base64/hex runs), class names + lengths onlyV13_POLICY_E2E_COMPLETE probes
Self-declared capability labels can't buy permissioncapabilityClassGate ENFORCE/AUDIT; LAB allowlist is floor-bound; labeling only RESTRICTSV13_POLICY_E2E_COMPLETE probes
Inter-agent channels stay on the declared graphallowedContacts directed edges; out-of-graph audited (a2a_contact), DENY blocks pre-factV13_WORKFLOW_E2E_COMPLETE probes
Delegations can't quietly exceed their taskOverreach audit: risk ceiling + approval gate + path scope, value-freeV13_WORKFLOW_E2E_COMPLETE probes
Benchmark scores can't sandbag the routerevidenceBacked flag (verifier-PASS rule) + fixed unscoredEvidenceWeight downweightV13_ROUTING_E2E_COMPLETE probes
Values never echoed in audit eventsTaint/contact/overreach events carry class names, ids and levels onlycode + suite checks
Paid models never fire by accidentUNKNOWN cost ⇒ DENY; allowPaid refused outside LAB; no automatic fallbackkeyless scenario gate 9/9
Destructive delegation is scopedblockedPaths > allowedPaths glob enforcement; DENY_ALL secret policysuite checks 15 (workflow)

🛡️ v1.3.1 review-hardening

Response to an external v1.3.0 review: 5 findings reproduced → fixed → proven (each with a failing test on the original code), plus outcome-based routing, evidence-bound verification, fast path/recovery, and a failure-injection harness. Full per-issue evidence — reproduction commands, root causes, before/ after outputs, remaining limits, and rollback steps (bash + PowerShell) — lives in docs/REVIEW-FIXES-v1.3.1.md.

#SeverityFinding (v1.3.0)Fix (v1.3.1)
AP1Paid/unknown-model LLM requests dispatched with no cost checkPre-dispatch deny at agent/request + llm/stream backstop; zero adapter calls on deny; UNKNOWN denied in production RM0; LAB exception contract kept
BP1Symlink inside allowedRoots escaped the file-hash verifierNative realpath validation of roots AND targets before any read; traversal/sibling-prefix/missing-file rejected; race reduced (not race-proof — documented)
CP1latestSelection shared across sessions/tasks (memory contamination)Selections bound to (session, task); unknown identity → empty; bounded LRU + cleanup on end/cancel/dispose
DP2copy_file {target: b.txt} misclassified as agent-to-agent contactTrusted comms-tool registry gates recipient extraction; post-fact emit stays detect-only
EP2JSON Schema additionalProperties:false silently ignored (false PASS)Deterministic validator: unsupported keywords → ERROR/UNAVAILABLE, never silent downgrade

Run it: bun run v131:verify (7 markers, 465 probes) — then read the doc before trusting this table.

📈 Benchmarks v1.3.1

Policy-enforcement delta across three labels — same runner, same datasets, thresholds frozen before evaluation, dev + held-out inputs disjoint:

LabelSetupResult
Aharness without Supreme30 cost-policy bypasses
BSupreme v1.3.0 (pre-review-fix)90 escapes (cost 30 · memory 15 · symlink 15 · A2A false-deny 15 · schema false-pass 15)
CSupreme v1.3.10 escapes — all kinds · benign pass 75/75 · overhead ≈ 10 ms/set (median 93 vs 82 ms)

All 6 thresholds PASS. Safety regressions (benign denials) block promotion.

  • Method, limits and cleanup record: benchmarks/BENCH-v1.3.1.md
  • Thresholds-then-evaluate: benchmarks/THRESHOLDS-v1.3.1.json · raw runs: benchmarks/runs/
  • Re-run yourself: bun real/bench-v131.mjs --label C --reps 1 → BENCH_C_THRESHOLDS_PASS
  • Astra (GPT-6): NOT_RUN — no verified public evaluation data; never fabricated (research note). This is a policy-enforcement benchmark, not a model-quality ranking.

🧬 v1.3 ASTRA-hardening features

Seven deterministic hardening features from the ASTRA-1 backlog (research/gpt6-astra-2026-09.md §7). No ML, no new deps — every feature is engine-checked in the keyless suite and proven end-to-end by bun run v13:verify (real engines + real pinned-cordis adapters). The shared label contract CapabilitySignal { capabilityClass?, cotVisibility? } is exported by supreme-policy and carried (never enforced) by the router.

supreme-policy — four features (config table)
Config keyDefaultMeaning
cotVisibilityProfiles{}routeId → expected CoT visibility. A route declared none never denies on cot_missing — ENFORCE downgrades to audit-only (empty-CoT models can't be coerced into producing a trace). Resolution: explicit signal > profile > verbose.
riskGatedCoTfalseENFORCE applies only to HIGH-risk tools (deterministic command/network/write name classifier); non-HIGH tools keep AUDIT.
denyCircumventionGuardtrueA same-shape retry of an already-denied call is refused with reason code deny_retry. The signature encodes argument NAMES + TYPES only — values can never enter it. First calls unaffected; resetDenyCircumvention(sessionId) is the escape hatch.
enableEncodingScanfalseAudit/deny ≥256-char base64/hex runs in tool arguments (encoding_blob; argument NAME + run LENGTH only). Extends the v1.2 taint surface: same event, same taintPolicy.
capabilityClassGate'OFF'Gate requests carrying capabilityClass: AUDIT records, ENFORCE denies unsanctioned classes. Unlabeled requests always pass untouched.
sanctionedCapabilityClasses / labCapabilityClassAllowlist[] / []Sanction lists; the LAB allowlist is additive and binds ONLY on the LAB floor. No implicit ROUTINE exemption — a self-declared label can only RESTRICT, never grant.
supreme-workflow-policy — two features (config table)
Config keyDefaultMeaning
agentContactPolicy / allowedContacts'LOG_ONLY' / []A2A contact graph: directed { from, to } edges of agent ids/roles (empty = inert). Out-of-graph spawn/message contacts are audited as a2a_contact; under 'DENY' the pre-fact tools/pre-execute waterfall refuses with a2a_contact_denied. Emit-mode seams are DETECT-only.
maxRiskLevel / approvalRequiredFor'HIGH' / []Overreach audit: delegations above the risk ceiling, listed task classes without an approval flag, or paths outside the v1.2 scope are audited as overreach_suspected (labels, levels, flags, config globs — never content).
supreme-router + supreme-benchmark — anti-sandbagging (config table)
Config keyPluginDefaultMeaning
requireEvidenceForScoresbenchmarkfalseScore claims without verifier-PASS evidence are flagged evidenceBacked: false on the score + run (flag only — scores never rewritten; re-evaluated when verification lands late).
unscoredEvidenceWeightrouter1FIXED multiplicative downweight for unevidenced benchmark claims (e.g. 0.5 halves such scores); ids + factors recorded on the decision + unscored_evidence events (ids only). 1 = off, back-compat.
—router—Carries capabilityClass / cotVisibility labels from candidates onto the selected RouteDecision (carrier, not enforcer).
Composition fragments (v1.3 posture)
Fragmentv1.3 keys
coredenyCircumventionGuard: true pinned (the one default-ON); everything else inherits OFF defaults
standardenableEncodingScan: true + capabilityClassGate: AUDIT — audit-only, cannot block
supremesame audit-only policy posture + requireEvidenceForScores: true + workflow keys pinned at behavior-preserving defaults
labenforcing demo: capabilityClassGate: ENFORCE + labCapabilityClassAllowlist, cotVisibilityProfiles + riskGatedCoT, declared contact graph + maxRiskLevel: MEDIUM, unscoredEvidenceWeight: 0.5

🧬 v1.2 governance features

Deterministic. No ML. No new runtime deps. Every feature binds to a real pinned upstream seam and ships with engine checks + boot-level proof (bun run v12:verify).

supreme-policy — unicode taint denial + CoT presence gate

Upstream freezes tool arguments after logging (wrappers may change only exec.signal), so the enforceable host-side posture is detect → audit → deny through the official tools/pre-execute seam ({ kind: 'deny', reason } — upstream materializes the error result; Supreme never fabricates tool output):

Config keyDefaultMeaning
enableUnicodeSanitizationtruescan tool arguments for zero-width / bidi-isolate / bidi-override / tag codepoints (U+200B–200F, U+2060–206F, U+202A–202E, U+FEFF, U+E0000–E007F)
logTaintAttemptstruerecord taint_detected events — class names only, values are NEVER echoed
taintPolicyLOG_ONLYDENY refuses the call before dispatch
reasoningTracePolicyOFFAUDIT records cot_missing when an assistant message carried no reasoning trace; ENFORCE additionally denies that session's tool calls (ENFORCE refused on the CORE floor)
supreme-router — RM0-first + effort pacing
Config keyDefaultMeaning
costFirsttruescore only the cheapest eligible cost class — FREE_CONFIRMED beats a rate-limited peer with better history; hard-gate evidence for ALL candidates preserved
effortPacing.enabledfalsedeterministic costClass → reasoningEffort mapping over the pinned agent/request seam (pinned DeepSeek levels: off / low / high / max)
effortPacing.escalateOnVerifierFailtrueone-step escalation (low → high) driven only by verifier FAIL evidence via reportVerifierOutcome() — never model self-confidence; PASS recovers
supreme-workflow-policy — surgical path scope + verifier-gated close
Config keyDefaultMeaning
allowedPaths / blockedPaths[] / []zero-dependency glob scope for delegations (** crosses segments, */? stay in-segment); blocked always wins; empty allowlist = unrestricted
requireVerifierPassOnClosefalseHIGH-risk tasks may only close with recorded verifier PASS evidence
supreme-memory-policy — bounded ledger + instinct-style gates
Config keyDefaultMeaning
ledgerEnabledfalseopt-in bounded, append-only JSONL note ledger (ledgerDir, ledgerFileName, ledgerMaxEntries) — credential-bearing notes rejected at admission
minConfidence0.7notes below this confidence never inject (ECC instincts analogue — recorded evidence quality, not self-assessment)
maxInjected6hard cap per selection
relevanceRankingtruedeterministic task-token-overlap ranking before priority (counting, not ANN)
supreme-benchmark — provenance binding + published schemas

Run records accept commitHash (40-hex sha or UNAVAILABLE) and irVersion — malformed values are rejected by validation, so routing evidence stays bound to the code that produced it.

  • schemas/suite-report.schema.json · benchmark-record.schema.json · ledger-note.schema.json — third parties can validate reports/records; a suite check keeps schemas and code from drifting.
  • Suite also runs config-key hygiene (every shipped YAML row validated against the plugin's real zod schema — the silent-strip trap stays closed), pinned-ref scan, and the six-surface security audit.

📦 Install as a dsh bundle

The repository IS the bundle: package.json declares dsh.bundle.patch → cordis.patch.yml, which inserts the seven frozen plugins as profile rows. Any profile can adopt Supreme through the official plugin flow:

# from a local checkout…
dsh plugin --profile <your-profile> add /path/to/dsh-supreme
# …or straight from GitHub
dsh plugin --profile <your-profile> add github:stadeummwt/dsh-supreme

# prove an install end-to-end (real CLI install + boot + layering checks)
bun run bundle:verify
# prove the v1.2 config surface end-to-end
bun run v12:verify
# prove the v1.3 ASTRA-hardening features end-to-end (all three verifiers)
bun run v13:verify

The bundle mounts the seven plugins with safe production defaults (PAID / TRIAL denied, commands/network off, zero router candidates). Extend candidates, project knowledge, and workflow limits from YOUR profile patch layer — the composer applies last write wins per row id, so user config always beats bundle defaults. The four support/fixture plugins are NOT part of the bundle: they never ship into user profiles.

Composition fragments ship under config/compositions/ — the bundle-world analogue of manifest-driven install profiles. Each fragment UPDATE-patches the bundle rows by id (whole-config replacement, disabled: true for rows outside the composition) and carries no name restatement, so it stays install-location-independent. Prove all four end-to-end: bun run composition:verify → COMPOSITIONS_E2E_COMPLETE.

Notes: dsh plugin add requires pnpm on PATH; installing from GitHub works without a prepare build because dist/ is committed. Fragment paths are relative to the dsh process working directory — override any row from your own patch layer.


🏗️ Architecture

┌────────────────────────────────────────────────────────────────────┐
│  Next.js dashboard (project app) — PROJECTION only, owns no state  │
│  GET/POST /api/supreme/*  (dev/LAB only)                           │
└──────────────────────────────┬─────────────────────────────────────┘
                               │ reads suite reports / triggers runs
┌──────────────────────────────▼─────────────────────────────────────┐
│  SUPREME PLUGIN LAYER (dsh-supreme/dist/plugins, project-owned)    │
│  policy · observability · benchmark · router · verifier ·          │
│  memory-policy · workflow-policy  (+ 4 support/fixture plugins)    │
│  Cordis conventions: name/inject/Config(Standard Schema)/apply     │
└──────────────────────────────┬─────────────────────────────────────┘
                               │ inject: official DSH service names
┌──────────────────────────────▼─────────────────────────────────────┐
│  DSH CORE (pinned upstream — never modified)                       │
│  ctx.llm · ctx.sessions · ctx.systemPrompt · ctx.tokenMeter ·      │
│  ctx.credentials · ctx.subagents · ctx.workflowEngine              │
│  Events: session/* · agent/request* · tools/execute ·              │
│          subagent/* · workflow/*                                   │
└────────────────────────────────────────────────────────────────────┘

Dependency direction (acyclic, enforced):

DSH core services    →  Supreme plugins        (injected seams)
supremePolicy        →  verifier, router, workflow-policy
supremeObservability →  router, verifier (optional), workflow-policy
supremeBenchmark     →  router                 (router reads history; benchmark NEVER depends on router)
supremeVerifier      →  workflow-policy        (verification evidence consulted)

Pinned upstream

ItemValue
Repositoryhttps://github.com/deepseek-ai/deepseek-harness
Pinned commitd347e703908d0406b7a7ef80e3a0e594d86b2215 (master, tag dsh-v0.1.3-alpha.1)
DSH version0.1.3-alpha.1
Vendored Cordis4.0.2 (vendor/cordis)
Upstream worktreekept pristine — UPSTREAM_CORE_MODIFIED = NO, patch count 0
ToolchainNode v24 (v24.19.0), pnpm 11.7.0, Bun 1.3.14 (bundler)

The pinned upstream checkout is read-only for this project. It is resolved at runtime: DSH_UPSTREAM_ROOT env override → sibling ../deepseek-harness → in-project node_modules/.upstream/deepseek-harness. Prefer the sibling location: some upstream builds (pnpm + declaration emit) reject checkouts nested under a node_modules directory. All Supreme code lives in project-owned paths.

Compositions (profiles)

ProfileBundleMounted Supreme plugins
supreme-minimalnone (bare Loader)minimal probe only
core@deepseek-ai/dsh-baseminimal probe, boot probe, supreme-policy (CORE)
standard@deepseek-ai/dsh-base+ observability, memory-policy, verifier
supreme@deepseek-ai/dsh-baseall 7 + fake-llm + gate-driver (SUPREME policy)
lab@deepseek-ai/dsh-baseall 7 + fake-llm + gate-driver, LAB-only overrides (allowPaid: true, allowCommands: true, maxConcurrentAgents: 4)

Full layer map + verified real-API evidence table: docs/architecture/ARCHITECTURE.md.


🚀 Build & verify from source

Prerequisites: Node ≥ 24, pnpm 11.7.0 (upstream build), Bun ≥ 1.3. Commands assume the repo root (dsh-supreme/ as published; inside the companion Next.js workspace the suite auto-detects both layouts).

# 1. Install dependencies
bun install

# 2. Clone the pinned DSH upstream (default lookup: sibling ../deepseek-harness;
#    any location works via DSH_UPSTREAM_ROOT — avoid nesting it under node_modules)
git clone https://github.com/deepseek-ai/deepseek-harness.git ../deepseek-harness
git -C ../deepseek-harness checkout d347e703908d0406b7a7ef80e3a0e594d86b2215

# 3. Build the pinned upstream libraries — official tsconfig graph, memory-batched
#    (one tsc -b over the 217-ref host graph needs ~4 GB headroom; the batched
#    runner keeps each invocation under 2 GB)
npm run build:upstream

# 4. Bundle every Supreme plugin to dist/ (one ESM file per plugin; zod external)
PLUGINS="supreme-policy supreme-observability supreme-benchmark supreme-router \
supreme-verifier supreme-memory-policy supreme-workflow-policy \
supreme-minimal-probe supreme-boot-probe supreme-gate-driver supreme-fake-llm"
for p in $PLUGINS; do
  bun build src/plugins/$p/index.ts \
    --outfile dist/plugins/$p/index.mjs \
    --format esm --target node --external zod
done

Each dist bundle externalizes only zod and Node builtins; @deepseek-ai/cordis appears solely as erased type imports. This exact command was verified to reproduce the committed dist/plugins/supreme-policy/index.mjs byte-for-byte.

Real boot (the only real-integration evidence)

# Boot any composition through the REAL pinned DSH Loader and dispose cleanly.
# --setup installs the profile under $DSH_HOME/profiles/<name>/ from config/.
node real/boot.mjs --profile supreme-minimal --setup
node real/boot.mjs --profile core         --setup
node real/boot.mjs --profile standard     --setup
node real/boot.mjs --profile supreme      --setup
node real/boot.mjs --profile lab          --setup

Each run prints one JSON result (bootMs, disposeMs, services presence map, gate results) and exits non-zero on any failure. Gate markers are appended under data/real/ — see the runbooks for expected markers per profile.

Suite execution

bun run suite            # full suite incl. 5 real boots (needs the built upstream)
bun run suite:json       # machine-readable SuiteReport
bun run suite:keyless    # Level A only — runs without the upstream; verdict stays
                         # PARTIAL (REAL_BOOT_SKIPPED, UPSTREAM_CHECKOUT_UNAVAILABLE)
bun run suite:keyless:ci # keyless with CI-friendly exit code: 0 iff verdict is PARTIAL
                         # with only the documented keyless blockers — any real
                         # failure (UNIT/leaks/hygiene/audit/schema) still fails

The suite exits 0 only when every mandatory gate passes (verdict: COMPLETE). Any failure prints the exact blocking gates.

HTTP API (dashboard projection — dev/LAB only)

The Next.js app exposes a thin, read-mostly projection over the suite. It owns no runtime state; runs live in an in-memory store (latest 20 runs) and suite execution is disabled in production (NODE_ENV=production returns 403 unless SUPREME_ENABLE_SUITE=1).

EndpointMethodBehavior
/api/supreme/statusGETSuite scope (frozen 7 plugins, compositions), upstream commit/cleanliness, DSH/cordis versions, runtime info. Always safe.
/api/supreme/reportGETLast SuiteReport from memory; 404 if no run yet (POST /api/supreme/suite/run first).
/api/supreme/suite/runPOSTExecutes the full suite (including 5 real boots). dev/LAB only — 403 in production without SUPREME_ENABLE_SUITE=1.
/api/supreme/suite/runs/:idGETOne run record (runId, startedAt, durationMs, full report); 404 for unknown ids.

Implementation: src/app/api/supreme/** + src/lib/supreme-suite.ts (project app, outside dsh-supreme/).


📤 Distribution (manual, owner-driven)

Repo policy: no pull requests are opened on third-party repositories on the owner's behalf. Prepared submission artifacts live in distribution/:

  • awesome-dsh-entry.yml — catalog-ready entry (single file, category security, validator-conformant keys only).
  • SUBMISSION-GUIDE.md — how listing on dsh-market actually works (it auto-feeds from the awesome-dsh-plugin catalog), the pre-flight gate checklist, the exact manual submission commands, and the npm-publish note.

The GitHub repo already carries the dsh-plugin topic and a dsh.bundle manifest, so the only remaining step for listing is the manual one-file PR the owner chooses to make.


📁 Directory layout
dsh-supreme/                      (repo root as published)
├── README.md                  ← this file
├── VISION.md                  ← original v1 project vision (frozen architecture contract)
├── CHANGELOG.md
├── AGENTS.md                  ← engineering rules for future agents
├── SOURCE-OF-TRUTH.md         ← upstream integrity record (historical + current)
├── LICENSE                    ← MIT (v1.2)
├── package.json               # suite/boot/build/verify scripts (zod + yaml deps)
├── assets/                    # README hero/footer SVGs (self-contained, dark + light)
├── schemas/                   # published JSON Schemas (suite report, benchmark, ledger)
├── benchmarks/                # v1.3.1 policy-enforcement benchmark (runner, datasets, thresholds, runs)
├── .github/workflows/ci.yml   # keyless + full suite on push/PR (v1.2)
├── config/
│   ├── supreme-minimal.cordis.yml   # bare-Loader probe gate
│   ├── core.cordis.yml              # CORE composition
│   ├── standard.cordis.yml          # STANDARD composition
│   ├── supreme.cordis.yml           # SUPREME composition (all 7)
│   ├── lab.cordis.yml               # LAB composition (LAB-only overrides)
│   ├── examples/                    # corrected config example (provenance noted)
│   └── compositions/                # 4 overlay fragments (core/standard/supreme/lab)
├── distribution/                    # manual submission artifacts (no auto-PRs)
│   ├── awesome-dsh-entry.yml        # catalog entry draft (one file)
│   └── SUBMISSION-GUIDE.md          # owner-driven listing walkthrough
├── real/
│   ├── supreme.mjs            # PLUG-AND-PLAY CLI: doctor · setup · verify (one command)
│   ├── lib/obs-proof.mjs      # shared deterministic observability proof (poll + flush + diagnosis)
│   ├── boot.mjs               # REAL DSH boot harness (Loader + root-fiber dispose)
│   ├── bundle-verify.mjs      # E2E: real CLI install + layering (BUNDLE_E2E_COMPLETE)
│   ├── composition-verify.mjs # E2E: 4 composition fragments (COMPOSITIONS_E2E_COMPLETE)
│   ├── v3-config-verify.mjs   # E2E: silent-strip proof (V3_CONFIG_REVIEW_EVIDENCE)
│   ├── v12-config-verify.mjs  # E2E: v1.2 config surface + probes (V12_E2E_COMPLETE)
│   ├── v13-policy-verify.mjs  # E2E: v1.3 policy features, 85 probes (V13_POLICY_E2E_COMPLETE)
│   ├── v13-workflow-verify.mjs# E2E: v1.3 A2A + overreach, 82 probes (V13_WORKFLOW_E2E_COMPLETE)
│   ├── v13-routing-verify.mjs # E2E: v1.3 labels + anti-sandbagging, 17 probes (V13_ROUTING_E2E_COMPLETE)
│   ├── v131-*.mjs             # E2E: 7 review-fix verifiers — cost-enforce · verifier-hardening ·
│   │                          # memory-isolation · a2a-falsepositive · evidence-binding ·
│   │                          # outcome-routing · failure-injection (465 probes total)
│   ├── bench-v131.mjs         # benchmark runner (--label A|B|C, --reps, --out)
│   └── build-batched.sh       # memory-batched official upstream build
├── dist/plugins/<name>/index.mjs    # bun-built ESM bundles loaded by the real Loader
├── data/
│   ├── observability/observability.jsonl   # runtime metadata log
│   ├── benchmark/benchmark.jsonl           # routing evidence store
│   └── real/*.markers.jsonl                # boot/gate markers (verified evidence)
├── src/
│   ├── plugins/               # engine.ts (pure logic) + index.ts (Cordis adapter) per plugin
│   │   ├── supreme-policy/  supreme-observability/  supreme-benchmark/
│   │   ├── supreme-router/  supreme-verifier/  supreme-memory-policy/
│   │   ├── supreme-workflow-policy/
│   │   └── supreme-minimal-probe/  supreme-boot-probe/  supreme-gate-driver/  supreme-fake-llm/
│   ├── suite/                 # runner.ts + cli.ts + engine-checks.ts + config-hygiene.ts
│   │                          # + surface-audit.ts + schema-contract.ts + harness.ts
│   └── harness/cordis-mini/   # Level-A lifecycle FIXTURE only (never cited as DSH proof)
├── research/                  # ECC dissection + ASTRA-1 + README v2 research (evidence base)
└── docs/
    ├── REVIEW-FIXES-v1.3.1.md # per-issue evidence for the v1.3.1 review hardening
    ├── architecture/ARCHITECTURE.md
    ├── decisions/ADR-0000 … ADR-0007
    └── runbooks/              # install, build, test, boot-*, upgrade-pinned-dsh, rollback

📖 Documentation map

DocContents
AGENTS.mdFrozen scope, upstream rules, Cordis conventions, ownership table, verification requirements
docs/architecture/ARCHITECTURE.mdLayer map, verified real-API evidence table, event seams, composition layering
docs/decisions/ADR-0000 (fixture history) + ADR-0001…0007 (one per major decision)
docs/runbooks/install, build, test, boot-core/standard/supreme/lab, upgrade-pinned-dsh, rollback
docs/REVIEW-FIXES-v1.3.1.mdThe 5 review findings: reproduction, root cause, fix, before/after, limits, rollback
benchmarks/BENCH-v1.3.1.mdBenchmark method, thresholds, results, cleanup record
research/ECC dissection (253,948★) + ASTRA-1 dissection + v3 plan review + README v2 research
SOURCE-OF-TRUTH.mdUpstream integrity record (historical + current)
Per-plugin READMEssrc/plugins/<name>/README.md — purpose, config tables, contracts, security boundaries

❓ FAQ

Does Supreme modify DeepSeek Harness?

No. The pinned upstream worktree stays pristine — UPSTREAM_CORE_MODIFIED = NO, patch count 0, re-verified on every suite run. Supreme is an ordinary Cordis plugin layer that consumes official services and event seams.

Is any of this AI-powered?

None. Every gate is deterministic code — counting, glob matching, string comparison, zod validation. That's why the router decides in ~0.02–0.03 ms and why results are reproducible on your machine, today.

Why does UNKNOWN cost deny the model?

Because an unclassified route is an unaudited spend path. supreme-policy treats it as a hard DENY; paid/trial classes require an explicit LAB-only override. RM0-first routing then prefers FREE_CONFIRMED candidates deterministically.

Can I use just the policy plugin?

Yes — that's the core fragment. Or standard for the daily-driver four. Fragments are one-line overlays on your own profile.

What if my config has a typo or an unknown key?

The v1.2 suite runs a config-key hygiene scan: every shipped YAML row is validated against the plugin's real zod schema, so the "boot passes but your governance keys were silently stripped" trap (proven live in V3_CONFIG_REVIEW_EVIDENCE) stays closed.

Does it work offline / air-gapped?

The six-surface audit, taint scanning, ledger and all suite checks are fully offline and deterministic. Real boots need the pinned upstream checked out locally — no network calls at runtime.

Why isn't Supreme listed in the dsh-market yet?

Listing requires a one-file PR to the catalog, and this repo's policy is that such PRs are made by the owner, manually (see distribution/SUBMISSION-GUIDE.md). Everything else is already prepared.


📜 Honest limitations

  • The real-loader path via real/boot.mjs is the only real-integration evidence; the Level-A lifecycle harness (src/harness/cordis-mini) is a fixture and is never cited as DSH proof.
  • Keyless suite verdict is honestly PARTIAL (REAL_BOOT_SKIPPED) without a built pinned upstream — it does not fake completeness.
  • The benchmark measures policy enforcement (bypass/escape counts), not model quality or intelligence; the Astra label is NOT_RUN by design until verified public data exists.
  • Deferred items (documented, not forgotten): HNSW-style memory indexing and Archify-style schema migration stay out of scope for the frozen seven.
  • Router candidates ship empty (zero-by-default): you add models from your own patch layer. Supreme governs choices; it does not preselect providers.

查看图片

Built proof-first. Bukti sebenar > klaim.

If Supreme hardened your harness, consider starring the repo — it helps other DSH users find governance tooling.

⬆ back to top

6supreme-memory-policysupremeMemoryPolicyMemory selection policy: confidence floor, injection cap, relevance ranking, bounded append-only note ledger (credential-bearing notes rejected at admission).
7supreme-workflow-policysupremeWorkflowPolicyWhen/how ctx.subagents / ctx.workflowEngine may run: limits, degradation ladder, glob path scoping (blocked beats allowed), verifier-gated close for HIGH-risk tasks, A2A contact graph + overreach audit.
HIGH-risk work can't skip verificationrequireVerifierPassOnClose evidence gateV12_E2E_COMPLETE probe
Supply chain stays pinnedExternal refs scanned; upstream commit + irVersion bound into run recordspinned-ref scan PASS
Your own audit, offlineSix-surface audit: prompts · hooks · MCP · permissions · secrets · agent filessuite check PASS
CHANGELOG.md
Version history with evidence markers per release