DeepSeek Harness Plugin Hub

发布与管理完整 Harness Profiles,发现适合你的插件。

探索

插件目录环境预设文档中心动态

社区

发布插件联系我们报告问题

相关链接

Plugin Hub GitHubDeepSeek Harness 官方项目系统状态隐私说明
© 2026 DeepSeek Harness Plugin HubPowered byPaxTech

独立、非官方社区项目,与 DeepSeek 官方无隶属、授权或背书关系。

Benchmark — DeepSeek Harness 插件(DSH Plugin)
DeepSeek Harness Plugin Hub
ProfilesPlugins分类动态文档登录管理 Profiles
ProfilesPlugins分类动态文档登录
← Plugins
B

dsh-benchmark

Benchmark

DSH 工具和插件的可复现确定性基准测试证据

插件会安装到这里;不确定时保持 web。

npx -y @deepseek-ai/dsh plugin --profile web add github:dongsheng123132/dsh-benchmark#3f48afaa5ad7bd4b1214e048bf1eae18f98b0cf0
README兼容性版本

兼容性与来源证明

Benchmark 以 dsh-benchmark 发布,当前版本为 0.2.1。Plugin Hub 会校验它的 manifest,并保存精确安装来源,便于复现安装结果。

DSH 兼容范围
*
运行环境
any
发布来源
github
Registry 更新时间
2026/9/7

版本

0.2.1stable
2026/9/7

相关插件

正在加载相关插件…

最新版
0.2.1
DSH
*
HMR
重启进程
Tree shaking
未声明可安全裁剪
解包体积
未提供
文件数
未提供
Surface
any
许可证
MIT
发布源
github
GitHub
★ 3
周下载
0
最近提交
2026/9/7
查看源码 ↗
README Badge

点击下方 Badge 复制 Markdown,粘贴到 README 即可。

这是你的 Plugin?认领权益 · 优先安全扫描

验证 package.json 声明的 GitHub 仓库,即可管理这个公开页面。认领后,Hub 会优先安排当前版本的安全扫描,并在通过后公开展示结果。

认领这个 Plugin →
报告问题

相关插件

继续浏览 developer-tools 分类下经过校验的插件。

Web App@deepseek-ai/dsh-web-appdsh 浏览器界面捆绑包:位于 dsh-base 之上的 Web 补丁层,加上运行时粘合插件(提供前端 dist、Web 界面提示符、bash 运行时变量和 URL 行)Sdk Minimal@deepseek-ai/dsh-sdk-minimal独立的最小 SDK 配置包:JSON-RPC、一个 DeepSeek 适配器、持久化 Shell 和 JSONL 会话Sdk App@deepseek-ai/dsh-sdk-appdsh SDK 配置包:基于 dsh-base 提供 stdio JSON-RPC 服务和进程生命周期管理Subagent Codex@deepseek-ai/dsh-subagent-codex基于官方 app-server 协议的一次性 Codex 子代理提供程序

README

dsh-benchmark

Reproducible, deterministic benchmark evidence for DeepSeek Harness tools and plugins.

This project deliberately does not duplicate dsh-batch-regression, which runs one shell command repeatedly for median/distribution statistics. dsh-benchmark defines an evidence protocol around fixed cases: explicit target and suite revisions, file-derived target fingerprints, bounded argv-only subprocesses, raw measurements, versioned deterministic scoring, content-addressed reports, and baseline regression comparison.

The first release evaluates commands and JSONL runners, not subjective LLM quality.

Version 0.2.0 is a formal Codex plugin and standalone proof-only MCP server, and uses the namespace export shape required by the stock DSH Web Loader. A real Cordis boot regression test guards that loader contract.

Adjacent benchmark skills often grade Skill or LLM quality. This project stays at the deterministic execution-evidence layer: fixed target revisions and cases, raw bounded measurements without raw business output, versioned scoring, content-addressed reports, and baseline regression decisions.

Evidence model

An explicit manifest freezes:

  • suite name and case revision;
  • target name, claimed revision, and files used to recompute its fingerprint;
  • executable, constrained working directory, warmup/repeat counts, timeout, output cap, and concurrency cap;
  • fixed argv and optional JSONL stdin for every case;
  • expected exit code, stdout/stderr SHA-256, and optional JSONL line count;
  • scorer version, minimum pass rate, output-stability rule, and maximum median-latency regression.

Each run records warmup and measured observations separately: duration in nanoseconds, exit code, signal, timeout/output-limit state, output byte counts and hashes, JSONL validity, and every expectation check. Raw argv, stdin, stdout, stderr, inherited environment, timestamps, and hostnames are excluded from reports.

Safety model

  • shell: false; no command strings or shell interpolation.
  • node maps to the current absolute process.execPath. Other executables must be explicit workspace-relative regular files; PATH lookup is not used.
  • cwd, target files, manifests, reports, and artifact directories cannot escape workspaceRoot through traversal or symlinks.
  • Child processes receive a minimal deterministic environment instead of inherited secrets.
  • Timeout, captured-output bytes, and concurrency are mandatory bounded manifest values.
  • Secret-bearing manifest fields such as tokens, cookies, authorization, credentials, and custom environment secrets are rejected.
  • Reports contain hashes and measurements, not command inputs or output bodies.
  • Artifact writes are restricted to explicit artifactDir, content addressed, exclusive, and verified by read-back SHA-256.

Run only trusted benchmark executables. The isolation above prevents accidental shell expansion and environment leakage; it is not an OS sandbox for malicious code.

Install in DSH

dsh plugin --profile benchmark add github:dongsheng123132/dsh-benchmark

The bundle registers:

  • dsh_benchmark_inspect — inspect protocol metadata and fingerprints without execution.
  • dsh_benchmark_run — run fixed cases and write a content-addressed report.
  • dsh_benchmark_compare — compare current and baseline reports with manifest thresholds.

MCP

.mcp.json declares a standalone stdio MCP server:

  • benchmark_manifest_lint validates an inline manifest and returns only identifiers, bounded policies and hashes of runner/case inputs.
  • benchmark_report_address recomputes the exact report SHA-256 and returns a bounded summary while rejecting raw-output and secret-bearing fields.

MCP accepts bounded inline JSON, never executes a command, and never reads or writes the filesystem. Actual benchmark execution remains available only through the workspace-bounded DSH tool and CLI surfaces.

CLI

dsh-benchmark inspect --root /workspace --manifest benchmark.json

dsh-benchmark run \
  --root /workspace \
  --manifest benchmark.json \
  --artifact-dir benchmark-artifacts

dsh-benchmark compare \
  --root /workspace \
  --manifest benchmark.json \
  --baseline benchmark-artifacts/baseline.json \
  --current benchmark-artifacts/current.json \
  --artifact-dir benchmark-comparisons

Exit code 0 means pass. 2 means a report/comparison was written but its scorer failed. 1 means a manifest or operational error.

Manifest example

examples/benchmark.example.json benchmarks a fixed JSONL runner. Run it from this repository:

node bin/dsh-benchmark.mjs run \
  --root . \
  --manifest examples/benchmark.example.json \
  --artifact-dir artifacts

Arguments and JSONL values can contain ordinary test data, but the report stores only their SHA-256 fingerprints. Do not place real secrets in a benchmark manifest.

Develop

npm test
npm run check
npm run smoke:plugin
npm run smoke:mcp
python C:/Users/ZhuanZ/.codex/skills/.system/plugin-creator/scripts/validate_plugin.py .

Requires Node.js 22+. No runtime dependency or install lifecycle script is used beyond the optional DSH tools SDK peer.

License

MIT