DeepSeek Harness Plugin Hub

发布与管理完整 Harness Profiles,发现适合你的插件。

探索

插件目录环境预设文档中心动态

社区

发布插件联系我们报告问题

相关链接

Plugin Hub GitHubDeepSeek Harness 官方项目系统状态隐私说明
© 2026 DeepSeek Harness Plugin HubPowered byPaxTech

独立、非官方社区项目,与 DeepSeek 官方无隶属、授权或背书关系。

Plugin Abtest — DeepSeek Harness 插件(DSH Plugin)
DeepSeek Harness Plugin Hub
ProfilesPlugins分类动态文档登录管理 Profiles
ProfilesPlugins分类动态文档登录
← Plugins
P

dsh-plugin-abtest

Plugin Abtest

用于 DSH 插件的配对实验和推广门槛。

插件会安装到这里;不确定时保持 web。

npx -y @deepseek-ai/dsh plugin --profile web add github:Morriaty-The-Murderer/dsh-plugin-abtest#17b4a91b7fe443ce591cd8b799f9b2b403b528c1
README兼容性版本
DSH Plugin A/B Test — paired experiments, auditable evidence, safer promotion

兼容性与来源证明

Plugin Abtest 以 dsh-plugin-abtest 发布,当前版本为 0.1.0。Plugin Hub 会校验它的 manifest,并保存精确安装来源,便于复现安装结果。

DSH 兼容范围
*
运行环境
any
发布来源
github
Registry 更新时间
2026/8/21

版本

0.1.0stable
2026/8/21

相关插件

正在加载相关插件…

最新版
0.1.0
DSH
*
HMR
重启进程
Tree shaking
未声明可安全裁剪
解包体积
未提供
文件数
未提供
Surface
any
许可证
MIT
发布源
github
GitHub
★ 0
周下载
0
最近提交
2026/8/26
查看源码 ↗
README Badge

点击下方 Badge 复制 Markdown,粘贴到 README 即可。

这是你的 Plugin?认领权益 · 优先安全扫描

验证 package.json 声明的 GitHub 仓库,即可管理这个公开页面。认领后,Hub 会优先安排当前版本的安全扫描,并在通过后公开展示结果。

认领这个 Plugin →
报告问题

相关插件

继续浏览 developer-tools 分类下经过校验的插件。

Web App@deepseek-ai/dsh-web-appdsh 浏览器界面捆绑包:位于 dsh-base 之上的 Web 补丁层,加上运行时粘合插件(提供前端 dist、Web 界面提示符、bash 运行时变量和 URL 行)Sdk Minimal@deepseek-ai/dsh-sdk-minimal独立的最小 SDK 配置包:JSON-RPC、一个 DeepSeek 适配器、持久化 Shell 和 JSONL 会话Sdk App@deepseek-ai/dsh-sdk-appdsh SDK 配置包:基于 dsh-base 提供 stdio JSON-RPC 服务和进程生命周期管理Subagent Codex@deepseek-ai/dsh-subagent-codex基于官方 app-server 协议的一次性 Codex 子代理提供程序

README

DSH Plugin A/B Test — paired experiments, auditable evidence, safer promotion

English · 简体中文

DSH Plugin A/B Test

Test a DSH plugin change on the same tasks before you ship it.

DSH Plugin A/B Test runs your current plugin (Control) and proposed change (Candidate) in isolated DSH environments, pairs their results case by case, and produces evidence you can review before release.

It helps answer three practical questions:

  • Did the Candidate improve task success?
  • Did that improvement come with a meaningful regression in tokens, latency, or tool errors?
  • Can someone else reproduce the result from the same inputs?

Every experiment ends with one of four deterministic outcomes: PROMOTE, REVIEW, REJECT, or INCONCLUSIVE. Even PROMOTE is an offline recommendation only—this project never changes your real DSH profile or publishes a plugin for you.

See the decision first

This is a real result from the repository's offline example, with unrelated fields omitted:

{
  "outcome": "PROMOTE",
  "validPairCount": 2,
  "invalidPairCount": 0,
  "triggeredRules": ["primary.superiority"]
}

Alongside the decision, you get raw session evidence, assertion results, pair-level deltas, and reports in JSON, Markdown, and HTML. The model does not choose the outcome; deterministic rules from the experiment manifest do.

Quick start

The current MVP runs from source and requires Node.js ^22.19.0 || >=24.0.0 and pnpm 11.19.0. The starter experiment uses an offline scripted provider, so no model API key is required.

pnpm install --frozen-lockfile

node --import tsx src/cli/bin.ts init --output ./my-experiment --json
node --import tsx src/cli/bin.ts freeze --manifest ./my-experiment/experiment.yml --output ./evidence --json
node --import tsx src/cli/bin.ts run --manifest ./my-experiment/experiment.yml --output ./evidence --json
node --import tsx src/cli/bin.ts decision --manifest ./my-experiment/experiment.yml --output ./evidence --json
node --import tsx src/cli/bin.ts report --manifest ./my-experiment/experiment.yml --output ./evidence --json

Open ./evidence/<experiment-id>/report.html to view the static report. To test your own plugin, edit experiment.yml, evals/cases.yml, and the two variant configurations created by init.

Understand the outcome

OutcomeWhat it meansTypical next step
PROMOTEEvidence is sufficient, quality meets the target, and guardrails passContinue through your human release process
REVIEWResults improved, but cost, latency, error rate, or variance needs judgmentReview the pair-level evidence
REJECTA hard gate failed, a critical case regressed, or the gain was too smallFix the Candidate and rerun
INCONCLUSIVEThere were too few valid pairs, exposure was not proven, or environments were not comparableComplete the evidence instead of treating it as a failure

Task success is the default primary metric. You can also guard token usage, P95 latency, and tool error rate. Thresholds, minimum valid pairs, repetitions, and concurrency all live in the manifest. See the manifest reference and decision rules for details.

Why the evidence is trustworthy

  • Paired tasks: Control and Candidate receive the same case, workspace fixture, and model parameters.
  • Balanced order: Pair order alternates to reduce fixed first-run bias.
  • Isolated environments: Each arm gets its own DSH_HOME, profile, workspace, session root, and frozen plugin artifact.
  • Traceable artifacts: Sources may be a local directory, tarball, exact npm version, or pinned GitHub commit; hashes are checked again before execution.
  • Proven exposure: Session events, tool calls, or plugin receipts show whether the target plugin actually participated.
  • Blind comparison: An optional comparator sees anonymous A/B outputs; identity is revealed only after comparison.
  • Honest infrastructure failures: Provider outages, corrupted sessions, and environment mismatches become invalid evidence or INCONCLUSIVE, not fake Candidate regressions.

Evidence is stored under one experiment directory:

<output>/<experiment-id>/
├── manifest.lock.json
├── control-artifact.json
├── candidate-artifact.json
├── pairs/<case-id>-<repetition>/
│   ├── pair.json
│   ├── measurement.json
│   ├── control/
│   └── candidate/
├── comparison.json
├── decision.json
├── report.md
└── report.html

Completed pairs are reused by later run commands, and partial state is never silently overwritten. If inputs change or evidence is damaged, use a new output root.

Connect to real DSH

Non-mock providers invoke the exact pinned version @deepseek-ai/dsh@0.1.0-rc.7. Add model credentials and other environment variables by name to extensions.environment_allowlist. Fingerprints and artifact metadata store only the allowlist hash, never the original values.

Plugin and test-command stdout, stderr, and session logs are retained as raw evidence, so integrations must still avoid printing secrets.

The CLI is a trusted local automation boundary and may run commands explicitly declared in the manifest. Optional DSH/Cordis tool entry points are model-facing, so they reject experiments containing command_test rather than becoming arbitrary command-execution tools.

The complete CLI includes init, validate, freeze, run, status, compare, decision, and report. Stable exit codes and the live-model smoke procedure are documented in the verification record.

Development

pnpm identity:check
pnpm lint
pnpm typecheck
pnpm test
pnpm test:integration
pnpm test:rename
pnpm build
pnpm pack --dry-run

This project is currently an MVP and is not published as an npm package. See the architecture and upstream audit for implementation and security boundaries.

Public names, the npm package name, and the CLI name are managed from project.identity.json. A real rename test prevents stale public identity from surviving a rename; see the identity guide. Stable protocol identifiers do not change with the project brand.