DeepSeek Harness Plugin Hub

发布与管理完整 Harness Profiles,发现适合你的插件。

探索

插件目录环境预设文档中心动态

社区

发布插件联系我们报告问题

相关链接

Plugin Hub GitHubDeepSeek Harness 官方项目系统状态隐私说明
© 2026 DeepSeek Harness Plugin HubPowered byPaxTech

独立、非官方社区项目,与 DeepSeek 官方无隶属、授权或背书关系。

DeepSeek Harness Plugin Hub
ProfilesPlugins分类动态文档登录管理 Profiles
ProfilesPlugins分类动态文档登录
← Plugins
T

dsh-tool-bandit-search

Tool Bandit Search

一个用于 DeepSeek Harness 的搜索工具,使用上下文老虎机(Thompson sampling)学习在快速搜索和深入搜索策略之间进行选择。

插件会安装到这里;不确定时保持 web。

npx -y @deepseek-ai/dsh plugin --profile web add github:siruignaw-sys/dsh-tool-bandit-search#089dd0d3e5b2b1686e66363905b8a393caa1d387
README兼容性版本

兼容性与来源证明

Tool Bandit Search 以 dsh-tool-bandit-search 发布,当前版本为 0.1.0。Plugin Hub 会校验它的 manifest,并保存精确安装来源,便于复现安装结果。

DSH 兼容范围
*
运行环境
any
发布来源
github
Registry 更新时间
2026/8/23

版本

0.1.0stable
2026/8/23

相关插件

正在加载相关插件…

最新版
0.1.0
DSH
*
HMR
重启进程
Tree shaking
未声明可安全裁剪
解包体积
未提供
文件数
未提供
Surface
any
许可证
MIT
发布源
github
GitHub
★ 2
周下载
0
最近提交
2026/8/22
查看源码 ↗
README Badge

点击下方 Badge 复制 Markdown,粘贴到 README 即可。

这是你的 Plugin?认领权益 · 优先安全扫描

验证 package.json 声明的 GitHub 仓库,即可管理这个公开页面。认领后,Hub 会优先安排当前版本的安全扫描,并在通过后公开展示结果。

认领这个 Plugin →
报告问题
Tool Bandit Search — DeepSeek Harness 插件(DSH Plugin)

相关插件

继续浏览 search-research 分类下经过校验的插件。

Browser Skill Dsh Plugin@wxg-prc-cpg/browser-skill-dsh-plugin向模型提供 BrowserSkill 浏览器自动化(browser_* 工具)的 DeepSeek Harness 工具插件Weknora@wxg-prc-cpg/dsh-weknora适用于 DeepSeek Harness (dsh) 的 WeKnora 知识检索工具:通过自有知识库进行语义搜索、文档阅读以及 RAG/代理回答。Free Searchdsh-free-searchDeepSeek Harness 的免费网页搜索:13 个引擎(Bing/DuckDuckGo/AnySearch/SearXNG/Exa/Tavily/Keenable/Firecrawl 无需密钥;Parallel/Perplexity/SerpBase/DeepSeek 需要密钥),支持时间筛选、平台搜索和 web_fetch,并提供网页设置界面。Find Plugindsh-find-plugin在代理中查找 DeepSeek Harness 插件——实时搜索 GitHub 上的 dsh-plugin 主题,并按星标数排序。

README

dsh-tool-bandit-search

A DeepSeek Harness plugin that replaces the standard web_search tool with a search tool that learns which search strategy to use through a contextual multi-armed bandit, instead of relying on a single hardcoded approach.

Why

Every web search has a tradeoff: a fast, narrow query gets you an answer quickly, but a broader, multi-angle query gets you better coverage at the cost of latency. Hardcoding one strategy means always overpaying for simple questions or always underdelivering on complex ones. This plugin lets the tool discover, from real usage, which strategy tends to pay off — and keeps adapting as conditions change.

How it works

The search tool has two internal strategies ("arms"):

  • quick — a single search query, capped at 5 results. Fast, good for simple factual lookups.
  • thorough — three query variants (the original plus two reframed angles) run in parallel and merged/deduplicated, capped at 10 results. Slower, better for open-ended or multi-perspective questions.

On every call, the plugin uses Thompson sampling to pick an arm: each arm has a Beta(α, β) distribution representing its estimated reward, the plugin samples from both distributions, and whichever sample is higher gets used. This naturally balances exploration (trying the less-proven arm occasionally) against exploitation (favoring the arm that's performed better so far).

After the call, a continuous reward in [0, 1] is computed from two components, weighted equally:

  • Quality — how many results came back, relative to that arm's own maximum (so a 5-of-5 "quick" result is scored the same as a 10-of-10 "thorough" result — neither arm is structurally favored by its own result cap).
  • Speed — how fast the call completed, calibrated against realistic search latency.

That reward updates the chosen arm's Beta distribution (α += reward, β += 1 − reward), so the bandit's beliefs shift a little after every single call — no separate training phase, no manual tuning.

The model never sees the two arms directly. It just calls search(query); the plugin decides internally which strategy to run.

Example output

[bandit-search] arm=quick reward=1.000 durationMs=4393 resultCount=5 stats={"quick":{"alpha":2,"beta":1},"thorough":{"alpha":1,"beta":1}} [bandit-search] arm=thorough reward=0.854 durationMs=8481 resultCount=10 stats={"quick":{"alpha":2,"beta":1},"thorough":{"alpha":1.85,"beta":1.15}} [bandit-search] arm=quick reward=0.000 durationMs=5777 resultCount=0 stats={"quick":{"alpha":2,"beta":2},"thorough":{"alpha":1.85,"beta":1.15}}

Each log line shows which arm was picked, the reward it earned, and the running Beta parameters for both arms — you can watch the bandit's confidence shift in real time as it accumulates evidence.

Install

dsh plugin --profile web add github:siruignaw-sys/dsh-tool-bandit-search

For local development against a cloned/edited copy instead:

dsh plugin --profile web add link:/absolute/path/to/dsh-tool-bandit-search

Either way, restart the Web UI (a fresh pnpm dsh web / dsh web, not just a new chat) after installing — bundle installs only take effect on the next boot, and the plugin's system-prompt instruction steering the model toward search over the built-in web_search tool only applies to sessions started after that.

Requirements

Runs on top of dsh's native ctx.web search service — no separate API key needed beyond whatever search provider your dsh profile already has configured (e.g. dsh-web-search-deepseek).

Known limitations

  • Bandit state is in-memory and resets on every restart. Persisting it via ctx.storage (which dsh already exposes) is a natural next step.
  • Reward is a heuristic, not a measure of actual answer quality — it captures result count and latency, not whether the results were relevant or correct. A stronger version might score reward against whether the model's final answer actually used the returned sources.
  • The model can still issue multiple search calls per turn even when thorough is already broadening internally — the plugin optimizes strategy per call, not the model's own multi-call behavior.
  • Built and tested against dsh's developer preview; the plugin/tool APIs may change before a stable release.

License

MIT