DeepSeek Harness Plugin Hub

发布与管理完整 Harness Profiles,发现适合你的插件。

探索

插件目录环境预设文档中心动态

社区

发布插件联系我们报告问题

相关链接

Plugin Hub GitHubDeepSeek Harness 官方项目系统状态隐私说明
© 2026 DeepSeek Harness Plugin HubPowered byPaxTech

独立、非官方社区项目,与 DeepSeek 官方无隶属、授权或背书关系。

Llm Rate Limit — DeepSeek Harness 插件(DSH Plugin)
DeepSeek Harness Plugin Hub
ProfilesPlugins分类动态文档登录管理 Profiles
ProfilesPlugins分类动态文档登录
← Plugins

dsh-llm-rate-limit

Llm Rate Limit

DeepSeek Harness 的大语言模型 API 速率限制、并发控制、排队和自适应冷却

插件会安装到这里;不确定时保持 web。

npx -y @deepseek-ai/dsh plugin --profile web add dsh-llm-rate-limit@0.1.1
README兼容性版本

兼容性与来源证明

Llm Rate Limit 以 dsh-llm-rate-limit 发布,当前版本为 0.1.1。Plugin Hub 会校验它的 manifest,并保存精确安装来源,便于复现安装结果。

DSH 兼容范围
*
运行环境
any
发布来源
npm
Registry 更新时间
2026/9/20

版本

0.1.1stable
2026/8/20
0.1.0stable
2026/8/20

相关插件

正在加载相关插件…

最新版
0.1.1
DSH
*
HMR
重启进程
Tree shaking
未声明可安全裁剪
解包体积
70.7 kB
文件数
18
Surface
any
许可证
MIT
发布源
npm
GitHub
★ 1
周下载
86
最近提交
2026/8/20
查看源码 ↗项目主页 ↗
README Badge

点击下方 Badge 复制 Markdown,粘贴到 README 即可。

这是你的 Plugin?认领权益 · 优先安全扫描

验证 package.json 声明的 GitHub 仓库,即可管理这个公开页面。认领后,Hub 会优先安排当前版本的安全扫描,并在通过后公开展示结果。

认领这个 Plugin →
报告问题

相关插件

继续浏览 models-usage 分类下经过校验的插件。

Usage@linxin666/dsh-usage用于 dsh Web GUI 的使用统计插件:检测各提供商的余额和编码计划配额,并提供实时令牌使用记录,同时在侧边栏条目中显示当前会话提供商今日的使用量Whale Widgetdsh-whale-widgetDSH Web 界面右下角的 DeepSeek 余额小鲸鱼挂件:余额/今日已用/峰谷定价、自定义泡泡点击序列(文本/余额/今日/峰谷/图片/随机语句与并列加权选择)、逐行样式与字体、悬浮快捷编辑、音效与每轮消耗、自定义角色/动图/音效、吸附与翻转自定义Usage Stats@ychris12138/dsh-usage-statsdsh Web GUI 的令牌使用热力图、提供商余额和订阅配额Codex Connectdsh-codex-connect用于 DeepSeek Harness 的 ChatGPT OAuth 和 Codex 模型。

README

dsh-llm-rate-limit

English | 中文

A DeepSeek Harness (DSH) plugin that prevents avoidable API rate-limit errors by pacing LLM requests before they reach the provider. It provides per-provider RPM limits, optional token budgets, concurrency control, bounded FIFO queuing, and adaptive cooldown for DeepSeek API, Volcengine Ark, and other DSH providers.

Use it when parallel agents, subagents, retries, or background requests are producing HTTP 429 errors, provider throttling, or traffic bursts.

Install from npm

Install the latest release into the Web profile:

dsh plugin --profile web add dsh-llm-rate-limit
dsh web

Pin a version for reproducible environments:

dsh plugin --profile web add dsh-llm-rate-limit@0.1.1

Install separately for Headless:

dsh plugin --profile headless add dsh-llm-rate-limit

GitHub installation is also supported:

dsh plugin --profile web add github:Asong6824/dsh-llm-rate-limit#v0.1.1

The bundled default protects deepseek-official with 30 requests per minute, burst 1, two concurrent requests, and a bounded queue.

Features

  • Provider-scoped requests-per-minute token buckets with configurable burst capacity.
  • Optional estimated-token-per-minute budgets with actual-usage reconciliation.
  • Concurrency limits and bounded FIFO queues with timeout and cancellation.
  • Adaptive cooldown for provider error codes, HTTP statuses, and Retry-After.
  • Explicit auxiliary-request shedding so background traffic does not block primary work.
  • Durable admission wait/start events for DSH session diagnostics.
  • Clean lifecycle disposal without abandoning queued or active requests.
  • Retry-aware admission: every dsh-llm-retry attempt is admitted independently; this plugin never retries requests itself.

Configure DeepSeek and Ark

Override the complete llm-rate-limit config in $DSH_HOME/profiles/<profile>/cordis.patch.yml:

- id: llm-rate-limit
  config:
    providers:
      deepseek-official:
        requests: { perMinute: 30, burst: 1 }
        maxConcurrentRequests: 2
        queue: { maxSize: 100, maxWaitMs: 300000, auxiliary: reject }
        cooldown:
          codes: [RATE_LIMIT, SERVER]
          statuses: [429, 529]
          initialDelayMs: 500
          maxDelayMs: 60000
          maxProviderDelayMs: 3600000
          jitterRatio: 0.1
      volcengine-ark-coding:
        requests: { perMinute: 30, burst: 1 }
        maxConcurrentRequests: 2
        queue: { maxSize: 100, maxWaitMs: 300000, auxiliary: reject }

Provider keys must exactly match GenerateOptions.provider. Optional token limiting adds:

tokens:
  perMinute: 1000000
  burst: 200000
  estimatedOutputTokens: 8192
  imageTokens: 1024

tokens.burst must be large enough for one complete request estimate. Omit tokens when a provider should have RPM and concurrency control without a local token ceiling.

How it works

Before each provider call, the plugin reserves request capacity, estimated token capacity, and a concurrency slot. Requests without capacity wait in FIFO order. Provider throttling responses activate a shared cooldown; successful responses reconcile estimated tokens with actual usage. The state is process-local and resets when DSH restarts.

The plugin deliberately does not provide distributed quotas, automatic retries, or provider failover.

Compatibility and links

  • Requires DeepSeek Harness 0.1.0-rc.8 or newer and Node.js 22.19 or newer.
  • npm package
  • GitHub releases
  • DSH plugins topic
  • Machine-readable summary

Development

pnpm install
pnpm run check

MIT