dsh-health-scheduler
English | 中文
Health monitoring, restart-pressure scoring, maintenance scheduling and action
decisions for DeepSeek Harness — it never restarts anything itself.
dsh-health-scheduler watches a DS-Hns machine, reduces everything it can measure to a
single restart_pressure number between 0 and 100, decides what should happen next, and
then asks somebody else to do it. Levels 1 and 2 of its action ladder are handed to a
worker-control adapter; levels 3 and 4 are requests posted to the separate
dsh-restart bundle, which owns restart
execution.
The pipeline is one direction only:
providers -> normalization -> rolling windows -> trend -> pressure -> policy
-> maintenance scheduler -> action adapters (worker control | restart request)
What it is / What it is not
| Concern | Owner | This plugin's role |
|---|
| Sensing device and runtime health | Health Scheduler | Owns it. Providers sample, normalize and retain history. |
| Judging how bad things are | Health Scheduler | Owns it. The pressure model and the policy engine live here. |
| Scheduling maintenance | Health Scheduler | Owns it. Window, target, deferral budget, safe point. |
Reducing load (THROTTLE, PAUSE_NEW_WORK) | Health Scheduler | Requests it through the worker-control adapter. |
| Actually restarting anything | dsh-restart | Does not own it. Emits a request and reads the answer. |
| Restart locks, rate limits, checkpoints, graceful shutdown | dsh-restart | Not here. |
| Relaunching after a crash, crash-loop breaker | Supervisor | Not here. |
| Task state, checkpoints, resume | DS-Hns Core | Not here. The safe-point query is a question, not storage. |
| A second Mega Core | nobody | Explicitly out of scope. |
Three promises follow from that table, and they are load-bearing:
- No restart execution. There is no
taskkill, no reboot, no process kill, no
child_process call that could reach the operating system's restart path. The only
execFile in the codebase runs a configured telemetry probe (src/providers/sources.ts).
- Missing telemetry is
unknown, never healthy. A dimension with no data scores
null, and its nominal weight is redistributed over the dimensions that do have data.
- A single point sample never drives a high-risk action. Every scored metric passes a
sustain gate, every metric has a rolling window behind it, and the two highest rungs of
the ladder are additionally debounced, gated on the maintenance window and gated on a
safe point.
Install
dsh plugin --profile <name> <pnpm args> forwards its remaining arguments to pnpm
inside the profile directory and then reconciles the profile's bundle list, so an
installed package that declares dsh.bundle joins the layer stack automatically. This
package declares dsh.bundle.patch: ./cordis.patch.yml.
From a local checkout of this repository:
# Windows PowerShell, from the checkout directory
dsh plugin --profile web add <path-to-checkout>\dsh-health-scheduler
# any platform, from the checkout directory (a bare "." is anchored to your cwd)
cd dsh-health-scheduler
dsh plugin --profile web add .
From a package name or a tarball:
dsh plugin --profile web add dsh-health-scheduler
dsh plugin --profile web add ./dsh-health-scheduler-0.1.0.tgz
What the bundle contributes
cordis.patch.yml inserts one row, id: health-scheduler, whose config block restates the
whole balanced preset with comments. The layer is applied after every earlier bundle and
before your own profile's cordis.patch.yml, so your overrides win.
A patch replaces the targeted row's whole config; it is not a deep merge. If you override
providerOptions in your profile patch, copy the entire nested object you want, not just the
one leaf you are changing. The plugin's own configuration resolution is a deep merge over the
selected preset, so a document you pass to the plugin directly behaves the way you would expect
— it is the profile patch layer that replaces.
plugin/manifest.json describes the same contract in machine-readable form: id, kind, entry
module, the install command and patch path, required and optional services, the settings
namespace, the three tool names, the six event names, the action levels this plugin applies
versus the ones it only requests, and what degrades when dsh-restart or worker control is
absent.
Verify before booting
--dump-config prints the composed profile tree without booting it. Use it to confirm
that the plugin row is present and that its configuration is what you wrote:
dsh --profile web --dump-config
--dump-default-config prints the same tree without your user layer and without any
--patch overlays, which makes it the fastest way to see what this plugin contributes.
Boot
dsh --profile web
dsh web is a hardcoded alias for --profile web. A successful load logs one line:
health-scheduler: monitoring started (7 providers, interval 15000 ms, preset balanced)
The git-install caveat
A git-hosted plugin builds on install through its prepare script, and pnpm blocks that
build until you allow it. When dsh plugin ... add git+https://… fails, the CLI prints the
exact key pnpm wants; add it under allowBuilds in
<profile directory>/pnpm-workspace.yaml and re-run the same command.
This plugin ships "prepack": "npm run build" and lib/ is gitignored, so:
- Installing from npm or a tarball needs no build permission. The published tarball
contains the built
lib/ produced by the prepack hook.
- Installing from a git URL always needs the
allowBuilds entry, because the checkout has
no lib/ and prepack must run tsc to create it.
- Installing from a local path behaves like a tarball when you have already run
npm run build, and needs that build otherwise.
Quick start
The plugin works with no configuration at all — the balanced preset is the default and
every leaf in it is a default, not a law. A minimal useful document is one line:
// profile package.json -> dsh.profile, or the plugin's settings namespace
{ "preset": "balanced" }
A more realistic first configuration turns on the scheduled maintenance window and points
the hardware provider at a thermal helper:
# $DSH_HOME/profiles/web/cordis.patch.yml
- id: health-scheduler
name: dsh-health-scheduler
config:
preset: balanced
maintenance:
enabled: true
targetTime: '04:00'
windowStart: '03:30'
windowEnd: '05:00'
maxDeferMs: 3600000
urgentOverridePressure: 92
safePointRequired: true
providerOptions:
hardware:
helperCommand: ['powershell', '-NoProfile', '-File', 'C:\\dsh\\gpu-temp.ps1']
helperTimeoutMs: 5000
statsFile:
paths: ['C:\\dsh\\telemetry\\metrics.json']
staleAfterMs: 120000
After a boot the plugin logs
health-scheduler: monitoring started (7 providers, interval 15000 ms, preset balanced)
and on demand the operator (or the model) reads the report, which begins like this:
Restart Pressure: 21 / 100
State: THROTTLED
Primary Cause: gpu_usage=95.00 pp scores 89/100, held 1260s
Telemetry coverage: 60%
Unknown dimensions: runtime, worker, computer_use_ui (not scored as healthy)
Maintenance: next window 03:30-05:00
Safe point: no safe-point source registered; readiness unknown
Capabilities: restart=unavailable, worker-control=unavailable
Every number there is traceable to a metric. There is no sentence in the output that is not
backed by a measurement — "the model thought it should restart" is not a possible reason
code.
How it works
1. Providers sample; nothing else touches the outside world
All health data enters through a HealthProvider. The pressure model and the policy engine
are forbidden from reaching for NVML, LibreHardwareMonitor, HWiNFO, a Windows API, worker
internals or Electron internals. A provider that cannot measure a metric omits the key;
absence is the only way to say "unknown", and it is deliberately not spelled 0.
2. Normalization rejects, clamps and sorts
normalizeSample turns a raw sample into the canonical vocabulary. A reading that is not a
finite number, or that is outside the metric's hard physical bound, is dropped and reported
as a violation. A ratio sent as a percentage (1.0 < v <= 1.5) is clamped to 1 with a
violation attached rather than silently accepted. The surviving bag is sorted by metric name
so snapshots are stable.
3. Rolling windows, not points
Every metric keeps raw samples for windows.rawMs (30 minutes, raised to the longest
statistics window if that is longer), one mean/max/min bucket per windows.aggregateBucketMs
(5 minutes) retained for windows.aggregateRetentionMs (24 hours), and per-day summaries
retained for windows.dailyRetentionMs (14 days). Statistics are computed over 5 min / 30 min
/ 2 h / 6 h windows. Memory is bounded by construction: the store does not grow with uptime.
4. Trend analysis with a polarity
A leak is not a value, it is a slope. TrendAnalyzer reports a slope only when there are at
least trend.minSamples samples spread over at least trend.minSpanMs, and only trusts it
when R² ≥ trend.minRSquared. The metric registry owns polarity, so a rising GPU
temperature is worsening and a falling recovery_rate is too.
5. Six dimensions, one number
Each metric ramps linearly from its warn endpoint (0 points) to its critical endpoint
(100 points), with the registry's polarity deciding which end is which. A dimension's score
blends its weighted mean with its worst member at WORST_WEIGHT = 0.5, so one critical
metric is not averaged away by five calm ones. The six dimension scores are then combined with
the configured weights.
6. Missing telemetry is unknown, and coverage says so
A dimension with no known metric scores null and is listed in unknownDimensions. Its
nominal weight is renormalized over the dimensions that do have data, and the renormalized
share actually backed by telemetry is published as coverage. A pressure reading at 40 %
coverage can never be mistaken for a full-confidence one, and every decision record carries
coverage_NNpct in its reason list.
7. Anti-flapping is not optional
Six independent mechanisms keep the plugin from becoming a source of churn:
| Mechanism | Default | What it stops |
|---|
| Sustain gate | per metric, e.g. 60 s for gpu_temp_c | A 30-second spike scoring at all. |
| Hysteresis | e.g. throttle enters at 55, exits at 45 | 55 -> throttle, 54 -> normal, 56 -> throttle. |
| Debounce | 2 consecutive evaluations | A single evaluation escalating to a high-risk level. |
| Dwell | 120 s | Any transition below restart level faster than the dwell time. |
minRepeatActionMs | 10 min | Re-issuing the same action sooner than allowed. |
| Three cooldowns | 5 / 30 / 60 min | Decision storms, including from a failing adapter. |
The cooldown starts when an adapter is actually invoked — including when it refuses or
throws. tests/scenarios.test.js drives a permanently critical machine for 30 minutes at a
15-second tick with a throwing restart adapter and asserts that at most three attempts reach
the adapter instead of one per tick.
8. Policy, then adapters
The policy engine emits at most one action per tick and never touches a process. Levels 1
and 2 go to WorkerControlAdapter; levels 3 and 4 become a RestartRequest for
RestartAdapter. Every applied action produces exactly one audit record with its pressure,
coverage, named drivers, reason codes and the adapter's answer.
Model-facing tools
Three tools are registered when the profile has a tool runtime. All three are read-only: none
of them exercises a capability.
health_status
Reports the current health of the DS-Hns runtime. Parameters: section
(full | pressure | maintenance | providers, optional, defaults to full).
health_status({ "section": "pressure" })
Restart Pressure: 21 / 100
State: THROTTLED
Coverage: 60%
Primary cause: gpu_usage=95.00 pp scores 89/100, held 1260s
[0.3] gpu_usage_critical: gpu_usage=95.00 pp scores 89/100, held 1260s
[0.2] cpu_usage_critical: cpu_usage=90.00 pp scores 71/100, held 1260s
[0.2] gpu_temp_c_critical: gpu_temp_c=88.00 °C scores 71/100, held 1260s
The full section is the whole report: pressure, state, primary cause, coverage, unknown
dimensions, maintenance summary, safe-point summary, capability states, a per-dimension
table, drivers, worsening trends, provider status, memory lines, the last five decisions and
the tick's warnings. The maintenance and providers sections are single-topic views of the
same snapshot.
health_history
Returns rolling-window statistics for one canonical metric.
| Parameter | Type | Required | Meaning |
|---|
metric | string | yes | A canonical metric name, e.g. gpu_temp_c. |
window_minutes | number | no | Restrict output to windows at or below this size. |
An unknown metric name throws with the full canonical list in the message.
{
"metric": "process_rss_bytes",
"windows": [
{
"window_minutes": 30,
"count": 120,
"mean": 2540000000,
"p95": 2870000000,
"max": 2900000000,
"latest": 2880000000,
"slope_per_hour": 1200000000,
"r_squared": 0.9821,
"consecutive_ms": 0,
"span_ms": 1785000
}
]
}
consecutive_ms is the sustain-gate duration: how long the metric has held its current band.
The full payload also carries min, earliest and median per window.
health_policy
Explains the decision policy. Parameters: action (explain | config | decisions,
optional, defaults to explain).
health_policy({ "action": "explain" })
Action ladder (enter/exit pressure):
1 THROTTLE enter >= 55, exit <= 45
2 PAUSE_NEW_WORK enter >= 70, exit <= 60
3 REQUEST_APP_RESTART enter >= 80, exit <= 68
4 REQUEST_SYSTEM_REBOOT enter >= 95, exit <= 85
Levels 1 and 2 are applied by this plugin through the worker-control adapter.
Levels 3 and 4 are *requests* handed to dsh-restart, which owns restart execution.
A restart request is additionally gated by the maintenance window and by getMaintenanceReadiness().
Dimension weights: time=0.15, thermal=0.2, memory=0.25, runtime=0.15, worker=0.15, computer_use_ui=0.1
Weights are renormalized over the dimensions that actually have telemetry; the fraction that does is reported as coverage.
config returns the resolved numbers as JSON (preset, thresholds, weights,
cooldowns, anti_flap, maintenance, windows_ms, sampling). decisions returns the
last 20 audit records as JSON.
Configuration
Configuration is a partial document. It is deep-merged onto the selected preset, so every
leaf you omit comes from the preset; arrays replace rather than merge. The plugin registers
the settings namespace health-scheduler when the profile provides a settings service,
and runs from the bundle patch alone when it does not.
One caveat belongs here rather than in the configuration reference: the profile patch layer
is not a deep merge. A cordis.patch.yml row replaces the targeted row's whole config, so an
override that restates only part of a nested object loses the rest. The deep merge described in
this section applies to the document the plugin actually receives. See
What the bundle contributes.
The complete key-by-key reference, including the stats-file format and the command-probe
format, is in docs/configuration.md.
| Group | Keys | Default |
|---|
| Master | enabled, preset | true, balanced |
sampling | intervalMs, trendIntervalMs, summaryIntervalMs, providerBackoffMs, providerBackoffMaxMs | 15 s, 60 s, 300 s, 30 s, 600 s |
windows | rawMs, windowsMs, aggregateBucketMs, aggregateRetentionMs, dailyRetentionMs | 30 min, [5m,30m,2h,6h], 5 min, 24 h, 14 d |
trend | minSamples, minSpanMs, minRSquared | 3, 5 min, 0.5 |
weights | time, thermal, memory, runtime, worker, computer_use_ui | 0.15 / 0.20 / 0.25 / 0.15 / 0.15 / 0.10 |
thresholds | one {enter, exit} band per action | 55/45, 70/60, 80/68, 95/85 |
metrics | per canonical metric: band, weight, sustainMs, trendPointsPerHour, trendCap | see docs/metrics.md |
cooldowns | throttleMs, maintenanceMs, escalationMs | 5 min, 30 min, 60 min |
throttle | concurrencyLimit, concurrencyFactor | null, 0.5 |
maintenance | enabled, targetTime, windowStart, windowEnd, maxDeferMs, urgentOverridePressure, allowAppRestart, safePointRequired |
A configuration that cannot be acted on is rejected loudly. resolveConfig throws a
ConfigError naming the dotted path; the plugin's apply catches it, logs it and continues
on the balanced preset rather than failing the boot.
How much history you get
Three horizons, all bounded, all configurable:
| Horizon | Where | Default | Answers |
|---|
| Raw samples | windows.rawMs | 30 min, raised to the longest windowsMs entry | Percentiles, slopes, consecutive-band duration. |
| Aggregate buckets | windows.aggregateRetentionMs, aggregateBucketMs | 24 h, 5 min buckets | Long-horizon movement without keeping raw points. |
| Daily summaries | windows.dailyRetentionMs | 14 days | "Was last Tuesday worse than today?" — one {count, mean, max, min} per metric per local day. |
Daily summaries are rolled up from the aggregate buckets on demand, so their cost is
proportional to the number of buckets, not to uptime. They are published on every
HealthSnapshot as dailySummaries and in the JSON payload as daily_summaries, and they are
frozen at local midnight boundaries so a day is comparable across a DST change.
Presets
All three presets start from the same neutral balanced document and differ in exactly four
places. PRESET_SCALES in src/core/presets.ts is the whole story:
conservative: { bands: 0.85, cooldown: 1.5, maintenance: 0.85 }
balanced: { bands: 1, cooldown: 1, maintenance: 1 }
aggressive: { bands: 1.15, cooldown: 0.7, maintenance: 1.3 }
scalePreset multiplies every ladder enter and exit by bands, so hysteresis keeps
its proportion instead of flattening; it multiplies the three cooldowns by cooldown; and it
multiplies maxDeferMs, minStateDwellMs and minRepeatActionMs by maintenance.
urgentOverridePressure is divided by bands, so a lower-pressure machine treats pressure
as urgent sooner. Scaled thresholds are clamped into 1 … 99, so no rung can be scaled out of
reach.
| Field | conservative | balanced | aggressive |
|---|
thresholds.throttle enter / exit | 47 / 38 | 55 / 45 | 63 / 52 |
thresholds.pause_new_work enter / exit | 60 / 51 | 70 / 60 | 81 / 69 |
thresholds.request_app_restart enter / exit | 68 / 58 | 80 / 68 | 92 / 78 |
thresholds.request_system_reboot enter / exit | 81 / 72 | 95 / 85 | 99 / 98 |
cooldowns.throttleMs | 7.5 min | 5 min | 3.5 min |
cooldowns.maintenanceMs | 45 min | 30 min | 21 min |
cooldowns.escalationMs | 90 min | 60 min | 42 min |
maintenance.maxDeferMs | 51 min | 60 min | 78 min |
maintenance.urgentOverridePressure | 108 | 92 | 80 |
antiFlap.minStateDwellMs | 102 s | 120 s | 156 s |
antiFlap.minRepeatActionMs | 8.5 min | 10 min | 13 min |
Note that a preset preserves the relative width of every hysteresis band rather than
flattening it, and entries are capped at 99 so no rung can be scaled out of reach: a
pressure above 100 does not exist, so an entry of 109 would silently delete that rung. On
aggressive the REQUEST_SYSTEM_REBOOT rung needs a perfectly saturated model
(entry 99, exit 98), which is exactly the "only under real duress" behaviour that preset is
meant to have. The metric table itself — every band, sustainMs and trend point — is
identical in all three presets.
The three presets are also shipped as JSON under presets/ for diffing, together
with a JSON Schema for a whole configuration document.
Capabilities and degradation
The plugin degrades; it does not fail. Each of these is a normal, well-typed outcome.
| Situation | What happens |
|---|
dsh-restart is not installed | UnavailableRestartAdapter reports capability: 'unavailable'. Monitoring and throttling continue. A restart decision is downgraded to PAUSE_NEW_WORK with reason restart_capability_unavailable. |
| No worker-control service is bound | UnavailableWorkerControlAdapter reports unavailable. A THROTTLE attempt returns applied: false with a detail explaining it; the tick continues. |
| A provider throws | That provider only is disabled for an exponential backoff (providerBackoffMs * 2^steps, capped at providerBackoffMaxMs, starting once consecutive failures reach providerFailureLimit). Its metrics simply stop arriving and become unknown. |
A provider returns degraded: true | The metrics it did measure stay authoritative; the note is surfaced in the snapshot's warnings. |
| A sensor does not exist | The metric key is omitted. The dimension may become unknown, coverage drops, and the reason list says so. |
| No stats file exists yet / it is stale | Nothing is reported and a detail names the paths that were looked for; a stale file's contents are refused and the sample is degraded with the staleness detail. |
| A configured helper command fails or times out | The probe error is reported as a degraded sample; already-measured metrics are kept. |
| The decision log is not writable | DecisionLog counts the failure and exposes it via lastError; the scheduler keeps ticking. |
| No safe-point source is registered | foldReadiness returns safe: null. With safePointRequired: true a restart request is blocked with reason safe_point_unknown — an unanswered question is not a yes. A source that throws or hangs contributes an unknown reading after a 1-second per-source budget. |
| The settings service or tool runtime is absent | A warning is logged and the plugin runs from the bundle patch alone, or registers no tools. A ConfigError from a bad user document is logged and the balanced preset is used instead of failing the boot. |
Paired with dsh-restart
The two plugins are separate packages with separate state directories, and nothing in
either one imports the other. That is deliberate: either can be uninstalled without
touching the other's code. What connects them is a contract and a directory, and you
assemble it in three steps.
1. Point both at the same state directory. dsh-restart writes its ticket,
heartbeat and ledger there; this plugin reads the ledger to learn whether the
supervisor has entered safe mode.
# this plugin's row
- id: health-scheduler
name: dsh-health-scheduler
config:
maintenance:
safePointRequired: true
# dsh-restart's row
- id: restart
name: dsh-restart
config:
safety:
checkpointRequired: true
Both default to $DSH_HOME/restart and $DSH_HOME/health-scheduler respectively when
no override is given; set storage.directory on both if you want them somewhere else.
The restart plugin's allowedSources must include dsh-health-scheduler, which it does
by default.
2. Give this plugin the two things it cannot make itself. A restart request is only
raised when something answers getMaintenanceReadiness() and something accepts a
RestartRequest. In production that is the harness kernel (which owns task state and
therefore owns the safe point) plus dsh-restart. In a profile where you have neither
wired, leave maintenance.enabled: false — that is the default — and this plugin will
throttle and pause rather than pretend it can restart.
To wire them programmatically, provide a RestartAdapter — the interface this plugin
defines and dsh-restart's manager satisfies — when you construct the scheduler:
import { HealthScheduler, UnavailableRestartAdapter } from 'dsh-health-scheduler'
import { RestartManager, TicketStore, RestartAuditLog } from 'dsh-restart'
// `requestApplicationRestart` / `requestSystemRestart` / `cancelPendingRestart` and
// the `capability` field are the whole contract; the adapter is the four-line bridge.
const restart = {
id: 'dsh-restart',
capability: 'available',
requestApplicationRestart: (request) => manager.requestApplicationRestart(request),
requestSystemRestart: (request) => manager.requestSystemRestart(request),
cancelPendingRestart: (requestId) => manager.cancelPendingRestart(requestId),
}
The RestartRequest this plugin builds is the one dsh-restart validates field for
field, including mode: 'application', reasonCode: 'RUNTIME_PRESSURE' and
checkpointRequired taken from maintenance.safePointRequired. If dsh-restart
answers accepted: false with CHECKPOINT_FAILED or SUPERVISOR_ABSENT, that answer is
recorded verbatim in the audit log and the action is reported as not applied — this
plugin does not retry around a refusal.
Two ways to hand it over. Passing restart into the scheduler (above) is the
supported one. The plugin's own apply also opportunistically looks for an object
named healthScheduler on the harness context and uses it if it structurally matches
RestartAdapter; that is a convenience for a composition that already publishes one,
not a contract, and dsh-restart does not publish anything under that name today. If
neither is present, UnavailableRestartAdapter is used and the plugin says so — it
monitors and throttles, and a restart decision is downgraded with reason
restart_capability_unavailable.
3. Register the safe point. A thin provider that answers for the harness kernel is
all it takes, and SafePointRegistry contains a provider that throws or hangs:
scheduler.safePoints.register({
id: 'dsh-core',
readiness: () => ({ source: 'dsh-core', safe: kernel.isIdle(), reason: kernel.reason(), estimatedState: 'idle' }),
})
With no source registered, getMaintenanceReadiness() returns safe: null, and that is
treated as not safe. A restart request with an unanswered safe-point question is
blocked with reason safe_point_unknown rather than proceed — the same rule that makes
a missing sensor unknown rather than healthy.
Provider telemetry matrix
Seven provider ids are registered by default. Be careful with this table: most of the
designed metrics have no native source, and the plugin says so instead of guessing.
| Provider id | Measured natively | Needs an external seam | Notes |
|---|
hardware | cpu_usage | cpu_temp_c, gpu_temp_c, gpu_usage, thermal_throttle, power_limit_hit | cpu_usage is a real delta of os.cpus() time counters and needs no privilege. gpu_usage is not native — there is no GPU counter here. All five thermal metrics arrive through providerOptions.hardware.helperCommand or a stats file. |
memory | ram_total_bytes, ram_available_bytes, ram_used_ratio, process_rss_bytes | commit_used_ratio, process_private_bytes, vram_used_ratio | From os.totalmem(), os.freemem() and the process RSS. process_rss_bytes is the process tree only when a platform helper supplied treeRssBytes; otherwise it is this process's RSS. |
runtime | uptime_seconds; handle_count when the runtime exposes process.getActiveResourcesInfo() | event_loop_latency_ms, worker_process_count, thread_count, restart_count, ipc_timeout_rate; heartbeat_delay_ms from a heartbeat file | The five middle metrics come from an injected RuntimeFeed. The plugin's own entry point passes EMPTY_RUNTIME_FEED, which answers null to everything — so with no integration they are always unknown. handle_count counts libuv handles plus timeouts, not OS handles. |
workers | nothing | active_workers, queued_tasks, task_latency_ms, timeout_rate, retry_rate, failure_rate, spawn_failure_rate, abnormal_exit_rate, queue_delay_ms | Stats file or providerOptions.statsFile.commands. The plugin does not instrument worker internals. |
|
Two consequences worth stating plainly: on a default install with no helper, no stats file and
no integration, only cpu_usage, the three RAM ratios and uptime_seconds have values, and
the report says the coverage is low. And CPU/GPU temperature is not measured natively —
there is no NVML, no WMI and no LibreHardwareMonitor binding inside this plugin. See
docs/configuration.md for a
working PowerShell helper.
Not implemented yet / Roadmap
Everything below is honestly absent from 0.1.0 rather than partially working.
- The
./startup subpath export points at files that are not built. package.json exports
./startup as ./lib/startup.js with ./lib/types/startup.d.ts; there is no src/startup.ts,
so that subpath cannot resolve. Nothing imports it and the bundle patch does not use it, so the
practical effect is a dead export rather than a broken install.
- No UI page. Every field the design's Health page needs — pressure, the per-dimension table,
per-metric values, maintenance, readiness, capabilities, provider status, trends, recent
decisions and the daily summaries — is in
HealthSnapshot and in the metricsSnapshot() JSON,
but no client plugin renders it.
- Aggregates are rolled up per day but never fitted.
RollingStore.dailySummaries() and
windows.dailyRetentionMs work, so "was last Tuesday worse than today?" is answerable. What is
missing is a trend read over that horizon: TrendAnalyzer fits raw points only, so a horizon
longer than windows.rawMs returns no samples despite 24 hours of retained aggregates.
providerOptions.memory.extraPids covers the launcher tree, not just this process. The sum is
refreshed in the background from the platform process list, so a leak in a separately-launched
worker shows up in process_rss_bytes too.
- Long-horizon trends are fitted from aggregate buckets. When the requested horizon reaches
further back than
windows.rawMs, TrendAnalyzer fits the retained bucket means instead of raw
samples, and the trend summary names the series it used (aggregate buckets). A 24-hour leak is
therefore visible on a machine whose raw horizon is 6 hours.
- The Long-Run test tier is not implemented. The design asks for 6 h / 12 h synthetic and a
24 h real-machine soak; the repository has unit, plugin, scheduler and synthetic-scenario tests
only. The 30-minute, 120-evaluation decision-storm bound is the closest thing to it.
git_operations_per_minute has no consumer. It is collected, carried in the time
dimension and weighted 0.5, but with no band and no trend term it contributes nothing, and
nothing yet uses it as a safe-point busy signal.
escalation_requested_at_maximum_pressure is a statement of condition, not of history. The
engine has no channel through which a previous application restart could report its outcome, so
the reason names what the pressure is doing rather than claiming a prior restart failed.
Documentation index
| Document | English | 中文 |
|---|
| Canonical metric registry: unit, polarity, bounds, default band, sustain, provider | docs/metrics.md | docs/metrics.zh.md |
The pressure model: dimensions, ramps, WORST_WEIGHT, trend, coverage, worked examples | docs/pressure-model.md | docs/pressure-model.zh.md |
| The action ladder, anti-flapping, cooldowns, dispatch, safe points, maintenance phases | docs/policies.md | docs/policies.zh.md |
| Every configuration key, the stats-file format, the command-probe format | docs/configuration.md | docs/configuration.zh.md |
| Acceptance criteria and design scenarios mapped to named tests | docs/acceptance.md | docs/acceptance.zh.md |
| Preset documents and the generated JSON Schema |
Development
npm install # dev dependencies only; the plugin has no runtime dependencies
npm run build # tsc -p tsconfig.json -> lib/
npm test # npm run build && node --test tests/*.test.js
npm run test:only # node --test tests/*.test.js, against the existing lib/
npm run typecheck # tsc -p tsconfig.json --noEmit
npm run presets # regenerate presets/*.json from lib/
npm run verify:artifacts # assert the built artifacts and the presets are coherent
TypeScript settings worth knowing: strict, noUncheckedIndexedAccess,
noUnusedLocals, noUnusedParameters, verbatimModuleSyntax, target: ES2023,
module: NodeNext. The engine (src/core, src/types, src/providers, src/adapters,
src/audit) has no dependency on the harness; only src/dsh/ knows about Cordis, and it
does so through the narrow structural interfaces in src/dsh/context.ts. That is what makes
the whole engine testable without a runtime. The suite is 152 tests across 8 files and passes
as committed:
node --test tests/*.test.js
# tests 152 / suites 27 / pass 152 / fail 0
node scripts/generate-presets.mjs --check
# ok presets/balanced.json matches PRESETS.balanced
# ok presets/conservative.json matches PRESETS.conservative
# ok presets/aggressive.json matches PRESETS.aggressive
# ok presets/schema.json matches the configuration schema
# ok 4 generated files are up to date
FAQ
Does it restart my machine?
No. It cannot. There is no restart execution in this plugin at all — no taskkill, no
reboot, no process kill. Levels 3 and 4 produce a RestartRequest object and hand it to
whatever RestartAdapter is on the context. With no dsh-restart installed that adapter
reports unavailable and the request is downgraded to PAUSE_NEW_WORK.
What if a temperature sensor is missing?
The metric is simply absent from the sample. cpu_temp_c and gpu_temp_c are never measured
natively, so on a default install they are always absent unless you configure a helper command
or a stats file. The thermal dimension then scores only what it does have (cpu_usage, and
anything the helper supplies); if it has nothing at all it is unknown, coverage drops, and
the report prints it under Unknown dimensions: … (not scored as healthy). It is never
scored as 0.
Why is my pressure 0 but state DEGRADED?
Because they answer different questions. restart_pressure can be 0 while pressure sits above
the throttle exit band — the state machine calls that DEGRADED, not HEALTHY. And
pressure: null (nothing measurable at all) also yields DEGRADED, on purpose: unknown is
not health. The reason list on the decision record names the gate that held.
How do I disable it?
Three ways, in increasing bluntness. Set enabled: false in the plugin configuration — it
still loads, registers its namespace and its tools, but collects nothing and starts no loop.
Or disable individual providers with disabledProviders: ["hardware"]. Or remove it from the
profile with dsh plugin --profile web remove dsh-health-scheduler, which also drops it from
the profile's bundle list. Uninstalling does not affect DS-Hns.
How much disk does it use?
Almost none, and it is bounded. The decision log is one decisions.jsonl under
<DSH_HOME>/health-scheduler (or storage.directory), rotated by rename once it exceeds
storage.maxLogBytes — 4 MiB by default, so the worst case is roughly 8 MiB with one .bak
beside it. Nothing else is written: the rolling history is in memory, bounded by
windows.rawMs and windows.aggregateRetentionMs. storage.enabled: false makes the plugin
strictly memory-only.
How do I add a custom sensor?
Write it into a stats file, or expose it as a command probe. Both speak the canonical metric
vocabulary, and non-canonical keys are reported rather than silently dropped. If your sensor
is genuinely new, the canonical registry must be extended — a provider can only report names
that exist in src/types/metrics.ts. See
docs/configuration.md.
Will it fight with dsh-restart?
It cannot fight, because it cannot act. It posts a RestartRequest with a caller-chosen
requestId and reads back accepted / rejected plus the restart side's lifecycle state.
Locks, rate limits, checkpoint tokens and crash-loop breaking all belong to dsh-restart;
this plugin's cooldowns pace only its requests, and a rejection starts the cooldown instead
of retrying immediately. Uninstalling dsh-restart leaves monitoring and throttling fully
functional.
Does it need admin rights?
No. It reads os.cpus(), os.totalmem(), os.freemem(), process.memoryUsage(),
process.uptime() and (when available) process.getActiveResourcesInfo(), reads and stats
files, and optionally runs the helper command you configure. A helper pointed at something
privileged inherits your privileges — that is your choice, not a requirement of the plugin,
which never escalates on its own.
Why is coverage only 60 %?
Because only four of the six dimensions have telemetry. Coverage is the share of nominal
weight backed by actual measurements. With no integration, runtime metrics other than
uptime, all worker metrics and all computer_use_ui metrics have no source, so coverage
sits near 0.6 when hardware and memory are reporting. That is the feature working: a
60 %-coverage pressure is labelled as such instead of pretending to be a full-confidence one.
Does it slow my machine down?
It ticks every 15 seconds by default and each tick is cheap: a few os calls, one stats-file
read at most once per 2 seconds, and at most one helper run per configured probe. Rolling
memory is bounded by construction, so nothing grows with uptime. The heaviest thing it can do is
the helper command you configure, which runs under a timeout you set (helperTimeoutMs, 5 s).
License
MIT © 2026 dsh-health-scheduler contributors. See LICENSE.
This is a community plugin. It is not affiliated with, sponsored by, or endorsed by
DeepSeek. "DeepSeek Harness" and "DS-Hns" are used descriptively to say what the plugin
integrates with.
Installing a plugin means running third-party code with your privileges — this one
included. Read SECURITY.md before you install it into a profile that can
reach production credentials or an unattended machine.