7.0 KiB
Worker/agent healthcheck contract
Applies to pulse-worker (cmd/worker) and pulse-agent (cmd/agent). Both
are currently foundation stubs (cmd/worker/main.go, cmd/agent/main.go)
that will gain real scheduling loops. This document is the contract the Go
implementation must satisfy so the compose-side healthcheck already wired in
deploy/compose.yaml reports true liveness instead of "process exists".
Why not an HTTP /healthz
pulse-api and pulse-web already expose a network port, so an HTTP
/healthz costs nothing extra there. pulse-worker and pulse-agent sit on
pulse-internal only and have no other reason to bind a port; opening one
solely for a healthcheck would be an unnecessary internal attack surface
(see apps/agent/AGENTS.md: "No generic shell/exec/file-write endpoint" and
the general least-privilege posture of ADR-0005/ADR-0010). A heartbeat file
on the already-mounted tmpfs /tmp needs no new port, no new capability,
and no new listener.
What the compose/image side already does
deploy/pulse-entrypoint.shis the containerENTRYPOINT. It writes the current unix time to/tmp/.pulse-started-atand thenexecs the real binary (PID 1 becomes the Go process). This exists only so the healthcheck script can grant a startup grace period independent of Docker'sstart_period, which suppresses unhealthy status but does not stop theHEALTHCHECK CMD(or its side effects) from running during that window.deploy/healthcheck-heartbeat.shis theHEALTHCHECK CMD(also wired explicitly indeploy/compose.yaml,interval=15s timeout=5s start_period=20s retries=3). It checks the mtime of the heartbeat file described below. If the file is missing for longer thanPULSE_HEARTBEAT_MAX_AGE_SECONDS(default 45s) after container start, or exists but its mtime is older than that same threshold, the script logs why and sendsSIGKILLto PID 1.docker compose up(non-swarm) does not restart a container merely because it is reported "unhealthy" — killing PID 1 turns that into a real container exit, whichrestart: unless-stoppedthen recovers automatically. Same-UIDSIGKILLneeds no Linux capability, so this works undercap_drop: [ALL].
You do not need to change either script or the Dockerfiles to implement this
contract — only the points below, inside cmd/worker / cmd/agent and
whatever internal packages they call.
What the Go side must implement
- Path. Write the heartbeat to the file named by the
PULSE_HEARTBEAT_FILEenvironment variable if set, otherwise/tmp/healthy. That path is on the container'stmpfs/tmpmount (already present indeploy/compose.yamlfor both services), so no new volume or mount is needed. - Cadence tied to real progress, not a free-running ticker. Update the
heartbeat only after the main scheduling loop completes an iteration (a
scheduler tick, a due-job scan, a lease renewal, a completed
discovery/capability-collection pass — whatever the loop's unit of work
is for that binary). Do not heartbeat from an independent goroutine
that ticks on a timer regardless of whether the main loop is stuck — that
would silently defeat the whole point of this contract (the exact failure
mode
kill -0 1had). - Interval. The loop must complete an iteration, and therefore heartbeat,
at least every 10 seconds even when there is no work to do (an empty
poll is still a completed iteration). This gives a comfortable margin
under the 45s staleness threshold enforced by
deploy/healthcheck-heartbeat.sh, tolerating a couple of missed/slow cycles before the container is killed. - First heartbeat. Write one heartbeat immediately after startup (config loaded, DB reachable) and before blocking on the first real unit of work, so a slow-but-healthy cold start is not mistaken for a hang.
- Bounded blocking calls. Every blocking call inside the loop iteration
(DB queries, Unraid/agent-protocol calls, probe dials, notification
sends) must use a
contextwith a timeout well under 10s, and long-running job execution must happen in a separate goroutine so the scheduler loop itself stays responsive. This is also required byapps/worker/AGENTS.md("Every job is idempotent, cancellable, bounded and observable") independent of this contract — the two requirements reinforce each other: if a call isn't bounded, the loop stalls, the heartbeat goes stale, and the container is killed and restarted, correctly surfacing the hang instead of hiding it. - Write semantics. A simple truncate-and-write (or an
os.Chtimestouch if content does not need to change) is sufficient; the healthcheck reads only the file's mtime, not its contents, so no locking or atomic rename is required for correctness. For operator debuggability, write the current UTC time in RFC 3339 (e.g.2026-08-04T12:00:03Z\n) as the file content sodocker exec <container> cat /tmp/healthyis meaningful. - Failure handling. If writing the heartbeat file itself fails (e.g. tmpfs write error), log it at error level and continue the loop — do not crash the process for a heartbeat I/O failure alone. A sustained write failure will naturally show up as staleness and get caught by the healthcheck.
- Shutdown. No special handling is required on
SIGTERM/graceful shutdown; leaving the heartbeat file in place is fine since the container is stopping anyway (service.WaitForStopalready handles the signal-driven shutdown path in bothcmd/worker/main.goandcmd/agent/main.go).
Tunables
| Env var | Default | Set where | Meaning |
|---|---|---|---|
PULSE_HEARTBEAT_FILE |
/tmp/healthy |
worker/agent process env (compose) | Path the Go process writes and the healthcheck script reads. |
PULSE_HEARTBEAT_MAX_AGE_SECONDS |
45 |
healthcheck script env (compose, if overridden) | Staleness threshold before the healthcheck force-restarts the container. |
PULSE_STARTED_AT_FILE |
/tmp/.pulse-started-at |
entrypoint/healthcheck script env (compose, if overridden) | Where the entrypoint records container start time, used only for startup grace. |
None of these are currently set as explicit environment variables in
deploy/compose.yaml (the defaults baked into the scripts are used); add
them there if the recommended values above ever need to change per
deployment.
Reviewer double-check
Self-terminating a container from inside its own healthcheck is an unusual
pattern. It was chosen because docker compose up (non-swarm, confirmed by
docs/operations/DEVELOPMENT_SETUP.md) does not auto-restart merely
"unhealthy" containers, and an external auto-heal watcher would need Docker
socket access, which ADR-0005 forbids host-wide. If that trade-off is
unacceptable, the alternative is to drop the kill -9 1 and rely on the
Unraid dashboard surfacing "unhealthy" status for manual operator action —
but that leaves the original "never restarts" finding only partially fixed
(visible, not self-healing).