# Compute nodes and the Node Agent A compute node is a GPU host that serves capabilities. The Node Agent runs there, reports inventory and telemetry, and claims work. **The control plane never dials a node.** The agent is outbound-only: it needs to reach the API, and the API needs no route back. That is what makes a node behind NAT, on a home network, or on a different site work without inbound exposure. ## Requirements on the node | | | | --- | --- | | NVIDIA driver | matching your GPU | | NVIDIA container runtime | required for the Runtime Worker | | Docker | 24+ | | Reachability | outbound HTTPS to the control plane | | Storage | artifact root, quarantine and agent state on durable storage | ## Enrolling 1. An operator mints a single-use token: ```bash curl -s -X POST -H "X-ModelForge-Admin-Token: $KEY" -H "Content-Type: application/json" \ -d '{"display_name":"GPU Node","expires_in_seconds":900}' \ http://127.0.0.1:8000/api/v1/admin/node-enrollments ``` 2. Configure the agent on the node: ```ini MODELFORGE_AGENT_CONTROL_PLANE_URL=https://modelforge.example.internal:8000 MODELFORGE_AGENT_ENROLLMENT_TOKEN= MODELFORGE_AGENT_HOSTNAME=gpu_node MODELFORGE_AGENT_ACCELERATOR_MODE=nvidia MODELFORGE_AGENT_TLS_VERIFY=true MODELFORGE_AGENT_STATE_VOLUME=/mnt/user/appdata/modelforge/agent-state ``` 3. Start it: ```bash export MODELFORGE_NODE_AGENT_IMAGE=modelforge-node-agent:1.1.1 docker compose -f docker-compose.yml -f docker-compose.node-agent.yml up -d node-agent ``` For production, set `MODELFORGE_NODE_AGENT_IMAGE` to the exact release tag or registry digest. The Compose file retains `build:` for local source builds, but its fallback is explicitly `local` and never the floating `latest` tag. The token is **single use and atomically claimed**. Sixty concurrent attempts against one token produce exactly one identity and one active credential — that is enforced by a conditional update whose row count decides the winner, not by timing. Verified on a genuinely fresh control plane before release: the node appeared once, `online`, with exactly one active credential. ## NVIDIA startup contract The v1.1.1 Node Agent runs on a digest-pinned Debian/glibc base. NVIDIA Container Toolkit injects the host's driver-facing `libnvidia-ml.so.1`, dynamic loader bindings and `/dev/nvidia*` devices; the image deliberately does not bundle a userspace driver library. The earlier v1.1.0 Alpine/musl image could import the pure-Python `pynvml` package but could not relocate GPU Node's glibc-linked NVML library, producing `NVMLError_LibraryNotFound`. `MODELFORGE_AGENT_ACCELERATOR_MODE` has three values: - `nvidia`: require NVML, at least one real GPU, and matching telemetry; canonical GPU Compose deployments use this value. - `cpu`: permit a legitimate CPU-only node without NVML. - `auto`: require NVML only when NVIDIA devices or `NVIDIA_VISIBLE_DEVICES` are observed. The NVIDIA preflight runs before enrollment. A missing library, empty device result, or missing telemetry exits with code 3 and one of `NVIDIA_NVML_UNAVAILABLE`, `NVIDIA_DEVICE_NOT_FOUND`, or `NVIDIA_TELEMETRY_UNAVAILABLE`; no empty-success inventory or fabricated telemetry is sent. A bounded diagnostic check that never enrolls is available as: ```bash docker run --rm --gpus all --network none \ -e MODELFORGE_AGENT_ACCELERATOR_MODE=nvidia \ modelforge-node-agent:1.1.1 modelforge-node-agent --preflight ``` ## Identity A node's identity is persisted and survives restarts. If a node reappears as a *new* identity, its state volume is not durable — fix that before anything else, because duplicate identities are the one thing the platform cannot untangle for you. `MODELFORGE_AGENT_IDENTITY_MODE=persisted` keeps the identity in `MODELFORGE_AGENT_STATE_VOLUME`. On Unraid, point that at `appdata`, not at a container-local path. ## Liveness `online → stale → offline` on the configured thresholds (`MODELFORGE_NODE_STALE_AFTER_SECONDS`, `MODELFORGE_NODE_OFFLINE_AFTER_SECONDS`; offline must exceed stale, and startup refuses a configuration where it does not). No work is assigned to a node that is not online. A disconnected agent stops publishing and the control plane says so rather than assuming the node is fine — verified by disconnecting a disposable agent for 100 seconds and confirming its heartbeat genuinely stopped advancing. ## Private CA If the control plane is behind a private certificate authority: ```bash docker compose -f docker-compose.yml \ -f docker-compose.node-agent.yml \ -f docker-compose.node-agent.private-ca.yml up -d node-agent ``` with `MODELFORGE_AGENT_CA_CERT_PATH` and `MODELFORGE_AGENT_CONTROL_PLANE_HOST_ADDRESS` set. Do not turn off `MODELFORGE_AGENT_TLS_VERIFY` to work around a certificate problem — that replaces a certificate error with a silent one. ## The Runtime Worker The worker executes the model. It is a separate service and a separate image, pinned by digest, built on the CUDA runtime base: ```bash docker compose -f docker-compose.yml -f docker-compose.runtime-worker.yml up -d runtime-worker ``` It reports its own protocol version, and the same four compatibility answers apply. ## Diagnosing an agent that will not report A failed publication names the status, the method and the path — never the response body, which can echo what the agent was publishing, and never the credential: ```text ConnectError on POST /api/v1/agent/heartbeat (backoff 1s…16s) ``` That distinguishes a revoked credential from a DNS failure. If the cause is not obvious, check that `MODELFORGE_AGENT_CONTROL_PLANE_URL` is a real address and not the example placeholder. ## Compatibility The control plane speaks agent protocol version 1 and accepts version 1. An agent reporting anything else is told which of `TOO_OLD`, `TOO_NEW` or `UNKNOWN` applies and what to do. There is no silent mismatch. See [COMPATIBILITY.md](COMPATIBILITY.md). v1.1.1 keeps protocol 1, so v1.0.0 and v1.1.0 agents remain protocol-compatible. NVIDIA nodes should run v1.1.1 to obtain the glibc/NVML fix and fail-closed startup contract.