6.0 KiB
Compute nodes and the Node Agent
A compute node is a GPU host that serves capabilities. The Node Agent runs there, reports inventory and telemetry, and claims work.
The control plane never dials a node. The agent is outbound-only: it needs to reach the API, and the API needs no route back. That is what makes a node behind NAT, on a home network, or on a different site work without inbound exposure.
Requirements on the node
| NVIDIA driver | matching your GPU |
| NVIDIA container runtime | required for the Runtime Worker |
| Docker | 24+ |
| Reachability | outbound HTTPS to the control plane |
| Storage | artifact root, quarantine and agent state on durable storage |
Enrolling
- An operator mints a single-use token:
curl -s -X POST -H "X-ModelForge-Admin-Token: $KEY" -H "Content-Type: application/json" \
-d '{"display_name":"GPU Node","expires_in_seconds":900}' \
http://127.0.0.1:8000/api/v1/admin/node-enrollments
- Configure the agent on the node:
MODELFORGE_AGENT_CONTROL_PLANE_URL=https://modelforge.example.internal:8000
MODELFORGE_AGENT_ENROLLMENT_TOKEN=<the token>
MODELFORGE_AGENT_HOSTNAME=gpu_node
MODELFORGE_AGENT_ACCELERATOR_MODE=nvidia
MODELFORGE_AGENT_TLS_VERIFY=true
MODELFORGE_AGENT_STATE_VOLUME=/mnt/user/appdata/modelforge/agent-state
- Start it:
export MODELFORGE_NODE_AGENT_IMAGE=modelforge-node-agent:1.1.1
docker compose -f docker-compose.yml -f docker-compose.node-agent.yml up -d node-agent
For production, set MODELFORGE_NODE_AGENT_IMAGE to the exact release tag or registry digest. The
Compose file retains build: for local source builds, but its fallback is explicitly local and
never the floating latest tag.
The token is single use and atomically claimed. Sixty concurrent attempts against one token produce exactly one identity and one active credential — that is enforced by a conditional update whose row count decides the winner, not by timing.
Verified on a genuinely fresh control plane before release: the node appeared once, online, with
exactly one active credential.
NVIDIA startup contract
The v1.1.1 Node Agent runs on a digest-pinned Debian/glibc base. NVIDIA Container Toolkit injects
the host's driver-facing libnvidia-ml.so.1, dynamic loader bindings and /dev/nvidia* devices;
the image deliberately does not bundle a userspace driver library. The earlier v1.1.0 Alpine/musl
image could import the pure-Python pynvml package but could not relocate GPU Node's glibc-linked
NVML library, producing NVMLError_LibraryNotFound.
MODELFORGE_AGENT_ACCELERATOR_MODE has three values:
nvidia: require NVML, at least one real GPU, and matching telemetry; canonical GPU Compose deployments use this value.cpu: permit a legitimate CPU-only node without NVML.auto: require NVML only when NVIDIA devices orNVIDIA_VISIBLE_DEVICESare observed.
The NVIDIA preflight runs before enrollment. A missing library, empty device result, or missing
telemetry exits with code 3 and one of NVIDIA_NVML_UNAVAILABLE, NVIDIA_DEVICE_NOT_FOUND, or
NVIDIA_TELEMETRY_UNAVAILABLE; no empty-success inventory or fabricated telemetry is sent. A
bounded diagnostic check that never enrolls is available as:
docker run --rm --gpus all --network none \
-e MODELFORGE_AGENT_ACCELERATOR_MODE=nvidia \
modelforge-node-agent:1.1.1 modelforge-node-agent --preflight
Identity
A node's identity is persisted and survives restarts. If a node reappears as a new identity, its state volume is not durable — fix that before anything else, because duplicate identities are the one thing the platform cannot untangle for you.
MODELFORGE_AGENT_IDENTITY_MODE=persisted keeps the identity in
MODELFORGE_AGENT_STATE_VOLUME. On Unraid, point that at appdata, not at a container-local path.
Liveness
online → stale → offline on the configured thresholds
(MODELFORGE_NODE_STALE_AFTER_SECONDS, MODELFORGE_NODE_OFFLINE_AFTER_SECONDS; offline must exceed
stale, and startup refuses a configuration where it does not).
No work is assigned to a node that is not online. A disconnected agent stops publishing and the control plane says so rather than assuming the node is fine — verified by disconnecting a disposable agent for 100 seconds and confirming its heartbeat genuinely stopped advancing.
Private CA
If the control plane is behind a private certificate authority:
docker compose -f docker-compose.yml \
-f docker-compose.node-agent.yml \
-f docker-compose.node-agent.private-ca.yml up -d node-agent
with MODELFORGE_AGENT_CA_CERT_PATH and MODELFORGE_AGENT_CONTROL_PLANE_HOST_ADDRESS set. Do not
turn off MODELFORGE_AGENT_TLS_VERIFY to work around a certificate problem — that replaces a
certificate error with a silent one.
The Runtime Worker
The worker executes the model. It is a separate service and a separate image, pinned by digest, built on the CUDA runtime base:
docker compose -f docker-compose.yml -f docker-compose.runtime-worker.yml up -d runtime-worker
It reports its own protocol version, and the same four compatibility answers apply.
Diagnosing an agent that will not report
A failed publication names the status, the method and the path — never the response body, which can echo what the agent was publishing, and never the credential:
ConnectError on POST /api/v1/agent/heartbeat (backoff 1s…16s)
That distinguishes a revoked credential from a DNS failure. If the cause is not obvious, check that
MODELFORGE_AGENT_CONTROL_PLANE_URL is a real address and not the example placeholder.
Compatibility
The control plane speaks agent protocol version 1 and accepts version 1. An agent reporting anything
else is told which of TOO_OLD, TOO_NEW or UNKNOWN applies and what to do. There is no silent
mismatch. See COMPATIBILITY.md.
v1.1.1 keeps protocol 1, so v1.0.0 and v1.1.0 agents remain protocol-compatible. NVIDIA nodes should run v1.1.1 to obtain the glibc/NVML fix and fail-closed startup contract.