Files

3.6 KiB

Node recovery

A compute node is two separate things during recovery: a persisted hardware identity, which is authoritative history, and a credential, which is secret material that hashes cannot reconstruct.

Identity

ComputeNode.key is the agent's persisted identity, written to the node's state volume and reused across restarts. A recovered node that still has its identity file keeps its node id, its accelerators, its storage roots and its history. Losing the credential does not lose the identity.

A node whose identity file is also gone enrols as a new node. That is correct — ModelForge cannot distinguish it from different hardware — and the operator disables the superseded row rather than letting two nodes claim the same machine.

Single-use enrolment

One enrolment token mints exactly one node identity. The claim is a conditional update whose row count decides:

update node_enrollments set used_at = now()
 where id = :id and used_at is null

This closed a real defect found by the M15 node-recovery rehearsal: the previous check read used_at and wrote it later, and two agent threads racing into ensure_enrolled burned one token on two compute-node rows for the same hardware. The Node Agent's publication and artifact loops now also share an enrolment lock, so it issues one enrol request rather than two.

Credential recovery

RESTORABLE_SECRET      restored from backup and immediately usable
ROTATABLE_SECRET       hash restores as authoritative state; the consumer needs a new secret
NON_EXPORTABLE_SECRET  never written to a backup at all

Node credentials are ROTATABLE_SECRET. Only secret_hash is stored, so a restore reinstates which credentials existed and which were revoked, but cannot hand the plaintext back to a node that lost it. Scenarios:

Scenario Recovery
database restored, node still holds its credential the restored hash matches; the node reconnects
credential lost on the node operator issues a fresh enrolment token; the node re-enrols and keeps its identity
credential believed stolen operator revokes it; the node re-enrols
node hardware replaced new identity, new enrolment; the old row is disabled

Revocation is permanent. A revoked credential stays refused after re-enrolment: the partial unique index uq_active_node_credential allows one active credential per node, and enrolment revokes the previous one rather than deleting it, so the revocation survives in history.

Recovery sequence

node unavailable
  → liveness reflects offline/stale; capability health degrades honestly
  → operator issues an enrolment token
  → agent re-enrols, keeping its persisted identity
  → inventory and telemetry are re-collected
  → scheduler reconciles against re-measured NVML truth
  → capability smoke

Nothing about the node's current state is restored from a backup. After a control-plane restore the node rows exist with their identities and history, but telemetry, leases and residency are empty until a live agent reports again.

Live evidence

Disaster rehearsal D used a disposable Node Agent identity; the production GPU Node credential was never revoked. The disposable node enrolled once (one enrol request, zero 403s), its credential was revoked and refused three times consecutively with 401, its credential file was deleted, and it re-enrolled on a fresh token as node 41a7530c-7750-47c9-92d7-78302ee664fd — the same persisted identity, with one active and one revoked credential, one enabled rehearsal node and one distinct hostname. The old credential still returned 401 afterwards.