Files
ModelForge/docs/operations/RUNBOOK_NODE_OFFLINE.md
T

1013 B

Runbook: production node offline or stale

Trigger

NODE_OFFLINE is pending/firing, or GPU Node freshness is STALE/over 90 seconds.

Diagnose

  1. Confirm API and Operations view timestamps; do not infer current health from an old snapshot.
  2. Inspect Node Agent liveness, last server-received heartbeat/telemetry and credential state.
  3. Check network/TLS reachability from GPU Node outward and the Node Agent service logs.
  4. Review the incident for capability alerts correlated as DOWNSTREAM.
  5. Verify ExampleRAG's external Ollama production baseline independently.

Act and recover

Acknowledge with an operator reason. Restart the Node Agent only through the established host runbook and only after checking unrelated workloads; never stop the GPU Node host. Do not promote a fallback model. Recovery requires new server-received telemetry, online liveness, resolved alerts and a retained timeline. Escalate if the credential is rejected, inventory identity changed or freshness oscillates.