22 lines
1013 B
Markdown
22 lines
1013 B
Markdown
# Runbook: production node offline or stale
|
|
|
|
## Trigger
|
|
|
|
`NODE_OFFLINE` is pending/firing, or GPU Node freshness is `STALE`/over 90 seconds.
|
|
|
|
## Diagnose
|
|
|
|
1. Confirm API and Operations view timestamps; do not infer current health from an old snapshot.
|
|
2. Inspect Node Agent liveness, last server-received heartbeat/telemetry and credential state.
|
|
3. Check network/TLS reachability from GPU Node outward and the Node Agent service logs.
|
|
4. Review the incident for capability alerts correlated as `DOWNSTREAM`.
|
|
5. Verify ExampleRAG's external Ollama production baseline independently.
|
|
|
|
## Act and recover
|
|
|
|
Acknowledge with an operator reason. Restart the Node Agent only through the established host
|
|
runbook and only after checking unrelated workloads; never stop the GPU Node host. Do not promote a
|
|
fallback model. Recovery requires new server-received telemetry, online liveness, resolved alerts
|
|
and a retained timeline. Escalate if the credential is rejected, inventory identity changed or
|
|
freshness oscillates.
|