Files
ModelForge/docs/operations/RUNBOOK_GPU_PRESSURE.md

20 lines
822 B
Markdown

# Runbook: GPU pressure
## Trigger
`GPU_PRESSURE`, `CAPABILITY_CAPACITY`, stale capacity, or sustained `HIGH`/`CRITICAL` scheduler
pressure.
## Diagnose
Compare NVML observed, ModelForge resident, leases, reserve, external and schedulable bytes. Confirm
freshness and inspect active placement/queue decisions. Ollama, Plex and Tdarr are external capacity;
ModelForge must not kill, pause or reconfigure them.
## Act and recover
Pause optional LAB/benchmark submissions through existing policy where justified. Let bounded QoS,
queueing and managed-idle eviction operate. Never manufacture an OOM for diagnosis. Resolve only
after fresh telemetry shows recovered headroom/pressure and queues drain without orphan leases.
Escalate unexplained VRAM drift, repeated crash/OOM or managed residency that does not reclaim.