# Runbook: GPU pressure ## Trigger `GPU_PRESSURE`, `CAPABILITY_CAPACITY`, stale capacity, or sustained `HIGH`/`CRITICAL` scheduler pressure. ## Diagnose Compare NVML observed, ModelForge resident, leases, reserve, external and schedulable bytes. Confirm freshness and inspect active placement/queue decisions. Ollama, Plex and Tdarr are external capacity; ModelForge must not kill, pause or reconfigure them. ## Act and recover Pause optional LAB/benchmark submissions through existing policy where justified. Let bounded QoS, queueing and managed-idle eviction operate. Never manufacture an OOM for diagnosis. Resolve only after fresh telemetry shows recovered headroom/pressure and queues drain without orphan leases. Escalate unexplained VRAM drift, repeated crash/OOM or managed residency that does not reclaim.