822 B
822 B
Runbook: GPU pressure
Trigger
GPU_PRESSURE, CAPABILITY_CAPACITY, stale capacity, or sustained HIGH/CRITICAL scheduler
pressure.
Diagnose
Compare NVML observed, ModelForge resident, leases, reserve, external and schedulable bytes. Confirm freshness and inspect active placement/queue decisions. Ollama, Plex and Tdarr are external capacity; ModelForge must not kill, pause or reconfigure them.
Act and recover
Pause optional LAB/benchmark submissions through existing policy where justified. Let bounded QoS, queueing and managed-idle eviction operate. Never manufacture an OOM for diagnosis. Resolve only after fresh telemetry shows recovered headroom/pressure and queues drain without orphan leases. Escalate unexplained VRAM drift, repeated crash/OOM or managed residency that does not reclaim.