3.2 KiB
Advanced GPU Scheduler
M10 turns the single-GPU admission boundary into one deterministic planner used by dry-runs and real requests. The planner consumes current NVML use, managed residency, only the unmaterialized part of active leases, authoritative envelopes, health, queue state, priority and versioned policy.
Accounting invariant
NVML is the physical truth and already includes loaded ModelForge models:
physical_used = max(observed_nvml_used, attributed_managed_resident)
unmanaged_external = max(0, observed_nvml_used - attributed_managed_resident)
future_lease = max(0, reserved - materialized)
schedulable = max(0, total - physical_used - future_leases - dynamic_reserve)
This deliberately does not subtract residency twice. Attribution is KNOWN, ESTIMATED or
UNKNOWN; missing or stale telemetry is UNKNOWN and has zero schedulable capacity. The invariant
delta is exposed for diagnosis without permitting negative external use.
The active m10-v1 policy uses the maximum of 1 GiB, 5% total VRAM and a 256 MiB runtime margin.
Deployment admission uses measured peak plus the maximum of 128 MiB or 10% measured peak. These
values are one versioned policy, not dispersed constants.
PlacementPlan
Every plan records capability, node and accelerator, eligible state, residents, unmanaged use,
reserve, requirement, priority, evictions, before/after headroom, reason codes and a deterministic
fingerprint. Verdicts are ADMIT, ADMIT_AFTER_EVICTION, QUEUE, REJECT_CAPACITY,
REJECT_HEALTH and REJECT_POLICY. Recent history is bounded to 500 records.
Pressure worsens immediately through NORMAL, ELEVATED, HIGH, CRITICAL. Recovery requires a
separate threshold and 30 seconds of stable evidence. At HIGH/ELEVATED, new background/LAB loads
wait; at CRITICAL, new admissions fail with EXTERNAL_GPU_PRESSURE. ModelForge never changes an
unmanaged process.
Operations and security
Operator controls are limited to dry-run, policy revision/LAB pause, managed residency policy, drain and idle unload. There is no shell, Docker socket, SSH or GPU-PID control. Ordinary clients select only a capability. Scheduler evidence contains bounded shapes and hashes, never content.
M11 validates the existing scheduler rather than adding a project shortcut. ExampleVision reference index requests use a background-bound client, while card and speech requests use interactive-bound clients. During a 40-reference indexing continuation, ten interactive vision requests all completed. Vision and ASR also ran under observed external GPU pressure without OOM; final unload returned managed residency and leases to zero.
M16 pressure behaviour
Under GPU pressure the scheduler rejects or queues rather than risking an out-of-memory failure. A capacity rejection is a correct outcome, not a failure, and the soak counts those separately from internal errors for exactly that reason.
External VRAM is measured as pressure and never reclaimed: Ollama, Plex, Tdarr and any other unmanaged process are observed, never stopped, paused, throttled or reconfigured.
Stale telemetry blocks admission rather than being extrapolated, and no_stale_gpu_lease asserts
no expired lease is still presented as a live reservation.