4.7 KiB
Runtime Plane
M9 multimodal extension
The runtime plane now has specialized transformers_trocr, transformers_siglip2, and transformers_whisper adapters. They use exact local artifacts, local_files_only, trust_remote_code=false, denied network access and typed modality payloads. Runtime environments declare supported modality/model types; RuntimeProfiles bind preprocessing/audio semantics; probes and serving record the same measured resource evidence and unload behavior as embedding/reranking workers.
Scope
M4 turns verified local bytes into reproducible technical execution evidence. M5 consumes only that evidence through typed deployment load/invoke/health/drain/unload jobs. It does not expose an inference service, schedule workloads, bind projects, benchmark quality, or activate production.
ArtifactSet -> CompatibilityAssessment -> RuntimeProfile -> RuntimeProbe -> DeploymentCandidate
Each arrow is a separately persisted gate. DeploymentCandidate.status=lab_ready and
production=false are the terminal M4 state; a separate exact M5 production approval and
CapabilityDeployment are required before routing.
The runtime worker remains outbound-only and private. M10 replaces its historical single resident slot with a bounded map keyed by deployment UUID. Load, invoke, health, drain and unload target one exact deployment and do not implicitly clear other proven residents. The worker reports a complete typed inventory including ArtifactSet, RuntimeProfile, environment fingerprint and 64-bit worker generation. CUDA allocator measurements prove process-owned model allocation/reclaim; NVML remains the total-GPU/external-pressure source.
Control-plane restart adopts only an exact inventory match. Worker absence reconciles database state to cold; an exact worker inventory may be adopted without duplicate load. Unknown GPU processes are never adopted and remain unmanaged external usage.
Boundaries
- The control plane owns evidence, approvals, typed leases, state transitions and audit events.
- A static Compose runtime worker on the selected node polls outbound with that node's credential.
- The worker receives a fixed probe type, exact registered relative artifact location, manifest, immutable profile and environment identity. It never receives a shell command.
- A child process owns the ML runtime and exits after the probe, making process death the final unload boundary. The parent remains responsive to cancellation and reports NVML evidence.
- Artifacts mount read-only; root filesystem is read-only;
/tmpis the only writable runtime filesystem; the worker is non-root, has no published ports, Docker socket, SSH material or host-root mount, and drops Linux capabilities.
Adapter lifecycle
RuntimeAdapter defines inspect_support, prepare, start, probe, health, stop, unload
and collect_runtime_facts. M4 executes only the SentenceTransformers adapter. Static inspectors
exist for Transformers, vLLM, llama.cpp, Diffusers and custom runners. Unsupported executable
adapters fail with RUNTIME_INCOMPATIBLE rather than falling through to arbitrary code.
Lifecycle and recovery
Probes move through queued, preparing, loading, healthchecking, ready, unloading, then
completed, failed or cancelled. Node-scoped leases are hashed at rest and expire. An expired
active lease becomes claimable by the same node, increments attempt_count, and gets a new token.
No active probe is represented as a deployment candidate.
Health
The worker reports distinct process, runtime, model and capability layers. M4 requires healthy
process/runtime/model plus a finite non-empty vector, offline-local evidence and successful VRAM
reclamation. Capability health remains not_routed_in_m4 by design.
M14 runtime observability
Runtime operational truth remains the typed ServingJob, health result, GPU lease and measured
residency records. M14 exports bounded operation/status counters, evaluates load/invoke/unload/health
success, detects a bounded crash loop and correlates capability degradation without sending commands
to workers. Observability persistence failure cannot block or silently reroute inference. See
docs/architecture/OBSERVABILITY.md and docs/operations/RUNBOOK_RUNTIME_FAILURE.md.
M16 fault behaviour
A Runtime Worker failure during load, inference or unload produces a typed failure for the request,
fences the worker generation so a late result from a dead generation cannot be committed, and
reconciles the lease and residency. A repeated crash raises RUNTIME_CRASH_LOOP.
The no_orphan_serving_work invariant asserts that no queued or leased job outlives its lease and
that no residency allocation sits on a disabled node, which is what "the worker failure was
reconciled" has to mean in practice.