2.0 KiB
Runtime Probes
Operator flow
- Review the exact ArtifactSet and static assessment.
- Record exact-set LAB execution approval.
- Queue the typed fixed-input probe for a selected node.
- Follow preparing/load/health/ready/unload progress.
- Inspect technical output and measured resource evidence.
The sole M4 embedding input is ModelForge runtime compatibility probe. It is retained only as a
SHA-256 in runtime results. This is a technical smoke probe, not a benchmark or quality evaluation.
Measurements
NVML samples are captured before load, after load, around inference and after child-process exit. Stored evidence includes absolute used/free VRAM, delta above baseline, peak, GPU utilization, temperature, power, load time, inference latency, batch and maximum sequence length. Reclamation is checked against the measured baseline plus a bounded tolerance.
Failure and cancellation
Failures are typed (MODEL_LOAD_FAILED, GPU_OOM, RUNTIME_CRASH, TIMEOUT, INVALID_OUTPUT,
HEALTHCHECK_FAILED, OFFLINE_LOAD_VIOLATION, GPU_MEMORY_NOT_RECLAIMED, and related gate
codes). Unexpected exception text is normalized at the child boundary. Cancellation terminates the
child, runs reclamation checks, reports cancelled, and never creates a candidate.
Offline guarantee
The image sets Hugging Face/Transformers offline variables. All loads use the exact local path,
local_files_only=true and trust_remote_code=false; an in-process socket guard rejects IP network
connections while model loading and inference run. No duplicate model download is permitted.
Logs and worker identity
Every claimed attempt receives a bounded runtime-worker://<node>/<probe>/attempt/<n> log
reference. The probe, node, lease attempt and child-process lifecycle together identify the worker
without exposing a general process-control API. Verbose runtime output is not stored in the probe
row; technical input is represented by its hash, credentials are redacted, and audit events remain
separate from diagnostic logs.