4.9 KiB
Observability architecture
Boundary
M14 observes the existing control-plane truth. It does not introduce a second GPU, node, scheduler or runtime collector and it never becomes a serving dependency. NVML values remain owned by the Node Agent; gateway requests, placements, leases, residency, jobs, evaluations, lifecycle and migrations remain owned by their existing durable journals.
Telemetry has one of five explicit kinds:
COUNTER: monotonic outcomes, such as requests or operations;GAUGE: current bounded state, such as queue depth or bytes;HISTOGRAM: distributions with fixed buckets, such as request duration;EVENT: durable transitions with actor and time;STATE: a typed current classification such asSTALEorFIRING.
Metric contract
Names are lowercase Prometheus identifiers prefixed with modelforge_; counters end in _total,
durations in _seconds and storage values in _bytes. /metrics is operator-authenticated and is
not published by Compose. It includes process metrics even when historical storage is unavailable.
| Plane | Authoritative source | Export examples |
|---|---|---|
| API | in-process middleware | modelforge_api_requests_total, request-duration histogram |
| Gateway | gateway_requests |
request outcomes, request and queue duration histograms |
| Scheduler | placements, jobs, capacity snapshots | decisions, queue depth, schedulable/external bytes |
| Runtime | serving_jobs |
operation outcomes |
| Node/GPU | Node Agent receive time and NVML latest | freshness, RAM, utilization, temperature, power |
| Storage | registered storage roots and node volumes | free bytes and retained capacity snapshots |
| Acquisition | artifact jobs and registry state | job states and integrity alerts |
| Evaluation | evaluation runs | run states |
| Lifecycle | lifecycle operations | operation stages and rollback failures |
| Migration | migration plans/cutovers | states and failure alerts |
Cardinality and content
Permitted labels are bounded operational classifications: capability, status, priority, route class, node display name, device index, verdict, operation, state and availability. Label values are single-line, escaped and limited to 128 characters. The registry rejects request IDs, UUIDs, digests, query/prompt fields, filenames, raw errors and other high-cardinality or sensitive keys.
Metrics never contain prompts, query/document text, images, audio, vectors, artifact paths or raw payloads. Request correlation remains in structured logs and bounded database records, not metric labels.
Time and failure semantics
Capacity snapshots carry observed_at, received_at, calculated freshness and KNOWN, STALE or
UNAVAILABLE. Central receive time determines liveness. Stale data is historical evidence, never
current capacity truth.
The minute-level polling service catches persistence and collection failures, rolls back its own
transaction and sets modelforge_observability_degraded=1. It does not alter gateway, scheduler or
runtime decisions. The metrics endpoint falls back to the in-memory process registry on database
failure. This is OBSERVABILITY_DEGRADED, not inference failure.
Access and audit
All operations and metrics routes use the operator credential. Policy/rule revisions, acknowledgements and maintenance windows are audited. Individual metric samples are deliberately not audited. No notification or external-process-control adapter is part of M14.
M15 recovery observability
Observability history is authoritative and restores with the database: SLO policy revisions and evaluations, alert rule revisions, alerts, alert history, maintenance windows, incidents and capacity snapshots and aggregates are all fingerprinted subjects.
Current alert state is not. Recovery reconciliation suppresses restored PENDING/FIRING alerts so
a rebuilt control plane re-derives which alerts are firing now, while their history is retained. In
the live rehearsal three restored firing alerts were suppressed and the rebuilt plane raised its own
after re-evaluating current state, with 187 restored capacity snapshots plus the ones it measured
itself.
Recovery readiness is an operational signal rather than a new production SLO family: M15 adds no SLO policies, only the five recovery alert rules.
M16 observability under fault
Monitoring is never a safety dependency. When observability persistence is unavailable, the
in-memory modelforge_observability_degraded gauge rises, the operations overview reports
OBSERVABILITY_DEGRADED, and serving, scheduling and recovery continue on their own authoritative
state. The M16 invariants deliberately read none of the observability tables, so the question "is
state still sound?" can be answered precisely when monitoring is the thing that failed.
Alerts raised during a chaos scenario are real alerts and must be resolved or acknowledged with their true cause. The release gate reports any alert left pending or firing at the end of a run.