Files
ModelForge/docs/operations/FAILURE_AND_HEALTH_MODEL.md
T

1.4 KiB

Failure and Health Model

Health layers

Process health

Is the worker/control-plane process running?

Runtime health

Can the adapter communicate with the worker and does its runtime/version match the deployment profile?

Hardware health

Was the node inventory collected, and are previously observed accelerators still reconcilable?

Remote nodes additionally report online, stale, offline or disabled. These states derive from central receive time, not the node clock. Stale/offline nodes are retained with last-good inventory; they are visibly unavailable for future scheduling rather than silently deleted.

Accelerator health

Is the accelerator present and observable? Unsupported optional metrics are not health failures.

Model health

Can the deployed model load and execute a known smoke inference?

Capability health

Can the current routing/policy for a capability satisfy its contract?

Project health

Are all required project capabilities available within acceptable SLO/policy?

Degradation examples

rag.reranking may allow a project-specific policy to continue retrieval without second-stage reranking.

rag.embedding may be configured as a hard dependency and fail requests rather than generate inconsistent vectors via an unapproved fallback.

Fallback semantics are part of capability/project policy, not implicit behavior in runtime code.