Files
ModelForge/docs/operations/FAILURE_AND_HEALTH_MODEL.md

44 lines
1.4 KiB
Markdown

# Failure and Health Model
## Health layers
### Process health
Is the worker/control-plane process running?
### Runtime health
Can the adapter communicate with the worker and does its runtime/version match the deployment profile?
### Hardware health
Was the node inventory collected, and are previously observed accelerators still reconcilable?
Remote nodes additionally report `online`, `stale`, `offline` or `disabled`. These states derive
from central receive time, not the node clock. Stale/offline nodes are retained with last-good
inventory; they are visibly unavailable for future scheduling rather than silently deleted.
### Accelerator health
Is the accelerator present and observable? Unsupported optional metrics are not health failures.
### Model health
Can the deployed model load and execute a known smoke inference?
### Capability health
Can the current routing/policy for a capability satisfy its contract?
### Project health
Are all required project capabilities available within acceptable SLO/policy?
## Degradation examples
`rag.reranking` may allow a project-specific policy to continue retrieval without second-stage reranking.
`rag.embedding` may be configured as a hard dependency and fail requests rather than generate inconsistent vectors via an unapproved fallback.
Fallback semantics are part of capability/project policy, not implicit behavior in runtime code.