44 lines
1.4 KiB
Markdown
44 lines
1.4 KiB
Markdown
# Failure and Health Model
|
|
|
|
## Health layers
|
|
|
|
### Process health
|
|
|
|
Is the worker/control-plane process running?
|
|
|
|
### Runtime health
|
|
|
|
Can the adapter communicate with the worker and does its runtime/version match the deployment profile?
|
|
|
|
### Hardware health
|
|
|
|
Was the node inventory collected, and are previously observed accelerators still reconcilable?
|
|
|
|
Remote nodes additionally report `online`, `stale`, `offline` or `disabled`. These states derive
|
|
from central receive time, not the node clock. Stale/offline nodes are retained with last-good
|
|
inventory; they are visibly unavailable for future scheduling rather than silently deleted.
|
|
|
|
### Accelerator health
|
|
|
|
Is the accelerator present and observable? Unsupported optional metrics are not health failures.
|
|
|
|
### Model health
|
|
|
|
Can the deployed model load and execute a known smoke inference?
|
|
|
|
### Capability health
|
|
|
|
Can the current routing/policy for a capability satisfy its contract?
|
|
|
|
### Project health
|
|
|
|
Are all required project capabilities available within acceptable SLO/policy?
|
|
|
|
## Degradation examples
|
|
|
|
`rag.reranking` may allow a project-specific policy to continue retrieval without second-stage reranking.
|
|
|
|
`rag.embedding` may be configured as a hard dependency and fail requests rather than generate inconsistent vectors via an unapproved fallback.
|
|
|
|
Fallback semantics are part of capability/project policy, not implicit behavior in runtime code.
|