7.7 KiB
Troubleshooting
Operational states you will actually meet, what each one means, and what to do. Most of these are the platform refusing to guess — a refusal with a name is a better outcome than a plausible answer built on stale information.
The control plane will not start
It exits with code 3 and a list of settings
ModelForge refused to start: 7 configuration problem(s) must be resolved.
- [MISSING_REQUIRED_SETTING] MODELFORGE_OPERATOR_API_KEY: production requires an operator API key
- [INSECURE_PRODUCTION_SETTING] MODELFORGE_DATABASE_URL: the database password is a well-known
development value
...
This is startup validation, and it is deliberate. Every problem names the setting at fault. Fix them and start again; see CONFIGURATION.md.
It only refuses in production. In development the same problems are logged and the process starts,
so you find out at startup rather than at first use.
INCOMPATIBLE_DATABASE
The schema revision in front of the application is not one this release supports, or PostgreSQL is
older than major 16. The message says which of TOO_OLD, TOO_NEW or UNKNOWN applies and what to
do. See COMPATIBILITY.md.
Compose refuses to render the project
required variable MODELFORGE_OPERATOR_API_KEY is missing a value
The production overlay takes secrets through ${VAR:?message}, so a missing value fails before a
single container starts.
FileNotFoundError on a manifest
The config/ directory is not mounted, or is mounted empty. The deployment mounts ./config from
the directory you run Compose in.
Serving
503 NO_ELIGIBLE_NODE
No online node can serve this capability. In order:
- Is a node enrolled and
enabled? - Is its liveness
online, notstaleoroffline? - Does it have the accelerator the runtime profile requires?
A node goes online → stale → offline on the configured thresholds. Nothing is assigned to a node
that is not online — the platform will not send work to a node it cannot currently see.
CAPACITY_CONSTRAINED
A node exists but has no headroom. This is a correct outcome, not a failure. The scheduler refuses rather than risking an out-of-memory failure that would take down work already running.
VRAM held by processes ModelForge does not manage — Ollama, Plex, Tdarr, anything else on the host — counts as external pressure. ModelForge observes it and never reclaims it: it will not stop, pause, throttle or reconfigure a workload it does not own.
401 CAPABILITY_NOT_AUTHORIZED
The credential is missing, invalid, revoked, expired, or not scoped to this capability. Status and error code are identical across those cases so there is no enumeration oracle; the human-readable message does distinguish them, which only helps a caller who already holds the secret.
Scope is exact. rag.embedding@1 does not grant rag.embedding@2, appended whitespace does not
widen a scope, and a project-bound client cannot serve another project's binding.
Stale telemetry blocks admission
Rather than extrapolating from an old reading, the scheduler refuses. A reading it does not have is not a reading it can guess.
Nodes
NVIDIA_NVML_UNAVAILABLE
The Node Agent is configured for NVIDIA but cannot initialize NVML. It exits with code 3 before
enrollment, heartbeat, inventory, or telemetry publication. Confirm the host driver works with
nvidia-smi, NVIDIA Container Toolkit is installed, the container receives /dev/nvidia*, and the
runtime injects libnvidia-ml.so.1. Do not copy a driver library into the image: it must match the
host kernel driver. The v1.1.1 release image is Debian/glibc specifically for this injection path.
NVIDIA_DEVICE_NOT_FOUND means NVML initialized but enumerated no device;
NVIDIA_TELEMETRY_UNAVAILABLE means inventory exists but a listed GPU has no matching telemetry.
Use MODELFORGE_AGENT_ACCELERATOR_MODE=cpu only for a genuinely CPU-only node.
In auto mode, an inherited non-empty NVIDIA_VISIBLE_DEVICES is affirmative GPU evidence; clear
it or select cpu explicitly on a CPU-only deployment.
NODE_STALE then NODE_OFFLINE
The agent has stopped reporting. The agent is outbound-only, so look at the agent side first: is the process running, can it reach the control-plane URL, is its credential still valid?
A failed publication names the status, method and path — never the response body and never the credential. If you see repeated failures with no cause, check that the agent's control-plane URL is not still a placeholder.
The node reappears with a new identity
It should not. Identity is persisted and enrolment is single use. If it does, the persisted identity
file is not on durable storage — check MODELFORGE_AGENT_STATE_VOLUME.
Artifacts
ARTIFACT_REHYDRATION_BLOCKED
An artifact cannot be re-acquired from its recorded upstream identity. ModelForge will not fall back to a different source: an artifact that is not the one recorded is not the artifact.
An artifact is stuck in quarantine
It failed a supply-chain check. That is the check working. Quarantine is not a staging area to be emptied — see MODEL_SUPPLY_CHAIN_POLICY.md.
Backups and recovery
BACKUP_STALE
The newest verified backup is older than the configured threshold. Take one, and find out why the schedule did not.
BACKUP_FAILED with DESTINATION_UNAVAILABLE
The backup root is missing, full or not writable. The backup failed closed and produced nothing — there is no partial recovery point to mistake for a real one.
A backup is FAILED with MANIFEST_INCOMPLETE
It was interrupted before completion and can never become restore eligible. Take a new one.
RESTORE_FAILED
A restore stopped mid-phase. It is never silently resumed; it needs operator review.
See RUNBOOK_DATABASE_RESTORE.md.
Point-in-time recovery
NOT_SUPPORTED in v1. Verified snapshot restore with a measured RPO is what the platform offers.
Continuous WAL recovery is deliberately out of scope.
Migrations
A cutover sits in an intermediate stage
Migration is the one subsystem that deliberately does not resolve itself after an interruption. External alias truth cannot be inferred from a crash, so the cutover is reported and waits for an operator working through the typed adapter. An upgrade does not resolve it either.
See RUNBOOK_MIGRATION_FAILURE.md.
REQUIRES_REINDEX
The embedding space changed. Vectors from the old space are not comparable with the new one, and no amount of care makes them so. The migration engine will not swap an alias underneath data that has not been reindexed.
Observability
OBSERVABILITY_DEGRADED
Monitoring persistence is unavailable. Serving, scheduling and recovery continue on their own authoritative state — monitoring is never a safety dependency. The platform invariants deliberately read none of the observability tables, so "is state still sound?" can be answered precisely when monitoring is the thing that failed.
DEFERRED_EXTERNAL_VALIDATION
A validation that depends on a service outside ModelForge could not run. It is recorded as deferred, never as a pass.
Getting more detail
Every error response carries code, message and correlation_id and nothing else. Search the
logs for the correlation id. A 500 response body is exactly Internal Server Error — no DSN, no
driver name, no SQL, no path, no traceback.