5.3 KiB
Disaster recovery
Canonical restore flow
fresh host / control plane
→ checkout the exact ModelForge commit recorded in the BackupSet
→ provision PostgreSQL of a compatible major version
→ verify the BackupSet (manifest hash, payload hashes, decryption, archive structure)
→ restore the database into an empty destination
→ run only the forward migrations this build knows
→ supply secrets; rotate what cannot be restored
→ start the control plane against the restored database
→ reconcile: clear current truth, suppress restored alerts, surface in-flight operations
→ register or re-enrol the node
→ rehydrate missing artifacts from their exact revisions
→ validate, compare fingerprints, measure RPO and RTO
→ operator declares recovery complete
RestorePlan
RestorePlan is first class: backup set, mode, target environment and label, redacted database
destination, artifact strategy, secret strategy, node strategy, expected ModelForge version,
preflight evidence and validation requirements. Plans are idempotent — the same backup, mode,
destination and label return the same plan rather than a second one.
VALIDATION isolated rehearsal; never a production environment
REPLACEMENT an operator-declared maintenance target
DISASTER_RECOVERY the only mode permitted to target production, and only when explicitly enabled
Preflight
Eleven checks run before any byte moves, and again at execution time:
backup_verified payload_hashes schema_compatibility
encryption_key destination_isolation destination_reachable
destination_empty postgres_compatibility disk_space
source_control_revision external_dependencies
A failing preflight moves the plan to PREFLIGHT_FAILED and start is refused with
restore_preflight_required.
RestoreOperation
The journal is durable and append-only. Each phase writes a RestoreOperationEvent with its
evidence, and phase durations accumulate into the measured RTO.
PLANNED → PREFLIGHT → RESTORING_DATABASE → RESTORING_CONFIGURATION
→ REHYDRATING_ARTIFACTS → RECONCILING → VALIDATING → READY
↘
FAILED | MANUAL_INTERVENTION_REQUIRED
Retry resumes the journalled operation rather than creating a second one. A completed database
phase is never replayed onto live data: the destination-empty guard turns a retry into
DESTINATION_NOT_EMPTY instead of a merge. A restore interrupted mid-phase is moved to
MANUAL_INTERVENTION_REQUIRED by startup reconciliation, never silently resumed.
Reconciliation
Reconciliation runs exactly once per restore operation, keyed by a marker audit event carrying that operation's id. Matching on the action alone would be wrong: a backup legitimately contains the reconciliation events of earlier restores, and an unscoped check made every later restore believe it had already reconciled.
Reconciliation clears every current-truth table, suppresses restored firing alerts so a stale alert
is not presented as present-day truth while its history is retained, and reports in-flight lifecycle
operations and migration cutovers by id and stage with the typed outcome
RECOVERY_RECONCILIATION_REQUIRED. It never guesses external state.
In-flight operations
A backup taken while a lifecycle promotion or a migration cutover is mid-flight restores that operation intact, with its idempotency key, expected version and generation. On startup the lifecycle reconciler conservatively rolls the incomplete operation back from its immutable snapshot exactly once; a second start changes nothing. Migration cutovers are deliberately not auto- resolved, because external alias truth cannot be inferred after a crash — they stay in their intermediate stage and are reported as requiring reconciliation.
RPO and RTO
Both are measured, never assumed.
RPO is the gap between the newest audit event in the restored database and the newest in the source at validation time; the restore point is the backup's completion timestamp. RTO is the sum of the measured phase durations.
Rehearsal discipline
A rehearsal that reuses the live database volume is not a rehearsal. docker-compose.dr.yml
provisions a separate PostgreSQL 17 with its own volume and a second control plane bound to it, and
every rehearsal starts from a destroyed and recreated volume. The production database, ExampleRAG's
Qdrant, ExampleVision's recognizer and external GPU workloads are never written to.
Rehearsal catalogue
| Scenario | Outcome | |
|---|---|---|
| A | Database loss | full restore into a fresh PostgreSQL, measured RTO/RPO |
| B | Corrupt backup | payload and manifest tampering both fail closed |
| C | Artifact loss | exact-revision rehydration into a disposable location |
| D | Node credential loss | revoke, re-enrol, identity preserved |
| E | Control-plane rebuild | isolated plane serving restored truth |
| F | In-flight lifecycle | reconciled exactly once |
| G | In-flight migration | surfaced, never auto-resolved |
| H | Observability history | history restored, current state re-derived |
See docs/quality/M15_VERIFICATION.md for the measured results and
docs/operations/RUNBOOK_FULL_DR.md for the operator procedure.