Files

5.3 KiB

Disaster recovery

Canonical restore flow

fresh host / control plane
  → checkout the exact ModelForge commit recorded in the BackupSet
  → provision PostgreSQL of a compatible major version
  → verify the BackupSet (manifest hash, payload hashes, decryption, archive structure)
  → restore the database into an empty destination
  → run only the forward migrations this build knows
  → supply secrets; rotate what cannot be restored
  → start the control plane against the restored database
  → reconcile: clear current truth, suppress restored alerts, surface in-flight operations
  → register or re-enrol the node
  → rehydrate missing artifacts from their exact revisions
  → validate, compare fingerprints, measure RPO and RTO
  → operator declares recovery complete

RestorePlan

RestorePlan is first class: backup set, mode, target environment and label, redacted database destination, artifact strategy, secret strategy, node strategy, expected ModelForge version, preflight evidence and validation requirements. Plans are idempotent — the same backup, mode, destination and label return the same plan rather than a second one.

VALIDATION          isolated rehearsal; never a production environment
REPLACEMENT         an operator-declared maintenance target
DISASTER_RECOVERY   the only mode permitted to target production, and only when explicitly enabled

Preflight

Eleven checks run before any byte moves, and again at execution time:

backup_verified          payload_hashes           schema_compatibility
encryption_key           destination_isolation    destination_reachable
destination_empty        postgres_compatibility   disk_space
source_control_revision  external_dependencies

A failing preflight moves the plan to PREFLIGHT_FAILED and start is refused with restore_preflight_required.

RestoreOperation

The journal is durable and append-only. Each phase writes a RestoreOperationEvent with its evidence, and phase durations accumulate into the measured RTO.

PLANNED → PREFLIGHT → RESTORING_DATABASE → RESTORING_CONFIGURATION
        → REHYDRATING_ARTIFACTS → RECONCILING → VALIDATING → READY
                                                     ↘
                      FAILED | MANUAL_INTERVENTION_REQUIRED

Retry resumes the journalled operation rather than creating a second one. A completed database phase is never replayed onto live data: the destination-empty guard turns a retry into DESTINATION_NOT_EMPTY instead of a merge. A restore interrupted mid-phase is moved to MANUAL_INTERVENTION_REQUIRED by startup reconciliation, never silently resumed.

Reconciliation

Reconciliation runs exactly once per restore operation, keyed by a marker audit event carrying that operation's id. Matching on the action alone would be wrong: a backup legitimately contains the reconciliation events of earlier restores, and an unscoped check made every later restore believe it had already reconciled.

Reconciliation clears every current-truth table, suppresses restored firing alerts so a stale alert is not presented as present-day truth while its history is retained, and reports in-flight lifecycle operations and migration cutovers by id and stage with the typed outcome RECOVERY_RECONCILIATION_REQUIRED. It never guesses external state.

In-flight operations

A backup taken while a lifecycle promotion or a migration cutover is mid-flight restores that operation intact, with its idempotency key, expected version and generation. On startup the lifecycle reconciler conservatively rolls the incomplete operation back from its immutable snapshot exactly once; a second start changes nothing. Migration cutovers are deliberately not auto- resolved, because external alias truth cannot be inferred after a crash — they stay in their intermediate stage and are reported as requiring reconciliation.

RPO and RTO

Both are measured, never assumed.

RPO is the gap between the newest audit event in the restored database and the newest in the source at validation time; the restore point is the backup's completion timestamp. RTO is the sum of the measured phase durations.

Rehearsal discipline

A rehearsal that reuses the live database volume is not a rehearsal. docker-compose.dr.yml provisions a separate PostgreSQL 17 with its own volume and a second control plane bound to it, and every rehearsal starts from a destroyed and recreated volume. The production database, ExampleRAG's Qdrant, ExampleVision's recognizer and external GPU workloads are never written to.

Rehearsal catalogue

Scenario Outcome
A Database loss full restore into a fresh PostgreSQL, measured RTO/RPO
B Corrupt backup payload and manifest tampering both fail closed
C Artifact loss exact-revision rehydration into a disposable location
D Node credential loss revoke, re-enrol, identity preserved
E Control-plane rebuild isolated plane serving restored truth
F In-flight lifecycle reconciled exactly once
G In-flight migration surfaced, never auto-resolved
H Observability history history restored, current state re-derived

See docs/quality/M15_VERIFICATION.md for the measured results and docs/operations/RUNBOOK_FULL_DR.md for the operator procedure.