Files
ModelForge/docs/architecture/POSTGRESQL_BACKUP_RESTORE.md
T

5.5 KiB

PostgreSQL backup and restore

Method

ModelForge owns its own backup rather than delegating to a filesystem snapshot:

pg_dump --format=custom --compress=6 --no-owner --no-privileges --serializable-deferrable

Custom format restores schema, data, sequences, constraints, indexes and extension dependencies through pg_restore, and --serializable-deferrable takes the whole dump in one snapshot that cannot see a partially applied transaction. Copying /var/lib/postgresql/data from a running server is never done.

The control-plane image ships pg_dump, pg_restore and psql pinned to the same major version as the server (PostgreSQL 17), installed from Alpine's signed 3.23 repository. The release build upgrades repository packages before adding that client and verifies the resulting image.

Recorded identity

Every backup records the server version and version number, the major version, the database name, the installed extensions, the backup tool version, the backup method and a redacted connection string, alongside the Alembic revision. A restore preflight compares source and destination major versions and refuses POSTGRES_VERSION_INCOMPATIBLE.

Historical M15 live evidence: server PostgreSQL 17.11, server_version_num 170011, extensions plpgsql 1.0, tool pg_dump (PostgreSQL) 17.11 (Debian 17.11-1.pgdg13+2), schema 20260827_0021. The v1.1 Alpine release image independently reports pg_dump (PostgreSQL) 17.11.

Point-in-time recovery

PITR: NOT_SUPPORTED

M15 implements verified logical snapshots. WAL archiving and continuous point-in-time recovery are deliberately out of scope for v1. The value is exported as POINT_IN_TIME_SUPPORT and rendered on the recovery dashboard so no operator has to infer it. The practical consequence is a finite RPO bounded by the backup interval, and the platform measures it rather than claiming zero.

Restore

A restore runs against a genuinely separate database. _same_database compares host, port and database name, and a plan targeting the control plane's own database is refused at creation with DESTINATION_NOT_ISOLATED. Production targets additionally require MODELFORGE_RESTORE_ALLOW_PRODUCTION_TARGET and a DISASTER_RECOVERY mode.

pg_restore --no-owner --no-privileges --exit-on-error --single-transaction

--single-transaction means a failed restore leaves nothing behind. The destination must be empty: DESTINATION_NOT_EMPTY is returned rather than merging into existing data, which is what makes a retry safe.

Schema compatibility

The build walks its own Alembic chain by down_revision — filename order is not migration order — and compares the backup's revision against it:

Backup revision Result
equals head restored as is
known, older than head restored, then migrated forward to head
unknown to this build SCHEMA_TOO_NEW, fail closed
build ships no chain SCHEMA_UNKNOWN, fail closed
backup records none preflight fails

Restoring a newer backup onto older code is never silently downgraded. The discovery resolves both the source checkout and the installed image layout; a packaging error that hides the chain produces SCHEMA_UNKNOWN rather than an accidental pass.

Validation

After the restore the destination is measured, not assumed: table, constraint, index and sequence counts, the Alembic head, the audit high-water mark, a bounded semantic fingerprint and its diff against the fingerprint captured at backup time, and the residual row count of every current-truth table.

The READY gate requires all of:

database_validated          restored table and constraint counts meet the plan's requirement
current_truth_reconciled    no stale lease, telemetry or inventory row survives
schema_current              the destination is at this build's Alembic head
audit_available             the audit trail restored with a non-zero sequence
audit_chain_intact          every event/hash/link and the durable checkpoint verify strictly
fingerprint_compatible      current m15.2, or the exact schema-0022 m15.1 migration rule
no_unknown_corruption       every fingerprint difference is explained

READY always performs a fresh strict target audit verification. A resumed operation cannot reuse the presence of a journaled VALIDATING or READY duration as that proof. Audit rows are read with a stable (sequence, id) cursor, so a corrupt duplicate sequence split at the 500-row page boundary is observed and rejected rather than skipped. A separate aggregate identity check requires total rows, distinct UUIDs and distinct sequences to agree before the chain can pass; this also catches a duplicate cursor tuple after constraints were removed from a tampered restored schema.

Live evidence from disaster rehearsal A: 121 tables, 1,759 constraints, 590 indexes, audit high-water mark 838, 108 of 112 fingerprinted tables byte-identical, and four explained differences.

Identity columns

audit_events.sequence is GENERATED BY DEFAULT AS IDENTITY, but AuditWriter always supplies the value explicitly, so the underlying sequence is never advanced. The dump carries its SEQUENCE SET and the restore reproduces the source state exactly — an insert that omits sequence collides identically on both the source and the restored database. This is pre-existing M0 behaviour faithfully preserved by the restore, not a recovery defect, and it is recorded here because a future milestone that starts relying on the identity default must fix it in the audit writer first.