Files
ModelForge/docs/UPGRADE.md
T

9.8 KiB

Upgrading to the current ModelForge RC candidate

RC schema note: the repository's temporary 1.2.1 version has not been released with schema 20260830_0024. Version selection/stamping is a separate pending release tranche. The following preflight is the exact contract for rehearsing an existing stable v1.2.1/schema-0022 database before this candidate can be released; it must not be read as a READY announcement.

Required v1.2.1/schema-0022 role provisioning before migration 0024

Schema 0024 changes the production authority boundary. The API may no longer run as the table owner or the bootstrap superuser. Provisioning is deliberately not hidden inside Alembic because an application migration must not mint or retain an administrator credential.

  1. Take and verify the pre-upgrade backup, then stop api, web, and every older writer. This is a downtime migration; the 0024 migration also takes an advisory lock and ACCESS EXCLUSIVE on audit_events, but the stopped-application contract prevents an old process from resuming with its former owner credential.

  2. Generate three distinct URL-safe secrets: retained PostgreSQL administrator, non-superuser modelforge migration/schema owner, and NOINHERIT modelforge_runtime API role. Set MODELFORGE_POSTGRES_ADMIN_USER to a new retained name such as modelforge_admin; set the three password variables and both owner/runtime SQLAlchemy URLs documented in CONFIGURATION.md.

  3. Recreate only the PostgreSQL container so those values are in its environment. It may report unhealthy until the new admin role exists; that is expected during this bounded downtime. While still connecting with the existing v1.2.1 modelforge bootstrap superuser, run:

    docker compose exec -T postgres psql -v ON_ERROR_STOP=1 \
      -U modelforge -d "$MODELFORGE_POSTGRES_DB" \
      < deploy/postgres/provision-existing-1.2.1.sql
    

    The script first creates the retained administrator, then preserves the existing modelforge role OID/table ownership while demoting it to a non-superuser, creates the distinct non-owner runtime role, and purges stale memberships that would still permit SET ROLE. Passwords are read from process environment; none is in the SQL file.

  4. Run only the one-shot migration service. Its environment contains MODELFORGE_MIGRATION_DATABASE_URL but not the runtime/admin URL:

    docker compose -f docker-compose.yml -f docker-compose.production.yml run --rm migrate
    

    Migration 0024 fails before schema mutation unless both roles exist, modelforge is the current directly authenticated non-superuser owner, modelforge_runtime is NOINHERIT/non-admin, and neither application role has an explicit role membership. The refusal explicitly says to run this v1.2.1-to-schema-0024 provisioning step.

  5. Start the API only after Alembic reaches 0024. Production startup checks live catalog truth: runtime role/ownership/admin flags, table/schema/database privileges, canonical function owner/SECURITY DEFINER/fixed search path/grants and exact body digest, PUBLIC execution, four exact origin-mode protection-trigger shapes, role memberships, and every unexpected executable non-system function. Any drift stops startup.

The live PostgreSQL privilege/function exercise is a managed-runner release gate. SQLite exercises only the defense-in-depth application Engine hooks and is never production proof.

Supported direct schema sources for this unreleased candidate are 20260827_0021, 20260828_0022, and 20260830_0023; the target is 20260830_0024. Older releases require an intermediate upgrade to v1.0.0.

What this upgrade actually changes

The migration chain first adds the 0022 terminal-node/decommission state when needed, then 0023 validates every existing node-enrollment and node-publisher scope before adding exact database constraints. Any row outside node.enroll or node.publish aborts the upgrade before constraint DDL; it is never silently rewritten.

That has a consequence worth stating plainly: rolling back across these schema revisions needs the verified pre-upgrade database restore. Older applications do not claim compatibility with the newer schema.

1. Preflight

python scripts/upgrade.py --database-url postgresql+psycopg://user:pass@host:5432/modelforge --plan

It changes nothing and reports:

Check Blocks the upgrade
PostgreSQL major version yes
Current schema revision is one this release supports yes
The schema is complete, not half-migrated yes
A verified backup exists yes (overridable)
That backup's age no, advisory
Serving work in flight no, advisory
Enabled nodes speak a supported agent protocol yes
A migration cutover sitting in an intermediate stage no, advisory

A refusal looks like this and is the tool working:

BLOCK current schema    20990101_0099 is TOO_NEW; upgrade the control plane, which is older
                        than the component reporting to it
BLOCK verified backup   no verified backup exists; an upgrade without a recovery point cannot be
                        undone if the schema turns out to be irreversible
BLOCK agent protocol    enabled nodes speak [2]; this release supports [1]; incompatible: [2]

refusing to upgrade: 3 blocker(s)

2. Take a backup, and verify it

The preflight refuses without one. If you have taken a backup by other means, --allow-without-backup proceeds — that override is explicit and recorded in the report, never a default.

See OPERATIONS.md.

3. Stop the application

docker compose -f docker-compose.yml -f docker-compose.production.yml stop api web

Leave PostgreSQL and Redis running. This upgrade requires downtime: the control plane is stopped while the new version starts. Rolling, zero-downtime upgrade is not supported in v1 and is not claimed. Record the measured downtime for your deployment.

4. Migrate

python scripts/upgrade.py --database-url postgresql+psycopg://user:pass@host:5432/modelforge --apply

For the current candidate this reports schema migrated: <source> -> 20260830_0024. Runtime compatibility begins only after the scope validation, audit-chain checkpoint and PostgreSQL role boundary reach 0024.

5. Start the new version

MODELFORGE_VERSION=1.2.1 \
MODELFORGE_API_IMAGE=modelforge-api:1.2.1 \
MODELFORGE_WEB_IMAGE=modelforge-web:1.2.1 \
docker compose -f docker-compose.yml -f docker-compose.production.yml up -d --no-build api web

Pin MODELFORGE_API_IMAGE and MODELFORGE_WEB_IMAGE to the exact release digests from the release manifest when the images are available through a registry. A local source build is an explicit alternative for development or rehearsal only: pass MODELFORGE_COMMIT, MODELFORGE_BUILT_AT and --build yourself. Never overwrite a release tag with such a local build.

6. Verify

curl -s http://127.0.0.1:8000/api/v1/version   # expect 1.2.1 and your commit
curl -s http://127.0.0.1:8000/api/v1/health/ready

Then compare the pre/post authoritative-state fingerprint: models, revisions, artifact sets, runtime profiles, projects, bindings, deployments, lifecycle and migration state, recovery state, policies, SLO state, node identities and credential hashes. Only the expected 0022 structures, 0023 scope constraints, 0024 audit boundary and legitimate reconciliation/audit activity may differ.

Two changes are expected and are not damage: the liveness monitor writes a NODE_BECAME_OFFLINE event for any node not reporting yet, and the alert engine records the transitions that follow. Both are the platform observing its new surroundings correctly.

If the upgrade fails

Symptom What it means What to do
Preflight blocks The upgrade never started; nothing changed Fix what it names and re-run
The API exits during startup Configuration was refused before any request was served; the exit code is 3 and the log lists every problem Fix the settings it names and start again
The API starts but is unhealthy Something beyond configuration Roll back, then investigate with the deployment up

Rolling back across schemas 0022, 0023 and 0024

Application rollback alone is not enough. Restore the verified backup taken in step 2, then start the previous application version — see RUNBOOK_DATABASE_RESTORE.md. The 0022 → 0021 Alembic downgrade is retained and release-tested to prove DDL mechanics on an isolated, tombstone-free copy; 0023 → 0022 only removes the two scope constraints. A 0024 downgrade is refused after v2 audit events exist because rewriting accepted audit history would destroy the evidence the boundary was designed to preserve. These DDL paths are not the production rollback promise; the 0022 → 0021 step removes terminal node tombstones and decommission-operation evidence. Production rollback therefore restores the verified pre-upgrade backup instead of discarding history created after the upgrade.

Agent compatibility

The control plane speaks agent protocol version 1 and accepts version 1. An agent reporting anything else is answered explicitly — TOO_OLD, TOO_NEW or UNKNOWN — and told what to do, rather than being silently accepted. The preflight blocks the upgrade when an enabled node speaks a protocol this release does not support, so you find out before the control plane restarts rather than after.

Node Agents do not need to be upgraded in lockstep because protocol 1 is unchanged. Pin and deploy the matching Node Agent image to align package and OCI provenance. Upgrading only the Node Agent does not require re-enrollment; the control plane applies the fixed-scope migration before serving. The image deliberately preserves runtime UID 100 and GID 101, so the existing durable agent-state volume retains its ownership contract.