9.8 KiB
Upgrading to the current ModelForge RC candidate
RC schema note: the repository's temporary
1.2.1version has not been released with schema20260830_0024. Version selection/stamping is a separate pending release tranche. The following preflight is the exact contract for rehearsing an existing stable v1.2.1/schema-0022 database before this candidate can be released; it must not be read as a READY announcement.
Required v1.2.1/schema-0022 role provisioning before migration 0024
Schema 0024 changes the production authority boundary. The API may no longer run as the table owner or the bootstrap superuser. Provisioning is deliberately not hidden inside Alembic because an application migration must not mint or retain an administrator credential.
-
Take and verify the pre-upgrade backup, then stop
api,web, and every older writer. This is a downtime migration; the 0024 migration also takes an advisory lock andACCESS EXCLUSIVEonaudit_events, but the stopped-application contract prevents an old process from resuming with its former owner credential. -
Generate three distinct URL-safe secrets: retained PostgreSQL administrator, non-superuser
modelforgemigration/schema owner, and NOINHERITmodelforge_runtimeAPI role. SetMODELFORGE_POSTGRES_ADMIN_USERto a new retained name such asmodelforge_admin; set the three password variables and both owner/runtime SQLAlchemy URLs documented inCONFIGURATION.md. -
Recreate only the PostgreSQL container so those values are in its environment. It may report unhealthy until the new admin role exists; that is expected during this bounded downtime. While still connecting with the existing v1.2.1
modelforgebootstrap superuser, run:docker compose exec -T postgres psql -v ON_ERROR_STOP=1 \ -U modelforge -d "$MODELFORGE_POSTGRES_DB" \ < deploy/postgres/provision-existing-1.2.1.sqlThe script first creates the retained administrator, then preserves the existing
modelforgerole OID/table ownership while demoting it to a non-superuser, creates the distinct non-owner runtime role, and purges stale memberships that would still permitSET ROLE. Passwords are read from process environment; none is in the SQL file. -
Run only the one-shot migration service. Its environment contains
MODELFORGE_MIGRATION_DATABASE_URLbut not the runtime/admin URL:docker compose -f docker-compose.yml -f docker-compose.production.yml run --rm migrateMigration 0024 fails before schema mutation unless both roles exist,
modelforgeis the current directly authenticated non-superuser owner,modelforge_runtimeis NOINHERIT/non-admin, and neither application role has an explicit role membership. The refusal explicitly says to run this v1.2.1-to-schema-0024 provisioning step. -
Start the API only after Alembic reaches 0024. Production startup checks live catalog truth: runtime role/ownership/admin flags, table/schema/database privileges, canonical function owner/SECURITY DEFINER/fixed search path/grants and exact body digest, PUBLIC execution, four exact origin-mode protection-trigger shapes, role memberships, and every unexpected executable non-system function. Any drift stops startup.
The live PostgreSQL privilege/function exercise is a managed-runner release gate. SQLite exercises only the defense-in-depth application Engine hooks and is never production proof.
Supported direct schema sources for this unreleased candidate are 20260827_0021,
20260828_0022, and 20260830_0023; the target is 20260830_0024. Older releases require an
intermediate upgrade to v1.0.0.
What this upgrade actually changes
The migration chain first adds the 0022 terminal-node/decommission state when needed, then 0023
validates every existing node-enrollment and node-publisher scope before adding exact database
constraints. Any row outside node.enroll or node.publish aborts the upgrade before constraint
DDL; it is never silently rewritten.
That has a consequence worth stating plainly: rolling back across these schema revisions needs the verified pre-upgrade database restore. Older applications do not claim compatibility with the newer schema.
1. Preflight
python scripts/upgrade.py --database-url postgresql+psycopg://user:pass@host:5432/modelforge --plan
It changes nothing and reports:
| Check | Blocks the upgrade |
|---|---|
| PostgreSQL major version | yes |
| Current schema revision is one this release supports | yes |
| The schema is complete, not half-migrated | yes |
| A verified backup exists | yes (overridable) |
| That backup's age | no, advisory |
| Serving work in flight | no, advisory |
| Enabled nodes speak a supported agent protocol | yes |
| A migration cutover sitting in an intermediate stage | no, advisory |
A refusal looks like this and is the tool working:
BLOCK current schema 20990101_0099 is TOO_NEW; upgrade the control plane, which is older
than the component reporting to it
BLOCK verified backup no verified backup exists; an upgrade without a recovery point cannot be
undone if the schema turns out to be irreversible
BLOCK agent protocol enabled nodes speak [2]; this release supports [1]; incompatible: [2]
refusing to upgrade: 3 blocker(s)
2. Take a backup, and verify it
The preflight refuses without one. If you have taken a backup by other means, --allow-without-backup
proceeds — that override is explicit and recorded in the report, never a default.
See OPERATIONS.md.
3. Stop the application
docker compose -f docker-compose.yml -f docker-compose.production.yml stop api web
Leave PostgreSQL and Redis running. This upgrade requires downtime: the control plane is stopped while the new version starts. Rolling, zero-downtime upgrade is not supported in v1 and is not claimed. Record the measured downtime for your deployment.
4. Migrate
python scripts/upgrade.py --database-url postgresql+psycopg://user:pass@host:5432/modelforge --apply
For the current candidate this reports schema migrated: <source> -> 20260830_0024. Runtime
compatibility begins only after the scope validation, audit-chain checkpoint and PostgreSQL role
boundary reach 0024.
5. Start the new version
MODELFORGE_VERSION=1.2.1 \
MODELFORGE_API_IMAGE=modelforge-api:1.2.1 \
MODELFORGE_WEB_IMAGE=modelforge-web:1.2.1 \
docker compose -f docker-compose.yml -f docker-compose.production.yml up -d --no-build api web
Pin MODELFORGE_API_IMAGE and MODELFORGE_WEB_IMAGE to the exact release digests from the release
manifest when the images are available through a registry. A local source build is an explicit
alternative for development or rehearsal only: pass MODELFORGE_COMMIT, MODELFORGE_BUILT_AT and
--build yourself. Never overwrite a release tag with such a local build.
6. Verify
curl -s http://127.0.0.1:8000/api/v1/version # expect 1.2.1 and your commit
curl -s http://127.0.0.1:8000/api/v1/health/ready
Then compare the pre/post authoritative-state fingerprint: models, revisions, artifact sets, runtime profiles, projects, bindings, deployments, lifecycle and migration state, recovery state, policies, SLO state, node identities and credential hashes. Only the expected 0022 structures, 0023 scope constraints, 0024 audit boundary and legitimate reconciliation/audit activity may differ.
Two changes are expected and are not damage: the liveness monitor writes a NODE_BECAME_OFFLINE
event for any node not reporting yet, and the alert engine records the transitions that follow. Both
are the platform observing its new surroundings correctly.
If the upgrade fails
| Symptom | What it means | What to do |
|---|---|---|
| Preflight blocks | The upgrade never started; nothing changed | Fix what it names and re-run |
| The API exits during startup | Configuration was refused before any request was served; the exit code is 3 and the log lists every problem | Fix the settings it names and start again |
| The API starts but is unhealthy | Something beyond configuration | Roll back, then investigate with the deployment up |
Rolling back across schemas 0022, 0023 and 0024
Application rollback alone is not enough. Restore the verified backup taken in step 2, then start the previous application version — see RUNBOOK_DATABASE_RESTORE.md. The 0022 → 0021 Alembic downgrade is retained and release-tested to prove DDL mechanics on an isolated, tombstone-free copy; 0023 → 0022 only removes the two scope constraints. A 0024 downgrade is refused after v2 audit events exist because rewriting accepted audit history would destroy the evidence the boundary was designed to preserve. These DDL paths are not the production rollback promise; the 0022 → 0021 step removes terminal node tombstones and decommission-operation evidence. Production rollback therefore restores the verified pre-upgrade backup instead of discarding history created after the upgrade.
Agent compatibility
The control plane speaks agent protocol version 1 and accepts version 1. An agent reporting anything
else is answered explicitly — TOO_OLD, TOO_NEW or UNKNOWN — and told what to do, rather than
being silently accepted. The preflight blocks the upgrade when an enabled node speaks a protocol
this release does not support, so you find out before the control plane restarts rather than after.
Node Agents do not need to be upgraded in lockstep because protocol 1 is unchanged. Pin and deploy the matching Node Agent image to align package and OCI provenance. Upgrading only the Node Agent does not require re-enrollment; the control plane applies the fixed-scope migration before serving. The image deliberately preserves runtime UID 100 and GID 101, so the existing durable agent-state volume retains its ownership contract.