192 lines
9.8 KiB
Markdown
192 lines
9.8 KiB
Markdown
# Upgrading to the current ModelForge RC candidate
|
|
|
|
> RC schema note: the repository's temporary `1.2.1` version has **not** been released with
|
|
> schema `20260830_0024`. Version selection/stamping is a separate pending release tranche. The
|
|
> following preflight is the exact contract for rehearsing an existing stable v1.2.1/schema-0022
|
|
> database before this candidate can be released; it must not be read as a READY announcement.
|
|
|
|
## Required v1.2.1/schema-0022 role provisioning before migration 0024
|
|
|
|
Schema 0024 changes the production authority boundary. The API may no longer run as the table owner
|
|
or the bootstrap superuser. Provisioning is deliberately not hidden inside Alembic because an
|
|
application migration must not mint or retain an administrator credential.
|
|
|
|
1. Take and verify the pre-upgrade backup, then stop `api`, `web`, and every older writer. This is a
|
|
downtime migration; the 0024 migration also takes an advisory lock and `ACCESS EXCLUSIVE` on
|
|
`audit_events`, but the stopped-application contract prevents an old process from resuming with
|
|
its former owner credential.
|
|
2. Generate three distinct URL-safe secrets: retained PostgreSQL administrator, non-superuser
|
|
`modelforge` migration/schema owner, and NOINHERIT `modelforge_runtime` API role. Set
|
|
`MODELFORGE_POSTGRES_ADMIN_USER` to a new retained name such as `modelforge_admin`; set the three
|
|
password variables and both owner/runtime SQLAlchemy URLs documented in `CONFIGURATION.md`.
|
|
3. Recreate only the PostgreSQL container so those values are in its environment. It may report
|
|
unhealthy until the new admin role exists; that is expected during this bounded downtime. While
|
|
still connecting with the existing v1.2.1 `modelforge` bootstrap superuser, run:
|
|
|
|
```bash
|
|
docker compose exec -T postgres psql -v ON_ERROR_STOP=1 \
|
|
-U modelforge -d "$MODELFORGE_POSTGRES_DB" \
|
|
< deploy/postgres/provision-existing-1.2.1.sql
|
|
```
|
|
|
|
The script first creates the retained administrator, then preserves the existing `modelforge`
|
|
role OID/table ownership while demoting it to a non-superuser, creates the distinct non-owner
|
|
runtime role, and purges stale memberships that would still permit `SET ROLE`. Passwords are
|
|
read from process environment; none is in the SQL file.
|
|
4. Run only the one-shot migration service. Its environment contains
|
|
`MODELFORGE_MIGRATION_DATABASE_URL` but not the runtime/admin URL:
|
|
|
|
```bash
|
|
docker compose -f docker-compose.yml -f docker-compose.production.yml run --rm migrate
|
|
```
|
|
|
|
Migration 0024 fails before schema mutation unless both roles exist, `modelforge` is the current
|
|
directly authenticated non-superuser owner, `modelforge_runtime` is NOINHERIT/non-admin, and
|
|
neither application role has an explicit role membership.
|
|
The refusal explicitly says to run this v1.2.1-to-schema-0024 provisioning step.
|
|
5. Start the API only after Alembic reaches 0024. Production startup checks live catalog truth:
|
|
runtime role/ownership/admin flags, table/schema/database privileges, canonical function
|
|
owner/SECURITY DEFINER/fixed search path/grants and exact body digest, PUBLIC execution, four
|
|
exact origin-mode protection-trigger shapes, role memberships, and every unexpected executable
|
|
non-system function. Any drift stops startup.
|
|
|
|
The live PostgreSQL privilege/function exercise is a managed-runner release gate. SQLite exercises
|
|
only the defense-in-depth application Engine hooks and is never production proof.
|
|
|
|
Supported direct schema sources for this unreleased candidate are `20260827_0021`,
|
|
`20260828_0022`, and `20260830_0023`; the target is `20260830_0024`. Older releases require an
|
|
intermediate upgrade to v1.0.0.
|
|
|
|
## What this upgrade actually changes
|
|
|
|
The migration chain first adds the 0022 terminal-node/decommission state when needed, then 0023
|
|
validates every existing node-enrollment and node-publisher scope before adding exact database
|
|
constraints. Any row outside `node.enroll` or `node.publish` aborts the upgrade before constraint
|
|
DDL; it is never silently rewritten.
|
|
|
|
That has a consequence worth stating plainly: **rolling back across these schema revisions needs
|
|
the verified pre-upgrade database restore.** Older applications do not claim compatibility with
|
|
the newer schema.
|
|
|
|
## 1. Preflight
|
|
|
|
```bash
|
|
python scripts/upgrade.py --database-url postgresql+psycopg://user:pass@host:5432/modelforge --plan
|
|
```
|
|
|
|
It changes nothing and reports:
|
|
|
|
| Check | Blocks the upgrade |
|
|
| --- | --- |
|
|
| PostgreSQL major version | yes |
|
|
| Current schema revision is one this release supports | yes |
|
|
| The schema is complete, not half-migrated | yes |
|
|
| A verified backup exists | yes (overridable) |
|
|
| That backup's age | no, advisory |
|
|
| Serving work in flight | no, advisory |
|
|
| Enabled nodes speak a supported agent protocol | yes |
|
|
| A migration cutover sitting in an intermediate stage | no, advisory |
|
|
|
|
A refusal looks like this and is the tool working:
|
|
|
|
```text
|
|
BLOCK current schema 20990101_0099 is TOO_NEW; upgrade the control plane, which is older
|
|
than the component reporting to it
|
|
BLOCK verified backup no verified backup exists; an upgrade without a recovery point cannot be
|
|
undone if the schema turns out to be irreversible
|
|
BLOCK agent protocol enabled nodes speak [2]; this release supports [1]; incompatible: [2]
|
|
|
|
refusing to upgrade: 3 blocker(s)
|
|
```
|
|
|
|
## 2. Take a backup, and verify it
|
|
|
|
The preflight refuses without one. If you have taken a backup by other means, `--allow-without-backup`
|
|
proceeds — that override is explicit and recorded in the report, never a default.
|
|
|
|
See [OPERATIONS.md](OPERATIONS.md#backups-and-recovery).
|
|
|
|
## 3. Stop the application
|
|
|
|
```bash
|
|
docker compose -f docker-compose.yml -f docker-compose.production.yml stop api web
|
|
```
|
|
|
|
Leave PostgreSQL and Redis running. **This upgrade requires downtime**: the control plane is stopped
|
|
while the new version starts. Rolling, zero-downtime upgrade is not supported in v1 and is not
|
|
claimed. Record the measured downtime for your deployment.
|
|
|
|
## 4. Migrate
|
|
|
|
```bash
|
|
python scripts/upgrade.py --database-url postgresql+psycopg://user:pass@host:5432/modelforge --apply
|
|
```
|
|
|
|
For the current candidate this reports `schema migrated: <source> -> 20260830_0024`. Runtime
|
|
compatibility begins only after the scope validation, audit-chain checkpoint and PostgreSQL role
|
|
boundary reach `0024`.
|
|
|
|
## 5. Start the new version
|
|
|
|
```bash
|
|
MODELFORGE_VERSION=1.2.1 \
|
|
MODELFORGE_API_IMAGE=modelforge-api:1.2.1 \
|
|
MODELFORGE_WEB_IMAGE=modelforge-web:1.2.1 \
|
|
docker compose -f docker-compose.yml -f docker-compose.production.yml up -d --no-build api web
|
|
```
|
|
|
|
Pin `MODELFORGE_API_IMAGE` and `MODELFORGE_WEB_IMAGE` to the exact release digests from the release
|
|
manifest when the images are available through a registry. A local source build is an explicit
|
|
alternative for development or rehearsal only: pass `MODELFORGE_COMMIT`, `MODELFORGE_BUILT_AT` and
|
|
`--build` yourself. Never overwrite a release tag with such a local build.
|
|
|
|
## 6. Verify
|
|
|
|
```bash
|
|
curl -s http://127.0.0.1:8000/api/v1/version # expect 1.2.1 and your commit
|
|
curl -s http://127.0.0.1:8000/api/v1/health/ready
|
|
```
|
|
|
|
Then compare the pre/post authoritative-state fingerprint: models, revisions, artifact sets,
|
|
runtime profiles, projects, bindings, deployments, lifecycle and migration state, recovery state,
|
|
policies, SLO state, node identities and credential hashes. Only the expected 0022 structures,
|
|
0023 scope constraints, 0024 audit boundary and legitimate reconciliation/audit activity may differ.
|
|
|
|
Two changes are expected and are not damage: the liveness monitor writes a `NODE_BECAME_OFFLINE`
|
|
event for any node not reporting yet, and the alert engine records the transitions that follow. Both
|
|
are the platform observing its new surroundings correctly.
|
|
|
|
## If the upgrade fails
|
|
|
|
| Symptom | What it means | What to do |
|
|
| --- | --- | --- |
|
|
| Preflight blocks | The upgrade never started; nothing changed | Fix what it names and re-run |
|
|
| The API exits during startup | Configuration was refused before any request was served; the exit code is 3 and the log lists every problem | Fix the settings it names and start again |
|
|
| The API starts but is unhealthy | Something beyond configuration | Roll back, then investigate with the deployment up |
|
|
|
|
### Rolling back across schemas 0022, 0023 and 0024
|
|
|
|
Application rollback alone is not enough. Restore the verified backup taken in step 2, then start
|
|
the previous application version — see
|
|
[RUNBOOK_DATABASE_RESTORE.md](operations/RUNBOOK_DATABASE_RESTORE.md). The 0022 → 0021 Alembic
|
|
downgrade is retained and release-tested to prove DDL mechanics on an isolated, tombstone-free
|
|
copy; 0023 → 0022 only removes the two scope constraints. A 0024 downgrade is refused after v2 audit
|
|
events exist because rewriting accepted audit history would destroy the evidence the boundary was
|
|
designed to preserve. These DDL paths are not the production rollback promise;
|
|
the 0022 → 0021 step removes terminal node tombstones and
|
|
decommission-operation evidence. Production rollback therefore restores the verified pre-upgrade
|
|
backup instead of discarding history created after the upgrade.
|
|
|
|
## Agent compatibility
|
|
|
|
The control plane speaks agent protocol version 1 and accepts version 1. An agent reporting anything
|
|
else is answered explicitly — `TOO_OLD`, `TOO_NEW` or `UNKNOWN` — and told what to do, rather than
|
|
being silently accepted. The preflight blocks the upgrade when an enabled node speaks a protocol
|
|
this release does not support, so you find out before the control plane restarts rather than after.
|
|
|
|
Node Agents do not need to be upgraded in lockstep because protocol 1 is unchanged. Pin and deploy
|
|
the matching Node Agent image to align package and OCI provenance. Upgrading only the Node Agent
|
|
does not require re-enrollment; the control plane applies the fixed-scope migration before serving.
|
|
The image deliberately preserves runtime UID 100 and GID 101, so the existing durable agent-state
|
|
volume retains its ownership contract.
|