Files

301 lines
16 KiB
Markdown

# Threat Model
## M9 binary modality boundary
Document, image and audio content introduces decompression, oversized-payload and privacy risks. M9 applies decoded-byte, batch, page, pixel and duration limits; validates canonical base64 and media types; denies runtime network access; forbids remote repository code; keeps content in the transient payload store; and excludes payloads from logs, audit and metrics. Specialized workers expose typed jobs only and cannot execute generic shell commands.
## M10 scheduler boundary
The scheduler exposes typed dry-run, managed residency-policy, drain and unload operations only.
Ordinary clients cannot choose a node, model or eviction target. External GPU processes are
observation-only and have no controls. Placement history retains bounded numeric shape/provenance
metadata and never payload content. Database accelerator locks and expiring typed leases prevent
concurrent overcommit; uncertain telemetry fails closed. Runtime isolation remains non-root,
outbound-only and offline with read-only artifacts, no Docker socket, SSH material, Hub token,
generic shell or cloud fallback.
## M11 project-consumer boundary
Operational clients bind to one non-deprecated project contract and one capability scope. The server,
not the payload, assigns interactive/background priority. Rotation atomically invalidates the prior
hash-only credential. Consumers cannot select a deployment, model, node, runtime or worker. Image,
audio and document bytes remain transient and are excluded from audit, metrics and logs; adapters
enforce decoded-size/dimension/duration bounds and do not fall back to cloud services. OCR discovery
retains quarantine, safe-serialization, license and `trust_remote_code=false` gates.
## M6 additions
- Capability-scope escalation is denied by hashed service credentials and fixed client scopes.
- Bulk-priority escalation is denied because priority comes from server-side identity state.
- Cross-space vector reuse is denied by response, profile, payload and collection space checks.
- Evaluation accepts typed IDs/scores only; no arbitrary code or document-body upload is executed.
- Shadow targeting exists only in operator evaluation composition and is absent from public search
parameters and production alias mutation.
## Assets
- host and GPU compute resources;
- project data processed by local inference;
- Hugging Face/access tokens;
- model artifacts and derived artifacts;
- project credentials to ModelForge;
- benchmark datasets that may contain sensitive project material;
- production deployment and routing state;
- audit history.
## Trust boundaries
1. Public internet / model hub → downloader
2. Downloader → quarantine storage
3. Quarantine → approved artifact store
4. Control plane → inference workers
5. Project → ModelForge gateway
6. Browser/operator → control-plane write API
7. Remote compute node agent → central control-plane agent API
## Primary threats
### Malicious model repository
Potential arbitrary code, pickle payloads, custom Python modules, dependency confusion or misleading metadata.
Mitigations:
- treat repository as untrusted;
- prefer safe serialization;
- inspect file inventory;
- hash artifacts;
- quarantine before approval;
- default `trust_remote_code=false`;
- no production exception without explicit approval and evidence;
- isolated runtime with no secrets and no unnecessary network.
### Compromised runtime endpoint
Direct runtime exposure may bypass ModelForge authorization or surface runtime control endpoints.
Mitigations:
- private runtime network;
- only gateway can invoke workers;
- reverse proxy/API boundary;
- runtime endpoints not bound to public/LAN interfaces unnecessarily.
### Model artifact substitution
Mitigations:
- pin exact revision;
- resolve to commit SHA;
- store digests;
- verify before use;
- immutable approved artifacts.
### Unauthorized promotion/deletion
Mitigations:
- explicit privileged actions;
- dependency checks;
- append-only audit;
- confirmation for destructive cleanup;
- rollback retention policy.
### Data exfiltration
Mitigations:
- inference workers default to no egress;
- no cloud fallback without explicit separate capability/policy;
- no Hugging Face token in inference workers;
- redact secrets from logs;
- benchmark datasets have controlled access.
### GPU denial of service
Mitigations:
- workload priority;
- GPU lease scheduler;
- queue bounds;
- memory envelopes;
- OOM recovery;
- production workloads preempt/deny benchmarks as appropriate.
### Rogue or compromised compute node agent
Threats include stolen enrollment secrets, cross-node publication, replayed telemetry, fake client
timestamps, secret leakage and use of the agent as a remote-execution foothold.
Mitigations:
- short-lived single-use enrollment tokens and node credentials are random and stored hash-only;
- enrollment is rate-limited and all create/use/revoke/disable transitions are audited;
- every report is bound to the credential's node identity and an explicit protocol version;
- persistent sequences reject replay/reordering and central receive time determines liveness;
- disabled/revoked nodes fail closed;
- remote deployments require HTTPS with certificate validation;
- the agent is non-root, outbound-only, exposes no port, mounts no Docker socket and implements no
generic command-execution surface;
- UI/list APIs never redisplay enrollment or node credentials.
### Untrusted runtime execution
M4 permits a narrowly scoped model load only after exact-set LAB approval. The runtime worker is a
non-root static container with read-only root/artifact filesystems, no published port, Docker socket,
SSH key, project secret or host-root mount. It accepts typed leases only, re-hashes every file before
load, uses local-files-only/offline loading, enforces `trust_remote_code=false`, blocks IP network
connections during load/inference, and exits the child process before verifying VRAM reclamation.
LAB_READY evidence never creates a production route.
## Initial negative tests
- attempt direct runtime access from outside private runtime network;
- try promoting a quarantined artifact;
- try production-running a remote-code-required artifact without approval;
- attempt deleting an artifact referenced by stable deployment;
- attempt switching an embedding capability without migration;
- enqueue large benchmark during production saturation;
- force worker crash/OOM and verify controlled recovery;
- verify no runtime container receives HF token or project secrets.
- reuse an enrollment token and use one node credential for another identity;
- publish malformed, incompatible, stale and replayed agent reports;
- disable/revoke a node and verify further publication is denied;
- inspect logs and list APIs for credential leakage.
## M0 implementation mapping
- `PolicyDefaults` rejects unsafe changes to remote code, exact revision, digest, egress, runtime
exposure, automatic promotion, rollback retention and benchmark priority defaults.
- Runtime profile contracts default remote code and network egress to false and fingerprint the full
launch configuration.
- Artifact provenance validates exact hexadecimal revision and SHA-256 identities; derived artifacts
require conversion lineage.
- Promotion guards reject unverified artifacts, missing evidence/approval/rollback and embedding
changes without a ready migration.
- The M4 runtime worker receives only its node credential through a read-only state mount. It has no
hub token, project secret, Docker socket or public port; Compose and GPU Node inspection verify this.
- HTTP requests receive correlation IDs and structured request logs; persistence defines hash-chain-
ready, append-only audit events. Authentication/authorization arrives with the gateway and must
precede any production write API.
## M12 lifecycle controls
All lifecycle reads and writes require the operator/admin credential; capability-client, node and
runtime credentials are denied. Approval evidence stores identifiers and bounded status facts, never
secrets or model/project content. CAS, idempotency, single-stable uniqueness and execution-time cleanup
rechecks limit promotion/rollback/delete races. Lifecycle cannot stop or reclaim VRAM from Ollama,
Plex or Tdarr, invoke arbitrary shell, mount Docker control or expose a worker. Runtime remains offline
with `trust_remote_code=false`.
## M13 migration controls
All migration endpoints require the operator/admin credential. Plans accept identifiers, hashes,
counts, policy facts and a closed adapter operation enum; they accept no code, SQL, shell command,
secret, vector, query or document body. The control plane has no SSH/Docker-socket migration surface.
External alias changes remain adapter-owned and must report exact identity and health before database
commit. Generations, CAS, idempotency, one-active-production constraints and stale-source checks
limit replay/races. Failed rollback or ambiguous restart truth enters manual intervention. Active and
rollback-retained migrations block cleanup, and external Ollama/Plex/Tdarr remain observation-only.
## M14 observability controls
Operations and `/metrics` require the operator credential and are not public Compose ports.
Metric definitions reject high-cardinality/sensitive labels and exports contain no prompt, query,
document, image, audio, vector, filename, digest, request ID, secret or raw error. Label values are
escaped and bounded. Policy/rule changes, acknowledgement and maintenance windows are audited, while
individual samples are not. The poller owns no SSH, Docker socket, runtime command or external GPU
control. Persistence failure degrades monitoring only; it cannot trigger fallback inference,
promotion, rollback, deletion, process termination or serving failure.
## M16 chaos and release-gate controls
M16 validated the boundaries above under fault and adversarial load, and turned each conclusion into
a build-enforced property.
- PostgreSQL, Redis and the operator console bind to loopback by default. Before M16 the
control-plane database was published on every interface behind a development password: a probe
from the host's LAN address reached it and enumerated all 121 tables, including provenance,
credential hashes and the audit trail. A test fails if that binding returns.
- The API, Node Agent and console containers drop all capabilities and set `no-new-privileges`; the
Node Agent and console additionally run read-only. The console serves a static build from an
unprivileged nginx rather than a development server running as root.
- No product route can inject a fault, run a command or open a shell. Faults are injected only from
the container runtime and from existing rehearsal seams. Subprocess use is allowlisted to the
recovery plane, argv is always a list, and static tests enforce both.
- Fifteen executable invariants cover duplicate production state, duplicate node identity, stale
leases, mixed embedding spaces, unevidenced commits, restores from unverified backups, revoked
credential reuse, operator scopes on capability clients, unsafe promoted artifacts, orphaned work,
hidden promotion and audit-chain integrity. They are read-only and never touch the observability
database, so they remain usable when monitoring is what failed.
- Capability scope matching is exact: case, whitespace, a second capability in the same string and
zero-width characters do not widen it, and a contract version is part of the identity. A client
bound to a project cannot serve another project's binding.
- Unknown, revoked, near-miss and empty credentials return an identical status and code, so there is
no enumeration oracle. Sixty invalid attempts produced sixty identical refusals with no
amplification and no effect on the valid credential.
- A single-use enrolment token survives a sixty-way concurrent storm with exactly one winner, and
revocation racing authentication never yields an accepted revoked credential.
- Adversarial input across malformed bodies, wrong types, oversized payloads, mass assignment,
traversal, injection, XSS and header injection produces typed 4xx responses and never a server
error. Error responses carry no traceback, SQL, driver name, DSN, path or credential, verified
including with the database stopped.
- A body containing `NaN` or `Infinity` previously crashed the validation error handler, turning a
422 into a server error on unauthenticated-shaped input. Rejected values are now rendered safely
and truncated.
- Backup encryption remains AES-256-GCM from `cryptography` with a fresh nonce per chunk, associated
data binding each chunk to its key id and index, and fail-closed decryption that leaves no
plaintext. Static tests assert all three.
- Supply chain: a CycloneDX SBOM read from the built images, image provenance bound to the source
commit, zero dependency advisories after upgrading the installer, no floating dependency
specifier, no privileged mode, no Docker socket, and Gitleaks clean over full history.
## M15 recovery controls
Backups and restores are operator-authenticated admin routes; there is no public recovery surface
and no default administrator password anywhere in ModelForge.
- Control-plane backups are AES-256-GCM encrypted at rest using `cryptography`; ModelForge designs
no cryptography of its own and the cipher sits behind an adapter so a KMS/HSM can replace the
local key. A wrong key fails closed and removes both the partial file and the destination, so no
readable plaintext is ever left behind.
- The backup encryption key, the operator API key and the Hugging Face token are never written into
a backup, manifest, log or audit event. The configuration manifest names them and their recovery
class only. Live checks confirmed none of them, nor any node credential or database password,
appears in 328 KB of API logs or in any audit detail.
- Credentials are stored as hashes. A restore reinstates which credentials existed and which were
revoked but cannot return plaintext; consumers rotate. Revocation survives re-enrolment.
- One enrolment token mints exactly one node identity. The single-use claim is an atomic conditional
update, closing a race in which two concurrent agent threads could create two compute-node rows
for the same hardware, and the agent serialises its own enrolment.
- Every backup and restore path resolves through an allowlisted root with traversal, drive-qualified
component, home-expansion and symlink-escape rejection. Archive handling validates every member
before extraction and refuses links, devices, absolute or traversing names, more than 10,000
members or more than 64 GiB declared.
- Restore destinations are parsed, identifier-validated and compared against the running control
plane's own database; a self-targeting or non-PostgreSQL destination is refused. Production
targets need an explicit disaster-recovery mode and an explicit deployment flag, and destructive
replacement lives in the operator CLI rather than a remote call. No recovery route executes
arbitrary shell input.
- Restore responses expose only a redacted DSN; `RestorePlanResponse` has no field carrying a
password, and the published OpenAPI document is asserted against that.
- Rehydration is not a security bypass: a recovered artifact re-enters quarantine, is re-hashed
against provenance, is re-checked against the security policy, keeps `trust_remote_code=false`
and does not inherit its predecessor's approval. A digest mismatch fails the recovery.
- Semantic fingerprints exclude every column whose name carries secret material, so a recovery
comparison can be exported without leaking credential hashes.
- Recovery reconciliation clears current truth rather than resurrecting it, and the restore gate
refuses `READY` while a stale lease, telemetry row or inventory run survives.
## M5 serving additions
- Gateway bearer secrets are one-time, scoped, hashed at rest, revocable and distinct from operator
and node credentials.
- Request text is held transiently in Redis and is absent from PostgreSQL, audit and structured logs.
- Capability resolution is allow-listed; clients cannot submit a model path, runtime or arbitrary
capability.
- Database accelerator locks, bounded queues, leases, expiry and a 1 GiB reserve limit resource
exhaustion/races. Unmanaged GPU processes are capacity, never ModelForge kill targets.
- The worker retains M4 offline flags, `trust_remote_code=false`, read-only artifacts, dropped
capabilities and no public port.