301 lines
16 KiB
Markdown
301 lines
16 KiB
Markdown
# Threat Model
|
|
|
|
## M9 binary modality boundary
|
|
|
|
Document, image and audio content introduces decompression, oversized-payload and privacy risks. M9 applies decoded-byte, batch, page, pixel and duration limits; validates canonical base64 and media types; denies runtime network access; forbids remote repository code; keeps content in the transient payload store; and excludes payloads from logs, audit and metrics. Specialized workers expose typed jobs only and cannot execute generic shell commands.
|
|
|
|
## M10 scheduler boundary
|
|
|
|
The scheduler exposes typed dry-run, managed residency-policy, drain and unload operations only.
|
|
Ordinary clients cannot choose a node, model or eviction target. External GPU processes are
|
|
observation-only and have no controls. Placement history retains bounded numeric shape/provenance
|
|
metadata and never payload content. Database accelerator locks and expiring typed leases prevent
|
|
concurrent overcommit; uncertain telemetry fails closed. Runtime isolation remains non-root,
|
|
outbound-only and offline with read-only artifacts, no Docker socket, SSH material, Hub token,
|
|
generic shell or cloud fallback.
|
|
|
|
## M11 project-consumer boundary
|
|
|
|
Operational clients bind to one non-deprecated project contract and one capability scope. The server,
|
|
not the payload, assigns interactive/background priority. Rotation atomically invalidates the prior
|
|
hash-only credential. Consumers cannot select a deployment, model, node, runtime or worker. Image,
|
|
audio and document bytes remain transient and are excluded from audit, metrics and logs; adapters
|
|
enforce decoded-size/dimension/duration bounds and do not fall back to cloud services. OCR discovery
|
|
retains quarantine, safe-serialization, license and `trust_remote_code=false` gates.
|
|
|
|
## M6 additions
|
|
|
|
- Capability-scope escalation is denied by hashed service credentials and fixed client scopes.
|
|
- Bulk-priority escalation is denied because priority comes from server-side identity state.
|
|
- Cross-space vector reuse is denied by response, profile, payload and collection space checks.
|
|
- Evaluation accepts typed IDs/scores only; no arbitrary code or document-body upload is executed.
|
|
- Shadow targeting exists only in operator evaluation composition and is absent from public search
|
|
parameters and production alias mutation.
|
|
|
|
## Assets
|
|
|
|
- host and GPU compute resources;
|
|
- project data processed by local inference;
|
|
- Hugging Face/access tokens;
|
|
- model artifacts and derived artifacts;
|
|
- project credentials to ModelForge;
|
|
- benchmark datasets that may contain sensitive project material;
|
|
- production deployment and routing state;
|
|
- audit history.
|
|
|
|
## Trust boundaries
|
|
|
|
1. Public internet / model hub → downloader
|
|
2. Downloader → quarantine storage
|
|
3. Quarantine → approved artifact store
|
|
4. Control plane → inference workers
|
|
5. Project → ModelForge gateway
|
|
6. Browser/operator → control-plane write API
|
|
7. Remote compute node agent → central control-plane agent API
|
|
|
|
## Primary threats
|
|
|
|
### Malicious model repository
|
|
|
|
Potential arbitrary code, pickle payloads, custom Python modules, dependency confusion or misleading metadata.
|
|
|
|
Mitigations:
|
|
|
|
- treat repository as untrusted;
|
|
- prefer safe serialization;
|
|
- inspect file inventory;
|
|
- hash artifacts;
|
|
- quarantine before approval;
|
|
- default `trust_remote_code=false`;
|
|
- no production exception without explicit approval and evidence;
|
|
- isolated runtime with no secrets and no unnecessary network.
|
|
|
|
### Compromised runtime endpoint
|
|
|
|
Direct runtime exposure may bypass ModelForge authorization or surface runtime control endpoints.
|
|
|
|
Mitigations:
|
|
|
|
- private runtime network;
|
|
- only gateway can invoke workers;
|
|
- reverse proxy/API boundary;
|
|
- runtime endpoints not bound to public/LAN interfaces unnecessarily.
|
|
|
|
### Model artifact substitution
|
|
|
|
Mitigations:
|
|
|
|
- pin exact revision;
|
|
- resolve to commit SHA;
|
|
- store digests;
|
|
- verify before use;
|
|
- immutable approved artifacts.
|
|
|
|
### Unauthorized promotion/deletion
|
|
|
|
Mitigations:
|
|
|
|
- explicit privileged actions;
|
|
- dependency checks;
|
|
- append-only audit;
|
|
- confirmation for destructive cleanup;
|
|
- rollback retention policy.
|
|
|
|
### Data exfiltration
|
|
|
|
Mitigations:
|
|
|
|
- inference workers default to no egress;
|
|
- no cloud fallback without explicit separate capability/policy;
|
|
- no Hugging Face token in inference workers;
|
|
- redact secrets from logs;
|
|
- benchmark datasets have controlled access.
|
|
|
|
### GPU denial of service
|
|
|
|
Mitigations:
|
|
|
|
- workload priority;
|
|
- GPU lease scheduler;
|
|
- queue bounds;
|
|
- memory envelopes;
|
|
- OOM recovery;
|
|
- production workloads preempt/deny benchmarks as appropriate.
|
|
|
|
### Rogue or compromised compute node agent
|
|
|
|
Threats include stolen enrollment secrets, cross-node publication, replayed telemetry, fake client
|
|
timestamps, secret leakage and use of the agent as a remote-execution foothold.
|
|
|
|
Mitigations:
|
|
|
|
- short-lived single-use enrollment tokens and node credentials are random and stored hash-only;
|
|
- enrollment is rate-limited and all create/use/revoke/disable transitions are audited;
|
|
- every report is bound to the credential's node identity and an explicit protocol version;
|
|
- persistent sequences reject replay/reordering and central receive time determines liveness;
|
|
- disabled/revoked nodes fail closed;
|
|
- remote deployments require HTTPS with certificate validation;
|
|
- the agent is non-root, outbound-only, exposes no port, mounts no Docker socket and implements no
|
|
generic command-execution surface;
|
|
- UI/list APIs never redisplay enrollment or node credentials.
|
|
|
|
### Untrusted runtime execution
|
|
|
|
M4 permits a narrowly scoped model load only after exact-set LAB approval. The runtime worker is a
|
|
non-root static container with read-only root/artifact filesystems, no published port, Docker socket,
|
|
SSH key, project secret or host-root mount. It accepts typed leases only, re-hashes every file before
|
|
load, uses local-files-only/offline loading, enforces `trust_remote_code=false`, blocks IP network
|
|
connections during load/inference, and exits the child process before verifying VRAM reclamation.
|
|
LAB_READY evidence never creates a production route.
|
|
|
|
## Initial negative tests
|
|
|
|
- attempt direct runtime access from outside private runtime network;
|
|
- try promoting a quarantined artifact;
|
|
- try production-running a remote-code-required artifact without approval;
|
|
- attempt deleting an artifact referenced by stable deployment;
|
|
- attempt switching an embedding capability without migration;
|
|
- enqueue large benchmark during production saturation;
|
|
- force worker crash/OOM and verify controlled recovery;
|
|
- verify no runtime container receives HF token or project secrets.
|
|
- reuse an enrollment token and use one node credential for another identity;
|
|
- publish malformed, incompatible, stale and replayed agent reports;
|
|
- disable/revoke a node and verify further publication is denied;
|
|
- inspect logs and list APIs for credential leakage.
|
|
|
|
## M0 implementation mapping
|
|
|
|
- `PolicyDefaults` rejects unsafe changes to remote code, exact revision, digest, egress, runtime
|
|
exposure, automatic promotion, rollback retention and benchmark priority defaults.
|
|
- Runtime profile contracts default remote code and network egress to false and fingerprint the full
|
|
launch configuration.
|
|
- Artifact provenance validates exact hexadecimal revision and SHA-256 identities; derived artifacts
|
|
require conversion lineage.
|
|
- Promotion guards reject unverified artifacts, missing evidence/approval/rollback and embedding
|
|
changes without a ready migration.
|
|
- The M4 runtime worker receives only its node credential through a read-only state mount. It has no
|
|
hub token, project secret, Docker socket or public port; Compose and GPU Node inspection verify this.
|
|
- HTTP requests receive correlation IDs and structured request logs; persistence defines hash-chain-
|
|
ready, append-only audit events. Authentication/authorization arrives with the gateway and must
|
|
precede any production write API.
|
|
|
|
## M12 lifecycle controls
|
|
|
|
All lifecycle reads and writes require the operator/admin credential; capability-client, node and
|
|
runtime credentials are denied. Approval evidence stores identifiers and bounded status facts, never
|
|
secrets or model/project content. CAS, idempotency, single-stable uniqueness and execution-time cleanup
|
|
rechecks limit promotion/rollback/delete races. Lifecycle cannot stop or reclaim VRAM from Ollama,
|
|
Plex or Tdarr, invoke arbitrary shell, mount Docker control or expose a worker. Runtime remains offline
|
|
with `trust_remote_code=false`.
|
|
|
|
## M13 migration controls
|
|
|
|
All migration endpoints require the operator/admin credential. Plans accept identifiers, hashes,
|
|
counts, policy facts and a closed adapter operation enum; they accept no code, SQL, shell command,
|
|
secret, vector, query or document body. The control plane has no SSH/Docker-socket migration surface.
|
|
External alias changes remain adapter-owned and must report exact identity and health before database
|
|
commit. Generations, CAS, idempotency, one-active-production constraints and stale-source checks
|
|
limit replay/races. Failed rollback or ambiguous restart truth enters manual intervention. Active and
|
|
rollback-retained migrations block cleanup, and external Ollama/Plex/Tdarr remain observation-only.
|
|
|
|
## M14 observability controls
|
|
|
|
Operations and `/metrics` require the operator credential and are not public Compose ports.
|
|
Metric definitions reject high-cardinality/sensitive labels and exports contain no prompt, query,
|
|
document, image, audio, vector, filename, digest, request ID, secret or raw error. Label values are
|
|
escaped and bounded. Policy/rule changes, acknowledgement and maintenance windows are audited, while
|
|
individual samples are not. The poller owns no SSH, Docker socket, runtime command or external GPU
|
|
control. Persistence failure degrades monitoring only; it cannot trigger fallback inference,
|
|
promotion, rollback, deletion, process termination or serving failure.
|
|
|
|
## M16 chaos and release-gate controls
|
|
|
|
M16 validated the boundaries above under fault and adversarial load, and turned each conclusion into
|
|
a build-enforced property.
|
|
|
|
- PostgreSQL, Redis and the operator console bind to loopback by default. Before M16 the
|
|
control-plane database was published on every interface behind a development password: a probe
|
|
from the host's LAN address reached it and enumerated all 121 tables, including provenance,
|
|
credential hashes and the audit trail. A test fails if that binding returns.
|
|
- The API, Node Agent and console containers drop all capabilities and set `no-new-privileges`; the
|
|
Node Agent and console additionally run read-only. The console serves a static build from an
|
|
unprivileged nginx rather than a development server running as root.
|
|
- No product route can inject a fault, run a command or open a shell. Faults are injected only from
|
|
the container runtime and from existing rehearsal seams. Subprocess use is allowlisted to the
|
|
recovery plane, argv is always a list, and static tests enforce both.
|
|
- Fifteen executable invariants cover duplicate production state, duplicate node identity, stale
|
|
leases, mixed embedding spaces, unevidenced commits, restores from unverified backups, revoked
|
|
credential reuse, operator scopes on capability clients, unsafe promoted artifacts, orphaned work,
|
|
hidden promotion and audit-chain integrity. They are read-only and never touch the observability
|
|
database, so they remain usable when monitoring is what failed.
|
|
- Capability scope matching is exact: case, whitespace, a second capability in the same string and
|
|
zero-width characters do not widen it, and a contract version is part of the identity. A client
|
|
bound to a project cannot serve another project's binding.
|
|
- Unknown, revoked, near-miss and empty credentials return an identical status and code, so there is
|
|
no enumeration oracle. Sixty invalid attempts produced sixty identical refusals with no
|
|
amplification and no effect on the valid credential.
|
|
- A single-use enrolment token survives a sixty-way concurrent storm with exactly one winner, and
|
|
revocation racing authentication never yields an accepted revoked credential.
|
|
- Adversarial input across malformed bodies, wrong types, oversized payloads, mass assignment,
|
|
traversal, injection, XSS and header injection produces typed 4xx responses and never a server
|
|
error. Error responses carry no traceback, SQL, driver name, DSN, path or credential, verified
|
|
including with the database stopped.
|
|
- A body containing `NaN` or `Infinity` previously crashed the validation error handler, turning a
|
|
422 into a server error on unauthenticated-shaped input. Rejected values are now rendered safely
|
|
and truncated.
|
|
- Backup encryption remains AES-256-GCM from `cryptography` with a fresh nonce per chunk, associated
|
|
data binding each chunk to its key id and index, and fail-closed decryption that leaves no
|
|
plaintext. Static tests assert all three.
|
|
- Supply chain: a CycloneDX SBOM read from the built images, image provenance bound to the source
|
|
commit, zero dependency advisories after upgrading the installer, no floating dependency
|
|
specifier, no privileged mode, no Docker socket, and Gitleaks clean over full history.
|
|
|
|
## M15 recovery controls
|
|
|
|
Backups and restores are operator-authenticated admin routes; there is no public recovery surface
|
|
and no default administrator password anywhere in ModelForge.
|
|
|
|
- Control-plane backups are AES-256-GCM encrypted at rest using `cryptography`; ModelForge designs
|
|
no cryptography of its own and the cipher sits behind an adapter so a KMS/HSM can replace the
|
|
local key. A wrong key fails closed and removes both the partial file and the destination, so no
|
|
readable plaintext is ever left behind.
|
|
- The backup encryption key, the operator API key and the Hugging Face token are never written into
|
|
a backup, manifest, log or audit event. The configuration manifest names them and their recovery
|
|
class only. Live checks confirmed none of them, nor any node credential or database password,
|
|
appears in 328 KB of API logs or in any audit detail.
|
|
- Credentials are stored as hashes. A restore reinstates which credentials existed and which were
|
|
revoked but cannot return plaintext; consumers rotate. Revocation survives re-enrolment.
|
|
- One enrolment token mints exactly one node identity. The single-use claim is an atomic conditional
|
|
update, closing a race in which two concurrent agent threads could create two compute-node rows
|
|
for the same hardware, and the agent serialises its own enrolment.
|
|
- Every backup and restore path resolves through an allowlisted root with traversal, drive-qualified
|
|
component, home-expansion and symlink-escape rejection. Archive handling validates every member
|
|
before extraction and refuses links, devices, absolute or traversing names, more than 10,000
|
|
members or more than 64 GiB declared.
|
|
- Restore destinations are parsed, identifier-validated and compared against the running control
|
|
plane's own database; a self-targeting or non-PostgreSQL destination is refused. Production
|
|
targets need an explicit disaster-recovery mode and an explicit deployment flag, and destructive
|
|
replacement lives in the operator CLI rather than a remote call. No recovery route executes
|
|
arbitrary shell input.
|
|
- Restore responses expose only a redacted DSN; `RestorePlanResponse` has no field carrying a
|
|
password, and the published OpenAPI document is asserted against that.
|
|
- Rehydration is not a security bypass: a recovered artifact re-enters quarantine, is re-hashed
|
|
against provenance, is re-checked against the security policy, keeps `trust_remote_code=false`
|
|
and does not inherit its predecessor's approval. A digest mismatch fails the recovery.
|
|
- Semantic fingerprints exclude every column whose name carries secret material, so a recovery
|
|
comparison can be exported without leaking credential hashes.
|
|
- Recovery reconciliation clears current truth rather than resurrecting it, and the restore gate
|
|
refuses `READY` while a stale lease, telemetry row or inventory run survives.
|
|
|
|
## M5 serving additions
|
|
|
|
- Gateway bearer secrets are one-time, scoped, hashed at rest, revocable and distinct from operator
|
|
and node credentials.
|
|
- Request text is held transiently in Redis and is absent from PostgreSQL, audit and structured logs.
|
|
- Capability resolution is allow-listed; clients cannot submit a model path, runtime or arbitrary
|
|
capability.
|
|
- Database accelerator locks, bounded queues, leases, expiry and a 1 GiB reserve limit resource
|
|
exhaustion/races. Unmanaged GPU processes are capacity, never ModelForge kill targets.
|
|
- The worker retains M4 offline flags, `trust_remote_code=false`, read-only artifacts, dropped
|
|
capabilities and no public port.
|