4.7 KiB
Soak testing
What a soak is for
A short run proves a path works. A soak proves it keeps working: that memory does not climb, that leases are released, that queues drain, that retention keeps row growth bounded, and that latency does not degrade as caches warm and evict.
python scripts/m16_soak.py --database-url "$MODELFORGE_RUNTIME_DATABASE_URL" --minutes 45 --secret <gateway-secret> --report soak.json
The database URL is mandatory and must use the non-owner runtime role. The harness contains no embedded credential or owner-role fallback.
Workload shape
Three gateway workers drive the real capability path — rag.embedding@1 through the authenticated
gateway to GPU Node — mixing three phases:
burst 3-6 requests back to back, 50-400 ms apart
steady single requests, 1.5-5 s apart
idle 8-25 s of nothing
The mix matters. A soak that only hammers never exercises keep-warm and unload, so it never observes a cold start. A soak that only idles never exercises admission control, so it never observes a capacity rejection. Batch sizes vary from one to four inputs so payload handling varies too.
Two read workers poll registry, projects, hardware, health, operations, alerts, lifecycle, migrations and recovery. One background worker drives the operational work the platform already performs for itself — capacity collection, SLO evaluation, alert evaluation — at a bounded rate.
What is measured
per kind requests, succeeded, capacity rejected, internal errors, status distribution
latency p50, p95, p99, mean, max
gateway cold starts, warm invocations, queue p95, inference p50 and p95, failure codes
resources container memory, CPU, PIDs and restart counts, before and after
database row growth per table, connection count
runtime active leases, serving jobs in flight, residency allocations
GPU used, free and utilisation per accelerator, plus scheduler pressure
invariants the full set, at every progress interval
Capacity rejections are counted separately from internal errors. A 429 or a QUEUE_FULL is the
scheduler refusing unsafe work, which is the platform behaving correctly; counting it as a failure
would reward a system that over-admits.
Stopping early
The soak aborts and reports partial evidence if an invariant is violated at any progress interval. Partial evidence honestly labelled is worth more than a full run whose result is already known to be invalid.
It should also be aborted manually on uncontrolled host pressure, degradation of an external workload, repeated OOM, data corruption or any unsafe production mutation.
Reading the result
queue drained queued, running and active leases all zero at the end,
unless a resident is intentionally kept warm
memory a flat or sawtooth profile is healthy; monotonic growth is investigated
row growth bounded by retention and aggregation, not linear in request count
restart counts unchanged, unless a restart was deliberately injected
Honesty rules
Report the duration actually achieved. A 45-minute soak is a 45-minute soak; it is not evidence about a week.
If a chaos scenario or an infrastructure change overlapped the run, say so and say which numbers it affected, rather than presenting the mixed result as a clean baseline. The M16 evidence includes one soak that was discarded for exactly this reason after container hardening restarted its datastores mid-run.
The disposable gateway client the soak uses is revoked afterwards, and the fixtures it created are removed.
Measuring latency honestly
The soak reports percentiles, so the harness must first prove that it is not the thing being
measured. Before any worker starts it times a plain TCP connect to the target and refuses to run if
the median exceeds MAX_CONNECT_FLOOR_MS (100 ms), printing the floor it measured and recording it
in the report as client_connect_floor_ms.
This exists because the first M16 soak reported a p50 near two seconds on every route including
/api/v1/health/live, with only a few milliseconds of spread. A constant with no tail is never queueing.
The cause was the harness addressing the API as localhost, which resolves to ::1 first on this
host; each request waited out the IPv6 connect timeout before falling back to IPv4, adding a fixed
~2,045 ms of client cost to every sample. Direct measurement, both returning 200: 2,043.3 ms median
to http://localhost:8000/api/v1/health/live against 4.7 ms to
http://127.0.0.1:8000/api/v1/health/live.
Address the platform by an address that resolves directly. Both M16 harnesses default to
http://127.0.0.1:8000 for this reason, not to localhost.