Files

4.7 KiB

Soak testing

What a soak is for

A short run proves a path works. A soak proves it keeps working: that memory does not climb, that leases are released, that queues drain, that retention keeps row growth bounded, and that latency does not degrade as caches warm and evict.

python scripts/m16_soak.py --database-url "$MODELFORGE_RUNTIME_DATABASE_URL" --minutes 45 --secret <gateway-secret> --report soak.json

The database URL is mandatory and must use the non-owner runtime role. The harness contains no embedded credential or owner-role fallback.

Workload shape

Three gateway workers drive the real capability path — rag.embedding@1 through the authenticated gateway to GPU Node — mixing three phases:

burst    3-6 requests back to back, 50-400 ms apart
steady   single requests, 1.5-5 s apart
idle     8-25 s of nothing

The mix matters. A soak that only hammers never exercises keep-warm and unload, so it never observes a cold start. A soak that only idles never exercises admission control, so it never observes a capacity rejection. Batch sizes vary from one to four inputs so payload handling varies too.

Two read workers poll registry, projects, hardware, health, operations, alerts, lifecycle, migrations and recovery. One background worker drives the operational work the platform already performs for itself — capacity collection, SLO evaluation, alert evaluation — at a bounded rate.

What is measured

per kind          requests, succeeded, capacity rejected, internal errors, status distribution
latency           p50, p95, p99, mean, max
gateway           cold starts, warm invocations, queue p95, inference p50 and p95, failure codes
resources         container memory, CPU, PIDs and restart counts, before and after
database          row growth per table, connection count
runtime           active leases, serving jobs in flight, residency allocations
GPU               used, free and utilisation per accelerator, plus scheduler pressure
invariants        the full set, at every progress interval

Capacity rejections are counted separately from internal errors. A 429 or a QUEUE_FULL is the scheduler refusing unsafe work, which is the platform behaving correctly; counting it as a failure would reward a system that over-admits.

Stopping early

The soak aborts and reports partial evidence if an invariant is violated at any progress interval. Partial evidence honestly labelled is worth more than a full run whose result is already known to be invalid.

It should also be aborted manually on uncontrolled host pressure, degradation of an external workload, repeated OOM, data corruption or any unsafe production mutation.

Reading the result

queue drained      queued, running and active leases all zero at the end,
                   unless a resident is intentionally kept warm
memory             a flat or sawtooth profile is healthy; monotonic growth is investigated
row growth         bounded by retention and aggregation, not linear in request count
restart counts     unchanged, unless a restart was deliberately injected

Honesty rules

Report the duration actually achieved. A 45-minute soak is a 45-minute soak; it is not evidence about a week.

If a chaos scenario or an infrastructure change overlapped the run, say so and say which numbers it affected, rather than presenting the mixed result as a clean baseline. The M16 evidence includes one soak that was discarded for exactly this reason after container hardening restarted its datastores mid-run.

The disposable gateway client the soak uses is revoked afterwards, and the fixtures it created are removed.

Measuring latency honestly

The soak reports percentiles, so the harness must first prove that it is not the thing being measured. Before any worker starts it times a plain TCP connect to the target and refuses to run if the median exceeds MAX_CONNECT_FLOOR_MS (100 ms), printing the floor it measured and recording it in the report as client_connect_floor_ms.

This exists because the first M16 soak reported a p50 near two seconds on every route including /api/v1/health/live, with only a few milliseconds of spread. A constant with no tail is never queueing. The cause was the harness addressing the API as localhost, which resolves to ::1 first on this host; each request waited out the IPv6 connect timeout before falling back to IPv4, adding a fixed ~2,045 ms of client cost to every sample. Direct measurement, both returning 200: 2,043.3 ms median to http://localhost:8000/api/v1/health/live against 4.7 ms to http://127.0.0.1:8000/api/v1/health/live.

Address the platform by an address that resolves directly. Both M16 harnesses default to http://127.0.0.1:8000 for this reason, not to localhost.