4.1 KiB
API Blueprint
M5 inference boundary
POST /api/v1/capabilities/rag.embedding@1/invoke— native capability response.POST /v1/embeddings— documented OpenAI embedding subset using aliasrag.embedding.GET /api/v1/capability-deploymentsandGET /api/v1/scheduler— public operational state./api/v1/admin/service-clients, production approval/promotion, request history and lifecycle actions require the operator credential./api/v1/agent/serving-*is node-credential-only and never project-facing.
All inference failures use typed codes. X-Correlation-ID and the response request ID support
end-to-end tracing without retaining request content.
This document defines intended API surfaces, not frozen implementation details.
Control-plane API
Operator/UI-facing resource surfaces:
GET /api/v1/system
GET /api/v1/health/live
GET /api/v1/health/ready
GET /api/v1/hardware/nodes
GET /api/v1/hardware
GET /api/v1/hardware/nodes/{id}
GET /api/v1/hardware/accelerators
GET /api/v1/hardware/accelerators/{id}
POST /api/v1/hardware/refresh
POST /api/v1/admin/node-enrollments
GET /api/v1/admin/node-enrollments
DELETE /api/v1/admin/node-enrollments/{id}
PATCH /api/v1/admin/hardware/nodes/{id}
DELETE /api/v1/admin/hardware/nodes/{id}/credential
POST /api/v1/admin/hardware/nodes/{id}/credential/rotate
POST /api/v1/admin/hardware/nodes/{id}/decommission/preview
POST /api/v1/admin/hardware/nodes/{id}/decommission
GET /api/v1/models
GET /api/v1/models/{id}
GET /api/v1/models/{id}/revisions
POST /api/v1/models/discover
POST /api/v1/models/{id}/download
GET /api/v1/artifacts
GET /api/v1/deployments
POST /api/v1/deployments
POST /api/v1/deployments/{id}/validate
POST /api/v1/deployments/{id}/promote
POST /api/v1/deployments/{id}/rollback
GET /api/v1/capabilities
GET /api/v1/projects
GET /api/v1/projects/{id}/bindings
GET /api/v1/benchmarks/suites
POST /api/v1/benchmarks/runs
GET /api/v1/recommendations
GET /api/v1/jobs
GET /api/v1/security/findings
GET /api/v1/audit/events
Remote compute-node agent surface (protocol v1):
POST /api/v1/agent/enroll
POST /api/v1/agent/heartbeat
PUT /api/v1/agent/inventory
PUT /api/v1/agent/telemetry
Agent endpoints accept no arbitrary command or runtime-launch payload. See NODE_AGENT_PROTOCOL.md.
All lifecycle-changing operations should become asynchronous jobs once work can exceed a normal request duration.
Project-facing capability gateway
Native surface:
POST /api/v1/capabilities/{capability-key}
Project identity must be authenticated. Capability version/channel may be selected only within policy.
Compatibility facades
Where semantically safe, provide OpenAI-compatible or other standard interfaces so existing projects can redirect their base URL to ModelForge without coupling to a runtime worker.
A compatibility facade must still route through project authentication, capability resolution, scheduler policy, telemetry and audit.
Error model
Normalized errors should identify the layer without exposing secrets:
capability_unavailabledeployment_unhealthyscheduler_capacity_exhaustedrequest_timeoutinvalid_capability_inputpolicy_deniedmigration_requiredproject_not_authorized
Runtime-specific stack traces remain internal.
Request bodies are rejected before parsing with request_body_too_large (HTTP 413) when their
declared or streamed byte boundary is exceeded. The receiver also allows at most 32 consecutive
empty request events, 128 empty events in total and 4096 body events; exceeding any event/progress
budget returns exactly one request_body_progress_exhausted error (HTTP 400) and stops receiving.
Finite empty frames below those limits, normal chunks and disconnect events retain normal ASGI
semantics.
Node-decommission conflicts use node_decommission_blocked, decommission_preview_stale,
node_generation_conflict or decommission_confirmation_mismatch, with blocker details in the
normal correlation-ID error envelope. There is no force parameter.