115 lines
4.1 KiB
Markdown
115 lines
4.1 KiB
Markdown
# API Blueprint
|
|
|
|
## M5 inference boundary
|
|
|
|
- `POST /api/v1/capabilities/rag.embedding@1/invoke` — native capability response.
|
|
- `POST /v1/embeddings` — documented OpenAI embedding subset using alias `rag.embedding`.
|
|
- `GET /api/v1/capability-deployments` and `GET /api/v1/scheduler` — public operational state.
|
|
- `/api/v1/admin/service-clients`, production approval/promotion, request history and lifecycle
|
|
actions require the operator credential.
|
|
- `/api/v1/agent/serving-*` is node-credential-only and never project-facing.
|
|
|
|
All inference failures use typed codes. `X-Correlation-ID` and the response request ID support
|
|
end-to-end tracing without retaining request content.
|
|
|
|
This document defines intended API surfaces, not frozen implementation details.
|
|
|
|
## Control-plane API
|
|
|
|
Operator/UI-facing resource surfaces:
|
|
|
|
```text
|
|
GET /api/v1/system
|
|
GET /api/v1/health/live
|
|
GET /api/v1/health/ready
|
|
GET /api/v1/hardware/nodes
|
|
GET /api/v1/hardware
|
|
GET /api/v1/hardware/nodes/{id}
|
|
GET /api/v1/hardware/accelerators
|
|
GET /api/v1/hardware/accelerators/{id}
|
|
POST /api/v1/hardware/refresh
|
|
POST /api/v1/admin/node-enrollments
|
|
GET /api/v1/admin/node-enrollments
|
|
DELETE /api/v1/admin/node-enrollments/{id}
|
|
PATCH /api/v1/admin/hardware/nodes/{id}
|
|
DELETE /api/v1/admin/hardware/nodes/{id}/credential
|
|
POST /api/v1/admin/hardware/nodes/{id}/credential/rotate
|
|
POST /api/v1/admin/hardware/nodes/{id}/decommission/preview
|
|
POST /api/v1/admin/hardware/nodes/{id}/decommission
|
|
GET /api/v1/models
|
|
GET /api/v1/models/{id}
|
|
GET /api/v1/models/{id}/revisions
|
|
POST /api/v1/models/discover
|
|
POST /api/v1/models/{id}/download
|
|
GET /api/v1/artifacts
|
|
GET /api/v1/deployments
|
|
POST /api/v1/deployments
|
|
POST /api/v1/deployments/{id}/validate
|
|
POST /api/v1/deployments/{id}/promote
|
|
POST /api/v1/deployments/{id}/rollback
|
|
GET /api/v1/capabilities
|
|
GET /api/v1/projects
|
|
GET /api/v1/projects/{id}/bindings
|
|
GET /api/v1/benchmarks/suites
|
|
POST /api/v1/benchmarks/runs
|
|
GET /api/v1/recommendations
|
|
GET /api/v1/jobs
|
|
GET /api/v1/security/findings
|
|
GET /api/v1/audit/events
|
|
```
|
|
|
|
Remote compute-node agent surface (protocol v1):
|
|
|
|
```text
|
|
POST /api/v1/agent/enroll
|
|
POST /api/v1/agent/heartbeat
|
|
PUT /api/v1/agent/inventory
|
|
PUT /api/v1/agent/telemetry
|
|
```
|
|
|
|
Agent endpoints accept no arbitrary command or runtime-launch payload. See `NODE_AGENT_PROTOCOL.md`.
|
|
|
|
All lifecycle-changing operations should become asynchronous jobs once work can exceed a normal request duration.
|
|
|
|
## Project-facing capability gateway
|
|
|
|
Native surface:
|
|
|
|
```text
|
|
POST /api/v1/capabilities/{capability-key}
|
|
```
|
|
|
|
Project identity must be authenticated. Capability version/channel may be selected only within policy.
|
|
|
|
## Compatibility facades
|
|
|
|
Where semantically safe, provide OpenAI-compatible or other standard interfaces so existing projects can redirect their base URL to ModelForge without coupling to a runtime worker.
|
|
|
|
A compatibility facade must still route through project authentication, capability resolution, scheduler policy, telemetry and audit.
|
|
|
|
## Error model
|
|
|
|
Normalized errors should identify the layer without exposing secrets:
|
|
|
|
- `capability_unavailable`
|
|
- `deployment_unhealthy`
|
|
- `scheduler_capacity_exhausted`
|
|
- `request_timeout`
|
|
- `invalid_capability_input`
|
|
- `policy_denied`
|
|
- `migration_required`
|
|
- `project_not_authorized`
|
|
|
|
Runtime-specific stack traces remain internal.
|
|
|
|
Request bodies are rejected before parsing with `request_body_too_large` (HTTP 413) when their
|
|
declared or streamed byte boundary is exceeded. The receiver also allows at most 32 consecutive
|
|
empty request events, 128 empty events in total and 4096 body events; exceeding any event/progress
|
|
budget returns exactly one `request_body_progress_exhausted` error (HTTP 400) and stops receiving.
|
|
Finite empty frames below those limits, normal chunks and disconnect events retain normal ASGI
|
|
semantics.
|
|
|
|
Node-decommission conflicts use `node_decommission_blocked`, `decommission_preview_stale`,
|
|
`node_generation_conflict` or `decommission_confirmation_mismatch`, with blocker details in the
|
|
normal correlation-ID error envelope. There is no force parameter.
|