# Project state ## Publication and Unraid deployment (2026-08-02) - Unraid deployment is live at `http://192.168.10.150:1236` from `/mnt/user/appdata/mobilityops`, Compose project `mobilityops`. - Deployment config commits: `07ab7a3`, `847cd05`, `e1a1c67`, `1e13943`. The accepted baseline `4bf9afbeff44088864e0844769d4dd0e4089d85b` remains intact. - All four services are healthy with zero restarts. Only web port 1236 is exposed to the LAN; n8n is loopback-only on host port 15678; API and PostgreSQL are internal. - Migrations are at `e7b08389f47f (head)` and deterministic seed counts match final acceptance. A Chrome smoke test covered every requested page and a real return; its n8n event succeeded on attempt 1. Browser console and recent service log scans were clean. - RAGcore is disabled in favor of the honest local demo provider. MCP Hub registration is disabled. n8n is initialized, published, healthy, and live-verified. - Local post-change gates: 66 backend tests, Ruff, mypy (44 files), and frontend production build all pass. Evidence is in `artifacts/deployment/unraid-summary.md`. - Published to the private Gitea repository `https://gitea.itworx.tech/Jens/MobilityOps`. `master` is the default branch; the full commit history and baseline commit are present; zero tags exist; remote hygiene is clean. `origin` uses the SSH clone URL supplied by Gitea. - Exact next action: none — repository publication and Unraid deployment are complete. ## Current milestone M7 — complete. All milestones (M0–M7) done, plus a full post-M7 final-acceptance audit (see below). See `artifacts/final-acceptance/summary.md` for the definitive acceptance evidence (supersedes `artifacts/evidence/final-summary.md`, which is kept as historical M7 evidence). ## Locked decisions - Product name: MobilityOps. - Fictitious tenant: Northstar Mobility Demo. - PoC only; all operational and knowledge data are synthetic. - Core stack and boundaries are defined in `CLAUDE.md` and `docs/03-architecture.md`. - RAGcore and ITWorx MCP Hub are external central services. - n8n receives post-commit events through an outbox dispatcher. - SQLAlchemy 2 declarative models cover the full domain model (`backend/app/models/`); enums are plain `String` columns validated at the Pydantic/service layer, not native PG enums (simpler migrations). - `backend/requirements.lock` is compiled inside a `python:3.12-slim` container (matches the Dockerfile base image) via `pip-compile --extra dev`; regenerate the same way if `pyproject.toml` changes. - Frontend dependencies pinned (no more `"latest"`); `package-lock.json` committed; Docker build uses `npm ci`. - Demo auth is a lightweight HMAC-signed cookie (`app/core/security.py`), not a real password/JWT flow — matches "Demo role buttons create an authenticated session; they do not bypass authorization middleware." Two fixed demo users (`USR-OPS` operations_manager, `USR-EMP` rental_employee) are created by the seed loader, not from a CSV (no `users.csv` in `seed/`). - Seed loader (`backend/app/seed_loader.py`) only supports `seed --reset` (always rebuilds); there is no incremental/idempotent-without-reset mode, since the acceptance criteria only require deterministic reset, not partial import. - `DataQualityIssue.entity_ref`/`related_ref` from the CSVs are resolved to `entity_type`/`entity_id` (UUID) at load time per the domain model; the original human-readable refs are kept in `evidence_json` (`entity_ref`, `related_refs`) since the API and UI need them and re-resolving UUID→public_ref on every read would be wasteful. - `backend/app/core/config.py` added `app_secret`, `session_cookie_name`, `session_ttl_seconds`, `seed_dir` (`/app/seed` in-container), `cors_allow_origins` (comma-separated string, not a list — simpler with pydantic-settings env parsing), `demo_today` (drives the dashboard's "Today" section against the deterministic anchor date, default `2026-08-01`). - `compose.yaml` api build context changed from `./backend` to repo root with `dockerfile: backend/Dockerfile`, so the image can `COPY seed ./seed` (seed CSVs are outside `backend/`). - Frontend: added `react-router-dom@7.18.2` (bumped from 6.x to clear two real advisories — open redirect + arbitrary constructor injection in v6). One residual `npm audit` finding (RSC-mode CSRF, GHSA-qwww-vcr4-c8h2) does not apply — this SPA never uses React Router's RSC/SSR mode. - Nav/pages built so far: Dashboard, Vehicles (list+detail with tabs), Bookings (list+detail), Audit. Data Quality, Knowledge and Automation nav items are intentionally omitted until M3/M5/M4 build the pages behind them — CLAUDE.md forbids dead routes/placeholders. - Return workflow (`app/services/returns.py`): the spec's "validate submitted reading against booking start reading" step was dropped as a hard rejection. For the seeded S1 scenario, a booking's `start_odometer_km` can already equal the vehicle's canonical odometer, so any regression-testing value would also be below the booking start, making a hard floor there indistinguishable from — and in conflict with — the documented soft-regression path. Only one odometer check exists now: submitted vs. the vehicle's *canonical* odometer (`vehicle.odometer_km`), matching the domain-model invariant verbatim ("a return with a lower submitted reading is recorded as an inspection and issue, while canonical odometer remains unchanged"). - Idempotency: new `idempotency_records` table (migration `e7b08389f47f`), unique on `idempotency_key`, keyed to `booking_id`. Same key + same booking replays the stored response; same key + different booking → 409 `IDEMPOTENCY_KEY_REUSED`; different key on an already-returned booking → 409 `INVALID_BOOKING_STATE`. Concurrency is enforced by `SELECT ... FOR UPDATE` on the booking row (re-checked for the idempotency record immediately after acquiring the lock, as a safety net for two simultaneous identical-key requests racing the pre-lock check). - `seed_loader.clear_all()` must delete `idempotency_records` before `bookings` (FK) — easy to forget when adding new booking-referencing tables; the ordering list at the top of `seed_loader.py` is the single place to update. - Inspection `public_ref` is assigned as `INSP-{count+1:04d}` from a live count query (not gap-safe, fine for a PoC single-writer demo, would need a sequence for real concurrency-safe numbering). - Found and fixed during browser verification (not caught by pytest, since it's a UI-only defect): `ReturnForm` originally held its own `result` state and was conditionally rendered only when `booking.status === "active"`; once the return succeeded the booking flipped to `returned` and React unmounted the form before the user ever saw the result panel. Fixed by lifting the result into `BookingDetail` (`ReturnResultPanel` is now a sibling, not nested in `ReturnForm`). Also found: `OutboxEvent.event_id`'s Python-side `default=uuid.uuid4` on the mapped_column only applies at flush/commit time, so reading `event.event_id` before `db.commit()` returned `None` (rendered as the literal string "None" in the result panel); fixed by assigning `event_id=uuid.uuid4()` explicitly at construction. Lesson: SQLAlchemy column `default=` callables are not available on the in-memory Python object until flush — never rely on the generated value for a same-transaction response body without an explicit `db.flush()` or an explicit Python-side assignment. - Operational note for this environment: `docker compose run --rm api ...` (used for tests/lint) only starts a throwaway one-off container — it does **not** update the long-running `api`/`web` service containers. After any code change meant to be verified live (browser, curl), `docker compose up -d --build ` is required, not just `docker compose build`. ## Completed evidence ### M0 — Reproducible foundation - Added `backend/app/core/db.py` (engine/session), `backend/app/models/*` (User, Customer, Vehicle, Booking, Inspection, MaintenanceRecord, DataQualityIssue, OutboxEvent, AuditEvent), Alembic config (`backend/alembic.ini`, `backend/alembic/env.py`) and initial migration `backend/alembic/versions/c9498525abb5_initial_schema.py`. - Commands run and verified from this checkout: - `docker compose build api` — OK - `docker compose run --rm api alembic upgrade head` — applied cleanly to empty DB, created 9 tables + `alembic_version`. - `docker compose run --rm api pytest -q` — 1 passed. - `docker compose run --rm api ruff check .` — All checks passed (added `extend-exclude = ["alembic/versions"]` to `backend/pyproject.toml` for autogenerated migration line length). - `docker compose up -d --build` — all 4 services healthy: `curl http://localhost:8128/health` → `{"status":"ok",...}`; `curl -o /dev/null -w "%{http_code}" http://localhost:1228/` → 200; `curl http://localhost:5678/healthz` → 200. - Fixed a real scaffold bug: `frontend/src/App.tsx` used `import.meta.env` without a `vite/client` types reference, which broke `npm run build` in Docker (works fine under plain `vite dev` because Vite injects the global at dev-time but `tsc -b` still type-checks it). Added `frontend/src/vite-env.d.ts`. - `make` is not installed in this Windows/git-bash shell — validated the underlying `docker compose ...` commands directly instead (Makefile targets are thin wrappers around them and are correct as written for a Linux/CI shell or WSL). - Known accepted gap: `npm audit` reports 1 moderate/1 high transitive `esbuild` advisory (dev-server-only, fixed only by a Vite 8 major bump); left as-is for the PoC, noted here rather than silently upgrading a major version. ### M1 — Operational core - Backend additions: `app/core/security.py` (HMAC-signed session cookies), `app/api/deps.py` (`get_current_user`, `require_operations_manager`), `app/core/errors.py` (`AppError` + the documented `{"error": {...}}` shape wired as a FastAPI exception handler for both `AppError` and `HTTPException`), `app/seed_loader.py`, `app/cli.py` (`python -m app.cli seed --reset`), `app/services/audit.py`, `app/schemas.py`, routers under `app/api/routers/` (`demo`, `dashboard`, `vehicles`, `bookings`, `audit`). - Frontend additions: React Router-based app shell (`src/App.tsx`, `src/components/Layout.tsx`, `src/components/RequireAuth.tsx`), `AuthContext`, typed `api` client (`src/api/client.ts`, `src/api/types.ts`), pages `Login`, `Dashboard`, `Vehicles`/`VehicleDetail`, `Bookings`/`BookingDetail`, `Audit`. Full responsive stylesheet (`src/styles.css`) covering nav collapse and table→card layout under 700px, visible focus states, no hover-only actions. - Commands run and verified from this checkout (container rebuilt each time to pick up code changes): - `docker compose run --rm api pytest -q` — **19 passed** (new: `test_seed.py`, `test_auth.py`, `test_dashboard.py`, `test_vehicles.py`, `test_bookings.py`, `test_audit.py`; tests seed the real Postgres via `reset_and_seed` in a session fixture, then exercise the FastAPI app through `TestClient`, not mocks). - `docker compose run --rm api ruff check .` — All checks passed (added `ignore = ["B008"]` — FastAPI's `Depends()`-as-default is idiomatic, not a real bug). - `npm run build` (local, Node 24) — clean `tsc -b && vite build`. - `docker compose up -d --build` then `docker compose exec api python -m app.cli seed --reset` — counts: `users:2 customers:180 vehicles:50 bookings:246 inspections:75 maintenance:40 data_quality_issues:15 workflow_runs:20`. - `curl` end-to-end: `POST /api/v1/demo/login` sets cookie and returns the user; unauthenticated `GET /api/v1/dashboard` → 401 with the documented error shape; authenticated dashboard/vehicle-detail return real seeded data (verified metrics `available:21 rented:11 cleaning:6 maintenance:5 blocked:7`, matching the 50 seeded vehicles). - Browser smoke test (Chrome via MCP) at desktop width: login page → Operations Manager login → Dashboard (metrics + attention items + today + recent automation all populated) → Vehicle detail `MO-016` (tabs render, "Needs attention" badge correct — it's `DQ-DEMO-OVERLAP`/`DQ-DEMO-STATUS`) → Booking detail `BK-DEMO-RETURN` (matches S1 scenario: vehicle `MO-024`, status `active`, start odometer `53610`). Responsive CSS (`@media max-width:700px`) was written and code-reviewed but the automated resize during this session didn't visibly reflect in the captured screenshot (likely a screenshot-timing quirk of the browser tool, not necessarily a real bug) — **treat the ≤360px layout as visually unverified** and re-check with a real device/DevTools emulation before final acceptance (M7). - Known accepted gap carried over from M0: `npm audit` residual `esbuild`/Vite-8 dev-server-only advisory. ### M2 — Vehicle return vertical slice - Backend additions: `app/models/idempotency.py` (`IdempotencyRecord`), migration `e7b08389f47f_idempotency_records`, `app/services/returns.py` (`register_vehicle_return` — full transaction: row locks, idempotency replay, inspection, canonical-odometer update or regression issue, vehicle status derivation, two audit events, `vehicle.returned.v1` outbox event matching `contracts/events.schema.json`, next-booking-risk lookup), `POST /api/v1/bookings/{public_ref}/return` wired in `app/api/routers/bookings.py` with required `Idempotency-Key` header. - Frontend additions: `components/ReturnForm.tsx` (form + `ReturnResultPanel`), wired into `pages/BookingDetail.tsx` (shown only when `booking.status === "active"`; result persists via lifted state after the booking flips to `returned`). - Commands run and verified from this checkout: - `docker compose run --rm api pytest -q` — **26 passed**, including `tests/test_return.py` (success/canonical-update, S1 regression scenario by name, damage→blocked, idempotent replay, reject-already-returned, missing-header validation, and a real multi-threaded concurrent-submission test against Postgres asserting exactly 1×201 + 2×409). - `docker compose run --rm api ruff check .` — All checks passed. - `npm run build` — clean. - `docker compose up -d --build` (all services) then `docker compose exec api python -m app.cli seed --reset`, then a full browser run of the S1 demo scenario against `BK-DEMO-RETURN`/`MO-024`: submitted 53000 km (below canonical 54820) → result panel showed `INSP-0076`, `resulting_vehicle_status: maintenance` (correctly derived, since canonical 54820 ≥ `next_service_km` 40000), `DQ-RET-0076` created, automation event queued with a real UUID, "no upcoming booking" risk; vehicle detail page confirmed odometer unchanged at 54,820 km and a "Needs attention" badge. - Both real defects listed above (form disappearing before showing its result; `event_id` reading as `None`) were **found via the browser run, not by pytest** — the test suite asserted on API response shape/values, not on what the UI actually rendered after a status transition. Worth remembering for M3+: UI state-after-mutation bugs need a browser check, not just API tests. ### M3 — Data Quality Workbench - `app/services/data_quality.py`: `run_scan()` implements all five rules and is called automatically at the end of `seed_loader.reset_and_seed()` (after `db.commit()` of the base seed), plus exposed as `POST /api/v1/data-quality/scan` (Operations Manager only). Idempotency is simplified from the doc's literal `(rule_type, entity_type, entity_id, evidence fingerprint)` to just `(rule_type, entity_type, entity_id)` while an issue is open — see rationale below. - Router `app/api/routers/data_quality.py`: `GET /issues` (filters status/rule_type/severity), `GET /issues/{ref}` (adds `entity_snapshot`/`related_snapshots` for the UI), `POST /issues/{ref}/defer`, `/reject`, `/merge-customers` (Operations Manager only — enforced via `require_operations_manager`), `POST /scan`. - **Real bug found and fixed during this milestone, before any browser check**: the first cut of DQ-03 (odometer regression) compared every historical *returned* booking's `end_odometer_km` against the vehicle's *current* `odometer_km`. Since the seed generator assigns `vehicle.odometer_km` independently of booking history (see `seed/generate_seed.py`), this is true for nearly every historical booking by construction (odometer is monotonically increasing over time, so all-but-the-latest reading is "below current") — it produced 51 false-positive issues out of 50 vehicles on first run. Fixed twice: first attempt (compare only the single most-recent booking against canonical) still produced the same problem because canonical itself is disconnected from booking history in this dataset; the working fix compares each vehicle's *own returned-booking sequence* against itself (each booking's end reading vs. the immediately preceding one, chronologically) — a self-consistency check that doesn't depend on the unrelated `vehicle.odometer_km` field at all. Final deterministic seed+scan totals: 15 CSV-seeded + 11 scan-discovered = **26** open/resolved `data_quality_issues` (breakdown: 14 vehicle_status_conflict, 5 missing_required_field, 3 possible_duplicate_customer, 3 odometer_regression, 1 booking_overlap). `tests/test_seed.py`'s exact-count assertion was updated from 15 to 26 accordingly — if the scan logic changes again, update that count. - Idempotency simplification rationale: the doc's fingerprint-based key would make the scan blind to issues it structurally can't compute a matching fingerprint for against the CSV-seeded rows (which don't carry a fingerprint field), producing duplicate issues for the same real-world problem (e.g. a second `MO-016` overlap issue next to the seeded `DQ-DEMO-OVERLAP`). Using `(rule_type, entity_type, entity_id)` alone while open is a stricter, safe simplification: it can never falsely suppress an issue for a *different* entity, and per-entity there's realistically only one meaningful open issue of a given rule type at a time for this PoC's scope. - Merge UI intentionally does **not** use `window.confirm()` — a native dialog blocks further automation/testing and isn't screen-reader-distinguishable from page content the same way a rendered `role="alertdialog"` panel is. Built an inline two-step confirm instead (`ReturnForm`-style pattern reused). - `AttentionItem` gained an `issue_ref` field (dashboard now links attention items straight to `/data-quality/{issue_ref}` instead of only to vehicles); dashboard attention list capped at 8 items (was unbounded, would have shown up to 26 with the richer scan). - Commands run and verified from this checkout: - `docker compose run --rm api pytest -q` — **35 passed** (new `tests/test_data_quality.py`: all five rule types present, scan idempotent on rerun, scan requires Operations Manager, S2/S4 issue-detail snapshots correct, defer→reject-on-closed 409, merge requires Operations Manager, merge rejects an unrelated survivor ref, full S2 merge scenario asserting rewiring + audit + replay-is-409). - `docker compose run --rm api ruff check .` — All checks passed. - `npm run build` — clean (had to fix two `possibly 'null'` TS errors from a closure-narrowing limitation — TS doesn't narrow `const` captured-by-closure across nested function boundaries when the value comes from an index/property expression; fixed by re-binding to explicitly-typed local consts right after the guard). - Full browser run: Data Quality list (26 open issues, filterable) → `DQ-DEMO-DUPLICATE` two-column compare (CUS-0012 vs CUS-0178, per-field diff highlighting only where they differ) → merge with inline confirm → issue flips to `resolved` → confirmed `customer_merged` audit event with correct actor/entity/correlation → `DQ-DEMO-OVERLAP` (non-duplicate type) renders evidence JSON + defer/reject, no dead compare UI shown for a rule type it doesn't apply to. ### M4 — n8n automation - `app/services/dispatcher.py`: background daemon thread (started/stopped via FastAPI `lifespan`, not an `on_event` hook) polling every `N8N_DISPATCH_INTERVAL_SECONDS` (default 3s). Claim step (`_claim_due_events`) is a short transaction using `SELECT ... FOR UPDATE SKIP LOCKED` that only flips `pending`→`delivering` and commits immediately; the HTTP call to n8n happens with **no open transaction**; the outcome is recorded in a separate short transaction. Exponential backoff `min(2**attempts, 60)` seconds, `N8N_MAX_ATTEMPTS=5` before a permanent `failed`. - Dispatcher reconstructs the wire event from `contracts/events.schema.json`'s exact fields (`event_id`, `event_type`, `occurred_at`, `correlation_id`, `aggregate`, `data`) rather than forwarding `OutboxEvent.payload_json` wholesale — that column also carries an internal `aggregate_ref` convenience key (used by dashboard/workflows list rendering) that the schema's `additionalProperties: false` would reject. - `POST /api/v1/integrations/n8n/return-callback` (`app/api/routers/integrations.py`): shared-secret auth via `X-Service-Token` header (`N8N_CALLBACK_TOKEN`, propagated to both `api` and `n8n` containers as `MOBILITYOPS_CALLBACK_TOKEN`); idempotent by `Idempotency-Key` (the event UUID) — checked by querying for an existing `AuditEvent` with that event ID in its metadata, **not** by `OutboxEvent.external_run_id`, because the dispatcher only sets that field *after* it gets n8n's final response, which happens *after* n8n has already called this callback mid-workflow — using `external_run_id` as the idempotency guard would have missed the exact redelivery case it's meant to catch. - `GET /api/v1/workflows` + `POST /api/v1/workflows/{event_id}/retry` (`app/api/routers/workflows.py`), both Operations Manager only. Retry only allowed from `failed`; sets `pending` + clears `next_attempt_at` so the live dispatcher picks it up on its next cycle (does not reset `attempts`, so the counter reflects true delivery history). - Automation nav + page (`pages/Automation.tsx`): table of all runs with status/attempts/last error, Retry button for `failed` rows, visible only to Operations Manager (matches backend authorization rather than just hiding a link). - **Two real bugs found and fixed, the second only by testing the actual live n8n round-trip, not by pytest**: 1. Seed-loaded `workflow_runs.csv` rows only ever got `payload_json = {"aggregate_ref": ...}` (no `correlation_id`/`aggregate`/`data`) — fine for M1–M3 since nothing read those keys yet, but once the dispatcher tried to *redeliver* a seeded row (i.e. the S5 manual-retry demo scenario) it crashed with `KeyError: 'correlation_id'`, leaving that event stuck in `delivering` forever (the crash happened before the outcome-recording transaction). Fixed in two places: `seed_loader.py` now builds the full schema-compliant envelope for every `workflow_runs.csv` row (matching what the live M2 return flow produces), and `dispatcher._deliver_one` now catches malformed-payload `KeyError`s defensively and resolves the row to `pending`/`failed` instead of leaving it orphaned — added `test_deliver_one_handles_malformed_payload_without_getting_stuck` as a regression test for the latter. 2. This n8n image (2.32.7) has dropped `N8N_BASIC_AUTH_ACTIVE` as a UI/API gate — it requires an actual owner account via the `/setup` flow before anything (including webhook registration reliability) works correctly. Also: `n8n import:workflow` requires the workflow JSON to have a top-level `"id"` field (added `"id": "mobilityops-return-processing"`) and **always deactivates** the imported workflow regardless of its `"active"` field — activation requires `n8n publish:workflow --id=` followed by a full n8n restart (documented in n8n 2.x CLI, not obvious from the docs pack). Did this manually this session via the CLI + browser setup wizard; **this is a one-time operational step that is not automated** — a truly clean checkout still needs someone to run `docker compose exec n8n n8n import:workflow --input=//imports/mobilityops-return-processing.json`, `docker compose exec n8n n8n publish:workflow --id=mobilityops-return-processing`, `docker compose restart n8n`, and complete the one-time owner setup at `http://localhost:5678/setup` (any email/password, no verification required) before the automation demo will work. `docs/17-runbook.md` should get this exact sequence in M7. - Commands run and verified from this checkout: - `docker compose run --rm api pytest -q` — **49 passed** (new `tests/test_dispatcher.py` — claim/deliver success/failure/backoff/exhaustion-to-failed/malformed-payload, all via `monkeypatch.setattr(dispatcher.httpx, "post", ...)`, no real network calls in tests; `tests/test_integrations.py` — callback auth, unknown-event 404, idempotent-by-event-ID with a real duplicate-call assertion; `tests/test_workflows.py` — role gating, retry-only-from-failed, S5 retry-and-audit). - `docker compose run --rm api ruff check .` — All checks passed. - `npm run build` — clean. - Full live round trip (not mocked): registered a real return on `BK-DEMO-RETURN` → outbox event queued → background dispatcher delivered it to the now-activated n8n workflow within its 3s poll interval → n8n called back into `/api/v1/integrations/n8n/return-callback` (200 OK, confirmed in `docker compose logs api`) → dispatcher's original POST received n8n's success response → event flipped to `succeeded` on attempt 1, visible on `/automation`. - S5 scenario end-to-end in the browser: seeded `BK-H-0020` (`failed`, 3 attempts, "Synthetic connection timeout to n8n") → clicked Retry → `pending` → within ~3s, live dispatcher delivered it through the real n8n instance → `succeeded`, 4 attempts. This is the full documented S5 scenario working for real, not simulated. ### M5 — RAGcore knowledge integration - `app/services/knowledge/__init__.py`: `KnowledgeProvider` Protocol (sync, not async — the rest of the backend is sync SQLAlchemy/FastAPI, so an async provider interface would have meant bridging paradigms for no benefit) with `health()`/`ask()`, plus `GroundedAnswer`/`SourceCard`/`KnowledgeHealth` Pydantic models matching `contracts/openapi.yaml`'s `GroundedAnswer` schema exactly. `get_knowledge_provider()` factory switches on `settings.knowledge_provider` ("demo" default, "ragcore" opt-in). - `app/services/knowledge/demo.py` — `DemoKnowledgeProvider`: parses the 10 `knowledge/procedures/*.md` files' YAML frontmatter (hand-rolled flat parser, not PyYAML — avoided adding a dependency for a 6-key flat block) and `## `-delimited sections at startup, then does **TF-IDF-weighted keyword retrieval** (not naive keyword counting) with light suffix-stripping stemming (`returns`→`return`, `damaged`→`damage`). This is extractive, not generative: it returns real excerpts and a templated answer sentence, never invented text. - **Real bug found and fixed by testing the actual S6 question, not by inspection**: naive flat keyword-overlap scoring (first cut) let the word "vehicle" — present in nearly every document's title — crowd out the actually-relevant `damage-procedure` document from the top-3 results for "What must I do when a vehicle returns with damage?", because generic words scored the same as distinctive ones. Fixed by computing corpus-wide IDF per token (`log((N+1)/(df+1)) + 1`) and weighting matches by it, so common terms contribute little and rare/distinctive terms (like "damage") dominate the ranking. Verified: the S6 question now returns `damage-procedure` and `vehicle-return-procedure` in the top 3, matching the documented expectation exactly. - `app/services/knowledge/ragcore.py` — `RAGcoreKnowledgeProvider`: real `httpx` adapter guessing a plausible REST contract (`GET /health`, `POST /api/v1/ask`) per `contracts/ragcore-contract-assumptions.md` (RAGcore is built separately; no live instance was reachable this session to verify against). Any connection error, timeout, or malformed response degrades to `evidence_state: "unavailable"` rather than raising — this is the adapter that actually exercises the architecture's "RAGcore failure disables knowledge answers only" reliability boundary. Not wired as the active provider by default; `KNOWLEDGE_PROVIDER=ragcore` would need a real, verified base URL to turn on. - `POST /api/v1/knowledge/questions` + `GET /api/v1/knowledge/status` (`app/api/routers/knowledge.py`). Audit event `knowledge_question_asked` logs `evidence_state`, `provider`, `source_ids`, and `question_length` only — **not** the question text itself, per `docs/12-security-and-audit.md` ("log question metadata and source IDs, not unnecessary full prompts"). - Knowledge nav + page (`pages/Knowledge.tsx`): chat-style question box, source cards (title/version/section/excerpt) prioritized over the answer text per `docs/06-ui-ux.md`, explicit `grounded`/`insufficient`/`unavailable` states with distinct visual treatment — never a fabricated-looking answer for the latter two. - Dockerfile now also `COPY knowledge ./knowledge`; added `KNOWLEDGE_DIR` setting (`/app/knowledge/procedures` in-container, same pattern as `SEED_DIR`) rather than deriving the path from `__file__` — simpler and doesn't break if the module moves. - Commands run and verified from this checkout: - `docker compose run --rm api pytest -q` — **57 passed** (new `tests/test_knowledge.py`: S6 grounded-with-expected-sources, unrelated question is honestly insufficient with no fabrication, demo provider health/document count, endpoint auth required, audit doesn't leak question text, RAGcore adapter degrades to unavailable on a simulated connection error). - `docker compose run --rm api ruff check .` — All checks passed. - `npm run build` — clean. - Full browser run of S6 end-to-end: asked "What must I do when a vehicle returns with damage?" on `/knowledge` → grounded answer citing "Vehicle return procedure" (2 sections) and "Damage handling procedure" with real excerpts. Also asked an unrelated question ("What is the weather forecast for tomorrow?") → correctly returned "Insufficient evidence" / "No matching procedure was found" with zero sources, confirming no fabrication. ### M6 — ITWorx MCP Hub publication - `app/api/routers/mcp_integrations.py`: four read-only endpoints under `/api/v1/integrations/mcp/` — `GET operations-summary`, `GET attention-vehicles` (query params `minimum_severity`/`date`/`limit` matching `contracts/mcp-tools.json`'s `inputSchema` exactly), `GET vehicles/{vehicle_ref}`, `POST search-knowledge` (the "narrow façade" the doc calls for — wraps M5's `get_knowledge_provider()` rather than re-implementing retrieval; the contract's tool has no MobilityOps `endpoint` field, only `routing.preferred: ragcore`, so this façade path is MobilityOps's own addition for when the Hub needs a single provider boundary, not literally specified by the contract). - Auth: new `require_mcp_service_token` dependency in `app/api/deps.py`, same shared-secret-header shape as the M4 n8n callback (`X-Service-Token` against `MCP_HUB_SERVICE_TOKEN`) plus an optional `X-Client-Id` header (defaults to `"unknown-mcp-client"`) used as the audit actor label — the Hub's actual client-identity header name is unknown (no live Hub to confirm against), so this is a reasonable guess documented here rather than assumed silently. - `McpVehicleDetailOut` deliberately omits `registration_number` and all customer data — narrower than the browser-facing `VehicleOut`/`VehicleDetailOut`, matching "no customer or vehicle database access" and the read-only/summary intent of an AI-facing tool. Test `test_vehicle_details_known_ref` asserts the field's absence explicitly so a future change can't silently widen the exposed surface. - Extracted `app/services/operations.py` (`compute_metrics`, `list_attention_vehicles`) out of `app/api/routers/dashboard.py` so the MCP operations-summary/attention-vehicles endpoints and the human dashboard share one query implementation instead of two copies that could drift — the same "do not duplicate retrieval logic" principle the doc states for the knowledge tool, applied here to the operational-summary tools too. - Every provider call writes an `AuditEvent` (`actor_type="service"`, `actor_label=X-Client-Id`, `action="mcp_tool_request"`, `metadata={tool, status}`) — MobilityOps's own record that its provider APIs were reached, independent of whatever central tool-call audit the Hub itself keeps (per `docs/10-mcp-hub-integration.md`'s audit section, the Hub owns the central log; this is the local corroborating one). - No write/mutation endpoints exist under the `/api/v1/integrations/mcp/` namespace at all (verified by `test_no_write_endpoints_exist_under_mcp_namespace` — POST/PUT/DELETE against the vehicle-details path all 404/405) — return registration, customer merge, and any booking/vehicle mutation are correctly absent, per the doc's explicit restriction list. - Commands run and verified from this checkout: - `docker compose run --rm api pytest -q` — **66 passed** (new `tests/test_mcp_integrations.py`: token-required, wrong-token 401, all four tools' happy paths, severity/limit filtering, 404 for unknown vehicle, `max_sources` respected, audit actor/action verified, write-method rejection). - `docker compose run --rm api ruff check .` — All checks passed. - Live `curl` verification against the running stack (no browser needed — these are service-to-service endpoints, not UI): missing header → 422; wrong token → 401; correct token → all four endpoints return correct data (`operations-summary` metrics match the dashboard; `attention-vehicles?minimum_severity=high` returned `MO-016`×2 and `MO-031`, all severity `high`; `vehicles/MO-016` returned the narrow read-only shape; `search-knowledge` with `max_sources=2` returned exactly 2 grounded sources for the S6 question). Confirmed via `GET /api/v1/audit?action=mcp_tool_request` that all four calls were recorded with correct `actor_type=service`, tool name, and status. ### M7 — Portfolio polish and final acceptance - **Automated clean-checkout migrations**: `backend/entrypoint.sh` now runs `alembic upgrade head` before starting uvicorn (Dockerfile `CMD` changed from `uvicorn ...` to `./entrypoint.sh`). Verified with a true `docker compose down -v` (all volumes wiped) → `docker compose up --build -d` → all 11 tables present, `/health` and web both green, all 66 backend tests pass, with zero manual migration step. - **n8n one-time setup scripted where it can be**: `make n8n-setup` runs the import/publish/restart sequence (previously three manual commands discovered ad hoc in M4). The owner-account creation itself cannot be scripted safely (it's an interactive one-time step in n8n 2.x's own onboarding, not a MobilityOps concern) — documented precisely in the rewritten `docs/17-runbook.md`, including the exact URL and that no email verification is required. Re-ran this full sequence from the wiped-volumes state this session and confirmed the S1 return → outbox → live n8n → callback → `succeeded` round trip works on a genuinely clean checkout, not just the already-provisioned stack from M0–M6. - **Playwright E2E** (`frontend/e2e/demo.spec.ts`, `frontend/playwright.config.ts`): one test automating the full 9-step documented demo script end-to-end against the live stack — login, dashboard metrics, open `BK-DEMO-RETURN`, register an odometer-regression return (S1), verify the quality issue + queued automation event, merge the duplicate-customer scenario (S2), ask the damage question and verify both expected source citations (S6), inspect audit entries, and verify responsive nav + no horizontal overflow at 360px width. **Passing.** This also resolves the "≤360px layout visually unverified" gap flagged back in M1 — verified both by this test's overflow assertion and by the `9-mobile-dashboard.png` screenshot (nav wraps into rows, metric tiles collapse to a 2-column grid, no horizontal scroll). - Added `frontend/e2e/_capture-screenshots.spec.ts` as evidence-generation tooling (underscore-prefixed, excluded from the default `playwright test` / `make e2e` run via `testIgnore` in the config — it calls `demo/reset`, which a real regression test shouldn't do as a side effect). Captured all 9 screenshots into `artifacts/evidence/screenshots/`. - Wrote `artifacts/evidence/architecture.md` (mermaid, as-built — distinguishes verified-live components from implemented-but-never-reached-a-real-instance ones, i.e. RAGcore and the MCP Hub) and `artifacts/evidence/final-summary.md` (commit, exact commands, test counts, screenshot index, RAGcore success/unavailable evidence — including a live-demonstrated unavailable case against an unreachable host, not just the unit test — n8n success/retry evidence, MCP sample calls, known limitations, truthful portfolio wording per `docs/16-portfolio-case-study.md`'s template). - Updated `README.md` (dropped stale "minimal bootable scaffold, not the finished application" wording and the old two-line quickstart in favor of `make demo` + a pointer to the runbook) and `docs/17-runbook.md` (full rewrite: exact bootstrap, n8n one-time setup, verification commands, required operational checks, recovery expectations). - Final placeholder/dead-UI sweep: `grep`'d the full `frontend/src` and `backend/app` trees for scaffold/TODO/FIXME/"must be replaced" markers — none found. `FILE_INDEX.md` was left as-is; it's the original build-pack's archive-completeness manifest (a historical snapshot), not a living index that needs to track every file added since — updating it would misrepresent what it's for. - Commands run and verified from this checkout (this milestone, cumulative across the whole build): - `docker compose run --rm api pytest -q` — **66 passed**, ruff clean. - `cd frontend && npm run build` — clean. - `cd frontend && npx playwright test` — **1 passed** (full demo script, live stack). - Full clean-checkout drill: `docker compose down -v` → `docker compose up --build -d` → `docker compose exec api python -m app.cli seed --reset` → `docker compose run --rm api pytest -q` (66 passed) → n8n owner setup + `make n8n-setup` → live S1 return round-tripped through the real n8n instance to `succeeded`. ### Final acceptance audit (post-M7) A dedicated release-readiness audit was run after M7 claimed completion, specifically to catch anything the milestone-by-milestone build might have missed by only ever validating each piece in isolation. - **Real gap found: `mypy` had never been run.** `mypy` is a declared dev dependency (`backend/pyproject.toml`) but was never wired into any milestone's validation loop — only `ruff` was. Running it cold surfaced **43 real type errors across 10 files**, all pre-existing (not introduced by this audit). Triaged and fixed all of them rather than suppressing: - `services/returns.py`: the vehicle lookup after acquiring `FOR UPDATE` could type as `Vehicle | None` with no runtime guard — added an explicit `if vehicle is None: raise AppError(..., 404)`. This was a genuine defensive-programming gap (a dangling FK would have crashed with an unhandled `AttributeError`/500 instead of a clean 404), not just a type annotation issue. - `api/routers/bookings.py`: same pattern for `db.get(Customer, ...)` / `db.get(Vehicle, ...)` in `get_booking` — added a guard raising 500 with a clear message instead of crashing on `None.public_ref`. - `api/deps.py` + `api/routers/demo.py`: `CurrentUser.role` is a `Literal[...]`, but `SessionPayload.role` (decoded from an HMAC-signed cookie) and `User.role` (a DB column) are both plain `str`. Pydantic validates this at runtime already (so it was never exploitable), but `get_current_user` now explicitly checks membership before constructing `CurrentUser`, turning a would-be unhandled `ValidationError` (500) into a clean 401 for a corrupted/tampered cookie — another real defensive improvement, not just a type-checker appeasement. - `api/routers/dashboard.py`, `api/routers/data_quality.py`: two instances of reusing one variable name for both a `Vehicle` and a `Customer` across an if/else branch, which is genuinely confusing to read regardless of what mypy thinks — renamed to distinct variables (`entity`/typed union in dashboard, `customer`/`vehicle` in the data-quality snapshot helper). - `services/data_quality.py`, `seed_loader.py`: `Booking.__table__.update()` / `Customer.__table__.update()` don't typecheck against SQLAlchemy 2.0's stubs (the `.__table__` accessor is typed as the more general `FromClause`, which doesn't declare `.update()`) — switched to the idiomatic `sqlalchemy.update(Model)` construct, which is both correctly typed and the more modern SQLAlchemy 2.0 style anyway. - Remaining handful (schemas.py's deprecated `conint()` → `Annotated[int, Field(...)]`, a `Sequence` vs `list` `.sort()` call, an `assert`-guarded None-narrowing after a `WHERE ... IS NOT NULL` filter mypy can't see through, `Result.rowcount` typing gaps) were either latent pydantic-v1-style API usage or genuine SQLAlchemy stub limitations — fixed with the idiomatic modern equivalent or a narrowly-scoped, commented `# type: ignore[...]` at the exact line, never a blanket suppression. - `make lint` now runs both `ruff check .` and `mypy app`; `mypy app` reports **zero errors across 44 source files**. - **No other defects found.** Re-ran the full journey matrix end-to-end against a genuinely wiped-volumes (`docker compose down -v`) clean checkout: all 66 backend tests, ruff, ✅; ran the demo login → dashboard → vehicle/booking detail → return workflow → invalid-mileage rejection (422, both a negative value and a non-numeric string) → data-quality issue review → duplicate-customer merge → audit trail → Knowledge Assistant → live n8n round trip → MCP Hub endpoint journeys directly via `curl` against the running stack, all correct. - **Degraded-mode behavior explicitly re-verified live** (not just unit-tested): stopped n8n with `docker compose stop n8n`, registered a return — it committed (`201`, booking flipped to `returned`) exactly as required; the outbox event stayed `pending` with real `ConnectError`s logged and exponential backoff (2 attempts over ~8s); restarted n8n and the dispatcher **self-healed** without any manual intervention, delivering the event to `succeeded` on attempt 5. RAGcore unavailable mode re-verified live against an unreachable host (`ConnectError` → `evidence_state: "unavailable"`, empty answer, no fabrication). MCP Hub unavailability is architecturally moot for MobilityOps — the Hub only ever calls *into* MobilityOps, so there is nothing on the MobilityOps side that can degrade if the Hub is down (only the reverse, "does an unavailable Hub break MobilityOps," which is trivially no since nothing here calls out to it). - **New test coverage added, no existing tests weakened**: `frontend/e2e/interactive-elements.spec.ts` (11 Playwright tests — all seven nav items, every filter on every list page, vehicle detail tabs, defer/reject, automation retry, knowledge form, role-switching, and role-based page restriction) plus the existing `demo.spec.ts` — **12/12 e2e tests passing** against the live stack. - Verified `.env` is `.gitignore`d and was never committed (`git ls-files` / `git log --all -p -- '*.env'` both empty); scanned full git history for AWS keys, private-key headers, and `sk-...`-style tokens — none found. Every `Settings` field in `backend/app/core/config.py` has a corresponding entry either directly in `.env.example` or is derived/wired through `compose.yaml` (a few purely-internal container-path constants like `SEED_DIR`/`KNOWLEDGE_DIR` are intentionally not operator-configurable and correctly absent from `.env.example`). - Grepped the full `frontend/src` and `backend/app` trees for TODO/FIXME/placeholder/ fake/stub/mock/"not implemented" markers — zero real hits (the two `placeholder=` matches are legitimate HTML input placeholder attributes). Confirmed dashboard metrics and all list-page data are 100% DB-backed (`compute_metrics` in `services/operations.py`, never a literal in frontend JSX). Confirmed every frontend route in `App.tsx` maps to an implemented page and every nav item maps to a real route — no dead routes. - Commands run and verified from this audit: - `docker compose run --rm api pytest -q` — **66 passed**. - `docker compose run --rm api ruff check .` — All checks passed. - `docker compose run --rm api mypy app` — **Success: no issues found in 44 source files** (0 errors, down from 43). - `cd frontend && npm run build` — clean (`tsc -b && vite build`). - `cd frontend && npx playwright test` — **12 passed** (`demo.spec.ts` + `interactive-elements.spec.ts`). - Full clean-checkout drill repeated from a fresh `docker compose down -v`: automatic migrations, seed, 66/66 tests, n8n owner setup + `make n8n-setup`, live return round-tripped through n8n to `succeeded`. - See `artifacts/final-acceptance/summary.md` for the complete evidence write-up (commands, exact outputs, demo access, deployment instructions, five-minute demo flow). ## Definition of done All eight milestones (M0–M7) are complete, and a dedicated post-M7 final-acceptance audit found and fixed one real category of gap (`mypy` never having been run) with zero regressions. `docs/14-testing-and-acceptance.md`'s clean-checkout acceptance list has been walked item by item against a genuinely wiped-volumes checkout, twice (once in M7, once in this audit), and `artifacts/final-acceptance/summary.md` is the authoritative final evidence document. The two items not fully closed — a live RAGcore instance and a live ITWorx MCP Hub instance — were never reachable in this environment; both integrations are implemented, unit/contract-tested, directly verified against MobilityOps's own API, and their unavailable-degradation paths are live-verified, but an actual round trip against real RAGcore/Hub instances remains unconfirmed and is documented as such rather than claimed. ## Known blockers None. External service credentials may be absent; use the documented demo/degraded providers. The n8n workflow-activation steps are a one-time manual setup requirement in this environment (owner-account creation via n8n's own `/setup` UI cannot be scripted safely), fully documented in `docs/17-runbook.md` and scripted where possible (`make n8n-setup`). RAGcore and the ITWorx MCP Hub itself were never reachable in this environment — both integrations are implemented and directly tested/curl-verified against MobilityOps's own API, but neither a real RAGcore instance nor a real Hub round trip was available to confirm end-to-end. ## Premium Control Rail UI transformation (2026-08-02) - Branch: `design/mobilityops-premium-ui`, branched from verified deployed revision `dfabb41582e302f45a3de826f85f531bf23dfc8b`; master history was not rewritten. - Audited every route at 1440, 1280, 768 and 390 px. Baseline findings and captures are in `docs/design/current-ux-audit.md` and `artifacts/design-validation/current/`. - Authored three twelve-screen product directions and generated representative Stitch anchors in project `17018847755558569017`: Control Rail, Dispatch Ledger and Service Atelier. Control Rail was selected and refined twice for hierarchy, accessibility and responsive implementation. Decision, screen inventory, tokens and exact Stitch IDs are in `docs/design/design-directions.md`, `docs/design/design-system.md` and `docs/design/stitch-manifest.md`. - Rebuilt the complete React interface around a responsive Control Rail shell: inline SVG icon/brand system, desktop rail, named landmarks, skip link, top bar, mobile bottom navigation, shared loading/error/empty states and reduced-motion support. - Redesigned all shipped pages. The dashboard now prioritizes persisted readiness, Attention and today's movements; booking results paginate at 25 rows; every responsive table retains field labels; integrations distinguish n8n evidence, live RAGcore health and the unconfigured MCP adapter without inventing status. - Return registration is now capture → review → result. A regression test proves the return endpoint is not called before confirmation; the existing idempotency and local commit/outbox contract is unchanged. - Final browser captures are in `artifacts/design-validation/implementation/`. DOM measurements and Playwright both prove no horizontal overflow at 390, 768, 1280 and 1440 px. See `docs/design/implementation-validation.md`. - Final validation commands from this branch: - `docker compose run --rm api pytest -q` — **66 passed**. - `docker compose run --rm api ruff check .` — **All checks passed**. - `docker compose run --rm api mypy app` — **0 issues in 44 files**. - `cd frontend && npm run lint` — clean TypeScript check. - `cd frontend && npm run build` — production build succeeded (59 modules; 238.66 kB JS, 36.54 kB CSS before gzip). - `cd frontend && playwright test --reporter=line` — **18 passed** including the full five-minute demo, every interactive route, return review semantics and four viewport overflow checks. - Review deployment updated at `http://192.168.10.150:1236` with persistent PostgreSQL and n8n volumes preserved. Deployed smoke: all ten authenticated routes plus login at desktop and mobile sizes rendered without alert state or horizontal overflow; browser console had zero warnings/errors; seven authenticated API paths returned 200; PostgreSQL/API were healthy and n8n `/healthz` returned `{"status":"ok"}`. - The repository compose file uses developer ports, so the existing review deployment's established bindings were restored after archive extraction: web `1236:80`, n8n `127.0.0.1:15678:5678`, and no host API port. This is deployment configuration only; no unrelated container was altered. - Exact next action after this record: commit and push the deployed evidence, update the deployed source-revision marker to that final documentation commit without rebuilding unchanged images, then hand the branch off for review. Do not merge master automatically.