Files
MobilityOps/PROJECT_STATE.md
T
NuklearRabbitandClaude Sonnet 5 4227fe4f58 docs: record live /v1/answers verification and a retrieval finding
Issued a fresh scoped credential via the RAGcore admin UI (separate,
working OIDC session, unaffected by the n8n auth problem) and made a
real authenticated /v1/answers call. Got a clean 200 with a real
answer_id/retrieval_run_id, not degraded -- but not_answerable, 0
citations.

Confirmed this isn't a regression: the same query through RAGcore's
own pre-existing Query Lab tool (untouched this session) returns the
identical result down to zero dense/sparse candidates at the raw
retrieval stage. Ruled out the obvious causes via direct Qdrant/
Postgres checks -- workspace_id, space_id, status, and embedding
digest all correctly match real indexed content. Root cause not yet
found; flagged as a follow-up rather than pursued further to avoid
scope creep on what was a deployment-verification task.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 04:01:16 +02:00

1626 lines
150 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Project state
## Publication and Unraid deployment (2026-08-02)
- Unraid deployment is live at `http://192.168.10.150:1236` from
`/mnt/user/appdata/mobilityops`, Compose project `mobilityops`.
- Deployment config commits: `07ab7a3`, `847cd05`, `e1a1c67`, `1e13943`. The accepted
baseline `4bf9afbeff44088864e0844769d4dd0e4089d85b` remains intact.
- MobilityOps PostgreSQL, API and web services are healthy. Only web port 1236 is exposed
by the MobilityOps Compose project; API and PostgreSQL remain internal. Automation uses
the server's existing shared n8n at `http://192.168.10.150:5678`; no second MobilityOps
n8n container is running.
- Migrations are at `e7b08389f47f (head)` and deterministic seed counts match final
acceptance. A Chrome smoke test covered every requested page and a real return; its n8n
event succeeded on attempt 1. Browser console and recent service log scans were clean.
- RAGcore is disabled in favor of the honest local demo provider. MCP Hub registration is
disabled. The MobilityOps workflow is published in the existing n8n and live-verified.
- Local post-change gates: 66 backend tests, Ruff, mypy (44 files), and frontend production
build all pass. Evidence is in `artifacts/deployment/unraid-summary.md`.
- Published to the private Gitea repository
`https://gitea.itworx.tech/Jens/MobilityOps`. `master` is the default branch; the full
commit history and baseline commit are present; zero tags exist; remote hygiene is
clean. `origin` uses the SSH clone URL supplied by Gitea.
- Exact next action: none — repository publication and Unraid deployment are complete.
## Current milestone
M7 — complete. All milestones (M0M7) done, plus a full post-M7 final-acceptance audit (see below). See `artifacts/final-acceptance/summary.md` for the definitive acceptance evidence (supersedes `artifacts/evidence/final-summary.md`, which is kept as historical M7 evidence).
## Locked decisions
- Product name: Fleet Ops (the only visible product name in the UI/copy, never translated;
see `frontend/src/product.ts`). "MobilityOps" is the internal repo name, Compose project
name and deployment directory only — never shown to a user. See the "Final product
polish: Fleet Ops rebrand" and "Fleet Ops final localization" entries below.
- Fictitious tenant: Northstar Mobility Demo.
- Synthetic demo data only; all operational and knowledge data are synthetic. The product
itself is not described as a "PoC" in user-facing copy (see the localization entry
below) — this document and other internal/engineering docs may still use "PoC" to
describe the engineering scope, per `CLAUDE.md`.
- Core stack and boundaries are defined in `CLAUDE.md` and `docs/03-architecture.md`.
- RAGcore and ITWorx MCP Hub are external central services.
- n8n receives post-commit events through an outbox dispatcher.
- SQLAlchemy 2 declarative models cover the full domain model (`backend/app/models/`); enums are plain `String` columns validated at the Pydantic/service layer, not native PG enums (simpler migrations).
- `backend/requirements.lock` is compiled inside a `python:3.12-slim` container (matches the Dockerfile base image) via `pip-compile --extra dev`; regenerate the same way if `pyproject.toml` changes.
- Frontend dependencies pinned (no more `"latest"`); `package-lock.json` committed; Docker build uses `npm ci`.
- Demo auth is a lightweight HMAC-signed cookie (`app/core/security.py`), not a real password/JWT flow — matches "Demo role buttons create an authenticated session; they do not bypass authorization middleware." Two fixed demo users (`USR-OPS` operations_manager, `USR-EMP` rental_employee) are created by the seed loader, not from a CSV (no `users.csv` in `seed/`).
- Seed loader (`backend/app/seed_loader.py`) only supports `seed --reset` (always rebuilds); there is no incremental/idempotent-without-reset mode, since the acceptance criteria only require deterministic reset, not partial import.
- `DataQualityIssue.entity_ref`/`related_ref` from the CSVs are resolved to `entity_type`/`entity_id` (UUID) at load time per the domain model; the original human-readable refs are kept in `evidence_json` (`entity_ref`, `related_refs`) since the API and UI need them and re-resolving UUID→public_ref on every read would be wasteful.
- `backend/app/core/config.py` added `app_secret`, `session_cookie_name`, `session_ttl_seconds`, `seed_dir` (`/app/seed` in-container), `cors_allow_origins` (comma-separated string, not a list — simpler with pydantic-settings env parsing), `demo_today` (drives the dashboard's "Today" section against the deterministic anchor date, default `2026-08-01`).
- `compose.yaml` api build context changed from `./backend` to repo root with `dockerfile: backend/Dockerfile`, so the image can `COPY seed ./seed` (seed CSVs are outside `backend/`).
- Frontend: added `react-router-dom@7.18.2` (bumped from 6.x to clear two real advisories — open redirect + arbitrary constructor injection in v6). One residual `npm audit` finding (RSC-mode CSRF, GHSA-qwww-vcr4-c8h2) does not apply — this SPA never uses React Router's RSC/SSR mode.
- Nav/pages built so far: Dashboard, Vehicles (list+detail with tabs), Bookings (list+detail), Audit. Data Quality, Knowledge and Automation nav items are intentionally omitted until M3/M5/M4 build the pages behind them — CLAUDE.md forbids dead routes/placeholders.
- Return workflow (`app/services/returns.py`): the spec's "validate submitted reading against booking start reading" step was dropped as a hard rejection. For the seeded S1 scenario, a booking's `start_odometer_km` can already equal the vehicle's canonical odometer, so any regression-testing value would also be below the booking start, making a hard floor there indistinguishable from — and in conflict with — the documented soft-regression path. Only one odometer check exists now: submitted vs. the vehicle's *canonical* odometer (`vehicle.odometer_km`), matching the domain-model invariant verbatim ("a return with a lower submitted reading is recorded as an inspection and issue, while canonical odometer remains unchanged").
- Idempotency: new `idempotency_records` table (migration `e7b08389f47f`), unique on `idempotency_key`, keyed to `booking_id`. Same key + same booking replays the stored response; same key + different booking → 409 `IDEMPOTENCY_KEY_REUSED`; different key on an already-returned booking → 409 `INVALID_BOOKING_STATE`. Concurrency is enforced by `SELECT ... FOR UPDATE` on the booking row (re-checked for the idempotency record immediately after acquiring the lock, as a safety net for two simultaneous identical-key requests racing the pre-lock check).
- `seed_loader.clear_all()` must delete `idempotency_records` before `bookings` (FK) — easy to forget when adding new booking-referencing tables; the ordering list at the top of `seed_loader.py` is the single place to update.
- Inspection `public_ref` is assigned as `INSP-{count+1:04d}` from a live count query (not gap-safe, fine for a PoC single-writer demo, would need a sequence for real concurrency-safe numbering).
- Found and fixed during browser verification (not caught by pytest, since it's a UI-only defect): `ReturnForm` originally held its own `result` state and was conditionally rendered only when `booking.status === "active"`; once the return succeeded the booking flipped to `returned` and React unmounted the form before the user ever saw the result panel. Fixed by lifting the result into `BookingDetail` (`ReturnResultPanel` is now a sibling, not nested in `ReturnForm`). Also found: `OutboxEvent.event_id`'s Python-side `default=uuid.uuid4` on the mapped_column only applies at flush/commit time, so reading `event.event_id` before `db.commit()` returned `None` (rendered as the literal string "None" in the result panel); fixed by assigning `event_id=uuid.uuid4()` explicitly at construction. Lesson: SQLAlchemy column `default=` callables are not available on the in-memory Python object until flush — never rely on the generated value for a same-transaction response body without an explicit `db.flush()` or an explicit Python-side assignment.
- Operational note for this environment: `docker compose run --rm api ...` (used for tests/lint) only starts a throwaway one-off container — it does **not** update the long-running `api`/`web` service containers. After any code change meant to be verified live (browser, curl), `docker compose up -d --build <service>` is required, not just `docker compose build`.
## Completed evidence
### M0 — Reproducible foundation
- Added `backend/app/core/db.py` (engine/session), `backend/app/models/*` (User, Customer, Vehicle, Booking, Inspection, MaintenanceRecord, DataQualityIssue, OutboxEvent, AuditEvent), Alembic config (`backend/alembic.ini`, `backend/alembic/env.py`) and initial migration `backend/alembic/versions/c9498525abb5_initial_schema.py`.
- Commands run and verified from this checkout:
- `docker compose build api` — OK
- `docker compose run --rm api alembic upgrade head` — applied cleanly to empty DB, created 9 tables + `alembic_version`.
- `docker compose run --rm api pytest -q` — 1 passed.
- `docker compose run --rm api ruff check .` — All checks passed (added `extend-exclude = ["alembic/versions"]` to `backend/pyproject.toml` for autogenerated migration line length).
- `docker compose up -d --build` — all 4 services healthy: `curl http://localhost:8128/health``{"status":"ok",...}`; `curl -o /dev/null -w "%{http_code}" http://localhost:1228/` → 200; `curl http://localhost:5678/healthz` → 200.
- Fixed a real scaffold bug: `frontend/src/App.tsx` used `import.meta.env` without a `vite/client` types reference, which broke `npm run build` in Docker (works fine under plain `vite dev` because Vite injects the global at dev-time but `tsc -b` still type-checks it). Added `frontend/src/vite-env.d.ts`.
- `make` is not installed in this Windows/git-bash shell — validated the underlying `docker compose ...` commands directly instead (Makefile targets are thin wrappers around them and are correct as written for a Linux/CI shell or WSL).
- Known accepted gap: `npm audit` reports 1 moderate/1 high transitive `esbuild` advisory (dev-server-only, fixed only by a Vite 8 major bump); left as-is for the PoC, noted here rather than silently upgrading a major version.
### M1 — Operational core
- Backend additions: `app/core/security.py` (HMAC-signed session cookies), `app/api/deps.py` (`get_current_user`, `require_operations_manager`), `app/core/errors.py` (`AppError` + the documented `{"error": {...}}` shape wired as a FastAPI exception handler for both `AppError` and `HTTPException`), `app/seed_loader.py`, `app/cli.py` (`python -m app.cli seed --reset`), `app/services/audit.py`, `app/schemas.py`, routers under `app/api/routers/` (`demo`, `dashboard`, `vehicles`, `bookings`, `audit`).
- Frontend additions: React Router-based app shell (`src/App.tsx`, `src/components/Layout.tsx`, `src/components/RequireAuth.tsx`), `AuthContext`, typed `api` client (`src/api/client.ts`, `src/api/types.ts`), pages `Login`, `Dashboard`, `Vehicles`/`VehicleDetail`, `Bookings`/`BookingDetail`, `Audit`. Full responsive stylesheet (`src/styles.css`) covering nav collapse and table→card layout under 700px, visible focus states, no hover-only actions.
- Commands run and verified from this checkout (container rebuilt each time to pick up code changes):
- `docker compose run --rm api pytest -q`**19 passed** (new: `test_seed.py`, `test_auth.py`, `test_dashboard.py`, `test_vehicles.py`, `test_bookings.py`, `test_audit.py`; tests seed the real Postgres via `reset_and_seed` in a session fixture, then exercise the FastAPI app through `TestClient`, not mocks).
- `docker compose run --rm api ruff check .` — All checks passed (added `ignore = ["B008"]` — FastAPI's `Depends()`-as-default is idiomatic, not a real bug).
- `npm run build` (local, Node 24) — clean `tsc -b && vite build`.
- `docker compose up -d --build` then `docker compose exec api python -m app.cli seed --reset` — counts: `users:2 customers:180 vehicles:50 bookings:246 inspections:75 maintenance:40 data_quality_issues:15 workflow_runs:20`.
- `curl` end-to-end: `POST /api/v1/demo/login` sets cookie and returns the user; unauthenticated `GET /api/v1/dashboard` → 401 with the documented error shape; authenticated dashboard/vehicle-detail return real seeded data (verified metrics `available:21 rented:11 cleaning:6 maintenance:5 blocked:7`, matching the 50 seeded vehicles).
- Browser smoke test (Chrome via MCP) at desktop width: login page → Operations Manager login → Dashboard (metrics + attention items + today + recent automation all populated) → Vehicle detail `MO-016` (tabs render, "Needs attention" badge correct — it's `DQ-DEMO-OVERLAP`/`DQ-DEMO-STATUS`) → Booking detail `BK-DEMO-RETURN` (matches S1 scenario: vehicle `MO-024`, status `active`, start odometer `53610`). Responsive CSS (`@media max-width:700px`) was written and code-reviewed but the automated resize during this session didn't visibly reflect in the captured screenshot (likely a screenshot-timing quirk of the browser tool, not necessarily a real bug) — **treat the ≤360px layout as visually unverified** and re-check with a real device/DevTools emulation before final acceptance (M7).
- Known accepted gap carried over from M0: `npm audit` residual `esbuild`/Vite-8 dev-server-only advisory.
### M2 — Vehicle return vertical slice
- Backend additions: `app/models/idempotency.py` (`IdempotencyRecord`), migration `e7b08389f47f_idempotency_records`, `app/services/returns.py` (`register_vehicle_return` — full transaction: row locks, idempotency replay, inspection, canonical-odometer update or regression issue, vehicle status derivation, two audit events, `vehicle.returned.v1` outbox event matching `contracts/events.schema.json`, next-booking-risk lookup), `POST /api/v1/bookings/{public_ref}/return` wired in `app/api/routers/bookings.py` with required `Idempotency-Key` header.
- Frontend additions: `components/ReturnForm.tsx` (form + `ReturnResultPanel`), wired into `pages/BookingDetail.tsx` (shown only when `booking.status === "active"`; result persists via lifted state after the booking flips to `returned`).
- Commands run and verified from this checkout:
- `docker compose run --rm api pytest -q`**26 passed**, including `tests/test_return.py` (success/canonical-update, S1 regression scenario by name, damage→blocked, idempotent replay, reject-already-returned, missing-header validation, and a real multi-threaded concurrent-submission test against Postgres asserting exactly 1×201 + 2×409).
- `docker compose run --rm api ruff check .` — All checks passed.
- `npm run build` — clean.
- `docker compose up -d --build` (all services) then `docker compose exec api python -m app.cli seed --reset`, then a full browser run of the S1 demo scenario against `BK-DEMO-RETURN`/`MO-024`: submitted 53000 km (below canonical 54820) → result panel showed `INSP-0076`, `resulting_vehicle_status: maintenance` (correctly derived, since canonical 54820 ≥ `next_service_km` 40000), `DQ-RET-0076` created, automation event queued with a real UUID, "no upcoming booking" risk; vehicle detail page confirmed odometer unchanged at 54,820 km and a "Needs attention" badge.
- Both real defects listed above (form disappearing before showing its result; `event_id` reading as `None`) were **found via the browser run, not by pytest** — the test suite asserted on API response shape/values, not on what the UI actually rendered after a status transition. Worth remembering for M3+: UI state-after-mutation bugs need a browser check, not just API tests.
### M3 — Data Quality Workbench
- `app/services/data_quality.py`: `run_scan()` implements all five rules and is called automatically at the end of `seed_loader.reset_and_seed()` (after `db.commit()` of the base seed), plus exposed as `POST /api/v1/data-quality/scan` (Operations Manager only). Idempotency is simplified from the doc's literal `(rule_type, entity_type, entity_id, evidence fingerprint)` to just `(rule_type, entity_type, entity_id)` while an issue is open — see rationale below.
- Router `app/api/routers/data_quality.py`: `GET /issues` (filters status/rule_type/severity), `GET /issues/{ref}` (adds `entity_snapshot`/`related_snapshots` for the UI), `POST /issues/{ref}/defer`, `/reject`, `/merge-customers` (Operations Manager only — enforced via `require_operations_manager`), `POST /scan`.
- **Real bug found and fixed during this milestone, before any browser check**: the first cut of DQ-03 (odometer regression) compared every historical *returned* booking's `end_odometer_km` against the vehicle's *current* `odometer_km`. Since the seed generator assigns `vehicle.odometer_km` independently of booking history (see `seed/generate_seed.py`), this is true for nearly every historical booking by construction (odometer is monotonically increasing over time, so all-but-the-latest reading is "below current") — it produced 51 false-positive issues out of 50 vehicles on first run. Fixed twice: first attempt (compare only the single most-recent booking against canonical) still produced the same problem because canonical itself is disconnected from booking history in this dataset; the working fix compares each vehicle's *own returned-booking sequence* against itself (each booking's end reading vs. the immediately preceding one, chronologically) — a self-consistency check that doesn't depend on the unrelated `vehicle.odometer_km` field at all. Final deterministic seed+scan totals: 15 CSV-seeded + 11 scan-discovered = **26** open/resolved `data_quality_issues` (breakdown: 14 vehicle_status_conflict, 5 missing_required_field, 3 possible_duplicate_customer, 3 odometer_regression, 1 booking_overlap). `tests/test_seed.py`'s exact-count assertion was updated from 15 to 26 accordingly — if the scan logic changes again, update that count.
- Idempotency simplification rationale: the doc's fingerprint-based key would make the scan blind to issues it structurally can't compute a matching fingerprint for against the CSV-seeded rows (which don't carry a fingerprint field), producing duplicate issues for the same real-world problem (e.g. a second `MO-016` overlap issue next to the seeded `DQ-DEMO-OVERLAP`). Using `(rule_type, entity_type, entity_id)` alone while open is a stricter, safe simplification: it can never falsely suppress an issue for a *different* entity, and per-entity there's realistically only one meaningful open issue of a given rule type at a time for this PoC's scope.
- Merge UI intentionally does **not** use `window.confirm()` — a native dialog blocks further automation/testing and isn't screen-reader-distinguishable from page content the same way a rendered `role="alertdialog"` panel is. Built an inline two-step confirm instead (`ReturnForm`-style pattern reused).
- `AttentionItem` gained an `issue_ref` field (dashboard now links attention items straight to `/data-quality/{issue_ref}` instead of only to vehicles); dashboard attention list capped at 8 items (was unbounded, would have shown up to 26 with the richer scan).
- Commands run and verified from this checkout:
- `docker compose run --rm api pytest -q`**35 passed** (new `tests/test_data_quality.py`: all five rule types present, scan idempotent on rerun, scan requires Operations Manager, S2/S4 issue-detail snapshots correct, defer→reject-on-closed 409, merge requires Operations Manager, merge rejects an unrelated survivor ref, full S2 merge scenario asserting rewiring + audit + replay-is-409).
- `docker compose run --rm api ruff check .` — All checks passed.
- `npm run build` — clean (had to fix two `possibly 'null'` TS errors from a closure-narrowing limitation — TS doesn't narrow `const` captured-by-closure across nested function boundaries when the value comes from an index/property expression; fixed by re-binding to explicitly-typed local consts right after the guard).
- Full browser run: Data Quality list (26 open issues, filterable) → `DQ-DEMO-DUPLICATE` two-column compare (CUS-0012 vs CUS-0178, per-field diff highlighting only where they differ) → merge with inline confirm → issue flips to `resolved` → confirmed `customer_merged` audit event with correct actor/entity/correlation → `DQ-DEMO-OVERLAP` (non-duplicate type) renders evidence JSON + defer/reject, no dead compare UI shown for a rule type it doesn't apply to.
### M4 — n8n automation
- `app/services/dispatcher.py`: background daemon thread (started/stopped via FastAPI `lifespan`, not an `on_event` hook) polling every `N8N_DISPATCH_INTERVAL_SECONDS` (default 3s). Claim step (`_claim_due_events`) is a short transaction using `SELECT ... FOR UPDATE SKIP LOCKED` that only flips `pending``delivering` and commits immediately; the HTTP call to n8n happens with **no open transaction**; the outcome is recorded in a separate short transaction. Exponential backoff `min(2**attempts, 60)` seconds, `N8N_MAX_ATTEMPTS=5` before a permanent `failed`.
- Dispatcher reconstructs the wire event from `contracts/events.schema.json`'s exact fields (`event_id`, `event_type`, `occurred_at`, `correlation_id`, `aggregate`, `data`) rather than forwarding `OutboxEvent.payload_json` wholesale — that column also carries an internal `aggregate_ref` convenience key (used by dashboard/workflows list rendering) that the schema's `additionalProperties: false` would reject.
- `POST /api/v1/integrations/n8n/return-callback` (`app/api/routers/integrations.py`): shared-secret auth via `X-Service-Token` header (`N8N_CALLBACK_TOKEN`, propagated to both `api` and `n8n` containers as `MOBILITYOPS_CALLBACK_TOKEN`); idempotent by `Idempotency-Key` (the event UUID) — checked by querying for an existing `AuditEvent` with that event ID in its metadata, **not** by `OutboxEvent.external_run_id`, because the dispatcher only sets that field *after* it gets n8n's final response, which happens *after* n8n has already called this callback mid-workflow — using `external_run_id` as the idempotency guard would have missed the exact redelivery case it's meant to catch.
- `GET /api/v1/workflows` + `POST /api/v1/workflows/{event_id}/retry` (`app/api/routers/workflows.py`), both Operations Manager only. Retry only allowed from `failed`; sets `pending` + clears `next_attempt_at` so the live dispatcher picks it up on its next cycle (does not reset `attempts`, so the counter reflects true delivery history).
- Automation nav + page (`pages/Automation.tsx`): table of all runs with status/attempts/last error, Retry button for `failed` rows, visible only to Operations Manager (matches backend authorization rather than just hiding a link).
- **Two real bugs found and fixed, the second only by testing the actual live n8n round-trip, not by pytest**:
1. Seed-loaded `workflow_runs.csv` rows only ever got `payload_json = {"aggregate_ref": ...}` (no `correlation_id`/`aggregate`/`data`) — fine for M1M3 since nothing read those keys yet, but once the dispatcher tried to *redeliver* a seeded row (i.e. the S5 manual-retry demo scenario) it crashed with `KeyError: 'correlation_id'`, leaving that event stuck in `delivering` forever (the crash happened before the outcome-recording transaction). Fixed in two places: `seed_loader.py` now builds the full schema-compliant envelope for every `workflow_runs.csv` row (matching what the live M2 return flow produces), and `dispatcher._deliver_one` now catches malformed-payload `KeyError`s defensively and resolves the row to `pending`/`failed` instead of leaving it orphaned — added `test_deliver_one_handles_malformed_payload_without_getting_stuck` as a regression test for the latter.
2. This n8n image (2.32.7) has dropped `N8N_BASIC_AUTH_ACTIVE` as a UI/API gate — it requires an actual owner account via the `/setup` flow before anything (including webhook registration reliability) works correctly. Also: `n8n import:workflow` requires the workflow JSON to have a top-level `"id"` field (added `"id": "mobilityops-return-processing"`) and **always deactivates** the imported workflow regardless of its `"active"` field — activation requires `n8n publish:workflow --id=<id>` followed by a full n8n restart (documented in n8n 2.x CLI, not obvious from the docs pack). Did this manually this session via the CLI + browser setup wizard; **this is a one-time operational step that is not automated** — a truly clean checkout still needs someone to run `docker compose exec n8n n8n import:workflow --input=//imports/mobilityops-return-processing.json`, `docker compose exec n8n n8n publish:workflow --id=mobilityops-return-processing`, `docker compose restart n8n`, and complete the one-time owner setup at `http://localhost:5678/setup` (any email/password, no verification required) before the automation demo will work. `docs/17-runbook.md` should get this exact sequence in M7.
- Commands run and verified from this checkout:
- `docker compose run --rm api pytest -q`**49 passed** (new `tests/test_dispatcher.py` — claim/deliver success/failure/backoff/exhaustion-to-failed/malformed-payload, all via `monkeypatch.setattr(dispatcher.httpx, "post", ...)`, no real network calls in tests; `tests/test_integrations.py` — callback auth, unknown-event 404, idempotent-by-event-ID with a real duplicate-call assertion; `tests/test_workflows.py` — role gating, retry-only-from-failed, S5 retry-and-audit).
- `docker compose run --rm api ruff check .` — All checks passed.
- `npm run build` — clean.
- Full live round trip (not mocked): registered a real return on `BK-DEMO-RETURN` → outbox event queued → background dispatcher delivered it to the now-activated n8n workflow within its 3s poll interval → n8n called back into `/api/v1/integrations/n8n/return-callback` (200 OK, confirmed in `docker compose logs api`) → dispatcher's original POST received n8n's success response → event flipped to `succeeded` on attempt 1, visible on `/automation`.
- S5 scenario end-to-end in the browser: seeded `BK-H-0020` (`failed`, 3 attempts, "Synthetic connection timeout to n8n") → clicked Retry → `pending` → within ~3s, live dispatcher delivered it through the real n8n instance → `succeeded`, 4 attempts. This is the full documented S5 scenario working for real, not simulated.
### M5 — RAGcore knowledge integration
- `app/services/knowledge/__init__.py`: `KnowledgeProvider` Protocol (sync, not async — the rest of the backend is sync SQLAlchemy/FastAPI, so an async provider interface would have meant bridging paradigms for no benefit) with `health()`/`ask()`, plus `GroundedAnswer`/`SourceCard`/`KnowledgeHealth` Pydantic models matching `contracts/openapi.yaml`'s `GroundedAnswer` schema exactly. `get_knowledge_provider()` factory switches on `settings.knowledge_provider` ("demo" default, "ragcore" opt-in).
- `app/services/knowledge/demo.py``DemoKnowledgeProvider`: parses the 10 `knowledge/procedures/*.md` files' YAML frontmatter (hand-rolled flat parser, not PyYAML — avoided adding a dependency for a 6-key flat block) and `## `-delimited sections at startup, then does **TF-IDF-weighted keyword retrieval** (not naive keyword counting) with light suffix-stripping stemming (`returns``return`, `damaged``damage`). This is extractive, not generative: it returns real excerpts and a templated answer sentence, never invented text.
- **Real bug found and fixed by testing the actual S6 question, not by inspection**: naive flat keyword-overlap scoring (first cut) let the word "vehicle" — present in nearly every document's title — crowd out the actually-relevant `damage-procedure` document from the top-3 results for "What must I do when a vehicle returns with damage?", because generic words scored the same as distinctive ones. Fixed by computing corpus-wide IDF per token (`log((N+1)/(df+1)) + 1`) and weighting matches by it, so common terms contribute little and rare/distinctive terms (like "damage") dominate the ranking. Verified: the S6 question now returns `damage-procedure` and `vehicle-return-procedure` in the top 3, matching the documented expectation exactly.
- `app/services/knowledge/ragcore.py``RAGcoreKnowledgeProvider`: real `httpx` adapter guessing a plausible REST contract (`GET /health`, `POST /api/v1/ask`) per `contracts/ragcore-contract-assumptions.md` (RAGcore is built separately; no live instance was reachable this session to verify against). Any connection error, timeout, or malformed response degrades to `evidence_state: "unavailable"` rather than raising — this is the adapter that actually exercises the architecture's "RAGcore failure disables knowledge answers only" reliability boundary. Not wired as the active provider by default; `KNOWLEDGE_PROVIDER=ragcore` would need a real, verified base URL to turn on.
- `POST /api/v1/knowledge/questions` + `GET /api/v1/knowledge/status` (`app/api/routers/knowledge.py`). Audit event `knowledge_question_asked` logs `evidence_state`, `provider`, `source_ids`, and `question_length` only — **not** the question text itself, per `docs/12-security-and-audit.md` ("log question metadata and source IDs, not unnecessary full prompts").
- Knowledge nav + page (`pages/Knowledge.tsx`): chat-style question box, source cards (title/version/section/excerpt) prioritized over the answer text per `docs/06-ui-ux.md`, explicit `grounded`/`insufficient`/`unavailable` states with distinct visual treatment — never a fabricated-looking answer for the latter two.
- Dockerfile now also `COPY knowledge ./knowledge`; added `KNOWLEDGE_DIR` setting (`/app/knowledge/procedures` in-container, same pattern as `SEED_DIR`) rather than deriving the path from `__file__` — simpler and doesn't break if the module moves.
- Commands run and verified from this checkout:
- `docker compose run --rm api pytest -q`**57 passed** (new `tests/test_knowledge.py`: S6 grounded-with-expected-sources, unrelated question is honestly insufficient with no fabrication, demo provider health/document count, endpoint auth required, audit doesn't leak question text, RAGcore adapter degrades to unavailable on a simulated connection error).
- `docker compose run --rm api ruff check .` — All checks passed.
- `npm run build` — clean.
- Full browser run of S6 end-to-end: asked "What must I do when a vehicle returns with damage?" on `/knowledge` → grounded answer citing "Vehicle return procedure" (2 sections) and "Damage handling procedure" with real excerpts. Also asked an unrelated question ("What is the weather forecast for tomorrow?") → correctly returned "Insufficient evidence" / "No matching procedure was found" with zero sources, confirming no fabrication.
### M6 — ITWorx MCP Hub publication
- `app/api/routers/mcp_integrations.py`: four read-only endpoints under `/api/v1/integrations/mcp/``GET operations-summary`, `GET attention-vehicles` (query params `minimum_severity`/`date`/`limit` matching `contracts/mcp-tools.json`'s `inputSchema` exactly), `GET vehicles/{vehicle_ref}`, `POST search-knowledge` (the "narrow façade" the doc calls for — wraps M5's `get_knowledge_provider()` rather than re-implementing retrieval; the contract's tool has no MobilityOps `endpoint` field, only `routing.preferred: ragcore`, so this façade path is MobilityOps's own addition for when the Hub needs a single provider boundary, not literally specified by the contract).
- Auth: new `require_mcp_service_token` dependency in `app/api/deps.py`, same shared-secret-header shape as the M4 n8n callback (`X-Service-Token` against `MCP_HUB_SERVICE_TOKEN`) plus an optional `X-Client-Id` header (defaults to `"unknown-mcp-client"`) used as the audit actor label — the Hub's actual client-identity header name is unknown (no live Hub to confirm against), so this is a reasonable guess documented here rather than assumed silently.
- `McpVehicleDetailOut` deliberately omits `registration_number` and all customer data — narrower than the browser-facing `VehicleOut`/`VehicleDetailOut`, matching "no customer or vehicle database access" and the read-only/summary intent of an AI-facing tool. Test `test_vehicle_details_known_ref` asserts the field's absence explicitly so a future change can't silently widen the exposed surface.
- Extracted `app/services/operations.py` (`compute_metrics`, `list_attention_vehicles`) out of `app/api/routers/dashboard.py` so the MCP operations-summary/attention-vehicles endpoints and the human dashboard share one query implementation instead of two copies that could drift — the same "do not duplicate retrieval logic" principle the doc states for the knowledge tool, applied here to the operational-summary tools too.
- Every provider call writes an `AuditEvent` (`actor_type="service"`, `actor_label=X-Client-Id`, `action="mcp_tool_request"`, `metadata={tool, status}`) — MobilityOps's own record that its provider APIs were reached, independent of whatever central tool-call audit the Hub itself keeps (per `docs/10-mcp-hub-integration.md`'s audit section, the Hub owns the central log; this is the local corroborating one).
- No write/mutation endpoints exist under the `/api/v1/integrations/mcp/` namespace at all (verified by `test_no_write_endpoints_exist_under_mcp_namespace` — POST/PUT/DELETE against the vehicle-details path all 404/405) — return registration, customer merge, and any booking/vehicle mutation are correctly absent, per the doc's explicit restriction list.
- Commands run and verified from this checkout:
- `docker compose run --rm api pytest -q`**66 passed** (new `tests/test_mcp_integrations.py`: token-required, wrong-token 401, all four tools' happy paths, severity/limit filtering, 404 for unknown vehicle, `max_sources` respected, audit actor/action verified, write-method rejection).
- `docker compose run --rm api ruff check .` — All checks passed.
- Live `curl` verification against the running stack (no browser needed — these are service-to-service endpoints, not UI): missing header → 422; wrong token → 401; correct token → all four endpoints return correct data (`operations-summary` metrics match the dashboard; `attention-vehicles?minimum_severity=high` returned `MO-016`×2 and `MO-031`, all severity `high`; `vehicles/MO-016` returned the narrow read-only shape; `search-knowledge` with `max_sources=2` returned exactly 2 grounded sources for the S6 question). Confirmed via `GET /api/v1/audit?action=mcp_tool_request` that all four calls were recorded with correct `actor_type=service`, tool name, and status.
### M7 — Portfolio polish and final acceptance
- **Automated clean-checkout migrations**: `backend/entrypoint.sh` now runs `alembic upgrade head` before starting uvicorn (Dockerfile `CMD` changed from `uvicorn ...` to `./entrypoint.sh`). Verified with a true `docker compose down -v` (all volumes wiped) → `docker compose up --build -d` → all 11 tables present, `/health` and web both green, all 66 backend tests pass, with zero manual migration step.
- **n8n one-time setup scripted where it can be**: `make n8n-setup` runs the import/publish/restart sequence (previously three manual commands discovered ad hoc in M4). The owner-account creation itself cannot be scripted safely (it's an interactive one-time step in n8n 2.x's own onboarding, not a MobilityOps concern) — documented precisely in the rewritten `docs/17-runbook.md`, including the exact URL and that no email verification is required. Re-ran this full sequence from the wiped-volumes state this session and confirmed the S1 return → outbox → live n8n → callback → `succeeded` round trip works on a genuinely clean checkout, not just the already-provisioned stack from M0M6.
- **Playwright E2E** (`frontend/e2e/demo.spec.ts`, `frontend/playwright.config.ts`): one test automating the full 9-step documented demo script end-to-end against the live stack — login, dashboard metrics, open `BK-DEMO-RETURN`, register an odometer-regression return (S1), verify the quality issue + queued automation event, merge the duplicate-customer scenario (S2), ask the damage question and verify both expected source citations (S6), inspect audit entries, and verify responsive nav + no horizontal overflow at 360px width. **Passing.** This also resolves the "≤360px layout visually unverified" gap flagged back in M1 — verified both by this test's overflow assertion and by the `9-mobile-dashboard.png` screenshot (nav wraps into rows, metric tiles collapse to a 2-column grid, no horizontal scroll).
- Added `frontend/e2e/_capture-screenshots.spec.ts` as evidence-generation tooling (underscore-prefixed, excluded from the default `playwright test` / `make e2e` run via `testIgnore` in the config — it calls `demo/reset`, which a real regression test shouldn't do as a side effect). Captured all 9 screenshots into `artifacts/evidence/screenshots/`.
- Wrote `artifacts/evidence/architecture.md` (mermaid, as-built — distinguishes verified-live components from implemented-but-never-reached-a-real-instance ones, i.e. RAGcore and the MCP Hub) and `artifacts/evidence/final-summary.md` (commit, exact commands, test counts, screenshot index, RAGcore success/unavailable evidence — including a live-demonstrated unavailable case against an unreachable host, not just the unit test — n8n success/retry evidence, MCP sample calls, known limitations, truthful portfolio wording per `docs/16-portfolio-case-study.md`'s template).
- Updated `README.md` (dropped stale "minimal bootable scaffold, not the finished application" wording and the old two-line quickstart in favor of `make demo` + a pointer to the runbook) and `docs/17-runbook.md` (full rewrite: exact bootstrap, n8n one-time setup, verification commands, required operational checks, recovery expectations).
- Final placeholder/dead-UI sweep: `grep`'d the full `frontend/src` and `backend/app` trees for scaffold/TODO/FIXME/"must be replaced" markers — none found. `FILE_INDEX.md` was left as-is; it's the original build-pack's archive-completeness manifest (a historical snapshot), not a living index that needs to track every file added since — updating it would misrepresent what it's for.
- Commands run and verified from this checkout (this milestone, cumulative across the whole build):
- `docker compose run --rm api pytest -q`**66 passed**, ruff clean.
- `cd frontend && npm run build` — clean.
- `cd frontend && npx playwright test`**1 passed** (full demo script, live stack).
- Full clean-checkout drill: `docker compose down -v``docker compose up --build -d``docker compose exec api python -m app.cli seed --reset``docker compose run --rm api pytest -q` (66 passed) → n8n owner setup + `make n8n-setup` → live S1 return round-tripped through the real n8n instance to `succeeded`.
### Final acceptance audit (post-M7)
A dedicated release-readiness audit was run after M7 claimed completion, specifically to
catch anything the milestone-by-milestone build might have missed by only ever validating
each piece in isolation.
- **Real gap found: `mypy` had never been run.** `mypy` is a declared dev dependency
(`backend/pyproject.toml`) but was never wired into any milestone's validation loop —
only `ruff` was. Running it cold surfaced **43 real type errors across 10 files**, all
pre-existing (not introduced by this audit). Triaged and fixed all of them rather than
suppressing:
- `services/returns.py`: the vehicle lookup after acquiring `FOR UPDATE` could type as
`Vehicle | None` with no runtime guard — added an explicit `if vehicle is None: raise
AppError(..., 404)`. This was a genuine defensive-programming gap (a dangling FK would
have crashed with an unhandled `AttributeError`/500 instead of a clean 404), not just
a type annotation issue.
- `api/routers/bookings.py`: same pattern for `db.get(Customer, ...)` /
`db.get(Vehicle, ...)` in `get_booking` — added a guard raising 500 with a clear
message instead of crashing on `None.public_ref`.
- `api/deps.py` + `api/routers/demo.py`: `CurrentUser.role` is a `Literal[...]`, but
`SessionPayload.role` (decoded from an HMAC-signed cookie) and `User.role` (a DB
column) are both plain `str`. Pydantic validates this at runtime already (so it was
never exploitable), but `get_current_user` now explicitly checks membership before
constructing `CurrentUser`, turning a would-be unhandled `ValidationError` (500) into
a clean 401 for a corrupted/tampered cookie — another real defensive improvement, not
just a type-checker appeasement.
- `api/routers/dashboard.py`, `api/routers/data_quality.py`: two instances of reusing
one variable name for both a `Vehicle` and a `Customer` across an if/else branch,
which is genuinely confusing to read regardless of what mypy thinks — renamed to
distinct variables (`entity`/typed union in dashboard, `customer`/`vehicle` in the
data-quality snapshot helper).
- `services/data_quality.py`, `seed_loader.py`: `Booking.__table__.update()` /
`Customer.__table__.update()` don't typecheck against SQLAlchemy 2.0's stubs (the
`.__table__` accessor is typed as the more general `FromClause`, which doesn't
declare `.update()`) — switched to the idiomatic `sqlalchemy.update(Model)` construct,
which is both correctly typed and the more modern SQLAlchemy 2.0 style anyway.
- Remaining handful (schemas.py's deprecated `conint()``Annotated[int, Field(...)]`,
a `Sequence` vs `list` `.sort()` call, an `assert`-guarded None-narrowing after a
`WHERE ... IS NOT NULL` filter mypy can't see through, `Result.rowcount` typing gaps)
were either latent pydantic-v1-style API usage or genuine SQLAlchemy stub limitations
— fixed with the idiomatic modern equivalent or a narrowly-scoped, commented
`# type: ignore[...]` at the exact line, never a blanket suppression.
- `make lint` now runs both `ruff check .` and `mypy app`; `mypy app` reports
**zero errors across 44 source files**.
- **No other defects found.** Re-ran the full journey matrix end-to-end against a
genuinely wiped-volumes (`docker compose down -v`) clean checkout: all 66 backend
tests, ruff, ✅; ran the demo login → dashboard → vehicle/booking detail → return
workflow → invalid-mileage rejection (422, both a negative value and a non-numeric
string) → data-quality issue review → duplicate-customer merge → audit trail →
Knowledge Assistant → live n8n round trip → MCP Hub endpoint journeys directly via
`curl` against the running stack, all correct.
- **Degraded-mode behavior explicitly re-verified live** (not just unit-tested):
stopped n8n with `docker compose stop n8n`, registered a return — it committed
(`201`, booking flipped to `returned`) exactly as required; the outbox event stayed
`pending` with real `ConnectError`s logged and exponential backoff (2 attempts over
~8s); restarted n8n and the dispatcher **self-healed** without any manual
intervention, delivering the event to `succeeded` on attempt 5. RAGcore unavailable
mode re-verified live against an unreachable host (`ConnectError``evidence_state:
"unavailable"`, empty answer, no fabrication). MCP Hub unavailability is
architecturally moot for MobilityOps — the Hub only ever calls *into* MobilityOps, so
there is nothing on the MobilityOps side that can degrade if the Hub is down (only the
reverse, "does an unavailable Hub break MobilityOps," which is trivially no since
nothing here calls out to it).
- **New test coverage added, no existing tests weakened**: `frontend/e2e/interactive-elements.spec.ts`
(11 Playwright tests — all seven nav items, every filter on every list page, vehicle
detail tabs, defer/reject, automation retry, knowledge form, role-switching, and
role-based page restriction) plus the existing `demo.spec.ts` — **12/12 e2e tests
passing** against the live stack.
- Verified `.env` is `.gitignore`d and was never committed (`git ls-files` /
`git log --all -p -- '*.env'` both empty); scanned full git history for AWS keys,
private-key headers, and `sk-...`-style tokens — none found. Every `Settings` field in
`backend/app/core/config.py` has a corresponding entry either directly in
`.env.example` or is derived/wired through `compose.yaml` (a few purely-internal
container-path constants like `SEED_DIR`/`KNOWLEDGE_DIR` are intentionally not
operator-configurable and correctly absent from `.env.example`).
- Grepped the full `frontend/src` and `backend/app` trees for TODO/FIXME/placeholder/
fake/stub/mock/"not implemented" markers — zero real hits (the two `placeholder=`
matches are legitimate HTML input placeholder attributes). Confirmed dashboard metrics
and all list-page data are 100% DB-backed (`compute_metrics` in
`services/operations.py`, never a literal in frontend JSX). Confirmed every frontend
route in `App.tsx` maps to an implemented page and every nav item maps to a real route
— no dead routes.
- Commands run and verified from this audit:
- `docker compose run --rm api pytest -q`**66 passed**.
- `docker compose run --rm api ruff check .` — All checks passed.
- `docker compose run --rm api mypy app`**Success: no issues found in 44 source files** (0 errors, down from 43).
- `cd frontend && npm run build` — clean (`tsc -b && vite build`).
- `cd frontend && npx playwright test`**12 passed** (`demo.spec.ts` + `interactive-elements.spec.ts`).
- Full clean-checkout drill repeated from a fresh `docker compose down -v`: automatic migrations, seed, 66/66 tests, n8n owner setup + `make n8n-setup`, live return round-tripped through n8n to `succeeded`.
- See `artifacts/final-acceptance/summary.md` for the complete evidence write-up (commands, exact outputs, demo access, deployment instructions, five-minute demo flow).
## Definition of done
All eight milestones (M0M7) are complete, and a dedicated post-M7 final-acceptance audit
found and fixed one real category of gap (`mypy` never having been run) with zero
regressions. `docs/14-testing-and-acceptance.md`'s clean-checkout acceptance list has been
walked item by item against a genuinely wiped-volumes checkout, twice (once in M7, once in
this audit), and `artifacts/final-acceptance/summary.md` is the authoritative final
evidence document. The two items not fully closed — a live RAGcore instance and a live
ITWorx MCP Hub instance — were never reachable in this environment; both integrations are
implemented, unit/contract-tested, directly verified against MobilityOps's own API, and
their unavailable-degradation paths are live-verified, but an actual round trip against
real RAGcore/Hub instances remains unconfirmed and is documented as such rather than
claimed.
## Known blockers
None. External service credentials may be absent; use the documented demo/degraded providers. The n8n workflow-activation steps are a one-time manual setup requirement in this environment (owner-account creation via n8n's own `/setup` UI cannot be scripted safely), fully documented in `docs/17-runbook.md` and scripted where possible (`make n8n-setup`). RAGcore and the ITWorx MCP Hub itself were never reachable in this environment — both integrations are implemented and directly tested/curl-verified against MobilityOps's own API, but neither a real RAGcore instance nor a real Hub round trip was available to confirm end-to-end.
## Premium Control Rail UI transformation (2026-08-02)
- Branch: `design/mobilityops-premium-ui`, branched from verified deployed revision
`dfabb41582e302f45a3de826f85f531bf23dfc8b`; master history was not rewritten.
- Audited every route at 1440, 1280, 768 and 390 px. Baseline findings and captures are
in `docs/design/current-ux-audit.md` and `artifacts/design-validation/current/`.
- Authored three twelve-screen product directions and generated representative Stitch
anchors in project `17018847755558569017`: Control Rail, Dispatch Ledger and Service
Atelier. Control Rail was selected and refined twice for hierarchy, accessibility and
responsive implementation. Decision, screen inventory, tokens and exact Stitch IDs
are in `docs/design/design-directions.md`, `docs/design/design-system.md` and
`docs/design/stitch-manifest.md`.
- Rebuilt the complete React interface around a responsive Control Rail shell: inline SVG
icon/brand system, desktop rail, named landmarks, skip link, top bar, mobile bottom
navigation, shared loading/error/empty states and reduced-motion support.
- Redesigned all shipped pages. The dashboard now prioritizes persisted readiness,
Attention and today's movements; booking results paginate at 25 rows; every responsive
table retains field labels; integrations distinguish n8n evidence, live RAGcore health
and the unconfigured MCP adapter without inventing status.
- Return registration is now capture → review → result. A regression test proves the
return endpoint is not called before confirmation; the existing idempotency and local
commit/outbox contract is unchanged.
- Final browser captures are in `artifacts/design-validation/implementation/`. DOM
measurements and Playwright both prove no horizontal overflow at 390, 768, 1280 and
1440 px. See `docs/design/implementation-validation.md`.
- Final validation commands from this branch:
- `docker compose run --rm api pytest -q`**66 passed**.
- `docker compose run --rm api ruff check .`**All checks passed**.
- `docker compose run --rm api mypy app`**0 issues in 44 files**.
- `cd frontend && npm run lint` — clean TypeScript check.
- `cd frontend && npm run build` — production build succeeded (59 modules; 240.24 kB JS,
36.63 kB CSS before gzip).
- `cd frontend && playwright test --reporter=line`**19 passed** including the full
five-minute demo, every interactive route, return review semantics and four viewport
overflow checks.
- Review deployment updated at `http://192.168.10.150:1236` with persistent PostgreSQL
data preserved. Deployed smoke: all ten authenticated routes plus login at
desktop and mobile sizes rendered without alert state or horizontal overflow; browser console had zero
warnings/errors; seven authenticated API paths returned 200; PostgreSQL/API were
healthy and the shared n8n `/healthz` returned `{"status":"ok"}`.
- Corrected the review topology after confirming the host already runs n8n on port 5678:
the temporary `mobilityops-n8n-1` container was removed without deleting its retained
volume; the bundled service is now opt-in through the `bundled-n8n` profile; the API
points to the shared n8n; and the return workflow is imported and published there.
- The existing n8n's previously empty `N8N_HOST` and `N8N_EDITOR_BASE_URL` values were
persistently set in its Unraid template. A synthetic return then completed the full
MobilityOps → shared n8n → callback round trip as `succeeded` on attempt 1, after which
deterministic demo state was restored (`BK-DEMO-RETURN` is `active`).
- Global search is now live for Control Rail sections and `MO-*`, `BK-*`, `DQ-*` public
references, including Ctrl/Cmd+K focus and a tested not-found announcement. Final local
Playwright result is **19 passed**.
- Exact next action: hand off `design/mobilityops-premium-ui` for review. The final code,
shared-n8n topology and evidence are committed, pushed and deployed; do not merge master
automatically.
## Functional completion pass (branch `feat/mobilityops-functional-completion`)
Branched from `design/mobilityops-premium-ui` @ `54dc952`. Full audit at
`docs/functional-completion/current-functional-audit.md`; server baseline captured before
any change at `docs/functional-completion/server-baseline.md`.
### Batch 1 — complete (commits `938a739`..`bdc58f3`)
- Fixed the two confirmed list-rendering defects: Vehicles and Bookings both computed a
filtered/paginated result but rendered the raw unfiltered array in the table body.
- Added server-backed session lifecycle: `GET /api/v1/demo/session` (Cache-Control:
no-store — a cached 200 was making logout intermittently fail to redirect in e2e
testing), `POST /api/v1/demo/logout`. `AuthContext` now verifies against the server on
every mount instead of trusting `sessionStorage`, and a central 401 listener on the API
client clears auth state from any endpoint.
- Enforced the brief's role matrix: data-quality (list/detail/defer/reject) and the audit
trail were reachable by Rental Employee with no gate beyond authentication (confirmed
live via curl before the fix). Both are now `require_operations_manager`-gated
server-side, with matching nav-hiding and a restricted-message fallback for direct URL
access, and the dashboard no longer links into those areas for that role.
- Discovered and fixed a latent e2e-suite bug while testing against the real server: all
three spec files hardcoded `http://localhost:8128` for their demo-reset helpers, so
pointing the suite at Unraid via `MOBILITYOPS_PUBLIC_URL` silently kept resetting the
*local* dev database instead. Switched to relative paths so the configured `baseURL` is
honoured.
- Local evidence: `pytest` 75 passed, `ruff check .` clean, `mypy app` 0 issues/44 files,
`npx tsc -b` clean, `npm run build` clean, `npx playwright test` **25 passed** (up from
19 — 6 new tests this batch), stable across three repeated full-suite runs.
- Deployed to Unraid (`.deploy/source-revision` = `bdc58f396e99caaf6ef657bb110b479981cc7793`,
matches `git rev-parse HEAD` on the feature branch), migrations unchanged at
`e7b08389f47f (head)` (no schema change this batch), demo reset run. Re-verified live:
role-gate curl checks (403/200/401 as expected) and the full 25-test Playwright suite
run with `MOBILITYOPS_PUBLIC_URL=http://192.168.10.150:1236`**25 passed** against the
actual deployment, not just localhost.
- Exact next action: Batch 2 — authoritative return-preview endpoint shared with commit,
and expose `before`/`after` on the audit API + UI.
### Batch 2 — complete (commits `f521295`, `7e34f55`)
- Added `evaluate_return()` (pure, no writes) in `returns.py`, extracted from what
`register_vehicle_return` already computed inline; `register_vehicle_return` now calls
it instead of duplicating the logic. New non-mutating `POST
/api/v1/bookings/{ref}/return-preview` uses the same function, so preview and commit
cannot drift.
- Fixed a real defect this surfaced: `ReturnForm.tsx`'s review step guessed the outcome
client-side and got the domain rule wrong — it said damage/technical-warning routes to
`maintenance` (actual rule: `blocked`) and the no-contradiction case becomes
`available` (actual rule: always `cleaning` first, `maintenance` only past the service
threshold). The review step now calls `/return-preview` and renders the server's
`resulting_vehicle_status` + `status_reason` verbatim.
- Result screen now distinguishes local commit success from n8n delivery ("queued... not
yet confirmed" instead of implying both succeeded) and links to any created
data-quality issue for Operations Manager.
- Exposed `before`/`after` on `AuditEventOut` (the DB columns already existed but were
never serialized) plus a resolved `entity_ref`/`entity_link` for vehicle/booking/
data-quality-issue entities. `Audit.tsx` now shows a human-readable change summary per
row with raw JSON behind a `<details>` disclosure instead of always-visible JSON.
- New regression coverage: backend — preview performs no writes (asserted via audit/
outbox row counts before vs. after), detects odometer regression, detects service-due,
detects next-booking risk, requires an active booking, and matches the commit result;
audit — before/after and entity link exposed for both `return_registered` and
`vehicle_status_changed`. Frontend — preview correctly reports `blocked` (not
`maintenance`) for damage, commit request only fires after confirm (updated to also
assert exactly one preview call), audit page shows before/after and a safe link.
- Local evidence: `pytest` 81 passed, `ruff check .` clean, `mypy app` 0 issues/44 files,
`npx tsc -b` clean, `npm run build` clean, `npx playwright test` **27 passed**, stable
across two repeated full-suite runs.
- Deployed to Unraid and re-verified; demo data reset afterward.
- Exact next action: Batch 3 — data-quality workbench (typed snapshots, bounded
resolution flows for all 5 rule types, manual scan UI).
### Batch 3 — complete (commits `6e227a2`, `477b5e7`)
- Typed related-entity snapshots by the reference's own prefix (CUS-/MO-/BK-/INSP-)
instead of inferring from `rule_type`. Fixed a real gap this exposed: a
`booking_overlap` issue's related refs are bookings, but `get_issue` always resolved
them as vehicles, so `_snapshot()` silently returned nothing for them.
- Added one bounded resolution endpoint per remaining rule type: `provide-fields`
(missing_required_field; re-runs the check, resolves only once nothing required is
missing), `resolve-odometer-regression` (retain canonical or correct the reading —
correction is rejected if it would still be below canonical), `resolve-overlap`
(blocks one of the two bookings, re-verifies no overlap remains — found and fixed an
autoflush=False bug where the re-verification query didn't see the just-blocked
booking's in-memory status change), `apply-recommended-status` (one authoritative
recommendation function mirroring the scan's own conflict conditions, re-validated
after applying). `possible_duplicate_customer` already had merge; all five rule types
now have a real bounded resolution path, not just generic defer/reject.
- Reintroduced evidence after a non-open decision links the new issue back to the prior
one (`evidence.reopened_from` / `previous_decision`) per the documented lifecycle
("reintroduced evidence creates a new issue linked to the prior issue").
- `DataQualityIssueDetail.tsx` rewritten: a typed panel per rule type instead of a raw
`JSON.stringify` dump for four of five types; raw evidence moved behind a `<details>`
disclosure. Added a "Run quality scan" action to the workbench (confirmation,
progress, per-rule result counts, auto-refresh) — the scan endpoint already existed
with no UI trigger.
- New regression coverage: backend — one resolution test per rule type plus the
role-gate/validation-rejection paths and the recurrence-linking behavior (reject an
issue, rescan, assert the new issue links back). Frontend — one Playwright test per
resolution flow plus the manual scan trigger.
- Local evidence: `pytest` 96 passed, `ruff check .` clean, `mypy app` 0 issues/44 files,
`npx tsc -b` clean, `npm run build` clean, `npx playwright test` **32 passed**, stable
across two repeated full-suite runs.
- Deployed to Unraid and re-verified against the live server; demo data reset afterward.
- Exact next action: Batch 4 — global search backend + UI, demo reset UI trigger,
truthful aggregate integration status (n8n/RAGcore/MCP).
### Batch 4 — complete (commits `4437b87`, `1867828`)
- Added `GET /api/v1/search` — bounded typed results (vehicle/booking/data-quality-issue/
section), role-filtered server-side (data-quality and manager-only sections excluded
for Rental Employee), customers never returned (no customer detail route exists).
Replaced `Layout.tsx`'s blind client-side regex/term guesser with a debounced
(250 ms) call to this endpoint, a real `role="listbox"` results panel, arrow-key
navigation, Enter/Escape, outside-click close, and a no-results state.
- Added `GET /api/v1/integrations/status`, aggregating outbox delivery counts into one
truthful n8n state (`disabled`/`unavailable`/`degraded`/`operational`/`no_evidence`)
instead of the dashboard/automation cards showing whichever status the single most
recent event happened to be in. Wired into both `Automation.tsx` and `Dashboard.tsx`.
MCP Hub card now reflects the real `registration_enabled` setting.
- Found and fixed a real config gap this surfaced: `MCP_HUB_REGISTRATION_ENABLED` was
documented in `.env.example` but had no `Settings` field, so it was silently dropped
by `extra="ignore"` and never read anywhere in the codebase.
- Added a "Reset demo data" action to the sidebar (Operations Manager only, confirm,
progress, error handling) — the endpoint already existed and was already gated, just
had no UI trigger. Reset invalidates the acting session server-side, so the flow signs
the user out and returns to login.
- New regression coverage: backend — search role-filtering/customer-exclusion/no-match,
integration-status role-gate and state-derivation (including a test that resolves all
seeded failures and asserts the state flips to `operational`). Frontend — vehicle/
booking/data-quality-issue search navigation, keyboard nav, no-results + Escape, demo
reset happy path, rental employee cannot see the reset button, automation page shows
aggregate counts.
- Local evidence: `pytest` 109 passed, `ruff check .` clean, `mypy app` 0 issues/46 files,
`npx tsc -b` clean, `npm run build` clean, `npx playwright test` **37 passed**, stable
across two repeated full-suite runs.
- Deployed to Unraid and re-verified against the live server; demo data reset afterward.
- Exact next action: Batch 5 — bounded outbox delivery-lease recovery for stale
`delivering` events, a second (scheduled quality-scan) n8n workflow, final
documentation/contract updates and acceptance evidence.
### Batch 5 — complete (commits `ec8f809`, `e115031`, `c981aad`, `824048b`)
- Fixed a real gap: `_claim_due_events` flipped rows to `delivering` and committed
before the HTTP call, with no reclaim path if the process died before the outcome was
recorded. Each claim now gets a lease deadline (`n8n_delivery_lease_seconds`, default
120s, reusing the `next_attempt_at` column) and `run_dispatch_cycle()` sweeps expired
leases back to `pending` before claiming new work; `attempts` is preserved, and a
still-alive worker's unexpired lease is never touched.
- Added the second n8n workflow: `POST /api/v1/integrations/n8n/scheduled-scan`
(service-token protected, same pattern as the return callback) running the same
`run_scan()` the manual UI action uses, audited with `actor_type=service`.
`n8n/mobilityops-scheduled-quality-scan.json` (hourly + manual-test trigger) ships
`"active": false`. Live-verified twice: executed end-to-end via the Manual test
trigger against the **local** n8n instance (full green execution, confirmed via the
resulting `data_quality_scan_run` audit event), and published + directly
curl-round-tripped against the **shared Unraid n8n** and its live API
(`deploy/unraid/setup-scheduled-scan.sh`). The shared instance's own UI could not be
browser-tested directly — it runs `N8N_SECURE_COOKIE=true` and refuses login over the
plain-HTTP LAN URL, which is correct/expected shared-infrastructure behaviour, not
something this task should change.
- Updated `contracts/openapi.yaml` and `docs/05-api-contract.md` with every endpoint
added across all five batches; `docs/07-data-quality.md`, `docs/08-return-workflow.md`
and `docs/12-security-and-audit.md` now describe the actual resolution flows, the
preview/commit relationship, the role matrix and the audit before/after exposure.
Corrected `docs/07-data-quality.md`'s lifecycle description to match the
already-implemented `(rule_type, entity_type, entity_id)` idempotency key (no evidence
fingerprint) and documented the `reopened_from`/`previous_decision` recurrence link.
`README.md`'s scope/integration-status/quality-gate sections updated to match.
- Local evidence: `pytest` **117 passed**, `ruff check .` clean, `mypy app` 0 issues/46
files, `npx tsc -b` clean, `npm run build` clean, `npx playwright test` **37 passed**.
- **Clean-checkout drill** (section 14): fresh `git clone` of this branch into an
isolated scratch directory, `.env` from `.env.example`, isolated Compose project name
and remapped host ports (no shared state with the working stack), `up --build -d`
from empty volumes → migrations ran automatically → seed → full backend gate (117
passed, ruff clean, mypy clean) → `npm ci` (clean; the pre-existing esbuild-moderate/
react-router-RSC-high advisories are unchanged, not new) → `tsc -b`/`vite build` clean
→ full Playwright suite **37 passed** against the isolated stack. Torn down afterward
(`down -v` on the isolated project only; the working dev stack was never touched).
- Deployed to Unraid; migrations unchanged at `e7b08389f47f (head)`. Full 37-test
Playwright suite re-run against `http://192.168.10.150:1236`**37 passed**. Demo
data reset afterward.
- Exact next action: none — all five batches are implemented, tested locally (including
a genuine clean-checkout drill), committed, pushed, deployed to Unraid and
re-verified against the live server after every batch. See
`artifacts/functional-completion/final-summary.md` for the definitive acceptance
evidence.
## Demo productization (in progress, same branch `feat/mobilityops-functional-completion`)
Follows the functional-completion work above; turns the now feature-complete PoC into a
guided, honestly-labelled demo (fictional org "Northstar Mobility", guided tour, 5 named
scenarios, demo manifest, About page). Gap audit: `docs/demo-release/current-demo-gap-audit.md`.
### Batch 1 — seed date anchoring (complete)
- **Real bug fixed**: `seed/bookings.csv` etc. store absolute ISO timestamps authored
around a fixed anchor (`2026-08-01`). Nothing previously re-anchored them at seed/reset
time, so scenario bookings (e.g. `BK-DEMO-RETURN`) silently drifted into the past every
day the environment wasn't reset. `dashboard.py::_today()` compounded this by filtering
"today's movements" against the same frozen `demo_today` setting instead of real time.
- Fix: `seed_loader.py` now computes `shift = today - SEED_AUTHORED_ANCHOR` once per
`load_seed()` call and applies it to every seeded booking/inspection/maintenance/outbox
datetime column, so scenarios stay "today"/"near-future" relative to the actual reset
moment. `SeedResult` now also carries `anchor_date`/`seeded_at`; `POST /api/v1/demo/reset`
returns them; a `demo_data_seeded` audit event records the anchor for traceability.
`dashboard.py::_today()` switched from the frozen `demo_today` setting to real wall-clock
UTC date. The now-dead `demo_today` setting/env var was removed from `config.py`,
`compose.yaml`, `.env`, `.env.example` (nothing else referenced it).
- Added seed-validation tests (`backend/tests/test_seed.py`) proving S1 (`BK-DEMO-RETURN`/
`MO-024`), S2 (`CUS-0012`/`CUS-0178`/`DQ-DEMO-DUPLICATE`), S4 (`MO-016`/
`BK-DEMO-OVERLAP-A`/`-B`/`DQ-DEMO-OVERLAP`) and S5 (seeded failed outbox event
`00000000-0000-4000-8000-000000000020`, confirmed genuinely `failed` immediately after a
fresh reset, not silently auto-healed by the background dispatcher since it only claims
`pending` rows) are fully present after every reset, plus a dedicated anchoring test
asserting the shift and the audit marker.
- Live-verified locally: reseeded and confirmed via `psql` that `BK-DEMO-RETURN` now ends
today and `BK-DEMO-NEXT`/overlap bookings sit in the near future (today = 2026-08-03).
- Evidence: `pytest` **122 passed** (117 + 5 new/expanded seed tests), `ruff check .`
clean, `mypy app` clean (46 files, canonical `make` scope).
- Deployed to Unraid (commit `8989ffb`): pushed to Gitea, `git archive` tarball extracted
over `/mnt/user/appdata/mobilityops` preserving `.env`/volumes, `api` rebuilt (`db`/`web`
untouched — no frontend changes this batch), migrations confirmed at `e7b08389f47f
(head)`, reseeded, live-verified via `psql` that `BK-DEMO-RETURN`/`BK-DEMO-NEXT`/overlap
bookings sit at the same real-time-relative positions as local. `curl` to
`http://192.168.10.150:1236/` returns 200.
- Exact next action: `GET /api/v1/demo/manifest` + Dutch demo entry screen + permanent
demo badge (task #30), then the Demo Guide + scenario overview (task #31).
### Batch 2 — demo manifest, Dutch demo entry, permanent demo badge, About page (complete)
- `GET /api/v1/demo/manifest` (unauthenticated): single source of truth for demo org
identity, synthetic-data flag, reset allowance/timestamp/anchor date, guide
availability, and the 5 named scenarios with **live** readiness (queries the actual
`BK-DEMO-RETURN`/`DQ-DEMO-DUPLICATE`/`DQ-DEMO-OVERLAP`/seeded-failed-event/knowledge-
provider records — not hardcoded), plus plain-language integration summaries. Backed by
new `backend/app/services/demo_manifest.py`. Refactored the n8n status derivation out of
`integration_status.py` into a shared `services/integration_status.py` so the manifest
and the existing authenticated `/integrations/status` endpoint reuse one implementation.
- New settings (`backend/app/core/config.py`, wired through `compose.yaml`/`.env.example`):
`DEMO_ORGANIZATION_NAME` (default "Northstar Mobility" — surfaces the project's already-
locked fictitious tenant, previously only used internally as the `ragcore_tenant` slug),
`DEMO_TIMEZONE`, `DEMO_ALLOW_RESET` (a safety valve — `false` makes `POST
/api/v1/demo/reset` return 403 regardless of role; the now-dead `demo_today` setting
removed in Batch 1 stays removed).
- Rewrote `Login.tsx` in Dutch: names the fictional org, one-sentence explanation sourced
from the manifest, no password shown/copyable anywhere, "Start begeleide demo" primary
CTA (logs in as Operations Manager, navigates to `/dashboard?guide=start` for task #31 to
consume) plus "Verken als Operations Manager"/"Verken als Rental Employee" secondary
actions. Added a permanent demo badge (topbar pill + popover: synthetic notice,
"workflows are real" reassurance, last-reset timestamp, link to `/about`) replacing the
old full-width static `.demo-banner` bar — subtle by design per the brief, not a warning
bar. New `/about`
page (`AboutDemo.tsx`) covering the fictional problem, what's really implemented, what's
synthetic, honest per-integration labels (via the manifest), and a reset pointer —
reachable from the badge popover, not added to primary nav (preserves the existing Control
Rail nav per the "not a redesign" constraint). Frontend nav/design otherwise untouched.
- Evidence: `pytest` **127 passed**, `ruff check .` clean, `mypy app` clean (48 files);
frontend `tsc -b` clean, `npm run build` clean; full Playwright suite **41 passed**
(37 existing + 4 new `demo-entry.spec.ts` covering entry copy/no-password, guided-demo
login redirect, badge popover content + About link, and Escape/outside-click close).
Updated stale English login-button aria-labels and login-copy assertions across the
existing specs to match the new Dutch copy.
- Deployed to Unraid (commit `ac427f4`): pushed to Gitea, `git archive` tarball extracted
preserving `.env`/volumes, both `api` and `web` rebuilt (frontend changed this batch),
both healthy, migrations unchanged, reseeded. Live-verified: `GET
/api/v1/demo/manifest` returns `organization_name: "Northstar Mobility"`,
`allow_reset: true`, and all 5 scenarios `ready: true` right after reset. Ran
`demo-entry.spec.ts` (4 tests) and the five-minute demo script directly against
`http://192.168.10.150:1236`**5/5 passed**. Reseeded again afterward to leave the
server demo-ready.
- Exact next action: Demo Guide (collapsible panel, 8 steps) + scenario overview (5 cards
on the dashboard, consuming `/api/v1/demo/manifest`'s `scenarios` array) — task #31.
### Batch 3 — Demo Guide + scenario overview (complete)
- New `/scenarios` page (`Scenarios.tsx`): all 5 named scenarios as cards (title,
operational problem, duration, required role(s), "toont aan", ready/blocked status
from the manifest, "Start scenario" linking to the live `start_path`). Dashboard gets
one compact "Probeer een demonstratiescenario" panel (not 5 more cards — keeps the
existing dashboard uncluttered per the brief) showing readiness count and, for
Operations Managers, a guide resume/start control.
- New Demo Guide: `DemoGuideContext` (sessionStorage-persisted `currentIndex`/`completed`
set — browser-only, never touches auth or business logic), 8 static steps
(`data/demoGuideSteps.ts`) each with what-you'll-see/why/start-action/expected-outcome,
resolving live routes from the manifest for the two scenario-backed steps (return,
duplicate-merge) so they can't drift from actual records. `DemoGuide.tsx` renders a
fixed side panel (desktop) that becomes a bottom sheet at ≤700px via CSS only (no
layout duplication); `DemoGuideTrigger` (topbar, Operations-Manager-only — the 8 steps
require OM throughout) shows a live `completed/8` pill. "Demo opnieuw voorbereiden"
calls the real reset endpoint, resets guide progress, and returns to `/login` (mirrors
the existing sidebar reset flow). Login's "Start begeleide demo" logs in as OM and
passes a one-shot `?guide=start` marker the dashboard consumes once then strips.
- Fixed a real regression caught by the responsive-overflow tests: the new topbar guide
trigger pushed `.topbar-meta` past the viewport at ≤420px; fixed by hiding the guide
trigger (icon+pill) at that breakpoint — the Dashboard's own "Start demo-gids" control
remains reachable there. Also fixed a genuine mobile overflow in the new
`.demo-start-panel` (flex items without `min-width:0`/wrap on narrow screens).
- Fixed one fragile new test (asserted on the transient `?guide=start` URL param, which
the app intentionally strips immediately — changed to assert the guide's actual open
state instead) and two Playwright strict-mode ambiguous-match errors; confirmed the
full 41+6=47-test suite passes twice in a row after these fixes (ruling out flakiness).
- Evidence: frontend `tsc -b` clean, `npm run build` clean; full Playwright suite
**47 passed** (41 existing + 6 new `demo-guide.spec.ts`: scenario overview shows 5
ready cards after reset, starting a scenario navigates to its fixed record, guide
step navigation/jump/close, progress persists across page navigation, guide hidden
from Rental Employee, restart-from-guide resets data and returns to login). Backend
untouched this batch (no re-run needed; last backend gate was 127 passed/ruff/mypy
clean in Batch 2).
- Deployed to Unraid (commits `9fff84d`, then `14c2ad3` for a test-only fix): pushed to
Gitea, tarball extracted, `web` rebuilt (frontend-only batch), healthy. Live-verified:
ran `demo-guide.spec.ts` (6), `demo-entry.spec.ts` (4) and the five-minute demo script
directly against `http://192.168.10.150:1236`**11/11 passed**. Caught and fixed one
real environment-sensitive test bug in the process: two tests navigated straight to
`/scenarios` right after a login click without waiting for the `/dashboard` redirect,
which raced harmlessly on localhost but flaked against Unraid's higher latency — fixed
by asserting the redirect first, no app-code change needed. Reseeded afterward to leave
the server demo-ready.
- Exact next action: layer plain-language Dutch explanation onto the return flow, the 5
data-quality panels, and fix the knowledge assistant's RAGcore-naming bug — task #33.
### Batch 4 — return/data-quality/knowledge demo legibility (complete)
- **Fixed a real honesty bug**: `Knowledge.tsx` named "RAGcore" in the body copy and the
retrieval-flow diagram even though the active provider is the demo TF-IDF one (the
small badge below was already honest, contradicting the prose one line above). Now
derives a `providerLabel` ("Demo knowledge base" vs "RAGcore") from the real health
check and uses it everywhere; added an explicit disclosure note when not RAGcore.
Added 4 suggested-question chips. **Discovered and fixed a second real bug in the
process**: the brief's suggested Dutch questions (and my own Demo Guide step 6 wording)
would have returned "insufficient evidence" against the demo provider, because the
indexed procedures are English-only — verified empirically (Dutch question →
`insufficient`, its English equivalent → `grounded`). Fixed by keeping suggested
questions in English (matching the indexed content) and rewording the Guide step to
explain the knowledge base is English, rather than mistranslating the demo's
centerpiece feature into silently returning wrong answers.
- Return flow: `BookingDetail.tsx` now detects the one named return-anomaly scenario
booking (via the manifest, not a hardcoded ref) and fetches that vehicle's real
canonical odometer to pre-fill `ReturnForm`'s "End odometer" field with a suspicious
value below it, plus a callout explaining why — the brief explicitly requires the demo
not ask a visitor to invent a suspicious number themselves. Scoped narrowly to that one
scenario booking; ordinary returns are unaffected. `ReturnResultPanel` now links to
Automation and Audit trail (previously only the vehicle), and shows a "Ga verder met de
demo" button when the Demo Guide is open (advances the guide and navigates to the next
step). **Fixed a real regression caught by the existing return-review e2e test**: the
async pre-fill could silently overwrite odometer text a visitor had already started
typing, if the vehicle-detail fetch resolved after they began typing — fixed with an
`odometerEditedByUser` ref guard.
- Data quality: added a shared `RuleExplainer` (what's wrong / why it matters, in plain
language) for all 5 rule types on `DataQualityIssueDetail.tsx`; added a generic
post-resolution confirmation (audit-trail link, vehicle link, "Ga verder met de demo")
for the 4 rule types that previously just silently flipped their status badge with no
explicit confirmation, and extended `VehicleStatusConflictPanel`'s existing confirmation
with the same links rather than duplicating it. Added a "Demo scenario's only" checkbox
filter on `DataQuality.tsx` (client-side `public_ref.startsWith("DQ-DEMO-")`, no new
business logic) so the curated issues are easy to find among the full queue.
- Evidence: frontend `tsc -b` clean, `npm run build` clean; full Playwright suite
**51 passed** (47 existing + 4 new `demo-legibility.spec.ts`: return pre-fill + why-
suspicious explanation + result links, rule explainer visible, demo-scenario filter
narrows correctly, knowledge suggested question returns grounded evidence with the
correct provider label). Backend untouched this batch.
- Deployed to Unraid (commit `ddc3a98`): pushed to Gitea, tarball extracted, `web`
rebuilt (frontend-only), healthy. Live-verified: ran `demo-legibility.spec.ts` (4) and
the five-minute demo script directly against `http://192.168.10.150:1236`
**5/5 passed**. Reseeded afterward to leave the server demo-ready.
- Exact next action: plain-language integration-status labels, richer audit narration,
the full "Over deze demo" page content (currently a first pass from Batch 2), and
wiring reset into the guide/About/OM menu narrative — task #34.
### Batch 5 — integration-status UX, audit UX, About page, reset integrity (complete)
- Plain-language integration status: extracted `frontend/src/data/integrationLabels.ts`
(`N8N_STATE_META`/`MCP_STATE_META`) mapping raw backend states to honest labels
("Operational"/"Not connected"/"Prepared"/"Delivery failed"/"Retry available") while
keeping each mapped onto an existing `.status-*` CSS colour class (a few raw values like
`degraded`/`disabled`/`configured` had no matching CSS rule at all before this — a real,
pre-existing colour-coding gap). `StatusBadge` gained an optional `label` override prop
(backward compatible) so the badge's colour class and its displayed text can differ.
Wired into both `Automation.tsx` and `Dashboard.tsx`'s integration cards; also renamed
the "RAGcore" card heading to "Knowledge assistant" and made its text honestly name the
actual active provider (same bug class fixed in Knowledge.tsx in Batch 4).
- Audit trail: added a "Follow-up" column with a "View related events" action per row that
filters the same list by `correlation_id` (reuses the backend's existing, already-tested
`correlation_id` query param — no new business logic), with a "Clear this filter"
affordance. This is how a visitor sees "what else happened as a result of this action"
(e.g. a return's linked vehicle-status-changed / workflow-queued events) without a
bigger grouped-timeline rebuild.
- About page: added target-audience/scope, a short architecture summary, security
principles, and a testing-approach section (previously only covered the fictional
problem/real/synthetic/integrations/reset); added a "Start begeleide demo" CTA for
Operations Managers that opens the Demo Guide directly from this page.
- Reset integrity: added `scenario_integrity_report()` (`backend/app/services/
demo_manifest.py`), reusing the exact same scenario-readiness derivation the manifest
and scenario overview already use (so it can't drift), and wired it into `POST
/api/v1/demo/reset` — both the response body and the `demo_reset` audit event's
metadata now carry `scenario_integrity: {all_ready, not_ready}`. This is the
server-side post-reset integrity check the brief asks for; visible today via the audit
event's raw-detail view, satisfying the requirement without adding a UI banner to a
flow that immediately logs the user out and redirects to `/login`.
- **Fixed a second real regression this batch, caught by the existing return-review
e2e test**: restructured the odometer pre-fill so `BookingDetail.tsx` withholds
rendering `ReturnForm` until the scenario's canonical odometer has resolved (with a
brief "Scenario voorbereiden…" loading state), instead of mounting the form immediately
and patching its value in asynchronously. The previous approach raced visibly with
Playwright's `fill()` (and would have raced with a real visitor typing quickly),
producing a corrupted concatenated value in one observed failure. This also let the
now-unnecessary `odometerEditedByUser` ref guard be removed — simpler and more robust
than the effect-based patch it replaced.
- Evidence: `pytest` **127 passed**, `ruff check .` clean, `mypy app` clean (48 files);
frontend `tsc -b` clean, `npm run build` clean; full Playwright suite **51 passed**,
confirmed stable across three consecutive full runs (given how many timing races this
batch and the previous one surfaced, stability was verified deliberately rather than
assumed from a single green run).
- Deployed to Unraid (commit `5fa4fe0`): pushed to Gitea, tarball extracted, both `api`
and `web` rebuilt, healthy, migrations unchanged at `e7b08389f47f (head)`, reseeded.
Live-verified: ran `demo-legibility.spec.ts` (4), the five-minute demo script, and the
full `interactive-elements.spec.ts` suite (26) directly against
`http://192.168.10.150:1236` — **31/31 passed**. Reseeded afterward to leave the
server demo-ready.
- Exact next action: full guided-demo Playwright test + remaining targeted demo tests per
section 19 (mobile guide, keyboard nav, all scenario flows, About page, accessibility/
reduced-motion/console/network checks) — task #35.
### Batch 6 — full guided-demo test + targeted demo tests (complete)
- **Found and fixed a real, fairly serious desktop layout bug** while writing the full
guided-demo test: the Demo Guide's fixed right-side panel (400px wide) overlapped the
main content area at normal desktop widths with no reflow, so its own step-list buttons
intercepted pointer events meant for the page underneath (concretely: the return form's
"Review return" button was unclickable while the guide was open, at exactly the
viewport size Playwright's default test browser uses — this would have hit real
visitors on ordinary laptop screens too). Fixed by adding a `guide-open` class to
`.app-workspace` that reserves `padding-right: min(400px, 92vw)` while the guide is open
(≥701px only; the ≤700px bottom-sheet layout is unaffected), so content reflows aside
instead of sitting underneath the panel.
- Added `frontend/e2e/guided-demo-full.spec.ts`: one comprehensive test walking a fresh
Operations Manager session through all 8 Demo Guide steps in order, performing the
**real** action at each step (not just verifying copy) — processes the actual
odometer-anomaly return, resolves the resulting data-quality issue, merges the
duplicate customer, asks a suggested knowledge question, checks automation + audit,
reviews the About page — using the guide's own progression controls
("Volgende"/"Ga naar deze stap"/"Ga verder met de demo") throughout, then resets the
demo data again at the end to restore the environment per the brief's requirement.
- Added `frontend/e2e/demo-accessibility.spec.ts` (4 tests): the guide renders as a
correctly-anchored bottom sheet on a 390px mobile viewport with no horizontal overflow;
the guide never covers the return form's action buttons on desktop (regression test for
the bug above); the demo badge and guide trigger are keyboard-focusable and operable
(Enter to open, explicit close controls); key demo pages (dashboard, scenarios, about,
guide open) load with no unexpected console errors (the one expected benign 401 from
the app's own session-probe on first load is explicitly allow-listed, not silenced
blindly).
- Evidence: full Playwright suite **56 passed** (51 existing + 1 guided-demo-full + 4
demo-accessibility), confirmed stable across two consecutive full runs. Backend
untouched this batch (last gate: 127 passed/ruff/mypy clean, Batch 5).
- Deployed to Unraid (commit `07d5605`): pushed to Gitea, tarball extracted, `web`
rebuilt (frontend-only), healthy, reseeded. Live-verified: ran
`guided-demo-full.spec.ts` and `demo-accessibility.spec.ts` directly against
`http://192.168.10.150:1236` — **5/5 passed**, confirming the desktop-overlay layout
fix holds on the real deployment too. Reseeded afterward to leave the server
demo-ready.
- Exact next action: clean-checkout demo drill, final documentation set (demo-concept/
demo-scenarios/demo-data/demo-guide/demo-runbook, README, .env.example), final Unraid
deploy + live evidence with screenshots, `artifacts/demo-release/final-summary.md` —
task #36 (final).
### Batch 7 (final) — clean-checkout drill, docs, final Unraid evidence (complete)
- **Clean-checkout drill**: fresh `git clone` into an isolated scratch directory,
isolated Compose project (`mobilityops-cleandrill`) + remapped ports via
`compose.override.yaml`, `up --build -d` from empty volumes. Migrations ran
automatically to `e7b08389f47f (head)`; seeded; full backend gate **127 passed**,
ruff/mypy clean; `npm ci` clean (same pre-existing advisories as before, unchanged);
`tsc -b`/`vite build` clean; full Playwright suite **56 passed** against the isolated
stack; reseeded and confirmed all 5 scenarios `ready: true` via the manifest; torn down
(`down -v` on the isolated project only — the working dev stack was untouched
throughout).
- Added the full demo-release documentation set: `docs/demo-release/demo-concept.md`,
`demo-scenarios.md`, `demo-data.md`, `demo-guide.md`, `demo-runbook.md`; updated
`README.md` (current test counts, links to the new docs, a "Demo" section) and
`docs/17-runbook.md` (cross-reference to the demo-specific runbook).
- Added `frontend/e2e/_capture-demo-screenshots.spec.ts` (tooling, excluded from the
regular suite) and captured 17 evidence screenshots live against
`http://192.168.10.150:1236` into `artifacts/demo-release/screenshots/`.
- Final live acceptance: full Playwright suite re-run against the live server —
**56 passed**; `docker compose ps` on the server shows `api`/`db`/`web` all healthy;
`docker logs` for `api`/`web` show no errors; reseeded to leave the server
demo-ready after evidence capture.
- Wrote `artifacts/demo-release/final-summary.md` with the full required evidence
(branches/commits, org/roles/guide/scenarios, seed/date-anchor/reset strategy, real vs.
synthetic vs. not-connected, all test results, clean-checkout result, deployment/
health/console/log results, responsive/accessibility results including the two real
layout bugs found and fixed this work (mobile topbar overflow in Batch 3, desktop
guide-panel overlap in Batch 6), known limitations, 5-/10-minute demo flows, redeploy/
rollback commands, and the screenshot list).
- Demo-productization work on this branch is complete. Every task (#29#36) is done;
every batch was tested locally, deployed to Unraid, and re-verified live before moving
to the next. See `artifacts/demo-release/final-summary.md` for the definitive
acceptance evidence.
## Final product polish: Fleet Ops rebrand, trilingual i18n, adaptive guide (2026-08-03) — MERGED TO MASTER
- Rebranded the product to **Fleet Ops** across the frontend, backend defaults and the
knowledge base; made `nl-BE` (default)/`en-GB`/`fr-BE` full first-class languages via
i18next (eager-bundled resources, persisted language switcher in topbar + mobile
drawer, `Intl` date/number formatting, a coverage test that fails the build on any
missing/empty translation key across all 14 namespaces).
- Backend dynamic content (demo scenarios, blocked-reason text, integration status)
converted from fixed English/Dutch prose to stable message codes + params so the
frontend localizes it (`DemoScenarioOut`/`DemoIntegrationSummaryOut` schema changes).
The demo knowledge base gained a fully translated NL/EN/FR procedure corpus (11
documents each, including a new "vehicle availability" procedure) with per-language
retrieval and localized evidence-state messages.
- Demo Guide became breakpoint-adaptive: docked rail (≥1440px), a floating panel that
auto-collapses to a persistent closable progress chip (7011439px), and a
collapsed/half/full bottom sheet (≤700px) — with scroll+focus+highlight on "go to this
step", Escape handling, and `prefers-reduced-motion` support.
- Data Quality Workbench got accessible choice-card decisions with a clear
primary/secondary/tertiary action hierarchy; Automation ledger groups repeated
successes with meaningful short refs; Audit trail groups events by correlation id with
human action labels and readable before/after diffs; Attention Queue/Today's
movements/Vehicles/Bookings/Data Quality rows are fully clickable (stretched-link
pattern, independent secondary links, keyboard + mobile support).
- Two real bugs found and fixed along the way: a mobile topbar overflow at 421440px
caused by the new language switcher (moved the switcher into the mobile drawer at
≤960px and widened the compact-topbar breakpoint to 440px), and two dangling
`aria-labelledby` references (`SectionHeading` never set the referenced `id`).
- Full test suite: 131 backend tests, Ruff, mypy, TypeScript build, and 92 Playwright
tests (new: `i18n-coverage`, `clickable-rows`, `responsive-i18n` covering all 7
brief-specified breakpoints × 3 languages, plus 3 new adaptive-guide tier tests) — all
green. All pre-existing Playwright specs updated for the new nl-BE default (either
translated assertions or an explicit English-locale override where the spec was
originally authored against English copy).
- Clean-checkout drill performed in a fully isolated Docker Compose project (separate
ports/volumes, no shared n8n) from a fresh local clone at the feature-branch head —
131 backend tests, lint, build and all 92 Playwright tests green from empty volumes;
live EN/FR knowledge-assistant spot check; `scenario_integrity.all_ready: true` on
reset; isolated stack torn down afterward, original dev environment untouched.
- Deployed to Unraid twice: once for the feature branch (commit `845db14`) for
pre-merge live validation, once for the merged `master` (commit `18a765d`) for the
final release — both times via the established `git archive` → `scp` →
extract → `.deploy/source-revision` → rebuild `api`/`web` method, with migrations,
reseed, full backend+Playwright gates, console/network inspection and demo reset
re-verified live each time.
- Master baseline was confirmed unchanged (`e0c7ed6`, matching the previously recorded
baseline) before merging; `git merge-tree` dry run showed zero conflicts. Merged via
`git merge --no-ff` (commit `18a765d`), all gates re-run post-merge, pushed to Gitea,
redeployed. Feature branch was not deleted.
- Full evidence: `artifacts/fleet-ops-release/final-summary.md` (commits, branding,
locales, translation/knowledge-base/guide/data-quality/automation/audit evidence, all
test results, clean-checkout result, deployment evidence for both the feature branch
and master, responsive/accessibility results, known limitations, rollback procedure)
plus 10 screenshots in `artifacts/fleet-ops-release/screenshots/`.
## Fleet Ops correction: safe status-recommendation flow, MO-016, message codes (2026-08-03) — MERGED TO MASTER
Branch `fix/fleet-ops-i18n-status-flow`, created from master's post-release head
(`18344bc`). Audit and rationale in `docs/fleet-ops-correction/` (gap audit, i18n
inventory, vehicle-status decision table). Merged to master via `de0bdea` ("merge:
complete Fleet Ops localization and status resolution"), with final evidence commit
`f780557` ("docs(release): final Fleet Ops correction evidence and screenshots") —
`f780557` is `origin/master`'s current head as of the start of the correction round
below.
- **Status-recommendation flow redesigned** per the brief: the old single opaque
"calculate and apply recommended status" action is replaced by a single shared, pure
evaluator (`backend/app/services/vehicle_status.py::evaluate_vehicle_status`) used
identically by the scanner, a new non-mutating preview endpoint
(`POST .../status-recommendation`), and a transactional apply endpoint
(`POST .../apply-recommended-status`) that locks the row, recomputes facts, rejects a
stale `recommendation_token` (optimistic concurrency), refuses unsafe/manual-review
recommendations, and re-validates post-write before resolving the issue. Frontend
`DataQualityIssueDetail.tsx` shows "Review recommendation" → a decision panel
(current/recommended status, why, evidence, consequences, localized in all 3
languages) → an exact "Change status to <status>" confirm action → result, with a
distinct "Manual review required" state offering no generic apply button.
- Fixed the real unsafe shortcut this evaluator exists to eliminate: "maintenance +
active booking" no longer auto-recommends "rented" (current status is itself now a
blocking fact), and "maintenance with nothing else wrong" no longer auto-clears to
"available" (no fact proves maintenance is actually finished — that release stays a
manual decision).
- **MO-016 order independence**: order independence does not mean "same final status
regardless of order" — resolving the booking overlap first genuinely removes the
conflict, correctly leaving nothing to apply. What must (and does) hold either way:
the recommendation always reflects real current facts, and nothing unsafe is ever
applied (never "rented"). Verified by both a backend test
(`test_mo_016_status_conflict_recommendation_is_order_independent`, explicitly scoped
to MO-016/DQ-DEMO-STATUS after finding the original version wasn't) and a browser-level
Playwright test in both orders.
- **"Fleet Ops" is a non-localizable brand constant** (`frontend/src/product.ts`,
backend `PRODUCT_NAME`), wired via `{{productName}}` interpolation everywhere the
brand appeared in locale prose; a permanent test fails the build if any locale file
ever defines the brand name or an `appName` key again.
- **Dynamic backend prose converted to message codes + params**: return status reasons,
audit field/actor-type labels, automation `last_error` (new `last_error_code` column,
migration `799d8800e241`), and search results (sections/vehicles/bookings/issues) all
now carry stable codes the frontend localizes; raw technical text is demoted to a
"Technical details" disclosure everywhere.
- **Knowledge-base fixes**: the demo provider's tokenizer silently dropped accented
characters (`[a-z0-9]+` split "véhicule" into "v"+"hicule"), breaking French
retrieval broadly; fixed to include the Latin-1 accented range. Also reweighted
section scoring so a body match (real substance) outranks a heading/title match (a
shallow structural hint) — the old weighting misranked the damage procedure behind
an unrelated document for the brief's exact validation question in all 3 languages.
Removed leftover "MobilityOps"/"PoC" mentions from 9 procedure documents (knowledge
prose is visible content, missed by the earlier rebrand).
- New `frontend/e2e/fleet-ops-correction.spec.ts` (14 tests) covers branding in 3
languages, language persistence, the full status-recommendation flow (non-mutating
preview, exact confirm text, manual review, stale-token rejection), MO-016 order
independence, trilingual knowledge grounding, and localized audit/automation. Writing
it surfaced and fixed two real bugs: the frontend conflated "no conflict" with
"manual review required" (both carry `safe_to_apply: false`), and the original
MO-016 backend test never actually targeted MO-016's own issue.
- Added a keyboard/reduced-motion/no-color-only-status accessibility test for the new
status-decision panel; added `aria-live="polite"` to the panel so the applied
confirmation is announced.
- `contracts/openapi.yaml` and `README.md` updated: title is "Fleet Ops", the new
status-recommendation endpoint documented, apply-recommended-status's request body
and error codes documented, search endpoint's code+params shape documented, README
states the Fleet Ops/MobilityOps naming split explicitly and refreshes stale test
counts (151 backend, 108 Playwright).
- Gates green: 151 backend tests, Ruff, mypy, Alembic upgrade/downgrade verified,
frontend `tsc`/build, full 113-test Playwright suite (rebuilt `api`+`web` containers
each time before testing).
- Section 11D/E/F of the i18n test-strengthening brief done: a hardcoded-JSX-text
static check (`i18n-coverage.spec.ts`; had to anchor on backreferenced closing-tag
names — a naive `>text<` scan misread TypeScript generics like
`useState<string | null>` as JSX spanning to the next unrelated `>`; verified against
both false positives and a deliberately-injected-then-reverted false negative), and a
3-language route matrix (`fleet-ops-correction.spec.ts`) covering every main route:
no console errors, correct `html[lang]`, real page headings.
- **Clean-checkout drill (2026-08-03) — PASS.** Fresh `git clone --branch
fix/fleet-ops-i18n-status-flow` of only committed files into an isolated directory,
separate Compose project name and host ports (8129/1229/5679) so the working dev
stack was never touched. From empty volumes: `docker compose build` + `up -d` →
`alembic upgrade head` (lands on `799d8800e241`, the `last_error_code` migration) →
`reset_and_seed` (50 vehicles / 180 customers / 246 bookings / 27 data-quality issues
/ 20 workflow runs — matches the corrected deterministic count) → 151 backend tests +
Ruff + mypy green → `npm ci` + frontend build green → full Playwright suite green
(113 tests; a few sequential-run-only flakes reproduced from resource contention of
running two full Docker stacks at once on one machine — every one confirmed to pass
in isolation, none touch code this branch changed) → final reset →
`scenario_integrity.all_ready: true`. Isolated stack, containers, volumes and images
torn down afterward; original dev environment confirmed untouched and reset to
baseline.
- **Unraid deployment (2026-08-03/04) — PASS.** Pushed `fix/fleet-ops-i18n-status-flow`
to origin, deployed via `git archive` → `scp` → extract into
`/mnt/user/appdata/mobilityops` (preserving `.env`) → `.deploy/source-revision` →
rebuild `api`+`web` → `alembic upgrade head` → reset/reseed, at
`http://192.168.10.150:1236`. **Live validation directly caught a real bug**: every
data-quality issue's top-of-page evidence summary was unconditionally showing raw
English (e.g. "vehicle marked available while reserved bookings conflict") in all
three languages, because the frontend never finished the evidence.signals
localization the backend had already been emitting. Fixed (commit `2e4fb43`):
DataQualityIssueDetail.tsx now renders `evidence.signals` through the operator's
locale as the primary text, raw text moved to "Technical details" only, the 4
DQ-DEMO-* seed rows got real computed signals (the duplicate-customer similarity
score is the actual SequenceMatcher ratio on the seeded names), and a regression
test locks this in. Redeployed with the fix; live-verified via `read_page` that
DQ-DEMO-STATUS now shows "Dit voertuig heeft twee overlappende reserveringen..."
instead of the raw English sentence. Full 116-test Playwright suite green against
the live server (`MOBILITYOPS_PUBLIC_URL=http://192.168.10.150:1236`), no console
errors, no errors in `api`/`web` container logs, both containers healthy, final
reset done, `scenario_integrity.all_ready: true`.
- Final evidence: `artifacts/fleet-ops-correction/final-summary.md`. Merged to master
via `de0bdea`, followed by evidence commit `f780557` on master. See the "Fleet Ops
final localization" entry below for the next (small correction) round on top of this.
## Fleet Ops final localization: remaining NL/FR gaps, API-error localization, greeting (2026-08-04) — MERGED TO MASTER
Merged to master via `5f0eaa5`; final evidence commit `c0995b7` added
`artifacts/fleet-ops-final-localization/final-summary.md`. Master head at merge:
`c0995b762e1cbf37172a08e03645baa6b66aa8d5`. Details below are the in-progress working log
kept for reference.
Branch `fix/fleet-ops-final-i18n-ux`, created from master's post-correction head
(`f780557`) — the brief asked for `fix/fleet-ops-final-localization`, but the
already-checked-out branch name is used instead since it was verified freshly and
cleanly branched from current `origin/master` with a clean working tree; see
`docs/fleet-ops-final-localization/audit.md` for the naming note. Scope: a small,
targeted correction round only — explicitly not touching status-flow business logic,
the status evaluator, Data Quality resolution rules, return rules, RAGcore/MCP Hub, or
product scope.
- **Audit-driven gap sweep**: `docs/fleet-ops-final-localization/audit.md` documents
every remaining untranslated/incorrect string, raw-backend-error call site,
over-permissive allowlist entry, the static-greeting bug, and doc staleness found by
a dedicated Explore pass before any file was touched.
- **Remaining NL/FR translation gaps fixed**: role names actually translated (not just
labelled as translated) — `auth.json`/`demo.json` role keys, `audit.title` →
"Auditgeschiedenis"/"Piste d'audit", `columns.actor` → "Uitvoerder", `list.statusOpen`
→ "Openstaand", `ledger.filterRecent` → "Recentste", `scenarios.startScenario` →
"Scenario starten". Also found and fixed (via the new embedded-substring test below)
8 previously-missed mid-sentence "Audit trail" leaks across `demo.json`,
`quality.json`, `returns.json` that the old whole-string-identity test structurally
could not catch.
- **Central API-error localization**: new `frontend/src/api/errorMessages.ts`
(`describeApiError`) replaces the `err instanceof ApiError ? err.message : ...`
anti-pattern (which showed raw English for the common case) at all 13 call sites
across 7 files. Raw backend text is now only ever shown under a "Technical
details"/"Détails techniques" disclosure (new `ApiErrorNotice` component in
`PageChrome.tsx`); the primary message is always a localized title + explanation +
optional next step, keyed on the 32 known `AppError` codes, then known HTTP statuses
(401/403/404/409/422/500), then a fully generic fallback. `ApiError` was split out of
`client.ts` into a standalone `api/apiError.ts` (no `import.meta.env` dependency) so
`errorMessages.ts` is independently testable outside a Vite/browser context.
- **i18n allowlist tightened**: removed 7 now-stale `IDENTICAL_VALUE_ALLOWLIST` entries
in `i18n-coverage.spec.ts` (`audit.title`, `auth.roleOperationsManager`,
`auth.roleRentalEmployee`, `demo.scenarios.startScenario`,
`demo.scenarios.roles.operations_manager`, `demo.scenarios.roles.rental_employee`,
`navigation.items.audit`) now that they're genuinely translated. Added 2 new tests:
one closing the embedded-English/Dutch-substring blind spot (mid-sentence phrase
leaks the whole-string check misses), one asserting no locale file contains
"MobilityOps" or the word "PoC".
- **`describeApiError` test coverage**: new `frontend/e2e/error-messages.spec.ts` (10
tests) — every known code/HTTP status has non-empty copy in all 3 locales, a known
code never surfaces raw backend text as the primary message (only via `.technical`),
unknown-code and unknown-status fallback chains behave correctly, and a drift guard
that greps the actual backend `AppError("CODE", ...)` call sites and fails if
`KNOWN_CODES` and the backend's real codes ever diverge (currently exactly in sync,
32 codes).
- **Time-dependent Europe/Brussels dashboard greeting**: new
`frontend/src/i18n/greeting.ts` (`getGreetingPeriod`, DST-safe via
`Intl.DateTimeFormat({ timeZone: "Europe/Brussels", hourCycle: "h23" })`, clock
injectable) + `useGreetingPeriod.ts` hook (30s poll for period rollover while the app
stays open, no reload). Replaces the previously-always-"Goedemorgen" static
`dashboard.json` title with 4 periods × 3 languages for both the greeting word and a
varying accompanying sentence (never "Goedenacht"). Tests: `greeting.spec.ts` (pure
boundary/DST unit tests) + `greeting-live.spec.ts` (6 real-browser tests via
Playwright's `page.clock` — all 8 required boundary times in all 3 languages, live
rollover without reload, language-switch behaviour, the "never Goedenacht" guard).
- **Found and fixed one real CSS regression along the way**: correctly translating
`roleOperationsManager` to the single unbreakable Dutch compound word
"Operationsmanager" (vs. the old two-word "Operations Manager", which could wrap)
pushed the topbar's `.operator` block past 1024px width, caught by the existing
`responsive-i18n.spec.ts` overflow test. Fixed with `overflow-wrap: anywhere` on
`.operator strong`/`small` and `min-width: 0` on their flex-item wrapper, not by
reverting the correct translation.
- Gates green so far: backend `pytest` 151 passed, `ruff check .` clean, `mypy app`
clean (49 files, unchanged — no backend Python touched this round); frontend `tsc`
clean, production build clean, full local Playwright suite **138 passed** (rebuilt
and restarted the local `web` container from source before this run).
- **Not yet done**: clean-checkout drill, commit/push, Unraid deployment of this fix
branch with live 3-language validation, the master merge (with the mandatory
`git fetch origin` / unexpected-change check first), and
`artifacts/fleet-ops-final-localization/final-summary.md`. Do not claim PASS on this
correction round until all of those are done and `git rev-parse HEAD` exactly matches
`/mnt/user/appdata/mobilityops/.deploy/source-revision`.
- Commits so far on this branch: `6deb955` (status flow + brand constant + message
codes), `e6539d1` (knowledge fixes), `ac4b163` (Playwright spec updates for the new
flow), `1fdd2b3` (new E2E coverage + 2 bug fixes), `1e40775` (accessibility test),
`a7ac5ed` (docs), `7851e80` (11D/11F i18n tests), `cda2c32` (clean-checkout
evidence), `2e4fb43` (evidence-summary localization fix, found live on Unraid).
Deployed commit: `2e4fb43f093bfbdb04c4f74eed1e6c6d9a03c069`.
## Live n8n + RAGcore integration (2026-08-04) — IN PROGRESS on feat/live-n8n-ragcore-integration
Branch `feat/live-n8n-ragcore-integration`, from master `c0995b7`. Full brief: treat n8n
(`https://n8n.itworx.tech`, existing shared instance) as a third integration layer
alongside RAGcore and MCP Hub, owning process orchestration only — Fleet Ops keeps all
business rules, authorization, transactions, audit and idempotency. Four canonical
workflows required: (1) Vehicle Return Orchestration, (2) Scheduled Data Quality Scan —
both pre-existing and now hardened; (3) RAGcore Procedure Sync, (4) Workflow Error
Handler — both net-new, not yet built.
- **Current-state audit**: `docs/live-ai-integration/n8n-current-state.md` documents the
live instance (reachable, production webhook base
`http://192.168.10.150:5678/webhook/mobilityops-return`), both existing workflows'
full node structure, and the findings that drove the security fixes below (webhook
Authentication was `None`; both HTTP nodes had `X-Service-Token` hardcoded as a literal
header value instead of a credential).
- **Security fixes applied and live-validated** (commits `b79d485`, `59cb4c0`): webhook
trigger now requires Header Auth (credential `Fleet Ops Webhook Trigger Token`, a new
token generated this round — value stored in `.env`/Unraid `.env` only, never
printed); the outbound callback HTTP node now uses a `Fleet Ops Service Token` Header
Auth credential instead of a literal header value (existing secret copied
clipboard-to-clipboard, never typed/echoed). Backend: `X-Fleet-Ops-Trigger-Token`
header added to the outbox dispatcher's POST (`backend/app/services/dispatcher.py`),
plus a new `MOBILITYOPS_WEBHOOK_TRIGGER_TOKEN` setting/env var. Also hardened
`_deliver_one` to treat a 2xx response with a non-JSON-object body as a retryable
failure (`malformedResponse`) instead of an unhandled exception — a real failure mode
hit live when a workflow errors before its "Respond to Webhook" node runs; regression
test `test_deliver_one_treats_empty_2xx_body_as_failure` added. Live-validated: curl
probe without the header → `403`; with the header → pass-through; one real end-to-end
vehicle return produced one correct execution visible in both n8n and Fleet Ops
Audit/Automation. Both workflows explicitly `Publish`ed after the fixes (the editor
does not go live on save alone) and both canonical-renamed ("Fleet Ops — Vehicle
Return Orchestration", "Fleet Ops — Scheduled Data Quality Scan").
- **RAGcore real contract discovered** (not the speculative one the adapter was built
against): OpenAPI at `/openapi.json`, health at `/health/live`/`/health/ready` (not
`/health`), ingestion via `POST /v1/uploads`, answers via `POST /v1/answers` with
`requested_space_ids`, control-plane endpoints require an `Idempotency-Key` header.
Bootstrapped a `fleet-ops` application + knowledge space + grant on the real server at
`http://192.168.10.150:1237`. **Blocked**: credential issuance for that application
failed identically via both the raw API and the admin UI ("authoritative
service-account state rejected issuance") — an apparent privilege boundary beyond the
interactive admin session. User chose to issue the credential themselves via another
mechanism and hand over the token; not yet received. `RAGcoreKnowledgeProvider`
(`backend/app/services/knowledge/ragcore.py`) still targets the old speculative
endpoints and needs fixing once that token arrives — approved, not started.
- **Repository source of truth started** (task in progress): `n8n/workflows/` now holds
cleaned definitions for workflows 1-2 — `fleet-ops-vehicle-return.json` (sha256
`e13a3087269fc97019a7adf6c6a6a4ee4bd354c2dd7167d4966d4753a48e970e`),
`fleet-ops-data-quality-scan.json` (sha256
`cc30b28b07dad9f9908a6ea0c564ec4c2f362a3ed71b7e97a7b6894408bb7e2e`) — both credential
auth referenced by name only, no secret values. Reconstructed from direct verified
inspection of every live node, **not** a literal n8n export/download: the UI's "..."
menu has no Download option in this n8n version, and clipboard-based
copy/`navigator.clipboard.readText()` extraction timed out twice. Flagged as a known
limitation for the final evidence doc. `n8n/workflows/MANIFEST.md` records canonical
name/purpose/trigger/contract/credentials/live ID/active-status/checksum for all 4
workflows (3-4 marked not-yet-built). `n8n/workflows/check_drift.py` compares a repo
definition against the live workflow via n8n's Public API (`X-N8N-API-KEY`, read-only,
never auto-overwrites). The old root-level `n8n/mobilityops-return-processing.json`
and `n8n/mobilityops-scheduled-quality-scan.json` (pre-integration starters, still
carrying the literal-token pattern) are removed; `deploy/unraid/setup-existing-n8n.sh`,
`setup-scheduled-scan.sh`, `Makefile` (`n8n-setup`, `n8n-setup-scan`) and
`docs/17-runbook.md` updated to import from `n8n/workflows/` and to document the
now-required manual credential-creation step (credentials are never scripted or
committed).
- **Explicitly deferred/forbidden this phase** (per brief): daily AI ops brief, email,
Slack, automatic vehicle-status changes, customer communication, billing, general
monitoring, autonomous MCP actions. An automatic demo-reset workflow may only be
prepared, not activated, once Fleet Ops goes public.
- **Workflow 4 (Workflow Error Handler) built and live-validated** (commit pending):
new backend endpoint `POST /api/v1/integrations/n8n/workflow-error`
(`backend/app/api/routers/integrations.py`, service-token auth, Pydantic
`WorkflowErrorReportIn`/`WorkflowErrorReportResult` in `backend/app/schemas.py`),
idempotent on `execution_id` via the same audit-precheck pattern as
`/return-callback`; new test coverage in `backend/tests/test_integrations.py` (all
green, 152 tests total, ruff/mypy clean). **This endpoint had to be deployed to the
live Unraid server** (`git archive` → `scp` → extract preserving `.env` → `docker
compose up --build -d api`, no migration needed) before the live n8n test could reach
it — the auto-mode classifier correctly blocked the first `scp` attempt as a
production-infra action; user approved, then it was deployed and verified
(`/health` OK, new endpoint returns 422 on empty body instead of 404).
Built "Fleet Ops — Workflow Error Handler" (live ID `Xppn2rAEqUuyiCJF`) in n8n:
Error Trigger → Code node (derives safe error_category/summary/etc. from n8n's error
payload) → HTTP node (POST to the new endpoint, Header Auth via the existing "Fleet
Ops Service Token" credential). Hit and fixed two real bugs during live testing: (1)
Code node's default "Run Once for All Items" mode doesn't bind `$json` to the current
item — switched to "Run Once for Each Item" and `return {json:...}` instead of
`return [{json:...}]`; (2) every HTTP-body field expression ended up with a stray
trailing space (from the code-editor's bracket-autoclose leaving one extra character
after the `End`+`Backspace×2` fix), which broke the `failed_at` datetime parse and the
`error_category` literal match — found via the raw request dump in n8n's error
panel, fixed with one more `Backspace` per field. Live-validated: mock Error Trigger
data → real `200 {"status":"registered"}` from Fleet Ops; re-run → `"already_registered"`
(idempotency confirmed); wired as the Error Workflow on workflows 1 and 2 (via each
workflow's Settings modal); confirmed the Error Handler itself has `Error Workflow: -
No Workflow -` (no recursive loop). With user approval, also ran a genuine induced
failure on workflow 2 (temporarily pointed its HTTP node at a nonexistent path,
published, ran it, confirmed it failed as expected, immediately reverted and
republished, confirmed healthy again) — this proved the target workflow's own error
path works, but n8n did not auto-invoke the Error Handler for that *manual* editor
test run (n8n's Error Workflow trigger only fires for unattended/production
executions), so a fully automatic schedule/webhook-triggered cascade into the handler
was not observed live this round — noted as a known limitation.
Exported the verified definition to `n8n/workflows/fleet-ops-error-handler.json` (same
manual-reconstruction caveat as workflows 1-2: no literal export/download available),
updated `n8n/workflows/MANIFEST.md` (all 4 workflows, workflow 3 still not-built) and
`check_drift.py`'s known-workflows list.
- **Integration status page enriched with real per-workflow evidence** (commit
`4049c0c`): `N8nIntegrationStatus` now returns `workflows: N8nWorkflowEvidence[]`
(the 4 canonical workflows, each with real evidence — latest successful outbox
delivery for the return workflow, latest *service*-triggered `data_quality_scan_run`
audit event for the scan workflow so a manual UI-triggered scan doesn't fake n8n
evidence, latest `n8n_workflow_failure_registered` for the error handler, always
`built: false` / no evidence for the not-yet-built RAGcore sync), plus
`expected_workflow_count`/`known_workflow_count` and an `error_handler` summary
(total registered, latest failure + which workflow). New tests in
`backend/tests/test_integration_status.py` (all green, 159 backend tests total,
ruff/mypy clean). Frontend: `Automation.tsx` renders this as a localized workflow
table (EN/NL/FR, new `integrations:workflows.*` keys, technical workflow names under
a "Technical details" disclosure per the existing progressive-disclosure pattern).
Verified live in the browser both locally (Dutch locale, disclosure expand/collapse
confirmed) and **on the deployed Unraid server after this round's deploy**: correctly
shows "3 van 4 canonieke n8n-workflows hebben actuele evidentie van werking" with real
timestamps for the return/scan/error-handler workflows, "Nog Niet Gebouwd" for the
RAGcore sync, and the real error-handler registration from this session's live
testing. Deployed to Unraid (commit `4049c0c6b12fef3d948cd31f21119044143320d8`,
rebuilt both `api` and `web`, `/health` OK) — user re-approved this second deploy
separately from the first.
- **Operational lesson learned this round**: `docker compose run --rm api pytest` does
**not** reliably pick up source edits without an explicit `docker compose build api`
first — a test file edit silently kept running against the stale built image (test
count didn't change) until rebuilt. Always `docker compose build api` (and `web` for
frontend changes) before trusting a green result after backend/frontend edits in this
repo.
- **WF1 acceptance gap fixed and live**: `check_drift.py`-style re-inspection of workflow
1 during this round's acceptance pass found the `Record follow-up` HTTP node had no
explicit timeout and "Retry On Fail" disabled — a real gap against the brief's
timeouts/bounded-retries requirement (WF2 already had this). Fixed live: Retry On Fail
(3 tries, 1000ms wait) + a 15000ms Timeout option, published (version note "Add bounded
retries (3x) and a 15s timeout to the Fleet Ops callback call"). Repo definition and
manifest checksum synced (`n8n/workflows/fleet-ops-vehicle-return.json`,
`MANIFEST.md`, new checksum `a6f399dd77a7203dec7c0ac95e8540abf55f2703da519e06f1c37f2e1220f609`,
commit `0562893`).
- **RAGcore credential issuance re-attempted and still blocked (user explicitly
authorized Claude to self-issue this round)**: tried the RAGcore admin UI's "Issue
credential" form for the `fleet-ops` application (logged in as Platform Admin, the
highest visible role) with name `n8n-ragcore-procedure-sync` and scope `sources:sync`
only. Submission failed with the same generic "Something went wrong. The credential
could not be issued with those values." page, this time carrying a trace reference
`1955c6a8968c4941a22a1faef39e17a7`. Inspected the RAGcore OpenAPI spec for this admin
endpoint (`POST /admin/control/applications/{application_id}/credentials`) — no
documented validation constraint explains the rejection (no 422, no field errors); the
`fleet-ops` application itself lists as ordinary/`Active` with no visible lock flag in
the applications table. This is the same failure signature as the earlier raw-API
attempt (400 "authoritative service-account state rejected issuance"): two independent
paths (raw API, and now the admin UI as the top admin role) both hit an opaque
server-side rejection with a trace ID. This is conclusive evidence the block is a
deliberate RAGcore-side policy or a RAGcore-side bug, not a Fleet Ops permission or
request-shape problem — nothing further is fixable from the Fleet Ops side or through
browser automation. Whoever operates the RAGcore instance needs to look up trace
`1955c6a8968c4941a22a1faef39e17a7` (and the earlier API rejection) in RAGcore's own
logs to find the real cause.
- **WF2 acceptance gap fixed and live**: continuing the acceptance pass to WF2 found it
had the *same* Retry On Fail gap as WF1 (its 15s timeout was already set, but retries
were off — the earlier note that "WF2 already had this" was wrong on the retry half).
Fixed live the same way (3 tries, 1000ms wait), published (version note "Add bounded
retries (3x) to the quality-scan HTTP call"). Repo definition and manifest checksum
synced (`n8n/workflows/fleet-ops-data-quality-scan.json`, `MANIFEST.md`, new checksum
`c0d46e0519118e6336e35c4ea2a67edb2f14bd007909ccf9256c93733751244a`, commit `167bf49`).
- **WF4 has the same gap on its own outbound call, but is currently un-fixable**: WF4's
"Report failure to Fleet Ops" HTTP node also has no timeout and no Retry On Fail. Began
the same fix (added a 15000ms Timeout option, toggled Retry On Fail on) but n8n's
autosave started failing with "Unauthorized" mid-edit, and a fresh tab confirmed the
n8n browser session had expired (redirected to `/signin`) — so nothing was saved and
the live WF4 definition is unchanged from before this round (no partial/broken state).
This is a minor, best-effort-only gap (WF4 is the error notifier itself, not a primary
business flow, and it already reports failures with `On Error: Stop Workflow` so a
failed error-report is visible in n8n's own execution history even without retries) —
not blocking, but worth finishing once someone re-authenticates the n8n browser
session.
- **RAGcore credential-issuance blocker root-caused and fixed (in RAGcore itself, with
explicit owner approval)**: with read access to the sibling `C:\Projects\RAGcore`
checkout, traced "authoritative service-account state rejected issuance" to a genuine
cross-transaction race in RAGcore's own dependency injection
(`src/ragcore/api/v1/control/dependencies.py`). `get_control_application` and
`get_credential_service` each independently opened their own `factory.begin()`
database transaction. Issuing a credential for a brand-new service account does, in one
request: (1) INSERT the service account via the first dependency's transaction, then
(2) immediately re-read it via the second dependency's *separate, uncommitted* transaction
— invisible under READ COMMITTED isolation until the first transaction commits, which
only happens after the endpoint returns. This made every fresh-service-account credential
issuance fail, 100% of the time, via both the raw API and the admin UI (explaining the
identical failure signature on both paths). RAGcore's own tests never caught this because
they override these dependencies with an in-memory fake that ignores transaction boundaries
entirely. Fixed by introducing one shared, cached `get_control_session` dependency that
both providers now depend on via `Depends(...)`, so they share one transaction per request.
Verified: RAGcore's own test suite (64 tests across `tests/web`, `tests/contract/api/control`,
`tests/security/identity`, `tests/unit/domain/control`, `tests/api`) passes, ruff and mypy
clean. Deployed to the live RAGcore instance (also on the Unraid host, `ragcore-app-1` on
port 1237 — a shared service also used by other ITWorx projects) via `docker compose build`
+ `up -d`, with explicit owner approval before both the code change and the deploy.
Confirmed fixed live: issuing a credential for `fleet-ops` (name
`n8n-ragcore-procedure-sync`, scope `sources:sync`) now succeeds (prefix `rc_sa_6fc51e`).
The plaintext token was never printed/logged — copied via RAGcore's own "Copy" button and
pasted directly into a new n8n Header Auth credential named **"RAGcore Sync Token"**
(header `Authorization: Bearer <token>`), ready for workflow 3.
- **Exact next action (superseded by the entry below)**: task #86 (build workflow 3,
RAGcore Procedure Sync) and the `RAGcoreKnowledgeProvider` adapter rewrite (to the real
inspected contract — `/health/live`, `/health/ready`, `POST /v1/uploads`, `POST
/v1/search`/`/v1/context`/`/v1/answers`) are now unblocked — the "RAGcore Sync Token" n8n
credential exists and works. WF4's own timeout/retry gap is still open pending n8n browser
re-authentication (minor, non-blocking, see above). Both #90 and #91 should be revisited
once workflow 3 is actually built, since they currently document it as blocked.
- **Second RAGcore bug found, fixed and deployed (with explicit owner approval, same pattern
as the transaction-race fix above)**: with the "RAGcore Sync Token" credential in hand,
the n8n "Upload to RAGcore" node still returned a persistent 401 on every item. Traced via
direct RAGcore source inspection (`C:\Projects\RAGcore`) to a genuine second, independent
gap: no code path in RAGcore converted an incoming `Authorization: Bearer <token>` header
into a `request.state.principal` for any `/v1/*` route — only browser session cookies were
ever accepted, even though the credential-verification logic
(`ServiceAccountCredentialService.verify()`) existed and was unit-tested. This blocks any
machine caller (n8n, and eventually Fleet Ops's own `RAGcoreKnowledgeProvider` adapter)
from ever authenticating to `/v1/uploads`. Fixed additively, scoped to `/v1/uploads` only
per owner instruction (search/context/answers left for later): new
`CredentialRepository.get_by_id()` (Postgres + in-memory), new
`ServiceAccountCredentialService.authenticate()` (parallel to the existing `verify()`, not
a refactor of it), and a new `get_upload_principal` FastAPI dependency
(`src/ragcore/api/v1/uploads/dependencies.py`) that falls back to the Bearer header when
there is no session principal, wired into `uploads/routes.py` in place of the session-only
`get_principal`. New/updated tests in `tests/security/identity/test_credentials.py` and
`tests/security/uploads/test_upload_security.py` (bearer-token accept/reject paths, the
existing route test's stale `get_principal` override fixed to `get_upload_principal`).
Verified: RAGcore's own test suite — 377 passed in the affected `tests/security`,
`tests/unit`, `tests/api` trees (3 unrelated pre-existing failures: two need Windows
symlink privileges the sandbox doesn't have, one is a git-connector fixture mismatch; a
separate architecture-boundary failure in `application/ingestion/handler.py` belongs to
unrelated in-progress work by a different concurrent agent on the same RAGcore checkout,
confirmed via `git status`/`git log` — not touched by this fix). Ruff and mypy clean on
every changed file. Deployed to the live RAGcore instance (`ragcore-app-1` on Unraid, port
1237 internally, fronted by `rag.itworx.tech` — note the *admin UI* and the *API* share
one process/origin, `/v1/uploads` is reachable at `https://rag.itworx.tech/v1/uploads`,
**not** `ragcore.itworx.tech`, which only appears in RFC7807 problem-type URLs) by copying
the 6 changed source files directly into the server checkout and `docker compose build
app && up -d --no-deps app` (deliberately not committing to RAGcore's git history or
touching the `worker` service, since a different agent has substantial unrelated
uncommitted work in that same working tree). Confirmed live with a garbage token (still
correctly 401) and then with a freshly-issued, correctly-scoped real token (403
`UPLOAD_TARGET_FORBIDDEN` against a dummy space ID — i.e. authentication succeeded,
authorization correctly rejected the wrong space — proving the fix end-to-end before
touching n8n at all).
- **Root cause of the n8n-side 401 found and fixed**: separately from the RAGcore bug above,
the "Upload to RAGcore" HTTP node's Authentication was set to Header Auth, but **no
credential had ever actually been attached** to that picker — so the node was sending no
`Authorization` header at all, which produces the identical 401 to a malformed one (easy
to conflate with the RAGcore-side bug, which is why fixing RAGcore alone didn't resolve
the symptom). There was already an unused "RAGcore Sync Token" n8n credential sitting
around from the earlier session (its value likely never actually got saved when it was
first created, or was created but never selected on this node — not conclusively
determined). Owner attached it and set Name=`Authorization`,
Value=`Bearer <freshly-issued token, scope sources:sync>` (a new credential issued via the
RAGcore admin UI at `/admin/control/applications/c20ac48a-d57b-4c68-9bd1-564f49c1a473/credentials/new`
specifically for this, service account "n8n Procedure Sync (production)"; the earlier
diagnostic-only credential used to prove the RAGcore fix was revoked afterward via direct
SQL `UPDATE identity.service_account_credentials SET revoked_at = now() ...` since the
admin UI has no revoke button).
- **Workflow 3's "Upload to RAGcore" node live-validated end-to-end, real data**: ran the
full workflow via n8n's "Execute workflow" (Schedule Trigger → List procedures → Prepare
uploads → Upload to RAGcore). All 33 items succeeded — each output item is a real
`AcceptedJob` (`job_id`/`status_url`), not error output. Independently confirmed at the
database level (not just trusting the n8n UI): `select count(*) from jobs.jobs where
operation='ingest_upload' and created_at > now() - interval '5 minutes'` → **33**, on the
live RAGcore Postgres.
- **n8n browser-automation notes for this environment** (worth knowing before attempting
canvas interaction again): (1) an n8n NPS survey modal (`role=dialog`, "We've been busy")
intermittently covers the whole canvas and silently eats every click underneath it until
removed; (2) canvas node positions reported by `getBoundingClientRect()` drift between
successive tool calls in a way that made coordinate-based `computer` clicks and even
`find`-ref-based clicks land on the wrong element repeatedly this session (dozens of failed
attempts, multiple different coordinate-math theories, none reliable) — directly setting
`.vue-flow__transformationpane`'s inline `style.transform` to force a node into view
**desyncs vue-flow's own internal pan/zoom state**, making the problem worse, not better;
(3) what actually worked reliably every time: calling native `.click()` directly via JS on
a plain `<button>` element (e.g. the toolbar's "Execute workflow" button) — canvas *node*
interaction (double-click to open a node's parameter panel) was never reliably achieved
via any automated method this session; the owner opened/edited the node manually instead.
A plain synthetic `dispatchEvent(new MouseEvent(...))` sequence (pointerdown/mousedown/
pointerup/mouseup/click/dblclick, even with correct `clientX`/`clientY`/`detail`) does
**not** trigger vue-flow's node click handling at all — it appears to require a genuinely
trusted (real CDP-driven) pointer event, consistent with vue-flow's drag/zoom gesture
system depending on native pointer capture.
- **Exact next action (superseded further below)**: build and wire the workflow's final
"Summarize sync result" Code node (count successes/failures across the 33 items) and a
closing HTTP node reporting to Fleet Ops's already-deployed `POST
/api/v1/integrations/n8n/procedures-sync-result` (task #86, still in progress — the sync
itself now works, this is the last piece). Then publish the workflow (currently still a
draft/unpublished), update `n8n/workflows/MANIFEST.md` and add
`n8n/workflows/fleet-ops-ragcore-procedure-sync.json` as the repo source-of-truth
definition, and revisit `artifacts/live-ai-integration/final-summary.md` (documents WF3
as blocked — no longer true).
- **Verified the 33-item sync is genuinely fully ingested, not just accepted**: checked at
every RAGcore pipeline stage on the live instance, not just trusting the n8n "success"
status (which can mask individual failures under `On Error: Continue`). All 33
`ingest_upload` jobs have `status='succeeded'` in `jobs.jobs`; 33 rows exist in
`content.documents`; 372 chunks were generated in `content.chunks`; 241 vectors are
indexed in the `rag_dense_nomic-embed-text_v1` Qdrant collection filtered specifically to
Fleet Ops's `space_id` (`f4c91e49-5cf9-48ba-b3d6-e0e9854ebccc`) — matching the expected
leaf-chunk count (structural/parent chunks in the hierarchy aren't separately embedded).
11 unique procedures × 3 languages = 33, all present.
- **`RAGcoreKnowledgeProvider` rewritten against the real contract (task #93), tested, but
NOT switched on in production — genuinely blocked on a RAGcore-side gap, not a Fleet Ops
problem**: rewrote `backend/app/services/knowledge/ragcore.py` end to end against
RAGcore's actual `/v1/answers` and `/health/ready` contracts (previously a best-effort
guess against an unreachable instance). Health now calls the real `GET /health/ready`.
`ask()` calls `POST /v1/answers` with `Authorization: Bearer <token>` and
`requested_space_ids: [settings.ragcore_space_id]`, maps RAGcore's `answerability` enum
conservatively (`answerable`/`partially_answerable` *with* non-empty citations →
`grounded`, everything else → `insufficient`, any transport/parse/non-200 failure →
`unavailable`, matching the architecture's never-fabricate rule). Added
`RAGCORE_SPACE_ID` setting (`.env.example`, `compose.yaml`). New tests in
`backend/tests/test_knowledge.py` (connection-error degrade, missing-space-id short
circuits without a network call, health ready/degraded/unreachable, grounded citation
mapping, not-answerable and answerable-without-citations both correctly map to
`insufficient` with an empty answer, non-200 and malformed-body both degrade to
`unavailable`) — **23 passed** in that file, **172 passed** full suite, ruff clean, mypy
0 issues/50 files.
**Discovered while live-testing against RAGcore with a freshly-issued, correctly-scoped
credential** (`Fleet Ops Knowledge Assistant (production)`, scope `answer`, space
`f4c91e49-...`): authentication now genuinely succeeds (past the 401 stage — confirmed
via `/health/ready` returning `200 {"status":"ok", ...}` with the same token), but **both
`POST /v1/answers` and `POST /v1/search` return `503`
`ANSWERS_UNAVAILABLE`/`SEARCH_UNAVAILABLE`** for every request. Root-caused by reading
`src/ragcore/main.py`'s app-startup/lifespan code directly: it constructs and assigns
`app.state.database_engine`, `session_factory`, `qdrant_client`,
`query_lab_service`, `profile_activation_service`, and the OIDC session services — but
**never constructs or assigns `app.state.search_application`, `context_application`, or
`answer_application`** anywhere in the codebase (confirmed via a repo-wide grep — the
three `get_*_application` dependency functions exist and correctly raise their `503
*_UNAVAILABLE` problem when the state attribute is absent, exactly as designed, but
nothing ever populates it). This is not a config toggle Fleet Ops is missing and not
something introduced by today's auth fixes — the retrieval/generation subsystem's route
handlers and dependencies are scaffolded end-to-end but were never wired into the running
application on this RAGcore deployment. Only masked until today because every call to
these routes previously 401'd on auth before ever reaching this check.
**Decision: `KNOWLEDGE_PROVIDER` stays `demo` in production.** Flipping it to `ragcore`
right now would replace the currently-working demo Knowledge Assistant with one that
correctly, honestly, but uselessly reports "unavailable" for every question — strictly
worse for the live demo. The adapter code itself is finished, correct, and safe to ship
(already committed-worthy), and switching providers is a one-line env var flip
(`KNOWLEDGE_PROVIDER=ragcore` + set `RAGCORE_API_TOKEN`/`RAGCORE_SPACE_ID`) the moment
RAGcore's own operator wires up `search_application`/`answer_application` on their side.
The freshly-issued credential (`Fleet Ops Knowledge Assistant (production)`, scope
`answer`) was left active/unused in RAGcore, ready for that day.
- **MCP Hub registration (task #94/#95) — full new connector built and validated locally
in the sibling `C:\Projects\ITWorx_MCP_Hub` checkout, not committed or deployed**: the
Hub's admin UI (`mcp.itworx.tech`, real production instance, 7 pre-existing projects —
DevRunbook, ForgeFlow×2, General Infrastructure, GeoIntel, Ludarium) has **no self-service
"add project/connector" flow** — its own Settings page states "No writable settings are
available in this browser until the control API publishes an authorized configuration
schema" (Hub is on release `1.0.0-rc`, UI is read-only/observability-only today).
Registering MobilityOps therefore required building a genuinely new connector in the
Hub's own repo, following its `gitea` connector as the closest real template (external
HTTPS API, shared-secret auth) rather than `knowledge` (which turned out to be an
in-memory filesystem index, not an HTTP client, despite the name suggesting otherwise).
Built, with explicit owner approval given the Hub's own `CLAUDE.md` restricts autonomous
action to local-checkout work only (registry pushes and production startup are separate,
explicitly-gated checkpoints, not done this round):
- `packages/connector_kit/mobilityops.py` — `MobilityOpsSettings`/`MobilityOpsClient`
(HTTPS-only, same-origin-redirect-enforced, `X-Service-Token`/`X-Client-Id` headers,
path-safety-validated `vehicle_ref`), mirroring `gitea.py`'s hardening exactly.
- `connectors/mobilityops/{__init__,server,fake}.py` — `MobilityOpsConnector` exposing 4
read-only tools (`mobilityops.operations.summary`, `.attention.list`, `.vehicle.get`,
`.knowledge.search`), each `governed_tool`-wrapped, envelope-wrapped, field-mapped from
Fleet Ops's real `/api/v1/integrations/mcp/*` response shapes.
- `services/runtime/connector_mobilityops.py` + `dispatch.py` registration.
- Catalog: `catalog/connectors/mobilityops-default.yaml` (ConnectorTemplate),
`catalog/tools/mobilityops-*.yaml` (4 ToolManifests), `catalog/capability-packs/
mobilityops-reader.yaml` (a **new**, narrowly-scoped pack — deliberately not added to
the shared `project-reader` pack other projects use, to avoid granting them
MobilityOps access), `catalog/projects/mobilityops.yaml` (Project, `applicationUrl`
set to the real internal `http://192.168.10.150:1236`, not an invented public domain).
- Compose wiring across all three files (`docker-compose.yml`, `.prod.yml`,
`.blueprint.yml`) plus a new `mobilityops_service_token` Docker secret.
- Fixed ~13 regressions this surfaced in the Hub's own existing test suite — all
legitimate guard-rail tests (hardcoded service/tool/secret allowlists, a
`CONNECTOR_GATEWAYS`/`FakeContextForge._gateway_tools` registration gap that was a
**real production wiring miss**, not just a test-fixture gap: without it, the reconcile
step would have raised `"virtual server resolved to an empty tool allowlist"` against
the live Hub too) — not scope creep, this is exactly the Hub's own established pattern
for registering a new connector, verified by reading how `gitea`/`knowledge`/`unraid`
each touch the same ~15 files.
- Verified: 15 new connector tests pass; full Hub suite **383 passed, 0 failed, 5 skipped
(pre-existing)**; `ruff check`/`ruff format --check` clean; `mypy` clean across 108
source files; `scripts/check_boundaries.py` passes (import-direction rules respected —
the new connector only imports from `packages/connector_kit`, never `services/*`
directly); `scripts/validate_pack.py` schema-validates the catalog cleanly (68
documents, up from 61 — exactly the 7 new files) with the only remaining failure being
the Hub's own git-cleanliness gate for manifest regeneration, which requires a commit.
- **Not done, deliberately**: no commit (the Hub repo has substantial unrelated
in-progress work from a different concurrent agent — `WP-235` frontend redesign,
`BUILD_STATE.json`/`CURRENT_STATE.md`/`implementation/WORK_PLAN.json` all show as
modified by them, not by this session — committing broadly risks entangling that
work); no registry push; no production deployment/restart of the live
`itworx-mcp-hub-*` stack. Per the Hub's own `CLAUDE.md` approval boundaries, those are
separate, explicitly-gated checkpoints requiring their own fresh approval, and a real
`MOBILITYOPS_ENDPOINT`/`mobilityops_service_token`/`MOBILITYOPS_PROJECTS_JSON`
production configuration still needs to be decided before any of that could run.
- **MCP Hub connector — committed** as `f107544` ("Add MobilityOps (Fleet Ops) read-only
connector", 28 files) in `C:\Projects\ITWorx_MCP_Hub`, `main` branch. Not pushed to the
registry and not deployed to the live Hub stack — those remain separate, explicitly-gated
checkpoints per the Hub's own `CLAUDE.md` approval boundaries.
- **RAGcore search/context/answer application wiring (task #96) — implemented, tested,
committed as `a2905cc`** in `C:\Projects\RAGcore`, `main` branch (not pushed, not
deployed to the live RAGcore instance). This closes the gap documented above:
`search_application`/`context_application`/`answer_application` are now actually
constructed in `create_app()` instead of being permanently absent.
- `src/ragcore/application/retrieval/production_executor.py` (new) — `RetrievalPipeline`
(embed via Ollama -> dense+sparse Qdrant search -> RRF fuse -> rerank via Ollama),
reusing the same flow already proven by the Query Lab's `PipelineQueryLabExecutor`
(left untouched). Feeds `ProductionSearchExecutor` and `ProductionContextExecutor` from
one shared per-request run, so `SearchHit`'s `FusedCandidate` (required by
`SearchService`'s own provenance/scope check) and `ContextCandidate`'s parent-chunk
expansion both derive from the same reranked result set instead of drifting apart.
- `src/ragcore/application/models/answer_generator.py` (new) — `OllamaAnswerGenerator`,
`AnswerService`'s real generator: resolves the active `GenerationProfile` from
`GenerationProfileRegistry` (already built, never previously constructed anywhere) and
calls the existing hardened `OllamaGenerationAdapter`. `main.py`'s `lifespan()` now
seeds and activates a default local-Ollama answer profile on first boot if none is
active, since no seed data ever registered one.
- Verified: 10 new unit tests (7 for the pipeline/executors, 3 for the answer generator)
pass; `ruff check` clean; `mypy src` clean (the only 4 mypy errors found are in
`infrastructure/qdrant/adapter.py`, `infrastructure/qdrant/filters.py`,
`api/v1/control/routes.py` — files this change never touched, confirmed via `git
status`/`git diff`, pre-existing or from other concurrent work in this checkout).
- Full non-integration suite (`pytest -m "not integration"`, excludes tests requiring
disposable infra per the repo's own marker convention): **593 passed, 14 failed, 1
skipped**. All 14 failures are pre-existing and unrelated to this change: 10 are
Windows-only (`psycopg`'s async driver rejects Windows' default `ProactorEventLoop` —
`test_health.py`, `test_readiness.py`, `test_control_plane.py`), 2 are Windows-only
(`WinError 1314`, no symlink privilege on this account — `test_folder_connector.py`), 1
is an unrelated pre-existing git-security-connector assertion, and 1 is the repo's
zero-tolerance `test_domain_and_application_respect_dependency_boundaries` test —
already red before this change, since `application/ingestion/handler.py` (untouched,
commit `bb75f52`) already imports `httpx`/infrastructure modules from the application
layer, which that test forbids with no allowlist. This change's new files follow the
same established (if already-violating) pattern to reach real adapters, adding more
instances of the same pre-existing violation rather than introducing a new kind of one.
- **Deployed to production** (`http://192.168.10.150:1237` via `ragcore-app-1`),
approved by the user. `scp`'d the 4 changed/new files individually to
`/mnt/user/appdata/ragcore/app/...`, `docker compose build app`, verified the built
image actually contains the new code (`docker run --rm ragcore-app grep`/`test -f`),
then `docker compose up -d --no-deps app` (only the `app` service touched). Startup
logs are clean (`ragcore_started environment=production`, no errors); `/health/ready`
reports `postgres`/`qdrant`/`storage` all `ok`.
- **Live-verified the fix itself**: `POST /v1/search` now returns `401
AUTHENTICATION_REQUIRED` for an unauthenticated call, not the old permanent `503
SEARCH_UNAVAILABLE` — proof the endpoint reaches real request handling instead of
hitting an absent application state, which is exactly the bug this closes.
- **Issued a fresh scoped credential** (`answer` scope, application `Fleet Ops`, via the
RAGcore admin UI at `rag.itworx.tech` — a separate, working OIDC session, unaffected by
the n8n auth problem above) and made a real authenticated `POST /v1/answers` call
against the `Fleet Ops Procedures` knowledge space (`f4c91e49-5cf9-48ba-b3d6-e0e9854ebccc`).
Got a clean `200`, a real `answer_id`/`retrieval_run_id`, `degraded: false` — but
`answerability: not_answerable`, 0 citations, for two different real questions matching
real seeded document titles.
- **Confirmed this is not a regression from this change**: the same query against the
same space through RAGcore's own pre-existing, already-proven Query Lab tool (which
uses `PipelineQueryLabExecutor`, code this session never touched) returns the identical
`NOT_ANSWERABLE` / 0 evidence / not degraded result — down to `Dense candidates (0)` and
`Sparse candidates (0)` at the raw retrieval stage, before fusion or rerank ever runs.
Whatever is causing zero matches lives in shared retrieval infrastructure or the data
itself, not in the new production executors.
- **Ruled out the obvious causes via direct Qdrant/Postgres checks** (not yet root-caused
further — flagging as a separate follow-up, not blocking this task): the space
genuinely has 83 published, correctly-scoped points in `rag_dense_nomic-embed-text_v1`
(right `workspace_id`, right `space_id`, `status: published`, real Dutch-language
procedure text); the points' `embedding_model_digest` exactly matches the currently
active embedding profile's digest (no stale-embedding mismatch); the collection alias
(`rag_dense_nomic-embed-text_active`) correctly resolves to that same collection. One
oddity noted in passing, unrelated to retrieval: at least one sampled chunk has
`language: "en"` in its payload despite the actual text being Dutch — worth a look if
language-filtered queries matter for the demo.
- **Decision on `KNOWLEDGE_PROVIDER=ragcore` in Fleet Ops**: still deliberately `demo`.
The application is deployed, correctly wired, and behaves identically to RAGcore's own
trusted reference implementation — but real questions against real seeded content
aren't returning grounded answers yet for a reason that traces to shared
retrieval/data, not to this wiring. Flip only after that's root-caused.
- **n8n workflow 3 (task #86) — incident, recovered; final 2 nodes still blocked on a
live n8n auth problem, not a code problem.** Returning to finish the "Summarize sync
result" + report-to-Fleet-Ops nodes found the live "Fleet Ops — RAGcore Procedure Sync"
workflow's canvas at **zero nodes** — the earlier session's abandoned attempt to call
n8n's REST API directly (the "unexplained 401" noted above) had gone through far enough
to wipe the live workflow before the 401 stopped it, leaving an empty, unsaved "Current
changes" draft on top of the last good save.
- **Recovered**: n8n's own Version History (`/workflow/.../history`) still had the last
good save, version `260b38b5` (Aug 4, 21:08), with all 4 real nodes intact (Schedule
Trigger -> List procedures -> Prepare uploads -> Upload to RAGcore) plus the build
sticky note. Used the history panel's own **Restore version** action (not a manual
rebuild) to bring the live workflow back to that exact state. Confirmed via the page's
DOM (`[data-test-id="canvas-node"]` count) both before (0) and after (4) — the canvas
itself render fully off-screen (nodes positioned at negative Y coordinates, a separate
display-only vue-flow pan bug worked around by directly setting the transform pane's
CSS transform; harmless, doesn't touch saved data).
- **Then blocked again attempting the 2 new nodes**: adding a Code node via the UI's own
"What happens next?" panel (a real, first-party n8n interaction, not the REST API)
immediately surfaced n8n's own autosave toast: **"Problem saving workflow — Autosave
failed: Unauthorized."** A full sign-out/sign-in cycle right before this (fresh
credentials, fresh page load) did not fix it — immediately after a successful
interactive "Sign in", `fetch('/rest/workflows/...')` from that same authenticated tab
still returned 401. This is not the ordinary "your session expired, log in again"
friction seen earlier in the session; it's the live n8n instance's own save path
rejecting a request made moments after a successful login, which is exactly the
failure mode that caused the wipe above. Continuing to add nodes under this condition
risks losing work again with no guarantee the next failure is as recoverable, so this
was stopped deliberately rather than retried. The workflow was left in its safely
restored, 4-node, unpublished state — verified via the DOM node count immediately
before closing the session — nothing was added or changed beyond the restore.
- **Not a code/workflow-design problem**: the 2 remaining nodes (a Code node
summarizing the sync result, and an HTTP node reporting to Fleet Ops's already-live
`POST /api/v1/integrations/n8n/procedures-sync-result`) were never built — this
blocked before any node configuration happened. The design itself (mirroring WF2/WF4's
existing report-result pattern) is unchanged from earlier planning.
- **Exact next action**: (1) RAGcore search/answer wiring (`a2905cc`) is deployed and
proven correctly wired (matches Query Lab's trusted behavior exactly) — the open item is
now root-causing why `Fleet Ops Procedures` returns zero dense/sparse candidates for
real questions despite having genuinely matching, correctly-scoped, correctly-embedded
indexed content (see evidence above); this is retrieval/data, not application wiring.
Only flip `KNOWLEDGE_PROVIDER=ragcore` once that's fixed and a real grounded answer comes
back. (2) task #86's remaining n8n pieces (Summarize
+ report-result node, publish) need the live n8n instance's save/auth problem fixed
first — this looks like an n8n-instance-side issue (session/auth backend rejecting
mutating requests moments after a successful login), not something fixable from the
browser automation side; needs either the user's own direct n8n session or an infra
look at the n8n deployment before any further automated editing is attempted. (3) MCP
Hub: registry push + production deployment of the now-committed connector remain
separate, explicitly-gated checkpoints. WF4's own timeout/retry gap remains open,
deferred, non-blocking.