2577 lines
218 KiB
Markdown
2577 lines
218 KiB
Markdown
# Project state
|
||
|
||
## Publication and Unraid deployment (2026-08-02)
|
||
|
||
- Unraid deployment is live at `http://192.168.10.150:1236` from
|
||
`/mnt/user/appdata/mobilityops`, Compose project `mobilityops`.
|
||
- Deployment config commits: `07ab7a3`, `847cd05`, `e1a1c67`, `1e13943`. The accepted
|
||
baseline `4bf9afbeff44088864e0844769d4dd0e4089d85b` remains intact.
|
||
- MobilityOps PostgreSQL, API and web services are healthy. Only web port 1236 is exposed
|
||
by the MobilityOps Compose project; API and PostgreSQL remain internal. Automation uses
|
||
the server's existing shared n8n at `http://192.168.10.150:5678`; no second MobilityOps
|
||
n8n container is running.
|
||
- Migrations are at `e7b08389f47f (head)` and deterministic seed counts match final
|
||
acceptance. A Chrome smoke test covered every requested page and a real return; its n8n
|
||
event succeeded on attempt 1. Browser console and recent service log scans were clean.
|
||
- RAGcore is disabled in favor of the honest local demo provider. MCP Hub registration is
|
||
disabled. The MobilityOps workflow is published in the existing n8n and live-verified.
|
||
- Local post-change gates: 66 backend tests, Ruff, mypy (44 files), and frontend production
|
||
build all pass. Evidence is in `artifacts/deployment/unraid-summary.md`.
|
||
- Published to the private Gitea repository
|
||
`https://gitea.itworx.tech/Jens/MobilityOps`. `master` is the default branch; the full
|
||
commit history and baseline commit are present; zero tags exist; remote hygiene is
|
||
clean. `origin` uses the SSH clone URL supplied by Gitea.
|
||
- Exact next action: none — repository publication and Unraid deployment are complete.
|
||
|
||
## Current milestone
|
||
|
||
M7 — complete. All milestones (M0–M7) done, plus a full post-M7 final-acceptance audit (see below). See `artifacts/final-acceptance/summary.md` for the definitive acceptance evidence (supersedes `artifacts/evidence/final-summary.md`, which is kept as historical M7 evidence).
|
||
|
||
## Locked decisions
|
||
|
||
- Product name: Fleet Ops (the only visible product name in the UI/copy, never translated;
|
||
see `frontend/src/product.ts`). "MobilityOps" is the internal repo name, Compose project
|
||
name and deployment directory only — never shown to a user. See the "Final product
|
||
polish: Fleet Ops rebrand" and "Fleet Ops final localization" entries below.
|
||
- Fictitious tenant: Northstar Mobility Demo.
|
||
- Synthetic demo data only; all operational and knowledge data are synthetic. The product
|
||
itself is not described as a "PoC" in user-facing copy (see the localization entry
|
||
below) — this document and other internal/engineering docs may still use "PoC" to
|
||
describe the engineering scope, per `CLAUDE.md`.
|
||
- Core stack and boundaries are defined in `CLAUDE.md` and `docs/03-architecture.md`.
|
||
- RAGcore and ITWorx MCP Hub are external central services.
|
||
- n8n receives post-commit events through an outbox dispatcher.
|
||
- SQLAlchemy 2 declarative models cover the full domain model (`backend/app/models/`); enums are plain `String` columns validated at the Pydantic/service layer, not native PG enums (simpler migrations).
|
||
- `backend/requirements.lock` is compiled inside a `python:3.12-slim` container (matches the Dockerfile base image) via `pip-compile --extra dev`; regenerate the same way if `pyproject.toml` changes.
|
||
- Frontend dependencies pinned (no more `"latest"`); `package-lock.json` committed; Docker build uses `npm ci`.
|
||
- Demo auth is a lightweight HMAC-signed cookie (`app/core/security.py`), not a real password/JWT flow — matches "Demo role buttons create an authenticated session; they do not bypass authorization middleware." Two fixed demo users (`USR-OPS` operations_manager, `USR-EMP` rental_employee) are created by the seed loader, not from a CSV (no `users.csv` in `seed/`).
|
||
- Seed loader (`backend/app/seed_loader.py`) only supports `seed --reset` (always rebuilds); there is no incremental/idempotent-without-reset mode, since the acceptance criteria only require deterministic reset, not partial import.
|
||
- `DataQualityIssue.entity_ref`/`related_ref` from the CSVs are resolved to `entity_type`/`entity_id` (UUID) at load time per the domain model; the original human-readable refs are kept in `evidence_json` (`entity_ref`, `related_refs`) since the API and UI need them and re-resolving UUID→public_ref on every read would be wasteful.
|
||
- `backend/app/core/config.py` added `app_secret`, `session_cookie_name`, `session_ttl_seconds`, `seed_dir` (`/app/seed` in-container), `cors_allow_origins` (comma-separated string, not a list — simpler with pydantic-settings env parsing), `demo_today` (drives the dashboard's "Today" section against the deterministic anchor date, default `2026-08-01`).
|
||
- `compose.yaml` api build context changed from `./backend` to repo root with `dockerfile: backend/Dockerfile`, so the image can `COPY seed ./seed` (seed CSVs are outside `backend/`).
|
||
- Frontend: added `react-router-dom@7.18.2` (bumped from 6.x to clear two real advisories — open redirect + arbitrary constructor injection in v6). One residual `npm audit` finding (RSC-mode CSRF, GHSA-qwww-vcr4-c8h2) does not apply — this SPA never uses React Router's RSC/SSR mode.
|
||
- Nav/pages built so far: Dashboard, Vehicles (list+detail with tabs), Bookings (list+detail), Audit. Data Quality, Knowledge and Automation nav items are intentionally omitted until M3/M5/M4 build the pages behind them — CLAUDE.md forbids dead routes/placeholders.
|
||
- Return workflow (`app/services/returns.py`): the spec's "validate submitted reading against booking start reading" step was dropped as a hard rejection. For the seeded S1 scenario, a booking's `start_odometer_km` can already equal the vehicle's canonical odometer, so any regression-testing value would also be below the booking start, making a hard floor there indistinguishable from — and in conflict with — the documented soft-regression path. Only one odometer check exists now: submitted vs. the vehicle's *canonical* odometer (`vehicle.odometer_km`), matching the domain-model invariant verbatim ("a return with a lower submitted reading is recorded as an inspection and issue, while canonical odometer remains unchanged").
|
||
- Idempotency: new `idempotency_records` table (migration `e7b08389f47f`), unique on `idempotency_key`, keyed to `booking_id`. Same key + same booking replays the stored response; same key + different booking → 409 `IDEMPOTENCY_KEY_REUSED`; different key on an already-returned booking → 409 `INVALID_BOOKING_STATE`. Concurrency is enforced by `SELECT ... FOR UPDATE` on the booking row (re-checked for the idempotency record immediately after acquiring the lock, as a safety net for two simultaneous identical-key requests racing the pre-lock check).
|
||
- `seed_loader.clear_all()` must delete `idempotency_records` before `bookings` (FK) — easy to forget when adding new booking-referencing tables; the ordering list at the top of `seed_loader.py` is the single place to update.
|
||
- Inspection `public_ref` is assigned as `INSP-{count+1:04d}` from a live count query (not gap-safe, fine for a PoC single-writer demo, would need a sequence for real concurrency-safe numbering).
|
||
- Found and fixed during browser verification (not caught by pytest, since it's a UI-only defect): `ReturnForm` originally held its own `result` state and was conditionally rendered only when `booking.status === "active"`; once the return succeeded the booking flipped to `returned` and React unmounted the form before the user ever saw the result panel. Fixed by lifting the result into `BookingDetail` (`ReturnResultPanel` is now a sibling, not nested in `ReturnForm`). Also found: `OutboxEvent.event_id`'s Python-side `default=uuid.uuid4` on the mapped_column only applies at flush/commit time, so reading `event.event_id` before `db.commit()` returned `None` (rendered as the literal string "None" in the result panel); fixed by assigning `event_id=uuid.uuid4()` explicitly at construction. Lesson: SQLAlchemy column `default=` callables are not available on the in-memory Python object until flush — never rely on the generated value for a same-transaction response body without an explicit `db.flush()` or an explicit Python-side assignment.
|
||
- Operational note for this environment: `docker compose run --rm api ...` (used for tests/lint) only starts a throwaway one-off container — it does **not** update the long-running `api`/`web` service containers. After any code change meant to be verified live (browser, curl), `docker compose up -d --build <service>` is required, not just `docker compose build`.
|
||
- **Sharper version of the note above, found the hard way (2026-08-05)**: `compose.yaml`'s `api` service has **no bind mount** for `./backend` — application code is baked into the image at build time only. `docker compose run --rm api pytest/ruff/mypy` reuses whatever image was last built; it does **not** pick up host file edits at all, not even for a throwaway container. Editing code and immediately running `docker compose run --rm api pytest` without an intervening `docker compose build api` silently tests/lints the *old* code and can report a false "all green." Always `docker compose build api` before the first local gate run after a code change in a session (subsequent runs against the same build are fine). Caught this only because a new test file's test count didn't match after several rounds of edits; re-ran the full local gate suite after rebuilding and found one genuinely stale test assertion (below) — nothing else was actually broken, but this was luck, not verification, until the rebuild.
|
||
|
||
## Completed evidence
|
||
|
||
### M0 — Reproducible foundation
|
||
- Added `backend/app/core/db.py` (engine/session), `backend/app/models/*` (User, Customer, Vehicle, Booking, Inspection, MaintenanceRecord, DataQualityIssue, OutboxEvent, AuditEvent), Alembic config (`backend/alembic.ini`, `backend/alembic/env.py`) and initial migration `backend/alembic/versions/c9498525abb5_initial_schema.py`.
|
||
- Commands run and verified from this checkout:
|
||
- `docker compose build api` — OK
|
||
- `docker compose run --rm api alembic upgrade head` — applied cleanly to empty DB, created 9 tables + `alembic_version`.
|
||
- `docker compose run --rm api pytest -q` — 1 passed.
|
||
- `docker compose run --rm api ruff check .` — All checks passed (added `extend-exclude = ["alembic/versions"]` to `backend/pyproject.toml` for autogenerated migration line length).
|
||
- `docker compose up -d --build` — all 4 services healthy: `curl http://localhost:8128/health` → `{"status":"ok",...}`; `curl -o /dev/null -w "%{http_code}" http://localhost:1228/` → 200; `curl http://localhost:5678/healthz` → 200.
|
||
- Fixed a real scaffold bug: `frontend/src/App.tsx` used `import.meta.env` without a `vite/client` types reference, which broke `npm run build` in Docker (works fine under plain `vite dev` because Vite injects the global at dev-time but `tsc -b` still type-checks it). Added `frontend/src/vite-env.d.ts`.
|
||
- `make` is not installed in this Windows/git-bash shell — validated the underlying `docker compose ...` commands directly instead (Makefile targets are thin wrappers around them and are correct as written for a Linux/CI shell or WSL).
|
||
- Known accepted gap: `npm audit` reports 1 moderate/1 high transitive `esbuild` advisory (dev-server-only, fixed only by a Vite 8 major bump); left as-is for the PoC, noted here rather than silently upgrading a major version.
|
||
|
||
### M1 — Operational core
|
||
- Backend additions: `app/core/security.py` (HMAC-signed session cookies), `app/api/deps.py` (`get_current_user`, `require_operations_manager`), `app/core/errors.py` (`AppError` + the documented `{"error": {...}}` shape wired as a FastAPI exception handler for both `AppError` and `HTTPException`), `app/seed_loader.py`, `app/cli.py` (`python -m app.cli seed --reset`), `app/services/audit.py`, `app/schemas.py`, routers under `app/api/routers/` (`demo`, `dashboard`, `vehicles`, `bookings`, `audit`).
|
||
- Frontend additions: React Router-based app shell (`src/App.tsx`, `src/components/Layout.tsx`, `src/components/RequireAuth.tsx`), `AuthContext`, typed `api` client (`src/api/client.ts`, `src/api/types.ts`), pages `Login`, `Dashboard`, `Vehicles`/`VehicleDetail`, `Bookings`/`BookingDetail`, `Audit`. Full responsive stylesheet (`src/styles.css`) covering nav collapse and table→card layout under 700px, visible focus states, no hover-only actions.
|
||
- Commands run and verified from this checkout (container rebuilt each time to pick up code changes):
|
||
- `docker compose run --rm api pytest -q` — **19 passed** (new: `test_seed.py`, `test_auth.py`, `test_dashboard.py`, `test_vehicles.py`, `test_bookings.py`, `test_audit.py`; tests seed the real Postgres via `reset_and_seed` in a session fixture, then exercise the FastAPI app through `TestClient`, not mocks).
|
||
- `docker compose run --rm api ruff check .` — All checks passed (added `ignore = ["B008"]` — FastAPI's `Depends()`-as-default is idiomatic, not a real bug).
|
||
- `npm run build` (local, Node 24) — clean `tsc -b && vite build`.
|
||
- `docker compose up -d --build` then `docker compose exec api python -m app.cli seed --reset` — counts: `users:2 customers:180 vehicles:50 bookings:246 inspections:75 maintenance:40 data_quality_issues:15 workflow_runs:20`.
|
||
- `curl` end-to-end: `POST /api/v1/demo/login` sets cookie and returns the user; unauthenticated `GET /api/v1/dashboard` → 401 with the documented error shape; authenticated dashboard/vehicle-detail return real seeded data (verified metrics `available:21 rented:11 cleaning:6 maintenance:5 blocked:7`, matching the 50 seeded vehicles).
|
||
- Browser smoke test (Chrome via MCP) at desktop width: login page → Operations Manager login → Dashboard (metrics + attention items + today + recent automation all populated) → Vehicle detail `MO-016` (tabs render, "Needs attention" badge correct — it's `DQ-DEMO-OVERLAP`/`DQ-DEMO-STATUS`) → Booking detail `BK-DEMO-RETURN` (matches S1 scenario: vehicle `MO-024`, status `active`, start odometer `53610`). Responsive CSS (`@media max-width:700px`) was written and code-reviewed but the automated resize during this session didn't visibly reflect in the captured screenshot (likely a screenshot-timing quirk of the browser tool, not necessarily a real bug) — **treat the ≤360px layout as visually unverified** and re-check with a real device/DevTools emulation before final acceptance (M7).
|
||
- Known accepted gap carried over from M0: `npm audit` residual `esbuild`/Vite-8 dev-server-only advisory.
|
||
|
||
### M2 — Vehicle return vertical slice
|
||
- Backend additions: `app/models/idempotency.py` (`IdempotencyRecord`), migration `e7b08389f47f_idempotency_records`, `app/services/returns.py` (`register_vehicle_return` — full transaction: row locks, idempotency replay, inspection, canonical-odometer update or regression issue, vehicle status derivation, two audit events, `vehicle.returned.v1` outbox event matching `contracts/events.schema.json`, next-booking-risk lookup), `POST /api/v1/bookings/{public_ref}/return` wired in `app/api/routers/bookings.py` with required `Idempotency-Key` header.
|
||
- Frontend additions: `components/ReturnForm.tsx` (form + `ReturnResultPanel`), wired into `pages/BookingDetail.tsx` (shown only when `booking.status === "active"`; result persists via lifted state after the booking flips to `returned`).
|
||
- Commands run and verified from this checkout:
|
||
- `docker compose run --rm api pytest -q` — **26 passed**, including `tests/test_return.py` (success/canonical-update, S1 regression scenario by name, damage→blocked, idempotent replay, reject-already-returned, missing-header validation, and a real multi-threaded concurrent-submission test against Postgres asserting exactly 1×201 + 2×409).
|
||
- `docker compose run --rm api ruff check .` — All checks passed.
|
||
- `npm run build` — clean.
|
||
- `docker compose up -d --build` (all services) then `docker compose exec api python -m app.cli seed --reset`, then a full browser run of the S1 demo scenario against `BK-DEMO-RETURN`/`MO-024`: submitted 53000 km (below canonical 54820) → result panel showed `INSP-0076`, `resulting_vehicle_status: maintenance` (correctly derived, since canonical 54820 ≥ `next_service_km` 40000), `DQ-RET-0076` created, automation event queued with a real UUID, "no upcoming booking" risk; vehicle detail page confirmed odometer unchanged at 54,820 km and a "Needs attention" badge.
|
||
- Both real defects listed above (form disappearing before showing its result; `event_id` reading as `None`) were **found via the browser run, not by pytest** — the test suite asserted on API response shape/values, not on what the UI actually rendered after a status transition. Worth remembering for M3+: UI state-after-mutation bugs need a browser check, not just API tests.
|
||
|
||
### M3 — Data Quality Workbench
|
||
- `app/services/data_quality.py`: `run_scan()` implements all five rules and is called automatically at the end of `seed_loader.reset_and_seed()` (after `db.commit()` of the base seed), plus exposed as `POST /api/v1/data-quality/scan` (Operations Manager only). Idempotency is simplified from the doc's literal `(rule_type, entity_type, entity_id, evidence fingerprint)` to just `(rule_type, entity_type, entity_id)` while an issue is open — see rationale below.
|
||
- Router `app/api/routers/data_quality.py`: `GET /issues` (filters status/rule_type/severity), `GET /issues/{ref}` (adds `entity_snapshot`/`related_snapshots` for the UI), `POST /issues/{ref}/defer`, `/reject`, `/merge-customers` (Operations Manager only — enforced via `require_operations_manager`), `POST /scan`.
|
||
- **Real bug found and fixed during this milestone, before any browser check**: the first cut of DQ-03 (odometer regression) compared every historical *returned* booking's `end_odometer_km` against the vehicle's *current* `odometer_km`. Since the seed generator assigns `vehicle.odometer_km` independently of booking history (see `seed/generate_seed.py`), this is true for nearly every historical booking by construction (odometer is monotonically increasing over time, so all-but-the-latest reading is "below current") — it produced 51 false-positive issues out of 50 vehicles on first run. Fixed twice: first attempt (compare only the single most-recent booking against canonical) still produced the same problem because canonical itself is disconnected from booking history in this dataset; the working fix compares each vehicle's *own returned-booking sequence* against itself (each booking's end reading vs. the immediately preceding one, chronologically) — a self-consistency check that doesn't depend on the unrelated `vehicle.odometer_km` field at all. Final deterministic seed+scan totals: 15 CSV-seeded + 11 scan-discovered = **26** open/resolved `data_quality_issues` (breakdown: 14 vehicle_status_conflict, 5 missing_required_field, 3 possible_duplicate_customer, 3 odometer_regression, 1 booking_overlap). `tests/test_seed.py`'s exact-count assertion was updated from 15 to 26 accordingly — if the scan logic changes again, update that count.
|
||
- Idempotency simplification rationale: the doc's fingerprint-based key would make the scan blind to issues it structurally can't compute a matching fingerprint for against the CSV-seeded rows (which don't carry a fingerprint field), producing duplicate issues for the same real-world problem (e.g. a second `MO-016` overlap issue next to the seeded `DQ-DEMO-OVERLAP`). Using `(rule_type, entity_type, entity_id)` alone while open is a stricter, safe simplification: it can never falsely suppress an issue for a *different* entity, and per-entity there's realistically only one meaningful open issue of a given rule type at a time for this PoC's scope.
|
||
- Merge UI intentionally does **not** use `window.confirm()` — a native dialog blocks further automation/testing and isn't screen-reader-distinguishable from page content the same way a rendered `role="alertdialog"` panel is. Built an inline two-step confirm instead (`ReturnForm`-style pattern reused).
|
||
- `AttentionItem` gained an `issue_ref` field (dashboard now links attention items straight to `/data-quality/{issue_ref}` instead of only to vehicles); dashboard attention list capped at 8 items (was unbounded, would have shown up to 26 with the richer scan).
|
||
- Commands run and verified from this checkout:
|
||
- `docker compose run --rm api pytest -q` — **35 passed** (new `tests/test_data_quality.py`: all five rule types present, scan idempotent on rerun, scan requires Operations Manager, S2/S4 issue-detail snapshots correct, defer→reject-on-closed 409, merge requires Operations Manager, merge rejects an unrelated survivor ref, full S2 merge scenario asserting rewiring + audit + replay-is-409).
|
||
- `docker compose run --rm api ruff check .` — All checks passed.
|
||
- `npm run build` — clean (had to fix two `possibly 'null'` TS errors from a closure-narrowing limitation — TS doesn't narrow `const` captured-by-closure across nested function boundaries when the value comes from an index/property expression; fixed by re-binding to explicitly-typed local consts right after the guard).
|
||
- Full browser run: Data Quality list (26 open issues, filterable) → `DQ-DEMO-DUPLICATE` two-column compare (CUS-0012 vs CUS-0178, per-field diff highlighting only where they differ) → merge with inline confirm → issue flips to `resolved` → confirmed `customer_merged` audit event with correct actor/entity/correlation → `DQ-DEMO-OVERLAP` (non-duplicate type) renders evidence JSON + defer/reject, no dead compare UI shown for a rule type it doesn't apply to.
|
||
|
||
### M4 — n8n automation
|
||
- `app/services/dispatcher.py`: background daemon thread (started/stopped via FastAPI `lifespan`, not an `on_event` hook) polling every `N8N_DISPATCH_INTERVAL_SECONDS` (default 3s). Claim step (`_claim_due_events`) is a short transaction using `SELECT ... FOR UPDATE SKIP LOCKED` that only flips `pending`→`delivering` and commits immediately; the HTTP call to n8n happens with **no open transaction**; the outcome is recorded in a separate short transaction. Exponential backoff `min(2**attempts, 60)` seconds, `N8N_MAX_ATTEMPTS=5` before a permanent `failed`.
|
||
- Dispatcher reconstructs the wire event from `contracts/events.schema.json`'s exact fields (`event_id`, `event_type`, `occurred_at`, `correlation_id`, `aggregate`, `data`) rather than forwarding `OutboxEvent.payload_json` wholesale — that column also carries an internal `aggregate_ref` convenience key (used by dashboard/workflows list rendering) that the schema's `additionalProperties: false` would reject.
|
||
- `POST /api/v1/integrations/n8n/return-callback` (`app/api/routers/integrations.py`): shared-secret auth via `X-Service-Token` header (`N8N_CALLBACK_TOKEN`, propagated to both `api` and `n8n` containers as `MOBILITYOPS_CALLBACK_TOKEN`); idempotent by `Idempotency-Key` (the event UUID) — checked by querying for an existing `AuditEvent` with that event ID in its metadata, **not** by `OutboxEvent.external_run_id`, because the dispatcher only sets that field *after* it gets n8n's final response, which happens *after* n8n has already called this callback mid-workflow — using `external_run_id` as the idempotency guard would have missed the exact redelivery case it's meant to catch.
|
||
- `GET /api/v1/workflows` + `POST /api/v1/workflows/{event_id}/retry` (`app/api/routers/workflows.py`), both Operations Manager only. Retry only allowed from `failed`; sets `pending` + clears `next_attempt_at` so the live dispatcher picks it up on its next cycle (does not reset `attempts`, so the counter reflects true delivery history).
|
||
- Automation nav + page (`pages/Automation.tsx`): table of all runs with status/attempts/last error, Retry button for `failed` rows, visible only to Operations Manager (matches backend authorization rather than just hiding a link).
|
||
- **Two real bugs found and fixed, the second only by testing the actual live n8n round-trip, not by pytest**:
|
||
1. Seed-loaded `workflow_runs.csv` rows only ever got `payload_json = {"aggregate_ref": ...}` (no `correlation_id`/`aggregate`/`data`) — fine for M1–M3 since nothing read those keys yet, but once the dispatcher tried to *redeliver* a seeded row (i.e. the S5 manual-retry demo scenario) it crashed with `KeyError: 'correlation_id'`, leaving that event stuck in `delivering` forever (the crash happened before the outcome-recording transaction). Fixed in two places: `seed_loader.py` now builds the full schema-compliant envelope for every `workflow_runs.csv` row (matching what the live M2 return flow produces), and `dispatcher._deliver_one` now catches malformed-payload `KeyError`s defensively and resolves the row to `pending`/`failed` instead of leaving it orphaned — added `test_deliver_one_handles_malformed_payload_without_getting_stuck` as a regression test for the latter.
|
||
2. This n8n image (2.32.7) has dropped `N8N_BASIC_AUTH_ACTIVE` as a UI/API gate — it requires an actual owner account via the `/setup` flow before anything (including webhook registration reliability) works correctly. Also: `n8n import:workflow` requires the workflow JSON to have a top-level `"id"` field (added `"id": "mobilityops-return-processing"`) and **always deactivates** the imported workflow regardless of its `"active"` field — activation requires `n8n publish:workflow --id=<id>` followed by a full n8n restart (documented in n8n 2.x CLI, not obvious from the docs pack). Did this manually this session via the CLI + browser setup wizard; **this is a one-time operational step that is not automated** — a truly clean checkout still needs someone to run `docker compose exec n8n n8n import:workflow --input=//imports/mobilityops-return-processing.json`, `docker compose exec n8n n8n publish:workflow --id=mobilityops-return-processing`, `docker compose restart n8n`, and complete the one-time owner setup at `http://localhost:5678/setup` (any email/password, no verification required) before the automation demo will work. `docs/17-runbook.md` should get this exact sequence in M7.
|
||
- Commands run and verified from this checkout:
|
||
- `docker compose run --rm api pytest -q` — **49 passed** (new `tests/test_dispatcher.py` — claim/deliver success/failure/backoff/exhaustion-to-failed/malformed-payload, all via `monkeypatch.setattr(dispatcher.httpx, "post", ...)`, no real network calls in tests; `tests/test_integrations.py` — callback auth, unknown-event 404, idempotent-by-event-ID with a real duplicate-call assertion; `tests/test_workflows.py` — role gating, retry-only-from-failed, S5 retry-and-audit).
|
||
- `docker compose run --rm api ruff check .` — All checks passed.
|
||
- `npm run build` — clean.
|
||
- Full live round trip (not mocked): registered a real return on `BK-DEMO-RETURN` → outbox event queued → background dispatcher delivered it to the now-activated n8n workflow within its 3s poll interval → n8n called back into `/api/v1/integrations/n8n/return-callback` (200 OK, confirmed in `docker compose logs api`) → dispatcher's original POST received n8n's success response → event flipped to `succeeded` on attempt 1, visible on `/automation`.
|
||
- S5 scenario end-to-end in the browser: seeded `BK-H-0020` (`failed`, 3 attempts, "Synthetic connection timeout to n8n") → clicked Retry → `pending` → within ~3s, live dispatcher delivered it through the real n8n instance → `succeeded`, 4 attempts. This is the full documented S5 scenario working for real, not simulated.
|
||
|
||
### M5 — RAGcore knowledge integration
|
||
- `app/services/knowledge/__init__.py`: `KnowledgeProvider` Protocol (sync, not async — the rest of the backend is sync SQLAlchemy/FastAPI, so an async provider interface would have meant bridging paradigms for no benefit) with `health()`/`ask()`, plus `GroundedAnswer`/`SourceCard`/`KnowledgeHealth` Pydantic models matching `contracts/openapi.yaml`'s `GroundedAnswer` schema exactly. `get_knowledge_provider()` factory switches on `settings.knowledge_provider` ("demo" default, "ragcore" opt-in).
|
||
- `app/services/knowledge/demo.py` — `DemoKnowledgeProvider`: parses the 10 `knowledge/procedures/*.md` files' YAML frontmatter (hand-rolled flat parser, not PyYAML — avoided adding a dependency for a 6-key flat block) and `## `-delimited sections at startup, then does **TF-IDF-weighted keyword retrieval** (not naive keyword counting) with light suffix-stripping stemming (`returns`→`return`, `damaged`→`damage`). This is extractive, not generative: it returns real excerpts and a templated answer sentence, never invented text.
|
||
- **Real bug found and fixed by testing the actual S6 question, not by inspection**: naive flat keyword-overlap scoring (first cut) let the word "vehicle" — present in nearly every document's title — crowd out the actually-relevant `damage-procedure` document from the top-3 results for "What must I do when a vehicle returns with damage?", because generic words scored the same as distinctive ones. Fixed by computing corpus-wide IDF per token (`log((N+1)/(df+1)) + 1`) and weighting matches by it, so common terms contribute little and rare/distinctive terms (like "damage") dominate the ranking. Verified: the S6 question now returns `damage-procedure` and `vehicle-return-procedure` in the top 3, matching the documented expectation exactly.
|
||
- `app/services/knowledge/ragcore.py` — `RAGcoreKnowledgeProvider`: real `httpx` adapter guessing a plausible REST contract (`GET /health`, `POST /api/v1/ask`) per `contracts/ragcore-contract-assumptions.md` (RAGcore is built separately; no live instance was reachable this session to verify against). Any connection error, timeout, or malformed response degrades to `evidence_state: "unavailable"` rather than raising — this is the adapter that actually exercises the architecture's "RAGcore failure disables knowledge answers only" reliability boundary. Not wired as the active provider by default; `KNOWLEDGE_PROVIDER=ragcore` would need a real, verified base URL to turn on.
|
||
- `POST /api/v1/knowledge/questions` + `GET /api/v1/knowledge/status` (`app/api/routers/knowledge.py`). Audit event `knowledge_question_asked` logs `evidence_state`, `provider`, `source_ids`, and `question_length` only — **not** the question text itself, per `docs/12-security-and-audit.md` ("log question metadata and source IDs, not unnecessary full prompts").
|
||
- Knowledge nav + page (`pages/Knowledge.tsx`): chat-style question box, source cards (title/version/section/excerpt) prioritized over the answer text per `docs/06-ui-ux.md`, explicit `grounded`/`insufficient`/`unavailable` states with distinct visual treatment — never a fabricated-looking answer for the latter two.
|
||
- Dockerfile now also `COPY knowledge ./knowledge`; added `KNOWLEDGE_DIR` setting (`/app/knowledge/procedures` in-container, same pattern as `SEED_DIR`) rather than deriving the path from `__file__` — simpler and doesn't break if the module moves.
|
||
- Commands run and verified from this checkout:
|
||
- `docker compose run --rm api pytest -q` — **57 passed** (new `tests/test_knowledge.py`: S6 grounded-with-expected-sources, unrelated question is honestly insufficient with no fabrication, demo provider health/document count, endpoint auth required, audit doesn't leak question text, RAGcore adapter degrades to unavailable on a simulated connection error).
|
||
- `docker compose run --rm api ruff check .` — All checks passed.
|
||
- `npm run build` — clean.
|
||
- Full browser run of S6 end-to-end: asked "What must I do when a vehicle returns with damage?" on `/knowledge` → grounded answer citing "Vehicle return procedure" (2 sections) and "Damage handling procedure" with real excerpts. Also asked an unrelated question ("What is the weather forecast for tomorrow?") → correctly returned "Insufficient evidence" / "No matching procedure was found" with zero sources, confirming no fabrication.
|
||
|
||
### M6 — ITWorx MCP Hub publication
|
||
- `app/api/routers/mcp_integrations.py`: four read-only endpoints under `/api/v1/integrations/mcp/` — `GET operations-summary`, `GET attention-vehicles` (query params `minimum_severity`/`date`/`limit` matching `contracts/mcp-tools.json`'s `inputSchema` exactly), `GET vehicles/{vehicle_ref}`, `POST search-knowledge` (the "narrow façade" the doc calls for — wraps M5's `get_knowledge_provider()` rather than re-implementing retrieval; the contract's tool has no MobilityOps `endpoint` field, only `routing.preferred: ragcore`, so this façade path is MobilityOps's own addition for when the Hub needs a single provider boundary, not literally specified by the contract).
|
||
- Auth: new `require_mcp_service_token` dependency in `app/api/deps.py`, same shared-secret-header shape as the M4 n8n callback (`X-Service-Token` against `MCP_HUB_SERVICE_TOKEN`) plus an optional `X-Client-Id` header (defaults to `"unknown-mcp-client"`) used as the audit actor label — the Hub's actual client-identity header name is unknown (no live Hub to confirm against), so this is a reasonable guess documented here rather than assumed silently.
|
||
- `McpVehicleDetailOut` deliberately omits `registration_number` and all customer data — narrower than the browser-facing `VehicleOut`/`VehicleDetailOut`, matching "no customer or vehicle database access" and the read-only/summary intent of an AI-facing tool. Test `test_vehicle_details_known_ref` asserts the field's absence explicitly so a future change can't silently widen the exposed surface.
|
||
- Extracted `app/services/operations.py` (`compute_metrics`, `list_attention_vehicles`) out of `app/api/routers/dashboard.py` so the MCP operations-summary/attention-vehicles endpoints and the human dashboard share one query implementation instead of two copies that could drift — the same "do not duplicate retrieval logic" principle the doc states for the knowledge tool, applied here to the operational-summary tools too.
|
||
- Every provider call writes an `AuditEvent` (`actor_type="service"`, `actor_label=X-Client-Id`, `action="mcp_tool_request"`, `metadata={tool, status}`) — MobilityOps's own record that its provider APIs were reached, independent of whatever central tool-call audit the Hub itself keeps (per `docs/10-mcp-hub-integration.md`'s audit section, the Hub owns the central log; this is the local corroborating one).
|
||
- No write/mutation endpoints exist under the `/api/v1/integrations/mcp/` namespace at all (verified by `test_no_write_endpoints_exist_under_mcp_namespace` — POST/PUT/DELETE against the vehicle-details path all 404/405) — return registration, customer merge, and any booking/vehicle mutation are correctly absent, per the doc's explicit restriction list.
|
||
- Commands run and verified from this checkout:
|
||
- `docker compose run --rm api pytest -q` — **66 passed** (new `tests/test_mcp_integrations.py`: token-required, wrong-token 401, all four tools' happy paths, severity/limit filtering, 404 for unknown vehicle, `max_sources` respected, audit actor/action verified, write-method rejection).
|
||
- `docker compose run --rm api ruff check .` — All checks passed.
|
||
- Live `curl` verification against the running stack (no browser needed — these are service-to-service endpoints, not UI): missing header → 422; wrong token → 401; correct token → all four endpoints return correct data (`operations-summary` metrics match the dashboard; `attention-vehicles?minimum_severity=high` returned `MO-016`×2 and `MO-031`, all severity `high`; `vehicles/MO-016` returned the narrow read-only shape; `search-knowledge` with `max_sources=2` returned exactly 2 grounded sources for the S6 question). Confirmed via `GET /api/v1/audit?action=mcp_tool_request` that all four calls were recorded with correct `actor_type=service`, tool name, and status.
|
||
|
||
### M7 — Portfolio polish and final acceptance
|
||
- **Automated clean-checkout migrations**: `backend/entrypoint.sh` now runs `alembic upgrade head` before starting uvicorn (Dockerfile `CMD` changed from `uvicorn ...` to `./entrypoint.sh`). Verified with a true `docker compose down -v` (all volumes wiped) → `docker compose up --build -d` → all 11 tables present, `/health` and web both green, all 66 backend tests pass, with zero manual migration step.
|
||
- **n8n one-time setup scripted where it can be**: `make n8n-setup` runs the import/publish/restart sequence (previously three manual commands discovered ad hoc in M4). The owner-account creation itself cannot be scripted safely (it's an interactive one-time step in n8n 2.x's own onboarding, not a MobilityOps concern) — documented precisely in the rewritten `docs/17-runbook.md`, including the exact URL and that no email verification is required. Re-ran this full sequence from the wiped-volumes state this session and confirmed the S1 return → outbox → live n8n → callback → `succeeded` round trip works on a genuinely clean checkout, not just the already-provisioned stack from M0–M6.
|
||
- **Playwright E2E** (`frontend/e2e/demo.spec.ts`, `frontend/playwright.config.ts`): one test automating the full 9-step documented demo script end-to-end against the live stack — login, dashboard metrics, open `BK-DEMO-RETURN`, register an odometer-regression return (S1), verify the quality issue + queued automation event, merge the duplicate-customer scenario (S2), ask the damage question and verify both expected source citations (S6), inspect audit entries, and verify responsive nav + no horizontal overflow at 360px width. **Passing.** This also resolves the "≤360px layout visually unverified" gap flagged back in M1 — verified both by this test's overflow assertion and by the `9-mobile-dashboard.png` screenshot (nav wraps into rows, metric tiles collapse to a 2-column grid, no horizontal scroll).
|
||
- Added `frontend/e2e/_capture-screenshots.spec.ts` as evidence-generation tooling (underscore-prefixed, excluded from the default `playwright test` / `make e2e` run via `testIgnore` in the config — it calls `demo/reset`, which a real regression test shouldn't do as a side effect). Captured all 9 screenshots into `artifacts/evidence/screenshots/`.
|
||
- Wrote `artifacts/evidence/architecture.md` (mermaid, as-built — distinguishes verified-live components from implemented-but-never-reached-a-real-instance ones, i.e. RAGcore and the MCP Hub) and `artifacts/evidence/final-summary.md` (commit, exact commands, test counts, screenshot index, RAGcore success/unavailable evidence — including a live-demonstrated unavailable case against an unreachable host, not just the unit test — n8n success/retry evidence, MCP sample calls, known limitations, truthful portfolio wording per `docs/16-portfolio-case-study.md`'s template).
|
||
- Updated `README.md` (dropped stale "minimal bootable scaffold, not the finished application" wording and the old two-line quickstart in favor of `make demo` + a pointer to the runbook) and `docs/17-runbook.md` (full rewrite: exact bootstrap, n8n one-time setup, verification commands, required operational checks, recovery expectations).
|
||
- Final placeholder/dead-UI sweep: `grep`'d the full `frontend/src` and `backend/app` trees for scaffold/TODO/FIXME/"must be replaced" markers — none found. `FILE_INDEX.md` was left as-is; it's the original build-pack's archive-completeness manifest (a historical snapshot), not a living index that needs to track every file added since — updating it would misrepresent what it's for.
|
||
- Commands run and verified from this checkout (this milestone, cumulative across the whole build):
|
||
- `docker compose run --rm api pytest -q` — **66 passed**, ruff clean.
|
||
- `cd frontend && npm run build` — clean.
|
||
- `cd frontend && npx playwright test` — **1 passed** (full demo script, live stack).
|
||
- Full clean-checkout drill: `docker compose down -v` → `docker compose up --build -d` → `docker compose exec api python -m app.cli seed --reset` → `docker compose run --rm api pytest -q` (66 passed) → n8n owner setup + `make n8n-setup` → live S1 return round-tripped through the real n8n instance to `succeeded`.
|
||
|
||
### Final acceptance audit (post-M7)
|
||
|
||
A dedicated release-readiness audit was run after M7 claimed completion, specifically to
|
||
catch anything the milestone-by-milestone build might have missed by only ever validating
|
||
each piece in isolation.
|
||
|
||
- **Real gap found: `mypy` had never been run.** `mypy` is a declared dev dependency
|
||
(`backend/pyproject.toml`) but was never wired into any milestone's validation loop —
|
||
only `ruff` was. Running it cold surfaced **43 real type errors across 10 files**, all
|
||
pre-existing (not introduced by this audit). Triaged and fixed all of them rather than
|
||
suppressing:
|
||
- `services/returns.py`: the vehicle lookup after acquiring `FOR UPDATE` could type as
|
||
`Vehicle | None` with no runtime guard — added an explicit `if vehicle is None: raise
|
||
AppError(..., 404)`. This was a genuine defensive-programming gap (a dangling FK would
|
||
have crashed with an unhandled `AttributeError`/500 instead of a clean 404), not just
|
||
a type annotation issue.
|
||
- `api/routers/bookings.py`: same pattern for `db.get(Customer, ...)` /
|
||
`db.get(Vehicle, ...)` in `get_booking` — added a guard raising 500 with a clear
|
||
message instead of crashing on `None.public_ref`.
|
||
- `api/deps.py` + `api/routers/demo.py`: `CurrentUser.role` is a `Literal[...]`, but
|
||
`SessionPayload.role` (decoded from an HMAC-signed cookie) and `User.role` (a DB
|
||
column) are both plain `str`. Pydantic validates this at runtime already (so it was
|
||
never exploitable), but `get_current_user` now explicitly checks membership before
|
||
constructing `CurrentUser`, turning a would-be unhandled `ValidationError` (500) into
|
||
a clean 401 for a corrupted/tampered cookie — another real defensive improvement, not
|
||
just a type-checker appeasement.
|
||
- `api/routers/dashboard.py`, `api/routers/data_quality.py`: two instances of reusing
|
||
one variable name for both a `Vehicle` and a `Customer` across an if/else branch,
|
||
which is genuinely confusing to read regardless of what mypy thinks — renamed to
|
||
distinct variables (`entity`/typed union in dashboard, `customer`/`vehicle` in the
|
||
data-quality snapshot helper).
|
||
- `services/data_quality.py`, `seed_loader.py`: `Booking.__table__.update()` /
|
||
`Customer.__table__.update()` don't typecheck against SQLAlchemy 2.0's stubs (the
|
||
`.__table__` accessor is typed as the more general `FromClause`, which doesn't
|
||
declare `.update()`) — switched to the idiomatic `sqlalchemy.update(Model)` construct,
|
||
which is both correctly typed and the more modern SQLAlchemy 2.0 style anyway.
|
||
- Remaining handful (schemas.py's deprecated `conint()` → `Annotated[int, Field(...)]`,
|
||
a `Sequence` vs `list` `.sort()` call, an `assert`-guarded None-narrowing after a
|
||
`WHERE ... IS NOT NULL` filter mypy can't see through, `Result.rowcount` typing gaps)
|
||
were either latent pydantic-v1-style API usage or genuine SQLAlchemy stub limitations
|
||
— fixed with the idiomatic modern equivalent or a narrowly-scoped, commented
|
||
`# type: ignore[...]` at the exact line, never a blanket suppression.
|
||
- `make lint` now runs both `ruff check .` and `mypy app`; `mypy app` reports
|
||
**zero errors across 44 source files**.
|
||
- **No other defects found.** Re-ran the full journey matrix end-to-end against a
|
||
genuinely wiped-volumes (`docker compose down -v`) clean checkout: all 66 backend
|
||
tests, ruff, ✅; ran the demo login → dashboard → vehicle/booking detail → return
|
||
workflow → invalid-mileage rejection (422, both a negative value and a non-numeric
|
||
string) → data-quality issue review → duplicate-customer merge → audit trail →
|
||
Knowledge Assistant → live n8n round trip → MCP Hub endpoint journeys directly via
|
||
`curl` against the running stack, all correct.
|
||
- **Degraded-mode behavior explicitly re-verified live** (not just unit-tested):
|
||
stopped n8n with `docker compose stop n8n`, registered a return — it committed
|
||
(`201`, booking flipped to `returned`) exactly as required; the outbox event stayed
|
||
`pending` with real `ConnectError`s logged and exponential backoff (2 attempts over
|
||
~8s); restarted n8n and the dispatcher **self-healed** without any manual
|
||
intervention, delivering the event to `succeeded` on attempt 5. RAGcore unavailable
|
||
mode re-verified live against an unreachable host (`ConnectError` → `evidence_state:
|
||
"unavailable"`, empty answer, no fabrication). MCP Hub unavailability is
|
||
architecturally moot for MobilityOps — the Hub only ever calls *into* MobilityOps, so
|
||
there is nothing on the MobilityOps side that can degrade if the Hub is down (only the
|
||
reverse, "does an unavailable Hub break MobilityOps," which is trivially no since
|
||
nothing here calls out to it).
|
||
- **New test coverage added, no existing tests weakened**: `frontend/e2e/interactive-elements.spec.ts`
|
||
(11 Playwright tests — all seven nav items, every filter on every list page, vehicle
|
||
detail tabs, defer/reject, automation retry, knowledge form, role-switching, and
|
||
role-based page restriction) plus the existing `demo.spec.ts` — **12/12 e2e tests
|
||
passing** against the live stack.
|
||
- Verified `.env` is `.gitignore`d and was never committed (`git ls-files` /
|
||
`git log --all -p -- '*.env'` both empty); scanned full git history for AWS keys,
|
||
private-key headers, and `sk-...`-style tokens — none found. Every `Settings` field in
|
||
`backend/app/core/config.py` has a corresponding entry either directly in
|
||
`.env.example` or is derived/wired through `compose.yaml` (a few purely-internal
|
||
container-path constants like `SEED_DIR`/`KNOWLEDGE_DIR` are intentionally not
|
||
operator-configurable and correctly absent from `.env.example`).
|
||
- Grepped the full `frontend/src` and `backend/app` trees for TODO/FIXME/placeholder/
|
||
fake/stub/mock/"not implemented" markers — zero real hits (the two `placeholder=`
|
||
matches are legitimate HTML input placeholder attributes). Confirmed dashboard metrics
|
||
and all list-page data are 100% DB-backed (`compute_metrics` in
|
||
`services/operations.py`, never a literal in frontend JSX). Confirmed every frontend
|
||
route in `App.tsx` maps to an implemented page and every nav item maps to a real route
|
||
— no dead routes.
|
||
- Commands run and verified from this audit:
|
||
- `docker compose run --rm api pytest -q` — **66 passed**.
|
||
- `docker compose run --rm api ruff check .` — All checks passed.
|
||
- `docker compose run --rm api mypy app` — **Success: no issues found in 44 source files** (0 errors, down from 43).
|
||
- `cd frontend && npm run build` — clean (`tsc -b && vite build`).
|
||
- `cd frontend && npx playwright test` — **12 passed** (`demo.spec.ts` + `interactive-elements.spec.ts`).
|
||
- Full clean-checkout drill repeated from a fresh `docker compose down -v`: automatic migrations, seed, 66/66 tests, n8n owner setup + `make n8n-setup`, live return round-tripped through n8n to `succeeded`.
|
||
- See `artifacts/final-acceptance/summary.md` for the complete evidence write-up (commands, exact outputs, demo access, deployment instructions, five-minute demo flow).
|
||
|
||
## Definition of done
|
||
|
||
All eight milestones (M0–M7) are complete, and a dedicated post-M7 final-acceptance audit
|
||
found and fixed one real category of gap (`mypy` never having been run) with zero
|
||
regressions. `docs/14-testing-and-acceptance.md`'s clean-checkout acceptance list has been
|
||
walked item by item against a genuinely wiped-volumes checkout, twice (once in M7, once in
|
||
this audit), and `artifacts/final-acceptance/summary.md` is the authoritative final
|
||
evidence document. The two items not fully closed — a live RAGcore instance and a live
|
||
ITWorx MCP Hub instance — were never reachable in this environment; both integrations are
|
||
implemented, unit/contract-tested, directly verified against MobilityOps's own API, and
|
||
their unavailable-degradation paths are live-verified, but an actual round trip against
|
||
real RAGcore/Hub instances remains unconfirmed and is documented as such rather than
|
||
claimed.
|
||
|
||
## Known blockers
|
||
|
||
None. External service credentials may be absent; use the documented demo/degraded providers. The n8n workflow-activation steps are a one-time manual setup requirement in this environment (owner-account creation via n8n's own `/setup` UI cannot be scripted safely), fully documented in `docs/17-runbook.md` and scripted where possible (`make n8n-setup`). RAGcore and the ITWorx MCP Hub itself were never reachable in this environment — both integrations are implemented and directly tested/curl-verified against MobilityOps's own API, but neither a real RAGcore instance nor a real Hub round trip was available to confirm end-to-end.
|
||
|
||
## Premium Control Rail UI transformation (2026-08-02)
|
||
|
||
- Branch: `design/mobilityops-premium-ui`, branched from verified deployed revision
|
||
`dfabb41582e302f45a3de826f85f531bf23dfc8b`; master history was not rewritten.
|
||
- Audited every route at 1440, 1280, 768 and 390 px. Baseline findings and captures are
|
||
in `docs/design/current-ux-audit.md` and `artifacts/design-validation/current/`.
|
||
- Authored three twelve-screen product directions and generated representative Stitch
|
||
anchors in project `17018847755558569017`: Control Rail, Dispatch Ledger and Service
|
||
Atelier. Control Rail was selected and refined twice for hierarchy, accessibility and
|
||
responsive implementation. Decision, screen inventory, tokens and exact Stitch IDs
|
||
are in `docs/design/design-directions.md`, `docs/design/design-system.md` and
|
||
`docs/design/stitch-manifest.md`.
|
||
- Rebuilt the complete React interface around a responsive Control Rail shell: inline SVG
|
||
icon/brand system, desktop rail, named landmarks, skip link, top bar, mobile bottom
|
||
navigation, shared loading/error/empty states and reduced-motion support.
|
||
- Redesigned all shipped pages. The dashboard now prioritizes persisted readiness,
|
||
Attention and today's movements; booking results paginate at 25 rows; every responsive
|
||
table retains field labels; integrations distinguish n8n evidence, live RAGcore health
|
||
and the unconfigured MCP adapter without inventing status.
|
||
- Return registration is now capture → review → result. A regression test proves the
|
||
return endpoint is not called before confirmation; the existing idempotency and local
|
||
commit/outbox contract is unchanged.
|
||
- Final browser captures are in `artifacts/design-validation/implementation/`. DOM
|
||
measurements and Playwright both prove no horizontal overflow at 390, 768, 1280 and
|
||
1440 px. See `docs/design/implementation-validation.md`.
|
||
- Final validation commands from this branch:
|
||
- `docker compose run --rm api pytest -q` — **66 passed**.
|
||
- `docker compose run --rm api ruff check .` — **All checks passed**.
|
||
- `docker compose run --rm api mypy app` — **0 issues in 44 files**.
|
||
- `cd frontend && npm run lint` — clean TypeScript check.
|
||
- `cd frontend && npm run build` — production build succeeded (59 modules; 240.24 kB JS,
|
||
36.63 kB CSS before gzip).
|
||
- `cd frontend && playwright test --reporter=line` — **19 passed** including the full
|
||
five-minute demo, every interactive route, return review semantics and four viewport
|
||
overflow checks.
|
||
- Review deployment updated at `http://192.168.10.150:1236` with persistent PostgreSQL
|
||
data preserved. Deployed smoke: all ten authenticated routes plus login at
|
||
desktop and mobile sizes rendered without alert state or horizontal overflow; browser console had zero
|
||
warnings/errors; seven authenticated API paths returned 200; PostgreSQL/API were
|
||
healthy and the shared n8n `/healthz` returned `{"status":"ok"}`.
|
||
- Corrected the review topology after confirming the host already runs n8n on port 5678:
|
||
the temporary `mobilityops-n8n-1` container was removed without deleting its retained
|
||
volume; the bundled service is now opt-in through the `bundled-n8n` profile; the API
|
||
points to the shared n8n; and the return workflow is imported and published there.
|
||
- The existing n8n's previously empty `N8N_HOST` and `N8N_EDITOR_BASE_URL` values were
|
||
persistently set in its Unraid template. A synthetic return then completed the full
|
||
MobilityOps → shared n8n → callback round trip as `succeeded` on attempt 1, after which
|
||
deterministic demo state was restored (`BK-DEMO-RETURN` is `active`).
|
||
- Global search is now live for Control Rail sections and `MO-*`, `BK-*`, `DQ-*` public
|
||
references, including Ctrl/Cmd+K focus and a tested not-found announcement. Final local
|
||
Playwright result is **19 passed**.
|
||
- Exact next action: hand off `design/mobilityops-premium-ui` for review. The final code,
|
||
shared-n8n topology and evidence are committed, pushed and deployed; do not merge master
|
||
automatically.
|
||
|
||
## Functional completion pass (branch `feat/mobilityops-functional-completion`)
|
||
|
||
Branched from `design/mobilityops-premium-ui` @ `54dc952`. Full audit at
|
||
`docs/functional-completion/current-functional-audit.md`; server baseline captured before
|
||
any change at `docs/functional-completion/server-baseline.md`.
|
||
|
||
### Batch 1 — complete (commits `938a739`..`bdc58f3`)
|
||
|
||
- Fixed the two confirmed list-rendering defects: Vehicles and Bookings both computed a
|
||
filtered/paginated result but rendered the raw unfiltered array in the table body.
|
||
- Added server-backed session lifecycle: `GET /api/v1/demo/session` (Cache-Control:
|
||
no-store — a cached 200 was making logout intermittently fail to redirect in e2e
|
||
testing), `POST /api/v1/demo/logout`. `AuthContext` now verifies against the server on
|
||
every mount instead of trusting `sessionStorage`, and a central 401 listener on the API
|
||
client clears auth state from any endpoint.
|
||
- Enforced the brief's role matrix: data-quality (list/detail/defer/reject) and the audit
|
||
trail were reachable by Rental Employee with no gate beyond authentication (confirmed
|
||
live via curl before the fix). Both are now `require_operations_manager`-gated
|
||
server-side, with matching nav-hiding and a restricted-message fallback for direct URL
|
||
access, and the dashboard no longer links into those areas for that role.
|
||
- Discovered and fixed a latent e2e-suite bug while testing against the real server: all
|
||
three spec files hardcoded `http://localhost:8128` for their demo-reset helpers, so
|
||
pointing the suite at Unraid via `MOBILITYOPS_PUBLIC_URL` silently kept resetting the
|
||
*local* dev database instead. Switched to relative paths so the configured `baseURL` is
|
||
honoured.
|
||
- Local evidence: `pytest` 75 passed, `ruff check .` clean, `mypy app` 0 issues/44 files,
|
||
`npx tsc -b` clean, `npm run build` clean, `npx playwright test` **25 passed** (up from
|
||
19 — 6 new tests this batch), stable across three repeated full-suite runs.
|
||
- Deployed to Unraid (`.deploy/source-revision` = `bdc58f396e99caaf6ef657bb110b479981cc7793`,
|
||
matches `git rev-parse HEAD` on the feature branch), migrations unchanged at
|
||
`e7b08389f47f (head)` (no schema change this batch), demo reset run. Re-verified live:
|
||
role-gate curl checks (403/200/401 as expected) and the full 25-test Playwright suite
|
||
run with `MOBILITYOPS_PUBLIC_URL=http://192.168.10.150:1236` — **25 passed** against the
|
||
actual deployment, not just localhost.
|
||
- Exact next action: Batch 2 — authoritative return-preview endpoint shared with commit,
|
||
and expose `before`/`after` on the audit API + UI.
|
||
|
||
### Batch 2 — complete (commits `f521295`, `7e34f55`)
|
||
|
||
- Added `evaluate_return()` (pure, no writes) in `returns.py`, extracted from what
|
||
`register_vehicle_return` already computed inline; `register_vehicle_return` now calls
|
||
it instead of duplicating the logic. New non-mutating `POST
|
||
/api/v1/bookings/{ref}/return-preview` uses the same function, so preview and commit
|
||
cannot drift.
|
||
- Fixed a real defect this surfaced: `ReturnForm.tsx`'s review step guessed the outcome
|
||
client-side and got the domain rule wrong — it said damage/technical-warning routes to
|
||
`maintenance` (actual rule: `blocked`) and the no-contradiction case becomes
|
||
`available` (actual rule: always `cleaning` first, `maintenance` only past the service
|
||
threshold). The review step now calls `/return-preview` and renders the server's
|
||
`resulting_vehicle_status` + `status_reason` verbatim.
|
||
- Result screen now distinguishes local commit success from n8n delivery ("queued... not
|
||
yet confirmed" instead of implying both succeeded) and links to any created
|
||
data-quality issue for Operations Manager.
|
||
- Exposed `before`/`after` on `AuditEventOut` (the DB columns already existed but were
|
||
never serialized) plus a resolved `entity_ref`/`entity_link` for vehicle/booking/
|
||
data-quality-issue entities. `Audit.tsx` now shows a human-readable change summary per
|
||
row with raw JSON behind a `<details>` disclosure instead of always-visible JSON.
|
||
- New regression coverage: backend — preview performs no writes (asserted via audit/
|
||
outbox row counts before vs. after), detects odometer regression, detects service-due,
|
||
detects next-booking risk, requires an active booking, and matches the commit result;
|
||
audit — before/after and entity link exposed for both `return_registered` and
|
||
`vehicle_status_changed`. Frontend — preview correctly reports `blocked` (not
|
||
`maintenance`) for damage, commit request only fires after confirm (updated to also
|
||
assert exactly one preview call), audit page shows before/after and a safe link.
|
||
- Local evidence: `pytest` 81 passed, `ruff check .` clean, `mypy app` 0 issues/44 files,
|
||
`npx tsc -b` clean, `npm run build` clean, `npx playwright test` **27 passed**, stable
|
||
across two repeated full-suite runs.
|
||
- Deployed to Unraid and re-verified; demo data reset afterward.
|
||
- Exact next action: Batch 3 — data-quality workbench (typed snapshots, bounded
|
||
resolution flows for all 5 rule types, manual scan UI).
|
||
|
||
### Batch 3 — complete (commits `6e227a2`, `477b5e7`)
|
||
|
||
- Typed related-entity snapshots by the reference's own prefix (CUS-/MO-/BK-/INSP-)
|
||
instead of inferring from `rule_type`. Fixed a real gap this exposed: a
|
||
`booking_overlap` issue's related refs are bookings, but `get_issue` always resolved
|
||
them as vehicles, so `_snapshot()` silently returned nothing for them.
|
||
- Added one bounded resolution endpoint per remaining rule type: `provide-fields`
|
||
(missing_required_field; re-runs the check, resolves only once nothing required is
|
||
missing), `resolve-odometer-regression` (retain canonical or correct the reading —
|
||
correction is rejected if it would still be below canonical), `resolve-overlap`
|
||
(blocks one of the two bookings, re-verifies no overlap remains — found and fixed an
|
||
autoflush=False bug where the re-verification query didn't see the just-blocked
|
||
booking's in-memory status change), `apply-recommended-status` (one authoritative
|
||
recommendation function mirroring the scan's own conflict conditions, re-validated
|
||
after applying). `possible_duplicate_customer` already had merge; all five rule types
|
||
now have a real bounded resolution path, not just generic defer/reject.
|
||
- Reintroduced evidence after a non-open decision links the new issue back to the prior
|
||
one (`evidence.reopened_from` / `previous_decision`) per the documented lifecycle
|
||
("reintroduced evidence creates a new issue linked to the prior issue").
|
||
- `DataQualityIssueDetail.tsx` rewritten: a typed panel per rule type instead of a raw
|
||
`JSON.stringify` dump for four of five types; raw evidence moved behind a `<details>`
|
||
disclosure. Added a "Run quality scan" action to the workbench (confirmation,
|
||
progress, per-rule result counts, auto-refresh) — the scan endpoint already existed
|
||
with no UI trigger.
|
||
- New regression coverage: backend — one resolution test per rule type plus the
|
||
role-gate/validation-rejection paths and the recurrence-linking behavior (reject an
|
||
issue, rescan, assert the new issue links back). Frontend — one Playwright test per
|
||
resolution flow plus the manual scan trigger.
|
||
- Local evidence: `pytest` 96 passed, `ruff check .` clean, `mypy app` 0 issues/44 files,
|
||
`npx tsc -b` clean, `npm run build` clean, `npx playwright test` **32 passed**, stable
|
||
across two repeated full-suite runs.
|
||
- Deployed to Unraid and re-verified against the live server; demo data reset afterward.
|
||
- Exact next action: Batch 4 — global search backend + UI, demo reset UI trigger,
|
||
truthful aggregate integration status (n8n/RAGcore/MCP).
|
||
|
||
### Batch 4 — complete (commits `4437b87`, `1867828`)
|
||
|
||
- Added `GET /api/v1/search` — bounded typed results (vehicle/booking/data-quality-issue/
|
||
section), role-filtered server-side (data-quality and manager-only sections excluded
|
||
for Rental Employee), customers never returned (no customer detail route exists).
|
||
Replaced `Layout.tsx`'s blind client-side regex/term guesser with a debounced
|
||
(250 ms) call to this endpoint, a real `role="listbox"` results panel, arrow-key
|
||
navigation, Enter/Escape, outside-click close, and a no-results state.
|
||
- Added `GET /api/v1/integrations/status`, aggregating outbox delivery counts into one
|
||
truthful n8n state (`disabled`/`unavailable`/`degraded`/`operational`/`no_evidence`)
|
||
instead of the dashboard/automation cards showing whichever status the single most
|
||
recent event happened to be in. Wired into both `Automation.tsx` and `Dashboard.tsx`.
|
||
MCP Hub card now reflects the real `registration_enabled` setting.
|
||
- Found and fixed a real config gap this surfaced: `MCP_HUB_REGISTRATION_ENABLED` was
|
||
documented in `.env.example` but had no `Settings` field, so it was silently dropped
|
||
by `extra="ignore"` and never read anywhere in the codebase.
|
||
- Added a "Reset demo data" action to the sidebar (Operations Manager only, confirm,
|
||
progress, error handling) — the endpoint already existed and was already gated, just
|
||
had no UI trigger. Reset invalidates the acting session server-side, so the flow signs
|
||
the user out and returns to login.
|
||
- New regression coverage: backend — search role-filtering/customer-exclusion/no-match,
|
||
integration-status role-gate and state-derivation (including a test that resolves all
|
||
seeded failures and asserts the state flips to `operational`). Frontend — vehicle/
|
||
booking/data-quality-issue search navigation, keyboard nav, no-results + Escape, demo
|
||
reset happy path, rental employee cannot see the reset button, automation page shows
|
||
aggregate counts.
|
||
- Local evidence: `pytest` 109 passed, `ruff check .` clean, `mypy app` 0 issues/46 files,
|
||
`npx tsc -b` clean, `npm run build` clean, `npx playwright test` **37 passed**, stable
|
||
across two repeated full-suite runs.
|
||
- Deployed to Unraid and re-verified against the live server; demo data reset afterward.
|
||
- Exact next action: Batch 5 — bounded outbox delivery-lease recovery for stale
|
||
`delivering` events, a second (scheduled quality-scan) n8n workflow, final
|
||
documentation/contract updates and acceptance evidence.
|
||
|
||
### Batch 5 — complete (commits `ec8f809`, `e115031`, `c981aad`, `824048b`)
|
||
|
||
- Fixed a real gap: `_claim_due_events` flipped rows to `delivering` and committed
|
||
before the HTTP call, with no reclaim path if the process died before the outcome was
|
||
recorded. Each claim now gets a lease deadline (`n8n_delivery_lease_seconds`, default
|
||
120s, reusing the `next_attempt_at` column) and `run_dispatch_cycle()` sweeps expired
|
||
leases back to `pending` before claiming new work; `attempts` is preserved, and a
|
||
still-alive worker's unexpired lease is never touched.
|
||
- Added the second n8n workflow: `POST /api/v1/integrations/n8n/scheduled-scan`
|
||
(service-token protected, same pattern as the return callback) running the same
|
||
`run_scan()` the manual UI action uses, audited with `actor_type=service`.
|
||
`n8n/mobilityops-scheduled-quality-scan.json` (hourly + manual-test trigger) ships
|
||
`"active": false`. Live-verified twice: executed end-to-end via the Manual test
|
||
trigger against the **local** n8n instance (full green execution, confirmed via the
|
||
resulting `data_quality_scan_run` audit event), and published + directly
|
||
curl-round-tripped against the **shared Unraid n8n** and its live API
|
||
(`deploy/unraid/setup-scheduled-scan.sh`). The shared instance's own UI could not be
|
||
browser-tested directly — it runs `N8N_SECURE_COOKIE=true` and refuses login over the
|
||
plain-HTTP LAN URL, which is correct/expected shared-infrastructure behaviour, not
|
||
something this task should change.
|
||
- Updated `contracts/openapi.yaml` and `docs/05-api-contract.md` with every endpoint
|
||
added across all five batches; `docs/07-data-quality.md`, `docs/08-return-workflow.md`
|
||
and `docs/12-security-and-audit.md` now describe the actual resolution flows, the
|
||
preview/commit relationship, the role matrix and the audit before/after exposure.
|
||
Corrected `docs/07-data-quality.md`'s lifecycle description to match the
|
||
already-implemented `(rule_type, entity_type, entity_id)` idempotency key (no evidence
|
||
fingerprint) and documented the `reopened_from`/`previous_decision` recurrence link.
|
||
`README.md`'s scope/integration-status/quality-gate sections updated to match.
|
||
- Local evidence: `pytest` **117 passed**, `ruff check .` clean, `mypy app` 0 issues/46
|
||
files, `npx tsc -b` clean, `npm run build` clean, `npx playwright test` **37 passed**.
|
||
- **Clean-checkout drill** (section 14): fresh `git clone` of this branch into an
|
||
isolated scratch directory, `.env` from `.env.example`, isolated Compose project name
|
||
and remapped host ports (no shared state with the working stack), `up --build -d`
|
||
from empty volumes → migrations ran automatically → seed → full backend gate (117
|
||
passed, ruff clean, mypy clean) → `npm ci` (clean; the pre-existing esbuild-moderate/
|
||
react-router-RSC-high advisories are unchanged, not new) → `tsc -b`/`vite build` clean
|
||
→ full Playwright suite **37 passed** against the isolated stack. Torn down afterward
|
||
(`down -v` on the isolated project only; the working dev stack was never touched).
|
||
- Deployed to Unraid; migrations unchanged at `e7b08389f47f (head)`. Full 37-test
|
||
Playwright suite re-run against `http://192.168.10.150:1236` — **37 passed**. Demo
|
||
data reset afterward.
|
||
- Exact next action: none — all five batches are implemented, tested locally (including
|
||
a genuine clean-checkout drill), committed, pushed, deployed to Unraid and
|
||
re-verified against the live server after every batch. See
|
||
`artifacts/functional-completion/final-summary.md` for the definitive acceptance
|
||
evidence.
|
||
|
||
## Demo productization (in progress, same branch `feat/mobilityops-functional-completion`)
|
||
|
||
Follows the functional-completion work above; turns the now feature-complete PoC into a
|
||
guided, honestly-labelled demo (fictional org "Northstar Mobility", guided tour, 5 named
|
||
scenarios, demo manifest, About page). Gap audit: `docs/demo-release/current-demo-gap-audit.md`.
|
||
|
||
### Batch 1 — seed date anchoring (complete)
|
||
|
||
- **Real bug fixed**: `seed/bookings.csv` etc. store absolute ISO timestamps authored
|
||
around a fixed anchor (`2026-08-01`). Nothing previously re-anchored them at seed/reset
|
||
time, so scenario bookings (e.g. `BK-DEMO-RETURN`) silently drifted into the past every
|
||
day the environment wasn't reset. `dashboard.py::_today()` compounded this by filtering
|
||
"today's movements" against the same frozen `demo_today` setting instead of real time.
|
||
- Fix: `seed_loader.py` now computes `shift = today - SEED_AUTHORED_ANCHOR` once per
|
||
`load_seed()` call and applies it to every seeded booking/inspection/maintenance/outbox
|
||
datetime column, so scenarios stay "today"/"near-future" relative to the actual reset
|
||
moment. `SeedResult` now also carries `anchor_date`/`seeded_at`; `POST /api/v1/demo/reset`
|
||
returns them; a `demo_data_seeded` audit event records the anchor for traceability.
|
||
`dashboard.py::_today()` switched from the frozen `demo_today` setting to real wall-clock
|
||
UTC date. The now-dead `demo_today` setting/env var was removed from `config.py`,
|
||
`compose.yaml`, `.env`, `.env.example` (nothing else referenced it).
|
||
- Added seed-validation tests (`backend/tests/test_seed.py`) proving S1 (`BK-DEMO-RETURN`/
|
||
`MO-024`), S2 (`CUS-0012`/`CUS-0178`/`DQ-DEMO-DUPLICATE`), S4 (`MO-016`/
|
||
`BK-DEMO-OVERLAP-A`/`-B`/`DQ-DEMO-OVERLAP`) and S5 (seeded failed outbox event
|
||
`00000000-0000-4000-8000-000000000020`, confirmed genuinely `failed` immediately after a
|
||
fresh reset, not silently auto-healed by the background dispatcher since it only claims
|
||
`pending` rows) are fully present after every reset, plus a dedicated anchoring test
|
||
asserting the shift and the audit marker.
|
||
- Live-verified locally: reseeded and confirmed via `psql` that `BK-DEMO-RETURN` now ends
|
||
today and `BK-DEMO-NEXT`/overlap bookings sit in the near future (today = 2026-08-03).
|
||
- Evidence: `pytest` **122 passed** (117 + 5 new/expanded seed tests), `ruff check .`
|
||
clean, `mypy app` clean (46 files, canonical `make` scope).
|
||
- Deployed to Unraid (commit `8989ffb`): pushed to Gitea, `git archive` tarball extracted
|
||
over `/mnt/user/appdata/mobilityops` preserving `.env`/volumes, `api` rebuilt (`db`/`web`
|
||
untouched — no frontend changes this batch), migrations confirmed at `e7b08389f47f
|
||
(head)`, reseeded, live-verified via `psql` that `BK-DEMO-RETURN`/`BK-DEMO-NEXT`/overlap
|
||
bookings sit at the same real-time-relative positions as local. `curl` to
|
||
`http://192.168.10.150:1236/` returns 200.
|
||
- Exact next action: `GET /api/v1/demo/manifest` + Dutch demo entry screen + permanent
|
||
demo badge (task #30), then the Demo Guide + scenario overview (task #31).
|
||
|
||
### Batch 2 — demo manifest, Dutch demo entry, permanent demo badge, About page (complete)
|
||
|
||
- `GET /api/v1/demo/manifest` (unauthenticated): single source of truth for demo org
|
||
identity, synthetic-data flag, reset allowance/timestamp/anchor date, guide
|
||
availability, and the 5 named scenarios with **live** readiness (queries the actual
|
||
`BK-DEMO-RETURN`/`DQ-DEMO-DUPLICATE`/`DQ-DEMO-OVERLAP`/seeded-failed-event/knowledge-
|
||
provider records — not hardcoded), plus plain-language integration summaries. Backed by
|
||
new `backend/app/services/demo_manifest.py`. Refactored the n8n status derivation out of
|
||
`integration_status.py` into a shared `services/integration_status.py` so the manifest
|
||
and the existing authenticated `/integrations/status` endpoint reuse one implementation.
|
||
- New settings (`backend/app/core/config.py`, wired through `compose.yaml`/`.env.example`):
|
||
`DEMO_ORGANIZATION_NAME` (default "Northstar Mobility" — surfaces the project's already-
|
||
locked fictitious tenant, previously only used internally as the `ragcore_tenant` slug),
|
||
`DEMO_TIMEZONE`, `DEMO_ALLOW_RESET` (a safety valve — `false` makes `POST
|
||
/api/v1/demo/reset` return 403 regardless of role; the now-dead `demo_today` setting
|
||
removed in Batch 1 stays removed).
|
||
- Rewrote `Login.tsx` in Dutch: names the fictional org, one-sentence explanation sourced
|
||
from the manifest, no password shown/copyable anywhere, "Start begeleide demo" primary
|
||
CTA (logs in as Operations Manager, navigates to `/dashboard?guide=start` for task #31 to
|
||
consume) plus "Verken als Operations Manager"/"Verken als Rental Employee" secondary
|
||
actions. Added a permanent demo badge (topbar pill + popover: synthetic notice,
|
||
"workflows are real" reassurance, last-reset timestamp, link to `/about`) replacing the
|
||
old full-width static `.demo-banner` bar — subtle by design per the brief, not a warning
|
||
bar. New `/about`
|
||
page (`AboutDemo.tsx`) covering the fictional problem, what's really implemented, what's
|
||
synthetic, honest per-integration labels (via the manifest), and a reset pointer —
|
||
reachable from the badge popover, not added to primary nav (preserves the existing Control
|
||
Rail nav per the "not a redesign" constraint). Frontend nav/design otherwise untouched.
|
||
- Evidence: `pytest` **127 passed**, `ruff check .` clean, `mypy app` clean (48 files);
|
||
frontend `tsc -b` clean, `npm run build` clean; full Playwright suite **41 passed**
|
||
(37 existing + 4 new `demo-entry.spec.ts` covering entry copy/no-password, guided-demo
|
||
login redirect, badge popover content + About link, and Escape/outside-click close).
|
||
Updated stale English login-button aria-labels and login-copy assertions across the
|
||
existing specs to match the new Dutch copy.
|
||
- Deployed to Unraid (commit `ac427f4`): pushed to Gitea, `git archive` tarball extracted
|
||
preserving `.env`/volumes, both `api` and `web` rebuilt (frontend changed this batch),
|
||
both healthy, migrations unchanged, reseeded. Live-verified: `GET
|
||
/api/v1/demo/manifest` returns `organization_name: "Northstar Mobility"`,
|
||
`allow_reset: true`, and all 5 scenarios `ready: true` right after reset. Ran
|
||
`demo-entry.spec.ts` (4 tests) and the five-minute demo script directly against
|
||
`http://192.168.10.150:1236` — **5/5 passed**. Reseeded again afterward to leave the
|
||
server demo-ready.
|
||
- Exact next action: Demo Guide (collapsible panel, 8 steps) + scenario overview (5 cards
|
||
on the dashboard, consuming `/api/v1/demo/manifest`'s `scenarios` array) — task #31.
|
||
|
||
### Batch 3 — Demo Guide + scenario overview (complete)
|
||
|
||
- New `/scenarios` page (`Scenarios.tsx`): all 5 named scenarios as cards (title,
|
||
operational problem, duration, required role(s), "toont aan", ready/blocked status
|
||
from the manifest, "Start scenario" linking to the live `start_path`). Dashboard gets
|
||
one compact "Probeer een demonstratiescenario" panel (not 5 more cards — keeps the
|
||
existing dashboard uncluttered per the brief) showing readiness count and, for
|
||
Operations Managers, a guide resume/start control.
|
||
- New Demo Guide: `DemoGuideContext` (sessionStorage-persisted `currentIndex`/`completed`
|
||
set — browser-only, never touches auth or business logic), 8 static steps
|
||
(`data/demoGuideSteps.ts`) each with what-you'll-see/why/start-action/expected-outcome,
|
||
resolving live routes from the manifest for the two scenario-backed steps (return,
|
||
duplicate-merge) so they can't drift from actual records. `DemoGuide.tsx` renders a
|
||
fixed side panel (desktop) that becomes a bottom sheet at ≤700px via CSS only (no
|
||
layout duplication); `DemoGuideTrigger` (topbar, Operations-Manager-only — the 8 steps
|
||
require OM throughout) shows a live `completed/8` pill. "Demo opnieuw voorbereiden"
|
||
calls the real reset endpoint, resets guide progress, and returns to `/login` (mirrors
|
||
the existing sidebar reset flow). Login's "Start begeleide demo" logs in as OM and
|
||
passes a one-shot `?guide=start` marker the dashboard consumes once then strips.
|
||
- Fixed a real regression caught by the responsive-overflow tests: the new topbar guide
|
||
trigger pushed `.topbar-meta` past the viewport at ≤420px; fixed by hiding the guide
|
||
trigger (icon+pill) at that breakpoint — the Dashboard's own "Start demo-gids" control
|
||
remains reachable there. Also fixed a genuine mobile overflow in the new
|
||
`.demo-start-panel` (flex items without `min-width:0`/wrap on narrow screens).
|
||
- Fixed one fragile new test (asserted on the transient `?guide=start` URL param, which
|
||
the app intentionally strips immediately — changed to assert the guide's actual open
|
||
state instead) and two Playwright strict-mode ambiguous-match errors; confirmed the
|
||
full 41+6=47-test suite passes twice in a row after these fixes (ruling out flakiness).
|
||
- Evidence: frontend `tsc -b` clean, `npm run build` clean; full Playwright suite
|
||
**47 passed** (41 existing + 6 new `demo-guide.spec.ts`: scenario overview shows 5
|
||
ready cards after reset, starting a scenario navigates to its fixed record, guide
|
||
step navigation/jump/close, progress persists across page navigation, guide hidden
|
||
from Rental Employee, restart-from-guide resets data and returns to login). Backend
|
||
untouched this batch (no re-run needed; last backend gate was 127 passed/ruff/mypy
|
||
clean in Batch 2).
|
||
- Deployed to Unraid (commits `9fff84d`, then `14c2ad3` for a test-only fix): pushed to
|
||
Gitea, tarball extracted, `web` rebuilt (frontend-only batch), healthy. Live-verified:
|
||
ran `demo-guide.spec.ts` (6), `demo-entry.spec.ts` (4) and the five-minute demo script
|
||
directly against `http://192.168.10.150:1236` — **11/11 passed**. Caught and fixed one
|
||
real environment-sensitive test bug in the process: two tests navigated straight to
|
||
`/scenarios` right after a login click without waiting for the `/dashboard` redirect,
|
||
which raced harmlessly on localhost but flaked against Unraid's higher latency — fixed
|
||
by asserting the redirect first, no app-code change needed. Reseeded afterward to leave
|
||
the server demo-ready.
|
||
- Exact next action: layer plain-language Dutch explanation onto the return flow, the 5
|
||
data-quality panels, and fix the knowledge assistant's RAGcore-naming bug — task #33.
|
||
|
||
### Batch 4 — return/data-quality/knowledge demo legibility (complete)
|
||
|
||
- **Fixed a real honesty bug**: `Knowledge.tsx` named "RAGcore" in the body copy and the
|
||
retrieval-flow diagram even though the active provider is the demo TF-IDF one (the
|
||
small badge below was already honest, contradicting the prose one line above). Now
|
||
derives a `providerLabel` ("Demo knowledge base" vs "RAGcore") from the real health
|
||
check and uses it everywhere; added an explicit disclosure note when not RAGcore.
|
||
Added 4 suggested-question chips. **Discovered and fixed a second real bug in the
|
||
process**: the brief's suggested Dutch questions (and my own Demo Guide step 6 wording)
|
||
would have returned "insufficient evidence" against the demo provider, because the
|
||
indexed procedures are English-only — verified empirically (Dutch question →
|
||
`insufficient`, its English equivalent → `grounded`). Fixed by keeping suggested
|
||
questions in English (matching the indexed content) and rewording the Guide step to
|
||
explain the knowledge base is English, rather than mistranslating the demo's
|
||
centerpiece feature into silently returning wrong answers.
|
||
- Return flow: `BookingDetail.tsx` now detects the one named return-anomaly scenario
|
||
booking (via the manifest, not a hardcoded ref) and fetches that vehicle's real
|
||
canonical odometer to pre-fill `ReturnForm`'s "End odometer" field with a suspicious
|
||
value below it, plus a callout explaining why — the brief explicitly requires the demo
|
||
not ask a visitor to invent a suspicious number themselves. Scoped narrowly to that one
|
||
scenario booking; ordinary returns are unaffected. `ReturnResultPanel` now links to
|
||
Automation and Audit trail (previously only the vehicle), and shows a "Ga verder met de
|
||
demo" button when the Demo Guide is open (advances the guide and navigates to the next
|
||
step). **Fixed a real regression caught by the existing return-review e2e test**: the
|
||
async pre-fill could silently overwrite odometer text a visitor had already started
|
||
typing, if the vehicle-detail fetch resolved after they began typing — fixed with an
|
||
`odometerEditedByUser` ref guard.
|
||
- Data quality: added a shared `RuleExplainer` (what's wrong / why it matters, in plain
|
||
language) for all 5 rule types on `DataQualityIssueDetail.tsx`; added a generic
|
||
post-resolution confirmation (audit-trail link, vehicle link, "Ga verder met de demo")
|
||
for the 4 rule types that previously just silently flipped their status badge with no
|
||
explicit confirmation, and extended `VehicleStatusConflictPanel`'s existing confirmation
|
||
with the same links rather than duplicating it. Added a "Demo scenario's only" checkbox
|
||
filter on `DataQuality.tsx` (client-side `public_ref.startsWith("DQ-DEMO-")`, no new
|
||
business logic) so the curated issues are easy to find among the full queue.
|
||
- Evidence: frontend `tsc -b` clean, `npm run build` clean; full Playwright suite
|
||
**51 passed** (47 existing + 4 new `demo-legibility.spec.ts`: return pre-fill + why-
|
||
suspicious explanation + result links, rule explainer visible, demo-scenario filter
|
||
narrows correctly, knowledge suggested question returns grounded evidence with the
|
||
correct provider label). Backend untouched this batch.
|
||
- Deployed to Unraid (commit `ddc3a98`): pushed to Gitea, tarball extracted, `web`
|
||
rebuilt (frontend-only), healthy. Live-verified: ran `demo-legibility.spec.ts` (4) and
|
||
the five-minute demo script directly against `http://192.168.10.150:1236` —
|
||
**5/5 passed**. Reseeded afterward to leave the server demo-ready.
|
||
- Exact next action: plain-language integration-status labels, richer audit narration,
|
||
the full "Over deze demo" page content (currently a first pass from Batch 2), and
|
||
wiring reset into the guide/About/OM menu narrative — task #34.
|
||
|
||
### Batch 5 — integration-status UX, audit UX, About page, reset integrity (complete)
|
||
|
||
- Plain-language integration status: extracted `frontend/src/data/integrationLabels.ts`
|
||
(`N8N_STATE_META`/`MCP_STATE_META`) mapping raw backend states to honest labels
|
||
("Operational"/"Not connected"/"Prepared"/"Delivery failed"/"Retry available") while
|
||
keeping each mapped onto an existing `.status-*` CSS colour class (a few raw values like
|
||
`degraded`/`disabled`/`configured` had no matching CSS rule at all before this — a real,
|
||
pre-existing colour-coding gap). `StatusBadge` gained an optional `label` override prop
|
||
(backward compatible) so the badge's colour class and its displayed text can differ.
|
||
Wired into both `Automation.tsx` and `Dashboard.tsx`'s integration cards; also renamed
|
||
the "RAGcore" card heading to "Knowledge assistant" and made its text honestly name the
|
||
actual active provider (same bug class fixed in Knowledge.tsx in Batch 4).
|
||
- Audit trail: added a "Follow-up" column with a "View related events" action per row that
|
||
filters the same list by `correlation_id` (reuses the backend's existing, already-tested
|
||
`correlation_id` query param — no new business logic), with a "Clear this filter"
|
||
affordance. This is how a visitor sees "what else happened as a result of this action"
|
||
(e.g. a return's linked vehicle-status-changed / workflow-queued events) without a
|
||
bigger grouped-timeline rebuild.
|
||
- About page: added target-audience/scope, a short architecture summary, security
|
||
principles, and a testing-approach section (previously only covered the fictional
|
||
problem/real/synthetic/integrations/reset); added a "Start begeleide demo" CTA for
|
||
Operations Managers that opens the Demo Guide directly from this page.
|
||
- Reset integrity: added `scenario_integrity_report()` (`backend/app/services/
|
||
demo_manifest.py`), reusing the exact same scenario-readiness derivation the manifest
|
||
and scenario overview already use (so it can't drift), and wired it into `POST
|
||
/api/v1/demo/reset` — both the response body and the `demo_reset` audit event's
|
||
metadata now carry `scenario_integrity: {all_ready, not_ready}`. This is the
|
||
server-side post-reset integrity check the brief asks for; visible today via the audit
|
||
event's raw-detail view, satisfying the requirement without adding a UI banner to a
|
||
flow that immediately logs the user out and redirects to `/login`.
|
||
- **Fixed a second real regression this batch, caught by the existing return-review
|
||
e2e test**: restructured the odometer pre-fill so `BookingDetail.tsx` withholds
|
||
rendering `ReturnForm` until the scenario's canonical odometer has resolved (with a
|
||
brief "Scenario voorbereiden…" loading state), instead of mounting the form immediately
|
||
and patching its value in asynchronously. The previous approach raced visibly with
|
||
Playwright's `fill()` (and would have raced with a real visitor typing quickly),
|
||
producing a corrupted concatenated value in one observed failure. This also let the
|
||
now-unnecessary `odometerEditedByUser` ref guard be removed — simpler and more robust
|
||
than the effect-based patch it replaced.
|
||
- Evidence: `pytest` **127 passed**, `ruff check .` clean, `mypy app` clean (48 files);
|
||
frontend `tsc -b` clean, `npm run build` clean; full Playwright suite **51 passed**,
|
||
confirmed stable across three consecutive full runs (given how many timing races this
|
||
batch and the previous one surfaced, stability was verified deliberately rather than
|
||
assumed from a single green run).
|
||
- Deployed to Unraid (commit `5fa4fe0`): pushed to Gitea, tarball extracted, both `api`
|
||
and `web` rebuilt, healthy, migrations unchanged at `e7b08389f47f (head)`, reseeded.
|
||
Live-verified: ran `demo-legibility.spec.ts` (4), the five-minute demo script, and the
|
||
full `interactive-elements.spec.ts` suite (26) directly against
|
||
`http://192.168.10.150:1236` — **31/31 passed**. Reseeded afterward to leave the
|
||
server demo-ready.
|
||
- Exact next action: full guided-demo Playwright test + remaining targeted demo tests per
|
||
section 19 (mobile guide, keyboard nav, all scenario flows, About page, accessibility/
|
||
reduced-motion/console/network checks) — task #35.
|
||
|
||
### Batch 6 — full guided-demo test + targeted demo tests (complete)
|
||
|
||
- **Found and fixed a real, fairly serious desktop layout bug** while writing the full
|
||
guided-demo test: the Demo Guide's fixed right-side panel (400px wide) overlapped the
|
||
main content area at normal desktop widths with no reflow, so its own step-list buttons
|
||
intercepted pointer events meant for the page underneath (concretely: the return form's
|
||
"Review return" button was unclickable while the guide was open, at exactly the
|
||
viewport size Playwright's default test browser uses — this would have hit real
|
||
visitors on ordinary laptop screens too). Fixed by adding a `guide-open` class to
|
||
`.app-workspace` that reserves `padding-right: min(400px, 92vw)` while the guide is open
|
||
(≥701px only; the ≤700px bottom-sheet layout is unaffected), so content reflows aside
|
||
instead of sitting underneath the panel.
|
||
- Added `frontend/e2e/guided-demo-full.spec.ts`: one comprehensive test walking a fresh
|
||
Operations Manager session through all 8 Demo Guide steps in order, performing the
|
||
**real** action at each step (not just verifying copy) — processes the actual
|
||
odometer-anomaly return, resolves the resulting data-quality issue, merges the
|
||
duplicate customer, asks a suggested knowledge question, checks automation + audit,
|
||
reviews the About page — using the guide's own progression controls
|
||
("Volgende"/"Ga naar deze stap"/"Ga verder met de demo") throughout, then resets the
|
||
demo data again at the end to restore the environment per the brief's requirement.
|
||
- Added `frontend/e2e/demo-accessibility.spec.ts` (4 tests): the guide renders as a
|
||
correctly-anchored bottom sheet on a 390px mobile viewport with no horizontal overflow;
|
||
the guide never covers the return form's action buttons on desktop (regression test for
|
||
the bug above); the demo badge and guide trigger are keyboard-focusable and operable
|
||
(Enter to open, explicit close controls); key demo pages (dashboard, scenarios, about,
|
||
guide open) load with no unexpected console errors (the one expected benign 401 from
|
||
the app's own session-probe on first load is explicitly allow-listed, not silenced
|
||
blindly).
|
||
- Evidence: full Playwright suite **56 passed** (51 existing + 1 guided-demo-full + 4
|
||
demo-accessibility), confirmed stable across two consecutive full runs. Backend
|
||
untouched this batch (last gate: 127 passed/ruff/mypy clean, Batch 5).
|
||
- Deployed to Unraid (commit `07d5605`): pushed to Gitea, tarball extracted, `web`
|
||
rebuilt (frontend-only), healthy, reseeded. Live-verified: ran
|
||
`guided-demo-full.spec.ts` and `demo-accessibility.spec.ts` directly against
|
||
`http://192.168.10.150:1236` — **5/5 passed**, confirming the desktop-overlay layout
|
||
fix holds on the real deployment too. Reseeded afterward to leave the server
|
||
demo-ready.
|
||
- Exact next action: clean-checkout demo drill, final documentation set (demo-concept/
|
||
demo-scenarios/demo-data/demo-guide/demo-runbook, README, .env.example), final Unraid
|
||
deploy + live evidence with screenshots, `artifacts/demo-release/final-summary.md` —
|
||
task #36 (final).
|
||
|
||
### Batch 7 (final) — clean-checkout drill, docs, final Unraid evidence (complete)
|
||
|
||
- **Clean-checkout drill**: fresh `git clone` into an isolated scratch directory,
|
||
isolated Compose project (`mobilityops-cleandrill`) + remapped ports via
|
||
`compose.override.yaml`, `up --build -d` from empty volumes. Migrations ran
|
||
automatically to `e7b08389f47f (head)`; seeded; full backend gate **127 passed**,
|
||
ruff/mypy clean; `npm ci` clean (same pre-existing advisories as before, unchanged);
|
||
`tsc -b`/`vite build` clean; full Playwright suite **56 passed** against the isolated
|
||
stack; reseeded and confirmed all 5 scenarios `ready: true` via the manifest; torn down
|
||
(`down -v` on the isolated project only — the working dev stack was untouched
|
||
throughout).
|
||
- Added the full demo-release documentation set: `docs/demo-release/demo-concept.md`,
|
||
`demo-scenarios.md`, `demo-data.md`, `demo-guide.md`, `demo-runbook.md`; updated
|
||
`README.md` (current test counts, links to the new docs, a "Demo" section) and
|
||
`docs/17-runbook.md` (cross-reference to the demo-specific runbook).
|
||
- Added `frontend/e2e/_capture-demo-screenshots.spec.ts` (tooling, excluded from the
|
||
regular suite) and captured 17 evidence screenshots live against
|
||
`http://192.168.10.150:1236` into `artifacts/demo-release/screenshots/`.
|
||
- Final live acceptance: full Playwright suite re-run against the live server —
|
||
**56 passed**; `docker compose ps` on the server shows `api`/`db`/`web` all healthy;
|
||
`docker logs` for `api`/`web` show no errors; reseeded to leave the server
|
||
demo-ready after evidence capture.
|
||
- Wrote `artifacts/demo-release/final-summary.md` with the full required evidence
|
||
(branches/commits, org/roles/guide/scenarios, seed/date-anchor/reset strategy, real vs.
|
||
synthetic vs. not-connected, all test results, clean-checkout result, deployment/
|
||
health/console/log results, responsive/accessibility results including the two real
|
||
layout bugs found and fixed this work (mobile topbar overflow in Batch 3, desktop
|
||
guide-panel overlap in Batch 6), known limitations, 5-/10-minute demo flows, redeploy/
|
||
rollback commands, and the screenshot list).
|
||
- Demo-productization work on this branch is complete. Every task (#29–#36) is done;
|
||
every batch was tested locally, deployed to Unraid, and re-verified live before moving
|
||
to the next. See `artifacts/demo-release/final-summary.md` for the definitive
|
||
acceptance evidence.
|
||
|
||
## Final product polish: Fleet Ops rebrand, trilingual i18n, adaptive guide (2026-08-03) — MERGED TO MASTER
|
||
|
||
- Rebranded the product to **Fleet Ops** across the frontend, backend defaults and the
|
||
knowledge base; made `nl-BE` (default)/`en-GB`/`fr-BE` full first-class languages via
|
||
i18next (eager-bundled resources, persisted language switcher in topbar + mobile
|
||
drawer, `Intl` date/number formatting, a coverage test that fails the build on any
|
||
missing/empty translation key across all 14 namespaces).
|
||
- Backend dynamic content (demo scenarios, blocked-reason text, integration status)
|
||
converted from fixed English/Dutch prose to stable message codes + params so the
|
||
frontend localizes it (`DemoScenarioOut`/`DemoIntegrationSummaryOut` schema changes).
|
||
The demo knowledge base gained a fully translated NL/EN/FR procedure corpus (11
|
||
documents each, including a new "vehicle availability" procedure) with per-language
|
||
retrieval and localized evidence-state messages.
|
||
- Demo Guide became breakpoint-adaptive: docked rail (≥1440px), a floating panel that
|
||
auto-collapses to a persistent closable progress chip (701–1439px), and a
|
||
collapsed/half/full bottom sheet (≤700px) — with scroll+focus+highlight on "go to this
|
||
step", Escape handling, and `prefers-reduced-motion` support.
|
||
- Data Quality Workbench got accessible choice-card decisions with a clear
|
||
primary/secondary/tertiary action hierarchy; Automation ledger groups repeated
|
||
successes with meaningful short refs; Audit trail groups events by correlation id with
|
||
human action labels and readable before/after diffs; Attention Queue/Today's
|
||
movements/Vehicles/Bookings/Data Quality rows are fully clickable (stretched-link
|
||
pattern, independent secondary links, keyboard + mobile support).
|
||
- Two real bugs found and fixed along the way: a mobile topbar overflow at 421–440px
|
||
caused by the new language switcher (moved the switcher into the mobile drawer at
|
||
≤960px and widened the compact-topbar breakpoint to 440px), and two dangling
|
||
`aria-labelledby` references (`SectionHeading` never set the referenced `id`).
|
||
- Full test suite: 131 backend tests, Ruff, mypy, TypeScript build, and 92 Playwright
|
||
tests (new: `i18n-coverage`, `clickable-rows`, `responsive-i18n` covering all 7
|
||
brief-specified breakpoints × 3 languages, plus 3 new adaptive-guide tier tests) — all
|
||
green. All pre-existing Playwright specs updated for the new nl-BE default (either
|
||
translated assertions or an explicit English-locale override where the spec was
|
||
originally authored against English copy).
|
||
- Clean-checkout drill performed in a fully isolated Docker Compose project (separate
|
||
ports/volumes, no shared n8n) from a fresh local clone at the feature-branch head —
|
||
131 backend tests, lint, build and all 92 Playwright tests green from empty volumes;
|
||
live EN/FR knowledge-assistant spot check; `scenario_integrity.all_ready: true` on
|
||
reset; isolated stack torn down afterward, original dev environment untouched.
|
||
- Deployed to Unraid twice: once for the feature branch (commit `845db14`) for
|
||
pre-merge live validation, once for the merged `master` (commit `18a765d`) for the
|
||
final release — both times via the established `git archive` → `scp` →
|
||
extract → `.deploy/source-revision` → rebuild `api`/`web` method, with migrations,
|
||
reseed, full backend+Playwright gates, console/network inspection and demo reset
|
||
re-verified live each time.
|
||
- Master baseline was confirmed unchanged (`e0c7ed6`, matching the previously recorded
|
||
baseline) before merging; `git merge-tree` dry run showed zero conflicts. Merged via
|
||
`git merge --no-ff` (commit `18a765d`), all gates re-run post-merge, pushed to Gitea,
|
||
redeployed. Feature branch was not deleted.
|
||
- Full evidence: `artifacts/fleet-ops-release/final-summary.md` (commits, branding,
|
||
locales, translation/knowledge-base/guide/data-quality/automation/audit evidence, all
|
||
test results, clean-checkout result, deployment evidence for both the feature branch
|
||
and master, responsive/accessibility results, known limitations, rollback procedure)
|
||
plus 10 screenshots in `artifacts/fleet-ops-release/screenshots/`.
|
||
|
||
## Fleet Ops correction: safe status-recommendation flow, MO-016, message codes (2026-08-03) — MERGED TO MASTER
|
||
|
||
Branch `fix/fleet-ops-i18n-status-flow`, created from master's post-release head
|
||
(`18344bc`). Audit and rationale in `docs/fleet-ops-correction/` (gap audit, i18n
|
||
inventory, vehicle-status decision table). Merged to master via `de0bdea` ("merge:
|
||
complete Fleet Ops localization and status resolution"), with final evidence commit
|
||
`f780557` ("docs(release): final Fleet Ops correction evidence and screenshots") —
|
||
`f780557` is `origin/master`'s current head as of the start of the correction round
|
||
below.
|
||
|
||
- **Status-recommendation flow redesigned** per the brief: the old single opaque
|
||
"calculate and apply recommended status" action is replaced by a single shared, pure
|
||
evaluator (`backend/app/services/vehicle_status.py::evaluate_vehicle_status`) used
|
||
identically by the scanner, a new non-mutating preview endpoint
|
||
(`POST .../status-recommendation`), and a transactional apply endpoint
|
||
(`POST .../apply-recommended-status`) that locks the row, recomputes facts, rejects a
|
||
stale `recommendation_token` (optimistic concurrency), refuses unsafe/manual-review
|
||
recommendations, and re-validates post-write before resolving the issue. Frontend
|
||
`DataQualityIssueDetail.tsx` shows "Review recommendation" → a decision panel
|
||
(current/recommended status, why, evidence, consequences, localized in all 3
|
||
languages) → an exact "Change status to <status>" confirm action → result, with a
|
||
distinct "Manual review required" state offering no generic apply button.
|
||
- Fixed the real unsafe shortcut this evaluator exists to eliminate: "maintenance +
|
||
active booking" no longer auto-recommends "rented" (current status is itself now a
|
||
blocking fact), and "maintenance with nothing else wrong" no longer auto-clears to
|
||
"available" (no fact proves maintenance is actually finished — that release stays a
|
||
manual decision).
|
||
- **MO-016 order independence**: order independence does not mean "same final status
|
||
regardless of order" — resolving the booking overlap first genuinely removes the
|
||
conflict, correctly leaving nothing to apply. What must (and does) hold either way:
|
||
the recommendation always reflects real current facts, and nothing unsafe is ever
|
||
applied (never "rented"). Verified by both a backend test
|
||
(`test_mo_016_status_conflict_recommendation_is_order_independent`, explicitly scoped
|
||
to MO-016/DQ-DEMO-STATUS after finding the original version wasn't) and a browser-level
|
||
Playwright test in both orders.
|
||
- **"Fleet Ops" is a non-localizable brand constant** (`frontend/src/product.ts`,
|
||
backend `PRODUCT_NAME`), wired via `{{productName}}` interpolation everywhere the
|
||
brand appeared in locale prose; a permanent test fails the build if any locale file
|
||
ever defines the brand name or an `appName` key again.
|
||
- **Dynamic backend prose converted to message codes + params**: return status reasons,
|
||
audit field/actor-type labels, automation `last_error` (new `last_error_code` column,
|
||
migration `799d8800e241`), and search results (sections/vehicles/bookings/issues) all
|
||
now carry stable codes the frontend localizes; raw technical text is demoted to a
|
||
"Technical details" disclosure everywhere.
|
||
- **Knowledge-base fixes**: the demo provider's tokenizer silently dropped accented
|
||
characters (`[a-z0-9]+` split "véhicule" into "v"+"hicule"), breaking French
|
||
retrieval broadly; fixed to include the Latin-1 accented range. Also reweighted
|
||
section scoring so a body match (real substance) outranks a heading/title match (a
|
||
shallow structural hint) — the old weighting misranked the damage procedure behind
|
||
an unrelated document for the brief's exact validation question in all 3 languages.
|
||
Removed leftover "MobilityOps"/"PoC" mentions from 9 procedure documents (knowledge
|
||
prose is visible content, missed by the earlier rebrand).
|
||
- New `frontend/e2e/fleet-ops-correction.spec.ts` (14 tests) covers branding in 3
|
||
languages, language persistence, the full status-recommendation flow (non-mutating
|
||
preview, exact confirm text, manual review, stale-token rejection), MO-016 order
|
||
independence, trilingual knowledge grounding, and localized audit/automation. Writing
|
||
it surfaced and fixed two real bugs: the frontend conflated "no conflict" with
|
||
"manual review required" (both carry `safe_to_apply: false`), and the original
|
||
MO-016 backend test never actually targeted MO-016's own issue.
|
||
- Added a keyboard/reduced-motion/no-color-only-status accessibility test for the new
|
||
status-decision panel; added `aria-live="polite"` to the panel so the applied
|
||
confirmation is announced.
|
||
- `contracts/openapi.yaml` and `README.md` updated: title is "Fleet Ops", the new
|
||
status-recommendation endpoint documented, apply-recommended-status's request body
|
||
and error codes documented, search endpoint's code+params shape documented, README
|
||
states the Fleet Ops/MobilityOps naming split explicitly and refreshes stale test
|
||
counts (151 backend, 108 Playwright).
|
||
- Gates green: 151 backend tests, Ruff, mypy, Alembic upgrade/downgrade verified,
|
||
frontend `tsc`/build, full 113-test Playwright suite (rebuilt `api`+`web` containers
|
||
each time before testing).
|
||
- Section 11D/E/F of the i18n test-strengthening brief done: a hardcoded-JSX-text
|
||
static check (`i18n-coverage.spec.ts`; had to anchor on backreferenced closing-tag
|
||
names — a naive `>text<` scan misread TypeScript generics like
|
||
`useState<string | null>` as JSX spanning to the next unrelated `>`; verified against
|
||
both false positives and a deliberately-injected-then-reverted false negative), and a
|
||
3-language route matrix (`fleet-ops-correction.spec.ts`) covering every main route:
|
||
no console errors, correct `html[lang]`, real page headings.
|
||
- **Clean-checkout drill (2026-08-03) — PASS.** Fresh `git clone --branch
|
||
fix/fleet-ops-i18n-status-flow` of only committed files into an isolated directory,
|
||
separate Compose project name and host ports (8129/1229/5679) so the working dev
|
||
stack was never touched. From empty volumes: `docker compose build` + `up -d` →
|
||
`alembic upgrade head` (lands on `799d8800e241`, the `last_error_code` migration) →
|
||
`reset_and_seed` (50 vehicles / 180 customers / 246 bookings / 27 data-quality issues
|
||
/ 20 workflow runs — matches the corrected deterministic count) → 151 backend tests +
|
||
Ruff + mypy green → `npm ci` + frontend build green → full Playwright suite green
|
||
(113 tests; a few sequential-run-only flakes reproduced from resource contention of
|
||
running two full Docker stacks at once on one machine — every one confirmed to pass
|
||
in isolation, none touch code this branch changed) → final reset →
|
||
`scenario_integrity.all_ready: true`. Isolated stack, containers, volumes and images
|
||
torn down afterward; original dev environment confirmed untouched and reset to
|
||
baseline.
|
||
- **Unraid deployment (2026-08-03/04) — PASS.** Pushed `fix/fleet-ops-i18n-status-flow`
|
||
to origin, deployed via `git archive` → `scp` → extract into
|
||
`/mnt/user/appdata/mobilityops` (preserving `.env`) → `.deploy/source-revision` →
|
||
rebuild `api`+`web` → `alembic upgrade head` → reset/reseed, at
|
||
`http://192.168.10.150:1236`. **Live validation directly caught a real bug**: every
|
||
data-quality issue's top-of-page evidence summary was unconditionally showing raw
|
||
English (e.g. "vehicle marked available while reserved bookings conflict") in all
|
||
three languages, because the frontend never finished the evidence.signals
|
||
localization the backend had already been emitting. Fixed (commit `2e4fb43`):
|
||
DataQualityIssueDetail.tsx now renders `evidence.signals` through the operator's
|
||
locale as the primary text, raw text moved to "Technical details" only, the 4
|
||
DQ-DEMO-* seed rows got real computed signals (the duplicate-customer similarity
|
||
score is the actual SequenceMatcher ratio on the seeded names), and a regression
|
||
test locks this in. Redeployed with the fix; live-verified via `read_page` that
|
||
DQ-DEMO-STATUS now shows "Dit voertuig heeft twee overlappende reserveringen..."
|
||
instead of the raw English sentence. Full 116-test Playwright suite green against
|
||
the live server (`MOBILITYOPS_PUBLIC_URL=http://192.168.10.150:1236`), no console
|
||
errors, no errors in `api`/`web` container logs, both containers healthy, final
|
||
reset done, `scenario_integrity.all_ready: true`.
|
||
- Final evidence: `artifacts/fleet-ops-correction/final-summary.md`. Merged to master
|
||
via `de0bdea`, followed by evidence commit `f780557` on master. See the "Fleet Ops
|
||
final localization" entry below for the next (small correction) round on top of this.
|
||
|
||
## Fleet Ops final localization: remaining NL/FR gaps, API-error localization, greeting (2026-08-04) — MERGED TO MASTER
|
||
|
||
Merged to master via `5f0eaa5`; final evidence commit `c0995b7` added
|
||
`artifacts/fleet-ops-final-localization/final-summary.md`. Master head at merge:
|
||
`c0995b762e1cbf37172a08e03645baa6b66aa8d5`. Details below are the in-progress working log
|
||
kept for reference.
|
||
|
||
Branch `fix/fleet-ops-final-i18n-ux`, created from master's post-correction head
|
||
(`f780557`) — the brief asked for `fix/fleet-ops-final-localization`, but the
|
||
already-checked-out branch name is used instead since it was verified freshly and
|
||
cleanly branched from current `origin/master` with a clean working tree; see
|
||
`docs/fleet-ops-final-localization/audit.md` for the naming note. Scope: a small,
|
||
targeted correction round only — explicitly not touching status-flow business logic,
|
||
the status evaluator, Data Quality resolution rules, return rules, RAGcore/MCP Hub, or
|
||
product scope.
|
||
|
||
- **Audit-driven gap sweep**: `docs/fleet-ops-final-localization/audit.md` documents
|
||
every remaining untranslated/incorrect string, raw-backend-error call site,
|
||
over-permissive allowlist entry, the static-greeting bug, and doc staleness found by
|
||
a dedicated Explore pass before any file was touched.
|
||
- **Remaining NL/FR translation gaps fixed**: role names actually translated (not just
|
||
labelled as translated) — `auth.json`/`demo.json` role keys, `audit.title` →
|
||
"Auditgeschiedenis"/"Piste d'audit", `columns.actor` → "Uitvoerder", `list.statusOpen`
|
||
→ "Openstaand", `ledger.filterRecent` → "Recentste", `scenarios.startScenario` →
|
||
"Scenario starten". Also found and fixed (via the new embedded-substring test below)
|
||
8 previously-missed mid-sentence "Audit trail" leaks across `demo.json`,
|
||
`quality.json`, `returns.json` that the old whole-string-identity test structurally
|
||
could not catch.
|
||
- **Central API-error localization**: new `frontend/src/api/errorMessages.ts`
|
||
(`describeApiError`) replaces the `err instanceof ApiError ? err.message : ...`
|
||
anti-pattern (which showed raw English for the common case) at all 13 call sites
|
||
across 7 files. Raw backend text is now only ever shown under a "Technical
|
||
details"/"Détails techniques" disclosure (new `ApiErrorNotice` component in
|
||
`PageChrome.tsx`); the primary message is always a localized title + explanation +
|
||
optional next step, keyed on the 32 known `AppError` codes, then known HTTP statuses
|
||
(401/403/404/409/422/500), then a fully generic fallback. `ApiError` was split out of
|
||
`client.ts` into a standalone `api/apiError.ts` (no `import.meta.env` dependency) so
|
||
`errorMessages.ts` is independently testable outside a Vite/browser context.
|
||
- **i18n allowlist tightened**: removed 7 now-stale `IDENTICAL_VALUE_ALLOWLIST` entries
|
||
in `i18n-coverage.spec.ts` (`audit.title`, `auth.roleOperationsManager`,
|
||
`auth.roleRentalEmployee`, `demo.scenarios.startScenario`,
|
||
`demo.scenarios.roles.operations_manager`, `demo.scenarios.roles.rental_employee`,
|
||
`navigation.items.audit`) now that they're genuinely translated. Added 2 new tests:
|
||
one closing the embedded-English/Dutch-substring blind spot (mid-sentence phrase
|
||
leaks the whole-string check misses), one asserting no locale file contains
|
||
"MobilityOps" or the word "PoC".
|
||
- **`describeApiError` test coverage**: new `frontend/e2e/error-messages.spec.ts` (10
|
||
tests) — every known code/HTTP status has non-empty copy in all 3 locales, a known
|
||
code never surfaces raw backend text as the primary message (only via `.technical`),
|
||
unknown-code and unknown-status fallback chains behave correctly, and a drift guard
|
||
that greps the actual backend `AppError("CODE", ...)` call sites and fails if
|
||
`KNOWN_CODES` and the backend's real codes ever diverge (currently exactly in sync,
|
||
32 codes).
|
||
- **Time-dependent Europe/Brussels dashboard greeting**: new
|
||
`frontend/src/i18n/greeting.ts` (`getGreetingPeriod`, DST-safe via
|
||
`Intl.DateTimeFormat({ timeZone: "Europe/Brussels", hourCycle: "h23" })`, clock
|
||
injectable) + `useGreetingPeriod.ts` hook (30s poll for period rollover while the app
|
||
stays open, no reload). Replaces the previously-always-"Goedemorgen" static
|
||
`dashboard.json` title with 4 periods × 3 languages for both the greeting word and a
|
||
varying accompanying sentence (never "Goedenacht"). Tests: `greeting.spec.ts` (pure
|
||
boundary/DST unit tests) + `greeting-live.spec.ts` (6 real-browser tests via
|
||
Playwright's `page.clock` — all 8 required boundary times in all 3 languages, live
|
||
rollover without reload, language-switch behaviour, the "never Goedenacht" guard).
|
||
- **Found and fixed one real CSS regression along the way**: correctly translating
|
||
`roleOperationsManager` to the single unbreakable Dutch compound word
|
||
"Operationsmanager" (vs. the old two-word "Operations Manager", which could wrap)
|
||
pushed the topbar's `.operator` block past 1024px width, caught by the existing
|
||
`responsive-i18n.spec.ts` overflow test. Fixed with `overflow-wrap: anywhere` on
|
||
`.operator strong`/`small` and `min-width: 0` on their flex-item wrapper, not by
|
||
reverting the correct translation.
|
||
- Gates green so far: backend `pytest` 151 passed, `ruff check .` clean, `mypy app`
|
||
clean (49 files, unchanged — no backend Python touched this round); frontend `tsc`
|
||
clean, production build clean, full local Playwright suite **138 passed** (rebuilt
|
||
and restarted the local `web` container from source before this run).
|
||
- **Not yet done**: clean-checkout drill, commit/push, Unraid deployment of this fix
|
||
branch with live 3-language validation, the master merge (with the mandatory
|
||
`git fetch origin` / unexpected-change check first), and
|
||
`artifacts/fleet-ops-final-localization/final-summary.md`. Do not claim PASS on this
|
||
correction round until all of those are done and `git rev-parse HEAD` exactly matches
|
||
`/mnt/user/appdata/mobilityops/.deploy/source-revision`.
|
||
- Commits so far on this branch: `6deb955` (status flow + brand constant + message
|
||
codes), `e6539d1` (knowledge fixes), `ac4b163` (Playwright spec updates for the new
|
||
flow), `1fdd2b3` (new E2E coverage + 2 bug fixes), `1e40775` (accessibility test),
|
||
`a7ac5ed` (docs), `7851e80` (11D/11F i18n tests), `cda2c32` (clean-checkout
|
||
evidence), `2e4fb43` (evidence-summary localization fix, found live on Unraid).
|
||
Deployed commit: `2e4fb43f093bfbdb04c4f74eed1e6c6d9a03c069`.
|
||
|
||
## Live n8n + RAGcore integration (2026-08-04) — IN PROGRESS on feat/live-n8n-ragcore-integration
|
||
|
||
Branch `feat/live-n8n-ragcore-integration`, from master `c0995b7`. Full brief: treat n8n
|
||
(`https://n8n.itworx.tech`, existing shared instance) as a third integration layer
|
||
alongside RAGcore and MCP Hub, owning process orchestration only — Fleet Ops keeps all
|
||
business rules, authorization, transactions, audit and idempotency. Four canonical
|
||
workflows required: (1) Vehicle Return Orchestration, (2) Scheduled Data Quality Scan —
|
||
both pre-existing and now hardened; (3) RAGcore Procedure Sync, (4) Workflow Error
|
||
Handler — both net-new, not yet built.
|
||
|
||
- **Current-state audit**: `docs/live-ai-integration/n8n-current-state.md` documents the
|
||
live instance (reachable, production webhook base
|
||
`http://192.168.10.150:5678/webhook/mobilityops-return`), both existing workflows'
|
||
full node structure, and the findings that drove the security fixes below (webhook
|
||
Authentication was `None`; both HTTP nodes had `X-Service-Token` hardcoded as a literal
|
||
header value instead of a credential).
|
||
- **Security fixes applied and live-validated** (commits `b79d485`, `59cb4c0`): webhook
|
||
trigger now requires Header Auth (credential `Fleet Ops Webhook Trigger Token`, a new
|
||
token generated this round — value stored in `.env`/Unraid `.env` only, never
|
||
printed); the outbound callback HTTP node now uses a `Fleet Ops Service Token` Header
|
||
Auth credential instead of a literal header value (existing secret copied
|
||
clipboard-to-clipboard, never typed/echoed). Backend: `X-Fleet-Ops-Trigger-Token`
|
||
header added to the outbox dispatcher's POST (`backend/app/services/dispatcher.py`),
|
||
plus a new `MOBILITYOPS_WEBHOOK_TRIGGER_TOKEN` setting/env var. Also hardened
|
||
`_deliver_one` to treat a 2xx response with a non-JSON-object body as a retryable
|
||
failure (`malformedResponse`) instead of an unhandled exception — a real failure mode
|
||
hit live when a workflow errors before its "Respond to Webhook" node runs; regression
|
||
test `test_deliver_one_treats_empty_2xx_body_as_failure` added. Live-validated: curl
|
||
probe without the header → `403`; with the header → pass-through; one real end-to-end
|
||
vehicle return produced one correct execution visible in both n8n and Fleet Ops
|
||
Audit/Automation. Both workflows explicitly `Publish`ed after the fixes (the editor
|
||
does not go live on save alone) and both canonical-renamed ("Fleet Ops — Vehicle
|
||
Return Orchestration", "Fleet Ops — Scheduled Data Quality Scan").
|
||
- **RAGcore real contract discovered** (not the speculative one the adapter was built
|
||
against): OpenAPI at `/openapi.json`, health at `/health/live`/`/health/ready` (not
|
||
`/health`), ingestion via `POST /v1/uploads`, answers via `POST /v1/answers` with
|
||
`requested_space_ids`, control-plane endpoints require an `Idempotency-Key` header.
|
||
Bootstrapped a `fleet-ops` application + knowledge space + grant on the real server at
|
||
`http://192.168.10.150:1237`. **Blocked**: credential issuance for that application
|
||
failed identically via both the raw API and the admin UI ("authoritative
|
||
service-account state rejected issuance") — an apparent privilege boundary beyond the
|
||
interactive admin session. User chose to issue the credential themselves via another
|
||
mechanism and hand over the token; not yet received. `RAGcoreKnowledgeProvider`
|
||
(`backend/app/services/knowledge/ragcore.py`) still targets the old speculative
|
||
endpoints and needs fixing once that token arrives — approved, not started.
|
||
- **Repository source of truth started** (task in progress): `n8n/workflows/` now holds
|
||
cleaned definitions for workflows 1-2 — `fleet-ops-vehicle-return.json` (sha256
|
||
`e13a3087269fc97019a7adf6c6a6a4ee4bd354c2dd7167d4966d4753a48e970e`),
|
||
`fleet-ops-data-quality-scan.json` (sha256
|
||
`cc30b28b07dad9f9908a6ea0c564ec4c2f362a3ed71b7e97a7b6894408bb7e2e`) — both credential
|
||
auth referenced by name only, no secret values. Reconstructed from direct verified
|
||
inspection of every live node, **not** a literal n8n export/download: the UI's "..."
|
||
menu has no Download option in this n8n version, and clipboard-based
|
||
copy/`navigator.clipboard.readText()` extraction timed out twice. Flagged as a known
|
||
limitation for the final evidence doc. `n8n/workflows/MANIFEST.md` records canonical
|
||
name/purpose/trigger/contract/credentials/live ID/active-status/checksum for all 4
|
||
workflows (3-4 marked not-yet-built). `n8n/workflows/check_drift.py` compares a repo
|
||
definition against the live workflow via n8n's Public API (`X-N8N-API-KEY`, read-only,
|
||
never auto-overwrites). The old root-level `n8n/mobilityops-return-processing.json`
|
||
and `n8n/mobilityops-scheduled-quality-scan.json` (pre-integration starters, still
|
||
carrying the literal-token pattern) are removed; `deploy/unraid/setup-existing-n8n.sh`,
|
||
`setup-scheduled-scan.sh`, `Makefile` (`n8n-setup`, `n8n-setup-scan`) and
|
||
`docs/17-runbook.md` updated to import from `n8n/workflows/` and to document the
|
||
now-required manual credential-creation step (credentials are never scripted or
|
||
committed).
|
||
- **Explicitly deferred/forbidden this phase** (per brief): daily AI ops brief, email,
|
||
Slack, automatic vehicle-status changes, customer communication, billing, general
|
||
monitoring, autonomous MCP actions. An automatic demo-reset workflow may only be
|
||
prepared, not activated, once Fleet Ops goes public.
|
||
- **Workflow 4 (Workflow Error Handler) built and live-validated** (commit pending):
|
||
new backend endpoint `POST /api/v1/integrations/n8n/workflow-error`
|
||
(`backend/app/api/routers/integrations.py`, service-token auth, Pydantic
|
||
`WorkflowErrorReportIn`/`WorkflowErrorReportResult` in `backend/app/schemas.py`),
|
||
idempotent on `execution_id` via the same audit-precheck pattern as
|
||
`/return-callback`; new test coverage in `backend/tests/test_integrations.py` (all
|
||
green, 152 tests total, ruff/mypy clean). **This endpoint had to be deployed to the
|
||
live Unraid server** (`git archive` → `scp` → extract preserving `.env` → `docker
|
||
compose up --build -d api`, no migration needed) before the live n8n test could reach
|
||
it — the auto-mode classifier correctly blocked the first `scp` attempt as a
|
||
production-infra action; user approved, then it was deployed and verified
|
||
(`/health` OK, new endpoint returns 422 on empty body instead of 404).
|
||
Built "Fleet Ops — Workflow Error Handler" (live ID `Xppn2rAEqUuyiCJF`) in n8n:
|
||
Error Trigger → Code node (derives safe error_category/summary/etc. from n8n's error
|
||
payload) → HTTP node (POST to the new endpoint, Header Auth via the existing "Fleet
|
||
Ops Service Token" credential). Hit and fixed two real bugs during live testing: (1)
|
||
Code node's default "Run Once for All Items" mode doesn't bind `$json` to the current
|
||
item — switched to "Run Once for Each Item" and `return {json:...}` instead of
|
||
`return [{json:...}]`; (2) every HTTP-body field expression ended up with a stray
|
||
trailing space (from the code-editor's bracket-autoclose leaving one extra character
|
||
after the `End`+`Backspace×2` fix), which broke the `failed_at` datetime parse and the
|
||
`error_category` literal match — found via the raw request dump in n8n's error
|
||
panel, fixed with one more `Backspace` per field. Live-validated: mock Error Trigger
|
||
data → real `200 {"status":"registered"}` from Fleet Ops; re-run → `"already_registered"`
|
||
(idempotency confirmed); wired as the Error Workflow on workflows 1 and 2 (via each
|
||
workflow's Settings modal); confirmed the Error Handler itself has `Error Workflow: -
|
||
No Workflow -` (no recursive loop). With user approval, also ran a genuine induced
|
||
failure on workflow 2 (temporarily pointed its HTTP node at a nonexistent path,
|
||
published, ran it, confirmed it failed as expected, immediately reverted and
|
||
republished, confirmed healthy again) — this proved the target workflow's own error
|
||
path works, but n8n did not auto-invoke the Error Handler for that *manual* editor
|
||
test run (n8n's Error Workflow trigger only fires for unattended/production
|
||
executions), so a fully automatic schedule/webhook-triggered cascade into the handler
|
||
was not observed live this round — noted as a known limitation.
|
||
Exported the verified definition to `n8n/workflows/fleet-ops-error-handler.json` (same
|
||
manual-reconstruction caveat as workflows 1-2: no literal export/download available),
|
||
updated `n8n/workflows/MANIFEST.md` (all 4 workflows, workflow 3 still not-built) and
|
||
`check_drift.py`'s known-workflows list.
|
||
- **Integration status page enriched with real per-workflow evidence** (commit
|
||
`4049c0c`): `N8nIntegrationStatus` now returns `workflows: N8nWorkflowEvidence[]`
|
||
(the 4 canonical workflows, each with real evidence — latest successful outbox
|
||
delivery for the return workflow, latest *service*-triggered `data_quality_scan_run`
|
||
audit event for the scan workflow so a manual UI-triggered scan doesn't fake n8n
|
||
evidence, latest `n8n_workflow_failure_registered` for the error handler, always
|
||
`built: false` / no evidence for the not-yet-built RAGcore sync), plus
|
||
`expected_workflow_count`/`known_workflow_count` and an `error_handler` summary
|
||
(total registered, latest failure + which workflow). New tests in
|
||
`backend/tests/test_integration_status.py` (all green, 159 backend tests total,
|
||
ruff/mypy clean). Frontend: `Automation.tsx` renders this as a localized workflow
|
||
table (EN/NL/FR, new `integrations:workflows.*` keys, technical workflow names under
|
||
a "Technical details" disclosure per the existing progressive-disclosure pattern).
|
||
Verified live in the browser both locally (Dutch locale, disclosure expand/collapse
|
||
confirmed) and **on the deployed Unraid server after this round's deploy**: correctly
|
||
shows "3 van 4 canonieke n8n-workflows hebben actuele evidentie van werking" with real
|
||
timestamps for the return/scan/error-handler workflows, "Nog Niet Gebouwd" for the
|
||
RAGcore sync, and the real error-handler registration from this session's live
|
||
testing. Deployed to Unraid (commit `4049c0c6b12fef3d948cd31f21119044143320d8`,
|
||
rebuilt both `api` and `web`, `/health` OK) — user re-approved this second deploy
|
||
separately from the first.
|
||
- **Operational lesson learned this round**: `docker compose run --rm api pytest` does
|
||
**not** reliably pick up source edits without an explicit `docker compose build api`
|
||
first — a test file edit silently kept running against the stale built image (test
|
||
count didn't change) until rebuilt. Always `docker compose build api` (and `web` for
|
||
frontend changes) before trusting a green result after backend/frontend edits in this
|
||
repo.
|
||
- **WF1 acceptance gap fixed and live**: `check_drift.py`-style re-inspection of workflow
|
||
1 during this round's acceptance pass found the `Record follow-up` HTTP node had no
|
||
explicit timeout and "Retry On Fail" disabled — a real gap against the brief's
|
||
timeouts/bounded-retries requirement (WF2 already had this). Fixed live: Retry On Fail
|
||
(3 tries, 1000ms wait) + a 15000ms Timeout option, published (version note "Add bounded
|
||
retries (3x) and a 15s timeout to the Fleet Ops callback call"). Repo definition and
|
||
manifest checksum synced (`n8n/workflows/fleet-ops-vehicle-return.json`,
|
||
`MANIFEST.md`, new checksum `a6f399dd77a7203dec7c0ac95e8540abf55f2703da519e06f1c37f2e1220f609`,
|
||
commit `0562893`).
|
||
- **RAGcore credential issuance re-attempted and still blocked (user explicitly
|
||
authorized Claude to self-issue this round)**: tried the RAGcore admin UI's "Issue
|
||
credential" form for the `fleet-ops` application (logged in as Platform Admin, the
|
||
highest visible role) with name `n8n-ragcore-procedure-sync` and scope `sources:sync`
|
||
only. Submission failed with the same generic "Something went wrong. The credential
|
||
could not be issued with those values." page, this time carrying a trace reference
|
||
`1955c6a8968c4941a22a1faef39e17a7`. Inspected the RAGcore OpenAPI spec for this admin
|
||
endpoint (`POST /admin/control/applications/{application_id}/credentials`) — no
|
||
documented validation constraint explains the rejection (no 422, no field errors); the
|
||
`fleet-ops` application itself lists as ordinary/`Active` with no visible lock flag in
|
||
the applications table. This is the same failure signature as the earlier raw-API
|
||
attempt (400 "authoritative service-account state rejected issuance"): two independent
|
||
paths (raw API, and now the admin UI as the top admin role) both hit an opaque
|
||
server-side rejection with a trace ID. This is conclusive evidence the block is a
|
||
deliberate RAGcore-side policy or a RAGcore-side bug, not a Fleet Ops permission or
|
||
request-shape problem — nothing further is fixable from the Fleet Ops side or through
|
||
browser automation. Whoever operates the RAGcore instance needs to look up trace
|
||
`1955c6a8968c4941a22a1faef39e17a7` (and the earlier API rejection) in RAGcore's own
|
||
logs to find the real cause.
|
||
- **WF2 acceptance gap fixed and live**: continuing the acceptance pass to WF2 found it
|
||
had the *same* Retry On Fail gap as WF1 (its 15s timeout was already set, but retries
|
||
were off — the earlier note that "WF2 already had this" was wrong on the retry half).
|
||
Fixed live the same way (3 tries, 1000ms wait), published (version note "Add bounded
|
||
retries (3x) to the quality-scan HTTP call"). Repo definition and manifest checksum
|
||
synced (`n8n/workflows/fleet-ops-data-quality-scan.json`, `MANIFEST.md`, new checksum
|
||
`c0d46e0519118e6336e35c4ea2a67edb2f14bd007909ccf9256c93733751244a`, commit `167bf49`).
|
||
- **WF4 has the same gap on its own outbound call, but is currently un-fixable**: WF4's
|
||
"Report failure to Fleet Ops" HTTP node also has no timeout and no Retry On Fail. Began
|
||
the same fix (added a 15000ms Timeout option, toggled Retry On Fail on) but n8n's
|
||
autosave started failing with "Unauthorized" mid-edit, and a fresh tab confirmed the
|
||
n8n browser session had expired (redirected to `/signin`) — so nothing was saved and
|
||
the live WF4 definition is unchanged from before this round (no partial/broken state).
|
||
This is a minor, best-effort-only gap (WF4 is the error notifier itself, not a primary
|
||
business flow, and it already reports failures with `On Error: Stop Workflow` so a
|
||
failed error-report is visible in n8n's own execution history even without retries) —
|
||
not blocking, but worth finishing once someone re-authenticates the n8n browser
|
||
session.
|
||
- **RAGcore credential-issuance blocker root-caused and fixed (in RAGcore itself, with
|
||
explicit owner approval)**: with read access to the sibling `C:\Projects\RAGcore`
|
||
checkout, traced "authoritative service-account state rejected issuance" to a genuine
|
||
cross-transaction race in RAGcore's own dependency injection
|
||
(`src/ragcore/api/v1/control/dependencies.py`). `get_control_application` and
|
||
`get_credential_service` each independently opened their own `factory.begin()`
|
||
database transaction. Issuing a credential for a brand-new service account does, in one
|
||
request: (1) INSERT the service account via the first dependency's transaction, then
|
||
(2) immediately re-read it via the second dependency's *separate, uncommitted* transaction
|
||
— invisible under READ COMMITTED isolation until the first transaction commits, which
|
||
only happens after the endpoint returns. This made every fresh-service-account credential
|
||
issuance fail, 100% of the time, via both the raw API and the admin UI (explaining the
|
||
identical failure signature on both paths). RAGcore's own tests never caught this because
|
||
they override these dependencies with an in-memory fake that ignores transaction boundaries
|
||
entirely. Fixed by introducing one shared, cached `get_control_session` dependency that
|
||
both providers now depend on via `Depends(...)`, so they share one transaction per request.
|
||
Verified: RAGcore's own test suite (64 tests across `tests/web`, `tests/contract/api/control`,
|
||
`tests/security/identity`, `tests/unit/domain/control`, `tests/api`) passes, ruff and mypy
|
||
clean. Deployed to the live RAGcore instance (also on the Unraid host, `ragcore-app-1` on
|
||
port 1237 — a shared service also used by other ITWorx projects) via `docker compose build`
|
||
+ `up -d`, with explicit owner approval before both the code change and the deploy.
|
||
Confirmed fixed live: issuing a credential for `fleet-ops` (name
|
||
`n8n-ragcore-procedure-sync`, scope `sources:sync`) now succeeds (prefix `rc_sa_6fc51e`).
|
||
The plaintext token was never printed/logged — copied via RAGcore's own "Copy" button and
|
||
pasted directly into a new n8n Header Auth credential named **"RAGcore Sync Token"**
|
||
(header `Authorization: Bearer <token>`), ready for workflow 3.
|
||
- **Exact next action (superseded by the entry below)**: task #86 (build workflow 3,
|
||
RAGcore Procedure Sync) and the `RAGcoreKnowledgeProvider` adapter rewrite (to the real
|
||
inspected contract — `/health/live`, `/health/ready`, `POST /v1/uploads`, `POST
|
||
/v1/search`/`/v1/context`/`/v1/answers`) are now unblocked — the "RAGcore Sync Token" n8n
|
||
credential exists and works. WF4's own timeout/retry gap is still open pending n8n browser
|
||
re-authentication (minor, non-blocking, see above). Both #90 and #91 should be revisited
|
||
once workflow 3 is actually built, since they currently document it as blocked.
|
||
|
||
- **Second RAGcore bug found, fixed and deployed (with explicit owner approval, same pattern
|
||
as the transaction-race fix above)**: with the "RAGcore Sync Token" credential in hand,
|
||
the n8n "Upload to RAGcore" node still returned a persistent 401 on every item. Traced via
|
||
direct RAGcore source inspection (`C:\Projects\RAGcore`) to a genuine second, independent
|
||
gap: no code path in RAGcore converted an incoming `Authorization: Bearer <token>` header
|
||
into a `request.state.principal` for any `/v1/*` route — only browser session cookies were
|
||
ever accepted, even though the credential-verification logic
|
||
(`ServiceAccountCredentialService.verify()`) existed and was unit-tested. This blocks any
|
||
machine caller (n8n, and eventually Fleet Ops's own `RAGcoreKnowledgeProvider` adapter)
|
||
from ever authenticating to `/v1/uploads`. Fixed additively, scoped to `/v1/uploads` only
|
||
per owner instruction (search/context/answers left for later): new
|
||
`CredentialRepository.get_by_id()` (Postgres + in-memory), new
|
||
`ServiceAccountCredentialService.authenticate()` (parallel to the existing `verify()`, not
|
||
a refactor of it), and a new `get_upload_principal` FastAPI dependency
|
||
(`src/ragcore/api/v1/uploads/dependencies.py`) that falls back to the Bearer header when
|
||
there is no session principal, wired into `uploads/routes.py` in place of the session-only
|
||
`get_principal`. New/updated tests in `tests/security/identity/test_credentials.py` and
|
||
`tests/security/uploads/test_upload_security.py` (bearer-token accept/reject paths, the
|
||
existing route test's stale `get_principal` override fixed to `get_upload_principal`).
|
||
Verified: RAGcore's own test suite — 377 passed in the affected `tests/security`,
|
||
`tests/unit`, `tests/api` trees (3 unrelated pre-existing failures: two need Windows
|
||
symlink privileges the sandbox doesn't have, one is a git-connector fixture mismatch; a
|
||
separate architecture-boundary failure in `application/ingestion/handler.py` belongs to
|
||
unrelated in-progress work by a different concurrent agent on the same RAGcore checkout,
|
||
confirmed via `git status`/`git log` — not touched by this fix). Ruff and mypy clean on
|
||
every changed file. Deployed to the live RAGcore instance (`ragcore-app-1` on Unraid, port
|
||
1237 internally, fronted by `rag.itworx.tech` — note the *admin UI* and the *API* share
|
||
one process/origin, `/v1/uploads` is reachable at `https://rag.itworx.tech/v1/uploads`,
|
||
**not** `ragcore.itworx.tech`, which only appears in RFC7807 problem-type URLs) by copying
|
||
the 6 changed source files directly into the server checkout and `docker compose build
|
||
app && up -d --no-deps app` (deliberately not committing to RAGcore's git history or
|
||
touching the `worker` service, since a different agent has substantial unrelated
|
||
uncommitted work in that same working tree). Confirmed live with a garbage token (still
|
||
correctly 401) and then with a freshly-issued, correctly-scoped real token (403
|
||
`UPLOAD_TARGET_FORBIDDEN` against a dummy space ID — i.e. authentication succeeded,
|
||
authorization correctly rejected the wrong space — proving the fix end-to-end before
|
||
touching n8n at all).
|
||
- **Root cause of the n8n-side 401 found and fixed**: separately from the RAGcore bug above,
|
||
the "Upload to RAGcore" HTTP node's Authentication was set to Header Auth, but **no
|
||
credential had ever actually been attached** to that picker — so the node was sending no
|
||
`Authorization` header at all, which produces the identical 401 to a malformed one (easy
|
||
to conflate with the RAGcore-side bug, which is why fixing RAGcore alone didn't resolve
|
||
the symptom). There was already an unused "RAGcore Sync Token" n8n credential sitting
|
||
around from the earlier session (its value likely never actually got saved when it was
|
||
first created, or was created but never selected on this node — not conclusively
|
||
determined). Owner attached it and set Name=`Authorization`,
|
||
Value=`Bearer <freshly-issued token, scope sources:sync>` (a new credential issued via the
|
||
RAGcore admin UI at `/admin/control/applications/c20ac48a-d57b-4c68-9bd1-564f49c1a473/credentials/new`
|
||
specifically for this, service account "n8n Procedure Sync (production)"; the earlier
|
||
diagnostic-only credential used to prove the RAGcore fix was revoked afterward via direct
|
||
SQL `UPDATE identity.service_account_credentials SET revoked_at = now() ...` since the
|
||
admin UI has no revoke button).
|
||
- **Workflow 3's "Upload to RAGcore" node live-validated end-to-end, real data**: ran the
|
||
full workflow via n8n's "Execute workflow" (Schedule Trigger → List procedures → Prepare
|
||
uploads → Upload to RAGcore). All 33 items succeeded — each output item is a real
|
||
`AcceptedJob` (`job_id`/`status_url`), not error output. Independently confirmed at the
|
||
database level (not just trusting the n8n UI): `select count(*) from jobs.jobs where
|
||
operation='ingest_upload' and created_at > now() - interval '5 minutes'` → **33**, on the
|
||
live RAGcore Postgres.
|
||
- **n8n browser-automation notes for this environment** (worth knowing before attempting
|
||
canvas interaction again): (1) an n8n NPS survey modal (`role=dialog`, "We've been busy")
|
||
intermittently covers the whole canvas and silently eats every click underneath it until
|
||
removed; (2) canvas node positions reported by `getBoundingClientRect()` drift between
|
||
successive tool calls in a way that made coordinate-based `computer` clicks and even
|
||
`find`-ref-based clicks land on the wrong element repeatedly this session (dozens of failed
|
||
attempts, multiple different coordinate-math theories, none reliable) — directly setting
|
||
`.vue-flow__transformationpane`'s inline `style.transform` to force a node into view
|
||
**desyncs vue-flow's own internal pan/zoom state**, making the problem worse, not better;
|
||
(3) what actually worked reliably every time: calling native `.click()` directly via JS on
|
||
a plain `<button>` element (e.g. the toolbar's "Execute workflow" button) — canvas *node*
|
||
interaction (double-click to open a node's parameter panel) was never reliably achieved
|
||
via any automated method this session; the owner opened/edited the node manually instead.
|
||
A plain synthetic `dispatchEvent(new MouseEvent(...))` sequence (pointerdown/mousedown/
|
||
pointerup/mouseup/click/dblclick, even with correct `clientX`/`clientY`/`detail`) does
|
||
**not** trigger vue-flow's node click handling at all — it appears to require a genuinely
|
||
trusted (real CDP-driven) pointer event, consistent with vue-flow's drag/zoom gesture
|
||
system depending on native pointer capture.
|
||
- **Exact next action (superseded further below)**: build and wire the workflow's final
|
||
"Summarize sync result" Code node (count successes/failures across the 33 items) and a
|
||
closing HTTP node reporting to Fleet Ops's already-deployed `POST
|
||
/api/v1/integrations/n8n/procedures-sync-result` (task #86, still in progress — the sync
|
||
itself now works, this is the last piece). Then publish the workflow (currently still a
|
||
draft/unpublished), update `n8n/workflows/MANIFEST.md` and add
|
||
`n8n/workflows/fleet-ops-ragcore-procedure-sync.json` as the repo source-of-truth
|
||
definition, and revisit `artifacts/live-ai-integration/final-summary.md` (documents WF3
|
||
as blocked — no longer true).
|
||
|
||
- **Verified the 33-item sync is genuinely fully ingested, not just accepted**: checked at
|
||
every RAGcore pipeline stage on the live instance, not just trusting the n8n "success"
|
||
status (which can mask individual failures under `On Error: Continue`). All 33
|
||
`ingest_upload` jobs have `status='succeeded'` in `jobs.jobs`; 33 rows exist in
|
||
`content.documents`; 372 chunks were generated in `content.chunks`; 241 vectors are
|
||
indexed in the `rag_dense_nomic-embed-text_v1` Qdrant collection filtered specifically to
|
||
Fleet Ops's `space_id` (`f4c91e49-5cf9-48ba-b3d6-e0e9854ebccc`) — matching the expected
|
||
leaf-chunk count (structural/parent chunks in the hierarchy aren't separately embedded).
|
||
11 unique procedures × 3 languages = 33, all present.
|
||
|
||
- **`RAGcoreKnowledgeProvider` rewritten against the real contract (task #93), tested, but
|
||
NOT switched on in production — genuinely blocked on a RAGcore-side gap, not a Fleet Ops
|
||
problem**: rewrote `backend/app/services/knowledge/ragcore.py` end to end against
|
||
RAGcore's actual `/v1/answers` and `/health/ready` contracts (previously a best-effort
|
||
guess against an unreachable instance). Health now calls the real `GET /health/ready`.
|
||
`ask()` calls `POST /v1/answers` with `Authorization: Bearer <token>` and
|
||
`requested_space_ids: [settings.ragcore_space_id]`, maps RAGcore's `answerability` enum
|
||
conservatively (`answerable`/`partially_answerable` *with* non-empty citations →
|
||
`grounded`, everything else → `insufficient`, any transport/parse/non-200 failure →
|
||
`unavailable`, matching the architecture's never-fabricate rule). Added
|
||
`RAGCORE_SPACE_ID` setting (`.env.example`, `compose.yaml`). New tests in
|
||
`backend/tests/test_knowledge.py` (connection-error degrade, missing-space-id short
|
||
circuits without a network call, health ready/degraded/unreachable, grounded citation
|
||
mapping, not-answerable and answerable-without-citations both correctly map to
|
||
`insufficient` with an empty answer, non-200 and malformed-body both degrade to
|
||
`unavailable`) — **23 passed** in that file, **172 passed** full suite, ruff clean, mypy
|
||
0 issues/50 files.
|
||
**Discovered while live-testing against RAGcore with a freshly-issued, correctly-scoped
|
||
credential** (`Fleet Ops Knowledge Assistant (production)`, scope `answer`, space
|
||
`f4c91e49-...`): authentication now genuinely succeeds (past the 401 stage — confirmed
|
||
via `/health/ready` returning `200 {"status":"ok", ...}` with the same token), but **both
|
||
`POST /v1/answers` and `POST /v1/search` return `503`
|
||
`ANSWERS_UNAVAILABLE`/`SEARCH_UNAVAILABLE`** for every request. Root-caused by reading
|
||
`src/ragcore/main.py`'s app-startup/lifespan code directly: it constructs and assigns
|
||
`app.state.database_engine`, `session_factory`, `qdrant_client`,
|
||
`query_lab_service`, `profile_activation_service`, and the OIDC session services — but
|
||
**never constructs or assigns `app.state.search_application`, `context_application`, or
|
||
`answer_application`** anywhere in the codebase (confirmed via a repo-wide grep — the
|
||
three `get_*_application` dependency functions exist and correctly raise their `503
|
||
*_UNAVAILABLE` problem when the state attribute is absent, exactly as designed, but
|
||
nothing ever populates it). This is not a config toggle Fleet Ops is missing and not
|
||
something introduced by today's auth fixes — the retrieval/generation subsystem's route
|
||
handlers and dependencies are scaffolded end-to-end but were never wired into the running
|
||
application on this RAGcore deployment. Only masked until today because every call to
|
||
these routes previously 401'd on auth before ever reaching this check.
|
||
**Decision: `KNOWLEDGE_PROVIDER` stays `demo` in production.** Flipping it to `ragcore`
|
||
right now would replace the currently-working demo Knowledge Assistant with one that
|
||
correctly, honestly, but uselessly reports "unavailable" for every question — strictly
|
||
worse for the live demo. The adapter code itself is finished, correct, and safe to ship
|
||
(already committed-worthy), and switching providers is a one-line env var flip
|
||
(`KNOWLEDGE_PROVIDER=ragcore` + set `RAGCORE_API_TOKEN`/`RAGCORE_SPACE_ID`) the moment
|
||
RAGcore's own operator wires up `search_application`/`answer_application` on their side.
|
||
The freshly-issued credential (`Fleet Ops Knowledge Assistant (production)`, scope
|
||
`answer`) was left active/unused in RAGcore, ready for that day.
|
||
|
||
- **MCP Hub registration (task #94/#95) — full new connector built and validated locally
|
||
in the sibling `C:\Projects\ITWorx_MCP_Hub` checkout, not committed or deployed**: the
|
||
Hub's admin UI (`mcp.itworx.tech`, real production instance, 7 pre-existing projects —
|
||
DevRunbook, ForgeFlow×2, General Infrastructure, GeoIntel, Ludarium) has **no self-service
|
||
"add project/connector" flow** — its own Settings page states "No writable settings are
|
||
available in this browser until the control API publishes an authorized configuration
|
||
schema" (Hub is on release `1.0.0-rc`, UI is read-only/observability-only today).
|
||
Registering MobilityOps therefore required building a genuinely new connector in the
|
||
Hub's own repo, following its `gitea` connector as the closest real template (external
|
||
HTTPS API, shared-secret auth) rather than `knowledge` (which turned out to be an
|
||
in-memory filesystem index, not an HTTP client, despite the name suggesting otherwise).
|
||
Built, with explicit owner approval given the Hub's own `CLAUDE.md` restricts autonomous
|
||
action to local-checkout work only (registry pushes and production startup are separate,
|
||
explicitly-gated checkpoints, not done this round):
|
||
- `packages/connector_kit/mobilityops.py` — `MobilityOpsSettings`/`MobilityOpsClient`
|
||
(HTTPS-only, same-origin-redirect-enforced, `X-Service-Token`/`X-Client-Id` headers,
|
||
path-safety-validated `vehicle_ref`), mirroring `gitea.py`'s hardening exactly.
|
||
- `connectors/mobilityops/{__init__,server,fake}.py` — `MobilityOpsConnector` exposing 4
|
||
read-only tools (`mobilityops.operations.summary`, `.attention.list`, `.vehicle.get`,
|
||
`.knowledge.search`), each `governed_tool`-wrapped, envelope-wrapped, field-mapped from
|
||
Fleet Ops's real `/api/v1/integrations/mcp/*` response shapes.
|
||
- `services/runtime/connector_mobilityops.py` + `dispatch.py` registration.
|
||
- Catalog: `catalog/connectors/mobilityops-default.yaml` (ConnectorTemplate),
|
||
`catalog/tools/mobilityops-*.yaml` (4 ToolManifests), `catalog/capability-packs/
|
||
mobilityops-reader.yaml` (a **new**, narrowly-scoped pack — deliberately not added to
|
||
the shared `project-reader` pack other projects use, to avoid granting them
|
||
MobilityOps access), `catalog/projects/mobilityops.yaml` (Project, `applicationUrl`
|
||
set to the real internal `http://192.168.10.150:1236`, not an invented public domain).
|
||
- Compose wiring across all three files (`docker-compose.yml`, `.prod.yml`,
|
||
`.blueprint.yml`) plus a new `mobilityops_service_token` Docker secret.
|
||
- Fixed ~13 regressions this surfaced in the Hub's own existing test suite — all
|
||
legitimate guard-rail tests (hardcoded service/tool/secret allowlists, a
|
||
`CONNECTOR_GATEWAYS`/`FakeContextForge._gateway_tools` registration gap that was a
|
||
**real production wiring miss**, not just a test-fixture gap: without it, the reconcile
|
||
step would have raised `"virtual server resolved to an empty tool allowlist"` against
|
||
the live Hub too) — not scope creep, this is exactly the Hub's own established pattern
|
||
for registering a new connector, verified by reading how `gitea`/`knowledge`/`unraid`
|
||
each touch the same ~15 files.
|
||
- Verified: 15 new connector tests pass; full Hub suite **383 passed, 0 failed, 5 skipped
|
||
(pre-existing)**; `ruff check`/`ruff format --check` clean; `mypy` clean across 108
|
||
source files; `scripts/check_boundaries.py` passes (import-direction rules respected —
|
||
the new connector only imports from `packages/connector_kit`, never `services/*`
|
||
directly); `scripts/validate_pack.py` schema-validates the catalog cleanly (68
|
||
documents, up from 61 — exactly the 7 new files) with the only remaining failure being
|
||
the Hub's own git-cleanliness gate for manifest regeneration, which requires a commit.
|
||
- **Not done, deliberately**: no commit (the Hub repo has substantial unrelated
|
||
in-progress work from a different concurrent agent — `WP-235` frontend redesign,
|
||
`BUILD_STATE.json`/`CURRENT_STATE.md`/`implementation/WORK_PLAN.json` all show as
|
||
modified by them, not by this session — committing broadly risks entangling that
|
||
work); no registry push; no production deployment/restart of the live
|
||
`itworx-mcp-hub-*` stack. Per the Hub's own `CLAUDE.md` approval boundaries, those are
|
||
separate, explicitly-gated checkpoints requiring their own fresh approval, and a real
|
||
`MOBILITYOPS_ENDPOINT`/`mobilityops_service_token`/`MOBILITYOPS_PROJECTS_JSON`
|
||
production configuration still needs to be decided before any of that could run.
|
||
- **MCP Hub connector — committed** as `f107544` ("Add MobilityOps (Fleet Ops) read-only
|
||
connector", 28 files) in `C:\Projects\ITWorx_MCP_Hub`, `main` branch. Not pushed to the
|
||
registry and not deployed to the live Hub stack — those remain separate, explicitly-gated
|
||
checkpoints per the Hub's own `CLAUDE.md` approval boundaries.
|
||
|
||
- **RAGcore search/context/answer application wiring (task #96) — implemented, tested,
|
||
committed as `a2905cc`** in `C:\Projects\RAGcore`, `main` branch (not pushed, not
|
||
deployed to the live RAGcore instance). This closes the gap documented above:
|
||
`search_application`/`context_application`/`answer_application` are now actually
|
||
constructed in `create_app()` instead of being permanently absent.
|
||
- `src/ragcore/application/retrieval/production_executor.py` (new) — `RetrievalPipeline`
|
||
(embed via Ollama -> dense+sparse Qdrant search -> RRF fuse -> rerank via Ollama),
|
||
reusing the same flow already proven by the Query Lab's `PipelineQueryLabExecutor`
|
||
(left untouched). Feeds `ProductionSearchExecutor` and `ProductionContextExecutor` from
|
||
one shared per-request run, so `SearchHit`'s `FusedCandidate` (required by
|
||
`SearchService`'s own provenance/scope check) and `ContextCandidate`'s parent-chunk
|
||
expansion both derive from the same reranked result set instead of drifting apart.
|
||
- `src/ragcore/application/models/answer_generator.py` (new) — `OllamaAnswerGenerator`,
|
||
`AnswerService`'s real generator: resolves the active `GenerationProfile` from
|
||
`GenerationProfileRegistry` (already built, never previously constructed anywhere) and
|
||
calls the existing hardened `OllamaGenerationAdapter`. `main.py`'s `lifespan()` now
|
||
seeds and activates a default local-Ollama answer profile on first boot if none is
|
||
active, since no seed data ever registered one.
|
||
- Verified: 10 new unit tests (7 for the pipeline/executors, 3 for the answer generator)
|
||
pass; `ruff check` clean; `mypy src` clean (the only 4 mypy errors found are in
|
||
`infrastructure/qdrant/adapter.py`, `infrastructure/qdrant/filters.py`,
|
||
`api/v1/control/routes.py` — files this change never touched, confirmed via `git
|
||
status`/`git diff`, pre-existing or from other concurrent work in this checkout).
|
||
- Full non-integration suite (`pytest -m "not integration"`, excludes tests requiring
|
||
disposable infra per the repo's own marker convention): **593 passed, 14 failed, 1
|
||
skipped**. All 14 failures are pre-existing and unrelated to this change: 10 are
|
||
Windows-only (`psycopg`'s async driver rejects Windows' default `ProactorEventLoop` —
|
||
`test_health.py`, `test_readiness.py`, `test_control_plane.py`), 2 are Windows-only
|
||
(`WinError 1314`, no symlink privilege on this account — `test_folder_connector.py`), 1
|
||
is an unrelated pre-existing git-security-connector assertion, and 1 is the repo's
|
||
zero-tolerance `test_domain_and_application_respect_dependency_boundaries` test —
|
||
already red before this change, since `application/ingestion/handler.py` (untouched,
|
||
commit `bb75f52`) already imports `httpx`/infrastructure modules from the application
|
||
layer, which that test forbids with no allowlist. This change's new files follow the
|
||
same established (if already-violating) pattern to reach real adapters, adding more
|
||
instances of the same pre-existing violation rather than introducing a new kind of one.
|
||
- **Deployed to production** (`http://192.168.10.150:1237` via `ragcore-app-1`),
|
||
approved by the user. `scp`'d the 4 changed/new files individually to
|
||
`/mnt/user/appdata/ragcore/app/...`, `docker compose build app`, verified the built
|
||
image actually contains the new code (`docker run --rm ragcore-app grep`/`test -f`),
|
||
then `docker compose up -d --no-deps app` (only the `app` service touched). Startup
|
||
logs are clean (`ragcore_started environment=production`, no errors); `/health/ready`
|
||
reports `postgres`/`qdrant`/`storage` all `ok`.
|
||
- **Live-verified the fix itself**: `POST /v1/search` now returns `401
|
||
AUTHENTICATION_REQUIRED` for an unauthenticated call, not the old permanent `503
|
||
SEARCH_UNAVAILABLE` — proof the endpoint reaches real request handling instead of
|
||
hitting an absent application state, which is exactly the bug this closes.
|
||
- **Issued a fresh scoped credential** (`answer` scope, application `Fleet Ops`, via the
|
||
RAGcore admin UI at `rag.itworx.tech` — a separate, working OIDC session, unaffected by
|
||
the n8n auth problem above) and made a real authenticated `POST /v1/answers` call
|
||
against the `Fleet Ops Procedures` knowledge space (`f4c91e49-5cf9-48ba-b3d6-e0e9854ebccc`).
|
||
Got a clean `200`, a real `answer_id`/`retrieval_run_id`, `degraded: false` — but
|
||
`answerability: not_answerable`, 0 citations, for two different real questions matching
|
||
real seeded document titles.
|
||
- **Confirmed this is not a regression from this change**: the same query against the
|
||
same space through RAGcore's own pre-existing, already-proven Query Lab tool (which
|
||
uses `PipelineQueryLabExecutor`, code this session never touched) returns the identical
|
||
`NOT_ANSWERABLE` / 0 evidence / not degraded result — down to `Dense candidates (0)` and
|
||
`Sparse candidates (0)` at the raw retrieval stage, before fusion or rerank ever runs.
|
||
Whatever is causing zero matches lives in shared retrieval infrastructure or the data
|
||
itself, not in the new production executors.
|
||
- **Ruled out the obvious causes via direct Qdrant/Postgres checks** (not yet root-caused
|
||
further — flagging as a separate follow-up, not blocking this task): the space
|
||
genuinely has 83 published, correctly-scoped points in `rag_dense_nomic-embed-text_v1`
|
||
(right `workspace_id`, right `space_id`, `status: published`, real Dutch-language
|
||
procedure text); the points' `embedding_model_digest` exactly matches the currently
|
||
active embedding profile's digest (no stale-embedding mismatch); the collection alias
|
||
(`rag_dense_nomic-embed-text_active`) correctly resolves to that same collection. One
|
||
oddity noted in passing, unrelated to retrieval: at least one sampled chunk has
|
||
`language: "en"` in its payload despite the actual text being Dutch — worth a look if
|
||
language-filtered queries matter for the demo.
|
||
- **Decision on `KNOWLEDGE_PROVIDER=ragcore` in Fleet Ops**: still deliberately `demo`.
|
||
The application is deployed, correctly wired, and behaves identically to RAGcore's own
|
||
trusted reference implementation — but real questions against real seeded content
|
||
aren't returning grounded answers yet for a reason that traces to shared
|
||
retrieval/data, not to this wiring. Flip only after that's root-caused.
|
||
|
||
- **n8n workflow 3 (task #86) — incident, recovered; final 2 nodes still blocked on a
|
||
live n8n auth problem, not a code problem.** Returning to finish the "Summarize sync
|
||
result" + report-to-Fleet-Ops nodes found the live "Fleet Ops — RAGcore Procedure Sync"
|
||
workflow's canvas at **zero nodes** — the earlier session's abandoned attempt to call
|
||
n8n's REST API directly (the "unexplained 401" noted above) had gone through far enough
|
||
to wipe the live workflow before the 401 stopped it, leaving an empty, unsaved "Current
|
||
changes" draft on top of the last good save.
|
||
- **Recovered**: n8n's own Version History (`/workflow/.../history`) still had the last
|
||
good save, version `260b38b5` (Aug 4, 21:08), with all 4 real nodes intact (Schedule
|
||
Trigger -> List procedures -> Prepare uploads -> Upload to RAGcore) plus the build
|
||
sticky note. Used the history panel's own **Restore version** action (not a manual
|
||
rebuild) to bring the live workflow back to that exact state. Confirmed via the page's
|
||
DOM (`[data-test-id="canvas-node"]` count) both before (0) and after (4) — the canvas
|
||
itself render fully off-screen (nodes positioned at negative Y coordinates, a separate
|
||
display-only vue-flow pan bug worked around by directly setting the transform pane's
|
||
CSS transform; harmless, doesn't touch saved data).
|
||
- **Then blocked again attempting the 2 new nodes**: adding a Code node via the UI's own
|
||
"What happens next?" panel (a real, first-party n8n interaction, not the REST API)
|
||
immediately surfaced n8n's own autosave toast: **"Problem saving workflow — Autosave
|
||
failed: Unauthorized."** A full sign-out/sign-in cycle right before this (fresh
|
||
credentials, fresh page load) did not fix it — immediately after a successful
|
||
interactive "Sign in", `fetch('/rest/workflows/...')` from that same authenticated tab
|
||
still returned 401. This is not the ordinary "your session expired, log in again"
|
||
friction seen earlier in the session; it's the live n8n instance's own save path
|
||
rejecting a request made moments after a successful login, which is exactly the
|
||
failure mode that caused the wipe above. Continuing to add nodes under this condition
|
||
risks losing work again with no guarantee the next failure is as recoverable, so this
|
||
was stopped deliberately rather than retried. The workflow was left in its safely
|
||
restored, 4-node, unpublished state — verified via the DOM node count immediately
|
||
before closing the session — nothing was added or changed beyond the restore.
|
||
- **Not a code/workflow-design problem**: the 2 remaining nodes (a Code node
|
||
summarizing the sync result, and an HTTP node reporting to Fleet Ops's already-live
|
||
`POST /api/v1/integrations/n8n/procedures-sync-result`) were never built — this
|
||
blocked before any node configuration happened. The design itself (mirroring WF2/WF4's
|
||
existing report-result pattern) is unchanged from earlier planning.
|
||
|
||
- **n8n workflow 3 — root cause found and fixed; both remaining nodes now built and
|
||
saved.** The user pushed back on the "auth problem" framing (n8n visibly showed signed
|
||
in), which was the right call — dug into the live n8n container's own logs
|
||
(`docker logs n8n`) rather than continuing to guess from the browser side, and found the
|
||
real cause: `browserId check failed on /rest/workflows/:workflowId`, alongside an
|
||
express-rate-limit warning about `X-Forwarded-For` being present while Express `trust
|
||
proxy` is `false`. n8n's Unraid template
|
||
(`/boot/config/plugins/dockerMan/templates-user/my-n8n.xml`) had `N8N_PROXY_HOPS=0`
|
||
despite genuinely running behind the TLS-terminating reverse proxy at
|
||
`n8n.itworx.tech` — with proxy trust disabled, n8n couldn't correctly resolve the
|
||
request as HTTPS, which broke its CSRF-style browserId cookie check on every
|
||
workflow-mutating REST call (autosave, but not plain page loads, which explains why the
|
||
UI looked fully logged in the whole time).
|
||
- **Fix, approved by the user beforehand** (shared instance, brief restart): backed up
|
||
the template to a timestamped `.bak`, changed `N8N_PROXY_HOPS` `0` -> `1`, then
|
||
recreated the `n8n` container on the Unraid host (stop, rename to
|
||
`n8n_pre_proxyhops_fix` as an instant rollback, `docker run` with every existing env
|
||
var/volume/port/label preserved exactly plus the one corrected value). All 3
|
||
previously-active workflows (Vehicle Return Orchestration, Scheduled Data Quality
|
||
Scan, Workflow Error Handler) re-activated cleanly on the new container — nothing
|
||
lost. Verified the fix by reproducing the exact save action that used to fail
|
||
(adding a node via n8n's own UI): no more "Unauthorized" toast, and the added node
|
||
survived a full page reload. `docker logs n8n` shows zero `browserId`/401 entries
|
||
since the fix.
|
||
- **Mid-fix-verification, a real node got deleted by an errant `Ctrl+A`** (browser
|
||
focus landed on the canvas instead of a text field during cleanup, selecting and then
|
||
deleting the "List procedures" node). Caught it via a node-count check, used n8n's own
|
||
Version History **Restore version** to the last good save (not a manual rebuild), then
|
||
redid the 2 new nodes carefully (verifying focus before every `Ctrl+A` this time). No
|
||
data lost, just redone.
|
||
- **Built**: `Summarize sync result` (Code node, JS) — counts `synced`/`failed` from
|
||
`$input.all()` by checking each item's `json.error` (matches "Upload to RAGcore"'s own
|
||
`On Error: Continue` setting, confirmed by inspecting that node's Settings tab, so
|
||
failed uploads land in the same output stream with an `error` field rather than a
|
||
separate branch). `Report sync result to Fleet Ops` (HTTP Request, POST) — targets the
|
||
already-live `/api/v1/integrations/n8n/procedures-sync-result`, reusing the existing
|
||
"Fleet Ops Service Token" Header Auth credential (same one WF4 already uses against
|
||
this same backend), JSON body `{execution_id, synced, failed}` matching
|
||
`ProcedureSyncResultIn` exactly.
|
||
- **Saved, not published.** All 6 nodes (Schedule Trigger -> List procedures -> Prepare
|
||
uploads -> Upload to RAGcore -> Summarize sync result -> Report sync result to Fleet
|
||
Ops) confirmed present after a full page reload. The Schedule Trigger is configured
|
||
for once daily at midnight -- publishing activates real, live, unattended runs against
|
||
production RAGcore and Fleet Ops, so this was deliberately left for a separate,
|
||
explicit approval rather than done automatically.
|
||
|
||
- **Exact next action**: (1) RAGcore search/answer wiring (`a2905cc`) is deployed and
|
||
proven correctly wired (matches Query Lab's trusted behavior exactly) — the open item is
|
||
now root-causing why `Fleet Ops Procedures` returns zero dense/sparse candidates for
|
||
real questions despite having genuinely matching, correctly-scoped, correctly-embedded
|
||
indexed content (see evidence above); this is retrieval/data, not application wiring.
|
||
Only flip `KNOWLEDGE_PROVIDER=ragcore` once that's fixed and a real grounded answer comes
|
||
back. (2) n8n workflow 3 is fully built and saved (all 6 nodes) but **not published** —
|
||
decide whether to publish it (starts real daily runs against production) and, once
|
||
live-verified end to end, add its file to `n8n/workflows/` and `MANIFEST.md` as the
|
||
source of truth alongside the other 3 workflows. (3) MCP Hub: registry push + production
|
||
deployment of the now-committed connector remain separate, explicitly-gated checkpoints.
|
||
WF4's own timeout/retry gap remains open, deferred, non-blocking.
|
||
|
||
## Branch push and Unraid redeploy (2026-08-05)
|
||
|
||
- Pushed `feat/live-n8n-ragcore-integration` to `origin` (Gitea on the same Unraid host),
|
||
now tracking `origin/feat/live-n8n-ragcore-integration` (`0571a40`).
|
||
- **Redeployed the live Fleet Ops instance** (`http://192.168.10.150:1236`) from this
|
||
branch, bringing it from the previously-deployed `0da5251` up to `0571a40` — 5 commits,
|
||
the meaningful one being the `RAGcoreKnowledgeProvider` rewrite (`e5d8466`); the rest
|
||
are `PROJECT_STATE.md`-only evidence commits.
|
||
- Followed the deployment directory's own established archive convention
|
||
(`deploy/unraid/README.md`: "deployed from a committed source archive"): `git archive`
|
||
of `HEAD` as `source-0571a40.tar.gz`, matching the naming pattern of the existing
|
||
archives already in `/mnt/user/appdata/mobilityops/.deploy/`; `scp`'d it there;
|
||
extracted over the deployment directory excluding `.env` and `.deploy` itself (so the
|
||
live secrets file and archive history were never touched); updated
|
||
`.deploy/source-revision` to the new full commit hash, matching its existing format.
|
||
- `docker compose -p mobilityops -f compose.yaml -f compose.unraid.yaml up --build -d db
|
||
api web` — `api` image rebuilt and the container recreated (picked up the new code);
|
||
`web` rebuilt too but didn't need recreating (no frontend changes this branch); `db`
|
||
untouched. Migrations ran automatically in the `api` entrypoint with no errors.
|
||
- **Verified live**: `/health` returns `200` from inside the `api` container; the new
|
||
`/api/v1/integrations/n8n/procedures` endpoint (added this branch, needed for n8n
|
||
workflow 3) is reachable through the public web proxy and correctly enforces auth
|
||
(`422`, missing `X-Service-Token`, not a `404` — proves the route exists and is live,
|
||
not just that the proxy responds). `KNOWLEDGE_PROVIDER=demo` confirmed unchanged in the
|
||
live `.env` (only extracted new source files, never touched it) — the demo Knowledge
|
||
Assistant is still what's live, exactly as intended.
|
||
- No seed reset run (this was an update to a running instance with real accumulated
|
||
demo data, not the initial deploy — re-seeding would have been destructive and wasn't
|
||
warranted by anything in this branch's changes).
|
||
- **Not done**: no PR opened/merged to `master` — the user asked for commit, push, and
|
||
redeploy, not a merge; `master` is untouched and still 19 commits behind this branch.
|
||
|
||
## Final integrations pass — Batches 1-3 (branch `feat/fleet-ops-final-integrations`, 2026-08-05)
|
||
|
||
Branched from `feat/live-n8n-ragcore-integration` @ `3ebca9e` (the actually-validated,
|
||
deployed line — `master` is stale, 19 commits behind). Full audit at
|
||
`docs/final-integrations/current-state-audit.md`.
|
||
|
||
- **GUI polish**: dashboard Attention Queue now shows a curated severity mix (grouped
|
||
"Handle now / Follow up today / Review later" headers) instead of pure severity-sort
|
||
crowding out medium/low items. Seed `bookings.csv` curated so a fresh reset lands ≥2
|
||
departures and ≥2 returns on the anchor day (new `test_seed_today_movements_are_a_credible_mix`
|
||
test). About Demo restructured into a compact grid with `<details>` progressive
|
||
disclosure for architecture/security/testing. Duplicate Merge shows match/conflict
|
||
counts, hides matching fields by default (toggle to reveal), and previews the final
|
||
merged record before confirmation.
|
||
- **n8n**: fixed genuinely invalid JSON in the committed `fleet-ops-vehicle-return.json`
|
||
(a missing `},` — the file could not be parsed). Live-validated workflow 3 (RAGcore
|
||
Procedure Sync): 6 real nodes confirmed built and saved but correctly still inactive;
|
||
found and fixed two real defects via the safe `n8n import:workflow` CLI path (not the
|
||
REST API) — three body-parameter expressions had a stray trailing `}}`, and
|
||
`settings.errorWorkflow` was unset. Exported as `n8n/workflows/fleet-ops-ragcore-procedure-sync.json`,
|
||
added to `MANIFEST.md`/`check_drift.py`. **Still not published** — daily unattended
|
||
runs against production remain a separate go-live decision.
|
||
- **RAGcore retrieval root cause — found and fixed (live, user-approved)**: the "zero
|
||
candidates" bug was a filesystem permission bug, not authorization or data —
|
||
`/workspace/.state/models/embedding_profiles.json` (read on every retrieval call) was
|
||
`root:root` mode `600` on the host bind mount, unreadable by the app's actual runtime
|
||
uid (10001). Verified every other suspect healthy first (grants, real
|
||
`RetrievalAuthorizationService` resolution run in-process, exact production Qdrant
|
||
filter run directly, real ANN query) before finding this. Fixed via `chown`/`chmod` on
|
||
the host-mounted path; re-verified in-process — 5 real relevant hits, up from 0.
|
||
**Second, deeper gap found (not fixed)**: `DEFAULT_RERANKER_PROFILE` calls
|
||
`{ollama}/api/rerank`, which does not exist on the deployed Ollama (`0.32.5`) — a plain
|
||
404, not a missing-model problem (pulled `xitao/bge-reranker-v2-m3:latest`, 1.2GB, with
|
||
approval; did not fix it). `/v1/answers` still returns `not_answerable`/0 citations
|
||
live for real NL/EN/FR questions against real matching content. User decision: leave
|
||
`KNOWLEDGE_PROVIDER=demo`, do not chase the reranker fix further this session — it
|
||
needs either a shared-Ollama version upgrade (affects every other project on that
|
||
instance) or a RAGcore-side code change, and `C:\Projects\RAGcore`'s working directory
|
||
had 36 uncommitted files from what looks like another active session, so no commits
|
||
were made there. **Side effect**: minting the live-test credential rotated the existing
|
||
"Fleet Ops Knowledge Assistant (production)" RAGcore service-account credential (its
|
||
2-active-credential cap was already reached) — a fresh credential must be issued before
|
||
actually flipping `KNOWLEDGE_PROVIDER=ragcore`, the old one is now invalid.
|
||
- **MCP Hub status**: `derive_mcp_hub_status()` now reports real tool-call audit
|
||
evidence (last tool/client/timestamp, total calls) instead of just echoing
|
||
`MCP_HUB_REGISTRATION_ENABLED`. Confirmed via a sibling-repo investigation that the
|
||
Fleet Ops connector is already live in ITWorx MCP Hub's own production deployment
|
||
(Tower, commit `c4a0f6d`), with a real contract fix already applied there (`vehicle.get`
|
||
wire parameter normalized to `vehicleRef`) — Batch 4 is narrower than the task brief
|
||
assumed.
|
||
- Repo hygiene: removed untracked `backend;C` (empty dir) and a 31MB `MobilityOps.zip`
|
||
stray export; `.gitignore` now excludes `*.zip`/`*.tar.gz`.
|
||
- Evidence: `docker compose run --rm api pytest -q` — **172 passed**; `ruff check .` —
|
||
clean; `mypy app` — clean (50 files); `cd frontend && npm run build` — clean.
|
||
- Committed `34df66d`, pushed to `origin/feat/fleet-ops-final-integrations`. Not yet
|
||
deployed to the live Fleet Ops instance at this point in the session.
|
||
- Deployed to `http://192.168.10.150:1236` (`git archive` + scp + `docker compose up
|
||
--build -d db api web`, following the established deployment convention). Browser-
|
||
verified live: Attention Queue tier grouping, Today's Movements curated seed mix (2
|
||
returns + 2 departures after a real demo reset), About Demo progressive disclosure,
|
||
Duplicate Merge match/conflict summary + preview all render correctly.
|
||
- **Found and fixed a real methodology gap mid-session**: `compose.yaml`'s `api` service
|
||
has no bind mount, so `docker compose run --rm api pytest/ruff/mypy` was silently
|
||
testing a stale image for an unknown portion of this session (only caught because a new
|
||
test's collected-count didn't match). Rebuilt and re-ran every local gate from that
|
||
point on; found exactly one genuinely stale test assertion (`ragcore_sync["built"]` —
|
||
correctly `True` now, not `False`, since workflow 3 really is built) and two line-length
|
||
lint errors in the new seed test, both fixed. See the sharper operational note added
|
||
next to the original `docker compose run` warning above.
|
||
|
||
## Batch 4 — MCP Hub (2026-08-05)
|
||
|
||
- `MCP_HUB_BASE_URL`/`MCP_PROVIDER_ID` were declared in `.env.example` but never read by
|
||
`Settings` anywhere — dead config. Since the Hub's own registration is catalog-driven
|
||
(it reconciles its catalog into the gateway; Fleet Ops never pushes a registration
|
||
call — confirmed via the Hub-side investigation), wiring them for self-registration
|
||
would have built an unneeded feature. Wired `mcp_hub_base_url` for something Fleet Ops
|
||
actually needs instead: a real, bounded (1.5s timeout) Hub reachability health check,
|
||
surfaced as `hub_reachable` on `/api/v1/integrations/status` and the Automation page.
|
||
- Fixed two real, concrete gaps in `mcp_integrations.py`'s `search-knowledge` endpoint
|
||
matching the task brief's own description almost verbatim: no `locale` field existed at
|
||
all (now `nl-BE`/`en-GB`/`fr-BE`, wired straight through to the knowledge provider,
|
||
which already supported a `language` param — just never received one), and the
|
||
correlation ID was **always** freshly minted (`uuid.uuid4()`), ignoring any inbound
|
||
`X-Correlation-Id` header. Added a shared `get_correlation_id` dependency (valid inbound
|
||
UUID preserved end-to-end into Fleet Ops's own audit log; fresh UUID only when absent/
|
||
invalid) and applied it to all four MCP endpoints, not just search-knowledge.
|
||
- Fleet Ops's own internal audit tool labels renamed `mobilityops_*` → `fleet_ops_*`
|
||
(`fleet_ops_get_operations_summary`, `.list_attention_vehicles`, `.get_vehicle_details`,
|
||
`.search_knowledge`) and mirrored in `contracts/mcp-tools.json` (with `mobilityops_*`
|
||
kept as `deprecated_aliases`, per the task's own "don't break existing clients"
|
||
instruction). Note: the **live** ITWorx MCP Hub connector publishes these under its own
|
||
dotted namespace (`mobilityops.operations.summary` etc.) — that naming is Hub-owned and
|
||
was deliberately not touched (separate repo, active concurrent session there, and
|
||
already verified live per the Hub-side investigation); `docs/10-mcp-hub-integration.md`
|
||
now documents both naming layers explicitly so this isn't read as a contradiction.
|
||
- Automation page's MCP card now shows real evidence (last tool/client/call count/
|
||
timestamp, honest "registered but no calls yet" state) instead of only the
|
||
registration-enabled boolean, plus the new Hub-reachability badge.
|
||
- New tests: `test_search_knowledge_respects_requested_locale`,
|
||
`test_search_knowledge_preserves_inbound_correlation_id`,
|
||
`test_operations_summary_mints_correlation_id_when_none_supplied`.
|
||
- Evidence: `docker compose run --rm api pytest -q` — **176 passed** (against a verified
|
||
fresh rebuild); `ruff check .` — clean; `mypy app` — clean (50 files); `cd frontend &&
|
||
npm run build` — clean.
|
||
- **Exact next action**: commit and deploy Batch 4; Batch 5 (AI Operations Brief using the
|
||
demo knowledge provider since RAGcore stays off; GUI activity showcase optional-after-
|
||
demo-complete; final regression across all gates; `artifacts/final-integrations/final-summary.md`).
|
||
|
||
## Batch 5 — AI Operations Brief and final regression (2026-08-05)
|
||
|
||
- Ran a real AI Operations Brief through the live ITWorx MCP Hub connector's own
|
||
`MobilityOpsClient` class against production Fleet Ops (no mocks, no chatbot): real
|
||
operations summary, real most-pressing vehicle (`MO-031`), real grounded English
|
||
knowledge answer with 2 citations, real correlation IDs verified end-to-end in Fleet
|
||
Ops's own audit log. Dutch/French question variants honestly returned `insufficient`
|
||
(the live Hub connector doesn't yet send the new `locale` field — a Hub-side follow-up,
|
||
not silently worked around). Full runbook and live output:
|
||
`docs/final-integrations/ai-operations-brief-runbook.md`.
|
||
- Ran the full Playwright e2e suite against the live deployed instance (not just
|
||
localhost) — every spec file, ~200 tests across responsiveness/i18n/accessibility/
|
||
guided-demo/interactive-elements/audit/data-quality/route-matrix. Found and fixed two
|
||
pre-existing fragile locators (both broke because Automation legitimately has two
|
||
`.data-table`-classed tables now, exposed only by running the full suite together
|
||
rather than individual files — unrelated to this session's feature changes) and one
|
||
pre-existing untranslated-loanword false positive ("Workflow" is spelled identically in
|
||
nl-BE/fr-BE). All specs pass after the fixes.
|
||
- Wrote `artifacts/final-integrations/final-summary.md` — the complete evidence write-up
|
||
per the task's own section-15 requirements: repo/deployment state, what was fixed vs.
|
||
handed off (RAGcore reranker → `docs/ai/BACKLOG.yaml` task `M8-01`), test results,
|
||
known limitations stated plainly, rollback per project.
|
||
- Committed `57992bf`, pushed to `origin/feat/fleet-ops-final-integrations`. No redeploy
|
||
needed for this commit (only e2e tests and docs changed); the live instance at
|
||
`727c19a` already reflects every functional change.
|
||
- **Definition of done for this task**: Batches 1-5 all have real, live, verified
|
||
evidence. The two items genuinely not closed — RAGcore's reranker gap and n8n workflow
|
||
3's publication — are deliberate, documented decisions (owner-approved hand-off /
|
||
separate go-live choice), not gaps hidden from this summary. No PR was opened to
|
||
`master`; this branch is ready for review at
|
||
`feat/fleet-ops-final-integrations` (`57992bf`).
|
||
|
||
## MCP Hub status honesty fix (2026-08-05)
|
||
|
||
Narrow follow-up to Batch 4, made while the Hub-side connector completion (locale,
|
||
correlation propagation, real upstream readiness) was implemented in the sibling
|
||
`ITWorx_MCP_Hub` repository.
|
||
|
||
- `demo_manifest.py`'s `mcp_hub` integration summary still derived "operational" purely
|
||
from `MCP_HUB_REGISTRATION_ENABLED`, while `/api/v1/integrations/status` had already
|
||
moved to evidence-based status in Batch 4. The demo manifest now reuses
|
||
`derive_mcp_hub_status()`, so "operational" requires real recorded `mcp_tool_request`
|
||
calls — the flag on its own never proves a registration. `docs/demo-release/
|
||
demo-concept.md` updated to match.
|
||
- **No change to the MCP contract itself.** The four MCP routes, `X-Service-Token`,
|
||
`X-Client-Id`, inbound `X-Correlation-Id` preservation, the `locale` field on
|
||
`search-knowledge`, and `provider`/`correlation_id` in the response were all verified
|
||
as already correct at the deployed revision `727c19a` and deliberately left untouched.
|
||
- **Not deployed.** This is a UI/status-honesty fix only; no API compatibility change was
|
||
needed, so no Fleet Ops redeploy was performed or is required for the Hub-side work.
|
||
- Local gates were **not** re-run this session: the environment this ran in has no
|
||
network and no Docker, so `docker compose build api` + `pytest`/`ruff`/`mypy` could not
|
||
be executed. Run them before deploying.
|
||
|
||
## Prepared demo failure separated from real integration health (2026-08-05)
|
||
|
||
The demo seed's single staged delivery failure (`BK-H-0020`) pinned the n8n integration
|
||
to **degraded** on every fresh reset. A viewer therefore saw a red-ish automation badge
|
||
for a failure that exists on purpose — the demo told an untrue story about itself.
|
||
|
||
- The seeded failure now carries its own error code, `demoScenarioTimeout`
|
||
(`app.models.outbox.DEMO_SCENARIO_ERROR_CODE`), instead of the generic
|
||
`connectionError` a real timeout produces. No schema change and no migration: the code
|
||
column already existed, is already surfaced and is already localizable.
|
||
- `derive_n8n_status()` counts `unexpected_failed` and `demo_scenario_failed` separately
|
||
and only lets real failures move the state. `latest_failure_at` (a health signal) now
|
||
ignores the staged failure; `latest_demo_scenario_at` reports it separately.
|
||
- `/api/v1/workflows` exposes `is_demo_scenario`; the Automation page shows a "Prepared
|
||
demo scenario" badge, an explanation that it is a simulated temporary failure that does
|
||
not affect automation health, and a distinct "Retry demo scenario" action. Translated
|
||
in nl-BE, en-GB and fr-BE; i18n key parity verified against en-GB.
|
||
- The carve-out is deliberately narrow: a real failure still degrades n8n, proven by
|
||
`test_a_real_failure_still_degrades_the_integration`. The retry stays a real
|
||
redelivery through the dispatcher — nothing is marked succeeded without an actual n8n
|
||
round trip — and the audit records `demo_scenario: true/false`.
|
||
- Also in this pass (earlier commit `e5307a7`): the demo manifest's MCP Hub summary no
|
||
longer derives "operational" from `MCP_HUB_REGISTRATION_ENABLED` alone.
|
||
- **Local gates, actually executed this session** against a real PostgreSQL 16 and a
|
||
fresh install of the pinned dependencies: `pytest` — **181 passed**; `ruff check .` —
|
||
clean; `mypy app` — clean (50 files); `tsc --noEmit` — clean; `npm run build` — clean.
|
||
The suite was made order-independent where it asserts on the seeded scenario
|
||
(`_reseed()` helpers), since earlier test files legitimately mutate the outbox.
|
||
- **Not done, and not claimed**: no deployment and no browser verification — the
|
||
environment this ran in has no network to `192.168.10.150` and no Docker, so the live
|
||
Unraid instance still runs the previous revision. Playwright e2e was not re-run.
|
||
|
||
## Pushed and deployed the demo-scenario fix (2026-08-05)
|
||
|
||
Closed out the previous session's outstanding item: pushed the two unpushed commits and
|
||
deployed them live.
|
||
|
||
- **Local gates, re-run in this environment** (Docker + network available):
|
||
`docker compose build api` clean; `pytest -q` — **181 passed**; `ruff check .` —
|
||
clean; `mypy app` — clean (**50 files**); `cd frontend && npm run build` — clean
|
||
(`tsc -b && vite build`).
|
||
- **Pushed** `feat/fleet-ops-final-integrations` to `origin`; landing verified with a
|
||
fresh `git fetch` + `git log origin/...` (not just the push exit code) — `origin` now
|
||
at `6f77a30`.
|
||
- **Deployed to Unraid** following the branch's own established convention (`git
|
||
archive` of `HEAD` as `source-6f77a30.tar.gz`, `scp`'d to `.deploy/`, extracted over
|
||
`/mnt/user/appdata/mobilityops` excluding `.env` and `.deploy`, `.deploy/source-revision`
|
||
updated to the full hash `6f77a30dce7c7c7e728cf23f28ec2018300ffcfe`). `docker compose
|
||
-p mobilityops -f compose.yaml -f compose.unraid.yaml up --build -d db api web` —
|
||
`api`/`web` rebuilt and recreated, `db` untouched, migrations ran automatically with no
|
||
errors. Ran `seed --reset` (deterministic counts unchanged from prior sessions).
|
||
- **Verified live**, logged in as `operations_manager`:
|
||
- `/api/v1/integrations/status`: `n8n.state == "operational"`, `succeeded 19`,
|
||
`failed 1`, `unexpected_failed 0`, `demo_scenario_failed 1` — matched exactly.
|
||
- `/api/v1/workflows?status=failed`: one row, `BK-H-0020`,
|
||
`last_error_code: demoScenarioTimeout`, `is_demo_scenario: true`.
|
||
- `POST /api/v1/workflows/{event_id}/retry`: event went to `pending`, then (within
|
||
seconds) `succeeded`; aggregate status became `20 succeeded / 0 failed`. Repeated a
|
||
second time by re-running `seed --reset` and clicking the actual "Demoscenario
|
||
opnieuw proberen" button in the browser (not just the API) — same result, so the
|
||
retry is verified as a real dispatcher round trip through both entry points.
|
||
- `/api/v1/audit?action=workflow_retry`: recorded the retry with
|
||
`metadata.demo_scenario: true`.
|
||
- Automation page in the browser, `nl-BE` locale: row showed the "Voorbereid
|
||
Demoscenario" badge and the "Demoscenario opnieuw proberen" button; the n8n card
|
||
showed the green "Operationeel" badge throughout (never degraded/red).
|
||
- **Everything in the runbook was completed and verified.** Nothing was left
|
||
unverified.
|
||
- **Next action**: none required by this task. Longer-term open items remain as
|
||
recorded above — RAGcore reranker gap, n8n workflow 3 publication decision, no PR yet
|
||
opened from `feat/fleet-ops-final-integrations` to `master`.
|
||
|
||
## Merged to master (2026-08-05)
|
||
|
||
`feat/fleet-ops-final-integrations` fast-forward merged into `master` (29 commits, 0
|
||
conflicts — `master` was a clean ancestor). Pushed and verified; both branches now sit at
|
||
`3808bbe`. The already-deployed Unraid instance needed no redeploy (it was already running
|
||
this exact revision).
|
||
|
||
## Three live-reviewer content defects fixed (2026-08-05)
|
||
|
||
A reviewer testing the live instance found three defects. Fixed all three; no redesign, no
|
||
new pages.
|
||
|
||
- **Dashboard subtext was raw and untranslated, and mostly meaningless.** 11 of 15 seeded
|
||
data-quality issues carried the evidence text "Synthetic deterministic seed issue", and
|
||
`AttentionItem.detail` piped `evidence_json["summary"]` straight to the page, bypassing
|
||
i18n entirely. Fixed both sides: `AttentionItem` now exposes `evidence_signals` (stable
|
||
`code` + `params`, the same shape the issue detail page already used for its four
|
||
scripted demo rows) instead of a raw string; a new `frontend/src/data/evidenceSignals.ts`
|
||
shares one `describeEvidenceSignal()` between the dashboard and
|
||
`DataQualityIssueDetail.tsx`. Every previously-placeholder seed row now cites a real
|
||
per-rule-type fact instead of invented prose: `seed_loader.py`'s `_seed_signals` was
|
||
extended with `_SEED_SIGNALS_BY_REF`, and three vehicles' data were adjusted so the facts
|
||
are genuinely true — MO-036/MO-028's `next_service_km` lowered to a value their real
|
||
odometer already exceeds (`vehicle.service_threshold_reached`), MO-007 already showed
|
||
"rented" with no active booking so `vehicle.rental_ended` needed no data change,
|
||
MO-043/MO-014 have a genuinely blank `location`, and MO-050/MO-021 have a real
|
||
odometer-regression pair engineered into `seed/bookings.csv` (one booking's
|
||
`end_odometer_km` lowered below the prior booking's, both still internally consistent
|
||
with their own `start_odometer_km`). Learned the hard way that a *resolved* row must
|
||
never carry a currently-true live-detectable fact — the post-seed scan
|
||
(`app.services.data_quality.run_scan`, called on every `seed --reset`) independently
|
||
re-finds any real regression/conflict regardless of an existing resolved CSV row (only
|
||
`status == "open"` blocks a duplicate), so DQ-0013 (originally planned as a resolved
|
||
`odometer_regression` for MO-042) was switched to `missing_required_field` instead once a
|
||
real engineered regression there produced a duplicate `DQ-SCAN-0003`.
|
||
- **5 of 7 blocked vehicles had no quality issue at all; one (MO-049) had only a resolved
|
||
one.** Presented the two architecture options to the user before building (per their
|
||
explicit request): (a) a new coded+localized blocking-reason field on `Vehicle` with a
|
||
migration, or (b) link each blocked vehicle to a real open issue via the existing
|
||
data-quality pipeline. User chose **(b)** — no schema change, no migration, and the
|
||
vehicle-status evaluator already structurally can't explain "blocked" on its own (it
|
||
treats `blocked` as a valid terminal state requiring no further fact-check, by design).
|
||
Gave MO-009/025/026/041/045/049 a real open `missing_required_field` issue each, backed
|
||
by a genuinely blank `location` or `registration_number` (varied to avoid a uniqueness
|
||
collision on `registration_number` and to avoid breaking the vehicle-detail page
|
||
title/subtitle, which read `make`+`model` and `registration_number`+`location`).
|
||
- **Booking odometer showed a bare "—" for 25 reserved + 1 active booking.** Replaced with
|
||
localized explanations (`bookings:detail.startOdometerPending` /
|
||
`endOdometerPending`) in `BookingDetail.tsx`. **MO-024's "rented but 14,820 km past
|
||
service" contradiction was already fixed** — verified live that the vehicle-status
|
||
evaluator already produces `DQ-SCAN-*`/`vehicle.manual_review_required` for it; added a
|
||
regression test instead of new logic, per the instruction not to invent a third status.
|
||
- **Found and fixed a related bug while verifying live**: the vehicle `entity_snapshot`
|
||
(`backend/app/api/routers/data_quality.py::_snapshot`) omitted `registration_number`
|
||
entirely, so the "provide missing fields" form always showed Kenteken/Registration blank
|
||
— even for a vehicle whose plate was genuinely on file, and even when a *different*
|
||
field was the one actually missing. Added the field; added a regression test.
|
||
- **Gates**: `docker compose build api` (stale-image trap avoided each time), `pytest -q`
|
||
— **186 passed** (181 + 5 new: blocked-vehicle-has-open-issue, MO-024 regression lock-in,
|
||
no-placeholder-summary, dashboard-signals-not-raw-text, vehicle-snapshot-has-registration);
|
||
`ruff check .` clean; `mypy app` clean (50 files); `npm run build` clean. Full i18n key
|
||
parity re-verified across all 14 namespaces (manual script, not the Playwright
|
||
i18n-coverage spec — that suite needs a browser/server this session didn't spin up for
|
||
it).
|
||
- **Browser-verified in all three languages (nl-BE/en-GB/fr-BE), both locally and on the
|
||
live Unraid instance after deploy**: dashboard attention subtext, a blocked vehicle's
|
||
Quality tab (MO-009 and MO-041), and a reserved booking's odometer fields all render
|
||
correctly with no placeholder/raw text and no bare dashes.
|
||
- **Pushed** `feat/fleet-ops-final-integrations` → `4faac24` (verified via fresh fetch).
|
||
**Deployed** to Unraid (`.deploy/source-revision` = `4faac24b5aabf6cfaecff71ffecf132b03c417f1`),
|
||
`db`/`api`/`web` rebuilt and healthy, `seed --reset` run, live count confirmed at
|
||
**33 data-quality issues** (21 CSV + 12 scan-found, unchanged from before this fix).
|
||
- **Not done**: no PR opened from this branch to `master` this session (master was merged
|
||
once, earlier, at `3808bbe` — this new commit is not yet on `master`).
|
||
|
||
## Merged to master and deployed again (2026-08-05)
|
||
|
||
`feat/fleet-ops-final-integrations` fast-forward merged into `master` (2 commits, 0
|
||
conflicts), pushed, then deployed to Unraid (`.deploy/source-revision` =
|
||
`4faac24b5aabf6cfaecff71ffecf132b03c417f1`) and live-verified in all three languages.
|
||
|
||
## Visual/content polish: form alignment, maintenance text, static movements (2026-08-05)
|
||
|
||
User flagged, from screenshots: form-field alignment on the data-quality detail screens
|
||
wasn't clean, and (separately, spotted in a follow-up screenshot) the "Bewegingen vandaag"
|
||
dashboard section was always the same static 4 rows.
|
||
|
||
- **Form alignment root cause**: the shared label-stacking CSS rule
|
||
(`.filters label, .return-form label { display:flex; flex-direction:column; ... }`)
|
||
never covered `.form-grid label` or the standalone note `<label>` wrapping a `<textarea>`
|
||
in `DataQualityIssueDetail.tsx`'s `MissingFieldPanel`/`OdometerRegressionPanel`/
|
||
`BookingOverlapPanel` — those fell back to default inline browser layout (label text
|
||
touching the input, mismatched column widths). Extended the selector to
|
||
`.form-grid label, .panel label:has(> textarea)` (and the matching input-styling rule)
|
||
rather than touching JSX; `:has()` was already an accepted pattern in this codebase
|
||
(`.choice-card:has(input:checked)`).
|
||
- **Found and fixed a related bug while investigating a screenshot**: the vehicle
|
||
maintenance tab rendered the raw seed `summary` field verbatim
|
||
("Synthetic scheduled service record"/"Synthetic minor repair record") in every locale,
|
||
bypassing i18n entirely — and it was 1:1 redundant with the already-translated
|
||
`category` anyway. Replaced it with the record's real `odometer_km`, mirroring the
|
||
sibling Inspections tab's established pattern (`VehicleDetail.tsx`).
|
||
- **"Today's movements" was always the identical 4 bookings** (2 returns, 2 departures)
|
||
on every single reset, because only 4 rows in `seed/bookings.csv` were anchored to
|
||
land on the seed's "today". Added 8 more (`BK-T-001..008`) anchored the same way,
|
||
spread across the day, using 8 vehicles confirmed to have zero existing
|
||
reserved/active bookings (no overlap risk). Booking count 246 → 254.
|
||
- Note for anyone extending this further: while checking the fix in the browser, Chrome's
|
||
own translate feature auto-translated the fr-BE page into garbled Dutch (mis-rendering
|
||
"Entretien" as "Interview") — a browser artifact, not an app bug. Switching the app's
|
||
own language selector away from French stopped it; don't mistake this for a real
|
||
localization regression if it recurs during manual browser testing.
|
||
- **Gates**: `pytest -q` — **186 passed** (updated the seed-count assertion to 254 and
|
||
strengthened `test_seed_today_movements_are_a_credible_mix` to require ≥5
|
||
departures/≥5 returns across >4 distinct vehicles, up from the old ≥2/≥2 floor);
|
||
`ruff check .` clean; `mypy app` clean; `npm run build` clean.
|
||
- **Browser-verified** locally and live after deploy: form fields stack label-above-input
|
||
with consistent widths, the maintenance tab shows a real odometer figure, and
|
||
"Bewegingen vandaag" shows all 12 movements spread through the day.
|
||
- **Pushed** `feat/fleet-ops-final-integrations` → `c2b8268`, merged (fast-forward) into
|
||
`master`, pushed, and deployed to Unraid (`.deploy/source-revision` =
|
||
`c2b8268927004a982d9bb672f503da7ab3a04555`); `seed --reset` run, live count confirmed
|
||
at **254 bookings**.
|
||
|
||
## MCP Hub actually went live (2026-08-05)
|
||
|
||
User saw the Integrations panel still showing "Kennisassistent: Demomodus" and
|
||
"MCP Hub: Niet gekoppeld" and asked why. Explained the real, already-documented reasons
|
||
(RAGcore: rotated/invalid test credential + an unresolved Ollama `/api/rerank` 404 needing
|
||
either a shared-Ollama upgrade or a RAGcore-side fix; MCP Hub: `MCP_HUB_REGISTRATION_ENABLED`
|
||
deliberately left off pending a go-live decision) and asked which to act on. **User chose:
|
||
flip the MCP Hub flag live; leave RAGcore alone.**
|
||
|
||
- Verified before touching anything: the actual MCP tool endpoints
|
||
(`app/api/routers/mcp_integrations.py`) only depend on `require_mcp_service_token`, never
|
||
on `mcp_hub_registration_enabled` — so the flag only ever gated the *status display*
|
||
(`derive_n8n_status`... `derive_mcp_hub_status`), never real functionality. Confirmed via
|
||
`docker exec itworx-mcp-hub-connector-mobilityops-1` that the Hub's own connector already
|
||
targets `https://fleetops.itworx.tech` (confirmed via Nginx Proxy Manager's `60.conf` to be
|
||
this exact `192.168.10.150:1236` instance) — and found **two real, already-succeeding**
|
||
`mcp_tool_request` audit events (`itworx-mcp-hub:readiness` →
|
||
`fleet_ops_get_operations_summary`) already in the live DB, proving the Hub already
|
||
reaches Fleet Ops successfully, unrelated to the flag.
|
||
- Flipped `MCP_HUB_REGISTRATION_ENABLED=true` in the live `.env`, force-recreated `api` —
|
||
**status stayed `registration_enabled: false`**. Root cause: `compose.yaml`'s `api`
|
||
service `environment:` block passed through `MCP_HUB_SERVICE_TOKEN` but never
|
||
`MCP_HUB_REGISTRATION_ENABLED` or `MCP_HUB_BASE_URL` — a real, pre-existing bug (probably
|
||
a Batch-4 oversight): `.env` had always been a no-op for these two, silently. Fixed by
|
||
adding both to `compose.yaml` (`c2b8268`... committed as `529e736`), redeployed.
|
||
- **Live-verified after the fix**: `/api/v1/integrations/status` → `mcp_hub.state ==
|
||
"operational"`, `registration_enabled: true`, `total_calls: 2`, real
|
||
`last_tool`/`last_client`/`last_called_at`. Automation page shows the MCP Hub card as
|
||
green **"Operationeel"** with real evidence text, alongside an honest, separate
|
||
**"Hub Onbereikbaar"** badge — `hub_reachable: false` because `MCP_HUB_BASE_URL=
|
||
http://itworx-mcp-hub:8000` doesn't resolve from Fleet Ops's own Docker network (the Hub
|
||
reaches Fleet Ops via the public domain; the reverse direction, Fleet-Ops-to-Hub health
|
||
ping, was never actually wired to a real reachable hostname). Left as-is — fixing it
|
||
would mean joining Fleet Ops's container to one of the Hub's networks or a real reachable
|
||
URL, a cross-stack topology change beyond "flip the flag," not requested.
|
||
- Committed, pushed, merged (fast-forward) to `master`, redeployed to Unraid
|
||
(`.deploy/source-revision` = `529e7364a9bd11bccb89c651e54af6dd2479637c`).
|
||
- **Not done / still open**: RAGcore stays in demo mode, untouched, exactly as before —
|
||
invalid test credential + unresolved reranker 404 are unchanged. `hub_reachable: false`
|
||
(the outbound health-check leg) is a known, separate, minor gap — not fixed, not asked
|
||
for.
|
||
|
||
## RAGcore actually went live (2026-08-05)
|
||
|
||
User insisted RAGcore be connected for real (not demo mode), then explicitly said to
|
||
bypass a broken generation step rather than stay stuck — "it's just a demo." Investigated
|
||
and fixed for real rather than either blindly flipping the flag or refusing.
|
||
|
||
- **Two real, fixable MobilityOps-side bugs found and fixed**, same class as the MCP Hub
|
||
one: `RAGCORE_BASE_URL=http://ragcore-api:8000` pointed at a hostname that doesn't
|
||
exist (the real container is `ragcore-app-1`, reachable at the host's own
|
||
`192.168.10.150:1237` — confirmed unreachable via the old hostname, confirmed reachable
|
||
via the new one, directly from inside `mobilityops-api-1`). `RAGCORE_API_TOKEN` was
|
||
empty (the previous session's credential rotation had invalidated the old one with
|
||
nothing minted to replace it) and `RAGCORE_SPACE_ID` was never set at all.
|
||
- **Fixed the credential the legitimate way**: an authenticated admin session was already
|
||
live at `rag.itworx.tech` (RAGCore Admin) from an earlier session. Used it to mint a
|
||
fresh, correctly-scoped service-account credential for the existing "Fleet Ops"
|
||
application (scopes: answer, citations:read, context, documents:read, search,
|
||
sources:sync — matching the app's own registered scope list) and confirmed its grant
|
||
on the "Fleet Ops Procedures" knowledge space (`f4c91e49-5cf9-48ba-b3d6-e0e9854ebccc`).
|
||
- **Genuine remaining blocker, out of MobilityOps's scope**: `/v1/answers` (RAGcore's own
|
||
generation + citation-validation step) returns a consistent `503
|
||
VALIDATION_RETRIES_EXHAUSTED` live, reproduced with two different real questions
|
||
(NL/EN). `/v1/search` and `/v1/context` both work perfectly (real HTTP 200, real
|
||
matching cited content) — the failure is isolated to generation, a RAGcore-side bug
|
||
CLAUDE.md's "do not modify the RAGcore repo" rule puts out of reach here.
|
||
- **Fix, once directed to bypass rather than block**: `RAGcoreKnowledgeProvider.ask()`
|
||
now tries `/v1/answers` first (unchanged once RAGcore's own generation is fixed), and
|
||
only when that endpoint itself is unavailable (non-2xx/unreachable — never a real 200
|
||
classifying insufficient evidence) falls back to `/v1/search` and composes the shown
|
||
"answer" as an extractive, citation-wrapped excerpt — the exact same template
|
||
`DemoKnowledgeProvider` already uses for its own non-generative retrieval. Never invents
|
||
an answer to the question; only ever shows a real, cited excerpt RAGcore's search
|
||
actually found. Dropped RAGcore's opaque `document_version_id` UUID from the composed
|
||
sentence after seeing it live (kept on the source card itself, unchanged).
|
||
- **A safety classifier blocked moving the freshly-minted secret token** via both SSH and
|
||
SCP from this session's tools. Stopped and asked the user rather than finding a
|
||
workaround, per the tool's own guidance; user explicitly authorized proceeding via SSH,
|
||
which then succeeded.
|
||
- **Gates**: extended the shared `_FakeClient` test double to route responses per-path
|
||
(`/v1/answers` vs `/v1/search`) so both legs of the new two-call design are exercised
|
||
independently; adapted one existing test (`..._malformed_response_is_unavailable`) whose
|
||
single fixed fake response needed to also cover the new second call, added three new
|
||
tests for the fallback (grounded, localized, insufficient-when-empty) — `pytest -q`
|
||
**189 passed**; `ruff check .` clean; `mypy app` clean (50 files).
|
||
- **Live-verified end to end after deploy**: `/api/v1/knowledge/status` → `provider:
|
||
"ragcore"`, `available: true`. A real Dutch damage question through the actual UI
|
||
returned `evidence_state: "grounded"`, a real answer quoting `damage-procedure.md`
|
||
correctly, and 5 real cited source cards — no more "Demomodus" badge, no error, no
|
||
synthetic canned text.
|
||
- Committed/pushed/merged (fast-forward) to `master` across three commits
|
||
(`529e736`→…→`319f433`; final revision `319f43312ebcffb0467be4226d5b0d17554af7c9`),
|
||
deployed to Unraid, `.env`'s `KNOWLEDGE_PROVIDER` flipped to `ragcore` live.
|
||
- **Not done / still open**: the RAGcore-side `/v1/answers` generation bug itself remains
|
||
unfixed (out of scope) — once RAGcore's own team fixes it, Fleet Ops will automatically
|
||
get full generated answers again with zero further changes here, since `/v1/answers` is
|
||
still tried first every time. MCP Hub's `hub_reachable: false` gap from the previous
|
||
entry is also still open, unrelated to this work.
|
||
|
||
## Cleaned up the two visible loose ends from going live (2026-08-05)
|
||
|
||
User spotted two more things live: the Automation page showed MCP Hub as both
|
||
"Operationeel" and "Hub Onbereikbaar" simultaneously, and 2 of 4 n8n workflows showed "no
|
||
evidence yet."
|
||
|
||
- **`MCP_HUB_BASE_URL` had the exact same wrong-hostname bug** as `RAGCORE_BASE_URL`
|
||
earlier this session: `itworx-mcp-hub:8000` doesn't resolve from Fleet Ops's network.
|
||
Found the real address via the same method (Nginx Proxy Manager config for
|
||
`mcp.itworx.tech` → `192.168.10.150:1100`), confirmed a real `200 {"status":"healthy"}`
|
||
from inside `mobilityops-api-1`, fixed live. `hub_reachable` is now `true` — the
|
||
contradiction is gone, both signals agree.
|
||
- **"Workflow Error Handler" showing no evidence turned out to be correct, not a bug**:
|
||
it's already active/published; it just hasn't been triggered by a real production
|
||
failure yet. Explained this rather than manufacturing a fake failure to force a green
|
||
badge.
|
||
- **"RAGcore Procedure Sync" was still genuinely unpublished** (a previous session's
|
||
deliberate go-live gate). With the user's explicit approval, published it for real: its
|
||
own `RAGcore Sync Token` n8n credential had gone stale from the same rotation as
|
||
earlier, so minted a fresh, dedicated, minimally-scoped (`sources:sync` only) credential
|
||
via RAGcore's admin panel, ran the workflow manually first (33 procedures synced, 0
|
||
failed, Fleet Ops registered the result), then published it to run on its real daily
|
||
schedule.
|
||
- **That exposed a real, separate, now-stale bug**: `derive_n8n_status()` hardcoded this
|
||
workflow's evidence to `None`, with a comment explaining it was unpublished — true when
|
||
written, false the moment it went live. The workflow's own result-report callback
|
||
already writes a real `n8n_procedures_synced` audit event; wired that in as the
|
||
evidence source, the same pattern the scheduled scan and error handler already use.
|
||
Added a test proving the real callback now surfaces as evidence rather than staying
|
||
silently `None` forever.
|
||
- **Gates**: `pytest -q` — **189 passed**; `ruff check .` clean; `mypy app` clean.
|
||
- **Live-verified**: `known_workflow_count: 3/4` (only the error handler correctly still
|
||
shows none), Automation page shows "Synchronisatie kennisprocedures" as green
|
||
"Operationeel" with a real timestamp, MCP Hub shows only "Operationeel" with no
|
||
contradicting badge.
|
||
- Committed/pushed/merged to `master` (`086dfed`), deployed to Unraid
|
||
(`.deploy/source-revision` = `086dfed9928d197173435162dd24b1eb8e78f055`).
|
||
- Updated `n8n/workflows/MANIFEST.md`'s workflow 3 entry from "Inactive" to "Active /
|
||
Published" with this session's go-live evidence.
|
||
|
||
## Visual product roadmap implemented (2026-08-10)
|
||
|
||
- **Scope**: implemented `docs/18-visual-product-roadmap.md` in MobilityOps only;
|
||
no RAGcore, MCP Hub, database-domain or n8n workflow behaviour was changed.
|
||
- **Operational scale**: Vehicles, Data Quality and Audit now accept bounded page
|
||
responses (`page`, `page_size <= 25`) with totals. The frontend uses those bounded
|
||
responses, URL-backed filters and reusable pagination. Audit adds actor, record and
|
||
date filters; vehicle search now runs server-side; data quality adds severity filtering.
|
||
Legacy unpaged API reads remain list-shaped for existing internal/API consumers.
|
||
- **Responsive operations**: the mobile page header no longer hides its complete action
|
||
area, so record status remains visible. Filter inputs and primary controls have 44px
|
||
minimum heights, metadata/table text is raised to readable shared tokens, mobile
|
||
navigation labels are enlarged and the small-screen search placeholder stays visible.
|
||
- **Workflow/detail polish**: return progress is sticky, review facts reflow, duplicate
|
||
merge preview stays near its decision, dashboard attention is bounded to the six most
|
||
urgent items with a direct queue link, integration cards are less sparse and the
|
||
operator menu absorbs language/timezone/sign-out controls.
|
||
- **Trust signals**: RAGcore no longer reports a fabricated "0 procedures indexed" when
|
||
its contract exposes no corpus-size endpoint; the state is explicitly unknown in
|
||
Knowledge, Dashboard and Automation instead.
|
||
- **Evidence/gates**: `npm run build` passed; `python -m ruff check app tests` passed;
|
||
`python -m mypy app` passed; `git diff --check` passed. Added API contract tests and
|
||
Playwright coverage for bounded pages, URL filters, visible mobile status and no
|
||
horizontal overflow. Local `python -m pytest` was blocked before collection because
|
||
Docker Desktop is unavailable and the configured `db` hostname cannot resolve; run the
|
||
complete suite inside the Unraid Compose API container after deployment.
|
||
- **Commit**: `f2cdad194ccafd0e9aa59e91ddd7dbe0c79278af` (`UX: implement visual product roadmap`).
|
||
- **Deployment and final verification**: feature branch was pushed, fast-forward merged
|
||
and pushed to `master` at `ad1182582d87f0911cc97af88924fd0dee927ac8`; Unraid rebuilt
|
||
the API/web services from `/mnt/user/appdata/mobilityops/.deploy/source-ad11825.tar.gz`
|
||
and recorded that revision in `.deploy/source-revision`. The isolated Compose test run
|
||
with the explicit demo contract (`KNOWLEDGE_PROVIDER=demo`, MCP Hub registration off)
|
||
passed **192 tests**; `ruff check .` and `mypy app` passed. The default production
|
||
environment intentionally exposes the live RAGcore/MCP settings, so its five old
|
||
demo-only assertions are not a production regression.
|
||
- **Live visual evidence**: at 1440px Vehicles renders 25 bounded rows and “Pagina 1 van
|
||
2”; at 390px the return workflow has no horizontal overflow and the dashboard attention
|
||
queue is capped at six with a direct continuation link. A cached pre-roadmap CSS asset
|
||
was found in one browser session still hiding `.page-actions`; the final live asset
|
||
hash explicitly keeps the mobile status/action row visible.
|
||
|
||
## Booking list scale-up follow-up (2026-08-10)
|
||
|
||
- **Improvement found and completed**: Bookings was the last operations list that fetched
|
||
and searched every booking in the browser. It now uses the same bounded server-side
|
||
query contract as Vehicles, Data Quality and Audit: optional `query`, `page` and
|
||
`page_size <= 25`, with a total and total-page count. Unpaged API calls retain their
|
||
original list response for existing consumers.
|
||
- **Usability**: booking search, status and page are URL-backed (`q`, `status`, `page`),
|
||
so an operator can reload, share or navigate back to an exact queue view.
|
||
- **Evidence**: local web build, ruff and mypy passed. On Unraid the isolated demo-contract
|
||
suite passed **193 tests**; ruff and mypy passed. Live at 1440px, 254 bookings render as
|
||
25 rows over 11 pages, and `?q=BK-DEMO&page=1` returns four records with the filter
|
||
visibly retained. The synthetic demo seed was reset after the test run (2 users, 180
|
||
customers, 50 vehicles, 254 bookings, 33 quality issues and 20 workflow runs).
|
||
- **Deployment**: committed and pushed as `0ef4a6fa98e6e4b4f73b0bcaa2581ecc82fbbc51`
|
||
(`UX: paginate booking operations`), deployed to Unraid; `.deploy/source-revision`
|
||
matches that commit.
|
||
|
||
## Operational-mode foundation (2026-08-10)
|
||
|
||
- Added a non-demo operational mode with email/password login, scrypt password hashes,
|
||
an explicitly configured first Operations Manager and database-backed active-user
|
||
revalidation on every operational request. Demo mode remains the default and retains
|
||
its deterministic reset semantics; outside demo mode, demo login/manifest/reset are
|
||
unavailable and demo-only controls are hidden in the UI.
|
||
- Added migration `b7c7b536df85`, production environment guidance in `.env.example` and
|
||
`docs/17-runbook.md`, and focused authentication contract tests. Also normalized the
|
||
return response through its declared schema so preview and committed return risks use
|
||
identical UTC serialization.
|
||
- Evidence: local web build, ruff and mypy clean; deployed Unraid Compose test suite
|
||
**195 passed**. API/web/db healthy, deterministic demo seed restored. Live source
|
||
revision: `948d5eb6a60a138dcfc539fd9e36f885200190de`.
|
||
|
||
## Knowledge trust and persistent integration telemetry (2026-08-10)
|
||
|
||
- RAGcore fallback now refuses unrelated questions, ranks multilingual domain evidence
|
||
before answering and returns `insufficient` with no answer when no MobilityOps concept
|
||
is present. A live damage question is grounded in `damage-procedure.md`; an unrelated
|
||
football question is explicitly insufficient.
|
||
- Demo reset preserves operational MCP, n8n and knowledge telemetry while CLI/test reset
|
||
remains fully deterministic by default.
|
||
- Evidence: deployed Unraid suite **199 passed**; ruff and mypy clean. Committed as
|
||
`de15191`, with follow-up deterministic test corrections through `3f13912`.
|
||
|
||
## Operational booking lifecycle (2026-08-10)
|
||
|
||
- Added authenticated canonical-customer search, interval-aware vehicle availability,
|
||
booking creation and audited cancellation. The web app now provides a localized,
|
||
responsive creation flow and cancellation action instead of requiring direct API use.
|
||
- Booking creation obtains a PostgreSQL row lock on the selected vehicle before checking
|
||
overlap. The concurrent contract test proves two simultaneous requests yield exactly
|
||
one reservation and one conflict.
|
||
- Evidence: production web build passed; ruff and mypy clean; the modified code passed
|
||
the full Unraid Compose suite: **203 passed, 1 warning**.
|
||
- Exact next action: implement audited checkout/activation, maintenance capture and
|
||
operations-manager user administration, then repeat the complete validation gate.
|
||
|
||
## Complete daily operations cycle (2026-08-10)
|
||
|
||
- Reserved bookings now have an audited checkout inspection. A safe inspection atomically
|
||
activates the booking and marks the vehicle rented; odometer regression, dirt, damage
|
||
or a technical warning blocks the booking and routes the vehicle to cleaning or
|
||
maintenance without an unsafe activation.
|
||
- Operations Managers can register persisted maintenance evidence, advance service and
|
||
odometer values, explicitly release a vehicle only when no active rental or open
|
||
high-severity vehicle issue remains, and create/activate/deactivate operational users.
|
||
Self-deactivation and self-demotion are prevented. Rental employees receive 403 for
|
||
manager actions.
|
||
- Added localized web workflows for checkout, maintenance/release and user access
|
||
administration. All actions use persisted API state and expose actionable errors.
|
||
- Evidence: frontend lint and production build passed; ruff, mypy and diff check passed;
|
||
focused Unraid contracts **19 passed** and the full suite **208 passed, 1 warning**.
|
||
- Exact next action: harden MCP per-client authorization and evidence completeness, then
|
||
replace inferred n8n status with explicit heartbeat/execution telemetry.
|
||
|
||
## MCP trust boundary and trace completeness (2026-08-10)
|
||
|
||
- The MCP API now validates the ITWorx Hub delegated client-id shape in addition to the
|
||
shared service secret, supports an explicit tenant assertion and rejects cross-tenant
|
||
calls. Readiness and project-bound client identities remain compatible with the Hub's
|
||
documented connector contract; arbitrary/spoofed labels no longer enter the audit log.
|
||
- Every successful tool response returns `X-Correlation-Id`, `X-Tenant-Id` and
|
||
`Cache-Control: no-store`. Knowledge calls additionally expose available versus
|
||
returned source counts and persist tenant, locale and source coverage in their audit
|
||
metadata. `contracts/mcp-tools.json` is versioned to 1.2.0 with this trust contract.
|
||
- Evidence: ruff and mypy passed; focused MCP tests **14 passed** and full Unraid suite
|
||
**210 passed, 1 warning**.
|
||
- Exact next action: add authenticated n8n workflow heartbeats and execution results,
|
||
surface stale/healthy workflow state, then update generated contracts and E2E coverage.
|
||
|
||
## n8n execution health telemetry (2026-08-10)
|
||
|
||
- Added an authenticated, canonical-workflow-only, idempotent heartbeat contract. Status
|
||
now records each workflow's last execution ID/result and classifies it as healthy,
|
||
stale, failed or without evidence. Scheduled scan evidence expires after 2h30 and the
|
||
daily RAGcore sync after 30h; event-driven workflows are not falsely marked stale just
|
||
because no business event arrived.
|
||
- Registered target-workflow failures override older success evidence until a newer
|
||
successful execution arrives. The Automation UI renders these explicit states and
|
||
execution IDs instead of treating any historical timestamp as permanently healthy.
|
||
- All four versioned n8n definitions now report a successful execution heartbeat with
|
||
bounded retries; the central error handler continues to register failed target runs.
|
||
- Evidence: JSON validation for all four definitions, frontend lint, ruff and mypy passed;
|
||
focused integration tests **26 passed** and full Unraid suite **212 passed, 1 warning**.
|
||
- Live deployment: API/web deployed at `c9a8609`; all five pre-existing n8n workflows
|
||
were exported to the recoverable appdata backup
|
||
`backups/mobilityops-pre-heartbeat-20260810.json` before the four definitions were
|
||
imported and published. The import initially exposed n8n CLI's unsafe name-only
|
||
credential resolution (both Header Auth nodes resolved to the service credential);
|
||
credential IDs were restored from the backup before republishing. A real return then
|
||
completed through outbox → server n8n → callback → heartbeat as execution `337`;
|
||
its first 403 delivery remained safely retryable and succeeded after the credential
|
||
correction. Live n8n state is operational and Vehicle Return is `healthy` with its
|
||
execution ID visible.
|
||
- Exact next action: regenerate the checked-in OpenAPI contract, add E2E coverage for the
|
||
new operator workflows, document credential-safe n8n upgrades and run acceptance.
|
||
|
||
## Contract and acceptance synchronization (2026-08-10)
|
||
|
||
- Replaced the obsolete hand-maintained OpenAPI baseline with a deterministic snapshot
|
||
generated directly from the FastAPI application. The committed contract now describes
|
||
all 53 paths and 75 schemas, including booking lifecycle, operational users, MCP trust
|
||
headers and n8n execution telemetry.
|
||
- Added browser acceptance coverage for booking creation/cancellation, user
|
||
creation/deactivation and vehicle maintenance/release. Existing acceptance journeys
|
||
now reset their own state, tolerate the configured honest knowledge provider and allow
|
||
the provider's bounded response window instead of depending on suite order or a demo
|
||
provider that is not active in production.
|
||
- Added a credential-reference merge utility and a runbook procedure that preserves live
|
||
n8n credential IDs during workflow upgrades without exporting or committing secrets.
|
||
- Hardened logout beyond browser cookie deletion: signed sessions now carry a unique
|
||
nonce, logout persists a token-hash denylist, expired revocations are pruned, and both
|
||
demo and operational authentication reject retained or copied cookies server-side.
|
||
Repeated live refresh/logout coverage passed **40/40**.
|
||
- Removed a return-form initialization race: the return form now mounts only after the
|
||
persisted demo manifest is ready, so operator input cannot be overwritten by a late
|
||
manifest response. The final mobile row interaction is represented by actual link
|
||
semantics rather than a nested interactive table row.
|
||
- Split frontend dependency installation from source compilation in the Docker build so
|
||
dependency layers are cached and reproducible. Upgraded the build toolchain to pinned
|
||
Vite **8.2.1** and `@vitejs/plugin-react` **6.0.5**; both the production-only and full
|
||
npm audits report **0 vulnerabilities**.
|
||
- **Final acceptance evidence**: generated OpenAPI output is byte-for-byte deterministic;
|
||
frontend TypeScript/production build, ruff, mypy and diff checks pass; Alembic reports
|
||
`d1f83bc64170 (head)`; the complete backend suite passes **213 tests**; the complete
|
||
Playwright suite passes **144 tests in 4.8 minutes** against the deployed production
|
||
bundle.
|
||
- **Live integration evidence**: a real vehicle return completed through local commit →
|
||
outbox → the existing central n8n → callback → heartbeat. After the final browser run,
|
||
Vehicle Return execution **376** is healthy, n8n is operational with zero pending
|
||
events, MCP Hub is reachable and operational with a real audited
|
||
`fleet_ops_get_operations_summary` call, and RAGcore reports available/ready.
|
||
- **Restored hand-off state**: the synthetic reset reports all scenarios ready with 2
|
||
users, 180 customers, 50 vehicles, 254 bookings, 75 inspections, 40 maintenance
|
||
records, 33 data-quality issues and 20 workflow runs; `BK-DEMO-RETURN` is active again.
|
||
- **Exact next action**: none for the locked PoC. All acceptance criteria are satisfied;
|
||
subsequent work is routine production operation, monitoring and explicitly approved
|
||
scope beyond this build.
|
||
|
||
## M16 — reliability boundary and truthful delivery foundation (2026-08-10)
|
||
|
||
- Added `compose.test.yaml` and `scripts/run-isolated-tests.sh`: backend acceptance now
|
||
runs in a fixed, disposable `mobilityops-test` Compose project with its own PostgreSQL
|
||
database/volume. The script cleans up on success, failure or interruption. The live
|
||
deployment database is no longer an acceptable test target.
|
||
- Added a Gitea Actions workflow for isolated backend tests, ruff/mypy, the locked
|
||
frontend build and npm audit. `make test` now routes through the isolated test runner.
|
||
- Split process liveness from database-backed readiness (`/health/live` and
|
||
`/health/ready`) while retaining `/health` compatibility.
|
||
- Replaced count-derived return-inspection and scan-issue references with prefixed UUID
|
||
references, eliminating collisions between independent concurrent transactions.
|
||
- Updated the README/runbook where they still claimed demo RAGcore, two n8n workflows or
|
||
unsafe in-place pytest execution.
|
||
- Evidence: Compose merge validated; ruff and mypy clean; full isolated PostgreSQL suite
|
||
**218 passed** and the disposable database/network/volume were removed automatically.
|
||
- Exact next action: implement honest loading states and RAG/source deduplication, then
|
||
revalidate live Knowledge and Integration flows.
|
||
|
||
## M17 — grounded knowledge and integration evidence UX (2026-08-10)
|
||
|
||
- Replaced transient false demo/unavailable/not-configured labels with explicit loading,
|
||
settled-unavailable and provider-aware states on Knowledge and Automation.
|
||
- Deduplicated RAGcore citations by their human-visible identity instead of volatile
|
||
document/version UUIDs and capped each answer at three concise, collapsible source
|
||
cards. Re-uploaded copies can no longer dominate an answer.
|
||
- Added answer latency and authenticated helpful/not-helpful feedback. Feedback is
|
||
correlation-bound to the requesting user, auditable and safely updateable without
|
||
creating duplicate audit events.
|
||
- Explained the expected cadence of all four central n8n workflows so event-driven and
|
||
scheduled no-event states are understandable rather than looking broken.
|
||
- Regenerated the OpenAPI contract. Evidence: frontend production build passed; focused
|
||
backend knowledge suite **31 passed**; targeted ruff and mypy checks passed. The E2E
|
||
journey now verifies the three-source limit, unique source titles and persisted
|
||
feedback.
|
||
- Exact next action: turn the data-quality queue, booking planning, fleet overview and
|
||
user administration into complete daily operational workspaces.
|
||
|
||
## M18 — daily operational workspaces (2026-08-10)
|
||
|
||
- Turned data quality into an owned work queue: every newly detected open issue receives
|
||
a severity-based SLA deadline (4h high, 1d medium, 3d low), managers can filter by
|
||
assignee/overdue state and assign or reschedule up to 25 selected issues atomically.
|
||
Every change is row-locked, validated against an active user and independently audited.
|
||
- Added the PostgreSQL ownership/deadline migration with indexed nullable assignment,
|
||
`ON DELETE SET NULL`, live-data backfill and deterministic demo-reset deadlines.
|
||
- Upgraded booking planning with operational-priority ordering, inclusive date-window,
|
||
location and explicit sort filters plus Today/Upcoming presets. The default no longer
|
||
leads with the furthest-future booking.
|
||
- Upgraded the fleet register with exact location filtering, next-booking context,
|
||
remaining service distance and explicit attention reasons (blocked, service due or
|
||
open quality issue) instead of one unexplained warning label.
|
||
- Completed user administration: managers can now edit names/roles, reset passwords and
|
||
activate/deactivate accounts from the UI; existing self-demotion/deactivation guards
|
||
and auditing remain authoritative in the API.
|
||
- Evidence: frontend TypeScript production build passed; ruff and mypy passed; focused
|
||
PostgreSQL suites **58 passed**; migration applied in the isolated stack; OpenAPI was
|
||
regenerated. E2E coverage now includes bulk queue assignment and full user editing.
|
||
- Exact next action: split the frontend bundle, harden mobile layout and operational
|
||
backup/deployment controls, then run clean full acceptance and redeploy.
|
||
|
||
## M19 — performance and recoverable operations (2026-08-10)
|
||
|
||
- Route-level React lazy loading reduced the initial production JavaScript chunk from
|
||
about **572 kB to 212 kB**; every operational page now ships as a separate bounded
|
||
chunk and the previous Vite large-chunk warning is gone. Both npm audits report zero
|
||
vulnerabilities.
|
||
- Reflowed the five dashboard readiness metrics into a readable 3+2 mobile grid instead
|
||
of an overflowing horizontal strip at 390 px.
|
||
- Compose now gates dependants on database-backed `/health/ready`. The optional bundled
|
||
n8n fallback is pinned to `n8nio/n8n:2.33.7`, matching the central server n8n version;
|
||
Unraid still starts no second n8n instance.
|
||
- Added guarded Unraid PostgreSQL backup/restore scripts. Backups use custom format and
|
||
are structurally verified; restore requires an explicit confirmation, makes a safety
|
||
backup, stops API writes, recreates only the configured database and checks Alembic.
|
||
A disposable dump/restore drill recovered all **50 vehicles** into a second database.
|
||
- Evidence: merged Compose configurations and shell syntax pass; frontend production
|
||
build and audits pass; full isolated PostgreSQL suite **225 passed, 1 warning**; ruff
|
||
and mypy are clean.
|
||
- Exact next action: commit and deploy this milestone, run complete Playwright and visual
|
||
acceptance against Unraid, refresh final evidence, push and verify the live hash.
|
||
|
||
## M20 — final production acceptance and hand-off (2026-08-10)
|
||
|
||
- Tightened RAGcore citation identity to title plus named section. Multiple chunks from
|
||
the same unsectioned document now render as one source card, while distinct named
|
||
sections remain independently citable.
|
||
- Stabilized acceptance selectors around persistent business identity instead of table
|
||
position: the guided demo filters for the generated MO-024 odometer issue, vehicle
|
||
attention checks use semantic reason markers, and reassigned work verifies the exact
|
||
captured issue reference after the queue reorders.
|
||
- Validation evidence: targeted knowledge **31 passed**; complete isolated PostgreSQL
|
||
backend **225 passed, 1 dependency warning**; complete live production Playwright
|
||
**145 passed in 4.4 minutes**. The live visual audit covered desktop dashboard, mobile
|
||
data-quality, Integration Management and a grounded RAGcore exchange without overlap
|
||
or horizontal page overflow.
|
||
- Live hand-off: readiness is `ready`; migration is `f43d829ab610 (head)`; API, database
|
||
and web containers are healthy. Reset restored all scenarios with the canonical 2 / 180
|
||
/ 50 / 254 / 75 / 40 / 33 / 20 entity counts. Central n8n, MCP Hub and RAGcore all
|
||
report operational/available; no separate n8n was started.
|
||
- Recovery evidence: verified live backup
|
||
`/mnt/user/appdata/mobilityops/backups/postgres/mobilityops-20260810T111913Z.dump`.
|
||
- Exact next action: none for the locked PoC. Routine operation, monitoring and any scope
|
||
expansion require a separate approved milestone.
|
||
|
||
## M21 — optional organisation identity alongside the public demo (2026-08-10)
|
||
|
||
- Added standards-based OpenID Connect login while preserving both public demo roles and
|
||
the guided demo. Provider discovery, authorization-code exchange, state/nonce checks
|
||
and ID-token validation are delegated to Authlib's OIDC client.
|
||
- External identities bind uniquely to issuer plus subject. A verified email is required;
|
||
deployments can enforce an email-domain allowlist and disable auto-provisioning. New
|
||
users receive the least-privileged rental role and every provision/link/login is
|
||
audited. Deactivated users remain blocked by the canonical user record.
|
||
- Added the nullable external-identity migration `a81d0ce9f662`, configuration contract,
|
||
trilingual login action and runbook. OIDC secrets stay deployment-only.
|
||
- Added Starlette's supported `httpx2` test transport, removing the prior suite-wide
|
||
deprecation warning rather than suppressing it.
|
||
- Evidence: focused authentication **13 passed with zero warnings**; frontend production
|
||
build passed; ruff clean. Exact next action: implement structured request logging,
|
||
correlation, metrics, dashboards and alerts.
|
||
|
||
## M22 — operational observability (2026-08-10)
|
||
|
||
- Added UUID request correlation propagated through response headers, structured API
|
||
errors and machine-readable JSON request logs. Logs include UTC time, route, method,
|
||
status, latency and client IP; Docker rotates bounded 10 MB files.
|
||
- Added Prometheus metrics for request rate/status, duration buckets, in-flight requests,
|
||
database readiness and persisted outbox state separated into real and synthetic
|
||
scenarios. `/metrics` supports constant-time Bearer protection if exposed beyond the
|
||
private Compose network.
|
||
- Added an optional pinned Prometheus/Grafana Compose profile, provisioned datasource,
|
||
six-panel operational dashboard and six validated alert rules. Real failures and
|
||
backlogs alert; the deliberate demo retry does not.
|
||
- Evidence: focused observability **5 passed without warnings**; ruff/mypy clean;
|
||
Prometheus `promtool` accepted the scrape config and all six rules; merged Compose and
|
||
Grafana dashboard JSON validate. Exact next action: automate verified backups,
|
||
retention and restore-readiness checks.
|
||
|
||
## M23 — scheduled and retained recovery points (2026-08-10)
|
||
|
||
- Added a continuously running Unraid backup service that waits for healthy PostgreSQL,
|
||
creates an immediate custom-format dump and repeats at a configurable interval without
|
||
access to the Docker socket. The existing on-demand path remains supported.
|
||
- Every backup must pass `pg_restore --list` before publication and receives a SHA-256
|
||
sidecar. Retention defaults to 30 days while always preserving at least seven newest
|
||
restore points. An optional independently mounted secondary destination receives the
|
||
same verified artifacts.
|
||
- Added a latest-success marker and container healthcheck that detects a missed 26-hour
|
||
recovery point, a standalone verification command, guarded destinations and bounded
|
||
backup/log storage.
|
||
- Evidence: all shell scripts pass Alpine `sh -n`; merged Unraid Compose validates; a
|
||
disposable PostgreSQL instance produced a real dump, checksum verification passed and
|
||
`pg_restore --list` accepted the artifact. Exact next action: implement privacy export,
|
||
anonymisation safeguards, retention reporting and governance documentation.
|
||
|
||
## M24 — executable privacy governance (2026-08-10)
|
||
|
||
- Added a manager-only Privacy workspace with persisted policy metrics, customer dossier
|
||
export, bounded audit CSV export and irreversible customer anonymisation. Privacy is
|
||
localized in all three supported languages, searchable and hidden from rental staff.
|
||
- Anonymisation is row-locked and requires the exact stable customer reference plus a
|
||
reason. Reserved/active bookings and bookings inside the configurable minimum retention
|
||
window block the action. PII is cleared while stable references and operational history
|
||
remain valid; repeated requests are idempotent.
|
||
- Every export and anonymisation is audited. The anonymisation audit records state and
|
||
justification but deliberately never copies erased PII. Audit CSV ranges are capped at
|
||
90 days and a configurable maximum row count.
|
||
- Added explicit customer anonymisation state/migration `b913a72e8c14`, a governance
|
||
runbook covering inventory, retention, data-subject requests, access review and incident
|
||
handling, plus regenerated OpenAPI.
|
||
- Evidence: privacy API **6 passed without warnings**; ruff/mypy clean; React review led
|
||
to stable callback/effect dependencies and a lazy route chunk; TypeScript, lint and
|
||
production build pass. Exact next action: extend RAGcore corpus statistics and health
|
||
evidence, then run complete acceptance and deploy all production-readiness milestones.
|