Ran a real AI Operations Brief via the live ITWorx MCP Hub connector's own
MobilityOpsClient against production Fleet Ops: real operations summary, real
most-pressing vehicle, real grounded knowledge answer with citations, real
correlation IDs verified end-to-end in Fleet Ops's own audit log. No write
actions performed. Runbook and full output in
docs/final-integrations/ai-operations-brief-runbook.md.
Ran the full Playwright e2e suite against the live deployed instance and fixed
two pre-existing fragile locators unrelated to this session's feature work
(both broke because Automation now legitimately has two tables sharing the
same generic selectors, exposed by running the full suite rather than
individual files) plus one pre-existing untranslated-loanword false positive.
All specs pass.
artifacts/final-integrations/final-summary.md has the complete evidence
write-up: repository/deployment state, what was fixed vs. handed off, test
results, and known limitations stated plainly.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Role names, audit/scenario labels, and status text were previously either
left in English or only partially translated:
- auth.json/demo.json role labels actually translated (not just labelled
as translated): Operationsmanager/Verhuurmedewerker,
Responsable des operations/Collaborateur de location.
- "Audit trail" -> Auditgeschiedenis/Piste d'audit (title, column header,
and every mid-sentence occurrence across demo.json, quality.json,
returns.json -- these embedded leaks were previously invisible to the
whole-string identity check).
- "Open" (status) -> Openstaand, "Recent" -> Recentste,
"Start scenario" -> Scenario starten / Demarrer le scenario.
Matching Playwright spec text updated in the same commit so the suite
never regresses through a broken intermediate state.
Live validation on the deployed fix branch caught a real bug: every data-quality
issue's top-of-page "Evidence summary" line rendered the raw, always-English legacy
evidence.summary string unconditionally -- in all three languages -- even though the
backend has been emitting structured, localizable evidence.signals for a while
(app/services/data_quality.py already documented this exact intent). The frontend
side of that conversion was never finished.
- DataQualityIssueDetail.tsx now renders evidence.signals through the operator's
locale as the primary summary; the raw evidence.summary string is only visible
inside "Technical details" (via the existing EvidenceDisclosure JSON dump).
- The four DQ-DEMO-* seed rows that anchor the guided demo's scripted scenarios now
carry real, accurate signals computed at seed time (duplicate-customer's similarity
score is the actual SequenceMatcher ratio on the seeded names, not invented) instead
of only a legacy English sentence.
- Rows with no structured signals (generic filler seed data) fall back to the raw
text rather than showing a blank summary; the one known placeholder string gets its
own localized rendering so it never displays as English filler either.
- New regression test: the vehicle_status_conflict evidence summary must show
localized text and must never contain the specific raw English sentence that was
live-visible before this fix, in all 3 languages.
151 backend tests, Ruff, mypy green; full local Playwright suite green (a couple of
sequential-run-only flakes, both confirmed to pass in isolation and unrelated to this
change).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- fleet-ops-correction.spec.ts: opens every main route in all 3 languages, asserting
no console errors, correct html[lang], and a real non-empty page heading (key parity
across locale files is already proven structurally elsewhere, so this focuses on what
only a live render can catch).
- i18n-coverage.spec.ts: a static scan for hardcoded JSX text bypassing t(...). A naive
`>text<` regex falsely flagged TypeScript generics everywhere (`useState<string |
null>(null)` was read as a "JSX tag" spanning to the next unrelated `>`) -- fixed by
requiring the closing tag name to backreference the opening one
(`<Tag>...</Tag>`), which generics can never satisfy. Verified against both false
positives (passes clean on the current codebase) and false negatives (deliberately
injected and reverted a hardcoded string to confirm it's caught).
Known pre-existing flake (unrelated to this branch, not touched by it): "logout
invalidates the server session so a refresh returns to login" in
interactive-elements.spec.ts occasionally fails only in the full sequential run,
never in isolation -- AuthContext.logout() clears local state and redirects before
awaiting the server-side cookie-clearing POST, a narrow race no human interaction
speed would ever hit. Noted as a known limitation, not fixed (out of this branch's
scope).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
New frontend/e2e/fleet-ops-correction.spec.ts covers section 12 of the brief:
branding (Fleet Ops visible, no MobilityOps/PoC leaks, in all 3 languages), the
language switcher persisting across reload, the full status-recommendation flow
(non-mutating preview, exact-status confirm button, manual review with no apply
button, stale-token rejection), MO-016 order independence at the browser level, the
knowledge base grounding the exact brief question in its own language, and localized
audit/automation content with raw codes only under "Technical details".
Writing these tests surfaced two real bugs:
- DataQualityIssueDetail.tsx conflated "no conflict" with "manual review required"
because both carry safe_to_apply: false (a no_conflict recommendation has nothing to
apply, so it's trivially "not safe to apply" without being unsafe). This showed a
false "manual review required" panel for MO-016 after its overlap was resolved,
instead of the correct "no change needed" state. Fixed by keying the branch on
manual_review_required alone.
- test_mo_016_status_conflict_recommendation_is_order_independent never actually
exercised MO-016: _first_open() returned whichever vehicle_status_conflict issue was
most recently detected (there are ~14 open after a reset), not necessarily
DQ-DEMO-STATUS, so the test's MO-016 assertions were trivially true regardless of
what the code under test did. Added _first_open_for_vehicle() and rewrote the test
to explicitly target MO-016, and to assert the behaviour order independence actually
requires: resolving the overlap first must correctly leave nothing to apply (the
vehicle already matches the facts), not literally the same end status as resolving
the conflict first.
151 backend tests, Ruff, mypy green; full 108-test Playwright suite green (two
transient, non-reproducible flakes confirmed to pass in isolation and unrelated to
this change).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>