Files
MobilityOps/docs/final-integrations/current-state-audit.md
T
NuklearRabbitandClaude Sonnet 5 34df66d28c M8: GUI polish, n8n workflow-3 fixes, RAGcore retrieval root-cause and fix
GUI: dashboard Attention Queue presents a curated severity mix instead of pure
severity-sort (grouped Now/Today/Later headers); Today's Movements seed data
curated so a fresh reset shows a credible day (2+ departures, 2+ returns), with
a new seed-integrity test; About Demo restructured into a compact grid with
progressive disclosure for technical sections; Duplicate Merge shows match/conflict
counts, hides matching fields by default, and previews the final merged record
before confirmation.

Repo hygiene: removed a stray empty `backend;C` directory and an untracked 31MB
zip export; `.gitignore` now excludes future archive exports.

n8n: fixed invalid JSON (a missing `},` between two node objects) in the committed
`fleet-ops-vehicle-return.json` -- the file could not be parsed. Live-validated
workflow 3 (RAGcore Procedure Sync): found and fixed a real defect (three body
parameters had a stray trailing `}}`) and a missing Error Workflow wiring, both
via the safe `n8n import:workflow` CLI path; exported the corrected, still-
inactive workflow as the new source of truth and updated MANIFEST.md/check_drift.py.
Publishing it (starts real daily unattended runs) remains a separate decision.

RAGcore: root-caused and fixed (live, approved) the "zero retrieval candidates"
bug -- a filesystem permission bug (`embedding_profiles.json` unreadable by the
app's own runtime user) that broke every retrieval call before it reached Qdrant.
Every other suspect (grants, scope resolution, Qdrant filters, embeddings) was
verified healthy first. Found a second, deeper gap: the reranker adapter calls
an Ollama HTTP route that does not exist on the deployed Ollama version, so
`/v1/answers` still returns `not_answerable`. `KNOWLEDGE_PROVIDER` stays `demo`
until that is resolved on the RAGcore side. Evidence-based MCP Hub integration
status (real tool-call audit history, not just a boolean flag) replaces the old
`configured`/`not_configured` guess. Full findings in
`docs/final-integrations/current-state-audit.md`.

Backend: 172 tests passing, ruff clean, mypy clean (50 files). Frontend: tsc
clean, production build clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 13:05:02 +02:00

153 lines
11 KiB
Markdown

# Current-state audit — Fleet Ops final integrations
Date: 2026-08-05. Compiled from direct repository inspection (git log/status/diff across
all three repos), `PROJECT_STATE.md` history, and read-only investigation of the sibling
repos' own state docs. No live server SSH/curl evidence is included in this pass yet —
see `integration-release-state.md` for the live-verification checklist as it is executed.
## Repository revisions at audit time
| Repo | Path | Branch | HEAD | Notes |
|---|---|---|---|---|
| Fleet Ops (MobilityOps) | `C:\Projects\MobilityOps` | `feat/fleet-ops-final-integrations` (new, branched from `feat/live-n8n-ragcore-integration`) | `3ebca9e` | `feat/live-n8n-ragcore-integration` was pushed to `origin` at `0571a40` and deployed live; `3ebca9e` (logo rebrand) is one commit ahead, not yet deployed. `master` is 19 commits behind and stale (localization-round only). |
| RAGcore | `C:\Projects\RAGcore` | `main` | `64a908a` | Up to date with `origin/main`. Uncommitted local work in progress (see below) — not Fleet-Ops-related, left untouched. |
| ITWorx MCP Hub | `C:\Projects\ITWorx_MCP_Hub` | `feature/wp240-final-acceptance` | `26e6bd8` (+ later `75bb16a`) | Contains `f107544` (MobilityOps connector) as a direct ancestor, plus a real contract fix (`96de385`, vehicleRef camelCase). Already deployed live to Tower at `c4a0f6d`. |
## Branch-name correction (recorded assumption)
The task brief names the working branch `feat/fleet-ops-final-integrations` as already
selected and "branched from the most recently validated, localized, deployed master
branch." That literal branch did not exist. `master` is in fact stale (19 commits behind,
last touched for a localization round only) — the actually-validated, deployed line of
work is `feat/live-n8n-ragcore-integration` (pushed to `origin`, deployed to
`http://192.168.10.150:1236` at `0571a40`, one commit behind current HEAD). Created
`feat/fleet-ops-final-integrations` from that branch's HEAD (`3ebca9e`) instead of from
`master`, since that satisfies the actual intent (continue from the validated/deployed
line) even though the literal branch name in the brief was inaccurate.
## What is actually already done (contradicts "not yet live" framing in places)
- **n8n**: 3 of 4 canonical workflows are live and active in the shared instance
(`n8n.itworx.tech`): Vehicle Return Orchestration, Scheduled Data Quality Scan,
Workflow Error Handler. The 4th, RAGcore Procedure Sync, has all 6 nodes built and
saved but is **not published** (deliberately left for an explicit activation decision,
since publishing starts real unattended daily runs against production). The root cause
of an earlier "auth"-looking failure (`N8N_PROXY_HOPS=0` behind the TLS-terminating
reverse proxy, breaking the browserId CSRF check on every mutating REST call) was found
and fixed at the infrastructure level (Unraid template), not worked around.
- **RAGcore**: deployed to `http://192.168.10.150:1237`, application wiring for
search/context/answer is real (commit `a2905cc`, confirmed present in RAGcore's own
history at `13 commits behind HEAD`). `KNOWLEDGE_PROVIDER` is still `demo` in Fleet Ops
because real queries against the "Fleet Ops Procedures" space return **zero dense and
zero sparse candidates** at the raw retrieval stage — confirmed not a Fleet-Ops-side
wiring bug (RAGcore's own trusted Query Lab tool reproduces the identical zero-candidate
result against the same space). Root cause not yet found as of this audit; ruled out so
far: point count/scoping (83 published, correctly scoped), embedding digest mismatch
(matches), collection alias resolution (resolves correctly). One separate, confirmed,
pre-existing bug: `_DEFAULT_LANGUAGE = "en"` is hardcoded in RAGcore's ingestion handler
— every chunk is stamped `language: "en"` regardless of actual content; RAGcore has never
done real language detection. Not the cause of zero candidates, but must be fixed for
trilingual retrieval (task 6A) once the space is answerable at all.
- **MCP Hub**: the Fleet Ops read-only connector (4 tools, `mobilityops.*`) is **already
live in production** on Tower (commit `c4a0f6d`), reachable via `fleetops.itworx.tech`,
end-to-end verified once already per the Hub's own `CLAUDE.md`/`BUILD_STATE.json`. A
real contract bug was found and fixed there (`vehicle.get`'s input schema disagreed with
the actual wire parameter name — `vehicle_ref` vs `vehicleRef`). What is **not** yet done:
the Hub's own formal production-acceptance checklist row for MobilityOps (`CON-P04`) has
not been executed, and Fleet Ops's own `MCP_HUB_REGISTRATION_ENABLED`/base-URL
configuration has not been confirmed as actually wired and flipped on from the Fleet Ops
side (open item for this audit's Batch 4).
## Confirmed contradictions to resolve (task section 3)
- "Twee versus vier n8n-workflows": resolved above — 3 active + 1 built-but-unpublished.
Canonical set is 4; only 3 are live.
- "Demo-provider versus live RAGcore": Fleet Ops is still on the demo knowledge provider
by deliberate, documented decision (not an oversight) pending the retrieval root cause.
- "MCP Hub-status": prior Fleet Ops docs (`.env.example`, `MCP_HUB_REGISTRATION_ENABLED`)
predate the Hub-side deployment and need reconciling against the fact that the connector
is already live on the Hub side.
- Repo hygiene: removed an untracked, empty `backend;C` directory and an untracked 31 MB
`MobilityOps.zip` stray export; added `*.zip`/`*.tar.gz` to `.gitignore`. No accidentally
committed `__pycache__`/`.pytest_cache`/`test-results` were found in git history.
## RAGcore retrieval root cause — found and partially fixed (2026-08-05, this session)
Investigated live against production (`192.168.10.150`, containers `ragcore-app-1`,
`ragcore-qdrant-1`, `ragcore-postgres-1`, `ollama`), read-only first, then two approved
live changes.
**Root cause #1 (FIXED): filesystem permission bug, not authorization/data.** Verified,
in order, that every earlier suspect was actually healthy: the `control.grants` row
(active, `editor` role, correct application/space), the real
`ControlPlaneAuthorizationInputsProvider` + `RetrievalAuthorizationService.resolve()` code
path run in-process against the live DB (resolves a non-empty `effective_space_ids`), the
exact production Qdrant filter run directly against the live collection (returns real
matching points), and a real ANN vector query under that filter (real hits, sensible
scores). The actual break: `/workspace/.state/models/embedding_profiles.json` — the file
`RetrievalPipeline.run()` reads on every single query to resolve the active embedding
profile — was owned by container-side `root:root` mode `600` on the bind-mounted
`/mnt/cache/appdata/ragcore/state/models` host path, while the real running app process
is uid 10001 (`ragcore`). Every retrieval call hit a `PermissionError` reading its own
state file before ever reaching Qdrant — a plain filesystem-ownership bug, invisible to
every DB/Qdrant-level check. **Fixed live**: `chown 10001:10001` +
`chmod 644`/`755` on that file/directory (approved by the user beforehand). Re-verified
in-process: `RetrievalPipeline.run()` now returns 5 real, relevant hits for an English
damage-procedure question (previously 0).
**Root cause #2 (found, NOT fixed — needs a design decision): reranking is
architecturally unavailable.** `DEFAULT_RERANKER_PROFILE.model_identifier` is
`bge-reranker-v2-m3:v1`, which was never actually present in Ollama's model list (0 of 14
installed models matched). With the user's approval, pulled a working GGUF
(`xitao/bge-reranker-v2-m3:latest`, 1.2 GB) into the shared Ollama instance. **This did
not fix reranking**: `OllamaRerankAdapter` posts to `{ollama_base_url}/api/rerank`, and
this Ollama server (version `0.32.5`) returns a plain `404` for that route — it has no
rerank endpoint at all. This is not a missing-model problem, it is that RAGcore's
reranker adapter was built against an Ollama HTTP API that does not exist in the deployed
version (matches the code's own comment that no reranker-profile registry or live
validation existed yet). The retrieval pipeline degrades gracefully on rerank failure
(RRF-fusion-only hits still returned, confirmed above), but the `/v1/answers` endpoint's
answerability classifier still returns `not_answerable`/0 citations for real NL/EN/FR
questions against real matching content, live-verified after fix #1 with a freshly
minted, correctly-scoped credential.
Options for #2, not decided yet: (a) find/confirm whether a newer Ollama version adds a
real `/api/rerank` route and upgrade the shared instance (affects every other project on
this Ollama — needs its own explicit approval and blast-radius review); (b) change
RAGcore's reranker adapter to call a route Ollama actually supports (e.g. score via
`/api/embed` + a manual similarity/cross-encoder computation, or drop the separate
rerank step and let the answerability classifier trust RRF-fused scores) — a RAGcore
code/design change, out of Fleet Ops's own mandate to decide unilaterally; (c) leave
`KNOWLEDGE_PROVIDER=demo` until RAGcore's own team/session resolves this.
**Side effect to flag**: minting the live-verification credential used `rotate=True` on
the existing "Fleet Ops Knowledge Assistant (production)" service account (a second
credential would have exceeded RAGcore's own 2-active-credential cap), which invalidates
whatever token was previously issued for that account. Since Fleet Ops is still on
`KNOWLEDGE_PROVIDER=demo`, this has no live user-facing impact today, but a fresh
credential must be issued and wired into Fleet Ops's `RAGCORE_API_TOKEN` at actual
cutover time — do not assume the old one still works.
**Concurrency note**: `C:\Projects\RAGcore` had substantial uncommitted local changes
from what appears to be a different, actively-running session (36 modified/untracked
files by the end of this investigation, including files this investigation also read).
No commits or file edits were made in that checkout this session precisely because of
that collision risk — the two live fixes above were applied directly to the running
containers/Ollama instance (approved), not to the RAGcore git repository. **Follow-up
required**: once the concurrent session's work lands, the reranker-profile fix (whichever
option above is chosen) still needs an actual code change + commit + redeploy in
`C:\Projects\RAGcore`, which was not safe to do mid-collision this session.
## Minimal remaining implementation order
1. Root-cause the RAGcore zero-candidate retrieval bug (blocks flipping `KNOWLEDGE_PROVIDER`
and blocks the trilingual live-acceptance and AI Operations Brief tasks).
2. Fix RAGcore's hardcoded `language: "en"` chunk metadata for trilingual retrieval.
3. Decide on and execute n8n workflow 3 publication, with live no-op-on-rerun verification.
4. Confirm/complete Fleet Ops-side MCP Hub registration wiring and run the Hub's own
CON-P04 acceptance row.
5. Build the AI Operations Brief runbook once RAGcore and MCP Hub are both live-green.
6. GUI polish batch (dashboard Today/Attention presentation, duplicate-merge presentation,
About Demo scannability, Demo Guide completion state).
7. Final regression gates and evidence write-up.