Files
MobilityOps/artifacts/live-ai-integration/final-summary.md
T
NuklearRabbit cf4d8e3649 docs: write final n8n + RAGcore integration evidence summary
Consolidates this effort's outcome across all four canonical workflows:
WF1/WF2 hardened and live, WF3 blocked on a RAGcore-side credential
rejection (with trace IDs for the operator to investigate), WF4 built
and live-validated with one open non-blocking follow-up. No credential
values or secrets included.
2026-08-04 17:07:54 +02:00

12 KiB

Live n8n + RAGcore integration — final evidence

No credential values, tokens, or secrets appear anywhere in this document. Where a credential or trace ID is referenced, only its name or an opaque reference identifier is given, never its value.

Commit

Built on branch feat/live-n8n-ragcore-integration, HEAD at commit aaa16305354d34f9c1f4d57253d33d9062c38faa ("docs: record WF2 retry fix and WF4's n8n-session-expiry blocker"). Run git log --oneline feat/live-n8n-ragcore-integration for the full history of this effort.

Scope

The brief required treating n8n (https://n8n.itworx.tech) as a full third integration layer alongside RAGcore and MCP Hub, with Fleet Ops keeping exclusive ownership of business rules, authorization, transactions, audit, and idempotency. Four canonical n8n workflows were required. The repository (n8n/workflows/*.json + MANIFEST.md + n8n/workflows/check_drift.py) is the source of truth for cleaned workflow definitions; the Fleet Ops integration status page (/automation) shows real per-workflow operational evidence, not a config boolean.

Result summary

# Workflow Status Live evidence this round
1 Fleet Ops — Vehicle Return Orchestration Live, hardened Timeout+bounded-retry gap found and fixed
2 Fleet Ops — Scheduled Data Quality Scan Live, hardened Same gap found and fixed
3 Fleet Ops — RAGcore Procedure Sync Blocked Not built — RAGcore rejects credential issuance (see below)
4 Fleet Ops — Workflow Error Handler Live, validated Mock + genuine induced-failure test; own hardening incomplete (see below)

Full per-workflow detail (purpose, trigger, event contract, required credentials, live workflow ID, checksum) is in n8n/workflows/MANIFEST.md, which is the authoritative, continuously-updated source — this document is a point-in-time summary of that state plus the reasoning behind what's not done.

Workflow 1 — Vehicle Return Orchestration

Live workflow ID mobilityops-return-processing. Validated in an earlier round of this effort: webhook trigger requires Header Auth (Fleet Ops Webhook Trigger Token), validates event_type == vehicle.returned.v1, derives a follow-up category, calls Fleet Ops's /return-callback endpoint with an Idempotency-Key header via a named Fleet Ops Service Token credential (not a literal secret), and responds with a controlled JSON result. Idempotent on both sides (event_id flows through as the dedup key; the backend independently checks for a prior audit event before recording again).

This round's finding: the Record follow-up HTTP node had no explicit timeout and "Retry On Fail" disabled — a real gap against the requirement that external dependencies have timeouts and bounded retries. Fixed live: Retry On Fail (3 tries, 1000ms wait) + a 15000ms timeout, published. Safe to retry because the callback is idempotent. Repo definition and manifest checksum synced (commit 0562893).

Attached to workflow 4 as its Error Workflow.

Workflow 2 — Scheduled Data Quality Scan

Live workflow ID mobilityops-scheduled-quality-scan. Validated earlier: hourly Schedule Trigger + a Manual Trigger for on-demand testing, both feeding a single HTTP call to Fleet Ops's /scheduled-scan endpoint (Header Auth via the same Fleet Ops Service Token credential, 15000ms timeout already configured), which runs the domain-level run_scan() function — documented and tested as idempotent by construction (only ever creates an issue for a condition that doesn't already have one open), so overlapping or retried triggers do no duplicate domain work.

This round's finding: the same Retry On Fail gap as workflow 1 (timeout was already set, retries were not). Fixed live the same way (3 tries, 1000ms wait), published. Repo definition and manifest checksum synced (commit 167bf49).

Attached to workflow 4 as its Error Workflow.

Workflow 3 — RAGcore Procedure Sync — blocked

Not built. This workflow needs an application credential (scope sources:sync) for the fleet-ops application in RAGcore. Two independent issuance attempts, in two separate rounds of this effort, both failed with an opaque server-side rejection:

  1. Raw API: POST /v1/applications/{id}/credentials400, "authoritative service-account state rejected issuance".
  2. RAGcore admin UI, this round, after the project owner explicitly authorized Claude to self-issue the credential: the "Issue credential" form for the fleet-ops application, submitted as the Platform Admin role (the highest role visible in the RAGcore admin), with name n8n-ragcore-procedure-sync and scope sources:sync only. Result: "Something went wrong. The credential could not be issued with those values.", trace reference 1955c6a8968c4941a22a1faef39e17a7.

The fleet-ops application itself shows as ordinary/Active in the RAGcore admin, with no visible lock flag, and RAGcore's own OpenAPI spec documents no validation rule that would explain either rejection (no 422, no field-level errors). Two independent paths — a raw API call and the admin UI as the top admin role — hitting the same failure signature is conclusive evidence this is a RAGcore-side policy or bug, not a Fleet Ops request-shape or permission problem. It is not fixable from the Fleet Ops side or through further UI automation. Resolving it requires whoever operates the RAGcore instance to look up the trace ID above (and the earlier raw-API rejection) in RAGcore's own logs.

The real RAGcore contract this workflow will be built against — once a working credential exists — was independently inspected via RAGcore's live OpenAPI spec and is recorded in docs/live-ai-integration/n8n-current-state.md and contracts/ragcore-contract-assumptions.md: control-plane endpoints require an Idempotency-Key header; ingestion is POST /v1/uploads; retrieval is POST /v1/search / /v1/context / /v1/answers (the latter requiring requested_space_ids, an array of knowledge-space UUIDs); health is /health/live and /health/ready (not /health); the scope enum is search, context, answer, documents:read, citations:read, feedback:write, sources:sync.

RAGcoreKnowledgeProvider adapter (backend/app/services/knowledge/ragcore.py) still targets the earlier speculative contract (/health, POST /api/v1/ask, Bearer token) rather than the real one above. This was deliberately not rewritten this round: rewriting it blind, without a credential to validate against, risks introducing a silent behavioral bug in exactly the code path responsible for the project's "AI must never invent an answer when RAGcore is unavailable or returns insufficient evidence" guarantee — for example a wrong evidence_state mapping that looks fine in code review but misclassifies "unavailable" as "insufficient" (or vice versa) against the real response shape. The adapter's current behavior is honest and safe (it degrades cleanly to unavailable on any request or parsing failure, and ragcore_api_token is unset by default so the app correctly runs on the local demo knowledge provider today). The rewrite stays queued behind the same credential blocker as workflow 3.

Workflow 4 — Workflow Error Handler

Live workflow ID Xppn2rAEqUuyiCJF. Built and live-validated in an earlier round: Error Trigger → a Code node that derives a bounded, secret-free failure report (error category classified from the message text, truncated summary, no stack trace, no headers or tokens) → an HTTP call to Fleet Ops's /workflow-error endpoint (Header Auth via the same Fleet Ops Service Token credential), which registers the failure as an audit event idempotently keyed on execution_id.

Set as the Error Workflow on both workflow 1 and workflow 2. Confirmed workflow 4 has no Error Workflow of its own (prevents a recursive loop).

Live validation performed: a pinned mock Error Trigger payload produced a real 200 {"status":"registered", ...} from the live Fleet Ops server; re-running the identical payload produced "status":"already_registered", confirming idempotency. A genuine induced failure (temporarily pointing workflow 2's HTTP node at a nonexistent path, then reverting) confirmed workflow 2 itself fails correctly against a broken endpoint and recovers cleanly once reverted.

Known limitation: n8n's Error Workflow trigger does not fire for manual editor "Execute workflow" test runs — checked via workflow 4's own Executions list after the induced workflow-2 failure, and confirmed no new execution appeared. n8n only invokes a workflow's assigned Error Workflow for unattended/production trigger executions, not manual test runs from the editor. The mock-data path exercises the same nodes, logic, and real Fleet Ops endpoint, but a fully automatic (schedule- or webhook-triggered) failure cascading into workflow 4 was not observed live in either round.

Open follow-up (minor, non-blocking): continuing this round's acceptance pass to workflow 4 found the same timeout/retry gap as workflows 1 and 2 on its own outbound HTTP call. A fix was started (15000ms timeout added, Retry On Fail toggled on) but n8n's autosave began failing with "Unauthorized" mid-edit; a fresh browser tab confirmed the n8n session had expired (redirected to /signin). Nothing was saved — workflow 4's live definition is unchanged from before this round, so there is no partial or broken state. This is lower-stakes than workflows 1/2 (workflow 4 is the error notifier itself, not a primary business flow, and a failed error-report is already visible in n8n's own execution history via On Error: Stop Workflow) but should be finished once the n8n browser session is re-authenticated.

Repository source of truth

n8n/workflows/ holds cleaned, credential-value-free JSON definitions for all built workflows, n8n/workflows/MANIFEST.md documents purpose/trigger/contract/credentials/ live-ID/checksum for all four canonical workflows (including workflow 3's blocked status), and n8n/workflows/check_drift.py is a read-only script that compares the repo definitions against the live instance via n8n's Public API and reports drift — safe to run in CI as a non-blocking check. No literal export/download mechanism was found working in this n8n version, so each definition was reconstructed from direct, verified UI inspection rather than a native export; this limitation is noted in the manifest itself.

Integration status page

/automation (Operations Manager only) surfaces real per-workflow evidence derived purely from Fleet Ops's own audit/outbox tables — no new dependency on n8n's API was added to the backend. Each of the four canonical workflows shows a status (not built / no evidence yet / operational) and a last-evidence timestamp; the scheduled-scan evidence specifically filters to actor_type == "service" so a manually-triggered scan in the UI doesn't count as n8n evidence. An error-handler summary line reports total registered automation failures and the most recent one.

Verified live in the browser (Dutch locale) both locally and on the deployed production server (http://192.168.10.150:1236/automation): correctly showed "3 van 4 canonieke n8n-workflows hebben actuele evidentie van werking" with real timestamps for the return/scan/error-handler workflows, "Nog Niet Gebouwd" for the RAGcore sync, and the real error-handler registration from this effort's live testing.

Deployments performed (all explicitly user-approved)

  1. Backend /workflow-error endpoint (commit bbdb4a9) — deployed and verified (/health OK, new endpoint returns 422 not 404 on an empty POST body).
  2. Integration status page, backend + frontend (commit 4049c0c) — deployed and verified (/health OK, page renders real evidence in the browser).

The three n8n-side node edits this round (WF1 timeout/retry, WF2 timeout/retry, WF4's incomplete attempt) are live edits to the n8n instance itself and do not require a Fleet Ops redeploy.

What's left

  1. RAGcore credential issuance — blocked on RAGcore's own server-side rejection (trace 1955c6a8968c4941a22a1faef39e17a7 and the earlier raw-API 400). Needs RAGcore's operator to investigate. Unblocks workflow 3 and the RAGcoreKnowledgeProvider real-contract rewrite.
  2. Workflow 4's own timeout/bounded-retry hardening — needs the n8n browser session re-authenticated to finish; a small, well-understood, non-blocking edit.
  3. Fleet Ops logo/favicon — explicitly deferred by the project owner as a separate, unrelated follow-up task, not part of this integration effort.