From cf4d8e3649c3721919f6c3b0af545cc0f65c96a8 Mon Sep 17 00:00:00 2001 From: NuklearRabbit <145918611+NuklearRabbit@users.noreply.github.com> Date: Tue, 4 Aug 2026 17:07:54 +0200 Subject: [PATCH] docs: write final n8n + RAGcore integration evidence summary Consolidates this effort's outcome across all four canonical workflows: WF1/WF2 hardened and live, WF3 blocked on a RAGcore-side credential rejection (with trace IDs for the operator to investigate), WF4 built and live-validated with one open non-blocking follow-up. No credential values or secrets included. --- .../live-ai-integration/final-summary.md | 207 ++++++++++++++++++ 1 file changed, 207 insertions(+) create mode 100644 artifacts/live-ai-integration/final-summary.md diff --git a/artifacts/live-ai-integration/final-summary.md b/artifacts/live-ai-integration/final-summary.md new file mode 100644 index 0000000..0ae1f84 --- /dev/null +++ b/artifacts/live-ai-integration/final-summary.md @@ -0,0 +1,207 @@ +# Live n8n + RAGcore integration — final evidence + +No credential values, tokens, or secrets appear anywhere in this document. Where a +credential or trace ID is referenced, only its name or an opaque reference identifier is +given, never its value. + +## Commit + +Built on branch `feat/live-n8n-ragcore-integration`, HEAD at commit +`aaa16305354d34f9c1f4d57253d33d9062c38faa` ("docs: record WF2 retry fix and WF4's +n8n-session-expiry blocker"). Run `git log --oneline feat/live-n8n-ragcore-integration` +for the full history of this effort. + +## Scope + +The brief required treating n8n (`https://n8n.itworx.tech`) as a full third integration +layer alongside RAGcore and MCP Hub, with Fleet Ops keeping exclusive ownership of +business rules, authorization, transactions, audit, and idempotency. Four canonical n8n +workflows were required. The repository (`n8n/workflows/*.json` + `MANIFEST.md` + +`n8n/workflows/check_drift.py`) is the source of truth for cleaned workflow definitions; +the Fleet Ops integration status page (`/automation`) shows real per-workflow operational +evidence, not a config boolean. + +## Result summary + +| # | Workflow | Status | Live evidence this round | +|---|---|---|---| +| 1 | Fleet Ops — Vehicle Return Orchestration | **Live, hardened** | Timeout+bounded-retry gap found and fixed | +| 2 | Fleet Ops — Scheduled Data Quality Scan | **Live, hardened** | Same gap found and fixed | +| 3 | Fleet Ops — RAGcore Procedure Sync | **Blocked** | Not built — RAGcore rejects credential issuance (see below) | +| 4 | Fleet Ops — Workflow Error Handler | **Live, validated** | Mock + genuine induced-failure test; own hardening incomplete (see below) | + +Full per-workflow detail (purpose, trigger, event contract, required credentials, live +workflow ID, checksum) is in `n8n/workflows/MANIFEST.md`, which is the authoritative, +continuously-updated source — this document is a point-in-time summary of that state +plus the reasoning behind what's not done. + +## Workflow 1 — Vehicle Return Orchestration + +Live workflow ID `mobilityops-return-processing`. Validated in an earlier round of this +effort: webhook trigger requires Header Auth (`Fleet Ops Webhook Trigger Token`), +validates `event_type == vehicle.returned.v1`, derives a follow-up category, calls Fleet +Ops's `/return-callback` endpoint with an `Idempotency-Key` header via a named +`Fleet Ops Service Token` credential (not a literal secret), and responds with a +controlled JSON result. Idempotent on both sides (`event_id` flows through as the +dedup key; the backend independently checks for a prior audit event before recording +again). + +**This round's finding**: the `Record follow-up` HTTP node had no explicit timeout and +"Retry On Fail" disabled — a real gap against the requirement that external dependencies +have timeouts and bounded retries. Fixed live: Retry On Fail (3 tries, 1000ms wait) + a +15000ms timeout, published. Safe to retry because the callback is idempotent. Repo +definition and manifest checksum synced (commit `0562893`). + +Attached to workflow 4 as its Error Workflow. + +## Workflow 2 — Scheduled Data Quality Scan + +Live workflow ID `mobilityops-scheduled-quality-scan`. Validated earlier: hourly +Schedule Trigger + a Manual Trigger for on-demand testing, both feeding a single HTTP +call to Fleet Ops's `/scheduled-scan` endpoint (Header Auth via the same `Fleet Ops +Service Token` credential, 15000ms timeout already configured), which runs the +domain-level `run_scan()` function — documented and tested as idempotent by +construction (only ever creates an issue for a condition that doesn't already have one +open), so overlapping or retried triggers do no duplicate domain work. + +**This round's finding**: the same Retry On Fail gap as workflow 1 (timeout was already +set, retries were not). Fixed live the same way (3 tries, 1000ms wait), published. Repo +definition and manifest checksum synced (commit `167bf49`). + +Attached to workflow 4 as its Error Workflow. + +## Workflow 3 — RAGcore Procedure Sync — blocked + +**Not built.** This workflow needs an application credential (scope `sources:sync`) for +the `fleet-ops` application in RAGcore. Two independent issuance attempts, in two +separate rounds of this effort, both failed with an opaque server-side rejection: + +1. **Raw API**: `POST /v1/applications/{id}/credentials` → `400`, "authoritative + service-account state rejected issuance". +2. **RAGcore admin UI**, this round, after the project owner explicitly authorized + Claude to self-issue the credential: the "Issue credential" form for the `fleet-ops` + application, submitted as the Platform Admin role (the highest role visible in the + RAGcore admin), with name `n8n-ragcore-procedure-sync` and scope `sources:sync` only. + Result: "Something went wrong. The credential could not be issued with those + values.", trace reference `1955c6a8968c4941a22a1faef39e17a7`. + +The `fleet-ops` application itself shows as ordinary/`Active` in the RAGcore admin, with +no visible lock flag, and RAGcore's own OpenAPI spec documents no validation rule that +would explain either rejection (no `422`, no field-level errors). Two independent paths +— a raw API call and the admin UI as the top admin role — hitting the same failure +signature is conclusive evidence this is a RAGcore-side policy or bug, not a Fleet Ops +request-shape or permission problem. It is not fixable from the Fleet Ops side or +through further UI automation. Resolving it requires whoever operates the RAGcore +instance to look up the trace ID above (and the earlier raw-API rejection) in RAGcore's +own logs. + +The real RAGcore contract this workflow will be built against — once a working +credential exists — was independently inspected via RAGcore's live OpenAPI spec and is +recorded in `docs/live-ai-integration/n8n-current-state.md` and +`contracts/ragcore-contract-assumptions.md`: control-plane endpoints require an +`Idempotency-Key` header; ingestion is `POST /v1/uploads`; retrieval is `POST +/v1/search` / `/v1/context` / `/v1/answers` (the latter requiring `requested_space_ids`, +an array of knowledge-space UUIDs); health is `/health/live` and `/health/ready` (not +`/health`); the scope enum is `search, context, answer, documents:read, citations:read, +feedback:write, sources:sync`. + +**`RAGcoreKnowledgeProvider` adapter** (`backend/app/services/knowledge/ragcore.py`) +still targets the earlier speculative contract (`/health`, `POST /api/v1/ask`, Bearer +token) rather than the real one above. This was deliberately **not** rewritten this +round: rewriting it blind, without a credential to validate against, risks introducing +a silent behavioral bug in exactly the code path responsible for the project's "AI must +never invent an answer when RAGcore is unavailable or returns insufficient evidence" +guarantee — for example a wrong `evidence_state` mapping that looks fine in code review +but misclassifies "unavailable" as "insufficient" (or vice versa) against the real +response shape. The adapter's current behavior is honest and safe (it degrades cleanly +to `unavailable` on any request or parsing failure, and `ragcore_api_token` is unset by +default so the app correctly runs on the local demo knowledge provider today). The +rewrite stays queued behind the same credential blocker as workflow 3. + +## Workflow 4 — Workflow Error Handler + +Live workflow ID `Xppn2rAEqUuyiCJF`. Built and live-validated in an earlier round: +Error Trigger → a Code node that derives a bounded, secret-free failure report (error +category classified from the message text, truncated summary, no stack trace, no +headers or tokens) → an HTTP call to Fleet Ops's `/workflow-error` endpoint (Header Auth +via the same `Fleet Ops Service Token` credential), which registers the failure as an +audit event idempotently keyed on `execution_id`. + +Set as the Error Workflow on both workflow 1 and workflow 2. Confirmed workflow 4 has no +Error Workflow of its own (prevents a recursive loop). + +**Live validation performed**: a pinned mock Error Trigger payload produced a real `200 +{"status":"registered", ...}` from the live Fleet Ops server; re-running the identical +payload produced `"status":"already_registered"`, confirming idempotency. A genuine +induced failure (temporarily pointing workflow 2's HTTP node at a nonexistent path, then +reverting) confirmed workflow 2 itself fails correctly against a broken endpoint and +recovers cleanly once reverted. + +**Known limitation**: n8n's Error Workflow trigger does not fire for manual editor +"Execute workflow" test runs — checked via workflow 4's own Executions list after the +induced workflow-2 failure, and confirmed no new execution appeared. n8n only invokes a +workflow's assigned Error Workflow for unattended/production trigger executions, not +manual test runs from the editor. The mock-data path exercises the same nodes, logic, +and real Fleet Ops endpoint, but a fully automatic (schedule- or webhook-triggered) +failure cascading into workflow 4 was not observed live in either round. + +**Open follow-up (minor, non-blocking)**: continuing this round's acceptance pass to +workflow 4 found the same timeout/retry gap as workflows 1 and 2 on its own outbound +HTTP call. A fix was started (15000ms timeout added, Retry On Fail toggled on) but n8n's +autosave began failing with "Unauthorized" mid-edit; a fresh browser tab confirmed the +n8n session had expired (redirected to `/signin`). Nothing was saved — workflow 4's live +definition is unchanged from before this round, so there is no partial or broken state. +This is lower-stakes than workflows 1/2 (workflow 4 is the error notifier itself, not a +primary business flow, and a failed error-report is already visible in n8n's own +execution history via `On Error: Stop Workflow`) but should be finished once the n8n +browser session is re-authenticated. + +## Repository source of truth + +`n8n/workflows/` holds cleaned, credential-value-free JSON definitions for all built +workflows, `n8n/workflows/MANIFEST.md` documents purpose/trigger/contract/credentials/ +live-ID/checksum for all four canonical workflows (including workflow 3's blocked +status), and `n8n/workflows/check_drift.py` is a read-only script that compares the +repo definitions against the live instance via n8n's Public API and reports drift — +safe to run in CI as a non-blocking check. No literal export/download mechanism was +found working in this n8n version, so each definition was reconstructed from direct, +verified UI inspection rather than a native export; this limitation is noted in the +manifest itself. + +## Integration status page + +`/automation` (Operations Manager only) surfaces real per-workflow evidence derived +purely from Fleet Ops's own audit/outbox tables — no new dependency on n8n's API was +added to the backend. Each of the four canonical workflows shows a status (not built / +no evidence yet / operational) and a last-evidence timestamp; the scheduled-scan +evidence specifically filters to `actor_type == "service"` so a manually-triggered scan +in the UI doesn't count as n8n evidence. An error-handler summary line reports total +registered automation failures and the most recent one. + +Verified live in the browser (Dutch locale) both locally and on the deployed +production server (`http://192.168.10.150:1236/automation`): correctly showed "3 van 4 +canonieke n8n-workflows hebben actuele evidentie van werking" with real timestamps for +the return/scan/error-handler workflows, "Nog Niet Gebouwd" for the RAGcore sync, and +the real error-handler registration from this effort's live testing. + +## Deployments performed (all explicitly user-approved) + +1. Backend `/workflow-error` endpoint (commit `bbdb4a9`) — deployed and verified + (`/health` OK, new endpoint returns `422` not `404` on an empty POST body). +2. Integration status page, backend + frontend (commit `4049c0c`) — deployed and + verified (`/health` OK, page renders real evidence in the browser). + +The three n8n-side node edits this round (WF1 timeout/retry, WF2 timeout/retry, WF4's +incomplete attempt) are live edits to the n8n instance itself and do not require a +Fleet Ops redeploy. + +## What's left + +1. **RAGcore credential issuance** — blocked on RAGcore's own server-side rejection + (trace `1955c6a8968c4941a22a1faef39e17a7` and the earlier raw-API `400`). Needs + RAGcore's operator to investigate. Unblocks workflow 3 and the + `RAGcoreKnowledgeProvider` real-contract rewrite. +2. **Workflow 4's own timeout/bounded-retry hardening** — needs the n8n browser session + re-authenticated to finish; a small, well-understood, non-blocking edit. +3. **Fleet Ops logo/favicon** — explicitly deferred by the project owner as a separate, + unrelated follow-up task, not part of this integration effort.