From a red nodeto a named cause.
The pipeline is itself an n8n workflow, so Insight runs on the platform it diagnoses. Scroll and each stage lights up with what it does. The one that's switched off says so.
01
Trigger
Three ways in, one pipeline. A monitored workflow's Error Trigger posts a thin payload with a per-instance token. The public page and the CLI send an uploaded execution, or an id plus your instance details.
- push
- Error Trigger → ingest webhook
- tenant
- ingest token, stored SHA-256 hashed
- public
- 5 requests / min per IP
02
Fetch
Given an id, the pipeline asks your instance for the full execution, every node's input and output. Transient failures on this call are retried with backoff. Uploads skip this step.
- endpoint
- GET /executions/{id}?includeData=true
- retries
- 3 tries, backoff
- upload
- skipped
03
Redact
The first thing that happens to the data. Secret-shaped values are stripped by field name and by shape, including the API key you just typed, before anything reaches a model or a database.
- when
- before the LLM, before storage
- how
- field names + credential shapes
- cli
- runs on your own machine
04
Short-circuit
Timeouts, resets, rate limits and gateway errors have a signature. When the error matches one and nothing else, the answer is “looks transient”, returned without spending a model call on it.
- matches
- ETIMEDOUT · 429 · 502–504
- result
- looks transient
- cost
- no LLM call
05[ switched off ]
Retrieve
A knowledge base of n8n-specific failure patterns, embedded and searched by the failing node's type and error. It is built and wired, and switched off until the Qdrant collection is seeded with real patterns. Until then the model diagnoses from the error and execution alone.
- store
- Qdrant + Hugging Face embeddings
- status
- built, disabled
- cli
- patterns carried in the prompt
06
Diagnose
One bounded call to Groq. The execution goes in as quoted, untrusted data, because an error message or an API response can carry text that looks like instructions. What comes back is structured: node, category, explanation, confidence, fix.
- model
- Llama 3.3 70B on Groq
- calls
- one per diagnosis
- input
- untrusted, framed as data
07
Calibrate and deliver
The confidence score decides the wording: under 0.40 a lead to verify, above 0.70 a specific fix. Public results are shown and not kept. Connected instances get a metadata row on the dashboard and a Slack alert.
- tiers
- 0.40 · 0.70
- stored
- metadata row, no raw payload
- alert
- Slack, connected instances
Diagnosedbefore youlook.
The public page answers one failure at a time. Connecting an instance means the next failure is diagnosed the moment it happens, and the answer is already in Slack when you open it.
| Step | You | Insight | Worth knowing |
|---|---|---|---|
| Connect | Base URL and an n8n API key | The key is encrypted at rest; you get a per-instance ingest token | Nothing on your instance changes yet |
| Scan | Insight lists every workflow on the instance | Each is flagged monitored or not | Read-only calls |
| Add workflow | One click on a workflow | Insight installs its error-workflow template once, activates it, and points that workflow's Error Workflow setting at it | Never touches the workflow's own nodes or connections |
| Fail | The workflow breaks in production | The template posts the execution id to Insight; the full pipeline runs | Your instance has to be reachable from the internet |
What's live,and whatisn't yet.
The headline number for a tool like this is accuracy with retrieval switched on. That number doesn't exist yet, so this page doesn't pretend it does.
| Diagnosis pipeline | ● live | Fetch, redact, short-circuit, diagnose and store, end to end |
|---|---|---|
| Public diagnose page | ● live | Upload or execution id, no account |
| Monitoring and Slack alerts | ● live | Connect, add workflow, diagnosed on failure |
| npx insight-n8n | ● live | Redaction runs locally; diagnosis runs locally with your Groq key, or on the hosted pipeline |
| Knowledge-base retrieval | ○ built, off | Waiting on a seeded Qdrant collection of real failure patterns |
| Accuracy scorecard | ○ not yet | The PRD targets 85% on a 25-case eval set; it hasn't been run with retrieval on, so there is no number to show |