October 6, 2026
n8n Error Handling and Retries: A Workflow You Can Copy
The one-paragraph version
n8n gives you three layers of error defense: per-node settings (retry on fail + what to do when retries run out), error branches (route failures to cleanup logic instead of killing the workflow), and the Error Trigger node (a separate workflow that fires whenever any workflow errors). The pattern that covers 95% of cases: enable "retry on fail" with a wait between attempts on every node that calls an external API, set failing nodes to continue with error output instead of stopping, and add one Error Trigger workflow that logs the failure and pings you. Below is the exact setup, setting by setting.
How n8n thinks about errors
Every n8n node has a Settings tab with two controls that matter:
- Retry on Fail — automatically re-runs the node when it errors. You set the number of attempts and the wait between them.
- On Error — what happens when retries are exhausted: Stop Workflow (default — the whole execution dies), Continue (regular output) (pretend it succeeded, pass along empty/partial data), or Continue (using error output) (route the error into a separate output branch you can handle).
Most n8n beginners leave everything on defaults: no retries, stop on first error. That's why their workflows die at 3am on a transient API hiccup. Ten minutes of configuration fixes it permanently.
Layer 1: Retries on every external call
For every node that touches an external API (HTTP Request, Gmail, Slack, Airtable, OpenAI, etc.), open Settings and configure:
- Retry on Fail: ON
- Max Tries: 3 (5 for notoriously flaky APIs)
- Wait Between Tries: 5,000 ms (5 seconds). For rate-limit-prone APIs, use 30,000-60,000 ms.
That's it. Transient failures — 429s, 503s, timeouts, the API having a bad minute — now resolve themselves. This single change eliminates the majority of 3am failures.
One caution: don't retry on data errors. If the API returns 400 "invalid email," retrying three times just wastes time. Retries are for transient failures (5xx, 429, timeouts). For 4xx errors, you want Layer 2.
Layer 2: Error branches for graceful degradation
Set On Error → Continue (using error output) on nodes where failure shouldn't kill the workflow. The node gains a second output — the error branch — which you wire to fallback logic.
The classic example — an AI enrichment step:
- HTTP Request (calls an AI API to summarize a lead) — On Error: Continue (error output), Retry on Fail: 3 tries.
- Success branch → continues to CRM update with the summary.
- Error branch → Set node that substitutes the raw lead notes as the "summary" → merges back into the CRM update.
The workflow never dies because the AI API hiccuped; the CRM just gets slightly less polished data that one time. Apply this pattern to every optional step: Slack notifications, enrichment calls, logging. Core steps (charging the customer, creating the record) should usually still stop the workflow on failure — you want to know about those loudly.
Layer 3: The Error Trigger — your safety net
The Error Trigger node starts a workflow whenever another workflow errors. Build one global error-handling workflow:
- Error Trigger node (no config needed — it catches errors from workflows you attach it to).
- Set node — extract: workflow name, node that failed, error message, timestamp, execution URL.
- IF node — is this a known-transient error (429/503/timeout)? If yes → log quietly. If no → continue.
- Slack / Email node — ping yourself with the details and a link to the failed execution.
- Google Sheets / Airtable node — append to an error log for weekly review.
Attach it via each workflow's settings (Error Workflow dropdown). Now no failure goes unnoticed: transient ones get logged, real ones page you with everything needed to debug.
Putting it together: the copyable pattern
For any production n8n workflow, apply this checklist:
- [ ] Every external-API node: Retry on Fail ON (3 tries, 5s wait minimum)
- [ ] Optional steps (notifications, enrichment): On Error → Continue (error output) with a fallback branch
- [ ] Critical steps (payments, record creation): On Error → Stop Workflow (fail loudly, don't fake success)
- [ ] Workflow settings: attach your Error Trigger workflow
- [ ] Add a data validation step before fragile nodes (e.g., an IF node checking "email contains @" before the email API call) — bad data is the #1 cause of non-transient failures
- [ ] Test it: temporarily point an HTTP node at a URL that returns 500 and watch the retry + error branch work
Testing your error handling (don't skip this)
Error handling you haven't tested is decoration. Two quick tests:
- Transient failure test: set an HTTP node's URL to
https://httpstat.us/503, run the workflow, confirm it retries 3 times then takes the error branch. Then restore the URL. - Bad data test: feed a record with a missing required field through the workflow, confirm the validation step catches it before the API call.
Five minutes of testing beats a month of wondering whether your safety net works.
Verdict
n8n's error model is genuinely good — retries, error branches, and a global error trigger cover everything from hiccups to catastrophes. The failure mode isn't the tooling, it's the defaults: out of the box, n8n stops dead on the first error. Spend ten minutes per workflow on the three layers above, test them once, and your automations survive the real world.
Based on n8n's node settings and Error Trigger as of October 2026. Setting names may shift between versions — the three-layer pattern doesn't.
Enjoyed this? Join the newsletter for one practical automation tip a week. No spam, no hype.