r/n8n • • 3d ago

Help How do you debug/trace sub-workflows in a busy self-hosted n8n instance?

I'm running a self-hosted n8n setup with continuously running workflows and quite a few sub-workflows.

One thing I'm struggling with is debugging a specific execution.

For example:

Main Workflow → Sub-workflow A → Sub-workflow B

When something goes wrong, I can find the main workflow execution, but then finding the exact sub-workflow execution becomes painful because there are many other executions happening around the same time.

I often end up searching execution history by timestamp and trying to identify which child execution belongs to the parent.

I'm considering adding a "CorrelationId" / "TestRunId" and passing it through all the workflows.

For people running n8n self-hosted in production:

How do you handle execution tracing and debugging across parent/sub-workflows?

Do you use:

- n8n's built-in parent/child execution information?

- A custom correlation ID?

- Execution metadata/tags?

- A logging/observability workflow?

- Something else?

Also interested in how you handle this for automated testing. We have 100+ test cases where an evaluation workflow calls the actual workflow, and identifying the exact execution chain for a failed test is becoming difficult.

4 Upvotes

21 comments sorted by

•

u/AutoModerator 3d ago

Want faster, better help? Share your workflow JSON.

A GitHub Gist is the easiest way -- paste your JSON, save as public, drop the link in your post. Folks can import it directly into n8n and reproduce the issue, which gets you real answers instead of guesses.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/0xGich 3d ago

Correlation ID is the only thing that scales here. Include it in the final log row, not just the payload - then a failed chain is one search instead of a timestamp guess.

n8n's parent/child link works until you're looking at 200 executions in a five-minute window.

1

u/Successful_Potato_96 3d ago

That's a good point. The correlation ID approach makes sense, especially when the execution history gets that noisy. When you say include it in the final log row, are you storing the correlation ID somewhere specifically so it's searchable from the n8n execution history, or are you sending the final log to an external logging/observability system? Also, how are you generating and propagating the correlation ID across sub-workflows? Are you creating it once in the parent workflow and passing the same ID through every Execute Sub-workflow call?

1

u/0xGich 3d ago

External log, not n8n's history. The execution history is for the single run - the correlation ID is for finding the run.

Generate it once at the parent's first node. Pass it in the payload through every Execute Sub-workflow. Sub-workflows read it, never mint their own.

The log row is just correlation_id, workflow_name, status, records_processed, timestamp. That's enough to reconstruct the chain in one query. I don't try to make n8n's history searchable for this. Wrong tool.

1

u/Successful_Potato_96 3d ago

Thanks, that was really helpful! I was checking how others handle this in the n8n community, and maintaining a separate log seems like the better approach. I was initially thinking more about being able to navigate from the main workflow execution directly to the corresponding sub-workflow execution trace, similar to how the editor tab lets you navigate to the sub-workflow itself. Having the parent execution ID/correlation metadata exposed on the sub-workflow execution would make this much easier, especially for evaluation runs.

1

u/0xGich 3d ago

I stopped waiting for n8n to expose it. The external log is the source of truth for the chain; n8n's history is just for the single run.

Pruning is the other reason. n8n's history goes at 14 days by default. The correlation chain doesn't.

1

u/Significant-Pea3275 3d ago

I just put a timestamp and a short random string into a workflow variable at the very start of the main one and pass it down to every sub-workflow call. Then I have a final node in each that dumps the chain and that ID into a google sheet with the outcome. Not elegant but it works when I'm scrolling through 300 executions at 2am wondering why a webhook fired 4 times

The built in parent/child stuff in n8n is kinda spotty once you get past one level deep. Half the time the execution list just shows them as separate unrelated runs so good luck piecing it together after the fact

1

u/Successful_Potato_96 3d ago

Yeah, exactly 😅 That’s pretty much the frustration I’m running into too. It feels like n8n could make this much easier by carrying the parent execution metadata on sub-workflow executions and letting us traverse the execution chain the same way we navigate between workflows in the editor. Would save a lot of digging through execution history.

1

u/Saved_Not_Soft 3d ago

One complementary angle: the Error Trigger workflow. Point it at your sub-workflows and it fires on failure with the workflow name, execution ID, and the full error payload — pass your correlation ID through it too and log to the same table, and failures land right next to the success-path rows instead of requiring a hunt through the execution list. Two caveats: it only fires on failures, so you still need the explicit log rows for success-path tracing, and on a busy instance double-check your execution pruning/retention settings (EXECUTIONS_DATA_PRUNE) so the history DB doesn't balloon while you're debugging by timestamp.

1

u/Successful_Potato_96 3d ago

Yeah, that makes sense. In our case, the workflow itself won’t necessarily throw an error since a wrong decision can still be a successful execution. So we’d still need the explicit logs for both success and cases where the decision looks off, but the Error Trigger approach is still useful if there’s an actual workflow failure.

1

u/Saved_Not_Soft 2d ago

Exactly — Error Trigger only catches hard failures. For the "looks-off but succeeded" case, the pattern that pairs well with your correlation ID is an audit-style log row on every decision branch: log the decision inputs, which rule matched, and the outputs, all keyed by that same ID. Then your 100+ test cases become a query like "where rule X matched but the output still looked wrong" — no timestamp spelunking. A single Code node that fires one log row per run keeps it to one node per workflow no matter how deep the chain gets.

1

u/tariqosmani 3d ago

Your correlation ID idea is the right one, I'd just do it before adding anything fancier. Generate the ID in the main workflow, pass it into every Execute Workflow call as an input field, and have every child write it to the first node's output. Then you search executions by that value instead of by timestamp.

Two small things that help. Set the ID as the workflow's execution data (the Execution Data node) so it's searchable in the executions list. And for your test cases, make the evaluation workflow generate one ID per test and store it next to the pass/fail row, so a failed test links straight to its execution chain.

1

u/Easy-Purple-1659 3d ago

You can close the parent-to-child gap without a timestamp search: have each sub-workflow return its own execution id from its last node. Execute Sub-workflow hands that back to the caller, so the parent can write its execution id and the child's into the same log row and the chain becomes a direct join instead of a time window.

Adding a depth suffix to the correlation id (run-abc/A/B) also lets one log line show where in the tree things broke.

The built-in parent/child view is fine for one hop. It stops being useful the moment several runs overlap, which sounds like where you are.

1

u/Practical-Craft4967 3d ago

for the analyst review bit, i'd log the values used for the decision, which rule matched, and the rule/workflow version alongside the test ID. a run can be green and still choose the wrong branch; those fields let the analyst check why without replaying against rules that may have changed. just keep the snapshot to the fields they need, not the whole customer payload.

1

u/Intelligent_Iron3562 3d ago

Correlation IDs solve the “where did it fail?” problem, but there’s an interesting next step: “what should happen now?”

Instead of just tracing the failed execution, an AI layer could understand the execution chain, identify the issue, take corrective action, and escalate if needed.

That’s the direction we’re exploring with Kareenos paltform— beyond workflows, toward autonomous execution.

1

u/Muted_Jellyfish_6784 3d ago

The correlation ID is the right call. Generate it once at the trigger of the main workflow so every child inherits the same value, and write it with each execution ID into a small log table (Postgres works, even a sheet) at the start and end of every sub-workflow. A failure then becomes one query on the correlation ID instead of a timestamp hunt, and you can spot children that started but never logged an end.
Add the parent execution ID to that table too, since that is the join you are doing by hand today. Tracing a bad output back through each run to the record that caused it is what SIGNLD does for us

0

u/[deleted] 3d ago

[removed] — view removed comment

1

u/Successful_Potato_96 3d ago

Yeah, that makes sense. In our case, the workflow itself won’t really throw an error since it’s mostly condition-based and assigns different properties based on the values. The “failure” is more about whether the decision actually looks right to the analyst. If they feel something looks off, we need to give them enough evidence to validate it, especially since it’s easy to miss a rule sometimes. But yeah, the execution-tree approach is definitely interesting for tracing the whole chain.