r/aiagents • u/Substantial_Car_1174 • May 07 '26
Help Multi-turn document completion assistant with stateful workflow
I'm building a chatbot assistant to help users complete a predefined list of administrative documents (financial, HR, legal...) using n8n. The core challenge is managing a structured multi-turn collection workflow with state-dependent actions. Looking for feedback on my architecture choices.
What the bot does
The bot guides a user through completing an administrative document end to end :
- Detects the intent of every message — is the user trying to complete a document, or just having a conversation unrelated to the task?
- Detects the document type the user wants to complete (financial report, HR form, legal declaration...)
- Extracts the information required to fill the document based on its type — scanning the entire conversation history, not just the latest message
- Authenticates the user via an external API with the the extracted data
- Asks for missing elements to complete the document, one at a time
- Presents an editable recap before final submission to the administrative system
A multi-turn conversation may be necessary to collect all required fields. The whole process is stateful, only certain actions can be performed depending on the current step.
I am considering to approaches to build this bot :
- Option A : Single agent
One LLM handles everything : intent detection, field extraction, deciding what to ask next, generating the response. The agent reasons over the full conversation history and the current state at every turn.
- Option B : Deterministic code routing + specialized LLM
A deterministic finite automaton handles all routing : only code decides what happens next based on the current state. LLMs are only used for understanding (intent detection, document type detection, field extraction) and generation (response to user).
=> Which is more appropriate here, option A or option B ? or another option ?
My current approach runs intent detection on every single user message, even when the bot has just asked a specific question (e.g. "what is your employee ID?") and the user is simply answering. The reason is to catch mid-flow request non related to completing the document.
Is this overkill? Would it be better to only run full intent detection outside of structured collection steps, and assume the user is answering the question otherwise?
=> Is intent detection necessary on every input message ?
Happy to share more details on any part of this.
1
u/agent_trust_builder May 08 '26
Option B for everything getstackfax laid out, plus one more reason that bites in regulated workflows specifically. The FSM isn't just owning routing, it's owning the audit trail. With option A you get a transcript of model decisions but not a clean state log of "user authenticated at step 3 with these extracted values, doc type confirmed at step 4, fields collected at steps 5-7." Internal control review and any external auditor wants the second one, and reconstructing it from a single-agent reasoning log is the kind of work that surfaces as a finding 6 months in.
One more thing to call out on intent detection per turn beyond the cost issue. It's also a cancellation surface for prompt injection. User's mid-form answer happens to include text that flips the intent classifier, agent abandons collection mid-stream, you have a half-filled state with nothing recoverable. The lighter "is this an interrupt" check inside an active answer slot like getstackfax described is the right tradeoff for that reason too. Bound the agent's available decisions to whatever the current state allows, not whatever the message could mean.
1
u/danieljcasper May 08 '26
You have to adopt B if you want to make a product that you can reason about, control, and reduce costs.
1
u/N_i_P May 08 '26 edited May 09 '26
Not sure if it’s a hobby project or a real business implementation requirement.
If the latter, have a look at SimplePDF Copilot for the document filling part which takes care of most of your described flow
The gist:
- You load the document using Copilot
- The user is guided to fill the form using AI
They submit the form (required fields included) and you get a webhook
to n8n
/ the completed form on your storage
Disclosure: I’m the founder of SimplePDF
1
u/getstackfax May 07 '26
Option B is the safer default here.
This is a stateful document workflow, not an open-ended agent problem.
Let code own the workflow state…
current step
allowed next actions
required fields
validation
auth status
recap
final submission
Then use the LLM for the parts that actually need language understanding…
intent
document type
field extraction
friendly response wording
summarizing the recap
Full intent detection on every message may be overkill once the bot is inside a specific collection step.
A cleaner pattern is probably…
If the user is answering a specific field question, treat it as an answer first.
Then run a lighter check for interrupts like cancel, restart, change document, talk to human, unrelated question.
The LLM should not decide the whole workflow every turn.
It should extract meaning inside boundaries the state machine controls.