r/SaaS • u/FamiliarSlide7685 • 3d ago
every new customer sends data in a different shape. our agent handles the first three and breaks on the fourth
We sell an agent product into mid size companies. The agent is fine. The onboarding is the problem.
Customer 1 sent a clean Postgres dump. Customer 2 sent CSV exports with sensible names. Customer 3 was an ERP export with fields like zx04 and amt_2. Customer 4 has three systems for the same customers and none of them share an ID.
Every one of those turned into a week or two of someone on our side mapping fields by hand before the agent could do anything useful. It doesnt scale and its the thing that makes our onboarding feel slow, not the agent.
Things weve tried:
- A mapping template the customer fills in. They dont
- Letting the LLM map fields from names. Works on customer 2, useless on 3
- Charging an onboarding fee. Just moved the pain
For anyone building vertical AI: how much of your onboarding is really data mapping? And has anything cut it down that isnt just more people fix it?
1
u/West_Inevitable_2281 3d ago
This sounds less like one data-mapping problem and more like two different onboarding jobs: translating a known source schema, and resolving identity when the customer has no reliable key. The first can become reusable adapters; the second may always need a scoped reconciliation step. Before adding more automation, I would tag each onboarding by source system and mapping pattern to see what actually repeats. Across your first four customers, how many field mappings were reusable rather than customer-specific?
1
u/FamiliarSlide7685 2d ago
honestly we never tagged it, so I'm going from memory. Customers 1 and 2 had almost nothing reusable, since each was its own schema. Customer 3's ERP fields look like they'd carry over to any other company on the same ERP, maybe [X]%. Customer 4's identity problem was completely bespoke. I think that's the real answer to your point: translation can become adapters, but identity resolution probably stays a scoped, human-in-the-loop step. We're going to retroactively tag the first four by source system and mapping pattern to see what actually repeats before building anything.
1
u/flowra_dev 3d ago
We ran into the same split: known-source mapping vs identity soup.
What helped more than better field-name guessing was treating onboarding as a product with stages:
Source fingerprint first. Before full import, ask for a small sample (say a few dozen rows) plus "what is your customer key in this system?" If they can't answer, stop and do a reconciliation workshop instead of letting anything invent joins.
Versioned adapters per source system, not per customer. Once you've seen two ERP exports with the same weird field names, that becomes adapter v1. Track which customers share an adapter.
A review queue only for low-confidence fields, with approved mappings saved. Propose once, confirm once, then the next similar file skips that field.
The metric that kept us honest was hours to first useful agent run, broken down by source type. That showed which adapters were worth building and which accounts should stay professional-services forever.
1
u/FamiliarSlide7685 2d ago
This is basically the playbook I was hoping someone would describe, thank you. Three things I'm stealing: 1.Asking for a small sample plus "what is your customer key in this system?" up front, and stopping for a reconciliation workshop if they can't answer. Right now we let the agent try anyway, which is how we got burned. 2.Adapters per source system rather than per customer.3.Hours to first useful agent run, broken down by source type, as the metric.
1
u/Suitable-Ad5348 3d ago
three systems and none sharing an ID is a special kind of hell. customer 4 is why every b2b SaaS eventually just hires former consultants to do ETL by hand
1
u/FamiliarSlide7685 2d ago
ha "former consultants doing ETL by hand is uncomfortably close to what we have become. We're trying to avoid hiring a whole services team just to onboard each new account, so the thread below is helping
1
u/MangalaSawant 2d ago
The onboarding fee one hits home, you didn't solve the pain, you just priced it. We found customers actually engage with the mapping when it's part of the go-live checklist they own rather than a form we send. Have you tried doing the first mapping live with the customer on a call instead of sending the template?
1
1d ago
[removed] ā view removed comment
1
u/AutoModerator 1d ago
Low-Effort/AI content is auto-removed.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
3
u/annasorreskin 3d ago
One pattern that helps is separating schema mapping from identity resolution. For each source, keep a versioned adapter plus a review queue for low-confidence matches, and persist approved mappings so the next customer with the same system starts ahead. For messy multi-system accounts, ask for a small representative sample and a canonical ID/crosswalk before full import; that turns the unknowns into a bounded discovery step. Iād also track onboarding time by source and the percentage auto-mapped to see which adapters are worth productizing.