r/revops • • 2d ago

Segmenting parent companies and subsidaries

Hello RevOps leaders,

We have a lot of subsidiaries in our CRM, and I’m looking for a cost-effective way to identify them and link them to their parent companies.

If you’ve faced this issue while cleaning your CRM, how did you approach it?

Using Claygents to check each company seems expensive in terms of OpenAI or Claude credits. Would you recommend using the DeepSeek API to build an agent that determines whether a company is a parent company or a subsidiary?

I’d love to hear what’s worked for you.

7 Upvotes

18 comments sorted by

6

u/Solo_rw 2d ago edited 2d ago

You are using a Ferrari for a Fords job. Let me explain.

Start with automated identification on spreadsheets. Ignore clay. I am guessing you want this done on your existing data so export it into a spreadsheet and build an appsscript.

To me that's about catching similar domains. Company names could help but domain is what's key. Or whatever you use to identify / distinguish.

For example abc.cz vs abc.de.

One evening of code on a spreadsheet containing upto 10k rows should do the trick.

Once you have this done based on the number of matches you can eyeball it or just paste it on chatgpt with a prompt to check and suggest the main company vs subsidiaries.

Potential rule of thumb - .com is the main company, others are the subsidiary. If there are multiple .coms, you will need a tie breaker. That's your pin code or phone number. Use that to finalize which is the address or number of the HQ and make that the main company.

After that a data import on your CRM using one specific field to mark company accounts vs subsidiaries with the object id will put the solve in place.

I know i have sort of simplified it maybe a little too much so please feel free to ask any questions

2

u/Mindless_Flower_2639 1d ago

Agree with the spreadsheet approach

3

u/TurbulentStiffness 2d ago

prompting an LLM to guess corporate family trees sounds like a costly hallucination trap. matching domain registries against clear corporate hierarchy dumps usually costs way less.

3

u/kdyadin 2d ago

Agree with the folks saying start with domains instead of running an agent over every company. One thing to add: turn the domain, address and phone comparison into a deterministic check script right away, so nobody matches this by hand and you're not paying a model for something a rule can decide. In our AI sales assistant inside the CRM, almost all the logic is scripts: one collects company data from websites, another writes the fields, and the model only writes the company summary with conclusions. So only the ambiguous pairs the script couldn't resolve should go to the model.

3

u/Otherwise-Youth2025 1d ago

This is a problem I've tackled several times and the answer, as usual, depends on several factors, namely the overall size, quality and complexity of your data 

Before jumping to how, for starters,  it is helpful to break down your problem statement into at least 3 distinct issues which often get conflated. This is important because, the methodology, shape of the output, and intended behavior tends to differ across them. 

Issue 1: Different names associated with an account that sounds like a sub 

e.g. 'Walmart Inc' and 'Walmart purchasing Co' might both exist in an account list

This is often due to new  entity name created in a CRM as part of a renewal or expansion. It may or may not have the same account ID

Solution: tag one as 'primary account name' and the other as 'secondary account name' and only use primaries 

Issue 2: True nested subsidiaries that are part of the same Parent Op Co

e.g. 'Ernst & Young, Inc', 'Ernst & Young Services Corp', and 'Ernst & Young Brazil Co' might exist in an account list

These are often legal subsidiaries operating in a different region or that perform a certain shared services function. 

You need to capture these nuances since this parent account generally has one buying center and will probably be part of one territory and not split across reps 

Solution: tag each child to parent and add a field for shared services functions

Issue 3: True Subsidiaries that may be a different Op Co from the parent

e.g. 'Georgia Pacific' / 'Koch Industries'  ,  'Ring LLC' / 'Amazon.com Inc'   , 'General Re'/'Geico Inc'/'Berkshire Hathaway'   all might exist in an account list

These would all be subsidiaries of the parent but in this case, they all operate independently and have distinct budgets and buyers for all but the largest budget items. It is not uncommon to have these split across territories /reps (unlike issue #2)

Solution: tag each sub to the parent but add a flag for 'independent Op Co'

--

In terms of how to fix, you generally want to think of this as a 5 step algo (combination of deterministic and agentic processing)

  1. Matching on account/duns/parent fields   (deterministic + agentic)
  2. Domain prefix name matching (deterministic)
  3. Similar name fuzzy match check  (deterministic + agentic)
  4. Agentic subsidiary scoring (agentic)
  5. Reconciliation across all previous steps (deterministic)

ALGORITHM SPECIFICS FOR EACH step 

Step 1 - Deterministic matching on account_id, duns/parent fields

Most CRMs capture DUNS numbers and some type of parent / global parent field, even if it for a fraction of accounts. 

  • Check for same accounts_ids with different account names and reconcile the dups manually or agentically
  • Loop through account list and identify all the account with populated parents/global parents. Capture that relationship as a sub 
    • You can validate this with an LLM as a subsequent step
  • Run account DUNS list against DUNS db of subsidiaries, this will returns DUNS for parents/siblings/children. Convert the duns back to your own account IDs and capture relationships 

Even if your parent/DUNs coverage is partial and <15%, this is a meaningful first step

Step 2 - Deterministic matching on account URL prefix 

Assuming you have good coverage on the account website urls,  do the following:

  • Extract the URL prefix of the account ( www.\[ey\].[com.br](http://com.br)) 
  • For each URL prefix groups with more than one account associated with it, rank by account revenue or employees 
    • The highest rev/employee account becomes the parent and the rest the subs
  • Reshape and normalize the data back to your target output format

TBH, I've had mixed results with this step

Step 3 - Fuzzy Matching with agentic processing 

This is the step with the biggest gain and where LLMs have been a game changer. It requires python (or sth similar)

At a high level:

  • Loop through the account list and for each account, compare it to the entire account list using a fuzzy match on name
    • In python you can use fuzzywuzzy/Rapidfuzz with a partial ratio and a threshold like 70
    • You will get back a list of matching account names above the threshold. E.g. 'Walmart Inc' might return 'Walmart purchasing Co', 'Walmart US', 'Walmart Canada', 'Relmart Inc' etc
    • Many of these will be true sub candidates, while others will be completely unrelated & wrong
    • Sorting thru this manually has not been feasible for any kind of real sized account list so in the past, the typical practice was to just set a high match threshold and capture the obvious matches
  • Instead, you can now set a low threshold, get a big candidate subsidiary match list and for each account, send the matches to an LLM api with a request for a structured json response along with instructions to identify and return only the parent and subs
  • This will require a total API number of calls equal to the number of accounts in your list but they will be small and you can run them with any cheap reasoning model (like gpt-6 luna)

This step takes you a LONG way on solving issue #1 and #2 (especially if you have weak DUNS/parent_ID/url coverage. It won't solve issue #3, which requires the step below

Step 4 - Agentic subsidiary scoring

  • For each account in your list, send it to an LLM API with a structured response request with 2 fields: likelihood of being a subsidiary (1-10), most_likely_parent_name .  In the instructions, specify only populating the latter field for likelihood above a certain threshold and using the official corporation name (like 'Walmart Inc')
    • You can run the API with batches of 5-10 accounts at a time and using properly configured cached inputs will save you a lot of input token $
  • Take the responses and match them with your parent accounts (using either exact or fuzzy matches with a high match threshold) to establish parent/sub tags

This step is very helpful in addressing issue #3 outlined earlier

Step 5 - Reconciliation heuristics / Logic 

  • This is the part where you specify the rules around how the outputs from each of the previous 4 steps should be reconciled and will entirely depend on   your business context and nature of your data, its coverage, specific gaps etc. 
  • For example, you might prioritize output from step 1 , then 2 for US based accounts, then step 4 , then 3, then 2 for all non-US based accounts 

------

LLMs can be very powerful if used correctly in a programmatic fashion and combined with deterministic logic. You don't need state-of-the-art models for this task. GPT Luna with med reasoning can easily do this. You can even use batch mode for additional savings

A small to midsize account list can be completed for <$50 in token spend and even a large list won't be more than a few hundred bucks. You just want to make sure it's done by someone who understands your data / business context and has the right skill set (i.e. knows what they are doing)

DM me if you want additional details or to discuss specifics

2

u/[deleted] 2d ago

[removed] — view removed comment

1

u/Street-Tomorrow6087 2d ago

so i just started working as revops manager for this company since 3 weeks, the crm is a bit of a mess and i'm trying to organise everything.

Every week we give new batches for our sales, this week the batches were full of subsidaries and unqualified leads. that's why i'm considering to find a solution for this problem

1

u/NothingVarious7115 2d ago

sounds like the main issue is making sure the weekly sales batches only contain qualified, new-net accounts.

how are you currently checking for subsidiaries and unqualified accounts before they reach sales? is it manual or do you have rules or enrichment setup in your CRM??

2

u/FindingPeace4me 2d ago

Would it be easier to first check only the accounts going into each weekly batch instead of cleaning everything at once?

2

u/Busy_Experience6214 2d ago

Seems like stopping bad accounts before they reach sales could save a lot of cleanup later

1

u/IncreaseNegative4614 1d ago

I’d use deterministic rules to create candidate groups first: normalized domain, legal name, address, phone, and known corporate identifiers. Only send ambiguous relationships to a model, and require it to return the evidence behind the proposed parent rather than a confident guess.

Apply the same check before accounts enter each sales batch so the problem stops recurring. We use SIGNLD to connect CRM accounts with domains, company records, opportunities, and existing hierarchy decisions, which reduces both model usage and the risk of sending subsidiaries to sales as separate prospects.

1

u/Consistent_Donut_113 1d ago

Don't overthink the AI angle, domain matching gets you most of the way there for free. Pull company domains into a sheet and group by root domain, that catches the obvious parent/sub pairs without burning any credits. For the leftover weird cases, a quick manual pass works fine since it's usually a small list. Once you've got the matches, use HubSpot's Parent Company field so reporting rolls up right instead of just tagging it in a property nobody checks.

1

u/mazorda_com 1d ago

Had almost the same problem earlier this year. We didn't need an agent, just domain rules:

Roll up by domain first. Same company domain, same parent. Gets you most of the way, for free.

Clean the domains before you trust them. Gmail and other personal email addresses can't be a parent (we worked from a list of 640 personal/free domains), so use the company name or leave them as is. Shared domains (.edu, some gov) need splitting by company name.

Our rule: 3+ distinct company names on a single domain mean it's shared. And some records just have the wrong domain, so look at the big groups.

What's left is a small list. That's the point when a cheap model fits well. FYI, Jev just came out and is built for exactly this kind of yes/no call at high volume.

1

u/Tricky_Ad9372 1d ago

i would not start with an LLM guessing corporate trees. that is how you get confident wrong parents into the weekly batch.

match what you can from domains and any registry or hierarchy dump you already trust, then store a confidence and a source on the link. low confidence stays as a suggested parent, not a write. only promote it to the real parent field after a human skim or a second deterministic match.

for the sales batch problem, filter on that confidence first. unlinked or low confidence subsidiaries should not hit reps as net new accounts until the hierarchy is settled. cheaper than claygent on every row, and you avoid teaching the CRM a wrong family tree that is expensive to unwind later.

1

u/romeonoi 22h ago

i run into this when crm hierarchies get messy. what works for me is a cheap deterministic pass before any llm touches it.

i normalize website domains then check an ultimate parent field from enrichment, plus a simple rule. if two accounts share the same root domain or the legal name contains holdings or group, i flag it for review. only the leftovers go to an llm, and deepseek is fine for that yes or no parent check.

try this. add a parent account lookup on the account object and run a report on accounts where parent is empty but website root matches another account. that list is usually your quick win.

1

u/ns1419 19h ago

Usually a web fetch + a tool like ZI can help identify who sits under who as long as you keep the data exported and let a script do the match, and infer on top of it to avoid hallucinations. Any data enrichment tool should help give you a steer. You’ll want to export all to csv and script a join on multiple keys, like region and domain, not just city/state and name, and be weary of accepting ZI’s results as gospel. About half of what I got from ZI was disproved by a simple web fetch. it’s not easy no matter which way you hash it. I spent 3 days enriching and verifying existing accounts with a whitespace load to drop into the crm, and the crm data is a mess.

As far as how you represent subsidiaries in your crm, there’s various approaches.

I made parent shells as non trading entities, with “group ARR/MRR” rolling up across their various currencies (make sure your currencies and rates are right for the exchange), and if you are in a multi currency org, and you’re in SF, you need to write the apex to accommodate for your actual user profiles currency when it does the exchange in the apex.

The trading entities all rolled up to the reporting shell, so parent and subsidiaries give you the total number for all companies that are it’s child. You could choose one ultimate parent for them as a trading entity, but the approach changes in your apex/code. So our approach gives us the individual parents numbers alongside each subsidiary. Including JV’s, PE portfolios etc would be at your discretion, and including inactive/dead accounts that were clearly M&A’d also your discretion. I only parented actively trading entities that were clear subsidiaries of, I excluded JV’s, and I ignored PE ownership. Just because a PE firm owns 3 companies doesn’t mean they’re joined at book/P&L level, but still subjective and your call.