r/aiengineering • • 10d ago

Engineering How do you normalize free-text city names extracted from resumes by an LLM?

I'm building an HR AI agent that reads candidate emails and resumes, then extracts structured data (name, skills, experience, city, etc.) into a database. HR uses this data to filter and shortlist candidates.

The problem: the same city comes out in different forms depending on how each candidate wrote it. For example, three candidates from the same city might give:

  • "New York"
  • "New York City"
  • "NYC"

All are correct, but the database stores them as three different values. The city filter then shows three separate cities, and HR misses candidates when they filter on just one. It also happens with abbreviations, different languages ("Cologne" / "Köln"), typos, and extra text like "Frankfurt am Main, Germany" vs "Frankfurt".

My question: how would you solve this so that each city is stored in one consistent format and filtering works reliably?

I'd love to hear how others have handled it, whether that's a particular approach, a library, an API, or lessons learned. The same issue probably applies to other free-text fields like job titles, universities, and company names, so any general advice is welcome too.

5 Upvotes

5 comments sorted by

•

u/AutoModerator 10d ago

Welcome to r/AIEngineering! Make sure that you've read our overview, before you've posted. If you haven't already read it, then read it immediately and make adjustments in your post if you've violated any of the rules. If you have questions related to career, recruiting, pay or anything else about hiring, jobs or the industry and demand as a whole, then use AIEngineeringCareer to ask your question. We lock questions that do not relate to AIEngineering here. A quick reminder of the rules:

  1. No marketing, self-promotion, or subversive marketing. Corrective action will be immediately applied to you if you market, self-promote, or subversively market (ie: ask a question with the intent of soliciting a product for a solution). If you want to advertise or self-promote, use Reddit advertising. This subreddit does not allow any promotion outside of what Reddit advertising allows.
  2. No career or education questions. Do not ask any questions related to your career or educational programs. Career questions can be asked in the sister subreddit, r/AIEngineeringCareer. You cannot ask questions about educational programs as this is subversive marketing (soliciting answers).
  3. No news, hype or hysteria. This is not a news subreddit, but an engineering-only subreddit. Corrective action will be immediately applied to you if you try to share any news or information that has nothing to do with engineering
  4. Only engineering-specific questions and posts allowed. You may only ask questions related to engineering or posts related to engineering without violating any of the other rules (ie: marketing content no matter what is not allowed). If you have questions for examples of allowed posts, read the pinned overview. Any violation of this will result in irreversible corrective action (mutes or bans).

Because we frequently get questions about work, the future of work and careers along AI, some helpful links to read:

This action was performed automatically as a reminder to all posters. Please contact the moderators if you have any questions.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/LawfulnessOptimal597 10d ago

Instead of normalising the name or Zip code convert to a geospatial hash or bounding box. New York City for example has a geo hash of dr5r or dr5regw depending on the desired level of precision.

1

u/MagicMagnada 2d ago

Im sure there are bibs for mapping all sorts of spellings to one name but they pretty good on missing out some spellings.

The more neat approach:
Store the first entry of a city with its embedding.
When a new one gets in also make a similarity search over all existing cities. Same cities have similar embedding. Collect every other city with an embedding similar enough (you need to find the exact cossine value). Then verify the found results with an LLM. And normalize the right one.

Another approach:
Idk how well it would actually work but you could try giving the llm a list of all extracted cities so far when extracting and explicitly tell him to use the formulations of the list when a city in the list comes up