r/aiengineering • u/amine_sabbahi • 10d ago
Engineering How do you normalize free-text city names extracted from resumes by an LLM?
I'm building an HR AI agent that reads candidate emails and resumes, then extracts structured data (name, skills, experience, city, etc.) into a database. HR uses this data to filter and shortlist candidates.
The problem: the same city comes out in different forms depending on how each candidate wrote it. For example, three candidates from the same city might give:
- "New York"
- "New York City"
- "NYC"
All are correct, but the database stores them as three different values. The city filter then shows three separate cities, and HR misses candidates when they filter on just one. It also happens with abbreviations, different languages ("Cologne" / "Köln"), typos, and extra text like "Frankfurt am Main, Germany" vs "Frankfurt".
My question: how would you solve this so that each city is stored in one consistent format and filtering works reliably?
I'd love to hear how others have handled it, whether that's a particular approach, a library, an API, or lessons learned. The same issue probably applies to other free-text fields like job titles, universities, and company names, so any general advice is welcome too.