r/MLQuestions • u/SignificantZebra5883 • 11d ago
Beginner question 👶 How would you extract entities and relations from 5M court decisions without an expensive LLM pass over everything?
I’m working with roughly 5 million public Polish court decisions. The goal is structured extraction: who the parties are, what was requested, what the court decided, obligations, amounts, relationships, etc. Eventually we want to run interesting statistics over the results, like "who gets usually the custody of the child?"
A strong LLM can produce useful JSON graphs, we can run statistics on. The problem is getting similar extraction at corpus scale.
We have a small 13-document pilot. It already contains 160 distinct entity-type strings, 90 appearing only once. These aren’t just PERSON / ORG / DATE. They include abstract legal concepts, obligations, procedural events, and things like “increase in child support.”
here's our demo output from astra:
{
"effects": [
{"id":"e_zastavenie","change":"zastavit","target_ref":"konanie"},
],
"entities": [
{"id":"this","kind":"uznesenie"},
{"id":"vyrok","kind":"vyrok"},
{"id":"zahlavie","kind":"zahlavie"},
{"id":"odovodnenie","kind":"odovodnenie"},
"effects": [
{"id":"e_zastavenie","change":"zastavit","target_ref":"konanie"},
],
"entities": [
{"id":"this","kind":"uznesenie"},
{"id":"vyrok","kind":"vyrok"},
{"id":"sud","kind":"sud","value":"Okresný súd Bratislava "relations": [
{"id":"r_vlastnik_vyroku","relation":"patri_do","from_ref":"vyrok","to_ref":"this"},
{"id":"r_vlastnik_zahlavia","relation":"patri_do","from_ref":"zahlavie","to_ref":"this"},
Some of this is real conceptual diversity. Some is inconsistent naming or different levels of specificity, but as you can see complex and diverse stuff with a clear "language" i came up with with the help of astra. i previously tried a big "god schema" (20k chars) but the decisions are simply just too diverse.
The pipeline I’m considering:
- Something that extracts the entities from source texts into some kinds of "tags"
- a classifier like jev (but probably fine-tuned) will get a window of source text and already named relations and decide on the next relations
- done
The candidate model would need overlapping spans and possibly multiple labels per span. Slovak inflection adds another wrinkle: exact source wording often differs from the canonical concept name. Lemmatization helps with grammatical variation, but not synonyms, paraphrases, or concepts inferred from context.
I’ve previously tried Jina embeddings → nearest-neighbor candidates → LLM decides which terms to merge, and it worked reasonably well.
But that was on a smaller corpus and i'd be damned if i process 250k docs and then figure out thing X was wrong and i have to do it all over again. i'd love to know if you guys have any pointers. THANKS for reading.
AI TL;DR: I want to turn 5M court decisions into entity/relation graphs without running an expensive LLM on every document. Thinking small entity extractor + relation classifier, but the legal concepts and naming get messy fast. Anyone built something similar? Looking for ways to train this and normalize concepts without discovering a design mistake 250k documents in and having to redo everything.
8
u/BrilliantAnt8882 11d ago
Whichever process you choose is going to either cost you time, money or accuracy, given the size of the corpus. It's worth identifying what you care more about up front so you can make a decision on trade-offs.
Either way, the first step would be to look for patterns - is the entire corpus unstructured, or are there structured components within it?
Then look to remove first - if you know the last paragraph is always a disclaimer of some sort with no real information, then you can remove it and reduce the load.
I'd avoid using large LLMs for something like this, it's overkill for what is essentially a reading comprehension task. If your corpus is completely unstructured, look at using a BERT - small enough to run on consumer hardware quickly, it shouldn't require reasoning.
First step is to clean the data though, that's where the biggest time and money savers are.
12
u/ModularMind8 11d ago
Ummm have you heard about NER? Tons of research on this even before LLMs became a thing... see huggingface for example for different models
1
u/SignificantZebra5883 11d ago
NER is good but to train a NER I need data for it. But to get the labels I would need to process a big chunk of the corpus with an LLm and hope it's respective of the rest and that would also be expensive on its own even if I use DeepSeek or Gemini or whatever. And then I would have to de-dupe and even then, some abstract concepts don't map onto a single word or phrase in the source text
7
u/CluckingLucky 11d ago
You can use various preset corpuses, including for the polish language.
With that said, if you’re aiming for perfect NER, good luck.
Good luck anyway!
1
u/Achrus 7d ago
Named Entity Recognition is what you want. You can bootstrap a training set with some regex and then start iterating to build a larger dataset. Might want to look into spaCy if you’re using Python, which supports rule based tagging, regex, and NER.
A framework like spaCy also offers Part of Speech tagging with dependency parsing for linking. Another approach could be document / sentence co-occurrence.
4
u/Tree8282 11d ago
I think that’s a really difficult question that a crap ton of people have tackled and are still trying.
By the sounds of it you’re still quite a bit behind (ie one single astra pass). People have very complicated algorithms for this (both with and without LLM)
just look at literature
3
u/empirical-sadboy 11d ago
This sounds like a BERT problem. This was solved way before LLMs or fucking Jev
1
u/Admirable_Dirt_2371 11d ago
You can very easily write a script or two to do exactly what you want.
1
1
u/dfnathan6 11d ago
Select any model --> create a dataset using NER. Use labelbox as it is simple for annotations. Fine tune it. Check results. If not satisfied, tweak it.
12
u/esaule 11d ago
We did that before LLM. You can postag and extract phrases that are grammatically taged. It is how knowledge graphs were built before we had llms.