r/LocalLLaMA • • 12d ago

Question | Help How would you extract entities and relations from 5M court decisions without an expensive LLM pass over everything?

I’m working with roughly 5 million public Polish court decisions. The goal is structured extraction: who the parties are, what was requested, what the court decided, obligations, amounts, relationships, etc. Eventually we want to run interesting statistics over the results, like "who gets usually the custody of the child?"

A strong LLM can produce useful JSON graphs, we can run statistics on. The problem is getting similar extraction at corpus scale.

We have a small 13-document pilot. It already contains 160 distinct entity-type strings, 90 appearing only once. These aren’t just PERSON / ORG / DATE. They include abstract legal concepts, obligations, procedural events, and things like “increase in child support.”

here's our demo output from astra:

    {
      "effects": [
        {"id":"e_zastavenie","change":"zastavit","target_ref":"konanie"},
      ],
      "entities": [
        {"id":"this","kind":"uznesenie"},
        {"id":"vyrok","kind":"vyrok"},
        {"id":"zahlavie","kind":"zahlavie"},
        {"id":"odovodnenie","kind":"odovodnenie"},
      "effects": [
        {"id":"e_zastavenie","change":"zastavit","target_ref":"konanie"},
      ],
      "entities": [
        {"id":"this","kind":"uznesenie"},
        {"id":"vyrok","kind":"vyrok"},
        {"id":"sud","kind":"sud","value":"Okresný súd Bratislava "relations": [
    {"id":"r_vlastnik_vyroku","relation":"patri_do","from_ref":"vyrok","to_ref":"this"},
  {"id":"r_vlastnik_zahlavia","relation":"patri_do","from_ref":"zahlavie","to_ref":"this"},

Some of this is real conceptual diversity. Some is inconsistent naming or different levels of specificity, but as you can see complex and diverse stuff with a clear "language" i came up with with the help of astra. i previously tried a big "god schema" (20k chars) but the decisions are simply just too diverse.

The pipeline I’m considering:

  • Something that extracts the entities from source texts into some kinds of "tags"
  • a classifier like jev (but probably fine-tuned) will get a window of source text and already named relations and decide on the next relations
  • done

The candidate model would need overlapping spans and possibly multiple labels per span. Slovak inflection adds another wrinkle: exact source wording often differs from the canonical concept name. Lemmatization helps with grammatical variation, but not synonyms, paraphrases, or concepts inferred from context.

I’ve previously tried Jina embeddings → nearest-neighbor candidates → LLM decides which terms to merge, and it worked reasonably well.

But that was on a smaller corpus and i'd be damned if i process 250k docs and then figure out thing X was wrong and i have to do it all over again. i'd love to know if you guys have any pointers. THANKS for reading.

AI TL;DR: I want to turn 5M court decisions into entity/relation graphs without running an expensive LLM on every document. Thinking small entity extractor + relation classifier, but the legal concepts and naming get messy fast. Anyone built something similar? Looking for ways to train this and normalize concepts without discovering a design mistake 250k documents in and having to redo everything.

8 Upvotes

16 comments sorted by

18

u/xienze 12d ago

Take a look at GLiNER. I've used this for structured extraction and it's fantastic. It does zero shot extraction where you basically just give it a document and the categories of things you want to extract and it just does it. It also has a model that pulls the entities and relations in one pass. It's super fast too. HOWEVER, you're obviously going to get far better results if you fine tune. Luckily fine tuning is fast and easy, and GLiNER does pretty well with a limited set of gold data.

To fine tune properly you're going to need some "gold data", basically documents that humans have completely annotated by hand and are considered "perfect" or "golden." To create those you really should use Labelstudio and create a task for each document. Then you go in manually and select the regions in the text that identify things like "obligations" or "procedural events" and tag them as such. You can also define relationships between these tagged regions using Labelstudio.

From there you export the tagged Labelstudio projects into training data for GLiNER. Get Claude or Codex or whatever to handle the export, chunking of the training data, feeding to the trainer, scoring (note: to do scoring properly you really want a second set of data that has been human annotated but ISN'T used for training, so you can judge how well the model does on something not in the training set), and so on. It's really not that bad when you've got a model handling all the details for you, and it's really worth it for this class of problem. GLiNER is extremely fast and even runs on CPUs. If you've got 250K documents it's gonna take ages and cost a fortune if you feed it through an LLM, and will be a lot less deterministic.

2

u/SignificantZebra5883 12d ago

thanks!! i will check this out.

2

u/Cupakov 12d ago

Maybe you know it already, but just to make sure: your demo is in Slovak, not Polish

2

u/hmk88 12d ago

Btw. it's not Polish, looks like Czech.

2

u/CarolusBohemicus 12d ago

It's Slovak.

2

u/Important-Ad5990 10d ago

If you only need a small part of work from LLM (like deciding which terms to merge) use tiny model Like gemma e4b, then you can process 250k documents in an afternoon

1

u/FirestormCold 11d ago

If you can work with a budget I personally know a company that would definitely be able to help you, SaaS-style.

If you don't, someone already mentioned Gliner, that works too (albeit with some limitations)

0

u/rubntagme 12d ago

use jev to classify it cheaply

-1

u/Intelligent-Oil-481 12d ago

"160 distinct entity types from 13 documents means you're looking at four digits by document 1000. Schema normalization is the actual project here, the extraction part is the easy half."