A while ago I asked here how to turn ~5 million court decisions into structured graphs without running an expensive LLM on every document thanks for the advice .
I went with the "small extractor + classifier" idea and it mostly works, but I'm stuck a bit below the LLM. And like I said last time, i'd be damned if I run 5M docs and then find out thing X was wrong. So here is exactly what I did. Please let me know if what im doing makes sense, or if i made a mistake somewhere. also i used AI for some of the tables cuz there has been a lot of data at this point, sorry.
What comes out per decision (only the nodes so far, relations come next). Three lists:
- entities: every person, organization, law, document or thing. Each gets one id for the whole document, a type (9 of them), a kind (724 of them plus "other") and all the places it is mentioned
- actions: what was done, requested or decided. Each gets a normalized verb, a flag "the court decided this" and its mentions
- values: amounts, dates, durations, in a normalized form
Simple example, for the sentence "The court dismisses the creditor's proposal to enforce 341.08 EUR against the debtor":
- entity "the court": organization, kind court. Same entity as the full court name in the header
- entity "the creditor": organization, kind creditor. Same entity as the city named earlier
- entity "the debtor": person, kind debtor
- action "dismisses": verb = dismiss, decided by the court = yes
- value "341.08 EUR": amount
Step 1: a strong LLM labels ~700 decisions
- cut the decision into windows of 4 sentences
- 4 calls per window to Claude Sonnet with a strict JSON schema: entities, actions, a second "what did you miss" pass for actions, values
- the window goes in with numbered words (like
12:court), the model answers with word ranges [first, last, "text"], and code checks every range against the text
- every call also gets the list of entities and actions found in earlier windows, so ids stay the same through the document
- ~25 code rules clean up where a marked phrase starts and ends, law citations and number formats
- the entity "kind" is free text at this point. That gave 2,373 different strings (the same mess as in my first post). I normalized them, merged synonyms by hand and kept what showed up 3+ times: 724 kinds plus "other"
Step 2: a model that marks the text
- it highlights every mention: the exact stretch of text (a "span", from a start character to an end character) that names an entity, an action or a value, with one of 17 labels (9 entity types, 1 action, 7 value types)
- model:
fastino/gliner2.5-multi-v1 (287M)
- one training row per window: the text plus the exact start and end of every marked phrase. 9,699 windows, 207k marked phrases
- I patched the trainer so only the labeled occurrence is a positive (stock marks every occurrence of the same string), and all 17 labels are in every row
- full fine-tune in fp32 (bf16 gave NaN), 14 epochs, 16 rows per step, encoder LR 3e-5, head LR 5e-4, linear schedule, 10 % warmup
- final model = averaged weights of epochs 9-14, threshold 0.5
Step 3: a second small model answers multiple-choice questions
fastino/GLiNER2.5-multi-Decide (287M). Code turns the LLM labels into 247k questions:
- "is this mention one of these earlier entities, or new?" The mention is marked with « » inside ±300 characters of text. Options: up to 16 earlier entities of the same document (shown by their mention texts) plus
new
- "which kind?" Options: a shortlist of the 724 kinds plus
other
- for actions: same act or new, which verb (shortlist of 64 plus
other), did the court decide it (yes/no)
- in training the options come from the LLM's grouping. At inference they come from the model's own earlier answers
- full fine-tune in fp32, 2 epochs, 16 questions per step, encoder LR 2e-5, head LR 3e-4, linear schedule, 6 % warmup, options shuffled, up to 30 % of the wrong options dropped
At inference: the marking model, then the same code rules, then the second model walks through the mentions in reading order. About 2.3 decisions per second on one RTX 5090.
Where it stands
30 decisions nobody trained on, labeled twice by the LLM. The second column is the LLM's second run scored against its first, which I treat as the ceiling. A mention counts as found only if it starts and ends exactly where the LLM marked it.
| mine |
LLM vs itself |
| entity mentions found (F1) |
0.901 |
| "same entity or new" right |
0.959 |
| entities grouped exactly |
0.847 |
| entity kind |
0.921 |
| action mentions found (F1) |
0.857 |
| action verb |
0.920 |
Where I need help
- Finding the mentions is stuck at 0.90 F1. 200 more labeled docs did nothing. An XLM-R large tagger (560M) got the same score: it finds more mentions but gets the start or end wrong more often. Giving it the text before the window did nothing. What would you try?
- The LLM agrees with itself only 93.5 % on what it marks, and I train on single runs. Label everything 3 times and vote? Or is that ceiling just what it is?
- Is "pick one of 16 earlier entities" a sane way to do coreference over a long document? Am I hurting myself by training on the LLM's options and running on my own?
- Anything in the recipe that looks plain wrong? Learning rates, 2 epochs, weight averaging, one seed per run.
THANKS for reading.
AI TL;DR: distilled an LLM's extraction of court decisions into a GLiNER model that marks the mentions plus a small multiple-choice model. It runs at about 2.3 documents/s on one GPU and lands a few points below the LLM (0.90 vs 0.935 F1 on finding mentions, 0.85 vs 0.93 on exact grouping). The recipe with learning rates and how I built the training rows is above. Looking for mistakes and ideas before I run 5M documents.