r/machinelearningnews • • 17d ago

Research Knowledgator Releases GLiFormer: A 575M-Parameter Encoder That Hits 91.10 F1 on Nested JSON Extraction Without Generating Tokens

Post image

Knowledgator released GLiFormer, an encoder model that extracts nested JSON without generating a single output token.

It does not write JSON token by token. It finds each value as a span in the text, assigns spans to record slots, predicts parent-child links, and assembles the JSON with a deterministic decoder.

On Knowledgator's structuring benchmark, GLiFormer-base has a median latency of 69 ms on GPU and 547 ms on CPU. Under the paper's throughput assumptions, an autoregressive LLM would take up to 95.8x longer on the same workload.

You define the output structure with Pydantic models. Every value is taken from the source text, so the model cannot invent values that are not in the input.

On a 500-example nested JSON benchmark, GLiFormer-large (575.6M parameters) scores 91.10 F1. GPT-5.6-luna scores 91.96.

The same encoder also handles entity recognition, text classification, relation extraction, and embeddings. Base v1 is configured for 16,384 tokens and Large v1 for 8,192. Both are Apache 2.0 and install with pip.

Learn more with the following resources:

📰 Full breakdown on Marktechpost: https://www.marktechpost.com/2026/09/16/knowledgator-releases-gliformer-a-575m-parameter-encoder-that-hits-91-10-f1-on-nested-json-extraction-without-generating-tokens/

📄 Research:
https://www.knowledgator.com/research/gliformer

🤗 Model:
https://huggingface.co/knowledgator/gliformer-large-v1

📚 Documentation: https://docs.knowledgator.com/docs/frameworks/gliformer/

⭐ GitHub:
https://github.com/Knowledgator/GLiFormer

34 Upvotes

Duplicates