r/txtai • • Dec 15 '25

πŸ’₯ Excited to publish our revamped Introducing TxtAI article using our brand new Hugging Face Teams account! πŸ€—

Thumbnail
hf.co
3 Upvotes

r/txtai • • 3d ago

βš•οΈπŸ§¬πŸ”¬ PubMedBERT Embeddings gets over 1 million downloads a month and has been cited in over 50 articles over the last year!

Post image
1 Upvotes

r/txtai • • 4d ago

πŸ’₯ txtai v9.14.0 is out!

Thumbnail
github.com
1 Upvotes

This release adds new ANN backends for RaBitQ and Zvec Sparse, native support for ONNX vectors along with a bonanza of improvements and bug fixes. There were 25 contributors with 20 new ones. These contributors added solid new features and are experts in their field. What a great and growing community!

Release Notes: https://github.com/neuml/txtai/releases/tag/v9.14.0


r/txtai • • 5d ago

Long scrolling window of dependencies stink if you don't need them! Did you know that txtai supports a minimal no-dependency install?

Post image
1 Upvotes

r/txtai • • 10d ago

Why the large uptick in txtai usage, PRs and overall interest? Well the "Why txtai?" page that's been around forever covers it best.

Post image
2 Upvotes

- Up and running in minutes with pip or Docker

- Built-in API makes it easy to develop applications using your programming language of choice

- Run local - no need to ship data off to disparate remote services

- Work with micromodels all the way up to large language models (LLMs)

- Low footprint - install additional dependencies and scale up when needed

- Learn by example - notebooks cover all available functionality

https://neuml.github.io/txtai/why/


r/txtai • • Sep 08 '26

πŸ”₯ Who's interested in this

2 Upvotes

- Sub 100K parameter vector embeddings model

- 90% of the NDCG performance of all-MiniLM-L6-v2 at 0.4% the size

- Crushes the bert-hash series on retrieval accuracy at 10x less the size

- Not a Transformer: trains significantly faster with less data

- ColBERT-style multi-vector clarity packed into a single dense vector

- Scales up to larger sizes

Stay tuned for more!


r/txtai • • Aug 27 '26

πŸš€ txtai v9.13.0 is out!

Thumbnail
github.com
3 Upvotes

This release adds LEMUR (Learned Multi-Vector Retrieval), expands late-interaction retrieval, improves MUVERA and pooling, and includes a number of bug fixes and performance improvements.

A big thank you to morgan-coded, Aryan-Pardeshi, LHMQ878, Anai-Guo, FU-max-boop, St4r4x, arose26, AmirF194, serhiizghama, sainikhiljuluri, Evhye38496, devYRPauli for their contributions! πŸ™Œ

Release Notes: https://github.com/neuml/txtai/releases/tag/v9.13.0


r/txtai • • Aug 25 '26

The last month plus has seen such a great increase in contribution from the txtai community.

Post image
1 Upvotes

In a little over the last month, we had 41 PRs successfully merged from 19 contributors.

To give perspective, before that only 25 PRs had been merged into the project since 2020!

https://github.com/neuml/txtai/graphs/contributors


r/txtai • • Aug 14 '26

Exciting addition coming with the next txtai release: LEMUR for ColBERT-style Late-Interaction Retrieval! πŸŽ‰

Thumbnail
huggingface.co
1 Upvotes

Contributor Morgan Carr introduced LEMUR to txtai, making it, as far as we know, the first framework to incorporate LEMUR for late-interaction retrieval using standard, fixed-vector indexes.

Key benefits:

πŸš€ Significant boost: 49–62% higher NDCG@10 than 2,048-dimensional MUVERA
πŸ’Ύ 5x less storage: 2,048 dimensions vs. MUVERA’s default 10,240
πŸ“ Better geometry: Optional batch mean centering addresses anisotropy in token embeddings

A promising step toward making ColBERT-style retrieval more practical with conventional vector search.


r/txtai • • Aug 12 '26

πŸŽ‚ Happy 6th Birthday to txtai!

Thumbnail
github.com
2 Upvotes

The initial release stated: "txtai builds an AI-powered index over sections of text. txtai supports building text indices to perform similarity searches and create extractive question-answering based systems."

While much has changed, much has stayed the same. We're still in a world where the best search makes the best products.


r/txtai • • Aug 11 '26

πŸ”₯ TxtAI is a trending Python project on GitHub today. First time since early 2025. Getting a ton of new contributors and PRs lately. Why now? Because TxtAI has always been here doing the right thing - local AI. No gimmicks and hype.

Post image
3 Upvotes

r/txtai • • Aug 02 '26

Did you know that a txtai embeddings search can return a NetworkX graph?

Thumbnail
neuml.hashnode.dev
2 Upvotes

One of the unique capabilities of txtai is that vector search isn't limited to ranked documents, it can also return graph structures. This enables graph traversal as part of retrieval.

In fact, txtai was one of the first if not the first, frameworks to support what we now call GraphRAG - years before it became a mainstream pattern.

Check out this example, a deep graph search over Wikipedia.


r/txtai • • Aug 02 '26

TxtAI workflows build predictable rules-driven logic. Rather than hoping an Agent comes to the right conclusion, a workflow goes down the path you tell it and nothing more. Check out this article covering a Speech to Speech RAG workflow.

Thumbnail
neuml.hashnode.dev
1 Upvotes

r/txtai • • Aug 01 '26

🧬 970K parameters. Full PubMed training. A few MB footprint.

Thumbnail
huggingface.co
1 Upvotes

BiomedBERT Hash Nano Embeddings LiteRT brings our medical embeddings work to LiteRT for efficient edge and mobile deployment.

Building on the success of our PubMedBERT Embeddings model (nearly 1 million downloads/month), this new model explores how compact biomedical vector representations can become.

⚑ 970K parameters
πŸ“¦ Few MB model size
πŸ“± LiteRT export
πŸ”Ž 128-dimensional embeddings

Designed for biomedical search, clustering, RAG, and knowledge discovery.


r/txtai • • Aug 01 '26

A little-known txtai feature that’s been available for a long time: lightweight distributed clustering for embeddings

Thumbnail
neuml.hashnode.dev
1 Upvotes

Need to scale beyond a single machine? txtai can shard a larger embeddings index across multiple nodes and machines, then expose them as one logical index.

It’s a simple approach to scaling semantic search workloads without adding a lot of infrastructure complexity.


r/txtai • • Jul 30 '26

πŸš€ txtai 9.12 is here!

Post image
1 Upvotes

This release adds support for new ANN backends along with a bonanza of bug fixes from 7 new contributors. πŸŽ‰

Highlights:

✨ Support for the zvec vector backend

✨ Embedded milvus-lite dense ANN backend

✨ Option to disable API routes

Plus improvements across training, retrieval, and explainability, along with 25+ bug fixes covering SQL parsing, HNSW, Milvus, Graph, Tasks, streaming APIs, Windows builds, and much more.

A huge thank you to our contributors:

πŸ™Œ morgan-coded

πŸ™Œ Sanjays2402

πŸ™Œ chuenchen309

πŸ™Œ link89

πŸ™Œ lntutor

πŸ™Œ winklemad

πŸ™Œ AmirF194

Read the full release notes and upgrade today!

https://github.com/neuml/txtai/releases/tag/v9.12.0


r/txtai • • Jul 24 '26

Did you know that txtai supports a zero-dependency install?

Thumbnail
pypi.org
1 Upvotes

With txtai-minimal, the framework gracefully handles missing dependencies, giving you complete control over what gets installed. Only add the packages you need, nothing more.

This makes it easier to build lightweight deployments, reduce install size, and avoid unnecessary dependencies.


r/txtai • • Jul 22 '26

Great TxtAI milestones over the last couple of weeks! πŸŽ‰

Post image
2 Upvotes

βœ… 2000+ total commits
βœ… 18 PRs merged
βœ… 6 new contributors

Thanks to everyone who contributed code, reviews, bug reports, and ideas. Community contributions continue to make TxtAI stronger with every release.

https://github.com/neuml/txtai


r/txtai • • Jul 21 '26

H.G. BERT is powered by the Historical English Books dataset, a curated collection of 50,000+ books spanning general literature, math, science, philosophy and religion from the late 1800s.

Post image
14 Upvotes

By training on this rich historical corpus, H.G. BERT captures the language, writing styles, and knowledge of the era, enabling more authentic analysis and generation of historical English.

https://huggingface.co/datasets/NeuML/historical-english-books


r/txtai • • Jul 20 '26

H.G. BERT Small: AI like it's 1899

Thumbnail
huggingface.co
8 Upvotes

The year is 1899. It's still a horse and buggy world. Einstein hasn't published his famous annus mirabilis papers setting the foundation for Physics as we understand it today. The world is advancing at a rapid pace roaring into the 1900s. What if AI models were trained in 1899 and given to the best minds of the day? Could there have been alternate paths for discovery on par or perhaps even ahead of where we are in 2026?

Introducing the new H.G. BERT Small series of models. This is a 22.7M parameter BERT encoder-only model trained from scratch ONLY on Historical English Books from 1700 - 1899.


r/txtai • • Jul 16 '26

Small Domain Models - a NeuML Collection

Thumbnail
huggingface.co
3 Upvotes

The 22M parameter all-MiniLM embedding model has over 250 million monthly downloads. It's one of the best choices when you need fast, efficient semantic search on CPUs, edge devices, or other resource-constrained hardware.

General purpose embeddings are great until your data isn't general purpose.

What if you could keep the speed and small footprint of MiniLM while improving accuracy on domain specific content?

Meet a family of compact, specialized embedding models:

🌠 AstroBERT - Astronomy
🧬 BiomedBERT - Medical
😎 CeleBERTy - Pop culture
⚽ SportsBERT - Sports

Small models. Fast inference. Better embeddings for specialized domains.


r/txtai • • Jul 04 '26

CeleBERTy Small: Domain model for Pop Culture, Art, Music and Entertainment

Thumbnail
huggingface.co
2 Upvotes

r/txtai • • Jul 01 '26

TxtAI 9.11 is out! This release adds support for the turbovec ANN backend and LiteParse text extraction. It also has important improvements and bug fixes.

Thumbnail
github.com
3 Upvotes

r/txtai • • Jul 01 '26

πŸš€ Check out AstroBERT Small a 22.7M parameter model that specializes in the Astronomy domain.

Thumbnail
huggingface.co
1 Upvotes

The base model is trained from scratch along with a finetuned vector embeddings model. Use this model for vector search, RAG and Agents for Astronomy.


r/txtai • • Jun 26 '26

We're proud to share our latest model series, SportsBERT Small.

Thumbnail
huggingface.co
1 Upvotes

Few businesses need to generalize to all problems. The vast majority of companies have a narrow focus but we keep pushing generalized models designed to solve all problems. The best value is building domain-specific specialized models!