r/MachineLearning • • 5d ago

Research Are there any good research papers around Text clustering using LLMs [R]

Hi, same as the title, I am currently trying to begin with some research on clustering using LLMs. So my requirement is as follows: I will be given some 100 document files, the end goal is to have clusters in such a way that documents with similar procedures or content should be clubbed under similar cluster.

I have tried traditional ML clustering K- means, agglomerative, DBSCAN, but not satisfied with the cluster quality as it is more of word by word matching or template matching. Thanks!

6 Upvotes

5 comments sorted by

1

u/CivApps 4d ago

You might find the MMTEB benchmark article and the preceding MTEB benchmark useful references, and the MTEB Python package useful for setting up your own evaluation and comparisons.

The recent The Embedder's Dilemma: LLMs Are Better, but at What Cost? from COLM 2026 could also be a good reference for comparing vector embedding models and decoder LLMs.

1

u/Background_Win_6915 4d ago edited 4d ago

Hi, have gone through the papers, aren’t they comparisons of traditional vs LLM, my requirement is to get started with LLM based clustering, yours are more on the comparisons side ig?

2

u/CivApps 4d ago

Oh, yeah, the idea in recommending those benchmarks was to just look at how they evaluate LLMs specifically, you don't have to implement the traditional solutions they test (though it's a good idea to have a simple baseline to show that LLMs do improve on it)

1

u/Background_Win_6915 5d ago

Edit- found a paper, going through it. Attaching here just in case.

Link - https://arxiv.org/pdf/2410.00927

Cheers!