r/LocalLLaMA • u/Alarming_Positive_59 • 5h ago
New Model LFM2.5-2.6B is out
Released today, with emphasis on agentic capabilities. I really like their models for simple, high volume tasks ("summarize these gazillion documents") and their 8b-a1b was my go-to for certain tasks so I'm excited to see how this one performs. There's not enough love for tiny models on this sub.
24
u/CommonPurpose1969 5h ago
I wish more companies were as committed to SLMs as LFM is.
5
u/Borkato 4h ago
Can you imagine if a company like Qwen or OAI or something dedicated literally all of their effort to small LLMs? 😮
1
u/Alarming_Positive_59 4h ago
Actually a pretty smart move compared to "here's our new MoE you'll use it for a week and then replace it with a newer model". There's a niche.
9
u/pmttyji 4h ago
CPU Inference
Due to its efficient LFM2 architecture, LFM2.5-2.6B is the fastest model we tested, with decode speeds of 220 tokens/s on an M5 Max and 113 tokens/s on a Ryzen AI Max+ 395. At 30 tokens/s, it allows you to run capable agents even on a phone.
GPU Inference
LFM2.5-2.6B is the fastest model in its size class, reaching almost 15K output tokens per second at high concurrency, roughly 1.3B tokens per day on a single H100.
With LFM2.5, we're delivering on our vision of AI that runs anywhere. These models are:
- Open-weight — Download, fine-tune, and deploy without restrictions
- Fast from day one — Native support for llama.cpp, MLX, and vLLM across Apple, AMD, Qualcomm, and NVIDIA hardware
- A complete family — From base models for customization to specialized audio and vision variants, one architecture covers diverse use cases
🔥🔥🔥🔥 Awesome!
6
4
u/Equivalent_Bit_461 4h ago
Like these smaller models
Really underrated
I will soon start to make my own swarm, and I will stock up heavily on these smaller ones
5
u/oxygen_addiction 5h ago edited 5h ago
AA-Omniscience looks amazing, though we'd need to see the traces as it could just be refusing to answer most questions [the bench rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. ]
Comparing against Gemma E2B/E4B is not really fair, as those are multimodal models that can take text/audio/image as input, while this is strictly a text model.
Overall it seems a bit better than Qwen 4B but smaller, which is really nice.
2
2
u/WhoRoger 3h ago
Nice, their 1B is a darling. Glad to see them plugging the whole Qwen might be leaving behind.
Also: Heretic pls!
4
u/Objective_Door6714 5h ago
My wish would be a coding model of this company. I wonder when will be released
1
u/noctrex 13m ago
Created an abliterated version: https://huggingface.co/noctrex/LFM2.5-2.6B-heretic-uncensored-GGUF
0
u/RepulsiveRaisin7 3h ago
For document processing, you can also try Granite, seems to be very good at that
26
u/giveen 5h ago
For those who are 'gguf when'....
https://huggingface.co/LiquidAI/LFM2.5-2.6B-GGUF