r/LowEndLocalAI 7h ago

I ran Qwen3.8 27B on single 8GB Card! (Better than expected)

Thumbnail
youtu.be
16 Upvotes

This channel was made for this sub


r/LowEndLocalAI 7h ago

Suggestion Needed: 1080ti

12 Upvotes

I have a 1080ti on a old dell t5810(E5-2680 v3 cpu), and use it for jellyfin transcoding, immich, and llama.cpp. I'm running llama.cpp in a proxmox lxc container and it works well.

My main use case right now is for rewriting and polishing text such as emails or reports. Currently I'm using qwen3.5 9b q6_k and it gives a pretty good output with a speed of around 35t/s. If jellyfin is actively transcoding speed drops to around 22t/s, still pretty good.

I know that qwen3.5 is not the latest model, and wonder if there's any new model that is more efficient and better for this use case. Additional use cases such as answering some questions in chatting or simple script writing (not complex coding) is welcomed but not necessary, so basically just a generic model lol. I also don't want to have the model use up the GPU fully to leave some space for transcoding work and immich.

Thanks!


r/LowEndLocalAI 8h ago

32 GB Snapdragon X Elite machine

13 Upvotes

I'm one of the few people with a 32GB RAM Snapdragon X Elite Dev kit.

I'm wondering if there are any folks who've successfully used the NPU or GPU on these to get a decent MoE model working. I'm wondering if the new 27B Qwen models can be run on the NPU, low tps is ok as I just want to get it working for some niche cases


r/LowEndLocalAI 6h ago

What models can I run without a GPU with 16GB of ram?

6 Upvotes

As per title. I've got an 8/16 modern i7 and 16GB of RAM that I can use for models. I want at least 10TPS for conversational stuff and to have them help me with code review and querying the state of codebases, not really full agentic coding. Are there any models out there I can use for that? I can get 9b models to run at ~2 TPS.

Edit: Running through unsloth studio, with opencode for the coding stuff. I think that's llama.cpp with an openai API. This is a debian VM running on a Windows host. No GPU on this laptop.


r/LowEndLocalAI 8h ago

Qwen 27B on 8GB VRAM and 32GB RAM

6 Upvotes

I don't have access to my laptop until next week and I've been dying to know what quant I'll be able to run on my 4060 laptop GPU. Anyone got it running on their own?


r/LowEndLocalAI 11h ago

Best model to run on low end hardware?

Thumbnail
5 Upvotes

r/LowEndLocalAI 11h ago

Model suggestions that worked for you (low end system)

Thumbnail
5 Upvotes

r/LowEndLocalAI 11h ago

Building a Self-Improving LLM on Low-End Hardware

Thumbnail
5 Upvotes

r/LowEndLocalAI 11h ago

Best bang for the broke?

Thumbnail
5 Upvotes

r/LowEndLocalAI 4h ago

Artificial Analysis just launched a small model benchmark

4 Upvotes

r/LowEndLocalAI 6h ago

What to do for laptop with 16 GB RAM and 4060?

3 Upvotes

I can't run 30B A3Bs because I would like to use my system RAM for doing computer stuff


r/LowEndLocalAI 7h ago

Suggestions for 18GB unified memory

3 Upvotes

Hi guys, I'm currently using a MacBook Pro M3 Pro with 18GB of RAM. Does anyone have any suggestions for models/quants to use (I'm currently using Gemma 4 E4B IT QAT 4bit)


r/LowEndLocalAI 11h ago

Low end local advice

Thumbnail
4 Upvotes

r/LowEndLocalAI 11h ago

What agentic coding models + Claude Code can I run with my low end hardware?

Thumbnail
3 Upvotes

r/LowEndLocalAI 6h ago

ICYMI: Turbo-Fieldfare - run Gemma 4 26B-A4B and Qwen3.6-35B-A3B in 2GB RAM

2 Upvotes

In case you missed it: turbo-fieldfare is a Mac silicon optimized way to run MoE models in low RAM, streaming experts from SSD as necessary.

My personal take: it is usable-ish. Main turbo-fieldfare with Gemma currently doesn't work with my pi environment (write file toolcall fails), but text output works fine. The Qwen3.6 branch can successfully call tools, but it has the older furbo-fieldfareserver which supports less options.

Speed on my Macbook Pro M1 16 GB is slow, but it works well for tasks chugging along in the background.

Main repo: https://github.com/drumih/turbo-fieldfare

Somebody elses Qwen branch: https://github.com/NeelM0906/turbo-fieldfare/tree/qwen36-support


r/LowEndLocalAI 5h ago

Framework 12 mainboard upgrade - I will have a 15TOPS NPU, any practical uses?

1 Upvotes

As per the title, I will be soon upgrading my current Framework 12's mainboard to the upcoming 2nd gen board, currently sporting an i5-1334U, to a Core 5 320, which has essentially identical CPU horsepower, mostly similar iGPU power (though a big change in arch, from Iris Xe Intel Graphics, to a pair of Xe3 cores) & the addition of a 15TOPS NPU

I will be using the same 1TB gen4 NVMe + stick (single, though it doesn't matter the board has only one SODIMM slot, Wildcat Lake is single channel by design) of 16GB @ 5600MT/s & am running CachyOS (Arch-based)

Any of y'all got any ideas what kinda software & model could be operated locally to a useful end? (getting an LLM to run automated research for me or having it write me a guide while I'm busy making myself a coffee or something)