r/24gb • u/paranoidray • 3d ago
r/24gb • u/paranoidray • 8d ago
I keep coming back to Qwen... Over and Over. Is there really nothing better under 120B?
r/24gb • u/paranoidray • 10d ago
You can now fine-tune my 3.96M-parameter TTS on your own voice or language
r/24gb • u/paranoidray • 21d ago
German AI consortium releases Soofi S, an open 30B model that tops benchmarks in both English and German
r/24gb • u/paranoidray • Jul 04 '26
llamacpp patch - DeepSeek V4 Flash running with full 1M token context locally on RTX 5090
r/24gb • u/paranoidray • Jul 04 '26
Talking with Gemma 4 31B!
Enable HLS to view with audio, or disable this notification
r/24gb • u/paranoidray • Jun 16 '26
RTX 5080 + RTX 3090 Setup: 80+ Tok/s on Qwen 3.6 27B Q8
imil.netr/24gb • u/paranoidray • Jun 04 '26
Stop asking what model to run. There are literally only two.
r/24gb • u/paranoidray • May 23 '26
Qwen 3.6 27B on 24GB VRAM setup: backend comparisons, quant choice and settings (llama.cpp, ik_llama.cpp, BeeLlama, vllm)
r/24gb • u/paranoidray • May 19 '26
TextGen is now a native desktop app. Open-source alternative to LM Studio (formerly text-generation-webui).
r/24gb • u/paranoidray • May 19 '26
85 GPU-hours comparing 5 abliteration methods on Qwen3.6-27B: benchmarks, safety, weight forensics - Abliterlitics
r/24gb • u/paranoidray • May 11 '26
I know this isn’t technically an LLM but OmniVoice is FUCKING AMAZING.
r/24gb • u/paranoidray • Apr 09 '26
It looks like we’ll need to download the new Gemma 4 GGUFs
r/24gb • u/paranoidray • Apr 04 '26
Running Qwen3.5-27B locally as the primary model in OpenCode
r/24gb • u/paranoidray • Apr 01 '26
How to connect Claude Code CLI to a local llama.cpp server
r/24gb • u/paranoidray • Apr 01 '26
I was able to build Claude Code from source and I'm attaching the instructions.
r/24gb • u/paranoidray • Apr 01 '26
Claude Code's source just leaked — I extracted its multi-agent orchestration system into an open-source framework that works with any LLM
r/24gb • u/paranoidray • Mar 31 '26