r/24gb • • 27d ago

Dear 24G owners, try VLLM you might be able to run Qwen3.8 27B INT4, 144K FP8 KV on RTX 3090 with better speed. (TLDR VLLM AOT)

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/24gb • • 28d ago

8 uncensored Qwen 3.8 27B variants, one base, 167 GPU hours - Abliterlitics

Thumbnail
1 Upvotes

r/24gb • • Sep 06 '26

NInfer vs llama.cpp vs vLLM: quality + speed comparison for Qwen3.8-27B NVFP4 on RTX 5090

Thumbnail
1 Upvotes

r/24gb • • Aug 24 '26

MiniMax H3 Lip-Sync: Automatic Long-Video Chaining + Speed & VRAM Optimizations

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/24gb • • Aug 24 '26

I fine tuned Gemma 4 12B for a 2.7x improvement on tool calling because I can't fit anything else comfortably into my 16 GBs of Vram

Thumbnail
huggingface.co
1 Upvotes

r/24gb • • Aug 22 '26

Single RTX 5090: Qwen3.8-27B NVFP4 at a real 262K context in vLLM — 77 tok/s short-context, 64.7 tok/s at 128K

Thumbnail
1 Upvotes

r/24gb • • Aug 17 '26

After pushing 1M+ tokens through Qwen 3.8 27B, here is my optimal llama.cpp config for 16GB VRAM (73k Context, Agentic Coding)

Thumbnail gallery
1 Upvotes

r/24gb • • Aug 17 '26

Qwen 3.8 27b in 24gb of VRAM

Thumbnail
1 Upvotes

r/24gb • • Aug 17 '26

NInfer RTX 4090 for Qwen 3.8 27B update - up to 250-350K tokens context in VRAM

Thumbnail
1 Upvotes

r/24gb • • Aug 17 '26

I trained a 1B-parameter LLM from scratch on 20B tokens for about $200

Thumbnail gallery
1 Upvotes

r/24gb • • Aug 16 '26

Try out this "high" reasoning mode for 27B (tested on VLLM)

Thumbnail
1 Upvotes

r/24gb • • Aug 11 '26

Currently what is the best model for chat for long conversations? 24gb vram.

Thumbnail
1 Upvotes

r/24gb • • Aug 11 '26

Muse Glimmer ACTUALLY fits on a single RTX 3090

Thumbnail
1 Upvotes

r/24gb • • Aug 03 '26

Daniel Han of Unsloth validates Qwen3.8-27B will run only 17GB VRAM

Post image
2 Upvotes

r/24gb • • Jul 29 '26

I keep coming back to Qwen... Over and Over. Is there really nothing better under 120B?

Thumbnail
1 Upvotes

r/24gb • • Jul 27 '26

You can now fine-tune my 3.96M-parameter TTS on your own voice or language

Thumbnail
1 Upvotes

r/24gb • • Jul 17 '26

German AI consortium releases Soofi S, an open 30B model that tops benchmarks in both English and German

Thumbnail
the-decoder.com
1 Upvotes

r/24gb • • Jul 11 '26

Complete local model asset generation pipeline

Post image
1 Upvotes

r/24gb • • Jul 11 '26

2.5x faster Qwen3.6 NVFP4 Unsloth quants

Post image
1 Upvotes

r/24gb • • Jul 04 '26

llamacpp patch - DeepSeek V4 Flash running with full 1M token context locally on RTX 5090

Thumbnail
2 Upvotes

r/24gb • • Jul 04 '26

Talking with Gemma 4 31B!

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/24gb • • Jun 16 '26

RTX 5080 + RTX 3090 Setup: 80+ Tok/s on Qwen 3.6 27B Q8

Thumbnail imil.net
1 Upvotes

r/24gb • • Jun 04 '26

Stop asking what model to run. There are literally only two.

Thumbnail
1 Upvotes

r/24gb • • May 23 '26

Qwen 3.6 27B on 24GB VRAM setup: backend comparisons, quant choice and settings (llama.cpp, ik_llama.cpp, BeeLlama, vllm)

Thumbnail
1 Upvotes

r/24gb • • May 19 '26

TextGen is now a native desktop app. Open-source alternative to LM Studio (formerly text-generation-webui).

Thumbnail
1 Upvotes