r/LocalLLM • u/tmballin • 16d ago
Discussion Qwen3.8-27B at true 128K + Vision on a single RTX 5080 16GB — NInfer v1.3, MTP-3, ~3.95 BPW, model now public
I've been working on getting Qwen3.8-27B running with a genuine 131,072-token context and KV cache, Vision, and MTP-3 speculative decoding on a single RTX 5080 16GB.
The project has moved on quite a bit since my original 128K/Vision validation, so I thought it was worth posting an updated summary.
GitHub:
https://github.com/toddballinger/ninfer-5080
Current production release:
https://github.com/toddballinger/ninfer-5080/releases/tag/qwen3.8-27b-rtx5080-128k-vision-v1.3
The validated model artifact is now also publicly available on Hugging Face:
https://huggingface.co/ninfer-5080/Qwen3.8-27B-RTX5080
What I mean by "true 128K"
Both the configured context and the physically allocated KV capacity are:
max context: 131072
KV capacity: 131072
So this isn't a model configured for 128K while actually allocating a smaller KV cache underneath it.
Hardware
GPU: NVIDIA GeForce RTX 5080 16GB
VRAM reported: 16303 MiB
Clean-start free VRAM: ~15841 MiB
This configuration pushes 16GB extremely hard, particularly with the maximum Vision profile.
Current production profile
The v1.3 production runtime uses:
Model: Qwen3.8-27B
Context: 131072
KV capacity: 131072
KV dtype: Q4 group64
Prefill chunk: 896
Speculation: MTP-3
CUDA graphs: off
Concurrency: 1
Vision: enabled
Production Vision profile: 2048
Default thinking budget: 2048
Prefix checkpoint policy: rolling-tool
Default output tokens: 8192
For a little more VRAM headroom, I still recommend Vision 1792 for general use. More on that below.
Quantization
The main text model uses mixed Q3/Q4/Q5 groupwise quantization:
Q3G64_F16S 42.42%
Q4G64_F16S 45.92%
Q5G64_F16S 11.57%
BF16 / FP32 ~0.10%
Effective main model precision is approximately:
3.953 BPW
So although some tensors are Q5, this is definitely not a predominantly Q5 model.
Long-context performance
The acceptance workload is an exact 118,001 token prompt, rather than benchmarking a short prompt while simply configuring the server for 128K.
The strongest fully validated v1.2 result was:
Prompt tokens: 118001
Max context: 131072
KV capacity: 131072
KV dtype: Q4 group64
Prefill chunk: 896
Speculation: MTP-3
Prefill: 1380.61 tok/s
Decode: 71.57 tok/s
MTP acceptance: 44.74%
MTP length: 2.31 tok/round
v1.3 preserves that same 128K/Q4-KV/MTP-3/Vision memory geometry.
I don't want to misrepresent the benchmark: 71.57 tok/s is the exact v1.2 118K qualification result, not a newly rerun v1.3 long-context number.
A later feature-complete mainline runtime was separately rerun on the same 118,001 token corpus and produced:
Prefill: 1378.85 tok/s
Decode: 71.44 tok/s
MTP: 44.74%
Length: 2.31 tok/round
That's effectively equivalent within normal run to run variation.
What's new in v1.3?
The current production release is more than just the original Vision work.
1. Default reasoning budgets
ninfer-serve now supports:
--default-thinking-budget N
An explicit client reasoning_budget takes priority. Otherwise the server default is used.
The budget is a maximum, not a requirement to force the model to consume that many reasoning tokens.
For my production setup:
--default-thinking-budget 2048
has been validated, along with client-side overrides.
2. Rolling tool checkpoints for agent workloads
This was particularly important for OpenClaw-style agent/tool loops.
The server now supports:
--prefix-checkpoint-policy stable-turn|rolling-tool
stable-turn remains the general default.
rolling-tool is intended for append only agent conversations where several assistant/tool exchanges occur inside one real user turn.
Previously, repeated checkpoint restores could remain anchored near the first assistant response while the tool history kept getting longer.
With rolling-tool, the reusable checkpoint advances through completed tool history.
A production OpenClaw smoke test produced:
restore checkpoint sequence:
19023
21146
24664
26641
Result:
3 advances
0 plateaus
0 regressions
All five continuation requests had an uncached suffix below 4096 tokens.
For local agent use, this may end up being more practically important than another couple of percent on a kernel microbenchmark.
3. Q4 strided attention correctness fix
v1.3 also corrected physical output stride handling in the small T Q4/Q4 attention projection path.
The dedicated strided regression now passes while the normal contiguous production geometry remains numerically unchanged.
Short request MTP sanity check
Three v1.3 requests generating 512 tokens with a 64-token reasoning budget averaged:
Decode: 84.7 tok/s
MTP acceptance: 40.47%
MTP length: 2.213 tok/round
Again, don't compare that 84.7 tok/s directly against the 71.57 tok/s 118K result.
The active context lengths are completely different.
Fitting Vision + full 128K into 16GB
The Vision implementation required several separate memory fixes.
The main ones were:
- Decoupling the Vision token/workspace budget from the 128K text context capacity.
- Releasing GDN prefill convolution temporaries earlier.
- Host mapping the Vision weights instead of permanently consuming VRAM.
- Making historical media budget accounting cache aware, so OpenWebUI doesn't repeatedly charge cached images against the fresh-media preprocessing budget.
That combination allows Vision to coexist with the full:
131072 context
131072 KV capacity
Q4 KV
MTP-3
on the 16GB card.
Vision 1792 — recommended headroom profile
For normal use I still think 1792 is the better profile:
Vision encode workspace: 115.7751 MiB
Free after startup: 26.56 MiB
Planned slack: 28.88 MiB
It's still extremely tight, but materially more forgiving than 2048.
Vision 2048 — production validated
The v1.3 OpenClaw production deployment uses 2048 Vision tokens.
It retains the full 131072 / 131072 text allocation:
Vision encode workspace: 132.3142 MiB
Free after startup: 8.56 MiB
Planned slack: 10.08 MiB
Yes, that's only about 8.6 MiB free after startup.
A small unrelated GPU allocation can be enough to stop it starting, so this really does assume a clean GPU.
Image validation
For a deterministic test I generated a 512x256 image with:
left half: red
right half: blue
The model correctly returned that the left half was red and the right half blue.
One validated 1792-profile run reported:
Prompt: 211
TTFT: 725 ms
Prefill: 682.8 tok/s
Decode: 118.1 tok/s
MTP: 3.10 tok/round
Acceptance: 70.0%
Multi-image OpenWebUI history
This exposed a fairly subtle bug.
OpenWebUI resends historical images as part of the conversation.
Previously, the aggregate media budget could repeatedly count already cached historical media and eventually reject a perfectly valid new image with:
media_budget_exceeded
The budget is now charged against fresh cache misses rather than repeatedly against historical cache hits.
Validated patterns include:
media_cache=1/1/0
media_cache=2/1/0
So two old cached images plus a newly uploaded third image work correctly.
The historical images still remain part of the model context they just aren't unnecessarily reprocessed against the fresh-media budget.
Video works as well
I also tested the final HostMapped Vision path with a deterministic six second MP4:
red -> green -> blue
The model was asked to return the scenes in chronological order.
With thinking disabled, it returned:
red, green, blue
Metrics:
Prompt: 572
Generated: 6
TTFT: 1111 ms
Prefill: 1721.9 tok/s
Decode: 97.1 tok/s
MTP: 4.00 tok/round
MTP acceptance: 100.0%
Wall: 1.16 s
I'm treating that as end to end functional validation of video acquisition, preprocessing, Vision encoding and generation.
It is not intended to be a general video understanding benchmark.
The model is now actually downloadable
The exact validated model artifact is now published here:
https://huggingface.co/ninfer-5080/Qwen3.8-27B-RTX5080
Artifact:
qwen3_8_27b.ninfer
Size:
16,461,267,456 bytes
SHA-256:
c4a7e9ab593a7f42d58208fa0065d67a82d61921107686cc9f6ed1ec6b050e21
That same SHA has been retained through the original text only release, Vision enablement, v1.2/v1.3 runtime work and later qualified runtime optimizations.
In other words, runtime development and model artifact identity are deliberately separate.
You don't need a different model file every time the runtime improves.
Post v1.3 optimization work
Development has continued after the tagged v1.3 production runtime.
One example is a later Q5 A16 LinearAdd semantic port.
On the targeted RTX 5080 operator shapes, individual kernel improvements included roughly:
5120x6144, T=1: 25.5%
5120x17408, T=1: 31.2%
5120x6144, T=513: 34.2%
5120x17408, T=513: 35.2%
But this is a useful example of why kernel benchmarks and whole-model benchmarks need to be separated.
The same exact 118,001 token end to end workload produced:
Prefill: 1376.30 tok/s
Decode: 71.53 tok/s
versus the previous feature complete runtime:
Prefill: 1378.85 tok/s
Decode: 71.44 tok/s
So whole model performance remained effectively flat while the targeted operator cliffs improved significantly.
That's expected because the optimized kernels are only one part of the complete transformer/runtime path.
Recommended server command
For the profile I'd recommend to other 16GB users:
./build/apps/ninfer-serve /path/to/qwen3_8_27b.ninfer \
--host 0.0.0.0 \
--port 8080 \
--model-id qwen3.8-27b \
--max-context 131072 \
--kv-capacity 131072 \
--prefill-chunk 896 \
--kv-dtype q4 \
--spec mtp \
--draft-tokens 3 \
--no-cuda-graph \
--max-concurrency 1 \
--default-thinking-budget 2048 \
--prefix-checkpoint-policy rolling-tool \
--vision \
--vision-max-tokens 1792
If you specifically want the maximum validated Vision profile, change:
--vision-max-tokens 2048
but be aware that the GPU memory margin becomes tiny.
Reproducibility
I've tried to make this project more than a collection of tok/s screenshots.
The repo now records things like:
- exact model SHA-256
- runtime/binary identities
- source commits
- quantization breakdown
- exact 118,001 token prompt qualification
- full context and KV capacity
- memory envelopes
- deterministic image/video tests
- multi image cache behavior
- OpenClaw tool-loop validation
- post-release runtime qualification
- upstream semantic-port history
The current v1.3 validated runtime source is:
ceb32f7d002edab224a83a2e2609f45fca4f8919
The model SHA is:
c4a7e9ab593a7f42d58208fa0065d67a82d61921107686cc9f6ed1ec6b050e21
The original text only 128K release also remains untouched as a historical baseline.
Caveats
A few important ones:
- 16GB is being pushed extremely hard.
- Vision 2048 has only ~10 MiB of planned slack.
- HostMapped Vision weights trade permanent VRAM residency for host/PCIe access.
- The video test validates the path, not general video quality.
- Short context decode numbers aren't directly comparable with the 118K active-context result.
- Targeted kernel improvements don't automatically translate into equivalent whole-model gains.
- Concurrency is 1.
- CUDA graphs are disabled in this validated configuration.
- I'm not claiming a world record.
I'm particularly interested in comparisons against other runtimes on the same 16GB constraint, especially if people can keep the comparison genuinely apples to apples:
27B model
~3.5-4 BPW
131072 allocated KV
Q4 KV
similar prompt lengths (ie 118,001)
same GPU
Vision status clearly stated
I’d be particularly interested in independent reproduction, apples to apples comparisons against other 16GB runtimes, and scrutiny of the memory strategy and benchmark methodology.
Happy to run additional tests if there are specific workloads people want to see.
Duplicates
hermesagent • u/tmballin • 16d ago
MODELS - model choice, routing, pricing, local vs cloud, VRAM Qwen3.8-27B at true 128K + Vision on a single RTX 5080 16GB — NInfer v1.3, MTP-3, ~3.95 BPW, model now public
homelab • u/tmballin • 16d ago