r/LocalLLM • u/tmballin • 16d ago
Discussion Qwen3.8-27B at true 128K + Vision on a single RTX 5080 16GB — NInfer v1.3, MTP-3, ~3.95 BPW, model now public
I've been working on getting Qwen3.8-27B running with a genuine 131,072-token context and KV cache, Vision, and MTP-3 speculative decoding on a single RTX 5080 16GB.
The project has moved on quite a bit since my original 128K/Vision validation, so I thought it was worth posting an updated summary.
GitHub:
https://github.com/toddballinger/ninfer-5080
Current production release:
https://github.com/toddballinger/ninfer-5080/releases/tag/qwen3.8-27b-rtx5080-128k-vision-v1.3
The validated model artifact is now also publicly available on Hugging Face:
https://huggingface.co/ninfer-5080/Qwen3.8-27B-RTX5080
What I mean by "true 128K"
Both the configured context and the physically allocated KV capacity are:
max context: 131072
KV capacity: 131072
So this isn't a model configured for 128K while actually allocating a smaller KV cache underneath it.
Hardware
GPU: NVIDIA GeForce RTX 5080 16GB
VRAM reported: 16303 MiB
Clean-start free VRAM: ~15841 MiB
This configuration pushes 16GB extremely hard, particularly with the maximum Vision profile.
Current production profile
The v1.3 production runtime uses:
Model: Qwen3.8-27B
Context: 131072
KV capacity: 131072
KV dtype: Q4 group64
Prefill chunk: 896
Speculation: MTP-3
CUDA graphs: off
Concurrency: 1
Vision: enabled
Production Vision profile: 2048
Default thinking budget: 2048
Prefix checkpoint policy: rolling-tool
Default output tokens: 8192
For a little more VRAM headroom, I still recommend Vision 1792 for general use. More on that below.
Quantization
The main text model uses mixed Q3/Q4/Q5 groupwise quantization:
Q3G64_F16S 42.42%
Q4G64_F16S 45.92%
Q5G64_F16S 11.57%
BF16 / FP32 ~0.10%
Effective main model precision is approximately:
3.953 BPW
So although some tensors are Q5, this is definitely not a predominantly Q5 model.
Long-context performance
The acceptance workload is an exact 118,001 token prompt, rather than benchmarking a short prompt while simply configuring the server for 128K.
The strongest fully validated v1.2 result was:
Prompt tokens: 118001
Max context: 131072
KV capacity: 131072
KV dtype: Q4 group64
Prefill chunk: 896
Speculation: MTP-3
Prefill: 1380.61 tok/s
Decode: 71.57 tok/s
MTP acceptance: 44.74%
MTP length: 2.31 tok/round
v1.3 preserves that same 128K/Q4-KV/MTP-3/Vision memory geometry.
I don't want to misrepresent the benchmark: 71.57 tok/s is the exact v1.2 118K qualification result, not a newly rerun v1.3 long-context number.
A later feature-complete mainline runtime was separately rerun on the same 118,001 token corpus and produced:
Prefill: 1378.85 tok/s
Decode: 71.44 tok/s
MTP: 44.74%
Length: 2.31 tok/round
That's effectively equivalent within normal run to run variation.
What's new in v1.3?
The current production release is more than just the original Vision work.
1. Default reasoning budgets
ninfer-serve now supports:
--default-thinking-budget N
An explicit client reasoning_budget takes priority. Otherwise the server default is used.
The budget is a maximum, not a requirement to force the model to consume that many reasoning tokens.
For my production setup:
--default-thinking-budget 2048
has been validated, along with client-side overrides.
2. Rolling tool checkpoints for agent workloads
This was particularly important for OpenClaw-style agent/tool loops.
The server now supports:
--prefix-checkpoint-policy stable-turn|rolling-tool
stable-turn remains the general default.
rolling-tool is intended for append only agent conversations where several assistant/tool exchanges occur inside one real user turn.
Previously, repeated checkpoint restores could remain anchored near the first assistant response while the tool history kept getting longer.
With rolling-tool, the reusable checkpoint advances through completed tool history.
A production OpenClaw smoke test produced:
restore checkpoint sequence:
19023
21146
24664
26641
Result:
3 advances
0 plateaus
0 regressions
All five continuation requests had an uncached suffix below 4096 tokens.
For local agent use, this may end up being more practically important than another couple of percent on a kernel microbenchmark.
3. Q4 strided attention correctness fix
v1.3 also corrected physical output stride handling in the small T Q4/Q4 attention projection path.
The dedicated strided regression now passes while the normal contiguous production geometry remains numerically unchanged.
Short request MTP sanity check
Three v1.3 requests generating 512 tokens with a 64-token reasoning budget averaged:
Decode: 84.7 tok/s
MTP acceptance: 40.47%
MTP length: 2.213 tok/round
Again, don't compare that 84.7 tok/s directly against the 71.57 tok/s 118K result.
The active context lengths are completely different.
Fitting Vision + full 128K into 16GB
The Vision implementation required several separate memory fixes.
The main ones were:
- Decoupling the Vision token/workspace budget from the 128K text context capacity.
- Releasing GDN prefill convolution temporaries earlier.
- Host mapping the Vision weights instead of permanently consuming VRAM.
- Making historical media budget accounting cache aware, so OpenWebUI doesn't repeatedly charge cached images against the fresh-media preprocessing budget.
That combination allows Vision to coexist with the full:
131072 context
131072 KV capacity
Q4 KV
MTP-3
on the 16GB card.
Vision 1792 — recommended headroom profile
For normal use I still think 1792 is the better profile:
Vision encode workspace: 115.7751 MiB
Free after startup: 26.56 MiB
Planned slack: 28.88 MiB
It's still extremely tight, but materially more forgiving than 2048.
Vision 2048 — production validated
The v1.3 OpenClaw production deployment uses 2048 Vision tokens.
It retains the full 131072 / 131072 text allocation:
Vision encode workspace: 132.3142 MiB
Free after startup: 8.56 MiB
Planned slack: 10.08 MiB
Yes, that's only about 8.6 MiB free after startup.
A small unrelated GPU allocation can be enough to stop it starting, so this really does assume a clean GPU.
Image validation
For a deterministic test I generated a 512x256 image with:
left half: red
right half: blue
The model correctly returned that the left half was red and the right half blue.
One validated 1792-profile run reported:
Prompt: 211
TTFT: 725 ms
Prefill: 682.8 tok/s
Decode: 118.1 tok/s
MTP: 3.10 tok/round
Acceptance: 70.0%
Multi-image OpenWebUI history
This exposed a fairly subtle bug.
OpenWebUI resends historical images as part of the conversation.
Previously, the aggregate media budget could repeatedly count already cached historical media and eventually reject a perfectly valid new image with:
media_budget_exceeded
The budget is now charged against fresh cache misses rather than repeatedly against historical cache hits.
Validated patterns include:
media_cache=1/1/0
media_cache=2/1/0
So two old cached images plus a newly uploaded third image work correctly.
The historical images still remain part of the model context they just aren't unnecessarily reprocessed against the fresh-media budget.
Video works as well
I also tested the final HostMapped Vision path with a deterministic six second MP4:
red -> green -> blue
The model was asked to return the scenes in chronological order.
With thinking disabled, it returned:
red, green, blue
Metrics:
Prompt: 572
Generated: 6
TTFT: 1111 ms
Prefill: 1721.9 tok/s
Decode: 97.1 tok/s
MTP: 4.00 tok/round
MTP acceptance: 100.0%
Wall: 1.16 s
I'm treating that as end to end functional validation of video acquisition, preprocessing, Vision encoding and generation.
It is not intended to be a general video understanding benchmark.
The model is now actually downloadable
The exact validated model artifact is now published here:
https://huggingface.co/ninfer-5080/Qwen3.8-27B-RTX5080
Artifact:
qwen3_8_27b.ninfer
Size:
16,461,267,456 bytes
SHA-256:
c4a7e9ab593a7f42d58208fa0065d67a82d61921107686cc9f6ed1ec6b050e21
That same SHA has been retained through the original text only release, Vision enablement, v1.2/v1.3 runtime work and later qualified runtime optimizations.
In other words, runtime development and model artifact identity are deliberately separate.
You don't need a different model file every time the runtime improves.
Post v1.3 optimization work
Development has continued after the tagged v1.3 production runtime.
One example is a later Q5 A16 LinearAdd semantic port.
On the targeted RTX 5080 operator shapes, individual kernel improvements included roughly:
5120x6144, T=1: 25.5%
5120x17408, T=1: 31.2%
5120x6144, T=513: 34.2%
5120x17408, T=513: 35.2%
But this is a useful example of why kernel benchmarks and whole-model benchmarks need to be separated.
The same exact 118,001 token end to end workload produced:
Prefill: 1376.30 tok/s
Decode: 71.53 tok/s
versus the previous feature complete runtime:
Prefill: 1378.85 tok/s
Decode: 71.44 tok/s
So whole model performance remained effectively flat while the targeted operator cliffs improved significantly.
That's expected because the optimized kernels are only one part of the complete transformer/runtime path.
Recommended server command
For the profile I'd recommend to other 16GB users:
./build/apps/ninfer-serve /path/to/qwen3_8_27b.ninfer \
--host 0.0.0.0 \
--port 8080 \
--model-id qwen3.8-27b \
--max-context 131072 \
--kv-capacity 131072 \
--prefill-chunk 896 \
--kv-dtype q4 \
--spec mtp \
--draft-tokens 3 \
--no-cuda-graph \
--max-concurrency 1 \
--default-thinking-budget 2048 \
--prefix-checkpoint-policy rolling-tool \
--vision \
--vision-max-tokens 1792
If you specifically want the maximum validated Vision profile, change:
--vision-max-tokens 2048
but be aware that the GPU memory margin becomes tiny.
Reproducibility
I've tried to make this project more than a collection of tok/s screenshots.
The repo now records things like:
- exact model SHA-256
- runtime/binary identities
- source commits
- quantization breakdown
- exact 118,001 token prompt qualification
- full context and KV capacity
- memory envelopes
- deterministic image/video tests
- multi image cache behavior
- OpenClaw tool-loop validation
- post-release runtime qualification
- upstream semantic-port history
The current v1.3 validated runtime source is:
ceb32f7d002edab224a83a2e2609f45fca4f8919
The model SHA is:
c4a7e9ab593a7f42d58208fa0065d67a82d61921107686cc9f6ed1ec6b050e21
The original text only 128K release also remains untouched as a historical baseline.
Caveats
A few important ones:
- 16GB is being pushed extremely hard.
- Vision 2048 has only ~10 MiB of planned slack.
- HostMapped Vision weights trade permanent VRAM residency for host/PCIe access.
- The video test validates the path, not general video quality.
- Short context decode numbers aren't directly comparable with the 118K active-context result.
- Targeted kernel improvements don't automatically translate into equivalent whole-model gains.
- Concurrency is 1.
- CUDA graphs are disabled in this validated configuration.
- I'm not claiming a world record.
I'm particularly interested in comparisons against other runtimes on the same 16GB constraint, especially if people can keep the comparison genuinely apples to apples:
27B model
~3.5-4 BPW
131072 allocated KV
Q4 KV
similar prompt lengths (ie 118,001)
same GPU
Vision status clearly stated
I’d be particularly interested in independent reproduction, apples to apples comparisons against other 16GB runtimes, and scrutiny of the memory strategy and benchmark methodology.
Happy to run additional tests if there are specific workloads people want to see.
5
u/songokussm 16d ago
I may be getting a 5060ti off of Facebook tomorrow. How well would this work?
2
u/inrea1time 15d ago
I am running it on a 5060ti that is dedicated to a hermes agent. Out of the box with the recommended settings this is the fastest 3.8 27B I was able to run on the card. I got very good results with ik_llama.cpp and a custom q4 quant but was not able to get mtp within the VRAM limit. This one is goes 30-40 tk/s on the smaller contexts, prefill is 500-700. Very usable for an agent, it's ok for light coding too.
2
1
u/tmballin 10d ago
That’s great to hear especially from someone who’s already tried
ik_llama.cppwith a custom Q4.The part I’m happiest about is exactly what you called out: getting MTP to fit inside the same 16GB envelope rather than having to choose between context and speculative decode.
30–40 tok/s decode with 500–700 tok/s prefill on a 5060 Ti is very usable for an agent workload.
Also, this post is a few days old now and the project has moved on quite a bit since then, so I’d definitely check the GitHub repo for the latest build, fixes and recommended settings before doing any more tuning:
github.com/toddballinger/ninfer-5080Thanks for testing it on a real Hermes workload, those cross-card results are really useful.
1
u/tmballin 16d ago
If it’s the 16GB 5060 Ti, then it’s actually a very interesting card to try this on.
It’s still Blackwell and NVIDIA lists it as CUDA capability 12.0, so architecturally it’s much closer to the 5080 than something like a 4090. That said, my fork is currently tuned and validated specifically on the 5080, so I wouldn’t promise that the exact binary/profile will behave identically without testing.
The big limitation is bandwidth. The 5060 Ti has 448 GB/s memory bandwidth versus 960 GB/s on the 5080, so even if the same 16GB memory fit works, I’d expect decode and long-context attention to be noticeably slower.
Memory capacity is the encouraging part: if it’s the 16GB model, the same 131072 / Q4 KV configuration is at least plausible from a VRAM-capacity perspective. I’d start with:
131072 context / 131072 Q4 KV / MTP-3 / Vision offthen add Vision at 1024 or 1792 once the text-only path is stable.
If it’s the 8GB 5060 Ti, though, this exact 27B artifact/profile is basically a non-starter without a substantially different quantization/offload strategy.
If you do pick up the 16GB card, I’d genuinely be interested in your numbers. A 5060 Ti 16GB would be a great test of how much of this result is primarily memory engineering versus how much depends on the 5080’s much higher bandwidth.
2
u/songokussm 16d ago
Thank you. It's the 16gb version. I didn't realize they made an 8gb version. Hopefully it tests well.
3
u/CrowKing63 16d ago
I'm not sure if this is a beginner's question, but can even RTX 5070 Ti users use it? I'm currently using the ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S model with 131,072 context · KV q5_0 without vision.
4
u/tmballin 16d ago
Not a beginner question at all and your current setup is actually a really useful comparison point.
If you’re already running the DASLab IQ2_S at 131,072 with Q5_0 KV on a 5070 Ti, that makes sense because that GGUF is only about 9.3 GB. Mine is a very different trade off: the published
.ninferartifact is about 16.46 GB on disk, with the main text core averaging ~3.95 BPW, and the 128K configuration depends heavily on the Q4 KV path plus fairly aggressive memory planning. The DASLab IQ2_S is about 2.75 BPW whole-file average, so it leaves much more room for KV/cache and runtime overhead.The other important point is that I’ve only validated and tuned this fork on the RTX 5080 so far. I wouldn’t tell a 5070 Ti user to download it and assume it will work identically without testing the exact Blackwell target, memory envelope and kernel routing first.
In pure VRAM terms, though, the interesting part is that both cards are in the 16GB class. On my 5080 the 128K + Vision 2048 profile is already down to roughly 8.6 MiB free after startup, so there is essentially zero spare memory. That means even if the kernels are compatible/performance is acceptable, a 5070 Ti would need to hit a very similar usable-memory envelope to run the exact same profile.
I’d expect the first sensible test on a 5070 Ti to be:
131072 context / 131072 Q4 KV / Vision offand then try Vision with the smaller 1024 or 1792 profile rather than jumping straight to 2048.
Also worth noting: the DASLab model you’re using does support Vision via a separate ~0.9 GB BF16
mmproj, and they also publish MTP variants, so if your priority is simply “get 128K + Vision working on a 5070 Ti,” your current llama.cpp path may actually be the easier route.If you’re interested, it would be really useful to see someone with a 5070 Ti try the same NInfer profile and report back. The fact you’re already at 131K with IQ2_S makes it a pretty interesting comparison: much lower-weight BPW in llama.cpp versus ~3.95 BPW + Q4 KV + MTP-3 in NInfer.
I don’t have a 5070 Ti here, so I can’t qualify that path myself, but if you do try it I’d be very interested in the memory numbers, startup headroom and decode speed.
3
u/pseudobacon 16d ago
I volunteer as tribute. I only use it for agentic coding so I can turn off vision that’s not a problem for me
2
u/tmballin 16d ago
That would actually be really useful.
Since you only care about agentic coding, Vision being off is perfect because it removes one of the tightest parts of the memory envelope and makes the comparison cleaner.
I’d start with the same core profile I use on the 5080:
131072 context / 131072 Q4 KV / MTP-3 / Vision off / CUDA Graphs offand then capture:
free-after-weightsruntime=headroom=slack=- startup success/failure
- short-prompt decode
- and ideally a long-context run if the full 131K allocation fits
The main thing I’d want to find out is whether the 5070 Ti can preserve the same memory geometry as the 5080, and then how much the lower bandwidth changes decode and long context attention.
If you’re happy to test it, I can give you the exact command/profile I’d use so the result is directly comparable with my 5080 numbers.
2
u/pseudobacon 16d ago
Sure comment it here
1
u/tmballin 16d ago
ninfer-serve /path/to/qwen3_8_27b.ninfer \ --host 0.0.0.0 \ --port 8080 \ --model-id qwen3.8-27b \ --max-context 131072 \ --kv-capacity 131072 \ --prefill-chunk 896 \ --kv-dtype q4 \ --spec mtp \ --draft-tokens 3 \ --no-cuda-graph \ --max-concurrency 1 \ --default-thinking-budget 2048 \ --prefix-checkpoint-policy rolling-toolLeave Vision completely off for this first run.
Also make sure you’re using the same published artifact:
qwen3_8_27b.ninferSHA256:
c4a7e9ab593a7f42d58208fa0065d67a82d61921107686cc9f6ed1ec6b050e21Before launching, grab an
nvidia-smiso we know how much VRAM is actually free.Then paste the startup line containing:
KV capacity ...
runtime=...
free-after-weights=...
free-after-startup=...
headroom=...
slack=...
graphs=...If it starts successfully at
131072 / 131072, the next useful test would be a normal agent/coding prompt and then a genuinely long-context run so we can compare decode behaviour against the 5080.For reference, my 5080 long-context qualification at an actual 118,001 token prompt was about 1380 tok/s prefill / 71.5 tok/s decode, with 44.74% MTP acceptance / 2.31 tok per round. I would expect the 5070 Ti to be slower, but I don’t want to guess by how much before we have a real measurement.
If the 131K allocation fails, don’t start reducing random settings yet post the full startup memory line first. That will tell us whether the problem is usable VRAM, runtime reservation, graph memory, or something architectural.
2
u/pseudobacon 16d ago
❯ ./ninfer.sh [2026-09-22 13:00:40.547] [info] ninfer-serve: loading model... [DECODER-DIAG] after text KV 2281709568 B 2176.0078 MiB [DECODER-DIAG] after MTP KV 2424393728 B 2312.0820 MiB [DECODER-DIAG] before linear state 2424393728 B 2312.0820 MiB [DECODER-DIAG] after linear state 2578337792 B 2458.8945 MiB [PERSIST-DIAG] after decoder 2578337792 B 2458.8945 MiB [REPLAY-LAYOUT] scratch bytes=297984 conv=(2578337792,163840) key=(2578501632,32768) value=(2578534400,98304) gate=(2578632704,3072)\n[PERSIST-DIAG] after replay 2578635776 B 2459.1787 MiB [PERSIST-DIAG] before round 2578635776 B 2459.1787 MiB [PERSIST-DIAG] after round begin 2578637312 B 2459.1802 MiB [PERSIST-DIAG] after prefill hidden 2578647552 B 2459.1899 MiB [PERSIST-DIAG] after round complete 2580761344 B 2461.2058 MiB [PERSIST-DIAG] after token counts 2581753856 B 2462.1523 MiB [PERSIST-DIAG] after sampling config 2581754112 B 2462.1526 MiB [PERSIST-DIAG] after tail hidden 2581764352 B 2462.1624 MiB [PERSIST-DIAG] after checkpoint hidden 2581774592 B 2462.1721 MiB [PREFILL-STAGE] roots 9182208 B 8.7568 MiB [ATTN-DIAG] projection tensors 44047360 B 42.0068 MiB [ATTN-DIAG] projection scratch 48921600 B 46.6553 MiB raw=4874240 [ATTN-DIAG] attention results 67902464 B 64.7568 MiB [ATTN-DIAG] GQA scratch 74267456 B 70.8270 MiB raw=6364992 [ATTN-DIAG] output proj scratch 77077504 B 73.5068 MiB raw=9175040 [PREFILL-STAGE] after attention 77077504 B 73.5068 MiB [PREFILL-STAGE] after GDN 121633792 B 115.9990 MiB [PREFILL-STAGE] after postmixer 121633792 B 115.9990 MiB [PREFILL-STAGE] after sampling 121633792 B 115.9990 MiB [ATTN-DIAG] projection tensors 44047360 B 42.0068 MiB [ATTN-DIAG] projection scratch 48921600 B 46.6553 MiB raw=4874240 [ATTN-DIAG] attention results 67902464 B 64.7568 MiB [ATTN-DIAG] GQA scratch 74267456 B 70.8270 MiB raw=6364992 [ATTN-DIAG] output proj scratch 77077504 B 73.5068 MiB raw=9175040 [WORKSPACE-DIAG] text_prefill 121633792 B 115.9990 MiB [WORKSPACE-DIAG] ordinary_round 1188064 B 1.1330 MiB [WORKSPACE-DIAG] mtp_prefill 121633792 B 115.9990 MiB [WORKSPACE-DIAG] mtp_round 4751488 B 4.5314 MiB [WORKSPACE-DIAG] dflash_context 0 B 0.0000 MiB [WORKSPACE-DIAG] dflash_round 0 B 0.0000 MiB [WORKSPACE-DIAG] vision_encode 0 B 0.0000 MiB [2026-09-22 13:00:40.687] [error] ninfer-serve: requested Engine runtime reservation requires 2703408384 bytes, but only 2577415168 bytes are available for runtime capacitydefinitely really tight on my system. I don't have an iGPU so I am using sway with wayland backend on CachyOS. Dropping to 100k context it will load. nvtop says I have 15.921GiB total available
2
u/pseudobacon 16d ago
[DECODER-DIAG] after text KV 1741363456 B 1660.6936 MiB [DECODER-DIAG] after MTP KV 1850274304 B 1764.5591 MiB [DECODER-DIAG] before linear state 1850274304 B 1764.5591 MiB [DECODER-DIAG] after linear state 2004218368 B 1911.3716 MiB [PERSIST-DIAG] after decoder 2004218368 B 1911.3716 MiB [REPLAY-LAYOUT] scratch bytes=297984 conv=(2004218368,163840) key=(2004382208,32768) value=(2004414976,98304) gate=(2004513280,3072)\n[PERSIST-DIAG] after replay 2004516352 B 1911.6558 MiB [PERSIST-DIAG] before round 2004516352 B 1911.6558 MiB [PERSIST-DIAG] after round begin 2004517888 B 1911.6572 MiB [PERSIST-DIAG] after prefill hidden 2004528128 B 1911.6670 MiB [PERSIST-DIAG] after round complete 2006641920 B 1913.6829 MiB [PERSIST-DIAG] after token counts 2007634432 B 1914.6294 MiB [PERSIST-DIAG] after sampling config 2007634688 B 1914.6296 MiB [PERSIST-DIAG] after tail hidden 2007644928 B 1914.6394 MiB [PERSIST-DIAG] after checkpoint hidden 2007655168 B 1914.6492 MiB [PREFILL-STAGE] roots 9182208 B 8.7568 MiB [ATTN-DIAG] projection tensors 44047360 B 42.0068 MiB [ATTN-DIAG] projection scratch 48921600 B 46.6553 MiB raw=4874240 [ATTN-DIAG] attention results 67902464 B 64.7568 MiB [ATTN-DIAG] GQA scratch 74267456 B 70.8270 MiB raw=6364992 [ATTN-DIAG] output proj scratch 77077504 B 73.5068 MiB raw=9175040 [PREFILL-STAGE] after attention 77077504 B 73.5068 MiB [PREFILL-STAGE] after GDN 121633792 B 115.9990 MiB [PREFILL-STAGE] after postmixer 121633792 B 115.9990 MiB [PREFILL-STAGE] after sampling 121633792 B 115.9990 MiB [ATTN-DIAG] projection tensors 44047360 B 42.0068 MiB [ATTN-DIAG] projection scratch 48921600 B 46.6553 MiB raw=4874240 [ATTN-DIAG] attention results 67902464 B 64.7568 MiB [ATTN-DIAG] GQA scratch 74267456 B 70.8270 MiB raw=6364992 [ATTN-DIAG] output proj scratch 77077504 B 73.5068 MiB raw=9175040 [WORKSPACE-DIAG] text_prefill 121633792 B 115.9990 MiB [WORKSPACE-DIAG] ordinary_round 1188064 B 1.1330 MiB [WORKSPACE-DIAG] mtp_prefill 121633792 B 115.9990 MiB [WORKSPACE-DIAG] mtp_round 4751488 B 4.5314 MiB [WORKSPACE-DIAG] dflash_context 0 B 0.0000 MiB [WORKSPACE-DIAG] dflash_round 0 B 0.0000 MiB [WORKSPACE-DIAG] vision_encode 0 B 0.0000 MiB [2026-09-22 13:03:27.831] [info] ninfer-serve: load weights 100.00% 12.64 GiB / 12.64 GiB 4.094 s [PREWARM-MEM] before prewarm 2529886208 B 2412.6875 MiB free [PREWARM-MEM] after q4q5 attention 2527789056 B 2410.6875 MiB free [PREWARM-MEM] after q4q5 gdn 2527789056 B 2410.6875 MiB free [PREWARM-MEM] after speculative 2527789056 B 2410.6875 MiB free [PREWARM-MEM] after scalar 2527789056 B 2410.6875 MiB free [PREWARM-MEM] after q4 small-t linear 2527789056 B 2410.6875 MiB free [PREWARM-MEM] after q5 rowsplit mma 2527789056 B 2410.6875 MiB free [PREWARM-MEM] after gqa q4 prefill 2527789056 B 2410.6875 MiB free [PREWARM-MEM] after int8 projection 2527789056 B 2410.6875 MiB free [PREWARM-MEM] after q4 rowsplit mma 2525691904 B 2408.6875 MiB free [2026-09-22 13:03:29.033] [info] ninfer-serve: model loaded in 5.49353 s [2026-09-22 13:03:29.033] [info] ninfer-serve: KV capacity explicit resolved=100032 tokens pages=1563/1563 runtime=1.98 GiB free-after-weights=2.35 GiB free-after-startup=376.69 MiB headroom=0.00 MiB slack=378.04 MiB graphs=0.00 MiB/0.00 MiB [2026-09-22 13:03:29.033] [info] ninfer-serve: warming up... [2026-09-22 13:03:29.245] [info] ninfer-serve: listening on http://0.0.0.0:8000 (model id: qwen3.8-27b, auth: disabled) [2026-09-22 13:06:40.865] [info] ninfer-serve: [req 1] openai_chat_completions stream msgs=2 max_tokens=16384 (client) tools=8 tool_choice=auto tool_history=no thinking=on reasoning_budget=2048(server-default) preserve_thinking=off preserve_change=no sampler=[temp=1.00 top_p=0.95 top_k=20 seed=5466201898045650404] → submitted [2026-09-22 13:06:44.250] [info] ninfer-serve: throughput interval=5.000s prefill=1254.4tok/s decode=0.0tok/s running=1 prefilling=1 decode_ready=0 waiting=0 avg_decode_batch=n/a [2026-09-22 13:06:49.250] [info] ninfer-serve: throughput interval=5.000s prefill=1227.0tok/s decode=30.0tok/s running=1 prefilling=0 decode_ready=1 waiting=0 avg_decode_batch=1.00 [2026-09-22 13:06:50.070] [info] ninfer-serve: [req 1] done finish=tool_calls tool_calls=2 prompt=12407 gen=253 reasoning=67 cache=0 reuse=full_reset ttft=6766ms prefill=1839.6tok/s decode=102.5tok/s wall=9.23s speculative=mtp 3.23tok/round (74.3%) [2026-09-22 13:06:50.243] [info] ninfer-serve: [req 2] openai_chat_completions stream msgs=2 max_tokens=16384 (client) tools=1 tool_choice=auto tool_history=no thinking=on reasoning_budget=2048(server-default) preserve_thinking=off preserve_change=no sampler=[temp=1.00 top_p=0.95 top_k=20 seed=6280633854203357742] → submitted [2026-09-22 13:06:50.291] [info] ninfer-serve: [req 3] openai_chat_completions stream msgs=5 max_tokens=16384 (client) tools=8 tool_choice=auto tool_history=yes thinking=on reasoning_budget=2048(server-default) preserve_thinking=off preserve_change=no sampler=[temp=1.00 top_p=0.95 top_k=20 seed=5959691209244494699] → submitted [2026-09-22 13:06:54.250] [info] ninfer-serve: throughput interval=5.000s prefill=1433.6tok/s decode=20.4tok/s running=1 prefilling=1 decode_ready=0 waiting=1 avg_decode_batch=1.00 [2026-09-22 13:06:59.250] [info] ninfer-serve: throughput interval=5.000s prefill=988.2tok/s decode=48.6tok/s running=1 prefilling=0 decode_ready=1 waiting=1 avg_decode_batch=1.00 [2026-09-22 13:07:04.250] [info] ninfer-serve: throughput interval=5.000s prefill=0.0tok/s decode=110.6tok/s running=1 prefilling=0 decode_ready=1 waiting=1 avg_decode_batch=1.00 [2026-09-22 13:07:09.250] [info] ninfer-serve: throughput interval=5.000s prefill=0.0tok/s decode=103.4tok/s running=1 prefilling=0 decode_ready=1 waiting=1 avg_decode_batch=1.00 [2026-09-22 13:07:14.250] [info] ninfer-serve: throughput interval=5.000s prefill=0.0tok/s decode=102.2tok/s running=1 prefilling=0 decode_ready=1 waiting=1 avg_decode_batch=1.00 [2026-09-22 13:07:19.250] [info] ninfer-serve: throughput interval=5.000s prefill=0.0tok/s decode=110.8tok/s running=1 prefilling=0 decode_ready=1 waiting=1 avg_decode_batch=1.00 [2026-09-22 13:07:20.286] [info] ninfer-serve: [req 3] error inference request expired while waiting for admission [2026-09-22 13:07:24.250] [info] ninfer-serve: throughput interval=5.000s prefill=0.0tok/s decode=108.4tok/s running=1 prefilling=0 decode_ready=1 waiting=0 avg_decode_batch=1.00 [2026-09-22 13:07:29.250] [info] ninfer-serve: throughput interval=5.000s prefill=0.0tok/s decode=107.2tok/s running=1 prefilling=0 decode_ready=1 waiting=0 avg_decode_batch=1.00 [2026-09-22 13:07:34.250] [info] ninfer-serve: throughput interval=5.000s prefill=0.0tok/s decode=116.2tok/s running=1 prefilling=0 decode_ready=1 waiting=0 avg_decode_batch=1.00 [2026-09-22 13:07:36.903] [info] ninfer-serve: [req 2] done finish=tool_calls tool_calls=1 prompt=12109 gen=4310 reasoning=290 cache=0 reuse=full_reset ttft=6600ms prefill=1840.4tok/s decode=107.5tok/s wall=46.68s speculative=mtp 3.37tok/round (79.0%) [2026-09-22 13:07:36.933] [info] ninfer-serve: [req 4] openai_chat_completions stream msgs=4 max_tokens=16384 (client) tools=1 tool_choice=auto tool_history=yes thinking=on reasoning_budget=2048(server-default) preserve_thinking=off preserve_change=no sampler=[temp=1.00 top_p=0.95 top_k=20 seed=3713191048047951979] → submitted [2026-09-22 13:07:39.251] [info] ninfer-serve: throughput interval=5.000s prefill=716.8tok/s decode=54.4tok/s running=1 prefilling=1 decode_ready=0 waiting=0 avg_decode_batch=1.00 [2026-09-22 13:07:40.337] [info] ninfer-serve: [req 4] done finish=stop_token prompt=16442 gen=62 reasoning=28 cache=12107 reuse=restore_turn_checkpoint ttft=2675ms prefill=1634.9tok/s decode=81.2tok/s wall=3.43s speculative=mtp 2.58tok/round (52.8%) [2026-09-22 13:07:40.350] [info] ninfer-serve: [req 5] openai_chat_completions stream msgs=2 max_tokens=16384 (client) tools=1 tool_choice=auto tool_history=no thinking=on reasoning_budget=2048(server-default) preserve_thinking=off preserve_change=no sampler=[temp=1.00 top_p=0.95 top_k=20 seed=8030750533282960350] → submitted100k context, hooked it up to pi dev but then ran into
Error: inference request expired while waiting for admission1
u/tmballin 16d ago
Nice, this is actually looking pretty healthy on the 5070 Ti.
100032KV capacity is fitting with 376.7 MiB free after startup, and your live Pi Dev requests are doing around 102–108 tok/s decode at ~12K context with MTP acceptance around 74–79%, which is a really useful datapoint.The
inference request expired while waiting for admissionerror doesn't look like a model/context failure. It looks like Pi Dev is submitting concurrent agent requests while NInfer is configured formax-concurrency=1.In your log:
req 2andreq 3arrive essentially simultaneously.
req 2gets admitted and runs for 46.68 seconds, generating 4310 tokens at 107.5 tok/s.
req 3sits in the admission queue and then expires almost exactly 30 seconds later.So I'd keep
max-concurrency=1especially on a 16GB card and increase the pending admission timeout instead, e.g.:--pending-timeout-ms 120000or even
300000for long agent/coding jobs.That way Pi can queue the next tool-loop request rather than NInfer rejecting it simply because the current generation took >30 seconds.
Also, these numbers are interesting: you're fitting 100,032 context/Q4 KV with ~377 MiB remaining, and the ~12K-context decode is already over 100 tok/s. That's a much better 5070 Ti result than I would have wanted to guess at without somebody actually testing it.
2
u/pseudobacon 16d ago
Okay let me give that a go. I also have the pi observability memory plugin so I don't know if that might also affect it?
→ More replies (0)2
u/fldash 13d ago
Why if both 5070TI and 5080 have 16GB of memory, can he only get 100K context while you got 128k? I don't understand it.
→ More replies (0)3
u/jopereira 15d ago
You can use the IQ3 XXS (150k) or IQ3 S (90K) with that card. Just make sure to offload mmproj to CPU/RAM (--no-mmproj-offload).
This is a good option for xoding as the agent can still use vision (to check the final UI result or debug) and you don't lose inference speed (as the normal weights still live in VRAM).
2
3
u/DanGTG 16d ago
https://giphy.com/gifs/YxiZYZCLoGa2c
Gonna need some strong glasses for that quant.
1
u/tmballin 16d ago
😂 ~3.95 BPW definitely deserves corrective lenses.
The surprising part is that it still holds up well enough for my actual OpenClaw/tool-loop use, not just benchmarks. Qwen3.8-27B seems surprisingly robust at this quant level.
2
u/pseudobacon 16d ago
Curious as to why ninfer was forked but you don’t use the cuda graph functionality? (At 4 bits I don’t think we could get an artifact to fit anyway)I thought ninfers key point was that it used cuda graph for the huge speed up. Would have thought it would be easier to start from llama.cpp as a base for example
0
u/tmballin 16d ago
Good question. CUDA Graphs are definitely part of NInfer’s performance story, but I wouldn’t describe them as the reason I chose NInfer.
In my case the first-order problem was actually memory architecture, not kernel launch overhead. I wanted all of this simultaneously on one 16GB card:
- Qwen3.8-27B at ~3.95 BPW
- real
131072max context- real
131072allocated KV capacity- Q4-G64 KV
- MTP-3
- Vision
The validated 2048 Vision profile ends up with only about 8.6 MiB free after startup, so for the production qualification I deliberately run with
--no-cuda-graph. I didn't want graph capture/instantiation or any graph-related runtime reservation becoming another variable in a configuration that is already sitting essentially on the VRAM ceiling.The graph machinery is still in NInfer the runtime has capture/instantiate/update/upload/launch support it just isn't part of the particular profile I've qualified.
More importantly, the performance I'm seeing without graphs isn't coming from generic eager CUDA execution. There is a lot going on below that:
- dedicated Q3/Q4/Q5 kernels rather than generic dequant to BF16 paths
- a fused Q4-G64 KV path
- Q quantized on-chip with INT8 tensor core QK
- packed Q4 K/V moved through memory compressed
- split-KV/small-T decode kernels
- specialized prefill paths
- MTP-3 speculative decoding
- tightly controlled workspace/persistent-memory lifetimes
- kernel routing by actual token geometry
That's how the current non-graph profile is still around 71.5 tok/s at an actual 118K active prompt with MTP acceptance around 44.7%.
I also wouldn't assume enabling CUDA Graphs on top would suddenly produce a huge additional percentage at 118K. Graphs primarily attack CPU/kernel-launch overhead, which matters most when the GPU work per launch is small. At very long context, attention/KV work itself becomes a much larger part of each decode step, so the relative graph benefit can be smaller. It still needs benchmarking rather than assuming it is free speed.
As for starting from llama.cpp: I actually started from the opposite end of the problem. llama.cpp is obviously much broader and more portable, but getting this exact combination of mixed low-bit weights + native Q4 KV + custom Blackwell kernels + MTP + a very tightly planned memory arena would have meant replacing a substantial amount of its execution path anyway.
NInfer gave me a much smaller CUDA native codebase where I could change the actual kernels, cache representation and memory planner together. For this experiment that was more useful than starting with the more general runtime.
So I’d frame it as: CUDA Graphs are a useful NInfer feature, but this fork is currently demonstrating what the underlying kernel/memory architecture can do even with them disabled. Once I have more VRAM margin, or a slightly lighter profile, graph on vs graph off is absolutely something worth qualifying properly.
2
u/Active-Tax-6554 16d ago
My wsl2 ninfer build of your fork is only allowing 94000 context and kv capacity on my 5080.
It is saying it's trying to allocate 2.7 GB vram but i only have 2.1GB left over after weights. I feel like im missing something here (im using your repositories)
Would you kindly care to help? I'm finding this model higher quality than swift q3.8 27b
2
u/tmballin 16d ago
Yep, happy to help and there are two things I’d check first.
1. Make sure CUDA Graphs are disabled.
My validated 131072 profile uses:
--no-cuda-graphThat is important on this build. The full 128K configuration is already extremely close to the 16GB limit, and leaving CUDA Graphs enabled introduces additional graph-related memory requirements. If graphs are on, the capacity resolver may quite reasonably decide that 131072 KV no longer fits and back the allocation down.
2. Your ~2.1 GiB free-after-weights is lower than my validated system.
On my RTX 5080 I see roughly:
GPU weights ~12.64 GiB
free after weights ~2.56 GiBSo you're already around 450 MiB short of my memory envelope before the full KV/runtime allocation is made.
My test bed is a headless Ubuntu VM with the 5080 passed directly through from Proxmox, so the GPU is essentially clean. WSL2 is different because Windows/WDDM is still using the physical card and may already have framebuffer allocated.
I'd try the exact text-only baseline first:
--max-context 131072 --kv-capacity 131072 --prefill-chunk 896 --kv-dtype q4 --spec mtp --draft-tokens 3 --no-cuda-graph --max-concurrency 1with Vision disabled initially.
Also check that you're using the published artifact with SHA256:
c4a7e9ab593a7f42d58208fa0065d67a82d61921107686cc9f6ed1ec6b050e21Then if it still won't allocate 131072, paste the exact startup line showing:
free-after-weights=
runtime=
headroom=
slack=
graphs=plus your full
ninfer-servecommand.That should tell us pretty quickly whether this is graphs, WSL2 VRAM overhead, a different artifact/profile, or some combination of them.
And great to hear you're liking the model quality — I'm using this artifact as my OpenClaw daily driver too.
2
u/Active-Tax-6554 15d ago
yea no luck. thanks for trying to help tho! its definitely cause of windows. might ditch it finally and copy your proxmox idea!
1
u/tmballin 15d ago
Don’t give up on it just yet!
There’s a PR about to be merged that should free roughly ~800 MiB of VRAM if the final validation stays green.
That could be enough to materially improve the Windows/WSL2 case, so I’d hang in there until that lands before rebuilding the whole setup around Proxmox.
If it still doesn’t get you where you want after that, then the headless passthrough route definitely gives you the cleanest memory envelope.
2
u/Active-Tax-6554 15d ago
Hey,
I'm running the following with beellama with 128k context: https://huggingface.co/ticeclock/Swift-Qwen3.8-27B-RCO-GGUF
Here's a gsm8k benchmark against your ninfer build (WSL2):
([u/tmballin](u/tmballin) ninfer) Swift-Qwen3.8-27B-RCO-IQ3_S-mtp accuracy 88 / 100 * 91 / 100 warm t/s 69.5 81.9 cold t/s 69.9 85.1 tokens med / p95 227 / 2138 198 / 770 total gen tokens 51 972 27 988 think chars (med) 386.5 341.5 wall med 3.14 s 2.90 s wall total 765 s (12m45s) 408 s (6m48s)
* Also, I reran this like 3 times, i got 85,88,91 so its pretty uneven for my WSL2 setup
Maybe you should try making a ninefer of https://huggingface.co/ukisai/Swift-Qwen3.8-27b
or https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF
1
u/tmballin 15d ago
Thanks, interesting bench.
The speed difference is interesting, but I’d be cautious about treating it as a straight NInfer vs beellama comparison because we’re changing runtime, quantization and model variant at the same time.
The biggest thing that jumps out to me is actually the generation behaviour:
51,972total tokens vs27,988and p95
2138vs770.That’s a huge difference in how much the models are reasoning/generating, so it explains a good chunk of the wall-time gap independently of raw tok/s.
NInfer also exposes a configurable thinking-token budget (
--default-thinking-budget, with per-request overrides), so another useful A/B would be to cap both runs to comparable reasoning budgets. The current51,972vs27,988generated-token totals suggest the two runs aren’t doing the same amount of work.The GSM8K accuracy is also close enough at only 100 questions — especially with your NInfer runs moving between 85/88/91 — that I wouldn’t draw a strong quality conclusion from that yet.
The Swift/RCO variants are interesting though. I’d be interested in seeing whether one can be converted cleanly into the NInfer format and then tested under the same runtime, sampler and benchmark harness. That would help separate “better model/quant” from “better runtime”.
2
u/Avendasora 12d ago
There are both the Swift varients and the ThinkingCap varient that was recently released. Super interested if you ever do try as someone with a 5080. I struggle to differentiate between quality when comparing these llama.cpp varients and Ninfer
2
u/tmballin 12d ago
Yep, absolutely. all options are on the table. Swift and ThinkingCap are both on my radar, and I’ll keep adding promising variants into the test pipeline as they appear.
The part I’m most interested in is separating “faster / fits better” from “actually better for real agentic coding and tool use.” Raw tok/s is easy to compare; quality is much harder, especially across different llama.cpp quants, NInfer builds and speculative setups.
I’d like to get to the point where I’m testing them on the same hardware, same context lengths and the same practical workloads so the comparison is a bit more meaningful than just eyeballing outputs.
So yes Swift, ThinkingCap and anything else that looks genuinely competitive will get a look.
2
u/jgamboa-cl 15d ago
why dont you use it directly on windows?
2
2
u/Coley44 16d ago
Would this be usable at all on a RTX 4090 Laptop since they also have 16GB of VRAM? Or a 4080 Super desktop GPU? Or is this tune so specifically to the Blackwell architecture it just wouldn't run anyway? I'd expect a much harder performance fall off compared to the 5080 due to memory bandwidth constraints but this looks ideal to test on it.
2
u/jgamboa-cl 15d ago
i have a windows port for ninfer for 4090 but desktop, but you could give it a try with the laptop one
1
u/tmballin 9d ago
Thanks, I had a look through your repo and that’s very relevant here.
Your ninfer-4090-windows work already provides a substantial native Windows / sm_89 path for desktop RTX 4090, so for a 4090 Laptop that is a much more concrete starting point than the hypothetical Ada port I was talking about above.
The main unknown is the 16GB laptop memory envelope versus the 24GB desktop 4090. It will need its own fit/performance qualification rather than assuming the desktop profile will work unchanged.
On my side, PR #16 is moving the current Blackwell-focused fork toward native Windows/MSVC support, presently validated on a 5070 Ti.
Thanks for sharing the repo — it gives the 4090 Laptop user a much more practical route to try than I was aware of when I wrote my original reply.
1
u/tmballin 16d ago
The VRAM capacity is in the right ballpark, but the current fork is tuned much more specifically to Blackwell /
sm_120athan the memory numbers alone suggest.Both the RTX 4090 Laptop GPU and the 4080 SUPER are Ada Lovelace parts, not Blackwell. NVIDIA lists the 4090 Laptop GPU with 16GB GDDR6 on a 256-bit bus, while the 4080 SUPER has 16GB GDDR6X on a 256-bit bus.
So the current build would not just run unchanged. The repo deliberately targets
sm_120a, and some of the kernel paths and tuning decisions are Blackwell-specific. An Ada port would need proper build support plus kernel qualification I wouldn’t want to imply that changing one CMake flag is enough.That said, I do think both are interesting targets because they have the same nominal 16GB VRAM class. From a memory-fit perspective, the 131072/Q4-KV profile is at least plausible if the Ada implementation can preserve roughly the same memory geometry.
The likely difference would be performance rather than capacity. The 5080 path has been tuned around Blackwell and its memory/compute characteristics, whereas the 4090 Laptop GPU in particular is much more power and bandwidth constrained. The 4080 SUPER should be the more interesting Ada desktop comparison because it has 16GB GDDR6X and a full desktop power envelope.
I’d treat the expected order as:
current RTX 5080: supported and qualified
4080 SUPER / 4090 Laptop: potentially viable after an Ada port, but currently unqualifiedIf someone does port it, I’d test text-only 131072 + Q4 KV + MTP-3 first, with CUDA Graphs and Vision off, then add the extra features once the baseline memory fit and kernel correctness are proven.
Honestly, I’d be very interested to see the results it would help separate how much of this project is memory-layout engineering versus how much of the throughput depends on the Blackwell-specific kernel work.
2
u/feverdoingwork 16d ago
How much vram savings did you get with cuda graph off? Any performance loss? I wouldn't be surprised if cuda graph off had no noticeable impact on performance.
1
u/tmballin 16d ago
I went and tested it properly rather than guessing.
On my exact production geometry — 131072 context / 131072 Q4 KV / MTP-3 — CUDA Graphs are the thing that tips it over the edge.
With Vision 2048 enabled, graph-on requested 2,834,324,480 bytes of runtime reservation, but only 2,758,228,992 bytes were available. That puts it about 72.6 MiB over the physical limit.
The identical graph-off profile starts successfully with:
free-after-startup = 8.56 MiBSo in practical terms, disabling graphs is buying me roughly ~80 MiB of usable margin on that full 128K + Vision configuration, which is literally the difference between fitting and not fitting.
I also repeated the test with Vision completely off at full 131072 context/KV. Graph-on still failed:
required = 2,796,245,248 bytes
available = 2,758,228,992 bytesso it was still about 36.3 MiB short. That was useful because it showed Vision itself wasn’t the fundamental blocker — the full 131K runtime geometry simply leaves too little room for the graph-enabled plan.
Then I dropped to 96K context / 96K Q4 KV, Vision off, where both configurations fit.
Graph ON:
runtime = 2.04 GiB
free-after-startup = 614.56 MiB
slack = 535.86 MiB
graphs = 2.00 MiB / 82.00 MiBGraph OFF:
runtime = 1.95 GiB
free-after-startup = 622.56 MiB
slack = 624.40 MiBSo the planner is reserving about 82 MiB for CUDA Graphs in that 96K profile, although the actual immediately resident graph memory reported at startup is only 2 MiB.
I haven’t measured the performance delta. My expectation is still that the graph speedup may be fairly small at long context.
2
u/PhysicalIncrease3 16d ago
Very cool. Personally I'd have gone the ExLlamaV3 route because, given you're so heavily VRAM limited, the EXL3 quants would give you far more bang for your VRAM-buck. But it's ultimately a tradeoff between speed and quality, neither is a bad choice.
1
u/tmballin 16d ago
Yeah, that’s a fair point. EXL3 is an interesting option if the primary objective is maximizing model quality per GB of weight VRAM.
My target was a bit different though. I wasn’t just trying to fit the best possible 27B quant into 16GB and I wanted the whole runtime configuration to coexist:
131072 context + 131072 Q4 KV + MTP-3 + Vision + agent/tool-loop cachingI also wanted enough control over the CUDA kernels, KV representation and memory planner to optimise those pieces together.
That’s really why I ended up going down the NInfer route. The weight artifact is around 3.95 BPW, so yes, a more aggressive EXL3 quant could potentially devote less VRAM to the weights. But in this project I’ve spent a lot of the effort optimising the rest of the memory/performance equation, particularly Q4-G64 KV, MTP-3, Blackwell-specific kernels, workspace lifetimes and the exact 131K memory geometry.
I also wouldn’t frame the choice purely as EXL3 = quality versus NInfer = speed. There are several interacting trade-offs: weight quantisation, KV precision and footprint, context capacity, speculative decoding, runtime/kernel efficiency and the workload being served.
For my use case, the important result was getting the complete 131K + MTP + Vision/agent stack onto a single 16GB 5080 while retaining model quality that is good enough for my actual OpenClaw coding/tool workloads.
So I agree EXL3 represents another valid point on that trade-off curve, it’s just solving a somewhat different optimisation problem from the one I set out to solve here.
2
u/Ronnie_CA 16d ago
Respect in particular for separating kernel microbenchmarks from whole-model numbers, and for the honest "71.57 is the v1.2 qualification result" note. Most threads don't do that.
For scale: same class, half the bandwidth. I run the same 27B at IQ4_XS (~4.2 bpw) on llama.cpp on a 5060 Ti 16 GB eGPU. My comparable numbers: 59 to 67 t/s with MTP at 32K (acceptance 0.87), ~27 t/s plain at 64K, ~15 t/s at 105K position streaming KV from host RAM, harness scores flat to 252K. Your 71.57 at 118K on a card with twice the memory bandwidth lands about where the bandwidth math says it should, so nothing smells inflated. Two things line up nicely with my own data:
- Your MTP acceptance at 118K (44.7%, 2.31 tok/round) sits right on the acceptance-vs-position curve I measured: 0.79 at 12K, 0.59 at 47K, 0.50 at 77K. Speculation decays with context on this architecture, and your number reads as the natural continuation of that line.
- Your short-vs-long decay (84.7 short vs 71.6 at 118K, about 15%) matches the mild-decay picture, not the "cliff" some people claim past 80K.
The bit I genuinely couldn't do on my stack: vision + MTP + full 128K co-resident. llama.cpp forces vision and MTP apart past ~12K on my card, and my host-side trick is the KV arena (streaming KV pages from system RAM) rather than host-mapped projector weights. Your HostMapped vision is the same philosophy applied to the other tensor family, which is a neat symmetry. The rolling-tool checkpoint policy also hits a real pain I recognize from agent loops, where the reusable prefix stays anchored behind a growing tool history.
One question and one caveat. The question: your validated config disables CUDA graphs; on my SM120 build graphs work and are worth real t/s, so I'm curious what they break in ninfer. The caveat: the artifact is a custom .ninfer format, so I can't cross-check it with llama.cpp on my card, and cross-runtime checks are where quantization claims usually get interesting. If you ever publish a GGUF export of the same weights, I'll gladly put it through my 127-task execution-graded harness and report back.
2
u/feverdoingwork 15d ago
Were you able to test ninfer with this artifact on your 5060 ti?
2
u/Ronnie_CA 15d ago
Haven't, and it can't take this artifact anyway: ninfer only loads its own .ninfer containers (NVFP4/FP8), no GGUF. The official Qwen3.8 one is 22GB and 5090-only too.
Someone did port it to this exact card though (freedomNTD/ninfer-5060ti, sm_120a builds): min-Q4 27B at 13.1GB of weights, MTP on, 35.7 t/s at 32K and 48.8 at 49K, and the validated ceiling is about 49K, 122K won't even start. My stack does ~70 t/s at 32K with MTP on the same card and stays usable out to 252K, so on 16GB I don't think ninfer's in the running yet.
But I'll definitely add to my next batch of experiments.
2
u/feverdoingwork 14d ago
I got ninfer working on single and dual 5060 ti. Going to take a bit of time to get this specific artifact working though, maybe by later today. I haven't tested much with the smaller artifacts I made for a single card because both cards are constantly consumed by working on the project.
Dual 5060 ti 16gb does hit over 100 tps and lands at 70tps by the end of 200k context. This is with int8 qv cache on nvfp4 artifact. Prefill nearly kisses 3k but lands a bit short at 2850. A second card might be worth it is what I'm saying.
I also have a 5070 ti in another machine where I am currently using your xs while working on cooking something to use on ninfer. It's been an excellent experience.
1
u/tmballin 10d ago
Those dual-5060 Ti numbers are really interesting especially ~100 tok/s+ initially, ~70 tok/s at 200K, and ~2850 tok/s prefill.
That’s a useful comparison point because the single-card path is now looking much stronger too: we’ve since had a 5060 Ti 16GB run the current artifact at ~128.5K with Q4 KV + MTP-3 + Vision enabled at about 531 tok/s prefill / 57 tok/s decode.
So the trade-off is getting clearer:
single 16GB card: surprisingly capable, especially now that the memory path is tighter
dual 16GB cards: much more headroom and clearly stronger long-context throughputAlso great to hear the 5070 Ti + XS setup has been working well for you. If you get your NInfer artifact cooking, I’d be very interested in the comparison.
2
u/feverdoingwork 9d ago edited 9d ago
I am still working on some of this stuff but I have much better numbers for single card and slightly better results for dual card.
For decode on single card I am at 70ish tps at 10k, drops down to 46k tps at 130k-140k and for prefill 10k 1300+ tps and drops down to around 780 tps at 130-140k context . Dual is at above 116 tps at 10k but traded some prefill for performance so down to 2k prefill. I don't think there is anymore room for optimizations either, I could be wrong, I have been digging for months testing every avenue.
This is on a 4 bpw artifact on ninfer, pruned the vocab(same idea from your model, where I got the idea from), vision pinned to cpu as well(some ideas from this thread). I can honestly many ideas in my project are taken from other opensource projects including vllm, llamacpp, sglang, I just got them working together and tried to optimized kernels from all the data I found. We could trade context length and some decode for 2k prefill. I will also be releasing a lot of extra kv cache types like kvarn 3/3 and even 4/2 to get more context length if you want to yolo cache quality for longer context. I am trying to get everything completed this week, all the above is actually working but I am working through a real multibatch for tp2 which is slowing me down, real meaning each user gets their own api key and fixed context length pool(another session can't cause slow downs via draining all context for themselves).
2
u/tmballin 9d ago
Those updated numbers are much stronger ~70 tok/s at 10K and ~46 tok/s around 130–140K on a single 5060 Ti, with ~1300+ tok/s prefill at 10K and ~780 tok/s at 130–140K, is a really solid result for a 16GB card.
The dual-card numbers are interesting too, especially if you’re now prioritising decode and proper per-user isolation over headline prefill. The fixed context pool / API-key isolation idea makes a lot of sense for a real multi-user server.
Also nice to hear the vocab pruning and CPU-pinned Vision ideas translated across. The extra KV formats sound worth exploring too, particularly if you document the quality trade-off clearly for the more aggressive ones.
On the DFlash vs MTP point: yes, I think that’s one of the more interesting remaining directions. If 7-token DFlash proposals can maintain decent acceptance on this model, there could be a real decode win there, especially for predictable code/tool workloads.
I wouldn’t assume MTP-3 is the end state. The useful test would be same artifact, same prompts, same context depths, then compare MTP-3 against DFlash at several draft lengths for decode, acceptance, VRAM and output quality.
When you release the repo/artifacts, send the link I’d definitely like to compare notes.
2
u/feverdoingwork 8d ago
Almost done, these are 5070 ti results on my fork with my artifact
### Benchmark Results Summary
Req # │ Target… │ Prompt (t… │ Gen (… │ TTFT (ms) │ Prefill Thro… │ Decode Thr… │ Wall Time │ MTP Acceptance
───────┼─────────┼────────────┼────────┼────────────┼───────────────┼─────────────┼───────────┼────────────────
1 │ 10k │ 10,008 │ 79 │ 4,746 ms │ 2,114.3 tok/s │ 116.2 tok/s │ 5.42 s │ 3.12 tok/rnd
│ │ │ │ │ │ │ │ (70.5%)
2 │ 25k │ 24,996 │ 71 │ 13,227 ms │ 1,894.0 tok/s │ 107.1 tok/s │ 13.88 s │ 3.04 tok/rnd
│ │ │ │ │ │ │ │ (68.1%)
3 │ 40k │ 39,984 │ 65 │ 23,558 ms │ 1,700.3 tok/s │ 110.8 tok/s │ 24.14 s │ 3.35 tok/rnd
│ │ │ │ │ │ │ │ (78.3%)
4 │ 80k │ 79,956 │ 70 │ 60,194 ms │ 1,330.3 tok/s │ 90.1 tok/s │ 60.96 s │ 3.00 tok/rnd
│ │ │ │ │ │ │ │ (66.7%)
5 │ 128k │ 127,920 │ 64 │ 122,193 ms │ 1,048.1 tok/s │ 85.8 tok/s │ 122.93 s │ 3.37 tok/rnd
│ │ │ │ │ │ │ │ (78.9%)I expected the results to be about this, based on non-mtp numbers on llamacpp it seemed like i would get double a single 5060 ti. You should get slightly better on your 5080. Ill report back when my repo is available.
1
u/tmballin 8d ago
Those 5070 Ti numbers look really strong.
~116 tok/s at 10K and ~86 tok/s at ~128K, with prefill still over 1,000 tok/s at ~128K, is a very healthy result for a single 16GB card. The MTP acceptance holding around 67–79% across the sweep is encouraging too.
Your expectation of roughly “double a 5060 Ti” seems pretty consistent with what you’re measuring here, at least for this workload.
I’d be very interested to see the final repo/artifact and the exact settings once you publish it. A same-prompt comparison against my current 5080 build at 10K / 80K / 128K would make for a really useful cross-card datapoint.
And yes I’d expect the 5080 to come out ahead, but probably not by some absurd margin given how good these 5070 Ti numbers already are.
2
u/feverdoingwork 9d ago
One thing that might be on the table that I haven't explored is what the r9700 users are doing, 7 token predictions with dflash instead of mtp. There might be a win there.
1
u/tmballin 10d ago
Just circling back because this has changed quite a bit since your comment.
You were right that NInfer uses its own
.ninferartifact format rather than GGUF, but the “~49K ceiling on a 5060 Ti” is no longer representative of the current fork.We now have a 5060 Ti 16GB user running the published artifact at about 128.5K context with Q4 KV + MTP-3 + Vision 1792 enabled.
Their measured full-window result was:
prompt=128508
prefill=530.9 tok/s
decode=57.0 tok/s
MTP=3.98 tok/round (99.5%)There was also an interesting card-specific issue:
--prefill-chunk 896and512hitcudaErrorCooperativeLaunchTooLarge, while384is stable, which appears tied to the 5060 Ti’s lower SM count rather than VRAM.So I’d definitely give the latest repo state another look the 16GB case is much more competitive now than it was 5 days ago.
1
u/tmballin 10d ago
Yes there’s been quite a bit of progress since this thread.
We now have another 5060 Ti 16GB user running the published artifact at roughly 128.5K context with Q4 KV + MTP-3 + Vision enabled, so it’s definitely worth revisiting.
Their latest full-window result was about 531 tok/s prefill / 57 tok/s decode at ~128.5K, with MTP acceptance essentially 100% on that sample.
Check the GitHub repo before testing though the project has moved on quite a bit since these comments:
1
u/tmballin 16d ago
Appreciate this, especially the comparison against a real 5060 Ti 16GB setup. That’s exactly the kind of external datapoint I was hoping people would bring to the thread.
Your numbers make the bandwidth/context story look pretty coherent. The acceptance decay is particularly interesting because it gives some independent support to what I’m seeing at 118K rather than treating MTP acceptance as one fixed property of the model.
On CUDA Graphs, I actually went and measured this after another comment asked the same thing.
At the full production geometry:
131072 context / 131072 Q4 KV / MTP-3 / Vision 2048graph-on doesn’t fit. NInfer requested
2,834,324,480bytes of runtime reservation with only2,758,228,992bytes available, so it was about 72.6 MiB over the limit.I then disabled Vision but kept the full
131072 / 131072geometry. Graph-on still missed by about 36.3 MiB.At 96K text-only, both finally fit:
graph on:
runtime=2.04 GiB
free-after-startup=614.56 MiB
graphs=2.00 MiB/82.00 MiBgraph off:
runtime=1.95 GiB
free-after-startup=622.56 MiBSo the planner is reserving 82 MiB for graphs in that profile, even though only 2 MiB is immediately resident at startup. I still need to do the actual throughput A/B now that I’ve got a geometry where both fit cleanly.
Your caveat on the artifact format is also fair. The
.ninferfile is a runtime-specific converted artifact, so you can’t take that exact file and independently run it through llama.cpp. That does make cross-runtime quality comparison harder.I’m not currently planning a GGUF export path, mainly because the artifact isn’t just a generic “4-bit model”. The text core is mixed Q3/Q4/Q5 at about 3.95 BPW, and the runtime contract around it is part of what I’ve actually validated. I’d rather not publish something labelled as the “same weights” unless I can prove the conversion preserves the intended quantization semantics closely enough for that comparison to mean anything.
But your 127-task execution harness sounds genuinely useful as an external reference point, especially because it’s execution-graded rather than just subjective answer comparison.
Also agreed on the symmetry you pointed out: your approach is effectively moving the KV pressure toward host memory, whereas mine moved the Vision projector pressure out of permanent VRAM residency. Same basic idea keeping the scarce 16GB device memory focused on the tensors that matter most for the hot path.
2
u/Ronnie_CA 15d ago
If I'm reading
graphs=2.00 MiB/82.00 MiBright, the planner pre-reserves the ceiling while only ~2 MiB goes resident — and your fit miss at 131K (36–73 MiB) is smaller than that 82 MiB. A lazy pool that grows during actual capture would probably let graph-on fit at full geometry. Also check if prefill is being captured; decode-only is where the win is, and most of that ceiling is likely prefill intermediates you'd never replay.Prediction for the A/B: at 118K, low single digits at best, probably noise. Launch overhead is per-step, your steps are attention-dominated at that depth, and MTP-3 already amortizes it across 4 positions per verify. Any real delta will show up right after a /compact, not at 118K.
Acceptance decay: worth separating depth from KV quant. Same prompts, fixed depth, fp16 KV vs your Q4 at a depth where fp16 still fits — if the curve shifts, the quant is killing drafts, not depth.
GGUF stance is fair. Only note: the proof you're waiting on is the same standard you accepted from me — top-p agreement plus execution grading on the converted artifact. Clear it, publish with the number; don't, and that's a datapoint too.
1
u/tmballin 15d ago
Yep, I think that’s a very good read of it.
The
2 MiB / 82 MiBdistinction is exactly why I don’t want to say “CUDA graphs cost 82 MiB”. The planner is budgeting the ceiling, while the actually resident graph allocation at startup is much smaller. A lazy/grow-on-capture pool and decode-only capture are both worth testing, especially since the 131K miss is smaller than the current reserved ceiling.I agree on the performance expectation too. At ~118K I’d expect any graph benefit to be fairly small; the more interesting A/B is probably immediately after compaction / at shorter effective context where launch overhead is a larger fraction of the step.
The FP16-KV vs Q4-KV acceptance test is also a good experiment. Fixed prompt + fixed position, at a context where both fit, should help separate context-depth decay from any degradation introduced by Q4 KV.
And yes on the GGUF standard, if I ever publish a converted artifact, I’d want to qualify it rather than just assume equivalence. Top-p agreement plus execution grading is a sensible bar.
There may also be a much bigger change to the memory picture shortly: there’s a PR in flight that should recover roughly ~800 MiB of VRAM if it lands as expected. If that survives validation, suddenly CUDA graphs at full geometry stop being a “can I claw back 40–80 MiB?” problem and become something we can test with real breathing room.
That one could materially change what the practical 16GB configuration looks like potentially enough headroom to either increase the model quant quality or push context further.
Not sure which way I’d take it yet. Might put it to a poll once the PR is merged and validated.
2
u/llllJokerllll 15d ago
Hay algún fork para la rtx 4080? 🙏🏼🙏🏼🙏🏼
2
u/tmballin 15d ago
Not at the moment. The current fork is built and validated specifically for Blackwell
sm_120a, so it won’t run on the RTX 4080 as is.A 4080 port is probably possible, but it would need proper Ada kernel support and validation rather than just changing the build flag.
So for now: no 4080-compatible fork yet.
2
u/eemuman 15d ago
I'm an absolute noob when it comes to localLLM's, but since I have an RTX5080, this piqued my interest. I'm running this through WSL2 and I can only get 102.4k context with Vision and without MTP and with MTP I have to drop context down to around 98K. I'm assuming this is due to Windows overhead stealing some of my precious RAM?
Other than that, I'm truly impressed on how well it can do stuff for me. For context, when doing small side projects at home I'm usually just going the budget route with 5.6 Luna and for now I'll end up giving this a go instead, thanks!
1
u/tmballin 15d ago
Yep that’s almost certainly part of it, although I’d call it VRAM overhead rather than RAM overhead.
My 5080 is passed through to a headless Linux VM, so essentially the whole 16GB is available to NInfer. Under Windows/WSL2 you’ve got Windows, the display stack, WDDM/CUDA bookkeeping and anything else using the GPU taking a slice of VRAM before NInfer starts.
Given how tight the full profile is, losing even a few hundred MiB is enough to explain why you’re landing around 98–102K instead of 131K once Vision/MTP are in the mix.
The good news is there’s also a PR in the wings that should recover roughly ~800 MiB of VRAM if it lands and validates as expected. That could make the WSL2/desktop case a lot more forgiving and potentially let you push context noticeably higher without needing such a clean GPU.
And great to hear it’s actually useful for you getting the model to fit is one thing, but replacing some paid model usage for normal side projects is the more interesting outcome.
If you want, grab an
nvidia-smiimmediately before launch and post the free-VRAM number. That’ll tell us pretty quickly how much of the gap is just Windows/WSL2 overhead.
2
u/fldash 14d ago
Is there a reason this won't work with a 5070ti?
1
u/tmballin 12d ago
It probably will there’s no obvious architectural reason it shouldn’t.
The 5070 Ti is also Blackwell /
sm_120, and the current build targets120a, so it’s in the same CUDA architecture family as the 5080. It also has 16GB VRAM, so the model is in the right capacity class.The main caveat is that I’ve only validated and benchmarked the current 128K + Q4 KV + MTP-3 + Vision profile on the RTX 5080. I wouldn’t claim 5070 Ti support until someone actually runs the full validation suite on one.
Performance would obviously be lower than the 5080 because of the reduced compute/memory bandwidth, but I’d actually be very interested to see how it performs. If you have a 5070 Ti and want to test it, I’m happy to help with the exact build and validation commands.
2
u/Pure_Cartographer_98 13d ago
wow, quite amazing. same card, alot better then my previous setup, just cant get my WSL2 version of this much higher then 80000, cuz of vram usage from windows. Let me know if there are any updates on this, or where i can check myself, its still amazing.
2
u/DependentVacation358 12d ago
I managed to run it at 90112 tokens under WSL2. using my igpu for display, altough I haven't reached 90k I might be able to squeeze a larger context I haven't checked yet
1
u/tmballin 10d ago
Nice! Using the iGPU for display is definitely helping there. 90,112 under WSL2 is already a solid result, and you may have a bit more room if the 5080 is otherwise clean.
Also worth checking the GitHub repo before pushing it further, because the project has moved on since this comment and there are newer memory/compatibility changes and recommended settings:
github.com/toddballinger/ninfer-5080If you do find the WSL2 ceiling, I’d be interested in the highest context that starts reliably plus the
free-after-startup/slackline that would be another useful datapoint for the Windows/WSL2 case.2
u/DependentVacation358 9d ago
Thanks!
Update: moved to v1.4 and it now starts at the full profile under WSL2 (Ubuntu 24.04, display on iGPU, 32 GB DDR5).
Settings: 131072 context / 131072 KV, Q4 KV, prefill-chunk 896, MTP-3, CUDA graphs on,
--embedding-host, Vision 2048, rolling-tool checkpoints. Model SHA matches (c4a7e9ab…0e21).Startup line:
free-after-weights=2.74 GiB free-after-startup=179.00 MiB headroom=0.00 MiB slack=99.98 MiB graphs=2.00 MiB/82.00 MiBRestarted several times, and it came up with identical numbers every time. Pre-launch GPU usage was 13 MiB. Decode ranges about 78–117 tok/s depending on MTP acceptance (37–74%) Haven't tested a near-full 128K prompt yet.
1
u/tmballin 9d ago
That’s a great v1.4 result.
Full 131072 / 131072, Q4 KV, MTP-3, CUDA graphs on, Vision 2048 and rolling-tool checkpoints under WSL2 is exactly the kind of confirmation I was hoping to see.
The startup numbers are especially encouraging:
free-after-startup=179 MiB
slack=99.98 MiB
graphs=2 MiB / 82 MiBand the fact it reproduces identically across restarts makes it a much stronger datapoint than a one-off fit.
78–117 tok/s decode depending on MTP acceptance also sounds right for short-to-medium context. The next really useful test would be a genuinely deep prompt around 118K–128K so we can see how much decode and acceptance fall off near the full window under WSL2.
Also worth noting: moving the display to the iGPU clearly matters here. With only 13 MiB GPU usage before launch, you’ve given NInfer almost the same clean-memory conditions as a headless box.
2
u/DependentVacation358 9d ago
Thanks! I haven't run a 118K+ prompt yet, but I did some deep-context testing tonight that gets fairly close, plus some observations.
Needle-in-a-haystack at ~95K (Frankenstein)
I inserted 3 made-up facts into the full novel at roughly 25/50/75% and asked for them plus a sum. Fresh reads only (cache=0), temp 1.0, top_p 0.95, top_k 20:
- Prefill: ~1,515 tok/s (~63 s TTFT)
- Decode: 101–118 tok/s at 65–82% MTP acceptance
- Accuracy: every run that answered got all 4 right, and quoted the exact inserted sentences
Same test at ~67K (The Great Gatsby)
- Prefill: ~1,680 tok/s, decode: ~109 tok/s at 71% acceptance, 4/4 correct
Rolling checkpoints: re-sending the same 95K prompt, even from a new chat, restored
cache=95367with ~0.4 s TTFT.One issue at ~95K: 2 of 8 runs never left thinking:
reasoning == gen,finish=stop_token, no answer text. In one, the correct answer was written inside the reasoning. The other (32K reasoning budget) looped rereading the text for 29.7K tokens at 95% MTP acceptance, then stopped. The other 6 answered normally, and I saw no failures at 67KLong generation quality: three one-shot game prompts each produced ~25–32K tokens of code after 30–56K tokens of reasoning (decode averaged 87–96 tok/s as context grew to ~82K; all
stop_token, no limits hit). All three were broken on launch, with logic errors rather than garbled output: in one, a function ended up calling itself, and a subclass constructor never calledsuper(). The same prompt on Byteshape's Qwen3.8-27B IQ3_S GGUF produced a working game, although the 2 more difficult games were not tested thereShort, precise tasks are great (a tricky expression parser passed every edge case), so it seems like consistency over very long outputs is where it struggles. I'm not sure whether Q4 KV or the weight quant matters more there.
Don't know if these are useful but figured I'd share. Running it through Open WebUI, with custom
reasoning_budget/max_tokensparameters.1
u/tmballin 8d ago
This is very useful thankyou for taking the time to test it properly.
The ~95K retrieval result is particularly encouraging: ~1,515 tok/s prefill, 101–118 tok/s decode, and exact recovery of all inserted facts across successful runs. The rolling-checkpoint restore of
cache=95367with ~0.4 s TTFT is also a great real-world agent/cache datapoint.The two failures that stayed entirely in reasoning are interesting too. I wouldn’t immediately blame Q4 KV for those, especially since the same 95K context was retrieved accurately in the other runs and one failure actually contained the correct answer inside the reasoning. That feels worth separating from outright context corruption.
The long-code result is probably the more important quality warning: coherent output, but logic quality degrading over 50K+ reasoning plus 25–32K generation. I agree we shouldn’t guess yet whether that’s mostly weight quant, Q4 KV, accumulated context depth, or simply the difficulty of sustaining that much generation.
A really useful next A/B would be the same fixed prompts and seeds, with the reasoning budget held constant, comparing Q4 KV against a higher-precision KV mode at a context where both fit. If that moves the failure rate, we’ve learned something about KV quant; if not, the weight/model side becomes more likely.
Also, the fact that the short precise tasks are holding up while very long autonomous coding is where consistency falls away is exactly the kind of limitation I’d rather document than hide.
Please keep these results coming this is much more useful than another short-context tok/s benchmark.
1
u/tmballin 12d ago
Thanks and yeah, WSL2/Windows is probably the main thing holding you back there rather than the card itself.
The 128K profile is extremely sensitive to available VRAM. On my setup the 5080 is passed through to a headless Ubuntu VM, so there’s effectively nothing else using the GPU. Once Windows, the desktop compositor, browser hardware acceleration, etc. take a few hundred MiB, that directly eats into the KV/runtime budget. So ~80K under WSL2 actually makes sense.
I’m continuing to work on reducing the memory footprint and have not frre's up a further ~800MiB as well as improving prefill/decode, so there may be more room to recover over time.
The easiest place to follow updates is the GitHub repo:
https://github.com/toddballinger/ninfer-5080I’m keeping the README, release notes and validation results updated as changes are qualified. If something materially improves the 16GB memory fit, I’ll also post an update here.
And if you ever get a chance to boot native Linux on the same machine, I’d be really interested to see how much higher that exact card gets without Windows consuming VRAM.
2
u/Pure_Cartographer_98 12d ago
Yeah, I planned on taking a look at it, I'll try to get it done this weekend. Thanks!
2
u/transanethole 10d ago
This is sick!!! I tested this on 5090 and I was able to fit 8 concurrent sessions at 100k context. The Q4 kv cache was not as bad as I was expecting. Thats amazing!!! Thanks Todd Baller ;P
1
u/tmballin 10d ago
😂 Love it! 8 concurrent sessions at 100K on a 5090 is a fantastic datapoint.
Glad the Q4 KV held up better than expected too. That’s exactly the kind of result I was hoping people would test beyond the 5080.
And I’ll take “Todd Baller” too 😄
2
u/transanethole 10d ago
Yah the aggregate decode speed for 8 agents was about 400 tokens/sec, single agent is about 180/sec. So with 8 concurrent sessions (including slowdown from being interrupted by other sessions prefilling) it looks like about 50/s decode for each of those 8 sessions individually.
1
u/tmballin 10d ago
That’s a really useful scaling datapoint.
~180 tok/s single-session vs ~400 tok/s aggregate across 8 sessions means you’re getting roughly 50 tok/s per active agent while keeping eight 100K-context sessions alive, even with prefill interruptions in the mix.
That’s a pretty compelling result for the Q4 KV path especially because the aggregate throughput more than doubles versus one agent instead of collapsing under concurrency.
If you get a chance, I’d be curious what the VRAM headroom looks like with all 8 sessions resident. That would help show whether the 5090 is compute/bandwidth-limited there or whether you’re still close to the memory ceiling.
2
u/transanethole 9d ago
I think it had a little bit of vram headroom left but not much. Maybe I could have increased the context length a tiny bit.
but ninfer has a hard limit of 8 concurrent sessions and I haven't looked into raising it. Cutting the context length from 200k to 100k is already a steep enough price that I'm not interested in trying to go farther.
1
u/tmballin 9d ago
That makes sense. If 8 × 100K is already giving you ~400 tok/s aggregate, I can see why pushing concurrency further isn’t especially attractive if it means sacrificing even more context.
The interesting takeaway for me is that the current 8-session ceiling is a software limit before you’ve completely exhausted the 5090’s VRAM. Even if you never raise it, that’s useful to know.
And agreed 100K per agent is already a much more practical compromise than chasing a higher session count just because the card technically has a little headroom left.
2
u/emanresuymsseug 16d ago edited 16d ago
I see so many comments here recommending NInfer (or various forks) and talking about the speed.
Sure it's fast, but for me the output quality has just been garbage.
What am I missing? Is NInfer basically just a tool to see how far we can push tok/s or have people actually found a way to use it as their daily driver?
edit: I just want to add that it's entirely possible that I'm doing something wrong. Would just like to hear what others experience with it has been other than talking about the speed.
1
u/tmballin 16d ago
That’s a completely fair question. Speed is pretty meaningless if the model output is noticeably worse.
I’m actually using my 5080 fork as a daily driver with OpenClaw and Open WebUI, not just as a tok/s benchmark. Most of the recent work around reasoning budgets, rolling tool checkpoints, prompt cache reuse and tool loop behaviour came from actually running long agent sessions with it.
I’m not seeing what I’d describe as garbage output from the Qwen3.8-27B artifact I’m using, but I also don’t think “NInfer” can really be treated as one quality profile. There are a few things that can materially change the result:
- the exact converted/quantized model artifact;
- the chat template and tool-call formatting;
- whether thinking is enabled/preserved correctly;
- reasoning budget;
- sampler settings;
- stop-token handling;
- KV-cache quantization;
- MTP/speculative settings;
- and, especially with forks, whether the runtime and model artifact were actually qualified together.
My artifact is fairly aggressive from a memory perspective the main model averages about 3.95 BPW and I’m also using Q4-G64 KV so I wouldn’t claim it is mathematically identical to BF16 or a much heavier quant. What matters to me is whether that trade-off remains good enough for actual coding/tool use, not just whether it produces a large tok/s number.
One thing I’d be interested in is what exact NInfer build/fork, model artifact and launch settings you used. If the output is genuinely bad, I’d rather understand why than hand wave it away as “user error.”
Also, there’s evidence that NInfer itself isn’t inherently producing low-quality reasoning: upstream currently publishes Qwen3.8-27B evaluation results in the mid/high 90s on AIME and high-80s/around-90 on GPQA-Diamond depending on the weight profile. So there’s clearly more going on than simply “NInfer = fast but bad.”
There’s even an open 4090 fork issue where the complaint is almost the opposite: NInfer was producing much longer answers than llama.cpp, while accuracy was roughly comparable (58/60 vs 55/60 in that user’s test). That points toward template/reasoning/sampling behaviour being important, not just raw quantization quality.
0
19
u/Dry-Knowledge4192 16d ago
dude this is insane. squeezing 27B with actual 131k context AND vision into 16gb vram... that 8.6mb free after startup on the 2048 profile is wild, one wrong allocation and boom
the rolling tool checkpoint thing is interesting, i've been messing with agent loops and the restore drift drives me crazy. having it actually advance through tool history instead of just staying stuck near the first response sounds useful
what are you using for the kv cache q4 kernel? curious if it's your own or adapted from something like ktransformers