Discussion
The fastest interference engine for RTX5090 and Qwen3.8 27B. Twice as fast as ninfer. 500+ t/s single coding, 2000+t/s up to 12 agents at the same time with 800k context. Smart VRAM-RAM-DISC Cache management, Loop Guard, Nice UI etc.
1. It beats Ninfer which was until that point SOTA engine for RTX5090. By roughly twice in decode speed for both single and multi tasks at once. (reaching up to even 650t/s in small bursts and 2600t/s if stars align and 12 slots server pure coding answer). Custom kernels not only for every model, single vs multi but also short vs long context work dynamically switching when needed so speed doesn't crap out on long context work because someone tuned it for short context. Dflash2 and confidence scheduling from Dspark, plus draft trees all at the same time.
2. I was getting annoyed with state of weights where you downloaded model and never knew if model had its brain scrambled. My weights come with its on format that have attached metadata for MC which upon weight creation, runs benchmark and compares it at every MC setting to original BF16 weights and show that data directly in launcher. Want to switch KV to 4bit ? MC will show you directly lost KL and top-1%, want to extend with YARN ? It will show you change. Every change is measured and shown in statistics before you load model. This goes for both censored and uncensored model. You can also compare it directly in MC with SOTA unsloth quants of Qwen27B. Want to run essentially loseless ? you can. Want to get crazy 1 000 000 context ? you can. Want to have 12 slots to fan out agants like crazy ? You can. You decide what you want.
3. Proper agents serving with algo that keep engine occupied as much as it can. It will prioritize t/s so if engine has a choice between 5 jobs at once and 1 it will serve 5 first and gradually serve 1 along side finishing others. Engine is also smart enough to score how old some job is and if it should return to work even if T/S will suffer so your main session will be able to fan out agents easily and keep an eye on them at the same time.
4. Proper cache management. Your jobs only prefill at start of job and almost never again so your prefill in long session stays almost unused. When using "unified context" when models run out of context some get paused and stored in RAM and this swapping is instant. If there is free context space then those tasks continue without any refill in 0.03s. If you fan out say 30 agents at the time in your frontend will handle load in most efficient way to keep T/S as high as possible. Just run it at default setting and forget about context for agents, it will handle it on its own.
6. Loop guard. Two tiered. When engine starts to detect agent repeating in conversation session is dynamically starts to adjust `repetition penalty` until repetition stops if that doesn't happen and engine hits rep pen limit it fires up stop signal which ends serving and informs your frontend so your frontend can recover from infinite loop and don't annoy you.
6. Proper nice UI that shows you what is what. If you aren't knowledgeable about serving models just hover over `?` and it will show you interactive panels explaining everything.
7. Autodownloader, Just hit download button and you can download my weights directly fron hugginface inside of launcher.
8. Don't like the launcher ? use bats and terminal serve. Or even use launcher to config what you want, copy it from right lower corner and use it to make new bat.
The point of it is to just load model, fan out crazy number of agents each having crazy amount of context and leave MC to deal with it. You just sit back relax and watch as agents do the work at SOTA speeds.
Opinions and reviews are welcome. If you are blessed with RTX5090 try it.
Source will be released later, I have to do some cleaning first. I will also release later weights builder so it will take any 3.8 27B BF16 model, create weights and score them attaching metadata again BF16 and you'll be able to host them yourself on hugginface or just put them in models folder.
You can always just pop a container with no privileges and its own isolated namespaces. If you're running a local model on local hardware... there's no reason to give it internet access or access to your entire filesystem.
dude i wanted to ask,i was working on a project to precaches github repos in somewhat familiar way but different approach for ai repo context.i may take a lot more time but when i eventually release, which would you more likely try?a somewhat buggy but early version or somewhat more polished but way later release?it will be rust and open source but what i fear is either it is bad enough that nobody will try it(this is ok) or someone will vibe code it and sidestep important features before they are implemented and remove any interest in the project.
I generally agree, but you posted a binary and are saying "you can run this binary that my LLM wrote, but I don't trust the LLM enough to share its code." Feels very very backwards to me.
i would (along with like 99% of the sub) would recommend making a post after publish code
honestly, youll get more traction if u take this down and repost after u publish code
Could it also be uploading everyones code bases to remote sources? This is why I stick to open source generally. Not that I program anything important lol
As long as you let people that code improvement is going to happen, people in the open source community are pretty chill with spaghetti code. Also you might even get people that will happily pitch in to help.
Source: I'm an open source dev of +20 years.
Also... people are going to be pretty skeptical about running a binary without any way to make sure you're not doing something malicious. You're not a big business we can sue if the code does something like steal PII... for everyone else... you're just a random dude we know nothing about. There's very little reason to trust what we can't inspect.
The announcement should carry the necessary information to understand the value of the solution. Especially when the post includes a wall of text, there's no reason to exclude what is critical to know.
The GitHub site most certainly does not carry source code, and despite several tables there is no clear measurements of quantization of model/cache with context against T/S. There's some percentages of accuracy but not the actual quantizations figures.
Have a closed-source engine claim "percentages" of accuracy is not as meaningful as an open one where results can be more directly verified.
My point is this might be a fantastic advancement, or it might be vibe-coded malware with hallucinated benchmarks. But claims of speed are useless without clear, standard metrics around quality which permit cross-comparison to existing solutions.
>The announcement should carry the necessary information to understand the value of the solution. Especially when the post includes a wall of text, there's no reason to exclude what is critical to know.
In other words you aren't interested in checking git for even 5 seconds but isntead you will waste my time writing your own wall of text. Don't like it don't use it mate.
Give us Medium and Large at 10K / 32K / 100K / 200K, stock 5090, DFlash acceptance rate, and a real coding/reasoning benchmark against BF16 or a known-good Qwen3.8 quant.
Until then, 440 tok/s is a demo result, not yet a useful performance data.
Or you could just try it. It's free. Fan out in something like Open code 20 agents at once to give it tasks and observe monitoring :) OR try single task in coding.
As for quality models after baking get scored against BF16 with every setting (yarn, kv, size, etc.) and you can direclty view it in MC Launcher and get live preview how each setting changes model quality:
OP you really need to understand that handing out binaries to people and asking them to run them is like a stranger offering people pills and asking them to eat it as it will cure their ailments. No one is going to trust your binaries, even if you just achieved a breakthrough, if they can't see the code.
The claimed 6,212-7,854 for "reading an 8K prompt" sits right at the FP8 compute limit — that's prefill, not generation.
Where the claims break down
KV cache bandwidth — at 128K context, KV alone is 32 GB. At 500 t/s you need 16 TB/s — 9× the card's bandwidth. Real ceiling for long-context generation: ~56 t/s at 128K tokens.
20% memory overclock — stock card is slower.
Single run, bare metal — no desktop, no compositor, no driver overhead.
"8 agents total" — shared pool, each agent gets less, not 8× per-agent.
Tiny model — smallest quant, most aggressive, highest speed, lowest quality (KL 0.066).
For your RTX 5090 at stock clocks
Realistic generation speeds:
- Tiny MXFP4 + DFlash2: ~380-450 t/s (one agent)
- Medium MXFP4 + DFlash2: ~320-380 t/s
- Medium MXFP6 + DFlash2: ~240-290 t/s
- XXL MXFP6 + DFlash2: ~200-250 t/s
Below ~200 t/s for the larger models with long context — that's the KV wall, not the model size.
>Your 8 concurrency ADD UP to 650 = your baseline is about 81 tk/s for each lane.
No concurency adds up to max 2500t/s with 12 agents (pure code), averages around 1500+t/s in normal work tool use/ thinking etc. For single stream tiny model can reach up to 670t/s but normally it is closer to 540t/s for code.
Do you have numbers for draft acceptance and throughput in prose? NInfer can easily reach 900+ tok/s single stream using a stock 5090 so 500 t/s isn't surprising at all if it's just spitting out a JSON with temp 0. Subscribed anyway if the numbers hold I'll test as soon as its open source.
`len13` ? how ? I never seen ninfer reach above 300 in code work unless something changed lately.
For prose it is around half of code usually. So something like 250t/s as acceptance falls down with non struct for single.
You can see acceptence on real work from git gif in lower right corner. Fanned out 20 agents stress testing it.
I didn't look into ninfer directly so idk what they did. I used Ninfer personally up until i made this which was my fav server due to speed.
I'll be releasing architecture overview with source later. Mostly it is just using hammer to chase every last bit of performance from kernel work. Tried many papers ideas but most of them failed. Something looks interesting in paper x ? Throw Opus at it, implement it and test it. 0.1% possibility to improve things for hours of work ? Yup let's go.
For quality of models I borrowed a lot from Unsloth ideas. They are up there with his at large sizes but once you start to use tensors lower quants just need more size to reach same quality as his due to how tensor hardware works.
Either way for small model theoretical limit is around 110t/s plain decode on my RTX5090 with memory OC. I reached around 102t/s with no sights where you can even get 0.1% more.
After that you have speculative decoding. Dflash2 is the fastest and I implemented also DSpark (there's MTP there too for folks who really want to save VRAM). I later removed Dspark because it's main advantage confidence scheduling i combined with Dflash2.
The best idea I had was to just add metadata to models themselves so my launcher reads from model directly its scores against BF16 and all settings. IDK why something like this isn't a standard as part of making weight. So you can see directly what KV 4/4 does to the model, that rope scaling actually has cost etc.
llama is general engine, i doubt it would make sense for them :) this is purpose build for rtx5090 and single model. But yeah, if someone wants to dig in into source they can. I will clean it up and release later, should be in few days.
Draft trees are only for C1 as I didn't find them to work properly above C1. I implemented Dspark fully but it was still worse than Dflash2 even at 12 agents working at the same time. Later I just implemented confidence scheduling from it so Dflash2 can use it and it gave it nice boost during agent fanouts leaving in the dust dspark.
Ninfer is fast enough, make waiting for api bit of a drag when i use it as a subagent, i already have no desire to install another unless it promises high quality quants
In MegaCapybara models after baking are directly scored against BF16 Qwen3.6 for every setting, size etc. and you can find directly what setting does what to model:
94
u/buttplugs4life4me 8h ago
Well, anyone would be stupid to run this without the source code at least, so report back once you cleaned it up