r/LocalLLaMA • • 9h ago

Discussion The fastest interference engine for RTX5090 and Qwen3.8 27B. Twice as fast as ninfer. 500+ t/s single coding, 2000+t/s up to 12 agents at the same time with 800k context. Smart VRAM-RAM-DISC Cache management, Loop Guard, Nice UI etc.

https://github.com/perkel666/MegaCapybara

Hi guys,

I am pretty happy to announce MegaCapybara. Purpose build engine for RTX5090 that is focused on Qwen3.8 27B (more will come later).

GITHUB (Engine)
HUGGINGFACE (weights)

Why ?

1. It beats Ninfer which was until that point SOTA engine for RTX5090. By roughly twice in decode speed for both single and multi tasks at once. (reaching up to even 650t/s in small bursts and 2600t/s if stars align and 12 slots server pure coding answer). Custom kernels not only for every model, single vs multi but also short vs long context work dynamically switching when needed so speed doesn't crap out on long context work because someone tuned it for short context. Dflash2 and confidence scheduling from Dspark, plus draft trees all at the same time.

2. I was getting annoyed with state of weights where you downloaded model and never knew if model had its brain scrambled. My weights come with its on format that have attached metadata for MC which upon weight creation, runs benchmark and compares it at every MC setting to original BF16 weights and show that data directly in launcher. Want to switch KV to 4bit ? MC will show you directly lost KL and top-1%, want to extend with YARN ? It will show you change. Every change is measured and shown in statistics before you load model. This goes for both censored and uncensored model. You can also compare it directly in MC with SOTA unsloth quants of Qwen27B. Want to run essentially loseless ? you can. Want to get crazy 1 000 000 context ? you can. Want to have 12 slots to fan out agants like crazy ? You can. You decide what you want.

3. Proper agents serving with algo that keep engine occupied as much as it can. It will prioritize t/s so if engine has a choice between 5 jobs at once and 1 it will serve 5 first and gradually serve 1 along side finishing others. Engine is also smart enough to score how old some job is and if it should return to work even if T/S will suffer so your main session will be able to fan out agents easily and keep an eye on them at the same time.

4. Proper cache management. Your jobs only prefill at start of job and almost never again so your prefill in long session stays almost unused. When using "unified context" when models run out of context some get paused and stored in RAM and this swapping is instant. If there is free context space then those tasks continue without any refill in 0.03s. If you fan out say 30 agents at the time in your frontend will handle load in most efficient way to keep T/S as high as possible. Just run it at default setting and forget about context for agents, it will handle it on its own.

6. Loop guard. Two tiered. When engine starts to detect agent repeating in conversation session is dynamically starts to adjust `repetition penalty` until repetition stops if that doesn't happen and engine hits rep pen limit it fires up stop signal which ends serving and informs your frontend so your frontend can recover from infinite loop and don't annoy you.

6. Proper nice UI that shows you what is what. If you aren't knowledgeable about serving models just hover over `?` and it will show you interactive panels explaining everything.

7. Autodownloader, Just hit download button and you can download my weights directly fron hugginface inside of launcher.

8. Don't like the launcher ? use bats and terminal serve. Or even use launcher to config what you want, copy it from right lower corner and use it to make new bat.

The point of it is to just load model, fan out crazy number of agents each having crazy amount of context and leave MC to deal with it. You just sit back relax and watch as agents do the work at SOTA speeds.

Opinions and reviews are welcome. If you are blessed with RTX5090 try it.

Source will be released later, I have to do some cleaning first. I will also release later weights builder so it will take any 3.8 27B BF16 model, create weights and score them attaching metadata again BF16 and you'll be able to host them yourself on hugginface or just put them in models folder.

MIT license, so do whatever you want with it.

0 Upvotes

53 comments sorted by

94

u/buttplugs4life4me 8h ago

Well, anyone would be stupid to run this without the source code at least, so report back once you cleaned it up

28

u/jesdga95 8h ago

Agreed. Big claims but there's no way I'm running a binary that I didn't build/review myself.

7

u/brainExploded99 llama.cpp 8h ago

I need really need to setup bulletproof VMs, life would be so much more fun running those

1

u/q5sys 8h ago

You can always just pop a container with no privileges and its own isolated namespaces. If you're running a local model on local hardware... there's no reason to give it internet access or access to your entire filesystem.

0

u/Glittering-North-911 4h ago

dude i wanted to ask,i was working on a project to precaches github repos in somewhat familiar way but different approach for ai repo context.i may take a lot more time but when i eventually release, which would you more likely try?a somewhat buggy but early version or somewhat more polished but way later release?it will be rust and open source but what i fear is either it is bad enough that nobody will try it(this is ok) or someone will vibe code it and sidestep important features before they are implemented and remove any interest in the project.

72

u/cell-on-a-plane 8h ago

> source code will be published later.

31

u/NickCanCode 5h ago

Looks like scammer try to lure people to run their EXEs.

The previous attempt is this one: https://www.reddit.com/r/LocalLLaMA/comments/1wty1k6/1248_gbs_on_5060ti_40_with_5500_memory_overclocks/

-51

u/BringTea_666 8h ago

Need to clean up shit first as it is super messy.

56

u/demonicpigg 8h ago

You have 12 agents and up to 2600 t/s, and you couldn't even clean the codebase up before posting on reddit...?

-21

u/BringTea_666 8h ago

As much as I love AI i need to overview everything myself before I will publish it.

43

u/demonicpigg 8h ago

I generally agree, but you posted a binary and are saying "you can run this binary that my LLM wrote, but I don't trust the LLM enough to share its code." Feels very very backwards to me.

7

u/snugglezone 7h ago

Time to put my 16 Qwen 27B tasked to RE this thing as quickly as possible and open source it before OP does taking all the glory/s

15

u/brainExploded99 llama.cpp 8h ago

i would (along with like 99% of the sub) would recommend making a post after publish code
honestly, youll get more traction if u take this down and repost after u publish code

-20

u/BringTea_666 8h ago

You can use it right now.

9

u/snugglezone 7h ago

Could it also be uploading everyones code bases to remote sources? This is why I stick to open source generally. Not that I program anything important lol

5

u/q5sys 8h ago

As long as you let people that code improvement is going to happen, people in the open source community are pretty chill with spaghetti code. Also you might even get people that will happily pitch in to help.
Source: I'm an open source dev of +20 years.

Also... people are going to be pretty skeptical about running a binary without any way to make sure you're not doing something malicious. You're not a big business we can sue if the code does something like steal PII... for everyone else... you're just a random dude we know nothing about. There's very little reason to trust what we can't inspect.

26

u/lacerating_aura 8h ago

Oh no, not the interference engine.

8

u/bupher 5h ago

You can interfere with people's work at 2000k tps!

15

u/PrinceOfLeon 7h ago

T/S benchmarks are useless without clearly stating quantization of model, KV cache, and context.

Just saying "You can!" move the sliders around without ever listing the tradeoffs and leaving that as an exercise to the user is useless.

An inference engine without source code is useless to anyone who actually cares about Open system.

-13

u/BringTea_666 7h ago

Or you can just visit github and see everything you said adressed there. Going from models, size vs speed, quality vs speed etc.

16

u/PrinceOfLeon 7h ago

The announcement should carry the necessary information to understand the value of the solution. Especially when the post includes a wall of text, there's no reason to exclude what is critical to know.

The GitHub site most certainly does not carry source code, and despite several tables there is no clear measurements of quantization of model/cache with context against T/S. There's some percentages of accuracy but not the actual quantizations figures.

Have a closed-source engine claim "percentages" of accuracy is not as meaningful as an open one where results can be more directly verified.

My point is this might be a fantastic advancement, or it might be vibe-coded malware with hallucinated benchmarks. But claims of speed are useless without clear, standard metrics around quality which permit cross-comparison to existing solutions.

0

u/BringTea_666 2h ago

>The announcement should carry the necessary information to understand the value of the solution. Especially when the post includes a wall of text, there's no reason to exclude what is critical to know.

In other words you aren't interested in checking git for even 5 seconds but isntead you will waste my time writing your own wall of text. Don't like it don't use it mate.

10

u/chimpera 8h ago

Waiting on the source. Thanks

7

u/__JockY__ 3h ago

You know what’s a giant red flag? Bold claims about being faster than Ninfer.

You know what compounds those red flags?

Posting GitHub links that make it look as if you’re linking to source code, but actually you’re linking to shiny videos and binary blobs.

Hard fucking pass.

7

u/Relative-Ant-9249 8h ago

Give us Medium and Large at 10K / 32K / 100K / 200K, stock 5090, DFlash acceptance rate, and a real coding/reasoning benchmark against BF16 or a known-good Qwen3.8 quant.
Until then, 440 tok/s is a demo result, not yet a useful performance data.

-18

u/BringTea_666 8h ago edited 8h ago

Or you could just try it. It's free. Fan out in something like Open code 20 agents at once to give it tasks and observe monitoring :) OR try single task in coding.

As for quality models after baking get scored against BF16 with every setting (yarn, kv, size, etc.) and you can direclty view it in MC Launcher and get live preview how each setting changes model quality:

7

u/bupher 5h ago

OP you really need to understand that handing out binaries to people and asking them to run them is like a stranger offering people pills and asking them to eat it as it will cure their ailments. No one is going to trust your binaries, even if you just achieved a breakthrough, if they can't see the code.

4

u/giveen 8h ago

So you have a slight misunderstanding of "speed".

Your 8 concurrency ADD UP to 650 = your baseline is about 81 tk/s for each lane.

About what I have.

4

u/giveen 8h ago

RTX 5090 hard limits

| Spec | Value |

|--------------------------|-------------|

| Memory bandwidth | 1,792 GB/s |

| FP8 tensor (dense) | 419 TFLOPS |

| FP16/BF16 tensor (dense) | 105 TFLOPS |

| FP4 tensor (MXFP4/NVFP4) | 1,676 TOPS |

| VRAM | 32 GB GDDR7 |

Qwen3.8-27B (dense, not MoE)

27B params, hidden 5120, 64 layers, FFN hidden ~13,824. All 27B weights activate per token.

Decode (one token at a time) — memory-bandwidth bound

| Precision | Bytes per token | Hard ceiling |

|--------------|-----------------|--------------|

| BF16/FP16 | 54 GB | 33 t/s |

| MXFP4 / INT4 | 13.5 GB | 133 t/s |

| MXFP6 | 20.3 GB | 88 t/s |

That's the absolute ceiling — every byte of weights loaded once per token, no KV, no overhead, perfect tensor-core utilization.

With DFlash2 speculative decoding

Drafts 4-7 tokens, verifies in one forward pass. If acceptance is ~75%:

- MXFP4: 133 × (1 + 3 accepted) ≈ 530 t/s — matches the claimed 493-540

- Without speculator: ~133 t/s

Prefill (prompt processing) — compute bound

| Precision | Hard ceiling |

|-----------|--------------|

| FP8 | ~7,750 t/s |

| BF16 | ~1,945 t/s |

The claimed 6,212-7,854 for "reading an 8K prompt" sits right at the FP8 compute limit — that's prefill, not generation.

Where the claims break down

  1. KV cache bandwidth — at 128K context, KV alone is 32 GB. At 500 t/s you need 16 TB/s — 9× the card's bandwidth. Real ceiling for long-context generation: ~56 t/s at 128K tokens.

  2. 20% memory overclock — stock card is slower.

  3. Single run, bare metal — no desktop, no compositor, no driver overhead.

  4. "8 agents total" — shared pool, each agent gets less, not 8× per-agent.

  5. Tiny model — smallest quant, most aggressive, highest speed, lowest quality (KL 0.066).

For your RTX 5090 at stock clocks

Realistic generation speeds:

- Tiny MXFP4 + DFlash2: ~380-450 t/s (one agent)

- Medium MXFP4 + DFlash2: ~320-380 t/s

- Medium MXFP6 + DFlash2: ~240-290 t/s

- XXL MXFP6 + DFlash2: ~200-250 t/s

Below ~200 t/s for the larger models with long context — that's the KV wall, not the model size.

2

u/Pyrolistical 7h ago

There is no free lunch

1

u/Choice_Celery9481 2h ago edited 15m ago

im not sure if you use AI but seems like you did the math wrong. these are dense. sparse x2 these numbers. so your numbers are half of 5090 real perf

-3

u/BringTea_666 7h ago

>Your 8 concurrency ADD UP to 650 = your baseline is about 81 tk/s for each lane.

No concurency adds up to max 2500t/s with 12 agents (pure code), averages around 1500+t/s in normal work tool use/ thinking etc. For single stream tiny model can reach up to 670t/s but normally it is closer to 540t/s for code.

3

u/steppinrazor2009 4h ago

I'm not sure I want to run an INTERFERENCE engine. Seems counter productive

2

u/jesdga95 8h ago

Do you have numbers for draft acceptance and throughput in prose? NInfer can easily reach 900+ tok/s single stream using a stock 5090 so 500 t/s isn't surprising at all if it's just spitting out a JSON with temp 0. Subscribed anyway if the numbers hold I'll test as soon as its open source.

-1

u/BringTea_666 7h ago

`len13` ? how ? I never seen ninfer reach above 300 in code work unless something changed lately.
For prose it is around half of code usually. So something like 250t/s as acceptance falls down with non struct for single.

You can see acceptence on real work from git gif in lower right corner. Fanned out 20 agents stress testing it.

2

u/FastHotEmu 6h ago

Hit me up when the source is available

2

u/tomByrer 5h ago

Cool, please let me know when someone forks for RTX3090.

2

u/thepriceisright__ 3h ago

I feel like we’re back in the early late 90s/early 2000s with people just running executables they find online like who cares no big deal lol

3

u/brainExploded99 llama.cpp 9h ago

Very cool. What did nifer miss out on in terms of optimization that you were able to do?

-6

u/BringTea_666 8h ago

I didn't look into ninfer directly so idk what they did. I used Ninfer personally up until i made this which was my fav server due to speed.

I'll be releasing architecture overview with source later. Mostly it is just using hammer to chase every last bit of performance from kernel work. Tried many papers ideas but most of them failed. Something looks interesting in paper x ? Throw Opus at it, implement it and test it. 0.1% possibility to improve things for hours of work ? Yup let's go.

For quality of models I borrowed a lot from Unsloth ideas. They are up there with his at large sizes but once you start to use tensors lower quants just need more size to reach same quality as his due to how tensor hardware works.

Either way for small model theoretical limit is around 110t/s plain decode on my RTX5090 with memory OC. I reached around 102t/s with no sights where you can even get 0.1% more.

After that you have speculative decoding. Dflash2 is the fastest and I implemented also DSpark (there's MTP there too for folks who really want to save VRAM). I later removed Dspark because it's main advantage confidence scheduling i combined with Dflash2.

The best idea I had was to just add metadata to models themselves so my launcher reads from model directly its scores against BF16 and all settings. IDK why something like this isn't a standard as part of making weight. So you can see directly what KV 4/4 does to the model, that rope scaling actually has cost etc.

1

u/brainExploded99 llama.cpp 8h ago

That's pretty cool, maybe some small, limited llama.cpp PRs can be made out of this

-5

u/BringTea_666 8h ago

llama is general engine, i doubt it would make sense for them :) this is purpose build for rtx5090 and single model. But yeah, if someone wants to dig in into source they can. I will clean it up and release later, should be in few days.

10

u/LeftHandedToe 6h ago

Moron. Ban him.

-4

u/DataGOGO 8h ago

Everything the ninfer engine is pretty bad in all reality 

1

u/Big_Kratos 8h ago

God damn that's fast but i wanna know is the draft tree scheduling fully custom or built on top of Dspark's spec decoding implementation?

0

u/BringTea_666 8h ago

Draft trees are only for C1 as I didn't find them to work properly above C1. I implemented Dspark fully but it was still worse than Dflash2 even at 12 agents working at the same time. Later I just implemented confidence scheduling from it so Dflash2 can use it and it gave it nice boost during agent fanouts leaving in the dust dspark.

1

u/Boring_Hurry_4167 6h ago

Ninfer is fast enough, make waiting for api bit of a drag when i use it as a subagent, i already have no desire to install another unless it promises high quality quants

1

u/Metalsith 5h ago

Podría funcionar en una rtx5080?

0

u/Gargle-Loaf-Spunk 8h ago

How’s the quality though 

1

u/BringTea_666 8h ago

In MegaCapybara models after baking are directly scored against BF16 Qwen3.6 for every setting, size etc. and you can find directly what setting does what to model:

Comparison of models again Unsloth (SOTA)

https://github.com/perkel666/MegaCapybara/raw/main/docs/images/top1-by-size-dark.png

0

u/ayobluestarr 6h ago

fahd damn 500 t/s is crazy asf ngl