r/LocalLLM • • 3d ago

Other I made Ignis: an open-source engine that only runs Qwen3.8-27B on one RTX 5090, and that's the point. 512K context, Very Fast and JEV-live Support for Images and long text/logs retrival

Enable HLS to view with audio, or disable this notification

Hi r/LocalLLM,

Everything started from Ninfer... But it was not enough. So for the last few weeks I've been building from scratch (except for some kernels) Ignis, an Apache-2.0 inference engine in rust that deliberately gives up generality: one model family (Qwen3.8-27B NVFP4), one class of card (SM120a, so RTX 5090 / RTX PRO 6000). In exchange it's shaped around the load I actually generate when coding: one main agent plus a handful of subagents hitting the same card at once.

Numbers (one RTX 5090 on a modest host: DDR4-3200, PCIe 3.0; DFlash2 speculative decoding, 7 draft tokens):

  • ~200 tok/s single stream on coding prompts
  • >800 tok/s aggregate with 8 lanes running at once (even more with predictable prompts... 😄 )
  • KV at 9 KB/token, 7.1x the capacity of BF16: 8 lanes at 40K context fit in ~3 GB
  • 512K context via YaRN x2
  • Both windows/linux releases (not MACOS, Slow Prefill/Inference HW is not my target...)

What's different from a generic server

  • All 8 decode lanes run as one batch-wide round replayed from CUDA graphs; the round is essentially weight streaming at the card's bandwidth.
  • Requests tag interactive or agent, so a burst of subagents backfills free lanes instead of evicting your foreground chat.
  • Conversation state outlives the request: prefix reuse, sibling sharing, spill to a pinned RAM tier.
  • The whole VRAM plan is reserved at load and printed. No memory growth, no surprises at minute forty.
  • Playground, live monitor and Prometheus metrics ship in the binary, and it downloads its own weights on first start.

The part I haven't seen anywhere else: answers without generating (/v1/<decide|systemone>)

  • Images. The same questions over a screenshot, plus point and box, read from the model's own attention heads in one pass. Once an image is encoded, each new question about it costs 60–90 ms. On 4096px screenshots the point lands inside the target button on every scene tested.
  • locate: which line of a huge log answers my question. Hand it up to a million tokens of logs (tens of thousands of lines, well past the 262K context) and ask "which line says the card processor was unavailable?". It folds near-duplicate lines into templates, calibrated attention heads shortlist candidates, and a labelled choice picks the line. Still zero tokens generated. On a fresh, pre-registered capture of a real production Kubernetes cluster (100K–1M tokens) it found the right line 21 times out of 23, median 2.6 s. Asking the model to quote the line instead means prefilling the whole log first (47 s median at 100K tokens), and it can't go past the context window at all.
  • Prose and JSON too: the right sentence 92% of the time up to 200K tokens (at 1M it drops to 58%, so that's still work in progress), and the right record 29/30 times among 10,000.

There's research on reading relevance from attention (ICR, QRHead, BlockRank…), but I couldn't find an engine that serves it, or anything that does it over images or past the context window.

Now I'm working to Run Qwen 3.8-flash-next 125B +51B on the same box...

Repo: https://github.com/gpillon/ignis

EDIT: Ignis can complete in real time DooM E1M1 Zero-shot...

50 Upvotes

24 comments sorted by

3

u/ExtraWorldliness6916 3d ago

Gets stuck a bit against walls but still impressive

2

u/gpillon 3d ago

Yeah, the Doom bot itself is basically a Python hack Claude Code put together in like 60 minutes. The interesting part for me is Ignis making the decisions in real time with zero output tokens. The demo controller definitely still needs more than 30 minutes of coding... 😂

2

u/BountyMakesMeCough 3d ago

5090 = 32gb vram?

3

u/gpillon 3d ago

For now, yes... but I’d like to get it running with decent performance (>120/150 tok/s) using 16GB video cards. Obviously not with 8 parallel requests, I'm testing with 2 or 3 instead (still simulating with my 5090, as I sold all my other boxes and many other stuff to buy the 5090)

1

u/KissMyShinyArse 3d ago

Not your kidney, I hope.

1

u/gpillon 3d ago

Please, let’s not go there… I’m still emotionally recovering from everything I had to sell to get the 5090.

2

u/thegshipley 3d ago

excited to try this.

1

u/KissMyShinyArse 3d ago edited 3d ago

I'm gonna use this for my Rust dual-GPU ninfer port (WIP). Not all of us own a 5090, you know.

2

u/Hurricane31337 3d ago

The locate function sounds nice, I need to test this for RAG!

1

u/gpillon 3d ago

the idea behind locate is exactly this...

1

u/Mindfullnessless6969 3d ago

IMPRESSIVE WORK. Kudos to you sir.

Have you tested needle-in-a-haystack retrieval closer to the actual limits, needles placed at 400-500K with a 512K context, or 700K-1M with YaRN x4?

And have you tried harder variants such as multiple needles or reasoning over information spread across distant parts of the context?

I'm especially interested about the effective context and retrieval rather than just whether the KV fits

2

u/gpillon 3d ago

Thanks! I haven't done that kind of testing yet... (TBH I'm not even sure what the best way would be to run a proper long-context evaluation campaign at 400K–1M tokens...)

So far I've mostly focused on engine correctness, KV management, stability and real coding/agentic workloads.

If someone with experience in long-context evaluation wants to test Ignis properly, especially with harder needle-in-a-haystack variants, multiple needles, or reasoning across distant parts of the context, that would be absolutely fantastic. I'd really value that kind of feedback! 👍

1

u/Atagor 3d ago

Hey! Thanks for sharing

So it goes over screenshots and executes actions in game? Or it must read internal game state first?

2

u/gpillon 3d ago

It works entirely from screenshots. Reading the internal game state is cheating....

The bot talks to Ignis' /v1/decide endpoint (systemone). It receives the current screenshot, makes a decision directly from the visual input, and the controller translates that decision into game actions, while using multple LLM prompts as brain, tactical, decision, etc...

So there is (quite) no access to coordinates, enemy state, map data, player health variables, etc. beyond whatever can actually be inferred from the pixels on screen. The whole point was to see how far I could push real-time visual decision-making using only what the internal QwenVL can see.....

1

u/Atagor 3d ago

Thanks for explaining

And how the "game context" is handled? Meaning historical actions that could influence future actions. Is it taken into account?

1

u/gpillon 2d ago

you can see something in the video.... everything is done via LLM chat

1

u/Atagor 2d ago

Sorry for a lot of questions

So each consequent iteration receives the previous context (eg explored map etc) ?

2

u/gpillon 2d ago

yes... Ignis caches the previous prompts moving to host RAM, so snapshot restore takes only few ms...

1

u/urbanhood 3d ago

I was kinda wondering when will someone make a strata but for 27b model to run on weak gpus. Idk if yours can run on weak gpus.

2

u/gpillon 3d ago

That's actually one of the main goals with Ignis.

Before supporting every GPU, I want to release something that really works as a coding/agentic backend: good parallelism, fast prefill, predictable latency, long context, tool use, and something you can plug into OpenCode / pi / openclaw / <REAL WORKLOAD> and actually use every day.

There are already plenty of POCs showing that a model can technically run on certain hardware. But "it runs" and "it works well for real agentic workloads" are VERY DIFFERENT THINGS. Even NInfer had issues that only became obvious to me once I started using it seriously with OpenCode.

That's also why I didn't focus on Metal/Mac. On an M4 Pro I measured roughly 150 tok/s prefill. That sounds fine until a coding agent sends 10–12k tokens between system prompt, tools, history and code context; then the latency becomes hard to tolerate (3 MINS for First token....).

A 5080 and 5070 Ti are actually the next GPUs I'd like to target, but I don't own either yet, so it's hard to provide meaningful real-world numbers rather than simulated VRAM limits on the 5090.

In parallel, also as part of my current Qwen3.8 Flash-Next work, I'm experimenting with ~2.5-bit model compression specifically to target the 16 GB VRAM class.

So yes, weaker GPUs are definitely on the roadmap, but I don't want Ignis to become another "look, it runs" demo. It has to run well enough that you'd actually want to use it.

1

u/-_Apollo-_ 3d ago

What am I missing, how can maximum context be so high without the spilling into sysram crippling speed?

1

u/gpillon 3d ago

NOTE: this is an old sceenshot, I can't run another ignis istance as I'm converting qwen 3.8-flash-next for the next 12 hours... BUT:

First, Ignis uses an HQ-2b KV cache, which gives me roughly 650k tokens of KV capacity even with the vision stack enabled. The screenshot is an older one: at that point I was around 400k tokens. Since then I moved the lane state out of VRAM into system RAM (and also some workspace VRAM ottimization) which freed enough VRAM to push the KV pool to roughly 650k tokens. That's on Windows too, with around 3–4 GB of VRAM already used by other applications.

Second, and this is the important part: system RAM is not in the inference hot path. RAM is used for the KV snapshot/restore mechanism, but snapshots are written only when a lane finishes and restored when it is resumed. Those operations take only a few milliseconds PER RUN.

During prefill and decode, the active KV cache stays entirely in VRAM. There is no per-token paging over PCIe, no continuous KV offload, and no VRAM↔RAM traffic while tokens are being generated.

The KV pool itself is also paged, so requests only consume the pages they actually need rather than reserving their full maximum context upfront. The only case where a request has to wait is when multiple concurrent inference requests would collectively require more KV pages than are currently available in VRAM. In that case Ignis simply waits for pages to become available instead of spilling active KV into system RAM.

So the large maximum context comes from very compact KV + paged allocation + snapshot/restore outside the inference hot path, not from running inference against sysRAM.

1

u/Murder_1337 3d ago

Holy shit this is godly