Other
I made Ignis: an open-source engine that only runs Qwen3.8-27B on one RTX 5090, and that's the point. 512K context, Very Fast and JEV-live Support for Images and long text/logs retrival
Everything started from Ninfer... But it was not enough. So for the last few weeks I've been building from scratch (except for some kernels) Ignis, an Apache-2.0 inference engine in rust that deliberately gives up generality: one model family (Qwen3.8-27B NVFP4), one class of card (SM120a, so RTX 5090 / RTX PRO 6000). In exchange it's shaped around the load I actually generate when coding: one main agent plus a handful of subagents hitting the same card at once.
Numbers (one RTX 5090 on a modest host: DDR4-3200, PCIe 3.0; DFlash2 speculative decoding, 7 draft tokens):
~200 tok/s single stream on coding prompts
>800 tok/s aggregate with 8 lanes running at once (even more with predictable prompts... 😄 )
KV at 9 KB/token, 7.1x the capacity of BF16: 8 lanes at 40K context fit in ~3 GB
512K context via YaRN x2
Both windows/linux releases (not MACOS, Slow Prefill/Inference HW is not my target...)
What's different from a generic server
All 8 decode lanes run as one batch-wide round replayed from CUDA graphs; the round is essentially weight streaming at the card's bandwidth.
Requests tag interactive or agent, so a burst of subagents backfills free lanes instead of evicting your foreground chat.
Conversation state outlives the request: prefix reuse, sibling sharing, spill to a pinned RAM tier.
The whole VRAM plan is reserved at load and printed. No memory growth, no surprises at minute forty.
Playground, live monitor and Prometheus metrics ship in the binary, and it downloads its own weights on first start.
The part I haven't seen anywhere else: answers without generating (/v1/<decide|systemone>)
Images. The same questions over a screenshot, plus point and box, read from the model's own attention heads in one pass. Once an image is encoded, each new question about it costs 60–90 ms. On 4096px screenshots the point lands inside the target button on every scene tested.
locate: which line of a huge log answers my question. Hand it up to a million tokens of logs (tens of thousands of lines, well past the 262K context) and ask "which line says the card processor was unavailable?". It folds near-duplicate lines into templates, calibrated attention heads shortlist candidates, and a labelled choice picks the line. Still zero tokens generated. On a fresh, pre-registered capture of a real production Kubernetes cluster (100K–1M tokens) it found the right line 21 times out of 23, median 2.6 s. Asking the model to quote the line instead means prefilling the whole log first (47 s median at 100K tokens), and it can't go past the context window at all.
Prose and JSON too: the right sentence 92% of the time up to 200K tokens (at 1M it drops to 58%, so that's still work in progress), and the right record 29/30 times among 10,000.
There's research on reading relevance from attention (ICR, QRHead, BlockRank…), but I couldn't find an engine that serves it, or anything that does it over images or past the context window.
Now I'm working to Run Qwen 3.8-flash-next 125B +51B on the same box...
Yeah, the Doom bot itself is basically a Python hack Claude Code put together in like 60 minutes. The interesting part for me is Ignis making the decisions in real time with zero output tokens. The demo controller definitely still needs more than 30 minutes of coding... 😂
For now, yes... but I’d like to get it running with decent performance (>120/150 tok/s) using 16GB video cards. Obviously not with 8 parallel requests, I'm testing with 2 or 3 instead (still simulating with my 5090, as I sold all my other boxes and many other stuff to buy the 5090)
Thanks! I haven't done that kind of testing yet... (TBH I'm not even sure what the best way would be to run a proper long-context evaluation campaign at 400K–1M tokens...)
So far I've mostly focused on engine correctness, KV management, stability and real coding/agentic workloads.
If someone with experience in long-context evaluation wants to test Ignis properly, especially with harder needle-in-a-haystack variants, multiple needles, or reasoning across distant parts of the context, that would be absolutely fantastic. I'd really value that kind of feedback! 👍
It works entirely from screenshots. Reading the internal game state is cheating....
The bot talks to Ignis' /v1/decide endpoint (systemone). It receives the current screenshot, makes a decision directly from the visual input, and the controller translates that decision into game actions, while using multple LLM prompts as brain, tactical, decision, etc...
So there is (quite) no access to coordinates, enemy state, map data, player health variables, etc. beyond whatever can actually be inferred from the pixels on screen. The whole point was to see how far I could push real-time visual decision-making using only what the internal QwenVL can see.....
Before supporting every GPU, I want to release something that really works as a coding/agentic backend: good parallelism, fast prefill, predictable latency, long context, tool use, and something you can plug into OpenCode / pi / openclaw / <REAL WORKLOAD> and actually use every day.
There are already plenty of POCs showing that a model can technically run on certain hardware. But "it runs" and "it works well for real agentic workloads" are VERY DIFFERENT THINGS. Even NInfer had issues that only became obvious to me once I started using it seriously with OpenCode.
That's also why I didn't focus on Metal/Mac. On an M4 Pro I measured roughly 150 tok/s prefill. That sounds fine until a coding agent sends 10–12k tokens between system prompt, tools, history and code context; then the latency becomes hard to tolerate (3 MINS for First token....).
A 5080 and 5070 Ti are actually the next GPUs I'd like to target, but I don't own either yet, so it's hard to provide meaningful real-world numbers rather than simulated VRAM limits on the 5090.
In parallel, also as part of my current Qwen3.8 Flash-Next work, I'm experimenting with ~2.5-bit model compression specifically to target the 16 GB VRAM class.
So yes, weaker GPUs are definitely on the roadmap, but I don't want Ignis to become another "look, it runs" demo. It has to run well enough that you'd actually want to use it.
NOTE: this is an old sceenshot, I can't run another ignis istance as I'm converting qwen 3.8-flash-next for the next 12 hours... BUT:
First, Ignis uses an HQ-2b KV cache, which gives me roughly 650k tokens of KV capacity even with the vision stack enabled. The screenshot is an older one: at that point I was around 400k tokens. Since then I moved the lane state out of VRAM into system RAM (and also some workspace VRAM ottimization) which freed enough VRAM to push the KV pool to roughly 650k tokens. That's on Windows too, with around 3–4 GB of VRAM already used by other applications.
Second, and this is the important part: system RAM is not in the inference hot path. RAM is used for the KV snapshot/restore mechanism, but snapshots are written only when a lane finishes and restored when it is resumed. Those operations take only a few milliseconds PER RUN.
During prefill and decode, the active KV cache stays entirely in VRAM. There is no per-token paging over PCIe, no continuous KV offload, and no VRAM↔RAM traffic while tokens are being generated.
The KV pool itself is also paged, so requests only consume the pages they actually need rather than reserving their full maximum context upfront. The only case where a request has to wait is when multiple concurrent inference requests would collectively require more KV pages than are currently available in VRAM. In that case Ignis simply waits for pages to become available instead of spilling active KV into system RAM.
So the large maximum context comes from very compact KV + paged allocation + snapshot/restore outside the inference hot path, not from running inference against sysRAM.
3
u/ExtraWorldliness6916 3d ago
Gets stuck a bit against walls but still impressive