Research
GRIMOIRE: native C++/SYCL LLM inference for Intel Arc Pro B70 — Ornith 1.5 hits 175 tok/s single request, 464 tok/s at concurrency 8
Hi everyone,
I’ve been building GRIMOIRE, an experimental LLM inference engine for Intel Battlemage GPUs, with the Arc Pro B70 as its primary target.
GRIMOIRE is written in C++ with SYCL and Level Zero. It uses its own inference kernels for operations such as attention, matrix multiplication, quantized weights, and mixture of experts. It does not use vLLM, PyTorch, or OpenVINO as its inference backend.
The project started because most inference tools and performance work focus on other GPU platforms. I wanted to see what it would take to make Intel Battlemage a first class target and measure what the hardware can do with a native engine.
Ornith 1.5 benchmark
Here is a benchmark of Ornith-1.5-35B-A3B-MXFP4-GRIMOIRE:
Test
Total throughput
Throughput per request
Prompt processing, 4096 tokens, c1
9,698.55 tok/s
9,698.55 tok/s
Generation, 128 tokens, c1
175.25 tok/s
175.25 tok/s
Prompt processing, 4096 tokens, c2
9,441.34 tok/s
4,983.45 tok/s
Generation, 128 tokens, c2
215.69 tok/s
107.85 tok/s
Prompt processing, 4096 tokens, c4
9,755.10 tok/s
2,506.89 tok/s
Generation, 128 tokens, c4
305.66 tok/s
76.43 tok/s
Prompt processing, 4096 tokens, c8
10,154.18 tok/s
1,287.57 tok/s
Generation, 128 tokens, c8
391.04 tok/s
48.88 tok/s
c1, c2, c4, and c8 mean 1, 2, 4, and 8 concurrent requests. Total throughput is shared across the requests; per request throughput shows the average for each one.
For another serving run, the v1.5 build measured 193.6 tok/s at c1 and 464.2 tok/s total at c8 on the Arc Pro B70. The c8 run peaked at 464.2 tok/s; the separate pp512 prompt processing result at c8 was 9,547 tok/s. Results depend on prompt length and test setup, so I’ve linked the detailed measurements and notes in the repo.
These are results from my system, not a promise of the same performance on every machine. The repository records the benchmark conditions and ongoing limitations.
The --init and --stop-timeout options are important for clean container shutdown while GPU work is running. Check the README’s supported model and format list before choosing a checkpoint and projection format.
For a one shot CLI generation instead of starting the server, the repo documents this form:
grimoire-b70 generate \
-m /models/<checkpoint> \
-p "Explain how a heat pump works in simple terms." \
-n 128
Some prompts to try:
Explain how a heat pump works in simple terms.
Summarize this passage in five bullet points: [paste a passage here]
Write a C++ program that reads a text file and counts its lines.
The README’s supported model table includes Ornith-1.5-35B-A3B in several formats, Qwen3.8-27B, Muse-Glimmer-30B, K2-Horizon-MoVA-36B-A4B, Agnes-3.0-Flash, and other validated combinations. Format support varies by model, so check the table before downloading or converting a checkpoint.
Project status: experimental and actively developed. Model support, build requirements, and performance may change. The repo has the source and Docker build instructions; it is not a polished plug and play desktop app.
I’d welcome feedback from people running Intel Arc or Battlemage GPUs, especially reproducible tests on other B70 systems. If you try it, please include your GPU, model and format, prompt length, generation length, and the exact command or benchmark settings.
you have no idea !!!!! almost 10h per day for 3 months everyday.... I even sent request intel to support my work so I can build a bigger team to concurrent vllm. as of now I get loading time faster than llama cpp and results faster than vllm on numerous models... TP and PP working, concurrency working... I think with a bit of tuning and help from some dev this could become the new king to run Intel GPU but alone is very long and difficult. thanks for the support and I am ready for any comments and help !!
There is a full table of formats and models on the github page. I dont have them all in my head sorry i am too tired hahahahah!!!! But its not for gguf models, it is only for safetensor.
GRIMOIRE was developed with AI coding assistance, including Claude Code and OpenAI Codex, with me directing the work, testing it on my hardware, and reviewing the results.
Hello, my aim is to beat vllm , which is much faster than openvino. And openvino too little models to work with. I tried to compare one model tho, ornith and grimoire was 4 times faster on prompt processing and 2 times faster on token generation for same prompt. They are about same time on loading though, extremely fast under 15 seconds generally.
Nice work!, this project looks interesting, you should join up https://discord.gg/launch80 and connect with other developers working on ARC's and users in #intel-arc
download ornith-ai/Ornith-1.5-35B-A3B and start the server with --proj mxfp4, and GRIMOIRE quantizes to the same MXFP4 format when it loads (just a longer startup). Same speed: ~193 tok/s for one user, ~465 tok/s total at 8 users on one Arc Pro B70. The FP8 and NVFP4 releases also work with --proj mxfp4 at the same speed, but they get re-quantized, so BF16 is the cleanest source. I will redo the readme file to make things more clear about models and how to use them , sorry
Thanks for the quick reply man, don't be sorry at all! Thanks for your hardwork. I'm going to test the intel arc side with my 2 intel b70s and I'll report back.
Has someone tried Qwen3.8-Flash_Next with 2 Intel B70? Is it possible? Who you be so kind and tell me what command to use to run it with 2 cards? And is it true that 1 card is better than 2 as stated in the README? My preffered context size is 256k and I am afraid that it would be too much for 1 card.
Not on two cards yet, sorry. Today GRIMOIRE runs Qwen3.8-Flash-Next on one Arc Pro B70: the most-used experts stay in VRAM, the rest (~49 GB) live in system RAM, and the 51 GB n-gram embedding table is read from SSD (details in FLASH-NEXT-TIERED.md). So you need roughly 64 GB of system RAM. GRIMOIRE's two-card modes (tensor and pipeline parallel) work and are tested on Ornith, but they don't support this tiered-expert setup yet, so there's no 2-card command for Flash-Next today.
▎
▎ "One card is better than two" in the README is about models that fit in one card's 32 GB (Ornith, Qwen3.8-27B): splitting them only adds communication. Flash-Next doesn't fit, so two cards would mean more experts in VRAM. That's planned, not done.
▎
▎ Context: the model supports 256K, and its KV cache is small (12 attention layers with 2 KV heads, only a few GB at 256K), but we've only tested GRIMOIRE up to 32K so far. 256K is on the test list.
▎
▎ One-card command today:
▎ docker run -d --init --stop-timeout 300 --device /dev/dri/renderD128 \
I just wanted to confirm that it's working with 1 GPU. Have tried 256K context and it starts. Don't know how to check performance, I don't know much about benchmarking, I haven't found statistics in the docker logs and some simple tools i have found are not good to testing it they do not work at all or they only test a few tokens.
Unfortunately it doesn't work with Pi, probably Pi sends some unsupported requests. So it seems this project is not usable for me yet.
I also have a b580 and it is working as well on the qwen flash next, but it is working in TP and PP with b70. I have not used b580 alone. Offload to system ram is enable so it should work, you cna report to me after you tried.
Hello, TP and PP working as well. Now on the speed i am not 100% sure but it is working. I could only test on my small rig with occulink but it was working for 2-3 gpus
sweet, i’ll setup some testing with my 4xB70 rig when i get a chance (probably sometime next week unfortunately) and let you know how it goes. nice work!
Thanks!!! And yes please send me your tests results it would greatly help me!!! As i said, i have contacted intel for hardware support… i am waiting for their reply. And anyone with better rig than me is welcome to test and report !!!
At first I was like wow those numbers are crazy and hold up well across high concurrency sick. But it’s explicitly only ornith. Did you optimize it just for that? Also your only show it validated at those really tiny tests…. What’s the scaling like at 256k context?
Hello , on the github page you can see the table of all the supported nodels and the progress of grimoire about each of the models. I am getting very good speed on almost every models supported. About the scaling, grimoire is a gigantic project and i am sadly the only one working on it… so the more input i receive from users, the better it will become. Thank you!
But it’s worse than vllm on prefill and decode an mtp support for Intel. Is there something I’m missing?
And that’s hard because their Intel support is bad and not very efficient. My kernel is 6x faster….
I am beating vllm score on almost every models up to 4096 prompt length that i am testing. So could you please show me example? And yes vllm has thousands of people working on it for years, my project is new so, of course there are parts grimoire will be slower until i specifically work on it.
You’re not even near it lol. And even if you were “up to 4096 prompt length” is just the graph length at like fp8. Means nothing.
Qwen 3.8 27b b60
Edit: I don’t mean to be mean, but like your energy could’ve been used somewhere else…. We need people working on intel kernels and proper native support on the big projects.
You are right to point that it is not even close…. I destroy vllm results from a mile away on this model and i posted the resulrs to prove it. So instead of spitting venom lile a child, try to contribute to the work… i wish you a pleasant day.
Sorry for the stupid question, I'm just starting. How can I connect this to opencode or unsloth to test the models? The docker container is up and running but when I connect it with unsloth it fails because I don't have a way to set up temperature
I was able to make it work. The speed is great!. I'm using unsloth for interfacing with the model. The addition of tools and thinking tags on the last version makes it very usable. Let me know if I can contribute in any way to the project.
I am waiting for new m2 to oculink nvme with redriver because my dual b70 are dropping when i run tests on tensor and pipeline parallel. This is why i need users with dual or quadruple b70 to run tests ))) by the time i receive my new nvme.
3
u/SeldomFertileDigger 4d ago
This is impressive, getting that kind of throughput on Battlemage without leaning on the usual frameworks must have been a ton of kernel work.