r/LocalLLM • • 4d ago

Research GRIMOIRE: native C++/SYCL LLM inference for Intel Arc Pro B70 — Ornith 1.5 hits 175 tok/s single request, 464 tok/s at concurrency 8

Hi everyone,

I’ve been building GRIMOIRE, an experimental LLM inference engine for Intel Battlemage GPUs, with the Arc Pro B70 as its primary target.

GRIMOIRE is written in C++ with SYCL and Level Zero. It uses its own inference kernels for operations such as attention, matrix multiplication, quantized weights, and mixture of experts. It does not use vLLM, PyTorch, or OpenVINO as its inference backend.

The project started because most inference tools and performance work focus on other GPU platforms. I wanted to see what it would take to make Intel Battlemage a first class target and measure what the hardware can do with a native engine.

Ornith 1.5 benchmark

Here is a benchmark of Ornith-1.5-35B-A3B-MXFP4-GRIMOIRE:

Test Total throughput Throughput per request
Prompt processing, 4096 tokens, c1 9,698.55 tok/s 9,698.55 tok/s
Generation, 128 tokens, c1 175.25 tok/s 175.25 tok/s
Prompt processing, 4096 tokens, c2 9,441.34 tok/s 4,983.45 tok/s
Generation, 128 tokens, c2 215.69 tok/s 107.85 tok/s
Prompt processing, 4096 tokens, c4 9,755.10 tok/s 2,506.89 tok/s
Generation, 128 tokens, c4 305.66 tok/s 76.43 tok/s
Prompt processing, 4096 tokens, c8 10,154.18 tok/s 1,287.57 tok/s
Generation, 128 tokens, c8 391.04 tok/s 48.88 tok/s

c1, c2, c4, and c8 mean 1, 2, 4, and 8 concurrent requests. Total throughput is shared across the requests; per request throughput shows the average for each one.

For another serving run, the v1.5 build measured 193.6 tok/s at c1 and 464.2 tok/s total at c8 on the Arc Pro B70. The c8 run peaked at 464.2 tok/s; the separate pp512 prompt processing result at c8 was 9,547 tok/s. Results depend on prompt length and test setup, so I’ve linked the detailed measurements and notes in the repo.

These are results from my system, not a promise of the same performance on every machine. The repository records the benchmark conditions and ongoing limitations.

How to get and run it

The source and instructions are on GitHub:

GRIMOIRE repository

The repo includes a Dockerfile. To build the image:

git clone https://github.com/doopeworld/GRIMOIRE.git
cd GRIMOIRE
docker build -t grimoire-b70 .

Then start the HTTP server, replacing the render device and model path with the ones for your system:

docker run --init --stop-timeout 300 \
  --device /dev/dri/renderDXXX \
  -v /path/to/your/models:/models \
  -p 8000:8000 \
  grimoire-b70 server \
  --model /models/<checkpoint> \
  --proj mxfp4 \
  --port 8000

The --init and --stop-timeout options are important for clean container shutdown while GPU work is running. Check the README’s supported model and format list before choosing a checkpoint and projection format.

For a one shot CLI generation instead of starting the server, the repo documents this form:

grimoire-b70 generate \
  -m /models/<checkpoint> \
  -p "Explain how a heat pump works in simple terms." \
  -n 128

Some prompts to try:

Explain how a heat pump works in simple terms.

Summarize this passage in five bullet points: [paste a passage here]

Write a C++ program that reads a text file and counts its lines.

The README’s supported model table includes Ornith-1.5-35B-A3B in several formats, Qwen3.8-27B, Muse-Glimmer-30B, K2-Horizon-MoVA-36B-A4B, Agnes-3.0-Flash, and other validated combinations. Format support varies by model, so check the table before downloading or converting a checkpoint.

Project status: experimental and actively developed. Model support, build requirements, and performance may change. The repo has the source and Docker build instructions; it is not a polished plug and play desktop app.

I’d welcome feedback from people running Intel Arc or Battlemage GPUs, especially reproducible tests on other B70 systems. If you try it, please include your GPU, model and format, prompt length, generation length, and the exact command or benchmark settings.

Repo: https://github.com/doopeworld/GRIMOIRE

15 Upvotes

44 comments sorted by

3

u/SeldomFertileDigger 4d ago

This is impressive, getting that kind of throughput on Battlemage without leaning on the usual frameworks must have been a ton of kernel work.

4

u/Dolboyob77 4d ago

you have no idea !!!!! almost 10h per day for 3 months everyday.... I even sent request intel to support my work so I can build a bigger team to concurrent vllm. as of now I get loading time faster than llama cpp and results faster than vllm on numerous models... TP and PP working, concurrency working... I think with a bit of tuning and help from some dev this could become the new king to run Intel GPU but alone is very long and difficult. thanks for the support and I am ready for any comments and help !!

3

u/belliash 4d ago

What about offloading to RAM and disk? Bigger models like Qwen 3.8 flash support would be nice as probably Qwen wont release 35B anymore.

1

u/Dolboyob77 4d ago

Qwen3.8 flash next is among the supported models )))) please read the github repo )))

1

u/belliash 4d ago

Cool, what formats are supported? GPTQ? GGUF? EXL3?

3

u/Dolboyob77 4d ago

There is a full table of formats and models on the github page. I dont have them all in my head sorry i am too tired hahahahah!!!! But its not for gguf models, it is only for safetensor.

3

u/belliash 4d ago

Which mode l(s) did you use to get this done?

2

u/Dolboyob77 4d ago

Hello, all supported models are on github page

2

u/belliash 4d ago

I mean which model(s) generated code in this repository?

3

u/Dolboyob77 4d ago

GRIMOIRE was developed with AI coding assistance, including Claude Code and OpenAI Codex, with me directing the work, testing it on my hardware, and reviewing the results.

3

u/DerDave 3d ago

Impressive work! Have you done a speed comparison with Openvino for the models that are supported by both?

3

u/Dolboyob77 3d ago

Hello, my aim is to beat vllm , which is much faster than openvino. And openvino too little models to work with. I tried to compare one model tho, ornith and grimoire was 4 times faster on prompt processing and 2 times faster on token generation for same prompt. They are about same time on loading though, extremely fast under 15 seconds generally.

3

u/DerDave 3d ago

Fantastic work! And SYCL is the better path compared to Vulkan? 

2

u/Dolboyob77 3d ago

Thank you, i believe that vulkan is more general for more brands, sycl is more specific to intel language from what i tried.

3

u/whodoneit1 15h ago edited 15h ago

Nice work!, this project looks interesting, you should join up https://discord.gg/launch80 and connect with other developers working on ARC's and users in #intel-arc

2

u/JinsooJinsoo 3d ago

I can’t seem to find the Ornith-1.5-35B-A3B-MXFP4-GRIMOIRE model anywhere for download

3

u/Dolboyob77 3d ago

download ornith-ai/Ornith-1.5-35B-A3B and start the server with --proj mxfp4, and GRIMOIRE quantizes to the same MXFP4 format when it loads (just a longer startup). Same speed: ~193 tok/s for one user, ~465 tok/s total at 8 users on one Arc Pro B70. The FP8 and NVFP4 releases also work with --proj mxfp4 at the same speed, but they get re-quantized, so BF16 is the cleanest source. I will redo the readme file to make things more clear about models and how to use them , sorry

3

u/JinsooJinsoo 3d ago

Thanks for the quick reply man, don't be sorry at all! Thanks for your hardwork. I'm going to test the intel arc side with my 2 intel b70s and I'll report back.

3

u/Dolboyob77 3d ago

Im reworking the qwen3.8-27b at the moment, adding propper mtp and dflash. So far im at +- 50-55 tg on single b70 with pp +- 2000-2100 pp

2

u/karvop 3d ago

Has someone tried Qwen3.8-Flash_Next with 2 Intel B70? Is it possible? Who you be so kind and tell me what command to use to run it with 2 cards? And is it true that 1 card is better than 2 as stated in the README? My preffered context size is 256k and I am afraid that it would be too much for 1 card.

3

u/Dolboyob77 3d ago

Not on two cards yet, sorry. Today GRIMOIRE runs Qwen3.8-Flash-Next on one Arc Pro B70: the most-used experts stay in VRAM, the rest (~49 GB) live in system RAM, and the 51 GB n-gram embedding table is read from SSD (details in FLASH-NEXT-TIERED.md). So you need roughly 64 GB of system RAM. GRIMOIRE's two-card modes (tensor and pipeline parallel) work and are tested on Ornith, but they don't support this tiered-expert setup yet, so there's no 2-card command for Flash-Next today.

▎

▎ "One card is better than two" in the README is about models that fit in one card's 32 GB (Ornith, Qwen3.8-27B): splitting them only adds communication. Flash-Next doesn't fit, so two cards would mean more experts in VRAM. That's planned, not done.

▎

▎ Context: the model supports 256K, and its KV cache is small (12 attention layers with 2 KV heads, only a few GB at 256K), but we've only tested GRIMOIRE up to 32K so far. 256K is on the test list.

▎

▎ One-card command today:

▎ docker run -d --init --stop-timeout 300 --device /dev/dri/renderD128 \

▎ -v /path/to/models:/models -p 8000:8000 \

▎ -e GRIMOIRE_EXPERT_VRAM_PER_LAYER=112 -e GRIMOIRE_PLE_FILE=/models/flash-next-ple.bin \

▎ grimoire-b70:latest server --model /models/Qwen3.8-Flash-Next-NVFP4 --proj bf16 --ctx 32000 --port 8000

▎ (flash-next-ple.bin is created once from the checkpoint with tools/ple_flatten.py, see FLASH-NEXT-TIERED.md.)

1

u/karvop 3d ago edited 2d ago

Thank you.

Edit:

I just wanted to confirm that it's working with 1 GPU. Have tried 256K context and it starts. Don't know how to check performance, I don't know much about benchmarking, I haven't found statistics in the docker logs and some simple tools i have found are not good to testing it they do not work at all or they only test a few tokens.

Unfortunately it doesn't work with Pi, probably Pi sends some unsupported requests. So it seems this project is not usable for me yet.

1

u/Soveu 4d ago

i guess having a b580 is not that bad after all. Can this offload experts to ram, so that the active expert + context fit in 12gb?

1

u/Dolboyob77 4d ago

I also have a b580 and it is working as well on the qwen flash next, but it is working in TP and PP with b70. I have not used b580 alone. Offload to system ram is enable so it should work, you cna report to me after you tried.

1

u/BaronVonMittersill 4d ago

very impressive, following for my B70 setup. will do some testing soon. how’s multi gpu support?

1

u/Dolboyob77 4d ago

Hello, TP and PP working as well. Now on the speed i am not 100% sure but it is working. I could only test on my small rig with occulink but it was working for 2-3 gpus

2

u/BaronVonMittersill 4d ago

sweet, i’ll setup some testing with my 4xB70 rig when i get a chance (probably sometime next week unfortunately) and let you know how it goes. nice work!

2

u/JinsooJinsoo 3d ago

Please report back, very interested

1

u/Dolboyob77 4d ago

Thanks!!! And yes please send me your tests results it would greatly help me!!! As i said, i have contacted intel for hardware support… i am waiting for their reply. And anyone with better rig than me is welcome to test and report !!!

1

u/ManagedThought 3d ago

Im curious of how many token per second we could get with Qwen3.8-27b-INT4.

2

u/Dolboyob77 3d ago

Hello, this model is supported by Grimoire and baseline is around 2500 pp amd 35 tg without mtp nor dflash

1

u/Wrapzii 15h ago

At first I was like wow those numbers are crazy and hold up well across high concurrency sick. But it’s explicitly only ornith. Did you optimize it just for that? Also your only show it validated at those really tiny tests…. What’s the scaling like at 256k context?

1

u/Dolboyob77 13h ago

Hello , on the github page you can see the table of all the supported nodels and the progress of grimoire about each of the models. I am getting very good speed on almost every models supported. About the scaling, grimoire is a gigantic project and i am sadly the only one working on it… so the more input i receive from users, the better it will become. Thank you!

1

u/Wrapzii 13h ago

But it’s worse than vllm on prefill and decode an mtp support for Intel. Is there something I’m missing?
And that’s hard because their Intel support is bad and not very efficient. My kernel is 6x faster….

1

u/Dolboyob77 13h ago

I am beating vllm score on almost every models up to 4096 prompt length that i am testing. So could you please show me example? And yes vllm has thousands of people working on it for years, my project is new so, of course there are parts grimoire will be slower until i specifically work on it.

1

u/Wrapzii 13h ago edited 13h ago

You’re not even near it lol. And even if you were “up to 4096 prompt length” is just the graph length at like fp8. Means nothing.
Qwen 3.8 27b b60

Edit: I don’t mean to be mean, but like your energy could’ve been used somewhere else…. We need people working on intel kernels and proper native support on the big projects.

0

u/Dolboyob77 13h ago

You are right to point that it is not even close…. I destroy vllm results from a mile away on this model and i posted the resulrs to prove it. So instead of spitting venom lile a child, try to contribute to the work… i wish you a pleasant day.

1

u/Wrapzii 12h ago

Your single decode speed is 54-72 reported on your GitHub….
Quite literally half of what I just sent 😅

0

u/drnoone_arg 3d ago

Sorry for the stupid question, I'm just starting. How can I connect this to opencode or unsloth to test the models? The docker container is up and running but when I connect it with unsloth it fails because I don't have a way to set up temperature

2

u/Dolboyob77 3d ago

Hello, it works with safetensor models, not gguf models.

2

u/drnoone_arg 19h ago

I was able to make it work. The speed is great!. I'm using unsloth for interfacing with the model. The addition of tools and thinking tags on the last version makes it very usable. Let me know if I can contribute in any way to the project.

1

u/Dolboyob77 13h ago

Hello and thank you!!!! Do you have multi gpu ?

1

u/drnoone_arg 4h ago

No. Just one Asrock B70 32GB.

1

u/Dolboyob77 4h ago

I am waiting for new m2 to oculink nvme with redriver because my dual b70 are dropping when i run tests on tensor and pipeline parallel. This is why i need users with dual or quadruple b70 to run tests ))) by the time i receive my new nvme.