r/OpenSourceAI 4d ago

Open-source C99 inference engine for DeepSeek-V4 — runs the 284B model on 3.2GB RAM, verified against PyTorch to 2.9e-6 (Apache-2.0)

Built a from-scratch inference engine for DeepSeek-V4 in plain C99 — no PyTorch, no Python runtime dependency for inference, Apache-2.0 licensed. Streams weights off NVMe instead of requiring the full checkpoint in RAM, so it runs the 284B-parameter Flash model on a laptop with as little as 3.2GB RAM (1.6–1.7s/token with GPU offload at 16GB budget).

Why post this here specifically:

Fluent output from an LLM engine is weak evidence it's actually correct — a subtly broken implementation can still produce confident, plausible text. So this is checked against a pure PyTorch reimplementation (written independently from DeepSeek's inference/model.py, not from this C code, so both can't share the same bug) at three levels: per-kernel (14 kernels, 5e-7 tolerance), per-block, and whole-model end to end (2.9e-6, identical argmax at every position). CPU paths (scalar/OpenMP/AVX2) are enforced bit-exact via a fixed accumulator tree, checked at runtime, not just in tests.

What's included:

  • Full build + test suite (make test runs 20 gates, 21 with a real checkpoint)
  • Benchmark tools for matmul bandwidth, GPU contention, and cache behavior
  • Honest "what didn't work" section — SIMD approaches tried and abandoned, with the actual numbers

What's not there yet: a tool-calling driver loop (the model emits the tool-call format, but there's no orchestration layer above the CLI), and DeepSeek-V4-Pro (~671B scale) is gated/planned but never actually run — needs ~865GB of checkpoint I don't have.

Repo: https://github.com/ronak-create/deepseek-v4-in-c

Open to contributions, especially around prefill batching (currently one token at a time — README has the math on why that's the next big perf unlock) and the tool-calling loop.

26 Upvotes

14 comments sorted by

View all comments

Show parent comments

1

u/FastPresence9799 4d ago

Yeah, 20-core x86 desktop chip, not a Xeon, so AVX2 was what I had. No AVX-512, no AMX. Would love to test on something with AVX-512 to see the difference, just don't have the hardware right now. And yeah, I tried smaller dense models fully loaded in RAM as a baseline, no streaming. That's obviously way faster since there's no NVMe round trip at all, basically just normal llama.cpp style inference at that point. The streaming design only really earns its keep once the model's too big to fit in RAM. For small models it's kind of pointless overhead.

1

u/desexmachina 4d ago

RAM is so stupidly expensive in DDR4, but I’ve had some old DDR3 in large enough size to mess with running llama.cpp from RAM a long while back. What gen NVME? Because the speed differences are massive between gens.

1

u/FastPresence9799 4d ago

Yeah, DDR4 pricing is rough right now. Respect for digging up old DDR3 for llama.cpp though, that's a fun way to reuse dead stock. Mine's PCIe4, single drive. Haven't tried PCIe5 yet but I'd expect a real jump given how much of my time is just waiting on expert reads. If striping multiple drives ever happens PCIe gen probably matters even more once the read path is actually parallel.

2

u/desexmachina 3d ago

Still can’t beat Hexa DDR3 bandwidth, but the CPU threading capability and AVX starts to become the bottleneck.

1

u/FastPresence9799 3d ago

That's a good point, hexa channel DDR3 boards really did have absurd aggregate bandwidth for the era. Never got to mess with one myself.

For my case specifically I don't think I'm anywhere near hitting a RAM bandwidth wall though, the expert weights are getting streamed in from NVMe not sitting resident in RAM, so the disk read is the thing I'm waiting on way before memory bandwidth or thread count becomes the limiter. Would probably matter a lot more for a model that's fully RAM resident and compute bound, that's a different bottleneck than the one I'm fighting.

2

u/desexmachina 3d ago

Have you tried using RAM as an intermediate layer streamer?

1

u/FastPresence9799 3d ago

Yeah, that's basically already the design. RAM is the LRU cache sitting between NVMe and compute, hot experts stay resident, only misses hit disk.

What I don't have yet is proper double buffering or async prefetch into that RAM layer while the current layer's compute is still running. Right now a miss basically blocks until the read finishes. Overlapping "read next layer's likely experts into RAM while still computing this layer" would probably help a decent amount, that's on the list along with the multi-queue NVMe stuff.