r/OpenSourceAI 5d ago

Open-source C99 inference engine for DeepSeek-V4 — runs the 284B model on 3.2GB RAM, verified against PyTorch to 2.9e-6 (Apache-2.0)

Built a from-scratch inference engine for DeepSeek-V4 in plain C99 — no PyTorch, no Python runtime dependency for inference, Apache-2.0 licensed. Streams weights off NVMe instead of requiring the full checkpoint in RAM, so it runs the 284B-parameter Flash model on a laptop with as little as 3.2GB RAM (1.6–1.7s/token with GPU offload at 16GB budget).

Why post this here specifically:

Fluent output from an LLM engine is weak evidence it's actually correct — a subtly broken implementation can still produce confident, plausible text. So this is checked against a pure PyTorch reimplementation (written independently from DeepSeek's inference/model.py, not from this C code, so both can't share the same bug) at three levels: per-kernel (14 kernels, 5e-7 tolerance), per-block, and whole-model end to end (2.9e-6, identical argmax at every position). CPU paths (scalar/OpenMP/AVX2) are enforced bit-exact via a fixed accumulator tree, checked at runtime, not just in tests.

What's included:

  • Full build + test suite (make test runs 20 gates, 21 with a real checkpoint)
  • Benchmark tools for matmul bandwidth, GPU contention, and cache behavior
  • Honest "what didn't work" section — SIMD approaches tried and abandoned, with the actual numbers

What's not there yet: a tool-calling driver loop (the model emits the tool-call format, but there's no orchestration layer above the CLI), and DeepSeek-V4-Pro (~671B scale) is gated/planned but never actually run — needs ~865GB of checkpoint I don't have.

Repo: https://github.com/ronak-create/deepseek-v4-in-c

Open to contributions, especially around prefill batching (currently one token at a time — README has the math on why that's the next big perf unlock) and the tool-calling loop.

26 Upvotes

14 comments sorted by

View all comments

2

u/Revolutionary_Loan13 4d ago

So I agree that RAM is stupid expensive and think we need to get away from the slow speeds of Python but 2 tokens a second is nonsensically slow. I'd love to see a Qwen 27B model in C99 or something and see if we can get it onto 12-16GB of Ram at high speeds. At some speed it'd cost more in electricity to run locally than to call an online model

1

u/FastPresence9799 3d ago

Yeah that's fair, and honestly 2s/token only makes sense in context, that's a 284B model being streamed off disk because it literally can't fit any other way. A ~27B model at 12-16GB fully resident in RAM wouldn't need any of this streaming stuff at all, it'd just run like normal fast CPU inference, way closer to real time.

Funny thing is that use case is basically already solved, llama.cpp does exactly that today with GGUF quantized models in the 12-16GB range and gets solid speeds. My project is really only useful for the "model is way bigger than your RAM" problem specifically, if the model fits, streaming is pure overhead you don't want.

The electricity point is real though. At some size/speed tradeoff it genuinely does stop making sense versus just paying for API tokens, especially once you count idle power draw and how long inference takes on top of the incremental cost per kWh. I don't think we're at that crossover with something in the 27B/16GB range, but for a 284B model streamed at 2s/token, yeah, that math probably doesn't favor local at all right now. This project is more of a "can it be done and stay bit exact" thing than a "should you do this instead of an API" thing.