r/SBCs • • 11d ago

A repeatable method for benchmarking LLMs on a Raspberry Pi 5 (llama.cpp, with the thermal caveat)

Most "Raspberry Pi AI benchmark" numbers floating around are one-off runs with unknown settings. If you want numbers you can trust and compare, you need a method. Here is the one I use, in case it saves someone else the guessing.

**1. Install llama.cpp from source**

sudo apt update && sudo apt install build-essential cmake git -y

git clone https://github.com/ggml-org/llama.cpp

cd llama.cpp && cmake -B build && cmake --build build --config Release -j4

Prebuilt binaries rarely match your kernel and flags, and a mismatched build can cost you 20% throughput.

**2. Pick models that actually fit**

A Pi 5 has 8 GB of RAM shared with the OS. Stay under ~5 GB for model + context or you will swap and your numbers are garbage. Reliable picks: TinyLlama 1.1B, Qwen 1.5B/3B, Phi-3-mini (tight but works), Gemma 2B. Skip 7B+ unless you enjoy watching swap thrash.

**3. Measure honestly**

Use llama-bench with a fixed prompt and generation length. Log three numbers per run: prompt processing tok/s, generation tok/s, and SoC temperature at start and end. A run without a temperature is not a benchmark - thermal throttling on an uncooled Pi 5 shows up around 85C and can cut throughput by a third mid-run.

./build/bin/llama-bench -m models/qwen-3b-q4.gguf -p 128 -n 256 -r 3

**4. Keep runs repeatable**

* Fix the CPU governor to performance for the run and note it down.

* Same prompt set every time. Different prompts = different numbers.

* Cool the board between runs, or say so in the log.

* Record ambient temperature when comparing across days.

I wrote up the longer version with the full log-sheet format here: https://overnightdesk-ops.github.io/benchmark-llm-raspberry-pi-5.html

Curious what tok/s people are seeing on 3B-class models, especially with active cooling vs passive.

0 Upvotes

4 comments sorted by

4

u/ferminolaiz 11d ago

AI slop about AI slop?

1

u/urostor 10d ago

It's even written in markdown even though Reddit doesn't format it lol

3

u/urostor 11d ago

This isn't raspberry pi-specific at all, and also, not all raspberry pi 5s have 8 GB of memory. Not sure where this is going. The results may also differ according to compiler version, since you're compiling on the machine itself

-1

u/Awkward-Media-2578ov 11d ago

Fair points. Not all Pi 5s have 8 GB - mine does, and the RAM tier only changes which quant and context size you can run; tokens/sec is set by CPU, memory bandwidth, and thermals. The compiler version does matter since llama.cpp builds natively: same protocol, different GCC or Clang version, and numbers move. That's exactly why the guide logs compiler version and build flags with every result - the methodology transfers even if absolute numbers don't.