Hardware: Raspberry Pi 5 Model B, 16GB RAM, 469GB NVMe SSD, ARM Cortex-A76 4 cores @ 1.5โ2.4GHz
Software: Ollama 0.32.1, CPU-only inference (no GPU), Debian 13 (trixie)
Test duration: ~7 hours, sequential testing (one model at a time: download โ test โ cleanup โ next)
I wanted to know: which Ollama models are actually usable on a Raspberry Pi 5? Not "technically runnable" โ actually usable, as in "you'd wait for the response without giving up."
Each model got 5 standardized prompts:
Timeout: 300 seconds per prompt. If a model didn't finish in 5 minutes, that test was marked as timeout.
โ = completed in time โฐ = timeout (300s) ๐ง = thinking mode
| Model | Params | Quant | Avg time | Think | EN | CS | Logic | Math | Code | RAM |
|---|---|---|---|---|---|---|---|---|---|---|
| gemma3:1b | 1.0B | Q4_K_M | 8s | โ | โ | โ | โ | โ | โ | 1.2GB |
| llama3.2:1b | 1.2B | Q8_0 | 9s | โ | โ | โ | โ | โ | โ | 1.7GB |
| tinyllama:1.1b | 1B | Q4_0 | 11s | โ | โ | โ | โ | โ | โ | 0.7GB |
| qwen2.5:0.5b | 0.5B | Q4_K_M | 12s | โ | โ | โ | โ | โ | โ | 0.6GB |
| llama3.2:3b | 3.2B | Q4_K_M | 14s | โ | โ | โ | โ | โ | โ | 2.8GB |
| qwen2.5:1.5b | 1.5B | Q4_K_M | 15s | โ | โ | โ | โ | โ | โ | 1.1GB |
| minicpm-v4.6:1b | 0.8B | Q4_K_M | 23s | โ | โ | โ | โ | โ | โ | 1.9GB |
| granite4.1:3b | 3.4B | Q4_K_M | 24s | โ | โ | โ | โ | โ | โ | 2.5GB |
| ministral-3:8b | 8.9B | Q4_K_M | 24s | ๐ง | โ | โ | โ | โ | โ | 6.1GB |
| gemma3:4b | 4.3B | Q4_K_M | 25s | โ | โ | โ | โ | โ | โ | 4.1GB |
| qwen2.5:3b | 3.1B | Q4_K_M | 26s | โ | โ | โ | โ | โ | โ | 2.3GB |
| codegemma:7b | 9B | Q4_0 | 30s | โ | โ | โ | โ | โ | โ | 7.0GB |
| translategemma:4b | 4.3B | Q4_K_M | 33s | โ | โ | โ | โ | โ | โ | 3.6GB |
| ministral-3:3b | 3.8B | Q4_K_M | 33s | โ | โ | โ | โ | โ | โ | 3.3GB |
| qwen3:0.6b | 0.8B | Q4_K_M | 36s | ๐ง | โ | โ | โ | โ | โ | 1.3GB |
| llama3.1:8b | 8.0B | Q4_K_M | 36s | โ | โ | โ | โ | โ | โ | 4.8GB |
| gemma3:12b | 12.2B | Q4_K_M | 41s | โ | โ | โ | โ | โ | โ | 9.3GB |
| minicpm-v4.5:8b | 8.2B | Q4_K_M | 43s | โ | โ | โ | โ | โ | โ | 6.0GB |
| lfm2.5:8b | 8.5B | Q4_K_M | 44s | ๐ง | โ | โ | โ | โ | โ | 4.9GB |
| qwen2.5:7b | 7.6B | Q4_K_M | 46s | โ | โ | โ | โ | โ | โ | 4.4GB |
| mistral:7b | 7.2B | Q4_K_M | 49s | โ | โ | โ | โ | โ | โ | 4.3GB |
| codegemma:2b | 3B | Q4_0 | 51s | โ | โ | โ | โ | โ | โ | 1.7GB |
| deepseek-r1:1.5b | 1.8B | Q4_K_M | 54s | ๐ง | โ | โ | โ | โ | โ | 1.3GB |
| gemma2:9b | 9.2B | Q4_0 | 78s | โ | โ | โ | โ | โ | โ | 8.0GB |
| nemotron-3-nano:4b | 4.0B | Q4_K_M | 79s | ๐ง | โ | โ | โ | โ | โ | 3.3GB |
| granite4.1:8b | 8.8B | Q4_K_M | 79s | โ | โ | โ | โ | โ | โ | 5.5GB |
| phi3:mini | 3.8B | Q4_0 | 81s | โ | โ | โฐ | โ | โ | โ | 4.2GB |
| qwen3:1.7b | 2.0B | Q4_K_M | 82s | ๐ง | โ | โ | โ | โ | โ | 1.6GB |
| starcoder2:3b | 3B | Q4_0 | 85s | โ | โ | โฐ | โ | โ | โ | 1.5GB |
| starcoder2:7b | 7B | Q4_0 | 125s | โ | โ | โ | โ | โ | โ | 3.7GB |
| ornith:9b | 9.0B | Q4_K_M | 127s | ๐ง | โ | โ | โ | โ | โ | 6.0GB |
| qwen3-vl:2b | 2.1B | Q4_K_M | 175s | ๐ง | โ | โฐ | โ | โ | โฐ | 2.8GB |
| deepseek-r1:7b | 7.6B | Q4_K_M | 180s | ๐ง | โ | โ | โฐ | โ | โฐ | 4.4GB |
| qwen3-vl:4b | 4.4B | Q4_K_M | 209s | ๐ง | โ | โ | โ | โฐ | โฐ | 4.3GB |
| qwen3:8b | 8.2B | Q4_K_M | 223s | ๐ง | โ | โ | โ | โฐ | โฐ | 5.3GB |
| qwen3.5:0.8b | 0.9B | Q8_0 | 250s | ๐ง | โฐ | โฐ | โฐ | โฐ | โ | 1.7GB |
| deepseek-r1:8b | 8.2B | Q4_K_M | 253s | ๐ง | โ | โ | โฐ | โฐ | โ | 5.7GB |
| qwen3.5:2b | 2.3B | Q8_0 | 257s | ๐ง | โฐ | โฐ | โฐ | โฐ | โ | 3.3GB |
| qwen3:4b | 4.0B | Q4_K_M | 268s | ๐ง | โ | โฐ | โฐ | โฐ | โฐ | 3.4GB |
| qwen3.5:4b | 4.7B | Q4_K_M | 281s | ๐ง | โ | โฐ | โฐ | โฐ | โ | 4.3GB |
| phi4:14b | 14.7B | Q4_K_M | 300s | โ | โฐ | โฐ | โฐ | โฐ | โฐ | 9.7GB |
| qwen3.5:9b | 9.7B | Q4_K_M | 300s | ๐ง | โฐ | โฐ | โฐ | โฐ | โฐ | 7.0GB |
Models with built-in "thinking mode" (qwen3, qwen3.5, deepseek-r1 series) generate long internal reasoning chains before answering. On a CPU-only Raspberry Pi 5, this is devastating:
The exceptions: ministral-3:8b (24s) and lfm2.5:8b (45s) both have thinking mode but are remarkably fast. These are the only thinking-mode models I'd consider usable on RPi 5.
Quantization matters enormously on CPU:
If a model only offers Q8_0 (like qwen3.5:0.8b and qwen3.5:2b), think twice before running on Pi.
This was the biggest surprise. On RPi 5 with 16GB RAM:
37 seconds for an 8B model on a $80 board is genuinely impressive. They use 4โ6GB RAM, leaving plenty for the OS and context window.
If I had to pick one model for daily use on RPi 5, it would be llama3.2:3b:
At 14 seconds, it's only 5 seconds slower than 1B models, but with 3x the parameters. That's the sweet spot.
Specialized coding models (codegemma, starcoder2) are noticeably slower than general models of similar size:
For coding on RPi 5, codegemma:7b is the best option at 30s, but it's heavy (7GB RAM).
gemma3:12b completed all tests in 42s average using 9.3GB RAM. That's 12.2 billion parameters running on a Raspberry Pi. Not fast, but it works and leaves enough RAM for the OS (~6GB free with 16GB total).
Both vision models tested (qwen3-vl:2b and qwen3-vl:4b) combine vision capability with thinking mode, making them extremely slow (176s and 210s). Several tests timed out. If you need vision on RPi 5, you'd need to find a non-thinking vision model.
The Raspberry Pi 5 with 16GB RAM is a surprisingly capable LLM host โ if you pick the right models.
The key insight from testing 42 models: model architecture matters more than model size. A well-optimized 8B model without thinking mode (llama3.1:8b at 37s) can be faster than a badly-optimized 0.8B model with thinking mode (qwen3.5:0.8b at 250s).
The "thinking mode" trend in newer models (qwen3, deepseek-r1) is great for quality on powerful hardware but makes them 3โ20x slower on CPU. If you're running on Pi, stick to models that answer directly without internal monologues.