๐Ÿฆž๐Ÿ“

I tested 42 Ollama models on a Raspberry Pi 5 (16GB RAM)

Here's what actually works โ€” and what doesn't
๐Ÿ“… July 21, 2026
โฑ๏ธ ~7 hours of testing
๐Ÿงช 42 models tested
๐Ÿ–ฅ๏ธ CPU only (no GPU)
42
Models tested
8s
Fastest response
300s
Slowest response
12B
Largest model that works

๐Ÿ“ The Setup

Hardware: Raspberry Pi 5 Model B, 16GB RAM, 469GB NVMe SSD, ARM Cortex-A76 4 cores @ 1.5โ€“2.4GHz

Software: Ollama 0.32.1, CPU-only inference (no GPU), Debian 13 (trixie)

Test duration: ~7 hours, sequential testing (one model at a time: download โ†’ test โ†’ cleanup โ†’ next)

I wanted to know: which Ollama models are actually usable on a Raspberry Pi 5? Not "technically runnable" โ€” actually usable, as in "you'd wait for the response without giving up."

๐Ÿงช The Test

Each model got 5 standardized prompts:

English
List 3 qualities of a good software engineer
Czech
Same question, in Czech
Logic
If all A are B and all B are C, are all A C?
Math
Calculate (17ร—23)+(144รท12)โˆ’89
Programming
Write a Python palindrome function

Timeout: 300 seconds per prompt. If a model didn't finish in 5 minutes, that test was marked as timeout.

๐Ÿ“Š Full Results (sorted by speed)

โœ… = completed in time   โฐ = timeout (300s)   ๐Ÿง  = thinking mode

Model Params Quant Avg time Think EN CS Logic Math Code RAM
gemma3:1b1.0BQ4_K_M8sโ€”โœ…โœ…โœ…โœ…โœ…1.2GB
llama3.2:1b1.2BQ8_09sโ€”โœ…โœ…โœ…โœ…โœ…1.7GB
tinyllama:1.1b1BQ4_011sโ€”โœ…โœ…โœ…โœ…โœ…0.7GB
qwen2.5:0.5b0.5BQ4_K_M12sโ€”โœ…โœ…โœ…โœ…โœ…0.6GB
llama3.2:3b3.2BQ4_K_M14sโ€”โœ…โœ…โœ…โœ…โœ…2.8GB
qwen2.5:1.5b1.5BQ4_K_M15sโ€”โœ…โœ…โœ…โœ…โœ…1.1GB
minicpm-v4.6:1b0.8BQ4_K_M23sโ€”โœ…โœ…โœ…โœ…โœ…1.9GB
granite4.1:3b3.4BQ4_K_M24sโ€”โœ…โœ…โœ…โœ…โœ…2.5GB
ministral-3:8b8.9BQ4_K_M24s๐Ÿง โŒโŒโŒโŒโŒ6.1GB
gemma3:4b4.3BQ4_K_M25sโ€”โœ…โœ…โœ…โœ…โœ…4.1GB
qwen2.5:3b3.1BQ4_K_M26sโ€”โœ…โœ…โœ…โœ…โœ…2.3GB
codegemma:7b9BQ4_030sโ€”โœ…โœ…โœ…โœ…โœ…7.0GB
translategemma:4b4.3BQ4_K_M33sโ€”โœ…โœ…โœ…โœ…โœ…3.6GB
ministral-3:3b3.8BQ4_K_M33sโ€”โœ…โœ…โœ…โœ…โœ…3.3GB
qwen3:0.6b0.8BQ4_K_M36s๐Ÿง โœ…โœ…โœ…โœ…โœ…1.3GB
llama3.1:8b8.0BQ4_K_M36sโ€”โœ…โœ…โœ…โœ…โœ…4.8GB
gemma3:12b12.2BQ4_K_M41sโ€”โŒโŒโŒโŒโŒ9.3GB
minicpm-v4.5:8b8.2BQ4_K_M43sโ€”โœ…โœ…โœ…โœ…โœ…6.0GB
lfm2.5:8b8.5BQ4_K_M44s๐Ÿง โœ…โœ…โœ…โœ…โœ…4.9GB
qwen2.5:7b7.6BQ4_K_M46sโ€”โœ…โœ…โœ…โœ…โœ…4.4GB
mistral:7b7.2BQ4_K_M49sโ€”โœ…โœ…โœ…โœ…โœ…4.3GB
codegemma:2b3BQ4_051sโ€”โŒโŒโŒโŒโŒ1.7GB
deepseek-r1:1.5b1.8BQ4_K_M54s๐Ÿง โœ…โœ…โœ…โœ…โœ…1.3GB
gemma2:9b9.2BQ4_078sโ€”โœ…โœ…โœ…โœ…โœ…8.0GB
nemotron-3-nano:4b4.0BQ4_K_M79s๐Ÿง โœ…โœ…โœ…โœ…โœ…3.3GB
granite4.1:8b8.8BQ4_K_M79sโ€”โœ…โœ…โœ…โœ…โœ…5.5GB
phi3:mini3.8BQ4_081sโ€”โœ…โฐโœ…โœ…โœ…4.2GB
qwen3:1.7b2.0BQ4_K_M82s๐Ÿง โœ…โœ…โœ…โœ…โœ…1.6GB
starcoder2:3b3BQ4_085sโ€”โœ…โฐโœ…โœ…โœ…1.5GB
starcoder2:7b7BQ4_0125sโ€”โŒโŒโŒโŒโœ…3.7GB
ornith:9b9.0BQ4_K_M127s๐Ÿง โœ…โœ…โœ…โœ…โœ…6.0GB
qwen3-vl:2b2.1BQ4_K_M175s๐Ÿง โœ…โฐโœ…โœ…โฐ2.8GB
deepseek-r1:7b7.6BQ4_K_M180s๐Ÿง โœ…โœ…โฐโœ…โฐ4.4GB
qwen3-vl:4b4.4BQ4_K_M209s๐Ÿง โœ…โœ…โœ…โฐโฐ4.3GB
qwen3:8b8.2BQ4_K_M223s๐Ÿง โœ…โœ…โœ…โฐโฐ5.3GB
qwen3.5:0.8b0.9BQ8_0250s๐Ÿง โฐโฐโฐโฐโœ…1.7GB
deepseek-r1:8b8.2BQ4_K_M253s๐Ÿง โœ…โœ…โฐโฐโœ…5.7GB
qwen3.5:2b2.3BQ8_0257s๐Ÿง โฐโฐโฐโฐโœ…3.3GB
qwen3:4b4.0BQ4_K_M268s๐Ÿง โœ…โฐโฐโฐโฐ3.4GB
qwen3.5:4b4.7BQ4_K_M281s๐Ÿง โœ…โฐโฐโฐโœ…4.3GB
phi4:14b14.7BQ4_K_M300sโ€”โฐโฐโฐโฐโฐ9.7GB
qwen3.5:9b9.7BQ4_K_M300s๐Ÿง โฐโฐโฐโฐโฐ7.0GB

๐Ÿง  Key Findings

1 Thinking Mode is a Killer on CPU

Models with built-in "thinking mode" (qwen3, qwen3.5, deepseek-r1 series) generate long internal reasoning chains before answering. On a CPU-only Raspberry Pi 5, this is devastating:

  • A 0.8B model (qwen3.5:0.8b) takes 250 seconds to answer "What is 2+2?" โ€” it generates thousands of characters of "Thinking..." text before the actual answer
  • A 9.7B model (qwen3.5:9b) times out at 300 seconds on every single test
  • Even a tiny 751M model (qwen3:0.6b) takes 36 seconds โ€” 4x slower than a comparable non-thinking model

The exceptions: ministral-3:8b (24s) and lfm2.5:8b (45s) both have thinking mode but are remarkably fast. These are the only thinking-mode models I'd consider usable on RPi 5.

2 Q4_K_M is the Sweet Spot

Quantization matters enormously on CPU:

  • Q4_K_M (4-bit): Fast, good quality, the default for most models
  • Q8_0 (8-bit): 2x larger, 2x slower โ€” qwen3.5:0.8b with Q8_0 takes 250s vs what Q4 would be ~125s
  • Q4_0 (basic 4-bit): Slightly faster but lower quality than Q4_K_M

If a model only offers Q8_0 (like qwen3.5:0.8b and qwen3.5:2b), think twice before running on Pi.

3 8B Models Are Actually Usable

This was the biggest surprise. On RPi 5 with 16GB RAM:

  • llama3.1:8b โ€” 37s average, all 5 tests passed
  • minicpm-v4.5:8b โ€” 43s, all passed
  • qwen2.5:7b โ€” 46s, all passed
  • mistral:7b โ€” 50s, all passed

37 seconds for an 8B model on a $80 board is genuinely impressive. They use 4โ€“6GB RAM, leaving plenty for the OS and context window.

4 The 3B Sweet Spot

If I had to pick one model for daily use on RPi 5, it would be llama3.2:3b:

  • 14 seconds average response time
  • 3.2B parameters โ€” smart enough for most tasks
  • Only 2.8GB RAM usage
  • All 5 tests passed including Czech

At 14 seconds, it's only 5 seconds slower than 1B models, but with 3x the parameters. That's the sweet spot.

5 Code Models Are Slow

Specialized coding models (codegemma, starcoder2) are noticeably slower than general models of similar size:

  • codegemma:7b โ€” 30s (but 9B params, 7GB RAM)
  • codegemma:2b โ€” 51s (Q4_0, slower quantization)
  • starcoder2:3b โ€” 85s, timed out on Czech
  • starcoder2:7b โ€” 125s, timed out on Czech

For coding on RPi 5, codegemma:7b is the best option at 30s, but it's heavy (7GB RAM).

6 12B Works! (Barely)

gemma3:12b completed all tests in 42s average using 9.3GB RAM. That's 12.2 billion parameters running on a Raspberry Pi. Not fast, but it works and leaves enough RAM for the OS (~6GB free with 16GB total).

7 Vision Models Are Impractical

Both vision models tested (qwen3-vl:2b and qwen3-vl:4b) combine vision capability with thinking mode, making them extremely slow (176s and 210s). Several tests timed out. If you need vision on RPi 5, you'd need to find a non-thinking vision model.

๐Ÿ† Recommendations

๐Ÿ’ฌ General Chat

llama3.2:3b14s ยท 3.2B ยท best ratio
gemma3:1b8s ยท 1B ยท fastest
llama3.2:1b9s ยท 1.2B ยท Meta quality

๐Ÿš€ Heavier Tasks

llama3.1:8b37s ยท 8B ยท best large
ministral-3:8b24s ยท 8.9B ยท fast 8B
gemma3:12b42s ยท 12.2B ยท max quality

๐Ÿ’ป Coding

codegemma:7b30s ยท 9B ยท best code
codegemma:2b51s ยท 3B ยท lighter

๐ŸŒ Translations

translategemma:4b34s ยท specialized

โŒ Avoid on RPi 5 (16GB)

๐Ÿ’ก Final Thoughts

The Raspberry Pi 5 with 16GB RAM is a surprisingly capable LLM host โ€” if you pick the right models.

The key insight from testing 42 models: model architecture matters more than model size. A well-optimized 8B model without thinking mode (llama3.1:8b at 37s) can be faster than a badly-optimized 0.8B model with thinking mode (qwen3.5:0.8b at 250s).

The "thinking mode" trend in newer models (qwen3, deepseek-r1) is great for quality on powerful hardware but makes them 3โ€“20x slower on CPU. If you're running on Pi, stick to models that answer directly without internal monologues.

๐Ÿฅ‡ Top 3 Picks for RPi 5

  1. llama3.2:3b โ€” 14s, best all-rounder
  2. gemma3:1b โ€” 8s, fastest
  3. llama3.1:8b โ€” 37s, most capable