Measured LLM inference speeds on Apple Silicon, with raw data (CC BY 4.0)
2 days ago
- Benchmarks for LLM inference on Apple Silicon were measured using a production fleet with Ollama API, fixed prompt, and 512-token generation budget.
- Results for M4 Mac mini (16 GB, 120 GB/s) include models from 3B to 14B parameters, with generation speeds ranging from 46.7 tok/s (Llama 3.2 3B) to 11.7 tok/s (Qwen 2.5 14B).
- Future benchmarks will cover M4 Pro, M4 Max, and M3 Ultra chips, with performance expected to scale with memory bandwidth.
- Methodology: Ollama 0.31.2 on macOS 15.3.1, default Q4_K_M quantization, 3 recorded runs after warm-up, median speeds reported. Raw data available under CC BY 4.0.