LLaMA Now Goes Faster on CPUs
13 hours ago
- Justine wrote 84 new matrix multiplication kernels for llamafile, improving prompt eval speed on CPU by 30% to 500% compared to llama.cpp.
- Kernels are optimized for ARMv8.2 (Raspberry Pi 5), Intel Alderlake, and AVX512 (Zen4) CPUs, achieving 2x faster speeds than MKL for matrices in L2 cache.
- Llamafile uses Cosmopolitan Libc to package llama.cpp as a single-file cross-platform binary running on six OSes.
- Performance gains are most dramatic for prompts with fewer than 1,000 tokens, making kernels a work in progress.
- New kernels focus on q8_0, f16, q4_1, q4_0, and f32 data types; quantization may become a bottleneck with these improvements.
- Benchmarks show significant speedups on Raspberry Pi 5, Intel i9-14900K, Mac Studio M2 Ultra, and AMD Threadripper 7995WX.
- A practical example: using TinyLLaMA as a spam filter via shell script, running in seconds on RPI5 and milliseconds on Intel.
- Technical approach involves outer loop unrolling and vectorized outer product, achieving up to 810 gigaflops on Alderlake i9-14900K.
- Kernels are implemented in C++ with a custom threading model that integrates seamlessly with llama.cpp's spinlock barrier.
- Improvements have been contributed upstream to llama.cpp under the MIT license.