Hasty Briefsbeta

Bilingual

LLaMA Now Goes Faster on CPUs

13 hours ago
  • Justine wrote 84 new matrix multiplication kernels for llamafile, improving prompt eval speed on CPU by 30% to 500% compared to llama.cpp.
  • Kernels are optimized for ARMv8.2 (Raspberry Pi 5), Intel Alderlake, and AVX512 (Zen4) CPUs, achieving 2x faster speeds than MKL for matrices in L2 cache.
  • Llamafile uses Cosmopolitan Libc to package llama.cpp as a single-file cross-platform binary running on six OSes.
  • Performance gains are most dramatic for prompts with fewer than 1,000 tokens, making kernels a work in progress.
  • New kernels focus on q8_0, f16, q4_1, q4_0, and f32 data types; quantization may become a bottleneck with these improvements.
  • Benchmarks show significant speedups on Raspberry Pi 5, Intel i9-14900K, Mac Studio M2 Ultra, and AMD Threadripper 7995WX.
  • A practical example: using TinyLLaMA as a spam filter via shell script, running in seconds on RPI5 and milliseconds on Intel.
  • Technical approach involves outer loop unrolling and vectorized outer product, achieving up to 810 gigaflops on Alderlake i9-14900K.
  • Kernels are implemented in C++ with a custom threading model that integrates seamlessly with llama.cpp's spinlock barrier.
  • Improvements have been contributed upstream to llama.cpp under the MIT license.