Gallery of Processor Cache Effects
3 days ago
- Processor caches (L1, L2, L3) are fast but small memories that store recently accessed data, and understanding their behavior is crucial for program performance.
- Memory accesses dominate loop performance; stepping through an array with strides less than the cache line size (64 bytes) touches the same number of cache lines, making stride 1 and 16 loops run at similar speeds.
- Cache lines (64 bytes) are fetched as a whole; accessing multiple values within the same cache line is cheap, which explains why small strides don't reduce runtime proportionally.
- L1 and L2 cache sizes (e.g., 32 kB and 4 MB) cause performance drops when an array exceeds those sizes, as shown by timing experiments.
- Instruction-level parallelism can be exploited by avoiding data dependencies; incrementing two independent array elements is faster than incrementing the same element twice in a loop.
- Cache associativity (e.g., 16-way set associative) can cause conflicts when repeatedly accessing more than 16 cache lines from the same set, leading to performance degradation for certain stride values.
- False cache line sharing on multi-core CPUs invalidates entire cache lines when different cores modify values on the same line, severely slowing concurrent updates.
- Hardware optimizations and unpredictability (e.g., memory bank effects) make it essential to measure and verify performance assumptions rather than relying solely on theory.