Dirty Optimization Secrets (C for Playdate)
14 hours ago
- The Playdate CPU is fast but memory access is slow; optimize for register operations and avoid frequent memory loads.
- Instruction cache is tiny (4KB on Rev A, 16KB on Rev B); use -Os and compress core code to fit, placing rare ops elsewhere.
- Use custom linker scripts and section attributes to place core code contiguously and control alignment for cache efficiency.
- Inspect symbols with nm to check code size and addresses for cache fitting and alignment.
- DTCM (tightly-coupled memory) is faster; use stack-allocated data or a custom low-stack pool for frequently accessed variables.
- ITCM allows copying code to TCM for faster execution; use special macros and linker symbols, but handle Thumb bit and calls carefully.
- Performance can vary randomly due to cache misalignment and branch prediction table conflicts; align code to 32 bytes and use ALIGN(1024) tricks.
- Prefetching may help with indirection, but results are uncertain; measure with high-precision timers.