AMD Advancing AI 2026: Talking CDNA5 with AMD's Alan Smith
5 hours ago
- AMD's CDNA5 architecture rebases from GCN to RDNA, aiming for modern architecture, better efficiency, and unified GPU roadmap for AI and gaming.
- CDNA5 uses chiplet design with two compute chiplet versions: one optimized for HPC (double precision) and one for AI (high throughput for vector/tensor ops).
- Wave64 support is deprecated in CDNA5; Wave32 is prioritized to reduce overhead and improve common-case performance.
- VGPR count per wave increased from 256 to 1024 to reduce register pressure, though physical register file size is smaller than RDNA.
- WGP cache has fixed 320 KB LDS and 64 KB vector data cache on MI455, but can be programmable in future versions.
- L2 cache redesign moves to client-side caches on base dies, eliminating kernel boundary flushes and increasing bandwidth, with 14 TB/s die-to-die bandwidth.
- Two base dies each have a global L2 cache (GL2) shared by multiple shader engines, improving data reuse and reducing cross-chiplet latency.