- AMD's software and hardware have significantly improved, shifting from a 0% to a high chance of success against Nvidia's CUDA moat, but key risks remain.
- Two major risks: 1) Helios rack production challenges due to cable design and Ethernet retimer requirements; 2) Lack of stable GPU clusters for internal software development and CI testing, slowing progress.
- AMD's MI455X silicon leads in logic and memory integration (2nm, 12 HBM4 stacks, 432GB), but microarchitecture and software still lag behind Nvidia's Rubin.
- AMD's software stack (ROCm, vLLM, SGLang) is improving rapidly, especially in single-node inference and documentation, but distributed inference (disaggregation, WideEP) needs more work.
- AI agents and open-source tools like ROCm.ai/GEAK are helping AMD close the software gap faster than traditional methods.
- AMD needs to prioritize stable CI/GPGPU clusters for internal teams, upstream distributed inference support, and ensure composability of optimizations across models.
- Anthropic and Microsoft/OpenAI have committed to AMD's MI455X Helios, with financial incentives like equity rebate discounts making TCO highly attractive.
- The CUDA moat is eroding as AI agents reduce the need for manual kernel engineering, benefiting AMD's open-source approach.