- Grok 4.5, GPT-5.5, Claude Opus 4.8, and Fable 5 were tested in a one-shot interactive app building challenge.
- Round 1 (3D Rubik's Cube): Claude Opus 4.8 and Fable 5 succeeded first try; Grok 4.5 required a retry; GPT-5.5 produced an incomplete cube.
- Round 2 (Particle Gravity Sandbox): All models delivered working apps, with GPT-5.5 considered most mesmerizing.
- Round 3 (Breakout Game): All four models produced fully playable, polished games on the first attempt.