Hasty Briefsbeta

Bilingual

GPT-5.6 vs. Claude Fable 5 for Physical AI, which performs best?

13 hours ago
  • Physical AI requires correct physics modeling; models can compile and run while being physically wrong, and agentic AI worsens this by having agents write and pass their own tests that may rely on flawed simplifications.
  • A study evaluated frontier AI models (OpenAI GPT 5.6 family and Anthropic Claude Fable 5) on five sealed modeling problems, including NASA's HL-20 flight vehicle, using a fixed Dyad AI harness to isolate model performance.
  • Claude Fable 5 achieved the highest difficulty-weighted score (0.889), swept all core problems, and led on the HL-20, but cost $9.60 per trial, while GPT 5.6 Sol offered best value at $1.74 per trial with strong verification.
  • The model's work-style fingerprints varied: Fable focused on verifying by attempting to falsify solutions, Sol emphasized precision but sometimes misread specs, Luna iterated heavily without external validation, and Terra economized but occasionally trimmed margins incorrectly.
  • The harness provides more leverage than the model choice; swapping the harness improved scores by 0.366 points (from 0.533 to 0.899), more than double the 0.162 gap between best and worst models, emphasizing the importance of domain-specific tooling.
  • The study concludes that for physical AI, selecting the right harness is the primary decision, followed by choosing the best affordable model within it, as the harness enforces proper verification and modeling practices.