- The author believes a subtle 'GPT-4 moment' has occurred that current benchmarks fail to capture, with major implications for industries and the economy.
- Chat-based evaluation of LLMs is insufficient; speed of response often matters more than perceived 'quality' for everyday use.
- Gemini 3 Pro excels at design, generating visually appealing prototypes that other models cannot replicate, effectively acting as a midlevel designer.
- Opus 4.5 significantly reduces the need for constant supervision during coding tasks, allowing longer autonomous work with fewer errors.
- These advances together represent an order-of-magnitude improvement in the product development lifecycle, enabling management of a cross-functional squad.
- Current benchmarks focus on knowledge retrieval and isolated pass/fail scenarios, ignoring qualitative factors like design taste and iterative collaboration.
- New benchmarks are needed to capture real-world usage patterns and qualitative aspects, which may explain the disconnect between benchmark scores and economic impact.