Hasty Briefsbeta

Bilingual

Are we in a GPT-4-style leap that evals can't see?

a day ago
  • The author believes a subtle 'GPT-4 moment' has occurred that current benchmarks fail to capture, with major implications for industries and the economy.
  • Chat-based evaluation of LLMs is insufficient; speed of response often matters more than perceived 'quality' for everyday use.
  • Gemini 3 Pro excels at design, generating visually appealing prototypes that other models cannot replicate, effectively acting as a midlevel designer.
  • Opus 4.5 significantly reduces the need for constant supervision during coding tasks, allowing longer autonomous work with fewer errors.
  • These advances together represent an order-of-magnitude improvement in the product development lifecycle, enabling management of a cross-functional squad.
  • Current benchmarks focus on knowledge retrieval and isolated pass/fail scenarios, ignoring qualitative factors like design taste and iterative collaboration.
  • New benchmarks are needed to capture real-world usage patterns and qualitative aspects, which may explain the disconnect between benchmark scores and economic impact.