Hasty Briefsbeta

Bilingual

A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

2 days ago
  • The playbook consists of an open-source model, proprietary task data, and a reinforcement-learning stage.
  • Bridgewater Associates trained an open-source model on labels from expert investors, reducing mistakes by 30% at lower cost.
  • Harvey used reinforcement learning on an open-weight model to create a legal agent outperforming GPT-5.5 and Claude Opus 4.8.
  • Intercom post-trained its own vertical support model on billions of customer interactions, resolving more issues cheaper than frontier models.