OpenAI has a LOT of work to do if they think Luna can compete with Jev
4 hours ago
- OpenAI's Decisions API, announced at DevDay, picks one answer from a list, enabling software branching on decisions.
- The API runs on a specialized version of GPT-6 Luna, a low-cost model, with a decision speed of 150 milliseconds.
- When Luna claims 99% confidence, its actual accuracy is only 68%, showing significant overconfidence.
- Luna's accuracy drops sharply with chained reasoning: from 93% at depth zero to 45% at five inferences, vs Jev's 89%.
- Luna's confidence calibration is poor, with an expected calibration error of 0.32, making thresholds unreliable.
- Log-probabilities from Luna are not hidden but indicate overconfidence, not secrecy, as the main issue.
- The Decisions API potentially provides a confidence score, but users should test on actual task difficulty before trusting.