Hasty Briefsbeta

Bilingual

Calling the AI bluff: Adding "Do not guess" cut made-up claims from 71% to 20%

7 hours ago
  • The 'Do not guess' instruction reduced made-up fields by AI models from 70.7% to 20.2% in a web extraction test.
  • A twin-page test was used where one page had the answer and the other omitted it, with a decoy value to trap guesses.
  • Firecrawl, a paid API, made up 24 out of 36 missing fields, outperformed by most models using the instruction.
  • A cheap checker (GPT-6 Luna) caught 38 of 49 made-up values and rejected no correct ones, offering a low-cost verification method.
  • The marketplace 'Earn an Honest Dollar' requires agents to return null for unknown fields to ensure trustworthy transactions.
  • Limitations include single runs per contestant, synthetic pages, free tier APIs, and exclusion of email traps.