- AI chatbots like ChatGPT, Gemini, and Grok failed to update predictions based on experimental evidence in a pen demonstration, sticking to incorrect assumptions.
- A study showed AI agents ignored evidence in 68% of scientific reasoning tasks, made unsupported claims in 53%, and used contradictory evidence to change output only 26% of the time.
- AI systems lack an iterative reasoning process similar to human scientists, often refusing to revise hypotheses despite clear evidence, limiting their reliability in science and medicine.
- Researchers developed a benchmark to evaluate AI agents' reasoning process rather than just outcomes, revealing gaps in their ability to incorporate new data transparently.