UC Berkeley Haas Research Reveals AI Chatbots Exhibit Human-Like Overconfidence, Raising Trust Concerns
New research from UC Berkeley Haas, led by Professor Don Moore, indicates that AI chatbots, including leading large language models from OpenAI, Google, Anthropic, and Meta, exhibit significant overconfidence in their responses. The study, which tested 11 LLMs, found that on average, these models reported 88% confidence but were correct only 79% of the time, resulting in a nine-point overconfidence gap. This overconfidence mirrors a human tendency, where models become more overconfident as questions become harder and underconfident on easier ones, a phenomenon known as the 'hard-easy effect.' Moore's research suggests that AI chatbots are often trained to be compliant, agreeable, and quick to provide answers rather than admitting uncertainty, a behavior similar to humans under pressure to appear competent. This leads to 'hallucinations' where chatbots confidently present made-up facts, cases, or statistics.