The provided text explores why large language models (LLMs) hallucinate, arguing that this behaviour stems from their training and evaluation processes. It posits that during pre-training, errors, including plausible falsehoods, arise naturally due to statistical pressures, even with error-free data. The authors draw a connection between generative errors in LLMs and misclassification in binary classification, illustrating how factors like arbitrary facts and poor model design contribute to these issues. Furthermore, the text contends that hallucinations persist post-training because current evaluation benchmarks predominantly employ binary grading schemes, which inadvertently reward guessing over admitting uncertainty. The article advocates for a socio-technical mitigation strategy, suggesting modifications to existing benchmarks to explicitly value uncertainty, thereby fostering more trustworthy AI systems.
