Arena study shows AI models prefer their own answers 58% of the time

1 week ago 17



When you ask an AI to judge a writing contest and one of the contestants is itself, things get awkward. New research examining self-preference bias in large language models finds that AI judges favor their own generated responses roughly 58% of the time in pairwise comparisons. The findings cut to the heart of a growing dependency in the AI industry: using models to evaluate other models. Platforms like Chatbot Arena, which pit AI systems against each other in blind head-to-head matchups, have become de facto benchmarks for measuring progress. The narcissism problem The bias, sometimes called egocentric or narcissistic bias, shows up consistently across the industry’s biggest names. OpenAI’s GPT, Anthropic’s Claude, and Google’s Gemini all exhibit the tendency when placed in judging roles. Research suggests that somewhere between 40% and 96% of the observed self-preference could stem from actual quality differences in the outputs being compared. In other words, some models might genuinely prefer their own answers because those answers are, by certain metrics, better. One way researchers have tried to tease this apart is through perplexity, a measure of how surprised a model is by a...

Read Entire Article