AI Search Agents Confirm Knowledge, Not Research, New Benchmarks Show
TL;DR. A new study reveals leading AI search agents prioritize existing knowledge over actual web research when responding to queries. - Researchers developed LiveBrowseComp to test models on recent events, forcing reliance on live information. - Models performed poorly on LiveBrowseComp without prior knowledge, exposing an "intrinsic knowledge dependence." - The findings challenge previous benchmark results, suggesting current evaluations overstate AI search capabilities.
- AI search agents, including GPT-5.4 and Kimi K2.6, frequently use the web to confirm pre-existing knowledge rather than to conduct new research.
- The Harbin Institute of Technology created LiveBrowseComp, a time-based benchmark, to test AI search on events less than 90 days old.
- When faced with unfamiliar, recent information, AI models' performance significantly declines, disrupting current benchmark rankings.
- Models often perform worse when given search tools but no relevant sources, indicating a tendency to search for self-affirming information.