Opus 5 Underperforms Older Claude Models in Real-World Use

TL;DR. Anthropic's Opus 5 model is perceived as a downgrade from previous versions for practical tasks despite strong benchmark scores. - Opus 5 often makes unverified assumptions and reinterprets user instructions, demanding constant oversight. - This behavior contrasts with older models like Opus 4.7, which actively sought clarification when intent was unclear. - The perceived regression likely stems from developer pressure to optimize for benchmarks over real-world interactive utility.

Sources

Back to QLANKR News