Opus 5 Underperforms Older Claude Models in Real-World Use
TL;DR. Anthropic's Opus 5 model is perceived as a downgrade from previous versions for practical tasks despite strong benchmark scores. - Opus 5 often makes unverified assumptions and reinterprets user instructions, demanding constant oversight. - This behavior contrasts with older models like Opus 4.7, which actively sought clarification when intent was unclear. - The perceived regression likely stems from developer pressure to optimize for benchmarks over real-world interactive utility.
- Opus 5 is described as more capable in benchmarks but less effective for practical work than Opus 4.7, 4.8, and Fable.
- The model's tendency to make assumptions and reinterpret plans without asking for clarification causes frustration for users.
- This behavior is speculated to be a consequence of balancing desires for self-improving AI and high benchmark scores, leading to models that make bold assumptions.
Sources
- Why does Opus 5 feel worse to work with? — mun-logadan.github.io