AI Long-Horizon Task Reliability Remains Elusive
TL;DR. AI struggles with long, multi-step tasks due to feedback and credit assignment challenges, limiting reliable autonomy to short durations. - An agent's task completion reliability decreases significantly as the number of steps increases, even with high per-step accuracy. - Current AI training environments are too clean, failing to prepare agents for the complexities of real-world, long-horizon work. - Job markets show a shift from AI-exposed roles for younger workers, though AI also fosters new company creation through cheaper execution.
- AI agents struggle with 'horizon problem' on tasks over 16 hours due to difficulty assigning credit for actions.
- A model 95% reliable per step sees success rates drop to 36% over 20 steps, indicating a capability wall for complex tasks.
- AI-exposed jobs for 22-25 year olds are shrinking by 3.8% annually, contrasting with overall job market stability.
- Cheaper AI execution enables more new business creation, but necessitates rapid labor reallocation to new roles.
Sources
- Theses on AI — smunshi.net