big-pickle Model Benchmarked on SWE Atlas Codebase QnA
TL;DR. OpenCode Zen's big-pickle model achieved a 50.8% task resolve rate on Scale AI's SWE Atlas Codebase QnA benchmark. - The model outperformed all other entries within its Mini-SWE-Agent scaffold class on the official leaderboard. - big-pickle also surpassed GPT models running on the older Codex scaffold, trailing only proprietary Claude models. - The benchmark was conducted using official harness and judge models, with some resource caveats noted.
- OpenCode Zen's big-pickle model scored 50.8% on Scale AI's SWE Atlas Codebase QnA benchmark.
- big-pickle topped all Mini-SWE-Agent scaffold entries on the official leaderboard.
- The model also outranked GPT entries using the Codex scaffold, placing behind only Claude Code models.
- The evaluation followed Scale's published protocol, using their official harness and judge model.
Sources
- Big Pickle on SWE Atlas – Codebase QnA — github.com