big-pickle Model Benchmarked on SWE Atlas Codebase QnA

TL;DR. OpenCode Zen's big-pickle model achieved a 50.8% task resolve rate on Scale AI's SWE Atlas Codebase QnA benchmark. - The model outperformed all other entries within its Mini-SWE-Agent scaffold class on the official leaderboard. - big-pickle also surpassed GPT models running on the older Codex scaffold, trailing only proprietary Claude models. - The benchmark was conducted using official harness and judge models, with some resource caveats noted.

Sources

Back to QLANKR News