LLMs' Hacking Ability Tested on Vulnerable App
TL;DR. Researchers spent
TL;DR. Researchers spent
,500 testing multiple LLMs' ability to hack a custom-built vulnerable book review application. - GPT 5.5 achieved the highest success rate, solving 7 out of 10 challenges with an average cost of $9.46 per successful solve. - DeepSeek-v4-Pro showed promise at a low cost, while other popular models like Claude and Gemini performed poorly. - The experiment highlights specific LLMs' current proficiency and cost-effectiveness in security vulnerability identification.