CladBench Evaluates LLM Performance on UK Building Regulations
TL;DR. CladBench, an open evaluation benchmark, assesses the accuracy of large language models on UK and EU building regulations, providing insights into their real-world professional application performance for built environment professionals. - CladBench evaluates LLMs on 536 questions across twelve categories related to building regulations. - Claude Opus 4.7 currently leads in accuracy, surpassing other tested models like GPT-5 and Gemini 2.5 Pro. - Detailed scores and confidence intervals for the benchmark results are publicly available.
- CladBench provides a specialized benchmark for LLMs on UK/EU built environment regulations.
- It features 536 questions in categories like EPC prediction, BREEAM eligibility, and Net Zero reasoning.
- Claude Opus 4.7 achieved the highest score, demonstrating superior performance in this domain.
- The benchmark addresses a gap in general LLM evaluations for industry-specific tasks.