Code Arena Benchmark for Web Development
Assessing Model Capabilities in Building Interactive Web Apps
نظرة سريعة
Code Arena, a benchmark by researchers from UC Berkeley, UC San Diego, and Carnegie Mellon, evaluates AI models' ability to create complete web applications from scratch, with rankings based on blind user votes reflecting real developer preferences.
ملخص مُنشأ بالذكاء الاصطناعي
Alibaba owns the South China Morning Post. Unlike traditional coding benchmarks such as HumanEval or SWE-bench, which rely on standardised tests, Code Arena users test how well models can independently build complete, interactive web applications from scratch, based on user prompts. Users then vote on anonymised outputs in blind comparisons, meaning the leaderboard closely reflects the preferences of real-world developers. The benchmark is run by Arena, an organisation founded by researchers from the University of California, Berkeley in collaboration with University of California San Diego and Carnegie Mellon University.







