Code Arena Benchmark for Web Development
Assessing Model Capabilities in Building Interactive Web Apps
L'essentiel
Code Arena, a benchmark by researchers from UC Berkeley, UC San Diego, and Carnegie Mellon, evaluates AI models' ability to create complete web applications from scratch, with rankings based on blind user votes reflecting real developer preferences.
Résumé généré par IA
Alibaba owns the South China Morning Post. Unlike traditional coding benchmarks such as HumanEval or SWE-bench, which rely on standardised tests, Code Arena users test how well models can independently build complete, interactive web applications from scratch, based on user prompts. Users then vote on anonymised outputs in blind comparisons, meaning the leaderboard closely reflects the preferences of real-world developers. The benchmark is run by Arena, an organisation founded by researchers from the University of California, Berkeley in collaboration with University of California San Diego and Carnegie Mellon University.






