عاجل
ESFIFA desiste de crear filial para comercializar el Mundial tras rechazo de confederacionesESConfirman la muerte de Nirmal Purja en una avalancha en Broad PeakESEl Gobierno reduce el descuento al carburante a 10 céntimos por litro a partir de agostoESTres miembros de la UME heridos al volcar su camión en Robledo de ChavelaESSánchez solicita reunión urgente de la UE tras la crisis migratoria en CeutaESHamas acepta plan de desarme en Gaza con condiciones; Israel lo rechazaESResidentes de Lomas de Chapultepec protestan contra restaurante-bar por certificado de uso de suelo presuntamente ilegalESUn ataque antisemita en Barcelona: la kipá que identificó a siete judíos como presas de una turbaESLa ola de calor dispara las averías de coches en España durante la operación salidaESSheinbaum recibirá a la sección 22 de Oaxaca el 13 de agostoESFIFA desiste de crear filial para comercializar el Mundial tras rechazo de confederacionesESConfirman la muerte de Nirmal Purja en una avalancha en Broad PeakESEl Gobierno reduce el descuento al carburante a 10 céntimos por litro a partir de agostoESTres miembros de la UME heridos al volcar su camión en Robledo de ChavelaESSánchez solicita reunión urgente de la UE tras la crisis migratoria en CeutaESHamas acepta plan de desarme en Gaza con condiciones; Israel lo rechazaESResidentes de Lomas de Chapultepec protestan contra restaurante-bar por certificado de uso de suelo presuntamente ilegalESUn ataque antisemita en Barcelona: la kipá que identificó a siete judíos como presas de una turbaESLa ola de calor dispara las averías de coches en España durante la operación salidaESSheinbaum recibirá a la sección 22 de Oaxaca el 13 de agosto
Newsgather
رجوعStudy: Frontier AI Agents Fail to Produce Original Research for Top Conferences
Study: Frontier AI Agents Fail to Produce Original Research for Top Conferences
تقنية
Decryptقبل 5 ساعاتتقنية2 د قراءة

Study: Frontier AI Agents Fail to Produce Original Research for Top Conferences

نظرة سريعة

A new study by researchers from Princeton, Stanford, and other institutions found that current frontier AI agents can complete many engineering tasks for AI research but failed to generate original scientific contributions worthy of acceptance at a top machine learning conference like NeurIPS.

ملخص مُنشأ بالذكاء الاصطناعي

لماذا يهم

Researchers evaluated frontier AI agents by giving them central research questions from two unpublished NeurIPS 2026 papers, providing resources like GPU, internet, and API credits, and having the original authors review the AI-generated papers.

حجم الخط

A new study found today's frontier AI agents could complete many of the engineering tasks required for AI research but failed to produce original work worthy of acceptance at a top machine learning conference.

In the study, "Can AI agents conduct open-ended AI research?" published on Wednesday, researchers from Princeton University, the UK AI Security Institute, Stanford University, the University of Toronto, and several academic and research organizations evaluated whether frontier AI agents could independently conduct original AI research.

“Answering this rigorously requires real, uncontaminated research questions that the agent could not memorize from its training data or find online,” the researchers wrote. “To satisfy these requirements, we rely on high-quality AI research that was not public at the time we conducted the experiments.”

The researchers gave AI agents the central research questions from two unpublished NeurIPS 2026 papers, preventing the systems from retrieving answers from training data or the web. Each agent received six days, thousands of dollars in API credits, GPU resources, internet access, and access to a virtual machine to produce a conference-quality paper. The resulting papers were then reviewed by the original authors of the unpublished research. Both were rejected.

The agents completed much of the engineering required for research, conducting literature reviews, debugging software, running experiments, managing GPU resources, and producing complete academic papers without human intervention. But reviewers concluded the systems failed to generate original scientific contributions worthy of publication at a top machine learning conference.

The authors said their evaluation better measures scientific reasoning than previous benchmarks because it tests open-ended research problems rather than predefined tasks.

The authors cautioned that the study examined only two research projects and acknowledged limitations, including the small sample size and the fact that the original researchers evaluated the AI-generated papers. They said the results suggest current frontier AI agents can automate many of the engineering tasks involved in research but continue to struggle with generating original scientific work.

The study comes as researchers continue to uncover surprising and sometimes risky behaviors in increasingly autonomous AI agents.

أسئلة مفتوحة

  • How will AI agent capabilities evolve in generating original research?
  • What specific limitations prevent AI from original scientific contributions?
  • How can evaluation methods for AI research be further refined?

مواضيع ذات صلة

This article was originally published by Decrypt.

أخبار ذات صلة

المزيد حول هذا الموضوعai agents