A new study challenges the industry premise that AI is close to improving itself with minimal human oversight. Researchers found that while AI agents can handle the technical engineering required for research, they fall short on open-ended investigation requiring judgment and creativity.
What Happened
The multi-institution team led by Princeton University's Peter Kirgis and Sayash Kapoor proposed a 'shadow evaluation' method to test whether AI agents could conduct genuine research. They asked Anthropic's Claude Opus 4.8, running on OpenClaw, to answer unpublished research questions from two papers submitted to NeurIPS 2026—one about controlling language model personas through weight editing, another about detecting unreliability in spreadsheet prediction models. The agents received six days, $3,000 in API credits, GPU resources, virtual computers, and web access to produce publication-quality work. The original paper authors then graded the agent-produced papers as they would a conference submission. Both were rejected. Researchers found the agents could review literature, run hundreds of experiments, and compile results. But they ran bizarre experiments on tiny synthetic datasets, struggled to write intelligibly, made no novel contributions, committed too quickly to unpromising approaches, and failed to incorporate feedback from subagents or external reviewing tools.
Why It Matters
The findings suggest that hyped timelines for automating AI research may be running ahead of evidence. Current agents excel at narrow, checkable tasks but lack the judgment needed for open-ended scientific inquiry—choosing hypotheses, deciding what evidence settles a question, and knowing when to pivot. The gap raises questions about how quickly the industry can achieve recursive self-improvement, where AI systems improve their own capabilities without human guidance.
The Bottom Line
The study indicates AI agents remain fundamentally limited in conducting original research at conference quality. While capable of engineering tasks, they lack the creativity, judgment, and flexibility researchers say are integral to building self-improving AI.