A Princeton led study used two unpublished papers from a top AI conference as a hidden test, and the agents solved the engineering but missed the research judgment.
In August 2026, a Princeton-led team published a sharper test of one of the AI industry's loudest promises, that its models are starting to improve themselves. The result, covered by MIT Technology Review and detailed in an arXiv preprint, is specific: the engineering works; the research judgment does not.
The study, led by Peter Kirgis and Sayash Kapoor, asks a different question from the ones AI vendors usually run their agents on. Instead of asking the agent to reproduce a published result whose answer key is in plain sight, the team asked the agent to answer two research questions drawn from high-quality, unpublished papers submitted to NeurIPS 2026, one of the field's top machine-learning conferences. The standard against which the agent was measured was not a leaderboard. It was the work of the human researcher who actually produced the paper. The authors call this "shadow evaluation," and the point is to remove the cheat sheet.
The model tested was Anthropic's Claude Opus 4.8, running on the open-source agent harness OpenClaw. The two test questions were both real research problems with engineering scaffolding an agent could in principle attempt. The first asked whether a large language model's "personas," the stylized roles it can adopt in conversation, could be controlled by editing the model's weights directly. The second asked how to design a detector for AI systems making predictions on spreadsheet-style data, a problem at the heart of every credit-scoring and fraud-detection deployment.
The agents could read the relevant literature, set up the code, run the experiments, and produce numbers. They could not pick the question, choose the right method when the obvious one failed, recognize when a result was an artifact, or write a paper that would survive peer review at a top conference. The authors describe the missing piece as research judgment, and the gap between engineering and research judgment is where the recursive-self-improvement promise lives or dies.
MIT Technology Review's coverage frames the result as evidence that timelines for automating AI research are running ahead of the evidence. That framing is the authors' interpretation, and it is worth treating as such. The underlying study is a preprint. The two test questions come from papers that are themselves still under NeurIPS 2026 review. A falsifiable test is one thing; a sample size of two is another.
There is a counterweight the wire coverage tends to leave out. Anthropic's own page on recursive self-improvement reports that as of May 2026, more than 80% of the code merged into Anthropic's production codebase was authored by Claude. That figure is real, but it is a measure of code generation, not research judgment. Writing a function that compiles is a different problem from choosing which experiment to run next. Both are sometimes called "AI improving AI." The shadow-evaluation result suggests the second half of that phrase is the one that breaks.
The lesson worth carrying into the next round of announcements is methodological. When a vendor, a research lab, or a headline says AI is now improving itself, the first question is what the evaluation measured. If the answer key was public, the test measured memorization and pattern-matching, not research. If the answer key was hidden, and the agent's work was held against a human researcher's actual paper, the result is closer to evidence.
A Hacker News discussion of the study raised a second question worth asking, before the preprint had been public for a week: what happens when the test is re-run against a frontier model released in the last month? Claude Opus 4.8 was state-of-the-art when the study was designed, but the gap between "what AI can do" and "what was true when this study ran" is closing faster than academic review cycles can keep up with. A single replication with a newer model would not invalidate the finding, but several replications, on the same hidden test set, would tell readers whether the engineering/research-judgment gap is a property of today's models or a more durable feature of the field.
The shadow-evaluation result, in other words, is not a forecast. It is a tool. The next time someone says the technology is now improving itself, the right response is not to accept or reject the claim on vibes. It is to ask which kind of test the claim survived, and which it did not.