A research test without an answer key

Most AI benchmarks need a target that software can score. Improve a model by a fixed amount. Reproduce a known result. Beat a baseline. Those tests matter, but a clear score also turns the job into hill-climbing: try something, read the number, keep what wins.

The new study used what its authors call a shadow evaluation. Researchers took the central questions from two unpublished NeurIPS 2026 submissions and gave each question to a well-resourced frontier agent. The agent could not look up the paper’s methods or findings because they were not public. Afterward, the people who had spent months on the original question graded the agent’s paper as conference reviewers.

Each main run had a six-day schedule, $3,000 in model credits, GPU access, a Linux machine, the open web, subagents, and several AI review tools. The researchers later granted a 24-hour extension and requested a clearer final draft. They also fixed a general scaffold bug. They did not choose the experiments or write the findings for the agent.

The result was blunt. The paper on language-model personas received 2 out of 6, or Reject. The paper on detecting distribution shifts in tabular foundation models received 1 out of 6, or Strong Reject. The original authors pointed to poorly motivated experiments, conclusions that outran the evidence, limited significance, and prose that made it hard to tell what mattered.

The engineering was real

It would be a mistake to read those scores as total failure. The agents handled work that would have sounded implausible a few years ago. They navigated unfamiliar code, scheduled compute-heavy experiments, managed GPU resources, assembled literature, generated plausible hypotheses, and produced complete manuscripts without a human conducting the experiments.

The expert reviewers found parts of the literature reviews useful. Both agents proposed early directions similar to ones the human researchers had considered. Each run also surfaced at least a small finding the original authors found relevant. The agents were not merely producing random academic-looking text.

That distinction matters outside a lab. An AI assistant can be genuinely valuable at gathering sources, setting up analysis, checking a long list of cases, and turning results into a draft. None of that proves it picked the right question, used enough evidence, or recognized when a neat result was too weak to carry the claim.

The study’s uncomfortable contribution is to put both facts on the same page: the assistant can remove days of technical toil and still leave the hardest decision untouched.

It heard ‘reject’ and kept polishing

The clearest failure was not that the agents missed criticism. They received plenty of it. Across 15 rounds of internal and external AI review summarized by the researchers, no review returned an acceptance. Several critiques anticipated the same problems the human authors later identified.

The agents mostly answered those critiques by narrowing claims, adding caveats, and revising prose. They did not often redesign the experiment or abandon the premise. Both retired their most ambitious research targets within the first ten hours. In the personas run, the agent had budgeted roughly 42 hours for open exploration and settled on a direction after about five.

The resource behavior told the same story. Both main runs ended after using less than half of their model-credit budgets, even though the systems could see the remaining spend and time. More money would not automatically have fixed the work. The revealing part is that the agents declared completion while their own review process still described the papers as weak.

This is a familiar failure in ordinary office work. A draft accumulates formatting, footnotes, and increasingly careful disclaimers because rewriting the document feels cheaper than admitting the underlying approach is wrong. AI makes that motion faster. It does not yet make the judgment easier.

Two papers do not settle whether AI can do research

The evidence is deliberately small. The study covers two open-ended empirical machine-learning questions and five runs in total: two pilots, two main runs, and one robustness run with a different model and scaffold. That final run reproduced most of the same failure patterns, which weakens the ‘one bad setup’ explanation without eliminating it.

The grading was not blind. The reviewers knew the papers were AI-generated and had already developed their own approaches to the questions. They could prefer their own framing. The researchers also made judgment calls when choosing the questions, configuring the systems, and interpreting the logs.

So the defensible claim is not that AI agents cannot conduct research. It is that these well-resourced agents did not produce top-conference-quality answers to two hard, open questions, even while completing much of the surrounding engineering. Incremental research with a crisp metric may be a better fit. Stronger models, better systems, or a human steering the pivots may change the result.

That is still useful evidence. The phrase ‘AI researcher’ compresses several jobs into one title. This study separates running experiments from deciding which experiment deserves another week.

Ivy would delegate the pile. Cass would count every rejected pile.

Ivy Chen sees a practical division of labor. Give the AI the literature table, environment setup, repeated runs, plot generation, and first draft. Keep a named researcher responsible for the hypothesis, the stop-or-pivot call, and the sentence that says what the evidence changes. If that researcher spends Friday night untangling a polished 40-page dead end, the system did not return much time.

Cass Bell is watching the reporting incentive. Once a complete paper becomes cheap, a team can generate many of them and publicize the one that passes review. An acceptance means less if nobody reports the rejected runs, the compute spent, or how often a human quietly rescued the question. The unit to watch is not ‘paper produced.’ It is ‘credible finding per full attempt.’

Those views pull in different directions for a reason. The tool is already useful enough to delegate serious work. It is also cheap enough to manufacture impressive-looking evidence about itself. A good trial needs both a real assignment and an honest denominator.

What to hand an AI research assistant this week

Start with work that has a visible check: collect every source matching a rule, reproduce an existing figure, run the same analysis across ten variants, compare outputs against a held-out set, or build a table that a researcher can spot-check. Let the assistant own the repetition.

Keep three decisions with a person: what result would actually change our mind, when the current approach has failed, and whether the finding matters outside this dataset. Write those questions before the run. If the agent returns a polished artifact without answering them, the work is not done.

Then save the failed paths. A negative result can be valuable, but only when the test had enough power and the team can see why the hypothesis failed. A caveat added after a weak experiment is not the same thing as better evidence.

The best near-term use of AI in research may be less dramatic than an autonomous scientist. It may be an assistant that clears enough setup and repetition for a human to spend more of the week on judgment. That only gives time back if the human is allowed to rethink the question—not assigned a second shift reviewing whatever the machine decided to finish.