VideoAI agents tried to do real science. Both papers got rejected.
TL;DR Given $3,000 and six days, AI agents ran hundreds of experiments but never spotted a dead end. The original authors' verdict: reject.
Picture the dream: you hand an AI agent a research question, a budget and a GPU, and a week later a paper comes out. A Princeton-led team actually tried it. The agent got six days, $3,000 in API credits and the central question of two unpublished NeurIPS 2026 papers. Then the real authors reviewed the results like any other submission.
The verdict: 2/6, reject. 1/6, strong reject. And it wasn't laziness. The agents wrote literature reviews the authors praised, ran hundreds of experiments without help, and even dropped claims that didn't hold up.
What they couldn't do was the thing that makes a scientist: notice a dead end and change the plan. They narrowed claims instead of rethinking, rated their own weak work as nearly good enough, ignored rules they'd acknowledged, and finished with half the budget unused.
- Poor judgment about what's publishable
- Uncreative when an idea failed
- Never backtracked from dead ends
- Bad at budgeting time and money
- Drifted from the instructions
Use agents as a very fast lab assistant. Keep the judgment calls (when to stop, when to pivot) for yourself.
Sources: Princeton-led study, NeurIPS 2026 papers. Caveats: only two papers, and reviewers knew they were AI-written.


