arrow_backBack to research feed
agentsPublished: July 29, 2026

Can AI agents conduct open-ended AI research? Early evidence from two case studies

By Peter Kirgis, Sayash Kapoor, Andrew Schwartz, Stephan Rabanser, David Africa, Konstantinos Voudouris, Viet Nguyen, Toby Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Yue Ling, Abhishek Shetty, Helen Toner, Gillian Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani, Arvind Narayanan

Research TL;DR

"Introduces shadow evaluations where AI agents attempt open-ended AI research; agents perform engineering but fail at creative research tasks, highlighting five failure modes."

Abstract

Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.

Technical Analysis & Implementation

Summary§

This paper introduces a novel evaluation method, shadow evaluations, to measure the ability of AI agents to conduct open-ended AI research. Unlike existing benchmarks that focus on narrow, verifiable tasks or rely on noisy peer review, shadow evaluations have an agent tackle the central research question of a high-quality unpublished paper, and the original authors assess the output. The authors ran two such evaluations on unpublished NeurIPS 2026 submissions, providing frontier agents with six days and thousands of dollars in compute. The agents completed all engineering tasks without human assistance but failed to make substantial progress on the research questions, leading to unambiguous rejection by the original authors.

Methodology§

Shadow evaluations involve the following steps:

1. Select a target paper: Choose an unpublished, high-quality paper with a clearly defined open-ended research question. 2. Define the agent task: The agent must attempt to answer the paper's research question, given only the problem description (not the solution). The agent can use any resources, including code, literature, and compute. 3. Run the agent: Deploy a state-of-the-art LLM agent with a scaffold (e.g., tool-use, planning) for a fixed time and compute budget. 4. Evaluate: The original authors review the agent's output using a structured rubric, including criteria such as correctness, novelty, and depth. The agent's work is scored on a 1-5 scale, with a score below 3 indicating rejection.

Key Findings§

In both case studies, the agents achieved perfect scores on engineering components (e.g., implementing baselines, writing code) but scored 1-2 on research insight, novelty, and overall contribution. The authors identified five recurring failure modes:

  • Poor judgment about the bar for publishable research: Agents produced trivial or incremental contributions without recognizing their low significance.
  • Uncreative responses to shortcomings: When initial experiments failed, agents made minor tweaks (e.g., hyperparameter adjustments) rather than proposing novel approaches.
  • Ineffective backtracking: Agents persisted on dead-end paths without recognizing the need to pivot.
  • Poor resource awareness: Agents spent disproportionate compute on low-impact experiments.
  • Instruction drift: Agents gradually deviated from the original research question, exploring tangential ideas.

A robustness check with a different model and scaffold reproduced these failures, suggesting the issues are fundamental to current agent capabilities.

Mathematical Formulation§

Let $P$ be the target paper with research question $Q$. The agent $A$ operates over a time horizon $T$ with compute budget $C$. The agent's output $O$ consists of a report $R$, code $K$, and experimental results $E$. The original authors evaluate $O$ based on criteria $c_1, c_2, \ldots, c_n$ (e.g., correctness, novelty, clarity) on a scale $[1,5]$. The overall score $S$ is a weighted sum:

$$S = \sum_{i=1}^n w_i \cdot \text{score}(c_i)$$

where $\sum w_i = 1$. A threshold $\theta$ (e.g., $\theta = 3$) determines acceptance if $S \geq \theta$.

Implementation Details§

The agents were built using a frontier LLM (e.g., GPT-4, Claude) with a scaffold that included:

  • A file system for code and data.
  • A web browser for literature search.
  • A Python execution environment.
  • A persistent memory for tracking progress.

Agents were given the paper's abstract and a description of the research question, but not the full paper or solution. They were instructed to produce a final report detailing their findings. The budget was $5000 and 6 days per run.

Code Illustration§

Below is a simplified Python script simulating the shadow evaluation loop:

import openai
import subprocess
import time

class ShadowEvaluation:
    def __init__(self, agent, paper_question, budget_dollars, time_days):
        self.agent = agent
        self.paper_question = paper_question
        self.budget = budget_dollars
        self.time_limit = time_days * 86400  # seconds
        self.start_time = time.time()
        self.cost = 0

    def run(self):
        # Agent loop
        while time.time() - self.start_time < self.time_limit and self.cost < self.budget:
            self.cost += self.agent.step(self.paper_question)
        return self.agent.final_report()

# Example usage
eval = ShadowEvaluation(agent=MyAgent(), paper_question="How to improve GAN training stability?", budget_dollars=5000, time_days=6)
report = eval.run()
print(report)

Conclusions§

The paper provides early evidence that current AI agents can handle the engineering aspects of AI research but lack the creativity, judgment, and strategic thinking required for open-ended scientific discovery. The shadow evaluation method offers a scalable and reliable alternative to existing benchmarks.

Interactive SEO Tool

Interactive LLM Token & Cost Calculator

Estimate token usage and model pricing. Enter your prompt below to see how it is parsed into tokens and calculate the exact API cost for different providers.

Context Window1,048,576 tokens
Visual Tokenizer Chunks
Language models do not read text like humans. Instead, they process text in chunks called tokens. A token can be a single character, a syllable, a word, or even part of a word (like the "ing" in "walking"). On average, 1 token is equivalent to about 4 characters or 0.75 words of English text.
Estimated Token Count124

Cost Breakdown (USD)

Input Cost (Prompt):$0.000093
Output Cost (Generated):$0.000465
Total Est. Cost:$0.000558
Context Window Capacity0.0118%

API Pricing Comparison (per Million Tokens)

ModelInputOutput
Gemini 3.6 Flash (batch)$0.75$3.75
Gemini 3.5 Flash Lite (batch)$0.15$1.25
Claude Sonnet 5 (batch)$1.00$5.00
Claude Fable 5 (batch)$5.00$25.00
MiniMax M3 (batch)$0.15$0.60
Claude Opus 4.8 (batch)$2.50$12.50
Gemini 3.5 Flash (batch)$0.75$4.50
Gemini 3.1 Flash Lite (batch)$0.13$0.75
GPT-5.5 (batch)$2.50$15.00
Claude Opus 4.7 (batch)$2.50$12.50
Lyria 3 Pro Preview$0.00$0.00
Lyria 3 Clip Preview$0.00$0.00
MiniMax M2.7$0.25$1.00
GPT-5.4 Nano$0.20$1.25
GPT-5.4 Nano (batch)$0.10$0.63
GPT-5.4 Mini (batch)$0.38$2.25
GPT-5.4$2.50$15.00
GPT-5.4 (batch)$1.25$7.50
GPT-5.2 (batch)$0.88$7.00
Qwen2.5 Coder 32B Instruct$0.66$1.00
Gemini 3.1 Pro Preview (batch)$1.00$6.00
Claude Opus 4.6 (batch)$2.50$12.50
Gemini 3 Flash Preview (batch)$0.25$1.50
Qwen3 VL 8B Instruct$0.12$0.46
MiniMax M1$0.55$2.20
Saba$0.20$0.60
o3 Mini High$1.10$4.40
Llama 3.3 70B Instruct$0.13$0.40
GPT-4o-mini$0.15$0.60
Claude Opus Latest$5.00$25.00
Qwen3.7 Flash$0.03$0.13
Hermes 3 405B Instruct$1.00$1.00
Claude Opus 4.5 (batch)$2.50$12.50
Hermes 4 70B$0.13$0.40
GPT-5.1 (batch)$0.63$5.00
Claude Haiku 4.5 (batch)$0.50$2.50
Kimi K3$3.00$15.00
GPT-5.6 Terra Pro$1.25$7.50
GPT-5.6 Sol Pro$5.00$30.00
Claude Sonnet 4.5$3.00$15.00
Claude Sonnet 4.5 (batch)$1.50$7.50
Qwen2.5 VL 72B Instruct$0.80$1.00
Claude Opus 5 (Fast)$10.00$50.00
Claude Opus 5$5.00$25.00
Qwen3 Next 80B A3B Thinking$0.15$1.20
GPT-5.6 Sol$5.00$30.00
GPT-5 (batch)$0.63$5.00
GPT-5 Mini (batch)$0.13$1.00
Grok 4.5$2.00$6.00
Claude Sonnet 5$2.00$10.00
Claude Opus 4$15.00$75.00
Claude Fable Latest$10.00$50.00
GPT-5 Nano (batch)$0.03$0.20
Claude Opus 4.1 (batch)$7.50$37.50
Gemini 2.5 Flash Lite (batch)$0.05$0.20
Gemini 2.5 Flash (batch)$0.15$1.25
Gemini 2.5 Pro (batch)$0.63$5.00
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)$0.25$1.50
Anthropic Claude Haiku Latest$1.00$5.00
Qwen2.5 7B Instruct$0.10$0.20
Llama 3.1 8B Instruct$0.05$0.08
Gemma 2 27B$0.65$0.65
Mixtral 8x22B Instruct$2.00$6.00
Kimi K2.7 Code$0.73$3.50
Nemotron 3 Ultra$0.60$3.60
Qwen3.6 Flash$0.19$1.13
GLM 4.5V$0.60$1.80
Morph V3 Large$0.90$1.90
Command R7B (12-2024)$0.04$0.15
Nano Banana 2 (Gemini 3.1 Flash Image)$0.50$3.00
Inflection 3 Productivity$2.50$10.00
GLM 5.2$0.63$1.96
MiniMax M3$0.30$1.20
GPT-5.4 Image 2$8.00$15.00
Kimi K2.6$0.65$2.72
Claude Opus 4.5$5.00$25.00
GPT-4o (2024-11-20)$2.50$10.00
o1$15.00$60.00
Step 3.7 Flash$0.20$1.15
Claude Opus 4.8 (Fast)$10.00$50.00
Gemma 4 26B A4B$0.07$0.34
Claude Sonnet 4$3.00$15.00
Gemini 2.5 Pro Preview 05-06$1.25$10.00
Claude Haiku 4.5$1.00$5.00
GPT-4 Turbo Preview$10.00$30.00
o3$2.00$8.00
Google Gemini Flash Latest$1.50$7.50
o4 Mini$1.10$4.40
Grok 4.20$1.25$2.50
Claude Opus 4.8$5.00$25.00
Gemini 3.1 Pro Preview Custom Tools$2.00$12.00
MoonshotAI Kimi Latest$2.90$15.00
Gemini 3.5 Flash$1.50$9.00
Laguna S 2.1$0.10$0.20
Gemini 3.5 Flash Lite$0.30$2.50
Muse Spark 1.1$1.25$4.25
GPT-5.6 Luna Pro$0.50$3.00
Claude Opus 4.7 (Fast)$30.00$150.00
Reka Flash 3$0.10$0.20
GPT-4o (2024-08-06)$2.50$10.00
GPT-5.6 Terra$1.25$7.50
GPT-5.5 Pro$30.00$180.00
Nano Banana 2 (Gemini 3.1 Flash Image Preview)$0.50$3.00
Claude Sonnet 4.6$3.00$15.00
Gemini 3.6 Flash$1.50$7.50
Gemini 3.1 Flash$0.25$1.50
Hy3$0.13$0.53
Laguna XS 2.1$0.06$0.12
Qwen3 VL 32B Instruct$0.10$0.42
GPT-5.6 Luna$0.50$3.00
GLM 4.6V$0.30$0.90
Codestral 2508$0.30$0.90
Command R (08-2024)$0.15$0.60
Qwen3 235B A22B Instruct 2507$0.09$0.55
Fugu Ultra$5.00$30.00
Ministral 3 8B 2512$0.15$0.15
Llama 4 Scout$0.10$0.30
Qwen2.5 72B Instruct$0.36$0.40
KAT-Coder-Air V2.5$0.15$0.60
GPT-4o (2024-05-13)$5.00$15.00
Nex-N2-Mini$0.03$0.10
KAT-Coder-Pro V2.5$0.74$2.96
Nano Banana Pro (Gemini 3 Pro Image)$2.00$12.00
GPT-4o Search Preview$2.50$10.00
Nova 2 Lite$0.30$2.50
o1-pro$150.00$600.00
Llama 4 Maverick$0.20$0.80
Gemma 3 27B$0.08$0.45
Qwen3 VL 8B Thinking$0.18$2.10
Laguna M.1$0.20$0.40
Llama 3 8B Instruct$0.14$0.14
Qwen-Plus$0.26$0.78
Mistral Large$2.00$6.00
Nex-N2-Pro$0.25$1.00
Grok 4.3$1.25$2.50
Granite 4.1 8B$0.05$0.10
Qwen3.7 Max$1.48$4.42
Grok Build 0.1$1.00$2.00
Qwen3 Next 80B A3B Instruct$0.10$1.10
Sonar Pro$3.00$15.00
GPT-3.5 Turbo (older v0613)$1.00$2.00
Claude 3.5 Sonnet v2$3.00$15.00
Gemini 3.1 Flash Lite$0.25$1.50
GPT Chat Latest$5.00$30.00
Mistral Medium 3.5$1.50$7.50
Sonar Deep Research$2.00$8.00
Claude 3 Haiku$0.25$1.25
GPT-5 Codex$1.25$10.00
Google Gemini Pro Latest$2.00$12.00
Qwen3 VL 235B A22B Thinking$0.40$4.00
Anthropic Claude Sonnet Latest$2.00$10.00
Qwen3 VL 235B A22B Instruct$0.21$1.90
Qwen3.5 Plus 2026-04-20$0.30$1.80
Sonar$1.00$1.00
MiMo-V2.5$0.14$0.28
MiMo-V2.5-Pro$0.43$0.87
GLM 5.1$0.97$3.04
Gemma 4 31B$0.10$0.34
Qwen3.6 Plus$0.33$1.95
Grok 4.20 Multi-Agent$1.25$2.50
Qwen3 30B A3B Instruct 2507$0.05$0.19
Gemini 3.1 Flash Lite Preview$0.25$1.50
GLM 4.5 Air$0.13$0.85
KAT-Coder-Pro V2$0.30$1.20
Reka Edge$0.10$0.10
GLM 5 Turbo$1.20$4.00
Nemotron 3 Super$0.09$0.40
Seed-2.0-Lite$0.25$2.00
GPT-5.4 Pro$30.00$180.00
GPT-5.3 Chat$1.75$14.00
Seed-2.0-Mini$0.10$0.40
Qwen3.5-122B-A10B$0.26$2.08
Qwen3 Max Thinking$0.78$3.90
Morph V3 Fast$0.80$1.20
GPT-4o$2.50$10.00
Qwen3.5-35B-A3B$0.14$1.00
Qwen3.5-27B$0.20$1.56
Qwen3.5 Plus 2026-02-15$0.26$1.56
MiniMax M2-her$0.30$1.20
Gemini 2.5 Pro Preview 06-05$1.25$10.00
GPT-3.5 Turbo 16k$3.00$4.00
Mistral Small 3$0.10$0.30
Gemini 3 Flash Preview$0.50$3.00
Mistral Small 4$0.15$0.60
GPT-5.3-Codex$1.75$14.00
GPT-5.5$5.00$30.00
Qwen3.5 397B A17B$0.39$2.34
GPT-5.2-Codex$1.75$14.00
Claude Fable 5$10.00$50.00
Qwen3.7 Plus$0.32$1.28
GLM 5$0.95$2.55
Qwen3 Coder Next$0.12$0.80
UI-TARS 7B$0.10$0.20
o4 Mini High$1.10$4.40
GPT-3.5 Turbo$0.50$1.50
Mistral Small 3.2 24B$0.10$0.30
o3 Pro$20.00$80.00
Gemma 3n 4B$0.06$0.12
Mistral Large 2407$2.00$6.00
Gemini 3.1 Pro Preview$2.00$12.00
GPT-5.2 Chat$1.75$14.00
Devstral 2 2512$0.40$2.00
GPT-5.1-Codex-Max$1.25$10.00
Step 3.5 Flash$0.10$0.30
Kimi K2.5$0.57$2.85
gpt-oss-20b$0.03$0.13
Claude Opus 4.1$15.00$75.00
WizardLM-2 8x22B$0.62$0.62
DeepSeek V3.2$0.27$0.40
GPT-4$30.00$60.00
Mistral Large 3$0.50$1.50
o4 Mini Deep Research$2.00$8.00
GPT-5 Mini$0.25$2.00
Qwen3 8B$0.12$0.46
Qwen Plus 0728 (thinking)$0.40$1.20
GLM 5V Turbo$1.20$4.00
Llama 3.3 70B Instruct$0.13$0.40
DeepSeek V3 0324$0.27$1.12
GPT Audio Mini$0.60$2.40
Yi-Lightning$0.15$0.30
Ministral 3 14B 2512$0.20$0.20
Qwen Plus 0728$0.26$0.78
DeepSeek V4 Pro$0.43$0.87
Voxtral Small 24B 2507$0.10$0.30
Mistral Nemo$0.02$0.03
Qwen3 Coder 30B A3B Instruct$0.07$0.27
GPT-4o-mini (2024-07-18)$0.15$0.60
GPT-5.4 Mini$0.75$4.50
Qwen3.5-Flash$0.07$0.26
MiniMax M2.5$0.15$0.90
GPT Audio$2.50$10.00
Mistral Small 3.1 24B$0.35$0.56
Solar Pro 3$0.15$0.60
Command R$0.15$0.60
GPT-5.1 Chat$1.25$10.00
GPT-5.1-Codex$1.25$10.00
Kimi K2 0711$0.57$2.30
Mistral Medium 3$0.40$2.00
Claude Opus 4.6$5.00$25.00
GLM 4.7 Flash$0.06$0.40
GPT-5$1.25$10.00
Gemini 3.1 Pro$2.00$12.00
Claude Opus 4.7$5.00$25.00
Seed 1.6$0.25$2.00
Gemini 2.5 Pro$1.25$10.00
GPT-4.1 Nano$0.10$0.40
Llama 3.2 11B Vision$0.34$0.34
Qwen3.6 35B A3B$0.14$1.00
Hy3 preview$0.06$0.21
Seed 1.6 Flash$0.07$0.30
Qwen3.6 Max Preview$1.03$6.16
Nemotron 3 Nano 30B A3B$0.05$0.20
MiniMax M2$0.26$1.02
o3 Deep Research$10.00$40.00
ERNIE 4.0$1.20$2.40
Nova Lite 1.0$0.06$0.24
GLM 4.7$0.40$1.75
Ministral 3 3B 2512$0.10$0.10
GPT-5.1$1.25$10.00
GLM 4.5$0.60$2.20
Qwen 2.5-Coder 32B$0.35$0.70
R1 0528$0.50$2.15
Llama Guard 4 12B$0.18$0.18
Doubao Pro$0.80$1.60
Qwen3 235B A22B Thinking 2507$0.30$3.00
GLM 4.6$0.50$2.00
Qwen3 Max$0.78$3.90
Qwen3 30B A3B$0.12$0.50
Kimi K2 Thinking$0.60$2.50
Gemma 3 4B$0.05$0.10
Sonar Pro Search$3.00$15.00
Qwen3.5-9B$0.10$0.15
Mercury 2$0.25$0.75
Nano Banana (Gemini 2.5 Flash Image)$0.30$2.50
Qwen3 VL 30B A3B Thinking$0.20$2.40
Qwen3 Coder 480B A35B$0.30$1.00
Gemini 2.5 Flash Lite$0.10$0.40
Qwen3 VL 30B A3B Instruct$0.15$0.60
o3 Mini$1.10$4.40
Palmyra X5$0.60$6.00
gpt-oss-safeguard-20b$0.07$0.30
Mixtral 8x22B$0.50$1.00
Llama 3.1 405B$0.80$0.80
Llama 3.2 1B Instruct$0.03$0.20
GPT-5.2 Pro$21.00$168.00
Granite 4.0 Micro$0.02$0.11
GPT-5 Pro$15.00$120.00
DeepSeek V3.2 Exp$0.27$0.41
Hunyuan A13B Instruct$0.14$0.57
Llama 3.1 8B$0.04$0.04
Qwen3.6 27B$0.30$2.00
GPT-5.2$1.75$14.00
Nova Premier 1.0$2.50$12.50
DeepSeek V3.1 Terminus$0.27$1.00
GPT-4o-mini Search Preview$0.15$0.60
Kimi K2 0905$0.60$2.50
GPT-5 Chat$1.25$10.00
Qwen 2.5 72B$0.40$0.80
Sonar Reasoning Pro$2.00$8.00
GPT-5 Image Mini$2.50$2.00
DeepSeek R1$0.70$2.50
Qwen3 32B$0.08$0.28
Gemma 3 12B$0.05$0.15
Command R+$2.50$10.00
Grok 4.20$1.25$2.50
Qwen3 30B A3B Thinking 2507$0.20$2.40
R1 Distill Llama 70B$0.80$0.80
DeepSeek V3$0.26$1.03
Llama 3.2 3B Instruct$0.05$0.33
DeepSeek V4 Flash$0.14$0.28
MiniMax M2.1$0.30$1.20
GPT-5.1-Codex-Mini$0.25$2.00
GPT-5 Image$10.00$10.00
Hermes 4 405B$1.00$3.00
GPT-3.5 Turbo Instruct$1.50$2.00
DeepSeek V3.1$0.25$0.95
Gemini 2.5 Flash$0.30$2.50
Gemini 2.5 Flash Lite Preview 09-2025$0.10$0.40
Qwen3 14B$0.23$0.91
Llama 3.1 70B Instruct$0.40$0.40
GPT-4 Turbo$10.00$30.00
Mistral Large 3 2512$0.50$1.50
MiniMax-01$0.20$1.10
Qwen3 Coder Plus$0.65$3.25
Qwen3 Coder Flash$0.20$0.97
Mistral Medium 3.1$0.40$2.00
GPT-4.1 Mini$0.40$1.60
Nova Pro 1.0$0.80$3.20
R1$0.70$2.50
Jamba Large 1.7$2.00$8.00
ERNIE 4.5 VL 424B A47B$0.42$1.25
Llama 4 Maverick$0.20$0.80
Phi 4$0.07$0.14
Mistral Large 2$0.60$1.80
Nova Micro 1.0$0.04$0.14
GPT-5 Nano$0.05$0.40
Llama 3.2 11B Vision Instruct$0.34$0.34
Inflection 3 Pi$2.50$10.00
Gemini 2.0 Flash$0.10$0.40
Hunyuan Pro$0.60$1.20
Nano Banana Pro (Gemini 3 Pro Image Preview)$2.00$12.00
gpt-oss-120b$0.04$0.17
Qwen3 235B A22B$0.46$1.82
GPT-4.1$2.00$8.00
Command A$2.50$10.00
Hermes 3 70B Instruct$0.70$0.70
SHARE RESEARCH:
INTEGRATED RECOMMENDATION

Accelerate your workflow with Araho

Need help choosing the right model for your product? We build AI-native MVPs.

Get your MVP built in weeks with top-tier AI developers.