arrow_backBack to research feed
agentsPublished: July 20, 2026

Automated Discovery Has No Universally Superior Harness

By Akshat Gupta, Jermaine Lei, Alexander Lu, Gopala Anumanchipalli, Leshem Choshen

Research TL;DR

"No fixed discovery harness is universally superior for LLM automated search; an adaptive allocation method pruning weak runs based on early progress outperforms fixed and ensemble baselines."

Abstract

Autonomous discovery systems such as OpenEvolve and TTT-Discover are often used as general-purpose harnesses. However, in practice these are composite systems combining several design choices about archives, parent selection, exploration, and budget allocation into a single recipe. Because discovery runs are expensive and inherently stochastic, existing harnesses are often compared using too few independent trials to distinguish key methodological improvements from run-to-run variance. We systematically decompose OpenEvolve-style evolutionary search and the TTT-Discover search harness into its constituent components and systematically evaluate 30 budget-matched harnesses across 12 model-problem pairs using more than 3.1 million LLM rollouts and repeated-trial statistical analysis. Our results show that discovery harnesses have a generalization problem: No fixed harness is reliably superior across the evaluated model-problem pairs, and variants of OpenEvolve generally underperform simpler alternatives. Thus, harness choice is better viewed as a hyperparameter rather than as a universal recipe, and should be tailored to the specific problem and underlying model. We also find that early discovery progress predicts final performance, and use this property to present a budget-matched adaptive-allocation experiment that starts multiple harnesses, prunes weak partial runs, and reallocates compute to stronger survivors, outperforming both commitment to a randomly sampled fixed harness and a non-adaptive harness ensemble. Together, these results motivate shifting from fixed harness selection to online adaptation guided by early performance. We release all run pools including baseline null distributions for every model-problem pair as reusable statistical infrastructure against for future harness proposals.

Technical Analysis & Implementation

Technical Summary§

Core Problem§

Automated discovery systems (e.g., OpenEvolve, TTT-Discover) combine multiple design choices (archives, parent selection, exploration, budget allocation) into a single recipe. Evaluating these harnesses is expensive and stochastic, leading to unreliable comparisons. The paper demonstrates that no fixed harness dominates across model-problem pairs, motivating online adaptation.

Methodology§

  • Decomposition: OpenEvolve-style evolutionary search and TTT-Discover are broken into components: archive strategy (none, random, quality-diversity), parent selection (random, tournament, fitness-proportional), exploration (crossover, mutation, LLM-based), and budget allocation (fixed, adaptive).
  • Evaluation: 30 budget-matched harnesses × 12 model-problem pairs (e.g., GPT-3.5, LLaMA on math reasoning, code generation) using >3.1M LLM rollouts. Each run repeated with independent trials for statistical rigor.
  • Key Finding: No harness consistently bests others; variants of OpenEvolve underperform simpler alternatives (e.g., random search with mutation). Early discovery progress strongly correlates with final performance ($\rho \approx 0.85$).

Adaptive Allocation Algorithm§

Leveraging early progress, the authors propose an adaptive procedure: 1. Start $K$ harnesses with equal budget $B/K$ (partial budget). 2. After $T$ steps, evaluate each harness's performance and prune the bottom fraction. 3. Reallocate remaining budget to surviving harnesses proportionally to their performance.

Formally, let $s_i^{(t)}$ be a score (e.g., best objective so far) for harness $i$ at time $t$. After pruning threshold $\tau$, survivors $S = \{i : s_i^{(t)} > \tau\}$. Remaining budget $B_\text{rem}$ is allocated: $$b_i = \frac{\exp(\beta s_i^{(t)})}{\sum_{j \in S} \exp(\beta s_j^{(t)})} \cdot B_\text{rem}$$

Implementation Code Snippet§

import numpy as np

def adaptive_allocation(harness_scores, budget_remaining, beta=1.0, prune_fraction=0.3):
    """
    harness_scores: dict {harness_id: score}
    budget_remaining: float, total remaining budget
    Returns: dict {harness_id: allocated_budget}
    """
    n = len(harness_scores)
    k_prune = int(n * prune_fraction)
    sorted_harnesses = sorted(harness_scores.items(), key=lambda x: x[1], reverse=True)
    survivors = sorted_harnesses[:n - k_prune]  # keep top (1-prune_fraction)

    scores = np.array([s for _, s in survivors])
    weights = np.exp(beta * scores)
    weights /= weights.sum()

    allocations = {}
    for (hid, _), w in zip(survivors, weights):
        allocations[hid] = w * budget_remaining
    return allocations

Results§

Adaptive allocation outperforms both committing to a random fixed harness and a non-adaptive ensemble (equal budget split), achieving higher median performance and lower variance. The paper releases run pools with null distributions for future comparisons.

Implications§

Harness selection should be treated as a hyperparameter tailored to model and problem. Online adaptation using early signals is a promising direction for automated discovery.

Interactive SEO Tool

Interactive LLM Token & Cost Calculator

Estimate token usage and model pricing. Enter your prompt below to see how it is parsed into tokens and calculate the exact API cost for different providers.

Context Window1,050,000 tokens
Visual Tokenizer Chunks
Language models do not read text like humans. Instead, they process text in chunks called tokens. A token can be a single character, a syllable, a word, or even part of a word (like the "ing" in "walking"). On average, 1 token is equivalent to about 4 characters or 0.75 words of English text.
Estimated Token Count124

Cost Breakdown (USD)

Input Cost (Prompt):$0.003720
Output Cost (Generated):$0.022320
Total Est. Cost:$0.026040
Context Window Capacity0.0118%

API Pricing Comparison (per Million Tokens)

ModelInputOutput
GPT-5.5 Pro$30.00$180.00
GLM 4.7 Flash$0.06$0.40
o3 Mini$1.10$4.40
GPT-5.2-Codex$1.75$14.00
MiniMax M1$0.55$2.20
GPT-4o-mini$0.15$0.60
Gemini 2.5 Flash$0.30$2.50
WizardLM-2 8x22B$0.62$0.62
DeepSeek V3.1$0.25$0.95
GPT-4$30.00$60.00
Hermes 3 405B Instruct$1.00$1.00
o3 Pro$20.00$80.00
Mistral Medium 3.1$0.40$2.00
Claude Sonnet 5$2.00$10.00
Claude Sonnet 4.5$3.00$15.00
Qwen Plus 0728 (thinking)$0.26$0.78
Claude Opus 4$15.00$75.00
o4 Mini$1.10$4.40
GPT-4.1 Mini$0.40$1.60
Claude Opus 4.5$5.00$25.00
Claude Opus 4.7 (Fast)$30.00$150.00
Gemini 3.1 Flash Lite$0.25$1.50
o1$15.00$60.00
GLM 4.5V$0.60$1.80
GPT-5 Chat$1.25$10.00
GPT-4o (2024-11-20)$2.50$10.00
Mistral Large 2407$2.00$6.00
GPT Chat Latest$5.00$30.00
GPT-5 Nano$0.05$0.40
Claude Sonnet 4.6$3.00$15.00
gpt-oss-120b$0.04$0.17
Qwen2.5 7B Instruct$0.04$0.10
GPT-5.3-Codex$1.75$14.00
Gemini 3.1 Pro Preview$2.00$12.00
MoonshotAI Kimi Latest$3.00$15.00
Llama 3.2 3B Instruct$0.05$0.34
Google Gemini Flash Latest$1.50$9.00
Qwen3.5 Plus 2026-02-15$0.26$1.56
Claude Haiku 4.5$1.00$5.00
GPT-5 Mini$0.25$2.00
GPT-5.6 Luna Pro$1.00$6.00
GPT-5.6 Luna$1.00$6.00
Gemini 3.1 Flash$0.25$1.50
Qwen2.5 72B Instruct$0.36$0.40
Command R (08-2024)$0.15$0.60
Mistral Nemo$0.02$0.03
GPT-4o-mini (2024-07-18)$0.15$0.60
GPT-4o (2024-05-13)$5.00$15.00
Llama 4 Maverick$0.20$0.80
KAT-Coder-Air V2.5$0.15$0.60
KAT-Coder-Pro V2.5$0.74$2.96
Mixtral 8x22B Instruct$2.00$6.00
Llama 3.2 11B Vision$0.34$0.34
Mistral Large$2.00$6.00
Kimi K2.6$0.68$3.42
Llama 3 8B Instruct$0.14$0.14
GPT-3.5 Turbo (older v0613)$1.00$2.00
Llama 4 Scout$0.10$0.30
GPT-4 Turbo Preview$10.00$30.00
GLM 5$0.95$2.55
Claude 3 Haiku$0.25$1.25
Qwen3 30B A3B Instruct 2507$0.10$0.30
Gemini 2.0 Flash$0.10$0.40
GLM 4.5 Air$0.13$0.85
MiniMax M2.7$0.25$1.00
GPT-5.4 Nano$0.20$1.25
Qwen3 Coder 480B A35B$0.30$1.00
UI-TARS 7B$0.10$0.20
GPT-5.5$5.00$30.00
Mistral Small 3$0.10$0.30
Qwen3 Coder Next$0.11$0.80
MiniMax M2-her$0.30$1.20
Command R+$2.50$10.00
Mistral Small 4$0.15$0.60
GLM 5 Turbo$1.20$4.00
Qwen3 Max Thinking$0.78$3.90
Gemini 2.5 Pro Preview 06-05$1.25$10.00
GPT-4o$2.50$10.00
Gemini 2.5 Pro Preview 05-06$1.25$10.00
Claude Fable 5$10.00$50.00
Qwen3.7 Plus$0.32$1.28
Claude Opus 4.8$5.00$25.00
DeepSeek V3.1 Terminus$0.27$1.00
Qwen3 30B A3B Thinking 2507$0.13$1.56
Mistral Small 3.2 24B$0.10$0.30
Gemma 3n 4B$0.06$0.12
o3$2.00$8.00
MiniMax M3$0.30$1.20
Grok 4.20$1.25$2.50
Step 3.7 Flash$0.20$1.15
Qwen3.7 Max$1.48$4.42
Step 3.5 Flash$0.10$0.30
Kimi K2.5$0.57$2.85
gpt-oss-20b$0.03$0.13
Claude Opus 4.1$15.00$75.00
DeepSeek V3.2$0.27$0.40
Llama 3.1 8B$0.04$0.04
Nano Banana Pro (Gemini 3 Pro Image Preview)$2.00$12.00
GPT-5.1$1.25$10.00
Gemini 3.5 Flash$1.50$9.00
GLM 5V Turbo$1.20$4.00
Grok 4.20 Multi-Agent$1.25$2.50
GPT-5 Image Mini$2.50$2.00
Qwen3 8B$0.12$0.46
ERNIE 4.0$1.20$2.40
Qwen3.6 Flash$0.19$1.13
DeepSeek V4 Pro$0.43$0.87
Grok 4.20$1.25$2.50
Mistral Large 3$0.50$1.50
DeepSeek V3 0324$0.27$1.12
o1-pro$150.00$600.00
Llama 3.3 70B Instruct$0.13$0.40
Claude Opus 4.7$5.00$25.00
GPT Audio$2.50$10.00
GPT Audio Mini$0.60$2.40
Yi-Lightning$0.15$0.30
Qwen Plus 0728$0.26$0.78
Qwen3 235B A22B Thinking 2507$0.30$3.00
Mistral Large 2$0.60$1.80
GPT-5.4 Mini$0.75$4.50
Seed-2.0-Mini$0.10$0.40
Qwen3.5-Flash$0.07$0.26
GPT-5.1 Chat$1.25$10.00
Grok 4.3$1.25$2.50
Command R$0.15$0.60
GPT-5.1-Codex$1.25$10.00
Kimi K2 0711$0.57$2.30
Llama 3.1 405B$0.80$0.80
Seed-2.0-Lite$0.25$2.00
Mistral Small 3.1 24B$0.35$0.56
Qwen3.5 397B A17B$0.39$2.34
MiniMax M2.5$0.15$0.90
Solar Pro 3$0.15$0.60
Claude Opus 4.6$5.00$25.00
GPT-5.6 Sol Pro$5.00$30.00
GPT-5.6 Sol$5.00$30.00
GPT-5.1-Codex-Max$1.25$10.00
Ministral 3 14B 2512$0.20$0.20
Laguna XS 2.1$0.06$0.12
GPT-5$1.25$10.00
Nex-N2-Mini$0.03$0.10
Mistral Medium 3$0.40$2.00
Fugu Ultra$5.00$30.00
Nano Banana 2 (Gemini 3.1 Flash Image)$0.50$3.00
Nex-N2-Pro$0.25$1.00
Nemotron 3 Ultra$0.60$3.60
Hy3 preview$0.06$0.21
Gemini 2.5 Pro$1.25$10.00
GPT-4.1 Nano$0.10$0.40
Grok 4.5$2.00$6.00
Seed 1.6 Flash$0.07$0.30
Granite 4.1 8B$0.05$0.10
Gemini 3.1 Pro$2.00$12.00
Llama 4 Maverick$0.20$0.80
Laguna M.1$0.20$0.40
MiniMax M2$0.30$1.20
Google Gemini Pro Latest$2.00$12.00
Qwen3 VL 32B Instruct$0.10$0.42
Qwen3.6 35B A3B$0.14$1.00
GLM 5.1$0.97$3.04
Gemma 4 26B A4B$0.07$0.34
Nano Banana 2 (Gemini 3.1 Flash Image Preview)$0.50$3.00
Qwen3.5-35B-A3B$0.14$1.00
Ministral 3 8B 2512$0.15$0.15
o3 Deep Research$10.00$40.00
o4 Mini Deep Research$2.00$8.00
Qwen3.6 Max Preview$1.04$6.24
GPT-5.4 Image 2$8.00$15.00
Claude Opus Latest$5.00$25.00
Nova Lite 1.0$0.06$0.24
DeepSeek V4 Flash$0.09$0.19
MiMo-V2.5-Pro$0.43$0.87
Gemma 3 4B$0.05$0.10
GLM 4.7$0.40$1.75
Gemini 3 Flash Preview$0.50$3.00
Qwen 2.5-Coder 32B$0.35$0.70
Ministral 3 3B 2512$0.10$0.10
MiMo-V2.5$0.14$0.28
R1 0528$0.50$2.15
Gemma 4 31B$0.12$0.37
Llama Guard 4 12B$0.18$0.18
Qwen3 30B A3B$0.13$0.52
GPT-5.4 Pro$30.00$180.00
GPT-5.4$2.50$15.00
Nano Banana (Gemini 2.5 Flash Image)$0.30$2.50
Qwen3 VL 30B A3B Thinking$0.13$1.56
Doubao Pro$0.80$1.60
Qwen3.6 Plus$0.33$1.95
GLM 4.6$0.50$2.00
Qwen3 Max$0.78$3.90
Reka Edge$0.10$0.10
Nemotron 3 Super$0.08$0.45
Hunyuan A13B Instruct$0.14$0.57
Qwen3.5-27B$0.26$2.60
Qwen3.5-122B-A10B$0.26$2.08
Gemini 3.1 Pro Preview Custom Tools$2.00$12.00
Mixtral 8x22B$0.50$1.00
GPT-5.6 Terra Pro$2.50$15.00
GPT-5.6 Terra$2.50$15.00
GLM 5.2$0.94$2.94
Claude Opus 4.8 (Fast)$10.00$50.00
Qwen3.5-9B$0.10$0.15
Mercury 2$0.25$0.75
GPT-5.3 Chat$1.75$14.00
Gemini 3.1 Flash Lite Preview$0.25$1.50
Seed 1.6$0.25$2.00
GPT-4.1$2.00$8.00
Hy3$0.14$0.58
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)$0.25$1.50
Nemotron 3 Nano 30B A3B$0.05$0.20
Nano Banana Pro (Gemini 3 Pro Image)$2.00$12.00
GPT-5.2 Pro$21.00$168.00
GLM 4.6V$0.30$0.90
Qwen3 VL 30B A3B Instruct$0.13$0.52
Codestral 2508$0.30$0.90
Qwen3 Coder 30B A3B Instruct$0.07$0.27
Nova 2 Lite$0.30$2.50
Claude Fable Latest$10.00$50.00
KAT-Coder-Pro V2$0.30$1.20
Palmyra X5$0.60$6.00
GPT-5 Pro$15.00$120.00
Anthropic Claude Sonnet Latest$2.00$10.00
Qwen3.5 Plus 2026-04-20$0.30$1.80
DeepSeek V3.2 Exp$0.27$0.41
Kimi K2 0905$0.60$2.50
Hunyuan Pro$0.60$1.20
Grok Build 0.1$1.00$2.00
Mistral Medium 3.5$1.50$7.50
Anthropic Claude Haiku Latest$1.00$5.00
Qwen3.6 27B$0.45$2.70
Nova Premier 1.0$2.50$12.50
Sonar Pro Search$3.00$15.00
Granite 4.0 Micro$0.02$0.11
Qwen3 VL 8B Instruct$0.12$0.46
GLM 4.5$0.60$2.20
Gemini 2.5 Flash Lite$0.10$0.40
Qwen3 32B$0.08$0.28
Gemma 3 12B$0.05$0.15
DeepSeek R1$0.70$2.50
Qwen 2.5 72B$0.40$0.80
Kimi K2.7 Code$0.82$3.75
Lyria 3 Pro Preview$0.00$0.00
GPT-5.2$1.75$14.00
Devstral 2 2512$0.40$2.00
Qwen3 235B A22B Instruct 2507$0.09$0.55
Command A$2.50$10.00
GPT-4o-mini Search Preview$0.15$0.60
GPT-4o (2024-08-06)$2.50$10.00
Claude 3.5 Sonnet v2$3.00$15.00
MiniMax M2.1$0.30$1.20
GPT-5.2 Chat$1.75$14.00
Qwen-Plus$0.26$0.78
DeepSeek V3$0.20$0.80
GPT-5.1-Codex-Mini$0.25$2.00
Command R7B (12-2024)$0.04$0.15
Llama 3.3 70B Instruct$0.13$0.40
Llama 3.1 8B Instruct$0.05$0.08
GPT-5 Image$10.00$10.00
Qwen2.5 Coder 32B Instruct$0.66$1.00
Llama 3.1 70B Instruct$0.40$0.40
Kimi K3$3.00$15.00
Muse Spark 1.1$1.25$4.25
Kimi K2 Thinking$0.60$2.50
Voxtral Small 24B 2507$0.10$0.30
gpt-oss-safeguard-20b$0.07$0.30
Qwen3 VL 8B Thinking$0.12$1.36
Hermes 4 70B$0.13$0.40
Hermes 4 405B$1.00$3.00
Jamba Large 1.7$2.00$8.00
Morph V3 Large$0.90$1.90
Morph V3 Fast$0.80$1.20
Gemini 2.5 Flash Lite Preview 09-2025$0.10$0.40
Qwen2.5 VL 72B Instruct$0.80$1.00
R1 Distill Llama 70B$0.80$0.80
R1$0.70$2.50
Qwen3 Coder Plus$0.65$3.25
Qwen3 Coder Flash$0.20$0.97
MiniMax-01$0.20$1.10
Qwen3 Next 80B A3B Thinking$0.10$0.78
Qwen3 Next 80B A3B Instruct$0.10$0.78
GPT-4o Search Preview$2.50$10.00
Lyria 3 Clip Preview$0.00$0.00
Qwen3 VL 235B A22B Thinking$0.26$2.60
Qwen3 VL 235B A22B Instruct$0.21$1.90
Qwen3 14B$0.12$0.24
Qwen3 235B A22B$0.46$1.82
o4 Mini High$1.10$4.40
Reka Flash 3$0.10$0.20
Gemma 2 27B$0.65$0.65
Sonar Reasoning Pro$2.00$8.00
Sonar Pro$3.00$15.00
Mistral Large 3 2512$0.50$1.50
GPT-5 Codex$1.25$10.00
ERNIE 4.5 VL 424B A47B$0.42$1.25
Claude Sonnet 4$3.00$15.00
Sonar Deep Research$2.00$8.00
Sonar$1.00$1.00
Phi 4$0.07$0.14
GPT-4 Turbo$10.00$30.00
GPT-3.5 Turbo Instruct$1.50$2.00
Gemma 3 27B$0.10$0.30
Saba$0.20$0.60
o3 Mini High$1.10$4.40
Llama 3.2 11B Vision Instruct$0.34$0.34
Nova Micro 1.0$0.04$0.14
GPT-3.5 Turbo 16k$3.00$4.00
Nova Pro 1.0$0.80$3.20
GPT-3.5 Turbo$0.50$1.50
Inflection 3 Pi$2.50$10.00
Inflection 3 Productivity$2.50$10.00
Llama 3.2 1B Instruct$0.03$0.20
Hermes 3 70B Instruct$0.70$0.70
SHARE RESEARCH:
INTEGRATED RECOMMENDATION

Accelerate your workflow with Araho

Need help choosing the right model for your product? We build AI-native MVPs.

Get your MVP built in weeks with top-tier AI developers.