arrow_backBack to research feed
agentsPublished: July 21, 2026

Agents in the Wild: Where Research Meets Deployment

By Grace Hui Yang, Pranav N. Venkit, Hooman Sedghamiz, Enrico Santus, Victor Dibia, Ioana Baldini

Research TL;DR

"Practical design patterns for deploying reliable LLM agents: verification pipelines, fallback mechanisms, and human-in-the-loop supervision, illustrated via pharmaceutical and financial case studies."

Abstract

Agentic systems large language model (LLM) based architectures capable of reasoning, planning, acting, and coordinating with tools and other agents are rapidly transitioning from research prototypes to production scale deployments across domains such as software engineering, scientific discovery, and finance. While academic work has emphasized benchmarks and algorithmic innovation, deployment raises new challenges around robustness, safety, and reliability. This tutorial brings together researchers and practitioners to explore advances in reasoning and planning, multi agent coordination, and evaluation, highlighting open challenges arising from deployment experience. Through applied case studies in pharmaceutical discovery and financial systems, we analyze common design patterns that make agentic systems successful, and discuss practical mitigation strategies for failure modes, such as verification pipelines, fallback mechanisms, and human in the loop supervision. Attendees will gain a comprehensive view of the field along with concrete design patterns, evaluation checklists, and templates for safe and reliable deployment across industries.

Technical Analysis & Implementation

Overview§

This tutorial paper synthesizes insights from deploying LLM-based agentic systems in production across software engineering, scientific discovery, and finance. It emphasizes that while academic research focuses on benchmarks and algorithmic innovations, real-world deployment demands robustness, safety, and reliability. The paper provides concrete design patterns, evaluation checklists, and templates for safe deployment.

Core Architecture of Agentic Systems§

Agentic systems are composed of an LLM core, planning/reasoning module, tool-use interface, and multi-agent coordination layer. A typical workflow involves: 1. Task decomposition: breaking a complex goal into sub-tasks. 2. Tool invocation: calling external APIs (e.g., code execution, databases, web search). 3. Multi-agent coordination: agents specializing in different functions communicate via a shared memory or message-passing. 4. Verification and fallback: output validation, retry logic, and human escalation.

Mathematical Formulation§

Let the LLM be represented as a function $f_{\theta}(x)$ that maps input $x$ (including prompts, context, tool outputs) to actions $a$. The agent's policy can be modeled as a Markov decision process: $$ \pi(a_t | s_t) = \text{softmax}\left( \frac{Q(s_t, a_t)}{\tau} \right) $$ where $s_t$ is the current state (context + history), $a_t$ is the action (e.g., generate text, call tool), and $\tau$ is temperature. Planning uses tree search with a value function $V(s)$ estimating future reward.

Key Design Patterns for Reliable Deployment§

1. Verification Pipelines: After each agent action, a separate verifier LLM or rule-based system checks correctness. If failure detected, action is retried or alternative path taken. 2. Fallback Mechanisms: When the agent fails repeatedly, escalate to a human operator. This is modeled as a threshold on confidence scores or number of retries. 3. Human-in-the-loop Supervision: For high-stakes decisions (e.g., financial transactions), require human approval before execution.

Case Study: Pharmaceutical Discovery§

Agents coordinate to perform molecular docking, toxicity prediction, and synthesis planning. A typical workflow:

  • Agent 1 (retrieval): searches chemical databases for candidate molecules.
  • Agent 2 (simulation): runs docking simulations using external tool (e.g., AutoDock).
  • Agent 3 (analysis): analyzes results and filters candidates.
  • Verification agent checks simulation outputs for convergence and validity.

Code Example: Simple Tool-Use Agent§

import openai
import json

def call_llm(prompt):
    response = openai.ChatCompletion.create(
        model="gpt-4",
        messages=[{"role": "system", "content": "You are a helpful assistant."}, 
                  {"role": "user", "content": prompt}]
    )
    return response["choices"][0]["message"]["content"]

def execute_tool(tool_name, args):
    if tool_name == "search_web":
        # perform web search
        return search_results
    elif tool_name == "run_code":
        # execute Python code
        return exec(args["code"])
    # ...

def agent_loop(goal):
    state = {"goal": goal, "history": []}
    while not goal_achieved(state):
        action = call_llm(f"Given state: {json.dumps(state)}, what is the next action? "
                         f"Options: think, search_web, run_code, final_answer")
        if action == "think":
            state["history"].append({"role": "assistant", "content": call_llm(state["goal"])})
        elif action in ["search_web", "run_code"]:
            args = call_llm(f"Provide arguments for {action}")
            result = execute_tool(action, args)
            state["history"].append({"role": "tool", "content": result})
        elif action == "final_answer":
            return call_llm(f"Based on history: {state['history']}, provide final answer")
        # verification
        if not verify(state["history"][-1]):
            state["history"].append({"role": "system", "content": "Verification failed, retry."})

Evaluation Checklist§

  • Robustness: test against adversarial prompts, distribution shift.
  • Safety: enforce output constraints (e.g., no harmful instructions).
  • Reliability: measure task success rate, latency, error rates.
  • Human oversight: log all actions, allow manual override.

Conclusion§

The paper provides a bridge between academic research and production deployment, offering concrete patterns for building trustworthy agentic systems. The key takeaway is that successful deployment requires systematic verification, fallback strategies, and human-in-the-loop mechanisms.

Interactive SEO Tool

Interactive LLM Token & Cost Calculator

Estimate token usage and model pricing. Enter your prompt below to see how it is parsed into tokens and calculate the exact API cost for different providers.

Context Window163,840 tokens
Visual Tokenizer Chunks
Language models do not read text like humans. Instead, they process text in chunks called tokens. A token can be a single character, a syllable, a word, or even part of a word (like the "ing" in "walking"). On average, 1 token is equivalent to about 4 characters or 0.75 words of English text.
Estimated Token Count124

Cost Breakdown (USD)

Input Cost (Prompt):$0.000031
Output Cost (Generated):$0.000118
Total Est. Cost:$0.000149
Context Window Capacity0.0757%

API Pricing Comparison (per Million Tokens)

ModelInputOutput
DeepSeek V3.1$0.25$0.95
Hermes 3 405B Instruct$1.00$1.00
GPT-4o-mini$0.15$0.60
Kimi K3$3.00$15.00
GLM 4.7 Flash$0.06$0.40
Mistral Medium 3.1$0.40$2.00
Muse Spark 1.1$1.25$4.25
Nano Banana 2 (Gemini 3.1 Flash Image)$0.50$3.00
MiniMax M1$0.55$2.20
GLM 5.2$0.79$2.47
Nemotron 3 Ultra$0.60$3.60
Gemini 2.5 Flash$0.30$2.50
Claude Opus 4.8 (Fast)$10.00$50.00
GPT-5.2-Codex$1.75$14.00
Claude Sonnet 4.5$3.00$15.00
Laguna S 2.1$0.10$0.20
Sonar Reasoning Pro$2.00$8.00
Gemini 3.6 Flash$1.50$7.50
Gemini 3.5 Flash-Lite$0.30$2.50
Claude Sonnet 5$2.00$10.00
Qwen Plus 0728 (thinking)$0.26$0.78
Claude Opus 4$15.00$75.00
o4 Mini$1.10$4.40
GPT-4.1 Mini$0.40$1.60
Reka Flash 3$0.10$0.20
Claude Opus 4.7 (Fast)$30.00$150.00
Gemini 3.1 Flash Lite$0.25$1.50
GPT Chat Latest$5.00$30.00
o1$15.00$60.00
GPT-4o (2024-11-20)$2.50$10.00
Claude Sonnet 4.6$3.00$15.00
Claude Opus 4.5$5.00$25.00
GLM 4.5V$0.60$1.80
GPT-5 Chat$1.25$10.00
Mistral Large 2407$2.00$6.00
GPT-5 Nano$0.05$0.40
gpt-oss-120b$0.04$0.17
MoonshotAI Kimi Latest$3.00$15.00
Google Gemini Flash Latest$1.50$7.50
GPT-5.5 Pro$30.00$180.00
Grok 4.20 Multi-Agent$1.25$2.50
Claude Haiku 4.5$1.00$5.00
Qwen2.5 7B Instruct$0.04$0.10
Llama 3.2 3B Instruct$0.05$0.34
Qwen3.5-27B$0.26$2.60
Qwen3.5-122B-A10B$0.26$2.08
Gemini 3.1 Pro Preview Custom Tools$2.00$12.00
GPT-5.3-Codex$1.75$14.00
GPT-4 Turbo Preview$10.00$30.00
GPT-3.5 Turbo Instruct$1.50$2.00
Gemini 3.1 Flash$0.25$1.50
GPT-5.6 Luna Pro$1.00$6.00
GPT-5.6 Luna$1.00$6.00
Gemini 3.1 Pro Preview$2.00$12.00
Qwen3.5 Plus 2026-02-15$0.26$1.56
Qwen2.5 Coder 32B Instruct$0.66$1.00
Qwen2.5 72B Instruct$0.36$0.40
Command R (08-2024)$0.15$0.60
GPT-4o (2024-08-06)$2.50$10.00
Mistral Nemo$0.02$0.03
GPT-4o-mini (2024-07-18)$0.15$0.60
KAT-Coder-Air V2.5$0.15$0.60
KAT-Coder-Pro V2.5$0.74$2.96
Llama 4 Maverick$0.20$0.80
GPT-4o (2024-05-13)$5.00$15.00
Kimi K2.6$0.68$3.42
GLM 4.6V$0.30$0.90
Nova 2 Lite$0.30$2.50
Qwen3 VL 8B Thinking$0.12$1.36
Llama 3.2 11B Vision$0.34$0.34
Llama 4 Scout$0.10$0.30
Llama 3 8B Instruct$0.14$0.14
Mixtral 8x22B Instruct$2.00$6.00
Mistral Large$2.00$6.00
GPT-3.5 Turbo (older v0613)$1.00$2.00
MiniMax M2.7$0.25$1.00
GPT-5.4 Nano$0.20$1.25
Sonar Pro$3.00$15.00
Sonar Deep Research$2.00$8.00
GLM 5$0.95$2.55
Qwen3 30B A3B Instruct 2507$0.10$0.30
Qwen3 Next 80B A3B Instruct$0.10$1.10
GLM 4.5 Air$0.13$0.85
Qwen3 Coder 480B A35B$0.30$1.00
Sonar$1.00$1.00
Claude 3 Haiku$0.25$1.25
Claude 3.5 Sonnet v2$3.00$15.00
MiniMax M2-her$0.30$1.20
GPT-5.5$5.00$30.00
GPT-3.5 Turbo 16k$3.00$4.00
Mistral Small 4$0.15$0.60
UI-TARS 7B$0.10$0.20
GLM 5 Turbo$1.20$4.00
Mistral Small 3$0.10$0.30
Qwen3 Max Thinking$0.78$3.90
Qwen3 Coder Next$0.11$0.80
Morph V3 Fast$0.80$1.20
Gemini 2.5 Pro Preview 06-05$1.25$10.00
GPT-4o$2.50$10.00
Claude Fable 5$10.00$50.00
Gemma 3n 4B$0.06$0.12
Qwen3.7 Plus$0.32$1.28
Gemini 2.5 Pro Preview 05-06$1.25$10.00
Claude Opus 4.8$5.00$25.00
DeepSeek V3.1 Terminus$0.27$1.00
o4 Mini High$1.10$4.40
Qwen3 30B A3B Thinking 2507$0.13$1.56
Mistral Small 3.2 24B$0.10$0.30
GPT-3.5 Turbo$0.50$1.50
o3 Pro$20.00$80.00
MiniMax M3$0.30$1.20
Step 3.7 Flash$0.20$1.15
Qwen3.7 Max$1.48$4.42
Step 3.5 Flash$0.10$0.30
Kimi K2.5$0.57$2.85
gpt-oss-20b$0.03$0.13
Claude Opus 4.1$15.00$75.00
o3$2.00$8.00
Llama 3.1 8B Instruct$0.05$0.08
WizardLM-2 8x22B$0.62$0.62
Gemini 3.5 Flash$1.50$9.00
GLM 5V Turbo$1.20$4.00
DeepSeek V3.2$0.27$0.40
Nano Banana Pro (Gemini 3 Pro Image Preview)$2.00$12.00
GPT-5.1$1.25$10.00
GPT-5 Image Mini$2.50$2.00
Qwen3 8B$0.12$0.46
GPT-4$30.00$60.00
Qwen3.6 Flash$0.19$1.13
DeepSeek V4 Pro$0.43$0.87
Mistral Large 3$0.50$1.50
Grok 4.20$1.25$2.50
GPT-5 Mini$0.25$2.00
DeepSeek V3 0324$0.27$1.12
o1-pro$150.00$600.00
Llama 3.3 70B Instruct$0.13$0.40
Qwen Plus 0728$0.26$0.78
Qwen3 235B A22B Thinking 2507$0.30$3.00
Claude Opus 4.7$5.00$25.00
GPT-5.4 Mini$0.75$4.50
Seed-2.0-Mini$0.10$0.40
Qwen3.5-Flash$0.07$0.26
GPT Audio$2.50$10.00
Yi-Lightning$0.15$0.30
GPT Audio Mini$0.60$2.40
Grok 4.3$1.25$2.50
GPT-5.1 Chat$1.25$10.00
Seed-2.0-Lite$0.25$2.00
Qwen3.5 397B A17B$0.39$2.34
MiniMax M2.5$0.15$0.90
GPT-5.1-Codex$1.25$10.00
Kimi K2 0711$0.57$2.30
Command R$0.15$0.60
Solar Pro 3$0.15$0.60
Mistral Small 3.1 24B$0.35$0.56
Mistral Medium 3$0.40$2.00
GPT-5.6 Sol Pro$5.00$30.00
Claude Opus 4.6$5.00$25.00
GPT-5.1-Codex-Max$1.25$10.00
GPT-5.6 Sol$5.00$30.00
Laguna XS 2.1$0.06$0.12
Nex-N2-Mini$0.03$0.10
Ministral 3 14B 2512$0.20$0.20
Fugu Ultra$5.00$30.00
Nex-N2-Pro$0.25$1.00
GPT-5$1.25$10.00
Gemini 2.5 Pro$1.25$10.00
Grok 4.5$2.00$6.00
Gemini 3.1 Pro$2.00$12.00
GPT-4.1 Nano$0.10$0.40
Granite 4.1 8B$0.05$0.10
Laguna M.1$0.20$0.40
Hy3 preview$0.06$0.21
Google Gemini Pro Latest$2.00$12.00
Seed 1.6 Flash$0.07$0.30
MiniMax M2$0.30$1.20
Llama 4 Maverick$0.20$0.80
Qwen3.6 35B A3B$0.14$1.00
Qwen3 VL 32B Instruct$0.10$0.42
Qwen3.6 Max Preview$1.04$6.24
GPT-5.4 Image 2$8.00$15.00
Claude Opus Latest$5.00$25.00
GLM 5.1$0.97$3.04
o3 Deep Research$10.00$40.00
o4 Mini Deep Research$2.00$8.00
Nova Lite 1.0$0.06$0.24
Gemma 4 26B A4B$0.07$0.34
Nano Banana 2 (Gemini 3.1 Flash Image Preview)$0.50$3.00
Qwen3.5-35B-A3B$0.14$1.00
Ministral 3 8B 2512$0.15$0.15
R1 0528$0.50$2.15
Qwen 2.5-Coder 32B$0.35$0.70
MiMo-V2.5-Pro$0.43$0.87
Llama Guard 4 12B$0.18$0.18
MiMo-V2.5$0.14$0.28
Qwen3 30B A3B$0.13$0.52
Gemma 4 31B$0.12$0.37
GLM 4.7$0.40$1.75
Gemini 3 Flash Preview$0.50$3.00
Ministral 3 3B 2512$0.10$0.10
Gemma 3 4B$0.05$0.10
GLM 4.6$0.50$2.00
Qwen3 Max$0.78$3.90
Qwen3.6 Plus$0.33$1.95
Reka Edge$0.10$0.10
Nemotron 3 Super$0.08$0.45
GPT-5.4 Pro$30.00$180.00
GPT-5.4$2.50$15.00
Nano Banana (Gemini 2.5 Flash Image)$0.30$2.50
Doubao Pro$0.80$1.60
Qwen3 VL 30B A3B Thinking$0.13$1.56
Mixtral 8x22B$0.50$1.00
GPT-5.6 Terra Pro$2.50$15.00
Qwen3.5-9B$0.10$0.15
GPT-5.6 Terra$2.50$15.00
Mercury 2$0.25$0.75
GPT-5.3 Chat$1.75$14.00
Gemini 3.1 Flash Lite Preview$0.25$1.50
Hunyuan A13B Instruct$0.14$0.57
o3 Mini$1.10$4.40
GPT-5.2 Pro$21.00$168.00
Qwen3 VL 30B A3B Instruct$0.13$0.52
Codestral 2508$0.30$0.90
Hy3$0.14$0.58
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)$0.25$1.50
Nano Banana Pro (Gemini 3 Pro Image)$2.00$12.00
KAT-Coder-Pro V2$0.30$1.20
Palmyra X5$0.60$6.00
Claude Fable Latest$10.00$50.00
Qwen3 Coder 30B A3B Instruct$0.07$0.27
Seed 1.6$0.25$2.00
Nemotron 3 Nano 30B A3B$0.05$0.20
GPT-4.1$2.00$8.00
Granite 4.0 Micro$0.02$0.11
Grok Build 0.1$1.00$2.00
Qwen3 VL 8B Instruct$0.12$0.46
Kimi K2 0905$0.60$2.50
Mistral Medium 3.5$1.50$7.50
Anthropic Claude Haiku Latest$1.00$5.00
GPT-5 Pro$15.00$120.00
Anthropic Claude Sonnet Latest$2.00$10.00
DeepSeek V3.2 Exp$0.27$0.41
Qwen3.5 Plus 2026-04-20$0.30$1.80
Qwen3.6 27B$0.45$2.70
Nova Premier 1.0$2.50$12.50
Sonar Pro Search$3.00$15.00
DeepSeek R1$0.70$2.50
Qwen 2.5 72B$0.40$0.80
GLM 4.5$0.60$2.20
Kimi K2.7 Code$0.82$3.75
Lyria 3 Pro Preview$0.00$0.00
Gemma 3 12B$0.05$0.15
Command A$2.50$10.00
Gemini 2.5 Flash Lite$0.10$0.40
GPT-4o-mini Search Preview$0.15$0.60
Qwen3 235B A22B Instruct 2507$0.09$0.55
Qwen3 32B$0.08$0.28
GPT-5.2$1.75$14.00
Devstral 2 2512$0.40$2.00
MiniMax M2.1$0.30$1.20
GPT-5.2 Chat$1.75$14.00
Command R+$2.50$10.00
GPT-5 Image$10.00$10.00
Qwen-Plus$0.26$0.78
Grok 4.20$1.25$2.50
GPT-5.1-Codex-Mini$0.25$2.00
DeepSeek V3$0.20$0.80
Command R7B (12-2024)$0.04$0.15
Llama 3.3 70B Instruct$0.13$0.40
Hermes 4 70B$0.13$0.40
DeepSeek V4 Flash$0.10$0.20
Llama 3.1 70B Instruct$0.40$0.40
GPT-4 Turbo$10.00$30.00
Kimi K2 Thinking$0.60$2.50
Hermes 4 405B$1.00$3.00
Jamba Large 1.7$2.00$8.00
Morph V3 Large$0.90$1.90
ERNIE 4.0$1.20$2.40
Voxtral Small 24B 2507$0.10$0.30
gpt-oss-safeguard-20b$0.07$0.30
Gemini 2.5 Flash Lite Preview 09-2025$0.10$0.40
Lyria 3 Clip Preview$0.00$0.00
Qwen3 VL 235B A22B Thinking$0.26$2.60
Qwen3 VL 235B A22B Instruct$0.21$1.90
GPT-4o Search Preview$2.50$10.00
Qwen2.5 VL 72B Instruct$0.80$1.00
R1 Distill Llama 70B$0.80$0.80
R1$0.70$2.50
Qwen3 Coder Plus$0.65$3.25
MiniMax-01$0.20$1.10
Qwen3 Coder Flash$0.20$0.97
Qwen3 Next 80B A3B Thinking$0.10$0.78
Phi 4$0.07$0.14
Mistral Large 3 2512$0.50$1.50
GPT-5 Codex$1.25$10.00
Claude Sonnet 4$3.00$15.00
Qwen3 14B$0.12$0.24
ERNIE 4.5 VL 424B A47B$0.42$1.25
Qwen3 235B A22B$0.46$1.82
Gemma 2 27B$0.65$0.65
Mistral Large 2$0.60$1.80
Llama 3.1 405B$0.80$0.80
Llama 3.1 8B$0.04$0.04
Gemma 3 27B$0.10$0.30
Saba$0.20$0.60
o3 Mini High$1.10$4.40
Nova Micro 1.0$0.04$0.14
Hermes 3 70B Instruct$0.70$0.70
Nova Pro 1.0$0.80$3.20
Inflection 3 Pi$2.50$10.00
Llama 3.2 11B Vision Instruct$0.34$0.34
Inflection 3 Productivity$2.50$10.00
Llama 3.2 1B Instruct$0.03$0.20
Gemini 2.0 Flash$0.10$0.40
Hunyuan Pro$0.60$1.20
SHARE RESEARCH:
INTEGRATED RECOMMENDATION

Accelerate your workflow with Araho

Need help choosing the right model for your product? We build AI-native MVPs.

Get your MVP built in weeks with top-tier AI developers.