arrow_backBack to research feed
agentsPublished: July 23, 2026

OpenForgeRL: Train Harness-native Agents in Any Environment

By Xiao Yu, Baolin Peng, Ruize Xu, Hao Zou, Qianhui Wu, Hao Cheng, Wenlin Yao, Nikhil Singh, Zhou Yu, Jianfeng Gao

Research TL;DR

"OpenForgeRL decouples training and inference for harness-based agents via a lightweight proxy that records model calls for RL training, orchestrated with Kubernetes for scalable rollouts."

Abstract

Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard to train end-to-end with open infrastructure, whose SFT/RL stacks cannot natively express stateful, multi-process harness inference. To address this, we present OpenForgeRL, an open-source framework for training harness-based agents end-to-end in diverse environments. OpenForgeRL achieves this with a lightweight proxy that serves the harness's model calls while recording them as training data for a standard RL codebase (e.g., veRL), and a Kubernetes orchestrator that runs each rollout in its own remote container, together enabling training on any harness in any environment at scale. By decoupling training and inference, OpenForgeRL allows researchers to easily train, study, and improve agents directly in the real harnesses and environments they are deployed with. We validate our framework across diverse, complex harnesses and environments, spanning tool/claw-based agents and multimodal GUI browser- and computer-use agents. Using only hundreds to a few thousand tasks, OpenForgeClaw reaches 31.7 pass^3 and 55.9 pass@3 on ClawEval and 33.7 on QwenClawBench. OpenForgeGUI reaches 37.7 on OSWorld-Verified, 63.0 on Online-Mind2Web, and 72.3 on WebVoyager. Both outperform open baselines of similar size on nearly all benchmarks, and in the GUI setting match or surpass models several times larger. Beyond benchmarks, we analyze how harness choice (e.g., ZeroClaw, OpenClaw, Codex) and RL shape agent behavior. We find that some harnesses are substantially harder to learn than others, and that RL improves agentic reliability, such as self-verification, tool coverage, and completing multi-step plans, though critical abilities such as error recovery remain weak.

Technical Analysis & Implementation

OpenForgeRL: Training Harness-Native Agents End-to-End§

OpenForgeRL is an open-source framework that enables end-to-end reinforcement learning (RL) training of agents that rely on complex inference harnesses (e.g., Claude Code, Codex, OpenClaw). The key challenge is that these harnesses are stateful, multi-process systems that are incompatible with standard RL training pipelines. OpenForgeRL solves this with two core components: a lightweight proxy and a Kubernetes orchestrator.

Methodological Core§

The proxy acts as a man-in-the-middle: it intercepts all model calls made by the harness (typically to an LLM API) and records the input/output pairs. These pairs form the training data. During RL training, the proxy serves the harness's original model calls using the current policy (which is being trained), and simultaneously logs the trajectories for RL optimization. This effectively decouples the harness's inference logic from the model training.

The Kubernetes orchestrator manages distributed rollouts: each rollout (a complete episode of the agent interacting with the environment) is executed in its own remote container. This allows scaling to many parallel environments. The orchestrator collects all recorded trajectories and sends them to the RL trainer (e.g., veRL) for policy updates.

Training Objective§

OpenForgeRL uses Proximal Policy Optimization (PPO) with a standard clipped surrogate objective:

$$\mathcal{L}^{CLIP}(\theta) = \mathbb{E}_{t} \left[ \min\left( r_t(\theta) \hat{A}_t, \text{clip}(r_t(\theta), 1-\epsilon, 1+\epsilon) \hat{A}_t \right) \right]$$

where $r_t(\theta) = \frac{\pi_\theta(a_t|s_t)}{\pi_{\theta_{\text{old}}}(a_t|s_t)}$ is the probability ratio, $\hat{A}_t$ is the generalized advantage estimate, and $\epsilon$ is a hyperparameter (typically 0.2). The trajectory data recorded by the proxy is used to compute $\hat{A}_t$ and the policy gradient.

Implementation Details§

Proxy Server. The proxy is a simple HTTP server that wraps the harness's model endpoint. It implements a custom /v1/chat/completions endpoint that returns completions from the current policy model (loaded into GPU memory) while logging the request/response pairs.

Kubernetes Integration. The orchestrator defines a rollout as a Kubernetes pod that runs the harness and environment inside a container. It uses a job queue to manage parallel pods, each of which communicates with the proxy. After completion, the logged data is stored in a shared volume (e.g., S3) and the trainer reads from it.

Code Snippet (Proxy Server).

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from flask import Flask, request, jsonify

app = Flask(__name__)
model = AutoModelForCausalLM.from_pretrained("policy_model_path")
tokenizer = AutoTokenizer.from_pretrained("policy_model_path")
model.half().cuda()

def log_trajectory(conversation, response):
    # Append to shared storage (e.g., Redis queue)
    pass

@app.route("/v1/chat/completions", methods=["POST"])
def chat_completions():
    data = request.get_json()
    messages = data["messages"]
    input_ids = tokenizer.apply_chat_template(messages, return_tensors="pt").cuda()
    with torch.no_grad():
        output_ids = model.generate(input_ids, max_new_tokens=1024, do_sample=True)
    response = tokenizer.decode(output_ids[0][input_ids.shape[-1]:], skip_special_tokens=True)
    log_trajectory(messages, response)
    return jsonify({"choices": [{"message": {"content": response}}]})

if __name__ == "__main__":
    app.run(host="0.0.0.0", port=8000)

Orchestrator Sketch (simplified).

from kubernetes import client, config
import uuid

config.load_kubeconfig()
v1 = client.CoreV1Api()

for task in task_queue:
    pod_name = f"rollout-{uuid.uuid4().hex[:8]}"
    pod_manifest = {
        "apiVersion": "v1",
        "kind": "Pod",
        "metadata": {"name": pod_name},
        "spec": {
            "containers": [{
                "name": "harness",
                "image": "openforgerl/harness:latest",
                "env": [{"name": "PROXY_URL", "value": "http://proxy-server:8000"}],
                "volumeMounts": [{"mountPath": "/data", "name": "shared-storage"}]
            }],
            "volumes": [{"name": "shared-storage", "persistentVolumeClaim": {"claimName": "rl-data"}}],
            "restartPolicy": "Never"
        }
    }
    v1.create_namespaced_pod(body=pod_manifest, namespace="rl")
    # wait for completion and collect logs

Evaluation and Findings§

OpenForgeRL was validated on two agent categories: tool/claw-based agents (using ZeroClaw, OpenClaw) and GUI agents (browser/computer use). With only hundreds to a few thousand tasks, training significantly improved performance: e.g., OpenForgeClaw achieved 31.7% pass^3 and 55.9% pass@3 on ClawEval, and OpenForgeGUI reached 37.7 on OSWorld-Verified. The framework allowed analyzing how harness choice affects learnability and how RL improves reliability (self-verification, tool coverage) but not error recovery.

Interactive SEO Tool

Interactive LLM Token & Cost Calculator

Estimate token usage and model pricing. Enter your prompt below to see how it is parsed into tokens and calculate the exact API cost for different providers.

Context Window262,144 tokens
Visual Tokenizer Chunks
Language models do not read text like humans. Instead, they process text in chunks called tokens. A token can be a single character, a syllable, a word, or even part of a word (like the "ing" in "walking"). On average, 1 token is equivalent to about 4 characters or 0.75 words of English text.
Estimated Token Count124

Cost Breakdown (USD)

Input Cost (Prompt):$0.000015
Output Cost (Generated):$0.000056
Total Est. Cost:$0.000071
Context Window Capacity0.0473%

API Pricing Comparison (per Million Tokens)

ModelInputOutput
Qwen3 VL 8B Instruct$0.12$0.46
Llama 3.3 70B Instruct$0.13$0.40
Qwen2.5 Coder 32B Instruct$0.66$1.00
Kimi K3$3.00$15.00
MiniMax M1$0.55$2.20
Saba$0.20$0.60
o3 Mini High$1.10$4.40
Claude Opus Latest$5.00$25.00
Claude Sonnet 4.5$3.00$15.00
Qwen Plus 0728 (thinking)$0.26$0.78
Hermes 3 405B Instruct$1.00$1.00
GPT-4o-mini$0.15$0.60
Qwen3 Next 80B A3B Thinking$0.10$0.78
Claude Opus 5 (Fast)$10.00$50.00
Claude Opus 5$5.00$25.00
Claude Opus 4$15.00$75.00
Qwen2.5 VL 72B Instruct$0.80$1.00
Qwen-Plus$0.26$0.78
R1 Distill Llama 70B$0.80$0.80
GLM 4.5V$0.60$1.80
Claude Opus 4.7 (Fast)$30.00$150.00
Morph V3 Large$0.90$1.90
Command R7B (12-2024)$0.04$0.15
Inflection 3 Productivity$2.50$10.00
Llama 3.1 8B Instruct$0.05$0.08
Gemma 2 27B$0.65$0.65
GPT-3.5 Turbo Instruct$1.50$2.00
GLM 5.1$0.97$3.04
Claude Sonnet 4.6$3.00$15.00
GPT-4o (2024-11-20)$2.50$10.00
Kimi K2.6$0.65$2.72
Claude Opus 4.5$5.00$25.00
Gemma 4 26B A4B$0.12$0.35
Gemma 4 31B$0.14$0.40
Qwen3.6 Plus$0.33$1.95
Claude Sonnet 4$3.00$15.00
Gemini 2.5 Pro Preview 05-06$1.25$10.00
o1$15.00$60.00
Gemini 3.1 Pro Preview Custom Tools$2.00$12.00
o3$2.00$8.00
o4 Mini$1.10$4.40
Claude Haiku 4.5$1.00$5.00
MoonshotAI Kimi Latest$3.00$15.00
Google Gemini Flash Latest$1.50$7.50
GPT-4 Turbo Preview$10.00$30.00
Qwen2.5 7B Instruct$0.04$0.10
Laguna S 2.1$0.10$0.20
Gemini 3.5 Flash Lite$0.30$2.50
Muse Spark 1.1$1.25$4.25
GPT-5.6 Luna Pro$1.00$6.00
Reka Flash 3$0.10$0.20
GPT-5.6 Terra Pro$2.50$15.00
GPT-4o (2024-08-06)$2.50$10.00
GPT-5.5 Pro$30.00$180.00
GPT-5.6 Terra$2.50$15.00
GPT-5.6 Sol Pro$5.00$30.00
GPT-5.6 Sol$5.00$30.00
Grok 4.5$2.00$6.00
Gemini 3.6 Flash$1.50$7.50
Hy3$0.13$0.53
Laguna XS 2.1$0.06$0.12
Gemini 3.1 Flash$0.25$1.50
Command R (08-2024)$0.15$0.60
GPT-5.6 Luna$1.00$6.00
Claude Sonnet 5$2.00$10.00
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)$0.25$1.50
Nano Banana 2 (Gemini 3.1 Flash Image)$0.50$3.00
Kimi K2.7 Code$0.75$3.50
Claude Fable Latest$10.00$50.00
Nemotron 3 Ultra$0.60$3.60
MiniMax M3$0.30$1.20
GLM 4.6V$0.30$0.90
Qwen2.5 72B Instruct$0.36$0.40
Nex-N2-Mini$0.03$0.10
KAT-Coder-Air V2.5$0.15$0.60
Fugu Ultra$5.00$30.00
Ministral 3 8B 2512$0.15$0.15
Llama 4 Scout$0.10$0.30
GPT-4o (2024-05-13)$5.00$15.00
Step 3.7 Flash$0.20$1.15
Nova 2 Lite$0.30$2.50
KAT-Coder-Pro V2.5$0.74$2.96
Nano Banana Pro (Gemini 3 Pro Image)$2.00$12.00
Llama 4 Maverick$0.20$0.80
GLM 5.2$0.68$2.14
Mixtral 8x22B Instruct$2.00$6.00
Nex-N2-Pro$0.25$1.00
Claude Opus 4.8 (Fast)$10.00$50.00
Qwen3 VL 8B Thinking$0.12$1.36
Claude Opus 4.8$5.00$25.00
Gemini 3.5 Flash$1.50$9.00
Grok 4.3$1.25$2.50
Granite 4.1 8B$0.05$0.10
Llama 3 8B Instruct$0.14$0.14
Mistral Large$2.00$6.00
Qwen3 Next 80B A3B Instruct$0.10$1.10
Qwen3.7 Max$1.48$4.42
Grok Build 0.1$1.00$2.00
Sonar Pro$3.00$15.00
Claude 3.5 Sonnet v2$3.00$15.00
GPT-3.5 Turbo (older v0613)$1.00$2.00
Gemini 3.1 Flash Lite$0.25$1.50
GPT Chat Latest$5.00$30.00
Mistral Medium 3.5$1.50$7.50
Laguna M.1$0.20$0.40
Anthropic Claude Haiku Latest$1.00$5.00
Sonar Deep Research$2.00$8.00
Claude 3 Haiku$0.25$1.25
Google Gemini Pro Latest$2.00$12.00
Anthropic Claude Sonnet Latest$2.00$10.00
Sonar$1.00$1.00
Qwen3.5 Plus 2026-04-20$0.30$1.80
Qwen3.6 Flash$0.19$1.13
Qwen3 VL 235B A22B Thinking$0.26$2.60
Qwen3 VL 235B A22B Instruct$0.21$1.90
GPT-5 Codex$1.25$10.00
MiMo-V2.5-Pro$0.43$0.87
Qwen3 30B A3B Instruct 2507$0.05$0.19
MiMo-V2.5$0.14$0.28
GPT-5.4 Image 2$8.00$15.00
Grok 4.20 Multi-Agent$1.25$2.50
Grok 4.20$1.25$2.50
Lyria 3 Pro Preview$0.00$0.00
Lyria 3 Clip Preview$0.00$0.00
MiniMax M2.7$0.25$1.00
GPT-5.4 Nano$0.20$1.25
KAT-Coder-Pro V2$0.30$1.20
Reka Edge$0.10$0.10
GLM 5 Turbo$1.20$4.00
Nemotron 3 Super$0.09$0.40
Seed-2.0-Lite$0.25$2.00
GLM 4.5 Air$0.13$0.85
GPT-5.4 Pro$30.00$180.00
GPT-5.4$2.50$15.00
Gemini 3.1 Flash Lite Preview$0.25$1.50
Morph V3 Fast$0.80$1.20
GPT-5.3 Chat$1.75$14.00
Seed-2.0-Mini$0.10$0.40
Nano Banana 2 (Gemini 3.1 Flash Image Preview)$0.50$3.00
Qwen3.5-122B-A10B$0.26$2.08
Qwen3 Max Thinking$0.78$3.90
GPT-4o$2.50$10.00
Qwen3.5-35B-A3B$0.14$1.00
Qwen3.5-27B$0.20$1.56
GPT-3.5 Turbo 16k$3.00$4.00
Qwen3.5 Plus 2026-02-15$0.26$1.56
MiniMax M2-her$0.30$1.20
Gemini 2.5 Pro Preview 06-05$1.25$10.00
GPT-5.5$5.00$30.00
Mistral Small 4$0.15$0.60
GPT-5.3-Codex$1.75$14.00
Mistral Small 3$0.10$0.30
Qwen3.5 397B A17B$0.39$2.34
GPT-5.2-Codex$1.75$14.00
Gemini 3 Flash Preview$0.50$3.00
UI-TARS 7B$0.10$0.20
o4 Mini High$1.10$4.40
Claude Fable 5$10.00$50.00
Qwen3.7 Plus$0.32$1.28
GLM 5$0.95$2.55
Qwen3 Coder Next$0.11$0.80
GPT-3.5 Turbo$0.50$1.50
Gemini 3.1 Pro Preview$2.00$12.00
GPT-5.2 Chat$1.75$14.00
Mistral Small 3.2 24B$0.10$0.30
o3 Pro$20.00$80.00
Gemma 3n 4B$0.06$0.12
Devstral 2 2512$0.40$2.00
Mistral Large 2407$2.00$6.00
GPT-5.1-Codex-Max$1.25$10.00
gpt-oss-20b$0.03$0.14
Claude Opus 4.1$15.00$75.00
WizardLM-2 8x22B$0.62$0.62
Step 3.5 Flash$0.10$0.30
Kimi K2.5$0.57$2.85
DeepSeek V3.2$0.27$0.40
Mistral Large 3$0.50$1.50
GPT-5 Mini$0.25$2.00
Qwen3 8B$0.12$0.46
GLM 5V Turbo$1.20$4.00
GPT-4$30.00$60.00
GPT Audio Mini$0.60$2.40
DeepSeek V3 0324$0.27$1.12
Llama 3.3 70B Instruct$0.13$0.40
Yi-Lightning$0.15$0.30
o1-pro$150.00$600.00
DeepSeek V4 Pro$0.43$0.87
Ministral 3 14B 2512$0.20$0.20
Qwen Plus 0728$0.26$0.78
Mistral Nemo$0.02$0.03
GPT-4o-mini (2024-07-18)$0.15$0.60
GPT-5.4 Mini$0.75$4.50
Qwen3.5-Flash$0.07$0.26
MiniMax M2.5$0.15$0.90
GPT Audio$2.50$10.00
Voxtral Small 24B 2507$0.10$0.30
Command R$0.15$0.60
Kimi K2 0711$0.57$2.30
Mistral Medium 3$0.40$2.00
Mistral Small 3.1 24B$0.35$0.56
GPT-5.1 Chat$1.25$10.00
Solar Pro 3$0.15$0.60
GPT-5.1-Codex$1.25$10.00
GLM 4.7 Flash$0.06$0.40
GPT-5$1.25$10.00
Gemini 3.1 Pro$2.00$12.00
Claude Opus 4.7$5.00$25.00
Claude Opus 4.6$5.00$25.00
Gemini 2.5 Pro$1.25$10.00
Qwen3.6 35B A3B$0.14$1.00
GPT-4.1 Nano$0.10$0.40
Llama 3.2 11B Vision$0.34$0.34
Hy3 preview$0.06$0.21
Seed 1.6 Flash$0.07$0.30
Seed 1.6$0.25$2.00
o3 Deep Research$10.00$40.00
Nemotron 3 Nano 30B A3B$0.05$0.20
ERNIE 4.0$1.20$2.40
Nova Lite 1.0$0.06$0.24
MiniMax M2$0.26$1.02
Qwen3.6 Max Preview$1.04$6.24
Qwen3 VL 32B Instruct$0.10$0.42
o4 Mini Deep Research$2.00$8.00
R1 0528$0.50$2.15
Llama Guard 4 12B$0.18$0.18
Qwen 2.5-Coder 32B$0.35$0.70
Ministral 3 3B 2512$0.10$0.10
GLM 4.7$0.40$1.75
GPT-5.1$1.25$10.00
Gemma 3 4B$0.05$0.10
Doubao Pro$0.80$1.60
Kimi K2 Thinking$0.60$2.50
Sonar Pro Search$3.00$15.00
GLM 4.6$0.50$2.00
Qwen3 Max$0.78$3.90
Qwen3 30B A3B$0.12$0.50
Nano Banana (Gemini 2.5 Flash Image)$0.30$2.50
Qwen3 VL 30B A3B Thinking$0.13$1.56
Hunyuan A13B Instruct$0.14$0.57
Qwen3.5-9B$0.10$0.15
Mercury 2$0.25$0.75
Qwen3 VL 30B A3B Instruct$0.15$0.60
Mixtral 8x22B$0.50$1.00
Llama 3.1 405B$0.80$0.80
Palmyra X5$0.60$6.00
gpt-oss-safeguard-20b$0.07$0.30
o3 Mini$1.10$4.40
GPT-5.2 Pro$21.00$168.00
GPT-5 Pro$15.00$120.00
Llama 3.1 8B$0.04$0.04
Llama 3.2 1B Instruct$0.03$0.20
Granite 4.0 Micro$0.02$0.11
DeepSeek V3.2 Exp$0.27$0.41
GPT-5.2$1.75$14.00
Nova Premier 1.0$2.50$12.50
DeepSeek V3.1 Terminus$0.27$1.00
Kimi K2 0905$0.60$2.50
GPT-4o-mini Search Preview$0.15$0.60
Qwen3.6 27B$0.30$2.00
DeepSeek R1$0.70$2.50
Qwen 2.5 72B$0.40$0.80
GPT-5 Image Mini$2.50$2.00
GPT-5 Chat$1.25$10.00
Qwen3 32B$0.08$0.28
Gemma 3 12B$0.05$0.15
Qwen3 30B A3B Thinking 2507$0.13$1.56
Hermes 4 70B$0.13$0.40
Command R+$2.50$10.00
Grok 4.20$1.25$2.50
DeepSeek V3$0.20$0.80
Llama 3.2 3B Instruct$0.05$0.33
GPT-5.1-Codex-Mini$0.25$2.00
GPT-5 Image$10.00$10.00
Hermes 4 405B$1.00$3.00
DeepSeek V4 Flash$0.14$0.28
MiniMax M2.1$0.30$1.20
Gemini 2.5 Flash Lite Preview 09-2025$0.10$0.40
DeepSeek V3.1$0.25$0.95
Llama 3.1 70B Instruct$0.40$0.40
Gemini 2.5 Flash$0.30$2.50
Qwen3 14B$0.23$0.91
GPT-4 Turbo$10.00$30.00
GPT-4.1 Mini$0.40$1.60
R1$0.70$2.50
MiniMax-01$0.20$1.10
Mistral Large 3 2512$0.50$1.50
Qwen3 Coder Plus$0.65$3.25
Qwen3 Coder Flash$0.20$0.97
Mistral Medium 3.1$0.40$2.00
Nova Pro 1.0$0.80$3.20
Mistral Large 2$0.60$1.80
Phi 4$0.07$0.14
ERNIE 4.5 VL 424B A47B$0.42$1.25
Jamba Large 1.7$2.00$8.00
Llama 4 Maverick$0.20$0.80
Nova Micro 1.0$0.04$0.14
Llama 3.2 11B Vision Instruct$0.34$0.34
Inflection 3 Pi$2.50$10.00
Gemini 2.0 Flash$0.10$0.40
Hunyuan Pro$0.60$1.20
GPT-5 Nano$0.05$0.40
Nano Banana Pro (Gemini 3 Pro Image Preview)$2.00$12.00
gpt-oss-120b$0.04$0.17
Codestral 2508$0.30$0.90
Qwen3 Coder 30B A3B Instruct$0.07$0.27
GLM 4.5$0.60$2.20
Qwen3 235B A22B Thinking 2507$0.30$3.00
Qwen3 Coder 480B A35B$0.30$1.00
Gemini 2.5 Flash Lite$0.10$0.40
Qwen3 235B A22B Instruct 2507$0.09$0.55
Qwen3 235B A22B$0.46$1.82
GPT-4.1$2.00$8.00
Command A$2.50$10.00
GPT-4o Search Preview$2.50$10.00
Gemma 3 27B$0.08$0.45
Sonar Reasoning Pro$2.00$8.00
Hermes 3 70B Instruct$0.70$0.70
SHARE RESEARCH:
INTEGRATED RECOMMENDATION

Accelerate your workflow with Araho

Need help choosing the right model for your product? We build AI-native MVPs.

Get your MVP built in weeks with top-tier AI developers.