An open-source, verifiable execution environment & GRPO training harness for coding agents. Built on dual Local/Docker sandboxes with AST anti-cheat and real test verifiers.
In GRPO, the model samples a group of rollouts (e.g. G = 4) on the same issue.
Rewards are normalized directly across the group using A_i = (r_i - ฮผ) / (ฯ + 1e-5), eliminating the need for a separate Critic/Value network.
| Rollout ID | Agent Policy Trajectory | PyTest Verification | AST Syntax | Raw Reward (r_i) | Normalized Advantage (A_i) | Policy Update Direction |
|---|
Train any open-weights coding model (e.g., Qwen2.5-Coder-0.5B-Instruct or DeepSeek-R1-Distill) using Hugging Face's official trl.GRPOTrainer:
from trl import GRPOTrainer, GRPOConfig
from trl_bridge import OpenEnvCodingAdapter
# 1. Initialize verifiable coding environment (zero docker required)
env = OpenEnvCodingAdapter(task_instance=my_swebench_task, use_docker=False)
# 2. Train with Hugging Face GRPOTrainer
trainer = GRPOTrainer(
model="Qwen/Qwen2.5-Coder-0.5B-Instruct",
reward_funcs=[env.reward_fn.evaluate],
args=GRPOConfig(
output_dir="./qwen-coder-rl",
learning_rate=1e-5,
num_generations=4,
max_prompt_length=1024,
)
)
trainer.train()