๐Ÿค— Hugging Face Space ๐Ÿš€ Verifiable RL Sandbox โš–๏ธ MIT Open Source

OpenCodingEnv

An open-source, verifiable execution environment & GRPO training harness for coding agents. Built on dual Local/Docker sandboxes with AST anti-cheat and real test verifiers.

โš™๏ธ Sandbox Controls

Test Suite (PyTest)
--
AST Syntax Anti-Cheat
--
Total Verifiable Reward (r_i)
--

๐Ÿ’ป Live Sandbox Terminal Output

$ Sandbox initialized. Ready for execution. Click 'Run Sandbox Rollout' to begin.

๐Ÿ“„ Candidate Git Diff (Unified Patch)

(No modifications made)

๐Ÿ“Š Group Relative Policy Optimization (GRPO) Advantage Breakdown

In GRPO, the model samples a group of rollouts (e.g. G = 4) on the same issue. Rewards are normalized directly across the group using A_i = (r_i - ฮผ) / (ฯƒ + 1e-5), eliminating the need for a separate Critic/Value network.

Rollout ID Agent Policy Trajectory PyTest Verification AST Syntax Raw Reward (r_i) Normalized Advantage (A_i) Policy Update Direction

๐Ÿ”Œ 1-Line Hugging Face TRL & OpenEnv Integration

Train any open-weights coding model (e.g., Qwen2.5-Coder-0.5B-Instruct or DeepSeek-R1-Distill) using Hugging Face's official trl.GRPOTrainer:

from trl import GRPOTrainer, GRPOConfig
from trl_bridge import OpenEnvCodingAdapter

# 1. Initialize verifiable coding environment (zero docker required)
env = OpenEnvCodingAdapter(task_instance=my_swebench_task, use_docker=False)

# 2. Train with Hugging Face GRPOTrainer
trainer = GRPOTrainer(
    model="Qwen/Qwen2.5-Coder-0.5B-Instruct",
    reward_funcs=[env.reward_fn.evaluate],
    args=GRPOConfig(
        output_dir="./qwen-coder-rl",
        learning_rate=1e-5,
        num_generations=4,
        max_prompt_length=1024,
    )
)

trainer.train()