Evolving terminal environments keeps training tasks challenging as agents improve
Fan et al. developed an environment-evolution method that modifies terminal tasks along three difficulty-related directions and schedules successive generations during training. Applied to two Qwen models using long-horizon reinforcement learning, the method improved Terminal-Bench 2.1 performance by 14.4 and 18.0 percentage points.