102 papers this week in Reinforcement learning12 active threadsbusiest: Context Engineering for Agentsdaily arXiv scan · 6am Brisbane

Reinforcement learning research

Latest Paper· Theory

Polylogarithmic Nash regret achieved for any finite matrix game

The authors study Nash regret minimization in unknown finite matrix games with bandit payoff feedback and observed opponent actions. They develop Optimistic Payoff Balancing (OPB), which achieves instance-dependent O(log^2 T) Nash regret against arbitrary adaptive opponents, including games with nonunique equilibria. This resolves the open problem posed by Maiti et al. (2025), extending their polylogarithmic guarantee from 2x2 games to arbitrary finite dimensions.

Yuheng Zhang
Reinforcement learning

Diffusion reward models capture multimodal preference that scalar scores miss

Wang et al. introduce DRM, a Diffusion Reward Model that recasts reward modeling as conditional density estimation over p(r|x,y) instead of point prediction. A lightweight Diffusion Transformer denoises Gaussian noise into a reward vector with no parametric family imposed, and N samples at inference form an empirical reward distribution that can be aggregated into a scalar, variance, or quantiles. Across five benchmarks, DRM matches or surpasses baselines under matched data and backbone, and recovers multimodal structure where conventional heads collapse to a single point.

today
Language models

Teacher alignment method closes capacity gap in reasoning distillation without discarding data

The authors propose Teacher Alignment to address the Gap Curse in reasoning distillation, where larger teacher models produce distributions too complex for smaller students to approximate. They develop TeacherGRPO, built on Group Relative Policy Optimization, which adapts the teacher via curriculum selective alignment and importance-adaptive length regularization. Experiments show consistent performance gains over existing baselines across diverse reasoning benchmarks.

today
Reinforcement learning

Three diffusion RL methods unified as divergence-constrained reward maximizers

A unified theoretical framework reveals that three prominent regression-based diffusion reinforcement learning methods—DiffusionNFT, FlowAWR, and RAM—are each solutions to a divergence-constrained reward-maximization problem, differing only in the convex generator defining the constraint. The authors identify approximations in prior work and propose DiffusionRFT, which uses an exact sparsemax projection, leading to faster convergence, more stable training, and top performance.

today
Theory

Unmodified posterior sampling achieves minimax regret in reinforcement learning

Goo and Hong prove that exact vanilla posterior sampling for reinforcement learning (PSRL) is minimax optimal in leading-order Bayesian regret for finite-horizon time-inhomogeneous tabular MDPs with unknown stochastic rewards, achieving the rate Õ(√(SAH³K)). They extend the same proof principle to linear-mixture MDPs, obtaining Õ(d√(H³K)). The key technical advance is a method to decouple the posterior-sampled model from its own continuation value using a common empirical transition reference and a Bellman-based variance argument.

today
Language models

Small models can steer stronger ones through shared reasoning traces

Liang and colleagues introduce Allspark, a framework for transferring reasoning behavior from a reinforcement-learned small model to stronger models through alternating text-based chains of thought. In experiments using Qwen models and ARC-AGI-2, the trained weak teacher improved some within-family and cross-family students, although the benefit depended on the inference configuration and required additional tokens.

today

Every paper read and written up by the research desk from the daily arXiv scan · threads are maintained lines of inquiry with running syntheses