24 papers this week in Theory12 active threadsbusiest: Context Engineering for Agentsdaily arXiv scan · 6am Brisbane

Theory research

Latest Paper· Theory

Polylogarithmic Nash regret achieved for any finite matrix game

The authors study Nash regret minimization in unknown finite matrix games with bandit payoff feedback and observed opponent actions. They develop Optimistic Payoff Balancing (OPB), which achieves instance-dependent O(log^2 T) Nash regret against arbitrary adaptive opponents, including games with nonunique equilibria. This resolves the open problem posed by Maiti et al. (2025), extending their polylogarithmic guarantee from 2x2 games to arbitrary finite dimensions.

Yuheng Zhang
Reinforcement learning

Three diffusion RL methods unified as divergence-constrained reward maximizers

A unified theoretical framework reveals that three prominent regression-based diffusion reinforcement learning methods—DiffusionNFT, FlowAWR, and RAM—are each solutions to a divergence-constrained reward-maximization problem, differing only in the convex generator defining the constraint. The authors identify approximations in prior work and propose DiffusionRFT, which uses an exact sparsemax projection, leading to faster convergence, more stable training, and top performance.

today
Theory

Unmodified posterior sampling achieves minimax regret in reinforcement learning

Goo and Hong prove that exact vanilla posterior sampling for reinforcement learning (PSRL) is minimax optimal in leading-order Bayesian regret for finite-horizon time-inhomogeneous tabular MDPs with unknown stochastic rewards, achieving the rate Õ(√(SAH³K)). They extend the same proof principle to linear-mixture MDPs, obtaining Õ(d√(H³K)). The key technical advance is a method to decouple the posterior-sampled model from its own continuation value using a common empirical transition reference and a Bellman-based variance argument.

today
Language models

Geometry-aware LoRA optimization improves convergence across supervised and reinforcement learning

Ding, Zazo and Hensman introduce Rotated Manifold Optimization, or RoM, an optimizer designed around a symmetry in low-rank adaptation: many pairs of LoRA factors represent exactly the same weight update. Across supervised fine-tuning and reinforcement learning experiments, RoM converged faster and reached lower held-out loss than the compared LoRA optimizers, while producing better or comparable downstream performance.

today

Every paper read and written up by the research desk from the daily arXiv scan · threads are maintained lines of inquiry with running syntheses