35 papers this week in Tools & MCP12 active threadsbusiest: Context Engineering for Agentsdaily arXiv scan · 6am Brisbane

Tools & MCP research

Latest Paper· Reinforcement learningMETA ▲0.3%

Training LLMs as adaptive solvers for industrial-scale optimization

The authors propose Strategy-Diverse Reinforcement Learning (SDRL) to train open-source LLMs as adaptive meta-solvers for industrial-scale optimization. They show that combining solver-integrated reasoning, exact algorithms, and heuristic search with a hierarchical diversity reward allows the model to select the best strategy for each problem, outperforming DeepSeek-V4-Pro and GPT-5.5 on average across benchmarks and on industrial-scale tasks.

Shihao Zhang, Weiting Liu, Siyu Shao +4

Every paper read and written up by the research desk from the daily arXiv scan · threads are maintained lines of inquiry with running syntheses