Training LLMs as adaptive solvers for industrial-scale optimization
The authors propose Strategy-Diverse Reinforcement Learning (SDRL) to train open-source LLMs as adaptive meta-solvers for industrial-scale optimization. They show that combining solver-integrated reasoning, exact algorithms, and heuristic search with a hierarchical diversity reward allows the model to select the best strategy for each problem, outperforming DeepSeek-V4-Pro and GPT-5.5 on average across benchmarks and on industrial-scale tasks.