Enumerating Tool Choices Beats Sampling Them in Genomic Reasoning
Liu et al. replace sampled reinforcement-learning updates with exact optimization over the complete, enumerable set of genomic tool combinations. Across five frozen reasoners and three benchmarks, their method outperformed GRPO in all 15 experimental settings, by 6.75 percentage points on average.
11 Sept 2026