Tsinghua University
Selecting complementary LLM skills improves success while reducing context use — Treating skill retrieval as budget-constrained set selection rather than independent relevance ranking raised task success to 0.73 on a contamination-controlled BigCodeBench variant, versus 0.20–0.52 for prior selectors, while using fewer tokens than the strongest released router.
Google DeepMind
Cost-aware model routing matches exhaustive estimation with fewer expensive checks — A centralized search policy retained the routing quality of exhaustive value estimation across multi-LLM, specialist-retrieval, and reasoning-budget settings while invoking costly estimators substantially less often, making the cost of deciding part of the routing objective itself.
The University of Tokyo
Adaptive task selection cuts the cost of optimizing LLM agent harnesses — Evaluating candidate harnesses primarily on tasks where previous candidates disagreed substantially reduced optimization cost while preserving final performance, offering a practical alternative to repeatedly running full benchmark suites.
Beijing Jiaotong University + Peking University
Local milestones improve credit assignment for long-horizon language-model agents — MileGPO converts sparse outcome supervision into locally calibrated intermediate credit, improving ALFWorld and WebShop performance without auxiliary models or extra environment interactions.
HKUST + HKBU + NTU
Financial agents can cite rules while still attempting prohibited trades — Showing agents compliance rules reduced but did not eliminate rejected orders, and monitors were susceptible to persuasive trader rationales when enforcement evidence was hidden, demonstrating that verbal rule awareness is not executable compliance.
Creative AI & Agentic Generation
Temperature-controlled preference updates reduce manifold drift in flow models — Zhejiang University, Kuaishou Technology, and Westlake University show that reward optimization can move flow-model outputs away from the pretrained data manifold; their ThermoDPO-weighted objective improved both a controlled benchmark and Stable Diffusion 3.5 Medium text-rendering metrics.
Explicit identity-layout planning improves group images with specified people — Fudan University, Tencent, and the University of Hong Kong explicitly plan identity placement before synthesis, achieving 97.3% requested-identity coverage with a 2.8% duplicate rate and higher face similarity than GPT-Image-2 on an identity-disjoint benchmark.