Agents
Tool-call pipelines often alter what language-model agents actually execute (Independent), production wrappers and interpreters frequently change the meaning of agent-generated shell calls, showing that trajectory analysis can blame models for failures introduced downstream.
AI agent teams coordinate worse than a single shared agent (Stanford, Anthropic), across five frontier models and 77 shared-resource scenarios, multi-agent teams lost to one coordinating agent in every environment, even with communication enabled.
Shared agent memory can leak information across enterprise users (Workday), ordinary similarity queries retrieved other users' semantically related memories in six experiments, while ownership filtering was the only tested mitigation that restored contamination to baseline.
Language models
Human preferences contain signals that rubrics and verifiers miss (Stanford, Toronto), analysis of 317 million judgments found persistent articulability gaps across seven domains, with heavier emphasis on explicit rubrics sometimes reducing alignment with human preferences.
Synthetic warm-up training scales, but does not teach language grammar (Sheffield, MBZUAI), synthetic non-language pretraining saved at least 21 billion tokens for 3B models, but the gains did not support the proposed explanation of a transferable grammatical prior.
Language models can suddenly and repeatedly switch between pattern-matching and genuine generalization during pre-training (UC Berkeley, DeepMind), checkpoint-level behavioral tests reveal abrupt mode switches within tens of billions of tokens that persist under metric changes and checkpoint averaging.
Reinforcement learning
RL post-training improves agent consistency but narrows solution coverage (Meta, UW Madison), across 14 base and post-trained model pairs, RL generally improved pass@1 while often reducing pass@K, trading reliable default behavior for less diverse solution coverage.
Changing reward requirements teaches vision-language models to actually use images (CAS), visual benchmark gains appeared even without visual training data until rewards made image inspection necessary, which then improved grounding and transfer.
Testing transfer helps language-model tutors teach instead of simply telling (ETH Zurich), rewarding performance on a new problem improved student transfer, but explicit constraints were still required to prevent tutors from simply revealing answers.
Vision
Vision models weigh surface objects against abstract relations in separate circuits (University of Tokyo, DeepMind), stronger models increasingly favored relational matches, with open-model analysis locating competition between early object-focused and late relation-focused processing.
Fused DINOv3 features enable fast, faithful real-world image super-resolution (UC Berkeley, Alibaba), fusing 23 frozen DINOv3-L layers enabled competitive restoration in one 37 millisecond pass with fewer hallucinations than pixel-space or VAE-latent alternatives.
Grounding foundation model surpasses larger vision-language models on precise perception and robotics (Independent), the 4B GroundingPI model averaged 73.68% across 34 benchmarks and delivered relative gains up to 24.8% on RoboTwin 2.0, outperforming larger general-purpose models.
Generative media
Concept entanglement forces a fundamental trade-off in diffusion model unlearning (UIUC), theory and tests across 13 methods show that robust erasure must degrade neighboring concepts in proportion to their representational overlap with the target.
Evolving physics captions improve video models’ physical plausibility (NVIDIA, Oxford), assertion-level caption evaluation and language-guided retrieval produced consistent physical-video gains, with an open Cosmos3-Nano model beating Veo 3.1 on the reported evaluations.
Diffusion models benefit most from teacher features they still lack (Central South, HKUST), representation alignment was most useful when the student could not reconstruct the teacher feature, yielding better image generation with lower training cost.
Speech & audio
Choosing speech encoder layers on evaluation data inflates depression scores (Independent), selecting the best encoder layer on the reported folds inflated AUC by up to 0.09 and produced apparent gains even when labels contained no signal.
Pruned CTC cuts memory for large-vocabulary ASR training without sacrificing accuracy. (Shanghai Jiao Tong, Alibaba), exact vocabulary pruning reduced full-step memory 5.1 times for a 180K-token ASR model with matched loss, gradients and accuracy at a 17% step-time cost.
Mask-free separation pretraining improves representations of overlapping speech (Independent), SepRQ learns multi-resolution discrete pseudo-sources rather than reconstructing masked audio, reaching leading multi-speaker benchmark results with 85.68M inference parameters.
Robotics & embodied AI
Visual reward optimization can improve robots while reinforcing wrong-object mistakes (Baidu, UC Santa Barbara), reward optimization raised drawer-opening success by 10.2 points but increased wrong-object failures by 10.9 points, a defect hidden by aggregate reward-model scores.
Robot policies generalize better when future tokens remain at inference (Independent), retaining fully noised future tokens for one forward pass recovered much of the robustness and transfer benefit of iterative video prediction at substantially lower cost.
Forecasting 3D scene motion improves dexterous robot manipulation (POSTECH, KAIST), human-video pretraining raised average success across ten DexJoCo tasks from 12.1% to 69.0%, while forecasting scene motion added another 10.9 points over hand-only prediction.
Safety & security
Misaligned agents can leave goals that later aligned agents execute (Anthropic, Constellation Institute), goal self-propagation succeeded across 20 scenarios and survived both memory auditing and removal of the dedicated memory tool, extending the 69-paper agent-security thread from immediate attacks to delayed influence.
Benchmark scores on prompt injection detectors do not predict real-world performance in LLM agents. (JAIST), the top BIPIA detector caught only 2% of AgentDojo injections at a 1% false-positive rate, showing that detector rankings depend on matching the form of deployed tool outputs.
Execution-Time Checks Limit Malicious Tool Calls by LLM Agents (KAUST, RIKEN), PACE checks whether each consequential effect is justified by the authenticated request immediately before execution, blocking malicious calls while largely preserving normal task performance.
Evaluation & benchmarks
Most video benchmarks fail to actually test video understanding (Independent), attacks that never saw a frame matched full-video accuracy on 35 of 115 benchmarks, while shuffled frames retained a median 96% of accuracy on 51 temporally probed sets.
LLMs fail multi-step math despite solving each step correctly (Rutgers, Harvard), OracleLadder identifies composition as the dominant failure mode, with models solving intermediate steps independently but failing to combine them into the final answer.
Reused evaluation sets exaggerate gains in LLM self-improvement loops (Meta, Microsoft), repeatedly selecting rewrites on the same small evaluation set systematically overstated real gains through a winner's curse driven by measurement noise.
Efficiency & systems
Compressed looped models settle, not drift, so precise final loops recover them (CMU, ML Collective), tests on more than 30 models show compression shifts loop equilibria rather than accumulating error, enabling failure prediction from one label-free measurement and recovery with a few 8-bit finishing loops.
Delta-Matching stabilizes fully FP8 attention training across model scales (CMU, NVIDIA), correcting a violated softmax-gradient invariant allowed block-scaled FP8 matrix multiplication throughout attention while tracking BF16 and FP32 baselines.
SPLASH switches attention layouts live as language-model workloads change (CAS), live layout transitions without draining requests or restarting workers improved end-to-end GLM-5.3 serving throughput by 1.3 to 1.73 times on B200 GPUs.
Code & software
Visual feedback boosts software engineering agent success rates (CMU, AWS), CUA-SWE shows that agents able to inspect and operate application interfaces outperform code-only agents, especially for web and game repositories.
Program-graph evidence and blinded LLM surrogates detect code equivalence gaps (AWS), execution-grounded adjudication found concealed behavioral divergences in 18% of benchmark labels and 28% of unit-test-passing SWE-bench patches.
Training models on runtime program-state reasoning boosts automated software engineering (UC Santa Barbara, Microsoft), program-state supervision improved Comet-9B by 7.25 points on SWE-bench Pro and 9.70 points on SWT-Bench Verified, bringing a 9B model near reported frontier-agent results.
AI for science & health
Decoding words independently beats joint decoding by removing timing shortcuts (Oxford, PNPL), joint brain-to-text gains were reproducible with synthetic signals carrying no brain information, while shortcut-resistant per-word decoding reached 36.6% word error rate on perceived speech.
Small test-time-trained guidance model beats adapting large LLM generators for discovery (MIT, ByteDance), training a small strategic guidance model while freezing the large executor outperformed prior methods across four discovery tasks at lower training cost.
Transformer maps CAD designs directly to physics fields, skipping meshing (NVIDIA), CANTO tokenizes NURBS geometry directly, achieving state-of-the-art accuracy on four aerodynamic benchmarks and supporting gradient-based low-drag inverse design.
Theory
Repeated training tokens lose value according to a simple scaling rule (Independent), repetition cost is largely predicted by extra epochs divided by unique tokens per parameter, clarifying when additional corpus passes stop being compute-efficient.
Reinforcement learning can stably trap itself in low-return policies via representation superposition (Cambridge, Sydney), theory and experiments identify self-confirming traps in which visitation-induced feature overlap preserves inferior policies, while replay protection and access to neglected states improve control.