32 papers this week12 active threadsbusiest: Memory Management for Agentsdaily arXiv scan · 6am Brisbane

Research

Latest Paper· AI & ML

Looped flows enable deeper reasoning by chaining denoising steps

Suleymanzade et al. propose looped flows, a method that trains recurrent neural networks for many-step inference using local denoising losses. This overcomes the difficulty of training early updates to support later ones when backpropagation is truncated. The authors report state-of-the-art accuracy among looped models on six reasoning benchmarks, including 58.8% on ARC-AGI-1 and 12.2% on ARC-AGI-2.

Ayhan Suleymanzade, Chanhyuk Lee, Floor Eijkelboom +3arXiv →
AI & ML

Diffusion models generate high-resolution elevation maps from low-resolution inputs guided by optical imagery

The authors propose a guided super-resolution approach for digital surface models (DSMs) using denoising diffusion, improving coarse 5 m DSMs to 0.5 m resolution by leveraging high-resolution optical imagery. Experiments on several Central European cities show the method produces DSMs with crisper building outlines and more detailed roof structures than conventional interpolation or filtering.

2 days ago · arXiv
Safety & Evals

Hidden prompt revisions inject cultural stereotypes into generated images

The authors introduce WORLDVIEW, a benchmark of 8,960 prompts spanning 15 languages and 31 language–context pairings, to examine how text-to-image systems rewrite requests before generation. Across DALL-E-3, Imagen-4, and GPT-Image-1.5, revisions marked non-Western and non-Anglophone settings more heavily than the United States and repeatedly reduced them to narrow, stereotypical vocabularies.

2 days ago · arXiv

The Field

Agent Harnesses

all →

Compressing Shared Sandbox Memory Makes High-Fanout Agents More Efficient

2 days ago

Embedded software agents need interaction contracts and continuous assurance

3 days ago

VikingRAG Cuts Token Use for Retrieval Over Structured Documents

3 days ago

Similarity-aware context windows improve model routing across multi-turn conversations

3 days ago

Tools & MCP

all →

Closed-loop synthesis efficiently trains language models to use external tools

5 days ago

Structured software interfaces outperform screenshot-and-click control for AI agents

17 days ago

Adaptive tool use improves video research agents’ accuracy and efficiency

18 days ago

Joint training helps small language models create and use tools

19 days ago

Multi-Agent Systems

all →

Backward Bayesian reasoning helps LLM agents resolve conflicting diagnoses

2 days ago

Task-specific hierarchies improve coordination in large embodied AI teams

3 days ago

LLM teammates retain performance after swaps but coordinate less efficiently

7 days ago

Shared infrastructure spread both cheating and resistance through an AI swarm

10 days ago

Creative Agents

all →

Identity preservation remains a distinct challenge for generative image models

8 days ago

Image-only pretraining helps build a strong open image generator

9 days ago

Coding agents combine generated imagery with editable web-based visual layouts

10 days ago

Language enables precise character and camera control in video worlds

12 days ago

Memory & Planning

all →

Agent-side memory steers stateless robots through long manipulation tasks

2 days ago

Kernel-managed memory personalizes multiple AI agents with shorter prompts

4 days ago

Agents can retain old facts without using them indiscriminately

4 days ago

Hierarchical memory trees improve long-context reasoning without model training

4 days ago

Safety & Evals

all →

Hidden prompt revisions inject cultural stereotypes into generated images

2 days ago

Deep research agents struggle with long, multimodal evidence chains

2 days ago

Refresh schedules, not cache age, determine agent answer staleness

3 days ago

Cross-device workflows expose major weaknesses in today’s GUI agents

4 days ago

RL for Agents

all →

Enumerating Tool Choices Beats Sampling Them in Genomic Reasoning

4 days ago

Synthetic rewards train agents to diagnose simulated advertising anomalies

4 days ago

Adaptive rollout trees broaden language models’ mathematical reasoning coverage

5 days ago

Factorized reinforcement learning improves vision-language models’ spatial reasoning

9 days ago

AI & ML

all →

Looped flows enable deeper reasoning by chaining denoising steps

2 days ago

Diffusion models generate high-resolution elevation maps from low-resolution inputs guided by optical imagery

2 days ago

Sparse mixture-of-experts models overfit repeated training data faster

2 days ago

Chain-of-thought reasoning traces are legible but not interpretable

8 days ago

Every paper read and written up by the research desk from the daily arXiv scan · threads are maintained lines of inquiry with running syntheses