Generative media

Generating images, video, 3D and audio: diffusion, flow matching, editing, generative pipelines

43 articles

Generative media

A method that generates consistent video observations across multiple vehicles in a shared driving scene

Meng et al. propose CoDrive, a cross-vehicle video generation framework that produces world-consistent observations for multiple vehicles in the same scene. By interleaving local and global self-attention and injecting camera trajectories in a shared coordinate system, the model significantly improves trajectory controllability and cross-agent geometric and instance consistency while maintaining visual quality.

3 Oct·2 min
Generative media

Gaussian mixtures improve distributional training for one-step image generators

Zhang, Shi, Liu and colleagues present a common mathematical account of distributional training, where generators learn by matching collections of real and generated features from frozen image encoders. Their resulting method, MGFlow, represents those distributions as Gaussian mixtures and reports stronger one-step generation than prior feature-distribution objectives on ImageNet and text-to-image benchmarks.

3 Oct·2 min
Generative media

Keeping token uncertainty improves discrete diffusion generation and fast sampling

Deschenaux et al. introduce Simplex Diffusion Models, which keep intermediate predictions as probability distributions rather than repeatedly collapsing them into categorical token choices. The authors derive closed-form reverse transitions, use ordinary cross-entropy training, and report gains over masked and uniform diffusion on code generation, including after aggressive sampling distillation.

3 Oct·2 min
Generative media

Training-free wrapper accelerates causal video diffusion to 50 FPS without quality loss

The authors present UnStep, a training-free wrapper that accelerates already-distilled causal video diffusion models by running them with fewer denoising steps and a limited temporal KV cache. Two inference-only mechanisms—clean-cache refinement (renoise and refine) and truncated SVD on attention projections—recover the quality lost by these reductions, while an optimized runtime stack further boosts throughput. UnStep achieves 50 frames per second on a single H100 GPU and 77 FPS on a GB200 GPU without quality degradation, all without retraining.

3 Oct·3 min
Generative media

Dynamic optimization loop aligns 3D generators with 2D diffusion priors to improve realism

The authors propose OREO, a framework that enhances visual fidelity of 3D generators by creating a dynamic optimization loop. They use a 2D image editing model to refine rendered views of 3D outputs on the fly, producing high-quality pseudo-targets. Then, a latent contrastive objective distills the improvements back into the 3D generator, resulting in significantly more realistic assets.

3 Oct·2 min
Generative media

Permutation-equivariant flow matching generates neural network weights without alignment

The authors introduce a flow-matching generative model whose velocity field is parameterized by a permutation-equivariant Graph Meta Network. This allows them to learn directly from collections of independently trained neural networks without the need for approximate neuron alignment. The method generates new networks that closely reproduce the joint statistics of accuracy, functional similarity, and weight similarity observed in training collections, and a single conditional model can produce task-specific networks for heterogeneous architectures and previously unseen hidden-width configurations.

3 Oct·3 min
Generative media

Staged AI agent improves generation of positive and negative space compositions

The authors present FaV-A, a multimodal agent that generates positive-negative space images through three connected stages. It first creates a base object, then analyzes its contours to propose secondary semantics for the negative space, and finally issues compositional instructions for the image generation model. In user studies and ablations, FaV-A produced more visually coherent and semantically aligned compositions than zero-shot MLLM baselines.

2 Oct·2 min
Generative media

Text-to-video model that learns composite physics and adapts to new dynamics

The authors propose CompAdapt, a diffusion-based text-to-video framework that improves physical consistency by modeling composite motions beyond simple single-type dynamics. The system translates natural language prompts into structured physical semantics, uses a dynamics-aware prior matching module for one-shot adaptation to novel environments, and employs a physics-aware latent feature fusion to maintain visual fidelity. On physics-focused benchmarks, CompAdapt outperforms both general T2V models and previous physics-constrained methods in physical plausibility while preserving high visual quality.

22 Sept·3 min
Generative media

Diffusion models generate high-resolution elevation maps from low-resolution inputs guided by optical imagery

The authors propose a guided super-resolution approach for digital surface models (DSMs) using denoising diffusion, improving coarse 5 m DSMs to 0.5 m resolution by leveraging high-resolution optical imagery. Experiments on several Central European cities show the method produces DSMs with crisper building outlines and more detailed roof structures than conventional interpolation or filtering.

13 Sept·2 min
Generative media

Identity preservation remains a distinct challenge for generative image models

The authors systematically benchmark three approaches to identity preservation in generative image models: encoding identity in the input context, using trainable subject-specific parameters (like LoRA), or maintaining a persistent identity layer. Their results show that persistent identity layers consistently reduce identity degradation across iterative edits, small subject scales, and multi-subject compositions, while preserving image quality and instruction adherence.

7 Sept·2 min
Generative media

World state registers enable consistent multi-agent video generation across views

The authors propose WorldWeaver, a streaming multi-agent video diffusion model that augments autoregressive rollout with cross-agent world state registers: learnable tokens that maintain shared world information, track individual agent status, and are updated after each generated chunk. The registers are grounded with supervision from agent status, global bird's-eye views, and scene text. Experiments in two-agent Minecraft show that explicit world-state modeling improves logical consistency and generation quality over baselines that only carry forward observation history.

24 July·3 min
Generative media

Skill evolution improves agent performance in image generation workflows

The authors introduce COMFYCLAW, an agentic harness that controls ComfyUI workflows for image generation. By representing workflow construction as typed graph editing, reverting invalid edits, and using a region-level vision-language model (VLM) verifier to provide actionable repair suggestions, COMFYCLAW evolves a skill library from past trajectories, errors, and verifier feedback. Across four benchmark splits, three agent models, and two image backbones, COMFYCLAW achieves the best average evaluation score in all six agent configurations, and human annotators prefer its outputs over variants without skill evolution.

2 July·3 min
Generative media

An agentic framework that fills in missing context for real-world image generation

The authors propose Qwen-Image-Agent to address the context gap where user requests for image generation are often underspecified, implicit, or depend on up-to-date knowledge. The framework uses Context-Aware Planning to identify missing details and Context Grounding to acquire them from reasoning, search, memory, and feedback. On the newly introduced IA-Bench benchmark and two other datasets, it achieves state-of-the-art performance against strong baselines.

25 June·2 min
Generative media

Agent framework auto-tunes video diffusion for 2x speedup

The authors present Sol Video Inference Engine, a training-free framework that uses parallel agent modules to optimize cache, sparse attention, token pruning, quantization, and kernel fusion for a given model-hardware-configuration triplet. Across three models (64B Cosmos3-Super, 22B LTX-2.3, 2B SANA-Video), the full stack achieves over 2x end-to-end acceleration while maintaining near-lossless VBench quality, with minimal human effort.

22 June·2 min
Generative media

Multi-agent orchestration generates 3D scenes from single images

The authors propose SceneConductor, a multi-agent framework that generates complete 3D scenes from a single input image. It decomposes the task into three stages—scene initialization, environment construction, and multi-agent refinement—and introduces a geometry-aware layout predictor trained with sparse point-map priors. The method outperforms prior approaches on geometric accuracy, spatial consistency, and perceptual realism across benchmark datasets.

7 June·2 min
Generative media

Multi-agent framework coordinates narrative and visual consistency for long-form video

The authors introduce ViMax, an agentic video generation framework that coordinates multiple specialized agents to produce long-form videos with narrative planning and visual consistency. By combining a hierarchical narrative engine with retrieval-augmented generation and dependency-aware visual tracking, ViMax maintains global story coherence and consistent character/environment states across scenes, addressing limitations of existing short-clip methods.

2 June·2 min
Generative media

Unified visual-generation agentic model outperforms larger closed-source models

The authors propose VisionCreator, a native visual-generation agentic model that unifies Understanding, Thinking, Planning, and Creation (UTPC) capabilities. End-to-end trained on a new dataset and optimized via Progressive Specialization Training and Virtual Reinforcement Learning, VisionCreator-8B/32B models outperform larger closed-source models on the new VisGenBench benchmark across multiple dimensions.

3 Mar·2 min