Adaptive reward routing improves joint audio-video diffusion training

The method dynamically redirects learning signals across model components and adjusts competing reward weights while retaining user-defined priorities.

Chinese Tech
Songlin Yang · Xiaotong Zhao · Jiacheng Zhang · Zhe Wang · Toyota Li · Eric Liu · +2 more

MMLab@HKUST · The Hong Kong University of Science and Technology · Tencent Video · The University of Hong Kong

Research Digest··3 min read
Yang and colleagues treat reinforcement learning for joint audio-video generation as a changing credit-assignment problem: the model must determine both where each reward should update the network and how strongly that reward should count.

The authors applied Adaptive Reward Routing to DiffusionNFT, a forward-process reinforcement-learning framework that directly supervises sampled diffusion or flow-matching timesteps.

Why this paper

From Tencent Video and 3 others

In one line

Adaptive Reward Routing dynamically adapts update locations and reward weights, improving joint audio-video diffusion across multiple metrics.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors (2 noted)
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.