The authors developed MaLiang-Harness to address the Program-to-Visual gap: code may run correctly while producing the wrong composition, appearance, or motion.
Stateful code revision improves programmable image and video generation
MaLiang-Harness links executable visual programs, rendered evidence, and revision history so multimodal models can inspect and correct their outputs.
Chinese Tech
Haoyu Zhao · Zihao Zhang · Xudong Wang · Jiaxi Gu · Zuxuan Wu · Yu-Gang Jiang · +1 more
National University of Singapore · Fudan University · Tencent
Research Digest··2 min read
Zhao et al.
Why this paper
From Tencent and 2 others · Released code
In one line
MaLiang-Harness bridges the program-to-visual gap by enabling MLLMs to iteratively construct, inspect, and revise executable visual programs for image and video generation.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§