One action-conditioned JEPA improves vision-language-action policies across all lifecycle stages

Juno unifies representation, prediction, and adaptation in a single model, raising SimplerEnv success to 68.5% and retaining 70-75% real-robot success under shifts that cripple the base policy.

Academic
Yuchen Zhu · Chenyi Xu · Yulin Zhang · Gang Xu · Wentao Zhu

University of Science and Technology of China · Eastern Institute of Technology, Ningbo · Hangzhou Dianzi University · ShanghaiTech University

Research Digest··2 min read
The authors present Juno, a framework that repurposes a single action-conditioned joint-embedding predictive architecture (JEPA) as a control-aligned backbone, a predictive teacher, and an adaptable dynamics model across the vision-language-action (VLA) lifecycle.

Juno pretrains a lightweight action-conditioned JEPA on the same embodiment-matched trajectories used for policy learning.

Why this paper

From University of Science and Technology of China and 3 others

In one line

Juno uses a single action-conditioned JEPA to improve VLA success rates across pretraining, policy learning, and deployment.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.

How we workSubscribe