Abstract
Hydra-0 uses action flow as a shared visual interface for generalist world modeling and robot control across diverse embodiments and tasks.
We introduce Hydra-0, a generalist world model conditioned on action flow, which represents robot actions as pixel motion. This shared visual interface enables generalist world modeling and control by learning action consequences across embodiments, tasks, environments, and video-generation backbones. Our best configuration achieves 90.4% lower robot-motion error and 60.2% lower object-motion error than our action-conditioned baseline, while supporting zero-shot composition and data-efficient adaptation. On the RoboLab benchmark, Hydra-0 achieves a Pearson correlation of r=0.96 between replayed and reference success rates. Finally, we uncover an emergent inverse mode of this interface: a world action model that predicts compatible robot motion from desired object flow transferred from a human demonstration. A trained action head maps the resulting latent features to executable actions without requiring task-specific expert robot demonstrations. Together, these results demonstrate the potential of action flow as a shared control interface connecting heterogeneous training data, open-loop policy evaluation, and robot control.
Community
A generalist world model conditioned on action flow: robot actions represented as pixel motion. At deployment it runs as a hybrid simulator:
a physics engine moves the robot, a learned video model predicts what the world does in response. One model trains across embodiments, simulates them, evaluates policies, and drives a real robot.This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Robot-Factored World Models via Robot Rendering (2026)
- JoyAI-RA 0.5: Scaling Robot Manipulation Learning via Dual Action Alignment (2026)
- WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory (2026)
- DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation (2026)
- Native Video-Action Pretraining for Generalizable Robot Control (2026)
- ContactFlow: A video action conditioning that transfers across embodiments (2026)
- Masked Visual Actions for Unified World Modeling (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.18077 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper