HiDream.ai has officially launched its embodied world model, HiDream-O1-Embodied. The company says the model strengthens robots’ physical perception and dynamic prediction, helping embodied intelligence achieve more advanced physical interaction. In the same release period, it entered the embodied intelligence evaluation platform RoboColiseum for the first time and topped the core “Disturbance Adaptation” sub-leaderboard with an average score of 0.692.
RoboColiseum’s testing method is deliberately demanding: it changes backgrounds, lighting, materials, robot initial states, camera positions, and image quality, while also rewriting instructions in diverse ways to test model stability and generalization in non-ideal environments. Disturbance adaptation is widely considered the most challenging of the four capability dimensions—not a greenhouse exam inside a lab, but a test of real ability amid various surprises.

Withstanding “imperfection” depends on a native omni-modal foundation. According to official descriptions, the model’s language understanding covers an equivalent instruction space of diverse verbs, sentence structures, and expressions, locking onto intent itself. Visually, it synthesizes multi-view information; when a single view is occluded or shifted, the system can still understand the scene through other views and continue execution. During training, it actively introduces non-ideal conditions such as lighting changes and image corruption, enabling the model to learn reliable judgments from limited clues.

The model’s confidence also comes from its data strategy. “Real foundation + generative augmentation” means high-quality real data—such as high-precision motion capture developed with Noitom—serves as the foundation, after which native omni-modal capabilities perform hundredfold-scale data expansion. Under strict physical constraints, it switches backgrounds, lighting, and object forms to produce massive training samples. The model acts as both examinee and question setter, with data and model driving each other as a flywheel. This also reflects broader AI industry trends: large-scale multimodal data processing and computing power consumption are now integral to embodied AI, alongside large language models, AI chips, and AI infrastructure.

Taking a longer view, this is HiDream’s second world model released in less than a month. Previously, the interactive world model HiDream-O1-World topped WBench’s Navi sub-leaderboard with 80.9. The interactive world model addresses “understanding and reasoning,” while the embodied world model addresses “operation and execution.” The two are complementary, pointing to the company’s bet on a unified world-model foundation.
From a technology narrative perspective, the enthusiasm around embodied world models and simulation evaluation reflects the industry’s anxiety over physical-world correctness: robots cannot merely obey commands in simulation; they must stably complete tasks under real lighting, occlusion, and interference. Training such models means processing massive multimodal data and consuming intensive compute. Evaluation, data, and training are forming a new closed