The first open-source MoE video foundation model for embodied intelligence.
The first open-source MoE video foundation model for embodied intelligence.
LingBot-Video is an MoE video foundation model and DiT pretraining paradigm for embodied intelligence, combining a task-unified Single-Stream Diffusion Transformer, sparse expert scaling, 70,000+ hours of embodiment-oriented data, and reward-aligned post-training to connect open-world generation with physically grounded action simulation.
General and Embodied Video Simulation
Explore scenes, materials, motion, and embodied tasks in one model, generating open‑world videos and future trajectories from text, image, and structured prompts.
Robotics and Embodied AI
Covers robotic arms, humanoids, quadrupeds, mobile platforms, and egocentric agents, generating action progress, contact states, environmental feedback, and task outcomes.
Material and Lighting Properties
Covers material appearance, surface detail, lighting, reflections, fluid motion, and deformation, preserving stable texture and plausible visual response under motion, contact, and camera changes.
Motion and Dynamics
Covers human, animal, sports, egocentric, and natural dynamics, generating coherent motion, stable poses, and realistic movement speed.
LingBot‑Video Core Model Features
Through Sparse MoE Scaling, Data Profiling Engine, Embodied Evaluation, and Action‑to‑Video Simulation, LingBot‑Video moves from open‑world video generation toward a physical‑world simulator.
Sparse MoE activates only a small subset of experts per token, allowing total model capacity to scale without proportionally increasing every denoising step.
MoE‑to‑dense speed ratio, computed as dense latency divided by MoE 30B‑A3B latency.
Sparse MoE activates only a small subset of experts per token, allowing total model capacity to scale without proportionally increasing every denoising step.
MoE‑to‑dense speed ratio, computed as dense latency divided by MoE 30B‑A3B latency.
Unlocking the Embodied Physical World
Across service, manufacturing, mobility, interaction, and robotic manipulation, LingBot‑Video can serve as an embodied video simulator for data synthesis, policy evaluation, and action planning.