LingBot-Video

The first open-source MoE video foundation model for embodied intelligence.

LingBot-Video

The first open-source MoE video foundation model for embodied intelligence.

LingBot-Video is an MoE video foundation model and DiT pretraining paradigm for embodied intelligence, combining a task-unified Single-Stream Diffusion Transformer, sparse expert scaling, 70,000+ hours of embodiment-oriented data, and reward-aligned post-training to connect open-world generation with physically grounded action simulation.

General and Embodied Video Simulation

Explore scenes, materials, motion, and embodied tasks in one model, generating open‑world videos and future trajectories from text, image, and structured prompts.

Material and Lighting Properties

Covers material appearance, surface detail, lighting, reflections, fluid motion, and deformation, preserving stable texture and plausible visual response under motion, contact, and camera changes.

Motion and Dynamics

Covers human, animal, sports, egocentric, and natural dynamics, generating coherent motion, stable poses, and realistic movement speed.

LingBot‑Video Core Model Features

Through Sparse MoE Scaling, Data Profiling Engine, Embodied Evaluation, and Action‑to‑Video Simulation, LingBot‑Video moves from open‑world video generation toward a physical‑world simulator.

Sparse MoE Scaling

LingBot‑Video uses a Single‑Stream DiT to model visual latents and condition tokens together, replacing dense FFNs with shared and top‑K routed experts to scale capacity while controlling per‑token computation; at 1M tokens, MoE 30B‑A3B reaches about 3.18x speedup over Dense-30B.

Sparse MoE activates only a small subset of experts per token, allowing total model capacity to scale without proportionally increasing every denoising step.

MoE‑to‑dense speed ratio, computed as dense latency divided by MoE 30B‑A3B latency.

Benchmark Results
Action‑to‑Video Simulation
Data Profiling Engine

Unlocking the Embodied Physical World

Across service, manufacturing, mobility, interaction, and robotic manipulation, LingBot‑Video can serve as an embodied video simulator for data synthesis, policy evaluation, and action planning.