LingBot-VLA 2.0
From Foundation to Application: Improving VLA Models in Practice
Scaling up Pre‑training Dataset
LingBot‑VLA 2.0 scales pre‑training with a broader data mixture: 50,000 hours of real robotic data and 10,000 hours of embodiment‑free egocentric manipulation data.
Data processing pipeline
The raw pool is split into robotic and egocentric streams, then filtered stage by stage. Gray bands represent rejected data; colored blocks are retained and passed forward.
Unified Action Space
LingBot‑VLA 2.0 aligns 20 robot embodiments into a shared action representation for scalable real‑robot training.
Model Design
Sparse MoE
LingBot‑VLA 2.0 adopts a sparse MoE architecture to achieve more efficient scaling. Under the same active parameter budget, MoE consistently attains lower training loss and validation error than its dense counterpart, demonstrating the superiority across both optimization and generalization metrics.

Perceptual Results of Dual‑query Distillation
We visualize current and future perceptual predictions across RGB, depth, and DINO representations.




















Real‑World Benchmark Results
We evaluate LingBot‑VLA 2.0 across tabletop bimanual manipulation and long‑horizon mobile manipulation in realistic scenes.
Tabletop Bimanual Manipulation
Overall average success rate and progress score over 9 GM-100 tasks under the generalist setting, where all tasks are jointly trained in a mixed multi‑task policy.
Long‑Horizon Mobile Manipulation
Progress score and success rate under in‑domain and out‑of‑domain settings.
It's showtime!
LingBot‑VLA 2.0 can adapt to multiple robotic embodiments and, when combined with pretraining over whole‑body degrees of freedom, achieves more proficient and stable performance across a wide range of daily‑life tasks, including high‑precision, articulated‑object, and long‑horizon mobile manipulation.