We present LingBot‑VLA, an embodied AI foundation model trained on approximately 20,000 hours of real‑world manipulation data spanning 9 mainstream dual‑arm robot configurations. We conducted a systematic evaluation across 4 robotic platforms, covering 100 tasks per configuration with 130 episodes used for post‑training adaptation per task. Experimental results demonstrate that LingBot‑VLA significantly outperforms existing solutions, validating its superior performance and extensive generalization capabilities. Furthermore, we developed an efficient post‑training toolchain that achieves a throughput of 261 samples per second on an 8-GPU setup, representing a 1.5× to 2.8× speedup over existing VLA codebases (depending on the underlying VLM). These features ensure the model is well‑suited for physical robot deployment. To advance the field of robot learning, we release our full code, base model, and benchmark data, aiming to support research on more challenging tasks and promote the standardization of evaluation protocols.
We meticulously curated 20,000 hours of real‑world robot training data, covering nine mainstream dual‑arm embodiments. To ensure precise annotation, videos were manually segmented into atomic actions and subsequently labeled with both global and sub‑task descriptions using VLM. Our codebase integrates advanced optimizations, including Fully Sharded Data Parallel (FSDP), mixed‑precision training, and operator fusion. This combination of massive high‑quality data and a highly efficient training infrastructure underpins the superior performance of LingBot‑VLA.
We conducted extensive real‑world validation across multiple robot embodiments, including Agibot G1, AgileX, and Galaxea R1Pro,and Leju KUAVO 4 Pro:
Each model was evaluated on 100 tasks per embodiment, with 130 trajectories collected per task for model training. Experimental results demonstrate that LingBot‑VLA significantly outperforms existing VLA baselines, exhibiting superior accuracy and generalization capabilities.
Real‑world Evaluation
| Platform | WALL-OSS | GROOT N1.6 | π0.5 | Ours w/o depth | Ours w/ depth | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| SR | PS | SR | PS | SR | PS | SR | PS | SR | PS | |
| Agibot G1(Wheeled) | 2.99% | 8.75% | 5.23% | 12.63% | 7.77% | 21.98% | 12.82% | 30.04% | 11.98% | 30.47% |
| AgileX(wheeled) | 2.26% | 8.16% | 3.26% | 10.52% | 17.20% | 34.82% | 15.50% | 36.31% | 18.93% | 40.36% |
| Galaxea R1Pro(wheeled) | 6.89% | 14.13% | 14.29% | 24.83% | 14.10% | 26.14% | 18.89% | 34.71% | 20.98% | 35.40% |
| Leju KUAVO 4 Pro(Bipedal) | 3.26% | 11.75% | 6.45% | 18.66% | 12.91% | 26.35% | 17.59% | 36.22% | 15.60% | 34.40% |
| Average | 3.85% | 10.70% | 7.31% | 16.66% | 13.00% | 27.32% | 16.20% | 34.32% | 16.87% | 35.16% |
Simulation Evaluation
(a). Clean Scenes
| π0.5 | Ours w/o depth | Ours w/ depth | |
|---|---|---|---|
| Average SR | 82.74% | 86.50% | 88.56% |
(b). Randomized Scenes
| π0.5 | Ours w/o depth | Ours w/ depth | |
|---|---|---|---|
| Average SR | 76.76% | 85.34% | 86.68% |
To explicitly capture spatial awareness within manipulation environments and further enhance the robot’s execution robustness, we apply a query‑based depth distillation method. Specifically, we introduce learnable queries corresponding to three‑view operational images, which are processed by the VLM and aligned with depth embeddings from LingBot‑Depth. This alignment effectively integrates depth information into LingBot‑VLA while maintaining training and inference efficiency. Extensive experiments, conducted across both real robot platforms and simulation, demonstrate that this integration enhances the manipulation performance of LingBot‑VLA.
After a quick fine‑tuning, our model can perform tasks based on relatively complex instructions, which involve numerous fine‑grained operations, such as cleaning and placing tableware.
After a quick fine‑tuning, our model can perform tasks based on relatively complex instructions, which involve numerous fine‑grained operations, such as cleaning and placing tableware.