LingBot-VLA

A Pragmatic VLA Foundation Model

Introduction

We present LingBot‑VLA, an embodied AI foundation model trained on approximately 20,000 hours of real‑world manipulation data spanning 9 mainstream dual‑arm robot configurations. We conducted a systematic evaluation across 4 robotic platforms, covering 100 tasks per configuration with 130 episodes used for post‑training adaptation per task. Experimental results demonstrate that LingBot‑VLA significantly outperforms existing solutions, validating its superior performance and extensive generalization capabilities. Furthermore, we developed an efficient post‑training toolchain that achieves a throughput of 261 samples per second on an 8-GPU setup, representing a 1.5× to 2.8× speedup over existing VLA codebases (depending on the underlying VLM). These features ensure the model is well‑suited for physical robot deployment. To advance the field of robot learning, we release our full code, base model, and benchmark data, aiming to support research on more challenging tasks and promote the standardization of evaluation protocols.

Real Machine Scaling Law
Average
Agibot G1
AgileX
Galaxea R1Pro
Leveraging massive real‑world pre‑training, we present the first systematic study of scaling laws for VLA models regarding their performance on physical robot tasks. By scaling the pre‑training data from 3,000 hours to 6,000, 13,000, 18,000, and finally 20,000 hours, we demonstrate that downstream success rates improve consistently and substantially. Notably, this performance trend shows no signs of saturation even at the 20,000-hour mark, suggesting that VLA models continue to benefit from increased data volume. These results provide empirical evidence of favorable scalability in real‑world robot learning, offering critical insights for future VLA development and large‑scale data curation.
Large‑scale Data Curation and Efficient Training Codebase
Figure 5. Training throughput analysis of the Qwen2.5-VL-3B-π model.
Figure 6. Training throughput analysis of the PaliGemma-3B-pt-224-π model.

We meticulously curated 20,000 hours of real‑world robot training data, covering nine mainstream dual‑arm embodiments. To ensure precise annotation, videos were manually segmented into atomic actions and subsequently labeled with both global and sub‑task descriptions using VLM. Our codebase integrates advanced optimizations, including Fully Sharded Data Parallel (FSDP), mixed‑precision training, and operator fusion. This combination of massive high‑quality data and a highly efficient training infrastructure underpins the superior performance of LingBot‑VLA.

Large‑scale Multi‑configuration Real‑world Testing

We conducted extensive real‑world validation across multiple robot embodiments, including Agibot G1, AgileX, and Galaxea R1Pro,and Leju KUAVO 4 Pro:
Each model was evaluated on 100 tasks per embodiment, with 130 trajectories collected per task for model training. Experimental results demonstrate that LingBot‑VLA significantly outperforms existing VLA baselines, exhibiting superior accuracy and generalization capabilities.

Depth‑Aware Robotic Manipulation

Real‑world Evaluation

PlatformWALL-OSSGROOT N1.6π0.5Ours w/o depthOurs w/ depth
SRPSSRPSSRPSSRPSSRPS
Agibot G1(Wheeled)2.99%8.75%5.23%12.63%7.77%21.98%12.82%30.04%11.98%30.47%
AgileX(wheeled)2.26%8.16%3.26%10.52%17.20%34.82%15.50%36.31%18.93%40.36%
Galaxea R1Pro(wheeled)6.89%14.13%14.29%24.83%14.10%26.14%18.89%34.71%20.98%35.40%
Leju KUAVO 4 Pro(Bipedal)3.26%11.75%6.45%18.66%12.91%26.35%17.59%36.22%15.60%34.40%
Average3.85%10.70%7.31%16.66%13.00%27.32%16.20%34.32%16.87%35.16%
Swipe to view

Simulation Evaluation

(a). Clean Scenes

π0.5Ours w/o depthOurs w/ depth
Average SR82.74%86.50%88.56%

(b). Randomized Scenes

π0.5Ours w/o depthOurs w/ depth
Average SR76.76%85.34%86.68%

To explicitly capture spatial awareness within manipulation environments and further enhance the robot’s execution robustness, we apply a query‑based depth distillation method. Specifically, we introduce learnable queries corresponding to three‑view operational images, which are processed by the VLM and aligned with depth embeddings from LingBot‑Depth. This alignment effectively integrates depth information into LingBot‑VLA while maintaining training and inference efficiency. Extensive experiments, conducted across both real robot platforms and simulation, demonstrate that this integration enhances the manipulation performance of LingBot‑VLA.

Efficient Adaptation Across Diverse Downstream Tasks for Robotic Manipulation
SR-π0.5
SR-LingBot-VLA
PR-π0.5
PR-LingBot-VLA
Empowered by large‑scale pre‑training across mainstream embodiments and comprehensive tasks, LingBot‑VLA exhibits robust general manipulation capabilities and can be efficiently transferred to diverse downstream robotic tasks. Our experiments demonstrate that LingBot‑VLA outperforms π0.5 in downstream tasks with less data, and this performance advantage continues to widen as data scales.
Interacting with transparent objects
After incorporating depth information, our model can better perceive transparent objects in the scene, such as glass vases, and follow instructions—such as "arrange flowers"—to complete the specified task.
Precise and complex operations

After a quick fine‑tuning, our model can perform tasks based on relatively complex instructions, which involve numerous fine‑grained operations, such as cleaning and placing tableware.