LingBot-VA 2.0

Native video‑action pretraining for generalizable robot control

LingBot‑VA 2.0 redesigns the video‑action model around robot control, from representation to deployment: it predicts ahead while acting and re‑grounds on real observations for real‑time closed‑loop control.
Task Demos
From fast‑paced air hockey and fragile chip picking to conveyor‑belt sorting and long‑horizon desk tidying, LingBot‑VA 2.0 keeps future visual prediction, latent‑action decoding, and real‑observation re‑grounding in one closed loop, demonstrating stable control across tasks.
Air Hockey
Chip Picking
Conveyor Belt
Desk Tidying
Quantitative Results
We evaluate LingBot‑VA 2.0 on four real‑world robot tasks — pen collection, fruit sorting, drawer tidying, and plate handover — plus the RoboTwin simulation benchmark under randomized settings, comparing success rate against LingBot‑VA 1.0 and π₀.₅.
Success Rate
LingBot-VA 2.0
LingBot-VA 1.0
π₀.₅
LingBot-VA 2.0's Capabilities

LingBot‑VA 2.0 keeps extending its abilities at deployment: Foresight Reasoning overlaps inference with execution to cut closed‑loop latency, consistency distillation with a three‑level acceleration stack delivers over 4x end‑to‑end speedup, a single demonstration video adapts the policy to new tasks via ICL, and a low‑frequency planner decomposes long‑horizon goals into executable chunks.

Foresight Reasoning: predict ahead while acting

LingBot‑VA 2.0 does not wait for a full video rollout before acting. A prediction stream prepares future latents while the robot executes the current action chunk, then re‑grounds the KV cache when new observations arrive, reducing closed‑loop latency and imagined‑rollout drift.

Pipeline: asynchronous prediction, execution, and observation grounding

When new real observations arrive, the inference branch first re‑grounds the cache with FDM, imagines the visual outcome of the currently executing action, and prepares the next action chunk before execution stalls.

t=0t=1t=2t=3t=1t=2t=3Obs 0VMIDMImagine 1Action 0Executing Action 0VM(FDM)Obs 0 CacheGrounded Obs 1 CacheVMIDMImagine 2Action 1Executing Action 1VM(FDM)Obs 0 CacheObs 1 CacheGrounded Obs 2 CacheVMIDMImagine 3Action 2
The asynchronous inference/execution timeline of Foresight Reasoning.
End‑to‑end speedup

By combining consistency distillation with a three‑level inference acceleration stack, including FP8 TensorRT compilation, long‑horizon attention optimization, and runtime overhead amortization, our system achieves over 4x end‑to‑end speedup while preserving smooth action control.

Preparing smooth playback
0%
30 Hz
Preparing smooth playback
0%
60 Hz
Preparing smooth playback
0%
150 Hz

Demonstration video as a dynamic visual example

ICL encodes human or robot demonstration videos into context that the policy can continuously attend to, letting the model understand new object layouts, motion timing, and task variants without retraining.

Scenario: multi‑object pick and place
Human demonstration video
Robot execution video

Low-frequency planning, high-frequency video-action control

The planner turns task goals and new observations into sub‑task context in the background, while the VA executor keeps acting at high frequency and reads the latest context at chunk boundaries.

Hierarchical Planning

Task goal, observations, and robot state form a loop: the planner updates context, the executor emits actions, and new observations feed the next planning step.

The planner uses the task goal and observations to generate sub‑task context; the low‑level executor emits actions and produces new observations.New observationObservationRobot StateTask GoalHigh‑Level PlannerVLMSub‑Task Contextinstruction / generation / sceneLow‑Level ExecutorVA · action chunks1.0-2.2...Robot Action
The planner updates task context at low frequency; the executor acts at high frequency and feeds new observations back into the loop.
How We Achieve This

LingBot‑VA 2.0 follows native video‑action pretraining: instead of retrofitting a generic video generator into a controller, it redesigns the tokenizer, causal modeling, capacity scaling, and training objectives around robot control. Pretraining settles representation, capacity, and data; the previous section covers acceleration and adaptation at inference time.

01
Semantic Visual-Action Tokenizer

The visual tokenizer is semantically aligned with a frozen perception encoder, and latent actions are self‑supervised by IDM / FDM from adjacent visual latents, keeping the representation useful for control.

02
Causal Video-Action Pretraining

A causal DiT jointly predicts future visual latents and latent actions under language and planner context, matching deployment where robots only see past observations.

03
Mixture-of-Experts Video Stream

The video stream uses sparse MoE routed layers (top-8 of 128 experts) while the action stream stays dense, scaling visual‑dynamics capacity without extra action‑decoding cost.

04
Multi-Chunk Prediction

Training jointly predicts next-1 / next-2 / next-3 future chunks, reducing myopic rollouts and error accumulation; ablations show faster convergence.