The visual tokenizer is semantically aligned with a frozen perception encoder, and latent actions are self‑supervised by IDM / FDM from adjacent visual latents, keeping the representation useful for control.
LingBot-VA 2.0
Native video‑action pretraining for generalizable robot control
LingBot‑VA 2.0 keeps extending its abilities at deployment: Foresight Reasoning overlaps inference with execution to cut closed‑loop latency, consistency distillation with a three‑level acceleration stack delivers over 4x end‑to‑end speedup, a single demonstration video adapts the policy to new tasks via ICL, and a low‑frequency planner decomposes long‑horizon goals into executable chunks.
Foresight Reasoning: predict ahead while acting
LingBot‑VA 2.0 does not wait for a full video rollout before acting. A prediction stream prepares future latents while the robot executes the current action chunk, then re‑grounds the KV cache when new observations arrive, reducing closed‑loop latency and imagined‑rollout drift.
When new real observations arrive, the inference branch first re‑grounds the cache with FDM, imagines the visual outcome of the currently executing action, and prepares the next action chunk before execution stalls.
By combining consistency distillation with a three‑level inference acceleration stack, including FP8 TensorRT compilation, long‑horizon attention optimization, and runtime overhead amortization, our system achieves over 4x end‑to‑end speedup while preserving smooth action control.



Demonstration video as a dynamic visual example
ICL encodes human or robot demonstration videos into context that the policy can continuously attend to, letting the model understand new object layouts, motion timing, and task variants without retraining.
Low-frequency planning, high-frequency video-action control
The planner turns task goals and new observations into sub‑task context in the background, while the VA executor keeps acting at high frequency and reads the latest context at chunk boundaries.
Task goal, observations, and robot state form a loop: the planner updates context, the executor emits actions, and new observations feed the next planning step.
LingBot‑VA 2.0 follows native video‑action pretraining: instead of retrofitting a generic video generator into a controller, it redesigns the tokenizer, causal modeling, capacity scaling, and training objectives around robot control. Pretraining settles representation, capacity, and data; the previous section covers acceleration and adaptation at inference time.
A causal DiT jointly predicts future visual latents and latent actions under language and planner context, matching deployment where robots only see past observations.
The video stream uses sparse MoE routed layers (top-8 of 128 experts) while the action stream stays dense, scaling visual‑dynamics capacity without extra action‑decoding cost.
Training jointly predicts next-1 / next-2 / next-3 future chunks, reducing myopic rollouts and error accumulation; ablations show faster convergence.