All models

Vision-language-action foundation model

LingBot-VLA

Robbyant's pragmatic vision-language-action model for generalist robotic manipulation across platforms.

Open-source code, paper, and model releasesLocal open source
OrganizationAnt Group / Robbyant
CategoryVision-language-action foundation model
AvailabilityGitHub, arXiv report, Hugging Face cards and collection, ModelScope checkpoints.
Last reviewed2026-05-05

01 / Overview

Embodied control, robot manipulation, instruction following, depth-aware perception, real-world post-training.

Robbyant's pragmatic vision-language-action model for generalist robotic manipulation across platforms.

LingBot-VLA is included because vision-language-action models explain how world-model-adjacent perception turns into task execution.

Its category value is strongest when paired with LingBot-VA: VLA policy framing and video-action world modeling are related but not interchangeable.

The page should keep emphasis on embodied control and benchmark evidence instead of stretching the term world model into a generic AI label.

02 / Strengths

Where it stands out

  • Extends physical-AI coverage from prediction into action-taking embodied systems.
  • Strong reproducible sources: code, arXiv paper, project page, downloadable checkpoints.
  • Shows Ant's embodied-AI stack spans simulation, robot-control modeling, VLA policies.

03 / Boundaries

What not to overclaim

  • No consumer world-building product; not HappyOyster, Marble, or Genie 3.
  • Evidence: robot manipulation benchmarks and releases, not open-ended world simulation.

04 / Practical path

How to evaluate it

  1. 01

    Confirm supported tasks, model variants, setup assumptions in GitHub.

  2. 02

    Check evaluation scope in arXiv before comparing robot-control world models.

  3. 03

    Verify checkpoint naming and availability in the Hugging Face collection.

Continue exploring

Compare the lane, then follow the evidence

Primary evidence

Sources