GitHub radar

A ready-made model that turns a photo and a command into robot-arm moves

This is a foundation model for robot arms: it takes a camera photo and a text command and outputs arm movements. Fine-tune it for your setup instead of training from scratch.

01Robbyant/lingbot-vla 1.8kPython

The model looks at a camera image, reads a text command, and outputs robot-arm movements.

When it helps

Useful if you work with robotics and want a ready base to fine-tune for your own arm. Not useful if you have no physical robot or your task isn't about arm control.

Pros

  • Trained on 20,000 hours of data from nine robots — a solid starting base
  • Training code runs 1.5–2.8x faster than comparable codebases
  • Weights are on Hugging Face and ModelScope; the fine-tuning guide is in the GitHub repo
  • A depth-distilled version exists alongside the regular one

Cons

  • Needs a physical robot and a GPU for training and running
  • Setup and fine-tuning require Python, conda, and editing config files

How to set it up — step by step

  1. 1Create a Python 3.12 environment with conda — a tool for managing Python environments.
  2. 2Clone the repository with git clone https://github.com/Robbyant/lingbot-vla.git
  3. 3Run bash install.sh to install the dependencies.
  4. 4Download the model weights from Hugging Face or ModelScope — AI model hosting platforms.
  5. 5Open the fine-tuning guide in the repo and prepare your data in LeRobot format.
  6. 6Send the agent a task to run training or evaluation for your robot arm.
Open on GitHub