GitHub radar
A ready-made model that turns a photo and a command into robot-arm moves
This is a foundation model for robot arms: it takes a camera photo and a text command and outputs arm movements. Fine-tune it for your setup instead of training from scratch.
The model looks at a camera image, reads a text command, and outputs robot-arm movements.
When it helps
Useful if you work with robotics and want a ready base to fine-tune for your own arm. Not useful if you have no physical robot or your task isn't about arm control.
Pros
- Trained on 20,000 hours of data from nine robots — a solid starting base
- Training code runs 1.5–2.8x faster than comparable codebases
- Weights are on Hugging Face and ModelScope; the fine-tuning guide is in the GitHub repo
- A depth-distilled version exists alongside the regular one
Cons
- Needs a physical robot and a GPU for training and running
- Setup and fine-tuning require Python, conda, and editing config files
How to set it up — step by step
- 1Create a Python 3.12 environment with conda — a tool for managing Python environments.
- 2Clone the repository with git clone https://github.com/Robbyant/lingbot-vla.git
- 3Run bash install.sh to install the dependencies.
- 4Download the model weights from Hugging Face or ModelScope — AI model hosting platforms.
- 5Open the fine-tuning guide in the repo and prepare your data in LeRobot format.
- 6Send the agent a task to run training or evaluation for your robot arm.
▌ More finds