Agentic Robotics: From LLM Tool-Calling to Vision-Language-Action Models
Wire an LLM agent to a real robot, then teach it to act
Wire an LLM agent to a real robot, then teach it to act. Anis Koubaa. 4 modules · 68 lessons · 5h · Intermediate. Audio in English, English (US), Arabic, Turkish, French.
A hands-on course on giving robots agency. You will connect a large language model to a ROS 2 robot through tool calls, collect a demonstration dataset, fine-tune and deploy a Vision-Language-Action model (SmolVLA), and finish with a planner-to-executor mission running end to end.
Everything is grounded in one simulated greenhouse robot and one concrete task — inspect the tomato rows — so the platform stays constant while the intelligence on top of it evolves lesson by lesson.
Audience: engineers and graduate students comfortable with Python. No robotics or machine-learning background is required. Ubuntu 24.04 is the supported platform.
All the code is included: the simulated robot, its skill servers, the agent scripts and every lesson demo come as a GitHub repository you clone and install in module 0.
Module 3 highlights: you open and read a real demonstration dataset (49 robot picks, 4,228 frames), see how a Vision-Language-Action model turns a camera image and an instruction into motion, fine-tune SmolVLA on that dataset with one command, and evaluate it honestly on episodes it never saw.
What you need for Module 3: modules 0 to 2 run on native Ubuntu 24.04 with 8 GB of RAM and no GPU. Module 3 adds the machine-learning stack (./setup/install.sh --full, about 4 GB). Its dataset lessons, 3.4 to 3.8, run without a GPU. Fine-tuning and running SmolVLA in lessons 3.11 and 3.12 needs an NVIDIA GPU with CUDA: the lesson recipe fits in under 4 GB of GPU memory, and the reference policy was trained in about two hours on a 16 GB laptop GPU. Without a GPU you can still follow those lessons and study their code and results. The Module 1 slide examples, such as the talker and listener, the AddTwoInts service and the Fibonacci action, illustrate ROS 2 concepts; the runnable robot code is in the repository's lab scripts and skill servers.
What you will learn
- Connect an LLM agent to a ROS 2 robot through tool calls
- Collect a demonstration dataset from teleoperation
- Fine-tune and deploy a Vision-Language-Action model (SmolVLA)
- Run a planner-to-executor mission end to end on a simulated Husky
- Clone and run the whole system from the course GitHub repository
Course content
- Welcome & Setup
- Introduction to ROS 2
- Agentic AI & Tool Calling
- From VLM to VLA
Instructor: Anis Koubaa