RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
TL;DR
RT-2 is the first large-scale Vision-Language-Action (VLA) model that directly applies internet-scale VLM knowledge to robotic control. By representing robot actions as text tokens and co-fine-tuning PaLI-X (55B) and PaLM-E (12B) on web and rob...
telos-robotics.hashnode.dev5 min read