Google launches RT-2 or Robotics Transformer 2, a “vision-language-action” model trained on text and images from the web that can output robotic actions
Our sneak peek into Google's new robotics model, RT-2, which melds artificial intelligence technology with robots.
Can you imagine those in a library helping with shelf-reading, re-shelving, and doing chats! Aided by A.I. Language Models, Google's Robots Are Getting Smart https://www.nytimes.com/...
Across all categories, we saw increased generalisation performance compared to previous baselines, such as on RT-1 models. We also evaluated RT-2 on a number of unseen objects and environments where it could successfully adapt to new situations: https://dpmd.ai/...
⚪ To explore RT-2's emergent capabilities in trials, we searched for tasks that require combining learnings from web data and the robot's experience. We then defined 3 skills it needed to show: 🔘 Symbol understanding 🔘 Reasoning 🔘 Human recognition https://dpmd.ai/... [image]
Yep, here we go... LLMs plugged into robots -> Aided by A.I. Language Models, Google's Robots Are Getting Smart “Google has recently begun plugging state-of-the-art language models into its robots, giving them the equivalent of artificial brains.” https://www.nytimes.com/... [ima…
Today, we announced 𝗥𝗧-𝟮: a first of its kind vision-language-action model to control robots. 🤖 It learns from both web and robotics data and translates this knowledge into generalised instructions. Find out more: https://dpmd.ai/... [video]