User profiles for Norman Di Palo
Norman Di PaloGoogle DeepMind Verified email at deepmind.com Cited by 3948 |
Keypoint action tokens enable in-context imitation learning in robotics
We show that off-the-shelf text-based Transformers, with no additional training, can perform
few-shot in-context visual imitation learning, mapping visual observations to action …
few-shot in-context visual imitation learning, mapping visual observations to action …
Towards a unified agent with foundation models
Language Models and Vision Language Models have recently demonstrated unprecedented
capabilities in terms of understanding human intentions, reasoning, scene understanding, …
capabilities in terms of understanding human intentions, reasoning, scene understanding, …
Gemini robotics: Bringing ai into the physical world
Recent advancements in large multimodal models have led to the emergence of remarkable
generalist capabilities in digital domains, yet their translation to physical agents such as …
generalist capabilities in digital domains, yet their translation to physical agents such as …
Open x-embodiment: Robotic learning datasets and rt-x models
Large, high-capacity models trained on diverse datasets have shown remarkable successes
on efficiently tackling downstream applications. In domains from NLP to Computer Vision, …
on efficiently tackling downstream applications. In domains from NLP to Computer Vision, …
Dinobot: Robot manipulation via retrieval and alignment with vision foundation models
We propose DINOBot, a novel imitation learning framework for robot manipulation, which
leverages the image-level and pixel-level capabilities of features extracted from Vision …
leverages the image-level and pixel-level capabilities of features extracted from Vision …
Language models as zero-shot trajectory generators
Large Language Models (LLMs) have recently shown promise as high-level planners for
robots when given access to a selection of low-level skills. However, it is often assumed that …
robots when given access to a selection of low-level skills. However, it is often assumed that …
Gemini robotics 1.5: Pushing the frontier of generalist robots with advanced embodied reasoning, thinking, and motion transfer
General-purpose robots need a deep understanding of the physical world, advanced reasoning,
and general and dexterous control. This report introduces the latest generation of the …
and general and dexterous control. This report introduces the latest generation of the …
R+ x: Retrieval and execution from everyday human videos
We present $\mathbf{R}+\mathbf{X}$ , a framework which enables robots to learn skills from
long, unlabelled, first-person videos of humans performing everyday tasks. Given a …
long, unlabelled, first-person videos of humans performing everyday tasks. Given a …
Learning multi-stage tasks with one demonstration via self-replay
In this work, we introduce a novel method to learn everyday-like multi-stage tasks from a
single human demonstration, without requiring any prior object knowledge. Inspired by the …
single human demonstration, without requiring any prior object knowledge. Inspired by the …
On the effectiveness of retrieval, alignment, and replay in manipulation
Imitation learning with visual observations is notoriously inefficient when addressed with end-to-end
behavioural cloning methods. In this letter, we explore an alternative paradigm …
behavioural cloning methods. In this letter, we explore an alternative paradigm …