User profiles for Norman Di Palo

Norman Di Palo

Google DeepMind
Verified email at deepmind.com
Cited by 3948

Keypoint action tokens enable in-context imitation learning in robotics

N Di Palo, E Johns - arXiv preprint arXiv:2403.19578, 2024 - arxiv.org
We show that off-the-shelf text-based Transformers, with no additional training, can perform
few-shot in-context visual imitation learning, mapping visual observations to action …

Towards a unified agent with foundation models

N Di Palo, A Byravan, L Hasenclever… - arXiv preprint arXiv …, 2023 - arxiv.org
Language Models and Vision Language Models have recently demonstrated unprecedented
capabilities in terms of understanding human intentions, reasoning, scene understanding, …

Gemini robotics: Bringing ai into the physical world

…, S Dasari, T Davchev, C Devin, N Di Palo… - arXiv preprint arXiv …, 2025 - arxiv.org
Recent advancements in large multimodal models have led to the emergence of remarkable
generalist capabilities in digital domains, yet their translation to physical agents such as …

Open x-embodiment: Robotic learning datasets and rt-x models

…, J Han, J Kim, JJ Lim, E Johns, N Di Palo… - … for Scalable Skill …, 2023 - openreview.net
Large, high-capacity models trained on diverse datasets have shown remarkable successes
on efficiently tackling downstream applications. In domains from NLP to Computer Vision, …

Dinobot: Robot manipulation via retrieval and alignment with vision foundation models

N Di Palo, E Johns - 2024 IEEE International Conference on …, 2024 - ieeexplore.ieee.org
We propose DINOBot, a novel imitation learning framework for robot manipulation, which
leverages the image-level and pixel-level capabilities of features extracted from Vision …

Language models as zero-shot trajectory generators

T Kwon, N Di Palo, E Johns - IEEE Robotics and Automation …, 2024 - ieeexplore.ieee.org
Large Language Models (LLMs) have recently shown promise as high-level planners for
robots when given access to a selection of low-level skills. However, it is often assumed that …

Gemini robotics 1.5: Pushing the frontier of generalist robots with advanced embodied reasoning, thinking, and motion transfer

…, T Davchev, MK Dave, C Devin, N Di Palo… - arXiv preprint arXiv …, 2025 - arxiv.org
General-purpose robots need a deep understanding of the physical world, advanced reasoning,
and general and dexterous control. This report introduces the latest generation of the …

R+ x: Retrieval and execution from everyday human videos

G Papagiannis, N Di Palo, P Vitiello… - 2025 IEEE International …, 2025 - ieeexplore.ieee.org
We present $\mathbf{R}+\mathbf{X}$ , a framework which enables robots to learn skills from
long, unlabelled, first-person videos of humans performing everyday tasks. Given a …

Learning multi-stage tasks with one demonstration via self-replay

N Di Palo, E Johns - Conference on Robot Learning, 2022 - proceedings.mlr.press
In this work, we introduce a novel method to learn everyday-like multi-stage tasks from a
single human demonstration, without requiring any prior object knowledge. Inspired by the …

On the effectiveness of retrieval, alignment, and replay in manipulation

N Di Palo, E Johns - IEEE Robotics and Automation Letters, 2024 - ieeexplore.ieee.org
Imitation learning with visual observations is notoriously inefficient when addressed with end-to-end
behavioural cloning methods. In this letter, we explore an alternative paradigm …