Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 100 results for author: Vondrick, C

Searching in archive cs. Search in all archives.
.
  1. arXiv:2606.21572  [pdf, ps, other

    cs.RO

    Robot Critics that Sweat the Small Stuff

    Authors: Sruthi Sudhakar, Junbang Liang, Sreehari Rammohan, Pavel Tokmakov, Richard Zemel, Carl Vondrick

    Abstract: Large vision-language models contain several priors about the world and object interactions, making them useful critics during inference to steer robot policies towards success. However, closed-loop robot manipulation requires judging small visual differences between success and failure, which remains a challenge for current VLMs. We introduce a method to fine-tune critics by constructing pairwise… ▽ More

    Submitted 19 June, 2026; originally announced June 2026.

  2. arXiv:2606.03148  [pdf, ps, other

    cs.CV

    $A^2$: Smaller Self-Supervised ViTs Localize Better than Larger Ones

    Authors: Sreehari Rammohan, Huy Ha, Carl Vondrick

    Abstract: Robust visual classification often depends on localizing the main foreground objects in an image while ignoring contextual distractors. Surprisingly, we find that the attention maps of smaller self-supervised ViTs localize foreground objects better than those of larger ViTs. However, we still need large ViTs, because they extract richer representations from each patch. To get the best of both worl… ▽ More

    Submitted 2 June, 2026; originally announced June 2026.

  3. arXiv:2605.09693  [pdf, ps, other

    cs.CV cs.AI cs.LG

    Do multimodal models imagine electric sheep?

    Authors: Santhosh Kumar Ramakrishnan, Carl Vondrick, Raja Giryes, Philipp Krähenbühl, Vladlen Koltun

    Abstract: Yes. We find that large multimodal models develop mental imagery when solving spatial puzzles, and they do imagine sheep when solving sheep puzzles. We fine-tune a Qwen3.5 VLM to solve twelve diverse visual reasoning tasks -- including tangram, jigsaw, sokoban, 3D mental rotation, and rush hour -- that require understanding geometry, spatial relationships, and the consequences of actions. By super… ▽ More

    Submitted 10 May, 2026; originally announced May 2026.

  4. arXiv:2602.12112  [pdf, ps, other

    cs.LG

    Few-Shot Design Optimization by Exploiting Auxiliary Information

    Authors: Arjun Mani, Carl Vondrick, Richard Zemel

    Abstract: Many real-world design problems involve optimizing an expensive black-box function $f(x)$, such as hardware design or drug discovery. Bayesian Optimization has emerged as a sample-efficient framework for this problem. However, the basic setting considered by these methods is simplified compared to real-world experimental setups, where experiments often generate a wealth of useful information. We i… ▽ More

    Submitted 12 February, 2026; originally announced February 2026.

  5. arXiv:2512.19648  [pdf, ps, other

    cs.CV

    4D Gaussian Splatting as a Learned Dynamical System

    Authors: Arnold Caleb Asiimwe, Carl Vondrick

    Abstract: We reinterpret 4D Gaussian Splatting as a continuous-time dynamical system, where scene motion arises from integrating a learned neural dynamical field rather than applying per-frame deformations. This formulation, which we call EvoGS, treats the Gaussian representation as an evolving physical system whose state evolves continuously under a learned motion law. This unlocks capabilities absent in d… ▽ More

    Submitted 22 December, 2025; originally announced December 2025.

  6. arXiv:2511.20544  [pdf, ps, other

    cs.CV cs.AI cs.LG

    New York Smells: A Large Multimodal Dataset for Olfaction

    Authors: Ege Ozguroglu, Junbang Liang, Ruoshi Liu, Mia Chiquier, Michael DeTienne, Wesley Wei Qian, Alexandra Horowitz, Andrew Owens, Carl Vondrick

    Abstract: While olfaction is central to how animals perceive the world, this rich chemical sensory modality remains largely inaccessible to machines. One key bottleneck is the lack of diverse, multimodal olfactory training data collected in natural settings. We present New York Smells, a large dataset of paired image and olfactory signals captured ``in the wild.'' Our dataset contains 7,000 smell-image pair… ▽ More

    Submitted 2 August, 2026; v1 submitted 25 November, 2025; originally announced November 2025.

    Comments: Project website at https://smell.cs.columbia.edu

  7. arXiv:2509.07680  [pdf, ps, other

    cs.CV cs.LG

    CAViAR: Critic-Augmented Video Agentic Reasoning

    Authors: Sachit Menon, Ahmet Iscen, Arsha Nagrani, Tobias Weyand, Carl Vondrick, Cordelia Schmid

    Abstract: Video understanding has seen significant progress in recent years, with models' performance on perception from short clips continuing to rise. Yet, multiple recent benchmarks, such as LVBench, Neptune, and ActivityNet-RTL, show performance wanes for tasks requiring complex reasoning on videos as queries grow more complex and videos grow longer. In this work, we ask: can existing perception capabil… ▽ More

    Submitted 9 September, 2025; originally announced September 2025.

  8. arXiv:2508.00795  [pdf, ps, other

    cs.RO

    Video Generators are Robot Policies

    Authors: Junbang Liang, Pavel Tokmakov, Ruoshi Liu, Sruthi Sudhakar, Paarth Shah, Rares Ambrus, Carl Vondrick

    Abstract: Despite tremendous progress in dexterous manipulation, current visuomotor policies remain fundamentally limited by two challenges: they struggle to generalize under perceptual or behavioral distribution shifts, and their performance is constrained by the size of human demonstration data. In this paper, we use video generation as a proxy for robot policy learning to address both limitations simulta… ▽ More

    Submitted 1 August, 2025; originally announced August 2025.

  9. arXiv:2505.00681  [pdf, ps, other

    cs.LG cs.CV

    MINERVA: Evaluating Complex Video Reasoning

    Authors: Arsha Nagrani, Sachit Menon, Ahmet Iscen, Shyamal Buch, Ramin Mehran, Nilpa Jha, Anja Hauth, Yukun Zhu, Carl Vondrick, Mikhail Sirotenko, Cordelia Schmid, Tobias Weyand

    Abstract: Multimodal LLMs are turning their focus to video benchmarks, however most video benchmarks only provide outcome supervision, with no intermediate or interpretable reasoning steps. This makes it challenging to assess if models are truly able to combine perceptual and temporal information to reason about videos, or simply get the correct answer by chance or by exploiting linguistic biases. To remedy… ▽ More

    Submitted 1 May, 2025; originally announced May 2025.

  10. arXiv:2504.12110  [pdf, ps, other

    cs.AI

    Towards LLM Agents for Earth Observation

    Authors: Chia Hsiang Kao, Wenting Zhao, Shreelekha Revankar, Samuel Speas, Snehal Bhagat, Rajeev Datta, Cheng Perng Phoo, Utkarsh Mall, Carl Vondrick, Kavita Bala, Bharath Hariharan

    Abstract: Earth Observation (EO) provides critical planetary data for environmental monitoring, disaster management, climate science, and other scientific domains. Here we ask: Are AI systems ready for reliable Earth Observation? We introduce \datasetnamenospace, a benchmark of 140 yes/no questions from NASA Earth Observatory articles across 13 topics and 17 satellite sensors. Using Google Earth Engine API… ▽ More

    Submitted 12 September, 2025; v1 submitted 16 April, 2025; originally announced April 2025.

    Comments: Accepted at ICML 2025 Workshop TerraBytes

  11. arXiv:2504.09737  [pdf, other

    cs.AI cs.CL cs.HC cs.LG

    Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025

    Authors: Nitya Thakkar, Mert Yuksekgonul, Jake Silberg, Animesh Garg, Nanyun Peng, Fei Sha, Rose Yu, Carl Vondrick, James Zou

    Abstract: Peer review at AI conferences is stressed by rapidly rising submission volumes, leading to deteriorating review quality and increased author dissatisfaction. To address these issues, we developed Review Feedback Agent, a system leveraging multiple large language models (LLMs) to improve review clarity and actionability by providing automated feedback on vague comments, content misunderstandings, a… ▽ More

    Submitted 13 April, 2025; originally announced April 2025.

    Comments: 30 pages, 7 figures

  12. arXiv:2504.08046  [pdf, other

    cs.CV

    Teaching Humans Subtle Differences with DIFFusion

    Authors: Mia Chiquier, Orr Avrech, Yossi Gandelsman, Berthy Feng, Katherine Bouman, Carl Vondrick

    Abstract: Scientific expertise often requires recognizing subtle visual differences that remain challenging to articulate even for domain experts. We present a system that leverages generative models to automatically discover and visualize minimal discriminative features between categories while preserving instance identity. Our method generates counterfactual visualizations with subtle, targeted transforma… ▽ More

    Submitted 15 May, 2025; v1 submitted 10 April, 2025; originally announced April 2025.

  13. arXiv:2502.10060  [pdf, other

    cs.CV cs.LG

    DiSciPLE: Learning Interpretable Programs for Scientific Visual Discovery

    Authors: Utkarsh Mall, Cheng Perng Phoo, Mia Chiquier, Bharath Hariharan, Kavita Bala, Carl Vondrick

    Abstract: Visual data is used in numerous different scientific workflows ranging from remote sensing to ecology. As the amount of observation data increases, the challenge is not just to make accurate predictions but also to understand the underlying mechanisms for those predictions. Good interpretation is important in scientific workflows, as it allows for better decision-making by providing insights into… ▽ More

    Submitted 14 February, 2025; originally announced February 2025.

  14. arXiv:2502.01980  [pdf, ps, other

    cs.LG cs.AI

    Generative Data Mining with Longtail-Guided Diffusion

    Authors: David S. Hayden, Mao Ye, Timur Garipov, Gregory P. Meyer, Carl Vondrick, Zhao Chen, Yuning Chai, Eric Wolff, Siddhartha S. Srinivasa

    Abstract: It is difficult to anticipate the myriad challenges that a predictive model will encounter once deployed. Common practice entails a reactive, cyclical approach: model deployment, data mining, and retraining. We instead develop a proactive longtail discovery process by imagining additional data during training. In particular, we develop general model-based longtail signals, including a differentiab… ▽ More

    Submitted 26 June, 2025; v1 submitted 3 February, 2025; originally announced February 2025.

    Comments: 20 pages

    Journal ref: Proceedings of the 42nd International Conference on Machine Learning (ICML), 2025

  15. arXiv:2412.02971  [pdf, other

    cs.CV

    MedAutoCorrect: Image-Conditioned Autocorrection in Medical Reporting

    Authors: Arnold Caleb Asiimwe, Dídac Surís, Pranav Rajpurkar, Carl Vondrick

    Abstract: In medical reporting, the accuracy of radiological reports, whether generated by humans or machine learning algorithms, is critical. We tackle a new task in this paper: image-conditioned autocorrection of inaccuracies within these reports. Using the MIMIC-CXR dataset, we first intentionally introduce a diverse range of errors into reports. Subsequently, we propose a two-stage framework capable of… ▽ More

    Submitted 3 December, 2024; originally announced December 2024.

  16. arXiv:2410.18969  [pdf, other

    cs.RO

    Self-Improving Autonomous Underwater Manipulation

    Authors: Ruoshi Liu, Huy Ha, Mengxue Hou, Shuran Song, Carl Vondrick

    Abstract: Underwater robotic manipulation faces significant challenges due to complex fluid dynamics and unstructured environments, causing most manipulation systems to rely heavily on human teleoperation. In this paper, we introduce AquaBot, a fully autonomous manipulation system that combines behavior cloning from human demonstrations with self-learning optimization to improve beyond human teleoperation p… ▽ More

    Submitted 24 October, 2024; originally announced October 2024.

    Comments: Project Page: https://aquabot.cs.columbia.edu/

  17. arXiv:2410.13851  [pdf, other

    cs.RO cs.CV cs.GR

    Differentiable Robot Rendering

    Authors: Ruoshi Liu, Alper Canberk, Shuran Song, Carl Vondrick

    Abstract: Vision foundation models trained on massive amounts of visual data have shown unprecedented reasoning and planning skills in open-world settings. A key challenge in applying them to robotic tasks is the modality gap between visual data and action data. We introduce differentiable robot rendering, a method allowing the visual appearance of a robot body to be directly differentiable with respect to… ▽ More

    Submitted 17 October, 2024; originally announced October 2024.

    Comments: Project Page: https://drrobot.cs.columbia.edu/

  18. arXiv:2409.00522  [pdf, other

    cs.CV

    EraseDraw: Learning to Draw Step-by-Step via Erasing Objects from Images

    Authors: Alper Canberk, Maksym Bondarenko, Ege Ozguroglu, Ruoshi Liu, Carl Vondrick

    Abstract: Creative processes such as painting often involve creating different components of an image one by one. Can we build a computational model to perform this task? Prior works often fail by making global changes to the image, inserting objects in unrealistic spatial locations, and generating inaccurate lighting details. We observe that while state-of-the-art models perform poorly on object insertion,… ▽ More

    Submitted 23 December, 2024; v1 submitted 31 August, 2024; originally announced September 2024.

  19. arXiv:2408.07147  [pdf, other

    cs.CV

    Controlling the World by Sleight of Hand

    Authors: Sruthi Sudhakar, Ruoshi Liu, Basile Van Hoorick, Carl Vondrick, Richard Zemel

    Abstract: Humans naturally build mental models of object interactions and dynamics, allowing them to imagine how their surroundings will change if they take a certain action. While generative models today have shown impressive results on generating/editing images unconditionally or conditioned on text, current methods do not provide the ability to perform object manipulation conditioned on actions, an impor… ▽ More

    Submitted 13 August, 2024; originally announced August 2024.

  20. arXiv:2406.16862  [pdf, other

    cs.RO cs.CV

    Dreamitate: Real-World Visuomotor Policy Learning via Video Generation

    Authors: Junbang Liang, Ruoshi Liu, Ege Ozguroglu, Sruthi Sudhakar, Achal Dave, Pavel Tokmakov, Shuran Song, Carl Vondrick

    Abstract: A key challenge in manipulation is learning a policy that can robustly generalize to diverse visual environments. A promising mechanism for learning robust policies is to leverage video generative models, which are pretrained on large-scale datasets of internet videos. In this paper, we propose a visuomotor policy learning framework that fine-tunes a video diffusion model on human demonstrations o… ▽ More

    Submitted 24 June, 2024; originally announced June 2024.

    Comments: Project page: https://dreamitate.cs.columbia.edu/

  21. arXiv:2406.14562  [pdf, other

    cs.CL cs.AI cs.CV

    Whiteboard-of-Thought: Thinking Step-by-Step Across Modalities

    Authors: Sachit Menon, Richard Zemel, Carl Vondrick

    Abstract: When presented with questions involving visual thinking, humans naturally switch reasoning modalities, often forming mental images or drawing visual aids. Large language models have shown promising results in arithmetic and symbolic reasoning by expressing intermediate reasoning in text as a chain of thought, yet struggle to extend this capability to answer text queries that are easily solved by v… ▽ More

    Submitted 20 June, 2024; originally announced June 2024.

    Comments: Project website: whiteboard.cs.columbia.edu/

  22. arXiv:2406.00955  [pdf, other

    cs.CV

    How Video Meetings Change Your Expression

    Authors: Sumit Sarin, Utkarsh Mall, Purva Tendulkar, Carl Vondrick

    Abstract: Do our facial expressions change when we speak over video calls? Given two unpaired sets of videos of people, we seek to automatically find spatio-temporal patterns that are distinctive of each set. Existing methods use discriminative approaches and perform post-hoc explainability analysis. Such methods are insufficient as they are unable to provide insights beyond obvious dataset biases, and the… ▽ More

    Submitted 2 June, 2024; originally announced June 2024.

    Comments: Project webpage is available at: https://facet.cs.columbia.edu

  23. arXiv:2405.14868  [pdf, other

    cs.CV cs.AI cs.LG cs.RO

    Generative Camera Dolly: Extreme Monocular Dynamic Novel View Synthesis

    Authors: Basile Van Hoorick, Rundi Wu, Ege Ozguroglu, Kyle Sargent, Ruoshi Liu, Pavel Tokmakov, Achal Dave, Changxi Zheng, Carl Vondrick

    Abstract: Accurate reconstruction of complex dynamic scenes from just a single viewpoint continues to be a challenging task in computer vision. Current dynamic novel view synthesis methods typically require videos from many different camera viewpoints, necessitating careful recording setups, and significantly restricting their utility in the wild as well as in terms of embodied AI applications. In this pape… ▽ More

    Submitted 5 July, 2024; v1 submitted 23 May, 2024; originally announced May 2024.

    Comments: Accepted to ECCV 2024. Project webpage is available at: https://gcd.cs.columbia.edu/

  24. arXiv:2404.09941  [pdf, other

    cs.CV cs.AI

    Evolving Interpretable Visual Classifiers with Large Language Models

    Authors: Mia Chiquier, Utkarsh Mall, Carl Vondrick

    Abstract: Multimodal pre-trained models, such as CLIP, are popular for zero-shot classification due to their open-vocabulary flexibility and high performance. However, vision-language models, which compute similarity scores between images and class labels, are largely black-box, with limited interpretability, risk for bias, and inability to discover new visual concepts not written down. Moreover, in practic… ▽ More

    Submitted 15 April, 2024; originally announced April 2024.

  25. arXiv:2403.10949  [pdf, other

    cs.CL cs.AI cs.LG

    SelfIE: Self-Interpretation of Large Language Model Embeddings

    Authors: Haozhe Chen, Carl Vondrick, Chengzhi Mao

    Abstract: How do large language models (LLMs) obtain their answers? The ability to explain and control an LLM's reasoning process is key for reliability, transparency, and future model developments. We propose SelfIE (Self-Interpretation of Embeddings), a framework that enables LLMs to interpret their own embeddings in natural language by leveraging their ability to respond to inquiries about a given passag… ▽ More

    Submitted 25 March, 2024; v1 submitted 16 March, 2024; originally announced March 2024.

  26. arXiv:2403.09566  [pdf, other

    cs.RO

    PaperBot: Learning to Design Real-World Tools Using Paper

    Authors: Ruoshi Liu, Junbang Liang, Sruthi Sudhakar, Huy Ha, Cheng Chi, Shuran Song, Carl Vondrick

    Abstract: Paper is a cheap, recyclable, and clean material that is often used to make practical tools. Traditional tool design either relies on simulation or physical analysis, which is often inaccurate and time-consuming. In this paper, we propose PaperBot, an approach that directly learns to design and use a tool in the real world using paper without human intervention. We demonstrated the effectiveness a… ▽ More

    Submitted 14 March, 2024; originally announced March 2024.

    Comments: Project Website: https://paperbot.cs.columbia.edu/

  27. arXiv:2402.10128  [pdf, other

    cs.CV cs.GR cs.LG

    GES: Generalized Exponential Splatting for Efficient Radiance Field Rendering

    Authors: Abdullah Hamdi, Luke Melas-Kyriazi, Jinjie Mai, Guocheng Qian, Ruoshi Liu, Carl Vondrick, Bernard Ghanem, Andrea Vedaldi

    Abstract: Advancements in 3D Gaussian Splatting have significantly accelerated 3D reconstruction and generation. However, it may require a large number of Gaussians, which creates a substantial memory footprint. This paper introduces GES (Generalized Exponential Splatting), a novel representation that employs Generalized Exponential Function (GEF) to model 3D scenes, requiring far fewer particles to represe… ▽ More

    Submitted 24 May, 2024; v1 submitted 15 February, 2024; originally announced February 2024.

    Comments: CVPR 2024 paper. project website https://abdullahamdi.com/ges

  28. arXiv:2401.14398  [pdf, other

    cs.CV cs.LG

    pix2gestalt: Amodal Segmentation by Synthesizing Wholes

    Authors: Ege Ozguroglu, Ruoshi Liu, Dídac Surís, Dian Chen, Achal Dave, Pavel Tokmakov, Carl Vondrick

    Abstract: We introduce pix2gestalt, a framework for zero-shot amodal segmentation, which learns to estimate the shape and appearance of whole objects that are only partially visible behind occlusions. By capitalizing on large-scale diffusion models and transferring their representations to this task, we learn a conditional diffusion model for reconstructing whole objects in challenging zero-shot cases, incl… ▽ More

    Submitted 25 January, 2024; originally announced January 2024.

    Comments: Website: https://gestalt.cs.columbia.edu/

  29. arXiv:2401.12970  [pdf, other

    cs.CL

    Raidar: geneRative AI Detection viA Rewriting

    Authors: Chengzhi Mao, Carl Vondrick, Hao Wang, Junfeng Yang

    Abstract: We find that large language models (LLMs) are more likely to modify human-written text than AI-generated text when tasked with rewriting. This tendency arises because LLMs often perceive AI-generated text as high-quality, leading to fewer modifications. We introduce a method to detect AI-generated content by prompting LLMs to rewrite text and calculating the editing distance of the output. We dubb… ▽ More

    Submitted 14 April, 2024; v1 submitted 23 January, 2024; originally announced January 2024.

    Comments: Accepted by ICLR 2024, Large Language Models, Detection

  30. arXiv:2312.06960  [pdf, other

    cs.CV cs.LG

    Remote Sensing Vision-Language Foundation Models without Annotations via Ground Remote Alignment

    Authors: Utkarsh Mall, Cheng Perng Phoo, Meilin Kelsey Liu, Carl Vondrick, Bharath Hariharan, Kavita Bala

    Abstract: We introduce a method to train vision-language models for remote-sensing images without using any textual annotations. Our key insight is to use co-located internet imagery taken on the ground as an intermediary for connecting remote-sensing images and language. Specifically, we train an image encoder for remote sensing images to align with the image encoder of CLIP using a large amount of paired… ▽ More

    Submitted 11 December, 2023; originally announced December 2023.

  31. arXiv:2310.10591  [pdf, other

    cs.CV

    Interpreting and Controlling Vision Foundation Models via Text Explanations

    Authors: Haozhe Chen, Junfeng Yang, Carl Vondrick, Chengzhi Mao

    Abstract: Large-scale pre-trained vision foundation models, such as CLIP, have become de facto backbones for various vision tasks. However, due to their black-box nature, understanding the underlying rules behind these models' predictions and controlling model behaviors have remained open challenges. We present a framework for interpreting vision transformer's latent tokens with natural language. Given a la… ▽ More

    Submitted 16 October, 2023; originally announced October 2023.

  32. arXiv:2309.05810  [pdf, other

    cs.CV cs.CR cs.LG cs.RO

    SHIFT3D: Synthesizing Hard Inputs For Tricking 3D Detectors

    Authors: Hongge Chen, Zhao Chen, Gregory P. Meyer, Dennis Park, Carl Vondrick, Ashish Shrivastava, Yuning Chai

    Abstract: We present SHIFT3D, a differentiable pipeline for generating 3D shapes that are structurally plausible yet challenging to 3D object detectors. In safety-critical applications like autonomous driving, discovering such novel challenging objects can offer insight into unknown vulnerabilities of 3D detectors. By representing objects with a signed distanced function (SDF), we show that gradient error s… ▽ More

    Submitted 11 September, 2023; originally announced September 2023.

    Comments: Accepted by ICCV 2023

  33. arXiv:2307.05663  [pdf, other

    cs.CV cs.AI

    Objaverse-XL: A Universe of 10M+ 3D Objects

    Authors: Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram Voleti, Samir Yitzhak Gadre, Eli VanderBilt, Aniruddha Kembhavi, Carl Vondrick, Georgia Gkioxari, Kiana Ehsani, Ludwig Schmidt, Ali Farhadi

    Abstract: Natural language processing and 2D vision models have attained remarkable proficiency on many tasks primarily by escalating the scale of training data. However, 3D vision tasks have not seen the same progress, in part due to the challenges of acquiring high-quality 3D data. In this work, we present Objaverse-XL, a dataset of over 10 million 3D objects. Our dataset comprises deduplicated 3D objects… ▽ More

    Submitted 11 July, 2023; originally announced July 2023.

  34. arXiv:2305.15399  [pdf, other

    cs.CV cs.AI cs.GR

    Sin3DM: Learning a Diffusion Model from a Single 3D Textured Shape

    Authors: Rundi Wu, Ruoshi Liu, Carl Vondrick, Changxi Zheng

    Abstract: Synthesizing novel 3D models that resemble the input example has long been pursued by graphics artists and machine learning researchers. In this paper, we present Sin3DM, a diffusion model that learns the internal patch distribution from a single 3D textured shape and generates high-quality variations with fine geometry and texture details. Training a diffusion model directly in 3D would induce la… ▽ More

    Submitted 20 February, 2024; v1 submitted 24 May, 2023; originally announced May 2023.

    Comments: Accepted to ICLR 2024. Project page: https://Sin3DM.github.io, Code: https://github.com/Sin3DM/Sin3DM

  35. arXiv:2305.03052  [pdf, other

    cs.CV cs.AI cs.LG cs.RO

    Tracking through Containers and Occluders in the Wild

    Authors: Basile Van Hoorick, Pavel Tokmakov, Simon Stent, Jie Li, Carl Vondrick

    Abstract: Tracking objects with persistence in cluttered and dynamic environments remains a difficult challenge for computer vision systems. In this paper, we introduce $\textbf{TCOW}$, a new benchmark and model for visual tracking through heavy occlusion and containment. We set up a task where the goal is to, given a video sequence, segment both the projected extent of the target object, as well as the sur… ▽ More

    Submitted 4 May, 2023; originally announced May 2023.

    Comments: Accepted at CVPR 2023. Project webpage is available at: https://tcow.cs.columbia.edu/

  36. arXiv:2305.01652  [pdf, other

    cs.CV

    Humans as Light Bulbs: 3D Human Reconstruction from Thermal Reflection

    Authors: Ruoshi Liu, Carl Vondrick

    Abstract: The relatively hot temperature of the human body causes people to turn into long-wave infrared light sources. Since this emitted light has a larger wavelength than visible light, many surfaces in typical scenes act as infrared mirrors with strong specular reflections. We exploit the thermal reflections of a person onto objects in order to locate their position and reconstruct their pose, even if t… ▽ More

    Submitted 2 May, 2023; originally announced May 2023.

    Comments: Website: https://thermal.cs.columbia.edu/

  37. arXiv:2304.06197  [pdf, other

    cs.LG physics.flu-dyn

    SURFSUP: Learning Fluid Simulation for Novel Surfaces

    Authors: Arjun Mani, Ishaan Preetam Chandratreya, Elliot Creager, Carl Vondrick, Richard Zemel

    Abstract: Modeling the mechanics of fluid in complex scenes is vital to applications in design, graphics, and robotics. Learning-based methods provide fast and differentiable fluid simulators, however most prior work is unable to accurately model how fluids interact with genuinely novel surfaces not seen during training. We introduce SURFSUP, a framework that represents objects implicitly using signed dista… ▽ More

    Submitted 8 September, 2023; v1 submitted 12 April, 2023; originally announced April 2023.

    Comments: Website: https://surfsup.cs.columbia.edu/

  38. arXiv:2303.11328  [pdf, other

    cs.CV cs.GR cs.RO

    Zero-1-to-3: Zero-shot One Image to 3D Object

    Authors: Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, Carl Vondrick

    Abstract: We introduce Zero-1-to-3, a framework for changing the camera viewpoint of an object given just a single RGB image. To perform novel view synthesis in this under-constrained setting, we capitalize on the geometric priors that large-scale diffusion models learn about natural images. Our conditional diffusion model uses a synthetic dataset to learn controls of the relative camera viewpoint, which al… ▽ More

    Submitted 20 March, 2023; originally announced March 2023.

    Comments: Website: https://zero123.cs.columbia.edu/

  39. arXiv:2303.08128  [pdf, other

    cs.CV

    ViperGPT: Visual Inference via Python Execution for Reasoning

    Authors: Dídac Surís, Sachit Menon, Carl Vondrick

    Abstract: Answering visual queries is a complex task that requires both visual processing and reasoning. End-to-end models, the dominant approach for this task, do not explicitly differentiate between the two, limiting interpretability and generalization. Learning modular programs presents a promising alternative, but has proven challenging due to the difficulty of learning both the programs and modules sim… ▽ More

    Submitted 14 March, 2023; originally announced March 2023.

    Comments: Website: https://viper.cs.columbia.edu/

  40. arXiv:2301.10939  [pdf, other

    cs.CV cs.CL cs.LG

    Affective Faces for Goal-Driven Dyadic Communication

    Authors: Scott Geng, Revant Teotia, Purva Tendulkar, Sachit Menon, Carl Vondrick

    Abstract: We introduce a video framework for modeling the association between verbal and non-verbal communication during dyadic conversation. Given the input speech of a speaker, our approach retrieves a video of a listener, who has facial expressions that would be socially appropriate given the context. Our approach further allows the listener to be conditioned on their own goals, personalities, or backgro… ▽ More

    Submitted 26 January, 2023; originally announced January 2023.

  41. arXiv:2212.07815  [pdf, other

    cs.CV

    Adversarially Robust Video Perception by Seeing Motion

    Authors: Lingyu Zhang, Chengzhi Mao, Junfeng Yang, Carl Vondrick

    Abstract: Despite their excellent performance, state-of-the-art computer vision models often fail when they encounter adversarial examples. Video perception models tend to be more fragile under attacks, because the adversary has more places to manipulate in high-dimensional data. In this paper, we find one reason for video models' vulnerability is that they fail to perceive the correct motion under adversar… ▽ More

    Submitted 12 December, 2022; originally announced December 2022.

  42. arXiv:2212.07016  [pdf, other

    cs.CV

    Understanding Zero-Shot Adversarial Robustness for Large-Scale Models

    Authors: Chengzhi Mao, Scott Geng, Junfeng Yang, Xin Wang, Carl Vondrick

    Abstract: Pretrained large-scale vision-language models like CLIP have exhibited strong generalization over unseen tasks. Yet imperceptible adversarial perturbations can significantly reduce CLIP's performance on new tasks. In this work, we identify and explore the problem of \emph{adapting large-scale models for zero-shot adversarial robustness}. We first identify two key factors during model adaption -- t… ▽ More

    Submitted 21 April, 2023; v1 submitted 13 December, 2022; originally announced December 2022.

  43. arXiv:2212.06202  [pdf, other

    cs.CV

    Doubly Right Object Recognition: A Why Prompt for Visual Rationales

    Authors: Chengzhi Mao, Revant Teotia, Amrutha Sundar, Sachit Menon, Junfeng Yang, Xin Wang, Carl Vondrick

    Abstract: Many visual recognition models are evaluated only on their classification accuracy, a metric for which they obtain strong performance. In this paper, we investigate whether computer vision models can also provide correct rationales for their predictions. We propose a ``doubly right'' object recognition benchmark, where the metric requires the model to simultaneously produce both the right labels a… ▽ More

    Submitted 22 March, 2023; v1 submitted 12 December, 2022; originally announced December 2022.

    Comments: Accepted at CVPR 2023

  44. arXiv:2212.06079  [pdf, other

    cs.CV

    Robust Perception through Equivariance

    Authors: Chengzhi Mao, Lingyu Zhang, Abhishek Joshi, Junfeng Yang, Hao Wang, Carl Vondrick

    Abstract: Deep networks for computer vision are not reliable when they encounter adversarial examples. In this paper, we introduce a framework that uses the dense intrinsic constraints in natural images to robustify inference. By introducing constraints at inference time, we can shift the burden of robustness from training to the inference algorithm, thereby allowing the model to adjust dynamically to each… ▽ More

    Submitted 3 June, 2023; v1 submitted 12 December, 2022; originally announced December 2022.

    Comments: Published in ICML 2023

  45. arXiv:2212.04412  [pdf, other

    cs.CV cs.LG

    Task Bias in Vision-Language Models

    Authors: Sachit Menon, Ishaan Preetam Chandratreya, Carl Vondrick

    Abstract: Incidental supervision from language has become a popular approach for learning generic visual representations that can be prompted to perform many recognition tasks in computer vision. We conduct an in-depth exploration of the CLIP model and show that its visual representation is often strongly biased towards solving some tasks more than others. Moreover, which task the representation will be bia… ▽ More

    Submitted 8 December, 2022; originally announced December 2022.

    Comments: First two authors contributed equally

  46. arXiv:2212.02978  [pdf, other

    cs.CV q-bio.TO

    Muscles in Action

    Authors: Mia Chiquier, Carl Vondrick

    Abstract: Human motion is created by, and constrained by, our muscles. We take a first step at building computer vision methods that represent the internal muscle activity that causes motion. We present a new dataset, Muscles in Action (MIA), to learn to incorporate muscle activity into human motion representations. The dataset consists of 12.5 hours of synchronized video and surface electromyography (sEMG)… ▽ More

    Submitted 20 March, 2023; v1 submitted 5 December, 2022; originally announced December 2022.

  47. arXiv:2212.00912  [pdf, other

    cs.LG cs.CR cs.CV

    Private Multiparty Perception for Navigation

    Authors: Hui Lu, Mia Chiquier, Carl Vondrick

    Abstract: We introduce a framework for navigating through cluttered environments by connecting multiple cameras together while simultaneously preserving privacy. Occlusions and obstacles in large environments are often challenging situations for navigation agents because the environment is not fully observable from a single camera view. Given multiple camera views of an environment, our approach learns to p… ▽ More

    Submitted 1 December, 2022; originally announced December 2022.

  48. arXiv:2211.11903  [pdf, other

    cs.RO cs.CV

    FLEX: Full-Body Grasping Without Full-Body Grasps

    Authors: Purva Tendulkar, Dídac Surís, Carl Vondrick

    Abstract: Synthesizing 3D human avatars interacting realistically with a scene is an important problem with applications in AR/VR, video games and robotics. Towards this goal, we address the task of generating a virtual human -- hands and full body -- grasping everyday objects. Existing methods approach this problem by collecting a 3D dataset of humans interacting with objects and training on this data. How… ▽ More

    Submitted 28 March, 2023; v1 submitted 21 November, 2022; originally announced November 2022.

    Comments: CVPR 2023 Camera-ready

  49. arXiv:2210.07183  [pdf, other

    cs.CV cs.LG

    Visual Classification via Description from Large Language Models

    Authors: Sachit Menon, Carl Vondrick

    Abstract: Vision-language models (VLMs) such as CLIP have shown promising performance on a variety of recognition tasks using the standard zero-shot classification procedure -- computing similarity between the query image and the embedded words for each category. By only using the category name, they neglect to make use of the rich context of additional information that language affords. The procedure gives… ▽ More

    Submitted 1 December, 2022; v1 submitted 13 October, 2022; originally announced October 2022.

  50. arXiv:2210.01322  [pdf, other

    cs.LG cs.AI cs.CV

    Representing Spatial Trajectories as Distributions

    Authors: Dídac Surís, Carl Vondrick

    Abstract: We introduce a representation learning framework for spatial trajectories. We represent partial observations of trajectories as probability distributions in a learned latent space, which characterize the uncertainty about unobserved parts of the trajectory. Our framework allows us to obtain samples from a trajectory for any continuous point in time, both interpolating and extrapolating. Our flexib… ▽ More

    Submitted 3 October, 2022; originally announced October 2022.

    Comments: Accepted to NeurIPS 2022