Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–15 of 15 results for author: Choutas, V

.
  1. arXiv:2607.23687  [pdf, ps, other

    cs.CV cs.GR

    GNM Head: A Generative aNthropometric Model of the human head

    Authors: Stylianos Ploumpis, Jan Bednarik, Gaspard Zoss, Ruslan Guseinov, Luca Prasso, Prashanth Chandran, Oliver Boyne, Vasileios Choutas, Timo Bolkart, Daoye Wang, Menglei Chai, Di Qiu, Sebastian Winberg, Gilles Rainer, Lewis Bridgeman, Delio Vicini, Jérémy Riviere, Yannick Boetzel, Alexander Koumis, Stylianos Moschoglou, Jay Busch, Cynthia Herrera, Jacob Still, Scott Ysebert, Peter Lincoln , et al. (5 additional authors not shown)

    Abstract: Parametric models of the human head are essential tools traditionally used in computer vision and graphics for animation, rendering, and reconstruction. More recently, they serve as crucial conditioning signals within generative large vision models, allowing for tight spatial control of generated imagery. However, existing publicly available models are typically limited in anatomical scope, modeli… ▽ More

    Submitted 18 August, 2026; v1 submitted 26 July, 2026; originally announced July 2026.

    Comments: The GNM is publicly available at: https://github.com/google/GNM

  2. arXiv:2606.29924  [pdf, ps, other

    cs.CV

    DCGrasp: Distance-aware Controllable Grasp Generation

    Authors: Hiroyasu Akada, Jesús Pérez, Emre Aksan, Vasileios Choutas, Cristian Romero, Alberto Garcia-Garcia, Vladislav Golyanik, Christian Theobalt, Thabo Beeler

    Abstract: Generating 3D hand-object interactions is essential for applications in robotics, XR, and synthetic data generation, where flexible controllability and strong generalization to diverse object geometries are required. However, existing methods rarely satisfy these requirements, limiting their practical applicability. We present DCGrasp, a distance-aware controllable grasp generation system built on… ▽ More

    Submitted 29 June, 2026; originally announced June 2026.

  3. arXiv:2606.15966  [pdf, ps, other

    cs.CV cs.GR

    VEPHand: View-Efficient Photometric Hand Performance Capture at Scale

    Authors: Zhengyang Shen, Kai-Hung Chang, Erroll Wood, Deying Kong, Bo Peng, Timo Bolkart, Jinlong Yang, Bowen Zhao, Danhang Tang, Sasa Petrovic, Emre Aksan, Jérémy Riviere, Vassilis Choutas, Delio Vicini, Jay Busch, Shichen Liu, Zhe Cao, Hugh Liu, JingJing Shen, Jonathan Taylor, Mingsong Dou

    Abstract: Robust, high-fidelity 3D hand capture, while fundamental to digital human creation, remains challenging with practical multi-view systems that balance rich photometry with the geometric ambiguities of reconstruction arising from limited viewpoint density. This paper presents an end-to-end pipeline for dynamic hand performance capture and registration, specifically designed for view-efficient setup… ▽ More

    Submitted 18 June, 2026; v1 submitted 14 June, 2026; originally announced June 2026.

    ACM Class: I.3.8; I.4.5

  4. arXiv:2412.05066  [pdf, other

    cs.CV cs.GR cs.RO

    BimArt: A Unified Approach for the Synthesis of 3D Bimanual Interaction with Articulated Objects

    Authors: Wanyue Zhang, Rishabh Dabral, Vladislav Golyanik, Vasileios Choutas, Eduardo Alvarado, Thabo Beeler, Marc Habermann, Christian Theobalt

    Abstract: We present BimArt, a novel generative approach for synthesizing 3D bimanual hand interactions with articulated objects. Unlike prior works, we do not rely on a reference grasp, a coarse hand trajectory, or separate modes for grasping and articulating. To achieve this, we first generate distance-based contact maps conditioned on the object trajectory with an articulation-aware feature representatio… ▽ More

    Submitted 25 March, 2025; v1 submitted 6 December, 2024; originally announced December 2024.

    Comments: CVPR2025

  5. arXiv:2411.15074  [pdf, other

    cs.CV cs.LG

    Learning to Stabilize Faces

    Authors: Jan Bednarik, Erroll Wood, Vasileios Choutas, Timo Bolkart, Daoye Wang, Chenglei Wu, Thabo Beeler

    Abstract: Nowadays, it is possible to scan faces and automatically register them with high quality. However, the resulting face meshes often need further processing: we need to stabilize them to remove unwanted head movement. Stabilization is important for tasks like game development or movie making which require facial expressions to be cleanly separated from rigid head motion. Since manual stabilization i… ▽ More

    Submitted 22 November, 2024; originally announced November 2024.

    Comments: Eurographics 2024

  6. arXiv:2312.16737  [pdf, other

    cs.CV

    HMP: Hand Motion Priors for Pose and Shape Estimation from Video

    Authors: Enes Duran, Muhammed Kocabas, Vasileios Choutas, Zicong Fan, Michael J. Black

    Abstract: Understanding how humans interact with the world necessitates accurate 3D hand pose estimation, a task complicated by the hand's high degree of articulation, frequent occlusions, self-occlusions, and rapid motions. While most existing methods rely on single-image inputs, videos have useful cues to address aforementioned issues. However, existing video-based 3D hand datasets are insufficient for tr… ▽ More

    Submitted 27 December, 2023; originally announced December 2023.

    Journal ref: WACV 2024

  7. arXiv:2304.10482  [pdf, other

    cs.CV cs.GR

    Reconstructing Signing Avatars From Video Using Linguistic Priors

    Authors: Maria-Paola Forte, Peter Kulits, Chun-Hao Huang, Vasileios Choutas, Dimitrios Tzionas, Katherine J. Kuchenbecker, Michael J. Black

    Abstract: Sign language (SL) is the primary method of communication for the 70 million Deaf people around the world. Video dictionaries of isolated signs are a core SL learning tool. Replacing these with 3D avatars can aid learning and enable AR/VR applications, improving access to technology and online media. However, little work has attempted to estimate expressive 3D avatars from SL video; occlusion, noi… ▽ More

    Submitted 20 April, 2023; originally announced April 2023.

  8. arXiv:2209.02250  [pdf, other

    cs.CV

    Spatio-Temporal Action Detection Under Large Motion

    Authors: Gurkirt Singh, Vasileios Choutas, Suman Saha, Fisher Yu, Luc Van Gool

    Abstract: Current methods for spatiotemporal action tube detection often extend a bounding box proposal at a given keyframe into a 3D temporal cuboid and pool features from nearby frames. However, such pooling fails to accumulate meaningful spatiotemporal features if the position or shape of the actor shows large 2D motion and variability through the frames, due to large camera motion, large actor shape def… ▽ More

    Submitted 25 October, 2022; v1 submitted 6 September, 2022; originally announced September 2022.

    Comments: 10 pages, 5 figures, 5 tables

  9. arXiv:2206.07036  [pdf, other

    cs.CV

    Accurate 3D Body Shape Regression using Metric and Semantic Attributes

    Authors: Vasileios Choutas, Lea Muller, Chun-Hao P. Huang, Siyu Tang, Dimitrios Tzionas, Michael J. Black

    Abstract: While methods that regress 3D human meshes from images have progressed rapidly, the estimated body shapes often do not capture the true human shape. This is problematic since, for many applications, accurate body shape is as important as pose. The key reason that body shape accuracy lags pose accuracy is the lack of data. While humans can label 2D joints, and these constrain 3D pose, it is not so… ▽ More

    Submitted 14 June, 2022; originally announced June 2022.

    Comments: First two authors contributed equally

    Journal ref: CVPR 2022

  10. arXiv:2112.11454  [pdf, other

    cs.CV

    GOAL: Generating 4D Whole-Body Motion for Hand-Object Grasping

    Authors: Omid Taheri, Vasileios Choutas, Michael J. Black, Dimitrios Tzionas

    Abstract: Generating digital humans that move realistically has many applications and is widely studied, but existing methods focus on the major limbs of the body, ignoring the hands and head. Hands have been separately studied, but the focus has been on generating realistic static grasps of objects. To synthesize virtual characters that interact with the world, we need to generate full-body motions and rea… ▽ More

    Submitted 16 March, 2023; v1 submitted 21 December, 2021; originally announced December 2021.

  11. arXiv:2111.14824  [pdf, other

    cs.CV

    Learning to Fit Morphable Models

    Authors: Vasileios Choutas, Federica Bogo, Jingjing Shen, Julien Valentin

    Abstract: Fitting parametric models of human bodies, hands or faces to sparse input signals in an accurate, robust, and fast manner has the promise of significantly improving immersion in AR and VR scenarios. A common first step in systems that tackle these problems is to regress the parameters of the parametric model directly from the input data. This approach is fast, robust, and is a good starting point… ▽ More

    Submitted 20 July, 2022; v1 submitted 29 November, 2021; originally announced November 2021.

    Comments: ECCV 2022

  12. arXiv:2105.05301  [pdf, other

    cs.CV

    Collaborative Regression of Expressive Bodies using Moderation

    Authors: Yao Feng, Vasileios Choutas, Timo Bolkart, Dimitrios Tzionas, Michael J. Black

    Abstract: Recovering expressive humans from images is essential for understanding human behavior. Methods that estimate 3D bodies, faces, or hands have progressed significantly, yet separately. Face methods recover accurate 3D shape and geometric details, but need a tight crop and struggle with extreme views and low resolution. Whole-body methods are robust to a wide range of poses and resolutions, but prov… ▽ More

    Submitted 15 October, 2021; v1 submitted 11 May, 2021; originally announced May 2021.

    Comments: 21 pages. The first two authors contributed equally to this work

  13. arXiv:2008.09062  [pdf, other

    cs.CV cs.GR

    Monocular Expressive Body Regression through Body-Driven Attention

    Authors: Vasileios Choutas, Georgios Pavlakos, Timo Bolkart, Dimitrios Tzionas, Michael J. Black

    Abstract: To understand how people look, interact, or perform tasks, we need to quickly and accurately capture their 3D body, face, and hands together from an RGB image. Most existing methods focus only on parts of the body. A few recent approaches reconstruct full expressive 3D humans from images using 3D body models that include the face and hands. These methods are optimization-based and thus slow, prone… ▽ More

    Submitted 20 August, 2020; originally announced August 2020.

    Comments: Accepted in ECCV'20. Project page: http://expose.is.tue.mpg.de

  14. arXiv:1908.06963  [pdf, other

    cs.CV

    Resolving 3D Human Pose Ambiguities with 3D Scene Constraints

    Authors: Mohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, Michael J. Black

    Abstract: To understand and analyze human behavior, we need to capture humans moving in, and interacting with, the world. Most existing methods perform 3D human pose estimation without explicitly considering the scene. We observe however that the world constrains the body and vice-versa. To motivate this, we show that current 3D human pose estimation methods produce results that are not consistent with the… ▽ More

    Submitted 20 August, 2019; originally announced August 2019.

    Comments: To appear in ICCV 2019

  15. arXiv:1904.05866  [pdf, other

    cs.CV

    Expressive Body Capture: 3D Hands, Face, and Body from a Single Image

    Authors: Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, Michael J. Black

    Abstract: To facilitate the analysis of human actions, interactions and emotions, we compute a 3D model of human body pose, hand pose, and facial expression from a single monocular image. To achieve this, we use thousands of 3D scans to train a new, unified, 3D model of the human body, SMPL-X, that extends SMPL with fully articulated hands and an expressive face. Learning to regress the parameters of SMPL-X… ▽ More

    Submitted 11 April, 2019; originally announced April 2019.

    Comments: To appear in CVPR 2019