Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–21 of 21 results for author: Hellwich, O

Searching in archive cs. Search in all archives.
.
  1. arXiv:2607.06176  [pdf, ps, other

    cs.CV

    Revisiting Scene Graph Generation from the Perspective of Detector-Conditioned Reachability

    Authors: Runfeng Qu, Pia K Bideau, Ole Hall, Julie Ouerfelli-Ethier, Klaus Obermayer, Olaf Hellwich

    Abstract: Scene graph generation (SGG) approaches can be broadly classified into detector-based and query-based methods according to their underlying reasoning mechanisms. However, the discrepancy in their predictive behaviors, induced by these distinct mechanisms, has not been systematically analyzed. In this work, we design a controlled experimental setup to examine prediction discrepancies from the persp… ▽ More

    Submitted 7 July, 2026; originally announced July 2026.

    Comments: Accepted by ECCV 2026

  2. arXiv:2605.05014  [pdf, ps, other

    cs.CV

    CARD: A Multi-Modal Automotive Dataset for Dense 3D Reconstruction in Challenging Road Topography

    Authors: Gasser Elazab, Frank Neuhaus, Tilman Koß, Malte Splietker, Aditya Date, Michael Unterreiner, Maximilian Jansen, Olaf Hellwich

    Abstract: Autonomous driving must operate across diverse surfaces to enable safe mobility. However, most driving datasets are captured on well-paved flat roads. Moreover, recent driving datasets primarily provide sparse LiDAR ground truth for images, which is insufficient for assessing fine-grained geometry in depth estimation and completion. To address these gaps, we introduce CARD, a multi-modal driving d… ▽ More

    Submitted 7 May, 2026; v1 submitted 6 May, 2026; originally announced May 2026.

    Comments: Accepted at CVPR 2026 (Highlight). Project page: https://card.content.cariad.digital

  3. arXiv:2605.01971  [pdf, ps, other

    cs.CV

    ProtoFair: Fair Self-Supervised Contrastive Learning via Pseudo-Counterfactual Pairs

    Authors: Marah Halawa, Olaf Hellwich

    Abstract: Self-supervised learning methods learn high-quality visual representations, yet recent studies show that these representations often capture demographic biases present in the training data. Existing fairness-aware methods address this by redesigning the self-supervised objective itself, limiting portability across the rapidly evolving landscape of self-supervised learning (SSL) frameworks. We prop… ▽ More

    Submitted 27 June, 2026; v1 submitted 3 May, 2026; originally announced May 2026.

    Comments: Paper accepted at ECCV 2026

  4. arXiv:2601.08728  [pdf, ps, other

    cs.CV

    Salience-SGG: Enhancing Unbiased Scene Graph Generation with Iterative Salience Estimation

    Authors: Runfeng Qu, Ole Hall, Pia K Bideau, Julie Ouerfelli-Ethier, Martin Rolfs, Klaus Obermayer, Olaf Hellwich

    Abstract: Scene Graph Generation (SGG) suffers from a long-tailed distribution, where a few predicate classes dominate while many others are underrepresented, leading to biased models that underperform on rare relations. Unbiased-SGG methods address this issue by implementing debiasing strategies, but often at the cost of spatial understanding, resulting in an over-reliance on semantic priors. We introduce… ▽ More

    Submitted 13 January, 2026; originally announced January 2026.

  5. arXiv:2512.04303  [pdf, ps, other

    cs.CV cs.AI

    Gamma-from-Mono: Road-Relative, Metric, Self-Supervised Monocular Geometry for Vehicular Applications

    Authors: Gasser Elazab, Maximilian Jansen, Michael Unterreiner, Olaf Hellwich

    Abstract: Accurate perception of the vehicle's 3D surroundings, including fine-scale road geometry, such as bumps, slopes, and surface irregularities, is essential for safe and comfortable vehicle control. However, conventional monocular depth estimation often oversmooths these features, losing critical information for motion planning and stability. To address this, we introduce Gamma-from-Mono (GfM), a lig… ▽ More

    Submitted 3 December, 2025; originally announced December 2025.

    Comments: Accepted in 3DV 2026

  6. Mouse Lockbox Dataset: Behavior Recognition for Mice Solving Lockboxes

    Authors: Patrik Reiske, Marcus N. Boon, Niek Andresen, Sole Traverso, Katharina Hohlbaum, Lars Lewejohann, Christa Thöne-Reineke, Olaf Hellwich, Henning Sprekeler

    Abstract: Machine learning and computer vision methods have a major impact on the study of natural animal behavior, as they enable the (semi-)automatic analysis of vast amounts of video data. Mice are the standard mammalian model system in most research fields, but the datasets available today to refine such methods focus either on simple or social behaviors. In this work, we present a video dataset of indi… ▽ More

    Submitted 17 June, 2025; v1 submitted 21 May, 2025; originally announced May 2025.

    Comments: Accepted and published (poster) at the CV4Animals: Computer Vision for Animal Behavior Tracking and Modeling workshop, in conjunction with Computer Vision and Pattern Recognition (CVPR) 2025

  7. arXiv:2501.11030  [pdf, other

    cs.CV

    Tracking Mouse from Incomplete Body-Part Observations and Deep-Learned Deformable-Mouse Model Motion-Track Constraint for Behavior Analysis

    Authors: Olaf Hellwich, Niek Andresen, Katharina Hohlbaum, Marcus N. Boon, Monika Kwiatkowski, Simon Matern, Patrik Reiske, Henning Sprekeler, Christa ThöneReineke, Lars Lewejohann, Huma Ghani Zada, Michael Brück, Soledad Traverso

    Abstract: Tracking mouse body parts in video is often incomplete due to occlusions such that - e.g. - subsequent action and behavior analysis is impeded. In this conceptual work, videos from several perspectives are integrated via global exterior camera orientation; body part positions are estimated by 3D triangulation and bundle adjustment. Consistency of overall 3D track reconstruction is achieved by intr… ▽ More

    Submitted 19 January, 2025; originally announced January 2025.

    Comments: 10 pages

    Journal ref: Reinhardt, Wolfgang; Huang, Hai (editors): Festschrift für Prof. Dr.-Ing. Helmut Mayer zum 60. Geburtstag, Institut für Geodäsie der Universität der Bundeswehr München, Vol. 101, 2024, pages 45 - 53

  8. arXiv:2411.19717  [pdf, other

    cs.CV cs.AI cs.LG cs.RO

    MonoPP: Metric-Scaled Self-Supervised Monocular Depth Estimation by Planar-Parallax Geometry in Automotive Applications

    Authors: Gasser Elazab, Torben Gräber, Michael Unterreiner, Olaf Hellwich

    Abstract: Self-supervised monocular depth estimation (MDE) has gained popularity for obtaining depth predictions directly from videos. However, these methods often produce scale invariant results, unless additional training signals are provided. Addressing this challenge, we introduce a novel self-supervised metric-scaled MDE model that requires only monocular video data and the camera's mounting position,… ▽ More

    Submitted 29 November, 2024; originally announced November 2024.

    Comments: Accepted at WACV 25, project page: https://mono-pp.github.io/

    Journal ref: Proceedings of the 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Tucson, AZ, USA, 26 February 2025, pp. 2777-2787

  9. arXiv:2409.02566  [pdf, other

    cs.CV

    How Do You Perceive My Face? Recognizing Facial Expressions in Multi-Modal Context by Modeling Mental Representations

    Authors: Florian Blume, Runfeng Qu, Pia Bideau, Martin Maier, Rasha Abdel Rahman, Olaf Hellwich

    Abstract: Facial expression perception in humans inherently relies on prior knowledge and contextual cues, contributing to efficient and flexible processing. For instance, multi-modal emotional context (such as voice color, affective text, body pose, etc.) can prompt people to perceive emotional expressions in objectively neutral faces. Drawing inspiration from this, we introduce a novel approach for facial… ▽ More

    Submitted 4 September, 2024; originally announced September 2024.

    Comments: GCPR 2024

  10. arXiv:2408.02766  [pdf, other

    cs.CV cs.LG

    ConDL: Detector-Free Dense Image Matching

    Authors: Monika Kwiatkowski, Simon Matern, Olaf Hellwich

    Abstract: In this work, we introduce a deep-learning framework designed for estimating dense image correspondences. Our fully convolutional model generates dense feature maps for images, where each pixel is associated with a descriptor that can be matched across multiple images. Unlike previous methods, our model is trained on synthetic data that includes significant distortions, such as perspective changes… ▽ More

    Submitted 5 August, 2024; originally announced August 2024.

  11. arXiv:2407.17209  [pdf, other

    cs.CV cs.AI cs.HC cs.LG

    Nonverbal Immediacy Analysis in Education: A Multimodal Computational Model

    Authors: Uroš Petković, Jonas Frenkel, Olaf Hellwich, Rebecca Lazarides

    Abstract: This paper introduces a novel computational approach for analyzing nonverbal social behavior in educational settings. Integrating multimodal behavioral cues, including facial expressions, gesture intensity, and spatial dynamics, the model assesses the nonverbal immediacy (NVI) of teachers from RGB classroom videos. A dataset of 400 30-second video segments from German classrooms was constructed fo… ▽ More

    Submitted 24 July, 2024; originally announced July 2024.

    Comments: 12 pages, 3 figures. Camera-ready version for the SAB 2024: 17th International Conference on the Simulation of Adaptive Behavior

    MSC Class: 68T45; 68T10; 68U10; 91E45 ACM Class: I.2.10; I.5.4; K.3.1

  12. arXiv:2404.10904  [pdf, other

    cs.CV

    Multi-Task Multi-Modal Self-Supervised Learning for Facial Expression Recognition

    Authors: Marah Halawa, Florian Blume, Pia Bideau, Martin Maier, Rasha Abdel Rahman, Olaf Hellwich

    Abstract: Human communication is multi-modal; e.g., face-to-face interaction involves auditory signals (speech) and visual signals (face movements and hand gestures). Hence, it is essential to exploit multiple modalities when designing machine learning-based facial expression recognition systems. In addition, given the ever-growing quantities of video data that capture human facial expressions, such systems… ▽ More

    Submitted 4 September, 2024; v1 submitted 16 April, 2024; originally announced April 2024.

    Comments: The paper will appear in the CVPR 2024 workshops proceedings

    Journal ref: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2024, pp. 4604-4614

  13. arXiv:2311.16829  [pdf, other

    cs.CV cs.LG

    Decomposer: Semi-supervised Learning of Image Restoration and Image Decomposition

    Authors: Boris Meinardus, Mariusz Trzeciakiewicz, Tim Herzig, Monika Kwiatkowski, Simon Matern, Olaf Hellwich

    Abstract: We present Decomposer, a semi-supervised reconstruction model that decomposes distorted image sequences into their fundamental building blocks - the original image and the applied augmentations, i.e., shadow, light, and occlusions. To solve this problem, we use the SIDAR dataset that provides a large number of distorted image sequences: each sequence contains images with shadows, lighting, and occ… ▽ More

    Submitted 28 November, 2023; originally announced November 2023.

  14. DIAR: Deep Image Alignment and Reconstruction using Swin Transformers

    Authors: Monika Kwiatkowski, Simon Matern, Olaf Hellwich

    Abstract: When taking images of some occluded content, one is often faced with the problem that every individual image frame contains unwanted artifacts, but a collection of images contains all relevant information if properly aligned and aggregated. In this paper, we attempt to build a deep learning pipeline that simultaneously aligns a sequence of distorted images and reconstructs them. We create a datase… ▽ More

    Submitted 17 October, 2023; originally announced October 2023.

  15. arXiv:2305.12036  [pdf, other

    cs.CV cs.GR cs.LG

    SIDAR: Synthetic Image Dataset for Alignment & Restoration

    Authors: Monika Kwiatkowski, Simon Matern, Olaf Hellwich

    Abstract: Image alignment and image restoration are classical computer vision tasks. However, there is still a lack of datasets that provide enough data to train and evaluate end-to-end deep learning models. Obtaining ground-truth data for image alignment requires sophisticated structure-from-motion methods or optical flow systems that often do not provide enough data variance, i.e., typically providing a h… ▽ More

    Submitted 19 May, 2023; originally announced May 2023.

  16. arXiv:2208.04201  [pdf, other

    cs.IR cs.CV cs.LG

    Content-Based Landmark Retrieval Combining Global and Local Features using Siamese Neural Networks

    Authors: Tianyi Hu, Monika Kwiatkowski, Simon Matern, Olaf Hellwich

    Abstract: In this work, we present a method for landmark retrieval that utilizes global and local features. A Siamese network is used for global feature extraction and metric learning, which gives an initial ranking of the landmark search. We utilize the extracted feature maps from the Siamese architecture as local descriptors, the search results are then further refined using a cosine similarity between lo… ▽ More

    Submitted 3 August, 2022; originally announced August 2022.

  17. Image-based Detection of Surface Defects in Concrete during Construction

    Authors: Dominik Kuhnke, Monika Kwiatkowski, Olaf Hellwich

    Abstract: Defects increase the cost and duration of construction projects as they require significant inspection and documentation efforts. Automating defect detection could significantly reduce these efforts. This work focuses on detecting honeycombs, a substantial defect in concrete structures that may affect structural integrity. We compared honeycomb images scraped from the web with images obtained from… ▽ More

    Submitted 6 December, 2022; v1 submitted 3 August, 2022; originally announced August 2022.

  18. arXiv:2207.08664  [pdf, other

    cs.CV

    Action-based Contrastive Learning for Trajectory Prediction

    Authors: Marah Halawa, Olaf Hellwich, Pia Bideau

    Abstract: Trajectory prediction is an essential task for successful human robot interaction, such as in autonomous driving. In this work, we address the problem of predicting future pedestrian trajectories in a first person view setting with a moving camera. To that end, we propose a novel action-based contrastive learning loss, that utilizes pedestrian action information to improve the learned trajectory e… ▽ More

    Submitted 18 July, 2022; originally announced July 2022.

    Comments: This paper will appear in the proceedings of The European Conference on Computer Vision (ECCV 2022)

  19. arXiv:2106.06073  [pdf, other

    cs.CV

    A modular framework for object-based saccadic decisions in dynamic scenes

    Authors: Nicolas Roth, Pia Bideau, Olaf Hellwich, Martin Rolfs, Klaus Obermayer

    Abstract: Visually exploring the world around us is not a passive process. Instead, we actively explore the world and acquire visual information over time. Here, we present a new model for simulating human eye-movement behavior in dynamic real-world scenes. We model this active scene exploration as a sequential decision making process. We adapt the popular drift-diffusion model (DDM) for perceptual decision… ▽ More

    Submitted 10 June, 2021; originally announced June 2021.

    Comments: Accepted for presentation at EPIC@CVPR2021 workshop, 4 pages, 2 figures

  20. arXiv:2008.07001  [pdf, other

    cs.CV cs.LG eess.IV

    Learning Disentangled Expression Representations from Facial Images

    Authors: Marah Halawa, Manuel Wöllhaf, Eduardo Vellasques, Urko Sánchez Sanz, Olaf Hellwich

    Abstract: Face images are subject to many different factors of variation, especially in unconstrained in-the-wild scenarios. For most tasks involving such images, e.g. expression recognition from video streams, having enough labeled data is prohibitively expensive. One common strategy to tackle such a problem is to learn disentangled representations for the different factors of variation of the observed dat… ▽ More

    Submitted 18 August, 2020; v1 submitted 16 August, 2020; originally announced August 2020.

    Comments: Accepted at ECCV2020 workshops

  21. Iterative Bilateral Filtering of Polarimetric SAR Data

    Authors: Olivier D'Hondt, Stéphane Guillaso, Olaf Hellwich

    Abstract: In this paper, we introduce an iterative speckle filtering method for polarimetric SAR (PolSAR) images based on the bilateral filter. To locally adapt to the spatial structure of images, this filter relies on pixel similarities in both spatial and radiometric domains. To deal with polarimetric data, we study the use of similarities based on a statistical distance called Kullback-Leibler divergence… ▽ More

    Submitted 1 November, 2013; originally announced November 2013.

    Comments: Available: http://ieeexplore.ieee.org/xpl/articleDetails.jsp?arnumber=6509975

    Journal ref: Selected Topics in Applied Earth Observations and Remote Sensing, IEEE Journal of (Volume:6, Issue: 3 ) 2013