Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 62 results for author: Jensfelt, P

.
  1. arXiv:2608.14266  [pdf, ps, other

    cs.RO cs.CV

    Accelerating Large-scale Bundle Adjustment for LiDAR Mapping via Parallel Computing

    Authors: Yixi Cai, Rundong Li, Yuhan Xie, Qingwen Zhang, Patric Jensfelt, Fu Zhang

    Abstract: LiDAR bundle adjustment is widely utilized in mapping to construct globally consistent point cloud maps. In this paper, we propose the first fully parallel computing framework to accelerate LiDAR bundle adjustment for large-scale mapping, incorporating three key techniques. First, we design an adaptive, asynchronous data loading strategy to efficiently process large-scale point cloud datasets on m… ▽ More

    Submitted 14 August, 2026; originally announced August 2026.

    Comments: Accepted by IEEE International Conference on Automation Science and Engineering (CASE), 2026

  2. arXiv:2608.10886  [pdf, ps, other

    cs.CV cs.RO

    GESTO: Human-Centric Spatio-Temporal Memory for Reasoning in Dynamic Scenes

    Authors: Ermanno Bartoli, Buwei He, Dennis Rotondi, Sebastian Koch, Federico Tombari, Kai O. Arras, Patric Jensfelt, Yixi Cai, Iolanda Leite

    Abstract: Robots operating in human environments need memories that capture not only what objects exist and where, but also how people use them over time and how individual interactions compose into goal-directed activities. Existing 4D scene graphs preserve object and place histories but omit activity structure, whereas activity representations are either not grounded in persistent 3D scenes or rely on ext… ▽ More

    Submitted 11 August, 2026; originally announced August 2026.

  3. arXiv:2607.15016  [pdf, ps, other

    cs.RO

    Risk-Aware Belief Control Barrier Functions over Random Finite Sets

    Authors: Shaohang Han, Gang Chen, Yixi Cai, Ignacio Torroba, Ivan Stenius, Patric Jensfelt, Javier Alonso-Mora, Jana Tumova

    Abstract: Ensuring robot safety in unknown, dynamic environments is a fundamental requirement. It involves inferring the states of an unknown and time-varying number of moving objects from noisy, incomplete measurements. We address safe control under the induced multi-object state uncertainty with a risk-aware belief control barrier function (BCBF) framework. The uncertainty is captured by a random finite s… ▽ More

    Submitted 16 July, 2026; originally announced July 2026.

  4. arXiv:2606.20189  [pdf, ps, other

    cs.CV cs.AI cs.RO

    HilDA: Hierarchical Distillation with Diffusion for Advancing Self-Supervised LiDAR Pre-training

    Authors: Maciej Wozniak, Jesper Ericsson, Hariprasath Govindarajan, Truls Nyberg, Thomas Gustafsson, Patric Jensfelt, Olov Andersson

    Abstract: Leveraging Vision Foundation Models (VFMs) for camera-to-LiDAR knowledge distillation offers a promising solution to the scarcity of annotated data needed to represent the immense geometric and kinematic diversity of real-world autonomous driving (AD). However, current approaches typically treat VFMs as black-box teachers, relying exclusively on frame-wise feature similarity. Consequently, they do… ▽ More

    Submitted 23 June, 2026; v1 submitted 18 June, 2026; originally announced June 2026.

    Comments: Accepted to ECCV 2026. Maciej and Jesper contributed equally

  5. arXiv:2604.13571  [pdf, ps, other

    cs.CV

    Radar-Informed 3D Multi-Object Tracking under Adverse Conditions

    Authors: Bingxue Xu, Emil Hedemalm, Ajinkya Khoche, Patric Jensfelt

    Abstract: The challenge of 3D multi-object tracking is achieving robustness in real-world applications, for example under adverse conditions and maintaining consistency as distance increases. To overcome these challenges, sensor fusion approaches that combine LiDAR, cameras, and radar have emerged. However, existing multimodal methods usually treat radar as another learned feature inside the network. When t… ▽ More

    Submitted 21 April, 2026; v1 submitted 15 April, 2026; originally announced April 2026.

    Comments: 7 pages, 5 figures

  6. arXiv:2604.09411  [pdf, ps, other

    cs.CV

    SynFlow: Scaling Up LiDAR Scene Flow Estimation with Synthetic Data

    Authors: Qingwen Zhang, Xiaomeng Zhu, Chenhan Jiang, Patric Jensfelt

    Abstract: Reliable 3D dynamic perception requires models that can anticipate motion beyond predefined categories, yet progress is hindered by the scarcity of dense, high-quality motion annotations. While self-supervision on unlabeled real data offers a path forward, empirical evidence suggests that scaling unlabeled data fails to close the performance gap due to noisy proxy signals. In this paper, we propos… ▽ More

    Submitted 28 July, 2026; v1 submitted 10 April, 2026; originally announced April 2026.

    Comments: ECCV 2026; 18 pages, 4 figures

  7. arXiv:2602.19053  [pdf, ps, other

    cs.CV cs.RO

    TeFlow: Enabling Multi-frame Supervision for Self-Supervised Feed-forward Scene Flow Estimation

    Authors: Qingwen Zhang, Chenhan Jiang, Xiaomeng Zhu, Yunqi Miao, Yushan Zhang, Olov Andersson, Patric Jensfelt

    Abstract: Self-supervised feed-forward methods for scene flow estimation offer real-time efficiency, but their supervision from two-frame point correspondences is unreliable and often breaks down under occlusions. Multi-frame supervision has the potential to provide more stable guidance by incorporating motion cues from past frames, yet naive extensions of two-frame objectives are ineffective because point… ▽ More

    Submitted 1 April, 2026; v1 submitted 22 February, 2026; originally announced February 2026.

    Comments: CVPR 2026; 16 pages, 8 figures

  8. arXiv:2511.12503  [pdf, ps, other

    cs.CV

    Visible Structure Retrieval for Lightweight Image-Based Relocalisation

    Authors: Fereidoon Zangeneh, Leonard Bruns, Amit Dekel, Alessandro Pieropan, Patric Jensfelt

    Abstract: Accurate camera pose estimation from an image observation in a previously mapped environment is commonly done through structure-based methods: by finding correspondences between 2D keypoints on the image and 3D structure points in the map. In order to make this correspondence search tractable in large scenes, existing pipelines either rely on search heuristics, or perform image retrieval to reduce… ▽ More

    Submitted 16 November, 2025; originally announced November 2025.

    Comments: Accepted at BMVC 2025

  9. arXiv:2510.18244  [pdf, ps, other

    cs.CV

    BlendCLIP: Bridging Synthetic and Real Domains for Zero-Shot 3D Object Classification with Multimodal Pretraining

    Authors: Ajinkya Khoche, Gergő László Nagy, Maciej Wozniak, Thomas Gustafsson, Patric Jensfelt

    Abstract: Zero-shot 3D object classification is crucial for real-world applications like autonomous driving, however it is often hindered by a significant domain gap between the synthetic data used for training and the sparse, noisy LiDAR scans encountered in the real-world. Current methods trained solely on synthetic data fail to generalize to outdoor scenes, while those trained only on real data lack the… ▽ More

    Submitted 20 October, 2025; originally announced October 2025.

    Comments: Under Review

  10. arXiv:2509.24966  [pdf, ps, other

    cs.CV

    Social 3D Scene Graphs: Modeling Human Actions and Relations for Interactive Service Robots

    Authors: Ermanno Bartoli, Dennis Rotondi, Buwei He, Patric Jensfelt, Kai O. Arras, Iolanda Leite

    Abstract: Understanding how people interact with their surroundings and each other is essential for enabling robots to act in socially compliant and context-aware ways. While 3D Scene Graphs have emerged as a powerful semantic representation for scene understanding, existing approaches largely ignore humans in the scene, also due to the lack of annotated human-environment relationships. Moreover, existing m… ▽ More

    Submitted 6 July, 2026; v1 submitted 29 September, 2025; originally announced September 2025.

    Comments: Equal contribution from E. Bartoli and D. Rotondi. Paper accepted at IROS 2026

  11. arXiv:2509.08764  [pdf, ps, other

    cs.CV

    ArgoTweak: Towards Self-Updating HD Maps through Structured Priors

    Authors: Lena Wild, Rafael Valencia, Patric Jensfelt

    Abstract: Reliable integration of prior information is crucial for self-verifying and self-updating HD maps. However, no public dataset includes the required triplet of prior maps, current maps, and sensor data. As a result, existing methods must rely on synthetic priors, which create inconsistencies and lead to a significant sim2real gap. To address this, we introduce ArgoTweak, the first dataset to comple… ▽ More

    Submitted 10 September, 2025; originally announced September 2025.

    Comments: ICCV 2025

  12. arXiv:2508.18506  [pdf, ps, other

    cs.CV

    DoGFlow: Self-Supervised LiDAR Scene Flow via Cross-Modal Doppler Guidance

    Authors: Ajinkya Khoche, Qingwen Zhang, Yixi Cai, Sina Sharif Mansouri, Patric Jensfelt

    Abstract: Accurate 3D scene flow estimation is critical for autonomous systems to navigate dynamic environments safely, but creating the necessary large-scale, manually annotated datasets remains a significant bottleneck for developing robust perception models. Current self-supervised methods struggle to match the performance of fully supervised approaches, especially in challenging long-range and adverse w… ▽ More

    Submitted 25 August, 2025; originally announced August 2025.

    Comments: Under Review

  13. arXiv:2508.17054  [pdf, ps, other

    cs.CV cs.RO

    DeltaFlow: An Efficient Multi-frame Scene Flow Estimation Method

    Authors: Qingwen Zhang, Xiaomeng Zhu, Yushan Zhang, Yixi Cai, Olov Andersson, Patric Jensfelt

    Abstract: Previous dominant methods for scene flow estimation focus mainly on input from two consecutive frames, neglecting valuable information in the temporal domain. While recent trends shift towards multi-frame reasoning, they suffer from rapidly escalating computational costs as the number of frames grows. To leverage temporal information more efficiently, we propose DeltaFlow ($Δ$Flow), a lightweight… ▽ More

    Submitted 22 December, 2025; v1 submitted 23 August, 2025; originally announced August 2025.

    Comments: NeurIPS 2025 Spotlight, 18 pages (10 main pages + 8 supp materail), 11 figures, code at https://github.com/Kin-Zhang/DeltaFlow

  14. arXiv:2507.17596  [pdf, ps, other

    cs.CV cs.AI cs.LG cs.RO

    PRIX: Learning to Plan from Raw Pixels for End-to-End Autonomous Driving

    Authors: Maciej K. Wozniak, Lianhang Liu, Yixi Cai, Patric Jensfelt

    Abstract: While end-to-end autonomous driving models show promising results, their practical deployment is often hindered by large model sizes, a reliance on expensive LiDAR sensors and computationally intensive BEV feature representations. This limits their scalability, especially for mass-market vehicles equipped only with cameras. To address these challenges, we propose PRIX (Plan from Raw Pixels). Our n… ▽ More

    Submitted 12 April, 2026; v1 submitted 23 July, 2025; originally announced July 2025.

    Comments: Accepted for Robotics and Automation Letters (RA-L) and will be presented at iROS 2026

  15. arXiv:2506.22336  [pdf, ps, other

    cs.CV

    MatChA: Cross-Algorithm Matching with Feature Augmentation

    Authors: Paula Carbó Cubero, Alberto Jaenal Gálvez, André Mateus, José Araújo, Patric Jensfelt

    Abstract: State-of-the-art methods fail to solve visual localization in scenarios where different devices use different sparse feature extraction algorithms to obtain keypoints and their corresponding descriptors. Translating feature descriptors is enough to enable matching. However, performance is drastically reduced in cross-feature detector cases, because current solutions assume common keypoints. This m… ▽ More

    Submitted 27 June, 2025; originally announced June 2025.

  16. arXiv:2504.07260  [pdf, other

    cs.CV

    Quantifying Epistemic Uncertainty in Absolute Pose Regression

    Authors: Fereidoon Zangeneh, Amit Dekel, Alessandro Pieropan, Patric Jensfelt

    Abstract: Visual relocalization is the task of estimating the camera pose given an image it views. Absolute pose regression offers a solution to this task by training a neural network, directly regressing the camera pose from image features. While an attractive solution in terms of memory and compute efficiency, absolute pose regression's predictions are inaccurate and unreliable outside the training domain… ▽ More

    Submitted 9 April, 2025; originally announced April 2025.

  17. arXiv:2504.01980  [pdf, other

    cs.RO cs.AI

    Information Gain Is Not All You Need

    Authors: Ludvig Ericson, José Pedro, Patric Jensfelt

    Abstract: Autonomous exploration in mobile robotics often involves a trade-off between two objectives: maximizing environmental coverage and minimizing the total path length. In the widely used information gain paradigm, exploration is guided by the expected value of observations. While this approach is effective under budget-constrained settings--where only a limited number of observations can be made--it… ▽ More

    Submitted 20 April, 2025; v1 submitted 28 March, 2025; originally announced April 2025.

    Comments: 9 pages, 6 figures, under review

  18. HiMo: High-Speed Objects Motion Compensation in Point Clouds

    Authors: Qingwen Zhang, Ajinkya Khoche, Yi Yang, Li Ling, Sina Sharif Mansouri, Olov Andersson, Patric Jensfelt

    Abstract: LiDAR point cloud is essential for autonomous vehicles, but motion distortions from dynamic objects degrade the data quality. While previous work has considered distortions caused by ego motion, distortions caused by other moving objects remain largely overlooked, leading to errors in object shape and position. This distortion is particularly pronounced in high-speed environments such as highways… ▽ More

    Submitted 30 November, 2025; v1 submitted 2 March, 2025; originally announced March 2025.

    Comments: 15 pages, 13 figures, Published in Transactions on Robotics (Volume 41)

  19. arXiv:2501.17821  [pdf, other

    cs.CV

    SSF: Sparse Long-Range Scene Flow for Autonomous Driving

    Authors: Ajinkya Khoche, Qingwen Zhang, Laura Pereira Sanchez, Aron Asefaw, Sina Sharif Mansouri, Patric Jensfelt

    Abstract: Scene flow enables an understanding of the motion characteristics of the environment in the 3D world. It gains particular significance in the long-range, where object-based perception methods might fail due to sparse observations far away. Although significant advancements have been made in scene flow pipelines to handle large-scale point clouds, a gap remains in scalability with respect to long-r… ▽ More

    Submitted 29 January, 2025; originally announced January 2025.

    Comments: 7 pages, 3 figures, accepted to International Conference on Robotics and Automation (ICRA) 2025

  20. arXiv:2410.04989  [pdf, other

    cs.CV

    Conditional Variational Autoencoders for Probabilistic Pose Regression

    Authors: Fereidoon Zangeneh, Leonard Bruns, Amit Dekel, Alessandro Pieropan, Patric Jensfelt

    Abstract: Robots rely on visual relocalization to estimate their pose from camera images when they lose track. One of the challenges in visual relocalization is repetitive structures in the operation environment of the robot. This calls for probabilistic methods that support multiple hypotheses for robot's pose. We propose such a probabilistic method to predict the posterior distribution of camera poses giv… ▽ More

    Submitted 7 October, 2024; originally announced October 2024.

    Comments: Accepted at IROS 2024

  21. arXiv:2409.11906  [pdf, other

    cs.RO

    Fusion in Context: A Multimodal Approach to Affective State Recognition

    Authors: Youssef Mohamed, Severin Lemaignan, Arzu Guneysu, Patric Jensfelt, Christian Smith

    Abstract: Accurate recognition of human emotions is a crucial challenge in affective computing and human-robot interaction (HRI). Emotional states play a vital role in shaping behaviors, decisions, and social interactions. However, emotional expressions can be influenced by contextual factors, leading to misinterpretations if context is not considered. Multimodal fusion, combining modalities like facial exp… ▽ More

    Submitted 18 September, 2024; originally announced September 2024.

  22. arXiv:2409.10178  [pdf, other

    cs.CV

    ExelMap: Explainable Element-based HD-Map Change Detection and Update

    Authors: Lena Wild, Ludvig Ericson, Rafael Valencia, Patric Jensfelt

    Abstract: Acquisition and maintenance are central problems in deploying high-definition (HD) maps for autonomous driving, with two lines of research prevalent in current literature: Online HD map generation and HD map change detection. However, the generated map's quality is currently insufficient for safe deployment, and many change detection approaches fail to precisely localize and extract the changed ma… ▽ More

    Submitted 16 September, 2024; originally announced September 2024.

    Comments: 17 pages, 3 figures

  23. arXiv:2407.19463  [pdf, other

    cs.RO

    HD-maps as Prior Information for Globally Consistent Mapping in GPS-denied Environments

    Authors: Waqas Ali, Patric Jensfelt, Thien-Minh Nguyen

    Abstract: In recent years, prior maps have become a mainstream tool in autonomous navigation. However, commonly available prior maps are still tailored to control-and-decision tasks, and the use of these maps for localization remains largely unexplored. To bridge this gap, we propose a lidar-based localization and mapping (LOAM) system that can exploit the common HD-maps in autonomous driving scenarios. Spe… ▽ More

    Submitted 28 July, 2024; originally announced July 2024.

  24. arXiv:2407.01702  [pdf, other

    cs.CV cs.RO

    SeFlow: A Self-Supervised Scene Flow Method in Autonomous Driving

    Authors: Qingwen Zhang, Yi Yang, Peizheng Li, Olov Andersson, Patric Jensfelt

    Abstract: Scene flow estimation predicts the 3D motion at each point in successive LiDAR scans. This detailed, point-level, information can help autonomous vehicles to accurately predict and understand dynamic changes in their surroundings. Current state-of-the-art methods require annotated data to train scene flow networks and the expense of labeling inherently limits their scalability. Self-supervised app… ▽ More

    Submitted 17 September, 2024; v1 submitted 1 July, 2024; originally announced July 2024.

    Comments: 25 pages (14 main pages + 11 supp materail), 5 figures, ECCV 2024

  25. Beyond the Frontier: Predicting Unseen Walls from Occupancy Grids by Learning from Floor Plans

    Authors: Ludvig Ericson, Patric Jensfelt

    Abstract: In this paper, we tackle the challenge of predicting the unseen walls of a partially observed environment as a set of 2D line segments, conditioned on occupancy grids integrated along the trajectory of a 360° LIDAR sensor. A dataset of such occupancy grids and their corresponding target wall segments is collected by navigating a virtual robot between a set of randomly sampled waypoints in a collec… ▽ More

    Submitted 13 June, 2024; originally announced June 2024.

    Comments: RA-L, 8 pages

    Journal ref: IEEE Robotics and Automation Letters (2024) pp. 2377-3766

  26. arXiv:2405.07283  [pdf, other

    cs.RO cs.CV

    BeautyMap: Binary-Encoded Adaptable Ground Matrix for Dynamic Points Removal in Global Maps

    Authors: Mingkai Jia, Qingwen Zhang, Bowen Yang, Jin Wu, Ming Liu, Patric Jensfelt

    Abstract: Global point clouds that correctly represent the static environment features can facilitate accurate localization and robust path planning. However, dynamic objects introduce undesired ghost tracks that are mixed up with the static environment. Existing dynamic removal methods normally fail to balance the performance in computational efficiency and accuracy. In response, we present BeautyMap to ef… ▽ More

    Submitted 12 May, 2024; originally announced May 2024.

    Comments: The first two authors are co-first authors. 8 pages, accepted by RA-L

  27. arXiv:2405.03633  [pdf, ps, other

    cs.CV cs.RO

    Neural Graph Map: Dense Mapping with Efficient Loop Closure Integration

    Authors: Leonard Bruns, Jun Zhang, Patric Jensfelt

    Abstract: Neural field-based SLAM methods typically employ a single, monolithic field as their scene representation. This prevents efficient incorporation of loop closure constraints and limits scalability. To address these shortcomings, we propose a novel RGB-D neural mapping framework in which the scene is represented by a collection of lightweight neural fields which are dynamically anchored to the pose… ▽ More

    Submitted 25 June, 2025; v1 submitted 6 May, 2024; originally announced May 2024.

    Comments: WACV 2025, Project page: https://kth-rpl.github.io/neural_graph_mapping/

  28. arXiv:2403.18649  [pdf, other

    cs.CV eess.SY

    Addressing Data Annotation Challenges in Multiple Sensors: A Solution for Scania Collected Datasets

    Authors: Ajinkya Khoche, Aron Asefaw, Alejandro Gonzalez, Bogdan Timus, Sina Sharif Mansouri, Patric Jensfelt

    Abstract: Data annotation in autonomous vehicles is a critical step in the development of Deep Neural Network (DNN) based models or the performance evaluation of the perception system. This often takes the form of adding 3D bounding boxes on time-sequential and registered series of point-sets captured from active sensors like Light Detection and Ranging (LiDAR) and Radio Detection and Ranging (RADAR). When… ▽ More

    Submitted 27 March, 2024; originally announced March 2024.

    Comments: Accepted to European Control Conference 2024

  29. arXiv:2403.17633  [pdf, other

    cs.CV cs.AI cs.RO

    UADA3D: Unsupervised Adversarial Domain Adaptation for 3D Object Detection with Sparse LiDAR and Large Domain Gaps

    Authors: Maciej K Wozniak, Mattias Hansson, Marko Thiel, Patric Jensfelt

    Abstract: In this study, we address a gap in existing unsupervised domain adaptation approaches on LiDAR-based 3D object detection, which have predominantly concentrated on adapting between established, high-density autonomous driving datasets. We focus on sparser point clouds, capturing scenarios from different perspectives: not just from vehicles on the road but also from mobile robots on sidewalks, which… ▽ More

    Submitted 21 October, 2024; v1 submitted 26 March, 2024; originally announced March 2024.

    Comments: Accepted for IEEE RA-L 2024

  30. arXiv:2403.11496  [pdf, other

    cs.RO cs.AI

    MCD: Diverse Large-Scale Multi-Campus Dataset for Robot Perception

    Authors: Thien-Minh Nguyen, Shenghai Yuan, Thien Hoang Nguyen, Pengyu Yin, Haozhi Cao, Lihua Xie, Maciej Wozniak, Patric Jensfelt, Marko Thiel, Justin Ziegenbein, Noel Blunder

    Abstract: Perception plays a crucial role in various robot applications. However, existing well-annotated datasets are biased towards autonomous driving scenarios, while unlabelled SLAM datasets are quickly over-fitted, and often lack environment and domain variations. To expand the frontier of these fields, we introduce a comprehensive dataset named MCD (Multi-Campus Dataset), featuring a wide range of sen… ▽ More

    Submitted 18 March, 2024; originally announced March 2024.

    Comments: Accepted by The IEEE/CVF Conference on Computer Vision and Pattern Recognition 2024

  31. arXiv:2403.01449  [pdf, other

    cs.RO cs.CV

    DUFOMap: Efficient Dynamic Awareness Mapping

    Authors: Daniel Duberg, Qingwen Zhang, MingKai Jia, Patric Jensfelt

    Abstract: The dynamic nature of the real world is one of the main challenges in robotics. The first step in dealing with it is to detect which parts of the world are dynamic. A typical benchmark task is to create a map that contains only the static part of the world to support, for example, localization and planning. Current solutions are often applied in post-processing, where parameter tuning allows the u… ▽ More

    Submitted 12 April, 2024; v1 submitted 3 March, 2024; originally announced March 2024.

    Comments: The first two authors hold equal contribution. 8 pages, 7 figures, project page https://kth-rpl.github.io/dufomap

  32. arXiv:2401.16122  [pdf, other

    cs.CV cs.RO

    DeFlow: Decoder of Scene Flow Network in Autonomous Driving

    Authors: Qingwen Zhang, Yi Yang, Heng Fang, Ruoyu Geng, Patric Jensfelt

    Abstract: Scene flow estimation determines a scene's 3D motion field, by predicting the motion of points in the scene, especially for aiding tasks in autonomous driving. Many networks with large-scale point clouds as input use voxelization to create a pseudo-image for real-time running. However, the voxelization process often results in the loss of point-specific features. This gives rise to a challenge in… ▽ More

    Submitted 29 January, 2024; originally announced January 2024.

    Comments: 7 pages, 4 figures, Code check https://github.com/KTH-RPL/deflow, accepted by ICRA 2024

  33. Transitional Grid Maps: Joint Modeling of Static and Dynamic Occupancy

    Authors: José Manuel Gaspar Sánchez, Leonard Bruns, Jana Tumova, Patric Jensfelt, Martin Törngren

    Abstract: Autonomous agents rely on sensor data to construct representations of their environments, essential for predicting future events and planning their actions. However, sensor measurements suffer from limited range, occlusions, and sensor noise. These challenges become more evident in highly dynamic environments. This work proposes a probabilistic framework to jointly infer which parts of an environm… ▽ More

    Submitted 4 November, 2024; v1 submitted 12 January, 2024; originally announced January 2024.

  34. arXiv:2310.04800  [pdf, other

    cs.CV cs.RO

    Towards Long-Range 3D Object Detection for Autonomous Vehicles

    Authors: Ajinkya Khoche, Laura Pereira Sánchez, Nazre Batool, Sina Sharif Mansouri, Patric Jensfelt

    Abstract: 3D object detection at long range is crucial for ensuring the safety and efficiency of self driving vehicles, allowing them to accurately perceive and react to objects, obstacles, and potential hazards from a distance. But most current state of the art LiDAR based methods are range limited due to sparsity at long range, which generates a form of domain gap between points closer to and farther away… ▽ More

    Submitted 20 May, 2024; v1 submitted 7 October, 2023; originally announced October 2023.

    Comments: Accepted to Intelligent Vehicle Symposium (IV) 2024

  35. arXiv:2307.07260  [pdf, other

    cs.RO cs.AI

    A Dynamic Points Removal Benchmark in Point Cloud Maps

    Authors: Qingwen Zhang, Daniel Duberg, Ruoyu Geng, Mingkai Jia, Lujia Wang, Patric Jensfelt

    Abstract: In the field of robotics, the point cloud has become an essential map representation. From the perspective of downstream tasks like localization and global path planning, points corresponding to dynamic objects will adversely affect their performance. Existing methods for removing dynamic points in point clouds often lack clarity in comparative evaluations and comprehensive analysis. Therefore, we… ▽ More

    Submitted 14 July, 2023; originally announced July 2023.

    Comments: Code check https://github.com/KTH-RPL/DynamicMap_Benchmark.git , 7 pages, accepted by ITSC 2023

  36. arXiv:2306.14589  [pdf, other

    cs.RO cs.HC

    Happily Error After: Framework Development and User Study for Correcting Robot Perception Errors in Virtual Reality

    Authors: Maciej K. Wozniak, Rebecca Stower, Patric Jensfelt, Andre Pereira

    Abstract: While we can see robots in more areas of our lives, they still make errors. One common cause of failure stems from the robot perception module when detecting objects. Allowing users to correct such errors can help improve the interaction and prevent the same errors in the future. Consequently, we investigate the effectiveness of a virtual reality (VR) framework for correcting perception errors of… ▽ More

    Submitted 26 June, 2023; originally announced June 2023.

    Comments: Accepted for IEEE RO-MAN 2023

  37. arXiv:2306.07344  [pdf, other

    cs.RO cs.CV

    Towards a Robust Sensor Fusion Step for 3D Object Detection on Corrupted Data

    Authors: Maciej K. Wozniak, Viktor Karefjards, Marko Thiel, Patric Jensfelt

    Abstract: Multimodal sensor fusion methods for 3D object detection have been revolutionizing the autonomous driving research field. Nevertheless, most of these methods heavily rely on dense LiDAR data and accurately calibrated sensors which is often not the case in real-world scenarios. Data from LiDAR and cameras often come misaligned due to the miscalibration, decalibration, or different frequencies of th… ▽ More

    Submitted 12 June, 2023; originally announced June 2023.

  38. arXiv:2301.08147  [pdf, other

    cs.CV cs.RO

    RGB-D-Based Categorical Object Pose and Shape Estimation: Methods, Datasets, and Evaluation

    Authors: Leonard Bruns, Patric Jensfelt

    Abstract: Recently, various methods for 6D pose and shape estimation of objects at a per-category level have been proposed. This work provides an overview of the field in terms of methods, datasets, and evaluation protocols. First, an overview of existing works and their commonalities and differences is provided. Second, we take a critical look at the predominant evaluation protocol, including metrics and d… ▽ More

    Submitted 19 January, 2023; originally announced January 2023.

    Comments: arXiv admin note: substantial text overlap with arXiv:2202.10346

  39. What you see is (not) what you get: A VR Framework for Correcting Robot Errors

    Authors: Maciej K. Wozniak, Rebecca Stower, Patric Jensfelt, Andre Pereira

    Abstract: Many solutions tailored for intuitive visualization or teleoperation of virtual, augmented and mixed (VAM) reality systems are not robust to robot failures, such as the inability to detect and recognize objects in the environment or planning unsafe trajectories. In this paper, we present a novel virtual reality (VR) framework where users can (i) recognize when the robot has failed to detect a real… ▽ More

    Submitted 17 January, 2023; v1 submitted 12 January, 2023; originally announced January 2023.

    Journal ref: HRI '23: Companion of the 2023 ACM/IEEE International Conference on Human-Robot Interaction

  40. arXiv:2301.02086  [pdf, other

    cs.CV cs.RO

    A Probabilistic Framework for Visual Localization in Ambiguous Scenes

    Authors: Fereidoon Zangeneh, Leonard Bruns, Amit Dekel, Alessandro Pieropan, Patric Jensfelt

    Abstract: Visual localization allows autonomous robots to relocalize when losing track of their pose by matching their current observation with past ones. However, ambiguous scenes pose a challenge for such systems, as repetitive structures can be viewed from many distinct, equally likely camera poses, which means it is not sufficient to produce a single best pose hypothesis. In this work, we propose a prob… ▽ More

    Submitted 5 January, 2023; originally announced January 2023.

  41. arXiv:2211.03900  [pdf, other

    cs.RO

    SLICT: Multi-input Multi-scale Surfel-Based Lidar-Inertial Continuous-Time Odometry and Mapping

    Authors: Thien-Minh Nguyen, Daniel Duberg, Patric Jensfelt, Shenghai Yuan, Lihua Xie

    Abstract: While feature association to a global map has significant benefits, to keep the computations from growing exponentially, most lidar-based odometry and mapping methods opt to associate features with local maps at one voxel scale. Taking advantage of the fact that surfels (surface elements) at different voxel scales can be organized in a tree-like structure, we propose an octree-based global map of… ▽ More

    Submitted 9 November, 2022; v1 submitted 7 November, 2022; originally announced November 2022.

  42. Semantic 3D Grid Maps for Autonomous Driving

    Authors: Ajinkya Khoche, Maciej K Wozniak, Daniel Duberg, Patric Jensfelt

    Abstract: Maps play a key role in rapidly developing area of autonomous driving. We survey the literature for different map representations and find that while the world is three-dimensional, it is common to rely on 2D map representations in order to meet real-time constraints. We believe that high levels of situation awareness require a 3D representation as well as the inclusion of semantic information. We… ▽ More

    Submitted 9 November, 2022; v1 submitted 3 November, 2022; originally announced November 2022.

    Comments: Submitted, accepted and presented at the 25th IEEE International Conference on Intelligent Transportation Systems (IEEE ITSC 2022)

  43. arXiv:2207.04880  [pdf, other

    cs.CV cs.RO

    SDFEst: Categorical Pose and Shape Estimation of Objects from RGB-D using Signed Distance Fields

    Authors: Leonard Bruns, Patric Jensfelt

    Abstract: Rich geometric understanding of the world is an important component of many robotic applications such as planning and manipulation. In this paper, we present a modular pipeline for pose and shape estimation of objects from RGB-D images given their category. The core of our method is a generative shape model, which we integrate with a novel initialization network and a differentiable renderer to en… ▽ More

    Submitted 11 July, 2022; originally announced July 2022.

    Comments: Accepted to IEEE Robotics and Automation Letters (and IROS 2022). Project page: https://github.com/roym899/sdfest

  44. arXiv:2205.02079  [pdf, other

    cs.CV cs.RO

    SDF-based RGB-D Camera Tracking in Neural Scene Representations

    Authors: Leonard Bruns, Fereidoon Zangeneh, Patric Jensfelt

    Abstract: We consider the problem of tracking the 6D pose of a moving RGB-D camera in a neural scene representation. Different such representations have recently emerged, and we investigate the suitability of them for the task of camera tracking. In particular, we propose to track an RGB-D camera using a signed distance field-based representation and show that compared to density-based representations, trac… ▽ More

    Submitted 4 May, 2022; originally announced May 2022.

    Comments: Accepted to the "Motion Planning with Implicit Neural Representations of Geometry" Workshop at ICRA 2022

  45. arXiv:2203.03385  [pdf, other

    cs.RO cs.CV

    FloorGenT: Generative Vector Graphic Model of Floor Plans for Robotics

    Authors: Ludvig Ericson, Patric Jensfelt

    Abstract: Floor plans are the basis of reasoning in and communicating about indoor environments. In this paper, we show that by modelling floor plans as sequences of line segments seen from a particular point of view, recent advances in autoregressive sequence modelling can be leveraged to model and predict floor plans. The line segments are canonicalized and translated to sequence of tokens and an attentio… ▽ More

    Submitted 7 March, 2022; originally announced March 2022.

    Comments: Submitted to IROS 2022. 7 pages, 6 figures

  46. arXiv:2202.10346  [pdf, other

    cs.CV cs.RO

    On the Evaluation of RGB-D-based Categorical Pose and Shape Estimation

    Authors: Leonard Bruns, Patric Jensfelt

    Abstract: Recently, various methods for 6D pose and shape estimation of objects have been proposed. Typically, these methods evaluate their pose estimation in terms of average precision, and reconstruction quality with chamfer distance. In this work we take a critical look at this predominant evaluation protocol including metrics and datasets. We propose a new set of metrics, contribute new annotations for… ▽ More

    Submitted 21 February, 2022; originally announced February 2022.

    Comments: 17 pages, 8 figures, submitted to IAS-17

  47. arXiv:2003.04749  [pdf, other

    cs.RO

    UFOMap: An Efficient Probabilistic 3D Mapping Framework That Embraces the Unknown

    Authors: Daniel Duberg, Patric Jensfelt

    Abstract: 3D models are an essential part of many robotic applications. In applications where the environment is unknown a-priori, or where only a part of the environment is known, it is important that the 3D model can handle the unknown space efficiently. Path planning, exploration, and reconstruction all fall into this category. In this paper we present an extension to OctoMap which we call UFOMap. UFOMap… ▽ More

    Submitted 10 March, 2020; originally announced March 2020.

    Comments: Project page: https://github.com/danielduberg/UFOMap

  48. arXiv:1912.03426  [pdf, other

    cs.CV cs.LG cs.RO

    Self-Supervised 3D Keypoint Learning for Ego-motion Estimation

    Authors: Jiexiong Tang, Rares Ambrus, Vitor Guizilini, Sudeep Pillai, Hanme Kim, Patric Jensfelt, Adrien Gaidon

    Abstract: Detecting and matching robust viewpoint-invariant keypoints is critical for visual SLAM and Structure-from-Motion. State-of-the-art learning-based methods generate training samples via homography adaptation to create 2D synthetic views with known keypoint matches from a single image. This approach, however, does not generalize to non-planar 3D scenes with illumination variations commonly seen in r… ▽ More

    Submitted 17 November, 2020; v1 submitted 6 December, 2019; originally announced December 2019.

  49. arXiv:1909.08812  [pdf, other

    cs.RO

    Flexible Disaster Response of Tomorrow -- Final Presentation and Evaluation of the CENTAURO System

    Authors: Tobias Klamt, Diego Rodriguez, Lorenzo Baccelliere, Xi Chen, Domenico Chiaradia, Torben Cichon, Massimiliano Gabardi, Paolo Guria, Karl Holmquist, Malgorzata Kamedula, Hakan Karaoguz, Navvab Kashiri, Arturo Laurenzi, Christian Lenz, Daniele Leonardis, Enrico Mingo Hoffman, Luca Muratore, Dmytro Pavlichenko, Francesco Porcini, Zeyu Ren, Fabian Schilling, Max Schwarz, Massimiliano Solazzi, Michael Felsberg, Antonio Frisoli , et al. (7 additional authors not shown)

    Abstract: Mobile manipulation robots have high potential to support rescue forces in disaster-response missions. Despite the difficulties imposed by real-world scenarios, robots are promising to perform mission tasks from a safe distance. In the CENTAURO project, we developed a disaster-response system which consists of the highly flexible Centauro robot and suitable control interfaces including an immersiv… ▽ More

    Submitted 19 September, 2019; originally announced September 2019.

    Comments: Accepted for IEEE Robotics and Automation Magazine (RAM), to appear December 2019

  50. arXiv:1909.07745  [pdf, other

    cs.RO cs.CV cs.LG

    Adversarial Feature Training for Generalizable Robotic Visuomotor Control

    Authors: Xi Chen, Ali Ghadirzadeh, Mårten Björkman, Patric Jensfelt

    Abstract: Deep reinforcement learning (RL) has enabled training action-selection policies, end-to-end, by learning a function which maps image pixels to action outputs. However, it's application to visuomotor robotic policy training has been limited because of the challenge of large-scale data collection when working with physical hardware. A suitable visuomotor policy should perform well not just for the t… ▽ More

    Submitted 17 September, 2019; originally announced September 2019.