-
Micro-DualNet: Dual-Path Spatio-Temporal Network for Micro-Action Recognition
Authors:
Naga VS Raviteja Chappa,
Evangelos Sariyanidi,
Lisa Yankowitz,
Gokul Nair,
Casey J. Zampella,
Robert T. Schultz,
Birkan Tunç
Abstract:
Micro-actions are subtle, localized movements lasting 1-3 seconds such as scratching one's head or tapping fingers. Such subtle actions are essential for social communication, ubiquitously used in natural interactions, and thus critical for fine-grained video understanding, yet remain poorly understood by current computer vision systems. We identify a fundamental challenge: micro-actions exhibit d…
▽ More
Micro-actions are subtle, localized movements lasting 1-3 seconds such as scratching one's head or tapping fingers. Such subtle actions are essential for social communication, ubiquitously used in natural interactions, and thus critical for fine-grained video understanding, yet remain poorly understood by current computer vision systems. We identify a fundamental challenge: micro-actions exhibit diverse spatio-temporal characteristics where some are defined by spatial configurations while others manifest through temporal dynamics. Existing methods that commit to a single spatio-temporal decomposition cannot accommodate this diversity. We propose a dual-path network that processes anatomically-grounded spatial entities through parallel Spatial-Temporal (ST) and Temporal-Spatial (TS) pathways. The ST path captures spatial configurations before modeling temporal dynamics, while the TS path inverts this order to prioritize temporal dynamics. Rather than fixed fusion, we introduce entity-level adaptive routing where each body part learns its optimal processing preference, complemented by Mutual Action Consistency (MAC) loss that enforces cross-path coherence. Extensive experiments demonstrate competitive performance on MA-52 dataset and state-of-the-art results on iMiGUE dataset. Our work reveals that architectural adaptation to the inherent complexity of micro-actions is essential for advancing fine-grained video understanding.
△ Less
Submitted 22 April, 2026;
originally announced April 2026.
-
Affect Decoding in Phonated and Silent Speech Production from Surface EMG
Authors:
Simon Pistrosch,
Kleanthis Avramidis,
Zhao Ren,
Tiantian Feng,
Jihwan Lee,
Monica Gonzalez-Machorro,
Anton Batliner,
Tanja Schultz,
Shrikanth Narayanan,
Björn W. Schuller
Abstract:
The expression of affect is integral to spoken communication, yet, its link to underlying articulatory execution remains unclear. Measures of articulatory muscle activity such as EMG could reveal how speech production is modulated by emotion alongside acoustic speech analyses. We investigate affect decoding from facial and neck surface electromyography (sEMG) during phonated and silent speech prod…
▽ More
The expression of affect is integral to spoken communication, yet, its link to underlying articulatory execution remains unclear. Measures of articulatory muscle activity such as EMG could reveal how speech production is modulated by emotion alongside acoustic speech analyses. We investigate affect decoding from facial and neck surface electromyography (sEMG) during phonated and silent speech production. For this purpose, we introduce a dataset comprising 2,780 utterances from 12 participants across 3 tasks, on which we evaluate both intra- and inter-subject decoding using a range of features and model embeddings. Our results reveal that EMG representations reliably discriminate frustration with up to 0.845 AUC, and generalize well across articulation modes. Our ablation study further demonstrates that affective signatures are embedded in facial motor activity and persist in the absence of phonation, highlighting the potential of EMG sensing for affect-aware silent speech interfaces.
△ Less
Submitted 19 March, 2026; v1 submitted 12 March, 2026;
originally announced March 2026.
-
Line congruences associated to Appell's hypergeometric functions of rank-4
Authors:
Matthew Ryan,
Michael T. Schultz
Abstract:
Line congruences are the genesis of important examples of transformations of projective surfaces, such as the Laplace transform. We survey and review results related to this historical subject, then derive original formulae for the Laplace transform of the entire rank-4 linear system associated to such an immersed projective surface. We apply our results to study the geometry of surfaces defined b…
▽ More
Line congruences are the genesis of important examples of transformations of projective surfaces, such as the Laplace transform. We survey and review results related to this historical subject, then derive original formulae for the Laplace transform of the entire rank-4 linear system associated to such an immersed projective surface. We apply our results to study the geometry of surfaces defined by Appell's hypergeometric functions of rank-4: namely, $F_2$ and $F_4$. We show that the sequence of Laplace invariants for each is determined respectively by the Euler-Poisson-Darboux equation for $F_2$, and Darboux's Harmonic equation for $F_4$. Further, we show the natural line congruences generated by the Laplace transforms of each constitute a $W$-congruence, an important example of line congruence in which a surface and its Laplace transform are simultaneously locally conformally equivalent.
△ Less
Submitted 23 February, 2026; v1 submitted 13 February, 2026;
originally announced February 2026.
-
A categorical perspective on extended metric-topological spaces
Authors:
Enrico Pasqualetto,
Timo Schultz,
Janne Taipalus
Abstract:
Motivated by the analysis and geometry of metric-measure structures in infinite dimensions, we study the category of extended metric-topological spaces, along with many of its distinguished subcategories (such as the one of compact spaces). One of the main achievements is the proof of the bicompleteness (i.e. of the existence of all small limits and colimits) of the aforementioned categories.
Motivated by the analysis and geometry of metric-measure structures in infinite dimensions, we study the category of extended metric-topological spaces, along with many of its distinguished subcategories (such as the one of compact spaces). One of the main achievements is the proof of the bicompleteness (i.e. of the existence of all small limits and colimits) of the aforementioned categories.
△ Less
Submitted 12 January, 2026;
originally announced January 2026.
-
Bitbox: Behavioral Imaging Toolbox for Computational Analysis of Behavior from Videos
Authors:
Evangelos Sariyanidi,
Gokul Nair,
Lisa Yankowitz,
Casey J. Zampella,
Mohan Kashyap Pargi,
Aashvi Manakiwala,
Maya McNealis,
John D. Herrington,
Jeffrey Cohn,
Robert T. Schultz,
Birkan Tunc
Abstract:
Computational measurement of human behavior from video has recently become feasible due to major advances in AI. These advances now enable granular and precise quantification of facial expression, head movement, body action, and other behavioral modalities and are increasingly used in psychology, psychiatry, neuroscience, and mental health research. However, mainstream adoption remains slow. Most…
▽ More
Computational measurement of human behavior from video has recently become feasible due to major advances in AI. These advances now enable granular and precise quantification of facial expression, head movement, body action, and other behavioral modalities and are increasingly used in psychology, psychiatry, neuroscience, and mental health research. However, mainstream adoption remains slow. Most existing methods and software are developed for engineering audiences, require specialized software stacks, and fail to provide behavioral measurements at a level directly useful for hypothesis-driven research. As a result, there is a large barrier to entry for researchers who wish to use modern, AI-based tools in their work. We introduce Bitbox, an open-source toolkit designed to remove this barrier and make advanced computational analysis directly usable by behavioral scientists and clinical researchers. Bitbox is guided by principles of reproducibility, modularity, and interpretability. It provides a standardized interface for extracting high-level behavioral measurements from video, leveraging multiple face, head, and body processors. The core modules have been tested and validated on clinical samples and are designed so that new measures can be added with minimal effort. Bitbox is intended to serve both sides of the translational gap. It gives behavioral researchers access to robust, high-level behavioral metrics without requiring engineering expertise, and it provides computer scientists a practical mechanism for disseminating methods to domains where their impact is most needed. We expect that Bitbox will accelerate integration of computational behavioral measurement into behavioral, clinical, and mental health research. Bitbox has been designed from the beginning as a community-driven effort that will evolve through contributions from both method developers and domain scientists.
△ Less
Submitted 19 December, 2025;
originally announced December 2025.
-
Concurrence: A dependence criterion for time series, applied to biological data
Authors:
Evangelos Sariyanidi,
John D. Herrington,
Lisa Yankowitz,
Pratik Chaudhari,
Theodore D. Satterthwaite,
Casey J. Zampella,
Jeffrey S. Morris,
Edward Gunning,
Robert T. Schultz,
Russell T. Shinohara,
Birkan Tunc
Abstract:
Measuring the statistical dependence between observed signals is a primary tool for scientific discovery. However, biological systems often exhibit complex non-linear interactions that currently cannot be captured without a priori knowledge or large datasets. We introduce a criterion for dependence, whereby two time series are deemed dependent if one can construct a classifier that distinguishes b…
▽ More
Measuring the statistical dependence between observed signals is a primary tool for scientific discovery. However, biological systems often exhibit complex non-linear interactions that currently cannot be captured without a priori knowledge or large datasets. We introduce a criterion for dependence, whereby two time series are deemed dependent if one can construct a classifier that distinguishes between temporally aligned vs. misaligned segments extracted from them. We show that this criterion, concurrence, is theoretically linked with dependence, and can become a standard approach for scientific analyses across disciplines, as it can expose relationships across a wide spectrum of signals (fMRI, physiological and behavioral data) without ad-hoc parameter tuning or large amounts of data.
△ Less
Submitted 22 April, 2026; v1 submitted 17 December, 2025;
originally announced December 2025.
-
Breathe with Me: Synchronizing Biosignals for User Embodiment in Robots
Authors:
Iddo Yehoshua Wald,
Amber Maimon,
Shiyao Zhang,
Dennis Küster,
Robert Porzel,
Tanja Schultz,
Rainer Malaka
Abstract:
Embodiment of users within robotic systems has been explored in human-robot interaction, most often in telepresence and teleoperation. In these applications, synchronized visuomotor feedback can evoke a sense of body ownership and agency, contributing to the experience of embodiment. We extend this work by employing embreathment, the representation of the user's own breath in real time, as a means…
▽ More
Embodiment of users within robotic systems has been explored in human-robot interaction, most often in telepresence and teleoperation. In these applications, synchronized visuomotor feedback can evoke a sense of body ownership and agency, contributing to the experience of embodiment. We extend this work by employing embreathment, the representation of the user's own breath in real time, as a means for enhancing user embodiment experience in robots. In a within-subjects experiment, participants controlled a robotic arm, while its movements were either synchronized or non-synchronized with their own breath. Synchrony was shown to significantly increase body ownership, and was preferred by most participants. We propose the representation of physiological signals as a novel interoceptive pathway for human-robot interaction, and discuss implications for telepresence, prosthetics, collaboration with robots, and shared autonomy.
△ Less
Submitted 16 December, 2025;
originally announced December 2025.
-
DBT-DINO: Towards Foundation model based analysis of Digital Breast Tomosynthesis
Authors:
Felix J. Dorfner,
Manon A. Dorster,
Ryan Connolly,
Oscar Gentilhomme,
Edward Gibbs,
Steven Graham,
Seth Wander,
Thomas Schultz,
Manisha Bahl,
Dania Daye,
Albert E. Kim,
Christopher P. Bridge
Abstract:
Foundation models have shown promise in medical imaging but remain underexplored for three-dimensional imaging modalities. No foundation model currently exists for Digital Breast Tomosynthesis (DBT), despite its use for breast cancer screening.
To develop and evaluate a foundation model for DBT (DBT-DINO) across multiple clinical tasks and assess the impact of domain-specific pre-training.
Sel…
▽ More
Foundation models have shown promise in medical imaging but remain underexplored for three-dimensional imaging modalities. No foundation model currently exists for Digital Breast Tomosynthesis (DBT), despite its use for breast cancer screening.
To develop and evaluate a foundation model for DBT (DBT-DINO) across multiple clinical tasks and assess the impact of domain-specific pre-training.
Self-supervised pre-training was performed using the DINOv2 methodology on over 25 million 2D slices from 487,975 DBT volumes from 27,990 patients. Three downstream tasks were evaluated: (1) breast density classification using 5,000 screening exams; (2) 5-year risk of developing breast cancer using 106,417 screening exams; and (3) lesion detection using 393 annotated volumes.
For breast density classification, DBT-DINO achieved an accuracy of 0.79 (95\% CI: 0.76--0.81), outperforming both the MetaAI DINOv2 baseline (0.73, 95\% CI: 0.70--0.76, p<.001) and DenseNet-121 (0.74, 95\% CI: 0.71--0.76, p<.001). For 5-year breast cancer risk prediction, DBT-DINO achieved an AUROC of 0.78 (95\% CI: 0.76--0.80) compared to DINOv2's 0.76 (95\% CI: 0.74--0.78, p=.57). For lesion detection, DINOv2 achieved a higher average sensitivity of 0.67 (95\% CI: 0.60--0.74) compared to DBT-DINO with 0.62 (95\% CI: 0.53--0.71, p=.60). DBT-DINO demonstrated better performance on cancerous lesions specifically with a detection rate of 78.8\% compared to Dinov2's 77.3\%.
Using a dataset of unprecedented size, we developed DBT-DINO, the first foundation model for DBT. DBT-DINO demonstrated strong performance on breast density classification and cancer risk prediction. However, domain-specific pre-training showed variable benefits on the detection task, with ImageNet baseline outperforming DBT-DINO on general lesion detection, indicating that localized detection tasks require further methodological development.
△ Less
Submitted 15 December, 2025;
originally announced December 2025.
-
CataractCompDetect: Intraoperative Complication Detection in Cataract Surgery
Authors:
Bhuvan Sachdeva,
Sneha Kumari,
Rudransh Agarwal,
Shalaka Kumaraswamy,
Niharika Singri Prasad,
Simon Mueller,
Raphael Lechtenboehmer,
Maximilian W. M. Wintergerst,
Thomas Schultz,
Kaushik Murali,
Mohit Jain
Abstract:
Cataract surgery is one of the most commonly performed surgeries worldwide, yet intraoperative complications such as iris prolapse, posterior capsule rupture (PCR), and vitreous loss remain major causes of adverse outcomes. Automated detection of such events could enable early warning systems and objective training feedback. In this work, we propose CataractCompDetect, a complication detection fra…
▽ More
Cataract surgery is one of the most commonly performed surgeries worldwide, yet intraoperative complications such as iris prolapse, posterior capsule rupture (PCR), and vitreous loss remain major causes of adverse outcomes. Automated detection of such events could enable early warning systems and objective training feedback. In this work, we propose CataractCompDetect, a complication detection framework that combines phase-aware localization, SAM 2-based tracking, complication-specific risk scoring, and vision-language reasoning for final classification. To validate CataractCompDetect, we curate CataComp, the first cataract surgery video dataset annotated for intraoperative complications, comprising 53 surgeries, including 23 with clinical complications. On CataComp, CataractCompDetect achieves an average F1 score of 70.63%, with per-complication performance of 81.8% (Iris Prolapse), 60.87% (PCR), and 69.23% (Vitreous Loss). These results highlight the value of combining structured surgical priors with vision-language reasoning for recognizing rare but high-impact intraoperative events. Our dataset and code will be publicly released upon acceptance.
△ Less
Submitted 24 November, 2025;
originally announced November 2025.
-
Synthetic approaches to Ricci flows
Authors:
Matthias Erbar,
Marco Flaim,
Eric Hupp,
Zhenhao Li,
Timo Schultz,
Karl-Theodor Sturm
Abstract:
We review different notions of synthetic Ricci flow that apply to time-dependent families of metric measure spaces and which are based on properties of the heat flow, ideas from optimal transport, and the asymptotic behaviour of volumes. Each notion equivalently characterises (weighted) Ricci flow for smooth families of weighted Riemannian manifolds. We discuss the features of the different notion…
▽ More
We review different notions of synthetic Ricci flow that apply to time-dependent families of metric measure spaces and which are based on properties of the heat flow, ideas from optimal transport, and the asymptotic behaviour of volumes. Each notion equivalently characterises (weighted) Ricci flow for smooth families of weighted Riemannian manifolds. We discuss the features of the different notions on various examples.
△ Less
Submitted 14 November, 2025;
originally announced November 2025.
-
Geometric aspects of rank-3 vector bundles over surfaces and 2-plane distributions on 5-manifolds
Authors:
Brandon P. Ashley,
Michael T. Schultz
Abstract:
We study geometric aspects of horizontal 2-plane distributions on the complement of the zero section in the 5-dimensional total space of a rank-3 vector bundle equipped with connection over a surface. We show that any surface in 3-dimensional projective space can be associated to such a geometric structure in 5-dimensions, and establish a dictionary between the projective differential geometry of…
▽ More
We study geometric aspects of horizontal 2-plane distributions on the complement of the zero section in the 5-dimensional total space of a rank-3 vector bundle equipped with connection over a surface. We show that any surface in 3-dimensional projective space can be associated to such a geometric structure in 5-dimensions, and establish a dictionary between the projective differential geometry of the surface and the growth vector of the 2-plane distribution.
△ Less
Submitted 11 December, 2025; v1 submitted 29 October, 2025;
originally announced October 2025.
-
EASELAN: An Open-Source Framework for Multimodal Biosignal Annotation and Data Management
Authors:
Rathi Adarshi Rammohan,
Moritz Meier,
Dennis Küster,
Tanja Schultz
Abstract:
Recent advancements in machine learning and adaptive cognitive systems are driving a growing demand for large and richly annotated multimodal data. A prominent example of this trend are fusion models, which increasingly incorporate multiple biosignals in addition to traditional audiovisual channels. This paper introduces the EASELAN annotation framework to improve annotation workflows designed to…
▽ More
Recent advancements in machine learning and adaptive cognitive systems are driving a growing demand for large and richly annotated multimodal data. A prominent example of this trend are fusion models, which increasingly incorporate multiple biosignals in addition to traditional audiovisual channels. This paper introduces the EASELAN annotation framework to improve annotation workflows designed to address the resulting rising complexity of multimodal and biosignals datasets. It builds on the robust ELAN tool by adding new components tailored to support all stages of the annotation pipeline: From streamlining the preparation of annotation files to setting up additional channels, integrated version control with GitHub, and simplified post-processing. EASELAN delivers a seamless workflow designed to integrate biosignals and facilitate rich annotations to be readily exported for further analyses and machine learning-supported model training. The EASELAN framework is successfully applied to a high-dimensional biosignals collection initiative on human everyday activities (here, table setting) for cognitive robots within the DFG-funded Collaborative Research Center 1320 Everyday Activity Science and Engineering (EASE). In this paper we discuss the opportunities, limitations, and lessons learned when using EASELAN for this initiative. To foster research on biosignal collection, annotation, and processing, the code of EASELAN is publicly available(https://github.com/cognitive-systems-lab/easelan), along with the EASELAN-supported fully annotated Table Setting Database.
△ Less
Submitted 17 October, 2025;
originally announced October 2025.
-
Machine Unlearning in Speech Emotion Recognition via Forget Set Alone
Authors:
Zhao Ren,
Rathi Adarshi Rammohan,
Kevin Scheck,
Tanja Schultz
Abstract:
Speech emotion recognition aims to identify emotional states from speech signals and has been widely applied in human-computer interaction, education, healthcare, and many other fields. However, since speech data contain rich sensitive information, partial data can be required to be deleted by speakers due to privacy concerns. Current machine unlearning approaches largely depend on data beyond the…
▽ More
Speech emotion recognition aims to identify emotional states from speech signals and has been widely applied in human-computer interaction, education, healthcare, and many other fields. However, since speech data contain rich sensitive information, partial data can be required to be deleted by speakers due to privacy concerns. Current machine unlearning approaches largely depend on data beyond the samples to be forgotten. However, this reliance poses challenges when data redistribution is restricted and demands substantial computational resources in the context of big data. We propose a novel adversarial-attack-based approach that fine-tunes a pre-trained speech emotion recognition model using only the data to be forgotten. The experimental results demonstrate that the proposed approach can effectively remove the knowledge of the data to be forgotten from the model, while preserving high model performance on the test set for emotion recognition.
△ Less
Submitted 22 December, 2025; v1 submitted 5 October, 2025;
originally announced October 2025.
-
An Introduction to Silent Paralinguistics
Authors:
Zhao Ren,
Simon Pistrosch,
Buket Coşkun,
Kevin Scheck,
Anton Batliner,
Björn W. Schuller,
Tanja Schultz
Abstract:
The ability to speak is an inherent part of human nature and fundamental to our existence as a social species. Unfortunately, this ability can be restricted in certain situations, such as for individuals who have lost their voice or in environments where speaking aloud is unsuitable. Additionally, some people may prefer not to speak audibly due to privacy concerns. For such cases, silent speech in…
▽ More
The ability to speak is an inherent part of human nature and fundamental to our existence as a social species. Unfortunately, this ability can be restricted in certain situations, such as for individuals who have lost their voice or in environments where speaking aloud is unsuitable. Additionally, some people may prefer not to speak audibly due to privacy concerns. For such cases, silent speech interfaces have been proposed, which focus on processing biosignals corresponding to silently produced speech. These interfaces enable synthesis of audible speech from biosignals that are produced when speaking silently and recognition aka decoding of biosignals into text that corresponds to the silently produced speech. While recognition and synthesis of silent speech has been a prominent focus in many research studies, there is a significant gap in deriving paralinguistic information such as affective states from silent speech. To fill this gap, we propose Silent Paralinguistics, aiming to predict paralinguistic information from silent speech and ultimately integrate it into the reconstructed audible voice for natural communication. This survey provides a comprehensive look at methods, research strategies, and objectives within the emerging field of silent paralinguistics.
△ Less
Submitted 25 August, 2025;
originally announced August 2025.
-
Measuring Dependencies between Biological Signals with Self-supervision, and its Limitations
Authors:
Evangelos Sariyanidi,
John D. Herrington,
Lisa Yankowitz,
Pratik Chaudhari,
Theodore D. Satterthwaite,
Casey J. Zampella,
Robert T. Schultz,
Russell T. Shinohara,
Birkan Tunc
Abstract:
Measuring the statistical dependence between observed signals is a primary tool for scientific discovery. However, biological systems often exhibit complex non-linear interactions that currently cannot be captured without a priori knowledge regarding the nature of dependence. We introduce a self-supervised approach, concurrence, which is inspired by the observation that if two signals are dependen…
▽ More
Measuring the statistical dependence between observed signals is a primary tool for scientific discovery. However, biological systems often exhibit complex non-linear interactions that currently cannot be captured without a priori knowledge regarding the nature of dependence. We introduce a self-supervised approach, concurrence, which is inspired by the observation that if two signals are dependent, then one should be able to distinguish between temporally aligned vs. misaligned segments extracted from them. Experiments with fMRI, physiological and behavioral signals show that, to our knowledge, concurrence is the first approach that can expose relationships across such a wide spectrum of signals and extract scientifically relevant differences without ad-hoc parameter tuning or reliance on a priori information, providing a potent tool for scientific discoveries across fields. However, dependencies caused by extraneous factors remain an open problem, thus researchers should validate that exposed relationships truly pertain to the question(s) of interest.
△ Less
Submitted 8 August, 2025; v1 submitted 29 July, 2025;
originally announced August 2025.
-
Examining the Effects of Human-Likeness of Avatars on Emotion Perception and Emotion Elicitation
Authors:
Shiyao Zhang,
Omar Faruk,
Robert Porzel,
Dennis Küster,
Tanja Schultz,
Hui Liu
Abstract:
An increasing number of online interaction settings now provide the possibility to visually represent oneself via an animated avatar instead of a video stream. Benefits include protecting the communicator's privacy while still providing a means to express their individuality. In consequence, there has been a surge in means for avatar-based personalization, ranging from classic human representation…
▽ More
An increasing number of online interaction settings now provide the possibility to visually represent oneself via an animated avatar instead of a video stream. Benefits include protecting the communicator's privacy while still providing a means to express their individuality. In consequence, there has been a surge in means for avatar-based personalization, ranging from classic human representations to animals, food items, and more. However, using avatars also has drawbacks. Depending on the human-likeness of the avatar and the corresponding disparities between the avatar and the original expresser, avatars may elicit discomfort or even hinder effective nonverbal communication by distorting emotion perception. This study examines the relationship between the human-likeness of virtual avatars and emotion perception for Ekman's six "basic emotions". Research reveals that avatars with varying degrees of human-likeness have distinct effects on emotion perception. High human-likeness avatars, such as human avatars, tend to elicit more negative emotional responses from users, a phenomenon that is consistent with the concept of Uncanny Valley in aesthetics, which suggests that closely resembling humans can provoke negative emotional responses. Conversely, a raccoon avatar and a shark avatar, known as cuteness, which exhibit moderate human similarity in this study, demonstrate a positive influence on emotion perception. Our initial results suggest that the human-likeness of avatars is an important factor for emotion perception. The results from the follow-up study further suggest that the cuteness of avatars and their natural facial status may also play a significant role in emotion perception and elicitation. We discuss practical implications for strategically conveying specific human behavioral messages through avatars in multiple applications, such as business and counseling.
△ Less
Submitted 3 August, 2025;
originally announced August 2025.
-
End-to-end Acoustic-linguistic Emotion and Intent Recognition Enhanced by Semi-supervised Learning
Authors:
Zhao Ren,
Rathi Adarshi Rammohan,
Kevin Scheck,
Sheng Li,
Tanja Schultz
Abstract:
Emotion and intent recognition from speech is essential and has been widely investigated in human-computer interaction. The rapid development of social media platforms, chatbots, and other technologies has led to a large volume of speech data streaming from users. Nevertheless, annotating such data manually is expensive, making it challenging to train machine learning models for recognition purpos…
▽ More
Emotion and intent recognition from speech is essential and has been widely investigated in human-computer interaction. The rapid development of social media platforms, chatbots, and other technologies has led to a large volume of speech data streaming from users. Nevertheless, annotating such data manually is expensive, making it challenging to train machine learning models for recognition purposes. To this end, we propose applying semi-supervised learning to incorporate a large scale of unlabelled data alongside a relatively smaller set of labelled data. We train end-to-end acoustic and linguistic models, each employing multi-task learning for emotion and intent recognition. Two semi-supervised learning approaches, including fix-match learning and full-match learning, are compared. The experimental results demonstrate that the semi-supervised learning approaches improve model performance in speech emotion and intent recognition from both acoustic and text data. The late fusion of the best models outperforms the acoustic and text baselines by joint recognition balance metrics of 12.3% and 10.4%, respectively.
△ Less
Submitted 10 July, 2025;
originally announced July 2025.
-
Beyond FACS: Data-driven Facial Expression Dictionaries, with Application to Predicting Autism
Authors:
Evangelos Sariyanidi,
Lisa Yankowitz,
Robert T. Schultz,
John D. Herrington,
Birkan Tunc,
Jeffrey Cohn
Abstract:
The Facial Action Coding System (FACS) has been used by numerous studies to investigate the links between facial behavior and mental health. The laborious and costly process of FACS coding has motivated the development of machine learning frameworks for Action Unit (AU) detection. Despite intense efforts spanning three decades, the detection accuracy for many AUs is considered to be below the thre…
▽ More
The Facial Action Coding System (FACS) has been used by numerous studies to investigate the links between facial behavior and mental health. The laborious and costly process of FACS coding has motivated the development of machine learning frameworks for Action Unit (AU) detection. Despite intense efforts spanning three decades, the detection accuracy for many AUs is considered to be below the threshold needed for behavioral research. Also, many AUs are excluded altogether, making it impossible to fulfill the ultimate goal of FACS-the representation of any facial expression in its entirety. This paper considers an alternative approach. Instead of creating automated tools that mimic FACS experts, we propose to use a new coding system that mimics the key properties of FACS. Specifically, we construct a data-driven coding system called the Facial Basis, which contains units that correspond to localized and interpretable 3D facial movements, and overcomes three structural limitations of automated FACS coding. First, the proposed method is completely unsupervised, bypassing costly, laborious and variable manual annotation. Second, Facial Basis reconstructs all observable movement, rather than relying on a limited repertoire of recognizable movements (as in automated FACS). Finally, the Facial Basis units are additive, whereas AUs may fail detection when they appear in a non-additive combination. The proposed method outperforms the most frequently used AU detector in predicting autism diagnosis from in-person and remote conversations, highlighting the importance of encoding facial behavior comprehensively. To our knowledge, Facial Basis is the first alternative to FACS for deconstructing facial expressions in videos into localized movements. We provide an open source implementation of the method at github.com/sariyanidi/FacialBasis.
△ Less
Submitted 30 May, 2025;
originally announced May 2025.
-
Global Context Is All You Need for Parallel Efficient Tractography Parcellation
Authors:
Valentin von Bornhaupt,
Johannes Grün,
and Justus Bisten,
Tobias Bauer,
Theodor Rüber,
Thomas Schultz
Abstract:
Whole-brain tractography in diffusion MRI is often followed by a parcellation in which each streamline is classified as belonging to a specific white matter bundle, or discarded as a false positive. Efficient parcellation is important both in large-scale studies, which have to process huge amounts of data, and in the clinic, where computational resources are often limited. TractCloud is a state-of…
▽ More
Whole-brain tractography in diffusion MRI is often followed by a parcellation in which each streamline is classified as belonging to a specific white matter bundle, or discarded as a false positive. Efficient parcellation is important both in large-scale studies, which have to process huge amounts of data, and in the clinic, where computational resources are often limited. TractCloud is a state-of-the-art approach that aims to maximize accuracy with a local-global representation. We demonstrate that the local context does not contribute to the accuracy of that approach, and is even detrimental when dealing with pathological cases. Based on this observation, we propose PETParc, a new method for Parallel Efficient Tractography Parcellation. PETParc is a transformer-based architecture in which the whole-brain tractogram is randomly partitioned into sub-tractograms whose streamlines are classified in parallel, while serving as global context for each other. This leads to a speedup of up to two orders of magnitude relative to TractCloud, and permits inference even on clinical workstations without a GPU. PETParc accounts for the lack of streamline orientation either via a novel flip-invariant embedding, or by simply using flips as part of data augmentation. Despite the speedup, results are often even better than those of prior methods. The code and pretrained model will be made public upon acceptance.
△ Less
Submitted 10 March, 2025;
originally announced March 2025.
-
A Modular Pipeline for 3D Object Tracking Using RGB Cameras
Authors:
Lars Bredereke,
Yale Hartmann,
Tanja Schultz
Abstract:
Object tracking is a key challenge of computer vision with various applications that all require different architectures. Most tracking systems have limitations such as constraining all movement to a 2D plane and they often track only one object. In this paper, we present a new modular pipeline that calculates 3D trajectories of multiple objects. It is adaptable to various settings where multiple…
▽ More
Object tracking is a key challenge of computer vision with various applications that all require different architectures. Most tracking systems have limitations such as constraining all movement to a 2D plane and they often track only one object. In this paper, we present a new modular pipeline that calculates 3D trajectories of multiple objects. It is adaptable to various settings where multiple time-synced and stationary cameras record moving objects, using off the shelf webcams. Our pipeline was tested on the Table Setting Dataset, where participants are recorded with various sensors as they set a table with tableware objects. We need to track these manipulated objects, using 6 rgb webcams. Challenges include: Detecting small objects in 9.874.699 camera frames, determining camera poses, discriminating between nearby and overlapping objects, temporary occlusions, and finally calculating a 3D trajectory using the right subset of an average of 11.12.456 pixel coordinates per 3-minute trial. We implement a robust pipeline that results in accurate trajectories with covariance of x,y,z-position as a confidence metric. It deals dynamically with appearing and disappearing objects, instantiating new Extended Kalman Filters. It scales to hundreds of table-setting trials with very little human annotation input, even with the camera poses of each trial unknown. The code is available at https://github.com/LarsBredereke/object_tracking
△ Less
Submitted 6 March, 2025;
originally announced March 2025.
-
Synthetic notions of Ricci flow for metric measure spaces
Authors:
Matthias Erbar,
Zhenhao Li,
Timo Schultz
Abstract:
We develop different synthetic notions of Ricci flow in the setting of time-dependent metric measure spaces based on ideas from optimal transport. They are formulated in terms of dynamic convexity and local concavity of the entropy along Wasserstein geodesics on the one hand and in terms of global and short-time asymptotic transport cost estimates for the heat flow on the other hand. We show that…
▽ More
We develop different synthetic notions of Ricci flow in the setting of time-dependent metric measure spaces based on ideas from optimal transport. They are formulated in terms of dynamic convexity and local concavity of the entropy along Wasserstein geodesics on the one hand and in terms of global and short-time asymptotic transport cost estimates for the heat flow on the other hand. We show that these properties characterise smooth (weighted) Ricci flows. Further, we investigate the relation between the different notions in the non-smooth setting of time-dependent metric measure spaces.
△ Less
Submitted 13 January, 2025;
originally announced January 2025.
-
Weakly Supervised Segmentation of Hyper-Reflective Foci with Compact Convolutional Transformers and SAM2
Authors:
Olivier Morelle,
Justus Bisten,
Maximilian W. M. Wintergerst,
Robert P. Finger,
Thomas Schultz
Abstract:
Weakly supervised segmentation has the potential to greatly reduce the annotation effort for training segmentation models for small structures such as hyper-reflective foci (HRF) in optical coherence tomography (OCT). However, most weakly supervised methods either involve a strong downsampling of input images, or only achieve localization at a coarse resolution, both of which are unsatisfactory fo…
▽ More
Weakly supervised segmentation has the potential to greatly reduce the annotation effort for training segmentation models for small structures such as hyper-reflective foci (HRF) in optical coherence tomography (OCT). However, most weakly supervised methods either involve a strong downsampling of input images, or only achieve localization at a coarse resolution, both of which are unsatisfactory for small structures. We propose a novel framework that increases the spatial resolution of a traditional attention-based Multiple Instance Learning (MIL) approach by using Layer-wise Relevance Propagation (LRP) to prompt the Segment Anything Model (SAM~2), and increases recall with iterative inference. Moreover, we demonstrate that replacing MIL with a Compact Convolutional Transformer (CCT), which adds a positional encoding, and permits an exchange of information between different regions of the OCT image, leads to a further and substantial increase in segmentation accuracy.
△ Less
Submitted 21 March, 2025; v1 submitted 10 January, 2025;
originally announced January 2025.
-
Deep Speech Synthesis from Multimodal Articulatory Representations
Authors:
Peter Wu,
Bohan Yu,
Kevin Scheck,
Alan W Black,
Aditi S. Krishnapriyan,
Irene Y. Chen,
Tanja Schultz,
Shinji Watanabe,
Gopala K. Anumanchipalli
Abstract:
The amount of articulatory data available for training deep learning models is much less compared to acoustic speech data. In order to improve articulatory-to-acoustic synthesis performance in these low-resource settings, we propose a multimodal pre-training framework. On single-speaker speech synthesis tasks from real-time magnetic resonance imaging and surface electromyography inputs, the intell…
▽ More
The amount of articulatory data available for training deep learning models is much less compared to acoustic speech data. In order to improve articulatory-to-acoustic synthesis performance in these low-resource settings, we propose a multimodal pre-training framework. On single-speaker speech synthesis tasks from real-time magnetic resonance imaging and surface electromyography inputs, the intelligibility of synthesized outputs improves noticeably. For example, compared to prior work, utilizing our proposed transfer learning methods improves the MRI-to-speech performance by 36% word error rate. In addition to these intelligibility results, our multimodal pre-trained models consistently outperform unimodal baselines on three objective and subjective synthesis quality metrics.
△ Less
Submitted 17 December, 2024;
originally announced December 2024.
-
Phase-Informed Tool Segmentation for Manual Small-Incision Cataract Surgery
Authors:
Bhuvan Sachdeva,
Naren Akash,
Tajamul Ashraf,
Simon Mueller,
Thomas Schultz,
Maximilian W. M. Wintergerst,
Niharika Singri Prasad,
Kaushik Murali,
Mohit Jain
Abstract:
Cataract surgery is the most common surgical procedure globally, with a disproportionately higher burden in developing countries. While automated surgical video analysis has been explored in general surgery, its application to ophthalmic procedures remains limited. Existing works primarily focus on Phaco cataract surgery, an expensive technique not accessible in regions where cataract treatment is…
▽ More
Cataract surgery is the most common surgical procedure globally, with a disproportionately higher burden in developing countries. While automated surgical video analysis has been explored in general surgery, its application to ophthalmic procedures remains limited. Existing works primarily focus on Phaco cataract surgery, an expensive technique not accessible in regions where cataract treatment is most needed. In contrast, Manual Small-Incision Cataract Surgery (MSICS) is the preferred low-cost, faster alternative in high-volume settings and for challenging cases. However, no dataset exists for MSICS. To address this gap, we introduce Sankara-MSICS, the first comprehensive dataset containing 53 surgical videos annotated for 18 surgical phases and 3,527 frames with 13 surgical tools at the pixel level. We benchmark this dataset on state-of-the-art models and present ToolSeg, a novel framework that enhances tool segmentation by introducing a phase-conditional decoder and a simple yet effective semi-supervised setup leveraging pseudo-labels from foundation models. Our approach significantly improves segmentation performance, achieving a $23.77\%$ to $38.10\%$ increase in mean Dice scores, with a notable boost for tools that are less prevalent and small. Furthermore, we demonstrate that ToolSeg generalizes to other surgical settings, showcasing its effectiveness on the CaDIS dataset.
△ Less
Submitted 3 December, 2024; v1 submitted 25 November, 2024;
originally announced November 2024.
-
Metric conditions that guarantee existence and uniqueness of Optimal Transport maps
Authors:
Shucheng Li,
Mattia Magnabosco,
Timo Schultz
Abstract:
We investigate metric conditions that allow to prove existence and uniqueness of a map solving the Monge problem between two marginals in a metric (measure) space, proving two main results. Firstly, we introduce a nonsmooth version of the Riemannian twist condition that we call local metric twist condition, showing, under this assumption on the cost function, existence and uniqueness of optimal tr…
▽ More
We investigate metric conditions that allow to prove existence and uniqueness of a map solving the Monge problem between two marginals in a metric (measure) space, proving two main results. Firstly, we introduce a nonsmooth version of the Riemannian twist condition that we call local metric twist condition, showing, under this assumption on the cost function, existence and uniqueness of optimal transport maps. Secondly, we prove the same result for cost equal to $d^2$ in a metric space $(X, d)$ satisfying a quantitative non-branching assumption, that we call locally-uniformly non-branching.
△ Less
Submitted 29 October, 2024;
originally announced October 2024.
-
Investigating Effective Speaker Property Privacy Protection in Federated Learning for Speech Emotion Recognition
Authors:
Chao Tan,
Sheng Li,
Yang Cao,
Zhao Ren,
Tanja Schultz
Abstract:
Federated Learning (FL) is a privacy-preserving approach that allows servers to aggregate distributed models transmitted from local clients rather than training on user data. More recently, FL has been applied to Speech Emotion Recognition (SER) for secure human-computer interaction applications. Recent research has found that FL is still vulnerable to inference attacks. To this end, this paper fo…
▽ More
Federated Learning (FL) is a privacy-preserving approach that allows servers to aggregate distributed models transmitted from local clients rather than training on user data. More recently, FL has been applied to Speech Emotion Recognition (SER) for secure human-computer interaction applications. Recent research has found that FL is still vulnerable to inference attacks. To this end, this paper focuses on investigating the security of FL for SER concerning property inference attacks. We propose a novel method to protect the property information in speech data by decomposing various properties in the sound and adding perturbations to these properties. Our experiments show that the proposed method offers better privacy-utility trade-offs than existing methods. The trade-offs enable more effective attack prevention while maintaining similar FL utility levels. This work can guide future work on privacy protection methods in speech processing.
△ Less
Submitted 17 October, 2024;
originally announced October 2024.
-
Speech as a Biomarker for Disease Detection
Authors:
Catarina Botelho,
Alberto Abad,
Tanja Schultz,
Isabel Trancoso
Abstract:
Speech is a rich biomarker that encodes substantial information about the health of a speaker, and thus it has been proposed for the detection of numerous diseases, achieving promising results. However, questions remain about what the models trained for the automatic detection of these diseases are actually learning and the basis for their predictions, which can significantly impact patients' live…
▽ More
Speech is a rich biomarker that encodes substantial information about the health of a speaker, and thus it has been proposed for the detection of numerous diseases, achieving promising results. However, questions remain about what the models trained for the automatic detection of these diseases are actually learning and the basis for their predictions, which can significantly impact patients' lives. This work advocates for an interpretable health model, suitable for detecting several diseases, motivated by the observation that speech-affecting disorders often have overlapping effects on speech signals. A framework is presented that first defines "reference speech" and then leverages this definition for disease detection. Reference speech is characterized through reference intervals, i.e., the typical values of clinically meaningful acoustic and linguistic features derived from a reference population. This novel approach in the field of speech as a biomarker is inspired by the use of reference intervals in clinical laboratory science. Deviations of new speakers from this reference model are quantified and used as input to detect Alzheimer's and Parkinson's disease. The classification strategy explored is based on Neural Additive Models, a type of glass-box neural network, which enables interpretability. The proposed framework for reference speech characterization and disease detection is designed to support the medical community by providing clinically meaningful explanations that can serve as a valuable second opinion.
△ Less
Submitted 16 September, 2024;
originally announced September 2024.
-
NeuroSpex: Neuro-Guided Speaker Extraction with Cross-Modal Attention
Authors:
Dashanka De Silva,
Siqi Cai,
Saurav Pahuja,
Tanja Schultz,
Haizhou Li
Abstract:
In the study of auditory attention, it has been revealed that there exists a robust correlation between attended speech and elicited neural responses, measurable through electroencephalography (EEG). Therefore, it is possible to use the attention information available within EEG signals to guide the extraction of the target speaker in a cocktail party computationally. In this paper, we present a n…
▽ More
In the study of auditory attention, it has been revealed that there exists a robust correlation between attended speech and elicited neural responses, measurable through electroencephalography (EEG). Therefore, it is possible to use the attention information available within EEG signals to guide the extraction of the target speaker in a cocktail party computationally. In this paper, we present a neuro-guided speaker extraction model, i.e. NeuroSpex, using the EEG response of the listener as the sole auxiliary reference cue to extract attended speech from monaural speech mixtures. We propose a novel EEG signal encoder that captures the attention information. Additionally, we propose a cross-attention (CA) mechanism to enhance the speech feature representations, generating a speaker extraction mask. Experimental results on a publicly available dataset demonstrate that our proposed model outperforms two baseline models across various evaluation metrics.
△ Less
Submitted 16 September, 2024; v1 submitted 4 September, 2024;
originally announced September 2024.
-
On the Role of Visual Grounding in VQA
Authors:
Daniel Reich,
Tanja Schultz
Abstract:
Visual Grounding (VG) in VQA refers to a model's proclivity to infer answers based on question-relevant image regions. Conceptually, VG identifies as an axiomatic requirement of the VQA task. In practice, however, DNN-based VQA models are notorious for bypassing VG by way of shortcut (SC) learning without suffering obvious performance losses in standard benchmarks. To uncover the impact of SC lear…
▽ More
Visual Grounding (VG) in VQA refers to a model's proclivity to infer answers based on question-relevant image regions. Conceptually, VG identifies as an axiomatic requirement of the VQA task. In practice, however, DNN-based VQA models are notorious for bypassing VG by way of shortcut (SC) learning without suffering obvious performance losses in standard benchmarks. To uncover the impact of SC learning, Out-of-Distribution (OOD) tests have been proposed that expose a lack of VG with low accuracy. These tests have since been at the center of VG research and served as basis for various investigations into VG's impact on accuracy. However, the role of VG in VQA still remains not fully understood and has not yet been properly formalized.
In this work, we seek to clarify VG's role in VQA by formalizing it on a conceptual level. We propose a novel theoretical framework called "Visually Grounded Reasoning" (VGR) that uses the concepts of VG and Reasoning to describe VQA inference in ideal OOD testing. By consolidating fundamental insights into VG's role in VQA, VGR helps to reveal rampant VG-related SC exploitation in OOD testing, which explains why the relationship between VG and OOD accuracy has been difficult to define. Finally, we propose an approach to create OOD tests that properly emphasize a requirement for VG, and show how to improve performance on them.
△ Less
Submitted 26 June, 2024;
originally announced June 2024.
-
Breaking Resource Barriers in Speech Emotion Recognition via Data Distillation
Authors:
Yi Chang,
Zhao Ren,
Zhonghao Zhao,
Thanh Tam Nguyen,
Kun Qian,
Tanja Schultz,
Björn W. Schuller
Abstract:
Speech emotion recognition (SER) plays a crucial role in human-computer interaction. The emergence of edge devices in the Internet of Things (IoT) presents challenges in constructing intricate deep learning models due to constraints in memory and computational resources. Moreover, emotional speech data often contains private information, raising concerns about privacy leakage during the deployment…
▽ More
Speech emotion recognition (SER) plays a crucial role in human-computer interaction. The emergence of edge devices in the Internet of Things (IoT) presents challenges in constructing intricate deep learning models due to constraints in memory and computational resources. Moreover, emotional speech data often contains private information, raising concerns about privacy leakage during the deployment of SER models. To address these challenges, we propose a data distillation framework to facilitate efficient development of SER models in IoT applications using a synthesised, smaller, and distilled dataset. Our experiments demonstrate that the distilled dataset can be effectively utilised to train SER models with fixed initialisation, achieving performances comparable to those developed using the original full emotional speech dataset.
△ Less
Submitted 29 May, 2025; v1 submitted 21 June, 2024;
originally announced June 2024.
-
Diff-ETS: Learning a Diffusion Probabilistic Model for Electromyography-to-Speech Conversion
Authors:
Zhao Ren,
Kevin Scheck,
Qinhan Hou,
Stefano van Gogh,
Michael Wand,
Tanja Schultz
Abstract:
Electromyography-to-Speech (ETS) conversion has demonstrated its potential for silent speech interfaces by generating audible speech from Electromyography (EMG) signals during silent articulations. ETS models usually consist of an EMG encoder which converts EMG signals to acoustic speech features, and a vocoder which then synthesises the speech signals. Due to an inadequate amount of available dat…
▽ More
Electromyography-to-Speech (ETS) conversion has demonstrated its potential for silent speech interfaces by generating audible speech from Electromyography (EMG) signals during silent articulations. ETS models usually consist of an EMG encoder which converts EMG signals to acoustic speech features, and a vocoder which then synthesises the speech signals. Due to an inadequate amount of available data and noisy signals, the synthesised speech often exhibits a low level of naturalness. In this work, we propose Diff-ETS, an ETS model which uses a score-based diffusion probabilistic model to enhance the naturalness of synthesised speech. The diffusion model is applied to improve the quality of the acoustic features predicted by an EMG encoder. In our experiments, we evaluated fine-tuning the diffusion model on predictions of a pre-trained EMG encoder, and training both models in an end-to-end fashion. We compared Diff-ETS with a baseline ETS model without diffusion using objective metrics and a listening test. The results indicated the proposed Diff-ETS significantly improved speech naturalness over the baseline.
△ Less
Submitted 11 May, 2024;
originally announced May 2024.
-
Twists, Eisenstein series, and Instantons in Local Mirror Symmetry
Authors:
Andreas Malmendier,
Michael T. Schultz
Abstract:
We propose a mechanism for computing the genus zero invariants of local Calabi-Yau fourfolds arising as the total space of the canonical bundle of a rank-1 Fano threefold. Our method relies heavily on modular parameterizations of the associated Landau-Ginzburg model, as well as extension regulator classes and higher normal functions studied by Doran and Kerr in this setting. The generating functio…
▽ More
We propose a mechanism for computing the genus zero invariants of local Calabi-Yau fourfolds arising as the total space of the canonical bundle of a rank-1 Fano threefold. Our method relies heavily on modular parameterizations of the associated Landau-Ginzburg model, as well as extension regulator classes and higher normal functions studied by Doran and Kerr in this setting. The generating function of the local invariants is proposed as a functional inverse of a weight-4 modular form of Eisenstein type that is naturally computed from the mirror data, in analogy with some known results for local threefolds that computes genus zero Gopakumar-Vafa invariants. Using the twist construction of Doran and the first named author, these two constructions are connected via Doran's generalized functional invariant map on the periods.
△ Less
Submitted 20 August, 2026; v1 submitted 12 March, 2024;
originally announced March 2024.
-
Correlated Rotational Alignment Spectroscopy: A New Tool for High-Resolution Spectroscopy and the Analysis of Heterogeneous Samples
Authors:
Thomas Schultz
Abstract:
Correlated rotational alignment spectroscopy correlates observables of ultrafast gas-phase spectroscopy with high-resolution, broad-band rotational Raman spectra. This article reviews the measurement principle of CRASY, existing implementations for mass-correlated measurements, and the potential for future developments. New spectroscopic capabilities are discussed in detail: Signals for individual…
▽ More
Correlated rotational alignment spectroscopy correlates observables of ultrafast gas-phase spectroscopy with high-resolution, broad-band rotational Raman spectra. This article reviews the measurement principle of CRASY, existing implementations for mass-correlated measurements, and the potential for future developments. New spectroscopic capabilities are discussed in detail: Signals for individual sample components can be separated even in highly heterogeneous samples. Isotopologue rotational spectra can be observed at natural isotope abundance. Fragmentation channels are readily assigned in molecular and cluster mass spectra. And finally, rotational Raman spectra can be measured with sub-MHz resolution, an improvement of several orders-of-magnitude as compared to preceding experiments.
△ Less
Submitted 6 March, 2024;
originally announced March 2024.
-
Is Open-Source There Yet? A Comparative Study on Commercial and Open-Source LLMs in Their Ability to Label Chest X-Ray Reports
Authors:
Felix J. Dorfner,
Liv Jürgensen,
Leonhard Donle,
Fares Al Mohamad,
Tobias R. Bodenmann,
Mason C. Cleveland,
Felix Busch,
Lisa C. Adams,
James Sato,
Thomas Schultz,
Albert E. Kim,
Jameson Merkow,
Keno K. Bressem,
Christopher P. Bridge
Abstract:
Introduction: With the rapid advances in large language models (LLMs), there have been numerous new open source as well as commercial models. While recent publications have explored GPT-4 in its application to extracting information of interest from radiology reports, there has not been a real-world comparison of GPT-4 to different leading open-source models.
Materials and Methods: Two different…
▽ More
Introduction: With the rapid advances in large language models (LLMs), there have been numerous new open source as well as commercial models. While recent publications have explored GPT-4 in its application to extracting information of interest from radiology reports, there has not been a real-world comparison of GPT-4 to different leading open-source models.
Materials and Methods: Two different and independent datasets were used. The first dataset consists of 540 chest x-ray reports that were created at the Massachusetts General Hospital between July 2019 and July 2021. The second dataset consists of 500 chest x-ray reports from the ImaGenome dataset. We then compared the commercial models GPT-3.5 Turbo and GPT-4 from OpenAI to the open-source models Mistral-7B, Mixtral-8x7B, Llama2-13B, Llama2-70B, QWEN1.5-72B and CheXbert and CheXpert-labeler in their ability to accurately label the presence of multiple findings in x-ray text reports using different prompting techniques.
Results: On the ImaGenome dataset, the best performing open-source model was Llama2-70B with micro F1-scores of 0.972 and 0.970 for zero- and few-shot prompts, respectively. GPT-4 achieved micro F1-scores of 0.975 and 0.984, respectively. On the institutional dataset, the best performing open-source model was QWEN1.5-72B with micro F1-scores of 0.952 and 0.965 for zero- and few-shot prompting, respectively. GPT-4 achieved micro F1-scores of 0.975 and 0.973, respectively.
Conclusion: In this paper, we show that while GPT-4 is superior to open-source models in zero-shot report labeling, the implementation of few-shot prompting can bring open-source models on par with GPT-4. This shows that open-source models could be a performant and privacy preserving alternative to GPT-4 for the task of radiology report classification.
△ Less
Submitted 19 February, 2024;
originally announced February 2024.
-
Correlating parent-fragment relationships in cluster photoionization
Authors:
Jong Chan Lee,
Begüm Rukiye Özer,
In Heo,
Thomas Schultz
Abstract:
Fragment signals in ordinary mass spectra carry no label to identify their parent molecule. By correlating mass signals with rotational Raman spectra, we created a method to label each ion signal with the spectroscopic fingerprint of its neutral parent molecule. In data for a carbon disulfide molecular cluster beam, we assigned 28 distinct ionization and fragmentation channels based on their mass-…
▽ More
Fragment signals in ordinary mass spectra carry no label to identify their parent molecule. By correlating mass signals with rotational Raman spectra, we created a method to label each ion signal with the spectroscopic fingerprint of its neutral parent molecule. In data for a carbon disulfide molecular cluster beam, we assigned 28 distinct ionization and fragmentation channels based on their mass-correlated rotational fingerprints. Unexpected observations included the formation of energetic S2 and SCCS cationic fragments from the CS2-dimer cluster and a significant CS3 signal, uncorrelated to the dimer. The large number of observed channels revealed a surprising complexity that could only be addressed with correlated spectroscopy and computer-aided correlation analysis.
△ Less
Submitted 13 February, 2024;
originally announced February 2024.
-
STAA-Net: A Sparse and Transferable Adversarial Attack for Speech Emotion Recognition
Authors:
Yi Chang,
Zhao Ren,
Zixing Zhang,
Xin Jing,
Kun Qian,
Xi Shao,
Bin Hu,
Tanja Schultz,
Björn W. Schuller
Abstract:
Speech contains rich information on the emotions of humans, and Speech Emotion Recognition (SER) has been an important topic in the area of human-computer interaction. The robustness of SER models is crucial, particularly in privacy-sensitive and reliability-demanding domains like private healthcare. Recently, the vulnerability of deep neural networks in the audio domain to adversarial attacks has…
▽ More
Speech contains rich information on the emotions of humans, and Speech Emotion Recognition (SER) has been an important topic in the area of human-computer interaction. The robustness of SER models is crucial, particularly in privacy-sensitive and reliability-demanding domains like private healthcare. Recently, the vulnerability of deep neural networks in the audio domain to adversarial attacks has become a popular area of research. However, prior works on adversarial attacks in the audio domain primarily rely on iterative gradient-based techniques, which are time-consuming and prone to overfitting the specific threat model. Furthermore, the exploration of sparse perturbations, which have the potential for better stealthiness, remains limited in the audio domain. To address these challenges, we propose a generator-based attack method to generate sparse and transferable adversarial examples to deceive SER models in an end-to-end and efficient manner. We evaluate our method on two widely-used SER datasets, Database of Elicited Mood in Speech (DEMoS) and Interactive Emotional dyadic MOtion CAPture (IEMOCAP), and demonstrate its ability to generate successful sparse adversarial examples in an efficient manner. Moreover, our generated adversarial examples exhibit model-agnostic transferability, enabling effective adversarial attacks on advanced victim models.
△ Less
Submitted 2 February, 2024;
originally announced February 2024.
-
On holomorphic conformal structures associated with lattice polarized K3 surfaces
Authors:
Andreas Malmendier,
Michael T. Schultz
Abstract:
We discuss the connection between Picard-Fuchs equations for certain families of lattice polarized K3 surfaces and the construction of integrable holomorphic conformal structures on their period domains. We then compute an explicit example of a locally conformally flat holomorphic metric associated with generic Jacobian Kummer surfaces, which allows for a novel description of the local variation o…
▽ More
We discuss the connection between Picard-Fuchs equations for certain families of lattice polarized K3 surfaces and the construction of integrable holomorphic conformal structures on their period domains. We then compute an explicit example of a locally conformally flat holomorphic metric associated with generic Jacobian Kummer surfaces, which allows for a novel description of the local variation of complex structure.
△ Less
Submitted 18 January, 2024;
originally announced January 2024.
-
Uncovering the Full Potential of Visual Grounding Methods in VQA
Authors:
Daniel Reich,
Tanja Schultz
Abstract:
Visual Grounding (VG) methods in Visual Question Answering (VQA) attempt to improve VQA performance by strengthening a model's reliance on question-relevant visual information. The presence of such relevant information in the visual input is typically assumed in training and testing. This assumption, however, is inherently flawed when dealing with imperfect image representations common in large-sc…
▽ More
Visual Grounding (VG) methods in Visual Question Answering (VQA) attempt to improve VQA performance by strengthening a model's reliance on question-relevant visual information. The presence of such relevant information in the visual input is typically assumed in training and testing. This assumption, however, is inherently flawed when dealing with imperfect image representations common in large-scale VQA, where the information carried by visual features frequently deviates from expected ground-truth contents. As a result, training and testing of VG-methods is performed with largely inaccurate data, which obstructs proper assessment of their potential benefits. In this study, we demonstrate that current evaluation schemes for VG-methods are problematic due to the flawed assumption of availability of relevant visual information. Our experiments show that these methods can be much more effective when evaluation conditions are corrected. Code is provided on GitHub.
△ Less
Submitted 15 February, 2024; v1 submitted 15 January, 2024;
originally announced January 2024.
-
Surface doping of rubrene single crystals by molecular electron donors and acceptors
Authors:
Christos Gatsios,
Andreas Opitz,
Dominique Lungwitz,
Ahmed E. Mansour,
Thorsten Schultz,
Dongguen Shin,
Sebastian Hammer,
Jens Pflaum,
Yadong Zhang,
Stephen Barlow,
Seth R. Marder,
Norbert Koch
Abstract:
The surface molecular doping of organic semiconductors can play an important role in the development of organic electronic or optoelectronic devices. Single-crystal rubrene remains a leading molecular candidate for applications in electronics due to its high hole mobility. In parallel, intensive research into the fabrication of flexible organic electronics requires the careful design of functional…
▽ More
The surface molecular doping of organic semiconductors can play an important role in the development of organic electronic or optoelectronic devices. Single-crystal rubrene remains a leading molecular candidate for applications in electronics due to its high hole mobility. In parallel, intensive research into the fabrication of flexible organic electronics requires the careful design of functional interfaces to enable optimal device characteristics. To this end, the present work seeks to understand the effect of surface molecular doping on the electronic band structure of rubrene single crystals. Our angle-resolved photoemission measurements reveal that the Fermi level moves in the band gap of rubrene depending on the direction of surface electron-transfer reactions with the molecular dopants, yet the valence band dispersion remains essentially unperturbed. This indicates that surface electron-transfer doping of a molecular single crystal can effectively modify the near-surface charge density, while retaining good charge-carrier mobility.
△ Less
Submitted 11 January, 2024;
originally announced January 2024.
-
Topological atom optics and beyond with knotted quantum wavefunctions
Authors:
Maitreyi Jayaseelan,
Joseph D. Murphree,
Justin T. Schultz,
Janne Ruostekoski,
Nicholas P. Bigelow
Abstract:
Atom optics demonstrates optical phenomena with coherent matter waves, providing a foundational connection between light and matter. Significant advances in optics have followed the realisation of structured light fields hosting complex singularities and topologically non-trivial characteristics. However, analogous studies are still in their infancy in the field of atom optics. Here, we investigat…
▽ More
Atom optics demonstrates optical phenomena with coherent matter waves, providing a foundational connection between light and matter. Significant advances in optics have followed the realisation of structured light fields hosting complex singularities and topologically non-trivial characteristics. However, analogous studies are still in their infancy in the field of atom optics. Here, we investigate and experimentally create knotted quantum wavefunctions in spinor Bose--Einstein condensates which display non-trivial topologies. In our work we construct coordinated orbital and spin rotations of the atomic wavefunction, engineering a variety of discrete symmetries in the combined spin and orbital degrees of freedom. The structured wavefunctions that we create map to the surface of a torus to form torus knots, Möbius strips, and a twice-linked Solomon's knot. In this paper we demonstrate striking connections between the symmetries and underlying topologies of multicomponent atomic systems and of vector optical fields--a realization of topological atom-optics.
△ Less
Submitted 15 December, 2023;
originally announced December 2023.
-
NeuroHeed: Neuro-Steered Speaker Extraction using EEG Signals
Authors:
Zexu Pan,
Marvin Borsdorf,
Siqi Cai,
Tanja Schultz,
Haizhou Li
Abstract:
Humans possess the remarkable ability to selectively attend to a single speaker amidst competing voices and background noise, known as selective auditory attention. Recent studies in auditory neuroscience indicate a strong correlation between the attended speech signal and the corresponding brain's elicited neuronal activities, which the latter can be measured using affordable and non-intrusive el…
▽ More
Humans possess the remarkable ability to selectively attend to a single speaker amidst competing voices and background noise, known as selective auditory attention. Recent studies in auditory neuroscience indicate a strong correlation between the attended speech signal and the corresponding brain's elicited neuronal activities, which the latter can be measured using affordable and non-intrusive electroencephalography (EEG) devices. In this study, we present NeuroHeed, a speaker extraction model that leverages EEG signals to establish a neuronal attractor which is temporally associated with the speech stimulus, facilitating the extraction of the attended speech signal in a cocktail party scenario. We propose both an offline and an online NeuroHeed, with the latter designed for real-time inference. In the online NeuroHeed, we additionally propose an autoregressive speaker encoder, which accumulates past extracted speech signals for self-enrollment of the attended speaker information into an auditory attractor, that retains the attentional momentum over time. Online NeuroHeed extracts the current window of the speech signals with guidance from both attractors. Experimental results demonstrate that NeuroHeed effectively extracts brain-attended speech signals, achieving high signal quality, excellent perceptual quality, and intelligibility in a two-speaker scenario.
△ Less
Submitted 26 July, 2023;
originally announced July 2023.
-
Anisotropic Fanning Aware Low-Rank Tensor Approximation Based Tractography
Authors:
Johannes Grün,
Jonah Sieg,
Thomas Schultz
Abstract:
Low-rank higher-order tensor approximation has been used successfully to extract discrete directions for tractography from continuous fiber orientation density functions (fODFs). However, while it accounts for fiber crossings, it has so far ignored fanning, which has led to incomplete reconstructions. In this work, we integrate an anisotropic model of fanning based on the Bingham distribution into…
▽ More
Low-rank higher-order tensor approximation has been used successfully to extract discrete directions for tractography from continuous fiber orientation density functions (fODFs). However, while it accounts for fiber crossings, it has so far ignored fanning, which has led to incomplete reconstructions. In this work, we integrate an anisotropic model of fanning based on the Bingham distribution into a recently proposed tractography method that performs low-rank approximation with an Unscented Kalman Filter. Our technical contributions include an initialization scheme for the new parameters, which is based on the Hessian of the low-rank approximation, pre-integration of the required convolution integrals to reduce the computational effort, and representation of the required 3D rotations with quaternions. Results on 12 subjects from the Human Connectome Project confirm that, in almost all considered tracts, our extended model significantly increases completeness of the reconstruction, while reducing excess, at acceptable additional computational cost. Its results are also more accurate than those from a simpler, isotropic fanning model that is based on Watson distributions.
△ Less
Submitted 3 July, 2023;
originally announced July 2023.
-
Measuring Faithful and Plausible Visual Grounding in VQA
Authors:
Daniel Reich,
Felix Putze,
Tanja Schultz
Abstract:
Metrics for Visual Grounding (VG) in Visual Question Answering (VQA) systems primarily aim to measure a system's reliance on relevant parts of the image when inferring an answer to the given question. Lack of VG has been a common problem among state-of-the-art VQA systems and can manifest in over-reliance on irrelevant image parts or a disregard for the visual modality entirely. Although inference…
▽ More
Metrics for Visual Grounding (VG) in Visual Question Answering (VQA) systems primarily aim to measure a system's reliance on relevant parts of the image when inferring an answer to the given question. Lack of VG has been a common problem among state-of-the-art VQA systems and can manifest in over-reliance on irrelevant image parts or a disregard for the visual modality entirely. Although inference capabilities of VQA models are often illustrated by a few qualitative illustrations, most systems are not quantitatively assessed for their VG properties. We believe, an easily calculated criterion for meaningfully measuring a system's VG can help remedy this shortcoming, as well as add another valuable dimension to model evaluations and analysis. To this end, we propose a new VG metric that captures if a model a) identifies question-relevant objects in the scene, and b) actually relies on the information contained in the relevant objects when producing its answer, i.e., if its visual grounding is both "faithful" and "plausible". Our metric, called "Faithful and Plausible Visual Grounding" (FPVG), is straightforward to determine for most VQA model designs.
We give a detailed description of FPVG and evaluate several reference systems spanning various VQA architectures. Code to support the metric calculations on the GQA data set is available on GitHub.
△ Less
Submitted 14 October, 2023; v1 submitted 24 May, 2023;
originally announced May 2023.
-
Visually Grounded VQA by Lattice-based Retrieval
Authors:
Daniel Reich,
Felix Putze,
Tanja Schultz
Abstract:
Visual Grounding (VG) in Visual Question Answering (VQA) systems describes how well a system manages to tie a question and its answer to relevant image regions. Systems with strong VG are considered intuitively interpretable and suggest an improved scene understanding. While VQA accuracy performances have seen impressive gains over the past few years, explicit improvements to VG performance and ev…
▽ More
Visual Grounding (VG) in Visual Question Answering (VQA) systems describes how well a system manages to tie a question and its answer to relevant image regions. Systems with strong VG are considered intuitively interpretable and suggest an improved scene understanding. While VQA accuracy performances have seen impressive gains over the past few years, explicit improvements to VG performance and evaluation thereof have often taken a back seat on the road to overall accuracy improvements. A cause of this originates in the predominant choice of learning paradigm for VQA systems, which consists of training a discriminative classifier over a predetermined set of answer options.
In this work, we break with the dominant VQA modeling paradigm of classification and investigate VQA from the standpoint of an information retrieval task. As such, the developed system directly ties VG into its core search procedure. Our system operates over a weighted, directed, acyclic graph, a.k.a. "lattice", which is derived from the scene graph of a given image in conjunction with region-referring expressions extracted from the question.
We give a detailed analysis of our approach and discuss its distinctive properties and limitations. Our approach achieves the strongest VG performance among examined systems and exhibits exceptional generalization capabilities in a number of scenarios.
△ Less
Submitted 15 November, 2022;
originally announced November 2022.
-
Molecular-Beam Spectroscopy with an Infinite Interferometer: Spectroscopic Resolution and Accuracy
Authors:
Thomas Schultz,
In Heo,
Jong Chan Lee,
Begüm Rukiye Özer
Abstract:
An interferometer with effectively infinite maximum optical path difference removes the dominant resolution limitation for interferometric spectroscopy. We present mass-correlated rotational Raman spectra that represent the world's highest resolution scanned interferometric data and discuss the current and expected future limitations in achievable spectroscopic performance.
An interferometer with effectively infinite maximum optical path difference removes the dominant resolution limitation for interferometric spectroscopy. We present mass-correlated rotational Raman spectra that represent the world's highest resolution scanned interferometric data and discuss the current and expected future limitations in achievable spectroscopic performance.
△ Less
Submitted 22 September, 2022;
originally announced September 2022.
-
Absolutely continuous and BV-curves in 1-Wasserstein spaces
Authors:
Ehsan Abedi,
Zhenhao Li,
Timo Schultz
Abstract:
We extend the result of Lisini (Calc Var Partial Differ Equ 28:85-120, 2007) on the superposition principle for absolutely continuous curves in $p$-Wasserstein spaces to the special case of $p=1$. In contrast to the case of $p>1$, it is not always possible to have lifts on absolutely continuous curves. Therefore, one needs to relax the notion of a lift by considering curves of bounded variation, o…
▽ More
We extend the result of Lisini (Calc Var Partial Differ Equ 28:85-120, 2007) on the superposition principle for absolutely continuous curves in $p$-Wasserstein spaces to the special case of $p=1$. In contrast to the case of $p>1$, it is not always possible to have lifts on absolutely continuous curves. Therefore, one needs to relax the notion of a lift by considering curves of bounded variation, or shortly BV-curves, and replace the metric speed by the total variation measure. We prove that any BV-curve in a 1-Wasserstein space can be represented by a probability measure on the space of BV-curves which encodes the total variation measure of the Wasserstein curve. In particular, when the curve is absolutely continuous, the result gives a lift concentrated on BV-curves which also characterizes the metric speed. The main theorem is then applied for the characterization of geodesics and the study of the continuity equation in a discrete setting.
△ Less
Submitted 16 December, 2023; v1 submitted 9 September, 2022;
originally announced September 2022.
-
Combining Image Space and q-Space PDEs for Lossless Compression of Diffusion MR Images
Authors:
Ikram Jumakulyyev,
Thomas Schultz
Abstract:
Diffusion MRI is a modern neuroimaging modality with a unique ability to acquire microstructural information by measuring water self-diffusion at the voxel level. However, it generates huge amounts of data, resulting from a large number of repeated 3D scans. Each volume samples a location in q-space, indicating the direction and strength of a diffusion sensitizing gradient during the measurement.…
▽ More
Diffusion MRI is a modern neuroimaging modality with a unique ability to acquire microstructural information by measuring water self-diffusion at the voxel level. However, it generates huge amounts of data, resulting from a large number of repeated 3D scans. Each volume samples a location in q-space, indicating the direction and strength of a diffusion sensitizing gradient during the measurement. This captures detailed information about the self-diffusion, and the tissue microstructure that restricts it. Lossless compression with GZIP is widely used to reduce the memory requirements. We introduce a novel lossless codec for diffusion MRI data. It reduces file sizes by more than 30% compared to GZIP, and also beats lossless codecs from the JPEG family. Our codec builds on recent work on lossless PDE-based compression of 3D medical images, but additionally exploits smoothness in q-space. We demonstrate that, compared to using only image space PDEs, q-space PDEs further improve compression rates. Moreover, implementing them with Finite Element Methods and a custom acceleration significantly reduces computational expense. Finally, we show that our codec clearly benefits from integrating subject motion correction, and slightly from optimizing the order in which the 3D volumes are coded.
△ Less
Submitted 5 April, 2023; v1 submitted 14 June, 2022;
originally announced June 2022.
-
A Primer for Telemetry Interfacing in Accordance with NASA Standards Using Low Cost FPGAs
Authors:
Jake A. McCoy,
Ted B. Schultz,
James H. Tutt,
Thomas Rogers,
Drew M. Miles,
Randall L. McEntaffer
Abstract:
Photon counting detector systems on sounding rocket payloads often require interfacing asynchronous outputs with a synchronously clocked telemetry (TM) stream. Though this can be handled with an on-board computer, there are several low cost alternatives including custom hardware, microcontrollers and field-programmable gate arrays (FPGAs). This paper outlines how a TM interface (TMIF) for detector…
▽ More
Photon counting detector systems on sounding rocket payloads often require interfacing asynchronous outputs with a synchronously clocked telemetry (TM) stream. Though this can be handled with an on-board computer, there are several low cost alternatives including custom hardware, microcontrollers and field-programmable gate arrays (FPGAs). This paper outlines how a TM interface (TMIF) for detectors on a sounding rocket with asynchronous parallel digital output can be implemented using low cost FPGAs and minimal custom hardware. Low power consumption and high speed FPGAs are available as commercial off-the-shelf (COTS) products and can be used to develop the main component of the TMIF. Then, only a small amount of additional hardware is required for signal buffering and level translating. This paper also discusses how this system can be tested with a simulated TM chain in the small laboratory setting using FPGAs and COTS specialized data acquisition products.
△ Less
Submitted 22 March, 2022;
originally announced March 2022.
-
On master test plans for the space of BV functions
Authors:
Francesco Nobili,
Enrico Pasqualetto,
Timo Schultz
Abstract:
We prove that on an arbitrary metric measure space a countable collection of test plans is sufficient to recover all $\rm BV$ functions and their total variation measures. In the setting of non-branching ${\sf CD}(K,N)$ spaces (with finite reference measure), we can additionally require these test plans to be concentrated on geodesics.
We prove that on an arbitrary metric measure space a countable collection of test plans is sufficient to recover all $\rm BV$ functions and their total variation measures. In the setting of non-branching ${\sf CD}(K,N)$ spaces (with finite reference measure), we can additionally require these test plans to be concentrated on geodesics.
△ Less
Submitted 10 September, 2021;
originally announced September 2021.
-
On the mixed-twist construction and monodromy of associated Picard-Fuchs systems
Authors:
Andreas Malmendier,
Michael T. Schultz
Abstract:
We use the mixed-twist construction of Doran and Malmendier to obtain a multi-parameter family of K3 surfaces of Picard rank $ρ\ge 16$. Upon identifying a particular Jacobian elliptic fibration on its general member, we determine the lattice polarization and the Picard-Fuchs system for the family. We construct a sequence of restrictions that lead to extensions of the polarization by two-elementary…
▽ More
We use the mixed-twist construction of Doran and Malmendier to obtain a multi-parameter family of K3 surfaces of Picard rank $ρ\ge 16$. Upon identifying a particular Jacobian elliptic fibration on its general member, we determine the lattice polarization and the Picard-Fuchs system for the family. We construct a sequence of restrictions that lead to extensions of the polarization by two-elementary lattices. We show that the Picard-Fuchs operators for the restricted families coincide with known resonant hypergeometric systems. Second, for the one-parameter mirror families of deformed Fermat hypersurfaces we show that the mixed-twist construction produces a non-resonant GKZ system for which a basis of solutions in the form of absolutely convergent Mellin-Barnes integrals exists whose monodromy we compute explicitly.
△ Less
Submitted 4 May, 2022; v1 submitted 15 August, 2021;
originally announced August 2021.