Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 152 results for author: Chung, J S

.
  1. arXiv:2608.09119  [pdf, ps, other

    cs.AI

    Motif 3: Technical Report

    Authors: Junghwan Lim, Joon Son Chung, Sungmin Lee, Wai Ting Cheung, Gihun Cho, Minsu Ha, Sangho Kang, Beomgyu Kim, Dongseok Kim, Jangwoong Kim, Taehyun Kim, Taewhan Kim, Jeesoo Lee, Jeongdoo Lee, Junhyeok Lee, Dongpin Oh, Hyeyeon Cho, Dahye Choi, Jaeheui Her, Hanbin Jung, Changjin Kang, Minjae Kim, Youngrok Kim, Hyukjin Kweon, Hongjoo Lee , et al. (2 additional authors not shown)

    Abstract: We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 billion activated per token. Each sparse MoE layer contains 384 routed experts, with eight selected per token. This fine-grained sparsity provides substantial expert capacity while limiting computation. Motif 3 is built around Grouped Differential Latent Attention (GDLA), which integra… ▽ More

    Submitted 10 August, 2026; originally announced August 2026.

  2. arXiv:2607.21179  [pdf, ps, other

    cs.CV

    Out of Sight, Still in Mind: Token Compression for Omni-LLMs

    Authors: Suho Yoo, Youngjoon Jang, Hyebin Cho, Joon Son Chung

    Abstract: The goal of this paper is to reduce the input token cost of Omni-modal large language models (Omni-LLMs) at inference time. Omni-LLMs reason jointly over audio, video and text, but the cost of the three streams is highly unbalanced: visual tokens account for the vast majority of the input, and are highly redundant. In this paper, we propose ReMo, a training-free framework that compresses visual to… ▽ More

    Submitted 23 July, 2026; originally announced July 2026.

    Comments: Preprint

  3. arXiv:2607.08039  [pdf, ps, other

    nucl-ex hep-ex

    A study of neutrinoless double electron capture in $^{40}$Ca from the AMoRE experiment

    Authors: AMoRE Collaboration, A. Agrawal, V. V. Alenkov, P. Aryal, J. Beyer, B. Bhandari, R. S. Boiko, K. Boonin, O. Buzanov, C. R. Byeon, N. Chanthima, M. K. Cheoun, J. S. Choe, Seonho Choi, S. Choudhury, J. S. Chung, F. A. Danevich, M. Djamal, D. Drung, C. Enss, A. Fleischmann, A. M. Gangapshev, L. Gastaldo, Y. M. Gavrilyuk, A. M. Gezhaev , et al. (85 additional authors not shown)

    Abstract: The search for neutrinoless double electron capture ($0ν\mathrm{2EC}$) provides a sensitive probe of lepton-number violation and the Majorana nature of neutrinos. We investigate the $0ν\mathrm{2EC}$ decay of $^{40}$Ca using cryogenic detectors equipped with metallic magnetic calorimeters in the AMoRE-I experiment. The analysis is based on a physics dataset corresponding to a total exposure of 7.32… ▽ More

    Submitted 8 July, 2026; originally announced July 2026.

    Comments: 7 pages, 3 figures

  4. arXiv:2606.27307  [pdf, ps, other

    cs.CV

    See & Sniff: Learning Visuo-Olfactory Representations

    Authors: Seongyu Kim, Seungwoo Lee, Hyeonggon Ryu, Joon Son Chung, Arda Senocak

    Abstract: While modern multimodal models integrate vision with language, audio, or touch, olfaction remains largely unexplored due to the lack of paired visuo-olfactory data. We introduce SmellNet-V, a scalable visuo-olfactory dataset built on the insight that odor identity is largely invariant to visual transformations within a semantic category. This allows us to synthetically pair smell-only samples with… ▽ More

    Submitted 25 June, 2026; originally announced June 2026.

    Comments: ECCV 2026. Project Page: https://mm.kaist.ac.kr/projects/SeeandSniff/

  5. arXiv:2606.21888  [pdf, ps, other

    eess.AS cs.SD

    ProsoCodec: Prosody-Oriented Speech Codec for Voice Conversion

    Authors: Jeongsoo Choi, Ji-Hoon Kim, Shujie Hu, Joon Son Chung

    Abstract: Neural speech codecs efficiently compress speech and have become a foundation for speech generation, but they are typically learned as holistic representations that intertwine linguistic content, speaker identity, and prosody. While this design is effective for zero-shot voice cloning, it hinders downstream tasks that require prosody preservation or transfer, such as voice conversion. To address t… ▽ More

    Submitted 20 June, 2026; originally announced June 2026.

    Comments: Interspeech 2026

  6. arXiv:2606.18924  [pdf, ps, other

    cs.SD

    Who Wins the Conflict? Mechanistic Interpretability of Text Bias in Audio LLMs

    Authors: Hyebin Cho, Suho Yoo, Jaehyuk Jang, Changick Kim, Joon Son Chung

    Abstract: While Audio Large Language Models (Audio LLMs) excel at multimodal understanding, they suffer from text dominance, a bias where models blindly favor text over acoustic evidence, causing hallucinations. However, the internal mechanisms underlying how these models behave when audio and textual inputs contradict each other remain unexplored. In this work, we present the first mechanistic analysis of… ▽ More

    Submitted 17 June, 2026; originally announced June 2026.

    Comments: Preprint

  7. arXiv:2606.15751  [pdf, ps, other

    cs.SD cs.LG cs.MM eess.AS

    Acoustic Prompting via Stage-wise Modulation for Few-Shot Learning in Audio Language Models

    Authors: Hyebin Cho, Jaehyuk Jang, Changick Kim, Joon Son Chung

    Abstract: Audio-Language Models (ALMs) have shown remarkable success in zero-shot audio classification by aligning audio waveforms with text. Recent efforts to improve downstream performance focus on learning optimal text prompts. However, previous approaches focus on the text encoder, leaving the potential of learnable prompts within the audio encoder unexplored. In this paper, we propose a novel framework… ▽ More

    Submitted 14 June, 2026; originally announced June 2026.

    Comments: Accepted to INTERSPEECH 2026

  8. arXiv:2606.08935  [pdf, ps, other

    cs.LG cs.AI

    PAI: Preserving Amplitude Information in Representation-Based Time-Series Anomaly Detection

    Authors: Kang Zhang, Wei Jian Lau, Shoushou Ren, Dong Lin, Joon Son Chung, Chuanhao Sun

    Abstract: Representation-based time-series anomaly detection algorithms significantly outperform other methods on diverse anomaly detection tasks. However, we notice that they suffer from a major limitation in our evaluation - their learned embeddings are often amplitude-agnostic. Losing amplitude information can degrade performance on amplitude related anomalies, and this failure is prevalent across all ex… ▽ More

    Submitted 7 June, 2026; originally announced June 2026.

    Comments: 15 pages

  9. arXiv:2606.03183  [pdf, ps, other

    cs.MM cs.CV cs.SD eess.AS

    Inference-Time Scaling for Joint Audio-Video Generation

    Authors: Jaemin Jung, Kyeongha Rho, Inkyu Shin, Joon Son Chung

    Abstract: Joint audio-video generation aims to synthesize realistic audio-video pairs that are both semantically aligned with text prompts and precisely synchronized. While existing joint audio-video generation models often require substantial training resources to improve fidelity, Inference-Time Scaling (ITS) has recently emerged as a promising training-free alternative in single-modality domains. However… ▽ More

    Submitted 2 June, 2026; originally announced June 2026.

    Comments: Accepted by Transactions on Machine Learning Research (TMLR). Project page: https://jung-jaemin.github.io/ITS-AVGen-Proj/

  10. arXiv:2605.15044  [pdf, ps, other

    cs.SD cs.AI cs.LG cs.MM eess.AS

    SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning

    Authors: KiHyun Nam, Jungwoo Heo, Siu Bae, Ha-Jin Yu, Joon Son Chung

    Abstract: As audio-first agents become increasingly common in physical AI, conversational robots, and screenless wearables, audio large language models (audio-LLMs) must integrate speaker-specific understanding to support user authorization, personalization, and context-aware interaction. This requires modeling who is speaking, how the voice sounds, and how recording conditions affect speaker cues. Conventi… ▽ More

    Submitted 14 May, 2026; originally announced May 2026.

  11. arXiv:2605.13071  [pdf, ps, other

    cs.NE

    FiTS: Interpretable Spiking Neurons via Frequency Selectivity and Temporal Shaping

    Authors: Jongmin Choi, Joon Son Chung

    Abstract: Spiking Neural Networks (SNNs) are a promising framework for event-driven temporal processing. Prior work has improved temporal modeling through richer neuron dynamics and network-level mechanisms such as recurrence and delays, but it remains unclear how individual spiking neurons should specialize within a network. In this work, we introduce FiTS, a spiking neuron that factorizes temporal computa… ▽ More

    Submitted 13 May, 2026; originally announced May 2026.

    Comments: 23 pages, 7 figures

  12. arXiv:2605.11605  [pdf, ps, other

    cs.CV cs.AI

    Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs

    Authors: Chaeyoung Jung, Kyeongha Rho, Joon Son Chung

    Abstract: Omnimodal Large Language Models (Omni-LLMs) incur substantial computational overhead due to the large number of multimodal input tokens they process, making token reduction essential for real-world deployment. Existing Omni-LLM pruning methods typically reduce this cost by selecting tokens that are important for the current query or strongly aligned with cross-modal cues. However, such strategies… ▽ More

    Submitted 12 May, 2026; originally announced May 2026.

  13. arXiv:2605.10815  [pdf, ps, other

    cs.AI eess.AS

    Probing Cross-modal Information Hubs in Audio-Visual LLMs

    Authors: Jihoo Jung, Chaeyoung Jung, Ji-Hoon Kim, Joon Son Chung

    Abstract: Audio-visual large language models (AVLLMs) have recently emerged as a powerful architecture capable of jointly reasoning over audio, visual, and textual modalities. In AVLLMs, the bidirectional interaction between audio and video modalities introduces intricate processing dynamics, necessitating a deeper understanding of their internal mechanisms. However, unlike extensively studied text-only or… ▽ More

    Submitted 11 May, 2026; v1 submitted 11 May, 2026; originally announced May 2026.

    Comments: Accepted by ICML 2026

  14. arXiv:2604.27866  [pdf, ps, other

    eess.AS cs.MM cs.SD

    LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition

    Authors: Doyeop Kwak, Jeongsoo Choi, Suyeon Lee, Joon Son Chung

    Abstract: We introduce LRS-VoxMM, an in-the-wild benchmark for audio-visual speech recognition (AVSR). The benchmark is derived from VoxMM, a dataset of diverse real-world spoken conversations with human-annotated transcriptions. We select AVSR-suitable samples and preprocess them in an LRS-style format for direct use in existing AVSR pipelines. Compared with commonly used benchmarks, LRS-VoxMM covers a mor… ▽ More

    Submitted 30 April, 2026; originally announced April 2026.

    Comments: Technical report for the LRS-VoxMM dataset release. Project page: https://mm.kaist.ac.kr/projects/voxmm

  15. arXiv:2604.11579  [pdf, ps, other

    cs.CV

    Seeing Through Touch: Tactile-Driven Visual Localization of Material Regions

    Authors: Seongyu Kim, Seungwoo Lee, Hyeonggon Ryu, Joon Son Chung, Arda Senocak

    Abstract: We address the problem of tactile localization, where the goal is to identify image regions that share the same material properties as a tactile input. Existing visuo-tactile methods rely on global alignment and thus fail to capture the fine-grained local correspondences required for this task. The challenge is amplified by existing datasets, which predominantly contain close-up, low-diversity ima… ▽ More

    Submitted 13 April, 2026; originally announced April 2026.

    Comments: CVPR 2026. Project page: https://mm.kaist.ac.kr/projects/SeeingThroughTouch/

  16. arXiv:2603.26113  [pdf, ps, other

    cs.MM cs.SD eess.AS

    Cinematic Audio Source Separation Using Visual Cues

    Authors: Kang Zhang, Suyeon Lee, Arda Senocak, Joon Son Chung

    Abstract: Cinematic Audio Source Separation (CASS) aims to decompose mixed film audio into speech, music, and sound effects, enabling applications like dubbing and remastering. Existing CASS approaches are audio-only, overlooking the inherent audio-visual nature of films, where sounds often align with visual cues. We present the first framework for audio-visual CASS (AV-CASS), leveraging visual context to e… ▽ More

    Submitted 27 March, 2026; originally announced March 2026.

    Comments: CVPR 2026. Project page: https://cass-flowmatching.github.io

  17. arXiv:2603.19697  [pdf, ps, other

    eess.AS cs.MM cs.SD

    Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction

    Authors: Doyeop Kwak, Suyeon Lee, Joon Son Chung

    Abstract: The goal of this paper is to provide a new perspective on audio-visual target speaker extraction (AV-TSE) by decoupling separation and target selection. Conventional AV-TSE systems typically integrate audio and visual features deeply to re-learn the entire separation process, which can act as a fidelity ceiling due to the noisy nature of in-the-wild audio-visual datasets. To address this, we propo… ▽ More

    Submitted 16 June, 2026; v1 submitted 20 March, 2026; originally announced March 2026.

    Comments: Accepted by Interspeech 2026; demo available https://plugandsteer.github.io

  18. arXiv:2603.14337  [pdf, ps, other

    cs.CV

    On the Nature of Attention Sink that Shapes Decoding Strategy in Omni-LLMs

    Authors: Suho Yoo, Youngjoon Jang, Joon Son Chung

    Abstract: The goal of this paper is to strengthen the reasoning of Omnimodal Large Language Models (Omni-LLMs) at inference time, without additional training. These models jointly process video, audio, and text, and given the large number of tokens they consume, how attention is routed across them is central to their behaviour. We focus specifically on attention sinks, tokens that absorb a disproportionate… ▽ More

    Submitted 7 May, 2026; v1 submitted 15 March, 2026; originally announced March 2026.

    Comments: Preprint

  19. arXiv:2603.12342  [pdf, ps, other

    eess.AS

    MamTra: A Hybrid Mamba-Transformer Backbone for Speech Synthesis

    Authors: Tan Dat Nguyen, Sangmin Bae, Joon Son Chung, Ji-Hoon Kim

    Abstract: Despite the remarkable quality of LLM-based text-to-speech systems, their reliance on autoregressive Transformers leads to quadratic computational complexity, which severely limits practical applications. Linear-time alternatives, notably Mamba, offer a potential remedy; however, they often sacrifice the global context essential for expressive synthesis. In this paper, we propose MamTra, an interl… ▽ More

    Submitted 19 June, 2026; v1 submitted 12 March, 2026; originally announced March 2026.

    Comments: Interspeech 2026

  20. arXiv:2602.05161  [pdf, ps, other

    hep-ex

    LiFE-SNS: LiF Experiment for keV-scale Sterile Neutrino Search

    Authors: Y. C. Lee, J. S. Chung, S. H. Choi, J. A. Jeon, D. H. Hwang, C. S. Kang, H. B. Kim, Ho Jong Kim, Hyeok Jun Kim, H. L. Kim, M. B. Kim, S. C. Kim, S. K. Kim, W. T. Kim, Y. H. Kim, Y. M. Kim, D. H. Kwon, D. Y. Lee, H. J. Lee, S. H. Lee, S. W. Lee, H. S. Lim, H. S. Park, K. R. Woo, J. Y. Yang , et al. (1 additional authors not shown)

    Abstract: The LiF Experiment for keV-scale Sterile Neutrino Search (LiFE-SNS) aims to probe sterile neutrinos through precision measurements of the tritium $β$ spectrum. Tritium nuclei are produced and embedded in LiF crystals via the ${}^{6}\mathrm{Li}(n,α){}^{3}\mathrm{H}$ reaction, allowing thermal calorimetric detection of $β$ decays with magnetic microcalorimeters (MMCs) operated at millikelvin tempera… ▽ More

    Submitted 4 February, 2026; originally announced February 2026.

  21. arXiv:2601.13143  [pdf, ps, other

    cs.LG

    FastAV: Efficient Token Pruning for Audio-Visual Large Language Model Inference

    Authors: Chaeyoung Jung, Youngjoon Jang, Seungwoo Lee, Joon Son Chung

    Abstract: In this work, we present FastAV, the first token pruning framework tailored for audio-visual large language models (AV-LLMs). While token pruning has been actively explored in standard large language models (LLMs) and vision-language models (LVLMs), its application to AV-LLMs has received little attention, even though multimodal integration substantially increases their token demands. To address t… ▽ More

    Submitted 19 January, 2026; originally announced January 2026.

  22. arXiv:2601.12802  [pdf, ps, other

    cs.SD eess.AS

    UNMIXX: Untangling Highly Correlated Singing Voices Mixtures

    Authors: Jihoo Jung, Ji-Hoon Kim, Doyeop Kwak, Junwon Lee, Juhan Nam, Joon Son Chung

    Abstract: We introduce UNMIXX, a novel framework for multiple singing voices separation (MSVS). While related to speech separation, MSVS faces unique challenges: data scarcity and the highly correlated nature of singing voices mixture. To address these issues, we propose UNMIXX with three key components: (1) musically informed mixing strategy to construct highly correlated, music-like mixtures, (2) cross-so… ▽ More

    Submitted 19 January, 2026; originally announced January 2026.

    Comments: Accepted by ICASSP 2026

  23. arXiv:2601.04658  [pdf, ps, other

    cs.SD cs.AI

    LAMB: LLM-based Audio Captioning with Modality Gap Bridging via Cauchy-Schwarz Divergence

    Authors: Hyeongkeun Lee, Jongmin Choi, KiHyun Nam, Joon Son Chung

    Abstract: Automated Audio Captioning aims to describe the semantic content of input audio. Recent works have employed large language models (LLMs) as a text decoder to leverage their reasoning capabilities. However, prior approaches that project audio features into the LLM embedding space without considering cross-modal alignment fail to fully utilize these capabilities. To address this, we propose LAMB, an… ▽ More

    Submitted 16 March, 2026; v1 submitted 8 January, 2026; originally announced January 2026.

    Comments: 5 pages, 2 figures; Accepted to ICASSP 2026

  24. arXiv:2512.20314  [pdf, ps, other

    eess.AS

    LP-CFM: Perceptual Invariance-Aware Conditional Flow Matching for Speech Modeling

    Authors: Doyeop Kwak, Youngjoon Jang, Joon Son Chung

    Abstract: The goal of this paper is to provide a new perspective on speech modeling by incorporating perceptual invariances such as amplitude scaling and temporal shifts. Conventional generative formulations often treat each dataset sample as a fixed representative of the target distribution. From a generative standpoint, however, such samples are only one among many perceptually equivalent variants within… ▽ More

    Submitted 23 December, 2025; originally announced December 2025.

  25. arXiv:2512.20296  [pdf, ps, other

    cs.CV cs.AI eess.AS eess.IV

    TAVID: Text-Driven Audio-Visual Interactive Dialogue Generation

    Authors: Ji-Hoon Kim, Junseok Ahn, Doyeop Kwak, Joon Son Chung, Shinji Watanabe

    Abstract: The objective of this paper is to jointly synthesize interactive videos and conversational speech from text and reference images. With the ultimate goal of building human-like conversational systems, recent studies have explored talking or listening head generation as well as conversational speech generation. However, these works are typically studied in isolation, overlooking the multimodal natur… ▽ More

    Submitted 23 December, 2025; originally announced December 2025.

    Comments: Project page: https://mm.kaist.ac.kr/projects/TAVID

  26. arXiv:2512.08040  [pdf, ps, other

    cs.CV

    Lost in Translation, Found in Embeddings: Sign Language Translation and Alignment

    Authors: Youngjoon Jang, Liliane Momeni, Zifan Jiang, Joon Son Chung, Gül Varol, Andrew Zisserman

    Abstract: Our aim is to develop a unified model for sign language understanding, that performs sign language translation (SLT) and sign-subtitle alignment (SSA). Together, these two tasks enable the conversion of continuous signing videos into spoken language text and also the temporal alignment of signing with subtitles -- both essential for practical communication, large-scale corpus construction, and edu… ▽ More

    Submitted 8 December, 2025; originally announced December 2025.

  27. arXiv:2512.02339  [pdf, ps, other

    cs.CV cs.AI

    Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision

    Authors: Chenshuang Zhang, Kang Zhang, Joon Son Chung, In So Kweon, Junmo Kim, Chengzhi Mao

    Abstract: Distinguishing visually similar objects by their motion remains a critical challenge in computer vision. Although supervised trackers show promise, contemporary self-supervised trackers struggle when visual cues become ambiguous, limiting their scalability and generalization without extensive labeled data. We find that pre-trained video diffusion models inherently learn motion representations suit… ▽ More

    Submitted 1 December, 2025; originally announced December 2025.

    Comments: Accepted at NeurIPS 2025

  28. arXiv:2512.00115  [pdf, ps, other

    cs.SD cs.CV cs.MM

    MoLT: Mixture of Layer-Wise Tokens for Efficient Audio-Visual Learning

    Authors: Kyeongha Rho, Hyeongkeun Lee, Jae Won Cho, Joon Son Chung

    Abstract: In this paper, we propose Mixture of Layer-Wise Tokens (MoLT), a parameter- and memory-efficient adaptation framework for audio-visual learning. The key idea of MoLT is to replace conventional, computationally heavy sequential adaptation at every transformer layer with a parallel, lightweight scheme that extracts and fuses layer-wise tokens only from the late layers. We adopt two types of adapters… ▽ More

    Submitted 27 November, 2025; originally announced December 2025.

    Comments: 10 pages, 5 figures

  29. arXiv:2510.24103  [pdf, ps, other

    cs.SD cs.AI cs.MM eess.AS

    Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation

    Authors: Kang Zhang, Trung X. Pham, Suyeon Lee, Axi Niu, Arda Senocak, Joon Son Chung

    Abstract: We present MGAudio, a novel flow-based framework for open-domain video-to-audio generation, which introduces model-guided dual-role alignment as a central design principle. Unlike prior approaches that rely on classifier-based or classifier-free guidance, MGAudio enables the generative model to guide itself through a dedicated training objective designed for video-conditioned audio generation. The… ▽ More

    Submitted 28 October, 2025; originally announced October 2025.

    Comments: accepted by NeurIPS 2025

  30. arXiv:2510.13281  [pdf, ps, other

    eess.AS cs.CL cs.LG

    Two Heads Are Better Than One: Audio-Visual Speech Error Correction with Dual Hypotheses

    Authors: Sungnyun Kim, Kangwook Jang, Sungwoo Cho, Joon Son Chung, Hoirin Kim, Se-Young Yun

    Abstract: This paper introduces a new paradigm for generative error correction (GER) framework in audio-visual speech recognition (AVSR) that reasons over modality-specific evidences directly in the language space. Our framework, DualHyp, empowers a large language model (LLM) to compose independent N-best hypotheses from separate automatic speech recognition (ASR) and visual speech recognition (VSR) models.… ▽ More

    Submitted 15 October, 2025; originally announced October 2025.

    Comments: Preprint work

  31. arXiv:2510.11330  [pdf, ps, other

    cs.SD cs.AI cs.CL cs.LG eess.AS

    Diffusion-Link: Diffusion Probabilistic Model for Bridging the Audio-Text Modality Gap

    Authors: KiHyun Nam, Jongmin Choi, Hyeongkeun Lee, Jungwoo Heo, Joon Son Chung

    Abstract: Contrastive audio-language pretraining yields powerful joint representations, yet a persistent audio-text modality gap limits the benefits of coupling multimodal encoders with large language models (LLMs). We present Diffusion-Link, a diffusion-based modality-bridging module that generatively maps audio embeddings into the text-embedding distribution. The module is trained at the output embedding… ▽ More

    Submitted 13 October, 2025; originally announced October 2025.

    Comments: 5 pages. Submitted to IEEE ICASSP 2026

  32. arXiv:2509.20802  [pdf, ps, other

    eess.AS cs.SD

    SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS

    Authors: Tan Dat Nguyen, Jaehun Kim, Ji-Hoon Kim, Shukjae Choi, Youshin Lim, Joon Son Chung

    Abstract: The goal of this paper is to introduce SPADE, a framework for Structured Pruning and Adaptive Distillation for Efficient Large Language Model-based text-to-speech (LLM-TTS). Recent LLM-TTS systems achieve strong controllability and zero-shot generalization, but their large parameter counts and high latency limit real-world deployment. SPADE addresses this by combining (i) a pruning step guided by… ▽ More

    Submitted 29 January, 2026; v1 submitted 25 September, 2025; originally announced September 2025.

    Comments: ICASSP 2026

  33. arXiv:2509.19881  [pdf, ps, other

    eess.AS cs.SD

    MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model

    Authors: The Hieu Pham, Tan Dat Nguyen, Phuong Thanh Tran, Joon Son Chung, Duc Dung Nguyen

    Abstract: Speech enhancement remains challenging due to the trade-off between efficiency and perceptual quality. In this paper, we introduce MAGE, a Masked Audio Generative Enhancer that advances generative speech enhancement through a compact and robust design. Unlike prior masked generative models with random masking, MAGE employs a scarcity-aware coarse-to-fine masking strategy that prioritizes frequent… ▽ More

    Submitted 13 March, 2026; v1 submitted 24 September, 2025; originally announced September 2025.

    Comments: ICASSP 2026

  34. arXiv:2509.19831  [pdf, ps, other

    eess.AS

    SCORE: Scaling audio generation using Standardized COmposite REwards

    Authors: Jaemin Jung, Jaehun Kim, Inkyu Shin, Joon Son Chung

    Abstract: The goal of this paper is to enhance Text-to-Audio generation at inference, focusing on generating realistic audio that precisely aligns with text prompts. Despite the rapid advancements, existing models often fail to achieve a reliable balance between perceptual quality and textual alignment. To address this, we adopt Inference-Time Scaling, a training-free method that improves performance by inc… ▽ More

    Submitted 24 September, 2025; originally announced September 2025.

  35. arXiv:2508.13992  [pdf, ps, other

    eess.AS cs.SD

    MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

    Authors: Sonal Kumar, Šimon Sedláček, Vaibhavi Lokegaonkar, Fernando López, Wenyi Yu, Nishit Anand, Hyeonggon Ryu, Lichang Chen, Maxim Plička, Miroslav Hlaváček, William Fineas Ellingwood, Sathvik Udupa, Siyuan Hou, Allison Ferner, Sara Barahona, Cecilia Bolaños, Satish Rahi, Laura Herrera-Alarcón, Satvik Dixit, Siddhi Patil, Soham Deshmukh, Lasha Koroshinadze, Yao Liu, Leibny Paola Garcia Perera, Eleni Zanou , et al. (9 additional authors not shown)

    Abstract: Audio comprehension-including speech, non-speech sounds, and music-is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio understanding to qualify as generally intelligent. However, evaluating auditory intelligence comprehensively remains challenging. To address this gap, we introduce MMAU-Pro, the most comprehensive and rigorously curated benc… ▽ More

    Submitted 19 August, 2025; originally announced August 2025.

  36. arXiv:2507.04559  [pdf, ps, other

    cs.CV

    MambaVideo for Discrete Video Tokenization with Channel-Split Quantization

    Authors: Dawit Mureja Argaw, Xian Liu, Joon Son Chung, Ming-Yu Liu, Fitsum Reda

    Abstract: Discrete video tokenization is essential for efficient autoregressive generative modeling due to the high dimensionality of video data. This work introduces a state-of-the-art discrete video tokenizer with two key contributions. First, we propose a novel Mamba-based encoder-decoder architecture that overcomes the limitations of previous sequencebased tokenizers. Second, we introduce a new quantiza… ▽ More

    Submitted 6 July, 2025; originally announced July 2025.

    Comments: Project website: https://research.nvidia.com/labs/dir/mamba-tokenizer/

  37. arXiv:2506.16231  [pdf, ps, other

    eess.AS cs.SD

    EDNet: A Versatile Speech Enhancement Framework with Gating Mamba Mechanism and Phase Shift-Invariant Training

    Authors: Doyeop Kwak, Youngjoon Jang, Seongyu Kim, Joon Son Chung

    Abstract: Speech signals in real-world environments are frequently affected by various distortions such as additive noise, reverberation, and bandwidth limitation, which may appear individually or in combination. Traditional speech enhancement methods typically rely on either masking, which focuses on suppressing non-speech components while preserving observable structure, or mapping, which seeks to recover… ▽ More

    Submitted 4 February, 2026; v1 submitted 19 June, 2025; originally announced June 2025.

    Comments: Accepted by IEEE Transactions on Audio, Speech and Language Processing. Copyright IEEE. The final version will appear in IEEE Xplore

  38. arXiv:2506.03020  [pdf, ps, other

    eess.AS

    InfiniteAudio: Infinite-Length Audio Generation with Consistency

    Authors: Chaeyoung Jung, Hojoon Ki, Ji-Hoon Kim, Junmo Kim, Joon Son Chung

    Abstract: This paper presents InfiniteAudio, a simple yet effective strategy for generating infinite-length audio using diffusion-based text-to-audio methods. Current approaches face memory constraints because the output size increases with input length, making long duration generation challenging. A common workaround is to concatenate short audio segments, but this often leads to inconsistencies due to the… ▽ More

    Submitted 3 June, 2025; originally announced June 2025.

  39. arXiv:2505.20899  [pdf, ps, other

    cs.CL cs.SD eess.AS

    Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing

    Authors: Jeongsoo Choi, Jaehun Kim, Joon Son Chung

    Abstract: This paper introduces a cross-lingual dubbing system that translates speech from one language to another while preserving key characteristics such as duration, speaker identity, and speaking speed. Despite the strong translation quality of existing speech translation approaches, they often overlook the transfer of speech patterns, leading to mismatches with source speech and limiting their suitabi… ▽ More

    Submitted 28 December, 2025; v1 submitted 27 May, 2025; originally announced May 2025.

    Comments: EMNLP 2025 Findings

  40. arXiv:2505.20873  [pdf, ps, other

    cs.CV

    Fork-Merge Decoding: Enhancing Multimodal Understanding in Audio-Visual Large Language Models

    Authors: Chaeyoung Jung, Youngjoon Jang, Jongmin Choi, Joon Son Chung

    Abstract: The goal of this work is to enhance balanced multimodal understanding in audio-visual large language models (AV-LLMs) by addressing modality bias without additional training. In current AV-LLMs, audio and video features are typically processed jointly in the decoder. While this strategy facilitates unified multimodal understanding, it may introduce modality bias, where the model tends to over-rely… ▽ More

    Submitted 30 September, 2025; v1 submitted 27 May, 2025; originally announced May 2025.

  41. arXiv:2505.20862  [pdf, ps, other

    cs.CV

    AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding

    Authors: Chaeyoung Jung, Youngjoon Jang, Joon Son Chung

    Abstract: Hallucination remains a major challenge in multimodal large language models (MLLMs). To address this, various contrastive decoding (CD) methods have been proposed that contrasts original logits with hallucinated logits generated from perturbed inputs. While CD has shown promise in vision-language models (VLMs), it is not well-suited for AV-LLMs, where hallucinations often emerge from both unimodal… ▽ More

    Submitted 30 September, 2025; v1 submitted 27 May, 2025; originally announced May 2025.

  42. arXiv:2505.19595  [pdf, ps, other

    eess.AS cs.SD

    Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment

    Authors: Jeongsoo Choi, Zhikang Niu, Ji-Hoon Kim, Chunhui Wang, Joon Son Chung, Xie Chen

    Abstract: The goal of this paper is to optimize the training process of diffusion-based text-to-speech models. While recent studies have achieved remarkable advancements, their training demands substantial time and computational costs, largely due to the implicit guidance of diffusion models in learning complex intermediate representations. To address this, we propose A-DMA, an effective strategy for Accele… ▽ More

    Submitted 30 May, 2025; v1 submitted 26 May, 2025; originally announced May 2025.

    Comments: Interspeech 2025

  43. arXiv:2505.16798  [pdf, ps, other

    eess.AS cs.AI

    SEED: Speaker Embedding Enhancement Diffusion Model

    Authors: KiHyun Nam, Jungwoo Heo, Jee-weon Jung, Gangin Park, Chaeyoung Jung, Ha-Jin Yu, Joon Son Chung

    Abstract: A primary challenge when deploying speaker recognition systems in real-world applications is performance degradation caused by environmental mismatch. We propose a diffusion-based method that takes speaker embeddings extracted from a pre-trained speaker recognition model and generates refined embeddings. For training, our approach progressively adds Gaussian noise to both clean and noisy speaker e… ▽ More

    Submitted 22 May, 2025; originally announced May 2025.

    Comments: Accepted to Interspeech 2025. The official code can be found at https://github.com/kaistmm/seed-pytorch

  44. arXiv:2505.09256  [pdf, other

    cs.CV

    Test-Time Augmentation for Pose-invariant Face Recognition

    Authors: Jaemin Jung, Youngjoon Jang, Joon Son Chung

    Abstract: The goal of this paper is to enhance face recognition performance by augmenting head poses during the testing phase. Existing methods often rely on training on frontalised images or learning pose-invariant representations, yet both approaches typically require re-training and testing for each dataset, involving a substantial amount of effort. In contrast, this study proposes Pose-TTA, a novel appr… ▽ More

    Submitted 14 May, 2025; originally announced May 2025.

  45. arXiv:2505.05343  [pdf, other

    cs.CV cs.SD eess.AS

    Hearing and Seeing Through CLIP: A Framework for Self-Supervised Sound Source Localization

    Authors: Sooyoung Park, Arda Senocak, Joon Son Chung

    Abstract: Large-scale vision-language models demonstrate strong multimodal alignment and generalization across diverse tasks. Among them, CLIP stands out as one of the most successful approaches. In this work, we extend the application of CLIP to sound source localization, proposing a self-supervised method operates without explicit text input. We introduce a framework that maps audios into tokens compatibl… ▽ More

    Submitted 8 May, 2025; originally announced May 2025.

    Comments: Journal Extension of WACV 2024 paper (arXiv:2311.04066). Code is available at https://github.com/swimmiing/ACL-SSL

  46. arXiv:2504.20629  [pdf, ps, other

    cs.CV cs.AI cs.MM

    AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation

    Authors: Jeongsoo Choi, Ji-Hoon Kim, Kim Sung-Bin, Tae-Hyun Oh, Joon Son Chung

    Abstract: In this paper, we address the task of multimodal-to-speech generation, which aims to synthesize high-quality speech from multiple input modalities: text, video, and reference audio. This task has gained increasing attention due to its wide range of applications, such as film production, dubbing, and virtual avatars. Despite recent progress, existing methods still suffer from limitations in speech… ▽ More

    Submitted 3 October, 2025; v1 submitted 29 April, 2025; originally announced April 2025.

    Comments: ACM Multimedia 2025

  47. arXiv:2504.02386  [pdf, other

    cs.CV eess.AS

    VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models

    Authors: Kim Sung-Bin, Jeongsoo Choi, Puyuan Peng, Joon Son Chung, Tae-Hyun Oh, David Harwath

    Abstract: We present VoiceCraft-Dub, a novel approach for automated video dubbing that synthesizes high-quality speech from text and facial cues. This task has broad applications in filmmaking, multimedia creation, and assisting voice-impaired individuals. Building on the success of Neural Codec Language Models (NCLMs) for speech synthesis, our method extends their capabilities by incorporating video featur… ▽ More

    Submitted 3 April, 2025; originally announced April 2025.

    Comments: https://voicecraft-dub.github.io/

  48. arXiv:2503.18880  [pdf, other

    cs.CV cs.SD eess.AS

    Seeing Speech and Sound: Distinguishing and Locating Audios in Visual Scenes

    Authors: Hyeonggon Ryu, Seongyu Kim, Joon Son Chung, Arda Senocak

    Abstract: We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches are typically limited to handling either speech or non-speech sounds independently, or at best, together but sequentially without mixing. This limitation prevents them from capturing… ▽ More

    Submitted 24 March, 2025; originally announced March 2025.

    Comments: CVPR 2025

  49. arXiv:2503.16956  [pdf, other

    eess.AS cs.AI cs.CV cs.SD

    From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech

    Authors: Ji-Hoon Kim, Jeongsoo Choi, Jaehun Kim, Chaeyoung Jung, Joon Son Chung

    Abstract: The objective of this study is to generate high-quality speech from silent talking face videos, a task also known as video-to-speech synthesis. A significant challenge in video-to-speech synthesis lies in the substantial modality gap between silent video and multi-faceted speech. In this paper, we propose a novel video-to-speech system that effectively bridges this modality gap, significantly enha… ▽ More

    Submitted 21 March, 2025; originally announced March 2025.

    Comments: CVPR 2025, demo page: https://mm.kaist.ac.kr/projects/faces2voices/

  50. arXiv:2503.03287  [pdf, other

    cs.CV

    Deep Understanding of Sign Language for Sign to Subtitle Alignment

    Authors: Youngjoon Jang, Jeongsoo Choi, Junseok Ahn, Joon Son Chung

    Abstract: The objective of this work is to align asynchronous subtitles in sign language videos with limited labelled data. To achieve this goal, we propose a novel framework with the following contributions: (1) we leverage fundamental grammatical rules of British Sign Language (BSL) to pre-process the input subtitles, (2) we design a selective alignment loss to optimise the model for predicting the tempor… ▽ More

    Submitted 5 March, 2025; originally announced March 2025.