Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 492 results for author: Kang, H

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.19669  [pdf, ps, other

    cs.CV cs.LG

    Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

    Authors: Haoqiang Kang, Yinpeng Chen, Luyang Liu, Jesper Sparre Andersen, Abhijit Ogale, Baochen Sun, Lichan Hong, Ed H. Chi

    Abstract: Latent reasoning has advanced multimodal reasoning through a two-stage training paradigm: (1) a helper image is encoded into latent tokens to teach visual chain-of-thought during a supervised fine-tuning (SFT) stage, and (2) these latent tokens are further refined with reward feedback during a reinforcement learning (RL) stage. In this paper, we identify two key limitations of this framework, one… ▽ More

    Submitted 20 August, 2026; originally announced August 2026.

    Report number: SM-2026-08-19

  2. Real-Time Control-Constrained DDP for Underactuated Balancing of Legged Robots

    Authors: SeongWon Nam, Hyunyong Lee, Hansol Kang, Jiman Park, Yeongwoo Son, Bumsu Yi, Jaeyoung Oh, Hyouk Ryeol Choi

    Abstract: This paper presents a real-time control-constrained Differential Dynamic Programming (DDP) framework for underactuated legged robots. To address the limitation of classical DDP in handling control constraints, we propose an Accelerated Projected Gradient (APG)-based control-constrained DDP (ABC-DDP), which efficiently computes constrained solutions and identifies active sets without repeated Karus… ▽ More

    Submitted 19 August, 2026; originally announced August 2026.

    Comments: Accepted for publication in IEEE Robotics and Automation Letters (RA-L)

    Journal ref: IEEE Robotics and Automation Letters (RA-L), 2026

  3. arXiv:2608.09467  [pdf, ps, other

    cs.CV cs.AI

    RecoverFly: A Failure-Aware Reinforcement Learning Post-Training Framework for Aerial Vision-Language Navigation

    Authors: Boxiong Wang, Hui Kang, Geng Sun, Jiahui Li, Chao Yu, Daxin Tian

    Abstract: Unmanned aerial vehicle vision-language navigation (UAV-VLN) requires agents to translate visual observations and language instructions into reliable flight actions in complex environments. Although recent end-to-end UAV vision-language-action (UAV-VLA) policies reduce reliance on separately designed perception, planning, and control modules, their behavior-cloning objectives provide limited corre… ▽ More

    Submitted 10 August, 2026; originally announced August 2026.

  4. arXiv:2608.07663  [pdf, ps, other

    cs.CV cs.AI cs.CL

    Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding

    Authors: Yeeun Choi, Youngbeom Yoo, Joon-Young Lee, Hyolim Kang, Seon Joo Kim

    Abstract: When videos extend from hours to days, directly processing them end-to-end becomes impractical for current Multi-modal Large Language Models (MLLMs). This ultra-long setting necessitates a two-stage paradigm: query-agnostic memory construction followed by retrieval-based inference. Prior work invests in complex memory construction to pre-model high-level relations in videos, despite not knowing th… ▽ More

    Submitted 7 August, 2026; originally announced August 2026.

    Comments: Accepted to ECCV 2026 (Oral). Project Page: https://choi-yeeun.github.io/MERIT/

  5. arXiv:2608.03450  [pdf, ps, other

    cs.MM cs.AI cs.CL cs.CV

    Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs

    Authors: Haoqian Kang, Liupeng Li, Kuofeng Gao, Jinpeng Wang, Zhenyu Lu, Bin Chen, Ke Chen, Yaowei Wang

    Abstract: Reasoning in Multimodal Large Language Models (MLLMs) requires both fine-grained visual perception and rigorous logical deduction. Explicit text-based Chain-of-Thought (CoT) is computationally expensive and prone to visual hallucinations, while existing latent reasoning methods typically require costly training. Furthermore, directly adapting training-free LLM reasoning mechanisms to the multimoda… ▽ More

    Submitted 4 August, 2026; originally announced August 2026.

    Comments: Accepted by ACM MM 2026. 10 pages, 6 figures, 5 tables

  6. arXiv:2608.02218  [pdf, ps, other

    cs.AI

    PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs

    Authors: Haojie Hu, Chenhao Dang, Yaojia Liu, Hengrui Kang, Conghui He, Weijia Li

    Abstract: Scientific poster construction compresses a long multimodal paper into a readable, editable canvas. Existing systems hide request-level failures by scoring only completed outputs; direct image generation is not element-editable, while coding-agent workflows are costly. PosterMELD is a template-conditioned multi-agent pipeline: capacity-aware slots guide writing before rendering, and deterministic… ▽ More

    Submitted 3 August, 2026; originally announced August 2026.

    Comments: 9 pages, 5 figures, and 4 tables. Code and resources are available at https://github.com/Shannon4Science/PosterMELD

  7. arXiv:2608.00353  [pdf, ps, other

    cs.AR

    Analyzing RV32/RV64 Trade-offs for FreeRTOS Latency on 8-Stage RISC-V Soft Processors

    Authors: Hyunwoo Kang, Geonwoo Yu, Jongwon Kim, Seungwoo You, Minchan Gil

    Abstract: Although many commercial RISC-V platforms provide real-time operating system support, practical examples that explain how to enable a preemptive RTOS on a custom bare-metal RISC-V soft processor remain limited, leaving the interaction between processor microarchitecture, interrupt handling, and RTOS context switching difficult to understand from simple hardware implementation examples. This paper… ▽ More

    Submitted 31 July, 2026; originally announced August 2026.

    Comments: 5 pages, 1 figure. Accepted for presentation at APCCAS 2026; to be presented on October 26, 2026

    ACM Class: C.1.0; B.7.1

  8. arXiv:2607.24327  [pdf

    cond-mat.mtrl-sci cs.CE physics.chem-ph

    Aligning Heterogeneous DFT Datasets: A Graph Neural Network Approach to Cross-Functional Formation Energies

    Authors: Yidong Huang, Tenglong Lu, Hanwen Kang, Junfeng Huang, Sheng Meng, Miao Liu

    Abstract: Heterogeneous density functional theory (DFT) calculations, particularly plane-wave implementations, introduce systematic formation energy errors ranging from tens to hundreds of meV/atom, depending on the selection of exchange-correlation functionals, kinetic energy cutoffs, pseudopotentials, and dispersion corrections. As demonstrated by the MatPES dataset, identical structures can exhibit an av… ▽ More

    Submitted 27 July, 2026; originally announced July 2026.

  9. arXiv:2607.24030  [pdf, ps, other

    cs.CL cs.SD

    MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition

    Authors: Sangmin Lee, Woojin Chung, Woongjib Choi, Hong-Goo Kang

    Abstract: Massively multilingual automatic speech recognition (ASR) models covering hundreds of languages must maintain robust performance across diverse linguistic and acoustic conditions. However, these models often encounter the curse of multilinguality, where model capacity is diluted across languages. To address this challenge, we propose Mixture of Language Group Experts (MoLGE), built upon speech sel… ▽ More

    Submitted 27 July, 2026; originally announced July 2026.

    Comments: Accepted to COLM 2026, Github: https://github.com/sanghyang00/molge

  10. arXiv:2607.22076  [pdf, ps, other

    cs.CR cs.SE

    PoCEvolve: Generating Proof-of-Concept Exploits from Security Patches with Vulnerability-Aware Prompt Evolution

    Authors: Duc Manh Tran, Ratnadira Widyasari, Ivana Clairine Irsan, Huihui Huang, Ting Zhang, Shar Lwin Khin, Ouh Eng Lieh, Hong Jin Kang, David Lo

    Abstract: Ideally, the detailed information about a vulnerability should be made available together with the fixing commit. In practice, however, such details often become available only long after the commit, even when a CVE has already been published. During this window, the patch is already public, so attackers can reverse-engineer it, yet defenders lack the details needed to assess exposure, prioritize,… ▽ More

    Submitted 24 July, 2026; originally announced July 2026.

  11. arXiv:2607.20062  [pdf, ps, other

    cs.CL

    Solar Open 2 Technical Report

    Authors: Sungrae Park, Sanghoon Kim, Gyoungjin Gim, Jungho Cho, Hyunwoong Ko, Minbyul Jeong, Minjeong Kim, Keunwoo Choi, Chaehun Shin, Chanwoong Yoon, Dongjun Kim, Eunwon Kim, Gyungin Shin, Hyeonju Lee, Hyungkyu Kang, Inseo Song, Jisu Bae, Jiyoon Han, Jiyun Lee, Joonkee Kim, Junyeop Lee, Mikyoung Cha, Sangwon Yu, Sehwan Joo, Seokyoon Kang , et al. (28 additional authors not shown)

    Abstract: We present Solar Open 2, a 250B-A15B Mixture-of-Experts language model built for long-horizon agentic tasks, scaled up from Solar Open 1 (Solar Open 100B). To hold entire agent trajectories in a single context, Solar Open 2 reaches a 1M-token window through a hybrid attention stack that interleaves one softmax layer among every three linear-attention layers, using no positional encoding and a gate… ▽ More

    Submitted 23 July, 2026; v1 submitted 22 July, 2026; originally announced July 2026.

  12. arXiv:2607.19669  [pdf, ps, other

    cs.CV

    A Unified Variational Framework for Deep Weakly Supervised Image Segmentation

    Authors: Yin King Chu, Lingfeng Li, Sung Ha Kang, Jianping Zhang, Xue-Cheng Tai

    Abstract: We propose a unified variational framework for image segmentation under sparse pixel-level supervision. Our method is based on a simplex-constrained Potts model with a smooth perimeter regularizer, yielding a convex, smooth energy functional that can be used as a training loss in weakly supervised deep learning paradigms or optimized efficiently using iterative methods. Sparse labels are incorpora… ▽ More

    Submitted 21 July, 2026; originally announced July 2026.

  13. arXiv:2607.19517  [pdf, ps, other

    cs.CV

    Crowd4D: Scene-Aware Monocular 4D Crowd Reconstruction

    Authors: Hongbo Kang, Tianyi Zhou, Qingyang Yang, Hongwei Wen, Jing Huang, Yu-Kun Lai, Kun Li

    Abstract: Recovering scene-consistent 4D crowd motion from monocular video in large-scale scenes remains challenging due to severe depth ambiguity and complex scene geometry. Existing monocular crowd reconstruction methods typically rely on single-plane assumptions, leading to unreliable metric scale and spatial drift under complex terrain. We propose Crowd4D, the first scene-aware 4D crowd reconstruction f… ▽ More

    Submitted 21 July, 2026; originally announced July 2026.

  14. arXiv:2607.17713  [pdf, ps, other

    cs.CL

    AEGIS: Awareness-Enhanced Guidance for Iterative Safeguard

    Authors: Kyungwon Park, Sangmin Lee, Heejae Chon, Hyungu Kang

    Abstract: Span-level rationales are often assumed to improve controllability in text detoxification, but it remains unclear when such guidance helps and when it introduces trade-offs. We present Awareness-Enhanced Guidance for Iterative Safeguard (AEGIS) as an exploratory framework for studying span-guided multilingual detoxification across English, Mandarin Chinese, and Korean. AEGIS combines span-level de… ▽ More

    Submitted 20 July, 2026; originally announced July 2026.

    Comments: 11 pages, 3 figures, 9 tables. Preprint

  15. arXiv:2607.17572  [pdf, ps, other

    cs.LG cs.CV eess.SY

    JAGG: Jacobian-Aggregated Group Gradient for Efficient GRPO Training of Diffusion Models

    Authors: Ruiyi Ding, Jie Li, He Kang, Ziyan Liu, Chengru Song, Yuan cheng

    Abstract: Group Relative Policy Optimization (GRPO) is a powerful reinforcement learning algorithm for aligning generative models with human preferences. While successful in large language models~\cite{shao2024deepseekmathpushinglimitsmathematical}, its extension to diffusion and flow matching models introduces a severe computational bottleneck: gradients must be back-propagated through the high-capacity Di… ▽ More

    Submitted 25 July, 2026; v1 submitted 20 July, 2026; originally announced July 2026.

    Comments: 21 pages

    ACM Class: I.4.5

  16. arXiv:2607.12287  [pdf, ps, other

    cs.RO

    Reducing Temporal Redundancy for Efficient Vision-Language-Action Inference

    Authors: Yuzhou Wu, Yuxin Zheng, Muchun Niu, Yishan Yang, Tianhao Liu, hanwen kang, Jiajian Jing, Linfeng Zhang, Chuan Wen

    Abstract: Vision-Language-Action (VLA) models exhibit strong generalization for robotic manipulation, yet their high inference latency limits real time deployment. We identify two primary sources of temporal redundancy in existing VLA pipelines: repeated visual encoding of highly similar consecutive frames and multi step iterative sampling in diffusion based policies. To address this, we propose a system le… ▽ More

    Submitted 13 July, 2026; originally announced July 2026.

    Comments: 13pages, 7 figuers

  17. arXiv:2607.09757  [pdf, ps, other

    cs.CL cs.AI

    RSRA: Training-Free Probing of Representation Sensitivity for Efficient LoRA Rank Allocation

    Authors: Jiaqi Liu, Haidong Kang, Qihui Zhao, Guo Yu, Jingchao Wang

    Abstract: Parameter-efficient fine-tuning enables large language models to adapt to downstream tasks with substantially lower computational and storage cost, and Low-Rank Adaptation (LoRA) is among its most widely used techniques. However, vanilla LoRA assigns a uniform rank to all adapted modules, while existing adaptive methods either incur additional optimization overhead or rely on static weights and lo… ▽ More

    Submitted 1 August, 2026; v1 submitted 5 July, 2026; originally announced July 2026.

    Comments: 14 pages, 5 figures

  18. arXiv:2607.07647  [pdf

    cond-mat.mtrl-sci cs.CE physics.comp-ph

    Are Machine Learning Interatomic Potentials Truly Practical? A Benchmark of 23 Mainstream Models

    Authors: Hanwen Kang, Tenglong Lu, Sheng Meng, Miao Liu

    Abstract: Most MLIP benchmarks reward static accuracy while ignoring inference efficiency and hardware scalability -- driving model bloat with unclear real-world value. We benchmark 23 mainstream open-source MLIPs on a low-cost NVIDIA DGX Spark (128 GB native memory, capped at 80 GB to mimic ordinary lab hardware), using a fixed 192-atom system under a unified ASE-based pipeline. We evaluate three dimension… ▽ More

    Submitted 8 July, 2026; originally announced July 2026.

  19. arXiv:2607.05197  [pdf, ps, other

    cs.SE

    Is Three the Magic Number? An Empirical Evaluation of LLM-Based Repair Loops

    Authors: Tobias Kiecker, Eik Reichmann, Hosung Kang, Gabin An, Lars Grunske

    Abstract: Iterative repair loops have become a core design pattern in LLM-based software engineering systems. These workflows repeatedly generate, validate, and repair artifacts using feedback such as compiler errors or test failures. Despite their widespread use, the impact of repair-loop iteration limits remains poorly understood, as most prior work adopts fixed, often arbitrary, repair budgets. We study… ▽ More

    Submitted 6 July, 2026; originally announced July 2026.

    Comments: 4 Pages (+1 for references), NIER Paper

  20. arXiv:2607.03744  [pdf, ps, other

    cs.AI

    Can Conversational Temporal Dynamics Improve Depression Detection in Dyads? A Preliminary Investigation in Multi-Modality Perspectives

    Authors: Hanie Kang, Huang-Cheng Chou, Sudarsana Reddy Kadiri, Shrikanth Narayanan

    Abstract: Automatic depression detection from clinical interviews typically models the semantic content and acoustic characteristics of participant speech. However, the interactional timing between the clinician and participant remains comparatively under-modeled. We investigate conversational temporal dynamics, specifically dyadic turn-pair timing, as a primary modality fused with self-supervised encoders.… ▽ More

    Submitted 4 July, 2026; originally announced July 2026.

    Comments: Submitted to SLT 2026

  21. arXiv:2607.02805  [pdf, ps, other

    cs.LG cs.AI cs.CL

    Training Hybrid Block Diffusion Language Models with Partial Bidirectionality

    Authors: Pranshu Chaturvedi, Parth Shroff, Tarun Suresh, Hangoo Kang, Kaiyue Wen

    Abstract: High-throughput long-context generation is one of the central challenges for large language models. Generation is typically memory-bandwidth-bound rather than compute-bound: each decoding step must stream the accumulated key/value (KV) cache from memory, so bandwidth demand grows with context length while only one token is emitted. Two parallel approaches have therefore emerged: reducing memory ac… ▽ More

    Submitted 2 July, 2026; originally announced July 2026.

    Comments: 16 pages, 3 figures

  22. arXiv:2607.00720  [pdf, ps, other

    cs.LG cs.AI

    Detecting the Undetectable: Enhancing Unsupervised time series Anomaly Detection via Active Learning

    Authors: Seung Hun Han, Hyeongwon Kang, Jinwoo Park, Pilsung Kang

    Abstract: Despite the increasing sophistication of industrial AI systems, the ability to reliably detect subtle and noisy anomalies in complex time series data remains a critical yet unresolved challenge. In large-scale industrial applications, labeling time series data is often prohibitively expensive and time-consuming, making unsupervised learning a practical and widely adopted approach. However, existin… ▽ More

    Submitted 1 July, 2026; originally announced July 2026.

  23. arXiv:2606.27187  [pdf, ps, other

    cs.CV cs.CL

    HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models

    Authors: Jiajun Wu, Haoyu Kang, Yining Sun, Jiacheng Hou, Heng Zhang, Danyang Zhang, Zhenjun Zhao, Haochi Zhang, Leixin Sun, Eric Hanchen Jiang, Yushan Li, Ruiyu Li, Mengkai Huang, Yan Gao, Xu Zhang, Guancheng Wan

    Abstract: Large vision-language models (LVLMs) have recently shown immense potential in automated content moderation, sparking growing interest in developing harmful-video benchmarks. However, we identify two primary limitations in existing works: 1) The multi-layered characteristics of harmful videos are overlooked. Existing benchmarks predominantly formulate evaluation as a binary classification task, fai… ▽ More

    Submitted 25 June, 2026; originally announced June 2026.

  24. arXiv:2606.25592  [pdf, ps, other

    cs.CV

    VPA-Guard: Defending and Benchmarking Image-to-Video Generation Against Visual Prompt Attacks

    Authors: Yining Sun, Haoyu Kang, Jiajun Wu, Heng Zhang, Danyang Zhang, Zhenjun Zhao, Haochen Han, Fangming Liu, Wai Kin Victor Chan, Alex Jinpeng Wang

    Abstract: Recent advancements in Image-to-Video (I2V) generation have transformed input images from simple appearance references into interactive control interfaces where visual cues such as arrows, sketches, and emojis orchestrate complex video dynamics with unprecedented controllability. However, these seemingly innocuous static cues can be interpreted by models as executable temporal instructions, unfold… ▽ More

    Submitted 24 June, 2026; originally announced June 2026.

    Comments: Dataset Page: https://huggingface.co/datasets/CSU-JPG/VVA-Bench

  25. arXiv:2606.24338  [pdf, ps, other

    cs.RO

    RoBoSR: Structured Scene Representations for Embodied Robotic Reasoning

    Authors: Kewei Hu, Wanchan Yu, Fangwen Chen, Jing Jiajian, Zimeng Li, Ying Wei, Tianhao Liu, Michael Zhang, Hanwen Kang

    Abstract: Despite rapid progress, embodied reasoning under real-world variability remains challenging. Existing approaches rely on demonstration-driven sequential biases, limiting flexibility in open-ended and long-horizon tasks that require structured reasoning over evolving states. We introduce RoBoSR, an intermediate structural representation that formulates manipulation as step-wise state transitions… ▽ More

    Submitted 23 June, 2026; originally announced June 2026.

  26. arXiv:2606.17730  [pdf, ps, other

    cs.CV

    ActWorld: From Explorable to Interactive World Model via Action-Aware Memory

    Authors: Zhexiao Xiong, Yizhi Song, Hao Kang, Qing Yan, Liming Jiang, Jenson Yang, Zhoujie Fu, Stathi Fotiadis, Angtian Wang, Zichuan Liu, Bo Liu, Yiding Yang, Xin Lu, Nathan Jacobs

    Abstract: Interactive world models aim to simulate environment dynamics under real-time user actions. However, their action vocabulary is largely confined to navigation: most actions correspond to motion (e.g., walk, turn, look around), while interaction with objects in the scene (e.g., pick up plates, open doors, or trigger physical responses) is either absent, restricted to game domains, or relegated to p… ▽ More

    Submitted 16 June, 2026; originally announced June 2026.

    Comments: Project page: https://interactwm.github.io/ActWorld

  27. arXiv:2606.15434  [pdf, ps, other

    cs.RO cs.HC eess.SY

    A Bilateral Teleoperation Framework for Dexterous Manipulation

    Authors: Stefano Dalla Gasperina, Dong Ho Kang, Haiyun Zhang, Aldo Galvan, Job D. Ramirez, Aaron Kim, Mark Helwig, Kazuto Yokoyama, Takahisa Ueno, Tetsuya Narita, Ann Majewicz-Fey, Ashish D. Deshpande, Luis Sentis

    Abstract: Dexterous teleoperation requires precise arm-hand coordination, low-latency feedback, and robust interaction in real-world contact-rich environments. This paper presents a modular bilateral teleoperation framework that integrates operator-side input interfaces with a robot-side dexterous hand and compliant robotic arm in a unified control architecture. The system supports position-based hand retar… ▽ More

    Submitted 13 June, 2026; originally announced June 2026.

    Comments: 4 pages, 7 figures, 1 appendix,

  28. arXiv:2606.15104  [pdf, ps, other

    cs.CV

    Text-Driven Fusion for Infrared and Visible Images: Achieving Image Scene Adaptation on Hyperbolic Space

    Authors: Huan Kang, Hui Li, Tianyang Xu, Tao Zhou, Xiao-Jun Wu, Josef Kittler

    Abstract: Infrared and visible image fusion aims to integrate complementary modalities, while existing Euclidean methods impose rigid distance metrics that distort multi-modal interactions and parent-to-child semantic hierarchies. To overcome these limitations, we introduce a text-driven fusion framework empowered by hyperbolic manifold learning. During training, BLIP-extracted text prompts serve as topolog… ▽ More

    Submitted 13 June, 2026; originally announced June 2026.

    Comments: 14 pages, 8 figures

    ACM Class: I.4

  29. arXiv:2606.14805  [pdf, ps, other

    cs.SE cs.AI

    Knowledge-Based Zero-Replay Debugging of Multi-Agent LLM Traces

    Authors: Dong Ho Kang, Hyeonjeong Cha, Daein Weon

    Abstract: Reliable operation of multi-agent large language model (LLM) systems depends on debugging long execution traces, where the few causally decisive events are buried in unstructured logs of messages, routes, memory writes, and tool calls. The standard tool is counterfactual replay (rewind, edit, and re-run the trajectory to measure each event's effect), but its cost grows linearly with the number of… ▽ More

    Submitted 11 June, 2026; originally announced June 2026.

    Comments: 21 pages, 1 figure, 6 tables. Submitted to Knowledge-Based Systems

    ACM Class: I.2.11; I.2.6; D.2.5

  30. arXiv:2606.14674  [pdf, ps, other

    cs.CL

    AgentSpec: Understanding Embodied Agent Scaffolds Through Controlled Composition

    Authors: Jixuan Chen, Jianzhi Shen, Haoqiang Kang, Zhi Hong, Qingyi Jiang, Soham Bose, Yiming Zhang, Leon Leng, Amit Vyas, Lingjun Mao, Siru Ouyang, Kun Zhou, Lianhui Qin

    Abstract: LLM agents are increasingly built not as single model calls, but as scaffolded systems that combine reasoning, memory, reflection, action execution, and learning. While such scaffolds often improve performance, they are often embedded in tightly coupled pipelines, making it difficult to isolate component contributions, compare alternative designs, or understand how module interactions shape agent… ▽ More

    Submitted 12 June, 2026; originally announced June 2026.

  31. arXiv:2606.12900  [pdf, ps, other

    cs.AI cs.CL cs.LG

    Zero-source LLM Hallucination Detection with Human-like Criteria Probing

    Authors: Jiahao Yang, Shuhai Zhang, Hailong Kang, Feng Liu, Qi Chen, Mingkui Tan

    Abstract: Large language models (LLMs) often hallucinate by generating factually incorrect or unfaithful content, posing significant risks to their safe use. Detecting such hallucinations is particularly challenging under the zero-source constraint, where no model internals or external references are available, and detection must rely solely on the textual query-answer pair. In this paper, we propose Human-… ▽ More

    Submitted 11 June, 2026; originally announced June 2026.

    Comments: Accepted at ICML 2026

  32. arXiv:2606.11681  [pdf, ps, other

    cs.CL cs.SD

    UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction

    Authors: Sangmin Lee, Eekgyun Ahn, Woongjib Choi, Hong-Goo Kang

    Abstract: We propose UR-BERT, a Romanized transcription-based text-to-speech (TTS) encoder for massively multilingual TTS systems. Conventional grapheme-to-phoneme (G2P)-based approaches are limited to around 100 languages due to the availability of reliable G2P resources. In contrast, UR-BERT scales to 495 languages by unifying diverse writing systems into a shared Romanization representation. To further e… ▽ More

    Submitted 11 June, 2026; v1 submitted 10 June, 2026; originally announced June 2026.

    Comments: Accepted to Interspeech 2026, Github: https://github.com/sanghyang00/ur-bert

  33. arXiv:2606.08465  [pdf, ps, other

    cs.FL cs.PF cs.PL cs.SE

    An Empirical Comparison of General Context-Free Parsers

    Authors: Huan Vo, Danushka Liyanage, Hong Jin Kang, Sasha Rubin, Rahul Gopinath

    Abstract: Parsing underpins a vast range of software engineering tasks, from compilers and static analyzers to language servers and fuzz testing tools. Yet most parsers deployed in practice are deterministic (LL or LR), forcing developers not only to contort their grammars to fit the parser, but to simplify the very languages they design sacrificing expressiveness for the sake of parseability. General conte… ▽ More

    Submitted 7 June, 2026; originally announced June 2026.

    MSC Class: 68Q42; 68N20 ACM Class: F.4.2; D.3.4

  34. arXiv:2606.06447  [pdf, ps, other

    cs.CL cs.LG

    Latent Reasoning with Normalizing Flows

    Authors: Guancheng Tu, Xiangjun Fu, Suhao Yu, Yao Tang, Haoqiang Kang, Lianhui Qin, Yizhe Zhang, Jiatao Gu

    Abstract: Large language models often improve reasoning by generating explicit chain-of-thought (CoT), demonstrating the importance of intermediate computation. However, textual CoT forces this computation through a discrete, serial, and communication-oriented token stream: each reasoning step must be verbalized before the model can proceed, even when the underlying update is semantic, uncertain, or only pa… ▽ More

    Submitted 4 June, 2026; originally announced June 2026.

  35. arXiv:2606.05179  [pdf, ps, other

    cs.CL

    Efficient Punctuation Restoration via Weighted Lookahead Scoring Method for Streaming ASR Systems

    Authors: Sungmook Woo, Hyungu Kang, Chanwoo Kim

    Abstract: Punctuation restoration improves ASR (Automatic Speech Recognition) readability. However streaming ASR requires online decisions with limited future context. In streaming ASR, the system predicts punctuation incrementally, which makes generation-based approaches prone to latency and alignment failures under boundary-wise evaluation. This paper proposes a non-autoregressive scoring method (no free-… ▽ More

    Submitted 17 April, 2026; originally announced June 2026.

    Comments: Accepted for presentation at The International Joint Conference on Neural Networks (IJCNN) 2026

  36. arXiv:2606.03180  [pdf, ps, other

    cs.CV cs.CL cs.LG

    GLINT: Sparsely Gated Vision-Language Alignment for Fine-Grained Radiology Representations

    Authors: Jonggwon Park, Seongeun Lee, Junhyun Park, Hannah Yun, Hyunwoong Kim, Sohyun Jeong, Hyewon Kang, Byungmu Yoon, Kyoyun Choi

    Abstract: Vision-language models (VLMs) for radiology have emerged as a scalable paradigm by leveraging image-report pairs naturally produced in clinical workflows. However, this pairing reveals a mismatch in scale: each finding occupies only a small region of the image, yet supervision is provided only at the global image-report level. This poses a central challenge: prior approaches spread weight densely… ▽ More

    Submitted 2 June, 2026; originally announced June 2026.

  37. arXiv:2606.00866  [pdf, ps, other

    cs.OS

    Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI

    Authors: Tian Xia, Hanchen Li, Zhifei Li, Xiaokun Chen, Hao Kang, Yifan Qiao, Yi Xu, Ion Stoica

    Abstract: Modern LLM serving systems increasingly host agentic workloads, whose sessions issue tens of model invocations interleaved with tool calls, accumulating KV cache that can be reused across steps. As requests' total KV cache size easily exceeds GPU HBM capacity, researchers offload them to CPU DRAM. However, tool-call durations span orders of magnitude, and the cost of transferring KV cache between… ▽ More

    Submitted 30 May, 2026; originally announced June 2026.

  38. arXiv:2605.31463  [pdf, ps, other

    cs.LG cs.AI cs.CL cs.DC

    PithTrain: A Compact and Agent-Native MoE Training System

    Authors: Ruihang Lai, Hao Kang, Haozhan Tang, Akaash R. Parthasarathy, Zichun Yu, Junru Shao, Todd C. Mowry, Chenyan Xiong, Tianqi Chen

    Abstract: Mixture-of-Experts (MoE) has become the dominant architecture for frontier language models. To meet this demand, production frameworks have built optimized MoE training stacks over years of engineering effort. Yet evolving these stacks for new architectures and system optimizations remains expensive. With the rise of AI coding agents, they could automate parts of training-framework development and… ▽ More

    Submitted 29 May, 2026; originally announced May 2026.

  39. arXiv:2605.30989  [pdf, ps, other

    cs.RO

    A study on a Real-Time VR-Based Teleoperation Framework for Manipulator in Dynamic Environment

    Authors: InGyu Choi, GeonYeong Go, SunWoo Ahn, HyoJae Kang, Min-Sung Kang

    Abstract: Robot teleoperation enables safe, non-contact task execution in hazardous environments where direct human access is difficult, and its application has expanded with recent VR technologies. Many VR teleoperation studies, however, have primarily served as data-collection tools for robot imitation learning, so they often do not explicitly address dynamic obstacles, workspace changes, or collision ris… ▽ More

    Submitted 29 May, 2026; originally announced May 2026.

    Comments: This manuscript has been submitted for possible publication

  40. arXiv:2605.30843  [pdf, ps, other

    cs.LG econ.EM

    A Lecture Note on Offline RL and IRL, Part II: Foundations of Inverse Reinforcement Learning and Dynamic Discrete Choice Models

    Authors: Enoch Hyunwook Kang

    Abstract: In the forward reinforcement-learning problem, the reward is fixed and known; the learner is asked to find a good policy or value function. Here we turn the question around. Given offline data generated by an expert, can we recover the reward the expert was optimizing? This is the inverse reinforcement learning problem, and remarkably, two communities, structural econometricians studying dynamic d… ▽ More

    Submitted 29 May, 2026; originally announced May 2026.

  41. arXiv:2605.30508  [pdf, ps, other

    cs.RO

    ARISTO Hand: Sensing-Driven Distal Hyperextension for Fine-Grained Manipulation

    Authors: Aaron Kim, Dong Ho Kang, Mark Helwig, Mingyo Seo, Kazuto Yokoyama, Tetsuya Narita, Luis Sentis

    Abstract: Manipulating thin objects requires precise contact geometry and reliable force perception, yet many anthropomorphic robotic hands lack the mechanical and sensing capabilities needed for such interactions. We present the ARISTO Hand, a tendon-driven robotic hand that integrates active distal hyperextension with a hybrid fingertip-sensing architecture that combines a rigid, nail-mounted force-torque… ▽ More

    Submitted 28 May, 2026; originally announced May 2026.

  42. arXiv:2605.30250  [pdf, ps, other

    cs.CV cs.GR

    Ambient-robust Inverse Rendering using Active RGB-NIR Imaging

    Authors: Hoon-Gyu Chung, Jinnyeong Kim, Hyunwoo Kang, Seung-Hwan Baek

    Abstract: Inverse rendering aims to reconstruct geometry and reflectance of objects from images. Despite recent progress, existing methods often produces inaccurate reconstructions that are sensitive to ambient illumination conditions. Here we introduce an ambient-robust inverse rendering method enabled by active RGB-NIR imaging. Our key insight is to leverage near-infrared (NIR) flash illumination-impercep… ▽ More

    Submitted 28 May, 2026; originally announced May 2026.

    Comments: 11 pages

  43. arXiv:2605.27006  [pdf, ps, other

    cs.LG cond-mat.dis-nn stat.ML

    Sampling Data with Chains of Forward-Backward Diffusion Steps

    Authors: Hyunmo Kang, Noam Itzhak Levi, Corinna Elena Wegner, Daniel J. Korchinski, Matthieu Wyart

    Abstract: Sampling from learned high-dimensional distributions is a foundational computational problem. We introduce U-turn chains: Markov chains obtained by iterating short forward-backward steps of a diffusion model, in which each step proposes a move that remains on the learned data manifold and, paired with a Metropolis-Hastings correction, samples from energy-modified targets. For synthetic languages,… ▽ More

    Submitted 26 May, 2026; originally announced May 2026.

  44. arXiv:2605.26552  [pdf, ps, other

    cs.LG cs.AI

    Aligning Few-Step Generative Models by Amortizing Sample-based Variational Inference

    Authors: Jaewoo Lee, Hyeongyu Kang, Dohyun Kim, Kyuil Sim, Woocheol Shin, Minsu Kim, Taeyoung Yun, Jeongjae Lee, Sanghyeok Choi, Tabitha Edith Lee, Jong Chul Ye, Jinkyoo Park

    Abstract: Aligning a few-step generative model is challenging, since existing alignment frameworks typically rely on restrictive assumptions: a tractable likelihood, a specific ODE/SDE solver, or a particular model family. We introduce FAV, Few-step Generative Models Alignment via Sample-based Variational Inference, a general alignment framework that requires only sample access to the generator and the refe… ▽ More

    Submitted 27 May, 2026; v1 submitted 26 May, 2026; originally announced May 2026.

    Comments: Under review

  45. arXiv:2605.25740  [pdf, ps, other

    cs.LG

    Latent Representation Alignment for Offline Goal-Conditioned Reinforcement Learning

    Authors: Hyungkyu Kang, Byeongchan Kim, Min-hwan Oh

    Abstract: Offline goal-conditioned reinforcement learning (GCRL) provides a practical framework for obtaining goal-reaching policies from fixed datasets. However, learning a reliable goal-conditioned value function in long-horizon tasks remains challenging. In this paper, we identify erroneous generalization in goal-conditioned value functions as a fundamental bottleneck, and demonstrate that appropriate in… ▽ More

    Submitted 25 May, 2026; originally announced May 2026.

    Comments: Accepted in ICML 2026

  46. arXiv:2605.23912  [pdf, ps, other

    cs.CL cs.AI cs.SD

    Raon-Speech Technical Report

    Authors: Beomsoo Kim, Changho Choi, Dohyun Kim, Dongki Lee, Ethan Ewer, Eunchong Kim, Gyeongman Kim, Haechan Kim, Hyeonghwan Kim, Inkyu Park, Jihun Yun, Jihwan Moon, Jiyun Kim, Joonghyun Bae, Junhyuck Kim, Minkyu Kim, Sehun Lee, Seungjun Chung, Sungwoo Cho, Dongmin Park, Dongwon Kim, Hara Kang, Jonghyun Lee, Keon Lee, Kangwook Lee , et al. (1 additional authors not shown)

    Abstract: We present Raon-Speech, a top-performing 9B-parameter speech language model (SpeechLM) for English and Korean speech understanding, answering, and generation, and Raon-SpeechChat, a high-performing full-duplex extension for natural real-time conversation. Raon-Speech successfully transforms a pre-trained LLM into a SpeechLM that both understands and generates speech while preserving strong text ca… ▽ More

    Submitted 8 April, 2026; originally announced May 2026.

  47. arXiv:2605.23655  [pdf, ps, other

    cs.CV cs.AI cs.LG cs.MM

    CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception

    Authors: Liupeng Li, Haoqian Kang, Zhenyu Lu, Jinpeng Wang, Bin Chen, Ke Chen, Yaowei Wang

    Abstract: High-resolution (HR) image perception presents a key bottleneck for multimodal large language models (MLLMs). While visual search offers a promising solution, existing methods struggle with the trade-off between coverage and efficiency. Visual expert-assisted search is efficient but prone to blind spots when proposals fail, whereas scan-based search guarantees coverage at the cost of computational… ▽ More

    Submitted 22 May, 2026; originally announced May 2026.

    Comments: Accepted by ICML 2026. 22 pages, 12 figures, 7 tables

  48. arXiv:2605.22658  [pdf, ps, other

    cs.CV cs.LG cs.MM eess.IV

    SegCompass: Exploring Interpretable Alignment with Sparse Autoencoders for Enhanced Reasoning Segmentation

    Authors: Zhenyu Lu, Liupeng Li, Jinpeng Wang, Haoqian Kang, Yan Feng, Ke Chen, Yaowei Wang

    Abstract: While large language models provide strong compositional reasoning, existing reasoning segmentation pipelines fail to transparently connect this reasoning to visual perception. Current methods, such as latent query alignment, are end-to-end yet opaque "black boxes". Conversely, textual localization readout is merely readable, not truly interpretable, often functioning as an unconstrained post-hoc… ▽ More

    Submitted 21 May, 2026; originally announced May 2026.

    Comments: Accepted by CVPR 2026. 15 pages, 9 figures, 6 tables

  49. arXiv:2605.20872  [pdf, ps, other

    cs.LG cs.AI cs.GR

    CAdam: Context-Adaptive Moment Estimation for 3D Gaussian Densification in Generative Distillation

    Authors: SeungJeh Chung, Geonho Park, Misong Kim, HyeongYeop Kang

    Abstract: Adaptive densification is the engine of 3D Gaussian Splatting (3DGS). However, when transposed to the optimization-based Generative Distillation paradigm, this reconstruction-native mechanism reveals fundamental limitations, resulting in inefficient representations cluttered with redundant primitives. We diagnose this failure as a Densification Dilemma stemming from the stochastic nature of genera… ▽ More

    Submitted 20 May, 2026; originally announced May 2026.

    Comments: Accepted to SIGGRAPH 2026 Conference Papers. 12 pages, 8 figures

  50. arXiv:2605.20865  [pdf, ps, other

    cs.LG cs.AI

    Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards

    Authors: Deokgyu Yoon, Hyungkyu Kang, Joongkyu Lee, Byeongchan Kim, Gyungin Shin, Sungrae Park, Min-hwan Oh

    Abstract: Reinforcement learning with verifiable rewards (RLVR) plays a pivotal role in improving the reasoning ability of large language models. However, widely used PPO surrogate objectives are fundamentally local, as they rely on a local approximation of the exact policy gradient objective. While this approximation improves stability by reducing the variance induced by importance sampling, it also introd… ▽ More

    Submitted 20 May, 2026; originally announced May 2026.