Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 713 results for author: Hwang, J

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.19861  [pdf, ps, other

    cs.AI cs.CL cs.LG

    PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents

    Authors: Seongjae Kang, Taehyung Yu, Sung Ju Hwang

    Abstract: Customer-service LLM agents must follow organizational policy when acting on a user's behalf. Compliance failures arise from either forbidden actions, such as granting an ineligible change, or omitted procedural requirements, such as identification or confirmation. Runtime safeguards can intervene on risky actions, but action-local checks do not guide an agent through a multi-step procedure. Workf… ▽ More

    Submitted 20 August, 2026; originally announced August 2026.

    Comments: 26 pages, 15 figures, including appendices

    ACM Class: I.2.11; I.2.7

  2. arXiv:2608.18009  [pdf, ps, other

    cs.CV

    Memory Tree Guided Key Frame Querying for Efficient 3D Question Answering

    Authors: Hsiang-Wei Huang, Fu-Chen Chen, Li-Wu Tsao, Cheng-Han Lee, Che-Chun Su, Lu Xia, Ronghui Peng, Jenq-Neng Hwang, Min Sun, Cheng-Hao Kuo

    Abstract: Answering questions accurately and efficiently in embodied scenarios presents significant challenges due to limited computational and memory resources for Vision Language Model (VLM) inference. Existing methods adopt visual search key frame retrieval method to select critical question-related key frames for VLM input. However, visual search methods are inefficient because they require visual searc… ▽ More

    Submitted 18 August, 2026; originally announced August 2026.

    Comments: ECCV 2026

  3. arXiv:2608.11546  [pdf, ps, other

    cs.CV

    Through Van Gogh's Eyes: Global Style Transfer with Diffusion Model

    Authors: Jeongha Lee, Yujin Kim, Ghazanfar Ali, Suhyun Kim, Jae-In Hwang

    Abstract: Artistic image synthesis aims to recreate the expressive visual identity of a target artist, yet existing methods often fail to capture an artist's global style. Conventional style transfer methods transfer the style of one or a few reference artworks to a content image in a One-to-One manner, making them effective for artwork-level stylization but limited in representing the broader stylistic dis… ▽ More

    Submitted 12 August, 2026; v1 submitted 11 August, 2026; originally announced August 2026.

    Comments: Published at ECCV 2026

  4. arXiv:2608.11038  [pdf, ps, other

    cs.DS cs.DM math.CO

    A 5/4 bound for graphic $s$-$t$ path TSP on subcubic graphs

    Authors: Junho Hwang

    Abstract: We study the graphic $s$-$t$ path TSP on subcubic graphs (maximum degree 3): given two vertices $s,t$, find a shortest walk from $s$ to $t$ that visits every vertex. Our main result is that the optimal $5/4$ coefficient is attained for every terminal pair -- including the difficult case where deleting both $s$ and $t$ disconnects the graph. Concretely, every pair of distinct vertices $s,t$ in a si… ▽ More

    Submitted 11 August, 2026; originally announced August 2026.

    Comments: 12 pages, 2 figures. To appear in the proceedings of WAOA 2026 (Springer LNCS)

    MSC Class: 68W25; 05C38; 05C85; 90C27 ACM Class: F.2.2; G.2.2

  5. arXiv:2608.07169  [pdf, ps, other

    cs.AI cs.LG

    Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory

    Authors: Taeil Kim, Kangsan Kim, Sung Ju Hwang

    Abstract: Memory systems have shown promise for improving agent performance, but their potential remains largely unexplored for small language models, which struggle to generate sufficient successful trajectories on their own. We propose Agent Memory Distillation (AMD), a training-free framework that transfers structured knowledge from a large teacher agent to a small student agent through hierarchical memo… ▽ More

    Submitted 7 August, 2026; originally announced August 2026.

    Comments: Under review

  6. arXiv:2608.05802  [pdf, ps, other

    cs.CL cs.LG

    On-Policy Delta Distillation for Multilingual Math Reasoning

    Authors: Byeongho Heo, Jaehui Hwang, Sangdoo Yun, Dongyoon Han

    Abstract: On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Policy Delta Distillation (OPD$^2$), for mathematical reasoning in English, Korean, and Japanese. OPD$^2$ improves OPD by using the probability gap between a post-trained… ▽ More

    Submitted 6 August, 2026; originally announced August 2026.

    Comments: 9 pages, 3 figures, 10 tables

  7. arXiv:2608.04505  [pdf, ps, other

    cs.CL

    K-EXAONE 2.0 Technical Report

    Authors: Eunbi Choi, Kibong Choi, Sehyun Chun, Seokhee Hong, Junwon Hwang, Hyojin Jeon, Ahra Jo, Hyunjik Jo, Yeonsik Jo, Minhyeok Jung, Doyoung Kim, Heegyu Kim, Joonkee Kim, Seonghwan Kim, Soyeon Kim, Sunkyoung Kim, Yireun Kim, Yongil Kim, Byungoh Ko, Changhun Lee, Dohaeng Lee, Haeju Lee, Jinsik Lee, Kyungmin Lee, Minwoo Lee , et al. (52 additional authors not shown)

    Abstract: This technical report presents K-EXAONE 2.0, an open-weight multilingual foundation model developed by LG AI Research as a step in our effort toward global frontier-scale foundation models. Rather than training from scratch, we upcycle K-EXAONE and expand its architecture, yielding a Mixture-of-Experts (MoE) model with 750B total parameters and approximately 37B activated per token---more than thr… ▽ More

    Submitted 5 August, 2026; originally announced August 2026.

  8. arXiv:2607.29059  [pdf, ps, other

    cs.CV

    Learning from Adversity: Semantic-Aware Mask Refinement through Adversarial Perturbation

    Authors: Beomyoung Kim, Sung Ju Hwang

    Abstract: Despite significant advances in image segmentation, even state-of-the-art models produce masks with imperfect boundaries, semantic inconsistencies, and structural errors. Mask refinement addresses these limitations, yet current approaches rely on simplistic synthetic noise that fails to capture the complex error patterns of real segmentation models. We introduce Phoenix, a novel framework that lev… ▽ More

    Submitted 31 July, 2026; originally announced July 2026.

    Comments: ECCV 2026

  9. Convolutional Neural Shading for High-Quality 3D Reconstruction from Multi-View Images

    Authors: Juheon Hwang, Taewan Kim, Heeseok Oh, Jiwoo Kang

    Abstract: We propose a convolutional neural shading (CNS), a novel pipeline to reconstruct high-quality 3D shapes from multi-view images. Several recent studies have used neural radiance fields and other neural differentiable rendering methods to understand 3D geometry. However, these approaches rely on single-point geometric information, such as positions and normals of the surface, leading to a lack of de… ▽ More

    Submitted 30 July, 2026; originally announced July 2026.

    Journal ref: Multimedia Systems, vol. 31, no. 4, pp. 296, July 2025

  10. Collaborative Feature Aggregation for Face Super-Resolution and Robust Re-Identification

    Authors: Juheon Hwang, Taewan Kim, Jiwoo Kang

    Abstract: We propose a novel collaborative approach for face super-resolution (SR) and robust person re-identification from sequential or multi-view facial images. Traditional SR methods often suffer from blurring and distortion in faces recovered from poor-quality images due to low resolution. Image- and video-based facial SR methods using facial landmarks or segmentation also have similar challenges. To o… ▽ More

    Submitted 30 July, 2026; originally announced July 2026.

    Journal ref: Multimedia Systems, vol. 31, no. 5, pp. 341, Aug. 2025

  11. arXiv:2607.26862  [pdf, ps, other

    cs.LG cs.AI

    ReCo: Reweighting GRPO Against Distributional Concentration

    Authors: Junoh Park, Junseo Hwang, Wonguk Cho, Taesup Kim

    Abstract: Group Relative Policy Optimization (GRPO) has become a standard reinforcement learning method for post-training language models. Recent work shows that GRPO can reduce the base model's reasoning capacity and underperform it in Pass@k when k is large, indicating reduced coverage of reasoning paths. We find that this reduction is associated with GRPO concentrating on responses that the base model al… ▽ More

    Submitted 29 July, 2026; originally announced July 2026.

  12. arXiv:2607.24555  [pdf, ps, other

    cs.LG cs.AI

    LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding

    Authors: Junsung Hwang

    Abstract: Serving large language models at long context is bottlenecked by the key-value (KV) cache, which is read in full at every decode step. Attention keys are locally low-rank though globally high-rank: shared low-rank bases discard page-specific directions that a page's own compact basis retains. LOCKS gives every page its own spectral summary (resident, about a tenth the cache's size), reconstructs w… ▽ More

    Submitted 27 July, 2026; originally announced July 2026.

  13. arXiv:2607.15161  [pdf, ps, other

    cs.LG cs.CL

    On-Policy Delta Distillation

    Authors: Byeongho Heo, Jaehui Hwang, Sangdoo Yun, Dongyoon Han

    Abstract: On-policy distillation is an alternative post-training method in reinforcement learning that alleviates the constraints imposed by reward models by providing token-level supervision from a teacher model. Although on-policy distillation has been studied and applied across various settings, its fundamental design remains underexplored. In this paper, we introduce a new distillation reward, termed th… ▽ More

    Submitted 16 July, 2026; originally announced July 2026.

    Comments: 19 pages, 4 figures, 12 tables

  14. arXiv:2607.11081  [pdf, ps, other

    cs.CV cs.AI

    Controlling Motion Transfer in Diffusion Transformers via Attention Heads

    Authors: Sunyoung Jung, Jiwoo Park, Yoonseok Choi, Kyobin Choo, Ming-Hsuan Yang, Seong Jae Hwang

    Abstract: Diffusion Transformers (DiTs) have advanced video generation with high-quality, temporally coherent results. However, extending them to motion transfer, which requires following reference motion while aligning with a target prompt, remains challenging due to limited understanding of motion and structure representations within DiTs. We analyze video DiTs at the attention-head level and identify dis… ▽ More

    Submitted 13 July, 2026; originally announced July 2026.

    Comments: Accepted to ECCV 2026, Project page: https://sunyj-hxppy.github.io/halo/

  15. Semantic Hardness Is Not Visual Hardness: Sign-Aware Hard Negative Mining for Sign Language Retrieval

    Authors: Junmyeong Lee, Chan Hur, ChangSu Choi, Sukmin Cho, Fitsum Gaim, Eui Jun Hwang, Hoyun Song, KyungTae Lim

    Abstract: Sign Language Retrieval (SLRet) enables efficient access to sign language content but remains fragile in fine-grained scenarios where visually similar signs must be distinguished. We show that this limitation does not stem from model capacity, but from ineffective hard negative supervision. Specifically, we formulate fine-grained retrieval failures as a negative distribution mismatch: semantically… ▽ More

    Submitted 10 July, 2026; originally announced July 2026.

    Comments: Accepted to ACL 2026 main

    Journal ref: Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2026, pages 28262-28277

  16. arXiv:2607.04127  [pdf, ps, other

    cs.CV cs.RO

    Real-Time LiDAR Gaussian Splatting SLAM

    Authors: Seungjun Tak, Yewon Jeon, Jaeik Hwang, SukMin Hwang, Seongbo Ha, Hyeonwoo Yu

    Abstract: We present a real-time LiDAR-based framework for Gaussian Splatting SLAM that tightly couples fast G-ICP registration with spherical rasterization-based dense mapping for large-scale sequences. Leveraging LiDAR geometry rather than appearance, we reuse tracking-estimated local covariances to initialize Gaussians with range-aware scales and to derive surface normals for geometry-aware map optimizat… ▽ More

    Submitted 5 July, 2026; originally announced July 2026.

    Comments: 18 pages, 5 figures

    MSC Class: 68T40

  17. arXiv:2607.03038  [pdf, ps, other

    cs.CV

    OmniDS: Dual-Stream Context Fusion for Omnidirectional Depth from Fisheye Cameras

    Authors: Chaesong Park, Jihyeon Hwang, Muyeol Sung, Jongwoo Lim

    Abstract: Omnidirectional depth estimation from multi-fisheye camera rigs is complicated by visibility conflicts: wide baselines cause different cameras to observe different portions, or even different faces, of the same object, so aggregating their features into a unified equirectangular (ERP) representation under fixed projection produces ambiguous matching evidence near occlusion boundaries and thin stru… ▽ More

    Submitted 3 July, 2026; originally announced July 2026.

    Comments: 15 pages, 3 figures, 3 tables

  18. arXiv:2607.01869  [pdf, ps, other

    cs.CV

    QWERTY: Training-Free Motion Control via Query-Warped Video Diffusion Transformers

    Authors: Kyobin Choo, Youngmin Kim, Hyunkyung Han, Geunrip Park, Chanyoung Kim, Sunyoung Jung, Seong Jae Hwang

    Abstract: Video diffusion transformers (DiTs) generate high-fidelity and temporally coherent videos, yet motion control remains implicit, primarily relying on text prompts. As a result, achieving desired motion often requires extensive prompt engineering and repeated resampling. While fine-tuning models with additional spatial prompts (e.g., bounding boxes or point trajectories) enables explicit control, it… ▽ More

    Submitted 2 July, 2026; originally announced July 2026.

    Comments: 37 pages, 18 figures, accepted at the European Conference on Computer Vision (ECCV) 2026

    ACM Class: I.4.9

  19. arXiv:2606.29225  [pdf, ps, other

    cs.AI cs.CL

    PolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM Agents

    Authors: Seongjae Kang, Taehyung Yu, Sung Ju Hwang

    Abstract: LLM agents handle user requests on behalf of organizations through tool calls and must follow the company policies stated in their system prompts. Prior work approaches this as a safeguarding problem -- external checks that block non-compliant agent actions. We argue that policy adherence is a broader problem: real workflows unfold across many turns, require explicit user confirmation and prerequi… ▽ More

    Submitted 28 June, 2026; originally announced June 2026.

    Comments: 20 pages, 8 figures

  20. arXiv:2606.24950  [pdf, ps, other

    cs.LG cs.AI

    MacroLens: A Multi-Task Benchmark for Contextual Financial Reasoning under Macroeconomic Scenarios

    Authors: Patara Trirat, Jin Myung Kwak, Jay Heo, Heejun Lee, Sung Ju Hwang

    Abstract: Financial decision-making is contextual: forecasting prices, valuing companies, and assessing event exposure weigh price history, accounting fundamentals, macroeconomic regime, and contemporaneous text. A benchmark over these four signals is hard to build because finance violates four assumptions of time-series evaluation: text must be gated by its publication date to prevent look-ahead, quarterly… ▽ More

    Submitted 23 June, 2026; originally announced June 2026.

    Comments: 25 pages, 3 figures

  21. arXiv:2606.21177  [pdf, ps, other

    eess.IV cs.AI cs.CV physics.med-ph

    Anatomically Consistent TMJ Disc Segmentation via Semantic Anchoring and Clinical Priors

    Authors: Dayun Ju, Chanyoung Kim, Sunyoung Jung, Hyo-Jung Jung, Chena Lee, Younjung Park, Seong Jae Hwang

    Abstract: Segmenting the temporomandibular joint (TMJ) disc from MRI is essential for accurate diagnosis of internal derangement, yet it remains unreliable in practice due to its small size, low contrast, and morphological variability. Existing methods, primarily adapted from general segmentation architectures, often produce fragmented or anatomically inconsistent masks, leading to unstable measurements of… ▽ More

    Submitted 19 June, 2026; originally announced June 2026.

    Comments: 10 pages, 3 figures

  22. arXiv:2606.21156  [pdf, ps, other

    cs.CV cs.AI

    Contrastive and Adaptive Multi-modal Masked Autoencoder for Spatial Transcriptomics

    Authors: Joohyeok Kim, Taejin Jeong, Jinyeong Kim, Seong Jae Hwang

    Abstract: The high cost of spatial transcriptomics (ST) has driven extensive studies into predicting gene expression directly from H&E histology images. However, this prediction task faces an inherent limitation, as tissue morphology alone provides insufficient information to fully resolve underlying gene expression. To address this limitation, a recent study leverages partial gene expression to guide the p… ▽ More

    Submitted 19 June, 2026; originally announced June 2026.

  23. arXiv:2606.20812  [pdf, ps, other

    cs.LG cs.AI

    B[FM]$^2$: Brain Foundation Model via Flow Matching with SplitUNet

    Authors: Jaedong Hwang, Kathleen Zhang, Wei Dai, Konstantinos Kontras, Maarten Vanmarcke, Maarten De Vos, Ila Fiete, Paul Pu Liang

    Abstract: EEG foundation models can learn generalizable representations from large-scale EEG corpora to enable single-backbone transfer across diverse clinical and brain-computer interface tasks. Existing models typically discretize the continuous multi-channel EEG waveform into patches or codebook tokens and train a transformer with masked self-supervision. Recognizing that this discretization fragments co… ▽ More

    Submitted 18 June, 2026; originally announced June 2026.

  24. arXiv:2606.20641  [pdf, ps, other

    cs.RO cs.AI cs.LG

    MAGNIFIED: RL Fine-tuning of Multimodal Large Language Models for Motion Planning

    Authors: Letian Chen, Yiren Lu, Justin Fu, Yichen Xie, Runsheng Xu, Jyh-Jing Hwang, Ben Sapp, Drago Anguelov

    Abstract: Multi-modal Large Language Models (MLLMs) have demonstrated remarkable capabilities in semantic understanding and common sense reasoning, making them promising candidates for solving planning problems in autonomous driving. However, the next-token text prediction objectives traditionally used in pre-training and supervised fine-tuning (SFT) of MLLMs may fall short of fulfilling the planning object… ▽ More

    Submitted 2 June, 2026; originally announced June 2026.

    Journal ref: ICRA 2026

  25. arXiv:2606.18620  [pdf, ps, other

    cs.CL cs.AI

    BCL: Bayesian In-Context Learning Framework for Information Extraction

    Authors: Haoliang Liu, Chengkun Cai, Xu Zhao, Han Zhu, Shizhou Huang, Xinglin Zhang, Tao Chen, Jenq-Neng Hwang, Zhang Huaping, Lei Li

    Abstract: Existing information extraction (IE) tasks increasingly adopt in-context learning (ICL) with large language models. However, current approaches either show inconsistent performance across model scales or lack systematic optimization and generalizability. Building on this, we propose BCL (Bayesian In-Context Learning Framework for Information Extraction), the first optimization framework that uses… ▽ More

    Submitted 16 June, 2026; originally announced June 2026.

    Comments: ACL 2026 Findings

  26. arXiv:2606.18066  [pdf, ps, other

    cs.LG

    NoiseTilt: Noise-Tilted Reverse Kernels for Diffusion Reward Alignment

    Authors: Jisung Hwang, Yunhong Min, Jaihoon Kim, I-Chao Shen, Minhyuk Sung

    Abstract: We introduce the Noise-Tilted Reverse Kernel (NTRK), a reward-guided diffusion sampler that injects reward gradients through the noise term, leaving the pretrained reverse kernel unchanged and requiring only a single sample per step. Reward-guided sampling at inference time has greatly expanded the versatility of pretrained diffusion models. Yet existing methods face a trade-off. Gradient-based gu… ▽ More

    Submitted 27 June, 2026; v1 submitted 16 June, 2026; originally announced June 2026.

    Comments: ECCV 2026

  27. arXiv:2606.14008  [pdf

    cs.CR cs.NI cs.PF eess.SY

    Pseudonym Scheme Based on Hybrid Certificates for Security Credential Management System in Vehicular Communications

    Authors: Abel C. H. Chen, F. J. Hwang, Yu-Chih Wei, Chin-Chen Chang, Bon-Yeh Lin

    Abstract: In recent years, the Institute of Electrical and Electronics Engineers (IEEE) and the European Telecommunications Standards Institute (ETSI) have developed a series of security communication standards for vehicular communications. These standards include mechanisms such as the Security Credential Management System (SCMS) and Butterfly Key Expansion (BKE) to protect vehicle privacy. However, these… ▽ More

    Submitted 11 June, 2026; originally announced June 2026.

    Journal ref: IEEE Canadian Journal of Electrical and Computer Engineering (2026)

  28. arXiv:2606.13655  [pdf, ps, other

    cs.CV cs.GR

    Flex4DHuman: Flexible Multi-view Video Diffusion for 4D Human Reconstruction

    Authors: Jen-Hao Cheng, Yipeng Wang, Hao Zhang, Gengshan Yang, Jenq-Neng Hwang

    Abstract: We present Flex4DHuman, a multi-view video diffusion model that transforms a monocular or sparse multi-view video of a dynamic subject into synchronized dense multi-view videos using only relative camera-pose conditioning. Unlike prior human-centric methods that rely on skeletons, depth maps, normals, or rendered target-view geometry, Flex4DHuman requires no explicit geometry priors and instead co… ▽ More

    Submitted 13 June, 2026; v1 submitted 11 June, 2026; originally announced June 2026.

    Comments: Project Page: https://andy-cheng.github.io/Flex4DHuman/

  29. arXiv:2606.13177  [pdf, ps, other

    cs.CL cs.AI cs.LG

    MemRefine: LLM-Guided Compression for Long-Term Agent Memory

    Authors: Minjae Kim, Jinheon Baek, Soyeong Jeong, Sung Ju Hwang

    Abstract: Large language model (LLM) agents are increasingly expected to operate over long-term interactions, where information from past dialogues must be preserved and recalled to support future tasks. However, as interactions accumulate, the memory store grows without bound and fills with redundant entries that inflate storage cost and degrade retrieval by crowding out the most useful evidence. Furthermo… ▽ More

    Submitted 11 June, 2026; originally announced June 2026.

  30. arXiv:2606.10423  [pdf, ps, other

    cs.CL

    WebChallenger: A Reliable and Efficient Generalist Web Agent

    Authors: Jayoo Hwang, Xiaowen Zhang, Vedant Padwal

    Abstract: Autonomous web navigation remains challenging for LLM agents, and the strongest generalist systems rely on proprietary reasoning models whose inference cost is prohibitive for the repetitive tasks where such agents would be most useful. We argue this gap stems not from insufficient model capability but from agent architectures that fail to replicate three human cognitive advantages: selective atte… ▽ More

    Submitted 9 June, 2026; originally announced June 2026.

  31. arXiv:2606.06361  [pdf, ps, other

    cs.CV

    Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them

    Authors: Woojung Han, Seil Kang, Youngjun Jun, Min-Hung Chen, Fu-En Yang, Seong Jae Hwang

    Abstract: Image-to-Video diffusion models leverage input images to generate visually stunning content, yet frequently produce motion that violates physical laws. We reveal a surprising finding: a 2-step generation often exhibits better physical consistency than a 50-step output from the same model. Through spectral analysis, we trace this to phase erosion during denoising; the phase degrades significantly (… ▽ More

    Submitted 17 June, 2026; v1 submitted 4 June, 2026; originally announced June 2026.

    Comments: ICML 2026

  32. arXiv:2606.05896  [pdf, ps, other

    cs.CV

    Resonant Minds: Closed-Loop Social Avatars with Theory of Mind

    Authors: Jianxu Shangguan, Jing Xu, Hang Ye, Xiaoxuan Ma, Yizhou Wang, Jenq-Neng Hwang, Wentao Zhu

    Abstract: Creating lifelike digital humans with genuine social intelligence requires unifying cognitive reasoning and multimodal generation within a coherent framework. Current approaches treat these as separate tasks: Large Language Models excel at dialogue but lack embodied expression, while diffusion-based talking head models achieve visual fidelity but ignore social cognition. To bridge this gap, we pro… ▽ More

    Submitted 22 June, 2026; v1 submitted 4 June, 2026; originally announced June 2026.

    Comments: Project page: https://resonantminds.github.io/

  33. arXiv:2606.04743  [pdf, ps, other

    cs.CL cs.AI cs.LG

    TIDE: Proactive Multi-Problem Discovery via Template-Guided Iteration

    Authors: Soyeong Jeong, Jinheon Baek, Minki Kang, Sung Ju Hwang

    Abstract: Agents are widely deployed as assistants over documents, tools, and code. However, they typically act only on explicit user requests, which surface only the problems the user has noticed, while many other important problems coexist, hidden in plain sight, within the broader user context, with their total number unknown in advance. We frame this as the task of discovering multiple hidden problems f… ▽ More

    Submitted 14 August, 2026; v1 submitted 3 June, 2026; originally announced June 2026.

  34. arXiv:2606.01694  [pdf, ps, other

    cs.CV cs.AI cs.LG cs.MM

    Understanding Identity Continuity in Thermal Video through Scene-Level Consistency

    Authors: Wei-Chieh Sun, Gyungmin Ko, Heejae Kwon, Hsiang-Wei Huang, Jenq-Neng Hwang

    Abstract: Thermal pedestrian MOT remains challenging because weak appearance cues and frequent detection interruptions cause severe trajectory fragmentation. We study whether lightweight post-processing can recover identity continuity without relying on heavy re-identification models or complex online association. Starting from a YOLOv8 and SORT baseline, we add a modular identity-repair backend consisting… ▽ More

    Submitted 1 June, 2026; originally announced June 2026.

    Comments: Accepted to CVPR 2026 Workshop on SVC. Published in CVPR Workshops proceedings

    ACM Class: I.4.8; I.4.9

    Journal ref: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2026, pp. 1411-1419

  35. arXiv:2606.00741  [pdf, ps, other

    cs.LG cs.AI stat.ML

    Quantum Tunneling-Aware Machine Learning: Physics-Derived Noise Models for Robust Deployment

    Authors: Uiwon Hwang, Jaeho Hwang

    Abstract: Transistor scaling is approaching a quantum-mechanical limit, as thin gate oxides induce electron leakage through quantum tunneling. Unlike conventional digital systems, AI inference can tolerate such errors provided their structure is modeled correctly. In this paper, we introduce quantum tunneling-aware machine learning (QTAML). We derive the deployment-time weight-error distribution from first… ▽ More

    Submitted 30 May, 2026; originally announced June 2026.

  36. How to Relieve Distribution Shifts in Semantic Segmentation for Off-Road Environments

    Authors: Ji-Hoon Hwang, Daeyoung Kim, Hyung-Suk Yoon, Dong-Wook Kim, Seung-Woo Seo

    Abstract: Semantic segmentation is crucial for autonomous navigation in off-road environments, enabling precise classification of surroundings to identify traversable regions. However, distinctive factors inherent to off-road conditions, such as source-target domain discrepancies and sensor corruption from rough terrain, can result in distribution shifts that alter the data differently from the trained cond… ▽ More

    Submitted 13 August, 2026; v1 submitted 28 May, 2026; originally announced May 2026.

    Comments: 8 pages, 6 figures. Accepted to IEEE Robotics and Automation Letters (RA-L)

    Journal ref: IEEE Robotics and Automation Letters, vol. 10, issue. 5, pp. 4500-4507, 2025

  37. arXiv:2605.29565  [pdf, ps, other

    cs.CV cs.RO

    From General Vision to Reliable Traversability Estimation: Adapting Vision Foundation Models for Unstructured Outdoor Environments

    Authors: Ji-Hoon Hwang, Jisung Bae, Dong-Wook Kim, Yeonkyu Lee, Seung-Woo Seo

    Abstract: Vision-based approaches have become the dominant paradigm for traversability estimation in unstructured outdoor environments, typically adapting vision foundation models (VFMs) via semantic segmentation supervision. However, this paradigm faces three fundamental challenges that undermine its reliability: the task-agnostic design of VFMs, the ambiguity of traversability annotations, and the discrep… ▽ More

    Submitted 28 May, 2026; originally announced May 2026.

    Comments: 8 pages, 5figures

  38. arXiv:2605.29250  [pdf, ps, other

    cs.CL cs.AI cs.IR cs.LG

    OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources

    Authors: Jinheon Baek, Soyeong Jeong, Sangwoo Park, Woongyeong Yeo, Minki Kang, Patara Trirat, Heejun Lee, Sung Ju Hwang

    Abstract: Real-world information needs require access to structurally diverse knowledge sources, from unstructured text and relational tables to knowledge graphs and property graphs. Existing retrievers, however, operate over one source at a time under a fixed query language, leaving the broader landscape of available knowledge fragmented behind incompatible interfaces. A natural attempt at unification woul… ▽ More

    Submitted 27 May, 2026; originally announced May 2026.

  39. arXiv:2605.28775  [pdf, ps, other

    cs.LG cs.AI cs.CL

    Learn from Weaknesses: Automated Domain Specialization for Small Computer-Use Agents

    Authors: Suji Kim, Kangsan Kim, Sung Ju Hwang

    Abstract: Computer-use agents (CUAs) have recently made substantial progress, but deploying a separate large expert for each software domain remains expensive. Small open computer-use agents are more practical specialization targets, but they remain substantially weaker and exhibit uneven domain-specific failures. A straightforward remedy is to synthesize large-scale training data for the target domain, yet… ▽ More

    Submitted 27 May, 2026; originally announced May 2026.

  40. arXiv:2605.28774  [pdf, ps, other

    cs.CL

    Agent Explorative Policy Optimization for Multimodal Agentic Reasoning

    Authors: Minki Kang, Shizhe Diao, Ryo Hachiuma, Sung Ju Hwang, Pavlo Molchanov, Yu-Chiang Frank Wang, Byung-Kwan Lee

    Abstract: Vision-language models with extended reasoning succeed on complex problems, but many real-world problems require external tools that internal reasoning alone often cannot resolve. Agentic reasoning therefore interleaves two behaviors with a structural asymmetry: thinking (the self-contained default) and tool use (a high-variance auxiliary acting). We refer to this asymmetry as the Thinking-Acting… ▽ More

    Submitted 27 May, 2026; originally announced May 2026.

    Comments: Project page: https://byungkwanlee.github.io/AXPO-page/

  41. arXiv:2605.25439  [pdf, ps, other

    cs.LG

    Missing Pattern Recognized Diffusion Imputation Model for Missing Not At Random

    Authors: Gyuwon Sim, Sumin Lee, Heesun Bae, Byeonghu Na, Doyun Kwon, Ju-Hee Hwang, Jae-Young Lim, Il-Chul Moon

    Abstract: Missing data frequently arises across diverse domains, including time-series and image domains. In the real world, missing occurrences often depend on the unobservable values themselves, which are referred to as Missing Not at Random (MNAR). In this work, we introduce the Missing Pattern Recognized Diffusion Imputation Model (PRDIM), a novel framework that explicitly captures the missing pattern a… ▽ More

    Submitted 25 May, 2026; originally announced May 2026.

  42. arXiv:2605.24776  [pdf, ps, other

    cs.CV

    How Noisy Poses Break Inverse Dynamics: Analysis and Mitigation for Video-Based Joint Torque Estimation

    Authors: Donghyun Kim, Chanyoung Kim, Eunseo Jeong, Youngjoong Kwon, Seong Jae Hwang

    Abstract: Recent advances in monocular 3D human pose estimation enable accurate body tracking from video. However, translating these kinematic estimates into physical quantities, such as joint torques, remains challenging due to noise amplification through inverse dynamics. In this work, we provide a systematic analysis of how pose estimation noise propagates through the inverse dynamics pipeline. We presen… ▽ More

    Submitted 23 May, 2026; originally announced May 2026.

  43. arXiv:2605.20258  [pdf, ps, other

    cs.LG cs.AI cs.CR

    It Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMs

    Authors: Sangwoo Park, Woongyeong Yeo, Seanie Lee, Yumin Choi, Hyomin Lee, Kangsan Kim, Jinheon Baek, Seong Joon Oh, Sung Ju Hwang

    Abstract: Contextual Integrity (CI) defines privacy not merely as keeping information hidden, but as governing information flows according to the norms of a given context. As large language models are increasingly deployed as personal agents handling sensitive workflows, adhering to CI becomes critical. However, even frontier models remain unreliable in making disclosure decisions, and existing mitigation s… ▽ More

    Submitted 18 May, 2026; originally announced May 2026.

    Comments: 28 pages, 16 figures

  44. arXiv:2605.20165  [pdf, ps, other

    cs.CV

    CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models

    Authors: Hsiang-Wei Huang, Junbin Lu, Kuang-Ming Chen, Jianxu Shangguan, Cheng-Yen Yang, Jenq-Neng Hwang

    Abstract: Vision-Language Models (VLMs) achieve strong performance on spatial question answering benchmarks, yet it remains unclear whether such gains reflect genuine spatial intelligence. We show that existing spatial VLMs lack basic camera motion understanding, a key component of spatial cognition. We propose the Spatial Narrative Score (SNS), an evaluation framework that requires VLMs to generate explici… ▽ More

    Submitted 19 May, 2026; originally announced May 2026.

    Comments: Code and model available at https://github.com/hsiangwei0903/CaMo

  45. arXiv:2605.18864  [pdf, ps, other

    cs.LG cs.AI cs.CL

    SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs

    Authors: Chanuk Lee, Minki Kang, Sung Ju Hwang

    Abstract: Recent studies observe that reinforcement learning with verifiable rewards (RLVR) reliably improves pass@1 on reasoning tasks, yet often fails to yield comparable gains in pass@k, raising the question of whether RLVR genuinely enables large language models to acquire novel reasoning abilities or merely enhances the efficiency of sampling reasoning modes already present in the base model. Prior ana… ▽ More

    Submitted 15 May, 2026; originally announced May 2026.

    Comments: Preprint

  46. arXiv:2605.17873  [pdf, ps, other

    cs.LG cs.AI cs.CL

    HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents

    Authors: Woongyeng Yeo, Yumin Choi, Taekyung Ki, Sung Ju Hwang

    Abstract: Training long-horizon LLM agents with reinforcement learning is challenging because sparse outcome rewards reveal whether a task succeeds, but not which intermediate actions caused the outcome or how they should be corrected. Recent methods alleviate this issue by generating rewards or textual hints from turn-level action-output signals, or by using feedback-conditioned self-distillation. However,… ▽ More

    Submitted 18 May, 2026; originally announced May 2026.

  47. arXiv:2605.15726  [pdf, ps, other

    cs.AI cs.CL

    Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR

    Authors: Chanuk Lee, Sangwoo Park, Minki Kang, Sung Ju Hwang

    Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as a scalable paradigm for improving the reasoning capabilities of large language models. However, its effectiveness is fundamentally limited by exploration: the policy can only improve on trajectories it has already sampled. While increasing the number of rollouts alleviates this issue, such brute-force scaling is computationally e… ▽ More

    Submitted 15 May, 2026; originally announced May 2026.

    Comments: 28 pages, 7 figures

  48. arXiv:2605.14705  [pdf, ps, other

    cs.CV

    Towards Continuous Sign Language Conversation from Isolated Signs

    Authors: Youngmin Kim, Kyobin Choo, Jiwoo Park, Minseo Kim, Chanyoung Kim, Junhyeok Kim, Seong Jae Hwang

    Abstract: Sign language is the primary language for many Deaf and Hard-of-Hearing (DHH) signers, yet most conversational AI systems still mediate interaction through spoken or written language. This spoken-language-centered interface can limit access for signers for whom spoken or written language is not the most accessible medium, motivating direct sign-to-sign conversational modeling. However, sentence-le… ▽ More

    Submitted 14 May, 2026; originally announced May 2026.

  49. arXiv:2605.14698  [pdf, ps, other

    cs.LG cs.AI

    NeuroAtlas: Benchmarking Foundation Models for Clinical EEG and Brain-Computer Interfaces

    Authors: Konstantinos Kontras, Trui Osselaer, Stylianos G. Mouslech, Angeliki-Ilektra Karaiskou, Guido Gagliardi, Thomas Strypsteen, Mohammad Hossein Badiei, Anku Rani, Maarten Vanmarcke, Miguel Bhagubai, Chanakya Ekbote, Jaedong Hwang, Christos Chatzichristos, Paul Pu Liang, Maarten De Vos

    Abstract: Foundation models (FMs) promise to extract unified representations that generalize across downstream tasks. They have emerged across fields, including electroencephalography (EEG), but it is less clear how effective they are in this particular field. Published evaluations differ in datasets, in the EEG-specific preprocessing that might influence reported results, and in the reported metrics, frequ… ▽ More

    Submitted 14 May, 2026; originally announced May 2026.

  50. arXiv:2605.14530  [pdf, ps, other

    cs.CV

    Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models

    Authors: Sujung Hong, Chanyong Yoon, Seong Jae Hwang

    Abstract: Large diffusion vision-language models (LDVLMs) have recently emerged as a promising alternative to autoregressive models, enabling parallel decoding for efficient inference and leveraging bidirectional attention for global context. Despite these advances, their behavior under long-form generation remains underexplored. In this work, we show that existing LDVLMs suffer from repetitive generation a… ▽ More

    Submitted 18 May, 2026; v1 submitted 14 May, 2026; originally announced May 2026.