-
A Dual-Expert Strategy Integrating LLMs to Mitigate Negative Transfer in Cross-Domain Sequential Recommendation
Authors:
Hyeongjun Yun,
Kihyuk Song,
Jaegul Choo,
Chung Park
Abstract:
Cross-Domain Sequential Recommendation (CDSR) predicts the next item a user will interact with based on their historical interaction sequences across multiple domains. Recent approaches leverage Large Language Models (LLMs) finetuned on textual representations of cross-domain user sequences to retrieve the recommended items, referred to as LLMRec. However, LLMRec primarily models the autoregressiv…
▽ More
Cross-Domain Sequential Recommendation (CDSR) predicts the next item a user will interact with based on their historical interaction sequences across multiple domains. Recent approaches leverage Large Language Models (LLMs) finetuned on textual representations of cross-domain user sequences to retrieve the recommended items, referred to as LLMRec. However, LLMRec primarily models the autoregressive patterns of token-level item texts, while overlooking item-level collaborative signals. This semantic misalignment often leads to distorted knowledge transfer across domains-termed negative transfer degrading performance in the CDSR task. To address this issue, we propose a novel LLM-based CDSR model, DuELRec: Domain-Gated Dual Experts with LLMs for Cross-Domain Sequential Recommendation. We propose a domain-gated dual-expert framework, equipped with an item-aware attention transformation module, which aggregates textual subtokens into item-level representations and enforces block-level attention masking. The single-domain expert restricts autoregressive attention to items within the same domain, while the cross-domain expert allows it across all domains. A gating mechanism adaptively fuses their outputs, using single-domain signals to reduce cross-domain noise that causes negative transfer. Second, we introduce a dual-sampling token-to-item contrastive learning objective that allows LLMs to capture the item-level collaborative signals from both single- and cross-domains. This is achieved by transforming token-level item texts into item-level representations and applying stochastic negative sampling from both single- and cross-domain item pools for contrastive learning. Extensive experiments on two real-world datasets across ten domains show that our model outperforms 26 state-of-the-art methods in recommendation performance.
△ Less
Submitted 24 August, 2026;
originally announced August 2026.
-
DeepSAGE: Stage-Aware Reinforcement Learning for Structured CBT Counseling Dialogue
Authors:
Qi Zhang,
Heajun An,
Prakriti Dumaru,
Sang Won Lee,
Lifu Huang,
Pamela J. Wisniewski,
Jin-Hee Cho
Abstract:
Large Language Model (LLM)-based counseling agents can generate fluent and supportive responses, but they often lack the structured, goal-directed progression required to conduct a coherent therapeutic session. We present DeepSAGE (Strategic AI Guidance Engine), a hybrid LLM--Deep Reinforcement Learning (DRL) framework for stage-aware counseling dialogue grounded in the first session of Cognitive…
▽ More
Large Language Model (LLM)-based counseling agents can generate fluent and supportive responses, but they often lack the structured, goal-directed progression required to conduct a coherent therapeutic session. We present DeepSAGE (Strategic AI Guidance Engine), a hybrid LLM--Deep Reinforcement Learning (DRL) framework for stage-aware counseling dialogue grounded in the first session of Cognitive Behavioral Therapy (CBT). DeepSAGE represents the session as eleven stages with explicit therapeutic objectives, with an external controller determines stage completion and the DRL model selects therapeutic intentions that guide LLM response generation. We evaluate DeepSAGE against six retrieval-, prompting-, stage-, and policy-based alternatives. DeepSAGE elicits higher simulated client engagement and openness and achieves the strongest balance of stage-goal completion and dialogue efficiency among stage-structured systems. Domain expert review further indicates that the generated conversations exhibit broadly plausible emotional trajectories and recognizable CBT processes. Because the evaluation relies primarily on simulated clients and model-based metrics, these findings demonstrate comparative dialogue-control improvements rather than clinical effectiveness. These results suggest that combining stage-structured dialogue with learned strategy selection is a promising approach for AI counseling, though clinical effectiveness, safety, and real-world utility require further human evaluation.
△ Less
Submitted 23 August, 2026;
originally announced August 2026.
-
PatchGate: Narrowing the Verbalization Gap with Intrinsic Object Inventories in Frozen Vision-Language Models
Authors:
Jihyung Ko,
Eunji Jung,
Hyeongsub Kim,
Ziseok Lee,
Jae Won Cho,
Sanghyun Jo,
Kyungsu Kim
Abstract:
Reliable image captioning in Vision-Language Models (VLMs) requires captions to be both precise and complete, avoiding unsupported object mentions while covering visible objects. Existing training-free methods primarily address the former requirement, suppressing unsupported object words by intervening on model-predicted mentions during generation. Because they operate only on objects the model is…
▽ More
Reliable image captioning in Vision-Language Models (VLMs) requires captions to be both precise and complete, avoiding unsupported object mentions while covering visible objects. Existing training-free methods primarily address the former requirement, suppressing unsupported object words by intervening on model-predicted mentions during generation. Because they operate only on objects the model is already likely to mention, visible objects omitted from the output remain difficult to recover. We propose PatchGate, a training-free framework that extracts prompt-free object evidence intrinsic to a frozen VLM before generation and uses it to narrow the gap between an intrinsic object set and final object mentions. In the first stage, Visual Evidence eXtraction (VEX) reads patch-level lexical evidence from the latter half of LM decoder layers and constructs an image-conditioned object set without any task prompt. In the second stage, Visual-Evidence Inclusion-Exclusion Decoding (VIED) uses this object evidence to calibrate decoding logits, promoting evidence-supported but under-verbalized objects and suppressing weakly supported but over-verbalized objects. On AMBER, PatchGate improves both sides of object-level reliability, increasing visible-object coverage from 49.4 to 56.0 (+13.4%) and reducing object hallucination by lowering CHAIR from 7.5 to 6.6 (-12.0%), without external detectors or fine-tuning and with one extra forward pass.
△ Less
Submitted 22 August, 2026;
originally announced August 2026.
-
Optical characterization and ionizing-radiation response of the BCF-20XL scintillating-wavelength-shifting fiber
Authors:
W. Bae,
L. Baker,
K. Chen,
J. Cho,
K. Lang,
C. Lee,
E. Liang,
C. Mathurin,
C. Murthy,
D. Myers,
S. Nguyen,
D. Phan,
M. Proga,
M. Zalikha,
J. Zey
Abstract:
We report optical characterization and ionizing-radiation measurements of the scintillating-wavelength-shifting (Sci-WLS) fiber BCF-20XL from Luxium Solutions. Attenuation lengths were obtained from transmitted emission spectra using a 3.0 m fiber with a spectrophotometer and described with a two-component attenuation model, yielding a long attenuation length of 6.70 m and a short attenuation leng…
▽ More
We report optical characterization and ionizing-radiation measurements of the scintillating-wavelength-shifting (Sci-WLS) fiber BCF-20XL from Luxium Solutions. Attenuation lengths were obtained from transmitted emission spectra using a 3.0 m fiber with a spectrophotometer and described with a two-component attenuation model, yielding a long attenuation length of 6.70 m and a short attenuation length of 0.16 m. Spectral measurements further show that the long attenuation component increases with wavelength, while the short attenuation component becomes less significant at longer wavelengths. The scintillation response was also characterized under radioactive alpha, beta, and gamma irradiation using SiPM-based readout, providing a benchmark measurement of the light output from BCF-20XL.
△ Less
Submitted 4 August, 2026;
originally announced August 2026.
-
EditPPT: Faithful Long-Deck Slide Editing via Structured Tool-Using Multi-Agent with Dual-Modal Validators
Authors:
Jiheon Kim,
Kyudan Jung,
Jaegul Choo
Abstract:
Automating slide editing requires simultaneously satisfying modification accuracy, preservation fidelity, and robustness to deck length. Existing LLM-based systems often fail on real-world presentation files because they rely on idealized intermediate representations or open-ended code generation, which are prone to cascading errors in long decks. We introduce EditPPT, a multi-agent framework that…
▽ More
Automating slide editing requires simultaneously satisfying modification accuracy, preservation fidelity, and robustness to deck length. Existing LLM-based systems often fail on real-world presentation files because they rely on idealized intermediate representations or open-ended code generation, which are prone to cascading errors in long decks. We introduce EditPPT, a multi-agent framework that reformulates slide editing as a constrained tool-selection problem. By executing localized shape-level operations through the native PowerPoint COM interface, EditPPT narrows the LLM action space while preserving the application-resolved structure of user-authored decks. By separating validation across modalities, our dual-modal validation provides more robust assessment of both instruction fidelity and visual quality. We also present DeckEdit-Bench, a benchmark with 28 human-authored decks, 582 slides, and 183 editing prompts across short, medium, and long deck tiers. Experiments show that EditPPT achieves a 99.5% execution rate, 88.7% slide-targeting F1, 82.5% instruction following, and 91.5% object preservation overall, while maintaining strong performance on long decks. Our code and benchmark are available at https://anonymous.4open.science/r/EditPPT-0E27/
△ Less
Submitted 29 June, 2026;
originally announced August 2026.
-
Denoised Variance-Based Pruning with Optimal Brain Bias Compensation
Authors:
Geon Tack Lee,
Jaegul Choo,
Kang Eun Jeon
Abstract:
Vision Transformers (ViTs) achieve state-of-the-art performance but carry massive computational overhead that restricts edge deployment. Although structural pruning has emerged as a key strategy to reduce these costs, existing methods often suffer from severe accuracy degradation or require expensive retraining. Recently, Variance-Based Pruning (VBP) introduced a promising paradigm by selecting ne…
▽ More
Vision Transformers (ViTs) achieve state-of-the-art performance but carry massive computational overhead that restricts edge deployment. Although structural pruning has emerged as a key strategy to reduce these costs, existing methods often suffer from severe accuracy degradation or require expensive retraining. Recently, Variance-Based Pruning (VBP) introduced a promising paradigm by selecting neurons based on activation variance; however, it remains limited by statistical noise in finite-sample activation covariance and reliance on bias-only updates that cannot fully account for structural reconstruction error. To address these limitations, we introduce Denoised Variance-Based Pruning with Optimal Brain Bias Compensation (DVBP + OB$^2$C). We leverage random matrix theory to filter noise from the activation covariance spectrum for robust neuron selection and mathematically prove that integrating mean-shift compensation into the Optimal Brain Compression objective reduces the layer-wise Hessian exactly to the activation covariance matrix. This enables an optimal, closed-form update of the remaining weights using the same statistics gathered for selection. Extensive experiments on DeiT, Swin, and ConvNeXt architectures demonstrate that DVBP + OB$^2$C achieves state-of-the-art training-free performance; at 50% MLP pruning, it retains over 90% of the original Top-1 accuracy on Small and Base variants, outperforming VBP by up to 29.46% (ConvNeXt-T) and 7.33% (Swin-S). The code is available at: https://github.com/geontackee/DVBP_OB2C.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Scintillation Properties of a Stilbene Crystal for Low-Mass Dark Matter Searches
Authors:
Se Hwan Lee,
Jaeyoung Cho,
D. Joseph Daniel,
Hongjoo Kim,
Young Ju Ko,
Jungho So,
In Soo Lee
Abstract:
Direct searches for low-mass WIMPs require sensitivity to low-energy nuclear recoils. Hydrogen-containing targets offer favorable scattering kinematics because low-mass WIMPs can transfer a larger fraction of their kinetic energy to hydrogen nuclei than to heavier nuclei. Trans-stilbene (t-stilbene) is a hydrogen-rich organic crystal that combines this kinematic advantage with efficient scintillat…
▽ More
Direct searches for low-mass WIMPs require sensitivity to low-energy nuclear recoils. Hydrogen-containing targets offer favorable scattering kinematics because low-mass WIMPs can transfer a larger fraction of their kinetic energy to hydrogen nuclei than to heavier nuclei. Trans-stilbene (t-stilbene) is a hydrogen-rich organic crystal that combines this kinematic advantage with efficient scintillation and pulse-shape discrimination (PSD).In this work, we characterize the scintillation decay behavior and effective light yield of a solution-grown t-stilbene crystal to evaluate its suitability for dark matter detection. The detector consists of a cylindrical crystal approximately 1.1 cm in diameter and 1.1 cm in length, optically coupled at opposite ends to two Hamamatsu R12669 photomultiplier tubes. The scintillation waveforms are described by a triple-exponential decay model, yielding fraction-weighted decay time constants of $8.3 \pm 0.2$ ns (fast), $20.1 \pm 0.3$ ns (medium), and $74.2 \pm 1.1$ ns (slow). The effective light yield, determined using a 59.54 keV $γ$ ray from an $^{241}$Am source, is $3.37 \pm 0.02$ photoelectrons per keV.These measurements establish the baseline performance of the detector and support further evaluation of t-stilbene as a target material for rare-event and low-mass dark matter searches.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Task Specialization Fine-Tuning for Contextual Reinforcement Learning
Authors:
Jianan Zhou,
Jung-Hoon Cho,
Tianyue Zhou,
Han Zheng,
Jie Zhang,
Roy Dong,
Yining Ma,
Cathy Wu
Abstract:
Contextual Reinforcement Learning (CRL) seeks to generalize classical RL by maximizing task coverage across a context space of related tasks. While prior works often train from scratch and rely on either multi-task learning for a single policy or strategically training multiple policies, we advocate for a unified alternative: pretraining a single policy with good initial performance, followed by f…
▽ More
Contextual Reinforcement Learning (CRL) seeks to generalize classical RL by maximizing task coverage across a context space of related tasks. While prior works often train from scratch and rely on either multi-task learning for a single policy or strategically training multiple policies, we advocate for a unified alternative: pretraining a single policy with good initial performance, followed by fine-tuning multiple policies for task specialization. This new paradigm, however, introduces unique challenges, such as heterogeneous marginal returns and sample inefficiency. This raises a critical research question: given a pretrained policy and a constrained budget, how much fine-tuning should each task region receive to enable sample-efficient CRL? To this end, we propose Task Specialization Fine-Tuning (TSFT), an online framework that predicts fine-tuning performance with a simple parametric model and exactly solves the resulting discrete budget allocation problem via integer linear programming. Extensive experiments across diverse decision domains, including combinatorial optimization, continuous control, and LLM fine-tuning, demonstrate that TSFT significantly outperforms baselines in task coverage and approaches oracle performance. Our work charts a new direction for model-based CRL, aligning with the modern pretrain-finetune era.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Claude-SpinDynamics: a cross-platform, dual-precision CPU/GPU micromagnetic simulator with native mumax3 script compatibility
Authors:
Chun-Yeol You,
Jaeyong Cho
Abstract:
Quantitative spintronics increasingly depends on a handful of GPU micromagnetic codes, all of which require NVIDIA hardware and single precision throughout, leaving researchers without such hardware unable to run even the standard validation problems. We report Claude-SpinDynamics (Claude-SD), a new open source micromagnetic simulator with a cross platform C++20 core (Windows and Linux) and a Pyth…
▽ More
Quantitative spintronics increasingly depends on a handful of GPU micromagnetic codes, all of which require NVIDIA hardware and single precision throughout, leaving researchers without such hardware unable to run even the standard validation problems. We report Claude-SpinDynamics (Claude-SD), a new open source micromagnetic simulator with a cross platform C++20 core (Windows and Linux) and a Python interface that closes this gap: a complete CPU build, validated by the same test suite as the GPU path, runs every unit test and uMAG standard problem with no accelerator at all, alongside GPU builds offering both single and double precision, a choice of two demagnetization FFT backends, and natively implemented spin-orbit, spin-transfer, and Zhang-Li torques, Dzyaloshinskii-Moriya interaction, and percell materials.Claude-SD natively interprets mumax3's .mx3 scripting language, so existing community scripts run unmodified; under matched conditions the two codes agree cell-by cell to single-precision round-off, and, together with mumax+ and OOMMF, to within 2% on the uMAG dynamic-switching standard problem, with MuMax-CO agreeing to mumax3 to float32 round-off on the same problem. Benchmarked head-to-head against these three codes, Claude-SD's single-precision build is the fastest solver on small and two-dimensional problems and remains competitive at the largest grid sizes, while its double-precision and dual-FFT-backend paths are unmatched among GPU micromagnetic codes. A GPU replicabatching extension further advances an entire ensemble of finite-temperature trajectories in a single kernel launch per step, giving one to two orders of magnitude of throughput over a per-trial loop while reproducing single-trajectory results to numerical round-off. The complete source is openly licensed and distributed with runnable example notebooks and documentation for independent reproduction.
△ Less
Submitted 15 August, 2026;
originally announced August 2026.
-
Splat-based Metal Artifact Reduction in Cone-Beam CT via Polychromatic Modeling
Authors:
Kiseok Choi,
Inchul Kim,
Jaemin Cho,
Hyeongjun Cho,
Min H. Kim
Abstract:
Cone-beam computed tomography (CBCT) enables volumetric reconstruction from X-ray projections, but suffers from severe artifacts--especially beam hardening--when imaging materials with high attenuation such as metals. These artifacts arise from the polychromatic nature of X-rays and are not properly addressed by conventional monochromatic reconstruction algorithms. While recent neural representati…
▽ More
Cone-beam computed tomography (CBCT) enables volumetric reconstruction from X-ray projections, but suffers from severe artifacts--especially beam hardening--when imaging materials with high attenuation such as metals. These artifacts arise from the polychromatic nature of X-rays and are not properly addressed by conventional monochromatic reconstruction algorithms. While recent neural representation-based methods offer improved reconstruction quality, they are computationally expensive and often impractical for deployment. We propose a novel physics-inspired, self-calibrating metal artifact reduction method that efficiently reconstructs 3D CBCT volumes while correcting beam hardening artifacts. Our method integrates a polychromatic X-ray projection model, material-dependent attenuation profiles, and system response modeling into a Gaussian Splatting framework. Unlike prior work, we eliminate the need for manual metal masks or strong prior assumptions, and we optimize both reconstruction parameters and X-ray spectral characteristics jointly during training. We further introduce a high-fidelity synthetic CBCT dataset generation pipeline validated on Monte-Carlo x-ray simulation toolbox and release new datasets with severe metal-induced artifacts to support the community. This is the first splat-based method for reducing beam hardening in CBCT. Extensive experiments on both synthetic and real-world datasets demonstrate that our method outperforms state-of-the-art approaches in artifact suppression and reconstruction accuracy.
△ Less
Submitted 13 August, 2026;
originally announced August 2026.
-
Context Blindness in DPO: Mitigating Object Hallucination in MLLMs via Context-Calibrated Preference Optimization
Authors:
Byungoh Ko,
Jinyoung Park,
Jongha Kim,
Jeehye Na,
Jaewon Cho,
Hyunwoo J. Kim
Abstract:
Multimodal large language models (MLLMs) have made rapid progress, yet they still exhibit object hallucination, generating plausible but incorrect descriptions that are inconsistent with the visual input. Direct Preference Optimization (DPO) mitigates this by training models to prefer non-hallucinated responses over hallucinated ones, and recent efforts further enrich the preference data with rele…
▽ More
Multimodal large language models (MLLMs) have made rapid progress, yet they still exhibit object hallucination, generating plausible but incorrect descriptions that are inconsistent with the visual input. Direct Preference Optimization (DPO) mitigates this by training models to prefer non-hallucinated responses over hallucinated ones, and recent efforts further enrich the preference data with relevant context. However, it remains unclear whether DPO actually leverages such context. To investigate this, we propose Contextual Preference Gain (CPG), a simple metric that measures how much a model's preference strengthens when relevant context is provided. We find that higher CPG consistently corresponds to lower hallucination, yet standard DPO and its variants exhibit only limited CPG, indicating that they underutilize contextual information and thus remain prone to hallucination. To address this, we propose Context-Calibrated DPO (C$^2$-DPO), which directly maximizes CPG while preserving the original preference ordering. Across multiple benchmarks, C$^2$-DPO substantially reduces hallucination without compromising general reasoning, relatively reducing the Object HalBench hallucination rate of Qwen2-VL-Instruct-2B by 36%. Code is available at https://github.com/mlvlab/C2-DPO
△ Less
Submitted 12 August, 2026;
originally announced August 2026.
-
Development and Initial Performance of an Upgraded NaI(Tl) Crystal Encapsulation for COSINE-100U
Authors:
Doohyeok Lee,
Jae Young Cho,
Chang Hyon Ha,
Eunju Jeon,
Hongjoo Kim,
Jinyoung Kim,
Kyungwon Kim,
SungHyun Kim,
Sun Kee Kim,
Won Kyung Kim,
Yeongduk Kim,
Young Ju Ko,
Hyunseok Lee,
Hyun Su Lee,
In Soo Lee,
Jaison Lee,
Seo Hyun Lee,
Seung Mok Lee,
Reina H. Maruyama,
Jong-Chul Park,
Kangsoon Park,
Kihong Park,
Se Dong Park,
Kyungmin Seo,
Min Ki Son
, et al. (1 additional authors not shown)
Abstract:
The COSINE-100 experiment was designed to test the DAMA/LIBRA annual-modulation claim using low-background NaI(Tl) detectors. For the COSINE-100U upgrade, we developed a new crystal-encapsulation system to increase light-collection efficiency while preserving long-term detector stability, thereby improving sensitivity to low-mass dark matter. The upgraded design eliminates the quartz optical windo…
▽ More
The COSINE-100 experiment was designed to test the DAMA/LIBRA annual-modulation claim using low-background NaI(Tl) detectors. For the COSINE-100U upgrade, we developed a new crystal-encapsulation system to increase light-collection efficiency while preserving long-term detector stability, thereby improving sensitivity to low-mass dark matter. The upgraded design eliminates the quartz optical windows used in COSINE-100 and directly couples the photomultiplier tubes (PMTs) to the crystal end faces through 2-mm-thick silicone optical pads, thereby reducing the number of optical interfaces. For the larger crystals, the crystal edges were beveled to guide scintillation light more efficiently onto 3-inch high-quantum-efficiency PMTs. The performance study uses 2462~h (102.6~days) of room-temperature COSINE-100U data and, for direct background comparisons, reference COSINE-100 data acquired near the end of operation. 698~h (29.1~days) of COSINE-100 data acquired near the end of operation in March 2023. All eight crystals showed higher light yields than in COSINE-100, with values ranging from 15.8 to 27.7~p.e./keV; six crystals exceeded 20~p.e./keV. The measured bulk-$α$ rates were lower than the COSINE-100 values and consistent with the expected time evolution of internal $^{210}$Pb, while the 1--2-MeV surface-$α$ rates were substantially reduced. The upgrade also restored two crystals that had previously been excluded from the COSINE-100 physics analysis because of poor optical performance. Independent validation tests demonstrated that the encapsulation remains mechanically robust and optically stable during long-term immersion in liquid scintillator at low temperature. This paper presents the encapsulation design, the room-temperature detector performance, and the reduction in surface-related backgrounds achieved at the Yemilab facility.
△ Less
Submitted 12 August, 2026;
originally announced August 2026.
-
Grid-Preserving Knowledge Distillation: Transferring Convolutional Inductive Bias to Vision Transformers under Data Scarcity
Authors:
Junyong Choi,
Cheolhyeon Park,
Jaehoon Cho
Abstract:
Vision Transformers demonstrate remarkable global modeling capacity but often underperform in data-scarce regimes. Distilling convolutional inductive biases from a CNN teacher provides an effective remedy while leaving the deployed model unchanged. However, general-purpose feature distillation transfers little in this setting. In CNN-to-CNN distillation, pooling, flattening, and logit-space projec…
▽ More
Vision Transformers demonstrate remarkable global modeling capacity but often underperform in data-scarce regimes. Distilling convolutional inductive biases from a CNN teacher provides an effective remedy while leaving the deployed model unchanged. However, general-purpose feature distillation transfers little in this setting. In CNN-to-CNN distillation, pooling, flattening, and logit-space projections remove the spatial grid that encodes locality and translation equivariance. Unlike a convolutional student, a ViT cannot readily reconstruct this structure on its own. In this paper, we propose iBKD, a distillation framework that preserves the spatial grid throughout the entire transfer process. Its core module, the Inductive Bias Attention Module, aggregates features from all student layers onto the teacher's grid using learned weights. It then enhances structural cues through channel and deformable spatial attention and injects them via convolutional cross-attention operating directly between spatial grids rather than token sets. The module is used only during training, leaving the deployed model as an unmodified ViT with no inference overhead. Across seven Transformer backbones and six data-scarce benchmarks, iBKD consistently outperforms both locality-guidance methods and general knowledge distillation baselines, with its advantage increasing as the amount of training data decreases.
△ Less
Submitted 13 August, 2026; v1 submitted 11 August, 2026;
originally announced August 2026.
-
V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control
Authors:
Donghu Kim,
Youngdo Lee,
Hojoon Lee,
Johan Obando-Ceron,
Byungkun Lee,
Aaron Courville,
Pablo Samuel Castro,
Jaegul Choo,
Clare Lyle
Abstract:
Improving sample efficiency remains a core challenge in reinforcement learning (RL), especially in real-world settings like robotics, where data collection is costly. This challenge is pronounced in visual RL, where high-dimensional inputs often obscure learning signals. While prior work in visual RL has focused on algorithmic solutions, such as better dynamics models or exploration strategies, re…
▽ More
Improving sample efficiency remains a core challenge in reinforcement learning (RL), especially in real-world settings like robotics, where data collection is costly. This challenge is pronounced in visual RL, where high-dimensional inputs often obscure learning signals. While prior work in visual RL has focused on algorithmic solutions, such as better dynamics models or exploration strategies, recent advances in state-based RL show that architectural design alone can lead to significant gains in sample efficiency. This raises an important question: Can these architectural principles transfer to visual RL? In response, we introduce V-Simba, a simple yet effective visual RL architecture inspired by the Simba architecture from state-based RL. Built on top of Soft Actor-Critic (SAC) with data augmentation, V-Simba modifies the architecture by adding normalization layers to stabilize training and using pointwise convolutions to reduce computation. Despite its simplicity, V-Simba matches or outperforms the state-of-the-art methods across the DMC, Adroit, and Meta-World benchmarks, while being more computationally efficient than DrQ-v2. We make our code publicly available at https://github.com/DAVIAN-Robotics/V-Simba.
△ Less
Submitted 7 August, 2026;
originally announced August 2026.
-
Neuro-Symbolic Closed-Loop Control of Laser Powder Bed Fusion with an In-Loop Ontology
Authors:
Gisuk Hong,
Jaebong Cho,
Hyunbo Cho
Abstract:
A geometry-conditioned, neuro-symbolic closed-loop architecture is proposed for laser powder bed fusion, in which a standards-aligned ontology operates inside the control loop and couples symbolic reasoning with statistical learning to set the targets of a constraint-aware predictive controller. The ontology links the process objectives and constraints to the signals a controller can observe, and…
▽ More
A geometry-conditioned, neuro-symbolic closed-loop architecture is proposed for laser powder bed fusion, in which a standards-aligned ontology operates inside the control loop and couples symbolic reasoning with statistical learning to set the targets of a constraint-aware predictive controller. The ontology links the process objectives and constraints to the signals a controller can observe, and a description-logic reasoner converts them into the references and bounds enforced on each scan. The demonstrated case is overhang dross, a quality limit on the melt pool depth, which governs quality yet cannot be measured during the build, is mapped through a geometry- and power-dependent depth-to-width ratio onto a bound on the observable width, with the ratio and its calibrated uncertainty supplied by a Gaussian process. The reasoner classifies each upcoming feature and selects the active constraints-adding a lack-of-fusion floor at overhangs, a monotone guard beyond the calibrated range, and an energy-density cap where a process window is declared while running only on changes of geometric context and otherwise leaving a single small quadratic program on the per-scan path. In an Eagar-Tsai surrogate calibrated to the NIST AM-Bench benchmark for IN625, the architecture eliminates the dross produced by a geometry-blind controller, holds dross at zero with only a small residual lack-of-fusion under dual scoring, degrades gracefully under deliberate plant mismatch, and retargets to new alloys and constraints by editing ontology data rather than code. The results establish architectural feasibility, experimental calibration of the ratio is the principal next step.
△ Less
Submitted 6 August, 2026;
originally announced August 2026.
-
Equipment-centric workpiece localization in near real-time using deep learning-based vision and event-driven finite state machines
Authors:
Dohyeon Kong,
Jaebong Cho,
Hyunbo Cho
Abstract:
Continuous workpiece localization is essential for traceability and process coordination in hot forging, but direct tracking is unreliable because of extreme temperatures, surface degradation, and irregular routing. This study presents an equipment-centric framework that infers workpiece locations from handling equipment observed by multiple static 2D cameras. The framework estimates floorplan-spa…
▽ More
Continuous workpiece localization is essential for traceability and process coordination in hot forging, but direct tracking is unreliable because of extreme temperatures, surface degradation, and irregular routing. This study presents an equipment-centric framework that infers workpiece locations from handling equipment observed by multiple static 2D cameras. The framework estimates floorplan-space 3D equipment coordinates and recognizes grasp and release activities. Event-driven finite state machines validate these activities as discrete handling events and continuously update workpiece states and locations. A keypoint-guided attention mechanism integrated into a 3D convolutional neural network improves activity recognition by focusing on functionally relevant equipment regions. Evaluation in an operational hot forging factory achieved 100\% event detection accuracy within a 33-second tolerance window, a mean localization error of 317.8 mm, and a mean system latency of 21 seconds. The framework connects vision-based perception with interpretable event-driven reasoning and supports visualization of workpiece transfers and quantitative analysis of equipment operations.
△ Less
Submitted 6 August, 2026;
originally announced August 2026.
-
Splat-Based Metal Artifact Reduction in Cone-Beam CT via Compact Attenuation Modeling
Authors:
Kiseok Choi,
Jaemin Cho,
Inchul Kim,
Min H. Kim
Abstract:
X-ray computed tomography (CT) suffers from severe metal artifacts when high-attenuation objects such as dental fillings or orthopedic implants are present. These artifacts originate from the polychromatic nature of X-rays, where attenuation varies strongly with photon energy and material composition, breaking the monochromatic assumption used by conventional reconstruction algorithms. Recent neur…
▽ More
X-ray computed tomography (CT) suffers from severe metal artifacts when high-attenuation objects such as dental fillings or orthopedic implants are present. These artifacts originate from the polychromatic nature of X-rays, where attenuation varies strongly with photon energy and material composition, breaking the monochromatic assumption used by conventional reconstruction algorithms. Recent neural rendering approaches attempt to address this mismatch through differentiable polychromatic projection models, but they still struggle with smoothness bias, loss of fine structures, and prohibitive computation when extended to large-scale cone-beam CT. We introduce a splat-based metal artifact reduction framework that incorporates a physically grounded polychromatic forward model into a continuous Gaussian representation for cone-beam CT. Each Gaussian encodes the energy-dependent attenuation of the underlying material using a compact material parameterization, which enables efficient joint optimization of geometric and material properties without relying on a metal mask. This compact attenuation formulation captures the essential variation across biological tissues and metallic implants, allowing our model to explain metal-induced nonlinearity while preserving high-frequency structure. Experiments on simulated and real cone-beam CT scans show that our method converges significantly faster and suppresses metal artifacts more effectively than existing reconstruction and neural field-based approaches.
△ Less
Submitted 5 August, 2026;
originally announced August 2026.
-
ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition
Authors:
Jooyeol Yun,
Jintae Park,
Hyesu Lim,
Junha Hyung,
Hyungjin Chung,
Jaegul Choo
Abstract:
Recovering an editable design file from a raster image is a common and costly bottleneck in modern design workflows, yet remains challenging since editability depends on recovering multi-modal attributes, such as typography, vector geometry, colors, grouping, and layer ordering. We present ReDesign, an agentic framework that grows an editable layer hierarchy by selecting and composing specialized…
▽ More
Recovering an editable design file from a raster image is a common and costly bottleneck in modern design workflows, yet remains challenging since editability depends on recovering multi-modal attributes, such as typography, vector geometry, colors, grouping, and layer ordering. We present ReDesign, an agentic framework that grows an editable layer hierarchy by selecting and composing specialized tools across modalities. To keep this long decision process reliable despite imperfect tool outputs, we introduce graceful verification at each expansion, which provides local accept, prune, or retry feedback that prevents error accumulation and avoids large scale reruns. To evaluate editability at scale, we introduce the Figma Edit Replay Benchmark, consisting of 909 raw Figma files and 14,796 controlled edit instructions that replay edits on reconstructed outputs. Across this benchmark and standard reconstruction metrics, ReDesign achieves strong visual fidelity while delivering the highest editability across layout, color, and text edits, outperforming layered decomposition baselines and serial tool use pipelines.
△ Less
Submitted 28 July, 2026;
originally announced July 2026.
-
Scalable No-Stockout Charging Scheduling for Battery Swapping Under Time-of-Use Prices
Authors:
Eunbin Cho,
Junki Cho,
Hakjin Lee,
Jaehoon Sim,
Junghoon Seo
Abstract:
A battery-swapping station must provide every arriving vehicle with a charged battery while minimizing the time-of-use cost of recharging returned units. Coordinating heterogeneous compatibility, vehicle-specific return times, and finite charger capacity requires service-aware recharge decisions across the planning horizon. We formulate a per-battery mixed-integer linear program that captures thes…
▽ More
A battery-swapping station must provide every arriving vehicle with a charged battery while minimizing the time-of-use cost of recharging returned units. Coordinating heterogeneous compatibility, vehicle-specific return times, and finite charger capacity requires service-aware recharge decisions across the planning horizon. We formulate a per-battery mixed-integer linear program that captures these operational features under a hard no-stockout constraint and derive a provably equivalent reduced form with fewer explicit binary variables. In the synthetic scaling study, a price-guided battery-path heuristic returned a full-service schedule for every instance; regime-level median solve times ranged from 0.24 to 8.0 seconds. Its median cost premiums were 7-8% over certified reference costs for small- and medium-scale instances, and its certified ex post optimality-gap upper bounds were 9-12% for large- and extra-large-scale instances. For each operational baseline, the certified reference schedules reduced charging-energy cost by 50-60% on instances that the baseline fully served and for which a certified reference was available. In a 30-day replay of 1,002 swaps recorded at a commercial station, the reduced-model and heuristic rolling controllers served every swap and reduced charging-energy cost by approximately 50% relative to immediate charging.
△ Less
Submitted 28 July, 2026; v1 submitted 26 July, 2026;
originally announced July 2026.
-
URHead: A Unified UV-Space Representation for Joint Mesh-3DGS Optimization in Head Avatars
Authors:
Seonghak Lee,
Junhee Cho,
Jisoo Park,
Min-Gyu Park,
Jongmin Lee,
Ju Hong Yoon,
Junseok Kwon
Abstract:
We present URHead, a unified representation for high-fidelity and animatable head avatars that fundamentally redefines mesh-Gaussian integration. While mesh-based methods offer precise geometric control but lack photorealistic detail, and Gaussian-based approaches achieve photorealism but suffer from poor structural consistency, existing hybrid solutions fail to fully leverage their complementary…
▽ More
We present URHead, a unified representation for high-fidelity and animatable head avatars that fundamentally redefines mesh-Gaussian integration. While mesh-based methods offer precise geometric control but lack photorealistic detail, and Gaussian-based approaches achieve photorealism but suffer from poor structural consistency, existing hybrid solutions fail to fully leverage their complementary strengths. Our key contribution is a UV-space unification where both representations share a common UV parameterization. Through joint optimization with adaptive gaussian sampling, our method automatically learns to disentangle and allocate appropriate roles to each component. URHead maintains full parametric controllability while preserving subject-specific details, and outperforms existing state-of-the-art methods in reconstruction quality and animation consistency.
△ Less
Submitted 4 August, 2026; v1 submitted 8 July, 2026;
originally announced July 2026.
-
PersonaTrail: Benchmarking Personalized Web Agents through Browsing Trails
Authors:
Seungbin Yang,
Chaewoon Ki,
Dohyun Lee,
Jaegul Choo,
ChaeHun Park
Abstract:
Recent advances in large language models have enabled web agents to autonomously execute complex tasks. In practice, users frequently provide underspecified instructions, requiring agents to infer the missing context from their raw browsing histories. Existing benchmarks fail to capture this form of personalization, as they either restrict tasks to fully explicit prompts or abstract web interactio…
▽ More
Recent advances in large language models have enabled web agents to autonomously execute complex tasks. In practice, users frequently provide underspecified instructions, requiring agents to infer the missing context from their raw browsing histories. Existing benchmarks fail to capture this form of personalization, as they either restrict tasks to fully explicit prompts or abstract web interaction history into simplified forms. To bridge this gap, we introduce PersonaTrail, a benchmark for personalized web agents operating in a managed open web environment. By leveraging realistic browsing trajectories as user history, PersonaTrail evaluates an agent's ability to infer user preferences and recall information from past browsing sessions. We further propose Preference-Aware Contextual Memory (PACMem), a framework that decomposes raw browsing histories into two types of structured memory: factual memories that summarize individual sessions and preference memories that distill recurring behavioral patterns. At inference time, the agent retrieves the most relevant entries from these memories to guide personalized navigation. Extensive experiments show that PACMem consistently outperforms existing memory-based baselines on both tasks.
△ Less
Submitted 4 August, 2026; v1 submitted 30 May, 2026;
originally announced July 2026.
-
ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion
Authors:
In Cho,
Jeonghwan Cho,
Mijin Yoo,
Gim Hee Lee,
Seon Joo Kim
Abstract:
3D Gaussian Splatting (3DGS) achieves high-quality novel-view synthesis by optimizing freely placed primitives in 3D and adaptively densifying them in under-reconstructed regions. However, this scene-adaptive capacity allocation is largely lost in existing feed-forward 3DGS methods, which commonly regress Gaussians at input pixels and lift them along camera rays. Such pixel-aligned formulations ma…
▽ More
3D Gaussian Splatting (3DGS) achieves high-quality novel-view synthesis by optimizing freely placed primitives in 3D and adaptively densifying them in under-reconstructed regions. However, this scene-adaptive capacity allocation is largely lost in existing feed-forward 3DGS methods, which commonly regress Gaussians at input pixels and lift them along camera rays. Such pixel-aligned formulations make the number and placement of primitives depend on image resolution and input viewpoints rather than scene complexity, resulting in dense and often redundant Gaussian sets. We present ATSplat, a feed-forward 3DGS framework that restores the adaptive allocation capability of 3DGS optimization through Adaptive 3D Tokens. ATSplat first lifts coarse patch-level depth and camera cues into sparse 3D anchor tokens, forming a compact scaffold of the scene. Each token is then regressed into local Gaussians with learnable 3D offsets, decoupling primitive placement from input image grids. An Adaptive Token Expansion module predicts a token-level uncertainty score, supervised by rendering error maps, and selectively expands high-uncertainty tokens through learnable expansion layers. This sparse-to-adaptive formulation enables ATSplat to concentrate primitives in challenging regions while maintaining a compact representation. Experiments on two representative datasets, RealEstate10K and DL3DV, show that ATSplat achieves state-of-the-art rendering quality while reducing the number of Gaussians by more than $5.7\times$ compared with dense feed-forward 3DGS methods. From 12 input images at $512 \times 960$ resolution, ATSplat completes reconstruction in less than a second using a single commercial GPU, and renders high-quality novel views at 1136 FPS ($512 \times 960$) with only 311K Gaussians.
△ Less
Submitted 28 July, 2026; v1 submitted 22 July, 2026;
originally announced July 2026.
-
Solar Open 2 Technical Report
Authors:
Sungrae Park,
Sanghoon Kim,
Gyoungjin Gim,
Jungho Cho,
Hyunwoong Ko,
Minbyul Jeong,
Minjeong Kim,
Keunwoo Choi,
Chaehun Shin,
Chanwoong Yoon,
Dongjun Kim,
Eunwon Kim,
Gyungin Shin,
Hyeonju Lee,
Hyungkyu Kang,
Inseo Song,
Jisu Bae,
Jiyoon Han,
Jiyun Lee,
Joonkee Kim,
Junyeop Lee,
Mikyoung Cha,
Sangwon Yu,
Sehwan Joo,
Seokyoon Kang
, et al. (28 additional authors not shown)
Abstract:
We present Solar Open 2, a 250B-A15B Mixture-of-Experts language model built for long-horizon agentic tasks, scaled up from Solar Open 1 (Solar Open 100B). To hold entire agent trajectories in a single context, Solar Open 2 reaches a 1M-token window through a hybrid attention stack that interleaves one softmax layer among every three linear-attention layers, using no positional encoding and a gate…
▽ More
We present Solar Open 2, a 250B-A15B Mixture-of-Experts language model built for long-horizon agentic tasks, scaled up from Solar Open 1 (Solar Open 100B). To hold entire agent trajectories in a single context, Solar Open 2 reaches a 1M-token window through a hybrid attention stack that interleaves one softmax layer among every three linear-attention layers, using no positional encoding and a gated delta rule extended to negative eigenvalues. To train at this scale under a fixed compute budget, we make training efficient in two ways: a stronger starting point, and higher-value data. For the starting point, we initialize Solar Open 2 from Solar Open 1, transferring the 5.69B-parameter shared skeleton that survives the architectural change and learning everything else through full pre-training. For the data, we curate for value per token: quality- and rarity-aware data curation and mixture-ratio optimization refine a 20T pool into a 10T mixture that, at equal token budget, outperforms the Solar Open 1 recipe. To build its agent skills, we train twelve domain specialists across purpose-built scenarios, then consolidate them into a single model by Multi-teacher On-Policy Distillation (MOPD). Against comparably sized open-weight models on English benchmarks, Solar Open 2 leads on MMLU-Pro, LiveCodeBench, and the APEX-Agents agentic suite, and stays competitive with the strongest (DeepSeek-V4-Flash and MiMo-V2.5) elsewhere. On Korean benchmarks, Solar Open 2 records the highest average of any model compared, including fast-tier closed APIs, and on Ko-GDPval, an in-house Korean officework-agent benchmark, it is competitive with DeepSeek-V4-Pro (1.6T) at less than a sixth of its size.
△ Less
Submitted 23 July, 2026; v1 submitted 22 July, 2026;
originally announced July 2026.
-
Howson property and finitely generated intersection problem for monogenic inverse semigroups
Authors:
Jung Won Cho,
Craig Miller,
Nik Ruškuc
Abstract:
An algebraic structure is said to have the Howson property if the intersection of any two finitely generated subalgebras is finitely generated. We explore the Howson property in the context of monogenic inverse semigroups. It is known, due to work of Jones and Trotter (1989) and Jones (2016), that every monogenic inverse semigroup has the Howson property considered as an inverse semigroup, i.e. wi…
▽ More
An algebraic structure is said to have the Howson property if the intersection of any two finitely generated subalgebras is finitely generated. We explore the Howson property in the context of monogenic inverse semigroups. It is known, due to work of Jones and Trotter (1989) and Jones (2016), that every monogenic inverse semigroup has the Howson property considered as an inverse semigroup, i.e. with respect to its inverse subsemigroups. In this paper, we consider monogenic inverse semigroups qua semigroups, i.e. we consider all their subsemigroups. We prove that every monogenic inverse semigroup possesses the Howson property in this broader sense, with the sole exception of the monogenic free inverse semigroup. For this exceptional case, we show that the problem of determining whether the intersection of two finitely generated subsemigroups is finitely generated is algorithmically decidable.
△ Less
Submitted 21 July, 2026;
originally announced July 2026.
-
DFT-p-FDMA Based Chirp Transmission in CP-OFDM for Unified ISAC Waveform Design
Authors:
Fabrizio Carpi,
Joonyoung Cho,
Kyeong Jin Kim,
Charlie Jianzhong Zhang
Abstract:
We propose an integrated sensing and communications (ISAC) framework that supports chirp signal transmission in CP-OFDM-based multiple access communication systems, enabling efficient coexistence of communication and sensing capabilities. Our framework employs the discrete Fourier transform phase rotated and permuted frequency division multiple access (DFT-p-FDMA) waveform to transmit chirp signal…
▽ More
We propose an integrated sensing and communications (ISAC) framework that supports chirp signal transmission in CP-OFDM-based multiple access communication systems, enabling efficient coexistence of communication and sensing capabilities. Our framework employs the discrete Fourier transform phase rotated and permuted frequency division multiple access (DFT-p-FDMA) waveform to transmit chirp signals using a portion of the frequency resources, while ensuring interference-free concurrent CP-OFDM data transmissions on other bands. We analyze the effective channel behavior under the DFT-p-FDMA waveform, characterizing how delays and Doppler shifts impact radar target echoes. We also show how processing multiple received symbols improves Doppler resolution in practical scenarios. Our framework allows flexible adjustment of range-Doppler resolution through optimized time-frequency resource allocation, offering a versatile solution for ISAC applications. Simulation results validate the framework's performance in delay and Doppler estimation, highlighting its potential to support ISAC in next-generation wireless networks.
△ Less
Submitted 16 July, 2026;
originally announced July 2026.
-
Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques
Authors:
Daehoon Gwak,
Minhyung Lee,
Junwoo Park,
Jaegul Choo
Abstract:
Diffusion large language models (dLLMs) offer a theoretical advantage in parallel generation over standard autoregressive models. However, parallel generation alone does not guarantee practical speedups. Realizing this efficiency requires specialized inference mechanisms, such as diffusion-aware caching and reuse. Consequently, as inference efficiency becomes a prerequisite for practical deploymen…
▽ More
Diffusion large language models (dLLMs) offer a theoretical advantage in parallel generation over standard autoregressive models. However, parallel generation alone does not guarantee practical speedups. Realizing this efficiency requires specialized inference mechanisms, such as diffusion-aware caching and reuse. Consequently, as inference efficiency becomes a prerequisite for practical deployment, recent research has actively explored acceleration techniques across algorithms, architectures, and systems. However, rigorous comparisons remain difficult, as end-to-end latency stems from intricate trade-offs between algorithmic, architectural, and system-level factors that are often conflated in existing benchmarks. In this survey, we introduce a unified latency decomposition framework for dLLMs to disentangle these factors and analyze their impact on inference speed in real deployments. Guided by this framework, we categorize acceleration techniques along three axes covering algorithmic innovations, architectural and system optimizations, and inference-time scaling. Finally, we provide guidelines for reproducible benchmarking and highlight open challenges for realizing the full potential of parallel generation.
△ Less
Submitted 14 July, 2026;
originally announced July 2026.
-
Bet on Features: Anytime-Valid and Feature-Aware Auditing of Conditional Quantile Forecasters
Authors:
Ivane Antonov,
Sohom Mukherjee,
Richard Pibernik,
Yo Joong Choe
Abstract:
Black-box conditional quantile forecasts are widely used for sequential decisions under asymmetric costs, such as inventory planning in supply chain management. Once deployed, such forecasters must be monitored continuously as data streams drift and regimes change; this invalidates standard, fixed-horizon backtests for calibration. Further, existing backtests do not take into account that the noti…
▽ More
Black-box conditional quantile forecasts are widely used for sequential decisions under asymmetric costs, such as inventory planning in supply chain management. Once deployed, such forecasters must be monitored continuously as data streams drift and regimes change; this invalidates standard, fixed-horizon backtests for calibration. Further, existing backtests do not take into account that the notion of calibration is, in fact, information-dependent: forecasts can look calibrated to an auditor with coarse information while being miscalibrated to an auditor with richer information. We develop a distribution-free and game-theoretic testing framework for continuously auditing black-box conditional quantile forecasters with non-i.i.d. losses, such that the resulting evidence process is powerful against predictably chosen alternatives specified by the features available to the auditor. We first formalize notions of conditional quantile calibration when different sets of features are available to the auditor, establishing that the coarseness of the auditor's information set determines the hardness of the testing problem. We then identify the sets of alternatives for which the auditor can achieve power, and focusing on contextual bets linear in the features, we derive finite-time detection guarantees for such alternatives, all without an i.i.d. assumption. The resulting evidence processes are interpretable at the feature level, as they quantify fine-grained, "feature-aware" evidence for miscalibration. We empirically validate these methods on simulated and real data, finding that a popular time series forecaster (Chronos-2) is highly miscalibrated w.r.t. multiple relevant features.
△ Less
Submitted 13 July, 2026;
originally announced July 2026.
-
See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models
Authors:
Byungkun Lee,
Dongyoon Hwang,
Dongjin Kim,
Hojoon Lee,
Minho Park,
Jaegul Choo
Abstract:
Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in the camera frame, creating a frame mismatch between where the scene is observed and where actions are defined. The mismatch is benign under a fixed viewpoint, where the policy can memorize…
▽ More
Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in the camera frame, creating a frame mismatch between where the scene is observed and where actions are defined. The mismatch is benign under a fixed viewpoint, where the policy can memorize a single observation-to-action mapping, but grows harder as large-scale datasets aggregate demonstrations across diverse camera setups and the policy must generalize this mapping across viewpoints. We address this mismatch with robot-centric pointmaps, images whose pixels store the 3D coordinates of scene points in the robot frame. Pointmaps provide robot-frame 3D geometry while preserving the dense H x W grid expected by pretrained 2D VLAs, so they integrate into existing VLAs with minimal architectural change. On RoboCasa, pointmaps improve both pi0.5 and SmolVLA and outperform representative camera-viewpoint and 3D-aware baselines. In real-robot experiments, their advantage over an RGB-only policy widens when the camera is moved to a placement unseen during training.
△ Less
Submitted 13 July, 2026;
originally announced July 2026.
-
MJ: Multi-turn LLM Jailbreaking via Decomposed Credit Assignment
Authors:
Junyoung Park,
Namgyu Park,
Sechan Lee,
Yoon-Chan Jhi,
Jihoon Cho,
Sangdon Park
Abstract:
Modern large language models (LLMs) operate in interactive multi-turn settings, making multi-turn jailbreaking a realistic threat model and an important setting for automated red teaming. A core challenge in learning multi-turn jailbreak attackers is credit assignment: different turns contribute differently to the final outcome, yet existing learning signals are often too coarse to identify their…
▽ More
Modern large language models (LLMs) operate in interactive multi-turn settings, making multi-turn jailbreaking a realistic threat model and an important setting for automated red teaming. A core challenge in learning multi-turn jailbreak attackers is credit assignment: different turns contribute differently to the final outcome, yet existing learning signals are often too coarse to identify their individual contributions. We propose decomposed credit GRPO (DC-GRPO), a unified turn-level credit assignment framework for Group Relative Policy Optimization in multi-turn jailbreak learning. DC-GRPO assigns a separate group-relative learning signal to each turn by combining immediate and future credit, avoiding the credit misassignment induced by broadcasting a single trajectory-level score across the dialogue. We instantiate this framework with static and dynamic weighting rules that differ in how the two credit sources are balanced while sharing the same turn-level structure. Across multiple victim LLMs and benchmarks, the dynamic- and static-weighted variants achieve average ASR5@3 scores of 98.26% and 97.88%, respectively, substantially outperforming the state-of-the-art methods, including SEMA (86.58%) and TROJail (86.23%). Their consistently strong performance indicates that the central empirical benefit comes from turn-level group-relative credit assignment rather than a particular weighting rule. Warning: This paper contains examples of harmful content.
△ Less
Submitted 13 July, 2026;
originally announced July 2026.
-
Strong First-Order Electroweak Phase Transitions and Gravitational Waves in the Normal Two-Higgs-Doublet Model: A Comparative Study of the Four Yukawa Types and Thermal Resummation Schemes
Authors:
Jin-Hwan Cho,
Dongjoo Kim,
Jinheung Kim,
Soojin Lee,
Jeonghyeon Song
Abstract:
We present a comprehensive global analysis of strong first-order electroweak phase transitions (SFOEWPTs) and their associated stochastic gravitational-wave (GW) backgrounds within the Normal Scenario of the $CP$-conserving Two-Higgs-Doublet Model (2HDM) with softly broken $Z_2$ symmetry, where the lighter $CP$-even scalar is identified as the observed $125~\text{GeV}$ Higgs boson. Across all four…
▽ More
We present a comprehensive global analysis of strong first-order electroweak phase transitions (SFOEWPTs) and their associated stochastic gravitational-wave (GW) backgrounds within the Normal Scenario of the $CP$-conserving Two-Higgs-Doublet Model (2HDM) with softly broken $Z_2$ symmetry, where the lighter $CP$-even scalar is identified as the observed $125~\text{GeV}$ Higgs boson. Across all four Yukawa structures (Type-I, II, X, and Y), we track the finite-temperature vacuum evolution, transition dynamics, and GW signatures. To quantify the theoretical uncertainty associated with thermal resummation, we perform a detailed comparison between the Parwani and Arnold--Espinosa prescriptions. While both schemes find that single-step paths overwhelmingly dominate successful transitions and consistently favor the Higgs alignment limit, the resulting SFOEWPT parameter space exhibits a pronounced scheme dependence. The Arnold-Espinosa prescription severely restricts the viable parameter space (with upper bounds on the heavy-scalar masses below approximately 800 GeV) and introduces an extreme parametric sensitivity that produces fragmented distributions and irregular voids in the heavy-scalar mass planes. In contrast, the more stable Parwani prescription allows heavy-scalar masses below $\sim 1.6~\text{TeV}$. We further identify highly restricted GW parameter regions capable of yielding a four-year LISA signal-to-noise ratio above 10, while demonstrating that the acoustic GW source is generically short-lived, leading to a substantial suppression of the predicted signal amplitude. Our results highlight the strong complementarity between future space-based GW observations and high-energy collider searches in probing the cosmological viability of the 2HDM.
△ Less
Submitted 10 July, 2026;
originally announced July 2026.
-
Dec-MARVEL: Decentralized Multi-Agent Exploration without Communication under Budget Constraints
Authors:
Janghyun Cho,
Jimmy Chiun,
Guillaume Sartoretti,
Changjoo Nam
Abstract:
Multi-UAV exploration is often constrained by unreliable communication, limited field-of-view sensing (e.g., lightweight onboard camera), and finite travel budgets that require each robot to reserve enough budget to return to its base. We present Dec-MARVEL, a decentralized budget-aware exploration framework for communication-free teams with directional sensing. Rather than exchanging maps, goals,…
▽ More
Multi-UAV exploration is often constrained by unreliable communication, limited field-of-view sensing (e.g., lightweight onboard camera), and finite travel budgets that require each robot to reserve enough budget to return to its base. We present Dec-MARVEL, a decentralized budget-aware exploration framework for communication-free teams with directional sensing. Rather than exchanging maps, goals, or messages, each robot coordinates through its incidental observations: any teammate trajectory within its field of view serves as a coordination signal. A graph-attention actor fuses local frontier geometry, teammate motion, and budget features to select return-feasible waypoint-heading actions. The actor is trained with phase-conditioned critics, a training-only task-oriented privileged critic, and a mixture-based budget curriculum. Across 900 held-out trials spanning three team sizes (2, 4, 8 robots) and three travel budgets (720, 800, 1024 meters) against four baselines, Dec-MARVEL achieves the highest or tied-highest exploration rate and lowest sensing overlap across all nine team-size budget configurations. Under our tightest 720m budget, it reaches 53%, 94%, and 100% success for 2, 4, and 8 robots, versus 37%, 83%, and 99% for the strongest baseline. Physical-robot experiments demonstrate successful sim-to-real transfer and real-world deployment of Dec-MARVEL.
△ Less
Submitted 13 July, 2026; v1 submitted 9 July, 2026;
originally announced July 2026.
-
A study of neutrinoless double electron capture in $^{40}$Ca from the AMoRE experiment
Authors:
AMoRE Collaboration,
A. Agrawal,
V. V. Alenkov,
P. Aryal,
J. Beyer,
B. Bhandari,
R. S. Boiko,
K. Boonin,
O. Buzanov,
C. R. Byeon,
N. Chanthima,
M. K. Cheoun,
J. S. Choe,
Seonho Choi,
S. Choudhury,
J. S. Chung,
F. A. Danevich,
M. Djamal,
D. Drung,
C. Enss,
A. Fleischmann,
A. M. Gangapshev,
L. Gastaldo,
Y. M. Gavrilyuk,
A. M. Gezhaev
, et al. (85 additional authors not shown)
Abstract:
The search for neutrinoless double electron capture ($0ν\mathrm{2EC}$) provides a sensitive probe of lepton-number violation and the Majorana nature of neutrinos. We investigate the $0ν\mathrm{2EC}$ decay of $^{40}$Ca using cryogenic detectors equipped with metallic magnetic calorimeters in the AMoRE-I experiment. The analysis is based on a physics dataset corresponding to a total exposure of 7.32…
▽ More
The search for neutrinoless double electron capture ($0ν\mathrm{2EC}$) provides a sensitive probe of lepton-number violation and the Majorana nature of neutrinos. We investigate the $0ν\mathrm{2EC}$ decay of $^{40}$Ca using cryogenic detectors equipped with metallic magnetic calorimeters in the AMoRE-I experiment. The analysis is based on a physics dataset corresponding to a total exposure of 7.32 kg$\cdot$yr from thirteen $^{40}$Ca$^{100}$MoO$_4$ crystals. No significant excess is observed, and a lower limit on the half-life is obtained as $T^{0ν}_{1/2} > 1.7 \times 10^{22}$ yr at 90$\%$ confidence level. An improved sensitivity is expected for the upcoming AMoRE-II experiment. These results demonstrate the potential of CaMoO$_4$ detectors to explore rare decay processes beyond the primary $^{100}$Mo $0νββ$ search program.
△ Less
Submitted 8 July, 2026;
originally announced July 2026.
-
Refractive-index tomography of opaque tissue from its own backscattered light
Authors:
Tran Dinh Hoang,
Jaecheol Cho,
Thi Van Anh Nguyen,
Eunyoung Seong,
Joowon Lim,
Jin Hee Hong,
Yongwoo Kwon,
Jun Wan Kim,
Juhee Yang,
Seokchan Yoon,
Sungsam Kang,
Wonshik Choi
Abstract:
The refractive index (RI) is an intrinsic, label-free marker of a living cell's dry mass and subcellular morphology, and hence of its physiological state. Its three-dimensional (3D) reconstruction has become a powerful way to study cells and tissues in their native state, spanning cell growth, drug response and disease diagnosis. Yet this capability rests on a fundamental constraint: the RI can be…
▽ More
The refractive index (RI) is an intrinsic, label-free marker of a living cell's dry mass and subcellular morphology, and hence of its physiological state. Its three-dimensional (3D) reconstruction has become a powerful way to study cells and tissues in their native state, spanning cell growth, drug response and disease diagnosis. Yet this capability rests on a fundamental constraint: the RI can be recovered only from light transmitted through the specimen, which demands optical access to both sides. The cells that matter most -- those within thick tissues, intact organs and living animals -- are therefore out of reach. A tissue, however, can illuminate its own cells from behind: light backscattered by intrinsic tissue structures beneath a cell carries the same transmission information a microscope would collect from the far side. Here we develop a divide-and-conquer inverse-scattering framework that recovers this transmission from the backscattering and reconstructs a cell's 3D RI. We demonstrate label-free, quantitative imaging of cells within an engineered tissue, and a living mouse through its intact skull, where we further quantify the dry mass of individual osteocytes in vivo. By removing the need for two-sided access, this reflection-only approach extends RI tomography into living tissue, enabling non-destructive, longitudinal imaging of cells in their native environment.
△ Less
Submitted 7 July, 2026;
originally announced July 2026.
-
Dynamic Congestion Pricing in Distribution Networks via a Convex-Analytic Bilevel Reformulation
Authors:
Reza Rahimi Baghbadorani,
Ali Nikseresht,
Jehum Cho,
Yashar Ghiassi-Farrokhfal
Abstract:
Dynamic congestion pricing is an important tool for managing congestion and coordinating distributed energy resources in active distribution networks. However, scalable mechanisms that preserve participant autonomy remain computationally challenging because the operator-resource interaction is naturally bilevel. This paper develops a convex-analytic framework in which a distribution system operato…
▽ More
Dynamic congestion pricing is an important tool for managing congestion and coordinating distributed energy resources in active distribution networks. However, scalable mechanisms that preserve participant autonomy remain computationally challenging because the operator-resource interaction is naturally bilevel. This paper develops a convex-analytic framework in which a distribution system operator computes dynamic congestion-price adders, while decentralized energy hubs schedule flexible demand, storage, local generation, renewable curtailment, and grid import/export. Unlike conventional single-level reformulations that replace lower-level problems by Karush-Kuhn-Tucker (KKT) conditions, complementarity constraints, and big-M linearizations, the proposed model represents follower feasibility and optimality through a Fenchel-Young equality involving the convex conjugate of an extended follower objective. The remaining bilinear price-response term is handled through a penalized difference-of-convex reformulation and sequential convex approximation. The method solves continuous convex subproblems and avoids the constraint-wise complementarity and branch-and-bound scaling of mixed-integer KKT reformulations; its main computational drivers are price-response dimension and conjugate evaluation rather than binary encodings of follower inequalities. On augmented IEEE 13- and 34-node feeders, it reduces congestion by 96.89% and 96.45%, respectively, approaches centralized full-information dispatch, certifies price-response consistency to numerical precision, and yields lower residual congestion than time-limited KKT incumbents within the computational budget.
△ Less
Submitted 3 July, 2026;
originally announced July 2026.
-
JointHOI: Jointly Generating Contact Maps Enhances Hand Object Interaction Generation
Authors:
Mingyeong Song,
Jungbin Cho,
Jisoo Kim,
Ananya Bal,
Kartik Sharma,
Youngjae Yu,
Laszlo A. Jeni,
Junhyug Noh
Abstract:
Text driven hand object interaction (HOI) generation is gaining attention for immersive applications and robotics, yet producing physically plausible interactions remains challenging. Even when individual motions appear natural, small contact errors can cause conspicuous artifacts such as floating and interpenetration. Prior methods mitigate these issues using explicit contact cues or implicit gra…
▽ More
Text driven hand object interaction (HOI) generation is gaining attention for immersive applications and robotics, yet producing physically plausible interactions remains challenging. Even when individual motions appear natural, small contact errors can cause conspicuous artifacts such as floating and interpenetration. Prior methods mitigate these issues using explicit contact cues or implicit grasp priors, but typically rely on multi stage pipelines and fail to model temporally evolving contact. We present JointHOI, a single stage diffusion framework that jointly generates 3D hand object motion and dynamic, distance based contact maps from text. By treating contact as an auxiliary inner modality, joint generation enables the model to learn contact motion coupling during training. At inference, contact guided sampling enforces consistency between generated contact maps and motion implied geometry, improving temporal stability and reducing penetration and floating. Experiments on GRAB and ARCTIC demonstrate consistent improvements in text adherence and physical plausibility over prior methods.
△ Less
Submitted 2 July, 2026;
originally announced July 2026.
-
Vitality-Aware Compression for Efficient Image-to-Shape Diffusion Transformers
Authors:
Jaeah Lee,
Hyunjin Kim,
Jaewoong Cho,
Gihyun Kwon
Abstract:
We propose the first compression approach for image-to-shape Diffusion Transformers (DiTs) that substantially reduces model size while preserving geometric fidelity. Despite remarkable progress in 3D shape generation, large DiT-based models remain computationally prohibitive in resource-constrained settings. Furthermore, it is difficult to directly transfer existing diffusion model compression str…
▽ More
We propose the first compression approach for image-to-shape Diffusion Transformers (DiTs) that substantially reduces model size while preserving geometric fidelity. Despite remarkable progress in 3D shape generation, large DiT-based models remain computationally prohibitive in resource-constrained settings. Furthermore, it is difficult to directly transfer existing diffusion model compression strategies developed for different domains to 3D generation, and prior 3D efficiency approaches focus primarily on inference speed rather than backbone compression. To address this limitation, we build a geometry-aware compression framework tailored to image-to-shape DiTs. Guided by the observation that 3D DiT layers exhibit non-uniform importance for geometry synthesis, we introduce a vitality-guided framework integrating structured pruning, adaptive quantization, and targeted fine-tuning. Our method achieves up to 66% model-size reduction across state-of-the-art image-to-3D models while maintaining synthesis fidelity comparable to full-sized counterparts. This highlights the potential of our framework as a plug-and-play solution for efficient 3D shape generation across diverse models.
△ Less
Submitted 30 June, 2026;
originally announced July 2026.
-
3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance
Authors:
Dongyoon Hwang,
Byungkun Lee,
Dongjin Kim,
Hyojin Jang,
Hoiyeong Jin,
Jueun Mun,
Minho Park,
Hojoon Lee,
Hyunseung Kim,
Jaegul Choo
Abstract:
Hierarchical Vision-Language-Action (VLA) models decouple high-level planning from low-level control to improve generalization in robot manipulation. Recent work in this paradigm uses 2D end-effector trajectories predicted by a Vision-Language Model (VLM) as explicit guidance for a downstream policy. However, state-of-the-art low-level policies operate in 3D metric space on point clouds, and feedi…
▽ More
Hierarchical Vision-Language-Action (VLA) models decouple high-level planning from low-level control to improve generalization in robot manipulation. Recent work in this paradigm uses 2D end-effector trajectories predicted by a Vision-Language Model (VLM) as explicit guidance for a downstream policy. However, state-of-the-art low-level policies operate in 3D metric space on point clouds, and feeding them 2D guidance that lacks depth forces each waypoint to be assigned the depth of whatever scene surface lies beneath it, producing geometrically distorted trajectories. We propose 3D HAMSTER, a hierarchical framework that closes this gap by having the planner directly output metrically reliable 3D trajectories. We augment a VLM with a dedicated depth encoder and a dense depth reconstruction objective to predict 3D waypoint sequences, which are directly integrated into a pointcloudbased low-level policy. Across 3D trajectory prediction, simulation, and real-world manipulation, 3D HAMSTER consistently outperforms proprietary VLMs and 2D-guided baselines, with the largest gains under appearance-altering shifts and unseen language, spatial, and visual conditions. The project page is available at https://davian-robotics.github.io/3D_HAMSTER/.
△ Less
Submitted 1 July, 2026; v1 submitted 30 June, 2026;
originally announced June 2026.
-
Can LLMs Imagine Moral Alternatives Beyond Binary Dilemmas?
Authors:
Jongchan Choi,
Nari Yang,
Sung Soo Park,
Jaemin Cho,
Han Seoyoung,
Haerin Shin,
Jun-Hyung Park
Abstract:
As large language models (LLMs) are increasingly deployed as moral advisors and agents, they need to address dilemmas between two competing values. However, existing research on LLMs with moral dilemmas overlooks a central aspect of human moral cognition: the ability to imagine alternatives that move beyond the given options. We introduce MoralAltDataset, a dataset of 307 moral dilemmas spanning n…
▽ More
As large language models (LLMs) are increasingly deployed as moral advisors and agents, they need to address dilemmas between two competing values. However, existing research on LLMs with moral dilemmas overlooks a central aspect of human moral cognition: the ability to imagine alternatives that move beyond the given options. We introduce MoralAltDataset, a dataset of 307 moral dilemmas spanning narrative Advisor dilemmas and AI-facing Agent dilemmas, each augmented with compromise and reframed alternatives. We first examine whether humans and LLMs shift their judgments when such alternatives are introduced. Across 15 LLMs, we find that compromise alternatives are often preferred over either original option, substantially reshaping moral choice. We then evaluate the quality of LLM-generated alternatives against human-authored ones using pairwise preference and expert-based criteria. Results show that LLM-generated alternatives are often preferred and better satisfy fine-grained structural and ethical criteria, while revealing trade-offs between structural quality and practical feasibility.
△ Less
Submitted 30 June, 2026;
originally announced June 2026.
-
Seeing Through the Weights: Privacy Leakage in Scene Coordinate Regression
Authors:
Oleksii Nasypanyi,
Jaemin Cho,
Utku Ozbulak,
Byungkon Kang,
Francois Rameau
Abstract:
Scene Coordinate Regression (SCR) methods are increasingly adopted for visual localization. In these approaches, the scene is implicitly encoded within a neural network that regresses a 3D world coordinate for each image pixel. Because the scene is represented only through the network parameters and not stored explicitly as images or maps, such methods are often assumed to be privacy-preserving. I…
▽ More
Scene Coordinate Regression (SCR) methods are increasingly adopted for visual localization. In these approaches, the scene is implicitly encoded within a neural network that regresses a 3D world coordinate for each image pixel. Because the scene is represented only through the network parameters and not stored explicitly as images or maps, such methods are often assumed to be privacy-preserving. In this work, we show that this assumption is incorrect in practice.
Specifically, we introduce a query-based attack that reconstructs the 3D geometry of the training environment from an SCR model under different levels of model access. To do so, we repeatedly query the model with batches of proxy images unrelated to the target scene to obtain dense pixel-wise 3D coordinates. Reliable points are identified through their stability under small input perturbations and can be further refined in a white-box setting. These stable points are accumulated across independent query batches to recover the scene geometry. From the recovered 3D representation, we also invert the network features to synthesize images from arbitrary viewpoints, revealing additional appearance information.
Experiments on indoor and outdoor datasets demonstrate that substantial portions of training environments can be reconstructed with high geometric fidelity. Beyond geometry, we also recover an approximate color appearance, which exposes recognizable layout and potentially sensitive scene elements. This directly contradicts claims in the literature that SCR representations are privacy-preserving by design, and reveals a real risk when such systems are deployed in private or security-critical spaces. The project page is available at https://jaeminch0.github.io/seeing-through-the-weights-privacy-leakage-in-scene-coordinate-regression.
△ Less
Submitted 30 June, 2026;
originally announced June 2026.
-
MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs
Authors:
Yuxuan Fan,
Gyusik Seo,
Jing Hao,
Jaemin Cho,
Mohit Bansal,
Jaehong Yoon
Abstract:
Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic meaning arises from deliberate combinations of visual, auditory, and narrative elements (e.g., fear amplified through claustrophobic framing, or grief conveyed through silence and lingering close-ups). True artistic understanding extends beyond recognizing what…
▽ More
Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic meaning arises from deliberate combinations of visual, auditory, and narrative elements (e.g., fear amplified through claustrophobic framing, or grief conveyed through silence and lingering close-ups). True artistic understanding extends beyond recognizing what is depicted to reasoning about why it is expressed through particular creative choices. Despite the strong progress of multimodal large language models (MLLMs), this critical aspect of artistic understanding remains underexplored, as existing benchmarks largely measure perceptual recognition while overlooking reasoning about creative intent. To address this gap, we introduce Musebench, a comprehensive benchmark designed to evaluate MLLMs on nuanced artistic understanding. It comprises 4,016 questions spanning cinematic arts, static visual arts, stage performing arts, and game arts, distilled from over 10K candidate video essays that pair professional commentary with visual demonstration. To capture the open-ended nature of artistic analysis at scale, the benchmark combines single-select and variable-option multi-select questions. All questions are generated and refined through a four-phase iterative pipeline combining shortcut filtering, adversarial distractors, and expert validation. Comprehensive zero-shot evaluation of 28 state-of-the-art MLLMs reveals that even the best-performing model achieves only 48.29% accuracy, substantially below human expert performance of 87.18%, exposing a significant gap in current models' creative domain expertise.
△ Less
Submitted 29 June, 2026;
originally announced June 2026.
-
Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation
Authors:
Atin Pothiraj,
Jaemin Cho,
Yue Zhang,
Elias Stengel-Eskin,
Mohit Bansal
Abstract:
Video generation models are increasingly capable of producing realistic videos, but they still struggle to generate videos that follow basic physical laws. Compounding this is a lack of reliable granular evaluation methods for localizing and specifying physical law violations in videos. We address this by introducing Physics Question Scene Graph (PQSG), a hierarchical question-based evaluation pip…
▽ More
Video generation models are increasingly capable of producing realistic videos, but they still struggle to generate videos that follow basic physical laws. Compounding this is a lack of reliable granular evaluation methods for localizing and specifying physical law violations in videos. We address this by introducing Physics Question Scene Graph (PQSG), a hierarchical question-based evaluation pipeline. PQSG evaluates generated videos by checking their faithfulness to a prompt across objects, actions, and adherence to physical laws using a graph-based hierarchy of questions generated by a vision-language model (VLM), guided by high-quality in-context examples. By representing questions as a graph, PQSG introduces logical dependencies within questions, ensuring that each query is contextually valid. Moreover, PQSG provides granular assessments of which qualities of the video violate physical plausibility constraints. We validate PQSG by creating FinePhyEval, a dataset with physics-based prompts and corresponding generated videos from diverse state-of-the-art video generation models (Sora 2, Veo 3, and Wan 2.1), with each video annotated across multiple categories by humans. Using FinePhyEval, we measure the correlation between PQSG's fine-grained scores and human judgments, showing higher overall correlations than prior work. We also find that PQSG ranks closed-source models higher than Wan 2.1 on physical realism. Lastly, we show that the annotations we provide in FinePhyEval can also be used for subtask evaluation: we benchmark two strong VLMs on generating and answering questions, finding that while models can create human-like questions, they still fall short of human performance in answering them.
△ Less
Submitted 23 June, 2026;
originally announced June 2026.
-
Only obscured yet luminous active galactic nuclei are closely associated with galaxy mergers: Direct observational evidence from type 2 active galactic nuclei
Authors:
Yongmin Yoon,
Yongjung Kim,
Dohyeong Kim,
Jaejun Cho,
Woowon Byun
Abstract:
To establish a more comprehensive understanding of the connection between galaxy mergers and active galactic nuclei (AGNs), it is essential to disentangle the contributions of intrinsic AGN luminosity and dust extinction to the merger-AGN connection. Since tidal features identified in deep images serve as direct evidence of recent mergers, we studied the fraction of AGN hosts with tidal features (…
▽ More
To establish a more comprehensive understanding of the connection between galaxy mergers and active galactic nuclei (AGNs), it is essential to disentangle the contributions of intrinsic AGN luminosity and dust extinction to the merger-AGN connection. Since tidal features identified in deep images serve as direct evidence of recent mergers, we studied the fraction of AGN hosts with tidal features ($f_T$) for a large sample of 748 type 2 AGNs at $z<0.063$. Specifically, we examined $f_T$ as a function of $E(B-V)$, derived from the Balmer decrement, and the internal-extinction-corrected luminosity of the [O III] $λ$5007 emission line ($L_{\text{[O III]}}$), which is a proxy for bolometric AGN luminosity. Our main finding is that $f_T$ is only significantly higher for AGNs that are simultaneously luminous and heavily dust-obscured. Specifically, AGNs with $\log L_{\text{[O III]}}\gtrsim41.5$ and $E(B-V)\gtrsim0.7$ exhibit a high $f_T$ of $\sim0.7$. In contrast, AGNs with either low luminosity ($\log L_{\text{[O III]}}\lesssim41.0$) or low dust obscuration ($E(B-V)\lesssim0.3$) show a low $f_T$ of $\lesssim0.2$. This trend suggests that galaxy mergers preferentially trigger AGNs that are simultaneously luminous and dust-obscured, whereas AGNs that are either luminous but unobscured or dust-obscured but less luminous are not strongly associated with merger-driven triggering. Based on several assumptions, our result can also be interpreted, despite certain caveats, within the framework of a merger-initiated evolution model for AGNs, suggesting that AGNs that are both obscured and luminous are temporally closer to merger events than those with lower luminosities and less dust obscuration.
△ Less
Submitted 23 June, 2026;
originally announced June 2026.
-
NGPS: Structure-Preserving Self-Supervised Denoising via Neighbor-Guided Patch Sampling
Authors:
Jaehyun Cho,
YoungJoon Yoo
Abstract:
Neighboring-slice self-supervised denoising is attractive for volumetric medical imaging, yet inter-slice misalignment breaks anatomical correspondence and often yields ghosting and blurred margins when adjacent slices are used naively as targets. We propose Neighbor-Guided Patch Sampling (NGPS), a lightweight framework that constructs neighboring supervision under local inter-slice misalignment w…
▽ More
Neighboring-slice self-supervised denoising is attractive for volumetric medical imaging, yet inter-slice misalignment breaks anatomical correspondence and often yields ghosting and blurred margins when adjacent slices are used naively as targets. We propose Neighbor-Guided Patch Sampling (NGPS), a lightweight framework that constructs neighboring supervision under local inter-slice misalignment without explicit registration. To avoid learning from misleading targets, prior methods commonly mask discrepant regions, but this stabilizes training at the cost of leaving a non-trivial portion of neighboring evidence unexploited, particularly around high-frequency anatomical boundaries. NGPS addresses this by decoupling structure matching from signal retrieval: for each masked location, it searches a local neighborhood for structurally similar candidate patches using a simple guide image (e.g., fast bilateral filtering), while retrieving the supervision signal directly from the raw noisy neighbor at the matched coordinates. By matching on a noise-attenuated guide while retrieving raw values from neighboring slices, NGPS constructs local pseudo targets without a learned registration module. Across the evaluated CT and synthetic-Rician MRI settings, NGPS improves fidelity and structure-sensitive metrics. Code is available at https://github.com/cv-cho/NGPS .
△ Less
Submitted 22 June, 2026;
originally announced June 2026.
-
Formalizing Task-Space Complexity for Zero-Shot Generalization
Authors:
Jung-Hoon Cho,
Heling Zhang,
Siqi Du,
Roy Dong,
Cathy Wu
Abstract:
Policies must operate across diverse conditions, yet a single policy is often conservative while fully adaptive schemes can be complex. We study zero-shot generalization in contextual dynamical systems and introduce a performance-centric, directional task dissimilarity--the signed divergence--that upper bounds the generalization gap from a source context to a target context. The signed divergence…
▽ More
Policies must operate across diverse conditions, yet a single policy is often conservative while fully adaptive schemes can be complex. We study zero-shot generalization in contextual dynamical systems and introduce a performance-centric, directional task dissimilarity--the signed divergence--that upper bounds the generalization gap from a source context to a target context. The signed divergence induces $\varepsilon$-tolerance sets that certify when a source policy class generalizes, and it yields a concrete notion of task-space complexity: the minimum number of source contexts needed so that every target context incurs at most $\varepsilon$ generalization gap. Under a mild local smoothness assumption on performance, the induced tolerance sets admit certified inner/outer balls and instance-dependent volume bounds on task-space complexity. In the finite-oracle setting, source selection reduces to set cover; a greedy strategy inherits the standard $H(n)$ approximation guarantee. Using a Mass-Spring-Damper system with linear-quadratic regulator (LQR) controllers and a nonlinear CartPole system with deep reinforcement learning controllers, we show that greedy selection achieves the same $\varepsilon$-coverage with fewer policies than uniform or random baselines. Our approach delivers a performance-based task similarity measure and practical certificates for building generalizable control with simple policies.
△ Less
Submitted 18 June, 2026;
originally announced June 2026.
-
Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement
Authors:
Kinam Kim,
Namiko Saito,
Heecheol Kim,
Katsushi Ikeuchi,
Jaegul Choo,
Yasuyuki Matsushita
Abstract:
Vision-Language-Action (VLA) models can generalize across diverse manipulation tasks, but their imitation-learning-based policies remain brittle in precise physical interactions due to compounding execution errors; Can a reinforcement learning policy trained purely in simulation improve the robustness of real-world VLAs zero-shot? Residual RL, which learns a corrective policy on top of a frozen VL…
▽ More
Vision-Language-Action (VLA) models can generalize across diverse manipulation tasks, but their imitation-learning-based policies remain brittle in precise physical interactions due to compounding execution errors; Can a reinforcement learning policy trained purely in simulation improve the robustness of real-world VLAs zero-shot? Residual RL, which learns a corrective policy on top of a frozen VLA, offers a natural framework, but existing approaches face a fundamental sim-to-real dilemma: privileged-state methods require lossy distillation for deployment; image-based methods suffer from the visual domain gap; and real-world RL is costly and unsafe. We propose an object-centric residual RL framework that refines VLA actions using object poses, enabling a compact observation space that transfers consistently between simulation and reality. To align the two domains, we additionally replay the same teleoperation demonstrations in simulation to train a sim counterpart of the real-world VLA. The residual RL policy is trained only in simulation with pose noise injection and dropout, and transfers zero-shot to the real robot. Across five manipulation tasks on a real Franka Research 3 (FR3) robot, our method improves the success rate from 42% to 76% zero-shot, and the improved rollouts can be further reused to retrain the base VLA for self-improvement without additional teleoperation. Project page: https://www.microsoft.com/en-us/research/articles/object-centric-residual-rl/
△ Less
Submitted 17 June, 2026;
originally announced June 2026.
-
Robust and Interpretable Adaptation of Equivariant Materials Foundation Models via Sparsity-promoting Fine-tuning
Authors:
Youngwoo Cho,
Seunghoon Yi,
Wooil Yang,
Sungmo Kang,
Young-woo Son,
Jaegul Choo,
Joonseok Lee,
Soo Kyung Kim,
Hongkee Yoon
Abstract:
Pre-trained materials foundation models, or machine learning interatomic potentials, leverage general physicochemical knowledge to effectively approximate potential energy surfaces. However, they often require domain-specific calibration due to physicochemical diversity as well as mismatches between practical computational settings and those used in constructing the pre-training data. To address t…
▽ More
Pre-trained materials foundation models, or machine learning interatomic potentials, leverage general physicochemical knowledge to effectively approximate potential energy surfaces. However, they often require domain-specific calibration due to physicochemical diversity as well as mismatches between practical computational settings and those used in constructing the pre-training data. To address this, we propose a sparsity-promoting fine-tuning method that selectively updates model parameters by exploiting the structural properties of E(3)-equivariant materials foundation models. On energy and force prediction tasks across molecular and crystalline benchmarks, our method matches or surpasses full fine-tuning and equivariant low-rank adaptation while updating only $\sim$3~\% of parameters, and in some cases as little as $\sim$0.5~\%. Beyond energy and force calibration, we further demonstrate task generalizability by applying our method to magnetic moment prediction and magnetism-aware total energy modeling. Finally, analysis of sparsity patterns reveals physically interpretable signatures, such as enhanced $d$-orbital contributions in transition metal systems. Overall, our results establish sparsity-promoting fine-tuning as a flexible and interpretable method for domain specialization of equivariant materials foundation models.
△ Less
Submitted 17 June, 2026;
originally announced June 2026.
-
Visualizing Uncertainty: Spatial Maps of Missing and Conflicting Evidence in Deep Learning
Authors:
Dong Hyun Jeong,
Feng Chen,
Jin-Hee Cho,
Lance M. Kaplan,
Audun Jøsang,
Soo-Yeon Ji
Abstract:
Understanding when and why deep neural networks are uncertain is crucial for deploying reliable machine learning systems in safety-critical domains. While existing uncertainty quantification methods provide scalar measures of model confidence, they offer limited insight into which spatial regions of an input contribute to different types of uncertainty. We propose a novel visualization framework,…
▽ More
Understanding when and why deep neural networks are uncertain is crucial for deploying reliable machine learning systems in safety-critical domains. While existing uncertainty quantification methods provide scalar measures of model confidence, they offer limited insight into which spatial regions of an input contribute to different types of uncertainty. We propose a novel visualization framework, Uncertainty Activation Map (UAM), that combines Evidential Deep Learning (EDL) with Full-Gradient Class Activation Mapping (FullGrad) to generate interpretable spatial uncertainty activation maps. Our approach distinguishes between two fundamental types of uncertainty: vacuity, representing lack of evidence, and dissonance, capturing conflicting evidence between competing hypotheses. By leveraging the complete gradient decomposition property of FullGrad and the principled uncertainty quantification of Subjective Logic, our method produces theoretically grounded visualizations that highlight specific image regions responsible for model uncertainty. With this framework, vacuity and dissonance activation maps are generated by computing belief-weighted attributions, enabling identification of where models lack knowledge versus where they encounter ambiguous evidence. Extensive evaluations across multiple benchmark datasets demonstrate that the proposed framework effectively addresses the critical gap between uncertainty quantification and explainability, providing intuitive visual feedback to assess model reliability in complex visual recognition tasks.
△ Less
Submitted 14 June, 2026;
originally announced June 2026.
-
AGORA: Can Deliberation and Governance Gates Absorb Participation Bias in Transit Planning?
Authors:
Jung-Hoon Cho,
Cathy Wu
Abstract:
Transit network design depends not only on the optimization algorithm but also on who shows up to the public hearing. Current practice often collects one-directional comments from self-selected attendees, leaving participant mix as an uncontrolled source of outcome variation. We present AGORA, a framework that holds the network, demand, and solver fixed while systematically varying meeting composi…
▽ More
Transit network design depends not only on the optimization algorithm but also on who shows up to the public hearing. Current practice often collects one-directional comments from self-selected attendees, leaving participant mix as an uncontrolled source of outcome variation. We present AGORA, a framework that holds the network, demand, and solver fixed while systematically varying meeting composition through stakeholder agents, structured deliberation, and governance gates. Across two standard benchmark networks at different scales, we find that (i) aggregate outcomes vary little across compositions, but on tail risk and fairness disparity, representative sampling still tends to outperform skewed compositions; (ii) without deliberation, composition produces no variation at all, showing that deliberation is the mechanism through which who attends affects outcomes; and (iii) governance gates compress cross-profile variance without shifting the average outcome on Mandl, but low acceptance on Mumford0 shows thresholds require instance-specific calibration. These findings reframe participation bias from an uncontrollable input to a process-design problem: even without guaranteed representative attendance, well-structured deliberation and governance criteria can substantially reduce how much outcomes depend on who is in the room.
△ Less
Submitted 31 May, 2026;
originally announced June 2026.
-
SheafStain: Sheaf-Theoretic Schrödinger Bridge for Spatially and Biologically Coherent Virtual Staining
Authors:
Hyeongyeol Lim,
Hongjun Yoon,
Eunjin Jang,
Daeky Jeong,
Won June Cho,
Hwamin Lee
Abstract:
Current virtual staining approaches offer the potential for time- and cost-efficient biomarker quantification in cancer diagnostics and prognostics. However, patch-wise inference for gigapixel whole slide images (WSIs) fails to maintain spatial continuity, yielding artifacts that cause catastrophic mismatches with ground-truth images. Although pathology Vision Foundation Models (VFMs) offer rich r…
▽ More
Current virtual staining approaches offer the potential for time- and cost-efficient biomarker quantification in cancer diagnostics and prognostics. However, patch-wise inference for gigapixel whole slide images (WSIs) fails to maintain spatial continuity, yielding artifacts that cause catastrophic mismatches with ground-truth images. Although pathology Vision Foundation Models (VFMs) offer rich representations, their self-attention causes varying global contexts to produce inconsistent embeddings for the same physical region. We formalize and validate this ``context contamination'' as a sheaf-theoretic problem where these embeddings form a presheaf that violates the gluing axiom. To address this, we propose SheafStain, a new approach that reinterprets VFM features as sheaf-like sections for spatially and biologically coherent virtual staining. Specifically, SheafStain integrates class and patch tokens into a Schrödinger Bridge framework as sheaf-like sections. While the class token anchors biological consistency, patch tokens form a per-position spatial map. A backbone co-pretrained on Hematoxylin \& Eosin (H\&E) and Immunohistochemistry (IHC) yields non-degenerate cross-stain stalks, so a single VFM feature space supervises both input conditioning and output stain alignment. Departing from prior work that evaluates on isolated $256 \times 256$ patches and either random-crops or resizes the $1024 \times 1024$ ground truth, we translate at $256 \times 256$ and evaluate on the stitched $1024 \times 1024$ outputs across HER2, ER, PR, and Ki-67. SheafStain demonstrates promising results against six prior methods while mitigating patch-boundary stitching artifacts. Code will soon be released.
△ Less
Submitted 10 June, 2026;
originally announced June 2026.
-
Heterophily-Aware Adaptive Knowledge Distillation for Hypergraph Neural Networks
Authors:
Joohee Cho,
David Yoon Suk Kang,
Yunyong Ko
Abstract:
Hypergraph knowledge distillation aims to retain the predictive performance of a hypergraph neural network (HNN) teacher while reducing inference costs through a lightweight student model. In this work, we observe that HNNs exhibit substantially lower prediction performance on heterophilic nodes connected through semantically diverse hyperedges, indicating that the reliability of teacher knowledge…
▽ More
Hypergraph knowledge distillation aims to retain the predictive performance of a hypergraph neural network (HNN) teacher while reducing inference costs through a lightweight student model. In this work, we observe that HNNs exhibit substantially lower prediction performance on heterophilic nodes connected through semantically diverse hyperedges, indicating that the reliability of teacher knowledge varies across nodes. Motivated by this observation, we propose HADES, a heterophily-aware adaptive distillation method for hypergraph neural networks. HADES quantifies node heterophily and leverages it as an estimate of teacher reliability to modulate the transfer of teacher knowledge during distillation. Experimental results on real-world hypergraphs demonstrate that HADES consistently improves student performance across different HNN teachers and distillation objectives. In many cases, the resulting student models surpass the predictive performance of their teachers while achieving up to 12.3 times faster inference.
△ Less
Submitted 7 June, 2026;
originally announced June 2026.