-
A Locally Tokenized Generative Model for Robust Time-Series Watermarking
Authors:
Dongbin Kim,
Geonwoo Shin,
Yujin Choi,
Soyeon Park,
Jaewook Lee
Abstract:
Watermarking is a central tool for provenance in generative models, yet its application to multivariate time series remains hindered by reliability failures under post-editing attacks. We show that existing detectors, which rely on globally coupled re-encoding, suffer from bidirectional drift of the null distribution: post-editing attacks can shift the z-score of non-watermarked samples in either…
▽ More
Watermarking is a central tool for provenance in generative models, yet its application to multivariate time series remains hindered by reliability failures under post-editing attacks. We show that existing detectors, which rely on globally coupled re-encoding, suffer from bidirectional drift of the null distribution: post-editing attacks can shift the z-score of non-watermarked samples in either direction, invalidating clean-calibrated thresholds. We argue that this instability is a property of the re-encoding, and that reliable detection requires each recovered unit to depend only on a bounded temporal neighborhood. Guided by this principle, we propose L-VQVAE, a generative model in which each discrete token is produced from a short contiguous window, and LVQMark, a watermarking method over this token space that combines logit-bias injection with robust re-encoding for attack-time detection. Experiments on four benchmarks spanning finance, energy, and neuroimaging show that our approach preserves generation quality while stabilizing both detection power and false-positive behavior under post-editing attacks.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
Proton emission half-lives and shape coexistence for $71 \leq Z \leq 83$ odd-$Z$ nuclei
Authors:
Yongbeom Choi,
Chang-Hwan Lee,
Youngman Kim
Abstract:
One-proton emission is a direct probe of nuclear structure near the proton drip line and plays a critical role in understanding exotic decay modes and nucleosynthesis processes. In this study, we investigate the half-lives of one-proton emitters for $71 \leq Z \leq 83$ odd-$Z$ nuclei by employing the WKB approximation with nuclear potentials obtained from the deformed relativistic Hartree-Bogoliub…
▽ More
One-proton emission is a direct probe of nuclear structure near the proton drip line and plays a critical role in understanding exotic decay modes and nucleosynthesis processes. In this study, we investigate the half-lives of one-proton emitters for $71 \leq Z \leq 83$ odd-$Z$ nuclei by employing the WKB approximation with nuclear potentials obtained from the deformed relativistic Hartree-Bogoliubov theory in continuum (DRHBc) and, for comparison, the relativistic continuum Hartree-Bogoliubov theory (RCHB). We first compare the calculated half-lives with available experimental data. The inclusion of quadrupole deformation via the DRHBc hardly contributes to improving the predictions of half-lives for the deformed nuclei. We find that all the studied nuclei exhibit ground states with $|β_{2,{\rm DRHBc}}| < 0.15$, and within this limited deformation range the spectroscopic factor provides the dominant contribution to the half-life, compared to the decay width. In particular, for nuclei exhibiting shape coexistence in DRHBc, such as $^{170}$Au, where the half-life varies significantly with the quadrupole deformation through its effect on the spectroscopic factor, we expect shape coexistence to exert a substantial influence on the variation of half-lives. Finally, we discuss the half-lives in consideration of shape coexistence. Our results indicate that the calculated half-life is governed not by the total-energy difference between coexisting minima but rather by the spectroscopic factor influenced by the deformation.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
SPADE: Self-Play in Adaptive Synthetic Executable Environments
Authors:
Bo Liu,
Simon Yu,
Yiding Jiang,
Ao Qu,
Andrew Zhao,
Zichen Liu,
Junsu Kim,
Zijian Zhou,
Seungone Kim,
Tongzheng Ren,
Mickel Liu,
Hanfei Yu,
Zhaorun Chen,
Weiyan Shi,
Paul Pu Liang,
Luke Zettlemoyer,
Yejin Choi,
Natasha Jaques
Abstract:
Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals. For language agents, existing training environment pools (hand-curated, statically synthesized, or frozen-verifier) keep the goal distribution fixed as the learner scales. We introduce SPADE (Self-Play in Adaptive Synthetic Executable Environments), a self-play RL framework in which a single LLM…
▽ More
Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals. For language agents, existing training environment pools (hand-curated, statically synthesized, or frozen-verifier) keep the goal distribution fixed as the learner scales. We introduce SPADE (Self-Play in Adaptive Synthetic Executable Environments), a self-play RL framework in which a single LLM plays two roles: an Environment Designer that writes complete, long-horizon training environments as executable code with an OpenAI Gym-style reset()/step() interface, and a Reasoning Agent that learns to act in them. Each is a stateful, multi-turn environment (state transitions, reward functions, and verification code), so one interface spans reasoning problems and multi-step agentic tool use. The Reasoning Agent's regret is estimated using the gap between its reward with and without privileged hints; in optimizing this regret signal the Environment Designer learns to target environments at the edge of the agent's capabilities while keeping them feasible. Through extensive experimentation, we find several components critical to success: grounding the Environment Designer on documents sampled from a large pretraining corpus, and giving it an accumulated environment memory. Scaling to 30B-parameter models, SPADE improves over the strongest fixed-environment baseline by +5.3 on average across eight held-out math, science, code, and reasoning benchmarks, and lifts the tool-use setting by +5.7 on BFCL-v4 multi-turn and +13.9 on ACEBench-Agent; on the games setting, the margin over the strongest baseline grows with model scale. By making environment design itself a learnable component, SPADE takes a concrete step toward open-ended self-improvement.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
Finite-energy weak solutions and relaxation for a compressible kinetic--fluid system with locally averaged Brinkman force
Authors:
Young-Pil Choi,
Roman Shvydkoy
Abstract:
We study a kinetic--fluid system in which a Vlasov or Vlasov--Fokker--Planck equation is coupled to the compressible Navier--Stokes equations with density-dependent viscosities through a locally averaged Brinkman force. The averaging is chosen in a conservative form so that the coupled system preserves the total momentum and satisfies a natural energy-dissipation balance, while avoiding the pointw…
▽ More
We study a kinetic--fluid system in which a Vlasov or Vlasov--Fokker--Planck equation is coupled to the compressible Navier--Stokes equations with density-dependent viscosities through a locally averaged Brinkman force. The averaging is chosen in a conservative form so that the coupled system preserves the total momentum and satisfies a natural energy-dissipation balance, while avoiding the pointwise evaluation of a possibly rough fluid velocity. The first main result is a conditional exponential relaxation estimate for sufficiently regular solutions. The estimate is proved under positive and negative density moment bounds and a Muckenhoupt $\calA_2$ condition on a power of the density, which replace uniform pointwise upper and lower bounds. The proof combines a modulated energy and hypocoercivity analysis with a compensating functional for the density fluctuation. A key ingredient is a weighted Calderón--Zygmund estimate for the associated elliptic corrector, which allows us to control the viscous contribution under the $\calA_2$ condition. The second main result concerns global finite-energy Bresch--Desjardins entropy weak solutions in admissible low-dimensional regimes. Using the additional entropy structure associated with the density-dependent viscosities, we verify the density assumptions required by the conditional relaxation theorem and obtain exponential relaxation for the weak solutions. In the diffusionless case, the particle distribution aligns toward a mono-kinetic state, while in the diffusive case the relaxation is towards a Maxwellian equilibrium.
△ Less
Submitted 13 August, 2026;
originally announced August 2026.
-
Early Exploration of the Scientific Discovery Space for the Habitable Worlds Observatory
Authors:
Courtney D. Dressing,
Danica Adams,
Evelyne Alecian,
Gagandeep Anand,
Giada Arney,
Sarah Gomes Aroucha Barbosa,
Martin Barstow,
Joanna K. Barstow,
Rachael L. Beaton,
Eduardo Bendek,
Svetlana Berdyugina,
Julie Biedermann,
Sarah Blunt,
Sanchayeeta Borthakur,
Kara Brugman,
Joseph N. Burchett,
Eric Burns,
Jenna M. Cann,
Ludmila Carone,
Cody A. Carr,
Richard Cartwright,
Renyue Cen,
Jean-yves Chaufray,
Pin Chen,
Lígia F Coelho
, et al. (302 additional authors not shown)
Abstract:
The Habitable Worlds Observatory (HWO) is a future NASA flagship mission concept identified by the Astro2020 Decadal Survey as the highest priority for large space missions. HWO should conduct "transformative astrophysics" and search for biosignatures in the atmospheres of approximately 25 potentially Earth-like planets. To further the early-stage development of HWO, NASA formed the Science, Techn…
▽ More
The Habitable Worlds Observatory (HWO) is a future NASA flagship mission concept identified by the Astro2020 Decadal Survey as the highest priority for large space missions. HWO should conduct "transformative astrophysics" and search for biosignatures in the atmospheres of approximately 25 potentially Earth-like planets. To further the early-stage development of HWO, NASA formed the Science, Technology, Architecture Review Team (START). In turn, START invited the scientific community to join working groups to explore the potential discovery space. In this paper, we present 70 science cases that resulted from this process. The cases address four scientific pillars: growth of galaxies (15 cases), evolution of the elements (13 cases), solar systems in context (32 cases), and living worlds (10 cases). Combined, they would address 27 of the 30 science questions and discovery areas identified by Astro2020. The 140 observing programs needed for the 70 investigations encompass a rich variety of spectroscopic (for 87% of science cases) and photometric (for 30%) observations extending from the UV to the NIR. Additionally, high-contrast and polarimetric capabilities would be needed for 34% and 27% of science cases, respectively. Access to UV wavelengths is critical: 83% of science cases need data at wavelengths <400 nm, and 26% extend to <100 nm. In the NIR, 26% of science cases need observations at wavelengths >=2000 nm. Pursuing the full portfolio of science would also necessitate precise astrometry for planet mass measurement, rapid response capabilities, a large instantaneous field of regard, non-sidereal tracking, saturation mitigation strategies, and high dynamic range.
△ Less
Submitted 11 August, 2026;
originally announced August 2026.
-
Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding
Authors:
Yeeun Choi,
Youngbeom Yoo,
Joon-Young Lee,
Hyolim Kang,
Seon Joo Kim
Abstract:
When videos extend from hours to days, directly processing them end-to-end becomes impractical for current Multi-modal Large Language Models (MLLMs). This ultra-long setting necessitates a two-stage paradigm: query-agnostic memory construction followed by retrieval-based inference. Prior work invests in complex memory construction to pre-model high-level relations in videos, despite not knowing th…
▽ More
When videos extend from hours to days, directly processing them end-to-end becomes impractical for current Multi-modal Large Language Models (MLLMs). This ultra-long setting necessitates a two-stage paradigm: query-agnostic memory construction followed by retrieval-based inference. Prior work invests in complex memory construction to pre-model high-level relations in videos, despite not knowing the downstream query at build time. We instead prioritize high-recall retrievability during memory building, and defer query-specific, high-level relation composition to inference time. To this end, we propose MERIT(Multi-key Episodic Retrieval with Inference-time Temporal expansion), a simple yet effective agentic framework for ultra-long video understanding. First, we formulate an episodic multi-key representation that enables precise retrieval of fine-grained memories through a simple key-matching mechanism. Second, we introduce a neighbor filtering mechanism to capture broader semantic context without the massive computational overhead of global memory construction. This is achieved by expanding the temporal scope exclusively around the retrieved segments at inference time. By leveraging simple key-matching with this on-demand temporal expansion, MERIT achieves state-of-the-art performance across three long-video benchmarks: EgoLifeQA, LVBench, and Video-MME (Long).
△ Less
Submitted 7 August, 2026;
originally announced August 2026.
-
Scylla: Observational Evidence for an Order of Magnitude in Dust Mass Opacity Evolution with ISM Density in the Large Magellanic Cloud
Authors:
Christina W. Lindberg,
Christopher J. R. Clark,
Claire E. Murray,
Julia Roman-Duval,
Caroline Bot,
Yumi Choi,
Roger E. Cohen,
Steven R. Goldman,
Karl D. Gordon,
Kristen B. W. McQuinn,
Elizabeth Tarantino,
Benjamin F. Williams,
Petia Yanchulova Merica-Jones,
Catherine Zucker
Abstract:
The emissivity of dust is known to vary greatly with radiative environment, density, grain chemistry, and geometry. Discrepancies between dust mass surface densities derived from far-infrared (FIR) emission and visible extinction persist across and within galaxies in the local Universe. Here, we use new extinction and emission measurements towards the LMC to show that this discrepancy is driven by…
▽ More
The emissivity of dust is known to vary greatly with radiative environment, density, grain chemistry, and geometry. Discrepancies between dust mass surface densities derived from far-infrared (FIR) emission and visible extinction persist across and within galaxies in the local Universe. Here, we use new extinction and emission measurements towards the LMC to show that this discrepancy is driven by the dust mass opacity evolving with the intrinsic density of the ISM, and that the ratio between FIR and optical dust mass opacity varies with gas surface density. These new findings imply that the dust mass opacity in the FIR could increase by nearly an order of magnitude (e.g., $κ_{160} = 0.3 - 6\ m^2\ kg^{-2}$) across over an order of magnitude of total hydrogen surface density $(Σ_H = 4 - 100 M_{\odot}\ pc^{-2})$, corroborating previous theoretical models for dust mass opacity evolution in the FIR, and providing new implications for emission-based dust mass estimates.
△ Less
Submitted 5 August, 2026;
originally announced August 2026.
-
When Absence Is Evidence: Evaluating Completeness-Sensitive Negative Reasoning in Large Language Models
Authors:
Byoungjae Min,
Kennedy Edemacu,
Sae-Hong Cho,
Yoonhyuk Choi,
Beakcheol Jang,
Jong Wook Kim
Abstract:
Large language models (LLMs) are often asked whether something is absent from a record, list, or retrieved context. Yet non-observation licenses a negative answer only when evidence completely covers the query scope; otherwise, the answer should remain unknown. We call this completeness-sensitive negative reasoning. We introduce CROWN-QA, comprising CROWN-Synth, a controlled paired core that fixes…
▽ More
Large language models (LLMs) are often asked whether something is absent from a record, list, or retrieved context. Yet non-observation licenses a negative answer only when evidence completely covers the query scope; otherwise, the answer should remain unknown. We call this completeness-sensitive negative reasoning. We introduce CROWN-QA, comprising CROWN-Synth, a controlled paired core that fixes the question and observed facts while varying only query-relative coverage, and CROWN-Real, a real-document contrast-set evaluation with controlled coverage variants. Across three LLM families, models show unstable closure judgments and substantial over-closure, failing to reliably distinguish a justified negative answer (Certified-Negative) from insufficient evidence (Unknown). The dominant CROWN-Synth failure is asymmetric: models often recognize implicitly complete evidence yet treat implicitly partial evidence as query-covering. Prompting redistributes errors between over- and under-closure rather than consistently resolving them. Structured certificate elicitation traces many errors to evidence-coverage mischaracterization. CROWN-Real shows that the core partial-coverage asymmetry persists on real-document content, while its strength and the balance between over- and under-closure vary by model, prompt, and source.
△ Less
Submitted 5 August, 2026;
originally announced August 2026.
-
K-EXAONE 2.0 Technical Report
Authors:
Eunbi Choi,
Kibong Choi,
Sehyun Chun,
Seokhee Hong,
Junwon Hwang,
Hyojin Jeon,
Ahra Jo,
Hyunjik Jo,
Yeonsik Jo,
Minhyeok Jung,
Doyoung Kim,
Heegyu Kim,
Joonkee Kim,
Seonghwan Kim,
Soyeon Kim,
Sunkyoung Kim,
Yireun Kim,
Yongil Kim,
Byungoh Ko,
Changhun Lee,
Dohaeng Lee,
Haeju Lee,
Jinsik Lee,
Kyungmin Lee,
Minwoo Lee
, et al. (52 additional authors not shown)
Abstract:
This technical report presents K-EXAONE 2.0, an open-weight multilingual foundation model developed by LG AI Research as a step in our effort toward global frontier-scale foundation models. Rather than training from scratch, we upcycle K-EXAONE and expand its architecture, yielding a Mixture-of-Experts (MoE) model with 750B total parameters and approximately 37B activated per token---more than thr…
▽ More
This technical report presents K-EXAONE 2.0, an open-weight multilingual foundation model developed by LG AI Research as a step in our effort toward global frontier-scale foundation models. Rather than training from scratch, we upcycle K-EXAONE and expand its architecture, yielding a Mixture-of-Experts (MoE) model with 750B total parameters and approximately 37B activated per token---more than three times the capacity of its predecessor. K-EXAONE 2.0 supports context lengths of up to 256K tokens and expands multilingual coverage from six to ten languages. Its training pipeline combines continual pre-training, difficulty-focused mid-training, and post-training to strengthen reasoning, agentic coding, multilingual capability, and safety grounded in Korean sociocultural contexts. Across nine evaluation categories selected to reflect the conditions of practical use, K-EXAONE 2.0 improves over K-EXAONE and remains competitive with open-weight models, showing its largest gains in agentic coding and long-context understanding and its clearest strengths in long-context retrieval and safety. Released under the Apache 2.0 license, K-EXAONE 2.0 enables the wider AI ecosystem to evaluate, deploy, adapt, and build upon it, while marking the beginning---rather than the endpoint---of our challenge toward the global frontier.
△ Less
Submitted 5 August, 2026;
originally announced August 2026.
-
Not All Redundant Tokens Are Alike: Analyzing Visual Token Pruning through Token Roles
Authors:
Hyeonyu Kim,
Sehwan Lim,
Youngwon Choi,
Taeyoun Kwon,
Jaejin Kim
Abstract:
Vision-language models (VLMs) process an image as a sequence of visual tokens, which creates a substantial computational bottleneck during inference. Recent visual token pruning methods address this issue by removing seemingly redundant tokens, yet it remains unclear how these pruning decisions relate to the functional roles of visual tokens. In this work, we analyze visual token pruning through t…
▽ More
Vision-language models (VLMs) process an image as a sequence of visual tokens, which creates a substantial computational bottleneck during inference. Recent visual token pruning methods address this issue by removing seemingly redundant tokens, yet it remains unclear how these pruning decisions relate to the functional roles of visual tokens. In this work, we analyze visual token pruning through the lens of token roles identified by EmbedLens. We first show that representative pruning methods exhibit distinct token-role biases, but these biases do not directly correlate with downstream performance. To better understand this behavior, we refine the token-role assignment procedure and evaluate role-protected pruning variants. Our results show that preserving non-alive tokens can sometimes maintain or improve performance, suggesting that tokens with weak direct semantic alignment may still affect model behavior under pruning. Our code is publicly available at https://github.com/jaykim9870/Not_All_Redundant_Tokens_Are_Alike.
△ Less
Submitted 5 August, 2026;
originally announced August 2026.
-
DP-MemView: A Memory Interface for Attribute-Level Transcript Privacy in Long-Term LLM Agents
Authors:
Jong Wook Kim,
Byoungjae Min,
Kennedy Edemacu,
Yoonhyuk Choi,
Sae-Hong Cho,
Beakcheol Jang
Abstract:
Long-term memory enables persistent personalization in LLM agents, but repeated memory-conditioned responses can cumulatively reveal protected attributes even when they are never stated explicitly. We formalize this threat as adaptive transcript privacy and introduce DP-MemView, a differentially private interface that privately selects public response-conditioning views and exposes those views---r…
▽ More
Long-term memory enables persistent personalization in LLM agents, but repeated memory-conditioned responses can cumulatively reveal protected attributes even when they are never stated explicitly. We formalize this threat as adaptive transcript privacy and introduce DP-MemView, a differentially private interface that privately selects public response-conditioning views and exposes those views---rather than raw memory---to the response LLM. Each private selection is charged to every protected attribute whose memory group intersects the read set. Per-attribute ledgers block any selection that would exceed its cap and return a fixed generic view instead. Under an explicit interface contract, we prove pure B_a-DP for the entire adaptive transcript. We also extend the result to stores that differ across multiple protected groups and bound how much observing the transcript can change an adversary's prior odds. We evaluate the online and preallocated modes with three response LLMs on a controlled adjacent-store benchmark and a public-corpus transfer track. Both modes keep transcript distinguishability near chance while preserving target-required personalization and overall response quality. Further diagnostics show that removing key safeguards causes mismatched output support, missing ledger charges, revealing side channels, or growing long-horizon leakage.
△ Less
Submitted 4 August, 2026;
originally announced August 2026.
-
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity
Authors:
Yongxi Zhou,
Junwei Yao,
Yuanzhe Liu,
Zihan Dong,
Wenbo Ye,
Jiaxi Wen,
Lai Yun Choi
Abstract:
A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is faithful: when an item's intent is held fixed and only its meaning-preserving surface form varies, does the canonical-form score estimate model behavior well, and how much of any variation is decoding/judge noise rather than signal? We instantiate thi…
▽ More
A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is faithful: when an item's intent is held fixed and only its meaning-preserving surface form varies, does the canonical-form score estimate model behavior well, and how much of any variation is decoding/judge noise rather than signal? We instantiate this in safety, a high-stakes setting with no gold label to average toward. To avoid prior confounds, we pre-author the reformulations (refusal-free, mostly non-LLM: machine back-translation and a Matrix-Language-Frame code-switch generator) so an identical surface form reaches every model, score all responses with one human-anchored, vendor-neutral judge (Claude, kappa = 0.86 vs. human on unsafe compliance, stable across languages, cross-checked by GPT-4o), and verify intent preservation. On 370 seeds x 5 surface forms x 5 models, no single transformation is uniformly most dangerous (6 of 20 per-transformation McNemar tests survive correction, most protective). Yet evaluating only the canonical prompt underestimates unsafe compliance: the union of unsafe outcomes across forms exceeds even the worst single form by 3.3-12.9 pp, with bootstrap 95% CIs excluding zero for all five models, and 5-13% of seeds safe on canonical are unsafe under some reformulation -- above a zero stochasticity floor (canonical resampled five times at temperature 0 gives 0/370 new exposures). The size of this gap is model-dependent (largest on Gemini 2.5 Pro). One form recovers only ~53% of a model's observed unsafe surface and about three reach 85% -- a redundancy characterization of this form set, not of a defined population. A benign control (XSTest) suggests the instability is bidirectional, though the benign and harmful pools are not item-matched. We release the dataset, code, and per-response labels.
△ Less
Submitted 1 August, 2026;
originally announced August 2026.
-
When Discovery Becomes a Storm: A ROS 2 Discovery Model for Wireless Robotic Networks
Authors:
Yeonwoo Choi,
Sanghoon Lee,
Kyung-Joon Park
Abstract:
In Robot Operating System 2 (ROS 2), Data Distribution Service (DDS) participants must discover one another before exchanging data. In wireless environments, delayed or lost discovery messages cause reliability timers to expire, triggering retransmissions that intensify channel contention and further delay the delivery of discovery messages. This self-reinforcing feedback can escalate into a disco…
▽ More
In Robot Operating System 2 (ROS 2), Data Distribution Service (DDS) participants must discover one another before exchanging data. In wireless environments, delayed or lost discovery messages cause reliability timers to expire, triggering retransmissions that intensify channel contention and further delay the delivery of discovery messages. This self-reinforcing feedback can escalate into a discovery storm. Existing models characterize discovery demand under fixed delivery conditions, but do not capture how shared-channel delay changes protocol state and generates further traffic. To address this issue, we present the first closed-loop analytical model of ROS 2 discovery that characterizes how delay-induced feedback amplifies retransmission overhead and leads to severe discovery storms. Our model represents channel contention as a shared service process, coupling message-delivery latency with receiver states and reliability timers. The model predicts both discovery completion time and per-class message counts. We validate the model through 1,350 experimental runs across 90 topology configurations. An open-loop airtime baseline captures only a fraction of the high-load completion time. The closed-loop model reproduces this rise and conservatively upper-bounds the observed high-load range. Guided by insights from the model, we further design a response-aware discovery policy that reduces mean discovery completion time by 25.3% to 39.7%.
△ Less
Submitted 3 August, 2026;
originally announced August 2026.
-
Correlated frailty model for analysis of genetic association in family studies
Authors:
Agnieszka Krol,
Virginie Rondeau,
Yun-Hee Choi,
Laurent Briollais
Abstract:
Family-based study designs allow the investigation of gene mutation effects on a disease risk by considering related family members. Some methods have been developed for testing sets of genetic variants in family studies but only very few can handle right-censored time-to-event data. We propose here a correlated frailty model for the analysis of a survival outcome related to cancer in presence of…
▽ More
Family-based study designs allow the investigation of gene mutation effects on a disease risk by considering related family members. Some methods have been developed for testing sets of genetic variants in family studies but only very few can handle right-censored time-to-event data. We propose here a correlated frailty model for the analysis of a survival outcome related to cancer in presence of familial correlations. These familial correlations are explained by a residual familial component specified by a kinship matrix and a region- or gene-based specific correlation structure modeled via identical-by-descent (IBD) probability matrix. The proposed approach is used to quantify and evaluate the association between a set of common single nucleotide polymorphism (SNPs) or rare variants (or both) from the same genomic region and a survival outcome, e.g. time to disease onset. The model's marginal likelihood is maximized using the Marquardt algorithm. We evaluated the method by simulations under various scenarios where we varied the family size, the strength of genetic associations from multiple rare variants and the presence or not of residual familial correlation. The results indicate that the correlated frailty model can be valuable in family cancer studies, for example to identify genomic regions significantly associated with the time to cancer onset.
△ Less
Submitted 3 August, 2026;
originally announced August 2026.
-
KoVRE: Training an Efficient Embedding Model for Korean Visual Document Retrieval
Authors:
Yongbin Choi,
Gyuho Shim,
Youngjoon Jang
Abstract:
Visual Document Retrieval (VDR) directly matches text queries against document images, preserving visual and structural information that may be lost during text extraction. However, existing VDR models and training resources remain predominantly English-centric, while many high-performing systems rely on massive backbones or storage-intensive multi-vector representations. To address these limitati…
▽ More
Visual Document Retrieval (VDR) directly matches text queries against document images, preserving visual and structural information that may be lost during text extraction. However, existing VDR models and training resources remain predominantly English-centric, while many high-performing systems rely on massive backbones or storage-intensive multi-vector representations. To address these limitations, we introduce KoVRE: Korean Visual Document Retrieval Embedding, a single-vector retriever for Korean visual documents, alongside a comprehensive training recipe. We train the model on 708,729 Korean and English query-page pairs using positive-aware hard-negative mining and conduct controlled analyses of training-data composition, hard-negative treatment, and reranker-based knowledge distillation. Across Korean visual document retrieval benchmarks, our 2B model substantially improves over the base backbone model, outperforming both its 8B single-vector counterpart and a strong multi-vector baseline. These results demonstrate that targeted bilingual supervision and our carefully designed training strategies can produce a highly effective Korean VDR model across diverse document domains, without requiring a scaled-up backbone or multi-vector representations.
△ Less
Submitted 2 August, 2026;
originally announced August 2026.
-
ABC methods for IoT Emitter Geolocalisation using LEO Satellite Doppler Measurements
Authors:
B. Ristic,
Y. Choi,
D. Y. Kim,
A. Hourani
Abstract:
We address the problem of passive localisation of a stationary, ground-level IoT radio emitter using Doppler frequency measurements collected by low-Earth orbit (LEO) satellites during an observation window. The problem is challenging because radio emission from low-cost IoT devices is affected by various compounding sources of measurement error, that collectively render the likelihood function in…
▽ More
We address the problem of passive localisation of a stationary, ground-level IoT radio emitter using Doppler frequency measurements collected by low-Earth orbit (LEO) satellites during an observation window. The problem is challenging because radio emission from low-cost IoT devices is affected by various compounding sources of measurement error, that collectively render the likelihood function intractable in a closed form. Hence, we apply and investigate the performance of Approximate Bayesian Computation (ABC) methods for this task. Numerical results demonstrate the statistical and computational performance of two ABC methods, rejection sampling ABC and sequential Monte Carlo ABC.
△ Less
Submitted 30 July, 2026;
originally announced July 2026.
-
LightRot: A Light-Weighted Rotation Scheme and Architecture for Accurate Low-Bit Large Language Model Inference
Authors:
Sangjin Kim,
Yuseon Choi,
Jungjun Oh,
Byeongcheol Kim,
Hoi-Jun Yoo
Abstract:
As large language models (LLMs) continue to demonstrate exceptional capabilities across various domains, the challenge of achieving energy-efficient and accurate inference becomes increasingly critical. This work presents LightRot, a lightweight rotation scheme and dedicated hardware accelerator designed for low-bit LLM inference. The proposed architecture integrates Grouped Local Rotation (GLR) a…
▽ More
As large language models (LLMs) continue to demonstrate exceptional capabilities across various domains, the challenge of achieving energy-efficient and accurate inference becomes increasingly critical. This work presents LightRot, a lightweight rotation scheme and dedicated hardware accelerator designed for low-bit LLM inference. The proposed architecture integrates Grouped Local Rotation (GLR) and Outlier Direction Aligning (ODA) algorithms with a hierarchical Fast Hadamard Transform (FHT)-based rotation unit to address key challenges in low-bit quantization, including the energy overhead of rotation operations. The proposed accelerator, implemented in a 28nm CMOS process, achieves a peak energy efficiency of 27.4 TOPS/W for 4-bit inference, surpassing prior state-of-the-art designs. Unlike conventional approaches that rely on higher-precision inference or evaluate on basic language modeling tasks like GPT-2, LightRot is optimized for advanced models such as LLaMA2-13B and LLaMA3-8B. Its performance is further validated on MT-Bench, demonstrating robust applicability to real-world conversational scenarios and redefining benchmarks for chat-based AI systems. By synergizing algorithmic innovations and hardware efficiency, this work sets a new paradigm for scalable, low-bit LLM inference, paving the way for sustainable AI advancements.
△ Less
Submitted 30 July, 2026;
originally announced July 2026.
-
GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference
Authors:
Sangjin Kim,
Yuseon Choi,
Byeongcheol Kim,
Jungjun Oh,
Hoi-jun Yoo
Abstract:
Low-bit quantization is essential for efficient LLM inference, and both rotation and fine-grained group quantization have shown individual promise. However, their combination often leads to accuracy degradation or hardware overhead due to a mismatch between the global nature of rotation and the localized behavior of group scaling. We propose GyRot, a quantization framework and hardware accelerator…
▽ More
Low-bit quantization is essential for efficient LLM inference, and both rotation and fine-grained group quantization have shown individual promise. However, their combination often leads to accuracy degradation or hardware overhead due to a mismatch between the global nature of rotation and the localized behavior of group scaling. We propose GyRot, a quantization framework and hardware accelerator that bridges this gap through algorithm-hardware co-design. GyRot introduces Coarse Rotation, Fine Grouping (CoRFiG) and Harmonic-Aligned Permutation (HAP) to enable cooperative integration of rotation and group quantization, enhancing quantizability while relaxing scaling factor precision. To further reduce hardware cost, we reformulate asymmetric quantization and introduce a zero-point rounding strategy that enables fully integer dequantization. Implemented on an INT4-based tensor PE architecture, GyRot achieves state-of-the-art 4-bit accuracy across LLaMA-family models, while delivering up to 3.4x speedup and 3.6x energy efficiency over baseline LLM accelerators. These results validate GyRot's practical effectiveness for scalable and energy-efficient LLM deployment.
△ Less
Submitted 30 July, 2026;
originally announced July 2026.
-
Incompressible Navier-Stokes-Fourier limit from a nonlinear quantum Fokker-Planck equation
Authors:
Young-Pil Choi,
Byung-Hoon Hwang,
Ju-Hwan Hyun
Abstract:
We derive the incompressible Navier-Stokes-Fourier limit from a nonlinear quantum Fokker-Planck equation with Bose-Einstein or Fermi-Dirac statistics. The model has a self-consistent collision structure, with the local density acting as the collision frequency and the bulk velocity and temperature determined by nonlinear quantum-weighted moments of the distribution. We work near a global quantum e…
▽ More
We derive the incompressible Navier-Stokes-Fourier limit from a nonlinear quantum Fokker-Planck equation with Bose-Einstein or Fermi-Dirac statistics. The model has a self-consistent collision structure, with the local density acting as the collision frequency and the bulk velocity and temperature determined by nonlinear quantum-weighted moments of the distribution. We work near a global quantum equilibrium under the diffusive scaling and keep the quantum parameter fixed. Uniform estimates with respect to the Knudsen number yield strong microscopic relaxation and identify the limiting infinitesimal quantum equilibrium. Using the local conservation laws, we prove the incompressibility condition, the Boussinesq relation, and strong compactness of the divergence-free velocity component and a quantum-adapted thermal mode, while the acoustic modes vanish locally by a dispersive estimate. The limiting viscous stress tensor and heat flux are identified by solving auxiliary equations for the linearized quantum Fokker-Planck operator and by expanding the local quantum equilibrium manifold. The resulting incompressible Navier-Stokes-Fourier system retains the effect of quantum statistics through its normalization constants and transport coefficients.
△ Less
Submitted 29 July, 2026;
originally announced July 2026.
-
SKY-Piano: A Multimodal Piano Performance Dataset
Authors:
Joonhyung Bae,
Dawon Park,
Taegyun Kwon,
Yoon-Seok Choi,
Hyeon Hur,
Satoshi Obata,
Shigeru Kai,
Yohei Wada,
Yu Takahashi,
Akira Maezawa,
Jaebum Park,
Jonghwa Park,
Juhan Nam
Abstract:
Music information retrieval research on piano performance increasingly involves diverse modalities of data and annotations beyond audio and MIDI. We present SKY-Piano, a multimodal piano performance dataset that includes 11 hours of performance recordings of motion, multi-view video, audio, MIDI from 7 professional and 12 amateur pianists along with MusicXML scores. The performance pieces were sel…
▽ More
Music information retrieval research on piano performance increasingly involves diverse modalities of data and annotations beyond audio and MIDI. We present SKY-Piano, a multimodal piano performance dataset that includes 11 hours of performance recordings of motion, multi-view video, audio, MIDI from 7 professional and 12 amateur pianists along with MusicXML scores. The performance pieces were selected considering playing technique, difficulty, and performer expertise on a shared core repertoire. The motion data include both hand and body motion, released in both flagged form, where samples lost to marker occlusion are marked as unreliable, and imputed form, where those gaps are reconstructed, together with Visual3D body-segment kinematics and other time-synchronized modalities. To easily browse different modalities of data at a glance, we provide an interactive web browser. In addition, we developed a fingering annotation model and tool for deriving pseudo fingering annotations from the MIDI and motion data. Lastly, we present MIDI-to-motion generation through a fine-tuning experiment as a use case of the dataset.
△ Less
Submitted 29 July, 2026;
originally announced July 2026.
-
Nonlinear quantum Fokker-Planck equation near equilibrium
Authors:
Young-Pil Choi,
Byung-Hoon Hwang,
Ju-Hwan Hyun
Abstract:
We investigate a nonlinear quantum Fokker--Planck equation with self-consistent collision frequency, bulk velocity, and temperature. In contrast to quantum Fokker--Planck equations with prescribed diffusion and friction coefficients, the macroscopic quantities are nonlinear functionals of the distribution function. The equation preserves mass, momentum, and kinetic energy, admits a quantum entropy…
▽ More
We investigate a nonlinear quantum Fokker--Planck equation with self-consistent collision frequency, bulk velocity, and temperature. In contrast to quantum Fokker--Planck equations with prescribed diffusion and friction coefficients, the macroscopic quantities are nonlinear functionals of the distribution function. The equation preserves mass, momentum, and kinetic energy, admits a quantum entropy dissipation structure, and propagates the Pauli admissible range in the fermionic case. Its collision operator is also formally connected to the quantum Landau equation. For the Cauchy problem in the three-dimensional whole space, we prove the global-in-time existence and uniqueness of strong solutions near a global quantum equilibrium. The proof is based on a perturbative macro--micro energy method that combines microscopic coercivity, estimates for nonlinear velocity moments, and a macroscopic dissipation argument. We further establish the propagation of nonnegativity and the fermionic Pauli upper bound. Under an additional negative Sobolev assumption on the initial perturbation, we obtain algebraic decay rates toward equilibrium.
△ Less
Submitted 28 July, 2026;
originally announced July 2026.
-
A Deep Look at the Ultra-Faint Milky Way Satellite Virgo III with Rubin Observatory Data Preview 2
Authors:
Aashay Pai,
Alex Drlica-Wagner,
Peter S. Ferguson,
William Cerny,
Chin Yi Tan,
Jeffrey L. Carlin,
Yumi Choi,
Željko Ivezić,
Lauren A. MacArthur,
Colin T. Slater,
Dan S. Taranu,
Brian Yanny
Abstract:
We analyze the ultra-faint Milky Way satellite Virgo III using data from the Vera C. Rubin Observatory Data Preview 2 (DP2). Virgo III was observed in the Rubin "Cosmic Treasure Chest" (M49) First Look field, which contains 924 visits in the u,g,r,i bands comprising ~10.5hrs of exposure time with LSSTCam. These data are considerably deeper than the majority of DP2, with a $5σ$ limiting magnitude t…
▽ More
We analyze the ultra-faint Milky Way satellite Virgo III using data from the Vera C. Rubin Observatory Data Preview 2 (DP2). Virgo III was observed in the Rubin "Cosmic Treasure Chest" (M49) First Look field, which contains 924 visits in the u,g,r,i bands comprising ~10.5hrs of exposure time with LSSTCam. These data are considerably deeper than the majority of DP2, with a $5σ$ limiting magnitude that approaches the expected 10-year depth of LSST (~25.2-26.5mag, depending on band). We report the morphological and stellar population parameters of Virgo III measured with the maximum-likelihood-based package ugali. The depth of the Rubin imaging yields more than a factor of four increase in the number of candidate member stars ($N_* = 114^{+11}_{-11}$) relative to the Virgo III discovery results ($N_* = 25^{+5}_{-4}$), enabling significantly more precise morphological constraints. Our best-fit parameters broadly agree with previous measurements, further confirming that Virgo III has properties that are consistent with an ultra-faint dwarf galaxy ($M_V = -2.72^{+0.49}_{-0.70}$; $r_{1/2} = 53^{+10}_{-8}$) located at a heliocentric distance of $D_\odot = 151^{+8}_{-8}$. We also demonstrate that the depth and photometric quality of the DP2 data are sufficient to separate metal-poor and metal-rich stars in color-color space. We further present period estimates for the three known RR Lyrae member stars derived from the DP2 forced photometry; we use theoretical Period-Luminosity-Metallicity (PLZ) and Period-Wesenheit-Metallicity (PWZ) relations to obtain independent distance estimates. We find that our period and distance estimates are broadly consistent with previous measurements for these RR Lyrae. These results demonstrate the power of LSST data for the discovery and characterization of ultra-faint dwarf galaxies and motivate future searches for new satellites across the southern sky.
△ Less
Submitted 14 August, 2026; v1 submitted 27 July, 2026;
originally announced July 2026.
-
Learning When to Reason for Text-to-SQL via SFT and DPO
Authors:
Soohyuk Jang,
Jiheum Yeom,
Nohil Park,
Sang Hun Kim,
Yoonyoung Choi,
Kiwook Bae,
Sungroh Yoon
Abstract:
Recent Text-to-SQL methods rely heavily on reasoning-centric paradigms such as Chain-of-Thought (CoT), achieving substantial gains on complex benchmarks at the cost of high inference-time overhead. However, a large fraction of real-world queries are simple lookups or aggregations that can be resolved without multi-step deduction, making forced reasoning wasteful. Thus, we propose AutoThinkSQL, a f…
▽ More
Recent Text-to-SQL methods rely heavily on reasoning-centric paradigms such as Chain-of-Thought (CoT), achieving substantial gains on complex benchmarks at the cost of high inference-time overhead. However, a large fraction of real-world queries are simple lookups or aggregations that can be resolved without multi-step deduction, making forced reasoning wasteful. Thus, we propose AutoThinkSQL, a framework that integrates an auto-thinking mechanism into both Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO) on Text-to-SQL. Our approach enables the model to dynamically bypass reasoning for simple queries while invoking deep CoT for complex queries. On Qwen3-Coder-30B-A3B, our method achieves consistent gains compared to the best counterpart baseline on both Spider and BIRD benchmarks while simultaneously reducing average output tokens by 24.6% and 18.3%, and average latency by 17.1% and 11.5% compared to CoT-only generation. Further analysis indicates that the model learns to align its reasoning decisions with query difficulty.
△ Less
Submitted 17 June, 2026;
originally announced July 2026.
-
Solar Open 2 Technical Report
Authors:
Sungrae Park,
Sanghoon Kim,
Gyoungjin Gim,
Jungho Cho,
Hyunwoong Ko,
Minbyul Jeong,
Minjeong Kim,
Keunwoo Choi,
Chaehun Shin,
Chanwoong Yoon,
Dongjun Kim,
Eunwon Kim,
Gyungin Shin,
Hyeonju Lee,
Hyungkyu Kang,
Inseo Song,
Jisu Bae,
Jiyoon Han,
Jiyun Lee,
Joonkee Kim,
Junyeop Lee,
Mikyoung Cha,
Sangwon Yu,
Sehwan Joo,
Seokyoon Kang
, et al. (28 additional authors not shown)
Abstract:
We present Solar Open 2, a 250B-A15B Mixture-of-Experts language model built for long-horizon agentic tasks, scaled up from Solar Open 1 (Solar Open 100B). To hold entire agent trajectories in a single context, Solar Open 2 reaches a 1M-token window through a hybrid attention stack that interleaves one softmax layer among every three linear-attention layers, using no positional encoding and a gate…
▽ More
We present Solar Open 2, a 250B-A15B Mixture-of-Experts language model built for long-horizon agentic tasks, scaled up from Solar Open 1 (Solar Open 100B). To hold entire agent trajectories in a single context, Solar Open 2 reaches a 1M-token window through a hybrid attention stack that interleaves one softmax layer among every three linear-attention layers, using no positional encoding and a gated delta rule extended to negative eigenvalues. To train at this scale under a fixed compute budget, we make training efficient in two ways: a stronger starting point, and higher-value data. For the starting point, we initialize Solar Open 2 from Solar Open 1, transferring the 5.69B-parameter shared skeleton that survives the architectural change and learning everything else through full pre-training. For the data, we curate for value per token: quality- and rarity-aware data curation and mixture-ratio optimization refine a 20T pool into a 10T mixture that, at equal token budget, outperforms the Solar Open 1 recipe. To build its agent skills, we train twelve domain specialists across purpose-built scenarios, then consolidate them into a single model by Multi-teacher On-Policy Distillation (MOPD). Against comparably sized open-weight models on English benchmarks, Solar Open 2 leads on MMLU-Pro, LiveCodeBench, and the APEX-Agents agentic suite, and stays competitive with the strongest (DeepSeek-V4-Flash and MiMo-V2.5) elsewhere. On Korean benchmarks, Solar Open 2 records the highest average of any model compared, including fast-tier closed APIs, and on Ko-GDPval, an in-house Korean officework-agent benchmark, it is competitive with DeepSeek-V4-Pro (1.6T) at less than a sixth of its size.
△ Less
Submitted 23 July, 2026; v1 submitted 22 July, 2026;
originally announced July 2026.
-
A Search for Exoplanets around Northern Circumpolar Stars X. The origin of radial velocity variations in the evolved star HD 216595
Authors:
Sang-Hee Kim,
Byeong-Cheol Lee,
Shenghong Gu,
Jae-Rim Koo,
Beomdu Lim,
Myeong-Gu Park,
Huan-Yu Teng,
Yeon-Ho Choi,
David Mkrtichian,
Tae-Yang Bang,
Hyeong-Ill Oh,
Heon-Young Chang
Abstract:
Detecting planetary companions around evolved stars, particularly asymptotic giant branch (AGB) stars, is challenging due to intrinsic stellar variability such as surface convection, pulsations, and mass loss, which can produce radial velocity (RV) signals that mimic Keplerian motion. We investigate the origin of long-period, low-amplitude RV variations observed in the AGB star HD 216595 based on…
▽ More
Detecting planetary companions around evolved stars, particularly asymptotic giant branch (AGB) stars, is challenging due to intrinsic stellar variability such as surface convection, pulsations, and mass loss, which can produce radial velocity (RV) signals that mimic Keplerian motion. We investigate the origin of long-period, low-amplitude RV variations observed in the AGB star HD 216595 based on high-resolution spectroscopic data spanning approximately 16 years obtained with the Bohyunsan Optical Astronomy Observatory Echelle Spectrograph (BOES) and the Las Cumbres Observatory Network of Robotic Echelle Spectrographs (NRES) instruments. The RV measurements reveal a statistically significant periodic signal at 567 days that can be described by a Keplerian model consistent with a substellar companion. However, no strong correlations are found between the RV variations and stellar activity indicators, including line bisectors, chromospheric activity, and photometric variability, although weak signals at similar timescales are present in some diagnostics. Given the stellar properties of HD 216595 and similarities to previously reported cases, the observed RV variations are likely related to intrinsic stellar processes, although a companion-induced origin cannot be definitively ruled out. Further progress will require improved diagnostics and more sophisticated modeling to disentangle stellar variability from genuine orbital signals.
△ Less
Submitted 22 July, 2026;
originally announced July 2026.
-
ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding
Authors:
Keuntae Kim,
Beomseok Lee,
Hyunwoo Kim,
Yong Suk Choi
Abstract:
Vision Language Models (VLMs) achieve strong reasoning with Chain-of-Thought (CoT) prompting but incur high sequential-generation cost, error accumulation, and limited self-correction. Diffusion Multimodal Large Language Models (dMLLMs) unmask tokens in an order-agnostic process, improving efficiency and enabling iterative refinement, yet their reasoning and how to enhance it remain underexplored.…
▽ More
Vision Language Models (VLMs) achieve strong reasoning with Chain-of-Thought (CoT) prompting but incur high sequential-generation cost, error accumulation, and limited self-correction. Diffusion Multimodal Large Language Models (dMLLMs) unmask tokens in an order-agnostic process, improving efficiency and enabling iterative refinement, yet their reasoning and how to enhance it remain underexplored. We propose a training-free method, Spatio-Temporal Token Veto (ST-Veto), which leverages the ability to observe all token positions at each diffusion step. Rather than relying only on current-step confidence, ST-Veto vetoes temporally unstable tokens via second-order Taylor prediction of confidence dynamics and filters weakly grounded tokens using image-attention mass, swapping them with safer candidates. Across multiple dMLLMs and multimodal reasoning benchmarks, ST-Veto consistently outperforms standard decoding policies and prior VLM reasoning methods, improving accuracy by up to 9% with no additional training or generation cost. Analyses show that ST-Veto steers generation toward higher-confidence, better-grounded paths.
△ Less
Submitted 20 July, 2026;
originally announced July 2026.
-
DRIFT: Difficulty-aware Rectified Flows for Through-plane MRI Super-Resolution
Authors:
Yoonseok Choi,
Eun-Gyu Ha,
Daniel Kim,
Mohammed A. Al-masni,
Ming-Hsuan Yang,
Dong-Hyun Kim
Abstract:
Magnetic Resonance Imaging (MRI) is often acquired with anisotropic resolution to reduce scan time, producing stair-step artifacts along the through-plane direction. In through-plane MRI super-resolution, an efficiency-fidelity trade-off arises: feed-forward regressors are fast but oversmooth at large slice-thicknesses, while sampling-based methods improve fidelity at high inference cost. We propo…
▽ More
Magnetic Resonance Imaging (MRI) is often acquired with anisotropic resolution to reduce scan time, producing stair-step artifacts along the through-plane direction. In through-plane MRI super-resolution, an efficiency-fidelity trade-off arises: feed-forward regressors are fast but oversmooth at large slice-thicknesses, while sampling-based methods improve fidelity at high inference cost. We propose DRIFT, a two-stage thickness-conditioned rectified flow framework for through-plane MRI super-resolution with continuous input slice-thickness. Stage 1 employs an Anatomical Projection Network (APN) to map low-resolution patches to a coarse high-resolution manifold, providing a deterministic anatomical initialization that shortens the residual transport of Stage 2 and stabilizes slice-wise refinement. Stage 2 refines details via rectified flow and introduces a Physics-Aware Difficulty (PAD) metric derived from slice-thickness induced through-plane bandwidth deficit to guide an Adaptive Integration Scheduler (AIS), allocating ODE steps by thickness. A Consistent Endpoint Trajectory Alignment (CETA) loss enforces thickness-consistent reconstructions. Experiments show that DRIFT outperforms super-resolution baselines while reducing inference cost. Code, models, and interactive demos are available at https://yoonseokchoi-ai.github.io/drift-eccv2026/.
△ Less
Submitted 18 July, 2026;
originally announced July 2026.
-
Controlling Motion Transfer in Diffusion Transformers via Attention Heads
Authors:
Sunyoung Jung,
Jiwoo Park,
Yoonseok Choi,
Kyobin Choo,
Ming-Hsuan Yang,
Seong Jae Hwang
Abstract:
Diffusion Transformers (DiTs) have advanced video generation with high-quality, temporally coherent results. However, extending them to motion transfer, which requires following reference motion while aligning with a target prompt, remains challenging due to limited understanding of motion and structure representations within DiTs. We analyze video DiTs at the attention-head level and identify dis…
▽ More
Diffusion Transformers (DiTs) have advanced video generation with high-quality, temporally coherent results. However, extending them to motion transfer, which requires following reference motion while aligning with a target prompt, remains challenging due to limited understanding of motion and structure representations within DiTs. We analyze video DiTs at the attention-head level and identify distinct heads specialized for motion and spatial structure. Based on this insight, we propose a head-aware controllable motion transfer framework that requires no parameter updates. Our method refines motion cues from motion-specialized heads via semantic correspondence guidance and preserves structure through selective feature injection. This head-level control not only enables accurate motion transfer but also provides an interpretable foundation for controllable video generation with DiTs.
△ Less
Submitted 13 July, 2026;
originally announced July 2026.
-
Finite-time breakdown of the Euler-alignment system for supercritical initial data
Authors:
Young-Pil Choi,
Eitan Tadmor
Abstract:
We study finite-time breakdown of classical solutions to the Euler-alignment system through the degeneration of the associated Lagrangian flow. This approach allows us to characterize singularity formation in terms of the loss of local invertibility of the flow and the resulting concentration of density along characteristics. For the case of constant communication kernels, we derive an explicit fo…
▽ More
We study finite-time breakdown of classical solutions to the Euler-alignment system through the degeneration of the associated Lagrangian flow. This approach allows us to characterize singularity formation in terms of the loss of local invertibility of the flow and the resulting concentration of density along characteristics. For the case of constant communication kernels, we derive an explicit formula for the flow and obtain an exact pointwise breakdown criterion in arbitrary dimension. In two dimensions, this criterion admits a closed-form reformulation in terms of the symmetric part of the initial velocity gradient and the initial vorticity. For general non-constant kernels, we derive sufficient conditions for finite-time degeneracy by combining a leading compressive mechanism with perturbative control of the nonlocal remainder. These conditions provide quantitative supercritical breakdown criteria in arbitrary dimension, complementing the existing subcritical global-regularity theory for multidimensional Euler-alignment systems.
△ Less
Submitted 8 July, 2026;
originally announced July 2026.
-
On-Orbit Calibration of Danuri/PolCam. II. Radiometric Calibration
Authors:
Kilho Baek,
Sungsoo S. Kim,
Minsup Jeong,
Young-Jun Choi
Abstract:
Danuri, South Korea's first lunar orbiter, was launched on August 5, 2022, and has successfully operated its two-year nominal mission phase. The wide-angle Polarimetric Camera (PolCam) onboard Danuri is the first instrument to conduct global polarimetric observations from lunar orbit. This paper presents the comprehensive radiometric calibration pipeline for PolCam's on-orbit data, consisting of d…
▽ More
Danuri, South Korea's first lunar orbiter, was launched on August 5, 2022, and has successfully operated its two-year nominal mission phase. The wide-angle Polarimetric Camera (PolCam) onboard Danuri is the first instrument to conduct global polarimetric observations from lunar orbit. This paper presents the comprehensive radiometric calibration pipeline for PolCam's on-orbit data, consisting of dark current removal, smear correction, and flat-fielding. Notably, PolCam's raw data exhibit severe smear artifacts induced by the frame-transfer CCD architecture, which significantly degrade both radiometric fidelity and the accuracy of polarimetric measurements. These smear artifacts have been effectively mitigated through a rigorous correction algorithm, restoring data quality to a level sufficient for scientific analysis and facilitating the precise derivation of the degree of linear polarization (DoLP). Finally, we present representative examples of polarimetric measurements to validate calibration performance. Although the current calibration focuses on restoring data quality for qualitative scientific analysis, these results clearly demonstrate the expected inverse relationship between intensity and polarization. The absolute photometric calibration required for quantitative DoLP analysis is reserved for a subsequent publication.
△ Less
Submitted 8 July, 2026;
originally announced July 2026.
-
PeTeR: Post-Training Robustification of Probabilistic Circuits
Authors:
Adrian Ciotinga,
Yeming Dai,
YooJung Choi
Abstract:
Probabilistic circuits (PCs) can model complex joint distributions while supporting exact and efficient computation of many inference queries. However, standard likelihood-based PC learning is vulnerable to overfitting and fragile generalization when confronted with data noise, small sample sizes, or distribution shifts. This can be mitigated using distributionally-robust optimization which consid…
▽ More
Probabilistic circuits (PCs) can model complex joint distributions while supporting exact and efficient computation of many inference queries. However, standard likelihood-based PC learning is vulnerable to overfitting and fragile generalization when confronted with data noise, small sample sizes, or distribution shifts. This can be mitigated using distributionally-robust optimization which consider worst-case distributions within a Wasserstein ball of the empirical distribution, but current methods are limited to training a model from scratch in this framework. Instead, we propose PeTeR: a novel, data-free post-training framework designed to robustify pre-trained PCs against distribution shifts without retraining from scratch. Empirical evaluations across multiple density estimation benchmarks demonstrate that PeTeR effectively robustifies baseline models against both random and adversarial perturbations, achieving competitive or superior performance to data-dependent robust learning baselines.
△ Less
Submitted 9 July, 2026; v1 submitted 8 July, 2026;
originally announced July 2026.
-
Constructing a Mock Galaxy Catalog for the All-sky SPECtroscopic Survey of Nearby Galaxies (A-SPEC) Using the Machine-assisted Semi-Simulation Model
Authors:
Dongkok Kim,
Yongseok Jo,
Ho Seong Hwang,
Ji-hoon Kim,
Juhan Kim,
Jaehyun Lee,
Hyeonguk Bahk,
Young-Man Choi,
Moo-Young Chun,
Sang-Hyun Chun,
Haeun Chung,
Kim Dachan,
Sungwook E. Hong,
Minhee Hyun,
Donghui Jeong,
Jae-Woo Kim,
Kang-Min Kim,
Yunjong Kim,
Jongwan Ko,
Minseong Kwon,
Ho-Gyu Lee,
Jong Chul Lee,
Yongseok Lee,
Hyunho Lim,
Heeyoung Oh
, et al. (4 additional authors not shown)
Abstract:
We present a methodology for constructing a mock galaxy catalog for the All-sky SPECtroscopic survey of nearby galaxies (A-SPEC) using the Machine-assisted Semi-Simulation Model. The model is trained on the cosmological magnetohydrodynamical simulation IllustrisTNG to predict baryonic properties of subhalos from dark-matter-only features and is applied to our own N-body simulation tailored to sati…
▽ More
We present a methodology for constructing a mock galaxy catalog for the All-sky SPECtroscopic survey of nearby galaxies (A-SPEC) using the Machine-assisted Semi-Simulation Model. The model is trained on the cosmological magnetohydrodynamical simulation IllustrisTNG to predict baryonic properties of subhalos from dark-matter-only features and is applied to our own N-body simulation tailored to satisfy the requirements of A-SPEC. We have improved the model's accuracy by introducing additional features such as subhalo anisotropy parameters and modified definitions of the subhalo environment, which result in the coefficient of determination R^2=0.96, 0.90, 0.70, 0.79 for stellar mass, gas mass, star formation rate, and gas metallicity, respectively. The resulting mock galaxies reproduce the luminosity-dependent clustering of the target galaxies when tuned to match the number density. We discuss avenues for further improvement, including the role of environment in the predictions. We release the mock galaxy catalog with the baryonic properties predicted from the model.
△ Less
Submitted 7 July, 2026;
originally announced July 2026.
-
Near itinerancy and slow singlet formation in the triangular lattice NaRuO2
Authors:
Charles C. Tam,
Alon Hendler Avidor,
Pritam Bhattacharyya,
Yongseong Choi,
Daniel Haskel,
Sven Luther,
Hlynur Gretarsson,
Liviu Hozoi,
Stephen D. Wilson
Abstract:
NaRuO$_2$ forms a delafossite-like structure that contains triangular sublattices of edge-sharing RuO$_6$ octahedra. It shows no evidence of magnetic order down to 100 mK and persistent spin fluctuations, suggestive of a quantum disordered magnetic ground state. In order to characterize the physical regime from which this disordered state arises, we use resonant inelastic X-ray scattering (RIXS) a…
▽ More
NaRuO$_2$ forms a delafossite-like structure that contains triangular sublattices of edge-sharing RuO$_6$ octahedra. It shows no evidence of magnetic order down to 100 mK and persistent spin fluctuations, suggestive of a quantum disordered magnetic ground state. In order to characterize the physical regime from which this disordered state arises, we use resonant inelastic X-ray scattering (RIXS) and X-ray absorption spectroscopy (XAS) at the Ru-$L_{2,3}$-edge, along with pulsed high-field magnetization to characterize both the local electronic structure and the magnetic interactions. Despite significant spin-orbit coupling inferred from XAS measurements, a spin-orbit exciton, characteristic of a spin-orbit assisted Mott insulator, was not observed with RIXS due to the presence of damped intraorbital excitations, which are characteristic of a metal. Corroborated by models of the high-field magnetization to a random singlet model, we propose a picture of a nearly itinerant system with strong magnetic and charge fluctuations that destabilize long-range magnetic order.
△ Less
Submitted 7 July, 2026; v1 submitted 6 July, 2026;
originally announced July 2026.
-
Leveraging Pathology Co-occurrence for Test-Time Adaptation in Chest X-Ray Diagnosis
Authors:
Woojin Jeong,
Yujin Choi,
Dongbin Kim,
Soyeon Park,
Jaewook Lee
Abstract:
Medical imaging models often degrade when deployed at new clinical sites due to differences in imaging equipment, protocols, and patient populations. Test-time adaptation (TTA) addresses this by updating a pretrained model using only unlabeled target data, without access to source data. However, existing TTA methods were designed for single-label classification on natural image benchmarks, minimiz…
▽ More
Medical imaging models often degrade when deployed at new clinical sites due to differences in imaging equipment, protocols, and patient populations. Test-time adaptation (TTA) addresses this by updating a pretrained model using only unlabeled target data, without access to source data. However, existing TTA methods were designed for single-label classification on natural image benchmarks, minimizing entropy uniformly across all samples without considering label dependencies. This overlooks a key property of multi-label medical imaging: pathologies do not occur independently but exhibit structured co-occurrence patterns. In this work, we propose Co-occurrence Weighted Adaptation (CoWA), which leverages disease co-occurrence patterns as a reliability signal for adaptation. CoWA estimates label co-occurrence structure from model predictions and downweights samples that deviate from expected patterns, enabling adaptation to rely more on consistent predictions while reducing the impact of noisy ones. We evaluate CoWA on chest X-ray benchmarks under domain shifts and demonstrate consistent improvements over established baselines.
△ Less
Submitted 9 July, 2026; v1 submitted 4 July, 2026;
originally announced July 2026.
-
Bottom of the Spectrum of Complete Kähler Metrics from Finite-Mass Plurisubharmonic Exhaustions
Authors:
Young-Jun Choi,
Jiwon Brandon Jeong
Abstract:
Let $Ω\subset\mathbb{C}^{n}$ be a bounded domain, and let $ρ:Ω\to[-1,0)$ be a smooth strictly plurisubharmonic exhaustion function. We consider the logarithmic potential $g=-\log(-ρ)$ and the associated complete Kähler metric $ω=dd^{c}g$. We prove that if $ρ$ satisfies the finite weighted Monge--Ampère mass condition $\int_Ω(-ρ)^{\varepsilon}(dd^{c}ρ)^{n}<+\infty$ for every $\varepsilon>0$, then t…
▽ More
Let $Ω\subset\mathbb{C}^{n}$ be a bounded domain, and let $ρ:Ω\to[-1,0)$ be a smooth strictly plurisubharmonic exhaustion function. We consider the logarithmic potential $g=-\log(-ρ)$ and the associated complete Kähler metric $ω=dd^{c}g$. We prove that if $ρ$ satisfies the finite weighted Monge--Ampère mass condition $\int_Ω(-ρ)^{\varepsilon}(dd^{c}ρ)^{n}<+\infty$ for every $\varepsilon>0$, then the bottom of the spectrum of the Laplace--Beltrami operator of $(Ω,ω)$ satisfies $λ_{0}(Δ_ω,Ω)=n^{2}$. The lower bound follows from the standard estimate applied to $g$, together with the inequality $|\partial g|_ω^{2}\le 1$. For the reverse inequality, for each $α>n/2$, we set $f=(-ρ)^α$ and prove that $f\in W^{1,2}(Ω,ω)$ if and only if $\int_Ω(-ρ)^{2α-n}(dd^{c}ρ)^{n}<+\infty$. Under the finite weighted Monge--Ampère mass condition, this allows us to let $α\downarrow n/2$ in the Rayleigh quotient and obtain the upper bound $λ_{0}(Δ_ω,Ω)\le n^{2}$. As an application, Cegrell's theorem gives a smooth strictly plurisubharmonic exhaustion with finite Monge--Ampère mass on every bounded hyperconvex domain; the associated complete Kähler metric constructed from this exhaustion therefore satisfies $λ_{0}(Δ_ω,Ω)=n^{2}$.
△ Less
Submitted 3 July, 2026;
originally announced July 2026.
-
Automated Data Readiness for Scientific AI
Authors:
Sean R. Wilkinson,
Valentine G. Anantharaj,
Jong Youl Choi,
Ketan Maheshwari,
Marshall McDonnell,
Massimiliano Lupo Pasini,
Polina Shpilker,
Renan Souza,
Patrick Widener,
Sarp Oral,
Wesley Brewer
Abstract:
Leadership computing facilities steward large-scale scientific datasets that routinely require substantial transformation before serving as AI training data. However, no existing framework fully unifies automated transformation, readiness assessment, provenance tracking, and agent-native deployment. We present REDI, an open-source framework that addresses this gap through a unified five-stage pipe…
▽ More
Leadership computing facilities steward large-scale scientific datasets that routinely require substantial transformation before serving as AI training data. However, no existing framework fully unifies automated transformation, readiness assessment, provenance tracking, and agent-native deployment. We present REDI, an open-source framework that addresses this gap through a unified five-stage pipeline (ingest, preprocess, transform, structure, and output) with per-stage instrumentation for reproducibility and deployment as an agent-callable skill; companion tool SetGo automates FAIR compliance and catalog publication. Evaluated across climate, proteomics, materials science, and nuclear fusion, REDI transforms all datasets from raw to AI-ready, with outputs validated against domain-expert references, and preliminary results show near-ideal parallel scaling to 100 nodes on Frontier for the climate case. Provenance-instrumented profiling reveals file I/O as the dominant pipeline cost, with format selection a first-order optimization lever. These results establish REDI as a cross-domain platform providing automated data readiness for scientific AI, transforming data preparation bottlenecks into reproducible, reusable community assets.
△ Less
Submitted 2 July, 2026;
originally announced July 2026.
-
Plaid-Like Spin Splitting and Chirality of Magnon Bands in Antiferromagnetic MnTe$_2$
Authors:
Dirk Wulferding,
Daehyeon An,
Jiwon Choi,
Dongmin Mun,
Youngsu Choi,
Sivasakthi Kuppusamy,
Sritharan Krishnamoorthi,
Raman Sankar,
Myung Joon Han,
Se Kwon Kim,
Kwang-Yong Choi
Abstract:
Altermagnets constitute an emerging class of magnetic materials that combine compensated antiferromagnetic order with spin-split excitations arising from crystalline symmetries. Despite strong theoretical interest, their experimental identification remains challenging. Here, we demonstrate that helicity- and angle-resolved Raman scattering measurements reveal reduced rotational symmetries of magno…
▽ More
Altermagnets constitute an emerging class of magnetic materials that combine compensated antiferromagnetic order with spin-split excitations arising from crystalline symmetries. Despite strong theoretical interest, their experimental identification remains challenging. Here, we demonstrate that helicity- and angle-resolved Raman scattering measurements reveal reduced rotational symmetries of magnons and a pronounced imbalance between left- and right-circular polarization channels, indicating momentum-dependent magnon handedness. First-principles DFT+$U$ calculations combined with linear spin-wave theory uncover a characteristic plaid-like spin-splitting structure in momentum space. The resulting magnon spin textures are dictated by the unconventional sublattice symmetries of MnTe$_2$ and closely emulate those of altermagnetic electronic bands. Our work provides evidence of chiral spin-wave excitations unique to this non-coplanar antiferromagnet.
△ Less
Submitted 2 July, 2026;
originally announced July 2026.
-
Domain Generalization via Text-Anchored Information Bottleneck
Authors:
Eunyi Lyou,
Yunjeong Choi,
Junho Lee,
Joonseok Lee
Abstract:
Visual recognition models often fail when deployed in new environments. Domain Generalization (DG) addresses this by learning representations that remain invariant to environment-specific variations. Recent approaches increasingly rely on large vision-language models, assuming that preserving their expressive visual representations improves robustness. However, we show that such visual expressiven…
▽ More
Visual recognition models often fail when deployed in new environments. Domain Generalization (DG) addresses this by learning representations that remain invariant to environment-specific variations. Recent approaches increasingly rely on large vision-language models, assuming that preserving their expressive visual representations improves robustness. However, we show that such visual expressiveness can instead propagate spurious cues that tie representations to the training environments, hindering invariant learning. We therefore discard visual guidance and instead treat the language embedding space as the primary source of domain invariance, naturally acting as an information bottleneck that preserves core semantics while suppressing domain-specific variations. Extensive experiments across diverse backbones exhibit state-of-the-art performance and further analyze what makes guidance effective for robust generalization. These findings shift the focus of DG from improving representations to designing supervision that enforces invariance.
△ Less
Submitted 1 July, 2026;
originally announced July 2026.
-
Concept Removal Guidance: Evidence-Calibrated Negative Guidance for Safe Diffusion Sampling
Authors:
Yoonseok Choi,
Chaeyoung Oh,
Hyunjun Choi,
Seokin Seo,
Kee-Eung Kim
Abstract:
Text-to-image diffusion models remain vulnerable to adversarial prompts that elicit disallowed content, motivating reliable inference-time controls. A popular approach is negative guidance, which subtracts a negative prompt direction with a fixed weight. However, it often forces a safety-fidelity trade-off, causing artifacts or prompt drift when over-applied and failing under attacks when under-ap…
▽ More
Text-to-image diffusion models remain vulnerable to adversarial prompts that elicit disallowed content, motivating reliable inference-time controls. A popular approach is negative guidance, which subtracts a negative prompt direction with a fixed weight. However, it often forces a safety-fidelity trade-off, causing artifacts or prompt drift when over-applied and failing under attacks when under-applied. Dynamic variants reweight guidance using posterior-odds signals, which can be brittle for open-vocabulary compositional prompts, while lightweight similarity-based methods ignore the evolving image evidence along the denoising trajectory. We introduce Concept Removal Guidance (CRG), a training-free method that estimates unwanted-concept presence at each diffusion step from the model's noise predictions, and adaptively calibrates negative guidance via a closed-form constrained update enforcing a target presence threshold while minimally perturbing the conditional trajectory. Across red-teaming benchmarks, CRG reduces attack success rates while preserving benign fidelity, and extends to additional suppression targets such as artist style and violence without fine-tuning or external classifiers.
△ Less
Submitted 29 June, 2026;
originally announced June 2026.
-
Heavy mesons from the QCD instanton vacuum beyond the static limit
Authors:
Ki-Hoon Hong,
Yongwoo Choi,
Nurmukhammad Rakhimov,
Hyun-Chul Kim
Abstract:
We study pseudoscalar heavy mesons in the QCD instanton vacuum beyond the static limit. Finite-mass effects in the heavy-light loop are encoded in a separable effective vertex built from a profile function $φ(\vec{p})$, kept distinct from the static Wilson-line form factor $F_Q^{(\infty)}(\vec{q})$ of the $m_Q\to\infty$ limit. The pseudoscalar two-point function fixes the residual mass $Λ$ and the…
▽ More
We study pseudoscalar heavy mesons in the QCD instanton vacuum beyond the static limit. Finite-mass effects in the heavy-light loop are encoded in a separable effective vertex built from a profile function $φ(\vec{p})$, kept distinct from the static Wilson-line form factor $F_Q^{(\infty)}(\vec{q})$ of the $m_Q\to\infty$ limit. The pseudoscalar two-point function fixes the residual mass $Λ$ and the residue-normalized meson-quark coupling, from which we evaluate the decay constant, the spin-independent kinetic matrix element, and the zero-recoil slope of the Isgur-Wise function at order $1/m_Q$. The subleading calculation is restricted to the kinetic (derivative) part of the HQET operators. For a representative vertex calibrated to the $B$-meson decay constant and the spin-averaged $B$-meson mass, we obtain $f_B = 186.8$~MeV, $Λ= 184.5$~MeV, $m_b^{\mathrm{eff}} = 5.04$~GeV, $λ_1^{(\partial)} = -0.922~\mathrm{GeV}^2$, and $ρ_{\mathrm{IW}}^2 = 1.105$. The kinetic contribution yields a mass shift of order $Λ/2$ and a sizable $1/m_Q$ current correction, indicating that the spin-independent nonperturbative $1/m_Q$ sector is a sensitive probe of the finite-mass heavy-light vertex.
△ Less
Submitted 26 June, 2026;
originally announced June 2026.
-
GeoFace: Consistent Multi-View Face Generation with Geometry-Constrained Diffusion
Authors:
Yeji Choi,
Jinhyeok Choi,
Jaewon Min,
Minkyung Kwon,
Jin Hyeon Kim,
Seungryong Kim
Abstract:
We present GeoFace, a geometry-constrained multi-view diffusion framework for consistent face generation from a single input. % While recent multi-view diffusion models achieve photorealistic synthesis at the per-view level, they lack an explicit mechanism to enforce a shared 3D structure across views, often leading to inconsistent geometry across viewpoints. To address this, GeoFace proposes a un…
▽ More
We present GeoFace, a geometry-constrained multi-view diffusion framework for consistent face generation from a single input. % While recent multi-view diffusion models achieve photorealistic synthesis at the per-view level, they lack an explicit mechanism to enforce a shared 3D structure across views, often leading to inconsistent geometry across viewpoints. To address this, GeoFace proposes a unified dual-stream framework for joint generation of multi-view RGB images and 3D face geometry, where the appearance and geometry streams interact through shared attention layers. To encourage the two streams to mutually constrain each other, we introduce a geometry-guided attention alignment loss that supervises the cross-attention between appearance and geometry tokens with 3D-consistent correspondences, enabling the appearance stream to correctly reference pose-invariant geometric cues for robust alignment across viewpoints. Geometry is represented as a canonical UV position map, derived from a FLAME mesh fitted to multi-view observations, serving as a view-invariant shared constraint across all generated views. Experiments on RenderMe-360 and NeRSemble demonstrate that GeoFace consistently outperforms existing methods in both visual quality and cross-view geometric consistency, facilitating more efficient 3D reconstruction.
△ Less
Submitted 25 June, 2026;
originally announced June 2026.
-
Confidence-Aware Tool Orchestration for Robust Video Understanding
Authors:
Yangfan He,
Yujin Choi,
Jaehong Yoon
Abstract:
Video reasoning language models implicitly assume that every input frame is equally reliable. This leads to what we term the Blind Trust Problem: under realistic perturbations such as motion blur, glare, or occlusion, frontier video reasoning models can suffer 15-30%p accuracy drops on real-world embodied benchmarks, while remaining unaware that their visual evidence has been degraded. To address…
▽ More
Video reasoning language models implicitly assume that every input frame is equally reliable. This leads to what we term the Blind Trust Problem: under realistic perturbations such as motion blur, glare, or occlusion, frontier video reasoning models can suffer 15-30%p accuracy drops on real-world embodied benchmarks, while remaining unaware that their visual evidence has been degraded. To address this challenge, we propose Robust-TO, an agentic video understanding framework that explicitly integrates per-frame trustworthiness into every stage of reasoning. Robust-TO organizes heterogeneous visual perception tools under a unified evidence interface. Each tool receives a sub-query derived from the original question and a set of trustworthy frames selected by the reliability-relevance score. It returns evidence in a shared format: a concrete prediction (e.g., a bounding box, motion trajectory, recognized text, or action label), temporal grounding, and a calibrated reliability score. During reasoning, these calibrated scores guide evidence weighting in a three-tier synthesis process (high/medium/low) and define a confidence-cost GRPO reward that jointly optimizes correctness, evidence reliability, and efficiency. On two video reasoning benchmarks spanning eight tasks, Robust-TO achieves 56.4% average accuracy on clean inputs, surpassing the strongest open-source baseline by 10.6%p and outperforming Gemini-2.5-Pro (46.2%). Under five realistic corruption types, Robust-TO maintains 54.3% average accuracy, 5.8%p above the strongest open-source baseline, while exhibiting the smallest clean-to-corrupted accuracy drop among all compared methods.
△ Less
Submitted 25 June, 2026;
originally announced June 2026.
-
The Interplay of Harness Design and Post-Training in LLM Agents
Authors:
Kyungmin Kim,
Youngbin Choi,
Seoyeon Lee,
Suhyeon Jun,
Dongwoo Kim,
Sangdon Park
Abstract:
Tool-integrated LLM agents are often wrapped within a harness: the scaffolding that determines which tools are exposed, how they are described, and what auxiliary information accompanies each per-step observation. While agents are routinely post-trained, this scaffolding is typically treated as a fixed engineering detail, with design effort limited to the training-free regime. Moreover, existing p…
▽ More
Tool-integrated LLM agents are often wrapped within a harness: the scaffolding that determines which tools are exposed, how they are described, and what auxiliary information accompanies each per-step observation. While agents are routinely post-trained, this scaffolding is typically treated as a fixed engineering detail, with design effort limited to the training-free regime. Moreover, existing post-training algorithms assume a static environment, even though tool environments and tasks often shift upon deployment. To address this gap, we extend $\texttt{ALFWorld}$ (i) to treat the harness as a controllable design dimension and (ii) to support evaluation under task and tool environment shifts. Building on this, we systematically analyze how the harness design influences post-training in both in-distribution and out-of-distribution (OOD) settings. We empirically show that harness-aware post-training not only improves in-distribution performance but also enables agents to robustly adapt to OOD settings. Under a harness with minimal design effort, post-training suffers a drastic performance drop under stronger tool environment shifts, further highlighting the importance of harness-aware post-training under such shifts.
△ Less
Submitted 24 June, 2026;
originally announced June 2026.
-
Intrinsic Defect Energetics and Fluorine Doping Effects in Li2CO3 and Li2O2: A First-Principles Study
Authors:
Youjeong Choi,
Tasuku Sugiura,
Keisuke Mukai,
Nanako Ishihara,
Shuji Nakanishi,
Teruyasu Mizoguchi
Abstract:
Lithium carbonate, Li2CO3, is a thermodynamically stable carbonate phase whose defect energetics are closely related to its stability and decomposition behavior in various lithium-based electrochemical systems. These properties of Li2CO3 are particularly important in lithium-oxygen battery environments. In these systems, Li2CO3 can form as a parasitic discharge product alongside Li2O2, the primary…
▽ More
Lithium carbonate, Li2CO3, is a thermodynamically stable carbonate phase whose defect energetics are closely related to its stability and decomposition behavior in various lithium-based electrochemical systems. These properties of Li2CO3 are particularly important in lithium-oxygen battery environments. In these systems, Li2CO3 can form as a parasitic discharge product alongside Li2O2, the primary discharge product, leading to performance degradation. However, compared with Li2O2, the intrinsic defect thermodynamics of Li2CO3 and how chemical doping modifies its defect energetics remain insufficiently understood. In this study, first-principles calculations were performed to systematically analyze the intrinsic point-defect energetics of Li2CO3 and to evaluate the effects of fluorine doping on vacancy formation energies in Li2CO3 and Li2O2. Intrinsic defect analysis reveals that defect behavior is predominantly governed by lithium-related defects. Upon fluorine doping, lithium and carbon vacancy formation energies decrease selectively in Li2CO3, partially destabilizing the carbonate framework, while a reduction in lithium vacancy formation energy is also observed in Li2O2. These results suggest that fluorine doping modulates the defect energetics of both discharge products, potentially providing a thermodynamic basis for controlling the stability of Li2CO3 and Li2O2 under thermodynamic conditions representative of lithium-oxygen batteries.
△ Less
Submitted 24 June, 2026;
originally announced June 2026.
-
Deep Imaging of Grus II and Horologium II: Structure and Extent of Two Ultra-Faint Milky Way Satellites
Authors:
Deepthi S. Prabhu,
David J. Sand,
Anirudh Chiti,
Burçin Mutlu-Pakdil,
Sasha N. Campana,
J. L. Carlin,
A. P. Ji,
Jaclyn Jensen,
C. E. Martínez-Vázquez,
Dennis Zaritsky,
A. B. Pace,
A. H. Riley,
D. Crnojević,
G. Limberg,
Laura Congreve Hunter,
Kristine Spekkens,
Michael G. Jones,
Amandine Doliva-Dolinsky,
Paul Bennet,
V. M. Placco,
Quinn O. Casey,
Guinevere Herron,
W. Cerny,
Nitya Kallivayalil,
Y. Choi
, et al. (7 additional authors not shown)
Abstract:
We present deep, wide-field Magellan/Megacam imaging of the ultra-faint Milky Way (MW) satellites Grus II (Gru II) and Horologium II (Hor II), with the aim of deriving improved constraints on their distances, luminosities, and structural parameters, while also searching for possible signs of tidal disturbance. Our photometry reaches approximately 3 magnitudes deeper than the discovery data, enabli…
▽ More
We present deep, wide-field Magellan/Megacam imaging of the ultra-faint Milky Way (MW) satellites Grus II (Gru II) and Horologium II (Hor II), with the aim of deriving improved constraints on their distances, luminosities, and structural parameters, while also searching for possible signs of tidal disturbance. Our photometry reaches approximately 3 magnitudes deeper than the discovery data, enabling robust measurements of these quantities. Both systems exhibit color-magnitude diagrams consistent with old ($\sim$12.5 Gyr), very metal-poor stellar populations. We find Gru II to be at a distance of $52.3 \pm 1.9$ kpc, with a half-light radius of $6.8 \pm 0.5$ arcmin (103 $\pm$ 9 pc), ellipticity $ε= 0.25 \pm 0.07$, and absolute magnitude $M_V = -4.07 \pm 0.50$ mag. Hor II is further away at a distance of $72.4^{+5.9}_{-5.5}$ kpc and more compact, with $r_h = 2.1 \pm 0.2$ arcmin (44$^{+6}_{-5}$ pc), $ε= 0.32^{+0.20}_{-0.16}$, and $M_V = -2.10 \pm 0.44$ mag. Both galaxies lie within the typical size-luminosity locus of MW ultra-faint dwarfs. Gru II shows an asymmetric morphology including multi-directional clumpy features, some of which may be suggestive of tidal disturbance. We further identify and spectroscopically confirm a new distant member just outside $3r_h$ in Gru II, providing independent evidence for member stars at large projected radii. In contrast, Hor II appears regular, with no significant extended structure detected to the surface-brightness limits of our data.
△ Less
Submitted 23 June, 2026;
originally announced June 2026.
-
Deciphering Fingerprints of 3D Molecular Surfaces for Accurate Epitope Prediction
Authors:
Fang Wu,
Weihao Xuan,
Jure Leskovec,
Yejin Choi,
Li Erran Li
Abstract:
Molecular surfaces encode the geometric and physicochemical patterns that determine antibody-antigen recognition, central to epitope prediction. However, existing methods rely on sequences or backbone structures and struggle to capture discontinuous, surface-driven epitopes. This study presents SurfBind, a surface-centric learning framework for epitope prediction that operates directly on molecula…
▽ More
Molecular surfaces encode the geometric and physicochemical patterns that determine antibody-antigen recognition, central to epitope prediction. However, existing methods rely on sequences or backbone structures and struggle to capture discontinuous, surface-driven epitopes. This study presents SurfBind, a surface-centric learning framework for epitope prediction that operates directly on molecular surface representations. SurfBind integrates geometric and physicochemical cues through a Transformer-based architecture with patch-level surface modeling, binder-aware cross-attention, and a hierarchical coarse-to-fine prediction paradigm. Experiments on challenging epitope identification benchmarks, including SAbDab and DB5.5, demonstrate that SurfBind achieves state-of-the-art performance and strong generalization across unseen antibodies and conformational states, highlighting the value of interaction-aware surface modeling for understanding the crucial mechanisms of protein-protein interactions.
△ Less
Submitted 22 June, 2026;
originally announced June 2026.
-
EXPO-SQL: Execution-based Clause-level Policy Optimization for Text-to-SQL
Authors:
Jaehoon Lee,
CheolWon Na,
Suyoung Bae,
Jin-Seop Lee,
Jihyung Lee,
YunSeok Choi,
Jee-Hyong Lee
Abstract:
Text-to-SQL enables users to query databases using natural language by generating executable SQL queries. Recent methods have increasingly adopted Large Language Models based reinforcement learning (RL) to leverage execution feedback for training. However, existing RL methods assign uniform query-level rewards to all clauses in a SQL query, treating correct and incorrect clauses equally. This coar…
▽ More
Text-to-SQL enables users to query databases using natural language by generating executable SQL queries. Recent methods have increasingly adopted Large Language Models based reinforcement learning (RL) to leverage execution feedback for training. However, existing RL methods assign uniform query-level rewards to all clauses in a SQL query, treating correct and incorrect clauses equally. This coarse-grained reward design leads to insufficient learning signals for correct SQL generation. To address this issue, we propose EXPO-SQL (EXecution-based clause-level Policy Optimization for Text-to-SQL) which provides fine-grained supervision through clause-level rewards. To assign clause-level rewards, our method identifies erroneous clauses by analyzing execution results, including error messages and clause-wise incremental execution. Experiments on widely-used Text-to-SQL benchmarks demonstrate that EXPO-SQL significantly outperforms existing supervised fine-tuning, prompting, and RL-based methods through fine-grained clause-level learning. Our code is available at https://github. com/jhn25/EXPO-SQL.
△ Less
Submitted 29 April, 2026;
originally announced June 2026.
-
Safe Few-Step Generation via Velocity Editing
Authors:
Yujin Choi,
Jaehong Yoon
Abstract:
Flow matching has recently emerged as a strong paradigm for state-of-the-art text-to-image (T2I) generation, enabling high-quality generation with a small number of sampling steps. As these models are increasingly integrated into real-world applications, ensuring safe and non-sensitive content generation has become a critical requirement. However, adapting safety and concept removal methods to thi…
▽ More
Flow matching has recently emerged as a strong paradigm for state-of-the-art text-to-image (T2I) generation, enabling high-quality generation with a small number of sampling steps. As these models are increasingly integrated into real-world applications, ensuring safe and non-sensitive content generation has become a critical requirement. However, adapting safety and concept removal methods to this new generation framework remains an open challenge. Specifically, prior methods largely rely on iterative trajectory steering across a number of denoising steps or on CLIP-centric prompt embedding manipulation. These design assumptions pose fundamental bottlenecks for safety in flow matching-based T2I generation, where limited sampling steps constrain iterative correction and modern context-aware text encoders diminish the effectiveness of embedding-level interventions. In this paper, we propose VESFlow, a training-free safety method tailored to flow matching with extremely few sampling steps. Leveraging the fact that flow matching models learn the marginal velocity, we directly edit the velocity field via a safe-conditional posterior. VESFlow steers the trajectory toward safe outputs while leaving the conditioning prompt unchanged. Building on the observation that VESFlow leaves outputs unchanged under benign prompts, we further introduce a risk score-based filtering that bypasses velocity editing to reduce computational cost while preserving benign prompt generation. Based on this filtering, we propose VESFlow+, a stronger variant of VESFlow that not only edits the velocity toward the safe direction, but also pushes it away from the unsafe direction. Experimental results show that VESFlow+ removes the target concept, reducing the attack success rate by NudeNet to 6.3% on Ring-A-Bell and 6.8% on MMA-Diffusion on the 4-step MeanFlow model, while preserving fidelity on benign prompts.
△ Less
Submitted 22 June, 2026;
originally announced June 2026.
-
TaLK: Text-attributed Graph Dataset Distillation via Coupling Language Model with Graph-Aware Kernel
Authors:
Yeongho Kim,
Yeonje Choi,
Kijung Shin
Abstract:
Text-attributed graphs (TAGs) are widely used in many real-world domains, and learning on TAGs requires jointly modeling text semantics and graph structure. A standard approach for modeling TAGs is to combine a language model (LM) and a graph neural network (GNN), but joint training is computationally expensive and difficult to scale. Dataset distillation is a promising way to reduce training cost…
▽ More
Text-attributed graphs (TAGs) are widely used in many real-world domains, and learning on TAGs requires jointly modeling text semantics and graph structure. A standard approach for modeling TAGs is to combine a language model (LM) and a graph neural network (GNN), but joint training is computationally expensive and difficult to scale. Dataset distillation is a promising way to reduce training costs, but existing methods are not well suited to TAGs because they are typically designed for a single modality or still require repeatedly training expensive LM-GNN models on the full dataset during distillation. To address this, we propose TaLK, an effective dataset distillation method for TAGs that couples an LM with a graph-aware neural tangent kernel.This design enables efficient dataset distillation, avoiding repeated joint training on the full dataset while reflecting both textual and structural information for effective TAG learning.Experiments on multiple TAG benchmarks show that TaLK consistently outperforms existing baselines and achieves up to 97% of full-dataset performance with only 1% synthetic data.
△ Less
Submitted 22 June, 2026;
originally announced June 2026.
-
CORTIS: Text-Only Adaptation of Spoken Language Models for Task-Oriented Voice Agents
Authors:
Youngwon Choi,
Hyeonyu Kim,
Taeyoun Kwon,
Donghyuk Jung,
Myeongkyun Cho
Abstract:
Task-oriented voice agents need to map spoken user requests to structured outputs such as semantic frames, executable actions, and function calls. A common approach is to cascade ASR with a text-based LLM, but transcription errors can propagate to downstream structured output generation, especially under noisy conditions. Spoken language models (SLMs) offer a direct speech-based alternative, yet a…
▽ More
Task-oriented voice agents need to map spoken user requests to structured outputs such as semantic frames, executable actions, and function calls. A common approach is to cascade ASR with a text-based LLM, but transcription errors can propagate to downstream structured output generation, especially under noisy conditions. Spoken language models (SLMs) offer a direct speech-based alternative, yet adapting them to new tasks typically requires paired speech-target annotations. Motivated by this gap, we present CORTIS, a text-only adaptation framework for task-oriented voice agents. CORTIS fine-tunes SLMs using text-form task supervision, enabling speech-based structured output generation at inference time without task-specific speech-target annotations during adaptation. We evaluate CORTIS on two Qwen2.5-Omni backbones and three task-oriented speech datasets, including an in-house product dataset, and compare it with matched ASR-LLM cascades trained with the same text-form task supervision. Results show that CORTIS performs competitively with matched cascades and offers clearer advantages under acoustic degradation, particularly in preserving high-level task semantics. These findings suggest that text-only fine-tuning of SLMs can serve as a practical adaptation strategy for voice agents when paired speech-target data are costly to collect.
△ Less
Submitted 19 June, 2026;
originally announced June 2026.