-
AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model
Authors:
Kwan Yun,
Serin Yoon,
Sunjin Jung,
Jung Eun Yoo,
Inyup Lee,
Junyong Noh
Abstract:
We present AnyTalk, a novel method for generating 3D speech animations for arbitrary characters without requiring any animation data. While existing audio-driven 3D speech animation methods rely on character-specific training data or laborious rigging/re-meshing, AnyTalk circumvents these limitations by leveraging recent video diffusion models trained on extensive video datasets. We first adapt a…
▽ More
We present AnyTalk, a novel method for generating 3D speech animations for arbitrary characters without requiring any animation data. While existing audio-driven 3D speech animation methods rely on character-specific training data or laborious rigging/re-meshing, AnyTalk circumvents these limitations by leveraging recent video diffusion models trained on extensive video datasets. We first adapt a pre-trained video diffusion model to a target character through our Character-specific Fine-tuning (\textit{CsF}) technique. By fine-tuning on rendered images of the 3D character paired with zeroed-out audio embeddings (representing "no motion"), we eliminate the need for animation data while preserving the motion prior of large-scale video diffusion model. We then uplift the resulting talking-head video into a 3D speech animation by estimating blendshape parameters through a proposed optimization process. AnyTalk enables lip-synced animations across diverse face meshes and blendshape configurations, significantly reducing manual effort and data requirements. We further enhance usability by distilling AnyTalk into a streamlined network, $\text{AnyTalk}_{RT}$, thereby enabling real-time performance. By leveraging talking-head video generation, our method broadens access to audio-driven speech animation technology for arbitrary characters. The code is publicly available at https://serin-yoon.github.io/projects/anytalk/.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
What You Ask is What You Ground: Bridging Question Intent to Temporal Evidence for Grounded VideoQA
Authors:
Jinhwan Seo,
Kyubeom Han,
Jumin Lee,
Junhyug Noh,
Sung-eui Yoon
Abstract:
We study a critical yet overlooked failure mode in Grounded Video Question Answering: question-invariant grounding, where models predict nearly identical temporal segments for different questions about the same video. We trace this behavior to two structural limitations in prior common designs: (i) modality isolation that fixes video representations before they receive question semantics, and (ii)…
▽ More
We study a critical yet overlooked failure mode in Grounded Video Question Answering: question-invariant grounding, where models predict nearly identical temporal segments for different questions about the same video. We trace this behavior to two structural limitations in prior common designs: (i) modality isolation that fixes video representations before they receive question semantics, and (ii) weak question injection inside the grounding module. To address this, we propose GroundFormer, which conditions video features on question intent before localization via learnable communication tokens that mediate directed visuo-lingual interaction. On top of the question-conditioned features, a factorized MIL cross-attention couples answer selection with temporal evidence under candidate-level supervision, while Gaussian smoothing converts peaked attention into temporally coherent segments. We further introduce a hierarchical multi-modal contrastive loss that aligns video, question, and answer embeddings across a two-pass training pipeline. GroundFormer achieves state-of-the-art grounded VideoQA performance on NExT-GQA and STAR, substantially improving question-discriminative temporal grounding.
△ Less
Submitted 16 August, 2026;
originally announced August 2026.
-
Ion-Scale Current Sheets Embedded in Reconnection Jet Shear Layer of the Near-Sun Heliospheric Current Sheet
Authors:
Dae-Young Lee,
Dooyoung Choi,
Kyung-Eun Choi,
Sung Jun Noh
Abstract:
Context. Magnetic reconnection in the heliospheric current sheet (HCS) plays an important role in restructuring the solar wind magnetic topology and generating plasma jets and magnetic islands. While large-scale signatures of HCS reconnection have been reported in many observational studies, the kinetic-scale structure embedded within reconnection regions remains less well understood. Aims. We inv…
▽ More
Context. Magnetic reconnection in the heliospheric current sheet (HCS) plays an important role in restructuring the solar wind magnetic topology and generating plasma jets and magnetic islands. While large-scale signatures of HCS reconnection have been reported in many observational studies, the kinetic-scale structure embedded within reconnection regions remains less well understood. Aims. We investigate the ion-scale currents sheets (CSs) embedded within an HCS reconnection region and their relationship to the flow-shear layer at the edge of a reconnection jet. Methods. We analyzed an HCS crossing observed by the Parker Solar Probe on March 29, 2024, using high-time-resolution magnetic field measurements. We focused on ion-scale magnetic transitions within two brief intervals of flow-shear layer at the edges of the reconnection jet and examined them in a local LMN coordinate system. Results. Twelve representative CSs are identified, whose duration is on average $\sim$0.06 sec, corresponding to spatial scales of only a few ion inertial lengths. They are classified into three types based on the behavior of the out-of-plane magnetic component $B_{M}$: (1) CSs showing clear bipolar $B_{M}$ variations without bifurcation in reconnecting-field ($B_{L}$), (2) CSs with both bipolar $B_{M}$ variations and bifurcated $B_{L}$ profiles characterized by a plateau structure, and (3) CSs where strong fluctuations obscure an otherwise expected bipolar signature. Conclusions. The reconnection jet shear layer in the HCS may serve as an active site that hosts a chain of ion-scale CSs. This provides new insight into the multiscale structure of HCS reconnection and suggests that flow shear layers may play an important role in generating secondary kinetic-scale structures.
△ Less
Submitted 15 July, 2026;
originally announced July 2026.
-
Learning Probabilistic Prompt for Continual Learning
Authors:
Hyekang Park,
Sanghoon Lee,
Geon Lee,
Jongyoun Noh,
Bumsub Ham
Abstract:
Continual learning aims to progressively learn from a sequence of tasks, each containing a disjoint subset of classes, while preserving previously learned knowledge. Prompt-based continual learning methods propose to learn a small set of parameters, i.e., prompts, by associating them with a query feature of an input image. These methods optimize the prompts, attempting to represent diverse pattern…
▽ More
Continual learning aims to progressively learn from a sequence of tasks, each containing a disjoint subset of classes, while preserving previously learned knowledge. Prompt-based continual learning methods propose to learn a small set of parameters, i.e., prompts, by associating them with a query feature of an input image. These methods optimize the prompts, attempting to represent diverse patterns of images. However, we have observed that existing prompt-based methods suffer from a prompt collapse problem, that is, the prompts tend to be highly similar to each other, thereby failing to capture the diverse data distributions in continual learning scenarios. To address this issue, we propose in this paper a novel prompt-based continual learning framework that captures diverse patterns of images across a sequence of tasks. To this end, we model each prompt as a probabilistic distribution and construct a mixture of these distributions, from which we sample diverse prompts. This enables our model to effectively capture highly diverse image distributions in the continual learning process. We also present a distribution regularization loss to prevent abrupt changes in the prompt distributions throughout the training process. We show extensive experimental results for continual learning on standard benchmarks, including ImageNet-R, CIFAR-100, and CUB-200, demonstrating the effectiveness of our framework.
△ Less
Submitted 6 July, 2026;
originally announced July 2026.
-
JointHOI: Jointly Generating Contact Maps Enhances Hand Object Interaction Generation
Authors:
Mingyeong Song,
Jungbin Cho,
Jisoo Kim,
Ananya Bal,
Kartik Sharma,
Youngjae Yu,
Laszlo A. Jeni,
Junhyug Noh
Abstract:
Text driven hand object interaction (HOI) generation is gaining attention for immersive applications and robotics, yet producing physically plausible interactions remains challenging. Even when individual motions appear natural, small contact errors can cause conspicuous artifacts such as floating and interpenetration. Prior methods mitigate these issues using explicit contact cues or implicit gra…
▽ More
Text driven hand object interaction (HOI) generation is gaining attention for immersive applications and robotics, yet producing physically plausible interactions remains challenging. Even when individual motions appear natural, small contact errors can cause conspicuous artifacts such as floating and interpenetration. Prior methods mitigate these issues using explicit contact cues or implicit grasp priors, but typically rely on multi stage pipelines and fail to model temporally evolving contact. We present JointHOI, a single stage diffusion framework that jointly generates 3D hand object motion and dynamic, distance based contact maps from text. By treating contact as an auxiliary inner modality, joint generation enables the model to learn contact motion coupling during training. At inference, contact guided sampling enforces consistency between generated contact maps and motion implied geometry, improving temporal stability and reducing penetration and floating. Experiments on GRAB and ARCTIC demonstrate consistent improvements in text adherence and physical plausibility over prior methods.
△ Less
Submitted 2 July, 2026;
originally announced July 2026.
-
An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios
Authors:
Hankyul Baek,
Jaewon Noh,
Sang Seo,
Yongsu Kim,
Gabriel Waikin Loh Matienzo,
Young Il Kim,
Ee Wei Seah,
Akriti Vij
Abstract:
AI agents are increasingly being adopted in enterprise and personal settings with access to emails, databases, documents, and other tools where they can read, update, and disseminate sensitive information. Much of prior research on data leakage risks in agents has focused on adversarial data exfiltration through prompt injections and jailbreaks. However, sensitive information may also be exposed d…
▽ More
AI agents are increasingly being adopted in enterprise and personal settings with access to emails, databases, documents, and other tools where they can read, update, and disseminate sensitive information. Much of prior research on data leakage risks in agents has focused on adversarial data exfiltration through prompt injections and jailbreaks. However, sensitive information may also be exposed during non-adversarial use, creating leakage risks even when users issue benign requests.
We report a joint evaluation by the Singapore AI Safety Institute and the Korea AI Safety Institute examining agent data leakage in 12 realistic, non-adversarial tasks spanning customer support, DevOps, web automation, and enterprise and personal productivity. The evaluation covers five risk types: lack of data awareness, audience awareness, policy compliance, data minimization, and access-boundary awareness. Both institutes tested a common set of scenarios mirroring real-world deployments using independent testing environments and task-specific LLM-judge rubrics.
Across the three tested agents, none achieved fully correct and fully safe execution across all scenarios. Successful task completion often coincided with data-handling failures such as accessing unnecessary information or disclosing information to inappropriate recipients, indicating that capability and data-handling safety should be evaluated separately. Qualitative review also revealed claim-action mismatches, simulation-aware behavior, user-simulator role reversal, and interpretation gaps in automated judging. Overall, the results indicate that operational data leakage is a first-order agent-safety concern distinct from adversarial exfiltration and provide a methodology for future evaluations of agent data-handling safety.
△ Less
Submitted 15 June, 2026;
originally announced June 2026.
-
CyberMaskQA: A Privacy-Aware Benchmark for Evaluating Large Language Models in Cybersecurity Question Answering
Authors:
Matilda Gaddi,
Jin Noh,
Onat Gungor,
Tajana Rosing
Abstract:
Large language models (LLMs) are increasingly applied to cybersecurity question answering (QA) for critical tasks such as incident response and vulnerability analysis. However, real-world operational contexts, including system logs and network configurations, inherently contain sensitive identifiers, e.g., IP addresses, host names, and user accounts. Processing this data with cloud-based models is…
▽ More
Large language models (LLMs) are increasingly applied to cybersecurity question answering (QA) for critical tasks such as incident response and vulnerability analysis. However, real-world operational contexts, including system logs and network configurations, inherently contain sensitive identifiers, e.g., IP addresses, host names, and user accounts. Processing this data with cloud-based models is often unsafe or infeasible in regulated environments. Furthermore, progress in privacy-preserving QA is hindered by the lack of annotated, context-rich datasets capable of jointly evaluating operational reasoning and privacy preservation. To address this gap, we introduce CYBERMASKQA, a privacy-aware QA benchmark covering key security domains. Unlike existing benchmarks that primarily test factual knowledge, CYBERMASKQA grounds questions in realistic organizational contexts with explicit causal dependencies among assets and privileges. Generated through a systematic pipeline, the dataset combines human-curated base scenarios with LLM-driven semantic expansion, annotating each instance with precise private entity labels to enable controlled information disclosure. Evaluations of QA accuracy and masking performance demonstrate the benchmark's utility for developing deployable, context-aware cybersecurity models and facilitating nuanced studies of privacy-utility trade-offs. Upon acceptance, we will release the dataset and the generation framework.
△ Less
Submitted 23 May, 2026;
originally announced May 2026.
-
Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition
Authors:
Junghyun Lee,
Hyunseo Kim,
Hanna Jang,
Junhyug Noh
Abstract:
Blended emotion recognition is challenging because emotions are often expressed as mixtures of subtle and overlapping multimodal cues rather than a single dominant signal. We propose a rank-aware multi-encoder framework that selectively combines complementary representations from diverse pre-extracted video and audio encoders. Our method projects heterogeneous encoder features into a shared latent…
▽ More
Blended emotion recognition is challenging because emotions are often expressed as mixtures of subtle and overlapping multimodal cues rather than a single dominant signal. We propose a rank-aware multi-encoder framework that selectively combines complementary representations from diverse pre-extracted video and audio encoders. Our method projects heterogeneous encoder features into a shared latent space, estimates sample-wise encoder importance through an attention-based gating module, and fuses only the top-n most informative encoders. To better model blended emotions, we decouple prediction into presence and salience heads and align them through probability-level fusion. We further incorporate feature-level unsupervised domain adaptation without pseudo-labeling to improve robustness under distribution shift. Experiments on the BlEmoRE challenge show that the proposed framework outperforms strong individual encoders and naïve multi-encoder fusion baselines. Our final system ranked 2nd in the competition, supporting the effectiveness of rank-aware selective fusion for fine-grained blended emotion recognition.
△ Less
Submitted 24 May, 2026; v1 submitted 20 May, 2026;
originally announced May 2026.
-
Skinned Motion Retargeting with Spatially Adaptive Interaction Guidance
Authors:
Soojin Choi,
Seokhyeon Hong,
Chaelin Kim,
Junghyun Nam,
Junhyuk Jeon,
Junyong Noh
Abstract:
Retargeting motion across characters with varying body shapes while preserving interaction semantics, such as self-contact and near-body proximity, remains a challenging problem. While recent geometry-aware approaches address this by maintaining spatial relationships between predefined corresponding regions, their reliance on static correspondences often struggles when the target character exhibit…
▽ More
Retargeting motion across characters with varying body shapes while preserving interaction semantics, such as self-contact and near-body proximity, remains a challenging problem. While recent geometry-aware approaches address this by maintaining spatial relationships between predefined corresponding regions, their reliance on static correspondences often struggles when the target character exhibits exaggerated body proportions. In this paper, we present a geometry-aware motion retargeting framework that preserves interaction semantics by performing proximity matching over spatially adaptive anchors. Unlike prior methods with static anchor definitions, the proposed method dynamically repositions anchors to reachable regions on the target character. This is achieved via a Transformer-based anchor refinement strategy that predicts anchor displacements and constrains the translated anchors to remain on the target character geometry through differentiable soft projection. By incorporating pose-dependent spatial structures from the source character, the adapted anchors provide structurally coherent guidance for interaction-aware retargeting. Conditioned on these anchors, a graph-based autoencoder predicts target skeletal motion that preserves the spatial configuration of the source. To encourage task-aligned optimization between anchor adaptation and motion retargeting, we adopt an alternating training scheme in which each module is optimized in turn. Through extensive evaluations, we demonstrate that our method outperforms state-of-the-art approaches in preserving interaction fidelity across diverse character geometries.
△ Less
Submitted 19 May, 2026;
originally announced May 2026.
-
Prompt Compression in Diffusion Large Language Models: Evaluating LLMLingua-2 on LLaDA
Authors:
Sterling Huang,
Abigayle Brown,
Jiyoo Noh,
Jiakang Xu,
Wantong Huo,
Kaung Myat Kyaw,
Jonathan Chan
Abstract:
Prompt compression reduces inference cost and context length in large language models, but prior evaluations focus mainly on autoregressive architectures. This study examines whether LLMLingua-2 transfers effectively to diffusion large language models (DLLMs), specifically LLaDA-8B-Instruct. We evaluate GSM8K, DUC2004, and ShareGPT using 250 prompts per dataset at an approximate 50\% compression r…
▽ More
Prompt compression reduces inference cost and context length in large language models, but prior evaluations focus mainly on autoregressive architectures. This study examines whether LLMLingua-2 transfers effectively to diffusion large language models (DLLMs), specifically LLaDA-8B-Instruct. We evaluate GSM8K, DUC2004, and ShareGPT using 250 prompts per dataset at an approximate 50\% compression ratio, covering mathematical reasoning, prompt reconstruction, and summarization. Outputs from original, compressed, and reconstructed prompts are compared using exact-match accuracy, BLEU, ROUGE, and BERTScore. Results show that high semantic preservation does not necessarily ensure stable downstream behavior in diffusion models. Summarization remains relatively robust, while mathematical reasoning degrades substantially despite high semantic similarity. Reconstruction further shows that semantically similar prompts may omit reasoning-critical information needed for stable denoising. Overall, compression failures are mainly driven by information omission rather than semantic drift, suggesting that autoregressive prompt compression methods may not transfer uniformly to DLLMs. These findings motivate diffusion-aware compression strategies.
△ Less
Submitted 11 July, 2026; v1 submitted 18 May, 2026;
originally announced May 2026.
-
Stylized Text-to-Motion Generation via Hypernetwork-Driven Low-Rank Adaptation
Authors:
Junhyuk Jeon,
Seokhyeon Hong,
Junyong Noh
Abstract:
Text-driven motion diffusion models are capable of generating realistic human motions, but text alone often struggles to express fine-level nuances of motion, commonly referred to as style. Recent approaches have tackled this challenge by attaching a style injection mechanism to a pretrained text-driven diffusion model. Existing stylization methods, however, either require style-specific fine-tuni…
▽ More
Text-driven motion diffusion models are capable of generating realistic human motions, but text alone often struggles to express fine-level nuances of motion, commonly referred to as style. Recent approaches have tackled this challenge by attaching a style injection mechanism to a pretrained text-driven diffusion model. Existing stylization methods, however, either require style-specific fine-tuning of existing models or rely on heavy ControlNet-based architectures, limiting efficiency and generalization to unseen styles. We propose a lightweight style conditioning framework that dynamically modulates a pretrained diffusion model through hypernetwork-generated LoRA parameters. A style reference motion is encoded into a global style embedding, which is mapped by a hypernetwork to low-rank updates applied at each denoising step of the diffusion model. By structuring the style latent space with a supervised contrastive loss, our framework reliably captures diverse stylistic attributes, improves generalization to unseen styles, and supports optimization-based guidance without requiring predefined style categories. Experiments on the HumanML3D and 100STYLE datasets show state-of-the-art stylization results, while achieving improved stylization for unseen styles.
△ Less
Submitted 13 May, 2026;
originally announced May 2026.
-
XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity
Authors:
Dasol Choi,
Eugenia Kim,
Jaewon Noh,
Sang Seo,
Eunmi Kim,
Myunggyo Oh,
Yunjin Park,
Brigitta Jesica Kartono,
Josef Pichlmeier,
Helena Berndt,
Sai Krishna Mendu,
Glenn Johannes Tungka,
Özlem Gökçe,
Suresh Gehlot,
Katherine Pratt,
Amanda Minnich,
Haon Park
Abstract:
Current LLM safety benchmarks are predominantly English-centric and often rely on translation, failing to capture country-specific harms. Moreover, they rarely evaluate a model's ability to detect culturally embedded sensitivities as distinct from universal harms. We introduce XL-SafetyBench. a suite of 5,500 test cases across 10 country-language pairs, comprising a Jailbreak Benchmark of country-…
▽ More
Current LLM safety benchmarks are predominantly English-centric and often rely on translation, failing to capture country-specific harms. Moreover, they rarely evaluate a model's ability to detect culturally embedded sensitivities as distinct from universal harms. We introduce XL-SafetyBench. a suite of 5,500 test cases across 10 country-language pairs, comprising a Jailbreak Benchmark of country-grounded adversarial prompts and a Cultural Benchmark where local sensitivities are embedded within innocuous requests. Each item is constructed via a multi-stage pipeline that combines LLM-assisted discovery, automated validation gates, and dual independent native-speaker annotators per country. To distinguish principled refusal from comprehension failure, we evaluate Attack Success Rate (ASR) alongside two complementary metrics we introduce: Neutral-Safe Rate (NSR) and Cultural Sensitivity Rate (CSR). Evaluating 10 frontier and 27 local LLMs reveals two key findings. First, jailbreak robustness and cultural awareness do not show a coupled relationship among frontier models, so a composite safety score obscures per-axis variation. Second, local models exhibit a near-linear ASR-NSR trade-off (r = -0.81), indicating that their apparent safety reflects generation failure rather than genuine alignment. XL-SafetyBench enables more nuanced, cross-cultural safety evaluation in the multilingual era.
△ Less
Submitted 7 May, 2026;
originally announced May 2026.
-
Understanding teens' self-beliefs when learning to construct and deconstruct AI/ML systems: Developing a survey instrument
Authors:
Luis Morales-Navarro,
Deborah Fields,
Michael T. Giang,
Daniel J. Noh,
Yasmin B. Kafai,
Danaé Metaxa
Abstract:
Despite growing calls to foster AI literacy, there are few available survey instruments designed for children and youth that study computational empowerment alongside construction and deconstruction activities. In such activities, learners' beliefs about their abilities and attributes can impact their engagement. In this paper, we introduce and validate a survey instrument with constructs related…
▽ More
Despite growing calls to foster AI literacy, there are few available survey instruments designed for children and youth that study computational empowerment alongside construction and deconstruction activities. In such activities, learners' beliefs about their abilities and attributes can impact their engagement. In this paper, we introduce and validate a survey instrument with constructs related to construction (creative expression and problem-solving self-beliefs) and deconstruction (auditing self-efficacy and fascination with auditing), along with more general self-beliefs related to design justice and the value of learning about AI/ML. We administered the instrument to 124 teenagers and assessed the six-factor structure of the instrument using confirmatory factor analysis. In addition to confirming the structure, we found that design justice beliefs strongly correlated with problem-solving, auditing self-efficacy, and creative expression.
△ Less
Submitted 24 April, 2026;
originally announced April 2026.
-
StyleID: A Perception-Aware Dataset and Metric for Stylization-Agnostic Facial Identity Recognition
Authors:
Kwan Yun,
Changmin Lee,
Ayeong Jeong,
Youngseo Kim,
Seungmi Lee,
Junyong Noh
Abstract:
Creative face stylization aims to render portraits in diverse visual idioms such as cartoons, sketches, and paintings while retaining recognizable identity. However, current identity encoders, which are typically trained and calibrated on natural photographs, exhibit severe brittleness under stylization. They often mistake changes in texture or color palette for identity drift or fail to detect ge…
▽ More
Creative face stylization aims to render portraits in diverse visual idioms such as cartoons, sketches, and paintings while retaining recognizable identity. However, current identity encoders, which are typically trained and calibrated on natural photographs, exhibit severe brittleness under stylization. They often mistake changes in texture or color palette for identity drift or fail to detect geometric exaggerations. This reveals the lack of a style-agnostic framework to evaluate and supervise identity consistency across varying styles and strengths. To address this gap, we introduce StyleID, a human perception-aware dataset and evaluation framework for facial identity under stylization. StyleID comprises two datasets: (i) StyleBench-H, a benchmark that captures human same-different verification judgments across diffusion- and flow-matching-based stylization at multiple style strengths, and (ii) StyleBench-S, a supervision set derived from psychometric recognition-strength curves obtained through controlled two-alternative forced-choice (2AFC) experiments. Leveraging StyleBench-S, we fine-tune existing semantic encoders to align their similarity orderings with human perception across styles and strengths. Experiments demonstrate that our calibrated models yield significantly higher correlation with human judgments and enhanced robustness for out-of-domain, artist drawn portraits. All of our datasets, code, and pretrained models are publicly available at https://kwanyun.github.io/StyleID_page/
△ Less
Submitted 23 April, 2026;
originally announced April 2026.
-
Enhancing Science Classroom Discourse Analysis through Joint Multi-Task Learning for Reasoning-Component Classification
Authors:
Jiho Noh,
Mukhesh Raghava Katragadda,
Raymond Carl,
Soon Lee
Abstract:
Analyzing the reasoning patterns of students in science classrooms is critical for understanding knowledge construction mechanism and improving instructional practice to maximize cognitive engagement, yet manual coding of classroom discourse at scale remains prohibitively labor-intensive. We present an automated discourse analysis system (ADAS) that jointly classifies teacher and student utterance…
▽ More
Analyzing the reasoning patterns of students in science classrooms is critical for understanding knowledge construction mechanism and improving instructional practice to maximize cognitive engagement, yet manual coding of classroom discourse at scale remains prohibitively labor-intensive. We present an automated discourse analysis system (ADAS) that jointly classifies teacher and student utterances along two complementary dimensions: Utterance Type and Reasoning Component derived from our prior CDAT framework. To address severe label imbalance among minority classes, we (1) stratify-resplit the annotated corpus, (2) apply LLM-based synthetic data augmentation targeting minority classes, and (3) train a dual-probe head RoBERTa-base classifier. A zero-shot GPT-5.4 baseline achieves macro-F1 of 0.467 on UT and 0.476 on RC, establishing meaningful upper bounds for prompt-only approaches motivating fine-tuning. Beyond classification, we conduct discourse pattern analyses including UTxRC co-occurrence profiling, Cognitive Complexity Index (CCI) computation per session, lag-sequential analysis, and IRF chain analysis, revealing that teacher Feedback-with-Question (Fq) moves are the most consistent antecedents of student inferential reasoning (SR-I). Our results demonstrate that LLM-based augmentation meaningfully improves UT minority-class recognition, and that the structural simplicity of the RC task makes it tractable even for lexical baselines.
△ Less
Submitted 15 August, 2026; v1 submitted 22 April, 2026;
originally announced April 2026.
-
Extensive Spatio-Temporal Chaos in Non-reciprocal Flocking
Authors:
Chul-Ung Woo,
Jae Dong Noh,
Heiko Rieger
Abstract:
Non-reciprocal interactions in active matter give rise to a multitude of fascinating phenomena among which are collective oscillatory states without intrinsic particle chirality and active turbulence. Here we show that in a paradigmatic model for non-reciprocal flocking, the two species Vicsek model, these two states coexist: chiral order for small flocks, and extensive spatiotemporal chaos for la…
▽ More
Non-reciprocal interactions in active matter give rise to a multitude of fascinating phenomena among which are collective oscillatory states without intrinsic particle chirality and active turbulence. Here we show that in a paradigmatic model for non-reciprocal flocking, the two species Vicsek model, these two states coexist: chiral order for small flocks, and extensive spatiotemporal chaos for large flocks, both separated by a finite-wavelength instability whose scale is set by the rotation radius of the chiral orbits. For system sizes larger than this length scale extensive spatiotemporal chaos unfolds, as manifested by an extensive number of positive Lyapunov exponents as well as of Floquet exponents, a finite correlation and chaotic length and a broad energy spectrum. Our results suggest that complex, turbulent behavior is a generic possibility in systems where particles or fields interact asymmetrically, and may have significant implications for understanding how non-reciprocal interactions could drive chaotic, fluid-like behavior in active matter.
△ Less
Submitted 6 July, 2026; v1 submitted 7 April, 2026;
originally announced April 2026.
-
Explicit Time-Frequency Dynamics for Skeleton-Based Gait Recognition
Authors:
Seoyeon Ko,
Yeojin Song,
Egene Chung,
Luca Quagliato,
Taeyong Lee,
Junhyug Noh
Abstract:
Skeleton-based gait recognizers excel at modeling spatial configurations but often underuse explicit motion dynamics that are crucial under appearance changes. We introduce a plug-and-play Wavelet Feature Stream that augments any skeleton backbone with time-frequency dynamics of joint velocities. Concretely, per-joint velocity sequences are transformed by the continuous wavelet transform (CWT) int…
▽ More
Skeleton-based gait recognizers excel at modeling spatial configurations but often underuse explicit motion dynamics that are crucial under appearance changes. We introduce a plug-and-play Wavelet Feature Stream that augments any skeleton backbone with time-frequency dynamics of joint velocities. Concretely, per-joint velocity sequences are transformed by the continuous wavelet transform (CWT) into multi-scale scalograms, from which a lightweight multi-scale CNN learns discriminative dynamic cues. The resulting descriptor is fused with the backbone representation for classification, requiring no changes to the backbone architecture or additional supervision. Across CASIA-B, the proposed stream delivers consistent gains on strong skeleton backbones (e.g., GaitMixer, GaitFormer, GaitGraph) and establishes a new skeleton-based state of the art when attached to GaitMixer. The improvements are especially pronounced under covariate shifts such as carrying bags (BG) and wearing coats (CL), highlighting the complementarity of explicit time-frequency modeling and standard spatio-temporal encoders.
△ Less
Submitted 3 April, 2026;
originally announced April 2026.
-
Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs
Authors:
Giyeong Oh,
Junghyun Lee,
Jaehyun Park,
Youngjae Yu,
Wonho Bae,
Junhyug Noh
Abstract:
Modern LLMs inherit strong priors from web-scale pretraining, which can limit the headroom of post-training data-selection strategies. While Active Preference Learning (APL) seeks to optimize query efficiency in online Direct Preference Optimization (DPO), the inherent richness of on-policy candidate pools often renders simple Random sampling a surprisingly formidable baseline. We evaluate uncerta…
▽ More
Modern LLMs inherit strong priors from web-scale pretraining, which can limit the headroom of post-training data-selection strategies. While Active Preference Learning (APL) seeks to optimize query efficiency in online Direct Preference Optimization (DPO), the inherent richness of on-policy candidate pools often renders simple Random sampling a surprisingly formidable baseline. We evaluate uncertainty-based APL against Random across harmlessness, helpfulness, and instruction-following settings, utilizing both reward models and LLM-as-a-judge proxies. We find that APL yields negligible improvements in proxy win-rates compared to Random. Crucially, we observe a dissociation where win-rate improves even as general capability -- measured by standard benchmarks -- degrades. APL fails to mitigate this capability collapse or reduce variance significantly better than random sampling. Our findings suggest that in the regime of strong pre-trained priors, the computational overhead of active selection is difficult to justify against the ``cheap diversity'' provided by simple random samples. Our code is available at https://github.com/BootsofLagrangian/random-vs-apl.
△ Less
Submitted 3 April, 2026;
originally announced April 2026.
-
SIREN: Spatially-Informed Reconstruction of Binaural Audio with Vision
Authors:
Mingyeong Song,
Seoyeon Ko,
Junhyug Noh
Abstract:
Binaural audio delivers spatial cues essential for immersion, yet most consumer videos are monaural due to capture constraints. We introduce SIREN, a visually guided mono to binaural framework that explicitly predicts left and right channels. A ViT-based encoder learns dual-head self-attention to produce a shared scene map and end-to-end L/R attention, replacing hand-crafted masks. A soft, anneale…
▽ More
Binaural audio delivers spatial cues essential for immersion, yet most consumer videos are monaural due to capture constraints. We introduce SIREN, a visually guided mono to binaural framework that explicitly predicts left and right channels. A ViT-based encoder learns dual-head self-attention to produce a shared scene map and end-to-end L/R attention, replacing hand-crafted masks. A soft, annealed spatial prior gently biases early L/R grounding, and a two-stage, confidence-weighted waveform-domain fusion (guided by mono reconstruction and interaural phase consistency) suppresses crosstalk when aggregating multi-crop and overlapping windows. Evaluated on FAIR-Play and MUSIC-Stereo, SIREN yields consistent gains on time-frequency and phase-sensitive metrics with competitive SNR. The design is modular and generic, requires no task-specific annotations, and integrates with standard audio-visual pipelines.
△ Less
Submitted 31 March, 2026;
originally announced March 2026.
-
ComVi: Context-Aware Optimized Comment Display in Video Playback
Authors:
Minsun Kim,
Dawon Lee,
Junyong Noh
Abstract:
On general video-sharing platforms like YouTube, comments are displayed independently of video playback. As viewers often read comments while watching a video, they may encounter ones referring to moments unrelated to the current scene, which can reveal spoilers and disrupt immersion. To address this problem, we present ComVi, a novel system that displays comments at contextually relevant moments,…
▽ More
On general video-sharing platforms like YouTube, comments are displayed independently of video playback. As viewers often read comments while watching a video, they may encounter ones referring to moments unrelated to the current scene, which can reveal spoilers and disrupt immersion. To address this problem, we present ComVi, a novel system that displays comments at contextually relevant moments, enabling viewers to see time-synchronized comments and video content together. We first map all comments to relevant video timestamps by computing audio-visual correlation, then construct the comment sequence through an optimization that considers temporal relevance, popularity (number of likes), and display duration for comfortable reading. In a user study, ComVi provided a significantly more engaging experience than conventional video interfaces (i.e., YouTube and Danmaku), with 71.9% of participants selecting ComVi as their most preferred interface.
△ Less
Submitted 27 March, 2026;
originally announced March 2026.
-
Building to Understand: Examining Teens' Technical and Socio-Ethical Pieces of Understandings in the Construction of Small Generative Language Models
Authors:
Luis Morales-Navarro,
Daniel J. Noh,
Lucianne Servat,
Carly Netting,
Yasmin B. Kafai,
Danaé Metaxa
Abstract:
The rising adoption of generative AI/ML technologies increases the need to support teens in developing AI/ML literacies. Child-computer interaction research argues that construction activities can support young people in understanding these systems and their implications. Recent exploratory studies demonstrate the feasibility of engaging teens in the construction of very small generative language…
▽ More
The rising adoption of generative AI/ML technologies increases the need to support teens in developing AI/ML literacies. Child-computer interaction research argues that construction activities can support young people in understanding these systems and their implications. Recent exploratory studies demonstrate the feasibility of engaging teens in the construction of very small generative language models (LMs). However, it is unclear how constructing such models may foster the development of teens' understanding of these systems from technical and socio-ethical perspectives. We conducted a week-long participatory design workshop in which sixteen teenagers constructed very small LMs to generate recipes, screenplays, and songs. Using thematic analysis, we identified technical and socio-ethical pieces of understandings that teens exhibited while designing generative LMs. This paper contributes (a) evidence of the kinds of pieces of understandings that teens have when constructing LMs and (b) a theory-backed framing to study novices' understandings of AI/ML systems.
△ Less
Submitted 15 April, 2026; v1 submitted 26 March, 2026;
originally announced March 2026.
-
The SPHEREx Ices Investigation: An Overview
Authors:
Gary J. Melnick,
Joseph L. Hora,
Matthew L. N. Ashby,
Volker Tolls,
Jaeyeong Kim,
Carey M. Lisse,
Roberta Paladini,
Michael W. Werner,
Jeong-Eun Lee,
Young-Jun Kim,
Miju Kang,
Yun-Ting Cheng,
James J. Bock,
Brendan P. Crill,
Ari Cukierman,
Olivier Dore,
Andreas Faisst,
Howard Hui,
Woong-Seob Jeong,
Chul-Hwan Kim,
Ho-Gyu Lee,
Jae-Joon Lee,
Daniel Masters,
Chi H. Nguyen,
Jinyoung Noh
, et al. (4 additional authors not shown)
Abstract:
SPHEREx is a NASA mission designed to perform an all-sky spectroscopic survey in the 0.75 - 5 $μ$m wavelength range. Its primary science objectives are to investigate: (1) inflationary cosmology, (2) the history of galaxy formation, and (3) the abundance of molecular ices - critical for prebiotic chemistry - found on the surfaces of interstellar dust grains within planet-forming regions. This pape…
▽ More
SPHEREx is a NASA mission designed to perform an all-sky spectroscopic survey in the 0.75 - 5 $μ$m wavelength range. Its primary science objectives are to investigate: (1) inflationary cosmology, (2) the history of galaxy formation, and (3) the abundance of molecular ices - critical for prebiotic chemistry - found on the surfaces of interstellar dust grains within planet-forming regions. This paper focuses on the third theme, the SPHEREx Ices investigation, for which SPHEREx is conducting a spectroscopic survey of nearly ten million preselected sources throughout the Milky Way and Magellanic Clouds to characterize their ice absorption features. By selecting targets based on infrared color, spatial isolation, and brightness, the Ices Investigation secures high-signal-to-noise spectra across a broad range of astrophysical environments that are relatively free of spectral contamination. Rather than attempting to decompose each spectrum into its individual ice components, the Ices Investigation prioritizes accurate measurements of the integrated optical depths of key molecular ice absorption features. This approach enables statistically powerful correlation studies between ice abundances and environmental parameters - including extinction, temperature, gas composition, radiation field strength, cosmic ray flux, and star formation activity. The data pipeline developed for this purpose incorporates machine learning for continuum estimation, drawing on both SPHEREx and ancillary datasets. Ultimately, the expansive spectral archive produced by SPHEREx, combined with targeted follow-up from facilities like JWST, will transform our understanding of Galactic ice formation, evolution, abundance and their inheritance into planetary systems and prebiotic inventories.
△ Less
Submitted 23 March, 2026;
originally announced March 2026.
-
SPHEREx Wide-Field Infrared Spectral Mapping of Interstellar Ices and Polycyclic Aromatic Hydrocarbons
Authors:
Joseph L. Hora,
Jinyoung K. Noh,
Gary J. Melnick,
Brandon S. Hensley,
Roberta Paladini,
Jeong-Eun Lee,
Matthew L. N. Ashby,
Volker Tolls,
Jaeyeong Kim,
Michael W. Werner,
James J. Bock,
Sean Bruton,
Shuang-Shuang Chen,
Tzu-Ching Chang,
Yi-Kuan Chiang,
Asantha Cooray,
Brendan P. Crill,
Ari J. Cukierman,
Olivier Doré,
Andreas L. Faisst,
Zhaoyu Huai,
Howard Hui,
Woong-Seob Jeong,
Miju Kang,
Phil M. Korngut
, et al. (10 additional authors not shown)
Abstract:
We present some of the first infrared spectral maps acquired by SPHEREx. These maps, which to our knowledge are the largest of their type ever compiled in the near-infrared, reveal multiple strong lines due to interstellar ices and polycyclic aromatic hydrocarbons (PAHs) throughout the Cygnus X and North American Nebula regions. The maps emphasize the strongest features arising from the 3 $μ$m H…
▽ More
We present some of the first infrared spectral maps acquired by SPHEREx. These maps, which to our knowledge are the largest of their type ever compiled in the near-infrared, reveal multiple strong lines due to interstellar ices and polycyclic aromatic hydrocarbons (PAHs) throughout the Cygnus X and North American Nebula regions. The maps emphasize the strongest features arising from the 3 $μ$m H$_2$O, 4.27 $μ$m CO$_2$, and 4.67 $μ$m CO lines and the 3.28 $μ$m PAH feature, all of which are detected over large areas with complex and filamentary spatial distributions. The ice absorption maps of H$_2$O and CO$_2$ in particular broadly trace dense, cold, and well-shielded regions across Cygnus X, consistent with the established picture of efficient ice formation in dense molecular clouds. The interstellar ice features are also detected abundantly in diffuse absorption over wide areas. The relative strength of the H$_2$O and CO$_2$ features varies among different lines of sight, indicating possible differences in local physical conditions or chemical variations. The 3.28 $μ$m PAH emission correlates with the emission from the 7.7 and 11.2 $μ$m features, but shows small differences that may trace the grain size distribution and variations in the ambient UV field. SPHEREx all-sky spectral imaging, of which only a small fraction is showcased in this work, will support numerous science investigations including the structure of the Galaxy, the physics of the interstellar medium, and the chemistry of stars.
△ Less
Submitted 12 March, 2026;
originally announced March 2026.
-
A Guideline-Aware AI Agent for Zero-Shot Target Volume Auto-Delineation
Authors:
Yoon Jo Kim,
Wonyoung Cho,
Jongmin Lee,
Han Joo Chae,
Hyunki Park,
Sang Hoon Seo,
Jae Myung Noh,
Kyungmi Yang,
Dongryul Oh,
Jin Sung Kim
Abstract:
Delineating the clinical target volume (CTV) in radiotherapy involves complex margins constrained by tumor location and anatomical barriers. While deep learning models automate this process, their rigid reliance on expert-annotated data requires costly retraining whenever clinical guidelines update. To overcome this limitation, we introduce OncoAgent, a novel guideline-aware AI agent framework tha…
▽ More
Delineating the clinical target volume (CTV) in radiotherapy involves complex margins constrained by tumor location and anatomical barriers. While deep learning models automate this process, their rigid reliance on expert-annotated data requires costly retraining whenever clinical guidelines update. To overcome this limitation, we introduce OncoAgent, a novel guideline-aware AI agent framework that seamlessly converts textual clinical guidelines into three-dimensional target contours in a training-free manner. Evaluated on esophageal cancer cases, the agent achieves a zero-shot Dice similarity coefficient of 0.842 for the CTV and 0.880 for the planning target volume, demonstrating performance highly comparable to a fully supervised nnU-Net baseline. Notably, in a blinded clinical evaluation, physicians strongly preferred OncoAgent over the supervised baseline, rating it higher in guideline compliance, modification effort, and clinical acceptability. Furthermore, the framework generalizes zero-shot to alternative esophageal guidelines and other anatomical sites (e.g., prostate) without any retraining. Beyond mere volumetric overlap, our agent-based paradigm offers near-instantaneous adaptability to alternative guidelines, providing a scalable and transparent pathway toward interpretability in radiotherapy treatment planning.
△ Less
Submitted 25 June, 2026; v1 submitted 10 March, 2026;
originally announced March 2026.
-
X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake Detection
Authors:
Youngseo Kim,
Kwan Yun,
Seokhyeon Hong,
Sihun Cha,
Colette Suhjung Koo,
Junyong Noh
Abstract:
The surge of highly realistic synthetic videos produced by contemporary generative systems has significantly increased the risk of malicious use, challenging both humans and existing detectors. Against this backdrop, we take a generator-side view and observe that internal cross-attention mechanisms in these models encode fine-grained speech-motion alignment, offering useful correspondence cues for…
▽ More
The surge of highly realistic synthetic videos produced by contemporary generative systems has significantly increased the risk of malicious use, challenging both humans and existing detectors. Against this backdrop, we take a generator-side view and observe that internal cross-attention mechanisms in these models encode fine-grained speech-motion alignment, offering useful correspondence cues for forgery detection. Building on this insight, we propose X-AVDT, a robust and generalizable deepfake detector that probes generator-internal audio-visual signals accessed via DDIM inversion to expose these cues. X-AVDT extracts two complementary signals: (i) a video composite capturing inversion-induced discrepancies, and (ii) an audio-visual cross-attention feature reflecting modality alignment enforced during generation. To enable faithful cross-generator evaluation, we further introduce MMDF, a new multimodal deepfake dataset spanning diverse manipulation types and rapidly evolving synthesis paradigms, including GANs, diffusion, and flow-matching. Extensive experiments demonstrate that X-AVDT achieves leading performance on MMDF and generalizes strongly to external benchmarks and unseen generators, outperforming existing methods with accuracy improved by 13.1%. Our findings highlight the importance of leveraging internal audio-visual consistency cues for robustness to future generators in deepfake detection.
△ Less
Submitted 9 March, 2026;
originally announced March 2026.
-
"You Can Actually Do Something'': Shifts in High School Computer Science Teachers' Conceptions of AI/ML Systems and Algorithmic Justice
Authors:
Daniel J. Noh,
Deborah A. Fields,
Yasmin B. Kafai,
Danaé Metaxa
Abstract:
The recent proliferation of artificial intelligence and machine learning (AI/ML) systems highlights the need for all people to develop effective competencies to interact with and examine AI/ML systems. We study shifts in five experienced high school CS teachers' understanding of AI/ML systems after one year of participatory design, where they co-developed lessons on AI auditing, a systematic metho…
▽ More
The recent proliferation of artificial intelligence and machine learning (AI/ML) systems highlights the need for all people to develop effective competencies to interact with and examine AI/ML systems. We study shifts in five experienced high school CS teachers' understanding of AI/ML systems after one year of participatory design, where they co-developed lessons on AI auditing, a systematic method to query AI/ML systems. Drawing on individual and group interviews, we found that teachers' perspectives became more situated, grounding their understanding in everyday contexts; more critical, reflecting growing awareness of harms; and more agentic, highlighting possibilities for action. Further, across all three perspectives, teachers consistently framed algorithmic justice through their role as educators, situating their concerns within their school communities. In the discussion, we consider the ways teachers' perspectives shifted, how AI auditing can shape these shifts, and the implications of these findings on AI literacy for both teachers and students.
△ Less
Submitted 27 March, 2026; v1 submitted 17 February, 2026;
originally announced February 2026.
-
Clutt3R-Seg: Sparse-view 3D Instance Segmentation for Language-grounded Grasping in Cluttered Scenes
Authors:
Jeongho Noh,
Tai Hyoung Rhee,
Eunho Lee,
Jeongyun Kim,
Sunwoo Lee,
Ayoung Kim
Abstract:
Reliable 3D instance segmentation is fundamental to language-grounded robotic manipulation. Its critical application lies in cluttered environments, where occlusions, limited viewpoints, and noisy masks degrade perception. To address these challenges, we present Clutt3R-Seg, a zero-shot pipeline for robust 3D instance segmentation for language-grounded grasping in cluttered scenes. Our key idea is…
▽ More
Reliable 3D instance segmentation is fundamental to language-grounded robotic manipulation. Its critical application lies in cluttered environments, where occlusions, limited viewpoints, and noisy masks degrade perception. To address these challenges, we present Clutt3R-Seg, a zero-shot pipeline for robust 3D instance segmentation for language-grounded grasping in cluttered scenes. Our key idea is to introduce a hierarchical instance tree of semantic cues. Unlike prior approaches that attempt to refine noisy masks, our method leverages them as informative cues: through cross-view grouping and conditional substitution, the tree suppresses over- and under-segmentation, yielding view-consistent masks and robust 3D instances. Each instance is enriched with open-vocabulary semantic embeddings, enabling accurate target selection from natural language instructions. To handle scene changes during multi-stage tasks, we further introduce a consistency-aware update that preserves instance correspondences from only a single post-interaction image, allowing efficient adaptation without rescanning. Clutt3R-Seg is evaluated on both synthetic and real-world datasets, and validated on a real robot. Across all settings, it consistently outperforms state-of-the-art baselines in cluttered and sparse-view scenarios. Even on the most challenging heavy-clutter sequences, Clutt3R-Seg achieves an AP@25 of 61.66, over 2.2x higher than baselines, and with only four input views it surpasses MaskClustering with eight views by more than 2x. The code is available at: https://github.com/jeonghonoh/clutt3r-seg.
△ Less
Submitted 12 February, 2026;
originally announced February 2026.
-
Improving Methodologies for LLM Evaluations Across Global Languages
Authors:
Akriti Vij,
Benjamin Chua,
Darshini Ramiah,
En Qi Ng,
Mahran Morsidi,
Naga Nikshith Gangarapu,
Sharmini Johnson,
Vanessa Wilfred,
Vikneswaran Kumaran,
Wan Sie Lee,
Wenzhuo Yang,
Yongsen Zheng,
Bill Black,
Boming Xia,
Frank Sun,
Hao Zhang,
Qinghua Lu,
Suyu Ma,
Yue Liu,
Chi-kiu Lo,
Fatemeh Azadi,
Isar Nejadgholi,
Sowmya Vajjala,
Agnes Delaborde,
Nicolas Rolin
, et al. (21 additional authors not shown)
Abstract:
As frontier AI models are deployed globally, it is essential that their behaviour remains safe and reliable across diverse linguistic and cultural contexts. To examine how current model safeguards hold up in such settings, participants from the International Network for Advanced AI Measurement, Evaluation and Science, including representatives from Singapore, Japan, Australia, Canada, the EU, Fran…
▽ More
As frontier AI models are deployed globally, it is essential that their behaviour remains safe and reliable across diverse linguistic and cultural contexts. To examine how current model safeguards hold up in such settings, participants from the International Network for Advanced AI Measurement, Evaluation and Science, including representatives from Singapore, Japan, Australia, Canada, the EU, France, Kenya, South Korea and the UK conducted a joint multilingual evaluation exercise. Led by Singapore AISI, two open-weight models were tested across ten languages spanning high and low resourced groups: Cantonese English, Farsi, French, Japanese, Korean, Kiswahili, Malay, Mandarin Chinese and Telugu. Over 6,000 newly translated prompts were evaluated across five harm categories (privacy, non-violent crime, violent crime, intellectual property and jailbreak robustness), using both LLM-as-a-judge and human annotation.
The exercise shows how safety behaviours can vary across languages. These include differences in safeguard robustness across languages and harm types and variation in evaluator reliability (LLM-as-judge vs. human review). Further, it also generated methodological insights for improving multilingual safety evaluations, such as the need for culturally contextualised translations, stress-tested evaluator prompts and clearer human annotation guidelines. This work represents an initial step toward a shared framework for multilingual safety testing of advanced AI systems and calls for continued collaboration with the wider research community and industry.
△ Less
Submitted 22 January, 2026;
originally announced January 2026.
-
Improving Methodologies for Agentic Evaluations Across Domains: Leakage of Sensitive Information, Fraud and Cybersecurity Threats
Authors:
Ee Wei Seah,
Yongsen Zheng,
Naga Nikshith,
Mahran Morsidi,
Gabriel Waikin Loh Matienzo,
Nigel Gay,
Akriti Vij,
Benjamin Chua,
En Qi Ng,
Sharmini Johnson,
Vanessa Wilfred,
Wan Sie Lee,
Anna Davidson,
Catherine Devine,
Erin Zorer,
Gareth Holvey,
Harry Coppock,
James Walpole,
Jerome Wynee,
Magda Dubois,
Michael Schmatz,
Patrick Keane,
Sam Deverett,
Bill Black,
Bo Yan
, et al. (45 additional authors not shown)
Abstract:
The rapid rise of autonomous AI systems and advancements in agent capabilities are introducing new risks due to reduced oversight of real-world interactions. Yet agent testing remains nascent and is still a developing science. As AI agents begin to be deployed globally, it is important that they handle different languages and cultures accurately and securely.
To address this, participants from T…
▽ More
The rapid rise of autonomous AI systems and advancements in agent capabilities are introducing new risks due to reduced oversight of real-world interactions. Yet agent testing remains nascent and is still a developing science. As AI agents begin to be deployed globally, it is important that they handle different languages and cultures accurately and securely.
To address this, participants from The International Network for Advanced AI Measurement, Evaluation and Science, including representatives from Singapore, Japan, Australia, Canada, the European Commission, France, Kenya, South Korea, and the United Kingdom have come together to align approaches to agentic evaluations.
This is the third exercise, building on insights from two earlier joint testing exercises conducted by the Network in November 2024 and February 2025. The objective is to further refine best practices for testing advanced AI systems.
The exercise was split into two strands: (1) common risks, including leakage of sensitive information and fraud, led by Singapore AISI; and (2) cybersecurity, led by UK AISI. A mix of open and closed-weight models were evaluated against tasks from various public agentic benchmarks. Given the nascency of agentic testing, our primary focus was on understanding methodological issues in conducting such tests, rather than examining test results or model capabilities. This collaboration marks an important step forward as participants work together to advance the science of agentic evaluations.
△ Less
Submitted 22 January, 2026;
originally announced January 2026.
-
Deep Learning Based Facial Retargeting Using Local Patches
Authors:
Yeonsoo Choi,
Inyup Lee,
Sihun Cha,
Seonghyeon Kim,
Sunjin Jung,
Junyong Noh
Abstract:
In the era of digital animation, the quest to produce lifelike facial animations for virtual characters has led to the development of various retargeting methods. While the retargeting facial motion between models of similar shapes has been very successful, challenges arise when the retargeting is performed on stylized or exaggerated 3D characters that deviate significantly from human facial struc…
▽ More
In the era of digital animation, the quest to produce lifelike facial animations for virtual characters has led to the development of various retargeting methods. While the retargeting facial motion between models of similar shapes has been very successful, challenges arise when the retargeting is performed on stylized or exaggerated 3D characters that deviate significantly from human facial structures. In this scenario, it is important to consider the target character's facial structure and possible range of motion to preserve the semantics assumed by the original facial motions after the retargeting. To achieve this, we propose a local patch-based retargeting method that transfers facial animations captured in a source performance video to a target stylized 3D character. Our method consists of three modules. The Automatic Patch Extraction Module extracts local patches from the source video frame. These patches are processed through the Reenactment Module to generate correspondingly re-enacted target local patches. The Weight Estimation Module calculates the animation parameters for the target character at every frame for the creation of a complete facial animation sequence. Extensive experiments demonstrate that our method can successfully transfer the semantic meaning of source facial expressions to stylized characters with considerable variations in facial feature proportion.
△ Less
Submitted 13 January, 2026;
originally announced January 2026.
-
Automated Domain Question Mapping (DQM) with Educational Learning Materials
Authors:
Jiho Noh,
Mukhesh Raghava Katragadda,
Dabae Lee
Abstract:
Concept maps have been widely utilized in education to depict knowledge structures and the interconnections between disciplinary concepts. Nonetheless, devising a computational method for automatically constructing a concept map from unstructured educational materials presents challenges due to the complexity and variability of educational content. We focus primarily on two challenges: (1) the lac…
▽ More
Concept maps have been widely utilized in education to depict knowledge structures and the interconnections between disciplinary concepts. Nonetheless, devising a computational method for automatically constructing a concept map from unstructured educational materials presents challenges due to the complexity and variability of educational content. We focus primarily on two challenges: (1) the lack of disciplinary concepts that are specifically designed for multi-level pedagogical purposes from low-order to high-order thinking, and (2) the limited availability of labeled data concerning disciplinary concepts and their interrelationships. To tackle these challenges, this research introduces an innovative approach for constructing Domain Question Maps (DQMs), rather than traditional concept maps. By formulating specific questions aligned with learning objectives, DQMs enhance knowledge representation and improve readiness for learner engagement. The findings indicate that the proposed method can effectively generate educational questions and discern hierarchical relationships among them, leading to structured question maps that facilitate personalized and adaptive learning in downstream applications.
△ Less
Submitted 11 January, 2026;
originally announced January 2026.
-
Collective behavior in the nonreciprocal multi-species Vicsek model
Authors:
Chul-Ung Woo,
Heiko Rieger,
Jae Dong Noh
Abstract:
We investigate collective behavior in a $Q$-species Vicsek model with a nonreciprocal velocity alignment interaction. This system is characterized by a constant phase shift $α$ in the inter-species velocity alignment rule. While the phase shift renders the interaction nonreciprocal, the system is globally invariant under any permutations of particle species, possessing Potts symmetry. The combinat…
▽ More
We investigate collective behavior in a $Q$-species Vicsek model with a nonreciprocal velocity alignment interaction. This system is characterized by a constant phase shift $α$ in the inter-species velocity alignment rule. While the phase shift renders the interaction nonreciprocal, the system is globally invariant under any permutations of particle species, possessing Potts symmetry. The combination of Potts symmetry and nonreciprocity gives rise to a rich phase diagram. The nonreciprocal phase shift generates either counter-clockwise or clockwise chirality. Potts symmetry can be broken spontaneously. Consequently, the system exhibits four distinct phases: A species-mixed chiral phase where particles perform counter-clockwise chiral motion with quasi-long-range order, a species separation phase where Potts symmetry is broken and species-separated particles form vortex cells with clockwise chirality, a coexistence phase, and a disordered phase, for $0<α<π$. We derive a Boltzmann equation and a hydrodynamic equation describing the system in the continuum limit, and present analytic arguments for the emergence of chirality and species separation.
△ Less
Submitted 11 August, 2026; v1 submitted 21 December, 2025;
originally announced December 2025.
-
Nonreciprocal yet Symmetric Multi-Species Active Matter: Emergence of Chirality and Species Separation
Authors:
Chul-Ung Woo,
Heiko Rieger,
Jae Dong Noh
Abstract:
Nonreciprocal active matter systems typically feature an asymmetric role among interacting agents, such as a pursuer-evader relationship. We propose a multi-species nonreciprocal active matter model that is invariant under permutations of the particle species. The nonreciprocal, yet symmetric, interactions emerge from a constant phase shift in the velocity alignment interactions, rather than from…
▽ More
Nonreciprocal active matter systems typically feature an asymmetric role among interacting agents, such as a pursuer-evader relationship. We propose a multi-species nonreciprocal active matter model that is invariant under permutations of the particle species. The nonreciprocal, yet symmetric, interactions emerge from a constant phase shift in the velocity alignment interactions, rather than from an asymmetric coupling matrix. This system possessing permutation symmetry displays rich collective behaviors, including a species-mixed chiral phase with quasi-long-range polar order and a species separation phase characterized by vortex cells. The system also displays a coexistence phase of the chiral and the species separation phases, in which intriguing dynamic patterns emerge. These rich collective behaviors are a consequence of the interplay between nonreciprocity and permutation symmetry.
△ Less
Submitted 21 December, 2025;
originally announced December 2025.
-
Beyond Softmax: Dual-Branch Sigmoid Architecture for Accurate Class Activation Maps
Authors:
Yoojin Oh,
Junhyug Noh
Abstract:
Class Activation Mapping (CAM) and its extensions have become indispensable tools for visualizing the evidence behind deep network predictions. However, by relying on a final softmax classifier, these methods suffer from two fundamental distortions: additive logit shifts that arbitrarily bias importance scores, and sign collapse that conflates excitatory and inhibitory features. We propose a simpl…
▽ More
Class Activation Mapping (CAM) and its extensions have become indispensable tools for visualizing the evidence behind deep network predictions. However, by relying on a final softmax classifier, these methods suffer from two fundamental distortions: additive logit shifts that arbitrarily bias importance scores, and sign collapse that conflates excitatory and inhibitory features. We propose a simple, architecture-agnostic dual-branch sigmoid head that decouples localization from classification. Given any pretrained model, we clone its classification head into a parallel branch ending in per-class sigmoid outputs, freeze the original softmax head, and fine-tune only the sigmoid branch with class-balanced binary supervision. At inference, softmax retains recognition accuracy, while class evidence maps are generated from the sigmoid branch -- preserving both magnitude and sign of feature contributions. Our method integrates seamlessly with most CAM variants and incurs negligible overhead. Extensive evaluations on fine-grained tasks (CUB-200-2011, Stanford Cars) and WSOL benchmarks (ImageNet-1K, OpenImages30K) show improved explanation fidelity and consistent Top-1 Localization gains -- without any drop in classification accuracy. Code is available at https://github.com/finallyupper/beyond-softmax.
△ Less
Submitted 5 November, 2025;
originally announced November 2025.
-
The SPHEREx Satellite Mission
Authors:
James J. Bock,
Asad M. Aboobaker,
Joseph Adamo,
Rachel Akeson,
John M. Alred,
Farah Alibay,
Matthew L. N. Ashby,
Yoonsoo P. Bach,
Lindsey E. Bleem,
Douglas Bolton,
David F. Braun,
Sean Bruton,
Sean A. Bryan,
Tzu-Ching Chang,
Shuang-Shuang Chen,
Yun-Ting Cheng,
James R. Cheshire IV,
Yi-Kuan Chiang,
Jean Choppin de Janvry,
Samuel Condon,
Walter R. Cook,
Asantha Cooray,
Brendan P. Crill,
Ari J. Cukierman,
Olivier Dore
, et al. (89 additional authors not shown)
Abstract:
SPHEREx, a NASA explorer satellite launched on 11 March 2025, is carrying out the first all-sky near-infrared spectral survey. The satellite observes in 102 spectral bands from 0.75 to 5.0 um with a resolving power ranging from 35 to 130 in 6.2 arcsecond pixels. The observatory obtains a 5-sigma depth of 19.5 - 19.9 AB mag for 0.75 to 3.8 um and 17.8 - 18.8 AB mag for 3.8 to 5.0 um after mapping t…
▽ More
SPHEREx, a NASA explorer satellite launched on 11 March 2025, is carrying out the first all-sky near-infrared spectral survey. The satellite observes in 102 spectral bands from 0.75 to 5.0 um with a resolving power ranging from 35 to 130 in 6.2 arcsecond pixels. The observatory obtains a 5-sigma depth of 19.5 - 19.9 AB mag for 0.75 to 3.8 um and 17.8 - 18.8 AB mag for 3.8 to 5.0 um after mapping the full sky four times over two years. Scientifically, SPHEREx will produce a large galaxy redshift survey over the full sky, intended to constrain the amplitude of inflationary non-Gaussianity. The observations will produce two deep spectral maps near the ecliptic poles that will use intensity mapping to probe the evolution of galaxies over cosmic history. By mapping the depth of infrared absorption features over the Galactic plane, SPHEREx will comprehensively survey the abundance and composition of water and other biogenic ice species in the interstellar medium. The initial data are rapidly released in the form of spectral images to the public. The project will release specialized data products over the life of the mission as the surveys proceed. The science team will also produce specialized spectral catalogs on planet-bearing and low-mass stars, solar system objects, and galaxy clusters 3 years after launch. We describe the design of the instrument and spacecraft, which flow from the core science requirements. Finally, we present an initial evaluation of the in-flight performance and key characteristics.
△ Less
Submitted 15 December, 2025; v1 submitted 4 November, 2025;
originally announced November 2025.
-
Pinpointing Trigger Moment for Grounded Video QA: Enhancing Spatio-temporal Grounding in Multimodal Large Language Models
Authors:
Jinhwan Seo,
Yoonki Cho,
Junhyug Noh,
Sung-eui Yoon
Abstract:
In this technical report, we introduce a framework to address Grounded Video Question Answering (GVQA) task for the ICCV 2025 Perception Test Challenge. The GVQA task demands robust multimodal models capable of complex reasoning over video content, grounding the resulting answers visually, and tracking the referenced objects temporally. To achieve this capability, our proposed approach decomposes…
▽ More
In this technical report, we introduce a framework to address Grounded Video Question Answering (GVQA) task for the ICCV 2025 Perception Test Challenge. The GVQA task demands robust multimodal models capable of complex reasoning over video content, grounding the resulting answers visually, and tracking the referenced objects temporally. To achieve this capability, our proposed approach decomposes the GVQA task into a three-stage pipeline: (1) Video Reasoning \& QA, (2) Spatio-temporal Grounding and (3) Tracking. Our key contribution is the introduction of a trigger moment, derived from our proposed CORTEX prompt, which pinpoints the single most visible frame of a target object to serve as a robust anchor for grounding and tracking. To this end, we achieve the HOTA score of 0.4968, which marks a significant improvement over the previous year's winning score of 0.2704 on GVQA task.
△ Less
Submitted 3 November, 2025;
originally announced November 2025.
-
Diffusion-Driven Two-Stage Active Learning for Low-Budget Semantic Segmentation
Authors:
Jeongin Kim,
Wonho Bae,
YouLee Han,
Giyeong Oh,
Youngjae Yu,
Danica J. Sutherland,
Junhyug Noh
Abstract:
Semantic segmentation demands dense pixel-level annotations, which can be prohibitively expensive - especially under extremely constrained labeling budgets. In this paper, we address the problem of low-budget active learning for semantic segmentation by proposing a novel two-stage selection pipeline. Our approach leverages a pre-trained diffusion model to extract rich multi-scale features that cap…
▽ More
Semantic segmentation demands dense pixel-level annotations, which can be prohibitively expensive - especially under extremely constrained labeling budgets. In this paper, we address the problem of low-budget active learning for semantic segmentation by proposing a novel two-stage selection pipeline. Our approach leverages a pre-trained diffusion model to extract rich multi-scale features that capture both global structure and fine details. In the first stage, we perform a hierarchical, representation-based candidate selection by first choosing a small subset of representative pixels per image using MaxHerding, and then refining these into a diverse global pool. In the second stage, we compute an entropy-augmented disagreement score (eDALD) over noisy multi-scale diffusion features to capture both epistemic uncertainty and prediction confidence, selecting the most informative pixels for annotation. This decoupling of diversity and uncertainty lets us achieve high segmentation accuracy with only a tiny fraction of labeled pixels. Extensive experiments on four benchmarks (CamVid, ADE-Bed, Cityscapes, and Pascal-Context) demonstrate that our method significantly outperforms existing baselines under extreme pixel-budget regimes. Our code is available at https://github.com/jn-kim/two-stage-edald.
△ Less
Submitted 25 October, 2025;
originally announced October 2025.
-
SteeringTTA: Guiding Diffusion Trajectories for Robust Test-Time-Adaptation
Authors:
Jihyun Yu,
Yoojin Oh,
Wonho Bae,
Mingyu Kim,
Junhyug Noh
Abstract:
Test-time adaptation (TTA) aims to correct performance degradation of deep models under distribution shifts by updating models or inputs using unlabeled test data. Input-only diffusion-based TTA methods improve robustness for classification to corruptions but rely on gradient guidance, limiting exploration and generalization across distortion types. We propose SteeringTTA, an inference-only framew…
▽ More
Test-time adaptation (TTA) aims to correct performance degradation of deep models under distribution shifts by updating models or inputs using unlabeled test data. Input-only diffusion-based TTA methods improve robustness for classification to corruptions but rely on gradient guidance, limiting exploration and generalization across distortion types. We propose SteeringTTA, an inference-only framework that adapts Feynman-Kac steering to guide diffusion-based input adaptation for classification with rewards driven by pseudo-label. SteeringTTA maintains multiple particle trajectories, steered by a combination of cumulative top-K probabilities and an entropy schedule, to balance exploration and confidence. On ImageNet-C, SteeringTTA consistently outperforms the baseline without any model updates or source data.
△ Less
Submitted 16 October, 2025;
originally announced October 2025.
-
Ultrafast exciton polaron dynamics in 2D Ruddlesden Popper lead halide perovskites
Authors:
Anirban Mondal,
Kwang Jin Lee,
Seungmin Lee,
Oui Jin Oh,
Myeongsam Jen,
Jun Hong Noh,
Jong Min Lim,
Minhaeng Cho
Abstract:
Two dimensional Ruddlesden Popper (2D) RP hybrid perovskites exhibit substantially higher chemical and structural stability than their three dimensional (3D) counterparts, positioning them as promising candidates for next generation optoelectronics. While quasiparticle dynamics in 3D perovskites are well studied, their 2D analogues remain comparatively underexplored. Here we systematically investi…
▽ More
Two dimensional Ruddlesden Popper (2D) RP hybrid perovskites exhibit substantially higher chemical and structural stability than their three dimensional (3D) counterparts, positioning them as promising candidates for next generation optoelectronics. While quasiparticle dynamics in 3D perovskites are well studied, their 2D analogues remain comparatively underexplored. Here we systematically investigate the branching, dynamics, and interactions of free excitons (FEs) and exciton polarons EPs in monolayer 2D RP perovskites using visible range femtosecond transient absorption TA spectroscopy. We prepared monolayer 2D RP perovskite thin films with varied organic spacers and distinct fabrication routes for comparative analysis. We find that the EP binding energy is 50 65 meV in (BA)2PbI4 and 37 39 meV in (PEA)2PbI4, consistent with spacer layer dependent coupling as corroborated by FTIR. We reveal a dynamic equilibrium between FEs and EPs that persists for tens of picoseconds. Notably, the TA signatures differ by fabrication route films from the newly developed process show weaker Auger annihilation and a reduced hot phonon bottleneck than those from the conventional route trends consistent with fewer traps and impurities in the former. Coupled rate equation modeling reproduces the transients and quantifies the processes of hot carrier relaxation, exciton exciton annihilation, exciton phonon coupling, and FE EP interconversion. These results demonstrate that the chemical synthetic process (fabrication route) and spacer choice significantly influence EP stability and population balance, offering practical levers for engineering ultrafast photophysics in 2D perovskites and guiding the design of advanced optoelectronic devices.
△ Less
Submitted 15 October, 2025;
originally announced October 2025.
-
Orbital Frontiers: Harnessing Higher Modes in Photonic Simulators
Authors:
Jiho Noh,
Julian Schulz,
Wladimir Benalcazar,
Christina Jörg
Abstract:
Photonic platforms have emerged as versatile and powerful classical simulators of quantum dynamics, providing clean, controllable optical analogs of extended structured (i.e., crystalline) electronic systems. While most realizations to date have used only the fundamental mode in each site, recent advances in structured light - particularly the use of higher-order spatial modes, including those wit…
▽ More
Photonic platforms have emerged as versatile and powerful classical simulators of quantum dynamics, providing clean, controllable optical analogs of extended structured (i.e., crystalline) electronic systems. While most realizations to date have used only the fundamental mode in each site, recent advances in structured light - particularly the use of higher-order spatial modes, including those with orbital angular momentum - are enabling richer dynamics and new functionalities. These additional degrees of freedom facilitate the emulation of phenomena ranging from topological band structures and synthetic gauge fields to orbitronics. In this perspective, we discuss how exploiting the internal structure of higher-order modes is reshaping the scope and capabilities of photonic platforms for simulating quantum phenomena.
△ Less
Submitted 6 October, 2025;
originally announced October 2025.
-
Vicinity-Guided Discriminative Latent Diffusion for Privacy-Preserving Domain Adaptation
Authors:
Jing Wang,
Wonho Bae,
Jiahong Chen,
Wenxu Wang,
Junhyug Noh
Abstract:
Recent work on latent diffusion models (LDMs) has focused almost exclusively on generative tasks, leaving their potential for discriminative transfer largely unexplored. We introduce Discriminative Vicinity Diffusion (DVD), a novel LDM-based framework for a more practical variant of source-free domain adaptation (SFDA): the source provider may share not only a pre-trained classifier but also an au…
▽ More
Recent work on latent diffusion models (LDMs) has focused almost exclusively on generative tasks, leaving their potential for discriminative transfer largely unexplored. We introduce Discriminative Vicinity Diffusion (DVD), a novel LDM-based framework for a more practical variant of source-free domain adaptation (SFDA): the source provider may share not only a pre-trained classifier but also an auxiliary latent diffusion module, trained once on the source data and never exposing raw source samples. DVD encodes each source feature's label information into its latent vicinity by fitting a Gaussian prior over its k-nearest neighbors and training the diffusion network to drift noisy samples back to label-consistent representations. During adaptation, we sample from each target feature's latent vicinity, apply the frozen diffusion module to generate source-like cues, and use a simple InfoNCE loss to align the target encoder to these cues, explicitly transferring decision boundaries without source access. Across standard SFDA benchmarks, DVD outperforms state-of-the-art methods. We further show that the same latent diffusion module enhances the source classifier's accuracy on in-domain data and boosts performance in supervised classification and domain generalization experiments. DVD thus reinterprets LDMs as practical, privacy-preserving bridges for explicit knowledge transfer, addressing a core challenge in source-free domain adaptation that prior methods have yet to solve.
△ Less
Submitted 11 November, 2025; v1 submitted 30 September, 2025;
originally announced October 2025.
-
Magnetic flux ropes within reconnection exhausts close to the centers of heliospheric current sheets near the Sun
Authors:
Dae-Young Lee,
Dooyoung Choi,
Kyung-Eun Choi,
Sung Jun Noh
Abstract:
Understanding the relationship between magnetic flux ropes and magnetic reconnection is fundamental to both space and astrophysical plasma studies. In this study, we report on two consecutive heliospheric current sheet (HCS) crossings by Parker Solar Probe (PSP), separated by ~10.5 hours, at a heliocentric distance of ~12 solar radii. For each crossing, we identified a series of flux ropes embedde…
▽ More
Understanding the relationship between magnetic flux ropes and magnetic reconnection is fundamental to both space and astrophysical plasma studies. In this study, we report on two consecutive heliospheric current sheet (HCS) crossings by Parker Solar Probe (PSP), separated by ~10.5 hours, at a heliocentric distance of ~12 solar radii. For each crossing, we identified a series of flux ropes embedded within reconnection exhausts on the sunward side of X-line. Their passage durations are <20sec, corresponding to spatial scales of a few thousands kilometers, still larger by three orders of magnitude than ion inertial length. This identification was possible particularly during intervals when PSP was closest to the HCS center. These flux ropes are distinguishable from the background exhausts by enhancements in magnetic field strength, significantly in the guide field component, travel speed slightly faster (typically by <10km/s) than surrounding outflows, and often accompanied by, though not always, increased density and reduced temperature. We attribute their origin to secondary reconnection within the exhausts and subsequent merging of smaller flux ropes into larger structures, consistent with predictions by various simulations. We suggest that such flux ropes are most readily identifiable at the HCS center where the background magnetic field is weakest so that the relative enhancement in flux rope field becomes most prominent. This observational advantage is particularly notable closer to the Sun where the high ambient magnetic field strength can otherwise obscure such structures unless the spacecraft trajectory remains within the HCS central region for a sufficient duration.
△ Less
Submitted 20 September, 2025;
originally announced September 2025.
-
Dual Interaction Network with Cross-Image Attention for Medical Image Segmentation
Authors:
Jeonghyun Noh,
Wangsu Jeon,
Jinsun Park
Abstract:
Medical image segmentation is a crucial method for assisting professionals in diagnosing various diseases through medical imaging. However, various factors such as noise, blurriness, and low contrast often hinder the accurate diagnosis of diseases. While numerous image enhancement techniques can mitigate these issues, they may also alter crucial information needed for accurate diagnosis in the ori…
▽ More
Medical image segmentation is a crucial method for assisting professionals in diagnosing various diseases through medical imaging. However, various factors such as noise, blurriness, and low contrast often hinder the accurate diagnosis of diseases. While numerous image enhancement techniques can mitigate these issues, they may also alter crucial information needed for accurate diagnosis in the original image. Conventional image fusion strategies, such as feature concatenation can address this challenge. However, they struggle to fully leverage the advantages of both original and enhanced images while suppressing the side effects of the enhancements. To overcome the problem, we propose a dual interactive fusion module (DIFM) that effectively exploits mutual complementary information from the original and enhanced images. DIFM employs cross-attention bidirectionally to simultaneously attend to corresponding spatial information across different images, subsequently refining the complementary features via global spatial attention. This interaction leverages low- to high-level features implicitly associated with diverse structural attributes like edges, blobs, and object shapes, resulting in enhanced features that embody important spatial characteristics. In addition, we introduce a multi-scale boundary loss based on gradient extraction to improve segmentation accuracy at object boundaries. Experimental results on the ACDC and Synapse datasets demonstrate the superiority of the proposed method quantitatively and qualitatively. Code available at: https://github.com/JJeong-Gari/DIN
△ Less
Submitted 7 September, 2025;
originally announced September 2025.
-
3DPillars: Pillar-based two-stage 3D object detection
Authors:
Jongyoun Noh,
Junghyup Lee,
Hyekang Park,
Bumsub Ham
Abstract:
PointPillars is the fastest 3D object detector that exploits pseudo image representations to encode features for 3D objects in a scene. Albeit efficient, PointPillars is typically outperformed by state-of-the-art 3D detection methods due to the following limitations: 1) The pseudo image representations fail to preserve precise 3D structures, and 2) they make it difficult to adopt a two-stage detec…
▽ More
PointPillars is the fastest 3D object detector that exploits pseudo image representations to encode features for 3D objects in a scene. Albeit efficient, PointPillars is typically outperformed by state-of-the-art 3D detection methods due to the following limitations: 1) The pseudo image representations fail to preserve precise 3D structures, and 2) they make it difficult to adopt a two-stage detection pipeline using 3D object proposals that typically shows better performance than a single-stage approach. We introduce in this paper the first two-stage 3D detection framework exploiting pseudo image representations, narrowing the performance gaps between PointPillars and state-of-the-art methods, while retaining its efficiency. Our framework consists of two novel components that overcome the aforementioned limitations of PointPillars: First, we introduce a new CNN architecture, dubbed 3DPillars, that enables learning 3D voxel-based features from the pseudo image representation efficiently using 2D convolutions. The basic idea behind 3DPillars is that 3D features from voxels can be viewed as a stack of pseudo images. To implement this idea, we propose a separable voxel feature module that extracts voxel-based features without using 3D convolutions. Second, we introduce an RoI head with a sparse scene context feature module that aggregates multi-scale features from 3DPillars to obtain a sparse scene feature. This enables adopting a two-stage pipeline effectively, and fully leveraging contextual information of a scene to refine 3D object proposals. Experimental results on the KITTI and Waymo Open datasets demonstrate the effectiveness and efficiency of our approach, achieving a good compromise in terms of speed and accuracy.
△ Less
Submitted 6 September, 2025;
originally announced September 2025.
-
OASIS: Object-based Analytics Storage for Intelligent SQL Query Offloading in Scientific Tabular Workloads
Authors:
Soon Hwang,
Junhyeok Park,
Junghyun Ryu,
Seonghoon Ahn,
Jeoungahn Park,
Jeongjin Lee,
Soonyeal Yang,
Jungki Noh,
Woosuk Chung,
Hoshik Kim,
Youngjae Kim
Abstract:
Computation-Enabled Object Storage (COS) systems, such as MinIO and Ceph, have recently emerged as promising storage solutions for post hoc, SQL-based analysis on large-scale datasets in High-Performance Computing (HPC) environments. By supporting object-granular layouts, COS facilitates column-oriented access and supports in-storage execution of data reduction operators, such as filters, close to…
▽ More
Computation-Enabled Object Storage (COS) systems, such as MinIO and Ceph, have recently emerged as promising storage solutions for post hoc, SQL-based analysis on large-scale datasets in High-Performance Computing (HPC) environments. By supporting object-granular layouts, COS facilitates column-oriented access and supports in-storage execution of data reduction operators, such as filters, close to where the data resides. Despite growing interest and adoption, existing COS systems exhibit several fundamental limitations that hinder their effectiveness. First, they impose rigid constraints on output data formats, limiting flexibility and interoperability. Second, they support offloading for only a narrow set of operators and expressions, restricting their applicability to more complex analytical tasks. Third--and perhaps most critically--they fail to incorporate design strategies that enable compute offloading optimized for the characteristics of deep storage hierarchies. To address these challenges, this paper proposes OASIS, a novel COS system that features: (i) flexible and interoperable output delivery through diverse formats, including columnar layouts such as Arrow; (ii) broad support for complex operators (e.g., aggregate, sort) and array-aware expressions, including element-wise predicates over array structures; and (iii) dynamic selection of optimal execution paths across internal storage layers, guided by operator characteristics and data movement costs. We implemented a prototype of OASIS and integrated it into the Spark analytics framework. Through extensive evaluation using real-world scientific queries from HPC workflows, OASIS achieves up to a 32.7% performance improvement over Spark configured with existing COS-based storage systems.
△ Less
Submitted 2 September, 2025;
originally announced September 2025.
-
Competition Between Controllable Non-Radiative and Intrinsic Radiative Second-Order Recombination in Halide Perovskites
Authors:
Dengyang Guo,
Alan R. Bowman,
Sebastian Gorgon,
Changsoon Cho,
Youngkwang Jung,
Jiashang Zhao,
Linjie Dai,
Jaewang Park,
Kyung Mun Yeom,
Satyawan Nagane,
Stuart Macpherson,
Weidong Xu,
Jun Hong Noh,
Sang Il Seok,
Tom Savenije,
Samuel D. Stranks
Abstract:
Halide perovskite solar cells have demonstrated a rapid increase in power conversion efficiencies. Understanding and mitigating remaining carrier losses in halide perovskites is now crucial to enable further increases to approach their practical efficiency limits. Whilst the most widely known non-radiative recombination from solar cells relates to carrier trapping and is first order in carrier den…
▽ More
Halide perovskite solar cells have demonstrated a rapid increase in power conversion efficiencies. Understanding and mitigating remaining carrier losses in halide perovskites is now crucial to enable further increases to approach their practical efficiency limits. Whilst the most widely known non-radiative recombination from solar cells relates to carrier trapping and is first order in carrier density, recent reports have revealed a non-radiative pathway that is second order. However, the origin and impact of this second-order process on devices remain unclear. Here, we understand this non-radiative second-order recombination (k2non) pathway by manipulating the charge carrier dynamics via controlling the bulk and surface conditions. By combining temperature-dependent spectroscopies, we demonstrate that the value of k2non depends on extrinsic factors, in contrast to intrinsic second-order recombination, which aligns with theoretical evaluations through van Roosbroeck-Shockley relations. Based on density functional theory simulations and Quasi-Fermi level calculations, we propose that shallow surface states are the primary origin of this second-order non-radiative component, contributing up to ~80 mV of the overall reduction in Voc at room temperature. This work reveals that carrier losses from two non-radiative recombination types (first and second order) are not linked, emphasizing the need for distinctive mitigation strategies targeting each type to unlock the full efficiency potential of perovskite solar cells.
△ Less
Submitted 21 August, 2025;
originally announced August 2025.
-
Zero-shot CT Super-Resolution using Diffusion-based 2D Projection Priors and Signed 3D Gaussians
Authors:
Jeonghyun Noh,
Hyun-Jic Oh,
Won-Ki Jeong
Abstract:
Computed tomography (CT) is important in clinical diagnosis, but acquiring high-resolution (HR) CT is constrained by radiation exposure risks. While deep learning-based super-resolution (SR) methods have shown promise for reconstructing HR CT from low-resolution (LR) inputs, supervised approaches require paired datasets that are often unavailable. Zero-shot methods address this limitation by opera…
▽ More
Computed tomography (CT) is important in clinical diagnosis, but acquiring high-resolution (HR) CT is constrained by radiation exposure risks. While deep learning-based super-resolution (SR) methods have shown promise for reconstructing HR CT from low-resolution (LR) inputs, supervised approaches require paired datasets that are often unavailable. Zero-shot methods address this limitation by operating on single LR inputs; however, they frequently fail to recover fine structural details due to limited LR information within individual volumes. To overcome these limitations, we propose a novel zero-shot 3D CT SR framework that integrates diffusion-based upsampled 2D projection priors into the 3D reconstruction process. Specifically, our framework consists of two stages: (1) LR CT projection SR, training a diffusion model on abundant X-ray data to upsample LR projections, thereby enhancing the scarce information inherent in the LR inputs. (2) 3D CT volume reconstruction, using 3D Gaussian splatting with our novel Negative Alpha Blending (NAB-GS), which models positive and negative Gaussian densities to learn signed residuals between diffusion-generated HR and upsampled LR projections. Our framework demonstrates superior quantitative and qualitative performance on two public datasets, and expert evaluations present the framework's clinical potential at 4x.
△ Less
Submitted 28 May, 2026; v1 submitted 20 August, 2025;
originally announced August 2025.
-
StyleMM: Stylized 3D Morphable Face Model via Text-Driven Aligned Image Translation
Authors:
Seungmi Lee,
Kwan Yun,
Junyong Noh
Abstract:
We introduce StyleMM, a novel framework that can construct a stylized 3D Morphable Model (3DMM) based on user-defined text descriptions specifying a target style. Building upon a pre-trained mesh deformation network and a texture generator for original 3DMM-based realistic human faces, our approach fine-tunes these models using stylized facial images generated via text-guided image-to-image (i2i)…
▽ More
We introduce StyleMM, a novel framework that can construct a stylized 3D Morphable Model (3DMM) based on user-defined text descriptions specifying a target style. Building upon a pre-trained mesh deformation network and a texture generator for original 3DMM-based realistic human faces, our approach fine-tunes these models using stylized facial images generated via text-guided image-to-image (i2i) translation with a diffusion model, which serve as stylization targets for the rendered mesh. To prevent undesired changes in identity, facial alignment, or expressions during i2i translation, we introduce a stylization method that explicitly preserves the facial attributes of the source image. By maintaining these critical attributes during image stylization, the proposed approach ensures consistent 3D style transfer across the 3DMM parameter space through image-based training. Once trained, StyleMM enables feed-forward generation of stylized face meshes with explicit control over shape, expression, and texture parameters, producing meshes with consistent vertex connectivity and animatability. Quantitative and qualitative evaluations demonstrate that our approach outperforms state-of-the-art methods in terms of identity-level facial diversity and stylization capability. The code and videos are available at [kwanyun.github.io/stylemm_page](kwanyun.github.io/stylemm_page).
△ Less
Submitted 15 August, 2025;
originally announced August 2025.
-
Understanding Human Daily Experience Through Continuous Sensing: ETRI Lifelog Dataset 2024
Authors:
Se Won Oh,
Hyuntae Jeong,
Seungeun Chung,
Jeong Mook Lim,
Kyoung Ju Noh,
Sunkyung Lee,
Gyuwon Jung
Abstract:
Improving human health and well-being requires an accurate and effective understanding of an individual's physical and mental state throughout daily life. To support this goal, we utilized smartphones, smartwatches, and sleep sensors to collect data passively and continuously for 24 hours a day, with minimal interference to participants' usual behavior, enabling us to gather quantitative data on d…
▽ More
Improving human health and well-being requires an accurate and effective understanding of an individual's physical and mental state throughout daily life. To support this goal, we utilized smartphones, smartwatches, and sleep sensors to collect data passively and continuously for 24 hours a day, with minimal interference to participants' usual behavior, enabling us to gather quantitative data on daily behaviors and sleep activities across multiple days. Additionally, we gathered subjective self-reports of participants' fatigue, stress, and sleep quality through surveys conducted immediately before and after sleep. This comprehensive lifelog dataset is expected to provide a foundational resource for exploring meaningful insights into human daily life and lifestyle patterns, and a portion of the data has been anonymized and made publicly available for further research. In this paper, we introduce the ETRI Lifelog Dataset 2024, detailing its structure and presenting potential applications, such as using machine learning models to predict sleep quality and stress.
△ Less
Submitted 17 July, 2025;
originally announced August 2025.
-
Kubo-Martin-Schwinger relation for energy eigenstates of SU(2)-symmetric quantum many-body systems
Authors:
Jae Dong Noh,
Aleksander Lasek,
Jade LeSchack,
Nicole Yunger Halpern
Abstract:
The fluctuation-dissipation theorem (FDT) is a fundamental result in statistical mechanics. It stipulates that, if perturbed out of equilibrium, a system responds at a rate proportional to a thermal-equilibrium property. Applications range from particle diffusion to electrical-circuit noise. To prove the FDT, one must prove that common thermal states obey a symmetry property, the Kubo-Martin-Schwi…
▽ More
The fluctuation-dissipation theorem (FDT) is a fundamental result in statistical mechanics. It stipulates that, if perturbed out of equilibrium, a system responds at a rate proportional to a thermal-equilibrium property. Applications range from particle diffusion to electrical-circuit noise. To prove the FDT, one must prove that common thermal states obey a symmetry property, the Kubo-Martin-Schwinger (KMS) relation. Energy eigenstates of certain quantum many-body systems were recently proved to obey the KMS relation. The proof relies on the eigenstate thermalization hypothesis (ETH), which explains how such systems thermalize internally. This KMS relation contains a finite-size correction that scales as the inverse system size. Non-Abelian symmetries conflict with the ETH, so a non-Abelian ETH was proposed recently. Using it, we derive a KMS relation for SU(2)-symmetric quantum many-body systems' energy eigenstates. The finite-size correction scales as usual under certain circumstances but can be polynomially larger in others, we argue. We support the ordinary-scaling result numerically, simulating a Heisenberg chain of 16-24 qubits. The numerics, limited by computational capacity, indirectly support the larger correction. This work helps extend into nonequilibrium physics the effort, recently of interest across quantum physics, to identify how non-Abelian symmetries may alter conventional thermodynamics.
△ Less
Submitted 24 April, 2026; v1 submitted 9 July, 2025;
originally announced July 2025.