-
Continual Reasoning Gym: Diagnosing and Harnessing Shared Reasoning in Continual RLVR
Authors:
Lirui Luo,
Guoxi Zhang,
Hongming Xu,
Rongqing Li,
Cong Fang,
Lifeng Fan
Abstract:
Reinforcement learning with verifiable rewards (RLVR) commonly post-trains reasoning models on multiple tasks, while rerunning multitask RLVR (MTRL) as new tasks are added makes capability expansion costly. We therefore study continual RLVR, which updates the existing model as each task arrives. The central question is whether a model updated this way can perform as well as a jointly trained model…
▽ More
Reinforcement learning with verifiable rewards (RLVR) commonly post-trains reasoning models on multiple tasks, while rerunning multitask RLVR (MTRL) as new tasks are added makes capability expansion costly. We therefore study continual RLVR, which updates the existing model as each task arrives. The central question is whether a model updated this way can perform as well as a jointly trained model. To answer this question, we introduce Continual Reasoning Gym, a continual-RLVR environment that organizes text and visual reasoning tasks into five task sequences. In this setting, we identify two key observations: Sequential RLVR exhibits modest forgetting, yet its final performance remains below that of MTRL. To understand the latter, we decompose final performance and show that forgetting accounts for only part of the gap. To explain the former, we identify shared reasoning: transferable reasoning structure allows training on one task to support others on average. We therefore introduce Continual Prompt Replay (CPR), which harnesses shared reasoning to improve learning on the arriving and future tasks by replaying previous-task prompts and regenerating their responses with the current policy. On average, only CPR reaches MTRL-level performance.
△ Less
Submitted 19 August, 2026; v1 submitted 19 August, 2026;
originally announced August 2026.
-
Astrophysical Graviton Squeezing Can Be Hidden in the Far-Field
Authors:
Cheng-Jun Fang,
Zong-Kuan Guo,
Zhen-Hong Lyu,
Jing Shu,
Yu-Heng Sun,
Zi-Zheng Zhou
Abstract:
While localized astrophysical sources can generate macroscopic graviton squeezing, their observable quantum signatures at far-field detectors remain unresolved. In this work, we investigate the propagation dynamics of the squeezed states using spatial quantum optics methods to evaluate correlation functions accessible to a local observer. Crucially, we reveal a severe kinematic conflict in same-co…
▽ More
While localized astrophysical sources can generate macroscopic graviton squeezing, their observable quantum signatures at far-field detectors remain unresolved. In this work, we investigate the propagation dynamics of the squeezed states using spatial quantum optics methods to evaluate correlation functions accessible to a local observer. Crucially, we reveal a severe kinematic conflict in same-cone measurements, which highly suppresses local quantum coherence. Consequently, these macroscopically squeezed states appear classically thermal to a single detector. Our results demonstrate that global squeezing does not guarantee local observability, and the measurable quantum signatures may be significantly weaker than what would be expected from the overall squeezing parameter of the state.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
EMAN: Optimization-Driven Capacity Growth through Path Emergence in Multi-Task Learning
Authors:
Chenlei Fang,
Jingchen Li,
Hongzong LI,
Qingyao Li,
Yixuan Zhang,
Huarui Wu,
Haobin Shi,
Chunjiang Zhao
Abstract:
Existing multi-task learning methods rely on hard sharing, multiple paths or experts, adaptive sharing, and dynamic expansion. However, their capacity changes are usually constrained by predefined structures or triggered by task boundaries and conflict signals. This raises a fundamental question: can a network start from exact single-path computation and grow a new independent path only when persi…
▽ More
Existing multi-task learning methods rely on hard sharing, multiple paths or experts, adaptive sharing, and dynamic expansion. However, their capacity changes are usually constrained by predefined structures or triggered by task boundaries and conflict signals. This raises a fundamental question: can a network start from exact single-path computation and grow a new independent path only when persistent optimization evidence appears? We propose the Emergent Modular Atomic Network (EMAN), an optimization-driven framework for exposing an antisymmetric growth direction through latent relative phases without instantiating a second path, and for monitoring multiple decision signals during training to transform local optimization evidence into a structural decision. EMAN materializes two equal-capacity independent paths only after certification. EMAN adaptively allocates shared and task-specific representation capacity to accommodate varying task requirements. Extensive experiments on controlled rank settings, PASCAL-Context, and NYUv2 validate its effectiveness, achieving improved performance at a competitive computational cost.
△ Less
Submitted 5 August, 2026;
originally announced August 2026.
-
Codegree Thresholds for $λ$-Choosability of Graphs
Authors:
Chunqiu Fang,
Rongxing Xu
Abstract:
Let $λ=\{k_1,\ldots,k_q\}$ be a partition, and let $|λ|=k_1+\cdots+k_q$. A $|λ|$-list assignment $L$ of a graph $G$ is a $λ$-assignment if its color set can be partitioned into $q$ disjoint sets $X_1,\ldots,X_q$ such that $|L(v)\cap X_i|=k_i$ for every vertex $v$ and every $i\in[q]$. This notion, introduced by Zhu [J. Combin. Theory Ser. B, 2020], puts ordinary coloring and list coloring in the sa…
▽ More
Let $λ=\{k_1,\ldots,k_q\}$ be a partition, and let $|λ|=k_1+\cdots+k_q$. A $|λ|$-list assignment $L$ of a graph $G$ is a $λ$-assignment if its color set can be partitioned into $q$ disjoint sets $X_1,\ldots,X_q$ such that $|L(v)\cap X_i|=k_i$ for every vertex $v$ and every $i\in[q]$. This notion, introduced by Zhu [J. Combin. Theory Ser. B, 2020], puts ordinary coloring and list coloring in the same framework. A theorem of Alon [Random Structures Algorithms, 2000] states that every graph with minimum degree $d$ has choice number at least $(1/2-o(1))\log_2d$. Saxton and Thomason [Invent. Math., 2015] later used the hypergraph container method to replace $1/2$ by the sharp constant $1$. It is natural to ask whether a similar phenomenon holds for every fixed partition $λ$. Minimum degree alone is not sufficient: balanced complete bipartite graphs have arbitrarily large minimum degree but are always $\{1,1\}$-choosable. We show that the appropriate replacement is the minimum $q$-codegree, defined for $|V(G)|\geq q$ by $δ_q(G)=\min\{|N_G(S)|:S\subseteq V(G),\,|S|=q\}$.
More precisely, for every partition $λ$ there exists an integer $d$ such that every graph $G$ with $δ_q(G)\geq d$ is not $λ$-choosable. Let $f(λ)$ be the least such $d$. For every fixed $q$, we prove $f(λ)\leq2^{(2q+o(1))|λ|}$ as $|λ|\to\infty$, while $f(λ)\geq(q+1)^{-1}(1+1/q)^{|λ|}$ for every $λ$. For the partition $\{k,\ldots,k\}$ with $q$ equal parts, we determine the threshold asymptotically: $f(\{k,\ldots,k\})=ρ_q^{-(1+o(1))k}$ as $k\to\infty$, where $ρ_q$ is the unique $x\in(0,1)$ satisfying $x=(1-x)^q$. When $q=1$, our result implies $\operatorname{ch}(G)\geq(1-o(1))\log_2δ(G)$.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
The Erdos-Mullin Five-Edge Intersection Problem
Authors:
Chengrui Fang,
Jianfeng Hou
Abstract:
For an $n$-vertex graph $G$ and a permutation $π$ of its vertex set, let $I_G(π)=|E(G)\cap E(πG)|$, and let $μ(G)=\min_π I_G(π)$. Let $f(n,k)$ be the minimum number of edges in an $n$-vertex graph $G$ satisfying $μ(G)\ge k$. Erdős recorded a construction of Mullin showing $f(n,5)\le 2n-2$ and asked whether equality holds for sufficiently large $n$. We prove that it does: $f(n,5)=2n-2$ for all suff…
▽ More
For an $n$-vertex graph $G$ and a permutation $π$ of its vertex set, let $I_G(π)=|E(G)\cap E(πG)|$, and let $μ(G)=\min_π I_G(π)$. Let $f(n,k)$ be the minimum number of edges in an $n$-vertex graph $G$ satisfying $μ(G)\ge k$. Erdős recorded a construction of Mullin showing $f(n,5)\le 2n-2$ and asked whether equality holds for sufficiently large $n$. We prove that it does: $f(n,5)=2n-2$ for all sufficiently large $n$. Equivalently, every sufficiently large $n$-vertex graph with at most $2n-3$ edges admits a relabelling with at most four common edges. The proof combines a quantitative exclusion of almost-universal vertices, a finite high-degree core with low-degree buffer vertices, list packing, and a sparse permutation version of the Lovász local lemma.
△ Less
Submitted 13 August, 2026;
originally announced August 2026.
-
Software Engineering for and with GUI Agent
Authors:
Shengcheng Yu,
Yuchen Ling,
Junyang Xing,
Quan Zhou,
Chunrong Fang,
Zhenyu Chen
Abstract:
GUI agents have advanced rapidly, producing a growing body of frameworks, benchmarks, and applications. However, this growth has outpaced the maturity of the field. GUI agents remain technically brittle, incompletely engineered, and insufficiently validated for sustained real-world use. They are evolving into closed-loop software systems. Within these systems, model reasoning is coupled with inter…
▽ More
GUI agents have advanced rapidly, producing a growing body of frameworks, benchmarks, and applications. However, this growth has outpaced the maturity of the field. GUI agents remain technically brittle, incompletely engineered, and insufficiently validated for sustained real-world use. They are evolving into closed-loop software systems. Within these systems, model reasoning is coupled with interface perception, execution feedback, recovery, and human oversight. This evolution calls for a software engineering perspective that remains largely absent from existing research. We address this gap by reviewing 336 GUI-agent papers from January 2018 to April 2026. Five research questions examine the research landscape, architectures, evaluation, software lifecycle concerns, and future opportunities. Our findings show that the field has expanded sharply since 2024, while mobile and web settings remain dominant. Architectures increasingly adopt modular perceive-reason-act loops, but recovery, human escalation, safety enforcement, and auditability remain underdeveloped. This architectural imbalance extends to evaluation. Evaluations are becoming more interactive, but they remain centered on task success and are difficult to compare across protocols. More broadly, existing studies provide limited support for testing beyond benchmarks and for maintaining agents after release. Observability, privacy engineering, and systematic human oversight are also underdeveloped. Together, these findings show that capability improvements alone cannot ensure deployment readiness. Future research should connect dependable execution with lifecycle-centered testing and reproducible evaluation. It should also integrate permission and privacy controls with cost-aware, human-centered governance. This integration is necessary to build dependable, maintainable, secure, and deployable GUI-agent systems.
△ Less
Submitted 10 August, 2026;
originally announced August 2026.
-
IRPol-Fuse: Energy-structure coordination for infrared polarization fusion under low visibility
Authors:
Zhuangfan Huang,
Chusheng Fang,
Xiaosong Li,
Yang Liua,
Xiaoqi Cheng,
Haishu Tan
Abstract:
Robust perception under low-visibility conditions requires fused imagery that jointly preserves infrared thermal saliency and polarization-derived structural details. However, existing infrared-polarization image fusion (IPIF) methods often overemphasize dominant infrared responses, causing weak yet informative polarization textures in dark regions to be suppressed. To address this issue, we propo…
▽ More
Robust perception under low-visibility conditions requires fused imagery that jointly preserves infrared thermal saliency and polarization-derived structural details. However, existing infrared-polarization image fusion (IPIF) methods often overemphasize dominant infrared responses, causing weak yet informative polarization textures in dark regions to be suppressed. To address this issue, we propose IRPol-Fuse, an energy-structure coordinated IPIF framework for challenging low-visibility scenarios. The proposed framework contains three key modules: Polarization Attention Fusion for adaptive infrared-polarization allocation, Infrared Highlight Injector for highlight-guided infrared preservation, and Polarization Texture Injector for polarization texture restoration and fine-detail recovery. We further construct LI-PI, a dedicated infrared-polarization evaluation dataset for low-visibility and visually concealed scenes. Experiments on LI-PI and the public LDDRS dataset demonstrate that IRPol-Fuse achieves favorable performance in thermal target preservation, structural detail recovery, and visual naturalness. Region-aware evaluation and downstream object detection further verify that the proposed energy-structure coordination strategy effectively preserves both infrared target saliency and polarization-derived structural information. Code is available at https://github.com/1hzf/IRPolar-Fuse .
△ Less
Submitted 7 August, 2026;
originally announced August 2026.
-
Autonomous Optimization of Complex Oxides for Thermochemical Fuel Production
Authors:
Shuiping Gong,
Mingcheng Li,
Han Hao,
Zhenhao Zhou,
Yi Li,
Xiaobo Liao,
Cheng Fang,
Jian Deng,
Jiangang He,
Wenpei Gao,
Yakun Yuan,
Chris Wolverton,
Tao Deng,
Chaochao Dun,
Runxia Cai,
Zhenpeng Yao
Abstract:
Two-step thermochemical fuel production, including H2O and CO2 splitting, offers a promising route to sustainable fuel manufacturing, with performance governed by redox-active oxides that enable cyclic reduction-oxidation reactions. Maximizing thermal-to-fuel conversion efficiency demands materials that simultaneously satisfy multiple stringent thermodynamic and kinetic targets. Addressing these r…
▽ More
Two-step thermochemical fuel production, including H2O and CO2 splitting, offers a promising route to sustainable fuel manufacturing, with performance governed by redox-active oxides that enable cyclic reduction-oxidation reactions. Maximizing thermal-to-fuel conversion efficiency demands materials that simultaneously satisfy multiple stringent thermodynamic and kinetic targets. Addressing these requirements has increasingly driven materials design toward complex, multi-cation oxides, such as mixed-cation fluorites, perovskites, and high-entropy oxides, wherein composition, defect chemistry, phase stability, and morphology should be co-optimized. This creates a challenging materials optimization problem that is poorly suited to traditional trial-and-error approaches. In this review, we argue that thermochemical fuel production provides a compelling frontier for autonomous materials design and optimization. We first examine why redox-active complex oxides are difficult to develop, owing to multidimensional phase spaces, harsh operating conditions, and competing functional targets. We then discuss how high-throughput computation, automated synthesis, characterization and testing, and machine learning can be integrated into closed-loop workflows to address these challenges. Building on broader oxide materials research, we organize recent progress into a capability roadmap for complex-oxide optimization, spanning compositionally diverse synthesis, operando characterization, robotic testing, operation-condition computation, and multi-objective optimization. Finally, we outline key experimental, computational, and data challenges for building self-improving materials development platforms for materials development in thermochemical fuel production.
△ Less
Submitted 7 August, 2026;
originally announced August 2026.
-
Unmasking Removal-Budget Confounding: A Matched Operating-Point Evaluation Framework for Adaptive Data Cleaning
Authors:
Wei-Hsiang Chen,
Pin-Hsuan Yu,
Chen-Hsuan Fang,
Jung-Hua Wang
Abstract:
Adaptive data-cleaning methods replace manual filtering thresholds with data-driven partitions. However, changing the partition granularity, the number of groups used to segment samples by estimated corruption risk, can implicitly shift the decision boundary and alter the overall number of removed samples. This creates a bias known as removal-budget confounding, where apparent gains in metrics lik…
▽ More
Adaptive data-cleaning methods replace manual filtering thresholds with data-driven partitions. However, changing the partition granularity, the number of groups used to segment samples by estimated corruption risk, can implicitly shift the decision boundary and alter the overall number of removed samples. This creates a bias known as removal-budget confounding, where apparent gains in metrics like precision or false-positive rate reflect a smaller removal budget rather than superior corruption discrimination. To address this evaluation bias, we introduce an operating-point-aware evaluation framework that evaluates methods using matched-budget and matched-recall controls alongside threshold-independent metrics (AUROC and AUPRC). We test this framework on a multi-cue adaptive cleaner redesign featuring a reweighted learning-difficulty cue, an auxiliary Euclidean-distance cue, and increased partition granularity intended to isolate clean-but-difficult samples. While naive evaluations (assessing configurations at their own induced operating points) suggest substantial performance improvements for the redesign, these gains disappear once operating points are equalized. False-positive decomposition reveals that clean-but-difficult samples primarily drive error counts at low corruption rates, become threshold-dependent at moderate corruption, and contribute negligibly under severe corruption. Experiments on CIFAR-10 and ImageNet-100 demonstrate that most performance differences observed in naive evaluation shrink or vanish at low-to-moderate corruption when operating points are matched. True ranking advantages only remain in specific low-prevalence settings and in high-recall regions under severe corruption. These findings highlight that adaptive cleaning methods must be benchmarked at matched operating points to ensure performance gains reflect genuine corruption discrimination.
△ Less
Submitted 6 August, 2026;
originally announced August 2026.
-
RxnCLF: Contrastive Transformation-Aware Reaction Foundation Model for Improved Reactivity Prediction
Authors:
Yiting Zheng,
Cheng Fang,
Anthony Donofrio,
Haote Li
Abstract:
Reaction yield prediction remains challenging because labeled data are scarce and reaction space is both combinatorially large and sparsely populated, limiting the generalization of existing reaction representations. String-, fingerprint-, and graph-based reaction encodings only partially capture chemical transformations, making accurate prediction difficult for reactions with complex substrates.…
▽ More
Reaction yield prediction remains challenging because labeled data are scarce and reaction space is both combinatorially large and sparsely populated, limiting the generalization of existing reaction representations. String-, fingerprint-, and graph-based reaction encodings only partially capture chemical transformations, making accurate prediction difficult for reactions with complex substrates. We propose reaction contrastive learning foundation (RxnCLF), a self-supervised contrastive framework for reaction representation learning. RxnCLF is built on a condensed reaction graph (CRG) that unifies reactant and product information into a single graph, enabling the model to learn explicit and enriched transformation structure rather than disconnected graphs. Pretrained on 1.7 million Pistachio reactions, RxnCLF learns a compact and continuous latent space that captures both reaction-center features and broader side chain contexts, making it transformation-aware and chemically interpretable. Fine-tuned on multiple yield prediction benchmarks, including Buchwald-Hartwig, Pd-catalyzed BH coupling, and proprietary HTE C-N coupling and amide formation datasets, RxnCLF consistently outperforms graph and sequence-based baselines, improving R2 and achieving the best performance overall. Our results highlight the promise of CRG-based RxnCLF as a scalable reaction foundation model, with the potential to generalize across broader reaction spaces and support diverse downstream reaction informatics tasks, including regioselectivity prediction, enantioselectivity prediction, and reaction condition optimization.
△ Less
Submitted 6 August, 2026;
originally announced August 2026.
-
Breaking Customized LLMs for Coding: Automated Red Teaming for Instruction Backdoor Attacks
Authors:
Yuchen Chen,
Wei Cheng,
Yuan Xiao,
Wising Sun,
Chunrong Fang,
Yang Liu,
Zhenyu Chen,
Baowen Xu
Abstract:
LLM customization platforms allow users to build task-specific models for code intelligence tasks by embedding instructions into system prompts, without modifying the underlying model parameters. While these platforms lower the barrier to developing customized LLMs, they also introduce a new attack surface: instruction backdoor attacks, in which adversaries implant hidden malicious behaviors into…
▽ More
LLM customization platforms allow users to build task-specific models for code intelligence tasks by embedding instructions into system prompts, without modifying the underlying model parameters. While these platforms lower the barrier to developing customized LLMs, they also introduce a new attack surface: instruction backdoor attacks, in which adversaries implant hidden malicious behaviors into customized instructions. However, existing attacks suffer from two key limitations. First, they often rely on explicit trigger patterns readily detected by platform-side or user-side inspection. Second, they require substantial manual effort to craft task-specific backdoored instructions, limiting their scalability.
In this paper, we propose ARIA, an automated red-teaming framework for crafting covert and effective backdoored instructions against customized LLMs. ARIA leverages an attacker LLM to iteratively generate and refine backdoored instructions, guided by structured feedback from the target LLM along three dimensions: stealthiness, clean-task utility, and backdoor effectiveness. We evaluate ARIA on three code intelligence tasks, using four representative LLMs, and compare it with three baseline attacks. Experimental results show that ARIA achieves the highest attack success rate of 0.945, while maintaining the best clean-task utility across all tasks. ARIA also generalizes well across programming languages and remains robust to generation temperature. Furthermore, ARIA significantly outperforms existing attacks in evading platform-side and user-side detection, achieving a false negative rate of up to 1.000, and stays effective against existing defense methods, demonstrating its strong generalizability and robustness.
△ Less
Submitted 6 August, 2026;
originally announced August 2026.
-
A Bayesian approach to the long-baseline neutrino oscillation sensitivity of DUNE
Authors:
DUNE Collaboration,
S. Abbaslu,
F. Abd Alrahman,
A. Abed Abud,
R. Acciarri,
M. A. Acero,
M. R. Adames,
G. Adamov,
M. Adamowski,
K. Adhikari,
C. Adriano,
K. Agudelo-Jaramillo,
F. Akbar,
F. Alemanno,
N. S. Alex,
L. Aliaga Soplin,
A. Alqaisi,
O. Alterkait,
A. Alton,
R. Alvarez,
T. Alves,
A. Aman,
H. Amar,
R. M. Amarinei,
P. Amedo
, et al. (1262 additional authors not shown)
Abstract:
The sensitivity of the Deep Underground Neutrino Experiment (DUNE) to neutrino oscillation is evaluated using a Bayesian Markov Chain Monte Carlo (MCMC) approach. This analysis uses the same underlying sensitivity inputs as previous DUNE studies [Eur. Phys. J. C 80, 978 (2020)], and therefore does not present updated DUNE sensitivities, but instead explores the additional inferences accessible usi…
▽ More
The sensitivity of the Deep Underground Neutrino Experiment (DUNE) to neutrino oscillation is evaluated using a Bayesian Markov Chain Monte Carlo (MCMC) approach. This analysis uses the same underlying sensitivity inputs as previous DUNE studies [Eur. Phys. J. C 80, 978 (2020)], and therefore does not present updated DUNE sensitivities, but instead explores the additional inferences accessible using a Bayesian approach. We present four-dimensional posterior probability distributions of the oscillation parameters, highlighting the breadth of correlation in the parameter space of interest, especially between $\sin^2 θ_{23}$ and $\sin^2 θ_{13}$. We exploit the flexibility of the Bayesian framework to incorporate parameter constraints post hoc and assess the impact of applying a reactor short-baseline $θ_{13}$ constraint. A significant increase in the sensitivity to the $θ_{23}$ octant is found when including the constraint. Posterior distributions of derived quantities can be easily constructed from MCMC results. This work presents the first study of DUNE's sensitivity to the Jarlskog invariant, $J$, a quantity that provides a parametrisation-independent measure of charge-parity violation in the leptonic sector.
△ Less
Submitted 4 August, 2026;
originally announced August 2026.
-
Endpoint boundedness of Orlicz-BMO commutators on Orlicz-Hardy type spaces
Authors:
Zixing Zhuang,
Chenglong Fang
Abstract:
Given a growth function $\varphi:[0,\infty)\rightarrow [0,\infty)$, it is established that the commutators generated by sublinear operators and Orlicz-$\mathrm{BMO}$ function $b$ are bounded from $H_{b}^{\varphi}(\mathbb{R}^{n})$ to $L^{1}(\mathbb{R}^{n})$, and from $H^{\varphi}(\mathbb{R}^{n})$ to $L^{1,\,\infty}(\rn)$, where $H_{b}^{\varphi}(\mathbb{R}^{n})$ is a specific subspace of Orlicz-Hard…
▽ More
Given a growth function $\varphi:[0,\infty)\rightarrow [0,\infty)$, it is established that the commutators generated by sublinear operators and Orlicz-$\mathrm{BMO}$ function $b$ are bounded from $H_{b}^{\varphi}(\mathbb{R}^{n})$ to $L^{1}(\mathbb{R}^{n})$, and from $H^{\varphi}(\mathbb{R}^{n})$ to $L^{1,\,\infty}(\rn)$, where $H_{b}^{\varphi}(\mathbb{R}^{n})$ is a specific subspace of Orlicz-Hardy space $H^{\varphi}(\mathbb{R}^{n})$ and sublinear operators include Lusin area integral, g-function, Marcinkiewicz integral and Bochner-Riesz mean operator. Under the assumptions $T^*1=0$ and $T^*b=0$, it is shown that the Orlicz-$\mathrm{BMO}$ commutator associated with the Bochner-Riesz mean operator admits endpoint boundedness from $H_{b}^{\varphi}(\mathbb{R}^{n})$ to $H^{1}(\mathbb{R}^{n})$. However, the commutators corresponding to other operators discussed in this paper do not possess the aforementioned endpoint boundedness, and a counterexample is provided to illustrate this point.
△ Less
Submitted 3 August, 2026;
originally announced August 2026.
-
ARES: Adaptive Reasoning-Effort Steering for PPA- and Cost-Aware RTL Optimization with LLM Agents
Authors:
Stef Cuyckens,
Mihaela Jivanescu,
Jun Yin,
Chao Fang,
Marian Verhelst
Abstract:
Large language model (LLM) agents optimize the power, performance, and area (PPA) of register-transfer-level (RTL) designs by iterating over edits, synthesis, and PPA analysis, paying a dollar cost for every LLM call. Prior agents report the quality reached without its normalized cost, attribute that quality to an engineered cross-design memory, and hold the reasoning effort of every call fixed. W…
▽ More
Large language model (LLM) agents optimize the power, performance, and area (PPA) of register-transfer-level (RTL) designs by iterating over edits, synthesis, and PPA analysis, paying a dollar cost for every LLM call. Prior agents report the quality reached without its normalized cost, attribute that quality to an engineered cross-design memory, and hold the reasoning effort of every call fixed. We propose Ares with three corresponding innovations. (1) We introduce a normalized dollar cost per LLM call reported alongside the figure of merit (FoM), enabling fair comparison across effort levels and optimizers. (2) Using this accounting, we find the construction of the long-term memory matters little. An engineered memory brings no dependable gain over a plain concatenation of the same experience. (3) We instead adapt the per-call reasoning effort by escalating to deeper reasoning only once progress at a lower effort stalls, via a patience counter fit on 21 training designs, allocating reasoning where it pays rather than uniformly across all iterations. On three test designs unseen during training, the effort policy lowers the FoM by 23-27% where the best fixed effort reaches 16-23%, at equal normalized cost. Ares closes up to 83% of the gap from an LLM-drafted multiply-accumulate unit to its highly hand-optimized counterpart, and reaches a 25% deeper FoM than state-of-the-art Dr. RTL at 12% of its tokens.
△ Less
Submitted 30 July, 2026;
originally announced July 2026.
-
MultiFixer: A Coordinator-Proposer Based Multi-Agent Framework For Fixing Multi-Hunk Bugs
Authors:
Haichuan Hu,
Chunrong Fang,
Ye Shang,
Jiawei Liu,
Weifeng Sun,
Guoqing Xie,
Chenxing Zhong,
Quanjun Zhang
Abstract:
Automated Program Repair (APR) has benefited greatly from Large Language Models (LLMs), but existing LLM-based APR methods still struggle with multi-hunk bugs that require coordinated changes across multiple locations. These bugs demand repository-level context understanding, repair-order scheduling, and effective hunk-level patch generation and selection. To address these challenges, we propose M…
▽ More
Automated Program Repair (APR) has benefited greatly from Large Language Models (LLMs), but existing LLM-based APR methods still struggle with multi-hunk bugs that require coordinated changes across multiple locations. These bugs demand repository-level context understanding, repair-order scheduling, and effective hunk-level patch generation and selection. To address these challenges, we propose MultiFixer, a novel Coordinator-Proposer based multi-agent framework for multi-hunk repair. MultiFixer performs tool-augmented bug analysis, constructs fine-grained repair context, iteratively generates patches through a Coordinator-Proposer architecture, and applies two-stage patch refinement for syntactic and semantic correctness. We evaluate MultiFixer on 835 bugs from Defects4J and three vulnerability benchmarks. On Defects4J, MultiFixer fixes 326 bugs, including 62 multi-method and 27 multi-file bugs, and outperforms prior APR baselines in the reported comparisons with the same base model. Moreover, MultiFixer also fixes 46 multi-hunk bugs among 95 unique fixes. When combined with Claude-3.5-Sonnet, MultiFixer repairs 420 bugs, establishing a new state of the art on Defects4J. On VUL4J, MultiFixer repairs 24 real-world vulnerabilities, including 5 multi-hunk cases. On the multi-hunk subsets of SEC-bench and PatchEval, MultiFixer fixes 11 and 19 vulnerabilities, respectively, outperforming all compared baselines under GPT-3.5. These results demonstrate the effectiveness of MultiFixer for multi-hunk repair.
△ Less
Submitted 29 July, 2026;
originally announced July 2026.
-
Execution-Grounded Security Testing for Coding Agents in Software Engineering Pipelines
Authors:
Yifei Ge,
Weisong Sun,
Jinkun Xiao,
Yuchen Chen,
Yebo Feng,
Peizhuo Lv,
Xia Feng,
Chunrong Fang,
Zhihong Zhao,
Zhenyu Chen,
Yang Liu
Abstract:
Coding agents are increasingly integrated into system operations, where their tool use can directly modify project artifacts, execution environments, and the underlying system. For example, if a coding agent inserts a hook into a system startup or configuration script, that change can persist after the interaction, be triggered later, and abuse delegated user or system privileges to modify the sys…
▽ More
Coding agents are increasingly integrated into system operations, where their tool use can directly modify project artifacts, execution environments, and the underlying system. For example, if a coding agent inserts a hook into a system startup or configuration script, that change can persist after the interaction, be triggered later, and abuse delegated user or system privileges to modify the system. This makes security testing a system problem: the key question is not only what the agent says, but what it actually does to the surrounding environment. We present an execution-grounded red-team testing framework for probing this execution-layer security boundary using observable sandbox evidence, including tool invocations, runtime traces, and file-system diffs. Our framework embeds target unsafe operations into routine software engineering workloads, including unit testing, regression testing, crash reproduction, and validation, and uses an execution oracle to guide refinement when an initial probe is rejected or fails. Across multiple agent frameworks and model backbones, our red-team workload reformulation substantially increases verified unsafe execution, reaching 73.61% on code carriers and 53.93% on text carriers. These results show that coding agents in system operations remain insecure under task disguise: once risky intent is hidden inside plausible engineering tasks, the agent can be induced to carry out unsafe actions on the surrounding system. More broadly, coding agents in system operations still demand stronger security testing and safeguards.
△ Less
Submitted 1 June, 2026;
originally announced July 2026.
-
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding
Authors:
Chao Fang,
Jun Yin,
Man Shi,
Marian Verhelst
Abstract:
With the rapid adoption of long-context large language models (LLMs), the continuously growing KV cache during decoding has become the critical memory bottleneck. To tackle this challenge, we propose HiKV, a novel algorithm-hardware co-design that exploits KV cache redundancy through hierarchical importance awareness. Algorithmically, HiKV compresses the KV cache at two granularities: Stage I evic…
▽ More
With the rapid adoption of long-context large language models (LLMs), the continuously growing KV cache during decoding has become the critical memory bottleneck. To tackle this challenge, we propose HiKV, a novel algorithm-hardware co-design that exploits KV cache redundancy through hierarchical importance awareness. Algorithmically, HiKV compresses the KV cache at two granularities: Stage I evicts unimportant tokens within a fixed budget, and Stage II further loads only the significant elements of each retained token, reaching compression ratios unattainable at a single granularity. Architecturally, we develop a dedicated accelerator centered on a reconfigurable importance sorter that switches between the distinct sorting datapaths each stage requires, unifying the two-stage acceleration in one circuit with minimal overhead. Evaluated on representative LLMs, HiKV achieves up to 7.95x speedup and 90% energy reduction in the attention computation over the vanilla KV cache baseline within negligible 1% accuracy loss. Under iso-accuracy constraints, HiKV outperforms state-of-the-art importance-based methods by achieving an additional 1.82~4.87x reduction in external memory accesses. These benefits are enabled by specialized hardware components that add only 8% to the system area.
△ Less
Submitted 24 July, 2026;
originally announced July 2026.
-
Insecure Coding Preferences in Long-Term Memory: Security Risks for LLM-based Code Generation
Authors:
Yuchen Chen,
Wei Cheng,
Yuan Xiao,
Zhou Yang,
Weifeng Sun,
Chunrong Fang,
Xiang Chen,
Baowen Xu,
David Lo,
Zhenyu Chen
Abstract:
LLM-based systems increasingly incorporate long-term memory to improve cross-session continuity. However, once insecure coding preferences are stored, they may silently influence security-critical decisions in subsequent generations. In this study, we conduct the first systematic empirical study on the impact of insecure coding preferences stored in long-term memory on the security of LLM-based co…
▽ More
LLM-based systems increasingly incorporate long-term memory to improve cross-session continuity. However, once insecure coding preferences are stored, they may silently influence security-critical decisions in subsequent generations. In this study, we conduct the first systematic empirical study on the impact of insecure coding preferences stored in long-term memory on the security of LLM-based code generation. We evaluate four LLMs (ChatGPT, Gemini, Qwen, and Grok) across five programming languages (Python, C, C++, Go, and JavaScript). Our results show that insecure memories significantly increase the risk of generating vulnerable code by 2.7-50.3 percentage points (pp). Moreover, they create a 5.4-14.0 percentage-point risk-warning gap, where warning-rate increases lag behind vulnerability-rate increases. Further analysis reveals that insecure memories are difficult to overwrite through normal interactions and can broadly influence model outputs even when prompts are phrased differently. Finally, we evaluate three mitigation strategies: security-requirement appending and memory storage reduce vulnerability rates by 19.7-33.6 pp but may degrade functional correctness by up to 15.9 pp; memory-level safety filtering achieves a 100\% detection rate on our evaluated risky memory entries and restores generation behavior to the without-memory baseline. Based on these findings, we provide actionable suggestions to improve the security of long-term memory in LLM-based code generation.
△ Less
Submitted 20 July, 2026;
originally announced July 2026.
-
Operation and performance of ProtoDUNE Dual Phase liquid argon time projection chamber
Authors:
DUNE Collaboration,
S. Abbaslu,
F. Abd Alrahman,
A. Abed Abud,
R. Acciarri,
L. P. Accorsi,
M. A. Acero,
M. R. Adames,
G. Adamov,
M. Adamowski,
K. Adhikari,
C. Adriano,
K. Agudelo-Jaramillo,
F. Akbar,
F. Alemanno,
N. S. Alex,
L. Aliaga Soplin,
A. Alqaisi,
M. Alrashed,
A. Alton,
R. Alvarez,
T. Alves,
A. Aman,
H. Amar,
R. Amarinei
, et al. (1341 additional authors not shown)
Abstract:
ProtoDUNE-DP was the largest ever built Liquid Argon Time Projection Chamber (LArTPC) operating in Dual-Phase (DP) mode, with a liquid target and charge read-out placed in the gas. It had an active volume of $6\times6\times6$\,m$^3$ corresponding to an active mass of 300\,t (total LAr mass of 720\,t), constructed at the CERN Neutrino Platform and took data from 2019 to 2020 with cosmic muons. In P…
▽ More
ProtoDUNE-DP was the largest ever built Liquid Argon Time Projection Chamber (LArTPC) operating in Dual-Phase (DP) mode, with a liquid target and charge read-out placed in the gas. It had an active volume of $6\times6\times6$\,m$^3$ corresponding to an active mass of 300\,t (total LAr mass of 720\,t), constructed at the CERN Neutrino Platform and took data from 2019 to 2020 with cosmic muons. In ProtoDUNE-DP the electric drift field is oriented in the vertical direction, causing the electrons to drift vertically towards the anode at the top. The ionization charge is then extracted into the gaseous argon above the liquid surface, amplified by Townsend avalanches, and collected by the charge readout planes. The detector experienced significant technical problems affecting the long-term operation of the Charge Readout Planes, formed by the Large Electron Multipliers, but other critical segments demonstrated required performance including the delivery of -300 kV to the TPC cathode, verification of replaceable charge read-out electronics, and operation of the photon detection system. ProtoDUNE-DP experience resulted in improved designs of the Vertical Drift LArTPC.
△ Less
Submitted 21 July, 2026; v1 submitted 17 July, 2026;
originally announced July 2026.
-
Explainable AI for Solar Flare Prediction: Quantitative Magnetic Field Analysis of Model-Focused Regions
Authors:
Z. Zheng,
Q. Hao,
C. Li,
P. F. Chen,
J. R. Hu,
M. D. Ding,
C. Fang
Abstract:
Solar flares are intense energy release events in the solar atmosphere that may pose significant space weather hazards, which makes developing reliable prediction models essential. Although deep learning methods, particularly convolutional neural networks (CNNs), demonstrate strong predictive performance when using solar magnetograms, their scientific credibility is undermined by a lack of physica…
▽ More
Solar flares are intense energy release events in the solar atmosphere that may pose significant space weather hazards, which makes developing reliable prediction models essential. Although deep learning methods, particularly convolutional neural networks (CNNs), demonstrate strong predictive performance when using solar magnetograms, their scientific credibility is undermined by a lack of physical interpretability. Explainable artificial intelligence (XAI) offers a potential solution. However, current XAI studies in solar flare prediction are largely qualitative and lack systematic, theory-based, quantitative validation. We present a quantitative XAI framework that can decipher the physical basis of CNN-based solar flare prediction models. Using gradient-weighted class activation mapping (Grad-CAM), we identify model-focused regions (MFRs) in solar magnetograms. Then, we perform two key analyses to evaluate the predictive capability of magnetic parameters derived from MFRs and to quantitatively characterize their magnetic complexity. Our results reveal a strong physical correlation between MFRs and flare occurrence. Specifically, magnetic features extracted from MFRs demonstrate high predictive power for flares. Flare-producing active regions are characterized by magnetically complex configurations that are dominated by a single polarity rather than by balanced or purely unipolar structures. This finding is consistent with established physical theories of magnetic systems prone to flares. Our results suggest that CNNs can learn physically meaningful representations when trained on large-scale observations. Integrating XAI with quantitative magnetic field analysis improves the physical interpretability of deep learning-based flare prediction models, making them useful tools for prediction and modeling investigation in solar physics.
△ Less
Submitted 17 July, 2026;
originally announced July 2026.
-
Understanding before Naming! Enhancing LLM-based Method Name Prediction with Code Summarization
Authors:
Wei Liu,
Weisong Sun,
Tingting Xu,
Hanwei Qian,
Yi Zhao,
Chunrong Fang,
Xia Feng
Abstract:
Method names are critical to software quality, affecting code comprehensibility, maintainability, and developer collaboration. However, manually designing meaningful method names is challenging. Method Name Prediction (MNP), which automatically generates method names from code snippets, has recently attracted attention. Although large language models (LLMs) show promising performance for MNP, two…
▽ More
Method names are critical to software quality, affecting code comprehensibility, maintainability, and developer collaboration. However, manually designing meaningful method names is challenging. Method Name Prediction (MNP), which automatically generates method names from code snippets, has recently attracted attention. Although large language models (LLMs) show promising performance for MNP, two challenges remain. First, existing evaluations mainly rely on token similarity metrics, which often fail to reflect human judgments of semantic quality. Second, current LLM-based MNP methods usually generate names through direct code-to-name mapping, which differs from the human process of understanding functionality before naming. To address these challenges, we conduct empirical studies on LLM-based evaluation and MNP strategies. We compare 6 metric-based evaluators, 5 LLM-based evaluators, and 6 human evaluators. Results show that LLM-based evaluators, especially DeepSeek-based evaluators, are more consistent with human judgments than traditional metrics. We further compare direct generation and summarization-and-refinement strategies. Results indicate that summarization and refinement generally improve the semantic quality of generated names. Case studies reveal three limitations: inaccurate summaries, semantic misalignment, and close semantic scores. Based on these findings, we propose SMNP, an MNP approach combining MNP-oriented summarization and chain-of-thought enhanced refinement. Experiments on 5 LLMs and 2 datasets demonstrate the effectiveness and robustness of SMNP.
△ Less
Submitted 14 July, 2026;
originally announced July 2026.
-
ReProAgent: Tool-Augmented Multi-Stage Agentic Generation of Bug Reproduction Tests from Issue Reports
Authors:
Quanjun Zhang,
Yi Zheng,
Ye Shang,
Weifeng Sun,
Haichuan Hu,
Chunrong Fang,
Zhenyu Chen,
Liang Xiao
Abstract:
Reproduction tests help developers confirm reported issues and provide executable feedback for issue resolution, yet issue reports in open-source projects rarely include such tests. Recent studies have explored generating issue reproduction tests from issue reports with large language models, but existing approaches largely rely on prompt-based pipelines that retrieve textual context and generate…
▽ More
Reproduction tests help developers confirm reported issues and provide executable feedback for issue resolution, yet issue reports in open-source projects rarely include such tests. Recent studies have explored generating issue reproduction tests from issue reports with large language models, but existing approaches largely rely on prompt-based pipelines that retrieve textual context and generate tests. This limits their ability to understand how reported issues behave in repository-scale codebases and to flexibly organize the construction of reproduction tests. In this paper, we propose ReProAgent, a multi-stage agent framework for reproduction test generation from issue reports. ReProAgent decomposes the task into four agent stages: bug localization, root cause analysis, test planning, and test generation. To support these stages, ReProAgent integrates task-specific tools for task decomposition and reflection, context retrieval from both textual sources and repository graphs, and runtime interaction with the execution environment. Experiments on SWT-bench-lite and SWT-bench-verified show that ReProAgent successfully reproduces 58.43% and 70.30% of issues, outperforming all baselines, with an average cost of $0.14 per instance. For example, when equipped with GPT-5-mini, ReProAgent exceeds OpenHands with the same backbone by 20.43 and 7.90 percentage points, respectively. ReProAgent also generalizes across multiple backbone LLMs and improves downstream issue resolution performance when integrated with existing repair approaches.
△ Less
Submitted 10 July, 2026;
originally announced July 2026.
-
Multi-Agent LLM Collaboration for Unit Test Generation via Human-Testing-Inspired Workflows
Authors:
Quanjun Zhang,
Ye Shang,
Siqi Gu,
Jianyi Zhou,
Chunrong Fang,
Zhenyu Chen,
Liang Xiao
Abstract:
Recently, the emergence of Large Language Models (LLMs) has spurred a surge of research into automated unit test generation, yielding impressive performance and reducing manual effort. However, existing LLM-based approaches still suffer from two major limitations: (1) they follow rigid, procedural workflows that underutilize the autonomous reasoning potential of LLMs, making it difficult to dynami…
▽ More
Recently, the emergence of Large Language Models (LLMs) has spurred a surge of research into automated unit test generation, yielding impressive performance and reducing manual effort. However, existing LLM-based approaches still suffer from two major limitations: (1) they follow rigid, procedural workflows that underutilize the autonomous reasoning potential of LLMs, making it difficult to dynamically adapt testing strategies based on real-time feedback; and (2) they rely on rule-based context extraction that is not tailored to test generation, failing to capture fine-grained code dependencies and test-specific knowledge required for deriving test requirements. In this paper, we propose TestAgent, an LLM-based test generation approach that addresses the above limitations by emulating human testing practices via a multi-agent collaboration mechanism. Particularly, TestAgent designs three specialized agents, namely a requirement planner, a test generator, and a test reviewer, to simulate how developers understand, construct, and validate unit tests. To unleash the autonomous capabilities of LLMs, we equip TestAgent with a set of tool APIs that can be invoked dynamically in an on-demand and adaptive manner. To further support repository-level reasoning, TestAgent constructs a test-specialized knowledge graph via static analysis, which captures code entities and their dependencies across the project and persistently stores testing artifacts (e.g., test reports and failure analyses) produced during generation. Experimental results show that TestAgent achieves 97.46% execution rate, 92.34% line coverage, 90.24% branch coverage, and 83.69% mutation score on six Java projects, outperforming LLM-based baselines across all metrics and achieving substantially higher mutation scores than search-based tools.
△ Less
Submitted 10 July, 2026;
originally announced July 2026.
-
SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling
Authors:
Jiahao Wang,
Kaizhan Lin,
Kaixi Zhang,
Jinbo Han,
Xingda Wei,
Sijie Shen,
Chenguang Fang,
Wenyuan Yu,
Rong Chen,
Haibo Chen
Abstract:
LLM scheduling is critical to serving, yet it remains unclear how well existing designs fit agentic serving--with LLM requests issued by agents instead of humans. This shifts the workload in two ways: (1) agents act only on complete responses, making the cluster's tokens per second (TPS) the primary goal and relaxing--not eliminating--per-token latency requirements; and (2) requests share much of…
▽ More
LLM scheduling is critical to serving, yet it remains unclear how well existing designs fit agentic serving--with LLM requests issued by agents instead of humans. This shifts the workload in two ways: (1) agents act only on complete responses, making the cluster's tokens per second (TPS) the primary goal and relaxing--not eliminating--per-token latency requirements; and (2) requests share much of their KV\$-reuse exceeds 80% of request tokens in a production trace from BAILIAN, versus 54-62% in chat.
This paper first contributes a systematic study of request scheduling for agents on two real-world traces. We find that to increase KV\$ reuse, existing schedulers overly prioritize routing requests to instances caching their KV\$, overloading a few while leaving the rest idle, capping TPS. We thus present two key insights: (1) load balance need not sacrifice all KV\$ reuse, thanks to the global-tier KV\$ store and (2) by utilizing the workload's intra-session locality, balancing a small fraction of requests--the first request in each agent session--suffices to balance the cluster without sacrificing most KV\$ reuse on local instances.
SMETRIC realizes these insights with balanced session-centric scheduling: it routes each session's first request purely for load balance and its follow-up requests in a cache-aware manner, preserving load balance and local reuse while keeping demand on the global tier low. Using the session turn information as the scheduling metric is deliberate: it is derived efficiently and accurately from the user inputs alone, so the scheduler stays clean and stateless. SMETRIC improves cluster TPS by 10-16% under prefill-decode colocation with a global store and prefill TPS by 2-34% under disaggregation over state-of-the-art schedulers, also with a better per-token latency.
△ Less
Submitted 9 July, 2026;
originally announced July 2026.
-
Rise From The Ashes: LLM-based Static Analysis for Deep Learning Framework Bugs
Authors:
Shaoyu Yang,
Haifeng Lin,
Chunrong Fang,
Xiang Chen,
Wei Cheng,
Jiawei Liu,
Yiyu Zhang,
Hongyu Liu,
Zhenyu Chen
Abstract:
Deep learning (DL) frameworks are critical AI infrastructures that often hide bugs with serious security implications. While dynamic approaches such as fuzzing are effective in uncovering these bugs, they require real test execution and incur high computational costs. Static analysis is a natural complement because it can detect bugs without runtime execution, offering fast and scalable testing. U…
▽ More
Deep learning (DL) frameworks are critical AI infrastructures that often hide bugs with serious security implications. While dynamic approaches such as fuzzing are effective in uncovering these bugs, they require real test execution and incur high computational costs. Static analysis is a natural complement because it can detect bugs without runtime execution, offering fast and scalable testing. Unfortunately, there is still limited work targeting static analysis for DL frameworks due to their multilingual architectures and tensor-related program state.
We present Phoenix, the first LLM-based static analysis technique for DL frameworks. Our key insight is that cross-language tensor flows in DL frameworks can be modeled, together with concrete code context, as a structured semantic bridge intermediate representation (SBIR) that LLMs can analyze for potential bugs in tensor semantic propagation. We implement this insight through a multi-agent workflow. A summarization agent first distills bug summaries from historical bug-fix patches and CWE rules. Guided by each summary, an extraction agent identifies bug-relevant repository symbols for code retrieval, and a generation agent synthesizes grounded SBIRs from the retrieved context. Finally, an analysis agent is leveraged to check SBIRs and report potential bugs. Our evaluation shows that Phoenix is a practical complement to dynamic DL framework testing for bug finding. To date, Phoenix has found 31 real new bugs in PyTorch for different heterogeneous hardware backends (Intel CPU, NVIDIA CUDA, and Apple MPS). Among them, 20 submitted bug-fixing patches have been merged into upstream.
△ Less
Submitted 1 July, 2026;
originally announced July 2026.
-
Ontology-Guided Evidence Path Inference for Multi-hop Knowledge Graph Question Answering
Authors:
Yongxue Shan,
Meihan Wu,
Cundi Fang,
Jie Peng,
Xiaodong Wang
Abstract:
Knowledge graph question answering (KGQA) aims to answer natural-language questions by reasoning over structured facts. Existing multi-hop KGQA methods mainly rely on topic-centered expansion, which faces two key challenges: the search space rapidly grows with noisy mixed-type paths, and retrieved paths may fail to satisfy the semantic constraints of complex questions. To address these challenges,…
▽ More
Knowledge graph question answering (KGQA) aims to answer natural-language questions by reasoning over structured facts. Existing multi-hop KGQA methods mainly rely on topic-centered expansion, which faces two key challenges: the search space rapidly grows with noisy mixed-type paths, and retrieved paths may fail to satisfy the semantic constraints of complex questions. To address these challenges, we propose OPI, an ontology-guided evidence path inference framework for multi-hop KGQA. OPI introduces a relation-centric ontology graph to capture the head-tail type constraints of relations, providing a compact interface for answer-side constraints. Based on this ontology graph, OPI first introduces a bidirectional retrieval mechanism by mapping the predicted answer type to compatible final-hop relations and combining topic-side prefix expansion with answer-side final-hop matching, thereby suppressing noisy mixed-type expansion. OPI further adopts an iterative refinement strategy to reassess retrieved paths and candidate answers under the question context, filtering type-compatible but question-irrelevant evidence for more reliable answer prediction. Experiments on WebQSP, CWQ, and MetaQA show that OPI substantially reduces the search space, improves Hit@1/F1 by 4.6/5.0 points on WebQSP and 8.9/3.3 points on CWQ over the strongest prior results, and achieves near-saturated Hit@1 on MetaQA with the retrieval module alone.
△ Less
Submitted 26 June, 2026;
originally announced June 2026.
-
Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling
Authors:
Kun Zhang,
Chenxin Fang,
Tao Chen,
Baiyang Song,
Yunhang Shen,
Yiyi Zhou,
Rongrong Ji
Abstract:
Long video understanding remains a daunting challenge for Multimodal Large Language Models (MLLMs) due to the excessive computation and memory footprint. Thus, keyframe selection is often adopted to mitigate this shortcoming, which however still suffers from low flexibility and high noise due to its hard sampling principle. In this paper, we define video frame selection as a problem of Quasi-Gauss…
▽ More
Long video understanding remains a daunting challenge for Multimodal Large Language Models (MLLMs) due to the excessive computation and memory footprint. Thus, keyframe selection is often adopted to mitigate this shortcoming, which however still suffers from low flexibility and high noise due to its hard sampling principle. In this paper, we define video frame selection as a problem of Quasi-Gaussian Sampling, and propose an adaptive and training-free approach termed AdaQ. Inspired by the 3-$σ$ rule of Gaussian distribution, the objective of AdaQ is to achieve the optimal 3-$σ$ interval for different examples, i.e., a smaller 3-$σ$ interval for the local query and a larger one for the global query, thereby facilitating robust and adaptive frame sampling. To validate AdaQ, we apply it to four MLLMs with three embedding models. The extensive experimental results not only show its obvious performance gains over the default MLLMs and the SOTA keyframe selection methods, e.g., helping Qwen3-VL-8B outperform GPT4o by 15.8% on average by using only 64 frames, but also confirm its superior robustness and high efficiency for long-video understanding, e.g., only 1 hyper-parameter needs to be set.
△ Less
Submitted 24 June, 2026; v1 submitted 23 June, 2026;
originally announced June 2026.
-
Tri-Efficient Transfer Learning for Point Cloud Videos
Authors:
Yiding Sun,
Dongxu Zhang,
Jihua Zhu,
Haozhe Cheng,
Zhengqiao Li,
Pengcheng Li,
Chaowei Fang,
Yonghao Dong,
Lin Chen
Abstract:
While point cloud foundation models have significantly advanced point cloud video understanding, existing parameter-efficient fine-tuning (PEFT) methods still suffer from two critical limitations: prohibitive annotation costs for large-scale point cloud datasets and severe memory bottlenecks. In this paper, we aim to mine richer supervision signals from existing data rather than blindly scaling da…
▽ More
While point cloud foundation models have significantly advanced point cloud video understanding, existing parameter-efficient fine-tuning (PEFT) methods still suffer from two critical limitations: prohibitive annotation costs for large-scale point cloud datasets and severe memory bottlenecks. In this paper, we aim to mine richer supervision signals from existing data rather than blindly scaling datasets. A further key principle is that the memory footprint of fine-tuning must be drastically reduced compared to full fine-tuning, which remains elusive for current PEFT techniques. Driven by these challenges, we identify three core desiderata: data-, parameter-, and memory efficiency, and present PoinTriE, a unified framework that excels along all three dimensions. For pre-training, pseudo-motion trajectories are synthesized via rigid transformations, paired with text corpora and 2D projections derived from raw point clouds. We then propose a Geometric-Motion Duality Network optimized via multimodal contrastive learning, rigid rotation prediction, and motion distribution divergence to produce dense self-supervision. During fine-tuning, we freeze the pretrained backbone and only update a lightweight Spatio-temporal Side Network built with LoRA units. Equipped with a gradient flow masking strategy, PoinTriE simultaneously reduces memory consumption and parameter overhead. Extensive experiments confirm that PoinTriE establishes new state-of-the-art results on action recognition and semantic segmentation tasks.
△ Less
Submitted 23 June, 2026;
originally announced June 2026.
-
Probing Nuclear Effects with Transverse Kinematic Imbalance in Muon-neutrino Induced Charged-Current $π^0$ Production on Argon with the MicroBooNE Detector
Authors:
MicroBooNE collaboration,
P. Abratenko,
D. Andrade Aldana,
J. Asaadi,
A. Ashkenazi,
S. Balasubramanian,
B. Baller,
A. Barnard,
G. Barr,
D. Barrow,
J. Barrow,
V. Basque,
J. Bateman,
B. Behera,
O. Benevides Rodrigues,
S. Berkman,
A. Bhat,
V. Bhelande,
M. Bhattacharya,
A. Binau,
M. Bishai,
A. Blake,
B. Bogart,
T. Bolton,
M. B. Brunetti
, et al. (170 additional authors not shown)
Abstract:
Neutrino-nucleus cross-section measurements are needed to improve interaction modeling and to enable precision neutrino oscillation measurements in upcoming experiments such as the Deep Underground Neutrino Experiment (DUNE), Hyper-Kamiokande, and the Short-Baseline Neutrino program. Baryon-resonance neutrino interactions constitute a dominant contribution near the peak of the DUNE neutrino energy…
▽ More
Neutrino-nucleus cross-section measurements are needed to improve interaction modeling and to enable precision neutrino oscillation measurements in upcoming experiments such as the Deep Underground Neutrino Experiment (DUNE), Hyper-Kamiokande, and the Short-Baseline Neutrino program. Baryon-resonance neutrino interactions constitute a dominant contribution near the peak of the DUNE neutrino energy spectrum. We present the first measurement of muon neutrino charged-current resonance-like interactions on argon using transverse kinematic imbalance variables with the MicroBooNE detector. These observables are highly sensitive to the modeling of final-state interactions. This measurement probes kinematic imbalances using the reconstructed momenta of the muon, leading proton, and neutral pion. A comprehensive characterization of the $π^0$-proton final state is presented; however, none of the models considered are able to simultaneously reproduce all measured observables.
△ Less
Submitted 22 June, 2026;
originally announced June 2026.
-
SparseCol: A 1320 BTOPS/W Precision-scalable NPU Exploiting Training-free Structured Bit-level Sparsity and Dynamic Dataflow
Authors:
Man Shi,
Vikram Jain,
Weijie Jiang,
Chao Fang,
Antony Joseph,
Wim Dehaene,
Marian Verhelst
Abstract:
Bit-serial computation enables sequential processing of data at the bit level, providing several advantages, such as scalable computational precision. This approach has gained significant attention, especially for exploiting bit-level sparsity in AI workloads. While current bit-serial processors leverage bit-level sparsity to eliminate the computation associated with zero bits, they face a fundame…
▽ More
Bit-serial computation enables sequential processing of data at the bit level, providing several advantages, such as scalable computational precision. This approach has gained significant attention, especially for exploiting bit-level sparsity in AI workloads. While current bit-serial processors leverage bit-level sparsity to eliminate the computation associated with zero bits, they face a fundamental trade-off: either they suffer from low memory-access and computation efficiency caused by irregular patterns of non-zero bits, or they incur substantial area overhead from complex online scheduling mechanisms required to reorganize bit-level data and preserve memory access and computation regularity. Therefore, we present the SparseCol processor, designed to harness extensive bit sparsity while maintaining high hardware utilization across various AI applications, including CNNs, RNNs, and transformers. In contrast to traditional methods, SparseCol exploits structured bit-level sparsity, denoted by bit-column sparsity, without requiring any re-training. Furthermore, SparseCol implements a dynamic dataflow architecture that tackles hardware under-utilization issues commonly found in existing bit-serial solutions. Fabricated in 16nm CMOS node, SparseCol delivers 1320 BTOPS/W (BTOPS represents Binary Tera-Operations Per Second, calculated as #W bits x #A bits TOPS) peak efficiency while maintaining accuracy, outperforming SotA sparse processors in terms of efficiency by 6.8x. Comprehensive evaluations on CNN classification tasks and transformer architectures demonstrate system-level efficiencies of 745.02 BTOPS/W and 850.5 BTOPS/W, respectively.
△ Less
Submitted 14 June, 2026;
originally announced June 2026.
-
Orchestrated Reality: From Role-Play to Living, Playable Game Worlds -- LLM-Driven World Simulation as a Parameterized-Action POMDP
Authors:
Yuhang Huang,
Chenmiao Li,
Chaowei Fang
Abstract:
Many games rely on storytelling combined with systems that track levelling, NPC behaviour, and consequence simulation; bridging tightly-authored narrative with deeply-simulated worlds -- most acute in sandbox and open-world settings -- has been prohibitively expensive. LLM-driven worlds open a new path: a single harness can coordinate numerical state, narrative voice, storytelling pacing, and rule…
▽ More
Many games rely on storytelling combined with systems that track levelling, NPC behaviour, and consequence simulation; bridging tightly-authored narrative with deeply-simulated worlds -- most acute in sandbox and open-world settings -- has been prohibitively expensive. LLM-driven worlds open a new path: a single harness can coordinate numerical state, narrative voice, storytelling pacing, and rule logic together. Realising this requires the LLM system to sustain a persistent world (who is where, what has just happened, what is currently true), which today's deployed systems do not: the narrative voice asserts state in free prose without any validated representation, so a fully autonomous game engine remains infeasible. We treat this as an architectural choice, not a limitation of language models, and report work in progress on a framework -- orchestrated reality -- that makes the world a canonical object owned by a singleton orchestration agent analogous to the tabletop-RPG Game Master (GM). We formalise an LLM-driven game world for a human player as a Parameterized-Action POMDP: state is a tree of canonical JSON entities, actions decompose as $a=(k, x_k)$ (a discrete intent kind plus structured JSON parameters), the agent observes only a narrative projection $o=O(s)$ of state, and the transition kernel $F$ is an LLM-driven Plan-Diff-Validate-Apply (PDVA) pipeline that commits schema-validated, content-hashed JSON deltas. We give the formal model, a JSON-state example, a worked single-turn example, and a catalogue of 15 illustrative incidents drawn from a real deployment showing the framework in action. Empirical validation through a planned human player study -- together with multi-NPC concurrent agency and deployment as an RL environment -- is situated as future work.
△ Less
Submitted 14 June, 2026;
originally announced June 2026.
-
Investigating Metamorphic Fuzz Oracle Enhancement via Large Language Models
Authors:
Ruixiang Qian,
Ding Yang,
Zengxu Chen,
Yuxuan Gao,
Chunrong Fang,
Chao Zhang,
Zhenyu Chen
Abstract:
Fuzz drivers are essential components of greybox fuzzing, as they encapsulate target interfaces, define test spaces, and largely determine fuzzing effectiveness. Existing fuzz drivers typically rely on crash-based oracles for security testing, overlooking library functionality and limiting bug detection capability.
In this paper, we present the first study on metamorphic-based fuzz oracle enhanc…
▽ More
Fuzz drivers are essential components of greybox fuzzing, as they encapsulate target interfaces, define test spaces, and largely determine fuzzing effectiveness. Existing fuzz drivers typically rely on crash-based oracles for security testing, overlooking library functionality and limiting bug detection capability.
In this paper, we present the first study on metamorphic-based fuzz oracle enhancement (MFOE), which augments existing fuzz drivers with metamorphic-based oracles derived from metamorphic relations (MRs). Since constructing and integrating such oracles requires substantial domain knowledge, automating MFOE is challenging. To address this challenge, we propose MetaFOE, an LLM-based framework that automatically generates and integrates metamorphic-based oracles.
We evaluate MetaFOE on OSS-Fuzz drivers using three modern LLMs and five prompt strategies. MetaFOE generates 3,475 MRs, of which 77.3% are applicable, and implements 12,351 meta drivers, with 6,228 being valid. After three hours of fuzzing, the valid meta drivers improve edge coverage by an average of 18.7% and trigger 1,528 unique crashes. Our results demonstrate both the effectiveness of metamorphic-based oracle enhancement and the feasibility of using LLMs to automate MFOE, providing valuable insights for advancing greybox fuzzing.
△ Less
Submitted 12 June, 2026;
originally announced June 2026.
-
First Measurement of Sub-GeV $ν_μ$ Charged-Current Coherent Pion Production on Argon in MicroBooNE
Authors:
MicroBooNE collaboration,
P. Abratenko,
D. Andrade Aldana,
J. Asaadi,
A. Ashkenazi,
S. Balasubramanian,
B. Baller,
A. Barnard,
G. Barr,
D. Barrow,
J. Barrow,
V. Basque,
J. Bateman,
B. Behera,
O. Benevides Rodrigues,
S. Berkman,
A. Bhat,
M. Bhattacharya,
V. Bhelande,
A. Binau,
M. Bishai,
A. Blake,
B. Bogart,
T. Bolton,
M. B. Brunetti
, et al. (167 additional authors not shown)
Abstract:
We report a measurement of the charged-current coherent pion production cross section on argon using the MicroBooNE liquid argon time projection chamber exposed to the Booster Neutrino Beam at Fermilab. The measurement uses the MicroBooNE data set corresponding to $1.26 \times 10^{21}$ protons on target with a mean neutrino energy of $0.8$~GeV. The flux-averaged cross section is measured to be…
▽ More
We report a measurement of the charged-current coherent pion production cross section on argon using the MicroBooNE liquid argon time projection chamber exposed to the Booster Neutrino Beam at Fermilab. The measurement uses the MicroBooNE data set corresponding to $1.26 \times 10^{21}$ protons on target with a mean neutrino energy of $0.8$~GeV. The flux-averaged cross section is measured to be $(9.1 \pm 1.2_{\text{stat}} \pm 1.2_\text{syst}) \times 10^{-40}\,\text{cm}^2/\text{Ar}$. This result represents the first measurement of charged-current coherent pion production on argon at sub-GeV neutrino energies. Due to its clean two-body kinematics, where the neutrino interacts coherently with the entire nucleus producing a forward muon and pion with no nuclear breakup, this process provides a useful tool for constraining neutrino flux uncertainties in current and future oscillation experiments such as DUNE.
△ Less
Submitted 11 June, 2026;
originally announced June 2026.
-
Securing Code Understanding: Detecting Natural Backdoor Vulnerability in Code Language Models
Authors:
Yuchen Chen,
Weisong Sun,
Haocheng Huang,
Yuan Xiao,
Chunrong Fang,
Yiran Zhang,
Tingting Xu,
Zhenpeng Chen,
An Guo,
Peizhuo Lv,
Xiaofang Zhang,
Zhenyu Chen,
Yang Liu,
Baowen Xu
Abstract:
Code Language Models (CodeLMs) have become integral to software engineering, significantly advancing code intelligence tasks. However, their widespread adoption has raised critical security concerns, particularly regarding susceptibility to backdoor attacks. Recent studies have uncovered naturally occurring backdoors, referred to as natural backdoors, in normally trained deep learning models. Desp…
▽ More
Code Language Models (CodeLMs) have become integral to software engineering, significantly advancing code intelligence tasks. However, their widespread adoption has raised critical security concerns, particularly regarding susceptibility to backdoor attacks. Recent studies have uncovered naturally occurring backdoors, referred to as natural backdoors, in normally trained deep learning models. Despite posing threats as serious as those introduced through data poisoning, security implications of natural backdoor vulnerabilities in CodeLMs remain poorly understood.
In this paper, we conduct a thorough empirical study of natural backdoor vulnerabilities in CodeLMs across various model architectures and code intelligence tasks. Specifically, we examine potential natural backdoor vulnerabilities across 44 scenarios, demonstrating that natural backdoors are prevalent and intrinsic to CodeLMs. We reveal differences between injected and natural backdoor vulnerabilities at both the model and parameter levels. We then analyze the transferability of natural backdoor vulnerabilities from three perspectives: datasets, model architectures, and shared knowledge. We further investigate the causes of natural backdoors from two aspects: training datasets and the model training procedure. We evaluate existing backdoor defense techniques, including pre-training, in-training, and post-training defenses, in mitigating natural backdoors. Finally, we propose ScanNBT, a novel detection method designed to improve comprehensive detection of natural backdoor vulnerabilities in CodeLMs. We aim for our findings to enhance understanding of these vulnerabilities and provide insights for strengthening CodeLM security against backdoor threats.
△ Less
Submitted 9 June, 2026;
originally announced June 2026.
-
Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation
Authors:
Yuchen Ling,
Shengcheng Yu,
Zhenyu Chen,
Chunrong Fang
Abstract:
Large language model (LLM) agents are rapidly moving from conversational interfaces to software components that plan, invoke tools, maintain memory, and act on external environments. This transition changes the nature of security risk. In agentic settings, failures are no longer limited to unsafe text generation. Untrusted content may redirect control flow, misuse tool privileges, corrupt persiste…
▽ More
Large language model (LLM) agents are rapidly moving from conversational interfaces to software components that plan, invoke tools, maintain memory, and act on external environments. This transition changes the nature of security risk. In agentic settings, failures are no longer limited to unsafe text generation. Untrusted content may redirect control flow, misuse tool privileges, corrupt persistent state, leak sensitive information, or trigger harmful external actions. At the same time, research on LLM agent security is expanding quickly but remains fragmented across attack families, defense layers, application domains, and evaluation settings. This paper synthesizes 247 papers through a lifecycle-based, systems-oriented framework that models agent security around the interaction of information flow, delegated authority, and persistent state. We organize the literature around four questions: how LLM agent security should be modeled, which threat surfaces and attack families dominate, what defenses have been proposed and with what tradeoffs, and how security claims are evaluated. We find that prompt injection and tool-mediated control-flow hijacking still dominate the field, while persistent state corruption and multi-agent propagation are becoming central emerging concerns. We further find that current defenses provide useful building blocks but remain weakly compositional, and that existing benchmarks still underrepresent long-horizon, stateful, and deployment-sensitive risks. We argue that secure LLM agents require explicit trust boundaries, principled privilege control, provenance-aware state management, and evaluation practices aligned with realistic operational settings.
△ Less
Submitted 9 June, 2026;
originally announced June 2026.
-
LentiAvatar: Pseudo-Multiview Reconstruction and Subpixel Prism Rendering for Real-Time Stereoscopic Communication
Authors:
Chufeng Fang,
Dongdong Teng,
Lilin Liu
Abstract:
Real-time stereoscopic video communication has long been a goal of immersive telepresence, yet practical systems still require specialized capture rigs or reduce remote users to a single portrait view. We present LentiAvatar, a Gaussian head-avatar system that connects monocular avatar capture with subpixel-encoded glasses-free lenticular display for real-time autostereoscopic communication. From…
▽ More
Real-time stereoscopic video communication has long been a goal of immersive telepresence, yet practical systems still require specialized capture rigs or reduce remote users to a single portrait view. We present LentiAvatar, a Gaussian head-avatar system that connects monocular avatar capture with subpixel-encoded glasses-free lenticular display for real-time autostereoscopic communication. From a monocular portrait video, LentiAvatar reconstructs a controllable head avatar and optimizes it for the lateral viewing zones induced by the display. The method uses natural head turns as pseudo-multiview (PMV) supervision to constrain regions that are otherwise weakly observed in monocular training, including hair, ears, jaw contours, and neck boundaries. Reliable side frames are yaw-binned, aligned to virtual cameras, and supervised within a strict head-and-hair domain; contour-aware losses and staged regularization further suppress ghosting, alpha leakage, and depth instability while preserving lateral detail. At runtime, LentiAvatar renders 32 virtual views and encodes them into a 4K lenticular raster with calibrated subpixel-routing masks. The live-tracker prototype sustains 10.65 FPS, and a subject-specific distilled driver raises the same display pipeline to 38.49 FPS.
△ Less
Submitted 15 June, 2026; v1 submitted 9 June, 2026;
originally announced June 2026.
-
An Adaptive Data cleaning Framework for Noisy Label Detection
Authors:
Chen-Hsuan Fang,
Wei-Hsinag Chen,
Pin-Hsuan Yu,
Jung-Hua Wang,
Tsung-Wei Pan
Abstract:
Deep neural networks (DNNs) excel in computer vision tasks given large annotated datasets. In real-world applications, however, labels are often corrupted by ambiguity, human error, or dynamic environments. Over-parameterized DNNs easily memorize these noisy labels during training, degrading model accuracy and generalization. Existing data-cleaning and sample-selection strategies often rely on man…
▽ More
Deep neural networks (DNNs) excel in computer vision tasks given large annotated datasets. In real-world applications, however, labels are often corrupted by ambiguity, human error, or dynamic environments. Over-parameterized DNNs easily memorize these noisy labels during training, degrading model accuracy and generalization. Existing data-cleaning and sample-selection strategies often rely on manually specified thresholds, prior knowledge of the noise ratio, or a single metric (either learning dynamics or geometric structure), making them unstable in complex data regimes. This paper proposes a self-adaptive data-cleaning framework that integrates local, global, and learning dynamics cues for robust noisy-label detection. Samples are mapped into a unified low-dimensional feature space through a modular feature concatenation paradigm. We provide two instantiations: a 2D metric integrating class-adaptive KNN-based local disagreement with k-means-based global centroid distance, and a 3D multi-metric that additionally incorporates a z-normalized score. Unlike conventional 1D Gaussian Mixture Models applied to a single scalar metric, our framework performs multi-metric clustering on the feature space to adaptively partition samples into clean-dominant and noise-dominant components without requiring manual thresholds or noise priors. Experiments on CIFAR-10, MNIST, and ImageNet-100 with 5% to 40% symmetric label noise show high recall across settings, including near-perfect recall (>=98%) on ImageNet-100 at 40% noise. Subsequent training yields accuracy gains across evaluated settings, especially under severe corruption on ImageNet-100. These findings suggest that multi-metric integration provides a threshold-free, practical, and low-tuning strategy for noisy label detection.
△ Less
Submitted 13 June, 2026; v1 submitted 5 June, 2026;
originally announced June 2026.
-
DiffSlack: Learning under Nonlinear Inequality Constraints via Learnable Slack Variables
Authors:
Ziqian Wang,
Chenxi Fang,
Zhen Zhang
Abstract:
Enforcing nonlinear inequality constraints in neural networks remains challenging, especially when the output is subject to many coupled constraints. Existing hard constraint methods often impose structural restrictions on the constraint set or introduce substantial computational overhead for large-scale nonlinear problems. Here, we propose DiffSlack, a differentiable projection layer for nonlinea…
▽ More
Enforcing nonlinear inequality constraints in neural networks remains challenging, especially when the output is subject to many coupled constraints. Existing hard constraint methods often impose structural restrictions on the constraint set or introduce substantial computational overhead for large-scale nonlinear problems. Here, we propose DiffSlack, a differentiable projection layer for nonlinear inequality-constrained neural prediction. DiffSlack reformulates inequalities as equalities with learnable slack variables, which are predicted as part of the augmented network output and provide a data-driven warm start for damped Gauss-Newton projection. The projection layer maps raw predictions onto the augmented feasible manifold while preserving end-to-end differentiability. A two-stage curriculum further stabilizes training and improves constraint satisfaction. We evaluate DiffSlack on vehicle path planning with 200 nonlinear inequality constraints from collision avoidance, curvature limits, and waypoint spacing. Compared with existing learning-based baselines, DiffSlack achieves a higher planning success rate and stronger geometric constraint satisfaction under a comparable inference budget. Ablation studies further show that the hard projection layer reduces sensitivity to supervision quality. Closed-loop tracking in CARLA and real-world vehicle experiments confirms the executability of the generated trajectories. These results demonstrate that DiffSlack provides a practical and scalable approach to embedding hard inequality constraints into neural networks for engineering applications.
△ Less
Submitted 3 June, 2026;
originally announced June 2026.
-
Stumbling Into AI Emotional Dependence: How Routine AI Interactions Reshape Human Connection
Authors:
Yaoxi Shi,
Cathy Mengying Fang,
Pattie Maez,
Amit Goldenberg
Abstract:
Public discourse and emerging policy typically assume that AI emotional support is a deliberate act: a lonely user consciously seeking comfort from a dedicated companion chatbot. In this paper, we draw on emerging empirical evidence and argue that this picture is inaccurate on two accounts, both in how AI emotional support arises and how it shapes future behavior. First, AI emotional support commo…
▽ More
Public discourse and emerging policy typically assume that AI emotional support is a deliberate act: a lonely user consciously seeking comfort from a dedicated companion chatbot. In this paper, we draw on emerging empirical evidence and argue that this picture is inaccurate on two accounts, both in how AI emotional support arises and how it shapes future behavior. First, AI emotional support commonly emerges incidentally within task-oriented interactions on general-purpose platforms, much as workplace friendships deepen through collaboration. Second, these incidental encounters are path-dependent: positive experiences of AI emotional support update people's beliefs about AI's emotional capabilities and redirect their choices for future emotional support, increasing preference for AI and decreasing preference for humans. We review recent evidence, including a large-scale longitudinal study conducted in collaboration with OpenAI, showing that daily five-minute conversations with an AI about personal issues over 28 days led to a 10.3% decrease in the preference for seeking support from humans and an 11.6% increase in the preference for AI. These findings suggest that current policy, focused on companion apps and isolated interactions, cannot adequately protect human connection. Instead, effective regulations should extend to general-purpose AI systems and address cumulative, trajectory-level changes in how people seek support. Recognizing how people stumble into AI emotional support and how those encounters redirect human connections over time is essential to safeguarding human well-being.
△ Less
Submitted 2 June, 2026;
originally announced June 2026.
-
Holographic complexity of de-Sitter black holes
Authors:
Chaoxi Fang,
Jiayue Yang,
Shao-Wen Wei,
Ming Zhang,
Robert B. Mann
Abstract:
We investigate holographic complexity within the Schwarzschild-de Sitter (SdS) black hole spacetime. Two distinct de Sitter holography prescriptions are examined: the static patch scheme restricted to the stretched horizon and the de Sitter/Conformal Field Theory (dS/CFT) correspondence scheme defined at asymptotic future and past infinities. We evaluate the Complexity equals Volume (CV) conjectur…
▽ More
We investigate holographic complexity within the Schwarzschild-de Sitter (SdS) black hole spacetime. Two distinct de Sitter holography prescriptions are examined: the static patch scheme restricted to the stretched horizon and the de Sitter/Conformal Field Theory (dS/CFT) correspondence scheme defined at asymptotic future and past infinities. We evaluate the Complexity equals Volume (CV) conjecture and extend the analysis to codimension-zero proposals, specifically Complexity equals Spacetime Volume (CV2.0) and Complexity equals Action (CA), through the Wheeler-DeWitt (WDW) patch we construct. The behaviors of the complexity in the static patch holography at late time and in the dS/CFT at infinite spacelike boundary coordinate are studied, respectively. We find that under both the CV and CV2.0 conjectures, the static patch holographic complexity and the dS/CFT holographic complexity consistently exhibit linear growth. Conversely, regarding the CA conjecture, the holographic complexity growth rates for both the static patch and the dS/CFT correspondence vanish. This behavior is attributed to the finiteness of the (regularized) action within the restricted WDW region. Furthermore, it is demonstrated that the complexity growth rate of the static patch scheme is identical to that in the dS/CFT scheme. This equivalence implies the existence of a unified description for bulk dynamics within de Sitter holography.
△ Less
Submitted 1 June, 2026;
originally announced June 2026.
-
RRP-Voice: A Longitudinal Dataset and Benchmark for Recurrent Respiratory Papillomatosis Detection
Authors:
Wenze Ren,
Ke-Han Lu,
Kai-Wei Chang,
Tiantian Feng,
Ching Fang,
Zhi-Chi Liao,
Dao Thi Hai Yen,
Syu-Siang Wang,
Yu Tsao,
Chi-Te Wang,
Shih-Hau Fang
Abstract:
Deep learning has advanced pathological voice detection rapidly, yet rare laryngeal diseases remain underexplored due to data scarcity. Recurrent Respiratory Papillomatosis (RRP) exemplifies this gap: an HPV-induced disease of the larynx in which patients oscillate between recurrence and post-surgical remission over the years. RRP demands continuous voice monitoring that existing cross-sectional c…
▽ More
Deep learning has advanced pathological voice detection rapidly, yet rare laryngeal diseases remain underexplored due to data scarcity. Recurrent Respiratory Papillomatosis (RRP) exemplifies this gap: an HPV-induced disease of the larynx in which patients oscillate between recurrence and post-surgical remission over the years. RRP demands continuous voice monitoring that existing cross-sectional corpora cannot support. We introduce the first longitudinal voice dataset for RRP, comprising recordings from 26 patients with up to ten years of follow-up. Each session pairs sustained vowels with sentence-level utterances, which are annotated by otolaryngologists and confirmed synchronously with laryngoscopy. Building on this resource, we establish a systematic benchmark spanning handcrafted features, end-to-end deep networks, self-supervised pretrained models, and recent audio large language models, all evaluated under session-level cross-validation with patient-level audit. Per-subject longitudinal analyses further confirm that the cross-sectional discriminative signal reflects laryngoscopic disease state rather than stable speaker attributes. This work lays a foundation for rare longitudinal pathological voice tasks in low-resource clinical settings.
△ Less
Submitted 31 May, 2026;
originally announced June 2026.
-
What Makes a Strong Model? A Unified Spectral Analysis of Knowledge Transfer over High-dimensional Linear Regression
Authors:
Wendao Wu,
Fangqing Zhang,
Haihan Zhang,
Cong Fang
Abstract:
Teacher-Student Knowledge Transfer (KT) is ubiquitous in modern machine learning, ranging from classical model compression via Knowledge Distillation (KD) to the emergent phenomenon of Weak-to-Strong (W2S) generalization. While existing studies offer isolated insights, a unified theoretical framework explaining the efficacy of KT across these disparate regimes remains lacking. In this work, we est…
▽ More
Teacher-Student Knowledge Transfer (KT) is ubiquitous in modern machine learning, ranging from classical model compression via Knowledge Distillation (KD) to the emergent phenomenon of Weak-to-Strong (W2S) generalization. While existing studies offer isolated insights, a unified theoretical framework explaining the efficacy of KT across these disparate regimes remains lacking. In this work, we establish a unified spectral analysis of SGD dynamics in high-dimensional linear regression, elucidating the efficiency of KT across seemingly disparate regimes. We characterize KT efficiency through two distinct mechanisms: \emph{Spectral Horizon Expansion} in KD, which enables the capture of statistically inaccessible high-frequency signals, and \emph{Spectral Denoising} in W2S, where the student acts as a filter for optimization noise. Our framework unifies these phenomena, revealing that the efficacy of transfer is governed by the interplay between implicit regularization and heterogeneous spectral learning speeds over the spectrum.
△ Less
Submitted 31 May, 2026;
originally announced June 2026.
-
Characterizing the energy resolution of the MicroBooNE LArTPC at the MeV scale using monoenergetic features of $^{208}$Tl decays
Authors:
MicroBooNE collaboration,
P. Abratenko,
D. Andrade Aldana,
J. Asaadi,
A. Ashkenazi,
S. Balasubramanian,
B. Baller,
A. Barnard,
G. Barr,
D. Barrow,
J. Barrow,
V. Basque,
J. Bateman,
B. Behera,
O. Benevides Rodrigues,
S. Berkman,
A. Bhat,
M. Bhattacharya,
V. Bhelande,
A. Binau,
M. Bishai,
A. Blake,
B. Bogart,
T. Bolton,
M. B. Brunetti
, et al. (167 additional authors not shown)
Abstract:
A detailed understanding of the capabilities and fidelity of low-energy reconstruction is crucial for taking advantage of MeV-scale neutrino physics opportunities in liquid argon time projection chambers (LArTPCs). This study presents a measurement of the resolution of reconstructed energy in the MicroBooNE LArTPC at $\approx 1.5$ MeV. The characterization is performed using monoenergetic signals…
▽ More
A detailed understanding of the capabilities and fidelity of low-energy reconstruction is crucial for taking advantage of MeV-scale neutrino physics opportunities in liquid argon time projection chambers (LArTPCs). This study presents a measurement of the resolution of reconstructed energy in the MicroBooNE LArTPC at $\approx 1.5$ MeV. The characterization is performed using monoenergetic signals generated by $2.614$ MeV $γ$-rays from $^{208}$Tl decays undergoing pair production in the detector. The resolution is found to be ($7.52 \pm 0.78 \text{(stat)} \pm 0.92 \text{(syst)}$)%. This value is consistent with the MicroBooNE simulation prediction of ($9.70 \pm 0.65 \text{(stat)}$)% at the $1.6 σ$ level. This study represents the first ever measurement of LArTPC energy resolution at the MeV scale and provides a pathway for monoenergetic energy calibrations in future experiments using LArTPC detectors.
△ Less
Submitted 17 August, 2026; v1 submitted 28 May, 2026;
originally announced May 2026.
-
Metasurfaces for neutral-atom trapping
Authors:
Chengyu Fang,
Minjeong Kim,
Mark Saffman,
Jennifer T. Choy,
Mikhail Kats
Abstract:
Trapped neutral atoms are one of the leading platforms for quantum information technologies, in particular for quantum computing, but scaling them to array sizes needed for utility-scale quantum computing is a major engineering challenge. Here we review optical metasurfaces as an enabling technology that provides fine control over the phase, amplitude, and polarization of light, with pixel counts…
▽ More
Trapped neutral atoms are one of the leading platforms for quantum information technologies, in particular for quantum computing, but scaling them to array sizes needed for utility-scale quantum computing is a major engineering challenge. Here we review optical metasurfaces as an enabling technology that provides fine control over the phase, amplitude, and polarization of light, with pixel counts far exceeding what is available with spatial light modulators (SLMs) and other active devices. The large pixel counts have recently led to demonstrations of arrays of optical tweezers with hundreds of thousands of sites and arrays of optical bottle-beams with complex three-dimensional trapping profiles. The flexibility and scalability of optical metasurfaces provides a route towards miniaturized, integrated, and highly scalable atomic experiments and instruments.
△ Less
Submitted 8 June, 2026; v1 submitted 28 May, 2026;
originally announced May 2026.
-
EvoRepair: Enhancing Vulnerability Repair Agents Through Experience-Based Self-Evolution
Authors:
Haichuan Hu,
Guoqing Xie,
Quanjun Zhang,
Jiawei Liu,
Shengcheng Yu,
Chunrong Fang,
Zhenyu Chen,
Liang Xiao
Abstract:
Large Language Models (LLMs) have shown promise for automated vulnerability repair (AVR), but they still face several limitations, including the lack of intra-vulnerability experience accumulation and the lack of cross-vulnerability experience reuse. As a result, LLMs may repeatedly make similar mistakes during iterative repair and underutilize valuable repair knowledge from historical vulnerabili…
▽ More
Large Language Models (LLMs) have shown promise for automated vulnerability repair (AVR), but they still face several limitations, including the lack of intra-vulnerability experience accumulation and the lack of cross-vulnerability experience reuse. As a result, LLMs may repeatedly make similar mistakes during iterative repair and underutilize valuable repair knowledge from historical vulnerabilities. To address these challenges, we propose EvoRepair, the first experience-based self-evolving AVR agent framework that enables LLMs to accumulate, refine, and leverage domain-specific knowledge across long-horizon vulnerability repairs. EvoRepair follows a cyclic learn-and-repair process that retrieves relevant past experiences to guide repair, extracts new experiences from repair trajectories, and updates an experience bank using quality-aware scoring. We evaluate EvoRepair against 12 representative vulnerability repair baselines on PATCHEVAL and SEC-bench using GPT-5-mini. Results show that EvoRepair achieves the best overall performance, reaching 93.47% on PATCHEVAL, 87.00% on SEC-bench, and 90.46% overall. In particular, EvoRepair outperforms latest LLM-based baseline LoopRepair by 39.56% and 33.50% on PATCHEVAL and SEC-bench, respectively, and surpasses IntentFix by 70.86% and 50.50%. Across both benchmarks, EvoRepair also exceeds the recent self-evolving agent Live-SWE-Agent by 6.98% overall. Additional transfer experiments on VUL4J further demonstrate the robustness of EvoRepair across models, programming languages, and datasets. These findings demonstrate that experience-based self-evolution substantially strengthens agentic AVR and goes beyond existing self-evolving techniques.
△ Less
Submitted 28 May, 2026;
originally announced May 2026.
-
Inversion of CHASE H$α$ Spectral Line during Solar Flares Based on RADYN Dataset via Deep Learning
Authors:
W. Xu,
Q. Hao,
Z. Zheng,
J. Hong,
J. Hu,
Y. Qiu,
C. Li,
M. D. Ding,
C. Fang
Abstract:
Solar flares represent one of the most intense forms of solar activity. Understanding the evolution of physical parameters in the solar atmosphere during flares is key to studying flare mechanisms and improving prediction capabilities. However, directly measuring quantities such as electron number density, temperature, and plasma velocity remains difficult. Here, we introduce a novel fully connect…
▽ More
Solar flares represent one of the most intense forms of solar activity. Understanding the evolution of physical parameters in the solar atmosphere during flares is key to studying flare mechanisms and improving prediction capabilities. However, directly measuring quantities such as electron number density, temperature, and plasma velocity remains difficult. Here, we introduce a novel fully connected neural network, trained on synthetic data from the Radiative Hydrodynamics Code (RADYN) simulations, to perform rapid inversion of physical parameters from H$α$ spectral profiles. The spectral data were processed to align with the observational resolution of the CHASE satellite, enabling seamless application of the model to real-world observations. Results demonstrate a high degree of consistency with RADYN simulations, achieving low errors under diverse flare conditions. Furthermore, we applied the developed model to analyze CHASE observations of a class X7.1 solar flare on October 1, 2024. The results reveal reasonable spatial and temporal evolution of key parameters throughout different flare phases. This work demonstrates the potential of deep learning techniques for fast and reliable spectral inversion, providing new tools for solar flare diagnostics based on H$α$ data.
△ Less
Submitted 25 May, 2026;
originally announced May 2026.
-
SSDAU: Structured Semantic Data Augmentation for Joint Entity and Relation Extraction
Authors:
Jiawei He,
Mengyu Shi,
Jiawei Liu,
Dong Sun,
Chunrong Fang,
Xikai Yang,
Zhijie Wang,
Lei Ma,
Zhenyu Chen
Abstract:
Joint Entity and Relation Extraction (JERE) is highly sensitive to training data quality, making data augmentation a natural way to improve generalization. However, existing augmentation methods often weaken entity relevance and disrupt semantic structure, limiting their effectiveness for JERE. In this paper, we propose \textbf{Structured Semantic Data Augmentation (SSDAU)}, a method designed to p…
▽ More
Joint Entity and Relation Extraction (JERE) is highly sensitive to training data quality, making data augmentation a natural way to improve generalization. However, existing augmentation methods often weaken entity relevance and disrupt semantic structure, limiting their effectiveness for JERE. In this paper, we propose \textbf{Structured Semantic Data Augmentation (SSDAU)}, a method designed to preserve triple-aware semantic structure during augmentation. SSDAU segments text by entity labels, captures semantic features through context-aware encoding, and restructures entity semantics to generate augmented data. To distinguish semantically similar entities, SSDAU combines contextualized embeddings with traditional similarity scores. To reduce topic inconsistency, we apply BERTopic-based filtering to remove irrelevant augmentations. We evaluate SSDAU on datasets with different annotation types and compare its performance on five representative JERE models against seven popular augmentation baselines. Experiments show that SSDAU generates semantically consistent data, is more robust to ambiguity than non-LLM methods (8.95\% vs. 23.58\% average relative F1 decrease), and significantly outperforms strong alternatives in most settings.
△ Less
Submitted 28 May, 2026; v1 submitted 22 May, 2026;
originally announced May 2026.
-
ThoughtTrace: Understanding User Thoughts in Real-World LLM Interactions
Authors:
Chuanyang Jin,
Binze Li,
Haopeng Xie,
Cathy Mengying Fang,
Tianjian Li,
Shayne Longpre,
Hongxiang Gu,
Maximillian Chen,
Tianmin Shu
Abstract:
Conversational AI has now reached billions of users, yet existing datasets capture only what people say, not what they think. We introduce ThoughtTrace, the first large-scale dataset that pairs real-world multi-turn human--AI conversations with users' self-reported thoughts: their reasons for sending prompts and reactions to assistant responses. ThoughtTrace comprises 1,058 users, 2,155 conversati…
▽ More
Conversational AI has now reached billions of users, yet existing datasets capture only what people say, not what they think. We introduce ThoughtTrace, the first large-scale dataset that pairs real-world multi-turn human--AI conversations with users' self-reported thoughts: their reasons for sending prompts and reactions to assistant responses. ThoughtTrace comprises 1,058 users, 2,155 conversations, 17,058 turns, and 10,174 thought annotations collected across 20 language models. Our analysis shows that ThoughtTrace captures long-horizon, topically diverse interactions, and that thoughts are semantically distinct from messages, difficult for frontier LLMs to infer from context, diverse in content, and tied to conversation stages. We further demonstrate the utility of thoughts for downstream modeling. First, thoughts improve user-behavior prediction as inference-time context. Second, thought-guided rewrites provide fine-grained alignment signals for training personalized assistants. Together, ThoughtTrace establishes user thoughts as a new data modality for studying the cognitive dynamics behind human--AI interaction and provides a foundation for building assistants that better understand and adapt to users' latent goals, preferences, and needs.
△ Less
Submitted 21 May, 2026; v1 submitted 19 May, 2026;
originally announced May 2026.
-
RIDE: Retinex-Informed Decoupling for Exposing Concealed Objects
Authors:
Chunming He,
Rihan Zhang,
Dingming Zhang,
Chengyu Fang,
Longxiang Tang,
Jingjia Feng,
Fengyang Xiao,
Sina Farsiu
Abstract:
Concealed Object Segmentation (COS) encompasses a family of dense-prediction tasks, including camouflaged object detection, polyp segmentation, transparent object detection, and industrial defect inspection, where targets are visually entangled with their surroundings through different physical mechanisms. Existing methods either operate directly on RGB images or employ \emph{heterogeneous} decomp…
▽ More
Concealed Object Segmentation (COS) encompasses a family of dense-prediction tasks, including camouflaged object detection, polyp segmentation, transparent object detection, and industrial defect inspection, where targets are visually entangled with their surroundings through different physical mechanisms. Existing methods either operate directly on RGB images or employ \emph{heterogeneous} decompositions (\eg, Fourier, wavelet) that redistribute spatial evidence across scale/frequency coefficients, making pixel-aligned cues less direct. We introduce a fundamentally different perspective: \textbf{homogeneous image decomposition} via Retinex theory, which factorizes an image into illumination and reflectance components within the \emph{same} spatial domain. Our key insight is that visual entanglement enforces appearance matching in the composite space, but this does \emph{not} necessitate simultaneous matching in both component spaces, a phenomenon we formalize as the \textbf{Discriminability Gap Theorem}. Crucially, we show that across diverse COS sub-tasks, the underlying physical processes systematically anti-correlate illumination and reflectance differences, yielding theoretical guarantees that Retinex decomposition preserves or strictly improves total foreground--background discriminability across the full physical regime, with anti-correlation maximizing the gain. Building on this, we propose \textbf{RIDE} comprising: (i) a Task-Driven Retinex Decomposition module that learns segmentation-optimal factorizations end-to-end; (ii) a Discriminability Gap Attention mechanism that adaptively exploits where decomposition helps; and (iii) a Camouflage-Breaking Contrastive loss operating in reflectance feature space.
△ Less
Submitted 14 May, 2026;
originally announced May 2026.
-
Probing Privacy Leaks in LLM-based Code Generation via Test Generation
Authors:
Yifei Ge,
Zhenpeng Chen,
Weisong Sun,
Yuchen Chen,
Chunrong Fang,
Juan Zhai,
Xiaofang Zhang,
Xia Feng,
Yang Liu,
Zhenyu Chen
Abstract:
The widespread availability of large-scale code datasets has fueled the rapid development of large language models (LLMs) for code-related tasks. These datasets may include sensitive personally identifiable information (PII), which can lead to privacy leakage when LLMs memorize and reproduce it. However, existing privacy-leakage detection methods rely on ad-hoc prompt construction (manually or aut…
▽ More
The widespread availability of large-scale code datasets has fueled the rapid development of large language models (LLMs) for code-related tasks. These datasets may include sensitive personally identifiable information (PII), which can lead to privacy leakage when LLMs memorize and reproduce it. However, existing privacy-leakage detection methods rely on ad-hoc prompt construction (manually or automatically designed). Therefore, they do not adequately approximate the real-world contexts in which PII appears in code corpora, making it difficult to extract realistic privacy leakage. In this paper, we propose a pipeline that simulates practical privacy-related code generation scenarios and adopts a test-driven strategy to elicit the memorized information from the generated test cases. We further introduce an automatically constructed privacy feature library that replaces manual prompt engineering by providing realistic templates and examples to guide test case generation. Large-scale experiments on 5 widely used LLMs show that our pipeline exposes more confirmed privacy leakage, achieving a 2.56 times increase in detected leakage compared to existing baselines.
△ Less
Submitted 14 May, 2026;
originally announced May 2026.