-
COSTA: A Cluster-Centric Paradigm for Annotation-Free Open-Set Semantic Segmentation of Aerial Point Clouds with Domain Shifts
Authors:
Yanghong Lin,
Li Fang,
Tianyu Li,
Shudong Zhou,
Wei Yao
Abstract:
Semantic segmentation of aerial point cloud is trapped in a generalization crisis under distinct domain shifts. While test-time adaptation offers a privacy-preserving and computationally efficient way to adapt pre-trained models to unlabeled target-domain data during inference, existing methods, bound to closed-set label assumptions and non-scalable point-wise segmentation pipelines, still struggl…
▽ More
Semantic segmentation of aerial point cloud is trapped in a generalization crisis under distinct domain shifts. While test-time adaptation offers a privacy-preserving and computationally efficient way to adapt pre-trained models to unlabeled target-domain data during inference, existing methods, bound to closed-set label assumptions and non-scalable point-wise segmentation pipelines, still struggle with semantic shifts. We ask: can we adapt any given pre-trained aerial point cloud segmentation model to a shifted target domain at the inference phase alone, without additional training, while segmenting target-specific categories beyond the source label space on demand? This paper introduces COSTA, which breaks this limitation by shifting from closed-set point-wise adaptation to cluster-centric open-set semantic propagation. Our core discovery is that, once effectively adapted at test time, the rich feature distribution of aerial point clouds can be distilled into a compact set of well-separated semantic centroids that are transferable across label spaces. COSTA leverages this to reformulate open-set semantic segmentation as a cluster-level propagating process: it first bridges the domain gap through proven test-time adaptation, then groups each batch of target-domain points into a small set of semantic clusters based on the similarity distribution in the adapted feature space, and finally propagates high-confidence pseudo labels obtained from an open-vocabulary vision-language model to all points through cluster-level voting. This cluster-centric paradigm enables test-time adaptation of aerial point clouds under significant domain gaps with mixed semantic shifts. With DALES as the source domain, COSTA enables on-demand segmentation across three aerial point cloud benchmarks with distinct domains and heterogeneous category spaces, achieving up to 70.09% mIoU under this new setting.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Dense-core approach to the Brualdi--Hoffman--Turán problem on odd wheels
Authors:
Longfei Fang,
Mingqing Zhai,
Yuhan Zhang
Abstract:
We present a unified presentation of the fixed-size adjacency-spectral extremal problem for odd wheels $W_{2k+1}$, where $k\geq2$ and $W_{2k+1}=K_1\vee C_{2k}$. The exceptional case $W_5$ and the general case $W_{2k+1}$, $k\ge3$, share the same dense-core reduction and edge-spectral stability, but have different rigidity structures. We prove that every $W_5$-free graph of sufficiently large size…
▽ More
We present a unified presentation of the fixed-size adjacency-spectral extremal problem for odd wheels $W_{2k+1}$, where $k\geq2$ and $W_{2k+1}=K_1\vee C_{2k}$. The exceptional case $W_5$ and the general case $W_{2k+1}$, $k\ge3$, share the same dense-core reduction and edge-spectral stability, but have different rigidity structures. We prove that every $W_5$-free graph of sufficiently large size $m$ satisfies $ρ(G)^2-ρ(G)\le m,$ with equality precisely for $K_{n,n}$ with a perfect matching embedded in each part, where $n$ is even and $m=n^2+n$. For any fixed $k\ge3$, every $W_{2k+1}$-free graph of sufficiently large size $m$ satisfies $ρ(G)^2-(k-1)ρ(G)\le m-\binom{k}{2},$ with equality precisely for $K_k\vee qK_1$ when $m=\binom{k}{2}+kq$. Our results completely settle a conjecture proposed by Yu, Li and Peng and, via a distinct approach, further strengthen known results concerning odd cycles, friendship graphs and odd fan graphs for sufficiently large $m.$ The proof combines the edge-spectral stability theorem, residual functions and the dense-core method.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Predictive Memory Localization: Forecasting Selective Intervention Paths from Internal Signals
Authors:
Jinhao Jing,
Tian Zeyu,
Lucas Qingyang Fang,
Zhisheng Chen,
Shuang Chen,
Yuhao Luo,
Qiannian Zhao
Abstract:
Activation steering turns localized representations into control directions, but localization alone does not reveal whether a direction has a selective operating regime. We introduce Predictive Memory Localization (PML), which treats the measured-grid intervention path as the predictive object of memory localization. PML separates random-calibrated target movement from semantic-neighbor and capabi…
▽ More
Activation steering turns localized representations into control directions, but localization alone does not reveal whether a direction has a selective operating regime. We introduce Predictive Memory Localization (PML), which treats the measured-grid intervention path as the predictive object of memory localization. PML separates random-calibrated target movement from semantic-neighbor and capability damage, and compares static localization and supervised geometry with a strength-disjoint low-dose causal response. Our frozen study covers 3,000 records from nine datasets and fourteen domains, yielding 30,000 distinct record-direction-layer paths and 210,000 distinct path-strength evaluations. At layer 7, the geometry-derived RFM/AGOP direction reaches 13.1% target-any and 12.3% clean-any, exceeding random by 3.6 and 3.4 percentage points under a record-paired bootstrap. Across record-, dataset-, and domain-grouped splits, responses at $|α|=0.1$ are the strongest signal for outcomes at disjoint strengths $|α|\in\{0.25,0.5\}$. On held-out records, a predictor-driven selector chooses a coefficient or abstains, improves utility and reduces semantic-neighbor damage relative to a train-tuned fixed-strength policy, and avoids most evaluations in a dense scan. Across three residual-norm-matched base models, learned directions retain selective-path gains and low-dose responses yield 0.801-0.828 record-held-out macro AUROC. PML therefore turns memory localization into a falsifiable forecast of margin-level selective outcomes and a risk-aware intervention decision.
△ Less
Submitted 14 August, 2026; v1 submitted 13 August, 2026;
originally announced August 2026.
-
All-electrical Coherent Control of a Single Rare-earth Spin Qubit
Authors:
Yaowu Liu,
Dasom Choi,
Stefano Reale,
Jeongmin Oh,
Seorhin Choi,
Lei Fang,
We-hyo Soe,
Arzhang Ardavan,
Andreas J. Heinrich,
Soo-hyon Phark,
Fabio Donati
Abstract:
Electrical control of single spin qubits is a major frontier for nanoscale, high-speed, and scalable quantum devices. Yet, extending it to highly shielded rare-earth 4f electrons remains an experimental challenge across solid-state platforms. Here we demonstrate all-electrical coherent control of a single Er electron spin, which is exchange-coupled to a nearby Ti atom. Scanning tunneling microscop…
▽ More
Electrical control of single spin qubits is a major frontier for nanoscale, high-speed, and scalable quantum devices. Yet, extending it to highly shielded rare-earth 4f electrons remains an experimental challenge across solid-state platforms. Here we demonstrate all-electrical coherent control of a single Er electron spin, which is exchange-coupled to a nearby Ti atom. Scanning tunneling microscopy-based electron spin resonance with three-dimensional magnetic-field control enables comprehensive mapping of the resonance and Rabi frequencies, revealing pronounced anisotropies in both the Er g-tensor and the Er-Ti exchange interaction. The electrical modulation of the anisotropic Er-Ti coupling results in an efficient drive of the Er spin, allowing us to achieve near-gigahertz Rabi frequencies - a ten-fold improvement over the present record for rare-earth spin qubits. By establishing anisotropic exchange as a general resource for electrically accessing shielded rare-earth spins, our results open a new route to ultrafast and local control of rare-earth spins in solid-state quantum devices.
△ Less
Submitted 8 August, 2026;
originally announced August 2026.
-
ControlRef: Efficient Layout-Guided Multi-Instance Generation via Anchored 4D-RoPE
Authors:
Yunkai Yang,
Yudong Zhang,
Xinying Chen,
Haoyuan Liang,
Yizhuo Niu,
Jinshuai Cheng,
Kunquan Zhang,
Liziyue Fang,
Weitao Wan,
Runmin Dong
Abstract:
Layout-guided multi-instance generation is essential for controllable image synthesis in Multi-Modal Diffusion Transformers (MM-DiTs). However, integrating this capability into unified architectures remains challenging. Prior frameworks rely on redundant full-resolution canvas padding and Shifted-RoPE to manage multiple reference images. This mechanism drastically inflates computational overhead f…
▽ More
Layout-guided multi-instance generation is essential for controllable image synthesis in Multi-Modal Diffusion Transformers (MM-DiTs). However, integrating this capability into unified architectures remains challenging. Prior frameworks rely on redundant full-resolution canvas padding and Shifted-RoPE to manage multiple reference images. This mechanism drastically inflates computational overhead for sparse layouts and disrupts critical low-frequency RoPE features, creating a severe spatial-frequency compromise that blurs absolute spatial correspondence. To overcome these limitations, we propose ControlRef, a highly efficient and precise multi-instance synthesis framework. ControlRef utilizes a Unified Instance-Layout Control (UILC) attention mask to strictly decouple inter-instance semantic interactions and enforce precise regional binding. To further promote region-level spatial alignment, we introduce Anchored 4D-RoPE, a novel positional encoding mechanism that directly anchors tokens to their absolute geometric centers. By pre-aligning reference images to their corresponding bounding box resolutions, physically anchoring both layout and reference tokens to their absolute geometric centers, and stacking the references along the z-axis, Anchored 4D-RoPE natively preserves spatial priors and mitigates the spatial-frequency compromise without lossy shifting. Extensive experiments demonstrate that ControlRef achieves state-of-the-art visual fidelity and localization accuracy, while concurrently slashing inference latency by over 80% in sparse layouts and reducing memory overhead by 50% in dense scenarios.
△ Less
Submitted 7 August, 2026;
originally announced August 2026.
-
Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence
Authors:
Ying Chen,
Weizhen Li,
Zhe Hu,
Zhenjiang Li,
Rui Jiang,
Zhifeng Gu,
Lihuang Fang,
Jiangping Liu,
Lei Yi,
Jie Chen
Abstract:
Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes the scene and physical state, continually renewing what must be perceived, reasoned about, and verified. Meeting these demands requires complementary capabilities that differ in supervision signals, prediction formats, and verification criteria. Ex…
▽ More
Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes the scene and physical state, continually renewing what must be perceived, reasoned about, and verified. Meeting these demands requires complementary capabilities that differ in supervision signals, prediction formats, and verification criteria. Existing approaches typically develop these capabilities against isolated, task-specific objectives, leaving open how they should be organized and integrated around execution as a whole. We present Capek 0.5, an embodied vision-language model built around an execution-centric capability taxonomy. Rather than organizing training by datasets or tasks, the taxonomy groups embodied capabilities according to their functional roles throughout execution and comprises four capability families: Spatial Reasoning, Temporal Understanding, Action Guidance, and State Verification. Each capability is first acquired by a dedicated specialist through reinforcement learning with verifiable rewards from a shared backbone, and the specialists are then consolidated into a single inference-time model through weight-space merging followed by routed policy-space distillation. We instantiate Capek 0.5 at the 2B and 35B-A3B scales and evaluate it from three complementary perspectives: comprehensive benchmark suites including Capek-StateBench, a new benchmark for state verification; a controlled study of capability retention from specialists to the unified model; and closed-loop evaluation in simulated embodied environments. Capek 0.5 improves the large majority of matched benchmark rows over its initialization, retains all four specialized capabilities in one checkpoint with quantified losses, and transfers to closed-loop embodied task execution.
△ Less
Submitted 6 August, 2026;
originally announced August 2026.
-
Clique supersaturation under a chromatic constraint below the Turán threshold
Authors:
Benju Wang,
Longfei Fang,
Jinlong Shu
Abstract:
A central theme in extremal graph theory is the supersaturation problem, which investigates the minimum number of copies of a target subgraph forced by prescribed edge conditions. This line of research goes back to Rademacher and Erdős for triangles, and was later extended to cliques by Lovász and Simonovits in the regime above the Turán threshold. Mubayi further extended this theory to color-crit…
▽ More
A central theme in extremal graph theory is the supersaturation problem, which investigates the minimum number of copies of a target subgraph forced by prescribed edge conditions. This line of research goes back to Rademacher and Erdős for triangles, and was later extended to cliques by Lovász and Simonovits in the regime above the Turán threshold. Mubayi further extended this theory to color-critical graphs. Below the Turán threshold, a closely related existence-threshold phenomenon arises in the non-$p$-partite setting: a classical result of Brouwer shows that, for $n\ge 2p+1$, every $n$-vertex non-$p$-partite $K_{p+1}$-free graph has at most $e(T_{n,p})-\lfloor n/p\rfloor+1$ edges.
Motivated by this threshold, we investigate a sharp clique-counting problem below the Turán threshold under the non-$p$-partite assumption. Let $p\ge 2$ and $s\ge 1$ be fixed integers. Let $Y_{n,p,s}$ be the graph obtained from $T_{n,p}$ by adding an edge inside a largest part and deleting all but $s$ of the edges from one endpoint of this new edge to a smallest part. Then $e(Y_{n,p,s})=e(T_{n,p})-\lfloor n/p\rfloor+s+1$. We prove that, for all sufficiently large $n$, every $n$-vertex non-$p$-partite graph $G$ with $e(G)\ge e(Y_{n,p,s})$ contains at least as many copies of $K_{p+1}$ as $Y_{n,p,s}$ does. The bound is sharp, as it is attained by the construction $Y_{n,p,s}$. Thus our result provides the exact clique-counting analogue of Brouwer's threshold for non-$p$-partite $K_{p+1}$-free graphs.
△ Less
Submitted 6 August, 2026;
originally announced August 2026.
-
Edge-spectral supersaturation for tripartite color-critical graphs
Authors:
Longfei Fang,
Huiqiu Lin,
Mingqing Zhai
Abstract:
We study edge-spectral supersaturation for two families of color-critical graphs with chromatic number three. For an integer $r\geq 1$, we define the spectral threshold \[ g_r(m):=\frac{r-1+\sqrt{4m-r^2+1}}{2}, \] which is the tight upper bound on the spectral radius of graphs avoiding $K_{s,t}^+$ (when $t+1\geq s\geq 3$) and $C_{2k+1}$ (when $r=k$), realized by split-graph constructions. First, l…
▽ More
We study edge-spectral supersaturation for two families of color-critical graphs with chromatic number three. For an integer $r\geq 1$, we define the spectral threshold \[ g_r(m):=\frac{r-1+\sqrt{4m-r^2+1}}{2}, \] which is the tight upper bound on the spectral radius of graphs avoiding $K_{s,t}^+$ (when $t+1\geq s\geq 3$) and $C_{2k+1}$ (when $r=k$), realized by split-graph constructions. First, let $t+1 \geq s\geq 3$ be fixed integers, and let $K_{s,t}^{+}$ be obtained by adding an edge to the part of size $s$ in $K_{s,t}$. We prove that every sufficiently large $m$-edge graph $G$ with $ρ(G)>g_{s-1}(m)$ contains $Ω(m^{(s+t-1)/2})$ copies of $K_{s,t}^{+}$. Second, for any fixed $k\geq 2$, the condition $ρ(G)>g_k(m)$ forces $N(C_{2k+1},G)=Ω(m^k).$ We also construct graphs showing that both lower bounds are tight up to constant factors. These results establish that exceeding the tight spectral Turán threshold $g_r(m)$ forces not just a single copy, but the optimal polynomial number of copies of these color-critical graphs. Thus, crossing the relevant split-graph spectral threshold forces the optimal polynomial order of copies, extending edge-spectral existence theorems to supersaturation results in the delicate three-chromatic regime.
△ Less
Submitted 11 August, 2026; v1 submitted 5 August, 2026;
originally announced August 2026.
-
Exploring Efficient Waveform Diffusion Models for Foley Sound Generation
Authors:
Runwu Shi,
Chang Li,
Jiahui Li,
Jiang Wang,
Yaozhong Kang,
Nabeela Khan,
Linghan Fang,
Benjamin Yen,
Takeshi Ashizawa,
Kazuhiro Nakadai
Abstract:
Recent advances in diffusion models have enabled high-fidelity Foley sound generation directly in the waveform space. Existing waveform diffusion models primarily rely on time-domain architectures, such as CNN-based U-Nets and DiffWave-style models, or frequency-domain Transformers modeling temporal dependencies. However, these systems are typically built with large model capacities and substantia…
▽ More
Recent advances in diffusion models have enabled high-fidelity Foley sound generation directly in the waveform space. Existing waveform diffusion models primarily rely on time-domain architectures, such as CNN-based U-Nets and DiffWave-style models, or frequency-domain Transformers modeling temporal dependencies. However, these systems are typically built with large model capacities and substantial computational costs, leaving compact and efficient waveform diffusion architectures largely underexplored. In this work, we introduce a Dual-Path (DP) architecture for waveform diffusion that performs dimension-wise self-attention along both subband and frame axes in the time-frequency domain. This DP design enables fine-grained temporal-spectral modeling while maintaining high efficiency. Based on the proposed DP backbone, we develop two variants: DP-DiT and DP-U-Net. Experiments on the DCASE and FSD-Kaggle2018 datasets demonstrate their superior performance. Notably, the 3M parameter variant achieves performance comparable to models with more than 50M parameters. Audio samples are available at https://samplesdemo.github.io/DP-Foley/.
△ Less
Submitted 31 July, 2026;
originally announced July 2026.
-
TransX: Scaling Transformer-based Recommendation via Behavioral and Serving Stream Crossings
Authors:
Da Xu,
Liyan Fang,
Divya Venugopalan,
Sunny Hsu,
Xukai Wang,
Rishav Roy Chowdhury,
Cindy Liang,
Nishant Satya Lakshmikanth
Abstract:
Modern industrial recommender systems (RecSys) increasingly adopt Transformer-based sequence models, with an emerging paradigm that frames recommendation as next-token prediction over a unified monolithic user sequence. However, collapsing heterogeneous data sources -- such as long-term user behaviors and real-time serving events -- into a single monolithic token stream that obscures their distinc…
▽ More
Modern industrial recommender systems (RecSys) increasingly adopt Transformer-based sequence models, with an emerging paradigm that frames recommendation as next-token prediction over a unified monolithic user sequence. However, collapsing heterogeneous data sources -- such as long-term user behaviors and real-time serving events -- into a single monolithic token stream that obscures their distinct causal roles and temporal characteristics, leading to inefficient modeling and elevated training and serving costs. We propose TransX, a production-oriented encoder-decoder architecture that reformulates recommendation as a sequence-to-sequence action transduction problem. TransX explicitly decouples behavior-stream modeling from serving-event modeling and conditions next-action decoding on scalable cross-attention between nearline behavior encodings and real-time serving representations. To enable low-latency, high-QPS deployment, TransX is co-designed with an amortized serving strategy that combines incremental behavior encoding with per-request key-value caching, rendering serving latency insensitive to behavior sequence length. Extensive offline experiments and large-scale online A/B tests on LinkedIn's recommender systems show that TransX consistently outperforms state-of-the-art DLRMs and sequential baselines, and delivers substantial CTR lift (+6.0%) and conversion gain (+4.4%) while maintaining serving costs comparable to existing production models where our co-designed serving strategy reduces online computation by approximately 80%.
△ Less
Submitted 30 July, 2026;
originally announced July 2026.
-
Stacking-dependent anisotropic altermagnetism in V$_{1/3}$NbS$_2$
Authors:
Chris J. Lygouras,
Nathan Prouse,
Jack H. Drouin,
Youzhe Chen,
Laura Garcia-Gassull,
Zili Feng,
Mingxuan Fu,
Lü Fang,
Alexander I. Kolesnikov,
Christina Hoffman,
Yiqing Hao,
Huibo Cao,
Maxime A. Siegler,
Robert J. Birgeneau,
Roser Valentí,
Satoru Nakatsuji,
Collin L. Broholm
Abstract:
We report profound impacts of the stacking sequence of triangular lattices of magnetic transition metal ions intercalated between the layers of the van der Waals material NbS$_2$. Using single crystal x-ray and neutron diffraction, and transport and magnetization measurements, we show there are two distinct polytypes of $\rm V_{1/3}NbS_2$ with disparate easy axes of magnetization and different ano…
▽ More
We report profound impacts of the stacking sequence of triangular lattices of magnetic transition metal ions intercalated between the layers of the van der Waals material NbS$_2$. Using single crystal x-ray and neutron diffraction, and transport and magnetization measurements, we show there are two distinct polytypes of $\rm V_{1/3}NbS_2$ with disparate easy axes of magnetization and different anomalous Hall responses. Self-consistent analysis of inelastic neutron scattering data provides evidence for oscillatory RKKY interactions that extend to 1 nm and stabilize quasi-collinear A-type altermagnetic orders in both polytypes though with perpendicular easy axes. The detailed stacking sequence of a bulk polytype crystal dramatically impact its macroscopic anomalous Hall response and magnetism, which suggests a new path to engineer the bulk properties of a layered three dimensional solid.
△ Less
Submitted 27 July, 2026;
originally announced July 2026.
-
DeFiScreener: Efficient DeFi Attack Pre-screening in Smart Contracts via Historical Case Matching
Authors:
Rui Cao,
Shaojing Fan,
Zhimei Sui,
Liming Fang,
Ziqi Yang,
Yingying Jiao,
Zhenguang Liu
Abstract:
Blockchain and its killer applications, particularly decentralized finance (DeFi), are gaining widespread adoption, with over 5,200 DeFi projects deployed on mainstream blockchains as of January 2026. At the same time, security risks in DeFi are becoming increasingly serious. However, existing DeFi detection tools usually cover only specific attack types, exhibiting severely limited detection cove…
▽ More
Blockchain and its killer applications, particularly decentralized finance (DeFi), are gaining widespread adoption, with over 5,200 DeFi projects deployed on mainstream blockchains as of January 2026. At the same time, security risks in DeFi are becoming increasingly serious. However, existing DeFi detection tools usually cover only specific attack types, exhibiting severely limited detection coverage.
In this paper, we argue that an effective way to address this gap is to pre-screen vulnerable instances from large volumes of smart contract functions and call sequences. This is motivated by a key phenomenon we term "perilous temporal asymmetry". Inspired by this, we propose DeFiScreener, the first automated pre-screening framework for DeFi attacks that uses historical exploit cases to identify potentially vulnerable functions and call sequences. Given the full source code of a target project, DeFiScreener builds Function Call Trees (FCTs) and generates semantic embeddings for each function using a large language model (LLM), allowing both program structure and function intent to be analyzed together. It then applies a dual-level screening process. At the function level, function embeddings are matched against an Attack Pattern Library of historically exploited functions. At the sequence level, the proposed Attack Pattern Oriented Monte Carlo Tree Search (APO-MCTS) efficiently explores the FCTs and screens vulnerable call sequences. The identified candidates are ultimately passed to an LLM for further interpretive and security analysis.
We empirically evaluate the DeFiScreener over datasets comprising 207 real-world DeFi attack incidents. Experimental results demonstrate that DeFiScreener achieves a remarkable 98.55% recall and 84.30% precision in attack pre-screening.
△ Less
Submitted 24 July, 2026;
originally announced July 2026.
-
EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization
Authors:
Lihuang Fang,
Yuchen Zou,
kebing Jin,
Jinghui Qin
Abstract:
Multimodal large language models (MLLMs) have achieved impressive performance in multimodal emotion recognition (MER) tasks and lifted MER to a new level that is complex emotion understanding with advanced video understanding abilities and natural language description. However, existing MLLM-based methods often use a fixed prompt to perceive the emotions, ignoring the dynamicity and complexity of…
▽ More
Multimodal large language models (MLLMs) have achieved impressive performance in multimodal emotion recognition (MER) tasks and lifted MER to a new level that is complex emotion understanding with advanced video understanding abilities and natural language description. However, existing MLLM-based methods often use a fixed prompt to perceive the emotions, ignoring the dynamicity and complexity of the emotion source in the multimodal inputs. To address these issues, we propose a novel Reinforcement Learning-based Dynamic Agent Specialization framework (\textbf{EmoAgent-R1}) to optimize the emotion recognition, reasoning, and generalization abilities of an MLLM with dynamic agent specialization based on reinforcement learning. Specifically, we first adopt a cold start strategy to endow an MLLM with preliminary emotion recognition, reasoning, and agent routing ability by training with synthetic answer-conditioned chain-of-thought data and agent routing data. Then, we further train the MLLM with reinforcement learning to perceive emotions in a two-step agentic workflow with agent selection and agent specialization. To effectively train EmoAgent-R1, we propose a novel Progressive Group-Relative Policy Optimization (P-GRPO) to combine group-based relative advantages with a PMI-inspired progressive token-level modulation to transform sparse rewards into fine-grained learning signals, mitigating the coarse-grained uniform credit assignment issue in GRPO. Extensive experiments on MER benchmarks demonstrate the superiority of our EmoAgent-R1 in stronger emotion reasoning performance and improved optimization stability.
△ Less
Submitted 25 July, 2026; v1 submitted 23 July, 2026;
originally announced July 2026.
-
Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning
Authors:
Lizhe Fang,
Weizhou Shen,
Tianyi Tang,
Yisen Wang
Abstract:
Large language models that generate step-by-step reasoning traces have achieved strong performance on complex tasks, and extending them to long-context settings has emerged as an important frontier. However, we identify a critical failure mode in this regime: \emph{repetitive copying}, where models extensively copy text from the input into their reasoning traces rather than productively solving th…
▽ More
Large language models that generate step-by-step reasoning traces have achieved strong performance on complex tasks, and extending them to long-context settings has emerged as an important frontier. However, we identify a critical failure mode in this regime: \emph{repetitive copying}, where models extensively copy text from the input into their reasoning traces rather than productively solving the problem. We show that this behavior is pervasive across frontier long-context LLMs and intensifies with context length. By separating each prompt into task-relevant key evidence and irrelevant distractor context, we further show that the root cause is insufficient grounding: models copy from the prompt indiscriminately, and those that fail to focus on key evidence are far more likely to answer incorrectly. Motivated by this diagnosis, we propose GEAR (Grounding Evidence-Aware Reward), a reward shaping method that augments the accuracy signal with a grounding reward for overlap with key evidence and a distractor penalty for overlap with irrelevant context. To enable GEAR on natural-language data, we develop an automated pipeline that constructs evidence-annotated training data from arbitrary documents. We validate GEAR across multiple model scales and benchmarks, showing consistent improvements of up to +4.6 average points over standard RL with accuracy-based rewards, with larger gains at longer contexts, while also reducing repetitive copying and thinking length. Our findings suggest that, even as long-context evaluation shifts from simple retrieval toward complex reasoning, accurate grounding in relevant evidence remains an indispensable capability with substantial room for improvement.
△ Less
Submitted 31 July, 2026; v1 submitted 21 July, 2026;
originally announced July 2026.
-
Spectral Radius Conditions for 3-Uniform Intersecting Families
Authors:
Lusheng Fang,
Guorong Gao,
An Chang
Abstract:
Let $M_k$ denote a matching of size $k$. The classical Erdős matching conjecture asks for the maximum number of edges of an intersecting $r$-graph without $M_k$. The csae for $k=2$, which is known as intersecting $r$-graph, is established by Erdős, Ko and Rado. Hilton and Milner further determine the maximum number of edges of a non-trivial intersecting $r$-graph, where the intersecting $r$-graph…
▽ More
Let $M_k$ denote a matching of size $k$. The classical Erdős matching conjecture asks for the maximum number of edges of an intersecting $r$-graph without $M_k$. The csae for $k=2$, which is known as intersecting $r$-graph, is established by Erdős, Ko and Rado. Hilton and Milner further determine the maximum number of edges of a non-trivial intersecting $r$-graph, where the intersecting $r$-graph $H$ is called non-trivial if $\cap_{e\in E(H)}e=\emptyset$. In this paper, we investigate the spectral analogues of the hpergraph matching problems and intersecting family problems. More precisely, for sufficiently large $n$, we determine respectively the maximum spectral radius of $M_{k+1}$-free and non-trivial intersecting $3$-graphs on $n$ vertices, and characterize the extremal hypergraphs.
△ Less
Submitted 9 July, 2026;
originally announced July 2026.
-
Distributionally Robust Optimization via Targeted Integral Probability Metrics for General Data Processes
Authors:
Lanran Fang,
Jianqiang Cheng,
Grani A. Hanasusanto,
Yijie Wang
Abstract:
Distributionally robust optimization (DRO) provides a principled framework for decision-making under distributional uncertainty. Classical data-driven DRO frameworks typically construct ambiguity sets from distributional information, such as moment constraints, divergence neighborhoods, or Wasserstein balls, specified before the downstream loss is considered. We propose a task-aware DRO framework…
▽ More
Distributionally robust optimization (DRO) provides a principled framework for decision-making under distributional uncertainty. Classical data-driven DRO frameworks typically construct ambiguity sets from distributional information, such as moment constraints, divergence neighborhoods, or Wasserstein balls, specified before the downstream loss is considered. We propose a task-aware DRO framework based on targeted integral probability metrics. The ambiguity set is defined directly through the loss functions induced by feasible decisions, thereby controlling the loss discrepancy between an adversarial distribution and a data-driven reference distribution. This construction leads to an expected hinge-constrained formulation that is equivalent to an infinitely constrained loss-discrepancy formulation. It also yields finite-sample guarantees that bypass the ambient curse of dimensionality: whenever an appropriate scalar pointwise concentration inequality is available for the induced loss estimator, the ambiguity radius can be calibrated at the canonical $\widetilde{\mathcal O}(N^{-1/2})$ rate after uniformization over the decision class. As a result, the framework applies broadly to settings including heavier-tailed sub-Weibull losses, Markovian data, outlier-corrupted data, and incomplete data. We derive exact infinite-dimensional dual reformulations, establish out-of-sample and excess-risk guarantees, and develop a conservative Monte Carlo approximation scheme with convergence and suboptimality guarantees. For piecewise affine losses, the sampled problems admit tractable conic reformulations. Numerical experiments in inventory management under heavy-tailed demand and regression with outlier corruption demonstrate strong out-of-sample performance relative to existing approaches.
△ Less
Submitted 6 July, 2026;
originally announced July 2026.
-
Two-way coupling in active suspensions suppresses particle accumulation and induces non-monotonic flow stabilization
Authors:
Miyi Wu,
Ziyue Yu,
Lei Fang
Abstract:
We experimentally study the two-way coupling between a swarm of centimetre-scale active swimmers (Artemia salina) and an electromagnetically driven quasi-two-dimensional cellular flow. The swimmer loading $N$ and the background forcing $E$ are varied independently across 85 conditions, and the coupled dynamics are characterized through Lagrangian diagnostics built on the attracting Lagrangian cohe…
▽ More
We experimentally study the two-way coupling between a swarm of centimetre-scale active swimmers (Artemia salina) and an electromagnetically driven quasi-two-dimensional cellular flow. The swimmer loading $N$ and the background forcing $E$ are varied independently across 85 conditions, and the coupled dynamics are characterized through Lagrangian diagnostics built on the attracting Lagrangian coherent structures (LCS) of the flow. In the dilute limit, swimmers accumulate onto the attracting LCS most strongly when their speed is comparable to the flow speed, recovering the mobility-selective accumulation predicted by one-way-coupled simulations at the fixed aspect ratio of A. salina. As the loading increases, this accumulation is progressively suppressed, a collective effect inaccessible to single-swimmer models. The back-action of the swarm on the flow is itself bidirectional and regime-dependent: at low forcing the swarm disorders the attracting-LCS web and scatters the topological critical points of the cellular pattern, whereas at high forcing it reorganizes and reinforces the skeleton. The temporal stability of the skeleton is correspondingly non-monotonic in both $N$ and $E$, with the conditions of maximal stability migrating systematically through the $(N,E)$ plane. This reinforcement of coherent structures has no counterpart in prior simulations, which reported a predominantly disruptive back-action. Resolving both directions of the coupling shows how collective activity can either erode or reinforce the transport skeleton of a structured flow.
△ Less
Submitted 10 August, 2026; v1 submitted 2 July, 2026;
originally announced July 2026.
-
DANTE-W: Diffuse Albedo Neural Texturing in the Wild
Authors:
Guangyu Wang,
Tianheng Lu,
Ruqi Huang,
Lu Fang
Abstract:
Classical mesh texturing techniques blend captured multi-view images directly, which inevitably suffer from baked-in shading and casted shadows that compromise visual fidelity during relighting. To circumvent this issue, we present a neural texturing framework, namely DANTE-W, to enable high-fidelity diffuse albedo texture recovery from unstructured image collections for large-scale, in-the-wild s…
▽ More
Classical mesh texturing techniques blend captured multi-view images directly, which inevitably suffer from baked-in shading and casted shadows that compromise visual fidelity during relighting. To circumvent this issue, we present a neural texturing framework, namely DANTE-W, to enable high-fidelity diffuse albedo texture recovery from unstructured image collections for large-scale, in-the-wild scenes, which integrates seamlessly with traditional 3D reconstruction pipelines. Given a reconstructed mesh and its surface parameterization, our method fuses view-space generative albedo priors into a coherent texture space via an expressive neural representation, while substantially enhancing fine-grained textural details through physically principled neural rendering. To comprehensively evaluate our method, we curate a benchmark dataset featuring diverse, fine-grained textures, comprising both real-world in-the-wild scenes and synthetic objects. Extensive experiments verify the effectiveness of our approach in reconstructing accurate albedo textures and boosting relighting fidelity. Project page: dante-wild.github.io.
△ Less
Submitted 26 June, 2026;
originally announced June 2026.
-
NOVA: A Verification-Aware Agent Harness for Architecture Evolution in Industrial Recommender Systems
Authors:
Shaohua Liu,
Liang Fang,
Yilong Sun,
Shudong Huang,
Qingsong Luo,
Shaoxin Liu,
Xiaoyang Chen,
Dongqiang Liu,
Chuangang Ma,
Zhenzhen Chai,
Henghuan Wang,
Shijie Quan,
Changyuan Cui,
Zhangbin Zhu,
Peng Chen,
Wei Xu,
Lei Xiao,
Haijie Gu,
Jie Jiang
Abstract:
Industrial advertising recommender systems are continually improved through architecture modifications, yet production iteration remains expert-intensive because coordinated changes to model topology, feature configuration, and interaction modules must satisfy strict interface, resource, and serving constraints. AutoML is limited to predefined search spaces, while generic coding agents verify runn…
▽ More
Industrial advertising recommender systems are continually improved through architecture modifications, yet production iteration remains expert-intensive because coordinated changes to model topology, feature configuration, and interaction modules must satisfy strict interface, resource, and serving constraints. AutoML is limited to predefined search spaces, while generic coding agents verify runnability rather than recommender-specific semantic validity. Executable candidates may therefore violate architectural contracts, while the lack of structured reuse of semantic diagnostics and evaluation outcomes can lead to repeated invalid or ineffective modifications.
We present NOVA, a verification-aware agent harness that organizes production architecture modification as multi-round search over concrete implementations within a fixed evaluation budget. At each round, NOVA generates multiple candidates under production constraints, rejects semantic violations, and ranks the valid survivors for local testing and offline evaluation. Across rounds, trajectory memory synthesizes semantic diagnostics, local-test outcomes, and offline metric changes into modification directions and forbidden patterns that guide subsequent search. Under the same maximum offline-evaluation budget for automated methods, NOVA achieves the highest effective pass rate, reaching 53.3% on ScaleUp and 51.7% on Literature-to-Production tasks. In a production A/B test covering 5% of traffic in an advertising system serving over one billion users, the selected Literature-to-Production candidate yields GMV gains of +1.25%, +1.70%, and +2.02% across three major pCVR objectives, with corresponding relative reductions in absolute pCVR bias of 58.8%, 66.7%, and 37.3%, respectively.
△ Less
Submitted 28 July, 2026; v1 submitted 25 June, 2026;
originally announced June 2026.
-
Knowledge Cascade: Reverse Knowledge Distillation on Nonparametric Multivariate Functional Estimation
Authors:
Luyang Fang,
Haoran Lu,
Yongkai Chen,
Wenxuan Zhong,
Ping Ma
Abstract:
As machine learning models and datasets continue to grow, developing complex models has become increasingly computationally demanding. Knowledge distillation reduces deployment cost by compressing a large, well-trained teacher model into a compact student model, but it does not address settings where constructing the teacher itself is the bottleneck. Motivated by this challenge, we introduce Knowl…
▽ More
As machine learning models and datasets continue to grow, developing complex models has become increasingly computationally demanding. Knowledge distillation reduces deployment cost by compressing a large, well-trained teacher model into a compact student model, but it does not address settings where constructing the teacher itself is the bottleneck. Motivated by this challenge, we introduce Knowledge Cascade (KCas), a reverse knowledge distillation framework that uses information from a small, inexpensive student model to guide the development of a more complex teacher model. Although this direction is counterintuitive because the teacher typically has greater representational capacity, we show that student-to-teacher transfer can be principled when supported by statistical scaling relationships. We first develop KCas for nonparametric multivariate functional estimation in reproducing kernel Hilbert spaces via smoothing splines, where selecting multiple smoothing parameters is a major computational bottleneck. KCas transfers student-selected smoothing parameters to the full-sample regime through asymptotic scaling laws, substantially reducing computational cost for high-dimensional and large-scale datasets while retaining theoretical guarantees. Beyond smoothing splines, we illustrate the same principle through kernel density estimation and deep learning hyperparameter transfer. Simulations and real-data experiments show that KCas achieves substantial computational savings while maintaining strong statistical performance, and can sometimes outperform the corresponding full-sample procedure.
△ Less
Submitted 24 June, 2026;
originally announced June 2026.
-
Decidability and Undecidability Results for LIA-Definable Impartial Combinatorial Games
Authors:
Shiguang Feng,
Liangda Fang,
Jiahao Luo,
Quanlong Guan
Abstract:
Combinatorial game theory is a branch of mathematics and theoretical computer science that studies deterministic games with perfect information and no elements of chance. The majority of combinatorial games are impartial and formalized in linear integer arithmetic, which we call LIA-definable impartial combinatorial games (ICGs). This paper studies decidability and undecidability questions for the…
▽ More
Combinatorial game theory is a branch of mathematics and theoretical computer science that studies deterministic games with perfect information and no elements of chance. The majority of combinatorial games are impartial and formalized in linear integer arithmetic, which we call LIA-definable impartial combinatorial games (ICGs). This paper studies decidability and undecidability questions for these games. We prove that deciding whether an LIA-definable ICG is terminating or cyclic is undecidable in general, while the corresponding questions become decidable for terminating LIA-definable ICGs. We also show that deciding whether an LIA formula exactly characterizes the set of winning, losing, or draw states of an LIA-definable ICG is undecidable in general and decidable for terminating LIA-definable ICGs. For state-level questions, deciding whether a state is winning or losing remains undecidable even for terminating LIA-definable ICGs, but becomes decidable for terminating finite-depth LIA-definable ICGs. Finally, deciding whether a state is draw for an LIA-definable ICG is undecidable in general and decidable for terminating LIA-definable ICGs.
△ Less
Submitted 23 June, 2026;
originally announced June 2026.
-
REBA: A Revealed Belief Automaton Framework for Online Planning in Continuous POMDPs
Authors:
Xiangwei Chen,
Lingling Fang,
Andreas Holzinger,
Liming Chen
Abstract:
Online planning in continuous partially observable Markov decision processes (POMDPs) using $ω$-regular specifications requires handling continuous belief dynamics within the finite symbolic memory in order to track temporal progress. Existing methods based on either direct search in belief space or predefined discrete abstractions suffer from drawbacks, e.g., lack of symbolic memory for long-hori…
▽ More
Online planning in continuous partially observable Markov decision processes (POMDPs) using $ω$-regular specifications requires handling continuous belief dynamics within the finite symbolic memory in order to track temporal progress. Existing methods based on either direct search in belief space or predefined discrete abstractions suffer from drawbacks, e.g., lack of symbolic memory for long-horizon logical progress or difficult to certify from noisy online beliefs. As such, obtaining reliable symbolic states online from continuous observations remains a challenge. To address this issue, we introduce the Revealed Belief Automaton (REBA), an event-driven framework that advances the research from global belief-space discretization to a fundamental new way of thinking, namely online certification of revelation events. Specifically, we propose an online revelation method that, through information-theoretic gates, can dynamically analyse and establish belief abstraction from the continuous belief space by discovering reliable anchors among noisy beliefs. We then develop an incremental topology adaptation mechanism over the certified anchors to realise the online finite Belief Automaton. By combining with the $ω$-regular specification, REBA is able to support formal parity policy synthesis without a predefined discrete abstraction, which in turn can guide the Monte Carlo Tree Search process to perform online search beyond its local horizon. In addition, we design an error decomposition analysis which can assess the effectiveness and reliability of this discrete guidance for the underlying continuous POMDP. Empirical evaluations in patrolling and navigation scenarios show that REBA matches or exceeds all evaluated baselines, with primary metric gains of +17.0\% to +47.4\% over state-of-the-art approaches.
△ Less
Submitted 20 June, 2026;
originally announced June 2026.
-
FLM-Occ: Feed-forward Likelihood Maximization for Efficient Indoor Occupancy Prediction
Authors:
Guangcheng Chen,
Lihuang Fang,
Huaqi Tao,
Yicheng He,
Li He,
Hong Zhang
Abstract:
Recent indoor occupancy prediction methods adopt Gaussian primitives as a sparse 3D representation for computational efficiency. However, their training relies on voxel classification, which imposes only local constraints and lacks global supervision on the distribution of the primitives. Therefore, they inevitably predict spurious primitives in empty regions, undermining both representational and…
▽ More
Recent indoor occupancy prediction methods adopt Gaussian primitives as a sparse 3D representation for computational efficiency. However, their training relies on voxel classification, which imposes only local constraints and lacks global supervision on the distribution of the primitives. Therefore, they inevitably predict spurious primitives in empty regions, undermining both representational and computational efficiency. To address this, we propose Feed-forward Likelihood Maximization (FLM), a novel framework that reformulates occupancy prediction as voxel distribution estimation. In FLM, a network is trained to predict a mixture model that maximizes the likelihood over ground-truth occupied voxels in a feed-forward manner. To enable end-to-end training of networks and voxelization of a standard mixture model, we define mixture weights as normalized primitive volumes to implicitly enforce simplex constraints and derive novel voxelization formulas. Based on FLM, our FLM-Occ, a novel method that is capable of relocating randomly initialized primitives over long distances to model a scene. On Occ-ScanNet, FLM-Occ achieves superior accuracy using only 32 superquadrics, 2.7% of the prior SoTA, while running 3.7 times faster.
△ Less
Submitted 19 June, 2026;
originally announced June 2026.
-
Confidence-Aware Automated Assessment of Student-Drawn Scientific Models
Authors:
Luyang Fang,
Yingchuan Zhang,
Jongchan Park,
Zhaoji Wang,
Ping Ma,
Xiaoming Zhai
Abstract:
Student-generated drawings are widely used in science education to assess learners' conceptual understanding in modeling-based tasks aligned with the Next Generation Science Standards (NGSS). However, scoring such drawings requires expert human judgment to interpret complex visual representations, making large-scale assessment costly to implement and sustain in classroom settings. In this work, we…
▽ More
Student-generated drawings are widely used in science education to assess learners' conceptual understanding in modeling-based tasks aligned with the Next Generation Science Standards (NGSS). However, scoring such drawings requires expert human judgment to interpret complex visual representations, making large-scale assessment costly to implement and sustain in classroom settings. In this work, we study automated scoring of student-generated scientific drawings using a vision-based model. We evaluate a Vision Transformer (ViT) with parameter-efficient adaptation and propose a confidence-aware scoring framework that derives response-level confidence from test-time predictive distributions. This confidence signal enables selective automation by scoring high-confidence responses automatically while deferring uncertain cases for human review. Experiments on six NGSS-aligned middle school assessment items show that the proposed approach improves scoring reliability while supporting a practical trade-off between automated coverage and scoring risk, highlighting the value of confidence-aware methods for trustworthy educational assessment.
△ Less
Submitted 18 June, 2026;
originally announced June 2026.
-
EyeMVP: OCT-Informed Fundus Representation Learning via Paired CFP--OCT Pretraining
Authors:
Zhuo Deng,
Ruiheng Zhang,
Ziheng Zhang,
Weihao Gao,
Yitong Li,
Qian Wang,
Lei Shao,
Jiaoyue Dong,
Zhixi Zeng,
Lijian Fang,
Haibo Wang,
Xiaobin Lin,
Tao Liu,
Zhicheng Du,
Zhengwei Zhang,
Lin Yang,
Zheng Gong,
Xinyu Zhao,
Zhenquan Wu,
Fang Li,
Zhiguang Zhou,
Guoming Zhang,
Sun Jing,
Han Lv,
Wenbin We
, et al. (1 additional authors not shown)
Abstract:
Color fundus photography (CFP) is the mainstay of large-scale retinal screening, but its diagnostic capacity is limited by the lack of depth-resolved structure, which optical coherence tomography (OCT) provides yet is less accessible at population scale. We present EyeMVP, a cross-modal retinal foundation model that uses paired CFP--OCT pretraining to learn OCT-informed CFP representations while r…
▽ More
Color fundus photography (CFP) is the mainstay of large-scale retinal screening, but its diagnostic capacity is limited by the lack of depth-resolved structure, which optical coherence tomography (OCT) provides yet is less accessible at population scale. We present EyeMVP, a cross-modal retinal foundation model that uses paired CFP--OCT pretraining to learn OCT-informed CFP representations while requiring only CFP at inference. Pretrained on 674,893 same-eye same-day CFP--OCT triples from 112,642 patients across eight hospitals, EyeMVP uses cross-modal masked reconstruction to enrich CFP features with OCT-associated supervision, and combines source-constrained cross-attention with CFP-derived structural masks to accommodate the non-aligned geometry of en-face CFP and cross-sectional OCT. Across 15 dataset-level settings spanning classification and segmentation, under both full-data and few-shot regimes, EyeMVP performs on par with or better than representative retinal foundation models, with consistent gains on macular and optic-nerve tasks; it attains AUROCs of 0.923 for macular edema and 0.867 for myopic macular schisis, two conditions poorly resolved in CFP. In an exploratory reader study, EyeMVP surpasses junior and intermediate ophthalmologists but not seniors on macular edema, while exceeding all groups on myopic macular schisis. These results indicate that cross-modal reconstruction can enrich CFP representations with OCT-associated supervision, offering a practical route to stronger CFP-based screening.
△ Less
Submitted 28 June, 2026; v1 submitted 13 June, 2026;
originally announced June 2026.
-
Violation-Informed Spatio-Temporal Adaptive Targeting Framework for EV-Driven Distribution System Expansion Planning
Authors:
Linhan Fang,
Xingpeng Li
Abstract:
The rapid adoption of electric vehicles (EVs) can cause severe voltage drops and line current overloads in distribution networks, creating an urgent need for scalable expansion planning methods. This paper proposes a computationally efficient violation-informed spatio-temporal adaptive targeting (STAT) framework for EV-driven distribution system expansion planning. The framework first identifies p…
▽ More
The rapid adoption of electric vehicles (EVs) can cause severe voltage drops and line current overloads in distribution networks, creating an urgent need for scalable expansion planning methods. This paper proposes a computationally efficient violation-informed spatio-temporal adaptive targeting (STAT) framework for EV-driven distribution system expansion planning. The framework first identifies potential voltage and current violations through a violation analysis model, and then mitigates them through a joint optimal expansion planning model that co-optimizes investment decisions for line reconductoring, shunt capacitors, and battery energy storage systems. To reduce computational burden, the proposed STAT-temporal criticality assessment (STAT-TCA) method extracts primitive stress events from annual operating data, derives an initial set of candidate planning horizons from signature-consistent segments, and selects a final transferable critical horizon set through cross-horizon validation based on optimization feasibility and cost. Meanwhile, the proposed STAT-adaptive spatial targeting (STAT-AST) method constructs device-specific spatial features for BESS and SC siting to retain compact yet high-impact candidate bus sets. Case studies on 33-bus and 240-bus distribution systems demonstrate that the proposed STAT framework can substantially reduce the temporal and spatial planning dimensions while preserving planning fidelity. Full-year validation further confirms that the resulting investment plans can eliminate EV-induced voltage and thermal violations while maintaining feasible BESS operations.
△ Less
Submitted 10 June, 2026;
originally announced June 2026.
-
Vector Map as Language: Toward Unified Remote Sensing Vector Mapping
Authors:
Yinglong Yan,
Yunkai Yang,
Haoyi Wang,
Wei Fu,
Linshan Wu,
Honghu Pan,
Shaobo Xia,
Shanghang Zhang,
Hao Chen,
Leyuan Fang
Abstract:
Remote sensing vector mapping aims to generate structured maps of geospatial entities, such as buildings, roads, and water bodies, from remote sensing imagery. In practice, vector maps usually contain multiple category layers and heterogeneous entity structures, requiring a unified model for diverse mapping needs. However, existing methods typically represent vector objects as polygons or graphs,…
▽ More
Remote sensing vector mapping aims to generate structured maps of geospatial entities, such as buildings, roads, and water bodies, from remote sensing imagery. In practice, vector maps usually contain multiple category layers and heterogeneous entity structures, requiring a unified model for diverse mapping needs. However, existing methods typically represent vector objects as polygons or graphs, making them suitable only for specific categories: polygons poorly capture topological relations, while graphs often blur instance boundaries. We observe that language, as a natural medium for human communication, offers a flexible and expressive representation that can accommodate heterogeneous map elements, including geometry, semantics, and topolog. Motivated by this insight, we propose Vector Map as Language (VecLang), a unified paradigm that reformulates multiclass vector mapping as structured text generation. VecLang encodes the common elements of different geospatial entities into a GeoJSON-like vector language, enabling cross-category modeling within a shared textual format. To generate this language reliably, we design a progressive vision-language mapping framework that first localizes vectorization units and then generates structured map elements. We further introduce Hierarchical Vector Language Optimization, which uses reinforcement learning to improve syntax validity, content fidelity, and map executability. We also build VecMap-Bench with 54K images and 800K instances, supporting training and evaluation across standard and generalization settings. Extensive experiments demonstrate that VecLang handles both single-class and multiclass vector mapping while achieving strong cross-dataset and open-vocabulary generalization. The model and dataset are publicly available at https://github.com/yyyyll0ss/VecLang.
△ Less
Submitted 9 June, 2026;
originally announced June 2026.
-
CapSenseBand: Sustaining Cross-Disciplinary Creativity When Stitches Must Meet Signals
Authors:
Sark Pangrui Xing,
Hongci Hu,
Lai Wei,
Le Fang,
Ziqian Bai,
Kinor Shou-xiang Jiang,
Stephen Jia Wang
Abstract:
Wearable sensing systems increasingly depend on textiles that are both materially wearable and electronically functional. Their design requires collaboration between textile designers, who reason through stitches, yarn behavior, and machine constraints, and interaction designers, who reason through electrodes, signal paths, and insulation. However, these forms of expertise do not easily translate…
▽ More
Wearable sensing systems increasingly depend on textiles that are both materially wearable and electronically functional. Their design requires collaboration between textile designers, who reason through stitches, yarn behavior, and machine constraints, and interaction designers, who reason through electrodes, signal paths, and insulation. However, these forms of expertise do not easily translate across disciplinary boundaries. This poster presents CapSenseBand, a knitted capacitive-sensing wristband developed through a research-through-design process organized around Analysis, Synthesis, and Detailing. We document an artifact chain spanning material swatches, a rapid wearable prototype, Paper Models as shared negotiation surfaces, a double-layer knitted structure, and an insulated Swept Frequency Capacitive Sensing breakout board. We show how Paper Models functioned as boundary objects, helping collaborators externalize intent, negotiate spatial and technical constraints, and preserve disciplinary expertise while converging on a shared design. We contribute a reusable swatch-to-sleeve pattern for material-centered HCI: keep discipline-specific probes open early, then converge through artifacts that make material, spatial, and electronic decisions legible before fabrication locks them in.
△ Less
Submitted 3 June, 2026;
originally announced June 2026.
-
Hosting Capacity Assessment and Enhancement for Edge Data Centers in Active Distribution Networks
Authors:
Linhan Fang,
Xingpeng Li
Abstract:
With the increasing demand for edge computing and AI-driven workloads, integrating small and medium-sized edge data centers into distribution networks has become increasingly important. This paper investigates the hosting capacity of distribution networks for data center integration and identifies the key physical mechanisms that limit the maximum allowable data center load. The baseline analysis…
▽ More
With the increasing demand for edge computing and AI-driven workloads, integrating small and medium-sized edge data centers into distribution networks has become increasingly important. This paper investigates the hosting capacity of distribution networks for data center integration and identifies the key physical mechanisms that limit the maximum allowable data center load. The baseline analysis shows that data center hosting capacity varies significantly across candidate buses due to network topology and electrical distance. Three dominant limiting mechanisms are identified: current-constrained locations, voltage-constrained locations, and mixed-constrained locations where both current loading and voltage deviation jointly affect hosting capacity. To increase the hosting capacity, this study evaluates multiple flexible resources, including battery energy storage systems (BESS), dispatchable distributed generators (DDG), and static synchronous compensators (STATCOM). Numerical results demonstrate that these resources provide complementary benefits through active power support, sustained local generation, and reactive power compensation, effectively expanding data center hosting capacity in distribution systems.
△ Less
Submitted 31 May, 2026;
originally announced June 2026.
-
Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors
Authors:
Luyang Fang,
Yongkai Chen,
Jiazhang Cai,
Ping Ma,
Wenxuan Zhong
Abstract:
Knowledge distillation is a powerful method for model compression, enabling the efficient deployment of complex deep learning models (teachers), including large language models. However, its underlying statistical mechanisms remain unclear, and uncertainty evaluation is often overlooked, especially in real-world scenarios requiring diverse teacher expertise. To address these challenges, we introdu…
▽ More
Knowledge distillation is a powerful method for model compression, enabling the efficient deployment of complex deep learning models (teachers), including large language models. However, its underlying statistical mechanisms remain unclear, and uncertainty evaluation is often overlooked, especially in real-world scenarios requiring diverse teacher expertise. To address these challenges, we introduce \textit{Multi-Teacher Bayesian Knowledge Distillation} (MT-BKD), where a distilled student model learns from multiple teachers within the Bayesian framework. Our approach leverages Bayesian inference to capture inherent uncertainty in the distillation process. We introduce a teacher-informed prior, integrating external knowledge from teacher models and task-specific training data, offering better generalization, robustness, and scalability. Additionally, an entropy-based weighting mechanism adaptively adjusts each teacher's influence, allowing the student to combine multiple sources of expertise effectively. MT-BKD enhances the interpretability of the student model's learning process, improves predictive accuracy, and provides uncertainty quantification. We validate MT-BKD on both synthetic and real-world tasks, including protein subcellular location prediction and image classification. Our experiments show improved performance and robust uncertainty quantification, highlighting the strengths of our MT-BKD framework.
△ Less
Submitted 27 May, 2026;
originally announced May 2026.
-
HoloFair: Unified T2I Fairness Evaluation and Fair-GRPO Debiasing
Authors:
Ruyi Chen,
Lu Zhou,
Xiaogang Xu,
Chiyu Zhang,
Jiafei Wu,
Liming Fang
Abstract:
Text-to-Image (T2I) models have made significant strides in visual realism and semantic consistency, yet they often perpetuate and amplify societal biases. Existing evaluation methods typically address only single-dimensional biases, lacking perspectives to uncover model biases at social-related deeper semantic levels. We introduce HoloFair, a comprehensive benchmark framework for multidimensional…
▽ More
Text-to-Image (T2I) models have made significant strides in visual realism and semantic consistency, yet they often perpetuate and amplify societal biases. Existing evaluation methods typically address only single-dimensional biases, lacking perspectives to uncover model biases at social-related deeper semantic levels. We introduce HoloFair, a comprehensive benchmark framework for multidimensional demographic bias analysis. Built upon our large-scale fairness-oriented dataset and the SpaFreq (Spatial-Frequency) attribute classifier, this framework proposes the Multi-attribute, Group-wise Bias Index (MGBI) metric, designed to assess both intrinsic diversity and conditional biases. Beyond evaluation, we further introduce Fair-GRPO, a reinforcement-learning-based debiasing method that alters the distribution of generative models through a designed multi-objective reward function. E.g., experiments on the SD3.5-Medium model demonstrate that Fair-GRPO significantly improves multidimensional fairness while maintaining high image quality. We also analyze potential reward hacking phenomena and provide corresponding mitigation strategies. Code and dataset are available at https://github.com/1059684669/HoloFair
△ Less
Submitted 23 May, 2026;
originally announced May 2026.
-
Code as Agent Harness
Authors:
Xuying Ning,
Katherine Tieu,
Dongqi Fu,
Tianxin Wei,
Zihao Li,
Yuanchen Bei,
Jiaru Zou,
Mengting Ai,
Zhining Liu,
Ting-Wei Li,
Lingjie Chen,
Yanjun Zhao,
Ke Yang,
Bingxuan Li,
Cheng Qian,
Gaotang Li,
Xiao Lin,
Zhichen Zeng,
Ruizhong Qiu,
Sirui Chen,
Yifan Sun,
Xiyuan Yang,
Ruida Wang,
Rui Pan,
Chenyuan Yang
, et al. (17 additional authors not shown)
Abstract:
Recent large language models (LLMs) have demonstrated strong capabilities in understanding and generating code, from competitive programming to repository-level software engineering. In emerging agentic systems, code is no longer only a target output. It increasingly serves as an operational substrate for agent reasoning, acting, environment modeling, and execution-based verification. We frame thi…
▽ More
Recent large language models (LLMs) have demonstrated strong capabilities in understanding and generating code, from competitive programming to repository-level software engineering. In emerging agentic systems, code is no longer only a target output. It increasingly serves as an operational substrate for agent reasoning, acting, environment modeling, and execution-based verification. We frame this shift through the lens of agent harnesses and introduce code as agent harness: a unified view that centers code as the basis for agent infrastructure. To systematically study this perspective, we organize the survey around three connected layers. First, we study the harness interface, where code connects agents to reasoning, action, and environment modeling. Second, we examine harness mechanisms: planning, memory, and tool use for long-horizon execution, together with feedback-driven control and optimization that make harness reliable and adaptive. Third, we discuss scaling the harness from single-agent systems to multi-agent settings, where shared code artifacts support multi-agent coordination, review, and verification. Across these layers, we summarize representative methods and practical applications of code as agent harness, spanning coding assistants, GUI/OS automation, embodied agents, scientific discovery, personalization and recommendation, DevOps, and enterprise workflows. We further outline open challenges for harness engineering, including evaluation beyond final task success, verification under incomplete feedback, regression-free harness improvement, consistent shared state across multiple agents, human oversight for safety-critical actions, and extensions to multimodal environments. By centering code as the harness of agentic AI, this survey provides a unified roadmap toward executable, verifiable, and stateful AI agent systems.
△ Less
Submitted 18 May, 2026;
originally announced May 2026.
-
MoASE++: Mixture of Activation Sparsity Experts with Domain-Adaptive On-policy Distillation for Continual Test Time Adaptation
Authors:
Ronyu Zhang,
Aosong Cheng,
Gaole Dai,
Yulin Luo,
Jiaming Liu,
Li Du,
Huanrui Yang,
Dan Wang,
Leyuan Fang,
Yuan Du,
Shanghang Zhang
Abstract:
Continual test-time adaptation adapts a source-pretrained model to non-stationary, unlabeled target streams while retaining past competence, yet texture-biased backbones risk error accumulation and catastrophic forgetting. Drawing inspiration from the process of decoupling shape and texture in the human visual system, we introduce MoASE, a plug-in mixture-of-experts that disentangles domain-agnost…
▽ More
Continual test-time adaptation adapts a source-pretrained model to non-stationary, unlabeled target streams while retaining past competence, yet texture-biased backbones risk error accumulation and catastrophic forgetting. Drawing inspiration from the process of decoupling shape and texture in the human visual system, we introduce MoASE, a plug-in mixture-of-experts that disentangles domain-agnostic structure from domain-specific texture using Activation Sparsity Experts with Spatial Differentiable Dropout, forming complementary high- and low-activation pathways, while high- and low-rank bottlenecks diversify representations. The Activation Sparsity Gate produces input-adaptive SDD thresholds for precise token selection, and the Domain-Aware Router assigns per-sample expert weights using texture-sensitive cues. To curb confirmation bias on unlabeled streams and stabilize supervision, we then introduce Domain-Adaptive On-Policy Distillation to constitute MoASE++, with an EMA-anchored on-policy reverse KL distillation and an augmentation policy conditioned on entropy and confidence that aligns predictions across the same views and improves the robustness-plasticity balance. Extensive experiments on classification (CIFAR-10/100-C, ImageNet-C) and semantic segmentation (Cityscapes->ACDC) demonstrate consistent state-of-the-art performance, offering a principled, controllable approach to continual adaptation in dynamic visual environments.
△ Less
Submitted 17 May, 2026;
originally announced May 2026.
-
NeuroMAS: Multi-Agent Systems as Neural Networks with Joint Reinforcement Learning
Authors:
Haoran Lu,
Luyang Fang,
Wenxuan Zhong,
Ping Ma
Abstract:
Multi-agent language systems are often built as hand-designed workflows, where agents are assigned semantic roles and communication protocols are specified in advance. We propose NeuroMAS, a method that first treats a multi-agent language system as a trainable and scalable neural-network-like architecture with LLM agents as nodes and intermediate textual signals as edges. In NeuroMAS, agent nodes…
▽ More
Multi-agent language systems are often built as hand-designed workflows, where agents are assigned semantic roles and communication protocols are specified in advance. We propose NeuroMAS, a method that first treats a multi-agent language system as a trainable and scalable neural-network-like architecture with LLM agents as nodes and intermediate textual signals as edges. In NeuroMAS, agent nodes are role-free but structure-aware: the topology only determines how information can flow in general, while reinforcement learning training determines how nodes communicate, specialize, and coordinate. This formulation shifts multi-agent design from workflow engineering toward architecture design, where depth, width, connectivity, and growth protocol become scalable sources of capability. Further, we provide a theoretical perspective showing why such modular textual computation is more parameter-efficient when tasks admit hierarchical decompositions. Experiments show that NeuroMAS improves significantly over both inference-time and trained multi-agent baselines. We further find that organizational scaling is path-dependent: larger systems can be challenging to train from scratch, but become feasible when grown progressively from smaller trained systems. These results suggest that learned neural multi-agent systems are a promising scaling axis for LLMs.
△ Less
Submitted 15 May, 2026;
originally announced May 2026.
-
UniER: A Unified Benchmark for Item-level and Path-level Exercise Recommendation
Authors:
Xinghe Cheng,
Guiyong Zhuang,
Yusheng Xie,
Jiapu Wang,
Yixin Liu,
Quanlong Guan,
Liangda Fang,
Shirui Pan
Abstract:
Personalized exercise recommendation dynamically aligns pedagogical resources with individual knowledge mastery, which is crucial for satisfying students' dynamic learning needs in modern education. The field is currently driven by two dominant paradigms: Item-Level Exercise Recommendation (ILER) optimizes for immediate single-step state transitions, while Path-Level Exercise Recommendation (PLER)…
▽ More
Personalized exercise recommendation dynamically aligns pedagogical resources with individual knowledge mastery, which is crucial for satisfying students' dynamic learning needs in modern education. The field is currently driven by two dominant paradigms: Item-Level Exercise Recommendation (ILER) optimizes for immediate single-step state transitions, while Path-Level Exercise Recommendation (PLER) constructs coherent learning paths to maximize cumulative gains. Despite sharing the same ultimate objective, disparate evaluation setups have kept these two lines of research isolated, hindering unified benchmarking and fair comparison. To fill the gap, in this paper, we present a Unified Benchmark for Exercise Recommendation (UniER), a comprehensive evaluation framework that unifies ILER and PLER. Specifically, we introduce Weighted Cognitive Gain (WCG) as a unified metric to measure cross-paradigm algorithmic performance. Our benchmark encompasses 9 datasets spanning four generation methods, facilitating the comparison of 18 representative ILER/PLER methods. Through multi-dimensional analyses covering effectiveness, generalizability, robustness, and efficiency, our results reveal the systematic dominance of PLER and expose the pedagogical failure of ILER's fragmented recommendations under extreme sparsity and noise. Furthermore, we provide an open-source codebase of UniER to foster reproducible research and outline potential directions for future investigations.
△ Less
Submitted 15 May, 2026;
originally announced May 2026.
-
Cross Modality Image Translation In Medical Imaging Using Generative Frameworks
Authors:
Giulia Romoli,
Alessia Capoccia,
Filippo Ruffini,
Francesco Di Feola,
Luca Boldrini,
Arturo Chiti,
Renato Cuocolo,
Tugba Akinci D'Antonoli,
Fatemeh Darvizeh,
Marcello Di Pumpo,
Bradley J. Erickson,
Liu Fang,
Deborah Fazzini,
Paola Feraco,
Fabrizia Gelardi,
Francesco Gossetti,
Ana Isabel Hernáiz Ferrer,
Michail E. Klontzas,
Seyedmehdi Payabvash,
Katrine Riklund,
Sara N. Strandberg,
Valerio Guarrasi,
Paolo Soda
Abstract:
Medical image-to-image (I2I) translation enables virtual scanning, i.e. the synthesis of a target imaging modality from a source one without additional acquisitions. Despite growing interest, most proposed methods operate on 2D slices, are evaluated on isolated tasks with different experimental set-ups and lack clinical validation. The primary contribution of this work is a reproducible, standardi…
▽ More
Medical image-to-image (I2I) translation enables virtual scanning, i.e. the synthesis of a target imaging modality from a source one without additional acquisitions. Despite growing interest, most proposed methods operate on 2D slices, are evaluated on isolated tasks with different experimental set-ups and lack clinical validation. The primary contribution of this work is a reproducible, standardized comparative evaluation of 3D I2I translation methods in oncological imaging, designed to standardize preprocessing, splitting, inference, and multi-level evaluation across heterogeneous clinical tasks. Within this framework, we compare seven generative models, three Generative Adversarial Networks (GANs: Pix2Pix, CycleGAN, SRGAN) and four latent generative models (Latent Diffusion Model, Latent Diffusion Model+ControlNet, Brownian Bridge, Flow Matching), across eleven datasets spanning three anatomical regions (head/neck, lung, pelvis) and four translation directions (cone-beam CT to CT, MRI to CT, CT to PET, MRI T2-weighted to T2-FLAIR), for a total of 77 experiments under uniform training, inference, and evaluation conditions. The results show that GANs outperform latent generative models across all tasks, with SRGAN achieving statistically significant superiority. Our lesion-level analysis reveals that all models struggle with small lesions and that, in CT to PET synthesis, models reproduce lesion shape more reliably than absolute uptake-related intensity. We also performed a Visual Turing test administered to 17 physicians, including 15 radiologists, which shows near-chance classification accuracy (56.7%), confirming that synthetic volumes are largely indistinguishable from real acquisitions, while exposing a dissociation between quantitative metrics and clinical preference.
△ Less
Submitted 13 May, 2026;
originally announced May 2026.
-
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments
Authors:
Chiyu Zhang,
Huiqin Yang,
Bendong Jiang,
Xiaolei Zhang,
Yiran Zhao,
Ruyi Chen,
Lu Zhou,
Xiaogang Xu,
Jiafei Wu,
Liming Fang,
Zhe Liu
Abstract:
The rapid proliferation of LLM-based autonomous agents in real operating system environments introduces a new category of safety risk beyond content safety: behavior jailbreak, where an adversary induces an agent to execute dangerous OS-level operations with irreversible consequences. Existing benchmarks either evaluate safety at the semantic layer alone, missing physical-layer harms, or fail to i…
▽ More
The rapid proliferation of LLM-based autonomous agents in real operating system environments introduces a new category of safety risk beyond content safety: behavior jailbreak, where an adversary induces an agent to execute dangerous OS-level operations with irreversible consequences. Existing benchmarks either evaluate safety at the semantic layer alone, missing physical-layer harms, or fail to isolate test cases, letting earlier runs contaminate later ones. We present LITMUS (LLM-agents In-OS Testing for Measuring Unsafe Subversion), a benchmark addressing both gaps via a semantic-physical dual verification mechanism and OS-level state rollback. LITMUS comprises 819 high-risk test cases organized into one harmful seed subset and six attack-extended subsets covering three adversarial paradigms (jailbreak speaking, skill injection, and entity wrapping), plus a fully automated multi-agent evaluation framework judging behavior at both conversational and OS-level physical layers. Evaluation across frontier agents reveals three findings: (1) current agents lack effective safety awareness, with strong models (e.g., Claude Sonnet 4.6) still executing 40.64% of high-risk operations; (2) agents exhibit pervasive Execution Hallucination (EH), verbally refusing a request while the dangerous operation has already completed at the system level, invisible to every prior semantic-only framework; and (3) skill injection and entity wrapping attacks achieve high success rates, exposing pronounced agent vulnerabilities. LITMUS provides the first standardized platform for reproducible, physically grounded behavioral safety evaluation of LLM agents in real OS environments.
△ Less
Submitted 11 May, 2026;
originally announced May 2026.
-
Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents
Authors:
Ying Chen,
Lihuang Fang,
Rui Jiang,
Mingxu Wang,
Zhifeng Gu,
Lei Yi,
Jie Chen
Abstract:
Standard embodied evaluations do not independently score whether an agent correctly commits to task completion at episode closure, a capacity we call terminal commitment. Behaviorally distinct failures--never completing the task, completing it but failing to stop, and reporting success without sufficient evidence--collapse into the same benchmark failure. We introduce VIGIL, an evaluation framewor…
▽ More
Standard embodied evaluations do not independently score whether an agent correctly commits to task completion at episode closure, a capacity we call terminal commitment. Behaviorally distinct failures--never completing the task, completing it but failing to stop, and reporting success without sufficient evidence--collapse into the same benchmark failure. We introduce VIGIL, an evaluation framework that makes terminal commitment independently measurable. Under VIGIL's default protocol, agents observe only egocentric RGB, receive no action-success signals, and must end each episode with a semantic report checked deterministically against hidden world state. This yields two separate scores: world-state completion (W) and benchmark success (B), where B additionally requires a correct terminal report. This decoupling makes four outcome categories distinguishable: missed execution, post-attainment drift, unsupported commitment, and verified success. Across 20 models on 1,000 frozen episodes, systems with comparable W differ by up to 19.7 pp in B: one model converts achieved states into correct reports, while another with near-identical execution drifts past the goal without closing. An action-feedback intervention further tests the separation: execution-oriented signals improve W broadly, yet commitment failures persist in models that do not already ground terminal reports in the achieved state. VIGIL provides a protocol that makes terminal commitment independently visible and scorable.
△ Less
Submitted 2 June, 2026; v1 submitted 9 May, 2026;
originally announced May 2026.
-
Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning
Authors:
Lingzhe Zhang,
Tong Jia,
Yunpeng Zhai,
Liancheng Fang,
Kening Zheng,
Hongyi Liu,
Xiaosong Huang,
Philip S. Yu,
Ying Li
Abstract:
Reinforcement fine-tuning (RFT) has become a core paradigm for post-training large language models, yet its training process remains highly fragile. Existing efforts mainly improve reliability at the system level or address specific issues in individual subproblems by modifying RFT algorithms. Despite their effectiveness, they largely overlook the problem of failure management at the training-proc…
▽ More
Reinforcement fine-tuning (RFT) has become a core paradigm for post-training large language models, yet its training process remains highly fragile. Existing efforts mainly improve reliability at the system level or address specific issues in individual subproblems by modifying RFT algorithms. Despite their effectiveness, they largely overlook the problem of failure management at the training-process level. When training goes wrong, practitioners still rely heavily on expert-driven manual inspection and correction, and automatic failure management for RFT remains largely unexplored. In this paper, we take a first step toward systematic failure management for reinforcement fine-tuning. To understand the empirical structure of RFT failures, we first construct RFT-FaultBench, the first benchmark for fine-grained failures in reinforcement fine-tuning, covering 5 fault families, 16 fault types, 779 training runs, 22,549 train-step records, and 1,457,288 trajectory-level records. Based on this benchmark, we conduct a comprehensive empirical study showing that RFT failures are both observable from training dynamics and distinguishable through their empirical fault fingerprints. Building on these findings, we propose RFT-FM, an automatic failure management framework for reinforcement fine-tuning that unifies anomaly detection, failure diagnosis, and auto remediation in a closed loop. Experimental results show that RFT-FaultBench is neither trivial nor saturated: it exhibits clear anomaly structure while still posing substantial challenges, especially under subtle fault settings. Moreover, RFT-FM shows strong capability in detecting, diagnosing, and mitigating RFT failures.
△ Less
Submitted 5 May, 2026;
originally announced May 2026.
-
Exact formulas for arbitrary order velocity-gradient moments in isotropic turbulence
Authors:
Tong Wu,
Chensheng Luo,
Le Fang,
Michael Wilczek
Abstract:
Statistical moments of velocity gradients provide fundamental information on the small-scale properties of turbulence. In this work, we propose a systematic method to derive exact expressions for statistical moments of arbitrary order for both longitudinal and transverse velocity gradients in isotropic turbulence. The approach is applicable to both compressible and incompressible flows and express…
▽ More
Statistical moments of velocity gradients provide fundamental information on the small-scale properties of turbulence. In this work, we propose a systematic method to derive exact expressions for statistical moments of arbitrary order for both longitudinal and transverse velocity gradients in isotropic turbulence. The approach is applicable to both compressible and incompressible flows and expresses the moments in terms of invariants of the velocity gradient tensor. The derivation combines isotropic tensor theory, orientational averaging, and an algorithmic implementation, enabling the computation of high-order moments in a unified framework. We show that longitudinal velocity gradient moments of order higher than three depend not only on $\mathrm{tr}(\boldsymbol{S}^2)$, which is proportional to the dissipation rate, but also on $\mathrm{tr}(\boldsymbol{S}^3)$, which reflects strain self-amplification, where $\boldsymbol{S}$ denotes the strain-rate tensor. The resulting theoretical expressions are validated through comparisons with existing theoretical results and direct numerical simulations.
△ Less
Submitted 23 April, 2026;
originally announced April 2026.
-
Do LLMs Need to See Everything? A Benchmark and Study of Failures in LLM-driven Smartphone Automation using Screentext vs. Screenshots
Authors:
Shiquan Zhang,
Tianyi Zhang,
Le Fang,
Simon D'Alfonso,
Hong Jia,
Vassilis Kostakos
Abstract:
With the rapid advancement of large language models (LLMs), mobile agents have emerged as promising tools for phone automation, simulating human interactions on screens to accomplish complex tasks. However, these agents often suffer from low accuracy, misinterpretation of user instructions, and failure on challenging tasks, with limited prior work examining why and where they fail. To address this…
▽ More
With the rapid advancement of large language models (LLMs), mobile agents have emerged as promising tools for phone automation, simulating human interactions on screens to accomplish complex tasks. However, these agents often suffer from low accuracy, misinterpretation of user instructions, and failure on challenging tasks, with limited prior work examining why and where they fail. To address this, we introduce DailyDroid, a benchmark of 75 tasks in five scenarios across 25 Android apps, spanning three difficulty levels to mimic everyday smartphone use. We evaluate it using text-only and multimodal (text + screenshot) inputs on GPT-4o and o4-mini across 300 trials, revealing comparable performance with multimodal inputs yielding marginally higher success rates. Through in-depth failure analysis, we compile a handbook of common failures. Our findings reveal critical issues in UI accessibility, input modalities, and LLM/app design, offering implications for future mobile agents, applications, and UI development.
△ Less
Submitted 20 April, 2026;
originally announced April 2026.
-
Net Load Forecasting Using Machine Learning with Growing Renewable Power Capacity Features: A Comparative Study of Direct and Indirect Methods
Authors:
Oluwafolajimi Samuel Bolusteve,
Linhan Fang,
Xingpeng Li
Abstract:
Renewable energy adoption has increased significantly over the past few years. However, with the increasing adoption of renewable energy, forecasting the net load has become a major challenge due to the inherent uncertainty associated with these renewable sources. To mitigate the impact of uncertainties, this study utilizes long short-term memory (LSTM) model and fully connected neural networks (F…
▽ More
Renewable energy adoption has increased significantly over the past few years. However, with the increasing adoption of renewable energy, forecasting the net load has become a major challenge due to the inherent uncertainty associated with these renewable sources. To mitigate the impact of uncertainties, this study utilizes long short-term memory (LSTM) model and fully connected neural networks (FCNN) to predict net load based on two independent approaches: the direct method and indirect method. While the conventional direct method directly forecasts the target net load, the indirect approach derives it by separately predicting total load and renewable energy generation. Furthermore, this study innovatively incorporates renewable energy capacity as an input feature to train the forecasting model. The indirect method for FCNN provided a better estimate than the direct method, and the indirect method for LSTM model gave the best prediction. These findings suggest that recurrent architectures like LSTM are particularly well-suited for net load forecasting applications, while the choice between direct and indirect methods depends on the specific neural network architecture employed. By advancing reliable forecasting tools for renewable energy integration, this work enhances grid resilience and accelerates the transition toward renewable-dominant power systems.
△ Less
Submitted 18 April, 2026;
originally announced April 2026.
-
A Parallel Approach to Counting Exact Covers Based on Decomposability Property
Authors:
Liangda Fang,
Yaohui Luo,
Delong Li,
Xuanxiang Huang,
Quanlong Guan
Abstract:
The exact cover problem is a classical NP-hard problem with broad applications in the area of AI. Algorithm DXZ is a method to count exact covers representing by zero-suppressed binary decision diagrams (ZBDDs). In this paper, we propose a zero-suppressed variant of decision decomposable negation normal form (in short, decision-ZDNNF), which is strictly more succinct than ZBDDs. We then design a n…
▽ More
The exact cover problem is a classical NP-hard problem with broad applications in the area of AI. Algorithm DXZ is a method to count exact covers representing by zero-suppressed binary decision diagrams (ZBDDs). In this paper, we propose a zero-suppressed variant of decision decomposable negation normal form (in short, decision-ZDNNF), which is strictly more succinct than ZBDDs. We then design a novel parallel algorithm, namely DXD, which constructs a decision-ZDNNF representing the set of all exact covers. Furthermore, we improve DXD by dynamically updating connected components. The experimental results demonstrate that the improved DXD algorithm outperforms all of state-of-the-art methods.
△ Less
Submitted 16 April, 2026;
originally announced April 2026.
-
FlowPalm: Optical Flow Driven Non-Rigid Deformation for Geometrically Diverse Palmprint Generation
Authors:
Yuchen Zou,
Huikai Shao,
Lihuang Fang,
Zhipeng Xiong,
Dexing Zhong
Abstract:
Recently, synthetic palmprints have been increasingly used as substitutes for real data to train recognition models. To be effective, such synthetic data must reflect the diversity of real palmprints, including both style variation and geometric variation. However, existing palmprint generation methods mainly focus on style translation, while geometric variation is either ignored or approximated b…
▽ More
Recently, synthetic palmprints have been increasingly used as substitutes for real data to train recognition models. To be effective, such synthetic data must reflect the diversity of real palmprints, including both style variation and geometric variation. However, existing palmprint generation methods mainly focus on style translation, while geometric variation is either ignored or approximated by simple handcrafted augmentations. In this work, we propose FlowPalm, an optical-flow-driven palmprint generation framework capable of simulating the complex non-rigid deformations observed in real palms. Specifically, FlowPalm estimates optical flows between real palmprint pairs to capture the statistical patterns of geometric deformations. Building on these priors, we design a progressive sampling process that gradually introduces the geometric deformations during diffusion while maintaining identity consistency. Extensive experiments on six benchmark datasets demonstrate that FlowPalm significantly outperforms state-of-the-art palmprint generation approaches in downstream recognition tasks. Project page: https://yuchenzou.github.io/FlowPalm/
△ Less
Submitted 10 April, 2026;
originally announced April 2026.
-
When Users Change Their Mind: Evaluating Interruptible Agents in Long-Horizon Web Navigation
Authors:
Henry Peng Zou,
Chunyu Miao,
Wei-Chieh Huang,
Yankai Chen,
Yue Zhou,
Hanrong Zhang,
Yaozu Wu,
Liancheng Fang,
Zhengyao Gu,
Zhen Zhang,
Kening Zheng,
Fangxin Wang,
Yi Nian,
Shanghao Li,
Wenzhe Fan,
Langzhou He,
Weizhi Zhang,
Xue Liu,
Philip S. Yu
Abstract:
As LLM agents transition from short, static problem solving to executing complex, long-horizon tasks in dynamic environments, the ability to handle user interruptions, such as adding requirement or revising goals, during mid-task execution is becoming a core requirement for realistic deployment. However, existing benchmarks largely assume uninterrupted agent behavior or study interruptions only in…
▽ More
As LLM agents transition from short, static problem solving to executing complex, long-horizon tasks in dynamic environments, the ability to handle user interruptions, such as adding requirement or revising goals, during mid-task execution is becoming a core requirement for realistic deployment. However, existing benchmarks largely assume uninterrupted agent behavior or study interruptions only in short, unconstrained language tasks. In this paper, we present the first systematic study of interruptible agents in long-horizon, environmentally grounded web navigation tasks, where actions induce persistent state changes. We formalize three realistic interruption types, including addition, revision, and retraction, and introduce InterruptBench, a benchmark derived from WebArena-Lite that synthesizes high-quality interruption scenarios under strict semantic constraints. Using a unified interruption simulation framework, we evaluate six strong LLM backbones across single- and multi-turn interruption settings, analyzing both their effectiveness in adapting to updated intents and their efficiency in recovering from mid-task changes. Our results show that handling user interruptions effectively and efficiently during long-horizon agentic tasks remains challenging for powerful large-scale LLMs. Code and dataset are available at https://github.com/HenryPengZou/InterruptBench.
△ Less
Submitted 1 April, 2026;
originally announced April 2026.
-
Locally Confident, Globally Stuck: The Quality-Exploration Dilemma in Diffusion Language Models
Authors:
Liancheng Fang,
Aiwei Liu,
Henry Peng Zou,
Yankai Chen,
Enze Ma,
Leyi Pan,
Chunyu Miao,
Wei-Chieh Huang,
Xue Liu,
Philip S. Yu
Abstract:
Diffusion large language models (dLLMs) theoretically permit token decoding in arbitrary order, a flexibility that could enable richer exploration of reasoning paths than autoregressive (AR) LLMs. In practice, however, random-order decoding often hurts generation quality. To mitigate this, low-confidence remasking improves single-sample quality (e.g., Pass@$1$) by prioritizing confident tokens, bu…
▽ More
Diffusion large language models (dLLMs) theoretically permit token decoding in arbitrary order, a flexibility that could enable richer exploration of reasoning paths than autoregressive (AR) LLMs. In practice, however, random-order decoding often hurts generation quality. To mitigate this, low-confidence remasking improves single-sample quality (e.g., Pass@$1$) by prioritizing confident tokens, but it also suppresses exploration and limits multi-sample gains (e.g., Pass@$k$), creating a fundamental quality--exploration dilemma. In this paper, we provide a unified explanation of this dilemma. We show that low-confidence remasking improves a myopic proxy for quality while provably constraining the entropy of the induced sequence distribution. To overcome this limitation, we characterize the optimal distribution that explicitly balances quality and exploration, and develop a simple Independent Metropolis--Hastings sampler that approximately targets this distribution during decoding. Experiments across a range of reasoning benchmarks including MATH500, AIME24/25, HumanEval, and MBPP show that our approach yields better exploration-quality tradeoff than both random and low-confidence remasking.
△ Less
Submitted 31 March, 2026;
originally announced April 2026.
-
Uniform Interpolation in Distributed Knowledge Modal Logics
Authors:
Kexu Wang,
Liangda Fang
Abstract:
Uniform interpolation is the property that, for any formula and set of atoms, there exists the strongest consequence omitting those atoms. It plays a central role in knowledge representation and reasoning tasks such as knowledge update and information hiding. This paper studies the uniform interpolation property in epistemic modal logics with distributed knowledge, which captures agents' collectiv…
▽ More
Uniform interpolation is the property that, for any formula and set of atoms, there exists the strongest consequence omitting those atoms. It plays a central role in knowledge representation and reasoning tasks such as knowledge update and information hiding. This paper studies the uniform interpolation property in epistemic modal logics with distributed knowledge, which captures agents' collective reasoning abilities. Building on the bisimulation-quantifier perspective, we extend the canonical-formula and literal-elimination framework of Fang, Liu, and van Ditmarsch to distributed knowledge settings and introduce the concept of collective $p$-bisimulation. We show that, for distributed knowledge modal logics $\mathsf{K}_n\mathbf{D}$, $\mathsf{D}_n\mathbf{D}$, and $\mathsf{T}_n\mathbf{D}$, every satisfiable canonical formula's uniform interpolant omitting an atom $p$ is exactly its remainder of eliminating $p$. Then, we provide a finer analysis for the transitive and Euclidean systems $\mathsf{K45}_n\mathbf{D}$, $\mathsf{KD45}_n\mathbf{D}$, and $\mathsf{S5}_n\mathbf{D}$, and prove that every formula of modal depth $k + 1$ has a uniform interpolant of modal depth $2 k + 1$. Thus, we prove the uniform interpolation property in all the six distributed knowledge modal logics. Finally, we generalize the results to some variants with propositional common knowledge and discuss the method's limitations.
△ Less
Submitted 30 March, 2026;
originally announced March 2026.
-
Closeby Habitable Exoplanet Survey (CHES). V. Planetary Parameters Derived from Angular Separation Variations
Authors:
Dongjie Tan,
Jianghui Ji,
Chunhui Bao,
Xiumin Huang,
Guo Chen,
Su Wang,
Yao Dong,
Jiacheng Liu,
Zi Zhu,
Haitao Li,
Junbo Zhang,
Liang Fang,
Dong Li,
Lei Deng
Abstract:
The Closeby Habitable Exoplanet Survey (CHES) aims to achieve microarcsecond-level astrometry of about one hundred nearby FGK-type stars within 10 parsecs to detect Earth-like planets. Such precision exceeds the capability of absolute astrometry relying on Gaia catalogs, whose positional accuracy degrades over time due to error propagation from stellar motion and epoch offsets, limiting their use…
▽ More
The Closeby Habitable Exoplanet Survey (CHES) aims to achieve microarcsecond-level astrometry of about one hundred nearby FGK-type stars within 10 parsecs to detect Earth-like planets. Such precision exceeds the capability of absolute astrometry relying on Gaia catalogs, whose positional accuracy degrades over time due to error propagation from stellar motion and epoch offsets, limiting their use in microarcsecond-level detection. Traditional relative astrometry depends on positional components along right ascension and declination, requiring precise knowledge of field rotation and satellite attitude, which introduces additional errors. To address this, we propose a new relative measurement model based solely on variations in the length of angular separation between the target and reference stars, independent of direction. The model incorporates effects such as proper motion, parallax, radial velocity, light aberration, gravitational lensing, and planetary perturbations, enabling reconstruction of planetary orbits and masses. This approach enhances measurement stability and precision, providing a framework that is not entirely dependent on the Gaia catalog and suitable for CHES and other future high-accuracy astrometric missions.
△ Less
Submitted 29 March, 2026;
originally announced March 2026.
-
TRACE: A Multi-Agent System for Autonomous Physical Reasoning for Seismology
Authors:
Feng Liu,
Jian Xu,
Xin Cui,
Xinghao Wang,
Zijie Guo,
Jiong Wang,
S. Mostafa Mousavi,
Xinyu Gu,
Hao Chen,
Ben Fei,
Lihua Fang,
Fenghua Ling,
Zefeng Li,
Lei Bai
Abstract:
Inferring physical mechanisms that govern earthquake sequences from geophysical observations remains a challenging task, particularly across tectonically distinct environments where similar seismic patterns can reflect different underlying processes. Current seismological processing and interpretation rely heavily on experts' choice of parameters and the synthesis of various seismological products…
▽ More
Inferring physical mechanisms that govern earthquake sequences from geophysical observations remains a challenging task, particularly across tectonically distinct environments where similar seismic patterns can reflect different underlying processes. Current seismological processing and interpretation rely heavily on experts' choice of parameters and the synthesis of various seismological products, limiting reproducibility and the formation of generalizable knowledge across settings. Here we present TRACE (Trans-perspective Reasoning and Automated Comprehensive Evaluator), a multi-agent system that combines large language model planning with formal seismological constraints to derive auditable, physically grounded mechanistic inferences from raw observations. Applied to the 2019 Ridgecrest sequence, TRACE autonomously identifies stress-perturbation-induced delayed triggering, resolving the cascading interaction between the Mw 6.4 and Mw 7.1 mainshocks. For the 2025 Santorini-Kolumbo volcanic eruption, the system identifies a structurally guided intrusion model, distinguishing episodic migration via fault channels from the continuous propagation expected in homogeneous crustal failure. By providing a generalizable infrastructure for deriving physical insights from seismic phenomena, TRACE advances the field from expert-dependent analysis toward knowledge-guided autonomous discovery in Earth sciences.
△ Less
Submitted 25 March, 2026; v1 submitted 22 March, 2026;
originally announced March 2026.
-
Flow-based Polynomial Chaos Expansion for Uncertainty Quantification in Power System Dynamic Simulation
Authors:
Le Fang,
Wangkun Xu,
Fei Teng
Abstract:
The large-scale integration of renewable energy sources introduces significant operational uncertainty into power systems. Although Polynomial Chaos Expansion (PCE) provides an efficient tool for uncertainty quantification (UQ) in power system dynamics, its accuracy depends critically on the faithful representation of input uncertainty, an assumption that is oftern violated in practice due to corr…
▽ More
The large-scale integration of renewable energy sources introduces significant operational uncertainty into power systems. Although Polynomial Chaos Expansion (PCE) provides an efficient tool for uncertainty quantification (UQ) in power system dynamics, its accuracy depends critically on the faithful representation of input uncertainty, an assumption that is oftern violated in practice due to correlated, non-Gaussian, and otherwise complex data distributions. In contrast to purely data-driven surrogates that often overlook rigorous input distribution modelling, this paper introduces flow-based PCE, a unified framework that couples expressive input modelling with efficient uncertainty propagation. Specifically, normalising flows are employed to learn an invertible transport map from a simple base distribution to the empirical joint distribution of uncertain inputs, and this map is then integrated directly into the PCE construction. In addition, the Map Smoothness Index (MSI) is introduced as a new metric to quantify the quality of the learned map, and smoother transformations are shown to yield more accurate PCE surrogates. The proposed Flow-based PCE framework is validated on benchmark dynamic models, including the IEEE 14-bus system and the Great Britain transmission system, under a range of uncertainty scenarios.
△ Less
Submitted 20 March, 2026;
originally announced March 2026.