-
MOSS-VL Technical Report
Authors:
Pengyu Wang,
Chenkun Tan,
Shaojun Zhou,
Qirui Zhou,
Yanxin Chen,
Xingyang He,
Huazheng Zeng,
Jijun Cheng,
Chenghao Wang,
Xiaomeng Qian,
Pengfei Wang,
Zhan Huang,
Shanqing Gao,
Wei Huang,
Longjun Cao,
Wu Ran,
Jie Liu,
Changtai Zhu,
Hongkai Wang,
Yixian Tian,
Chenghao Liu,
Zhen Ye,
Xinghao Wang,
Botian Jiang,
Guoguo Feng
, et al. (7 additional authors not shown)
Abstract:
We present MOSS-VL, an open vision-language model family that treats real-time interaction -- perceiving while it speaks -- as a first-class capability. It is co-designed across the stack: the language decoder attends to vision only through gated cross-attention, so the model can naturally see incoming frames while generating; a synthesized interaction corpus supervises when to speak, when to stay…
▽ More
We present MOSS-VL, an open vision-language model family that treats real-time interaction -- perceiving while it speaks -- as a first-class capability. It is co-designed across the stack: the language decoder attends to vision only through gated cross-attention, so the model can naturally see incoming frames while generating; a synthesized interaction corpus supervises when to speak, when to stay silent, and when to revise; and a staged curriculum concentrates all real-time-specific training in one light final stage over a strong offline foundation. Offline, MOSS-VL-Instruct is competitive at comparable scale and leads temporal-reasoning video sets. Across four streaming benchmarks, MOSS-VL-Realtime posts the best average on three (second on the fourth) among open-source streaming models, sweeping the three subsets that squarely test proactive behavior -- 66.0 vs. 37.5 for the best baseline on OmniMMI Proactive Alerting. With 11.3B parameters but visual tokens outside the decoded sequence, MOSS-VL widens its time-to-first-token advantage over same-backbone Qwen3-VL-8B from 2.8x to 5.1x as visual context grows. We release all five checkpoints, the training curriculum, and the real-time inference code at https://github.com/OpenMOSS/MOSS-VL.
△ Less
Submitted 15 August, 2026;
originally announced August 2026.
-
Rethinking Factor Sharing in Federated LoRA: A Rank-Aware Adaptive Approach
Authors:
Xinyi Xu,
Bingnan Xiao,
Shuang Qin,
Gang Feng,
Tony Q. S. Quek
Abstract:
Low-rank adaptation (LoRA) represents large language model (LLM) updates with two compact matrix factors, i.e., $A$ and $B$, providing an efficient way to fine-tune large models in federated learning paradigm. Inspired by the asymmetric roles of the LoRA factors, we study whether $A$ should be shared across clients while $B$ remains client-specific (Share-A/Local-B), or whether $B$ should instead…
▽ More
Low-rank adaptation (LoRA) represents large language model (LLM) updates with two compact matrix factors, i.e., $A$ and $B$, providing an efficient way to fine-tune large models in federated learning paradigm. Inspired by the asymmetric roles of the LoRA factors, we study whether $A$ should be shared across clients while $B$ remains client-specific (Share-A/Local-B), or whether $B$ should instead be shared while $A$ remains client-specific (Share-B/Local-A). With a least-squares surrogate, we reveal that Share-A/Local-B requires the client-specific LoRA update matrices to use a common rank-$r$ input-side space, whereas Share-B/Local-A requires a common rank-$r$ output-side space. The two strategies therefore incur different projection residuals, indicating that the preferred strategy is the one with the smaller aggregate residual across clients. With this insight, we propose Federated Adaptive Factor Sharing Low-Rank Adaptation (FedAS-LoRA), which selects the sharing side before training to enhance fine-tuning performance. To enable adaptive factor selection before training, we design a Rank-Aware Shared-Subspace Sufficiency (RSS) metric, which effectively assesses whether a shared rank-$r$ input subspace is sufficient for the local data distributions using representations extracted from a frozen LLM backbone. Experiments across different tasks, data distributions, LoRA ranks, and participation settings confirm the effectiveness of RSS and the superior performance of FedAS-LoRA.
△ Less
Submitted 10 August, 2026;
originally announced August 2026.
-
HarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent Harnesses
Authors:
Luan Zhang,
Ruochen Zhou,
Dandan Song,
Zhengyu Chen,
Yuhang Tian,
Jun Yang,
Huipeng Ma,
Chenhao Li,
Guangyuan Feng,
Xudong Li,
Yizhou Jin,
Yan Xu
Abstract:
Harness design plays a critical role in agent performance by shaping how large language models (LLMs) perceive, reason over, and act within executable environments. Recent work has proposed automatic harness evolution, which iteratively improves the harness from agent--environment interactions. However, existing methods often overfit to the evolution tasks, rely exclusively on trajectory-derived s…
▽ More
Harness design plays a critical role in agent performance by shaping how large language models (LLMs) perceive, reason over, and act within executable environments. Recent work has proposed automatic harness evolution, which iteratively improves the harness from agent--environment interactions. However, existing methods often overfit to the evolution tasks, rely exclusively on trajectory-derived signals, and optimize harness components jointly, causing interference across components. We propose HarnessCompass, a novel automatic harness evolution framework built around constrained evolution, proactive feedback, and component-wise optimization. HarnessCompass first enforces global constraints on evolution, restricting modifications to task-agnostic harness changes that generalize beyond the evolution tasks. It then augments trajectory-derived evidence with proactive first-person feedback from the agent about harness usage, yielding richer signals for evolution. Finally, it decouples the optimization of different harness components before consolidating them into a unified harness, reducing cross-component interference while preserving component synergy. On SWE-bench Verified with GPT-5.4, HarnessCompass improves Pass@1 from 54\% to 66\% in only 5 evolution iterations, outperforming AHE in both effectiveness and evolution efficiency. In addition, the evolved harness transfers effectively to held-out tasks and other models, demonstrating substantially stronger generalization than prior automatic harness evolution methods.
△ Less
Submitted 3 August, 2026;
originally announced August 2026.
-
Symmetry-Based Microscopic Theory of the Unconventional Pairing Mechanism in La$_5$Ni$_3$O$_{11}$
Authors:
Guan-Hao Feng,
Jun Quan
Abstract:
Recent experiments report high-temperature superconductivity in the hybrid nickelate $\mathrm{La}_5\mathrm{Ni}_3\mathrm{O}_{11}$, which is composed of alternating stacks of bilayer $\mathrm{La}_3\mathrm{Ni}_2\mathrm{O}_7$ and monolayer $\mathrm{La}_2\mathrm{NiO}_4$. However, the superconducting transition temperature $T_c \approx 64~\mathrm{K}$ for $\mathrm{La}_5\mathrm{Ni}_3\mathrm{O}_{11}$ is re…
▽ More
Recent experiments report high-temperature superconductivity in the hybrid nickelate $\mathrm{La}_5\mathrm{Ni}_3\mathrm{O}_{11}$, which is composed of alternating stacks of bilayer $\mathrm{La}_3\mathrm{Ni}_2\mathrm{O}_7$ and monolayer $\mathrm{La}_2\mathrm{NiO}_4$. However, the superconducting transition temperature $T_c \approx 64~\mathrm{K}$ for $\mathrm{La}_5\mathrm{Ni}_3\mathrm{O}_{11}$ is remarkably lower than the $80~\mathrm{K}$ observed for pressurized $\mathrm{La}_3\mathrm{Ni}_2\mathrm{O}_7$. Thus, an unified microscopic theory is required to address the difference in the pairing mechanisms between these systems. Here, we develop a phenomenological symmetry-based approach to systematically analyze the low-energy physics in $\mathrm{La}_5\mathrm{Ni}_3\mathrm{O}_{11}$, which is obtained by a charge self-consistent density functional theory plus dynamical mean-field theory method. We show that the superconductivity in $\mathrm{La}_5\mathrm{Ni}_3\mathrm{O}_{11}$ exhibits a two-gap nature, consisting of a leading interlayer pairing between the $d_{z^2}$ orbitals and a subleading intralayer pairing between the $d_{x^2-y^2}$ orbitals. The reduction of $T_c$ can be attributed to the diminished contribution of the interlayer pairing, as reflected by the hopping parameter ratio $|t_{\perp}^z/t_{\parallel}^{x}|$. Base on this unified picture, we discuss the possible pairing mechanism and the role of $γ$ pocket for the superconductivity in the bilayer NiO$_2$ planes of nickelate superconductors.
△ Less
Submitted 21 July, 2026; v1 submitted 20 July, 2026;
originally announced July 2026.
-
Time-Frequency Consistency Learning for Robust Speech Deepfake Detection
Authors:
Jun Xue,
Zhuolin Yi,
Yanzhen Ren,
Yihuan Huang,
Jiayu Xiong,
Yi Chai,
Guanxiang Feng,
Jiajun Liu,
Tong Zhang
Abstract:
Recently, speech deepfake detection (SDD) has achieved significant progress. However, its robustness evaluation remains largely confined to controlled additive noise scenarios, lacking systematic investigation of the complex distortions introduced by acoustic front-end (AFE) processing pipelines in real-world deployments. In this work, we simulate a unified AFE pipeline comprising acoustic echo ca…
▽ More
Recently, speech deepfake detection (SDD) has achieved significant progress. However, its robustness evaluation remains largely confined to controlled additive noise scenarios, lacking systematic investigation of the complex distortions introduced by acoustic front-end (AFE) processing pipelines in real-world deployments. In this work, we simulate a unified AFE pipeline comprising acoustic echo cancellation, noise suppression, automatic gain control, and voice activity detection (VAD), and conduct a comprehensive evaluation of current state-of-the-art models. The results show that the nonlinear and time-frequency coupled distortions introduced by AFE significantly degrade detection performance. To address this issue, we propose a Time-Frequency Consistency Learning (TFCL) framework, which aims to learn invariant spoofing representations that remain stable before and after AFE processing. We observe that AFE not only introduces temporal misalignment (e.g., segment-level shifts caused by VAD), but also weakens or distorts critical frequency-domain cues. To this end, TFCL employs an attention-driven soft alignment mechanism to capture cross-temporal dependencies, along with frequency-domain structural consistency constraints to enforce feature invariance. As a result, the model is able to maintain stable representations under both temporal perturbations and spectral distortions. Extensive experimental results demonstrate that the proposed method effectively mitigates the performance degradation caused by AFE processing, significantly improving the robustness of SDD in real-world scenarios. The code is available at https://github.com/JunXue-tech/TFCL.
△ Less
Submitted 28 July, 2026; v1 submitted 20 July, 2026;
originally announced July 2026.
-
Channel Knowledge Empowered Finite-Blocklength Rate-Splitting Transmission for High-Mobility Autonomous Driving
Authors:
Yi Wang,
Yingyang Chen,
Feng Bai,
Li Wang,
Gang Feng
Abstract:
To meet the extended ultra-low latency and high reliability (xURLLC) requirements for autonomous driving systems, multiple access schemes must operate reliably in high-mobility and complex propagation environments. Recently, rate-splitting multiple access (RSMA) has emerged as a promising multi-user transmission framework, showing robustness in dynamic situations where imperfect and outdated chann…
▽ More
To meet the extended ultra-low latency and high reliability (xURLLC) requirements for autonomous driving systems, multiple access schemes must operate reliably in high-mobility and complex propagation environments. Recently, rate-splitting multiple access (RSMA) has emerged as a promising multi-user transmission framework, showing robustness in dynamic situations where imperfect and outdated channel state information (CSI) is prevalent.Moreover, the advanced sensing, localization, and on-board computation capabilities of autonomous driving vehicles facilitate the construction of a channel knowledge map (CKM), which is a key enabler for environment-aware communications in future 6G networks.In this work, we propose a CKM empowered finite-blocklength (FBL) RSMA for downlink autonomous driving system. The location-dependent large-scale channel information provided by CKM is exploited in RSMA to develop a refined rate-splitting design. The min-rate performance of FBL rate splitting is analyzed explicitly to ensure user fairness. We derive a new and tight closed-form bound for the private-stream ergodic rate. Combined with the closed-form common-stream expression, an efficient optimization design of rate-splitting ratios has been formulated. Numerical results show that the CKM empowered FBL RSMA outperforms space-division multiple access (SDMA) and non-orthogonal multiple access (NOMA), particularly in high-mobility scenarios. Its performance is improved by a data-based CKM, which provides more accurate large-scale channel information than model-based approaches and enables more precise common-stream allocation. The results also reveal that RSMA is sensitive to errors in large-scale channel knowledge, emphasizing the importance of accurate CKM information for optimal rate-splitting.
△ Less
Submitted 11 July, 2026;
originally announced July 2026.
-
Single-sideband-interference twin-field quantum key distribution without global phase locking
Authors:
Xingjian Li,
Bingkun Wang,
Jianyong Hu,
Jianqiang Liu,
Shuxiao Wu,
Guosheng Feng,
Zhixing Qiao,
Changgang Yang,
Ruiyun Chen,
Chengbing Qin,
Guofeng Zhang,
Liantuan Xiao,
Suotang Jia
Abstract:
Twin-field quantum key distribution (TF QKD) can overcome the fundamental rate loss limit of repeaterless quantum links, but its practical deployment has long been hindered by the requirement of global phase locking between two independent lasers. By revisiting the fundamental principles of optical interference, this work reveals that interference in TF QKD inherently relies only on the instantane…
▽ More
Twin-field quantum key distribution (TF QKD) can overcome the fundamental rate loss limit of repeaterless quantum links, but its practical deployment has long been hindered by the requirement of global phase locking between two independent lasers. By revisiting the fundamental principles of optical interference, this work reveals that interference in TF QKD inherently relies only on the instantaneous phase alignment of two independent optical pulses at the moment they temporally overlap, rather than on continuous global phase synchronization. Guided by this insight, we propose and demonstrate a single-sideband-interference TF-QKD protocol that eliminates global phase locking. Each user employs an IQ modulator to generate a weak single sideband as the quantum signal, while the intrinsically phase-correlated optical carrier propagates as a real-time phase reference. Carrier interference at the receiver enables real-time phase extraction and feedback compensation for the sidebands. Unlike prior no phase locking approaches requiring second- or microsecond-level coherence, in principle, our scheme reduces this requirement to nanoseconds. We achieve 98% interference visibility over 100.8 km fibre and secure key rates surpassing the PLOB bound in the high-loss regime, providing a simpler route towards practical long-distance quantum communication networks.
△ Less
Submitted 26 June, 2026;
originally announced June 2026.
-
Shigatse Astronomical Site Testing. I. Cloud-cover Climatology and Selected Local Meteorological Conditions
Authors:
Baiyu Zhang,
Hejun Yang,
Xiaojun Dong,
Lingling Wang,
Juean Luobu,
Minfeng Gu,
Xiyan Peng,
Hao Luo,
Yindun Mao,
ZhaoXiang Qi,
Basangzeren,
Qihang He,
Guojie Feng,
Chunhai Bai,
Ali Esamdin,
Wenbo Gu,
Siqi Wang,
Zihuang Cao
Abstract:
As the first paper in a Shigatse astronomical site-testing series, we present a multi-source assessment of cloud cover and selected local meteorological conditions at the Shigatse 40 m site on the southern Tibetan Plateau. The study combines CALIPSO-GOCCP active-lidar climatology, ISCCP HXG passive-satellite cloud fields, conventional total-cloud-amount observations from the Shigatse Meteorologica…
▽ More
As the first paper in a Shigatse astronomical site-testing series, we present a multi-source assessment of cloud cover and selected local meteorological conditions at the Shigatse 40 m site on the southern Tibetan Plateau. The study combines CALIPSO-GOCCP active-lidar climatology, ISCCP HXG passive-satellite cloud fields, conventional total-cloud-amount observations from the Shigatse Meteorological Station, and on-site Weather Station measurements. Together, these records characterize Shigatse as a southern-plateau monsoon-transition cloud regime: the active-lidar climatology gives a moderate-to-low annual cloud fraction, and the cloudier months are concentrated in the June--September monsoon interval. In GOCCP, the annual mean cloud fraction is 42.1%, while the October--May low-cloud season has a mean cloud fraction of 26.3%, compared with 73.7% during the June--September monsoon interval. ISCCP gives higher absolute cloud fractions but supports the same seasonal phase and local spatial placement. The aligned 1988--2013 meteorological-station record gives a total-cloud-amount <=40% fraction of 80.7% during October--May, rising to 90.7% in the November--January core, and decreasing to 39.9% during June--September. The 2024--2025 Weather Station archive further shows high fractions of valid samples satisfying the adopted meteorological criteria during the low-cloud months: 92.6% for the October--May night-time proxy and 94.6% for the corresponding 24 h samples. These results identify Shigatse as a measured lower-latitude southern-plateau cloud-cover reference within China's site-testing network, with a well-defined October--May low-cloud observing period and a Shigatse--Ali low-cloud corridor for subsequent regional site testing.
△ Less
Submitted 14 June, 2026;
originally announced June 2026.
-
Large language models selectively converge with human-shared neural semantic representations
Authors:
Chen Hong,
Ximing Shao,
Gangyi Feng
Abstract:
Interpersonal communication requires building shared semantics that enable listeners to understand speakers' meanings from their unfolding language, but the dimensional structure of this shared neural representation remains unclear. LLMs increasingly approximate human language capability and neural responses, raising the question of whether they capture the same semantic structure shared between h…
▽ More
Interpersonal communication requires building shared semantics that enable listeners to understand speakers' meanings from their unfolding language, but the dimensional structure of this shared neural representation remains unclear. LLMs increasingly approximate human language capability and neural responses, raising the question of whether they capture the same semantic structure shared between human brains. Here, we combined storytelling-listening pseudo-hyperscanning MEG with dimension-resolved interbrain encoding modeling to compare human- and LLM-derived accounts of shared neural semantic representations. Content words from the speaker's narratives were rated by humans and five recent LLMs along ten semantic dimensions (i.e., perception, motor, space, time, socialness, animacy, emotion, attention, causality, and drive). We tested whether these dimensions explained speaker-listener neural synchronization (NS) beyond acoustic and phonological features. Both human- and LLM-derived semantic spaces explained NS, but these shared semantics are better characterized as a multidimensional neural structure rather than a single global signal. These patterns also predicted individual differences in listeners' story comprehension, linking neural alignment to cognition. However, comparable overall prediction concealed systematic differences in representational geometry. Larger LLMs aligned more closely and showed greater overlap with humans in semantic structure and NS, but this was incomplete and dimension-dependent. The largest divergences emerged for dimensions closely tied to agency, affect, and social experience. These findings show that LLMs capture substantial components of human shared neural semantics, but their alignment is selective. Larger or more capable models improve the approximation, whereas socially and affectively grounded dimensions are captured only partially.
△ Less
Submitted 9 June, 2026;
originally announced June 2026.
-
Exploring the Scale and Diversity of Speech Anti-spoofing Datasets: Experiments and Analysis
Authors:
Zhuolin Yi,
Jun Xue,
Yanzhen Ren,
Yihuan Huang,
Yi Chai,
Daixian Li,
Guanxiang Feng,
Jiajun Liu
Abstract:
The scale of speech anti-spoofing datasets has grown exponentially over the past decade, driven by the assumption that larger data leads to better performance. However, it remains unclear whether indiscriminate scaling commensurately improves model generalization. This study challenges the "scale-first" paradigm by decoupling the impacts of training data scale versus diversity. Through experiments…
▽ More
The scale of speech anti-spoofing datasets has grown exponentially over the past decade, driven by the assumption that larger data leads to better performance. However, it remains unclear whether indiscriminate scaling commensurately improves model generalization. This study challenges the "scale-first" paradigm by decoupling the impacts of training data scale versus diversity. Through experiments on representative datasets, we report two key findings: (1) Larger is not always better. Expanding data scale excessively under fixed generation methods yields negligible returns and may even degrade cross-domain generalization due to overfitting.(2) Diversity outweighs scale. A smaller composite training set featuring diverse attacks significantly outperforms larger-scale datasets with limited diversity in cross-dataset evaluations. We conclude that future dataset construction should prioritize the diversity of generation methods over scale to effectively enhance model generalization.
△ Less
Submitted 6 June, 2026;
originally announced June 2026.
-
Synthetic Data Alone is Enough? Rethinking Data Scarcity in Pediatric Rare Disease Recognition
Authors:
Ganlin Feng,
Yuxi Long,
Erin Lou,
Lianghong Chen,
Zihao Jing,
Pingzhao Hu,
Wei Xu
Abstract:
Children with rare genetic diseases often exhibit distinctive facial phenotypes, yet developing computer vision systems for early diagnosis remains challenging due to extreme data scarcity, privacy constraints, and limited data sharing in pediatric settings. These challenges not only hinder automated diagnosis but also restrict the availability of visual resources for clinical genetic counseling.…
▽ More
Children with rare genetic diseases often exhibit distinctive facial phenotypes, yet developing computer vision systems for early diagnosis remains challenging due to extreme data scarcity, privacy constraints, and limited data sharing in pediatric settings. These challenges not only hinder automated diagnosis but also restrict the availability of visual resources for clinical genetic counseling. While prior work has shown that synthetic data can augment real datasets and preserve phenotype-level semantics, it remains unclear whether synthetic data alone is sufficient for learning in ultra-low-resource pediatric settings. In this work, we study the synthetic-only regime for pediatric rare disease recognition. Under a controlled experimental setup, models are trained exclusively on phenotype-aware synthetic facial images at increasing scales. We find that synthetic-only training achieves performance comparable to real-data-only baselines at sufficient scale across multiple backbones, suggesting that high-fidelity synthetic data can approximate clinically meaningful distributions. These findings together further enable the use of synthetic pediatric facial images as privacy-preserving resources for genetic education and counseling, supporting clinician training and patient communication. Our results highlight the potential of computer vision to improve data efficiency and expand accessible visual tools in children's healthcare.
△ Less
Submitted 21 May, 2026;
originally announced May 2026.
-
Quantum compressed sensing
Authors:
Jianyong Hu,
Wei Li,
Shuxiao Wu,
Liwen Zhang,
Yongchuang Sun,
Jiazhao Tian,
Guosheng Feng,
Zhixing Qiao,
Jianqiang Liu,
Changgang Yang,
Ruiyun Chen,
Chengbing Qin,
Guofeng Zhang,
Liantuan Xiao,
Suotang Jia
Abstract:
How many measurements are fundamentally required to capture a signal. Shannon's information theory established the bedrock of this question in 1948, the Nyquist Shannon theorem set the first answer, and compressed sensing (CS) rewrote it in 2006 by reducing the required measurement number to M = O(Klog(N/K)) for a K sparse signal. Here, we propose quantum compressed sensing (QCS), a paradigm that…
▽ More
How many measurements are fundamentally required to capture a signal. Shannon's information theory established the bedrock of this question in 1948, the Nyquist Shannon theorem set the first answer, and compressed sensing (CS) rewrote it in 2006 by reducing the required measurement number to M = O(Klog(N/K)) for a K sparse signal. Here, we propose quantum compressed sensing (QCS), a paradigm that reframes signal acquisition as a unitary quantum evolution. By encoding high dimensional signal information into a single quantum probe state, then introducing domain-alignment evolution,a physically realizable unitary transformation that maps the sparse basis directly onto the measurement basis. QCS executes the support-set search at the quantum level without consuming measurement trials. The logarithmic penalty vanishes, compressing the required measurement number from the classical bound to M =O(K) and reducing reconstruction from ill posed optimization to linear estimation. We experimentally validate QCS using frequency and time domain sparse signals, confirming that the measurement number scales linearly with sparsity and decouples entirely from the signal dimension. Our work provides a physical pathway toward ultimate information acquisition efficiency, with broad implications for sensing, imaging, and communication.
△ Less
Submitted 15 May, 2026;
originally announced May 2026.
-
Concurrency without Model Changes: Future-based Asynchronous Function Calling for LLMs
Authors:
Guangyu Feng,
Huanzhi Mao,
Prabal Dutta,
Joseph E. Gonzalez
Abstract:
Function calling, also known as tool use, is a core capability of modern LLM agents but is typically constrained by synchronous execution semantics. Under these semantics, LLM decoding is blocked until each function call completes, resulting in increasing end-to-end latency. In this work, we introduce AsyncFC, a pure execution-layer framework that decouples LLM decoding from function execution, en…
▽ More
Function calling, also known as tool use, is a core capability of modern LLM agents but is typically constrained by synchronous execution semantics. Under these semantics, LLM decoding is blocked until each function call completes, resulting in increasing end-to-end latency. In this work, we introduce AsyncFC, a pure execution-layer framework that decouples LLM decoding from function execution, enabling overlap between model decoding and function execution as well as inter-function parallelism when dependencies permit. AsyncFC layers over existing models and unmodified function implementations, requiring no fine-tuning or changes to the standard synchronous function-calling protocol. Across standard function-calling benchmarks and adapted software engineering benchmarks, AsyncFC significantly reduces end-to-end task completion time while preserving task accuracy. Furthermore, these results reveal that LLMs possess a native capability to reason over symbolic futures that represent unresolved execution results, enabling an asynchronous paradigm for model-tool interaction.
△ Less
Submitted 14 May, 2026;
originally announced May 2026.
-
Predictive and feedback signals differently shape the formation of group-level and individualized language representations
Authors:
Shuguang Yang,
Shaoyun Yu,
Xin Jiang,
Suiping Wang,
Gangyi Feng
Abstract:
Adults vary greatly in how effectively they learn a new language, but the signals driving the learning processes and individual differences remain unclear. Over seven days, we tracked behavioral learning and collected fMRI data from 102 adults as they learned an artificial language with corrective feedback. We trained matched transformer models with prediction, feedback, or combined objectives and…
▽ More
Adults vary greatly in how effectively they learn a new language, but the signals driving the learning processes and individual differences remain unclear. Over seven days, we tracked behavioral learning and collected fMRI data from 102 adults as they learned an artificial language with corrective feedback. We trained matched transformer models with prediction, feedback, or combined objectives and compared their internal representations to brain activity. Representations derived from the prediction-focused model accounted for the largest share of unique neural variance at the group level, despite the human task being feedback-based. Throughout model training, both objectives showed a shift in brain-model alignment from sensory to higher-order language and associative networks, indicating abstraction processing. Conversely, neural patterns related to the feedback model were most useful for predicting individual generalization outcomes on Day 7. These findings support a multi-signal model of adult language learning, in which prediction shapes a common neural learning architecture across learners, whereas feedback-related mechanisms better explain individual differences over time.
△ Less
Submitted 10 May, 2026;
originally announced May 2026.
-
The Stellar Abundances and Galactic Evolution Survey (SAGES). V. The First Data Release of the DDO51 Band
Authors:
Qiqian Zhang,
Zhou Fan,
Gang Zhao,
Kai Xiao,
Wei Wang,
Hongrui Gu,
Jie Zheng,
Jingkun Zhao,
Chun Li,
Yuqin Chen,
Haibo Yuan,
Haining Li,
Kefeng Tan,
Yihan Song,
Ali Luo,
Nan Song,
Yujuan Liu,
Yaqian Wu,
Ali Esamdin,
Hubiao Niu,
Jinzhong Liu,
Guojie Feng,
Yu Zhang
Abstract:
We present the first public data release of DDO51 band from the Stellar Abundances and Galactic Evolution Survey (SAGES), based on Nanshan One-meter Wide-field Telescope (NOWT) observations obtained between 2023 September and 2024 January. This release initiates the DDO51-band component of the survey, covering $\sim$ 2,500 deg$^2$ of the northern sky and including more than 10 million sources. The…
▽ More
We present the first public data release of DDO51 band from the Stellar Abundances and Galactic Evolution Survey (SAGES), based on Nanshan One-meter Wide-field Telescope (NOWT) observations obtained between 2023 September and 2024 January. This release initiates the DDO51-band component of the survey, covering $\sim$ 2,500 deg$^2$ of the northern sky and including more than 10 million sources. The DDO51 filter is centered near the \ion{Mg}{1}~$b$ triplet and the adjacent MgH feature, offering sensitivity to stellar surface gravity. The data reduction pipeline incorporates an improved astrometric solution anchored to Gaia DR3 and a photometric calibration strategy tied to synthetic photometry from Gaia XP spectra. These procedures yield a point-source depth of $\sim$18.9 mag at S/N$\sim$10 and an internal photometric precision $\approx$6-7 mmag at the bright end. A preliminary color--color analysis using Gaia broadband photometry confirms the expected sensitivity of the DDO51 band to stellar surface gravity, demonstrating a clear photometric separation between dwarf and giant sequences for late-type stars. This dataset, when combined with existing SAGES photometry in other bands, provides a crucial tool for disentangling the substructures of the Milky Way. All data products from this release upon publication will be available.
△ Less
Submitted 9 May, 2026;
originally announced May 2026.
-
Quantum Compressed Sensing Enables Image Classification with a Single Photon
Authors:
Yanshan Fan,
Jianyong Hu,
Shuxiao Wu,
Zhixing Qiao,
Guosheng Feng,
Changgang Yang,
Jianqiang Liu,
Ruiyun Chen,
Chengbing Qin,
Guofeng Zhang,
Liantuan Xiao,
Suotang Jia
Abstract:
Image classification is a core task of intelligent sensing, conventionally follows a sequential imaging then processing pipeline. However, redundant high-dimensional image reconstruction is inherently inefficient, especially in photon limited scenarios. Here we report a photon level image classification method using quantum compressed sensing, which reformulates the classification task as a sparse…
▽ More
Image classification is a core task of intelligent sensing, conventionally follows a sequential imaging then processing pipeline. However, redundant high-dimensional image reconstruction is inherently inefficient, especially in photon limited scenarios. Here we report a photon level image classification method using quantum compressed sensing, which reformulates the classification task as a sparse signal measurement problem directly oriented toward class labels. By exploiting the parallelism of photonic quantum superposition states, a single photon can be encoded the complete spatial information of a high-dimensional image. Through a diffractive deep neural network, we physically construct a dedicated measurement basis aligned with the class space, enabling signal-dependent adaptive compressive measurement. Ideally, our method can extract class information via a single quantum projective measurement, reducing the required number of measurements from the logarithmic scaling O(Klog(N/K)) of classical compressed sensing to the constant-order information-theoretic limit M = K = 1. Experimental results show that a classification accuracy of 69.0% can be achieved by using a single-photon detection event as the decision criterion, while it increases to 95.0% with four-photon detection events. This work demonstrates image classification at the energy efficiency limit and introduces a measurement as decision framework. It provides a foundation for intelligent sensing systems that operate under extreme photon budgets and harsh environments.
△ Less
Submitted 28 April, 2026;
originally announced April 2026.
-
R-CoT: A Reasoning-Layer Watermark via Redundant Chain-of-Thought in Large Language Models
Authors:
Ziming Zhang,
Li Li,
Guorui Feng,
Hanzhou Wu,
Xinpeng Zhang
Abstract:
Large language models (LLMs) are widely deployed in multiple scenarios due to reasoning capabilities. In order to prevent the models from being misused, watermarking is generally employed to ensure ownership. However, most existing watermarking methods rely on superficial modifications to the model's output distribution, rendering the watermark vulnerable to perturbation and removal. To overcome t…
▽ More
Large language models (LLMs) are widely deployed in multiple scenarios due to reasoning capabilities. In order to prevent the models from being misused, watermarking is generally employed to ensure ownership. However, most existing watermarking methods rely on superficial modifications to the model's output distribution, rendering the watermark vulnerable to perturbation and removal. To overcome this challenge, this paper introduces a reasoning-layer framework termed Redundant Chain-of-Thought (R-CoT), which embeds watermarks into the reasoning path. A dual-trajectory optimization mechanism based on GRPO enables the native and the watermark reasoning path to coexist within a shared parameter space, internalizing the watermark as a distinct reasoning policy. Therefore, the watermark is embedded into the model's stable reasoning path, avoiding the watermark failure caused by output-level perturbations. Experimental results show that, compared with existing methods, R-CoT achieves high watermark effectiveness and strong robustness. Under fine-tuning and other post-training operations, the true positive rate (TPR) consistently remains above 95%, exhibiting only marginal degradation.
△ Less
Submitted 28 April, 2026;
originally announced April 2026.
-
POLAR-PIC: A Holistic Framework for Matrixized PIC with Co-Designed Compute, Layout, and Communication
Authors:
Yizhuo Rao,
Xingjian Cui,
Shangzhi Pang,
Jiabin Xie,
Guangnan Feng,
Jinhui Wei,
Ziyan Zhang,
Languang Gao,
Zhenyu Wang,
Zhiguang Chen,
Yutong Lu
Abstract:
Particle-in-Cell (PIC) simulations are fundamental to plasma physics but often suffer from limited scalability due to particle-grid interaction bottlenecks and particle redistribution costs. Specifically, the particle-grid interaction computations have not taken full advantage of the emerging Matrix Processing Units (MPUs), the particle motion introduces irregular memory accesses, and the bulk-syn…
▽ More
Particle-in-Cell (PIC) simulations are fundamental to plasma physics but often suffer from limited scalability due to particle-grid interaction bottlenecks and particle redistribution costs. Specifically, the particle-grid interaction computations have not taken full advantage of the emerging Matrix Processing Units (MPUs), the particle motion introduces irregular memory accesses, and the bulk-synchronous redistribution further destroys long-term data locality thereby limiting parallel efficiency. To address these inefficiencies, we present POLAR-PIC, a co-designed framework for large-scale PIC simulations that (i) reformulates Field Interpolation into an MPU-friendly outer-product form, (ii) maintains a physically ordered particle layout to preserve memory contiguity, and (iii) overlaps particle communication with Deposition to hide redistribution overhead. The evaluation on the pilot system of an Exascale supercomputer demonstrates that POLAR-PIC accelerates the entire particle-processing phase by up to 10.9x in uniform plasma and 4.4x in real-world laser-ion acceleration scenarios compared to the native WarpX reference pipeline on LX2. Ablation studies reveal that the speedups achieved by Interpolation and Deposition are 8.0x and 13.2x, respectively, and the asynchronous communication design sustains a 99.1% overlap ratio. In cross-platform comparisons, POLAR-PIC achieves 13.2% of theoretical peak efficiency on the CPU-based LS system, while WarpX reaches 9.6% on NVIDIA A800 GPUs. Notably, the scalability evaluation demonstrates that POLAR-PIC maintains 67.5% weak scaling efficiency on over 2 million cores under high-migration dynamic workloads, highlighting the importance of holistic co-design for future matrix-centric HPC systems.
△ Less
Submitted 21 April, 2026;
originally announced April 2026.
-
Distributed Resilient Fixed-Time Control for Cooperative Output Regulation of MASs over Directed Graphs under DoS Attacks
Authors:
Wenji Cao,
Lu Liu,
Dan Zhang,
Gang Feng
Abstract:
This paper addresses the problem of fixed-time cooperative output regulation for linear multi-agent systems over directed graphs under denial-of-service attacks. A novel distributed resilient fixed-time controller is developed that comprises a distributed resilient fixed-time observer taking general directed graphs into consideration, and a distributed resilient fixed-time control law for each age…
▽ More
This paper addresses the problem of fixed-time cooperative output regulation for linear multi-agent systems over directed graphs under denial-of-service attacks. A novel distributed resilient fixed-time controller is developed that comprises a distributed resilient fixed-time observer taking general directed graphs into consideration, and a distributed resilient fixed-time control law for each agent. The proposed controller neither depends on Laplacian symmetry nor requires strong connectivity and detail-balanced condition, in contrast to existing distributed resilient fixed-time controllers. Under the proposed controller, the regulated outputs converge to zero in a fixed time with its upper bound independent of the initial states of the multi-agent system. Ultimately, the efficacy of the proposed controller is demonstrated via a simulation example.
△ Less
Submitted 14 April, 2026;
originally announced April 2026.
-
In-Place Test-Time Training
Authors:
Guhao Feng,
Shengjie Luo,
Kai Hua,
Ge Zhang,
Di He,
Wenhao Huang,
Tianle Cai
Abstract:
The static ``train then deploy" paradigm fundamentally limits Large Language Models (LLMs) from dynamically adapting their weights in response to continuous streams of new information inherent in real-world tasks. Test-Time Training (TTT) offers a compelling alternative by updating a subset of model parameters (fast weights) at inference time, yet its potential in the current LLM ecosystem is hind…
▽ More
The static ``train then deploy" paradigm fundamentally limits Large Language Models (LLMs) from dynamically adapting their weights in response to continuous streams of new information inherent in real-world tasks. Test-Time Training (TTT) offers a compelling alternative by updating a subset of model parameters (fast weights) at inference time, yet its potential in the current LLM ecosystem is hindered by critical barriers including architectural incompatibility, computational inefficiency and misaligned fast weight objectives for language modeling. In this work, we introduce In-Place Test-Time Training (In-Place TTT), a framework that seamlessly endows LLMs with Test-Time Training ability. In-Place TTT treats the final projection matrix of the ubiquitous MLP blocks as its adaptable fast weights, enabling a ``drop-in" enhancement for LLMs without costly retraining from scratch. Furthermore, we replace TTT's generic reconstruction objective with a tailored, theoretically-grounded objective explicitly aligned with the Next-Token-Prediction task governing autoregressive language modeling. This principled objective, combined with an efficient chunk-wise update mechanism, results in a highly scalable algorithm compatible with context parallelism. Extensive experiments validate our framework's effectiveness: as an in-place enhancement, it enables a 4B-parameter model to achieve superior performance on tasks with contexts up to 128k, and when pretrained from scratch, it consistently outperforms competitive TTT-related approaches. Ablation study results further provide deeper insights on our design choices. Collectively, our results establish In-Place TTT as a promising step towards a paradigm of continual learning in LLMs.
△ Less
Submitted 7 April, 2026;
originally announced April 2026.
-
RDFace: A Benchmark Dataset for Rare Disease Facial Image Analysis under Extreme Data Scarcity and Phenotype-Aware Synthetic Generation
Authors:
Ganlin Feng,
Yuxi Long,
Hafsa Ali,
Erin Lou,
Fahad Butt,
Qian Liu,
Yang Wang,
Pingzhao Hu
Abstract:
Rare diseases often manifest with distinctive facial phenotypes in children, offering valuable diagnostic cues for clinicians and AI-assisted screening systems. However, progress in this field is severely limited by the scarcity of curated, ethically sourced facial data and the high similarity among phenotypes across different conditions. To address these challenges, we introduce RDFace, a curated…
▽ More
Rare diseases often manifest with distinctive facial phenotypes in children, offering valuable diagnostic cues for clinicians and AI-assisted screening systems. However, progress in this field is severely limited by the scarcity of curated, ethically sourced facial data and the high similarity among phenotypes across different conditions. To address these challenges, we introduce RDFace, a curated benchmark dataset comprising 456 pediatric facial images spanning 103 rare genetic conditions (average 4.4 samples per condition). Each ethically verified image is paired with standardized metadata. RDFace enables the development and evaluation of data-efficient AI models for rare disease diagnosis under real-world low-data constraints. We benchmark multiple pretrained vision backbones using cross-validation and explore synthetic augmentation with DreamBooth and FastGAN. Generated images are filtered via facial landmark similarity to maintain phenotype fidelity and merged with real data, improving diagnostic accuracy by up to 13.7% in ultra-low-data regimes. To assess semantic validity, phenotype descriptions generated by a vision-language model from real and synthetic images achieve a report similarity score of 0.84. RDFace establishes a transparent, benchmark-ready dataset for equitable rare disease AI research and presents a scalable framework for evaluating both diagnostic performance and the integrity of synthetic medical imagery.
△ Less
Submitted 3 April, 2026;
originally announced April 2026.
-
MERIT: Memory-Enhanced Retrieval for Interpretable Knowledge Tracing
Authors:
Runze Li,
Kedi Chen,
Guwei Feng,
Mo Yu,
Jun Wang,
Wei Zhang
Abstract:
Knowledge Tracing (KT) models students' evolving knowledge states to predict future performance, serving as a foundation for personalized education. While traditional deep learning models achieve high accuracy, they often lack interpretability. Large Language Models (LLMs) offer strong reasoning capabilities but struggle with limited context windows and hallucinations. Furthermore, existing LLM-ba…
▽ More
Knowledge Tracing (KT) models students' evolving knowledge states to predict future performance, serving as a foundation for personalized education. While traditional deep learning models achieve high accuracy, they often lack interpretability. Large Language Models (LLMs) offer strong reasoning capabilities but struggle with limited context windows and hallucinations. Furthermore, existing LLM-based methods typically require expensive fine-tuning, limiting scalability and adaptability to new data. We propose MERIT (Memory-Enhanced Retrieval for Interpretable Knowledge Tracing), a training-free framework combining frozen LLM reasoning with structured pedagogical memory. Rather than updating parameters, MERIT transforms raw interaction logs into an interpretable memory bank. The framework uses semantic denoising to categorize students into latent cognitive schemas and constructs a paradigm bank where representative error patterns are analyzed offline to generate explicit Chain-of-Thought (CoT) rationales. During inference, a hierarchical routing mechanism retrieves relevant contexts, while a logic-augmented module applies semantic constraints to calibrate predictions. By grounding the LLM in interpretable memory, MERIT achieves state-of-the-art performance on real-world datasets without gradient updates. This approach reduces computational costs and supports dynamic knowledge updates, improving the accessibility and transparency of educational diagnosis.
△ Less
Submitted 3 March, 2026;
originally announced March 2026.
-
AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization
Authors:
Dailan He,
Guanlin Feng,
Xingtong Ge,
Yi Zhang,
Bingqi Ma,
Guanglu Song,
Yu Liu,
Hongsheng Li
Abstract:
Streaming autoregressive (AR) video generators combined with few-step distillation achieve low-latency, high-quality synthesis, yet remain difficult to align via reinforcement learning from human feedback (RLHF). Existing SDE-based GRPO methods face challenges in this setting: few-step ODEs and consistency model samplers deviate from standard flow-matching ODEs, and their short, low-stochasticity…
▽ More
Streaming autoregressive (AR) video generators combined with few-step distillation achieve low-latency, high-quality synthesis, yet remain difficult to align via reinforcement learning from human feedback (RLHF). Existing SDE-based GRPO methods face challenges in this setting: few-step ODEs and consistency model samplers deviate from standard flow-matching ODEs, and their short, low-stochasticity trajectories are highly sensitive to initialization noise, rendering intermediate SDE exploration ineffective. We propose AR-CoPO (AutoRegressive Contrastive Policy Optimization), a framework that adapts the Neighbor GRPO contrastive perspective to streaming AR generation. AR-CoPO introduces chunk-level alignment via a forking mechanism that constructs neighborhood candidates at a randomly selected chunk, assigns sequence-level rewards, and performs localized GRPO updates. We further propose a semi-on-policy training strategy that complements on-policy exploration with exploitation over a replay buffer of reference rollouts, improving generation quality across domains. Experiments on Self-Forcing demonstrate that AR-CoPO improves both out-of-domain generalization and in-domain human preference alignment over the baseline, providing evidence of genuine alignment rather than reward hacking.
△ Less
Submitted 29 June, 2026; v1 submitted 18 March, 2026;
originally announced March 2026.
-
\HI 21-cm Line Properties of the Nearby LIRG IRAS 04296+2923
Authors:
Guixiang Feng,
Zhongzu Wu,
Chuanpeng Zhang,
Ming Zhu
Abstract:
We present an analysis of archival Very Large Array (VLA) and Five-hundred-meter Aperture Spherical radio Telescope (FAST) \HI\ 21 cm data, together with archival multi-band radio continuum observations, of the nearby luminous infrared galaxy IRAS~04296+2923. The system, located behind the Taurus dark cloud at a distance of $\sim$29 Mpc, forms a small galaxy group consisting of five members as rev…
▽ More
We present an analysis of archival Very Large Array (VLA) and Five-hundred-meter Aperture Spherical radio Telescope (FAST) \HI\ 21 cm data, together with archival multi-band radio continuum observations, of the nearby luminous infrared galaxy IRAS~04296+2923. The system, located behind the Taurus dark cloud at a distance of $\sim$29 Mpc, forms a small galaxy group consisting of five members as revealed by the \HI\ imaging. IRAS~04296+2923 has a close companion, HI~0432+2926, with a projected separation of $\sim$40 kpc, a small line-of-sight velocity difference of $Δ$ v = 26 km s$^{-1}$, and comparable total \HI\ masses of order $10^{9}$~$M_{\odot}$. Both galaxies exhibit regular \HI\ velocity fields and characteristic double-horn profiles in the VLA and FAST data, accompanied by only subtle asymmetries and extended \HI\ structures, indicating rotation-dominated kinematics with early signs of weak tidal interaction. Radio continuum emission is detected only from IRAS~04296+2923 and is confined to its nuclear region, consistent with previous studies. Modeling of its multi-band radio spectrum reveals a significant contribution from free--free emission at high frequencies ($>$30 GHz) and a high FIR-to-radio flux ratio ($q_{8.4}\simeq3.2$), implying a young, dust-obscured nuclear starburst. Taken together, the regular \HI\ kinematics, the small velocity offset, and the group-scale environment favor an interpretation in which IRAS~04296+2923 and HI~0432+2926 form a gravitationally bound, orbiting galaxy pair embedded in a small group, rather than an advanced merger. In this context, the luminous infrared galaxy (LIRG) nature of IRAS~04296+2923 is more plausibly driven by internal processes, such as bar-induced gas inflow, possibly modulated by long-timescale, low-level tidal interactions with nearby group companions.
△ Less
Submitted 12 March, 2026;
originally announced March 2026.
-
Developing Foundation Models for Universal Segmentation from 3D Whole-Body Positron Emission Tomography
Authors:
Yichi Zhang,
Le Xue,
Wenbo Zhang,
Lanlan Li,
Feiyang Xiao,
Yuchen Liu,
Xiaohui Zhang,
Hongwei Zhang,
Shuqi Wang,
Gang Feng,
Liling Peng,
Xin Gao,
Yuanfan Xu,
Yuan Qi,
Kuangyu Shi,
Hong Zhang,
Yuan Cheng,
Mei Tian,
Zixin Hu
Abstract:
Positron emission tomography (PET) is a key nuclear medicine imaging modality that visualizes radiotracer distributions to quantify in vivo physiological and metabolic processes, playing an irreplaceable role in disease management. Despite its clinical importance, the development of deep learning models for quantitative PET image analysis remains severely limited, driven by both the inherent segme…
▽ More
Positron emission tomography (PET) is a key nuclear medicine imaging modality that visualizes radiotracer distributions to quantify in vivo physiological and metabolic processes, playing an irreplaceable role in disease management. Despite its clinical importance, the development of deep learning models for quantitative PET image analysis remains severely limited, driven by both the inherent segmentation challenge from PET's paucity of anatomical contrast and the high costs of data acquisition and annotation. To bridge this gap, we develop generalist foundational models for universal segmentation from 3D whole-body PET imaging. We first build the largest and most comprehensive PET dataset to date, comprising 11041 3D whole-body PET scans with 59831 segmentation masks for model development. Based on this dataset, we present SegAnyPET, an innovative foundational model with general-purpose applicability to diverse segmentation tasks. Built on a 3D architecture with a prompt engineering strategy for mask generation, SegAnyPET enables universal and scalable organ and lesion segmentation, supports efficient human correction with minimal effort, and enables a clinical human-in-the-loop workflow. Extensive evaluations on multi-center, multi-tracer, multi-disease datasets demonstrate that SegAnyPET achieves strong zero-shot performance across a wide range of segmentation tasks, highlighting its potential to advance the clinical applications of molecular imaging.
△ Less
Submitted 12 March, 2026;
originally announced March 2026.
-
How Well Do Current Speech Deepfake Detection Methods Generalize to the Real World?
Authors:
Daixian Li,
Jun Xue,
Yanzhen Ren,
Zhuolin Yi,
Yihuan Huang,
Guanxiang Feng,
Yi Chai
Abstract:
Recent advances in speech synthesis and voice conversion have greatly improved the naturalness and authenticity of generated audio. Meanwhile, evolving encoding, compression, and transmission mechanisms on social media platforms further obscure deepfake artifacts. These factors complicate reliable detection in real-world environments, underscoring the need for representative evaluation benchmarks.…
▽ More
Recent advances in speech synthesis and voice conversion have greatly improved the naturalness and authenticity of generated audio. Meanwhile, evolving encoding, compression, and transmission mechanisms on social media platforms further obscure deepfake artifacts. These factors complicate reliable detection in real-world environments, underscoring the need for representative evaluation benchmarks. To this end, we introduce ML-ITW (Multilingual In-The-Wild), a multilingual dataset covering 14 languages, seven major platforms, and 180 public figures, totaling 28.39 hours of audio. We evaluate three detection paradigms: end-to-end neural models, self-supervised feature-based (SSL) methods, and audio large language models (Audio LLMs). Experimental results reveal significant performance degradation across diverse languages and real-world acoustic conditions, highlighting the limited generalization ability of existing detectors in practical scenarios. The ML-ITW dataset is publicly available.
△ Less
Submitted 5 March, 2026;
originally announced March 2026.
-
Clinical-Injection Transformer with Domain-Adapted MAE for Lupus Nephritis Prognosis Prediction
Authors:
Yuewen Huang,
Zhitao Ye,
Guangnan Feng,
Fudan Zheng,
Xia Gao,
Yutong Lu
Abstract:
Lupus nephritis (LN) is a severe complication of systemic lupus erythematosus that affects pediatric patients with significantly greater severity and worse renal outcomes compared to adults. Despite the urgent clinical need, predicting pediatric LN prognosis remains unexplored in computational pathology. Furthermore, the only existing histopathology-based approach for LN relies on multiple costly…
▽ More
Lupus nephritis (LN) is a severe complication of systemic lupus erythematosus that affects pediatric patients with significantly greater severity and worse renal outcomes compared to adults. Despite the urgent clinical need, predicting pediatric LN prognosis remains unexplored in computational pathology. Furthermore, the only existing histopathology-based approach for LN relies on multiple costly staining protocols and fails to integrate complementary clinical data. To address these gaps, we propose the first multimodal computational pathology framework for three-class treatment response prediction (complete remission, partial response, and no response) in pediatric LN, utilizing only routine PAS-stained biopsies and structured clinical data. Our framework introduces two key methodological innovations. First, a Clinical-Injection Transformer (CIT) embeds clinical features as condition tokens into patch-level self-attention, facilitating implicit and bidirectional cross-modal interactions within a unified attention space. Second, we design a decoupled representation-knowledge adaptation strategy using a domain-adapted Masked Autoencoder (MAE). This strategy explicitly separates self-supervised morphological feature learning from pathological knowledge extraction. Additionally, we introduce a multi-granularity morphological type injection mechanism to bridge distilled classification knowledge with downstream prognostic predictions at both the instance and patient levels. Evaluated on a cohort of 71 pediatric LN patients with KDIGO-standardized labels, our method achieves a three-class accuracy of 90.1% and an AUC of 89.4%, demonstrating its potential as a highly accurate and cost-effective prognostic tool.
△ Less
Submitted 3 March, 2026;
originally announced March 2026.
-
veScale-FSDP: Flexible and High-Performance FSDP at Scale
Authors:
Zezhou Wang,
Youjie Li,
Zhiqi Lin,
Jiacheng Yang,
Cong Xie,
Guanyu Feng,
Zheng Zhong,
Ziyue Huang,
Hongyu Zhu,
Zhi Zhang,
Yanghua Peng,
Xin Liu
Abstract:
Fully Sharded Data Parallel (FSDP), also known as Zero Redundancy Optimizer (ZeRO), is widely used for large-scale model training, because of its memory efficiency and minimal intrusion on model code. However, existing FSDP systems rely on fixed element-wise or row-wise sharding formats that conflict with block-structured computations. As a result, they struggle to support modern structure-aware t…
▽ More
Fully Sharded Data Parallel (FSDP), also known as Zero Redundancy Optimizer (ZeRO), is widely used for large-scale model training, because of its memory efficiency and minimal intrusion on model code. However, existing FSDP systems rely on fixed element-wise or row-wise sharding formats that conflict with block-structured computations. As a result, they struggle to support modern structure-aware training methods, including block-wise quantization and non-element-wise optimizers such as Shampoo and Muon. In addition, today's implementations incur communication and memory overheads that degrade efficiency at the scale of tens of thousands of GPUs. We introduce veScale-FSDP, a novel FSDP system that combines RaggedShard, a flexible sharding format, with a structure-aware planning algorithm to deliver both flexibility and performance. veScale-FSDP enables zero-copy FSDP communications and natively supports block-wise quantization and non-element-wise optimizers, achieving 5% to 66% higher throughput and 16% to 30% lower memory usage than existing FSDP systems, while scaling efficiently to tens of thousands of GPUs.
△ Less
Submitted 21 April, 2026; v1 submitted 25 February, 2026;
originally announced February 2026.
-
Unlocking Multimodal Document Intelligence: From Current Triumphs to Future Frontiers of Visual Document Retrieval
Authors:
Yibo Yan,
Jiahao Huo,
Guanbo Feng,
Mingdong Ou,
Yi Cao,
Xin Zou,
Shuliang Liu,
Yuanhuiyi Lyu,
Yu Huang,
Jungang Li,
Kening Zheng,
Xu Zheng,
Philip S. Yu,
James Kwok,
Xuming Hu
Abstract:
With the rapid proliferation of multimodal information, Visual Document Retrieval (VDR) has emerged as a critical frontier in bridging the gap between unstructured visually rich data and precise information acquisition. Unlike traditional natural image retrieval, visual documents exhibit unique characteristics defined by dense textual content, intricate layouts, and fine-grained semantic dependenc…
▽ More
With the rapid proliferation of multimodal information, Visual Document Retrieval (VDR) has emerged as a critical frontier in bridging the gap between unstructured visually rich data and precise information acquisition. Unlike traditional natural image retrieval, visual documents exhibit unique characteristics defined by dense textual content, intricate layouts, and fine-grained semantic dependencies. This paper presents the first comprehensive survey of the VDR landscape, specifically through the lens of the Multimodal Large Language Model (MLLM) era. We begin by examining the benchmark landscape, and subsequently dive into the methodological evolution, categorizing approaches into three primary aspects: multimodal embedding models, multimodal reranker models, and the integration of Retrieval-Augmented Generation (RAG) and Agentic systems for complex document intelligence. Finally, we identify persistent challenges and outline promising future directions, aiming to provide a clear roadmap for future multimodal document intelligence.
△ Less
Submitted 22 March, 2026; v1 submitted 23 February, 2026;
originally announced February 2026.
-
Universality Reconsidered: Rethinking the Validation of Foundation Models for General-Purpose 3D Medical Segmentation
Authors:
Yichi Zhang,
Le Xue,
Feiyang Xiao,
Wenbo Zhang,
Gang Feng,
Chenguang Zheng,
Yuan Qi,
Yuan Cheng,
Zixin Hu
Abstract:
Foundation models have emerged as a transformative paradigm in 3D medical imaging, with the promise of unified quantitative analysis across diverse targets and imaging modalities. Yet the prevailing conception of universality remains incomplete. Current models are predominantly developed and evaluated on datasets largely concentrated around a limited set of imaging modalities and anatomical region…
▽ More
Foundation models have emerged as a transformative paradigm in 3D medical imaging, with the promise of unified quantitative analysis across diverse targets and imaging modalities. Yet the prevailing conception of universality remains incomplete. Current models are predominantly developed and evaluated on datasets largely concentrated around a limited set of imaging modalities and anatomical regions. In this Perspective, we evaluate representative 3D segmentation foundation models using paired whole-body structural and functional imaging data. Our analysis reveals a substantial gap between benchmark-reported performance and real-world generalization, with marked degradation on previously unseen data and particularly severe failures on functional imaging modalities. These findings suggest that current foundation models remain far from achieving true universality. We argue that progress requires not only scaling models and datasets, but also a reconsideration of how universality is defined and validated, extending evaluation beyond regional structural benchmarks toward whole-body structural and functional imaging. Our observations highlight the need to distinguish benchmark success from genuine clinical generalization. Bridging this gap will be essential for translating foundation models from controlled evaluation settings to real-world medical practice.
△ Less
Submitted 22 July, 2026; v1 submitted 7 February, 2026;
originally announced February 2026.
-
Pruning for Generalization: A Transfer-Oriented Spatiotemporal Graph Framework
Authors:
Zihao Jing,
Yuxi Long,
Ganlin Feng
Abstract:
Multivariate time series forecasting in graph-structured domains is critical for real-world applications, yet existing spatiotemporal models often suffer from performance degradation under data scarcity and cross-domain shifts. We address these challenges through the lens of structure-aware context selection. We propose TL-GPSTGN, a transfer-oriented spatiotemporal framework that enhances sample e…
▽ More
Multivariate time series forecasting in graph-structured domains is critical for real-world applications, yet existing spatiotemporal models often suffer from performance degradation under data scarcity and cross-domain shifts. We address these challenges through the lens of structure-aware context selection. We propose TL-GPSTGN, a transfer-oriented spatiotemporal framework that enhances sample efficiency and out-of-distribution generalization by selectively pruning non-optimized graph context. Specifically, our method employs information-theoretic and correlation-based criteria to extract structurally informative subgraphs and features, resulting in a compact, semantically grounded representation. This optimized context is subsequently integrated into a spatiotemporal convolutional architecture to capture complex multivariate dynamics. Evaluations on large-scale traffic benchmarks demonstrate that TL-GPSTGN consistently outperforms baselines in low-data transfer scenarios. Our findings suggest that explicit context pruning serves as a powerful inductive bias for improving the robustness of graph-based forecasting models.
△ Less
Submitted 3 February, 2026;
originally announced February 2026.
-
Drinfeld Isomorphism for Novel Quantum Affine Algebra of Type $A_{1}^{(1)}$
Authors:
Rushu Zhuang,
Ge Feng,
Naihong Hu
Abstract:
In this paper, we first review the definition of the novel quantum affine algebra \(U_{\textbf{q}}(\widehat{\mathfrak{sl}}_2)\) of type \(A_{1}^{(1)}\) given in \cite{FHZ, HZhuang}. Furthermore, by introducing \(Ω\)-invariant generating functions, we construct the Drinfeld realization \(U^{D}_{\textbf{q}}(\widehat{\mathfrak{sl}}_2)\) of this algebra, and prove that \(U_{\textbf{q}}(\widehat{\mathf…
▽ More
In this paper, we first review the definition of the novel quantum affine algebra \(U_{\textbf{q}}(\widehat{\mathfrak{sl}}_2)\) of type \(A_{1}^{(1)}\) given in \cite{FHZ, HZhuang}. Furthermore, by introducing \(Ω\)-invariant generating functions, we construct the Drinfeld realization \(U^{D}_{\textbf{q}}(\widehat{\mathfrak{sl}}_2)\) of this algebra, and prove that \(U_{\textbf{q}}(\widehat{\mathfrak{sl}}_2)\) and \(U^{D}_{\textbf{q}}(\widehat{\mathfrak{sl}}_2)\) are algebraically isomorphic, which is known as the Drinfeld Isomorphism.
△ Less
Submitted 28 January, 2026;
originally announced January 2026.
-
Selecting and Testing Asset Pricing Models: A Stepwise Approach
Authors:
Guanhao Feng,
Wei Lan,
Hansheng Wang,
Jun Zhang
Abstract:
The asset pricing literature emphasizes factor models that minimize pricing errors but overlooks unselected candidate factors that could enhance the performance of test assets. This paper proposes a framework for factor model selection and testing by (i) selecting the optimal model that spans the joint efficient frontier of test assets and all candidate factors, and (ii) testing pricing performanc…
▽ More
The asset pricing literature emphasizes factor models that minimize pricing errors but overlooks unselected candidate factors that could enhance the performance of test assets. This paper proposes a framework for factor model selection and testing by (i) selecting the optimal model that spans the joint efficient frontier of test assets and all candidate factors, and (ii) testing pricing performance on both test assets and unselected candidate factors. Our framework updates a baseline model (e.g., CAPM) sequentially by adding or removing factors based on asset pricing tests. Ensuring model selection consistency, our framework utilizes the asset pricing duality: minimizing cross-sectionally unexplained pricing errors aligns with maximizing the Sharpe ratio of the selected factor model. Empirical evidence shows that workhorse factor models fail asset pricing tests, whereas our proposed 8-factor model is not rejected and exhibits robust out-of-sample performance.
△ Less
Submitted 15 January, 2026;
originally announced January 2026.
-
Matrix-PIC: Harnessing Matrix Outer-product for High-Performance Particle-in-Cell Simulations
Authors:
Yizhuo Rao,
Xingjian Cui,
Jiabin Xie,
Shangzhi Pang,
Guangnan Feng,
Jinhui Wei,
Zhiguang Chen,
Yutong Lu
Abstract:
Particle-in-Cell (PIC) simulations spend most of their execution time on particle--grid interactions, where fine-grained atomic updates become a major bottleneck on traditional many-core CPUs. Recent CPU architectures integrate specialized Matrix Processing Units (MPUs) that efficiently support matrix outer-product operations, offering new opportunities to overcome this limitation. Leveraging this…
▽ More
Particle-in-Cell (PIC) simulations spend most of their execution time on particle--grid interactions, where fine-grained atomic updates become a major bottleneck on traditional many-core CPUs. Recent CPU architectures integrate specialized Matrix Processing Units (MPUs) that efficiently support matrix outer-product operations, offering new opportunities to overcome this limitation. Leveraging this architectural shift, this work focuses on redesigning the current deposition step of PIC simulations under a matrix-centric execution model.
We present MatrixPIC, the first holistic co-design of the deposition kernel, data layout, and incremental particle sorting tailored to the hybrid MPU--VPU SIMD model on modern CPUs. MatrixPIC introduces: (i)~a block-matrix formulation of the current deposition algorithm that maps naturally to MPU outer-product primitives; (ii)~a hybrid execution pipeline that combines MPU-based high-density accumulation with VPU-based data preparation and control flow; and (iii)~an $O(1)$-amortized incremental sorter based on a gapped packed-memory array to preserve data locality for efficient MPU execution.
Evaluated on a next-generation HPC platform, MatrixPIC achieves significant performance gains. In Laser-Wakefield Acceleration (LWFA) simulations, it delivers up to $2.63\times$ speedup in total runtime. For third-order deposition, the core kernel is accelerated by $8.7\times$ over the baseline and $2.0\times$ over the best hand-optimized VPU implementation. Moreover, MatrixPIC reaches $83.08\%$ of theoretical CPU peak performance, nearly $2.8\times$ higher than a highly optimized CUDA kernel on a data center GPU. These results demonstrate the effectiveness of matrix-oriented co-design for accelerating PIC simulations on emerging CPU architectures.
△ Less
Submitted 13 January, 2026;
originally announced January 2026.
-
How Would Oblivious Memory Boost Graph Analytics on Trusted Processors?
Authors:
Jiping Yu,
Xiaowei Zhu,
Kun Chen,
Guanyu Feng,
Yunyi Chen,
Xiaoyu Fan,
Wenguang Chen
Abstract:
Trusted processors provide a way to perform joint computations while preserving data privacy. To overcome the performance degradation caused by data-oblivious algorithms to prevent information leakage, we explore the benefits of oblivious memory (OM) integrated in processors, to which the accesses are unobservable by adversaries. We focus on graph analytics, an important application vulnerable to…
▽ More
Trusted processors provide a way to perform joint computations while preserving data privacy. To overcome the performance degradation caused by data-oblivious algorithms to prevent information leakage, we explore the benefits of oblivious memory (OM) integrated in processors, to which the accesses are unobservable by adversaries. We focus on graph analytics, an important application vulnerable to access-pattern attacks. With a co-design between storage structure and algorithms, our prototype system is 100x faster than baselines given an OM sized around the per-core cache which can be implemented on existing processors with negligible overhead. This gives insights into equipping trusted processors with OM.
△ Less
Submitted 30 December, 2025;
originally announced December 2025.
-
Panel Coupled Matrix-Tensor Clustering Model with Applications to Asset Pricing
Authors:
Liyuan Cui,
Guanhao Feng,
Yuefeng Han,
Jiayan Li
Abstract:
We tackle the challenge of estimating grouping structures and factor loadings in asset pricing models, where traditional regressions struggle due to sparse data and high noise. Existing approaches, such as those using fused penalties and multi-task learning, often enforce coefficient homogeneity across cross-sectional units, reducing flexibility. Clustering methods (e.g., spectral clustering, Lloy…
▽ More
We tackle the challenge of estimating grouping structures and factor loadings in asset pricing models, where traditional regressions struggle due to sparse data and high noise. Existing approaches, such as those using fused penalties and multi-task learning, often enforce coefficient homogeneity across cross-sectional units, reducing flexibility. Clustering methods (e.g., spectral clustering, Lloyd's algorithm) achieve consistent recovery under specific conditions but typically rely on a single data source. To address these limitations, we introduce the Panel Coupled Matrix-Tensor Clustering (PMTC) model, which simultaneously leverages a characteristics tensor and a return matrix to identify latent asset groups. By integrating these data sources, we develop computationally efficient tensor clustering algorithms that enhance both clustering accuracy and factor loading estimation. Simulations demonstrate that our methods outperform single-source alternatives in clustering accuracy and coefficient estimation, particularly under moderate signal-to-noise conditions. Empirical application to U.S. equities demonstrates the practical value of PMTC, yielding higher out-of-sample total $R^2$ and economically interpretable variation in factor exposures.
△ Less
Submitted 29 December, 2025;
originally announced December 2025.
-
SN 2022ngb: A faint, slowly evolving Type IIb supernova with a low-mass envelope
Authors:
J. -W. Zhao,
S. Benetti,
Y. -Z. Cai,
A. Pastorello,
N. Elias-Rosa,
A. Reguitti,
G. Valerin,
Z. -Y. Wang,
E. Cappellaro,
G. -F. Feng,
A. Fiore,
B. Fitzpatrick,
M. Fraser,
J. Isern,
E. Kankare,
T. Kravtsov,
B. Kumar,
P. Lundqvist,
K. Matilainen,
S. Mattila,
P. A. Mazzali,
S. Moran,
P. Ochner,
Z. -H. Peng,
T. M. Reynolds
, et al. (13 additional authors not shown)
Abstract:
An extensive photometric and spectroscopic follow-up campaign of the Type IIb SN 2022ngb is presented in the article. Through detailed modeling of this dataset, we aim to constrain the key physical parameters of the explosion, infer the nature of the progenitor star and its environment, and probe the dynamical properties of the ejecta. We analyze photometric and spectroscopic data of SN 2022ngb. B…
▽ More
An extensive photometric and spectroscopic follow-up campaign of the Type IIb SN 2022ngb is presented in the article. Through detailed modeling of this dataset, we aim to constrain the key physical parameters of the explosion, infer the nature of the progenitor star and its environment, and probe the dynamical properties of the ejecta. We analyze photometric and spectroscopic data of SN 2022ngb. By constructing and modeling the bolometric light curve with semi-analytic models, we estimate the primary explosion parameters. The spectroscopic data are compared with those of well-studied SNe IIb and NLTE models to constrain the properties of the progenitor and the structure of the resulting ejecta. SN 2022ngb is a low-luminosity SN IIb with a peak bolometric luminosity of L_bol = 7.76 (+1.15/-1.00) x 10^41 erg/s and a V-band rising time of 24.32 +/- 0.50 days. Light curve modeling indicates an ejecta mass of ~2.9-3.2 M_sun, an explosion energy of ~1.4 x 10^51 erg, and a low synthesized 56Ni mass of ~0.045 M_sun. Nebular phase spectra exhibit asymmetric line profiles, pointing to a non-spherical explosion and an anisotropic distribution of radioactive material. Our analysis reveals a relatively compact stripped-envelope progenitor with a pre-SN mass of approximately 4.7 M_sun (corresponding to a 15-16 M_sun ZAMS star). Our analysis suggests that SN 2022ngb originated from the explosion of a moderate-mass relatively compact, stripped-envelope star in a binary system. The asymmetries inferred from the nebular phase spectral line features suggest a non-spherical explosion.
△ Less
Submitted 24 February, 2026; v1 submitted 10 December, 2025;
originally announced December 2025.
-
AuthenLoRA: Entangling Stylization with Imperceptible Watermarks for Copyright-Secure LoRA Adapters
Authors:
Fangming Shi,
Li Li,
Kejiang Chen,
Guorui Feng,
Xinpeng Zhang
Abstract:
Low-Rank Adaptation (LoRA) offers an efficient paradigm for customizing diffusion models, but its ease of redistribution raises concerns over unauthorized use and the generation of untraceable content. Existing watermarking techniques either target base models or verify LoRA modules themselves, yet they fail to propagate watermarks to generated images, leaving a critical gap in traceability. Moreo…
▽ More
Low-Rank Adaptation (LoRA) offers an efficient paradigm for customizing diffusion models, but its ease of redistribution raises concerns over unauthorized use and the generation of untraceable content. Existing watermarking techniques either target base models or verify LoRA modules themselves, yet they fail to propagate watermarks to generated images, leaving a critical gap in traceability. Moreover, traceability watermarking designed for base models is not tightly coupled with stylization and often introduces visual degradation or high false-positive detection rates. To address these limitations, we propose AuthenLoRA, a unified watermarking framework that embeds imperceptible, traceable watermarks directly into the LoRA training process while preserving stylization quality. AuthenLoRA employs a dual-objective optimization strategy that jointly learns the target style distribution and the watermark-induced distribution shift, ensuring that any image generated with the watermarked LoRA reliably carries the watermark. We further design an expanded LoRA architecture for enhanced multi-scale adaptation and introduce a zero-message regularization mechanism that substantially reduces false positives during watermark verification. Extensive experiments demonstrate that AuthenLoRA achieves high-fidelity stylization, robust watermark propagation, and significantly lower false-positive rates compared with existing approaches. Open-source implementation is available at: https://github.com/ShiFangming0823/AuthenLoRA
△ Less
Submitted 26 November, 2025;
originally announced November 2025.
-
Neighbor GRPO: Contrastive ODE Policy Optimization Aligns Flow Models
Authors:
Dailan He,
Guanlin Feng,
Xingtong Ge,
Yazhe Niu,
Yi Zhang,
Bingqi Ma,
Guanglu Song,
Yu Liu,
Hongsheng Li
Abstract:
Group Relative Policy Optimization (GRPO) has shown promise in aligning image and video generative models with human preferences. However, applying it to modern flow matching models is challenging because of its deterministic sampling paradigm. Current methods address this issue by converting Ordinary Differential Equations (ODEs) to Stochastic Differential Equations (SDEs), which introduce stocha…
▽ More
Group Relative Policy Optimization (GRPO) has shown promise in aligning image and video generative models with human preferences. However, applying it to modern flow matching models is challenging because of its deterministic sampling paradigm. Current methods address this issue by converting Ordinary Differential Equations (ODEs) to Stochastic Differential Equations (SDEs), which introduce stochasticity. However, this SDE-based GRPO suffers from issues of inefficient credit assignment and incompatibility with high-order solvers for fewer-step sampling. In this paper, we first reinterpret existing SDE-based GRPO methods from a distance optimization perspective, revealing their underlying mechanism as a form of contrastive learning. Based on this insight, we propose Neighbor GRPO, a novel alignment algorithm that completely bypasses the need for SDEs. Neighbor GRPO generates a diverse set of candidate trajectories by perturbing the initial noise conditions of the ODE and optimizes the model using a softmax distance-based surrogate leaping policy. We establish a theoretical connection between this distance-based objective and policy gradient optimization, rigorously integrating our approach into the GRPO framework. Our method fully preserves the advantages of deterministic ODE sampling, including efficiency and compatibility with high-order solvers. We further introduce symmetric anchor sampling for computational efficiency and group-wise quasi-norm reweighting to address reward flattening. Extensive experiments demonstrate that Neighbor GRPO significantly outperforms SDE-based counterparts in terms of training cost, convergence speed, and generation quality.
△ Less
Submitted 18 March, 2026; v1 submitted 21 November, 2025;
originally announced November 2025.
-
Multi-network Topology Underlying Individual Language Learning Success
Authors:
Peilun Song,
Shuguang Yang,
Xiujuan Geng,
Zhenzhong Gan,
Suiping Wang,
Gangyi Feng
Abstract:
Adult language learning varies greatly among individuals. Traditionally associated with frontotemporal language regions, this variability is increasingly seen as stemming from distributed brain networks. However, the role of these networks and their topological organization in explaining these differences remains unclear. We hypothesize that graph-theory-based network analysis of intrinsic multimo…
▽ More
Adult language learning varies greatly among individuals. Traditionally associated with frontotemporal language regions, this variability is increasingly seen as stemming from distributed brain networks. However, the role of these networks and their topological organization in explaining these differences remains unclear. We hypothesize that graph-theory-based network analysis of intrinsic multimodal connectivities across multiple networks explains overall and component-specific variations in language learning. We tested this in 101 healthy adults who underwent resting-state fMRI, structural MRI, and diffusion tensor imaging before seven days of six artificial language training tasks. We identified one dominant general learning component shared across tasks and five task-specific ones. Cross-validated predictive models used multimodal multi-network graph-theoretic metrics to predict final learning outcomes (LO) and rates (LR). We significantly predicted the LO and LR of the general component, which were primarily contributed by dorsal attention and frontoparietal networks. Nodal local efficiency was the most consistent predictor, with additional contributions from node clustering coefficient and network centrality for LR, highlighting local robustness, mesoscale network segregation, and global influence in explaining individual differences. Only task-specific word learning LO was predictable, relying on default mode and frontoparietal hubs with high betweenness centrality and efficiency. These findings demonstrate that intrinsic network topologies underlie differences in language learning success, supporting a multiple-systems hypothesis in which attentional-control networks interact with default and subcortical systems to shape learning trajectories. This advances mechanistic understanding and paves the way for personalized language education.
△ Less
Submitted 18 November, 2025;
originally announced November 2025.
-
CLIP4VI-ReID: Learning Modality-shared Representations via CLIP Semantic Bridge for Visible-Infrared Person Re-identification
Authors:
Xiaomei Yang,
Xizhan Gao,
Sijie Niu,
Fa Zhu,
Guang Feng,
Xiaofeng Qu,
David Camacho
Abstract:
This paper proposes a novel CLIP-driven modality-shared representation learning network named CLIP4VI-ReID for VI-ReID task, which consists of Text Semantic Generation (TSG), Infrared Feature Embedding (IFE), and High-level Semantic Alignment (HSA). Specifically, considering the huge gap in the physical characteristics between natural images and infrared images, the TSG is designed to generate tex…
▽ More
This paper proposes a novel CLIP-driven modality-shared representation learning network named CLIP4VI-ReID for VI-ReID task, which consists of Text Semantic Generation (TSG), Infrared Feature Embedding (IFE), and High-level Semantic Alignment (HSA). Specifically, considering the huge gap in the physical characteristics between natural images and infrared images, the TSG is designed to generate text semantics only for visible images, thereby enabling preliminary visible-text modality alignment. Then, the IFE is proposed to rectify the feature embeddings of infrared images using the generated text semantics. This process injects id-related semantics into the shared image encoder, enhancing its adaptability to the infrared modality. Besides, with text serving as a bridge, it enables indirect visible-infrared modality alignment. Finally, the HSA is established to refine the high-level semantic alignment. This process ensures that the fine-tuned text semantics only contain id-related information, thereby achieving more accurate cross-modal alignment and enhancing the discriminability of the learned modal-shared representations. Extensive experimental results demonstrate that the proposed CLIP4VI-ReID achieves superior performance than other state-of-the-art methods on some widely used VI-ReID datasets.
△ Less
Submitted 3 August, 2026; v1 submitted 13 November, 2025;
originally announced November 2025.
-
Realization of Thread Level Parallelism on Quantum Devices
Authors:
Keren Li,
Zidong Lin,
Zheng An,
Guanru Feng,
Zipeng Wu,
Shiyao Hou,
Jingen Xiang
Abstract:
Scaling up quantum devices is a central challenge for realizing practical quantum computation. Modular quantum architectures promise scalability, yet experiments to date have relied on either $\sim\!10^{3}$-qubit monolithic chips or fragile interconnects with high loss. Here, we introduce a classical linkage scheme that merges multiple independent quantum processing units (QPUs) into a single logi…
▽ More
Scaling up quantum devices is a central challenge for realizing practical quantum computation. Modular quantum architectures promise scalability, yet experiments to date have relied on either $\sim\!10^{3}$-qubit monolithic chips or fragile interconnects with high loss. Here, we introduce a classical linkage scheme that merges multiple independent quantum processing units (QPUs) into a single logical device, enabling thread-level parallelism (TLP). Theoretically, we show that quantum routines with product-state inputs and low-rank entangling layers can be re-expressed in an efficient parallelizable form. Experimentally, we validate this architecture on clusters comprising up to sixteen benchtop nuclear magnetic resonance (NMR) quantum nodes. A four-qubit Greenberger-Horne-Zeilinger (GHZ) state is partitioned into parallel two-qubit subcircuits, achieving a fidelity of $93.8\,\%$ with respect to the ideal state. A non-Hermitian evolution, implemented via a truncated Cauchy integral on Hermitian Hamiltonians, reproduces exact observables with high accuracy. Our results demonstrate that classical links suffice to scale up the logical size of quantum computations and realize general, non-unitary channels on today's hardware, opening an experimentally accessible route toward software-defined, clustered quantum accelerators.
△ Less
Submitted 7 November, 2025;
originally announced November 2025.
-
PETWB-REP: A Multi-Cancer Whole-Body FDG PET/CT and Radiology Report Dataset for Medical Imaging Research
Authors:
Le Xue,
Gang Feng,
Wenbo Zhang,
Yichi Zhang,
Lanlan Li,
Shuqi Wang,
Liling Peng,
Sisi Peng,
Xin Gao
Abstract:
Publicly available, large-scale medical imaging datasets are crucial for developing and validating artificial intelligence models and conducting retrospective clinical research. However, datasets that combine functional and anatomical imaging with detailed clinical reports across multiple cancer types remain scarce. Here, we present PETWB-REP, a curated dataset comprising whole-body 18F-Fluorodeox…
▽ More
Publicly available, large-scale medical imaging datasets are crucial for developing and validating artificial intelligence models and conducting retrospective clinical research. However, datasets that combine functional and anatomical imaging with detailed clinical reports across multiple cancer types remain scarce. Here, we present PETWB-REP, a curated dataset comprising whole-body 18F-Fluorodeoxyglucose (FDG) Positron Emission Tomography/Computed Tomography (PET/CT) scans and corresponding radiology reports from 490 patients diagnosed with various malignancies. The dataset primarily includes common cancers such as lung cancer, liver cancer, breast cancer, prostate cancer, and ovarian cancer. This dataset includes paired PET and CT images, de-identified textual reports, and structured clinical metadata. It is designed to support research in medical imaging, radiomics, artificial intelligence, and multi-modal learning.
△ Less
Submitted 5 November, 2025;
originally announced November 2025.
-
Interpretable Model-Aware Counterfactual Explanations for Random Forest
Authors:
Joshua S. Harvey,
Guanchao Feng,
Sai Anusha Meesala,
Tina Zhao,
Dhagash Mehta
Abstract:
Despite their enormous predictive power, machine learning models are often unsuitable for applications in regulated industries such as finance, due to their limited capacity to provide explanations. While model-agnostic frameworks such as Shapley values have proved to be convenient and popular, they rarely align with the kinds of causal explanations that are typically sought after. Counterfactual…
▽ More
Despite their enormous predictive power, machine learning models are often unsuitable for applications in regulated industries such as finance, due to their limited capacity to provide explanations. While model-agnostic frameworks such as Shapley values have proved to be convenient and popular, they rarely align with the kinds of causal explanations that are typically sought after. Counterfactual case-based explanations, where an individual is informed of which circumstances would need to be different to cause a change in outcome, may be more intuitive and actionable. However, finding appropriate counterfactual cases is an open challenge, as is interpreting which features are most critical for the change in outcome. Here, we pose the question of counterfactual search and interpretation in terms of similarity learning, exploiting the representation learned by the random forest predictive model itself. Once a counterfactual is found, the feature importance of the explanation is computed as a function of which random forest partitions are crossed in order to reach it from the original instance. We demonstrate this method on both the MNIST hand-drawn digit dataset and the German credit dataset, finding that it generates explanations that are sparser and more useful than Shapley values.
△ Less
Submitted 31 October, 2025;
originally announced October 2025.
-
Epitaxial Electrodeposition of Fe with Controlled In-Plane Variants for Reversible Metal Anode in Aqueous Electrolyte
Authors:
Chenxi Sui,
Ching-Tai Fu,
Guangxia Feng,
Yuqi Li,
Junyan Li,
Gangbin Yan,
Po-Chun Hsu,
Steven Chu,
Yi Cui
Abstract:
The development of reversible metal anodes is a key challenge for advancing aqueous battery technologies, particularly for scalable and safe stationary energy storage applications. Here we demonstrate a strategy to realize epitaxial electrodeposition of iron (Fe) on single-crystal copper (Cu) substrates in aqueous electrolytes. We compare the electrodeposition behavior of Fe on polycrystalline and…
▽ More
The development of reversible metal anodes is a key challenge for advancing aqueous battery technologies, particularly for scalable and safe stationary energy storage applications. Here we demonstrate a strategy to realize epitaxial electrodeposition of iron (Fe) on single-crystal copper (Cu) substrates in aqueous electrolytes. We compare the electrodeposition behavior of Fe on polycrystalline and single-crystalline Cu substrates, revealing that the latter enables highly uniform, dense, and crystallographically aligned Fe growth. Comprehensive electron backscatter diffraction (EBSD) and X-ray diffraction (XRD) analysis confirms the formation of Fe with specific out-of-plane and in-plane orientations, including well-defined rotational variants. Our findings highlight that epitaxial electrodeposition of Fe can suppress dendritic growth and significantly enhance Coulombic efficiency during plating/stripping cycles. This approach bridges fundamental crystallography with practical electrochemical performance, providing a pathway toward high-efficiency aqueous batteries utilizing Earth-abundant materials.
△ Less
Submitted 12 October, 2025;
originally announced October 2025.
-
dInfer: An Efficient Inference Framework for Diffusion Language Models
Authors:
Yuxin Ma,
Lun Du,
Lanning Wei,
Kun Chen,
Qian Xu,
Kangyu Wang,
Guofeng Feng,
Guoshan Lu,
Lin Liu,
Xiaojing Qi,
Xinyuan Zhang,
Zhen Tao,
Haibo Feng,
Ziyun Jiang,
Ying Xu,
Zenan Huang,
Yihong Zhuang,
Haokai Xu,
Jiaqi Hu,
Zhenzhong Lan,
Junbo Zhao,
Jianguo Li,
Da Zheng
Abstract:
Diffusion-based large language models (dLLMs) have emerged as a promising alternative to autoregressive (AR) LLMs, leveraging denoising-based generation to enable inherent parallelism. Even more and more open-sourced dLLM models emerge, yet their widespread adoption remains constrained by the lack of a standardized and efficient inference framework. We present dInfer, an efficient and extensible f…
▽ More
Diffusion-based large language models (dLLMs) have emerged as a promising alternative to autoregressive (AR) LLMs, leveraging denoising-based generation to enable inherent parallelism. Even more and more open-sourced dLLM models emerge, yet their widespread adoption remains constrained by the lack of a standardized and efficient inference framework. We present dInfer, an efficient and extensible framework for dLLM inference. dInfer decomposes the inference pipeline into four modular components--model, diffusion iteration manager, decoding strategy, and KV-cache manager--and integrates novel algorithms for each component alongside system-level optimizations. Through this combination of algorithmic innovations and system enhancements, dInfer achieves substantial efficiency gains without compromising output quality on LLaDA-MoE. At batch size 1, it surpasses 1,100 tokens per second on HumanEval and averages over 800 tokens per second across six benchmarks on $8\times$ H800 GPUs. Compared to prior systems, dInfer delivers a $10\times$ speedup over Fast-dLLM while maintaining similar model performance. Even compared to the AR model (with a comparable number of activation parameters and performance) QWen2.5-3B, which is highly optimized with the latest vLLM inference engine, dInfer still delivers a $2$-$3\times$ speedup. The implementation of dInfer is open-sourced at https://github.com/inclusionAI/dInfer.
△ Less
Submitted 22 October, 2025; v1 submitted 9 October, 2025;
originally announced October 2025.
-
Higher-Order Feature Attribution: Bridging Statistics, Explainable AI, and Topological Signal Processing
Authors:
Kurt Butler,
Guanchao Feng,
Petar Djuric
Abstract:
Feature attributions are post-training analysis methods that assess how various input features of a machine learning model contribute to an output prediction. Their interpretation is straightforward when features act independently, but it becomes less clear when the predictive model involves interactions, such as multiplicative relationships or joint feature contributions. In this work, we propose…
▽ More
Feature attributions are post-training analysis methods that assess how various input features of a machine learning model contribute to an output prediction. Their interpretation is straightforward when features act independently, but it becomes less clear when the predictive model involves interactions, such as multiplicative relationships or joint feature contributions. In this work, we propose a general theory of higher-order feature attribution, which we develop on the foundation of Integrated Gradients (IG). This work extends existing frameworks in the literature on explainable AI. When using IG as the method of feature attribution, we discover natural connections to statistics and topological signal processing. We provide several theoretical results that establish the theory, and we validate our theory on a few examples.
△ Less
Submitted 28 January, 2026; v1 submitted 7 October, 2025;
originally announced October 2025.
-
Relief of EGFR/FOS-downregulated miR-103a by loganin alleviates NF-kappaB-triggered inflammation and gut barrier disruption in colitis
Authors:
Yan Li,
Teng Hui,
Xinhui Zhang,
Zihan Cao,
Ping Wang,
Shirong Chen,
Ke Zhao,
Yiran Liu,
Yue Yuan,
Dou Niu,
Xiaobo Yu,
Gan Wang,
Changli Wang,
Yan Lin,
Fan Zhang,
Hefang Wu,
Guodong Feng,
Yan Liu,
Jiefang Kang,
Yaping Yan,
Hai Zhang,
Xiaochang Xue,
Xun Jiang
Abstract:
Due to the ever-rising global incidence rate of inflammatory bowel disease (IBD) and the lack of effective clinical treatment drugs, elucidating the detailed pathogenesis, seeking novel targets, and developing promising drugs are the top priority for IBD treatment. Here, we demonstrate that the levels of microRNA (miR)-103a were significantly downregulated in the inflamed mucosa of ulcerative coli…
▽ More
Due to the ever-rising global incidence rate of inflammatory bowel disease (IBD) and the lack of effective clinical treatment drugs, elucidating the detailed pathogenesis, seeking novel targets, and developing promising drugs are the top priority for IBD treatment. Here, we demonstrate that the levels of microRNA (miR)-103a were significantly downregulated in the inflamed mucosa of ulcerative colitis (UC) patients, along with elevated inflammatory cytokines (IL-1beta/TNF-alpha) and reduced tight junction protein (Occludin/ZO-1) levels, as compared with healthy control objects. Consistently, miR-103a deficient intestinal epithelial cells Caco-2 showed serious inflammatory responses and increased permeability, and DSS induced more severe colitis in miR-103a-/- mice than wild-type ones. Mechanistic studies unraveled that c-FOS suppressed miR-103a transcription via binding to its promoter, then miR-103a-targeted NF-kappaB activation contributes to inflammatory responses and barrier disruption by targeting TAB2 and TAK1. Notably, the traditional Chinese medicine Cornus officinalis (CO) and its core active ingredient loganin potently mitigated inflammation and barrier disruption in UC by specifically blocking the EGFR/RAS/ERK/c-FOS signaling axis, these effects mainly attributed to modulated miR-103a levels as the therapeutic activities of them were almost completely shielded in miR-103a KO mice. Taken together, this work reveals that loganin relieves EGFR/c-FOS axis-suppressed epithelial miR-103a expression, thereby inhibiting NF-kappaB pathway activation, suppressing inflammatory responses, and preserving tight junction integrity in UC. Thus, our data enrich mechanistic insights and promising targets for UC treatment.
△ Less
Submitted 5 October, 2025;
originally announced October 2025.
-
ActiveUMI: Robotic Manipulation with Active Perception from Robot-Free Human Demonstrations
Authors:
Qiyuan Zeng,
Chengmeng Li,
Jude St. John,
Zhongyi Zhou,
Junjie Wen,
Guorui Feng,
Yichen Zhu,
Yi Xu
Abstract:
We present ActiveUMI, a framework for a data collection system that transfers in-the-wild human demonstrations to robots capable of complex bimanual manipulation. ActiveUMI couples a portable VR teleoperation kit with sensorized controllers that mirror the robot's end-effectors, bridging human-robot kinematics via precise pose alignment. To ensure mobility and data quality, we introduce several ke…
▽ More
We present ActiveUMI, a framework for a data collection system that transfers in-the-wild human demonstrations to robots capable of complex bimanual manipulation. ActiveUMI couples a portable VR teleoperation kit with sensorized controllers that mirror the robot's end-effectors, bridging human-robot kinematics via precise pose alignment. To ensure mobility and data quality, we introduce several key techniques, including immersive 3D model rendering, a self-contained wearable computer, and efficient calibration methods. ActiveUMI's defining feature is its capture of active, egocentric perception. By recording an operator's deliberate head movements via a head-mounted display, our system learns the crucial link between visual attention and manipulation. We evaluate ActiveUMI on six challenging bimanual tasks. Policies trained exclusively on ActiveUMI data achieve an average success rate of 70\% on in-distribution tasks and demonstrate strong generalization, retaining a 56\% success rate when tested on novel objects and in new environments. Our results demonstrate that portable data collection systems, when coupled with learned active perception, provide an effective and scalable pathway toward creating generalizable and highly capable real-world robot policies.
△ Less
Submitted 1 October, 2025;
originally announced October 2025.
-
10-W Sub-100-fs Ultrafast Cr:ZnS/ZnSe MOPA System enabled by doping gradient engineering
Authors:
Guangzi Feng,
Xiyue Zhang,
Yuchen Wang,
Weibo Wu,
Gianluca Galzerano,
Qing Wang,
Ting Yu,
Yujie Peng,
Jintai Fan,
Benxue Jiang,
Yuxin Leng,
Long Zhang
Abstract:
We report on a high-power mid-infrared femtosecond master oscillator power amplifier (MOPA) system, employing Cr:ZnS and Cr:ZnSe polycrystals with fine-tuned doping profiles. Based on the soft-aperture Kerr-lens mode-locking in the soliton regime, the seed oscillator generates ~40-fs pulses with a repetition rate ~173 MHz with an average power close to 400 mW. The amplification process of the seed…
▽ More
We report on a high-power mid-infrared femtosecond master oscillator power amplifier (MOPA) system, employing Cr:ZnS and Cr:ZnSe polycrystals with fine-tuned doping profiles. Based on the soft-aperture Kerr-lens mode-locking in the soliton regime, the seed oscillator generates ~40-fs pulses with a repetition rate ~173 MHz with an average power close to 400 mW. The amplification process of the seed pulse train is investigated in depth in a single-pass configuration for both Cr:ZnS and Cr:ZnSe crystal rods. For further power scaling, a dual-stage MOPA system has been implemented, generating pulse trains with an average power up to 10.4 W, limited only by the pump source, with a re-compressed pulse duration of 78 fs using a dispersion compensator comprising chirped mirrors and sapphire plates. This work paves the way for further power scaling of mid-infrared Cr:ZnS/ZnSe ultrafast laser systems without moving parts for applications in material processing, remote sensing and medicine.
△ Less
Submitted 16 September, 2025;
originally announced September 2025.