-
Understanding Multilingual Medical ASR Adaptation Through Layer-Wise Analysis
Authors:
Souranil Kahali,
Rituparna Bose,
Abner Hernandez,
Tomas Arias-Vergara,
Andreas Maier,
Ning Ma,
Paula Andrea Perez-Toro
Abstract:
Medical automatic speech recognition (MedASR) requires adaptation to specialised terminology, limited annotated clinical data, and multilingual use cases. Although large-scale pretrained ASR models such as Whisper achieve strong generalisation, their behaviour after medical and multilingual adaptation remains insufficiently understood beyond word error rate (WER). This paper investigates how multi…
▽ More
Medical automatic speech recognition (MedASR) requires adaptation to specialised terminology, limited annotated clinical data, and multilingual use cases. Although large-scale pretrained ASR models such as Whisper achieve strong generalisation, their behaviour after medical and multilingual adaptation remains insufficiently understood beyond word error rate (WER). This paper investigates how multilingual medical adaptation reshapes the internal representations of Whisper models through layer-wise encoder analysis. We compare zero-shot decoding, English-only fine-tuning, German-only diagnostic fine-tuning, two-stage EN->EN+DE continuation, and direct EN+DE fine-tuning across Whisper model sizes. Fine-tuning substantially improves MedASR performance, but the best model depends on the adaptation setting: Whisper-Medium gives the lowest English WER (7.72%) and the lowest combined EN+DE WER under direct EN+DE training (26.30%); German-only Whisper-Large-v3 gives the lowest German WER (44.96%), but as a within-corpus diagnostic on 86 single-speaker training utterances rather than robust generalisation. Layer-wise analysis of the two-stage Whisper-Small trajectory shows that English medical fine-tuning produces the dominant encoder shift, whereas multilingual continuation largely preserves the adapted representation space. Domain and language information remain highly recoverable across layers, while linearly recoverable error-predictive cues weaken as WER improves.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning
Authors:
Kai Chen,
Jifeng Ding,
Ning Ding,
Jiaye Ge,
Lixin Gu,
Yicheng Gu,
Qipeng Guo,
Ermo Hua,
Haian Huang,
Haozheng Hou,
Jie Hou,
Xiangyu Hong,
Che Jiang,
Minxi Jin,
Cheng Liang,
Dahua Lin,
Dawei Liu,
Kuikun Liu,
Chengqi Lv,
Haijun Lv,
Han Lv,
Ningsheng Ma,
Biqing Qi,
Jianmin Qian,
Shiya Su
, et al. (22 additional authors not shown)
Abstract:
We introduce Mobius-v0, an architecture that comprises a globally shared Memory (FFN) that stores knowledge vectors and multiple Reasoners (Self-Attn) that iteratively achieve compositional reasoning. Using hidden states as cache and carrier, reasoners repeatedly query memory for required knowledge-vectors, while the knowledge is transmitted back to reasoning operators. Through this knowledge-reas…
▽ More
We introduce Mobius-v0, an architecture that comprises a globally shared Memory (FFN) that stores knowledge vectors and multiple Reasoners (Self-Attn) that iteratively achieve compositional reasoning. Using hidden states as cache and carrier, reasoners repeatedly query memory for required knowledge-vectors, while the knowledge is transmitted back to reasoning operators. Through this knowledge-reasoning-separation architecture, Mobius achieves better knowledge compression and reasoning efficiency. Built upon Mobius-v0 architecture: 1) Our 7B model trained-from-scratch achieves similar downstream score as a 7B Transformer baseline with 62.6% of baseline's training data. 2) Our Intern-S2-Mobius, continually-pretrained from Qwen3.5-35B, achieves similar downstream score while delivering nearly 4x end-to-end inference speedup.
△ Less
Submitted 14 August, 2026;
originally announced August 2026.
-
Intern-S2-Preview: Scientific Agentic Foundation Model
Authors:
Lei Bai,
Jiaqi Cao,
Chiyu Chen,
Guanzhou Chen,
Kai Chen,
Guangran Cheng,
Erfei Cui,
Xuanlang Dai,
Shengyuan Ding,
Shangheng Du,
Yanhui Duan,
Yue Fan,
Youqing Fang,
Quan Gan,
Yuanyuan Gao,
Jiaye Ge,
Lixin Gu,
Yuzhe Gu,
Qipeng Guo,
Junjun He,
Xin Hong,
Ming Hu,
Zhouqi Hua,
Haian Huang,
Junhao Huang
, et al. (100 additional authors not shown)
Abstract:
Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. We present Intern-S2-Preview, a series of scientific agentic foundation models designed to support multimodal scientific understanding, reasoning, generation, and long-horizon tas…
▽ More
Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. We present Intern-S2-Preview, a series of scientific agentic foundation models designed to support multimodal scientific understanding, reasoning, generation, and long-horizon tasks. The training pipeline begins with scientific multimodal pre-training over rendered scientific documents, interleaved image-text data, and diverse scientific corpora. Starting from the pretrained checkpoint, we apply a unified post-training pipeline consisting of supervised fine-tuning, scalable multi-task reinforcement learning (RL), black- and white-box agentic RL, and on-policy distillation. This pipeline is supported by practical techniques that improve rollout and training stability and efficiency, including partial rollout with off-policy correction, adaptive length regularization, online speculative decoding, robust multi-task optimization, and trace-aware experience assembly for agentic tasks. At the architecture level, Intern-S2-Preview-397B extends time series modelling from efficient long-sequence understanding to numerical forecasting, while Memory Decoder is studied as a separate memory-augmented path for rapid scientific specialization without modifying the frozen 397B backbone. Evaluations across scientific, multimodal, agentic, and general-purpose benchmarks show that Intern-S2-Preview-397B achieves competitive or leading results in multiple settings. The time series modules improve scientific signal understanding and forecasting on SciTS, while the separate Intern-MemDec-4B extension improves the Biology-Instructions average score from 56.92 to 60.32 without modifying the frozen 397B backbone.
△ Less
Submitted 13 August, 2026;
originally announced August 2026.
-
Adaptive Source-Channel Coding for Bi-static Integrated Sensing and Semantic Communications
Authors:
Haotian Wang,
Dan Wang,
Xiaodong Xu,
Chuan Huang,
Hao Chen,
Nan Ma,
Ping Zhang
Abstract:
Semantic communication (SemCom) has emerged as a new paradigm to facilitate the performance of integrated sensing and communication systems in 6G, due to its potential to enhance transmission efficiency by transmitting task-relevant semantic features rather than raw bits. However, most of the existing works mainly focus on sensing data compression to reduce the subsequent communication overheads,…
▽ More
Semantic communication (SemCom) has emerged as a new paradigm to facilitate the performance of integrated sensing and communication systems in 6G, due to its potential to enhance transmission efficiency by transmitting task-relevant semantic features rather than raw bits. However, most of the existing works mainly focus on sensing data compression to reduce the subsequent communication overheads, without considering the integrated transmission framework for both the SemCom and sensing tasks. This paper proposes a sensing-aware adaptive source-channel coding (SA-ASCC) and beamforming design framework for bi-static integrated sensing and SemCom (ISSC) systems by jointly optimizing the coding rate for SemCom task and the transmit beamforming for both the SemCom and sensing tasks. Specifically, an end-to-end semantic distortion function is approximated by deriving an upper bound composing of source and channel coding induced components, and then a hybrid Cramér-Rao bound (HCRB) is derived for target position under imperfect time synchronization due to the transceiver deployed at different places in our considered bi-static ISSC system. To characterize the achievable region between SemCom and sensing performance, a distortion minimization problem is formulated by considering the HCRB threshold, channel uses, and power budget, which is non-convex due to the coupled design variables and the mixed-integer program. Subsequently, an alternating optimization (AO) algorithm is proposed to decompose this problem into the model selection and joint rate and beamforming optimization subproblems, which are solved by the exhaustive search method and the combination of successive convex approximation and fractional programming, respectively. Finally, simulation results demonstrate that the proposed scheme outperforms the DJSCC-WF-ZF and BPG-WF-ZF benchmarks.
△ Less
Submitted 11 August, 2026;
originally announced August 2026.
-
REATS: LLM Reasoning-based Ensemble Learning for Adaptive Time Series Forecasting
Authors:
Xu Zhang,
Chang Xu,
Hui Sun,
Nan Ma,
Zijian Zhang,
Peng Wang,
Wei Wang,
Li Zhao
Abstract:
Due to the diversity of real-world time series, no single forecasting model consistently dominates across all samples. Ensemble learning addresses this by combining complementary model strengths, yet existing methods rely on fixed rules or black-box models based solely on numerical inputs, failing to leverage LLM reasoning for interpretable weighting decisions. We propose REATS, which leverages LL…
▽ More
Due to the diversity of real-world time series, no single forecasting model consistently dominates across all samples. Ensemble learning addresses this by combining complementary model strengths, yet existing methods rely on fixed rules or black-box models based solely on numerical inputs, failing to leverage LLM reasoning for interpretable weighting decisions. We propose REATS, which leverages LLM reasoning capabilities as an intelligent ensemble router that jointly processes textual temporal pattern descriptions and numerical features to produce interpretable, sample-adaptive ensemble weights through chain-of-thought reasoning. To enable effective LLM-based ensembling, we study its key design choices and propose: (i) a structured input pipeline that transforms raw time series into hybrid textual--numerical representations with fixed token cost, enabling rule-based chain-of-thought construction without API dependency, augmented with retrieved similar-sample priors; (ii) a diverse multi-row weight supervision scheme coupled with a token-efficient percentage-table format that reduces numerical complexity and mitigates LLM hallucinations; and (iii) a two-stage fine-tuning framework combining SFT with GRPO, where a reciprocal reward mapping transforms the continuous unbounded MSE gap into bounded signals with amplified near-oracle sensitivity, addressing the uniform sensitivity and outlier-dominated advantage compression inherent in naive reward designs for regression-based GRPO. Experiments on eight benchmarks demonstrate that REATS outperforms competitive ensemble baselines while providing natural language explanations and demonstrating strong transfer learning and out-of-domain generalization to unseen candidate models.
△ Less
Submitted 10 August, 2026;
originally announced August 2026.
-
Hybrid-Field Sparse Channel Representation and Recovery for XL-RIS-Assisted mmWave MIMO Systems
Authors:
Wenkai Liu,
Nan Ma,
Jianqiao Chen,
Hongtao Zhang,
Ping Zhang
Abstract:
Extremely large-scale reconfigurable intelligent surface (XL-RIS)-assisted communication is regarded as a key enabling technology for future 6G networks. However, hybrid-field channel estimation for XL-RIS-assisted systems is challenging due to the high-dimensional cascaded channel and the coexistence of far-field and near-field propagation. In this case, traditional full-dimensional sparse recove…
▽ More
Extremely large-scale reconfigurable intelligent surface (XL-RIS)-assisted communication is regarded as a key enabling technology for future 6G networks. However, hybrid-field channel estimation for XL-RIS-assisted systems is challenging due to the high-dimensional cascaded channel and the coexistence of far-field and near-field propagation. In this case, traditional full-dimensional sparse recovery methods require a large cascaded dictionary and suffer from severe computational and storage burdens. To address these challenges, we develop a double-timescale channel estimation framework that decouples sparse dictionary representation and recovery. Then, by exploiting the quasi-static property of the channel at the base station (BS) and RIS side, we propose a Dirichlet kernel-based off-grid dictionary compression (DK-ODC) scheme for sparse representation, which reduces the dimension of the corresponding dictionary as well as mitigates BS-side angular off-grid error. Furthermore, for the dynamic channel at the user equipment (UE) and RIS side, we propose a subspace-aware incremental variational Bayesian learning (SI-VBL) algorithm, which enables incremental learning of sparse channels by exploiting the identified low-dimensional subspace and pruning threshold. Analysis and simulation results confirm that the proposed framework avoids full-dimensional Bayesian recovery and achieves a favorable tradeoff among estimation accuracy, computational complexity, and storage overhead.
△ Less
Submitted 26 July, 2026;
originally announced August 2026.
-
Field-Selected Topological Buffering in a Disordered Skyrmion Crystal
Authors:
Wenyu Su,
Nvsen Ma,
Chen Cheng,
Hong-Gang Luo
Abstract:
Quenched disorder can disrupt crystalline order without immediately destroying the topology of its constituent textures, but the relation between these processes in skyrmion crystals remains unclear. Using large-scale simulations of a triangular-lattice chiral magnet with random DM interactions, we show that the magnetic field selects between two disordering routes. At high fields, global translat…
▽ More
Quenched disorder can disrupt crystalline order without immediately destroying the topology of its constituent textures, but the relation between these processes in skyrmion crystals remains unclear. Using large-scale simulations of a triangular-lattice chiral magnet with random DM interactions, we show that the magnetic field selects between two disordering routes. At high fields, global translational coherence is lost at a weak-disorder scale, while sixfold bond-orientational order survives to a larger disorder strength and the total topological charge remains nearly locked up to a substantially larger scale. The resulting interval defines a topological buffer containing a Bragg-glass- like skyrmion regime followed by a skyrmion-glass regime. Finite-size scaling, spatial correlations, defect statistics, and spin autocorrelations support their distinct structural and glassy character. At lower fields, bond-orientational disordering nearly coincides with topological reconstruction, eliminating the skyrmion-glass window and contracting the buffer. These results identify the magnetic field as a control knob for separating crystalline disordering from topological-charge loss and establish topological buffering as a mechanism by which topological textures can remain robust in structurally disordered media.
△ Less
Submitted 30 July, 2026;
originally announced July 2026.
-
Restoration Flow Matching-Based Channel Refinement and Equalization Correction for MIMO Semantic Communications
Authors:
Wenkai Liu,
Nan Ma,
Jianqiao Chen,
Xiaodong Xu,
Meixia Tao,
Ping Zhang
Abstract:
In multiple-input multiple-output (MIMO) semantic communication, imperfect channel state information (CSI) and equalization mismatch can seriously degrade semantic reconstruction quality. To address this issue, we propose a unified restoration flow matching (RFM)-based framework for channel refinement and equalization correction. Specifically, the channel RFM (CRFM) module is developed to refine t…
▽ More
In multiple-input multiple-output (MIMO) semantic communication, imperfect channel state information (CSI) and equalization mismatch can seriously degrade semantic reconstruction quality. To address this issue, we propose a unified restoration flow matching (RFM)-based framework for channel refinement and equalization correction. Specifically, the channel RFM (CRFM) module is developed to refine the coarse channel, thereby improving channel estimation accuracy. Based on the refined channel, the developed semantic RFM (SRFM) module is employed to correct the residual distortions in the post-equalization latent space. The key idea is to formulate the two cascaded inverse problems of channel estimation and equalization as the unified conditional restoration task, in which the learned conditional velocity field guides the perturbed distribution towards the target distribution. To enhance the robustness of these two modules under various distortion conditions, we develop a dual-anchor perturbation training strategy that jointly learns near-manifold refinement and large-error correction, and implement inference through a few-step deterministic ordinary differential equation (ODE) solver. Extensive experiments on MIMO channels and visual semantic transmission tasks demonstrate that the proposed scheme improves key metrics for channel estimation and semantic reconstruction quality. Moreover, compared with representative diffusion-based generative baselines, the proposed method requires fewer sampling steps.
△ Less
Submitted 26 July, 2026;
originally announced July 2026.
-
Backend-Aware Graph Learning for Denoising Outcome Distributions in Quantum Program Testing
Authors:
Ning Ma,
Jun Dai,
Heng Li
Abstract:
Testing quantum programs on NISQ (Noisy Intermediate-Scale Quantum) backends is challenging because the noise disturbs outcome distributions and can affect pass/fail decisions. We present Q-BRIDGE, a graph learning-based approach that converts noisy observations into denoised distributions suitable for oracle-based verification. Q-BRIDGE uses a graph transformer architecture to encode a transpiled…
▽ More
Testing quantum programs on NISQ (Noisy Intermediate-Scale Quantum) backends is challenging because the noise disturbs outcome distributions and can affect pass/fail decisions. We present Q-BRIDGE, a graph learning-based approach that converts noisy observations into denoised distributions suitable for oracle-based verification. Q-BRIDGE uses a graph transformer architecture to encode a transpiled quantum circuit, capturing the characteristics of its gates and their connectivity; the physical backend information is encoded together with the logical structure of the circuit. An additional conditioning layer, based on FiLM (Feature-Wise Linear Modulation), takes the encoding as input and integrates noisy observations to produce denoised outcomes. We evaluate Q-BRIDGE on 23 IBM noise backends and 6 circuit families representative of practical workloads. In the first setting, we train a separate Q-BRIDGE model for each backend; in the second setting, we train a single general model shared across all backends. Across both settings, Q-BRIDGE outperforms the state-of-the-art baseline in noise mitigation by a large margin. In testing scenarios with noisy executions, Q-BRIDGE achieves 93.97%-94.90% precision and 82.50%-83.51% recall in detecting bug-induced test failures, significantly outperforming the state-of-the-art baseline. These results indicate that considering the graph structure of the transpiled circuits and the physical characteristics of specific quantum backends is a practical route to more reliable noise-aware quantum program testing.
△ Less
Submitted 25 July, 2026;
originally announced July 2026.
-
Contrasting $Γ$- and K-Valley Moiré Physics in Twisted Monolayer/Bilayer WSe$_2$
Authors:
Jackson Kuklin,
Ning Mao,
Milan Mandigo-Stoba,
Edgar Elias,
Tianci Song,
Connor Engel,
Pola Pietrzkowski,
Kenji Watanabe,
Takashi Taniguchi,
Daniel Rhodes,
Yang Zhang,
Qianhui Shi
Abstract:
Electronic orbital character plays a central role in determining electronic correlations, spin-orbit coupling, dimensionality, and ultimately the quantum phases of condensed-matter systems. Two-dimensional moiré materials have emerged as highly tunable platforms for exploring correlated phenomena, but the role of orbital degrees of freedom remains largely unexplored. Here, we identify twisted mono…
▽ More
Electronic orbital character plays a central role in determining electronic correlations, spin-orbit coupling, dimensionality, and ultimately the quantum phases of condensed-matter systems. Two-dimensional moiré materials have emerged as highly tunable platforms for exploring correlated phenomena, but the role of orbital degrees of freedom remains largely unexplored. Here, we identify twisted monolayer/bilayer WSe$_2$ as a platform in which displacement-field tuning enables moiré physics to be realized in both the $K$ and $Γ$ valleys. The distinct orbital characters of these valleys give rise to contrasting correlated phases at moiré filling factors $ν=1$ and $ν=1/3$. At $ν=1$, the $K$-valley state is a weak insulator, consistent with an antiferromagnetic state near a van Hove singularity in the intermediate-coupling regime, similar to that observed in twisted bilayer WSe$_2$. In contrast, the $Γ$-valley state exhibits a pronounced Pomeranchuk effect, consistent with proximity to a Mott transition. At $ν=1/3$, the $K$ valley hosts a robust generalized Wigner crystal, whereas the $Γ$-valley state lies near the crystallization boundary and again exhibits a Pomeranchuk effect, with localization enhanced by increasing temperature or magnetic field. Our work highlights the importance of orbital character in defining quantum phases in moiré systems, and identify the $Γ$ valley as a promising platform for exploring correlated phenomena near quantum phase transitions, where competing phases and enhanced fluctuations may give rise to unconventional phases.
△ Less
Submitted 20 July, 2026;
originally announced July 2026.
-
Interference-Enhanced Large Electron-Phonon Coupling from Raman-active Breathing Modes in Moiré Semiconductors
Authors:
Ning Mao,
Shaozheng Wang,
Cheng Xu,
Xumin Chang,
Kenji Watanabe,
Takashi Taniguchi,
Claudia Felser,
Shengwei Jiang,
Yang Zhang
Abstract:
Superconductivity was recently observed in twisted WSe2 and MoTe2, raising a central question: is the pairing driven by electronic correlations, by phonons, or by both? Answering it requires determining the electron-phonon coupling (EPC) in these moiré semiconductors, whose calculation in realistic supercells of thousands of atoms lies beyond the reach of direct first-principles methods. Here we c…
▽ More
Superconductivity was recently observed in twisted WSe2 and MoTe2, raising a central question: is the pairing driven by electronic correlations, by phonons, or by both? Answering it requires determining the electron-phonon coupling (EPC) in these moiré semiconductors, whose calculation in realistic supercells of thousands of atoms lies beyond the reach of direct first-principles methods. Here we combine filling-dependent Raman spectroscopy with machine-learning first-principles calculations to obtain the EPC mode by mode in supercells of up to tens of thousands of atoms. Raman reveals only a few moiré phonons whose frequencies shift strongly with filling; we trace this to an interference selection rule: a phonon couples strongly only when its displacement texture matches the static lattice-reconstruction pattern, and is otherwise suppressed by destructive interference. The rule selects the low- and high-frequency breathing modes seen in Raman and makes the coupling peak at large twist angles, near those at which superconductivity appears. Lattice-reconstruction interference thus emerges as an organizing principle for moiré EPC, pointing to a substantial, potentially dominant, phonon contribution to large-angle pairing.
△ Less
Submitted 13 July, 2026;
originally announced July 2026.
-
WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos
Authors:
Jiahao Liu,
Zhongpu Xia,
Shuai Tian,
Huangrui Li,
Yuhang Zheng,
Ning Ma,
Xin Fu,
Xiaotian Liu,
Jing Li,
Yixian Li,
ShangQing Zhou,
Zebin Xing,
Linbo Wang,
Chaoyue Li,
Haoran Li,
Dongbin Zhao
Abstract:
Generalizable robot policies typically rely on action-labeled robot demonstrations, which are expensive to collect and difficult to scale. In contrast, large-scale human and robot videos contain rich physical interactions but often lack executable robot action labels. We present WALA, a framework for learning executable latent actions from both action-labeled demonstrations and action-free videos.…
▽ More
Generalizable robot policies typically rely on action-labeled robot demonstrations, which are expensive to collect and difficult to scale. In contrast, large-scale human and robot videos contain rich physical interactions but often lack executable robot action labels. We present WALA, a framework for learning executable latent actions from both action-labeled demonstrations and action-free videos. WALA first pretrains a semantic-geometric latent action model from videos by modeling the evolution between current observations and sparsely sampled future observations. Instead of reconstructing raw pixels, WALA predicts future deltas in the DINOv3 feature space and dense depth space, preserving task-relevant semantic and geometric structure while reducing sensitivity to appearance details. During policy training, the pretrained encoder provides stable latent action targets, and the decoder serves as a trainable latent world model. The latent actions generated by the vision-language backbone are jointly supervised by robot action prediction, latent action target matching, and future dynamics prediction. This enables action-labeled demonstrations to provide executable control supervision, while action-free videos contribute dynamics supervision without requiring robot action annotations. Experiments show that WALA achieves strong performance on RoboTwin, sets a new state-of-the-art result on RoboCasa with 75.2% average success, and improves both policy performance and generalization in real-world manipulation tasks.
△ Less
Submitted 13 July, 2026;
originally announced July 2026.
-
Half state at $ν_{tot}$ = -1/2 and its transition in Decoupled Twisted Double Bilayer Graphene
Authors:
Ning Ma,
Kenji Watanabe,
Takashi Taniguchi,
Mitali Banerjee
Abstract:
The origin of the fractional state at $ν$ = 1/2 observed in double-layer quantum Hall systems has been under debate for decades. Because of the variation of bilayer charge distribution and interlayer tunneling strength, the half-filling state can be attributed to a two-component(2C) or a one-component(1C) origin, which corresponds to Halperin state and Pffafian state, respectively. Here we report…
▽ More
The origin of the fractional state at $ν$ = 1/2 observed in double-layer quantum Hall systems has been under debate for decades. Because of the variation of bilayer charge distribution and interlayer tunneling strength, the half-filling state can be attributed to a two-component(2C) or a one-component(1C) origin, which corresponds to Halperin state and Pffafian state, respectively. Here we report the magnetotransport measurement in decoupled twisted double bilayer graphene(TDBG), which has been proved to be a promising platform for double quantum Hall system. Fractional quantum hall states in both odd and even denominator fillings are observed. We also found that the half-filling state occurs at zero displacement field at $ν_{tot}$ = -1/2, which is theoretically consistent with two-component Halperin-Laughlin (Ψ331) state. Moreover, we report the transition from two-component state at zero D field to one-component non-Abelian state by tunning displacement field. Our observation of the half filling state and its transition from 2C to 1C state provides the tunability of decoupled twisted double bilayer graphene and shed light on the understanding of the ground states at half-filling factor in the double quantum Hall system.
△ Less
Submitted 7 July, 2026;
originally announced July 2026.
-
Correlated Insulating States in Twisted Double Bilayer Graphene Enhanced by Interfacial Effect on CrOCl
Authors:
Ning Ma,
Zekang Zhou,
Chiara Cocchi,
Maurice Bal,
Maarten van Delft,
Kenji Watanabe,
Takashi Taniguchi,
Steffen Wiedmann,
Jian-Hao Chen,
Mitali Banerjee
Abstract:
Interaction between different two dimensional materials can give rise to many exotic physical phenomena which are rarely observed in intrinsic materials. Recently, several theoretical and experimental works have revealed that magnetic proximity effect between pristine graphene and magnetic substrates can lead to the emergence of quantum anomalous Hall states and quantum spin Hall states. However,…
▽ More
Interaction between different two dimensional materials can give rise to many exotic physical phenomena which are rarely observed in intrinsic materials. Recently, several theoretical and experimental works have revealed that magnetic proximity effect between pristine graphene and magnetic substrates can lead to the emergence of quantum anomalous Hall states and quantum spin Hall states. However, interplay between correlated states in graphene-based systems and magnetic materials has seldom been studied. Here we perform the transport measurement at ultrahigh magnetic field of twisted double bilayer graphene (TDBG) on CrOCl (COC) substrate, which is an antiferromagnetic material. Instead of a magnetic-exchange effect on graphene, we observe an enhanced correlated insulating state at half-filling factor of TDBG as a result of the charge-transfer process between TDBG and COC. The temperature and magnetic field dependence of this enhanced state are further studied. Our results demonstrate the influence of charge-related effect at the interface, and shed a light on a new route for manipulating the correlated states in graphene-based moiré systems using interfacial engineering.
△ Less
Submitted 7 July, 2026;
originally announced July 2026.
-
LLM-Empowered Multimodal Fusion Framework for Autonomous Driving: Semantic Enhancement and Channel-Adaptive Design
Authors:
Wen Wang,
Yaping Sun,
Yejun He,
Hao Chen,
Zhiyong Chen,
Xiaodong Xu,
Nan Ma,
Shuguang Cui
Abstract:
Vision-radar fusion is central to robust autonomous driving, combining dense visual semantics with precise range and velocity measurements from radar. However, real-world fusion quality is fundamentally challenged by dynamically varying input quality, stemming from occlusion, adverse weather, and channel noise. To address this, we re-frame the problem from static data fusion to channel-aware seman…
▽ More
Vision-radar fusion is central to robust autonomous driving, combining dense visual semantics with precise range and velocity measurements from radar. However, real-world fusion quality is fundamentally challenged by dynamically varying input quality, stemming from occlusion, adverse weather, and channel noise. To address this, we re-frame the problem from static data fusion to channel-aware semantic reasoning and propose a Large Language Model-centric Semantic-layer Channel-aware Integrated Perception (LM-SCIP) framework. It places a Large Language Model (LLM) as a central reasoning core to fuse a local visual stream with a quality-varying external radar stream used to cover perception-blind spots. Concretely, LM-SCIP couples a hierarchical radar-vision encoder with a Channel-Adaptive Semantic Module (CASM) that maps link indicators into a "Channel Prompt" to dynamically gate external radar features. A parameter-efficient, LoRA-tuned LLM, in conjunction with a heterogeneous Mixture-of-Experts (H-MoE), then arbitrates between local visual cues and the channel-conditioned radar context. Finally, a decoupled multi-task decoder outputs localization, trajectory forecasting, and image reconstruction. Experiments on nuScenes and VIRAT validate our approach. On nuScenes, under a controlled toggle of radar input, LM-SCIP reduces localization RMSE by 40.0% versus a vision-only baseline. On VIRAT, the model attains a 0.214m localization RMSE and 0.179m minFDE (k=1). These results reveal that the proposed LM-SCIP enables a robust vision-dominant fallback at low SNR and synergistic fusion at high SNR.
△ Less
Submitted 2 July, 2026;
originally announced July 2026.
-
MxGLUT: A Reconfigurable LUT-Centric Broadcast Dataflow Accelerator for Mixed-Precision GEMM
Authors:
Weiyu Zhou,
Chen Ding,
Mingyuan Liu,
Liangyu Gan,
Yukun Feng,
Hao Jia,
Haoming Chu,
Lirong Zheng,
Ning Ma,
Yuxiang Huan
Abstract:
Large language model (LLM) inference suffers from growing inefficiency across the prefill and decode phases, especially under weight-only quantization, where activations remain in FP8 while weights are compressed to low-bit integers. Existing LUT-based accelerators mainly target FP8-INT4 computation and still rely on separate floating-point (FP) datapaths for attention GEMM operations, leading to…
▽ More
Large language model (LLM) inference suffers from growing inefficiency across the prefill and decode phases, especially under weight-only quantization, where activations remain in FP8 while weights are compressed to low-bit integers. Existing LUT-based accelerators mainly target FP8-INT4 computation and still rely on separate floating-point (FP) datapaths for attention GEMM operations, leading to redundant hardware and non-unified mixed-precision execution. Moreover, their static dataflows are poorly matched to the distinct prefill and decode phases. To address these challenges, we propose MxGLUT, a reconfigurable LUT-centric broadcast (RLB) dataflow accelerator built on mixed-precision LUT-based processing elements (MxLPEs). Guided by a unified LUT-based execution framework, MxGLUT organizes both FP8-INT4 and FP8-FP8 GEMMs under a single LUT-based compute mechanism without dedicated FP multipliers or additional FP datapaths, and further adopts the RLB dataflow that localizes heavy partial-sum accumulation during the prefill phase and exploits weight reuse in the decode phase. Synthesized in UMC $28\,\mathrm{nm}$ CMOS at $200~\mathrm{MHz}$, MxGLUT reduces multiplier area by up to $56.92\%$ and power by up to $77.07\%$ and $78.35\%$ in FP8-INT4 and FP8-FP8 modes, respectively. At the accelerator level, MxGLUT achieves an area efficiency of $0.492~\mathrm{TFLOPS/mm^2}$ and an energy efficiency of $11.58~\mathrm{TFLOPS/W}$, while adding native FP8-FP8 support incurs only $2.57\%$ and $3.34\%$ reductions in area and energy efficiency, respectively, relative to the FP8-INT4-only FIGLUT baseline. Across the Llama family, MxGLUT achieves up to $2.16\times$ and $1.49\times$ latency speedup, and reduces normalized energy to $0.44\times$ and $0.71\times$ in prefill and decode, respectively, with at most $1.70\%$ perplexity increase.
△ Less
Submitted 1 July, 2026;
originally announced July 2026.
-
ClinRAG-GRAPH: Clinical-prior Retrieval-Augmented Graph Model with Domain Adversarial Learning for Breast pCR Prediction
Authors:
Yaofei Duan,
Yuhao Huang,
Tianyu Zhang,
Yuan Gao,
Luyi Han,
Xin Wang,
Xinyu Xie,
Xinglong Liang,
Chunyao Lu,
Muzhen He,
Patrick Pang,
Yue Sun,
Ning Mao,
Tao Tan,
Ritse Mann
Abstract:
Neoadjuvant chemotherapy (NAC) response prediction is clinically important for treatment stratification in breast cancer. However, robust pre-treatment pathological complete response (pCR) prediction remains challenging due to insufficient cross-modal modeling, multicenter imaging heterogeneity, and weak evidence-grounded interpretability. We propose ClinRAG-GRAPH, a Clinically informed Retrieval-…
▽ More
Neoadjuvant chemotherapy (NAC) response prediction is clinically important for treatment stratification in breast cancer. However, robust pre-treatment pathological complete response (pCR) prediction remains challenging due to insufficient cross-modal modeling, multicenter imaging heterogeneity, and weak evidence-grounded interpretability. We propose ClinRAG-GRAPH, a Clinically informed Retrieval-Augmented Generation Graph framework, for pre-treatment pCR prediction from DCE-MRI, structured clinical variables, and biopsy-derived pathological biomarkers. ClinRAG-GRAPH constructs an intra-patient clinical-prior graph and applies a prior-guided relation-aware graph convolutional network for structured multimodal representation learning. To improve cross-center robustness, we introduce a dual-branch domain-adversarial learning strategy to suppress protocol-related MRI bias while preserving pCR-relevant features. To enhance interpretability, we further incorporate large language model (LLM)-driven subgraph RAG module that retrieves clinically analogous historical cases and integrates retrieved evidence for pCR inference. We assemble a large-scale multicenter NAC breast cancer cohort for extensive validation, drawing from two public sources and three in-house centers.Results show that ClinRAG-GRAPH achieves AUCs of 0.815 on the internal test set and 0.774/0.712 on two external test sets, demonstrating robust pre-treatment pCR prediction across centers. The code is available at the anonymized https://github.com/miccai26-1181/ClinRAG-GRAPH.
△ Less
Submitted 1 July, 2026;
originally announced July 2026.
-
Selenium direct doping obtained high-performance-n-type Bi2Te3-based thermoelectric materials with a wide temperature range
Authors:
Zhiyuan Liu,
Junjie Ma,
Zhaopeng Zeng,
Ni Ma,
Qian Ba,
Di Zhang,
Zhe Tao,
Ailin Xia
Abstract:
The article reports on a series of n-type Bi2Te3-based thermoelectric materials prepared via a high-temperature melting combined with annealing process. The effects of Se doping content and annealing process on the carrier concentration, suppression of the bipolar effect, and thermoelectric performance of the materials were systematically investigated. The experimental results provide valuable ref…
▽ More
The article reports on a series of n-type Bi2Te3-based thermoelectric materials prepared via a high-temperature melting combined with annealing process. The effects of Se doping content and annealing process on the carrier concentration, suppression of the bipolar effect, and thermoelectric performance of the materials were systematically investigated. The experimental results provide valuable reference for researchers in this field.
△ Less
Submitted 27 June, 2026;
originally announced June 2026.
-
SA-RA-JSCC: SNR-Adaptive and Semantic-Rate-Aware Joint Source-Channel Coding
Authors:
Shitong Zhang,
Yaping Sun,
Hao Chen,
Xiaoyi Li,
Bo Gu,
Xiaodong Xu,
Nan Ma
Abstract:
In joint source-channel coding (JSCC)-based semantic communication systems, achieving stable and reliable image semantic transmission under channel constraints remains a key challenge. In most channel adaptation modules, the signal-to-noise ratio (SNR) is often injected into each layer of a channel-adaptation model in an independent and layer-wise manner, which undermines global coordination acros…
▽ More
In joint source-channel coding (JSCC)-based semantic communication systems, achieving stable and reliable image semantic transmission under channel constraints remains a key challenge. In most channel adaptation modules, the signal-to-noise ratio (SNR) is often injected into each layer of a channel-adaptation model in an independent and layer-wise manner, which undermines global coordination across layers. Therefore, consistent noise-robust representations may fail to be learned throughout the model. To address this problem, we propose SA-RA-JSCC, a novel channel-adaptive JSCC model. SA-RA-JSCC maps SNR into a unified semantic vector in the feature space and then applies a one-shot global reweighting to the encoded features, thereby enabling globally consistent and learnable channel adaptation. Moreover, in order to further enhance the anti-channel capability of semantic information, a semantic-rate-aware module is introduced, enabling the adaptive policy to respond simultaneously to fluctuations in channel quality and changes in semantic-rate constraints, thereby enhancing global network coordination and channel adaptivity. Extensive experiment results across multiple channels and datasets demonstrate that SA-RA-JSCC significantly outperforms existing semantic communication models in terms of reconstruction metrics such as PSNR and MS-SSIM, exhibiting stronger robustness across a broad range of SNR regimes.
△ Less
Submitted 16 June, 2026;
originally announced June 2026.
-
From Signals to Patterns: Non-Invasive Tuberculosis Detection from Cough Audio using Bandit Weighted Hyperbolic Prototypes
Authors:
Mohd Mujtaba Akhtar,
Girish,
Sanjam Wadhwa,
Muskaan Singh,
Ning Ma
Abstract:
In this study, we focus on cough-based tuberculosis screening (CBTS) and hypothesize that fusing speech/audio foundation representations with spectral descriptors will yield stronger screening performance. We expect this fusion to reveal complementary strengths: spectral features preserve fine-grained short-time acoustic detail in cough signals, while foundation embeddings capture higher-level tem…
▽ More
In this study, we focus on cough-based tuberculosis screening (CBTS) and hypothesize that fusing speech/audio foundation representations with spectral descriptors will yield stronger screening performance. We expect this fusion to reveal complementary strengths: spectral features preserve fine-grained short-time acoustic detail in cough signals, while foundation embeddings capture higher-level temporal and event-level patterns learned from large-scale pretraining. To this end, we propose COBALT, a novel fusion framework based on codebook-aligned hyperbolic prototypes and bandit-style reliability weighting to integrate heterogeneous representations effectively. Using the CODA TB DREAM Challenge benchmark, COBALT consistently outperforms individual representations and a concatenation baseline, achieving the best overall performance when fusing MFCC with PaSST thereby establishing a new state-of-the-art on the benchmark.
△ Less
Submitted 20 June, 2026; v1 submitted 15 June, 2026;
originally announced June 2026.
-
STCC: A Unified Source-Channel Semantic Token Coding Framework for Semantic Communications
Authors:
Zhicheng Bao,
Chen Dong,
Sen Wang,
Long Liu,
Nan Ma,
Hao Chen,
Xiaodong Xu,
Yinqiu Liu,
Ping Zhang
Abstract:
Deep Joint Source-Channel Coding (JSCC) has emerged as a promising paradigm for overcoming the ``cliff effect" in wireless communications. However, existing Deep JSCC frameworks operate directly on raw analog data such as image pixels rather than the discrete semantic tokens that foundation models require. Moreover, traditional systems employ fixed, hand-designed constellations that treat all toke…
▽ More
Deep Joint Source-Channel Coding (JSCC) has emerged as a promising paradigm for overcoming the ``cliff effect" in wireless communications. However, existing Deep JSCC frameworks operate directly on raw analog data such as image pixels rather than the discrete semantic tokens that foundation models require. Moreover, traditional systems employ fixed, hand-designed constellations that treat all tokens equally, leading to catastrophic random errors under channel noise. In this paper, the Semantic Token Codebook Communication (STCC) is proposed as a unified source-channel semantic token coding framework designed to transmit the discrete semantic tokens of foundation models over noisy channels. The core of STCC is the Semantic Token Codec (STC). It accepts discrete tokens as input, which maintains compatibility with foundation models while employing a residual multiple layer perceptron, i.e., MLP-based encoder that learns geometrically structured constellations optimized with a triple-loss objective. This learned mapping forces the channel topology to align with the semantic embedding space, ensuring that channel noise results in topological errors rather than random corruption. This phenomenon is theoretically and empirically characterized, identifying ``Semantic Drift" in symbolic modalities and ``Structural Distortion" in perceptual modalities, where errors shift predictions to semantically or structurally similar tokens. Extensive experiments demonstrate that STCC significantly outperforms traditional systems in low-SNR regimes, effectively converting channel noise into semantic variations without requiring receiver-side modification.
△ Less
Submitted 10 June, 2026;
originally announced June 2026.
-
A nuclear clock based on $^{229}$Th
Authors:
Beichen Huang,
Gaowei Yan,
Qi Xiao,
Wenhao Bu,
Zhen Zhang,
Chengchun Zhao,
Chao Yan,
Zhi-Ang Chen,
Peixiong Zhang,
Gleb Penyazkov,
Zhenhai Zhan,
Lingfeng Yan,
Yuefei Wang,
Lin Li,
Shanming Li,
Xiaobo Qian,
Xuegang Liu,
Qiange He,
Taoxiang Sun,
Haochen Tian,
Binkun Lu,
Ningyuan Ma,
Juxian Li,
Yanzhang Wu,
Qiaorui Gong
, et al. (13 additional authors not shown)
Abstract:
Atomic clocks have made time and frequency the most precisely measured quantities in physics, progressing from microwave standards that realize the SI second to optical clocks that now reach unprecedented levels of precision. A nuclear clock would shift the frequency reference from an electronic transition to the uniquely low-lying, laser-accessible isomeric transition in the $^{229}$Th nucleus, o…
▽ More
Atomic clocks have made time and frequency the most precisely measured quantities in physics, progressing from microwave standards that realize the SI second to optical clocks that now reach unprecedented levels of precision. A nuclear clock would shift the frequency reference from an electronic transition to the uniquely low-lying, laser-accessible isomeric transition in the $^{229}$Th nucleus, offering a route to compact, robust timekeeping and sensitive tests of fundamental physics. However, turning recent advances in spectroscopy of the $^{229}$Th nuclear resonance into clock operation requires the nuclear transition to serve as a stable discriminator for steering a traceable oscillator. Here we demonstrate the operation of a $^{229}$Th nuclear clock by stabilizing a continuous-wave narrow-linewidth 148.4 nm vacuum-ultraviolet (VUV) laser to a resolved nuclear transition in a solid-state host. This clock operation is enabled by fast frequency discrimination based on phototube photocurrent readout of the transmitted VUV power. The 10 $μ$W VUV laser, generated by four-wave mixing in cadmium vapour, provides a high-signal-to-noise absorption signal from a home-grown $^{229}$Th:CaF$_2$ crystal, allowing the laser to be locked to a weakly temperature-sensitive nuclear transition. The clock reaches a fractional frequency instability of $2\times10^{-12}/\sqrt{τ/s} $, where $τ$ is the averaging time. Remarkably, nuclear-clock frequencies measured with two distinct crystals agree at the $10^{-13}$ level, demonstrating the reproducibility of solid-state nuclear frequency references. By making a laser-addressed atomic nucleus an operational clock reference, this work extends quantum metrology from electronic to nuclear transitions, and opens a new platform for compact clocks, solid-state nuclear quantum sensors and precision tests of fundamental physics.
△ Less
Submitted 7 June, 2026;
originally announced June 2026.
-
LiAuto-GeoX: Efficient Grounded Driving Transformer
Authors:
Jiawei Lian,
Haoyi Sun,
Yang Wu,
Lifu Mu,
Siyuan Wang,
Le Hui,
Ning Mao,
Tao Wei,
Pan Zhou,
Kun Zhan,
Jian Yang
Abstract:
Dense 3D reconstruction has demonstrated immense potential for spatial understanding, yet its viability as a real-time, onboard representation for autonomous driving remains an open challenge. Existing large-scale visual geometry models typically require substantial computational resources and lack the long-range geometric fidelity, surround-view consistency, and real-time efficiency demanded by d…
▽ More
Dense 3D reconstruction has demonstrated immense potential for spatial understanding, yet its viability as a real-time, onboard representation for autonomous driving remains an open challenge. Existing large-scale visual geometry models typically require substantial computational resources and lack the long-range geometric fidelity, surround-view consistency, and real-time efficiency demanded by dynamic driving environments. To bridge this gap, we present \textbf{LiAuto-GeoX}, an efficient grounded driving transformer designed for deployable, ego-centric 3D scene understanding. Our approach begins by learning a high-capacity driving geometry model from large-scale surround-view data, utilizing sparse LiDAR priors to provide robust geometric grounding in distant, ambiguous, or structure-sparse regions. We then instantiate this capability into a highly compact 155M-parameter onboard model through a novel geometry-preserving distillation framework. This framework employs mask-guided depth-aware distillation to retain fine-grained metric structures by emphasizing geometrically informative regions, and relative-pose relational distillation to enforce cross-view spatial consistency through pose-induced geometric relations. Extensive evaluations reveal that \textbf{LiAuto-GeoX} runs at 220 FPS on KITTI while maintaining high-fidelity dense reconstruction, enabling real-time deployment. The learned geometry transfers seamlessly to downstream autonomy tasks, achieving 90.6 PDMS in trajectory prediction, 24.63 mIoU in occupancy prediction, and 47.67 IoU in future-frame prediction. These all demonstrate that efficient dense 3D reconstruction can transcend its traditional role as a perception target to serve as a scalable, foundational geometric representation for next-generation autonomous driving.
△ Less
Submitted 12 June, 2026; v1 submitted 4 June, 2026;
originally announced June 2026.
-
Knowledge Index of Noah's Ark
Authors:
Sheng Jin,
Minghao Liu,
Yunze Xiao,
Zeqi Zhou,
Heli Qi,
Yifan Yao,
Meishu Song,
Kaijing Ma,
Xuan Zhang,
Sicong Jiang,
Yizhe Li,
Ningshan Ma,
Jie Wei,
Ziniu Li,
Minglai Yang,
Bangya Liu,
Yiming Liang,
Xiao Fang,
Qingcheng Zeng,
Jiarui Liu,
Rui Yang,
Shen Yan,
Wenhao Huang,
Jiaheng Liu,
Zihan Wang
, et al. (2 additional authors not shown)
Abstract:
Knowledge benchmarks for LLMs face three issues: scaling-driven designs that do not operationalize disciplinary representativeness; flat-payment annotation that permits lazy consensus; and unaudited ranking instability under bounded test budgets. We introduce KINA, an 899-item benchmark across 261 fine-grained disciplines, with two formal results. First, we cast representativeness as a coverage-st…
▽ More
Knowledge benchmarks for LLMs face three issues: scaling-driven designs that do not operationalize disciplinary representativeness; flat-payment annotation that permits lazy consensus; and unaudited ranking instability under bounded test budgets. We introduce KINA, an 899-item benchmark across 261 fine-grained disciplines, with two formal results. First, we cast representativeness as a coverage-style objective over expert-elicited anchors and operationalize disciplinary representativeness through a proxy, yielding a (1-1/e) greedy approximation (Proposition 1); the guarantee applies to the proxy, not to population representativeness. Second, we prove a bonus-on-bar tournament weakly FOSD-dominates flat payment in released-review quality, with incentive-compatibility threshold B > Delta C / Delta p_min (Theorem 1). Evaluating 42 models from 13 labs, the top model, Gemini-3.1-Pro-Preview, reaches 53.17%, followed by Claude-Opus-4.6 at 49.92% and GPT-5.4 at 48.55%, leaving substantial headroom below saturation. The full leaderboard shows a tiered structure rather than a smooth total order: a small frontier tier lies above 48%, a dense strong-model tier spans roughly 38-45%, and low-performing models remain only modestly above the 10% chance baseline. Tool augmentation adds up to 5.17 points across the five tool-use evaluations, with gains varying substantially across models. We report bootstrap ranking-stability statistics to make bounded-budget variance explicit and to discourage over-interpretation of adjacent ranks.
△ Less
Submitted 4 June, 2026; v1 submitted 3 June, 2026;
originally announced June 2026.
-
Benchmarking Visual State Tracking in Multimodal Video Understanding
Authors:
Sihyun Yu,
Nanye Ma,
Pinzhi Huang,
Hyunseok Lee,
Shusheng Yang,
June Suk Choi,
Ellis Brown,
Oscar Michel,
Boyang Zheng,
Jinwoo Shin,
Saining Xie
Abstract:
Understanding a video requires more than recognizing isolated moments, as humans continuously track entities, states, and events over time. This capacity for visual state tracking is fundamental to video understanding, yet remains underexplored in current evaluations of Multimodal Large Language Models (MLLMs). We introduce Visual STAte Tracking benchmark (VSTAT), a video-based benchmark designed…
▽ More
Understanding a video requires more than recognizing isolated moments, as humans continuously track entities, states, and events over time. This capacity for visual state tracking is fundamental to video understanding, yet remains underexplored in current evaluations of Multimodal Large Language Models (MLLMs). We introduce Visual STAte Tracking benchmark (VSTAT), a video-based benchmark designed to diagnose visual state tracking in MLLMs. VSTAT consists of 834 clips drawn from both synthetic and real-world videos, paired with 1,500 questions that cannot be answered from any single frame or short segment, requiring continuous perception and integration of events across the entire video stream. Despite their strong performance on existing video benchmarks, we find that state-of-the-art MLLMs perform far below humans and only modestly above answer-prior baselines. To analyze this gap, we compare MLLMs' thinking traces with the underlying video stream to understand why and when MLLMs fail on VSTAT. We find that MLLMs reason and track correctly in text, but fail at visually perceiving the events they need to track. Finally, our preliminary evaluation suggests that recent agentic approaches, including MLLM-based video agents and coding agents, do not readily resolve these failures, still falling short on VSTAT.
△ Less
Submitted 2 June, 2026;
originally announced June 2026.
-
SparseStreet: Sparse Gaussian Splatting for Real-Time Street Scene Simulation
Authors:
Qingpo Wuwu,
Xiaobao Wei,
Peng Chen,
Nan Huang,
Zhongyu Zhao,
Hao Wang,
Ming Lu,
Ningning Ma,
Shanghang Zhang
Abstract:
While 3D Gaussian Splatting has shown promising results in street scene reconstruction, existing methods require massive numbers of Gaussian primitives to capture fine details, leading to prohibitive storage costs and slow rendering speeds. We observe that dynamic objects (e.g., vehicles and pedestrians) demand high-fidelity representations to maintain temporal consistency, while static background…
▽ More
While 3D Gaussian Splatting has shown promising results in street scene reconstruction, existing methods require massive numbers of Gaussian primitives to capture fine details, leading to prohibitive storage costs and slow rendering speeds. We observe that dynamic objects (e.g., vehicles and pedestrians) demand high-fidelity representations to maintain temporal consistency, while static background regions often contain substantial redundancy. Motivated by this, we propose SparseStreet, a general compression framework specifically designed for street scenes. First, we introduce a node-based learnable pruning strategy that systematically removes low-contributing Gaussian primitives while preserving visually critical regions. Second, after the scene representation stabilizes, we apply background compression, further reducing redundancy in static regions. Our method effectively preserves the geometry and appearance of dynamic objects while significantly reducing the total number of Gaussian primitives. Extensive experiments on the Waymo and nuScenes demonstrate that SparseStreet achieves up to 80% compression ratio with minimal quality degradation, enabling resource-efficient, high-fidelity dynamic scene reconstruction. Project website: https://sparsestreet.github.io/.
△ Less
Submitted 2 June, 2026;
originally announced June 2026.
-
Cosmos 3: Omnimodal World Models for Physical AI
Authors:
NVIDIA,
:,
Aditi,
Niket Agarwal,
Arslan Ali,
Jon Allen,
Martin Antolini,
Adeline Aubame,
Alisson Azzolini,
Junjie Bai,
Maciej Bala,
Yogesh Balaji,
Josh Bapst,
Aarti Basant,
Mukesh Beladiya,
Mohammad Qazim Bhat,
Zaid Pervaiz Bhat,
Dan Blick,
Vanni Brighella,
Han Cai,
Tiffany Cai,
Eric Cameracci,
Jiaxin Cao,
Yulong Cao,
Mark Carlson
, et al. (271 additional authors not shown)
Abstract:
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-transformers architecture. By supporting highly flexible input-output configurations, Cosmos 3 seamlessly unifies critical modalities for Physical AI -- effectively subsuming vision-language models, video generators, worl…
▽ More
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-transformers architecture. By supporting highly flexible input-output configurations, Cosmos 3 seamlessly unifies critical modalities for Physical AI -- effectively subsuming vision-language models, video generators, world simulators, and world-action models into a single framework. Our evaluation demonstrates that Cosmos 3 establishes a new state-of-the-art across a diverse suite of understanding and generation tasks, demonstrating omnimodal world models as scalable, general-purpose backbones for embodied agents. Our post-trained Cosmos 3 models were ranked as the best open-source Text-to-Image and Image-to-Video models by Artificial Analysis, and the best policy model by RoboArena at the time the technical report was written. To accelerate open research and deployment in Physical AI, we make our code, model checkpoints, curated synthetic datasets, and evaluation benchmark available under the Linux Foundation's OpenMDW-1.1 License at https://github.com/nvidia/cosmos and https://huggingface.co/collections/nvidia/cosmos3. The project website is available at https://research.nvidia.com/labs/cosmos-lab/cosmos3.
△ Less
Submitted 23 June, 2026; v1 submitted 1 June, 2026;
originally announced June 2026.
-
1/9 Magnetization Plateau in a Classical Kagome Ising Ferromagnet with Competing Further-Neighbor Interactions
Authors:
Yixin Guan,
Kan Zhao,
Nvsen Ma
Abstract:
The two-dimensional kagome lattice is a paradigmatic platform for exploring geometrically frustrated magnetism. While the nearest-neighbor ferromagnetic Ising model on this lattice is theoretically trivial, competing further-neighbor interactions can reintroduce severe frustration. In this work, we systematically investigate a classical kagome Ising model with ferromagnetic nearest-neighbor (J1) a…
▽ More
The two-dimensional kagome lattice is a paradigmatic platform for exploring geometrically frustrated magnetism. While the nearest-neighbor ferromagnetic Ising model on this lattice is theoretically trivial, competing further-neighbor interactions can reintroduce severe frustration. In this work, we systematically investigate a classical kagome Ising model with ferromagnetic nearest-neighbor (J1) and antiferromagnetic second- (J2) and third-neighbor (J3) couplings using simulated annealing Monte Carlo methods. We demonstrate that while J2 couplings merely suppress the conventional ferromagnetic order, the inclusion of J3 fundamentally reconstructs the low-temperature phase diagram. This extended geometric frustration stabilizes a novel ordered phase characterized by a robust 1/9 magnetization plateau and a massively enlarged 3 by 3 magnetic supercell. Crucially, this fractional ordered phase manifests as a stability plateau in the phase diagram, where its critical temperature becomes nearly independent of the coupling strength J3. We also calculate the corresponding static spin structure factor, revealing a distinct Z6-symmetric reciprocal-space signature for experimental identification. Our findings reveal that complex fractional magnetic orders can emerge purely from classical geometric frustration induced by competing extended interactions, providing a distinct mechanism for understanding fractionally ordered states in real frustrated magnets.
△ Less
Submitted 22 May, 2026;
originally announced May 2026.
-
OpenCompass: A Universal Evaluation Platform for Large Language Models
Authors:
Maosong Cao,
Kai Chen,
Haodong Duan,
Yixiao Fang,
Zhiwei Fei,
Tong Gao,
Ge Jiaye,
Mo Li,
Hongwei Liu,
Junnan Liu,
Yuan Liu,
Chengqi Lyu,
Han Lyu,
Ningsheng Ma,
Zerun Ma,
Yu Sun,
Zhiyong Wu,
Linchen Xiao,
Zhuozhi Xiong,
Jun Xu,
Haochen Ye,
Zhaohui Yu,
Yike Yuan,
Songyang Zhang,
Yufeng Zhao
, et al. (5 additional authors not shown)
Abstract:
In recent years, the field of artificial intelligence has undergone a paradigm shift from task-specific small-scale models to general-purpose large language models (LLMs). With the rapid iteration of LLMs, objective, quantitative, and comprehensive evaluation of their capabilities has become a critical link in advancing technological development. Currently, the mainstream static benchmark dataset-…
▽ More
In recent years, the field of artificial intelligence has undergone a paradigm shift from task-specific small-scale models to general-purpose large language models (LLMs). With the rapid iteration of LLMs, objective, quantitative, and comprehensive evaluation of their capabilities has become a critical link in advancing technological development. Currently, the mainstream static benchmark dataset-based evaluation methods face challenges such as the diversity of task types, inconsistent evaluation criteria, and fragmentation of data and processing workflows, making it difficult to efficiently conduct cross-domain and large-scale model evaluation. To address the aforementioned issues, this paper proposes and open-sources OpenCompass, a one-stop, scalable, and high-concurrency-supported general-purpose LLM evaluation platform. Adhering to the design philosophy of modularization and component decoupling, the platform boasts three core advantages: high compatibility, flexibility, and high concurrency. The core architecture of OpenCompass comprises five key components: the Configuration System, Task Partitioning Module, Execution and Scheduling Module, Task Execution Unit, and Result Visualization Module. Its workflow provides rule-based, LLM-as-a-Judge, and cascaded evaluators to adapt to the requirements of different task scenarios. Supporting mainstream benchmark datasets across multiple domains, including knowledge, reasoning, computation, science, language, code, etc., the platform offers a unified and efficient LLM evaluation tool for both academia and industry, facilitating the accurate identification of strengths and weaknesses of LLMs as well as their subsequent optimization.
△ Less
Submitted 7 June, 2026; v1 submitted 18 May, 2026;
originally announced May 2026.
-
Quantum geometry induced anomalous chiral transport and hidden symmetry breaking in centrosymmetric 2M-WS2
Authors:
Hang Cui,
Shao-Bo Liu,
Erqing Wang,
Mingxiang Pan,
Yuqiang Fang,
Ning Ma,
Wenlong Liu,
Di Chen,
Yu Zhang,
Yuanjun Song,
Tingting Hao,
Jiankun Li,
Jian Cui,
Ya Feng,
Haiwen Liu,
Fuqiang Huang,
Huaqing Huang,
X. -C. Xie,
Jian-Hao Chen
Abstract:
Chirality, a widely existing material property in nature involving the breaking of the left-right symmetry, has profound influences in various fields of natural sciences. Nonlinear response, such as electronic magnetochiral anisotropy (eMChA), has been recognized as a sensitive probe for the effects of symmetry breaking and nontrivial quantum geometries in solids. So far, observations of eMChA hav…
▽ More
Chirality, a widely existing material property in nature involving the breaking of the left-right symmetry, has profound influences in various fields of natural sciences. Nonlinear response, such as electronic magnetochiral anisotropy (eMChA), has been recognized as a sensitive probe for the effects of symmetry breaking and nontrivial quantum geometries in solids. So far, observations of eMChA have primarily been limited to inversion-symmetry broken materials. Here, we report a remarkable chiral transport in centrosymmetric candidate topological superconductor 2M-WS2 flakes observed via second-harmonic generation under an out-of-plane magnetic field. More importantly, the eMChA becomes significant around the crossover temperature TFL ~ 25 K from the Fermi liquid (FL) to strange metal (SM) in the normal state, which interestingly echoes with the anomalously large Nernst response at the same temperature in bulk 2M-WS2. These observations reveal a direct correspondence between the nonlinear response, Nernst response, and FL-SM transition in 2M-WS2. Theoretical analysis indicates that nontrivial quantum geometry is behind the simultaneous response of eMChA and Nernst effects in 2M-WS2 and the contribution from the orbital magnetic moment at the Fermi surface becomes significant during the FL-SM transition. Based on first-principles calculations, a thick-layer-sliding mechanism with minimal energy gain in 2M-WS2 provides one possibility for the generation of such nontrivial quantum geometry. The intertwined physics of remarkable eMChA, Nernst response, and FL-SM transition make 2M-WS2 a rare quantum platform to study the chiral transport and unexplored phenomena in strange metals, which may shed light on the trans-century, unresolved scientific issue in unconventional high-temperature superconductivity.
△ Less
Submitted 18 May, 2026;
originally announced May 2026.
-
GRACE: Gradient-aligned Reasoning Data Curation for Efficient Post-training
Authors:
Junjie Li,
Ziao Wang,
NingXuan Ma,
Jianghong Ma,
Xiaofeng Zhang
Abstract:
Existing reasoning data curation pipelines score whole samples, treating every intermediate step as equally valuable. In reality, steps within a trace contribute very unevenly, and selecting reasoning data well requires assessing them individually. We present GRACE, a gradient-aligned curation method that views each reasoning trace as a sequence of optimization events and scores every step by two…
▽ More
Existing reasoning data curation pipelines score whole samples, treating every intermediate step as equally valuable. In reality, steps within a trace contribute very unevenly, and selecting reasoning data well requires assessing them individually. We present GRACE, a gradient-aligned curation method that views each reasoning trace as a sequence of optimization events and scores every step by two complementary signals: its alignment with the answer-oriented gradient direction, and its consistency with the preceding reasoning trajectory. Step-level scores are aggregated into a sample-level value for subset selection, using only the model's internal optimization signals and no external reward models or step annotations. To make this scalable, GRACE introduces a representation-level gradient proxy that estimates step-level alignment from token-level upstream signals in a single forward pass. Post-training Qwen3-VL-2B-Instruct on MMathCoT-1M, GRACE reaches 108.8% of the full-data performance with 20% of the data and retains 100.2% with only 5%, with subsets that transfer effectively across model backbones.
△ Less
Submitted 13 May, 2026;
originally announced May 2026.
-
MLGIB: Multi-Label Graph Information Bottleneck for Expressive and Robust Message Passing
Authors:
Chaokai Wu,
Haofu Shi,
Ningxuan Ma,
Jianghong Ma,
Xiaofeng Zhang
Abstract:
Graph Neural Networks (GNNs) suffer from over-squashing in deep message passing, where information from exponentially growing neighborhoods is compressed into fixed-dimensional representations. We show that this issue becomes a distinct failure mode in multi-label graphs: neighboring nodes often share only limited labels while differing across many irrelevant ones, causing predictive signals to be…
▽ More
Graph Neural Networks (GNNs) suffer from over-squashing in deep message passing, where information from exponentially growing neighborhoods is compressed into fixed-dimensional representations. We show that this issue becomes a distinct failure mode in multi-label graphs: neighboring nodes often share only limited labels while differing across many irrelevant ones, causing predictive signals to be diluted by noisy label information. To address this challenge, we propose the Multi-Label Graph Information Bottleneck (MLGIB), which formulates multi-label message passing as constrained information transmission under irrelevant label noise. MLGIB balances expressiveness and robustness by preserving predictive label signals while suppressing irrelevant noise. Specifically, it constructs a Markovian dependence space and derives tractable variational bounds, where the lower bound maximizes mutual information with target labels and the upper bound constrains redundant source information. These bounds lead to an end-to-end label-aware message-passing architecture. Extensive experiments on multiple benchmarks demonstrate consistent improvements over existing methods, validating the effectiveness and generality of the proposed framework.
△ Less
Submitted 14 May, 2026; v1 submitted 13 May, 2026;
originally announced May 2026.
-
Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions
Authors:
Diancheng Kang,
Zheyuan Liu,
Ningshan Ma,
Yue Huang,
Zhaoxuan Tan,
Meng Jiang
Abstract:
Activation steering controls language model behavior by adding directions to internal representations at inference time, but standard residual-stream steering can fail in stateful dialogue. We identify KV-cache contamination as a key failure mode: steered token states are stored and repeatedly reused, turning a local perturbation into cumulative coherence degradation. To address this challenge, we…
▽ More
Activation steering controls language model behavior by adding directions to internal representations at inference time, but standard residual-stream steering can fail in stateful dialogue. We identify KV-cache contamination as a key failure mode: steered token states are stored and repeatedly reused, turning a local perturbation into cumulative coherence degradation. To address this challenge, we propose Gated Cropped Attention-Delta steering (GCAD), which extracts steering signals from system-prompt contributions to self-attention and applies them with token-level gating. Across persona-steering experiments, GCAD preserves trait control while substantially improving long-horizon coherence. On the main multi-turn benchmark, GCAD improves average coherence drift from -18.6 to -1.9 and raises turn-10 trait expression from 78.0 to 93.1. These results suggest that activation steering becomes more reliable when interventions follow the prompt-mediated pathways that models already use for behavioral control.
△ Less
Submitted 14 May, 2026; v1 submitted 11 May, 2026;
originally announced May 2026.
-
A Breast Vision Pathology Foundation Model for Real-world Clinical Utility
Authors:
Yingxue Xu,
Zhengyu Zhang,
Xiuming Zhang,
Mengwei Xu,
Fengtao Zhou,
Yihui Wang,
Jiabo Ma,
Yi Xin,
Danyi Li,
Chengyu Lu,
Zhijian Cen,
Ying Tan,
Qingbing Yao,
Qi Wang,
Zizhao Gao,
Yong Zhang,
Jingjing Chen,
Feifei Liu,
Qian Xu,
Yi Dai,
Hongxuan Tan,
Cheng Jin,
Huajun Zhou,
Zhengrui Guo,
Ling Liang
, et al. (10 additional authors not shown)
Abstract:
Pathology foundation models have shown strong retrospective performance, but whether such systems can support clinically relevant use remains unclear. This challenge is particularly important in breast cancer, where pathological assessment serves as the gold standard for diagnosis and guides treatment planning, surgical decision-making and risk stratification across pre-, intra- and post-operative…
▽ More
Pathology foundation models have shown strong retrospective performance, but whether such systems can support clinically relevant use remains unclear. This challenge is particularly important in breast cancer, where pathological assessment serves as the gold standard for diagnosis and guides treatment planning, surgical decision-making and risk stratification across pre-, intra- and post-operative stages. Here we present \textbf{BRAVE}, a breast-adaptive pathology foundation model developed and evaluated using a total resource of 101,638 breast whole-slide images from 32 sources across Asia, Europe and North America. We assessed BRAVE across 34 tasks in 82 cohorts spanning pre-operative biopsy, intra-operative frozen section and post-operative resection, using an evidence chain comprising retrospective benchmarking, clinically challenging scenarios, workflow-oriented clinical impact simulations, prospective observational validation with the thresholds locked in the retrospective cohorts and crossover pathologist-AI interaction studies. Across these settings, BRAVE supported practical roles in the clinical workflow, including safe exclusion of low-risk cases from routine review, AI-assisted second-review rescue of initially missed positives and prioritization of cases for further assessment. In prospective validation across three centres, BRAVE excluded 76.9% of negative biopsy cases (NPV 0.953) and 70.1% of negative frozen-section cases (NPV 0.973), and triaged 78.8% of post-operative subtyping cases as high-confidence clear-cut cases (NPV 1.000). In reader studies, AI assistance improved balanced accuracy from 88.5% to 95.1% (OR 3.14, P<0.001), with better efficiency, confidence and inter-rater agreement. BRAVE-derived scores also independently predicted disease-free survival (adjusted HR 4.79, P<0.001) and overall survival (adjusted HR 8.14, P<0.001).
△ Less
Submitted 6 May, 2026;
originally announced May 2026.
-
Where's the Plan? Locating Latent Planning in Language Models with Lightweight Mechanistic Interventions
Authors:
Nicole Ma,
Nick Rui
Abstract:
We study planning site formation in language models -- where internal representations of structurally-constrained future tokens form during the forward pass, and whether they causally drive generation. Using rhyming-couplet completion as a clean test of forward-looking constraint, we apply two lightweight methods (linear probing and activation patching) across Qwen3, Gemma-3, and Llama-3 at more t…
▽ More
We study planning site formation in language models -- where internal representations of structurally-constrained future tokens form during the forward pass, and whether they causally drive generation. Using rhyming-couplet completion as a clean test of forward-looking constraint, we apply two lightweight methods (linear probing and activation patching) across Qwen3, Gemma-3, and Llama-3 at more than ten scales. Probing shows that future-rhyme information is linearly decodable at the line boundary, with signal that strengthens with scale in all three families. Activation patching reveals that only Gemma-3-27B causally relies on this encoding, exhibiting a handoff in which the causal driver migrates from the rhyme word to the line boundary around layer 30. Every other model we test conditions on the rhyme word throughout generation, with near-zero causal effect at the line boundary despite strong probe signal. We localize the Gemma-3-27B handoff to five attention heads through two-stage path patching that recover ~90% of the rhyme-routing capacity at the newline.
△ Less
Submitted 11 June, 2026; v1 submitted 8 May, 2026;
originally announced May 2026.
-
Observation of the Magnus Nonlinear Hall effect from Chiral Weyl Monopoles
Authors:
Heda Zhang,
Nikolai Peshcherenko,
Ning Mao,
Nianlong Zou,
Jiaqiang Yan,
Claudia Felser,
Yang Zhang
Abstract:
The nonlinear Hall effect (NLHE) connects crystalline symmetry to quantum geometry, offering a probe of band topology beyond linear transport. While most studies have focused on the Berry curvature dipole in low-symmetry crystals, mechanisms that directly probe Berry monopoles in higher-symmetry chiral lattices remain unexplored. Here, we report the observations of the NLHE in the chiral Weyl semi…
▽ More
The nonlinear Hall effect (NLHE) connects crystalline symmetry to quantum geometry, offering a probe of band topology beyond linear transport. While most studies have focused on the Berry curvature dipole in low-symmetry crystals, mechanisms that directly probe Berry monopoles in higher-symmetry chiral lattices remain unexplored. Here, we report the observations of the NLHE in the chiral Weyl semimetal CoSi, a platform where the Berry curvature dipole is symmetry-forbidden. By employing focused ion beam-fabricated crossbar devices, we detect a robust second-harmonic Hall voltage under zero magnetic field, hosting all key signatures of the NLHE. Theoretical analysis attributes the nonlinear Hall conductivity to skew scattering of self-rotating electron wave packets, whose chirality is dictated by the underlying band topology, a process reminiscent of the classical Magnus effect. Furthermore, the NLHE signal exhibits a temperature-dependent sign reversal, and a strong, linearly field-dependent modulation that grows with carrier mobility, directly reflecting the topological Weyl nodes distribution near the Fermi level. These findings establish CoSi as a platform for Berry monopole-driven nonlinear transport, demonstrating a skew-scattering route to topological nonlinear Hall responses that bypasses conventional symmetry constraints.
△ Less
Submitted 30 April, 2026;
originally announced April 2026.
-
The Chameleon's Limit: Investigating Persona Collapse and Homogenization in Large Language Models
Authors:
Yunze Xiao,
Vivienne J. Zhang,
Chenghao Yang,
Ningshan Ma,
Weihao Xuan,
Jen-tse Huang
Abstract:
Applications based on large language models (LLMs), such as multi-agent simulations, require population diversity among agents. We identify a pervasive failure mode we term \emph{Persona Collapse}: agents each assigned a distinct profile nonetheless converge into a narrow behavioral mode, producing a homogeneous simulated population. To quantify persona collapse, we propose a framework that measur…
▽ More
Applications based on large language models (LLMs), such as multi-agent simulations, require population diversity among agents. We identify a pervasive failure mode we term \emph{Persona Collapse}: agents each assigned a distinct profile nonetheless converge into a narrow behavioral mode, producing a homogeneous simulated population. To quantify persona collapse, we propose a framework that measures how much of the persona space a population occupies (Coverage), how evenly agents spread across it (Uniformity), and how rich the resulting behavioral patterns are (Complexity). Evaluating ten LLMs on personality simulation (BFI-44), moral reasoning, and self-introduction, we observe persona collapse along two axes: (1) Dimensions: a model can appear diverse on one axis yet structurally degenerate on another, and (2) Domains: the same model may collapse the most in personality yet be the most diverse in moral reasoning. Furthermore, item-level diagnostics reveal that behavioral variation tracks coarse demographic stereotypes rather than the fine-grained individual differences specified in each persona. Counter-intuitively, \textbf{the models achieving the highest per-persona fidelity consistently produce the most stereotyped populations}. We release our toolkit and data to support population-level evaluation of LLMs.
△ Less
Submitted 27 April, 2026;
originally announced April 2026.
-
Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding
Authors:
Zhixuan Wu,
Quanxing Zha,
Teng Wang,
Genbao Xu,
Wenyuan Gu,
Wei Rao,
Nan Ma,
Bo Cheng,
Soujanya Poria
Abstract:
Video understanding requires identifying and reasoning over semantically discriminative visual objects across frames, yet existing object-agnostic solutions struggle to effectively handle substantial object variations over time. To address this, we introduce Chain-of-Glimpse, a search-guided progressive object-grounded reasoning framework that explicitly anchors each reasoning step to specific vis…
▽ More
Video understanding requires identifying and reasoning over semantically discriminative visual objects across frames, yet existing object-agnostic solutions struggle to effectively handle substantial object variations over time. To address this, we introduce Chain-of-Glimpse, a search-guided progressive object-grounded reasoning framework that explicitly anchors each reasoning step to specific visual evidence regions, enabling compositional and multi-step decision-making. Formally, Chain-of-Glimpse formulates video reasoning as a step-by-step process that incrementally builds spatially grounded traces around task-relevant visual objects, thereby mitigating over-reliance on saliency-driven cues. Specifically, Chain-of-Glimpse features a search-guided controller, optimized via reinforcement learning with a format reward that significantly incentivizes grounding capability, to iteratively ground visual evidence regions and form reliable reasoning trajectories, yielding accurate and interpretable multi-step decisions. Extensive evaluations on both in domain NExTQA and out-of-domain Video-Holmes, CG-Bench Reasoning, and VRBench benchmarks demonstrate consistent performance gains, robustness and generalization of Chain-of-Glimpse across diverse video reasoning tasks.
△ Less
Submitted 15 May, 2026; v1 submitted 16 April, 2026;
originally announced April 2026.
-
Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models
Authors:
Haoyi Sun,
Xiaoxiao Wang,
Ning Mao,
Qian Wang,
Lifu Mu,
Wen Zheng,
Tao Wei,
Wei Chen
Abstract:
Vision-Language Models (VLMs) have shown remarkable capabilities in joint vision-language understanding, but their large scale poses significant challenges for deployment in resource-constrained scenarios. Knowledge Distillation (KD) offers a viable way to improve model capabilities without increasing model size or data requirements, making deployment more efficient. However, applying KD to VLMs i…
▽ More
Vision-Language Models (VLMs) have shown remarkable capabilities in joint vision-language understanding, but their large scale poses significant challenges for deployment in resource-constrained scenarios. Knowledge Distillation (KD) offers a viable way to improve model capabilities without increasing model size or data requirements, making deployment more efficient. However, applying KD to VLMs is challenged by modality-specific supervision: although multimodal knowledge in VLMs is fused within the language space, current methods supervise each modality separately without explicitly addressing multimodal alignment, leading to inconsistent multimodal knowledge transfer. To address this, we propose Switch-KD, a visual-switch distillation framework that unifies vision-language knowledge transfer within a shared text-probability space. Switch-KD comprises two key components: (1) Visual-Switch Distillation, which switches the student's visual outputs into the teacher's language pathway to construct cross-modal probabilistic references for implicit visual knowledge transfer; and (2) Dynamic Bi-directional Logits Difference (DBiLD) loss, which adaptively aligns informative probability regions while preserving the distributional structures of teacher and student through bidirectional supervision. Guided by Switch-KD, a 0.5B TinyLLaVA effectively distills rich multimodal knowledge from its 3B teacher, yielding an average improvement of 3.6 points across 10 multimodal benchmarks without any architectural modification.
△ Less
Submitted 16 April, 2026;
originally announced April 2026.
-
Say Something Else: Rethinking Contextual Privacy as Information Sufficiency
Authors:
Yunze Xiao,
Wenkai Li,
Xiaoyuan Wu,
Ningshan Ma,
Yueqi Song,
Weihao Xuan
Abstract:
LLM agents increasingly draft messages on behalf of users, yet users routinely overshare sensitive information and disagree on what counts as private. Existing systems support only suppression (omitting sensitive information) and generalization (replacing information with an abstraction), and are typically evaluated on single isolated messages, leaving both the strategy space and evaluation settin…
▽ More
LLM agents increasingly draft messages on behalf of users, yet users routinely overshare sensitive information and disagree on what counts as private. Existing systems support only suppression (omitting sensitive information) and generalization (replacing information with an abstraction), and are typically evaluated on single isolated messages, leaving both the strategy space and evaluation setting incomplete. We formalize privacy-preserving LLM communication as an \textbf{Information Sufficiency (IS)} task, introduce \textbf{free-text pseudonymization} as a third strategy that replaces sensitive attributes with functionally equivalent alternatives, and propose a \textbf{conversational evaluation protocol} that assesses strategies under realistic multi-turn follow-up pressure. Across 792 scenarios spanning three power-relation types (institutional, peer, intimate) and three sensitivity categories (discrimination risk, social cost, boundary), we evaluate seven frontier LLMs on privacy at two granularities, covertness, and utility. Pseudonymization yields the strongest privacy\textendash utility tradeoff overall, and single-message evaluation systematically underestimates leakage, with generalization losing up to 16.3 percentage points of privacy under follow-up.
△ Less
Submitted 7 April, 2026;
originally announced April 2026.
-
Generative Channel Knowledge Base With Environmental Information for Joint Source-Channel Coding in Semantic Communications
Authors:
Xudong Long,
Hao Chen,
Dan Wang,
Chen Qiu,
Nan Ma,
Xiaodong Xu,
Yubin Zhao
Abstract:
Semantic knowledge bases are regarded as a promising technology for upcoming 6G communications. However, existing studies mainly focus on source-side semantic modeling while overlooking the structural impact of propagation environments on semantic transmission performance. To address this issue, we propose a generative channel knowledge base (CKB) with environmental information to facilitate joint…
▽ More
Semantic knowledge bases are regarded as a promising technology for upcoming 6G communications. However, existing studies mainly focus on source-side semantic modeling while overlooking the structural impact of propagation environments on semantic transmission performance. To address this issue, we propose a generative channel knowledge base (CKB) with environmental information to facilitate joint source-channel coding (JSCC) in semantic communications (SemCom) systems. First, to enable the construction of the CKB, an environment-aware dataset is established by collecting spatial position information, global image features, fine-grained semantic features, and the corresponding channel matrices. A region-of-interest (ROI)-based filtering algorithm is further designed to remove semantic components that are irrelevant to signal propagation. Second, a Transformer-based generative framework is developed to learn the mapping between multidimensional environmental information and channel matrices. A self-attention mechanism is introduced to adaptively fuse heterogeneous features, enabling the construction of a structured CKB. Third, a CKB-driven JSCC SemCom architecture is proposed, where the generated channel knowledge is injected into both of the encoder and decoder to jointly exploit source semantics and channel-environment priors in an end-to-end manner. Experimental results demonstrate that the proposed multidimensional feature fusion method achieves a channel matrix estimation error at the $10^{-3}$ level. Moreover, the CKB-driven JSCC SemCom framework integrated into SemCom systems significantly outperforms existing benchmark schemes in terms of transmission performance.
△ Less
Submitted 6 April, 2026;
originally announced April 2026.
-
Interpretation of Crystal Energy Landscapes with Kolmogorov-Arnold Networks
Authors:
Gen Zu,
Ning Mao,
Claudia Felser,
Yang Zhang
Abstract:
Characterizing crystalline energy landscapes is essential to predicting thermodynamic stability, electronic structure, and functional behavior. While machine learning (ML) enables rapid property predictions, the "black-box" nature of most models limits their utility for generating new scientific insights. Here, we introduce Kolmogorov-Arnold Networks (KANs) as an interpretable framework to bridge…
▽ More
Characterizing crystalline energy landscapes is essential to predicting thermodynamic stability, electronic structure, and functional behavior. While machine learning (ML) enables rapid property predictions, the "black-box" nature of most models limits their utility for generating new scientific insights. Here, we introduce Kolmogorov-Arnold Networks (KANs) as an interpretable framework to bridge this gap. Unlike conventional neural networks with fixed activation functions, KANs employ learnable functions that reveal underlying physical relationships. We developed the Element-Weighted KAN, a composition-only model that achieves state-of-the-art accuracy in predicting formation energy, band gap, and work function across large-scale datasets. Crucially, without any explicit physical constraints, KANs uncover interpretable chemical trends aligned with the periodic table and quantum mechanical principles through embedding analysis, correlation studies, and principal component analysis. These results demonstrate that KANs provide a powerful framework with high predictive performance and scientific interpretability, establishing a new paradigm for transparent, chemistry-based materials informatics.
△ Less
Submitted 6 April, 2026;
originally announced April 2026.
-
AD-CARE: A Guideline-grounded, Modality-agnostic LLM Agent for Real-world Alzheimer's Disease Diagnosis with Multi-cohort Assessment, Fairness Analysis, and Reader Study
Authors:
Wenlong Hou,
Sheng Bi,
Guangqian Yang,
Lihao Liu,
Ye Du,
Hanxiao Xue,
Juncheng Wang,
Yuxiang Feng,
Yue Xun,
Nanxi Yu,
Ning Mao,
Mo Yang,
Yi Wah Eva Cheung,
Ling Long,
Kay Chen Tan,
Lequan Yu,
Xiaomeng Ma,
Shaozhen Yan,
Shujun Wang
Abstract:
Alzheimer's disease (AD) is a growing global health challenge as populations age, and timely, accurate diagnosis is essential to reduce individual and societal burden. However, real-world AD assessment is hampered by incomplete, heterogeneous multimodal data and variability across sites and patient demographics. Although large language models (LLMs) have shown promise in biomedicine, their use in…
▽ More
Alzheimer's disease (AD) is a growing global health challenge as populations age, and timely, accurate diagnosis is essential to reduce individual and societal burden. However, real-world AD assessment is hampered by incomplete, heterogeneous multimodal data and variability across sites and patient demographics. Although large language models (LLMs) have shown promise in biomedicine, their use in AD has largely been confined to answering narrow, disease-specific questions rather than generating comprehensive diagnostic reports that support clinical decision-making. Here we expand LLM capabilities for clinical decision support by introducing AD-CARE, a modality-agnostic agent that performs guideline-grounded diagnostic assessment from incomplete, heterogeneous inputs without imputing missing modalities. By dynamically orchestrating specialized diagnostic tools and embedding clinical guidelines into LLM-driven reasoning, AD-CARE generates transparent, report-style outputs aligned with real-world clinical workflows. Across six cohorts comprising 10,303 cases, AD-CARE achieved 84.9% diagnostic accuracy, delivering 4.2%-13.7% relative improvements over baseline methods. Despite cohort-level differences, dataset-specific accuracies remain robust (80.4%-98.8%), and the agent consistently outperforms all baselines. AD-CARE reduced performance disparities across racial and age subgroups, decreasing the average dispersion of four metrics by 21%-68% and 28%-51%, respectively. In a controlled reader study, the agent improved neurologist and radiologist accuracy by 6%-11% and more than halved decision time. The framework yielded 2.29%-10.66% absolute gains over eight backbone LLMs and converges their performance. These results show that AD-CARE is a scalable, practically deployable framework that can be integrated into routine clinical workflows for multimodal decision support in AD.
△ Less
Submitted 26 March, 2026;
originally announced March 2026.
-
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale
Authors:
Yicheng Zou,
Dongsheng Zhu,
Lin Zhu,
Tong Zhu,
Yunhua Zhou,
Peiheng Zhou,
Xinyu Zhou,
Dongzhan Zhou,
Zhiwang Zhou,
Yuhao Zhou,
Bowen Zhou,
Zhanping Zhong,
Zhijie Zhong,
Haiteng Zhao,
Penghao Zhao,
Xiaomeng Zhao,
Zhiyuan Zhao,
Yechen Zhang,
Jin Zhang,
Wenwei Zhang,
Hongjie Zhang,
Zhuo Zhang,
Wenlong Zhang,
Bo Zhang,
Chao Zhang
, et al. (152 additional authors not shown)
Abstract:
We introduce Intern-S1-Pro, the first one-trillion-parameter scientific multimodal foundation model. Scaling to this unprecedented size, the model delivers a comprehensive enhancement across both general and scientific domains. Beyond stronger reasoning and image-text understanding capabilities, its intelligence is augmented with advanced agent capabilities. Simultaneously, its scientific expertis…
▽ More
We introduce Intern-S1-Pro, the first one-trillion-parameter scientific multimodal foundation model. Scaling to this unprecedented size, the model delivers a comprehensive enhancement across both general and scientific domains. Beyond stronger reasoning and image-text understanding capabilities, its intelligence is augmented with advanced agent capabilities. Simultaneously, its scientific expertise has been vastly expanded to master over 100 specialized tasks across critical science fields, including chemistry, materials, life sciences, and earth sciences. Achieving this massive scale is made possible by the robust infrastructure support of XTuner and LMDeploy, which facilitates highly efficient Reinforcement Learning (RL) training at the 1-trillion parameter level while ensuring strict precision consistency between training and inference. By seamlessly integrating these advancements, Intern-S1-Pro further fortifies the fusion of general and specialized intelligence, working as a Specializable Generalist, demonstrating its position in the top tier of open-source models for general capabilities, while outperforming proprietary models in the depth of specialized scientific tasks.
△ Less
Submitted 2 April, 2026; v1 submitted 26 March, 2026;
originally announced March 2026.
-
OAHuman: Occlusion-Aware 3D Human Reconstruction from Monocular Images
Authors:
Yuanwang Yang,
Hongliang Liu,
Muxin Zhang,
Nan Ma,
Jingyu Yang,
Yu-Kun Lai,
Kun Li
Abstract:
Monocular 3D human reconstruction in real-world scenarios remains highly challenging due to frequent occlusions from surrounding objects, people, or image truncation. Such occlusions lead to missing geometry and unreliable appearance cues, severely degrading the completeness and realism of reconstructed human models. Although recent neural implicit methods achieve impressive results on clean input…
▽ More
Monocular 3D human reconstruction in real-world scenarios remains highly challenging due to frequent occlusions from surrounding objects, people, or image truncation. Such occlusions lead to missing geometry and unreliable appearance cues, severely degrading the completeness and realism of reconstructed human models. Although recent neural implicit methods achieve impressive results on clean inputs, they struggle under occlusion due to entangled modeling of shape and texture. In this paper, we propose OAHuman, an occlusion-aware framework that explicitly decouples geometry reconstruction and texture synthesis for robust 3D human modeling from a single RGB image. The core innovation lies in the decoupling-perception paradigm, which addresses the fundamental issue of geometry-texture cross-contamination in occluded regions. Our framework ensures that geometry reconstruction is perceptually reinforced even in occluded areas, isolating it from texture interference. In parallel, texture synthesis is learned exclusively from visible regions, preventing texture errors from being transferred to the occluded areas. This decoupling approach enables OAHuman to achieve robust and high-fidelity reconstruction under occlusion, which has been a long-standing challenge in the field. Extensive experiments on occlusion-rich benchmarks demonstrate that OAHuman achieves superior performance in terms of structural completeness, surface detail, and texture realism, significantly improving monocular 3D human reconstruction under occlusion conditions.
△ Less
Submitted 15 March, 2026;
originally announced March 2026.
-
Systematic study of superheavy nuclei within a microscopic collective Hamiltonian: Impact of quantum shape fluctuations
Authors:
X. Q. Yang,
R. Y. Hu,
R. N. Mao,
J. Xiang,
Z. P. Li
Abstract:
The even-even superheavy nuclei with $104 \leqslant Z \leqslant 126$ and $N\leqslant 258$ have been investigated using a microscopic five-dimensional collective Hamiltonian (5DCH) based on constrained triaxial relativistic Hartree-Bogoliubov calculations with the PC-PK1 density functional. The 5DCH approach effectively captures the characteristic of isospin dependence of nuclear binding energies,…
▽ More
The even-even superheavy nuclei with $104 \leqslant Z \leqslant 126$ and $N\leqslant 258$ have been investigated using a microscopic five-dimensional collective Hamiltonian (5DCH) based on constrained triaxial relativistic Hartree-Bogoliubov calculations with the PC-PK1 density functional. The 5DCH approach effectively captures the characteristic of isospin dependence of nuclear binding energies, two-nucleon separation energies, and $α$-decay energies across isotopic chains and demonstrates consistent accuracy as $Z$ increases, underscoring the model's predictive power. The collective potentials, average quadrupole deformations, and characteristic collective observables: $E(2^+_1)$, $R_{42}$, and $B(E2; 2^+_1\to 0^+_1)$ reveal a shape transition from well-prolate deformation around $N=150$ and $N=210$ to medium-deformed $γ$-soft shape around $N=176$ and $N=246$, and finally to a spherical shape near $N=184$ and $N=258$ for the isotopic chains with $104\leqslant Z\leqslant 118$. Oblate deformations are favored for $Z\geqslant 120$ isotopes around $N=178$. Remarkably, for a substantial range of transitional superheavy nuclei with $N\gtrsim184$ and $N\gtrsim240$, no $0^+$ states bounded by the fission saddles are predicted within their very shallow potential wells due to quantum shape fluctuations (QSFs). Additionally, sharp variations predicted for two-neutron separation energies $S_{2n}$ and $α$-decay energies $Q_α$ at $N=184$ and $258$ in mean-field calculations are significantly reduced and shifted to $N=182$ and $256$ in the 5DCH calculations, which is caused by the rapid evolution of the dynamical correlation energies related to QSFs around the nuclear spherical shells.
△ Less
Submitted 10 March, 2026;
originally announced March 2026.
-
Unlocking High-Fidelity Analog Joint Source-Channel Coding on Standard Digital Transceivers
Authors:
Shumin Yao,
Hao Chen,
Yaping Sun,
Nan Ma,
Xiaodong Xu,
Qinglin Zhao,
Shuguang Cui
Abstract:
Analog joint source-channel coding (JSCC) has demonstrated superior performance for semantic communications through graceful degradation across channel conditions. However, a fundamental hardware-software mismatch prevents deployment on modern digital physical layers (PHYs): analog JSCC generates continuous-valued symbols requiring infinite waveform diversity, while digital PHYs produce a finite s…
▽ More
Analog joint source-channel coding (JSCC) has demonstrated superior performance for semantic communications through graceful degradation across channel conditions. However, a fundamental hardware-software mismatch prevents deployment on modern digital physical layers (PHYs): analog JSCC generates continuous-valued symbols requiring infinite waveform diversity, while digital PHYs produce a finite set of discrete waveforms and employ non-differentiable operations that break end-to-end gradient flow. Existing solutions either fundamentally limit representation granularity or require impractical white-box PHY access. We introduce D2AJSCC, a novel framework enabling high-fidelity analog JSCC deployment on standard digital PHYs. Our approach exploits orthogonal frequency-division multiplexing's parallel subcarrier structure as a waveform synthesizer: computational PHY inversion determines input bitstreams that orchestrate subcarrier amplitudes and phases to emulate ideal analog waveforms. To enable end-to-end training despite non-differentiable PHY operations, we develop ProxyNet-a differentiable neural surrogate of the communication link that provides uninterrupted gradient flow while preventing JSCC degeneration. Simulation results for image transmission over WiFi PHY demonstrate that our system achieves near-ideal analog JSCC performance with graceful degradation across SNR conditions, while baselines exhibit cliff effects or catastrophic failures. By enabling next-generation semantic transmission on legacy infrastructure without hardware modification, our framework promotes sustainable network evolution and bridges the critical gap between analog JSCC's theoretical promise and practical deployment on ubiquitous digital hardware.
△ Less
Submitted 9 March, 2026;
originally announced March 2026.
-
TempoFit: Plug-and-Play Layer-Wise Temporal KV Memory for Long-Horizon Vision-Language-Action Manipulation
Authors:
Jun Sun,
Boyu Yang,
Jiahao Zhang,
Ning Ma,
Chencheng Wu,
Siqing Zhang,
Yiou Huang,
Qiufeng Wang,
Shan Liang,
Yaran Chen
Abstract:
Pretrained Vision-Language-Action (VLA) policies have achieved strong single-step manipulation, but their inference remains largely memoryless, which is brittle in non-Markovian long-horizon settings with occlusion, state aliasing, and subtle post-action changes. Prior approaches inject history either by stacking frames, which scales visual tokens and latency while adding near-duplicate pixels, or…
▽ More
Pretrained Vision-Language-Action (VLA) policies have achieved strong single-step manipulation, but their inference remains largely memoryless, which is brittle in non-Markovian long-horizon settings with occlusion, state aliasing, and subtle post-action changes. Prior approaches inject history either by stacking frames, which scales visual tokens and latency while adding near-duplicate pixels, or by learning additional temporal interfaces that require (re-)training and may break the original single-frame inference graph. We present TempoFit, a training-free temporal retrofit that upgrades frozen VLAs through state-level memory. Our key insight is that prefix attention K/V already form a model-native, content-addressable runtime state; reusing them across timesteps introduces history without new tokens or trainable modules. TempoFit stores layer-wise FIFO prefix K/V at selected intermediate layers, performs parameter-free K-to-K retrieval with Frame-Gap Temporal Bias (FGTB), a fixed recency bias inspired by positional biases in NLP, to keep decisions present-dominant, and injects the retrieved context via pre-attention residual loading with norm-preserving rescaling to avoid distribution shift under frozen weights. On LIBERO-LONG, TempoFit improves strong pretrained backbones by up to +4.0% average success rate while maintaining near-real-time latency, and it transfers consistently to CALVIN and real-robot long-horizon tasks.
△ Less
Submitted 8 March, 2026;
originally announced March 2026.
-
Evaluating the Search Agent in a Parallel World
Authors:
Jiawei Chen,
Xintian Shen,
Lihao Zheng,
Lifu Mu,
Haoyi Sun,
Ning Mao,
Hao Ma,
Tao Wei,
Pan Zhou,
Kun Zhan
Abstract:
Integrating web search tools has significantly extended the capability of LLMs to address open-world, real-time, and long-tail problems. However, evaluating these Search Agents presents formidable challenges. First, constructing high-quality deep search benchmarks is prohibitively expensive, while unverified synthetic data often suffers from unreliable sources. Second, static benchmarks face dynam…
▽ More
Integrating web search tools has significantly extended the capability of LLMs to address open-world, real-time, and long-tail problems. However, evaluating these Search Agents presents formidable challenges. First, constructing high-quality deep search benchmarks is prohibitively expensive, while unverified synthetic data often suffers from unreliable sources. Second, static benchmarks face dynamic obsolescence: as internet information evolves, complex queries requiring deep research often degrade into simple retrieval tasks due to increased popularity, and ground truths become outdated due to temporal shifts. Third, attribution ambiguity confounds evaluation, as an agent's performance is often dominated by its parametric memory rather than its actual search and reasoning capabilities. Finally, reliance on specific commercial search engines introduces variability that hampers reproducibility. To address these issues, we propose a novel framework, Mind-ParaWorld, for evaluating Search Agents in a Parallel World. Specifically, MPW samples real-world entity names to synthesize future scenarios and questions situated beyond the model's knowledge cutoff. A ParaWorld Law Model then constructs a set of indivisible Atomic Facts and a unique ground-truth for each question. During evaluation, instead of retrieving real-world results, the agent interacts with a ParaWorld Engine Model that dynamically generates SERPs grounded in these inviolable Atomic Facts. We release MPW-Bench, an interactive benchmark spanning 19 domains with 1,608 instances. Experiments across three evaluation settings show that, while search agents are strong at evidence synthesis given complete information, their performance is limited not only by evidence collection and coverage in unfamiliar search environments, but also by unreliable evidence sufficiency judgment and when-to-stop decisions-bottlenecks.
△ Less
Submitted 27 April, 2026; v1 submitted 4 March, 2026;
originally announced March 2026.
-
Force-Aware Residual DAgger via Trajectory Editing for Precision Insertion with Impedance Control
Authors:
Yiou Huang,
Ning Ma,
Weichu Zhao,
Zinuo Liu,
Jun Sun,
Qiufeng Wang,
Yaran Chen
Abstract:
Imitation learning (IL) has shown strong potential for contact-rich precision insertion tasks. However, its practical deployment is often hindered by covariate shift and the need for continuous expert monitoring to recover from failures during execution. In this paper, we propose Trajectory Editing Residual Dataset Aggregation (TER-DAgger), a scalable and force-aware human-in-the-loop imitation le…
▽ More
Imitation learning (IL) has shown strong potential for contact-rich precision insertion tasks. However, its practical deployment is often hindered by covariate shift and the need for continuous expert monitoring to recover from failures during execution. In this paper, we propose Trajectory Editing Residual Dataset Aggregation (TER-DAgger), a scalable and force-aware human-in-the-loop imitation learning framework that mitigates covariate shift by learning residual policies through optimization-based trajectory editing. This approach smoothly fuses policy rollouts with human corrective trajectories, providing consistent and stable supervision. Second, we introduce a force-aware failure anticipation mechanism that triggers human intervention only when discrepancies arise between predicted and measured end-effector forces, significantly reducing the requirement for continuous expert monitoring. Third, all learned policies are executed within a Cartesian impedance control framework, ensuring compliant and safe behavior during contact-rich interactions. Extensive experiments in both simulation and real-world precision insertion tasks show that TER-DAgger improves the average success rate by over 37\% compared to behavior cloning, human-guided correction, retraining, and fine-tuning baselines, demonstrating its effectiveness in mitigating covariate shift and enabling scalable deployment in contact-rich manipulation.
△ Less
Submitted 23 July, 2026; v1 submitted 4 March, 2026;
originally announced March 2026.