-
Magnetically Self-Sealed MR Haptic Actuator With PWM-Based Excitation and High-Fidelity Torque Control
Authors:
Dong Qiang,
Tian Yuan,
Song Yang,
Kequan Xia,
Thomas Reddyhoff,
Yikun Zhang,
Cheng Cheng,
Min Yu
Abstract:
Accurate and stable torque rendering is essential for safe and perceptive human--machine interaction. Magnetorheological fluid (MRF)-based actuators offer a compact and rapidly controllable solution for haptic feedback, but their practical implementation requires reliable fluid sealing, low-hysteresis excitation, accurate torque control, and stable long-duration operation. This article presents an…
▽ More
Accurate and stable torque rendering is essential for safe and perceptive human--machine interaction. Magnetorheological fluid (MRF)-based actuators offer a compact and rapidly controllable solution for haptic feedback, but their practical implementation requires reliable fluid sealing, low-hysteresis excitation, accurate torque control, and stable long-duration operation. This article presents an integrated MRF haptic system featuring a compact magnetically self-sealed rotary actuator, low-hysteresis PWM operation, high-fidelity model-based torque rendering, and stable performance during long-time operation. Magnetostatic simulation guides the arrangement of magnetic and nonmagnetic materials to focus flux in the multidisk torque and permanent-magnet sealing regions, enabling a maximum 600 N$\cdot$mm/A output. Experiments show that higher PWM frequencies reduce hysteresis and improve repeatability. At 10 kHz, the response is represented by a nonlinear model that varies with the direction and speed of torque change. The real-time controller combines feedforward, hysteresis compensation, PI feedback, and sliding-mode correction. Compared with PID, it reduces square-wave overshoot, undershoot, and steady-state RMSE by 77.4\%, 61.9\%, and 68.3\%, respectively. It tracks sinusoidal and biomechanics-model-based references, and a 1.5-h test shows only a 2.5 $^\circ$C rise near the coil with no clear tracking loss. This high-fidelity torque rendering will fundamentally transform human--robot collaboration by making interactions safer, more efficient, and more intuitive.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
Demystifying Oversmoothing in Sheaf Neural Networks: An Index-Theoretic Criterion
Authors:
Junwen Dong,
Yuhan Peng,
Hao Li,
Huitao Feng,
Kelin Xia
Abstract:
To combat oversmoothing in Graph Convolutional Networks, Sheaf Neural Networks (SNNs) were proposed as a generalization by equipping the graph with a sheaf structure and replacing the graph Laplacian with a sheaf Laplacian $\mathcal{L}$. Existing analyses connect sheaf diffusion to oversmoothing via the harmonic space ($\ker\mathcal{L}$), taking its absolute dimension as an indicator of anti-overs…
▽ More
To combat oversmoothing in Graph Convolutional Networks, Sheaf Neural Networks (SNNs) were proposed as a generalization by equipping the graph with a sheaf structure and replacing the graph Laplacian with a sheaf Laplacian $\mathcal{L}$. Existing analyses connect sheaf diffusion to oversmoothing via the harmonic space ($\ker\mathcal{L}$), taking its absolute dimension as an indicator of anti-oversmoothing capacity. However, absolute dimension alone is not a reliable measure: certain sheaf configurations inflate $\dim \ker \mathcal{L}$ while their harmonic sections remain entirely constant, without enriching discriminative capacity. We instead introduce the first relative, geometric approach, yielding a precise characterisation of anti-oversmoothing capacity. Under natural conditions on stalk transportation and global sheaf structure, we establish an index-theoretic comparison criterion showing that one sheaf's harmonic space genuinely contains another's beyond trivial inflation. We illustrate this with a concrete instance and further introduce \textit{GyroSheaf}, a sheaf with curved gyrovector-space stalks, extending the criterion to the non-linear setting via local tangent-space linearization. Experiments across ten models confirm the theoretical criterion: sheaf models violating the criterion collapse despite possessing index jumps, while compliant models maintain depth-stable representations.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Critical behavior and critical exponents of rotating QCD matter
Authors:
Kai Xiao,
Fei Sun,
Shuang Li,
Xun Chen
Abstract:
We investigate the thermodynamic properties and critical behavior of rotating strongly interacting matter within the two-flavor Nambu--Jona-Lasinio (NJL) model in the mean-field approximation. The phase structure and the critical endpoint (CEP) are determined in the temperature--angular velocity \((T,ω)\) plane. By analyzing the singular behavior of thermodynamic observables near the CEP, we extra…
▽ More
We investigate the thermodynamic properties and critical behavior of rotating strongly interacting matter within the two-flavor Nambu--Jona-Lasinio (NJL) model in the mean-field approximation. The phase structure and the critical endpoint (CEP) are determined in the temperature--angular velocity \((T,ω)\) plane. By analyzing the singular behavior of thermodynamic observables near the CEP, we extract the corresponding effective critical exponents characterizing the scaling behavior of the specific heat density, the rotational polarization discontinuity, the rotational susceptibility, and the critical-isotherm behavior of the rotational polarization. The obtained exponents approach the expected mean-field values and satisfy the corresponding scaling relations, indicating that the rotational degree of freedom does not alter the underlying mean-field critical scaling behavior within the present framework. These results provide a systematic characterization of rotation-induced critical phenomena and establish a basis for further studies of rotating QCD matter beyond the mean-field approximation.
△ Less
Submitted 13 August, 2026;
originally announced August 2026.
-
Monolithic integration of optically anisotropic GeSe-based films on GaAs by templated solid-phase epitaxy
Authors:
Kira J. Martin,
Autumn Y. Lee,
Pranav Mahaadev,
Pooja D. Reddy,
Kelly Xiao,
Tri Nguyen,
Ashlee M. García,
Aaron M. Lindenberg,
Kunal Mukherjee
Abstract:
Layered IV-VI semiconductors such as GeSe exhibit strong in-plane optical anisotropy, making them promising candidates for polarization-sensitive photonic devices. However, realizing these properties in scalable platforms requires heteroepitaxial integration on technologically relevant substrates like GaAs. Direct growth of GeSe is complicated by its glass formation at low temperatures and high va…
▽ More
Layered IV-VI semiconductors such as GeSe exhibit strong in-plane optical anisotropy, making them promising candidates for polarization-sensitive photonic devices. However, realizing these properties in scalable platforms requires heteroepitaxial integration on technologically relevant substrates like GaAs. Direct growth of GeSe is complicated by its glass formation at low temperatures and high vapor pressure at elevated temperatures. To overcome this, we develop a method for ex-situ solid-phase epitaxy utilizing a SnSe buffer and offcut GaAs substrate to enable single-orientation crystalline GeSe films. Using polarized reflection measurements, we find that stabilizing a single-in-plane-orientation results in a 2x increase in anisotropic response between the armchair and zigzag directions. This work provides a new integration route to harness the anisotropic optical properties of GeSe and its alloys for polarization-sensitive technologies.
△ Less
Submitted 4 August, 2026;
originally announced August 2026.
-
TRNet: Topography-Guided Frequency Rectification and Structure-Aware Decoding for Multimodal Paddy Rice Segmentation
Authors:
Kaiwen Xiao,
Chunlong Fu,
Liping Zheng,
Yanfeng Su
Abstract:
Mapping paddy rice from very-high-resolution imagery in mountainous and hilly regions is difficult because terrain alters optical appearance and increases confusion with visually similar vegetation. We present TRNet for 0.5-m GaoJing-1 red--green--blue (RGB) imagery, a 5-m TanDEM-X digital elevation model (DEM), and derived slope. Separate visual and terrain encoders preserve modality-specific fea…
▽ More
Mapping paddy rice from very-high-resolution imagery in mountainous and hilly regions is difficult because terrain alters optical appearance and increases confusion with visually similar vegetation. We present TRNet for 0.5-m GaoJing-1 red--green--blue (RGB) imagery, a 5-m TanDEM-X digital elevation model (DEM), and derived slope. Separate visual and terrain encoders preserve modality-specific features. At an early encoder stage, Topographic Energy-Spectral Rectification applies terrain-conditioned low-frequency modulation and asymmetric high-frequency regulation to suppress steep-slope clutter and conditionally enhance compatible low-slope rice cues. The Topography-guided Paddy Structure Decoder combines semantic, rice--background boundary, and interior cues, using coarse terrain as context. Experiments used an Area A internal test set and held-out Area B, which had steeper terrain and lower rice prevalence. TRNet achieved rice intersection-over-union (IoU) values of 85.10\% and 80.68\%, exceeding the original Dual-Encoder U-Net by 9.15 and 18.83 percentage points, respectively. Ablation and slope-stratified results linked these gains to frequency rectification, structure learning, and fewer steep-terrain false positives. The results support coarse topography as a contextual prior for very-high-resolution paddy rice mapping.
△ Less
Submitted 4 August, 2026;
originally announced August 2026.
-
Sample-half-inserted quantum interferometer
Authors:
Wei Li,
Tao Xie,
Yu-Hang Luo,
Kang Zheng,
Meiyu Peng,
Hui Yang,
Chunling Ding,
Chen-Zhi Yuan,
Omar S. Magana-Loaiza,
Keyu Xia,
Ryosuke Shimizu,
Hui Jing,
Chenglong You,
Rui-Bo Jin
Abstract:
Quantum technologies have been widely recognized as unprecedented opportunities for ultra-high precision metrology. As a celebrated example in modern quantum optics, the Hong-Ou-Mandel (HOM) interferometer is well-known for enabling temporal resolutions on the attosecond scale. However, the relatively low Fisher information per trial in ordinary HOM measurements typically necessitates tens of thou…
▽ More
Quantum technologies have been widely recognized as unprecedented opportunities for ultra-high precision metrology. As a celebrated example in modern quantum optics, the Hong-Ou-Mandel (HOM) interferometer is well-known for enabling temporal resolutions on the attosecond scale. However, the relatively low Fisher information per trial in ordinary HOM measurements typically necessitates tens of thousands of repetitions to achieve such precision. Here, we propose and demonstrate a sample-half-inserted HOM (SHOM) interferometer, which enhances the Fisher information by five orders of magnitude in a single interference event. By introducing an asymmetric photon-sample interaction, the SHOM configuration produces a distinctive dip-bump-dip interference structure, converting what was previously viewed as an artifact into a helpful metrological resource. Experimentally, we measured the optical path difference with an average precision of 4.09 nm (13.63 as) and an average accuracy of 1.22 nm (4.07 as) using $O(10^7)$ photons. Our results establish SHOM interferometry as an efficient phase-insensitive approach, not only paving the way toward practical quantum-enhanced thickness measurement for transparent materials, but also serving as an elegant strategy to improve the performance of various quantum devices.
△ Less
Submitted 4 August, 2026;
originally announced August 2026.
-
Effects of the inflaton mass on the pre-inflationary dynamics and primordial power spectra in loop quantum cosmology
Authors:
Kui Xiao,
Abolhassan Mohammadi,
Hongwei Tan
Abstract:
Considering the mass parameters as free phenomenological parameters, we investigate inflation driven by a canonical scalar field with different potentials in LQC. We find that a smaller mass value makes it easier to obtain a sufficient number of e-folds for the quadratic potential, while for the Starobinsky potential, more e-folds are obtained for larger mass values. By estimating the probability…
▽ More
Considering the mass parameters as free phenomenological parameters, we investigate inflation driven by a canonical scalar field with different potentials in LQC. We find that a smaller mass value makes it easier to obtain a sufficient number of e-folds for the quadratic potential, while for the Starobinsky potential, more e-folds are obtained for larger mass values. By estimating the probability of slow-roll inflation, we find that it remains very close to unity for all considered cases. We compute the primordial power spectrum using three approaches: the dressed metric, the hybrid, and the alternative mass function approaches for the case that the kinetic energy dominated at bounce. The resulting power spectrum for both potentials and different mass values shows the same qualitative pattern in all three approaches. It is suppressed at small $k$, amplified and oscillating over an intermediate range, and nearly scale-invariant at large $k$. The approaches mainly differ in how fast the power spectrum converges to this regime, with the hybrid approach converging the fastest, followed by the alternative mass function and dressed metric approaches. We find that the pivot scale $k_\star$ is highly sensitive to the inflaton mass, changing by more than an order of magnitude for a mass variation of only a few percent. Using the tensor power spectrum at $k_\star$, we calculate the scalar spectral index $n_s$ and tensor-to-scalar ratio $r$, finding good agreement with current data. Finally, the scalar power spectra are fed into the CAMB code to obtain the angular power spectrum and compare with the Planck 2018 data and the best-fit $Λ$CDM model. All three approaches are consistent with data at high multipoles, while at low multipoles the hybrid approach gives the closest agreement and the dressed metric approach shows the largest deviation.
△ Less
Submitted 2 August, 2026;
originally announced August 2026.
-
Towards optimal photometric calibration of digital astronomical plates with deep learning
Authors:
Mingyang Ma,
Haibo Yuan,
Lin Yang,
Kai Xiao,
Bowen Huang,
Shiyin Shen,
Zhengjun Shang,
Yong Yu,
Meiting Yang,
Zhenghong Tang,
Jianhai Zhao
Abstract:
Photometric calibration of digitized photographic plates is commonly modeled with separable magnitude-, color-, and position-dependent terms, but this separability can break down when image quality varies across the field in a magnitude-dependent way, leaving coupled spatial systematics in the residuals. We introduce a deep-learning calibration framework, the Multi-Feature Fused Network (MFF-Net),…
▽ More
Photometric calibration of digitized photographic plates is commonly modeled with separable magnitude-, color-, and position-dependent terms, but this separability can break down when image quality varies across the field in a magnitude-dependent way, leaving coupled spatial systematics in the residuals. We introduce a deep-learning calibration framework, the Multi-Feature Fused Network (MFF-Net), which takes instrumental magnitude, color, and pixel coordinates as input and learns a single nonlinear correction that jointly captures their coupled dependencies. Tests on 1{,}200 digitized Chinese plates show that MFF-Net consistently outperforms the MYX25 method (Ma et al. 2025), improving the 5th--95th percentile precision from 0.11--0.26~mag to 0.08--0.18~mag and delivering an approximately factor-of-two gain for bright sources. The learned correction largely removes the magnitude--position coupling seen in post-calibration residual maps, enabling higher-precision plate photometry and more reliable use of large historical plate archives.
△ Less
Submitted 2 August, 2026;
originally announced August 2026.
-
FibVLA: An Efficient Temporal Vision-Language-Action Model with Fibonacci Sampling
Authors:
Li Lin,
Wujun Xu,
Weiwei Meng,
Kaiwen Xia,
Kang Hao Cheong,
Shuai Wang
Abstract:
Vision-language-action models (VLAs), which leverage the cognition of multimodal information to infer physical-world actions, provide a generalized solution for embodied AI applications. Conventional VLAs usually concentrate on current digital cognition. While some efforts are made to enhance VLAs' reasoning capabilities by capturing temporal information, encoding the long-context history causes a…
▽ More
Vision-language-action models (VLAs), which leverage the cognition of multimodal information to infer physical-world actions, provide a generalized solution for embodied AI applications. Conventional VLAs usually concentrate on current digital cognition. While some efforts are made to enhance VLAs' reasoning capabilities by capturing temporal information, encoding the long-context history causes an efficiency-decreasing issue. To reconcile the conflict between capturing temporal information and maintaining inference efficiency in VLAs, this paper introduces FibVLA, an efficient framework featuring temporal perception of long-context history. Specifically, we leverage logarithmic hindsight sampling to both proprioceptive states and visual frames to capture long-term temporal dependencies with minimal redundancy. For the action expert, we introduce the flow matching to produce action distributions, and the Fibonacci recurrent inference strategy to generate long-range planning steps based on real-time closed-loop feedback. Experiments demonstrate that FibVLA significantly improves action smoothness and success rates without retraining large-scale visual encoders. Efficiency analysis demonstrates superior real-time responsiveness compared to video-based baselines in real-world evaluations.
△ Less
Submitted 31 July, 2026;
originally announced July 2026.
-
GPT-Red: Automated Red Teaming via Self-Play at Scale
Authors:
Eric Wallace,
Christopher A. Choquette-Choo,
Nikhil Kandpal,
Sam Toyer,
Dylan Hunn,
Stephanie Lin,
Yuxin Wen,
Xiangyu Qi,
Christopher Wolff,
Zizhao Wang,
Milad Nasr,
Sicheng Zhu,
Chuan Guo,
Juan Felipe Cerón Uribe,
Kaiwen Wang,
Aiden Low,
Kai Xiao,
Kai Chen
Abstract:
We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially train GPT-5.6, our most robust model to prompt injections to date. To create GPT-Red, we design a scalable self-play algorit…
▽ More
We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially train GPT-5.6, our most robust model to prompt injections to date. To create GPT-Red, we design a scalable self-play algorithm where the model is tasked with attacking a diverse population of simultaneously-trained defender agents. We train the model on realistic red-teaming environments using compute on the same scale as some of our largest RL post-training runs, making it the single-largest LLM safety training run ever documented. GPT-Red excels at red-teaming: it reliably breaks our past models up to GPT-5.5, it finds more successful attacks than human red-teamers, and it generalizes to held-out environments, defender models, and harnesses. In the future, we expect that as we improve the robustness of each new GPT model, it will in turn will provide better learning signal for \textit{even stronger} red-teamer agents, thus unlocking a self-improvement flywheel.
△ Less
Submitted 28 July, 2026;
originally announced July 2026.
-
Thermodynamics-Informed Input Reparameterization for Neural Prediction of Real-Fluid Thermodynamic Properties in Supercritical Combustion
Authors:
Haoze Zhang,
Han Li,
Ke Xiao,
Yangchen Xu,
Runze Mao,
Zhi X. Chen
Abstract:
Real-fluid thermodynamic property evaluation is a major computational cost in supercritical combustion simulations. In the enthalpy-based pressure-correction formulation, the closure evaluates temperature T, density $ρ$, and compressibility coefficient $ψ$ from the solver state (h,p,Y) through enthalpy-temperature inversion and repeated real-fluid equation-of-state evaluations. Neural-network surr…
▽ More
Real-fluid thermodynamic property evaluation is a major computational cost in supercritical combustion simulations. In the enthalpy-based pressure-correction formulation, the closure evaluates temperature T, density $ρ$, and compressibility coefficient $ψ$ from the solver state (h,p,Y) through enthalpy-temperature inversion and repeated real-fluid equation-of-state evaluations. Neural-network surrogates offer fixed-cost inference, but direct mapping from (h,p,Y) to $(T,ρ,ψ)$ must capture the enthalpy-temperature relation and non-ideal equation-of-state response, resulting in a complex regression problem. This work introduces a thermodynamics-informed input reparameterization strategy, termed target-aligned input reparameterization (TAIR). TAIR replaces the raw enthalpy coordinate of each property network with a target-matched thermodynamic coordinate: the temperature network uses a temperature estimate obtained by inverting a constant-$c_p$ ideal-gas mixture enthalpy approximation, whereas the density and compressibility networks use an ideal-gas density estimate. These algebraic transformations use only solver-available variables and species constants, guiding the networks to learn real-fluid departures from ideal-gas baselines rather than reconstructing the full closure from raw enthalpy. The method is assessed using supercritical methane-oxygen counterflow flame data against a raw-input baseline and target-inconsistent cross-reparameterization controls. TAIR reduces held-out RMSE by factors of about 1.5, 2.0, and 7.5 for T, $ρ$, and $ψ$, respectively. For an unseen strain-rate flame within the augmented thermodynamic envelope, the corresponding factors are 3.6, 14.5, and 6.0. The target-inconsistent controls perform worse, indicating that the gains arise from thermodynamically matched input design rather than generic preprocessing.
△ Less
Submitted 21 July, 2026;
originally announced July 2026.
-
Unconventional superconductivity in ScIr$_2$ chiral crystal with a kagome lattice
Authors:
Keqi Xia,
Jianzhou Zhao,
Igor Plokhikh,
Marisa Medarde,
Yang Xu,
Qingfeng Zhan,
Dariusz Jakub Gawryluk,
Toni Shiroka,
Tian Shang
Abstract:
Materials with a kagome lattice host exotic quantum phenomena driven by the interplay between band topology, spin-orbit coupling, magnetism, and electronic correlations. While magnetism of kagome materials has been widely investigated, their unconventional superconductivity (SC) remains largely unexplored due to the limited availability of suitable materials. Here, we report evidence of unconventi…
▽ More
Materials with a kagome lattice host exotic quantum phenomena driven by the interplay between band topology, spin-orbit coupling, magnetism, and electronic correlations. While magnetism of kagome materials has been widely investigated, their unconventional superconductivity (SC) remains largely unexplored due to the limited availability of suitable materials. Here, we report evidence of unconventional SC in the ScIr$_{2-x}$Si$_{x}$ family by combining muon-spin spectroscopy measurements with band-structure calculations. The parent ScIr$_2$ undergoes a structural phase transition from a high-$T$ cubic- to a low-$T$ rhombohedral phase, while the Ir kagome layer remains, albeit slightly distorted. Although the structural transition is suppressed by Si substitution, the superconducting pairing of ScIr$_{2-x}$Si$_{x}$ remains well described by a two-gap model. Since at least one of the gaps has nodes, this indicates an unconventional SC. Its unconventional nature can be explained by the distinct flat bands occurring near the Fermi level, leading to strong electronic correlations in the ScIr$_{2-x}$Si$_{x}$ family. Moreover, the low-$T$ phase of ScIr$_2$ exhibits an Ir chiral chain; therefore, it can be classified as a topological chiral crystal. Overall, the unusual properties of the ScIr$_{2-x}$Si$_{x}$ family make it an interesting, albeit rare, system for studying the interplay between unconventional SC, flat bands, and chirality.
△ Less
Submitted 18 July, 2026;
originally announced July 2026.
-
The existence of $k$-convex hypersurface for a class of Hessian curvature equations
Authors:
Kang Xiao,
Jiabao Gong
Abstract:
This article investigates the existence of closed, star-shaped hypersurfaces for a class of Hessian curvature equations. By combining a priori estimates with the continuity method, we establish the existence and uniqueness of $k$-convex hypersurfaces for both nonhomogeneous and homogeneous Hessian curvature equations, and by establishing a constant rank theorem, we prove that the resulting $k$-con…
▽ More
This article investigates the existence of closed, star-shaped hypersurfaces for a class of Hessian curvature equations. By combining a priori estimates with the continuity method, we establish the existence and uniqueness of $k$-convex hypersurfaces for both nonhomogeneous and homogeneous Hessian curvature equations, and by establishing a constant rank theorem, we prove that the resulting $k$-convex hypersurfaces are strictly convex.
△ Less
Submitted 16 July, 2026;
originally announced July 2026.
-
MemPoison: Uncovering Persistent Memory Threats and Structural Blind Spots in LLM Agents
Authors:
Jifeng Gao,
Kang Xia,
Yi Zhang,
Xiaobin Hong,
Mingkai Lin,
Xingshen Wei,
Wenzhong Li,
Sanglu Lu
Abstract:
Persistent external memory enhances agent continuity but introduces persistent security vulnerabilities: adversarial content can be injected via standard interaction channels, retained across turns, and later distort downstream behavior. To address this challenge, we propose MemPoison, a comprehensive benchmark and analysis framework featuring 1227 hand-validated cases across four attack types, th…
▽ More
Persistent external memory enhances agent continuity but introduces persistent security vulnerabilities: adversarial content can be injected via standard interaction channels, retained across turns, and later distort downstream behavior. To address this challenge, we propose MemPoison, a comprehensive benchmark and analysis framework featuring 1227 hand-validated cases across four attack types, three injection channels, and three representative memory substrates, evaluated on seven open-weight and three closed-weight model families. We introduce a three-tier taxonomy: (L1) direct single-record corruption, (L2) compositional multi-record corruption and (L3) context-triggered dormant corruption. Our evaluations reveal a distinct defense frontier: while baseline write-time defenses, such as consistency checks, substantially suppress direct L1 attacks, they fail to reliably suppress L2 and L3 attacks. Through mechanistic influence decomposition (MID), we demonstrate structural blind spots in write-time defenses, which admit seemingly benign records that later become harmful through joint retrieval composition or trigger-conditioned activation. Our findings advocate for shifting from static filtering to adaptive, context-sensitive memory defense strategies.
△ Less
Submitted 16 July, 2026;
originally announced July 2026.
-
Multi-Knob Switchable Chiral Superconductivity Quartet in Rhombohedral Graphene
Authors:
Zhenqi Hua,
Shenyong Ye,
Phatthanon Pattanakanvijit,
Gang Shi,
Tonghang Han,
Emily Aitken,
Jixiang Yang,
Junseok Seo,
Haoyang Liu,
Ran Hao,
Kaitai Xiao,
Jiaxing Guo,
Vo Tien Phong,
Kenji Watanabe,
Takashi Taniguchi,
Chunli Huang,
Cyprian Lewandowski,
Long Ju,
Peng Xiong,
Zhengguang Lu
Abstract:
Chiral superconductors break orbital time-reversal symmetry and may host topological quasiparticles with non-Abelian statistics. In rhombohedral graphene, superconductivity develops from a spin-valley-polarized quarter-metal (QM) parent state and features unique magnetic hysteresis of resistance that indicates orbital time-reversal-symmetry-breaking. Exploring and controlling the full spin-valley…
▽ More
Chiral superconductors break orbital time-reversal symmetry and may host topological quasiparticles with non-Abelian statistics. In rhombohedral graphene, superconductivity develops from a spin-valley-polarized quarter-metal (QM) parent state and features unique magnetic hysteresis of resistance that indicates orbital time-reversal-symmetry-breaking. Exploring and controlling the full spin-valley flavors of such superconductivity could enable novel superconducting and topological devices, but have remained unexplored. Here we report transport measurements on rhombohedral hexalayer graphene (R6G), which reveal a new superconducting state (SCH) that is induced by an out-of-plane magnetic field, in addition to chiral superconductivity (CSC) similar to those observed in thinner layers. This SCH state emerges above 0.8 T, persists up to 1.6 T and can be switched on/off by magnetic field $H_\perp$, carrier density $n$, and gate displacement field $D$. Quantum oscillations and anomalous Hall measurements show that SCH stems from a field-induced quarter-metal (QM$'$) parent phase, which carries orbital magnetization opposite to that of the zero-field QM. Across the full $(n, D, H_\perp)$ parameter space, superconductivity can be realized from all four spin-valley isospin flavors, establishing a switchable chiral-superconductor quartet in R6G. We interpret the parent-state switching as arising from competition between a Kane-Mele-like spin-valley splitting and magnetic-field coupling to spin-valley-dependent magnetic moments. Our work establishes rhombohedral graphene as a multi-knob platform for different isospin-polarized superconductivities, which enables programmable superconducting networks with possible Majorana modes along domain walls.
△ Less
Submitted 7 July, 2026;
originally announced July 2026.
-
ChainSWE: Benchmarking Coding Agents on Multi-Bug Software Maintenance
Authors:
Qirui Jin,
Lingching Tung,
Kenan Li,
Qiyang Shi,
Yushi She,
Huanzhong Jia,
Harrison Zhao,
Kejing Xia,
Zhenbang Du,
Yikai Zhang,
Jiaxin Pei,
Zhenyu Zhang,
Zhen Qi,
Yuyan Duan,
Wenke Lee,
Zijian Jin
Abstract:
Language model (LM) agents are increasingly deployed to maintain codebases over extended periods, fixing streams of related defects while carrying context from one fix to the next. Yet existing software engineering (SWE) benchmarks evaluate models one bug at a time: the repository is reset, the codebase is re-read, and a single self-contained issue is graded in isolation. This setting collapses a…
▽ More
Language model (LM) agents are increasingly deployed to maintain codebases over extended periods, fixing streams of related defects while carrying context from one fix to the next. Yet existing software engineering (SWE) benchmarks evaluate models one bug at a time: the repository is reset, the codebase is re-read, and a single self-contained issue is graded in isolation. This setting collapses a continuous maintenance workflow into a series of independent sessions, ignoring the cumulative dependencies that make real-world bug fixing challenging. To bridge this gap, we introduce ChainSWE, the first benchmark for evaluating agents on sequential, dependent bug fixes within a shared codebase. We collect chronological chains of 304 issues across 54 Python projects, mined from six SWE-bench-family datasets. Our evaluation across a range of agents and models reveals a consistent performance drop by up to 70% as the chain length increases.
△ Less
Submitted 1 July, 2026;
originally announced July 2026.
-
Communicability-Inspired Positional Encoding (CIPE)
Authors:
Yipeng Zhang,
Zhongtian Sun,
Pietro Liò,
Kelin Xia
Abstract:
Positional encodings (PEs) are essential for Transformers. Yet designing effective PEs for non-Euclidean graphs remains challenging. Such encodings should ideally induce an Attention-Compatible Geometry for self-attention: not merely describing graph structure, but defining a geometry whose inner products reflect meaningful structural relatedness. To realize this geometry, we propose Communicabili…
▽ More
Positional encodings (PEs) are essential for Transformers. Yet designing effective PEs for non-Euclidean graphs remains challenging. Such encodings should ideally induce an Attention-Compatible Geometry for self-attention: not merely describing graph structure, but defining a geometry whose inner products reflect meaningful structural relatedness. To realize this geometry, we propose Communicability-Inspired Positional Encoding (CIPE), built from communicability, a measure between pairs of nodes that aggregates contributions from paths of all lengths. By construction, CIPE inner products recover communicability, converting global multi-path connectivity into an attention-ready similarity geometry. For practical Transformer training, we introduce dimensionality alignment, mapping graph-size-dependent CIPE representations to prescribed dimensions while faithfully preserving the induced geometry. Empirically, CIPE improves structure-agnostic Transformers by 35.5% on average across seven benchmarks, outperforming representative PEs; it also consistently improves structure-biased graph Transformers, where competing PEs often yield only marginal benefits. These results position CIPE as a principled framework for attention-compatible graph positional encodings.
△ Less
Submitted 23 June, 2026;
originally announced June 2026.
-
Fractional Magnonic Frequency Combs
Authors:
Qian-Nan Huang,
Zhiping Xue,
Xudong Wang,
Yanmeng Lei,
Lihui Bai,
Ke Xia,
Gerrit E. W. Bauer,
Tao Yu
Abstract:
Magnonic frequency combs (MFCs) are spectacular phenomena in microwave-driven high-quality magnets. Like the equally spaced prongs in a comb, conventional \textit{integer} MFCs are sharp resonances with an equal and constant frequency difference. Here we report \textit{fractional} MFCs in a high-quality magnetic sphere that emerges when adding a low-power, precisely detuned microwave to the main d…
▽ More
Magnonic frequency combs (MFCs) are spectacular phenomena in microwave-driven high-quality magnets. Like the equally spaced prongs in a comb, conventional \textit{integer} MFCs are sharp resonances with an equal and constant frequency difference. Here we report \textit{fractional} MFCs in a high-quality magnetic sphere that emerges when adding a low-power, precisely detuned microwave to the main drive that compresses the frequency spacings to a rational fraction of the original comb, generating high-density spectral grids with hundreds of lines. The theoretical analysis finds that parametric three-magnon scattering is the dominant non-linear process that reproduces the observation well. This mechanism is unique to magnets: it does not exist in an optomechanical system, where the Kerr and optical nonlinearities govern comb formation at a much higher power input. Since our platform operates as a frequency ``vernier caliper" with much higher sensitivity than integer MFCs, it has application potential in precision metrology.
△ Less
Submitted 23 June, 2026;
originally announced June 2026.
-
PACT: Privileged Trace Co-Training for Multi-Turn Tool-Use Agents
Authors:
Zhenbang Du,
Jun Luo,
Zhiwei Zheng,
Xiangchi Yuan,
Kejing Xia,
Dachuan Shi,
Qirui Jin,
Qijia He,
Shaofeng Zou,
Yingbin Liang,
Wenke Lee
Abstract:
Multi-turn tool-use agents must reason, call tools, and adapt to observations across several interaction turns. Post-training such agents is challenging, as reinforcement learning often suffers from sparse rewards and weak credit assignment despite matching the prompt-only inference setting, while supervised fine-tuning on expert traces provides dense process supervision but can over-constrain the…
▽ More
Multi-turn tool-use agents must reason, call tools, and adapt to observations across several interaction turns. Post-training such agents is challenging, as reinforcement learning often suffers from sparse rewards and weak credit assignment despite matching the prompt-only inference setting, while supervised fine-tuning on expert traces provides dense process supervision but can over-constrain the model to fixed trajectories. To tackle this, we propose PACT, a Privileged trAce Co-Training framework for multi-turn tool-use agents. The key idea is to use expert traces only as training-time optimization signals rather than rollout-time hints. PACT keeps rollout generation prompt-only, then uses expert traces to guide optimization through two complementary signals: a trace-conditioned RL surrogate that evaluates prompt-only rollouts under expert-trace context, and a component-aware SFT loss that supervises reasoning prefixes and tool-calls with annealed strength. To reduce over-reliance on the training-only trace context, PACT further introduces a prompt-only anchoring. We also provide a latent-trace view that connects the two trace-based objectives and explains how expert traces can guide optimization without being used during rollout generation. Experiments on FTRL, BFCL, and ToolHop show that PACT consistently improves over strong SFT- and RL-based baselines, highlighting the value of privileged trace co-training for multi-turn tool-use learning.
△ Less
Submitted 15 June, 2026;
originally announced June 2026.
-
FuseChain: Runtime Evidence Reconstruction for Software Supply-Chain Attacks
Authors:
Zhuoran Tan,
Yutian Tang,
Jeremy Singer,
Christos Anagnostopoulos,
Ke Xiao
Abstract:
Software supply-chain (SSC) attacks are increasingly multi-stage, cross-source, and temporally distributed. A single attack campaign may leave weak and fragmented traces across multi-source telemetry that captures different granularities and perspectives of runtime behavior. Existing runtime detection systems often analyze these sources independently, making it difficult to identify low-frequency…
▽ More
Software supply-chain (SSC) attacks are increasingly multi-stage, cross-source, and temporally distributed. A single attack campaign may leave weak and fragmented traces across multi-source telemetry that captures different granularities and perspectives of runtime behavior. Existing runtime detection systems often analyze these sources independently, making it difficult to identify low-frequency attack evidence or reconstruct the temporal context in which it appears. We present FUSECHAIN, a runtime detection framework that represents multi-source software supply-chain telemetry as a temporal heterogeneous provenance graph over a unified event-time axis. By aligning package/runtime traces, process events, network telemetry, DNS/HTTP metadata, and security alerts on a unified temporal graph, FuseChain captures cross-source dependencies and sparse attack evidence that may be ambiguous within any individual source. It learns anomaly-centric temporal representations from benign-prefix telemetry and performs deployable attack-stage reconstruction through a lightweight decoder on top of a frozen anomaly backbone. Our experiments show that jointly optimizing anomaly detection and stage prediction is ineffective under sparse and imbalanced runtime supply-chain telemetry. Across seven SSC attack scenarios, FuseChain improves deployable stage reconstruction from 0.369 to 0.881 Stage Recall@500 with a frozen-backbone decoder, while adaptive retrieval further increases observable-stage recall from 0.524 to 0.655 without modifying the detector. These results highlight the deployable value of decoupling runtime SSC anomaly detection from downstream attack-stage interpretation.
△ Less
Submitted 14 June, 2026;
originally announced June 2026.
-
An experimental study on the heat transport in porous media convection
Authors:
Jing Dong,
Lu Zhang,
Ke-Qing Xia
Abstract:
We investigate the heat transport in porous media convection over a wide Rayleigh--Darcy number range of $26.8\leq Ra\leq 2.62\times 10^5$, and a Darcy number range of $6.18\times10^{-7}\leq Da\leq 1.21\times 10^{-5}$. In the experiments, we employ 3D-printed lattice structures as the solid porous matrix and water as the working fluid. Quantitative analyses of the porous medium Nusselt number…
▽ More
We investigate the heat transport in porous media convection over a wide Rayleigh--Darcy number range of $26.8\leq Ra\leq 2.62\times 10^5$, and a Darcy number range of $6.18\times10^{-7}\leq Da\leq 1.21\times 10^{-5}$. In the experiments, we employ 3D-printed lattice structures as the solid porous matrix and water as the working fluid. Quantitative analyses of the porous medium Nusselt number $Nu_m$ and local temperature statistics reveal that the present system undergoes a transition through five distinct regimes: I. Conduction, II. Convection, III. Oscillation, IV. Transition, V. Classical Rayleigh--Bénard convection. This transitional process bridges the gap between Rayleigh--Darcy-like behaviour and Rayleigh--Bénard-like behaviour in porous media convection. By varying the permeability of the matrix, we further examine the role of the Darcy number $Da$, which turns out to have a profound impact on the transitional processes across different regimes. Flow field measurements reveal that the flow structures within Regime IV and Regime V evolve from several horizontally stacked convection rolls to a single-roll structure, and the pore-scale Reynolds number both exceeds unity in these two regimes. Finally, we report the corresponding phase diagram in the $Ra$-$Da$ space.
△ Less
Submitted 4 June, 2026;
originally announced June 2026.
-
Multiple critical Froude numbers for the centrifugal effects on heat transport in rotating Rayleigh-Bénard convection
Authors:
Zhi-Cong Kang,
Guang-Yu Ding,
Lu Zhang,
Ke-Qing Xia
Abstract:
The influence of centrifugal effects in rotating Rayleigh-Benard convection is investigated using direct numerical simulations. We find that the Nusselt number decreases beyond a critical Froude number, Fr_c*. This critical value depends on both the Rayleigh number Ra and the aspect ratio Gamma, following power-law scalings with each parameter. We interpret Fr_c* as the onset of centrifugal effect…
▽ More
The influence of centrifugal effects in rotating Rayleigh-Benard convection is investigated using direct numerical simulations. We find that the Nusselt number decreases beyond a critical Froude number, Fr_c*. This critical value depends on both the Rayleigh number Ra and the aspect ratio Gamma, following power-law scalings with each parameter. We interpret Fr_c* as the onset of centrifugal effects within the thermal boundary layers. This interpretation is supported by the thickening of the boundary layers and a reduction in the planar heat flux. We compare Fr_c* with two previously proposed critical Froude numbers. The first, Fr_Hu, marks the onset of centrifugal effects in the bulk, as evidenced by changes in local heat flux and radial vortex motion. For Fr_Hu < Fr < Fr_c*, centrifugal effects primarily redistribute heat within the bulk and have little influence on the global heat transfer. The second, Fr_Horn, is based on a global force-balance argument. The similar dependence of Fr_c* and Fr_Horn on the aspect ratio suggests a close connection between the global force balance and the onset of centrifugal effects in the thermal boundary layers. These results demonstrate that centrifugal forcing influences the bulk flow and the thermal boundary layers differently in rotating Rayleigh-Benard convection. While relatively weak centrifugal forcing modifies the bulk dynamics, substantially stronger forcing is required to alter boundary-layer properties and global heat transport.
△ Less
Submitted 4 June, 2026;
originally announced June 2026.
-
UNIVID: Unified Vision-Language Model for Video Moderation
Authors:
Kejuan Yang,
Yizhuo Zhang,
Mingyuan Du,
Yue Zhang,
Dixin Zheng,
Kaili Zhao,
Yang Xiao,
Hanzhong Liang,
Kenan Xiao
Abstract:
Global-scale video moderation faces a dual challenge: the need for fine-grained multi-modal reasoning and the demand for interpretable outputs to support downstream enforcement. Traditional moderation systems often rely on fragmented black-box classifiers that are difficult to maintain and lack transparency. In this paper, we present UNIVID, a UNIfied VIsion-language model for video moDeration. Un…
▽ More
Global-scale video moderation faces a dual challenge: the need for fine-grained multi-modal reasoning and the demand for interpretable outputs to support downstream enforcement. Traditional moderation systems often rely on fragmented black-box classifiers that are difficult to maintain and lack transparency. In this paper, we present UNIVID, a UNIfied VIsion-language model for video moDeration. Unlike standard classification models, UNIVID generates policy-aware captions that serve as an interpretable intermediate representation, enabling human-verifiable decisions and multi-task reusability. While existing open-source and commercial VLMs often suffer from safety-guardrail refusals and lack fine-grained policy alignment, we develop a specialized training data recipe that combines expert human-refined labels with synthetic data to align the model with our safety guidelines. By integrating UNIVID as the core captioner, we design a novel end-to-end video moderation system that reduces violation leakage by 42.7% and overkill rate by 37.0% relatively. Meanwhile, by replacing over 1,000 policy-specific models with a single UNIVID backbone, we recycled extensive computation resources while reducing engineering maintenance overhead. To our knowledge, this is one of the first reports of a high-efficiency captioning VLM successfully supporting industrial-scale moderation and cross-functional business.
△ Less
Submitted 4 June, 2026;
originally announced June 2026.
-
Scale-dependent force balance governs transition to the geostrophic regime in liquid metal rotating convection
Authors:
Shao-Peng Yang,
Lin Sun,
Guang-Yu Ding,
Ke-Qing Xia,
Yi-Chao Xie
Abstract:
Rotating convection in low-Prandtl-number liquid metal drives dynamo action in the Earth's outer core and is central to planetary interior dynamics. It has been proposed that flow regime transitions in rotating convection are controlled by competition between the thermal and Ekman boundary layers. However, through laboratory experiments and direct numerical simulations of rotating liquid-metal con…
▽ More
Rotating convection in low-Prandtl-number liquid metal drives dynamo action in the Earth's outer core and is central to planetary interior dynamics. It has been proposed that flow regime transitions in rotating convection are controlled by competition between the thermal and Ekman boundary layers. However, through laboratory experiments and direct numerical simulations of rotating liquid-metal convection, we find that this mechanism breaks down in the low-Prandtl-number regime. Here we show that increasing rotation reorganises the bulk flow: the large-scale circulation is suppressed and replaced by smaller-scale structures, producing a characteristic horizontal length scale $\ell$. Transitions to the geostrophic regime are then governed by a buoyancy--Coriolis balance defined on $\ell$ rather than by the boundary-layer crossing. This scale-dependent mechanism also yields heat-transport scalings that depart from boundary-layer-based predictions in the geostrophic regime. Our results reveal a distinct route to the geostrophic regime in low-Prandtl-number rotating convection with implications for rotating liquid metal flows in planetary interiors.
△ Less
Submitted 3 June, 2026;
originally announced June 2026.
-
Scale-invariance and characteristic length scale for the large-scale vortices in geostrophic convective turbulence with friction
Authors:
Guang-Yu Ding,
Tian-Yi Pei,
Hang-Yu Zhu,
Ke-Qing Xia
Abstract:
In geostrophic convective turbulence, large-scale vortices (LSVs) emerge through upscale energy transfer and are commonly regulated by large-scale friction. Yet the role of friction in setting the LSV size remains poorly understood. Here we perform direct numerical simulations of rotating Rayleigh-Benard convection with a linear friction term $α\mathbf{u}$. Contrary to the classical prediction…
▽ More
In geostrophic convective turbulence, large-scale vortices (LSVs) emerge through upscale energy transfer and are commonly regulated by large-scale friction. Yet the role of friction in setting the LSV size remains poorly understood. Here we perform direct numerical simulations of rotating Rayleigh-Benard convection with a linear friction term $α\mathbf{u}$. Contrary to the classical prediction $L_α\simα^{-3/2}$ obtained from the Kraichnan-Leith-Batchelor (KLB) theory, we find that the LSV radius follows $R_{LSV}\simα^{-1/2}$. This discrepancy originates from the energy spectrum of the barotropic (2D) manifold, which exhibits $E_{2D}(k)\sim k^{-3}$ over the range of upscale energy transfer, rather than the canonical $k^{-5/3}$ scaling. To explain this behavior, we analyze the energy pathways of the barotropic manifold and show that the inverse transfer is strongly nonlocal, coupling a broad range of intermediate scales directly to the cutoff scale. We propose that this coupling leads to a balance between the local and large-scale shear strain rates, resulting in a scale-invariant coarse-grained vorticity. The resulting prediction $E_{2D}(k)\sim k^{-3}$ is supported by circulation statistics exhibiting $\langle|Γ(r)|\rangle\sim r^2$. The observed $k^{-3}$ spectrum naturally yields the scaling $R_{LSV}\simα^{-1/2}$. These results provide a physical interpretation for the widely observed $k^{-3}$ spectrum in condensation-dominated turbulence and suggest that LSV-size estimates based on the classical $k^{-5/3}$ spectrum may be significantly biased in geophysical and astrophysical flows.
△ Less
Submitted 1 June, 2026;
originally announced June 2026.
-
Symmetry Breaking and Restoration in Turbulent Thermal Convection Arises from the Competition Between Advection and Buoyancy
Authors:
Guang-Yu Ding,
Fang Xu,
Ke-Qing Xia
Abstract:
Spontaneous symmetry breaking (SSB) remains poorly understood in thermal convection, but hints may be found from its restoration. We hereby compare the two convection systems: experiments with polymer additives, and simulations with linear friction. We observe the restoration of similar symmetric flows in both these systems. Additionally, restoration coincides with enhanced, time-symmetric velocit…
▽ More
Spontaneous symmetry breaking (SSB) remains poorly understood in thermal convection, but hints may be found from its restoration. We hereby compare the two convection systems: experiments with polymer additives, and simulations with linear friction. We observe the restoration of similar symmetric flows in both these systems. Additionally, restoration coincides with enhanced, time-symmetric velocity-buoyancy correlation, and a sharp drop in the normalized buoyancy-response time. These results indicate buoyancy predominance: velocity is statistically slaved to buoyancy and preferentially remains vertical. The predominance of buoyancy provides a local orientation mechanism, which is necessary for restoring the symmetry of the system. Conversely, this orientation mechanism is lost locally in canonical convective flows, thus SSB naturally occurs in Rayleigh-Bénard convection. Our results suggest that the breaking and restoration of symmetry in thermal convection are both attributable to the competition between advection and buoyancy.
△ Less
Submitted 1 June, 2026;
originally announced June 2026.
-
Subcritical transition to turbulence in buoyancy-driven flows with multiple hysteresis loops under quasi-one-dimensional confinement
Authors:
Lu Zhang,
Ke-Qing Xia
Abstract:
We present both static and quasi-static direct numerical simulations of Rayleigh-Bénard convection in a quasi-one-dimensional domain, revealing for the first time a clear subcritical transition to turbulence in a buoyancy-driven flow. Within a narrow range of Rayleigh number (Ra), three coexisting flow states are identified: steady convection, oscillatory chaos, and intermittent turbulence. The tr…
▽ More
We present both static and quasi-static direct numerical simulations of Rayleigh-Bénard convection in a quasi-one-dimensional domain, revealing for the first time a clear subcritical transition to turbulence in a buoyancy-driven flow. Within a narrow range of Rayleigh number (Ra), three coexisting flow states are identified: steady convection, oscillatory chaos, and intermittent turbulence. The transitions between these states are accompanied by abrupt jumps in both the Nusselt number (Nu) and Reynolds number (Re), the key global transport quantities in buoyancy-driven flows. Additionally, they exhibit pronounced hysteresis, forming three distinct hysteresis loops in the Nu-Ra plane: normal, reverse, and anomalous loops. More importantly, we show that the steady convection state is linearly stable against infinitesimal perturbations but can transition to intermittent turbulence when subjected to finite-amplitude disturbances, which is a defining hallmark of subcriticality. Thus, contrary to the prevailing view that the transition from convection to turbulence is supercritical, our results demonstrate that buoyancy-driven turbulence can emerge via a subcritical route, paving the way for a unified framework that describes instability mechanisms in both buoyancy-driven and shear-driven flows.
△ Less
Submitted 29 May, 2026;
originally announced May 2026.
-
Nanoscopic Multiplexing Optical Data Storage via Chip Fabrication
Authors:
Junyu Guan,
Quanshen Shen,
Bowen Tong,
Hanzhi Wang,
Zeyu Gao,
Hanyu Zhang,
Jingyang Zhou,
Zihua Chai,
Dong Liu,
Ya Wang,
Kangwei Xia
Abstract:
The accelerating growth of global data generation demands data storage platforms that offer high capacity, long lifespan, and low energy consumption beyond the limits of electronic memory technologies. Optical storage provides an attractive alternative. However, its density is fundamentally constrained by the optical diffraction limit and the limited scalability from the point-by-point laser writi…
▽ More
The accelerating growth of global data generation demands data storage platforms that offer high capacity, long lifespan, and low energy consumption beyond the limits of electronic memory technologies. Optical storage provides an attractive alternative. However, its density is fundamentally constrained by the optical diffraction limit and the limited scalability from the point-by-point laser writing, as well as thermal accumulation during high-speed writing. Here, we introduce a large-scale optical data storage scheme that is compatible with the progress in chip fabrication by combining electron-beam lithography (EBL) and ion implantation to deterministically encode high-density data. The approach achieves precise control of ion number and spatial distribution, enabling multi-bit grayscale encoding and wavelength division multiplexing with chip-scale patterning over millimeter areas. Wavelength-selective readout is performed using downconversion and upconversion fluorescence detection, allowing crosstalk-free retrieval of multiplexed data channels. We further develop a neural network-based super-resolution algorithm that reconstructs data beyond the diffraction limit, further increasing the effective storage density. Using this integrated framework, we achieve an optical data density of 10 Gbit/cm$^2$ with high fidelity. Our results establish a micro/nano-fabrication-compatible route to large-scale, high-density optical memory and provide a foundation for next-generation cold data optical storage technologies.
△ Less
Submitted 28 May, 2026;
originally announced May 2026.
-
Periodic Topological Deep Learning for Polymer Design and Discovery
Authors:
Yasharth Yadav,
Tze Kwang Gerald Er,
Atsushi Goto,
Kelin Xia
Abstract:
Polymers underpin applications across energy, healthcare, and materials science, yet their vast chemical space makes systematic discovery challenging. Most machine learning approaches represent polymers as molecular graphs of a single repeating unit, thereby missing both the periodicity of polymer chains and many-body interactions beyond pairwise bonds. We introduce Periodic-TDL, a deep learning f…
▽ More
Polymers underpin applications across energy, healthcare, and materials science, yet their vast chemical space makes systematic discovery challenging. Most machine learning approaches represent polymers as molecular graphs of a single repeating unit, thereby missing both the periodicity of polymer chains and many-body interactions beyond pairwise bonds. We introduce Periodic-TDL, a deep learning framework built on periodic Vietoris-Rips complexes that capture many-body interactions across multiple spatial scales, followed by a hierarchical simplicial message-passing (HSMP) encoder that propagates information from long-range interactions to covalent bonds, yielding representations enriched by higher-order topological features. Periodic-TDL outperforms all state-of-the-art models across polymer property prediction tasks spanning electronic, optical, physical, and thermal targets. Furthermore, we quantitatively validate how ester-to-amide substitution and $α$-methylation enhance thermal stability. Using a computationally synthesized dataset of 48,208 structures-generated via systematic substitution of acrylate and acrylamide polymers-we observed a mean $T_g$ increase of $\sim 55^\circ$C for ester-to-amide substitutions and $\sim 14^\circ$C for backbone $α$-methylation across matched polymer pairs. To verify these predicted trends, we use our Periodic-TDL model to analyze six novel polymer pairs from independent experimental measurements, including three newly synthesized polymers previously unreported in the literature. The experimental data successfully confirmed the model's predictions. Ultimately, these findings demonstrate that Periodic-TDL captures the underlying physical effects of specific functional group modifications, rather than merely optimizing predictive performance on benchmark datasets.
△ Less
Submitted 16 August, 2026; v1 submitted 26 May, 2026;
originally announced May 2026.
-
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
Authors:
Aili Chen,
Aonian Li,
Baichuan Zhou,
Bangwei Gong,
Binyang Jiang,
Boji Dan,
Changhao Zhang,
Changqing Yu,
Chao Wang,
Cheng Ma,
Cheng Zhong,
Cheng Zhu,
Chengjun Xiao,
Chengyi Yang,
Chengyu Du,
Chenyang Zhang,
Chi Zhang,
Chuangyi Huang,
Chunhao Zhang,
Chunhui Du,
Chunyu Zhao,
Congchao Guo,
Da Chen,
Deming Ding,
Dianjun Sun
, et al. (193 additional authors not shown)
Abstract:
We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The flagship M2 contains 229.9B total parameters with only 9.8B activated per token. Designed end-to-end for agentic deployment, the M2 series rests on three components: (i) agent-driven data pipelines producing large-scale…
▽ More
We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The flagship M2 contains 229.9B total parameters with only 9.8B activated per token. Designed end-to-end for agentic deployment, the M2 series rests on three components: (i) agent-driven data pipelines producing large-scale, verifiable trajectories across agentic coding and agentic cowork, each grounded in an executable workspace and an artifact-aligned reward; (ii) Forge, a scalable agent-native RL system that adapts to long-horizon agent trajectories, paired with windowed-FIFO scheduling, prefix-tree merging, inference optimization, and a clean training-inference-agent decoupling that supports both white-box and black-box agents; (iii) the latest M2.7 checkpoint takes an early step toward self-evolution -- autonomously debugging training runs and modifying its own scaffold. Across M2 through M2.7, this combination translates a mini-activation footprint into frontier-tier performance on agentic coding, deep search, office-task, and reasoning benchmarks.
△ Less
Submitted 30 July, 2026; v1 submitted 25 May, 2026;
originally announced May 2026.
-
HyLoVQA: Dynamic Hypernetwork-Generated Low-Rank Adaptation for Continual Visual Question Answering
Authors:
Yiran Wang,
Chenyi Xiong,
Ziyue Qin,
Miao Zhang,
Kui Xiao,
Zhifei Li
Abstract:
Continual Visual Question Answering (VQA) requires learning from non-stationary streams of visual inputs and questions while preserving past knowledge. Most prior methods adapt by updating a largely shared parameter set. This often leads to cross-level task interference, hindering accurate adaptation to the current task and object. To address this limitation, we propose HyLoVQA. It maintains a dri…
▽ More
Continual Visual Question Answering (VQA) requires learning from non-stationary streams of visual inputs and questions while preserving past knowledge. Most prior methods adapt by updating a largely shared parameter set. This often leads to cross-level task interference, hindering accurate adaptation to the current task and object. To address this limitation, we propose HyLoVQA. It maintains a drift-resilient memory bank of anchors. The bank stores the content of visual objects and textual tasks, and they are updated using current input features. Conditioned on retrieved anchors, a hypernetwork generates lightweight Low-Rank Adaptation (LoRA) adapters. This ensures parameter efficiency, allowing the model to adapt to each task and object dynamically. Additionally, we formulate an alignment loss that aligns semantic discrepancies in the feature space with functional changes in the parameter space, thereby constraining LoRA adapters to remain focused on the current task and object. Extensive experiments on VQA v2 and NExT-QA under both standard and compositional settings demonstrate the superiority of HyLoVQA over prior state-of-the-art methods.
△ Less
Submitted 21 May, 2026;
originally announced May 2026.
-
High-performance linear-scaling electronic structure method via chromatic superposition states
Authors:
Zhikang Jiang,
Zhizhi Xiao,
Mingfa Tang,
Weiyu Li,
Zhaoru Sun,
Ke Xia,
Youqi Ke
Abstract:
We introduce a high-performance linear-scaling electronic structure method that employs chromatic superposition states (CSS) as a low-dimensional, high-fidelity representation, which can be orders of magnitude smaller than the full Hilbert space. Grounded in the system's finite correlation length, the CSS representation aggregates the uncorrelated orbitals into a single basis via a graph-coloring…
▽ More
We introduce a high-performance linear-scaling electronic structure method that employs chromatic superposition states (CSS) as a low-dimensional, high-fidelity representation, which can be orders of magnitude smaller than the full Hilbert space. Grounded in the system's finite correlation length, the CSS representation aggregates the uncorrelated orbitals into a single basis via a graph-coloring scheme, and is independent of the system size yet accurately preserves all sparse operators in solving the Kohn-Sham equations. The projection onto CSSs is efficiently computed by employing the block-Lanczos Krylov method which features high hardware efficiency and linear-scaling cost, enabling fast calculation of large-scale Kohn-Sham density matrix. We show that this method already outperforms previous linear-scaling density matrix purification method by more than one order of magnitude in computational speed at even small scale, while preserving high accuracy. The practical utility of the CSS method is demonstrated through molecular dynamics simulation of a 10000 $H_2O$, and self-consistent calculation of a 1-million $H_2O$ with modest resources.
△ Less
Submitted 20 May, 2026;
originally announced May 2026.
-
Enhancing Graph-Based SLAM in GNSS-Denied environments by leveraging leg odometry
Authors:
Léon Perruchot-Triboulet,
Luc Jaulin,
Kai Xiao
Abstract:
Autonomous navigation in GNSS-denied environments remains a core challenge for legged robots, where exteroceptive sensors such as LiDAR are prone to elevation drift in geometrically sparse or repetitive scenes. We present a factor graph architecture that augments the LIO-SAM framework with a parallel kinematic lane driven by proprioceptive leg odometry, coupled to the main LiDAR-inertial lane via…
▽ More
Autonomous navigation in GNSS-denied environments remains a core challenge for legged robots, where exteroceptive sensors such as LiDAR are prone to elevation drift in geometrically sparse or repetitive scenes. We present a factor graph architecture that augments the LIO-SAM framework with a parallel kinematic lane driven by proprioceptive leg odometry, coupled to the main LiDAR-inertial lane via an identity relative pose constraint with a selective noise model. Applied to a Linxai D50 quadruped platform across two outdoor loops totaling over one kilometer, our approach reduces elevation drift from over 30m to under 30cm and enables convergence in a scene where the baseline pipeline fails entirely. These results suggest that proprioceptive data, already computed onboard for gait control, constitutes a lightweight and effective vertical anchor for SLAM in GNSS-denied settings.
△ Less
Submitted 19 May, 2026;
originally announced May 2026.
-
CopT: Contrastive On-Policy Thinking with Continuous Spaces for General and Agentic Reasoning
Authors:
Dachuan Shi,
Hanlin Zhu,
Xiangchi Yuan,
Wanjia Zhao,
Kejing Xia,
Wen Xiao,
Wenke Lee
Abstract:
Chain-of-thought (CoT) is a standard approach for eliciting reasoning capabilities from large language models (LLMs). However, the common CoT paradigm treats thinking as a prerequisite for answering, which can delay access to plausible answers and incur unnecessary token costs even when the model is able to identify an answer before extended thinking, a behavior known as performative reasoning. In…
▽ More
Chain-of-thought (CoT) is a standard approach for eliciting reasoning capabilities from large language models (LLMs). However, the common CoT paradigm treats thinking as a prerequisite for answering, which can delay access to plausible answers and incur unnecessary token costs even when the model is able to identify an answer before extended thinking, a behavior known as performative reasoning. In this paper, we introduce CopT, a reformulated reasoning pipeline that reverses the usual order of thinking and answering. Instead of thinking before answering, CopT first elicits a draft answer and then invokes subsequent on-policy thinking conditioned on its own draft answer for reflection and correction. To assess whether the draft answer should be trusted, CopT recasts continuous embeddings as inference-time contrastive verifiers. Specifically, it contrasts the model's support for the same generated tokens under discrete-token inputs and continuous-embedding inputs, yielding a sequence-level reverse KL estimator for answer reliability. Our analysis shows that under certain assumptions, the expected estimate equals the mutual information between the unresolved latent state and the emitted answer token, explaining why it captures answer-relevant uncertainty rather than arbitrary uncertainty in the latent state. When the answer is deemed insufficiently reliable, CopT performs further on-policy thinking, where a second KL estimator dynamically controls draft-answer visibility, preserving useful partial information while reducing the risk of being misled by unreliable content. Across mathematics, coding, and agentic reasoning tasks, CopT improves peak accuracy by up to 23% and reduces token usage by up to 57% at comparable or higher accuracy, without any additional training. The code is available at https://github.com/sdc17/CopT.
△ Less
Submitted 19 May, 2026;
originally announced May 2026.
-
The Stellar Abundances and Galactic Evolution Survey (SAGES). V. The First Data Release of the DDO51 Band
Authors:
Qiqian Zhang,
Zhou Fan,
Gang Zhao,
Kai Xiao,
Wei Wang,
Hongrui Gu,
Jie Zheng,
Jingkun Zhao,
Chun Li,
Yuqin Chen,
Haibo Yuan,
Haining Li,
Kefeng Tan,
Yihan Song,
Ali Luo,
Nan Song,
Yujuan Liu,
Yaqian Wu,
Ali Esamdin,
Hubiao Niu,
Jinzhong Liu,
Guojie Feng,
Yu Zhang
Abstract:
We present the first public data release of DDO51 band from the Stellar Abundances and Galactic Evolution Survey (SAGES), based on Nanshan One-meter Wide-field Telescope (NOWT) observations obtained between 2023 September and 2024 January. This release initiates the DDO51-band component of the survey, covering $\sim$ 2,500 deg$^2$ of the northern sky and including more than 10 million sources. The…
▽ More
We present the first public data release of DDO51 band from the Stellar Abundances and Galactic Evolution Survey (SAGES), based on Nanshan One-meter Wide-field Telescope (NOWT) observations obtained between 2023 September and 2024 January. This release initiates the DDO51-band component of the survey, covering $\sim$ 2,500 deg$^2$ of the northern sky and including more than 10 million sources. The DDO51 filter is centered near the \ion{Mg}{1}~$b$ triplet and the adjacent MgH feature, offering sensitivity to stellar surface gravity. The data reduction pipeline incorporates an improved astrometric solution anchored to Gaia DR3 and a photometric calibration strategy tied to synthetic photometry from Gaia XP spectra. These procedures yield a point-source depth of $\sim$18.9 mag at S/N$\sim$10 and an internal photometric precision $\approx$6-7 mmag at the bright end. A preliminary color--color analysis using Gaia broadband photometry confirms the expected sensitivity of the DDO51 band to stellar surface gravity, demonstrating a clear photometric separation between dwarf and giant sequences for late-type stars. This dataset, when combined with existing SAGES photometry in other bands, provides a crucial tool for disentangling the substructures of the Milky Way. All data products from this release upon publication will be available.
△ Less
Submitted 9 May, 2026;
originally announced May 2026.
-
Membership Inference Attacks on Vision-Language-Action Models
Authors:
Yuefeng Peng,
Mingzhe Li,
Kejing Xia,
Renhao Zhang,
Amir Houmansadr
Abstract:
Membership inference attacks (MIAs) have been extensively studied in large language models (LLMs) and vision-language models (VLMs), yet their implications for vision-language-action (VLA) models remain largely unexplored. VLA models differ from standard LLMs and VLMs in several important ways: they are often fine-tuned for many epochs on relatively small embodied datasets, operate over constraine…
▽ More
Membership inference attacks (MIAs) have been extensively studied in large language models (LLMs) and vision-language models (VLMs), yet their implications for vision-language-action (VLA) models remain largely unexplored. VLA models differ from standard LLMs and VLMs in several important ways: they are often fine-tuned for many epochs on relatively small embodied datasets, operate over constrained and structured action spaces, and expose action outputs that can be observed as executable behaviors and temporally correlated trajectories. These characteristics suggest a distinct and potentially more informative attack surface for membership inference. In this work, we present the first systematic study of MIAs against VLA systems. We formalize two membership inference settings for VLA models: sample-level inference over individual transition samples and trajectory-level inference over complete embodied demonstrations. We further develop a suite of attack methods under multiple access regimes, including strict black-box access. Our attacks exploit both classic MIA signals, such as token likelihood, and VLA-specific signals, such as observable action errors and temporal motion patterns. Across multiple VLA benchmarks and representative VLA models, these attacks achieve strong inference performance, showing that VLA models are highly vulnerable to membership inference. Notably, black-box attacks based only on generated actions achieve strong performance, highlighting a practical privacy risk for deployed embodied AI systems. Our findings reveal a previously underexplored privacy risk in robotic and embodied AI, and underscore the need for dedicated privacy evaluation and defenses for VLA models.
△ Less
Submitted 7 May, 2026;
originally announced May 2026.
-
Full-Spectrum Graph Neural Networks: Expressive and Scalable
Authors:
Xiaohan Wang,
Deyu Bo,
Longlong Li,
Kelin Xia
Abstract:
It is well established that spectral graph neural networks (GNNs) can universally approximate node signals; however, their expressive power remains bounded by the 1-dimensional Weisfeiler-Lehman test, which is mirrored in their lack of universality for higher-order signals. To go beyond this bound, we propose the Full-Spectrum GNNs (FSpecGNNs), a second-order generalization of classical spectral G…
▽ More
It is well established that spectral graph neural networks (GNNs) can universally approximate node signals; however, their expressive power remains bounded by the 1-dimensional Weisfeiler-Lehman test, which is mirrored in their lack of universality for higher-order signals. To go beyond this bound, we propose the Full-Spectrum GNNs (FSpecGNNs), a second-order generalization of classical spectral GNNs. FSpecGNN advances spectral filtering from two perspectives: (1) it lifts signals from the node domain to the node-pair domain; and (2) it extends the univariate spectral filter over eigenvalues to a bivariate filter over eigenvalue pairs. We show that classical spectral GNNs arise as a diagonal special case of FSpecGNNs, and prove that FSpecGNNs can be at most as expressive as Local 2-GNN while universally approximating node-pair signals, the latter being particularly beneficial for heterophilic graph learning. Moreover, FSpecGNN admits scalable implementations that avoid explicit node-pair-level computations; combined with a low-rank approximation that reduces full-spectrum convolution to a combination of polynomial spectral filters, it enables learning on large graphs. Empirically, FSpecGNN validates the predicted expressivity and delivers strong performance on heterophilic benchmarks.
△ Less
Submitted 24 May, 2026; v1 submitted 7 May, 2026;
originally announced May 2026.
-
ParkingScenes: A Structured Dataset for End-to-End Autonomous Parking in Simulation Scenes
Authors:
Haonan Chen,
Kaiwen Xiao,
Bin Tian,
Jun Fu
Abstract:
Autonomous parking remains a critical yet challenging task in intelligent driving systems, particularly within constrained urban environments where maneuvering space is limited and precise control is essential. While recent advances in end-to-end learning have shown great promise, the lack of high-quality, structured datasets tailored for parking scenarios remains a significant bottleneck.To addre…
▽ More
Autonomous parking remains a critical yet challenging task in intelligent driving systems, particularly within constrained urban environments where maneuvering space is limited and precise control is essential. While recent advances in end-to-end learning have shown great promise, the lack of high-quality, structured datasets tailored for parking scenarios remains a significant bottleneck.To address this gap, we present ParkingScenes, a comprehensive multimodal dataset specifically designed for end-to-end autonomous parking in simulated scenes. Built on the CARLA simulator, ParkingScenes features structured parking trajectories generated by a Hybrid A* planner and a Model Predictive Controller (MPC), providing accurate and reproducible supervision signals. The dataset includes 16 reverse-in and 6 parallel parking scenarios, each executed under two pedestrian conditions (present and absent), resulting in 704 structured episodes and approximately 105000 frames. Each scenario is repeated 16 times to ensure consistent coverage. Each frame contains synchronized data from four RGB cameras, four depth sensors, vehicle motion states, and Bird's-Eye View (BEV) representations, enabling rich multimodal fusion and context-aware learning. To demonstrate the utility of our dataset, we compare models trained on ParkingScenes with those trained on unstructured, manually collected simulation data under identical conditions. Results show significant improvements in performance, underscoring the effectiveness of structured supervision for robust and accurate parking policy learning. By releasing both the dataset and the collection framework, ParkingScenes establishes a scalable and reproducible benchmark for advancing learning-based autonomous parking systems. The dataset and collection framework will be released at: https://github.com/haonan-ai/ParkingScenes
△ Less
Submitted 20 April, 2026;
originally announced April 2026.
-
Filter Design for Estimating the Stellar Metallicity of Metal-poor Stars from Gaia XP Spectra
Authors:
Ruifeng Shi,
Yang Huang,
Kai Xiao,
Chuanjie Zheng,
Bowen Zhang,
Hongrui Gu,
Xinyi Li,
Huiling Chen
Abstract:
The estimation of stellar atmospheric parameters for large-scale samples, particularly metal-poor stars, is a cornerstone of Galactic archaeology. In this work, we optimized a photometric filter design tailored to measuring stellar metallicities for very metal-poor stars with [Fe/H]$< -1$.The optimal configurations consist of a central wavelength $λ_{\rm c}$ = 3960 Angstrom with a bandwidth $Δλ$ =…
▽ More
The estimation of stellar atmospheric parameters for large-scale samples, particularly metal-poor stars, is a cornerstone of Galactic archaeology. In this work, we optimized a photometric filter design tailored to measuring stellar metallicities for very metal-poor stars with [Fe/H]$< -1$.The optimal configurations consist of a central wavelength $λ_{\rm c}$ = 3960 Angstrom with a bandwidth $Δλ$ = 80 Angstrom for giant stars, and $λ_{\rm c} $= 3920 Angstrom with $Δλ$ = 80 Angstrom for dwarf stars. By applying these optimized filters to synthetic photometry derived from Gaia XP spectra, we inferred metallicities for both populations. Both internal and external validations demonstrate high precision across a wide metallicity range: 0.18-0.19 dex for $-2 \le \rm [Fe/H] \le -1$, 0.23-0.33 dex for $-3 \le \rm [Fe/H] \le -2$, and approximately 0.39 dex for the most metal-poor regime, successfully extending down to $\rm [Fe/H] \approx -4$ for giant stars, $\rm [Fe/H] \approx -3.3$ for dwarf stars. Finally, we present a catalog of approximately 14.5 million metal-poor stars with robust $\rm [Fe/H]$ measurements, along with more than ten thousand red giant ultra metal-poor candidates with $\rm [Fe/H] < -4.0$, providing a valuable resource for exploring the early formation and chemical evolution of the Milky Way.
△ Less
Submitted 23 April, 2026;
originally announced April 2026.
-
Sheaf Neural Networks on SPD Manifolds: Second-Order Geometric Representation Learning
Authors:
Yuhan Peng,
Junwen Dong,
Yuzhi Zeng,
Hao Li,
Ce Ju,
Huitao Feng,
Diaaeldin Taha,
Anna Wienhard,
Kelin Xia
Abstract:
Graph neural networks face two fundamental challenges rooted in the linear structure of Euclidean vector spaces: (1) Current architectures represent geometry through vectors (directions, gradients), yet many tasks require matrix-valued representations that capture relationships between directions-such as how atomic orientations covary in a molecule. These second-order representations are naturally…
▽ More
Graph neural networks face two fundamental challenges rooted in the linear structure of Euclidean vector spaces: (1) Current architectures represent geometry through vectors (directions, gradients), yet many tasks require matrix-valued representations that capture relationships between directions-such as how atomic orientations covary in a molecule. These second-order representations are naturally captured by points on the symmetric positive definite matrices (SPD) manifold; (2) Standard message passing applies shared transformations across edges. Sheaf neural networks address this via edge-specific transformations, but existing formulations remain confined to vector spaces and therefore cannot propagate matrix-valued features. We address both challenges by developing the first sheaf neural network operates natively on the SPD manifold. Our key insight is that the SPD manifold admits a Lie group structure, enabling well-posed analogs of sheaf operators without projecting to Euclidean space. Theoretically, we prove that SPD-valued sheaves are strictly more expressive than Euclidean sheaves: they admit consistent configurations (global sections) that vector-valued sheaves cannot represent, directly translating to richer learned representations. Empirically, our sheaf convolution transforms effectively rank-1 directional inputs into full-rank matrices encoding local geometric structure. Our dual-stream architecture achieves SOTA on 6/7 MoleculeNet benchmarks, with the sheaf framework providing consistent depth robustness.
△ Less
Submitted 31 May, 2026; v1 submitted 22 April, 2026;
originally announced April 2026.
-
$R^2$-dLLM: Accelerating Diffusion Large Language Models via Spatio-Temporal Redundancy Reduction
Authors:
Zhenbang Du,
Kejing Xia,
Xinrui Zhong,
Yonggan Fu,
Nicolai Oswald,
Binfei Ji,
Brucek Khailany,
Pavlo Molchanov,
Yingyan Lin
Abstract:
Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to autoregressive generation by enabling parallel token prediction. However, practical dLLM decoding still suffers from high inference latency, which limits deployment. In this work, we observe that a substantial part of this inefficiency comes from recurring redundancy in the decoding process, including spatial redund…
▽ More
Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to autoregressive generation by enabling parallel token prediction. However, practical dLLM decoding still suffers from high inference latency, which limits deployment. In this work, we observe that a substantial part of this inefficiency comes from recurring redundancy in the decoding process, including spatial redundancy caused by confidence clusters and positional ambiguity, and temporal redundancy caused by repeatedly remasking predictions that have already stabilized. Motivated by these patterns, we propose $R^{2}$-dLLM, a unified framework for reducing decoding redundancy from both inference and training perspectives. At inference time, we introduce training-free decoding rules that aggregate local confidence and token predictions, and finalize temporally stable tokens to avoid redundant decoding steps. We further propose a redundancy-aware supervised fine-tuning pipeline that aligns the model with efficient decoding trajectories and reduces reliance on manually tuned thresholds. Experiments demonstrate that $R^{2}$-dLLM consistently reduces the number of decoding steps by up to 88\% compared to existing decoding strategies, while maintaining competitive generation quality across different models and tasks. These results validate that decoding redundancy is a central bottleneck in dLLMs, and that explicitly reducing it yields substantial practical efficiency gains. Our code and models are available at https://github.com/GATECH-EIC/R2-dLLM.
△ Less
Submitted 1 June, 2026; v1 submitted 20 April, 2026;
originally announced April 2026.
-
Nonmonotonic Scaling of the Anomalous Hall Effect in a Bicollinear Antiferromagnet
Authors:
Ruifeng Wang,
Chi Fang,
Ilya Kostanovski,
Ke Xiao,
Felix Küster,
Jenny Davern,
Naoto Nagaosa,
Stuart S. P. Parkin
Abstract:
An anomalous Hall effect (AHE) in antiferromagnetic (AF) systems with no net magnetization is of considerable interest for both fundamental physics and spintronic applications. Of particular interest is the two-dimensional van der Waals antiferromagnet FeTe that has an unusual fully magnetically compensated bicollinear AF structure and exhibits pronounced Kondo interaction leading to strong band r…
▽ More
An anomalous Hall effect (AHE) in antiferromagnetic (AF) systems with no net magnetization is of considerable interest for both fundamental physics and spintronic applications. Of particular interest is the two-dimensional van der Waals antiferromagnet FeTe that has an unusual fully magnetically compensated bicollinear AF structure and exhibits pronounced Kondo interaction leading to strong band renormalization. Here, we investigate the AHE in epitaxial FeTe thin films grown by molecular beam epitaxy. A large anomalous Hall conductivity is exhibited below the Neel temperature (T_N ~ 60 K) and, strikingly, becomes nonlinear at high fields within a narrow temperature window around 49 K, deviating from conventional AHE scaling behavior versus its longitudinal conductivity. Linear fits reveal a pronounced negative peak in the intercept, accompanied by a field-induced canted magnetic moment. The AHE responses are related to the Berry curvature derived from FeTe's topological band structure, highlighting the intricate interplay between topology, magnetism, and electronic transport.
△ Less
Submitted 14 April, 2026;
originally announced April 2026.
-
Investigating the intrinsic anomalous Hall effect in MnPt3 topological semimetal
Authors:
Jing Meng,
Hongru Wang,
Kun Zheng,
Yuhao Wang,
Zheng Li,
Bocheng Yu,
Haoyu Lin,
Keqi Xia,
Jingzhong Luo,
Zengyao Wang,
Xiaoyan Zhu,
Baiqing Lv,
Yaobo Huang,
Jie Ma,
Yang Xu,
Shijing Gong,
Tian Shang,
Qingfeng Zhan
Abstract:
The cubic Cu$_3$Au-type $X$Pt$_3$ family ($X$ = V, Cr, and Mn) is a topological semimetal characterized by anti-crossing gapped nodal lines near the Fermi level, which give rise to significant Berry curvatures and thus to the anomalous Hall effect (AHE). Among the three members, CrPt$_3$ has been experimentally verified to exhibit a large anomalous Hall conductivity (AHC), while its counterparts M…
▽ More
The cubic Cu$_3$Au-type $X$Pt$_3$ family ($X$ = V, Cr, and Mn) is a topological semimetal characterized by anti-crossing gapped nodal lines near the Fermi level, which give rise to significant Berry curvatures and thus to the anomalous Hall effect (AHE). Among the three members, CrPt$_3$ has been experimentally verified to exhibit a large anomalous Hall conductivity (AHC), while its counterparts MnPt$_3$ and VPt$_3$ remain largely unexplored. Here, a series of MnPt$_3$ thin films with varying thicknesses (20--70 nm) was epitaxially grown on the MgO substrates using magnetron sputtering and was systematically investigated by magnetization, electrical resistivity, and Hall resistivity measurements. MnPt$_3$ films undergo a ferromagnetic transition at a Curie temperature $T_\mathrm{C}$, which increases as the film thickness increases, reaching $\sim$ 344 K for the 70-nm-thick film. All the anomalous Hall transport properties of MnPt$_3$ films, including the resistivity, conductivity, and angle, exhibit a strong correlation with their magnetic properties. The scaling analysis suggests that the intrinsic Berry-curvature mechanism dominates the observed AHE, while the extrinsic contributions are much smaller. The intrinsic AHC increases as the film thickness increases, while the extrinsic AHC is thickness-independent. Such an enhanced intrinsic AHC in the MnPt$_3$ films is most likely attributed to the strain effect, implying that it serves as an effective method to tune the electronic band topology in the $X$Pt$_3$ topological semimetal.
△ Less
Submitted 7 April, 2026;
originally announced April 2026.
-
An End-to-End Approach for Fixing Concurrency Bugs via SHB-Based Context Extractor
Authors:
Zhuang Li,
Qiuping Yi,
Keyang Xiao,
Zongcheng Ji,
Hongliang Liang
Abstract:
With the rise of multi-core processors and distributed systems, concurrent programming has become essential yet challenging, primarily due to the non-deterministic nature of thread execution. Manually addressing concurrency bugs is time-consuming and error-prone. Automated Program Repair techniques provide a promising solution. However, developing an end-to-end concurrency bug repair tool is parti…
▽ More
With the rise of multi-core processors and distributed systems, concurrent programming has become essential yet challenging, primarily due to the non-deterministic nature of thread execution. Manually addressing concurrency bugs is time-consuming and error-prone. Automated Program Repair techniques provide a promising solution. However, developing an end-to-end concurrency bug repair tool is particularly challenging. Most existing tools rely on the assumption that bug-related information is readily available or that concurrency bug contexts are ideally extracted, which is often impractical in real-world scenarios. This paper introduces ConFixAgent, an LLM-driven agent capable of fixing various types of concurrency bugs in an end-to-end manner, eliminating the need for any prior bug-related information. Specifically, we propose a novel context extraction approach designed for concurrency bug repair, utilizing Static Happens-Before Graphs to identify bug-relevant sections.We implemented ConFixAgent and evaluated it across multiple benchmark sets. Our extensive experiments demonstrate that ConFixAgent significantly outperforms state-of-the-art tools in addressing diverse types of concurrency bugs, with its context extraction method markedly enhancing the accuracy of LLM-generated repair solutions.
△ Less
Submitted 7 April, 2026;
originally announced April 2026.
-
GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding
Authors:
Rong Fan,
Kaiyan Xiao,
Minghao Zhu,
Liuyi Wang,
Kai Dai,
Zhao Yang
Abstract:
Video temporal grounding (VTG) is a critical task in video understanding and a key capability for extending video large language models (Vid-LLMs) to broader applications. However, existing Vid-LLMs rely on uniform frame sampling to extract video information, resulting in a sparse distribution of key frames and the loss of crucial temporal cues. To address this limitation, we propose Grounded Visu…
▽ More
Video temporal grounding (VTG) is a critical task in video understanding and a key capability for extending video large language models (Vid-LLMs) to broader applications. However, existing Vid-LLMs rely on uniform frame sampling to extract video information, resulting in a sparse distribution of key frames and the loss of crucial temporal cues. To address this limitation, we propose Grounded Visual Token Sampling (GroundVTS), a Vid-LLM architecture that focuses on the most informative temporal segments. GroundVTS employs a fine-grained, query-guided mechanism to filter visual tokens before feeding them into the LLM, thereby preserving essential spatio-temporal information and maintaining temporal coherence. Futhermore, we introduce a progressive optimization strategy that enables the LLM to effectively adapt to the non-uniform distribution of visual features, enhancing its ability to model temporal dependencies and achieve precise video localization. We comprehensively evaluate GroundVTS on three standard VTG benchmarks, where it outperforms existing methods, achieving a 7.7-point improvement in mIoU for moment retrieval and 12.0-point improvement in mAP for highlight detection. Code is available at https://github.com/Florence365/GroundVTS.
△ Less
Submitted 2 April, 2026;
originally announced April 2026.
-
Robust Flat Magnetoresistivity in D0$_3$-Fe$_3$Ga Driven by Chiral Anomaly
Authors:
Ruoqi Wang,
Xinyang Li,
Bo Zhao,
Haofu Wen,
Xin Gu,
Shijun Yuan,
Langsheng Ling,
Chuanying Xi,
Ze Wang,
Kunquan Hong,
Liang Ma,
Ke Xia,
Taishi Chen,
Jinlan Wang
Abstract:
Topologically non-trivial nodes emerging from flat-band crossings not only enhance unconventional topological responses but also play a fundamental role in exploring correlation-driven topological physics. Here, we report the exceptionally robust chiral-anomaly-dominated transport in D0_3-Fe_3Ga. First, we observe a combination of positive and negative magnetoresistance, ideal planar longitudinal…
▽ More
Topologically non-trivial nodes emerging from flat-band crossings not only enhance unconventional topological responses but also play a fundamental role in exploring correlation-driven topological physics. Here, we report the exceptionally robust chiral-anomaly-dominated transport in D0_3-Fe_3Ga. First, we observe a combination of positive and negative magnetoresistance, ideal planar longitudinal magnetoresistance (PLMR), and the planar Hall effect (PHE). Second, ultra-low-temperature resistivity exhibits pronounced non-Fermi-liquid (NFL) behavior, accompanied by the emergence of giant intrinsic anomalous Hall conductivity (AHC), in excellent agreement with our DFT calculations, which confirm the existence of tilted Weyl points arising from crossings of nearly three-dimensional (3D) flat bands. Most remarkably, we detect an exceptionally robust flat magnetoresistance (flat-MR) that persists without decay up to 33 T. This set of phenomena provides strong evidence that the Fermi level intersects the flattened Weyl crossings, offering confirmation of a topological flat-band semimetal. D0_3-Fe_3Ga presents a promising magnetic platform for quantum device innovations.
△ Less
Submitted 30 March, 2026;
originally announced March 2026.
-
Evaluating a Data-Driven Redesign Process for Intelligent Tutoring Systems
Authors:
Qianru Lyu,
Conrad Borchers,
Meng Xia,
Karen Xiao,
Paulo F. Carvalho,
Kenneth R. Koedinger,
Vincent Aleven
Abstract:
Past research has defined a general process for the data-driven redesign of educational technologies and has shown that in carefully-selected instances, this process can help make systems more effective. In the current work, we test the generality of the approach by applying it to four units of a middle-school mathematics intelligent tutoring system that were selected not based on suitability for…
▽ More
Past research has defined a general process for the data-driven redesign of educational technologies and has shown that in carefully-selected instances, this process can help make systems more effective. In the current work, we test the generality of the approach by applying it to four units of a middle-school mathematics intelligent tutoring system that were selected not based on suitability for redesign, as in previous work, but on topic. We tested whether the redesigned system was more effective than the original in a classroom study with 123 students. Although the learning gains did not differ between the conditions, students who used the Redesigned Tutor had more productive time-on-task, a larger number of skills practiced, and greater total knowledge mastery. The findings highlight the promise of data-driven redesign even when applied to instructional units *not* selected as likely to yield improvement, as evidence of the generality and wide applicability of the method.
△ Less
Submitted 30 March, 2026;
originally announced March 2026.
-
Semantic-Aware Interruption Detection in Spoken Dialogue Systems: Benchmark, Metric, and Model
Authors:
Kangxiang Xia,
Bingshen Mu,
Xian Shi,
Jin Xu,
Lei Xie
Abstract:
Achieving natural full-duplex interaction in spoken dialogue systems (SDS) remains a challenge due to the difficulty of accurately detecting user interruptions. Current solutions are polarized between "trigger-happy" VAD-based methods that misinterpret backchannels and robust end-to-end models that exhibit unacceptable response delays. Moreover, the absence of real-world benchmarks and holistic me…
▽ More
Achieving natural full-duplex interaction in spoken dialogue systems (SDS) remains a challenge due to the difficulty of accurately detecting user interruptions. Current solutions are polarized between "trigger-happy" VAD-based methods that misinterpret backchannels and robust end-to-end models that exhibit unacceptable response delays. Moreover, the absence of real-world benchmarks and holistic metrics hinders progress in the field. This paper presents a comprehensive frame-work to overcome these limitations. We first introduce SID-Bench, the first benchmark for semantic-aware interruption detection built entirely from real-world human dialogues. To provide a rigorous assessment of the responsiveness-robustness trade-off, we propose the Average Penalty Time (APT) metric, which assigns a temporal cost to both false alarms and late responses. Building on this framework, we design an LLM-based detection model optimized through a novel training paradigm to capture subtle semantic cues of intent. Experimental results show that our model significantly outperforms mainstream baselines, achieving a nearly threefold reduction in APT. By successfully resolving the long-standing tension between speed and stability, our work establishes a new state-of-the-art for intelligent interruption handling in SDS. To facilitate future research, SID-Bench and the associated code are available at: https://github.com/xkx-hub/SID-bench.
△ Less
Submitted 25 March, 2026;
originally announced March 2026.
-
Quality Over Clicks: Iterative Reinforcement Learning for Early-Stage E-Commerce Query Suggestion
Authors:
Qi Sun,
Kejun Xiao,
Huaipeng Zhao,
Tao Luo,
Xiaoyi Zeng
Abstract:
Existing dialogue systems rely on query suggestion to enhance user engagement. Recent approaches mainly optimize generative models using click-through rate (CTR) models to align with user preferences. However, these methods are less effective in early-stage deployment scenarios, where click feedback is sparse and insufficient for training a reliable CTR model. To bridge this gap, we propose QualEQ…
▽ More
Existing dialogue systems rely on query suggestion to enhance user engagement. Recent approaches mainly optimize generative models using click-through rate (CTR) models to align with user preferences. However, these methods are less effective in early-stage deployment scenarios, where click feedback is sparse and insufficient for training a reliable CTR model. To bridge this gap, we propose QualEQS, a quality-first iterative reinforcement learning framework for e-commerce query suggestion. We formalize actionable suggestion quality along three dimensions that directly affect downstream usability: answerability, factuality, and information gain. To continuously improve from online traffic without click supervision, we further propose group-level disagreement among candidate suggestions to identify ambiguous query contexts and mine hard training cases for iterative refinement. We also introduce EQS-Benchmark, a dataset of 16,949 real-world e-commerce queries for offline training and evaluation. Experiments show that our quality-based offline metrics correlate strongly with online performance, providing a practical evaluation recipe for sparse-feedback deployment. In both offline and online settings, QualEQS consistently outperforms strong baselines, yielding a 6.81% improvement in online ChatPV in a real-world enterprise-level conversational shopping assistant system.
△ Less
Submitted 18 June, 2026; v1 submitted 24 March, 2026;
originally announced March 2026.
-
Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding
Authors:
Xianjin Wu,
Dingkang Liang,
Tianrui Feng,
Kui Xia,
Yumeng Zhang,
Xiaofan Li,
Xiao Tan,
Xiang Bai
Abstract:
While Multimodal Large Language Models demonstrate impressive semantic capabilities, they often suffer from spatial blindness, struggling with fine-grained geometric reasoning and physical dynamics. Existing solutions typically rely on explicit 3D modalities or complex geometric scaffolding, which are limited by data scarcity and generalization challenges. In this work, we propose a paradigm shift…
▽ More
While Multimodal Large Language Models demonstrate impressive semantic capabilities, they often suffer from spatial blindness, struggling with fine-grained geometric reasoning and physical dynamics. Existing solutions typically rely on explicit 3D modalities or complex geometric scaffolding, which are limited by data scarcity and generalization challenges. In this work, we propose a paradigm shift by leveraging the implicit spatial prior within large-scale video generation models. We posit that to synthesize temporally coherent videos, these models inherently learn robust 3D structural priors and physical laws. We introduce VEGA-3D (Video Extracted Generative Awareness), a plug-and-play framework that repurposes a pre-trained video diffusion model as a Latent World Simulator. By extracting spatiotemporal features from intermediate noise levels and integrating them with semantic representations via a token-level adaptive gated fusion mechanism, we enrich MLLMs with dense geometric cues without explicit 3D supervision. Extensive experiments across 3D scene understanding, spatial reasoning, and embodied manipulation benchmarks demonstrate that our method outperforms state-of-the-art baselines, validating that generative priors provide a scalable foundation for physical-world understanding. Code is publicly available at https://github.com/H-EmbodVis/VEGA-3D.
△ Less
Submitted 17 July, 2026; v1 submitted 19 March, 2026;
originally announced March 2026.