-
Quantum Horizon Tadpole and Emergence of a de Sitter Interior
Authors:
Chong-Sun Chu
Abstract:
In a recent paper \cite{Chu:2026dhx}, the quantum stability of the fuzzy sphere was established at large but finite $N$. In addition, a tadpole was identified for the scaling fluctuation mode of the matrix geometry. In this paper, we show that the tadpole generates a positive surface tension and tends to contract the sphere if the horizon is left isolated. However, when the horizon is coupled to g…
▽ More
In a recent paper \cite{Chu:2026dhx}, the quantum stability of the fuzzy sphere was established at large but finite $N$. In addition, a tadpole was identified for the scaling fluctuation mode of the matrix geometry. In this paper, we show that the tadpole generates a positive surface tension and tends to contract the sphere if the horizon is left isolated. However, when the horizon is coupled to gravity, the tadpole requires the bulk geometry to adjust according to the junction condition of general relativity. Assuming a static, spherically symmetric and non-singular vacuum interior, pure de Sitter space is selected. The matching also fixes an intriguing large-$N$ soldering relation between the de Sitter time inside and the Schwarzschild time outside. Potential cosmological implications are discussed.
△ Less
Submitted 21 August, 2026;
originally announced August 2026.
-
MR-IQA-2: Faithful Image Quality Reflection via Fine-Grained Credit Assignment
Authors:
Yuan li,
Youyuan Lin,
Chenhui Chu,
Shin'ya Nishida
Abstract:
Multimodal large language models (MLLMs) have shown strong potential for image quality assessment (IQA) by improving consistency between quality ratings and their underlying reasoning. However, most approaches supervise reasoning through human-provided ratings and rarely examine whether it faithfully reflects image quality. Rating accuracy alone does not ensure faithful reasoning; a shared reward…
▽ More
Multimodal large language models (MLLMs) have shown strong potential for image quality assessment (IQA) by improving consistency between quality ratings and their underlying reasoning. However, most approaches supervise reasoning through human-provided ratings and rarely examine whether it faithfully reflects image quality. Rating accuracy alone does not ensure faithful reasoning; a shared reward also obscures supervision sources and may reinforce unfaithful reasoning when a correct rating occurs by chance. To improve the faithfulness and reliability of blind IQA, we aim to (1) decouple credit assignment for reasoning and rating and (2) provide verifiable supervision for faithful reasoning. We introduce MR-IQA-2, an actor-editor-judge framework that operationalizes reasoning-editing-reflection. The actor generates quality reasoning for an input image, and the editor revises the image according to the identified quality factors. A frozen judge compares the original and edited images and provides reflective supervision for the actor's reasoning. MR-IQA-2 further uses fine-grained credit assignment to decouple reasoning and rating supervision. Judge feedback supervises reasoning, whereas human ratings supervise the predicted rating. Masked token-specific updates distinguish these signals while preserving the causal relation from reasoning to rating. Across IQA benchmarks, MR-IQA-2 achieves competitive rating alignment with humans. Visual reflection also enables richer and more faithful visual understanding beyond rating, which may inform image-quality optimization and related downstream tasks. Code is available at https://github.com/RobinY99/MR-IQA-2.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
Quantum Stability of the Fuzzy Sphere Black Hole Horizon
Authors:
Chong-Sun Chu
Abstract:
We establish the perturbative stability of the fuzzy-sphere black-hole horizon in the large-$N$ matrix quantum mechanics of \cite{Chu:2024qil}. Large-$N$ counting shows that the leading quantum correction to the fluctuation spectrum is the planar bosonic one-loop contribution, while higher-loop bosonic effects and the leading fermionic one-loop contribution are parametrically suppressed. The class…
▽ More
We establish the perturbative stability of the fuzzy-sphere black-hole horizon in the large-$N$ matrix quantum mechanics of \cite{Chu:2024qil}. Large-$N$ counting shows that the leading quantum correction to the fluctuation spectrum is the planar bosonic one-loop contribution, while higher-loop bosonic effects and the leading fermionic one-loop contribution are parametrically suppressed. The classical spectrum contains a tachyonic $(1,2)$ mode and a marginal $(2,3)$ mode; we show that both acquire positive quantum curvature and are stabilized. For the general spectrum, we find that a factorization structure of the Hessian controls the leading quantum curvature and renders it non-negative. Quantum effects dominate the stabilization of generic low-angular-momentum modes, whereas classical curvature dominates at high angular momentum. The two effects become comparable in the crossover regime $L \sim \sqrt{N}$, where both must be retained. Under smoothness and non-saturation assumptions supported by finite-$N$ numerical tests, positivity of the quantum fluctuation spectrum is thus established for sufficiently large but finite $N$. This local curvature stability is compatible with non-perturbative monopole tunneling. Intriguingly, the same $L\sim\sqrt{N}$ mesoscopic angular momentum range also dominates the tunneling process, suggesting a distinguished IR--UV crossover regime in the quantum dynamics of the black-hole horizon.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Asymptotic Behavior and Error Bounds for Fisher-KPP Equations on the Real Half-Line
Authors:
Chu Chu,
M. W. Wong
Abstract:
We study the Fisher--KPP equation on the half-line under Dirichlet,Neumann, and Robin boundary conditions. For the autonomous logistic equation, we identify bounded stationary profiles converging to $1$ and obtain exponential far-field comparison estimates. We prove local uniform convergence of nontrivial Neumann solutions to $1$. Assuming local uniform convergence of the Robin solution to its sta…
▽ More
We study the Fisher--KPP equation on the half-line under Dirichlet,Neumann, and Robin boundary conditions. For the autonomous logistic equation, we identify bounded stationary profiles converging to $1$ and obtain exponential far-field comparison estimates. We prove local uniform convergence of nontrivial Neumann solutions to $1$. Assuming local uniform convergence of the Robin solution to its stationary profile, we derive asymptotic Neumann--Robin comparison estimates. We then consider small time-periodic Neumann and Robin boundary forcing. Under exponential stability of the homogeneous linearized semigroup, we construct a locally unique small periodic lifted mild solution and obtain a first-order expansion with a uniform $O(\varepsilon^2)$ remainder in $C_0([0,\infty))$.
△ Less
Submitted 13 August, 2026; v1 submitted 27 July, 2026;
originally announced August 2026.
-
VSWR-Resilient Mm-Wave and Cm-Wave PAs for Large-Scale Phased Arrays
Authors:
Chenhao Chu,
Masoud Pashaeifar,
Filippo Svelto,
Hua Wang
Abstract:
Large-scale mm-Wave and cm-Wave phased arrays have become central to wireless communication and sensing systems, including terrestrial 5G/6G links and base stations, non-terrestrial networks (NTNs), satellite communication (SATCOM), radar, and relay applications. In dense arrays, power amplifiers (PAs) and antenna array interact with each other: array radiation depends on the amplitude/phase of PA…
▽ More
Large-scale mm-Wave and cm-Wave phased arrays have become central to wireless communication and sensing systems, including terrestrial 5G/6G links and base stations, non-terrestrial networks (NTNs), satellite communication (SATCOM), radar, and relay applications. In dense arrays, power amplifiers (PAs) and antenna array interact with each other: array radiation depends on the amplitude/phase of PA output signals, while the load impedance experienced by each PA varies with frequency, scan angle, and element position. This active antenna impedance, described as voltage standing wave ratio (VSWR) variation, arises mainly from antenna inter-element mutual coupling and is shaped by package/interconnect parasitics. Each PA can deviate from its optimum large-signal operating condition, degrading output power, power gain, power-added efficiency, AM-AM/AM-PM, and reliability margin. These variations further affect array EIRP consistency, EVM headroom, link budget, thermal density, and beamforming calibration, complicating PA design for wideband, wide-scan-angle arrays. This review introduces the origins of antenna VSWR and antenna-PA interactions connecting the antenna reflection coefficient $Γ_{\mathrm{ant}}$ and PA output matching $S_{22}$ to delivered-power and transmitted-phase variation through the $S_{22}Γ_{\mathrm{ant}}$ dependence. Reverse-coupled excitation and reverse intermodulation distortion (RIMD) are discussed. This motivates PA designs that achieve simultaneous output and loadline matching (SOLM), enabling a small output reflection coefficient $|S_{22}|$ without significantly compromising large-signal performance. Recent mm-Wave and cm-Wave VSWR-resilient integrated PA techniques and demonstrations are reviewed. Finally, challenges and opportunities for compact, load-insensitive, energy-efficient, high-power-density, and calibration-scalable integrated PAs are discussed.
△ Less
Submitted 31 July, 2026;
originally announced August 2026.
-
From Horizon Microstates to the Black Hole Membrane
Authors:
Chong-Sun Chu
Abstract:
The membrane paradigm represents a black-hole horizon by a fictitious conducting surface. We derive a microscopic electromagnetic membrane from black-hole matrix quantum mechanics. The fuzzy-sphere horizon carries a Berry monopole, placing its fundamental fermionic partons in lowest-Landau-level states with Ohmic and Hall responses. Off-diagonal bifundamental modes connecting the horizon and exter…
▽ More
The membrane paradigm represents a black-hole horizon by a fictitious conducting surface. We derive a microscopic electromagnetic membrane from black-hole matrix quantum mechanics. The fuzzy-sphere horizon carries a Berry monopole, placing its fundamental fermionic partons in lowest-Landau-level states with Ohmic and Hall responses. Off-diagonal bifundamental modes connecting the horizon and exterior matrix blocks become tachyonic near the horizon and condense, dynamically coupling the horizon gauge field to the exterior Maxwell field. In the low-frequency regime $ωR\ll1$, the fixed-parton transport description predicts frequency- and helicity-dependent reflectivity. At larger frequency, real parton excitations require a black-hole $S$-matrix. Ohm's law then fixes the inclusive absorption probability; unitarity bounds the local absorption cross section by the horizon area, with the classical conductivity $1/4π$ saturating this maximal-absorption bound.
△ Less
Submitted 30 July, 2026;
originally announced July 2026.
-
Risk-Routed Implicit Boundary Refinement for Robust Ultrasound Image Segmentation
Authors:
Jingguo Qu,
Xinyang Han,
Xiang Wang,
Yuqi Yang,
Tonghuan Xiao,
Sheng Ning,
Jing Qin,
Ann Dorothy King,
Winnie Chiu-Wing Chu,
Jing Cai,
Michael Ying
Abstract:
Medical ultrasound (US) image segmentation faces significant challenges due to speckle noise, low-contrast boundaries, acoustic shadowing, and acquisition variation across operators and clinical centers. Although encoder-decoder and transformer-based networks have achieved strong performance, many methods recover boundary details through dense decoders or larger backbones, which may still produce…
▽ More
Medical ultrasound (US) image segmentation faces significant challenges due to speckle noise, low-contrast boundaries, acoustic shadowing, and acquisition variation across operators and clinical centers. Although encoder-decoder and transformer-based networks have achieved strong performance, many methods recover boundary details through dense decoders or larger backbones, which may still produce over-smoothed contours or unstable predictions under external distribution shifts. In this article, we propose Risk-routed Implicit Boundary Refinement (RIBR), a compact segmentation framework that uses implicit neural representation as a risk-routed residual correction rather than an unconstrained full-mask predictor. RIBR combines boundary-refinement implicit residuals, risk-routed residual control, and geometry- and speckle-aware boundary regularization to refine uncertain contours while suppressing non-boundary oscillations. Evaluation on nine US datasets covering lymph nodes, breast lesions, thyroid nodules, and prostate shows that RIBR achieves the best overall macro-average and consistently reduces boundary error across grouped and organ-specific comparisons under a compact parameter budget. These findings suggest that controlled implicit residual learning is a practical strategy for resource-constrained and boundary-sensitive US segmentation. Source code is available at https://github.com/jinggqu/ribr.
△ Less
Submitted 23 July, 2026;
originally announced July 2026.
-
SoK: Adversarial Robustness of the Variational Quantum Eigensolver via Red-Teaming
Authors:
Ahmed Azaz Humdoon,
Cheng Chu,
Lei Jiang,
Qian Lou,
Mengxin Zheng
Abstract:
The Variational Quantum Eigensolver (VQE) is a leading algorithm for estimating molecular ground-state energies on near-term quantum hardware, with applications spanning quantum chemistry, materials science, and drug discovery. As VQE workloads are increasingly deployed through cloud-based ``VQE-as-a-service'' pipelines, they become exposed to adversaries such as compromised service components, ma…
▽ More
The Variational Quantum Eigensolver (VQE) is a leading algorithm for estimating molecular ground-state energies on near-term quantum hardware, with applications spanning quantum chemistry, materials science, and drug discovery. As VQE workloads are increasingly deployed through cloud-based ``VQE-as-a-service'' pipelines, they become exposed to adversaries such as compromised service components, malicious co-tenants, or insiders in the transpilation stack, any of which can corrupt results before they reach the user. A range of attacks on variational quantum circuits has been proposed, but each has been studied in isolation: some on quantum classifiers with accuracy-based metrics, others on variational quantum algorithms with energy-error metrics. This lack of a common evaluation setup makes their relative severity difficult to compare and leaves the security of VQE poorly characterized. In this work, we present \textbf{VQE-AdvBench}, the first unified red-teaming benchmark for the Variational Quantum Eigensolver, systematizing these attacks under a single evaluation protocol to rigorously assess VQE's adversarial robustness. We organize attacks along a black-, gray-, and white-box access taxonomy, and evaluate seven representative attack scenarios -- the QTrojan circuit backdoor, the QDoor parameter backdoor, parameter-space adaptations of FGSM and PGD, and three QNBAD noise-induced variants -- over a fixed molecule-ansatz-backend-metric configuration, on H$_2$ and H$_3^+$ across five noise-calibrated IBM backends. Our results reveal a clear severity ordering: noise-induced attacks that manipulate the Zero-Noise Extrapolation (ZNE) pipeline are the most damaging (up to 8.84$\times$ error amplification), followed by the QTrojan circuit-level backdoor (7.52$\times$), while the QDoor parameter-level backdoor is the least effective, yielding only marginal amplification (up to 1.37$\times$).
△ Less
Submitted 21 July, 2026;
originally announced July 2026.
-
Hardware Robustness of Sample-Based Quantum Diagonalization
Authors:
Ahatesham Bhuiyan,
Cheng Chu,
Qian Lou,
Mengxin Zheng
Abstract:
Sample-based Quantum Diagonalization (SQD) is a hybrid quantum-classical method that replaces variational optimization with a self-consistent recovery loop over QPU samples. Although SQD is considered robust to noisy samples and imperfect classical inputs, its robustness across practical deployment choices has not been systematically analyzed. As a result, shot budgets, qubit layouts, noise mitiga…
▽ More
Sample-based Quantum Diagonalization (SQD) is a hybrid quantum-classical method that replaces variational optimization with a self-consistent recovery loop over QPU samples. Although SQD is considered robust to noisy samples and imperfect classical inputs, its robustness across practical deployment choices has not been systematically analyzed. As a result, shot budgets, qubit layouts, noise mitigation strategies, and the coupled-cluster singles and doubles (CCSD) amplitudes that initialize the ansatz are often chosen without clear empirical guidance. We analyze SQD robustness on IBM Heron hardware across these dimensions. Structured CCSD-amplitude perturbations, including complete zeroing, produce only modest energy shifts from the clean baseline. Differences across layouts and noise-mitigation settings are large in the first recovery iteration but narrow within a few iterations. Accuracy saturates at moderate shot budgets, while very large budgets slightly worsen recovered energies, likely because working-set selection limits the value of additional samples. These results identify where SQD provides genuine deployment robustness and where its limits remain.
△ Less
Submitted 20 July, 2026;
originally announced July 2026.
-
CutBackdoor: A Circuit Cut Triggered Backdoor Attack on Variational Quantum Algorithms
Authors:
Ahatesham Bhuiyan,
Hoang Ngo,
Cheng Chu,
Qian Lou,
Lei Jiang,
My T. Thai,
Mengxin Zheng
Abstract:
Variational Quantum Algorithms (VQAs) are a leading paradigm for near-term quantum computing, combining parameterized quantum circuits with classical optimization across quantum chemistry, combinatorial optimization, and quantum machine learning. Since real-world VQA deployments routinely require circuits that exceed available hardware capacity, quantum circuit cutting has become an indispensable…
▽ More
Variational Quantum Algorithms (VQAs) are a leading paradigm for near-term quantum computing, combining parameterized quantum circuits with classical optimization across quantum chemistry, combinatorial optimization, and quantum machine learning. Since real-world VQA deployments routinely require circuits that exceed available hardware capacity, quantum circuit cutting has become an indispensable execution strategy, and pre-trained parameters are increasingly distributed through public repositories, introducing supply-chain security risks that have received little attention. Prior quantum backdoor attacks either introduce detectable circuit modifications or depend on device-specific noise, and none consider circuit cutting as an attack surface. We present CutBackdoor, the first parameter-supply-chain backdoor that uses cut circuit execution from CutQC as the deployment-time trigger against VQAs. Under noisy finite-shot circuit-cut execution, poisoned parameters preserve full-circuit validation performance while substantially increasing cut-path reconstruction error, without any circuit modification. The trigger activates when a resource-limited victim responds to a qubit-capacity mismatch by invoking the cutting workflow, requiring no attacker presence at deployment. We provide a theoretical analysis and empirically validate it across varying shot budgets. Evaluation across multiple VQA benchmarks on IBM quantum backends demonstrates cut-path energy amplification of $1.3\times$ to $2.9\times$ \revA{over clean baselines on the VQE and VQD benchmarks while maintaining small stealthiness error on the full-circuit path. The cut-path gap persists across the evaluated backends and cut placements under matched compilation; Zero-Noise Extrapolation provides only partial mitigation, and the diagonal-cost QAOA benchmark
delineates the attack's structural boundary
△ Less
Submitted 20 July, 2026;
originally announced July 2026.
-
MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for Parameter-Efficient Reasoning in Compact Models
Authors:
Martino M. L. Pulici,
Cuong Xuan Chu,
Evgeny Kharlamov,
Zifeng Ding,
Volker Tresp,
Yunpu Ma
Abstract:
Large language models achieve strong reasoning performance, but often at prohibitive training cost - a challenge that is especially acute for compact models ($\leq 4 \, \mathrm{B}$ parameters) trained under limited budgets. We introduce MADA-RL, a post-training framework that specializes compact models into generator and critic roles and trains them with a debate-aware learning signal, fine-tuning…
▽ More
Large language models achieve strong reasoning performance, but often at prohibitive training cost - a challenge that is especially acute for compact models ($\leq 4 \, \mathrm{B}$ parameters) trained under limited budgets. We introduce MADA-RL, a post-training framework that specializes compact models into generator and critic roles and trains them with a debate-aware learning signal, fine-tuning only a small subset of parameters via LoRA adapters. Our central contribution is a counterfactual critic advantage: a dynamic, role-conditioned baseline that redefines the critic's advantage as its reward minus the generator ensemble's per-instance accuracy. This explicitly optimizes critics to improve over generator consensus rather than to merely reproduce a correct answer, yielding more targeted credit assignment than static mean-reward normalization. At deployment, the specialized agents are composed in a lightweight multi-round protocol. Across five mathematical reasoning benchmarks, MADA-RL raises the accuracy of the DeepSeek-R1-Distill-Qwen-1.5B model from $39.9 \, \%$ to $41.9 \, \%$ ($+2.0$ points, $p < 0.001$) using $16$ times fewer trainable parameters than fully fine-tuned baselines, placing it on the accuracy-trainable-parameter Pareto front. It approaches, but does not surpass, the strongest baselines (DeepScaleR, STILL-3), which are trained on substantially larger datasets; we analyse this gap and the associated inference-time cost directly. A controlled study isolates the source of MADA-RL's gains: the counterfactual advantage produces the highest critic improvement rate of any model evaluated, indicating that trained critics learn to correct generator errors rather than to imitate them.
△ Less
Submitted 20 July, 2026;
originally announced July 2026.
-
Quantum Horizon and Quantum Membrane Paradigm from Black Hole Quantum Mechanics
Authors:
Chong-Sun Chu
Abstract:
We develop a microscopic quantum membrane paradigm from the matrix quantum mechanics of black holes [1]. It was proposed that a quantum black hole is described by a fuzzy sphere together with a half-filled Fermi sea of horizon partons. We show that the topology of the fuzzy sphere induces a Berry monopole, providing a microscopic origin for the monopole appearing in the tunneling description of th…
▽ More
We develop a microscopic quantum membrane paradigm from the matrix quantum mechanics of black holes [1]. It was proposed that a quantum black hole is described by a fuzzy sphere together with a half-filled Fermi sea of horizon partons. We show that the topology of the fuzzy sphere induces a Berry monopole, providing a microscopic origin for the monopole appearing in the tunneling description of the decay of quantum black hole. Because the partons couple to the fuzzy sphere as fundamental degrees of freedom, they are sensitive to this monopole and form Lowest-Landau-Level (LLL) states rather than ordinary propagating two-dimensional fermions. Their guiding-center dynamics generates Ohmic, Hall, and polarization currents on the horizon. To couple these currents to an external electromagnetic field, we introduce a two-block matrix configuration comprising black-hole, environmental, and bifundamental link sectors. Off-diagonal link modes become tachyonic near the fuzzy sphere and condense, dynamically locking the horizon gauge field to the boundary value of the external Maxwell field. The parton current thereby becomes a physical boundary source for the exterior field. Our construction replaces the fictitious membrane of the classical paradigm with a dynamical quantum membrane populated by microscopic LLL degrees of freedom. The resulting boundary condition generalizes the classical ingoing membrane condition and yields explicit quantum, frequency-dependent, and helicity-dependent corrections, offering a possible probe of the microscopic quantum structure of the horizon.
△ Less
Submitted 30 July, 2026; v1 submitted 12 July, 2026;
originally announced July 2026.
-
Chiral-Structured Superconductors TrX4 (Tr = Rh, Ir; X = Ge, Si): A Platform for Mixed-Parity Pairing and Topological States
Authors:
Zhenhai Yu,
Yunguan Ye,
Yuwei Zhou,
Chaoyang Chu,
Congcong Le,
Lin Wu,
Jian Yuan,
Tong Shi,
Qingxin Dong,
Jinggeng Zhao,
Wei Xia,
Xiangqi Liu,
Xia Wang,
Bosen Wang,
Jinguang Cheng,
Yanhang Ma,
Xianxin Wu,
Xiangang Wan,
Huiqiu Yuan,
Yanfeng Guo
Abstract:
Chiral-structured superconductors, with simultaneous broken mirror and inversion symmetries, promote unconventional superconductivity through parity-mixing mechanisms. Yet a few bulk chiral-structured superconductors are known, partly due to the difficulty in directly determining their atomic-scale chirality. Here we report three chiral-structured superconductors, , RhGe4, IrGe4, and IrSi4, synthe…
▽ More
Chiral-structured superconductors, with simultaneous broken mirror and inversion symmetries, promote unconventional superconductivity through parity-mixing mechanisms. Yet a few bulk chiral-structured superconductors are known, partly due to the difficulty in directly determining their atomic-scale chirality. Here we report three chiral-structured superconductors, , RhGe4, IrGe4, and IrSi4, synthesized under high pressure, with Tc values of about 1.6 K, 1.1 K, and 2.5 K, respectively.Using atomic resolution Cs-corrected scanning transmission electron microscopy (STEM) combined with X-ray diffraction characterizations, we directly confirm their chiral structure (space group P3121). This real space imaging approach overcomes ambiguities in traditional diffraction based methods. These materials exhibit type-II superconductivity, and the enhancement of spin-orbit coupling (SOC) leads to the emergence of mixed parity pairing. Calculations also reveal symmetry protected Weyl points near the Fermi level, which is robust against the SOC. Our work not only expands the family of chiral-structured superconductors but also demonstrates the indispensable role of STEM in directly determining chiral crystal structures. These materials thus offer a clean platform to explore the interplay among structural chirality, SOC, mixed parity superconductivity, and topological quantum phenomena.
△ Less
Submitted 7 July, 2026;
originally announced July 2026.
-
LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension
Authors:
Shunya Kato,
Taiki Miyanishi,
Shuhei Kurita,
Mahiro Ukai,
Nakamasa Inoue,
Chenhui Chu
Abstract:
Egocentric videos capture rich and diverse human-object interactions and have emerged as a fundamental resource for understanding human activities related to objects. In this context, Video Referring Expression Comprehension (Video REC), the task of localizing the temporal and spatial extent of a referred object in video frames given a natural language query, plays a key role in linking textual de…
▽ More
Egocentric videos capture rich and diverse human-object interactions and have emerged as a fundamental resource for understanding human activities related to objects. In this context, Video Referring Expression Comprehension (Video REC), the task of localizing the temporal and spatial extent of a referred object in video frames given a natural language query, plays a key role in linking textual descriptions to observed objects in untrimmed egocentric recordings. However, existing egocentric Video REC benchmarks primarily focus on short video clips, where some target object appears densely within frames. Such settings do not reflect real-world egocentric recordings, which are long-form, untrimmed, and characterized by sparse object occurrences and complex activity transitions. To address this limitation, we introduce LongEgoRefer, a novel and challenging benchmark constructed from long-form videos in the Ego4D dataset. LongEgoRefer contains 1,498 referring expressions with an average video duration of 45 minutes. The benchmark exhibits extreme target sparsity, detailed linguistic descriptions, and complex human-object interactions embedded in long, dynamic egocentric narratives. Consequently, it defines a demanding spatio-temporal grounding problem that requires models to identify both when an event occurs and where the referred object appears within extended video sequences. We evaluate existing Video REC approaches, including training-free baselines based on vision-language models combined with Grounded SAM2. Extensive experiments show that even advanced baselines and current state-of-the-art models struggle significantly on LongEgoRefer. These results highlight the intrinsic difficulty of long-form egocentric spatio-temporal grounding and emphasize the need for more robust video understanding models.
△ Less
Submitted 2 July, 2026;
originally announced July 2026.
-
MR-IQA: A Unified Margin View of Regression and Ranking for Blind Image Quality Assessment
Authors:
Yuan Li,
Youyuan Lin,
Zitang Sun,
Yung-Hao Yang,
Kiyofumi Miyoshi,
Chenhui Chu,
Shin'ya Nishida
Abstract:
Blind image quality assessment (BIQA) is commonly built on two basic learning paradigms: regression and ranking. Regression calibrates absolute scores, whereas ranking recovers quality structure from ordinal relations. Although joint regression-ranking supervision often improves BIQA, the relation between the two paradigms remains largely empirical and underexplored. In this work, we revisit what…
▽ More
Blind image quality assessment (BIQA) is commonly built on two basic learning paradigms: regression and ranking. Regression calibrates absolute scores, whereas ranking recovers quality structure from ordinal relations. Although joint regression-ranking supervision often improves BIQA, the relation between the two paradigms remains largely empirical and underexplored. In this work, we revisit what underlies regression and ranking and identify pairwise relational distance, termed quality margin, as their common bridge. Our derivation shows that, at the objective-optimization level, both paradigms fit quality margins: regression fits margins induced by score endpoints, while ranking fits transformed or sign-level margins through preference probabilities. Motivated by this insight, we propose MR-IQA, a direct quality-margin optimization framework for reinforcement learning (RL)-based BIQA. MR-IQA samples quality scores and optimizes pairwise margin errors as policy rewards, thereby modeling quality structure more explicitly. Experiments on six BIQA benchmarks show competitive general performance, and controlled comparisons demonstrate that MR-IQA achieves the strongest average PLCC/SRCC over regression- or ranking-based RL methods. Our findings provide a new insight into unifying regression and ranking, offering a theoretical basis for understanding quality-structure modeling in BIQA and beyond. Code is available at https://github.com/RobinY99/MR-IQA.
△ Less
Submitted 30 June, 2026; v1 submitted 29 June, 2026;
originally announced June 2026.
-
Automated Quality Assessment of Geospatial Vector Data: A GeoAI Approach using Spatial Representation Learning
Authors:
Hao Li,
Chen Chu,
Filip Biljecki,
Cyrus Shahabi,
Wenwen Li
Abstract:
Geospatial vector data quality is a foundational research topic in GIS, yet classic rule-based quality assessment algorithms often struggle with diverse urban morphologies and massive data volumes. Recently, Geospatial Artificial Intelligence (GeoAI) shows promising potential for automating geospatial analysis, while its application to native vector data remains largely underexplored. To fill this…
▽ More
Geospatial vector data quality is a foundational research topic in GIS, yet classic rule-based quality assessment algorithms often struggle with diverse urban morphologies and massive data volumes. Recently, Geospatial Artificial Intelligence (GeoAI) shows promising potential for automating geospatial analysis, while its application to native vector data remains largely underexplored. To fill this research gap, we proposed Topo4Vec, an automated GeoAI framework, designed for scalable vector data quality assessment via advanced Spatial Representation Learning (SRL). Specifically, Topo4Vec relax the labor-intensive manual annotation process via topological error simulation, such as overlapping polygons and street network connectivity errors e.g., overshoots and undershoots. Then, it leverages state-of-the-art SRL approaches to encode complex, native vector geometries (e.g., polylines and polygons) into a latent space where topological errors are isolated from valid ones. A systematic performance evaluation across three study areas (Los Angeles, Munich, and Singapore) demonstrates the effectiveness and robustness of Topo4Vec, achieving a peak accuracy of 0.99 for detecting overlapping building footprints and 0.60 for overshoots and undershoots in street networks. Moreover, lessons learned from Topo4Vec shed a promising light into a scalable and autonomous GeoAI approach for large-scale vector data consistency and quality monitoring within the fast-growing geospatial data ecosystems. The code and data used in the paper are made openly available in https://figshare.com/s/612148eeb4bccadbd715.
△ Less
Submitted 23 June, 2026;
originally announced June 2026.
-
Sentence-Level Contextual Entrainment in Large Language Models
Authors:
Yang Liu,
Chenhui Chu
Abstract:
Contextual entrainment, which is a newly discovered phenomenon in large language models (LLMs), refers to the tendency of a model to assign higher probabilities to tokens that appear in its context. In this work, we extend this phenomenon from the token level to the sentence level by examining the per-token mean log-probability of a sentence instead of the probabilities of individual tokens. We in…
▽ More
Contextual entrainment, which is a newly discovered phenomenon in large language models (LLMs), refers to the tendency of a model to assign higher probabilities to tokens that appear in its context. In this work, we extend this phenomenon from the token level to the sentence level by examining the per-token mean log-probability of a sentence instead of the probabilities of individual tokens. We investigate sentence-level contextual entrainment across 26 LLMs from seven families and two datasets, which cover both subjective and objective tasks. We find that sentence-level contextual entrainment exists. This means that the sentences in the prompt (even if they are counterfactual statements) can significantly increase their probability during model inference time. As the model size increases, contextual entrainment gradually decreases. We also find that contextual entrainment is controlled by 2% to 4% of the attention heads. Turning off these attention heads can effectively mitigate contextual entrainment without hurting the model's performance.
△ Less
Submitted 22 June, 2026;
originally announced June 2026.
-
Scaling LLM Reasoning from Minimal Labels: A Semi-Supervised Framework with a Lightweight Verifier
Authors:
Keizo Kato,
Chenhui Chu,
Yugo Murawaki,
Sado Kurohashi
Abstract:
For the development of Large language models (LLMs), recent approaches to generating pseudo intermediate reasoning have shown remarkable progress. But they typically rely on large numbers of correctly annotated answers to assess reasoning quality. This paper presents a semi-supervised framework that scales reasoning learning from minimal supervision, turning reasoning verification itself into a da…
▽ More
For the development of Large language models (LLMs), recent approaches to generating pseudo intermediate reasoning have shown remarkable progress. But they typically rely on large numbers of correctly annotated answers to assess reasoning quality. This paper presents a semi-supervised framework that scales reasoning learning from minimal supervision, turning reasoning verification itself into a data creation mechanism. We train a lightweight reasoning-correctness classifier on only a few labeled samples, which judges whether intermediate reasoning traces generated by an LLM are valid. Furthermore, an entropy-based confidence threshold filters out unreliable samples, and the remaining high-confidence reasoning traces are used to fine-tune the model. Experiments on Verifiable Math Problems (Orca-Math subset) and Question Answering on Image Scene Graphs (GQA) with Visual Programming show that our method achieves accuracy comparable to using 10-15x more labeled data. Ablation analyses confirm that both the classifier and entropy filtering are essential for scalable and noise-resistant pseudo-labeling. By replacing expensive answer-level supervision with lightweight reasoning verification, our method provides a practical path toward constructing large-scale reasoning resources and paves the way for future autonomous reasoning systems that learn from minimal human input.
△ Less
Submitted 15 June, 2026;
originally announced June 2026.
-
Medical world models: representing medical states, modelling clinical dynamics and guiding intervention policies
Authors:
Ke Liu,
Mengxuan Li,
Yanyi Bao,
Tianyun Zhang,
Chong Chu,
Jiajun Bu,
Haishuai Wang
Abstract:
Medical diagnosis and treatment are dynamic processes in which patient states evolve over time and clinical interventions alter future outcomes. Although current medical AI can detect disease, estimate risk and generate reports, many systems still return static labels or scores, offering limited insight into how illness may progress or how alternative interventions may reshape its trajectory. Medi…
▽ More
Medical diagnosis and treatment are dynamic processes in which patient states evolve over time and clinical interventions alter future outcomes. Although current medical AI can detect disease, estimate risk and generate reports, many systems still return static labels or scores, offering limited insight into how illness may progress or how alternative interventions may reshape its trajectory. Medical world models adapt the world-model idea from artificial intelligence to healthcare by learning internal simulators of patient-state dynamics. Their long-term goal is to help clinicians anticipate deterioration, compare treatment-conditioned futures and tailor care to individual patients. Yet relevant work remains scattered across foundation models, longitudinal modelling, disease simulation, treatment-effect estimation, reinforcement learning and digital twins. To bridge this gap, this review outlines a roadmap for advancing medical AI from isolated diagnosis and prediction toward medical world models that simulate disease evolution and support intervention decisions. This roadmap is organized around three coupled capabilities: patient-state construction, clinical dynamics modelling and intervention decision support. Across representative systems, the comparison highlights what each capability contributes and how partial components can be integrated into more mature perception--dynamics--planning systems. Finally, we identify the challenges involved in turning plausible rollouts into clinically useful simulators. Related literature is available at https://github.com/1999kevin/awesome_medical_world_models.
△ Less
Submitted 15 June, 2026;
originally announced June 2026.
-
Artificial Intelligence for Power-Converter-Rich Electrical Systems: A Review
Authors:
Pengfeng Lin,
Yuan Gao,
Yuxi Tang,
Muhammad Waqas Qaisar,
Peifeng Hui,
Chuanlin Zhang,
Miao Zhu,
Xiaoyong,
Cao,
Chia-chi Chu,
Peng Wang
Abstract:
Power-converter-rich electrical systems, formed by renewable generation, electrified transportation, and inverter-based resources, exhibit strongly nonlinear dynamics, multi-physics design tradeoffs, fast control requirements, and growing reliability and cybersecurity constraints. These characteristics strain workflows that rely only on physics-based modeling, sequential optimization, and rule-bas…
▽ More
Power-converter-rich electrical systems, formed by renewable generation, electrified transportation, and inverter-based resources, exhibit strongly nonlinear dynamics, multi-physics design tradeoffs, fast control requirements, and growing reliability and cybersecurity constraints. These characteristics strain workflows that rely only on physics-based modeling, sequential optimization, and rule-based operation. This paper reviews artificial intelligence (AI) for power-converter-rich electrical systems through a life-cycle and deployment-readiness perspective. The literature is organized across converter design, real-time control, system-level operation, and compliance-oriented governance. For design, we examine surrogate modeling, topology and parameter synthesis, EMI/EMC-aware optimization, reliability-oriented design, and knowledge-assisted workflows. For control, we compare supervised learning, reinforcement learning, learning-augmented predictive control, and safety-constrained learning according to their role in closed-loop implementation. For operations, we focus on microgrid coordination, forecasting, distribution-system observability, privacy-preserving coordination, and cyber-resilient operation where converter-interfaced resources shape the operating problem. Across these stages, the review emphasizes deployment-critical gaps, including stability certification, constraint satisfaction, interpretability, extrapolation, data efficiency, sim-to-real transfer, embedded latency, cybersecurity, privacy, and standards alignment. The resulting taxonomy is intended to clarify where AI is already useful as an engineering support tool and where further validation is needed before autonomous or safety-critical deployment.
△ Less
Submitted 20 June, 2026; v1 submitted 14 June, 2026;
originally announced June 2026.
-
Can LLMs extract scientific consensus? A case study in high-temperature superconductivity
Authors:
Mouyang Cheng,
Wenhao He,
Zhuotao Jin,
Bowen Yu,
Ju Li,
Boris Kozinsky,
Yao Wang,
Pavel Volkov,
Liangzi Deng,
Ching-Wu Chu,
Xiao-Gang Wen,
Mingda Li
Abstract:
Scientific knowledge is increasingly dispersed across vast and heterogeneous scientific literature, where important claims are often implicit, evolving, and internally debated. While large language models (LLMs) have shown impressive performance in information extraction and summarization, their ability to recover latent scientific consensus remains unclear. Here, we investigate this problem in th…
▽ More
Scientific knowledge is increasingly dispersed across vast and heterogeneous scientific literature, where important claims are often implicit, evolving, and internally debated. While large language models (LLMs) have shown impressive performance in information extraction and summarization, their ability to recover latent scientific consensus remains unclear. Here, we investigate this problem in the context of high-temperature superconductivity (HTS), a long-standing and highly debated topic in condensed matter physics, as a challenging testbed. Using near 18,000 highly-cited publications over the past seven decades, we construct a structured knowledge graph linking competing superconducting mechanisms, material families, evidential modalities, and citation relations. We find that LLM-extracted representations recover coherent and physically interpretable structures, including family-dependent mechanism profiles, evidence-specific correlations, and citation-mediated temporal evolution of scientific beliefs. Ablation studies on LLM further show that the global structure remains robust across prompting, decoding, and model variations. Our results suggest that LLMs can indeed serve as scalable tools for deciphering scientific knowledge in domains characterized by competing interpretations and evolving knowledge.
△ Less
Submitted 25 May, 2026;
originally announced June 2026.
-
OneReason Technical Report
Authors:
OneRec Team,
Biao Yang,
Boyang Ding,
Chenglong Chu,
Dunju Zang,
Fei Pan,
Han Li,
Hao Jiang,
Honghui Bao,
Huanjie Wang,
Jian Liang,
Jiangxia Cao,
Jiao Ou,
Jiaxin Deng,
Jinghao Zhang,
Kun Gai,
Lu Ren,
Peiru Du,
Pengfei Zheng,
Rongzhou Zhang,
Ruiming Tang,
Shiyao Wang,
Siyang Mao,
Siyuan Lou,
Teng Shi
, et al. (59 additional authors not shown)
Abstract:
Generative recommendation models in the OneRec family have been widely deployed in many real-world services, such as short-video, live-streaming, advertising, and e-commerce. However, these generative models can only benefit from the scaling advantage, while their reasoning ability is hard to activate, since we cannot construct meaningful Chain-of-Thought (CoT) sequences consisting of itemic token…
▽ More
Generative recommendation models in the OneRec family have been widely deployed in many real-world services, such as short-video, live-streaming, advertising, and e-commerce. However, these generative models can only benefit from the scaling advantage, while their reasoning ability is hard to activate, since we cannot construct meaningful Chain-of-Thought (CoT) sequences consisting of itemic tokens only. Inspired by the success of the reasoning-style ``think before answer'' paradigm in the LLM field, we conduct preliminary studies (i.e., OneRec-Think, OpenOneRec) to explore reasoning capability in generative recommendation. Nevertheless, we notice an unexpected phenomenon: the thinking mode does not show advantages over the non-thinking mode. Drawing insights from recent findings on CoT robustness in multi-modal language models, we argue that effective reasoning in recommendation rests on two factors: perception, the ability to ground itemic tokens in their underlying language semantics, and cognition, the ability to reorganize a user's behavior sequence into coherent latent interest points. We therefore propose OneReason, which includes: (1) strong itemic token perception in pre-training, (2) a three-level cognition-enhanced CoT format for recommendation tasks in SFT, and (3) a specialize-then-unify training recipe in RL to enhance the thinking ability.
△ Less
Submitted 4 June, 2026;
originally announced June 2026.
-
From Symbolic to Geometric: Enabling Spatial Reasoning in Large Language Models
Authors:
Chen Chu,
Bita Azarijoo,
Li Xiong,
Khurram Shafique,
Cyrus Shahabi
Abstract:
Recent large language models (LLMs) often appear to exhibit spatial reasoning ability; however, this capability is largely \emph{symbolic}, arising from pattern matching over spatial language rather than true \emph{geometric} reasoning over space. Because LLMs operate on discrete tokens, they lack native support for continuous spatial representations, explicit geometric computation, and structured…
▽ More
Recent large language models (LLMs) often appear to exhibit spatial reasoning ability; however, this capability is largely \emph{symbolic}, arising from pattern matching over spatial language rather than true \emph{geometric} reasoning over space. Because LLMs operate on discrete tokens, they lack native support for continuous spatial representations, explicit geometric computation, and structured spatial operators. To address this limitation, we introduce the \emph{Spatial Language Model (SLM)}, the first multimodal LLM that treats location information as a first-class modality and enables geometric spatial reasoning within the model's inference process. SLM directly operates on learned spatial representations rather than textual descriptions of spatial relations. To support effective training, we construct a \emph{Spatial Instruction Dataset} that aligns spatial representations, atomic geometric operations, and natural language instructions. We further propose a new benchmark named \emph{SpatialEval}, which is designed to evaluate spatial reasoning across attributes, distance, topology, and relative-position tasks. Extensive experiments show that SLM significantly outperforms existing LLM-based approaches that rely on symbolic reasoning via prompt engineering or textual abstraction, demonstrating the benefits of integrating geometric spatial representations for robust spatial reasoning.
Our instruction dataset, evaluation benchmark, model training codes, and models' checkpoints can be found at:
\hyperlink{https://github.com/chuchen2017/SLM}{https://github.com/chuchen2017/SLM}.
△ Less
Submitted 2 June, 2026;
originally announced June 2026.
-
LeAP: Learnable Adaptive Permutation for Feature Selection in Heterogeneous and Sparse Recommender Systems
Authors:
Yihong Huang,
Chen Chu,
Fei Chen,
Yu Lin,
Ruiduan Li,
Zhihao Li
Abstract:
Modern industrial recommender systems rely on thousands of heterogeneous features -- ranging from low-dimensional scalars (e.g., statistical value) to high-dimensional embeddings (e.g., user-id embeddings, MLP representations) -- to achieve high-precision predictions. Given the immense computational costs associated with training, efficient feature selection is critical. However, existing methods…
▽ More
Modern industrial recommender systems rely on thousands of heterogeneous features -- ranging from low-dimensional scalars (e.g., statistical value) to high-dimensional embeddings (e.g., user-id embeddings, MLP representations) -- to achieve high-precision predictions. Given the immense computational costs associated with training, efficient feature selection is critical. However, existing methods encounter three primary bottlenecks: (1) they typically assume uniform feature dimensions or require costly mapping to a fixed size; (2) they struggle with extreme sparsity, where the majority of features (e.g., 99%+) remain at default values; and (3) traditional permutation-based approaches are computationally prohibitive in large-scale settings.
To address these challenges, we propose LeAP (Learnable Adaptive Permutation), a novel, model-agnostic plug-in module for feature selection. LeAP transforms the inefficient random permutation process into a learnable mechanism, significantly accelerating the evaluation of feature importance. In addition, we introduce an adaptive regularization strategy tailored for heterogeneous dimensions and extreme sparsity, enabling superior feature importance ranking results across asymmetric input spaces. Experiments on four public recommendation datasets demonstrate that LeAP achieves state-of-the-art performance. Furthermore, LeAP has been deployed in a large-scale industrial search ranking model with over a billion daily requests and a 2TB model parameter scale. In this real-world scenario involving 12,000+ total feature dimensions, LeAP successfully identified and removed over 3,600 redundant dimensions without performance degradation, which is 2 to 10 times the ability of compared baseline methods.
△ Less
Submitted 2 June, 2026; v1 submitted 31 May, 2026;
originally announced June 2026.
-
Learning Mid-circuit Measurement Backaction from Three Repeated Measurements
Authors:
Chia-Tung Chu,
Su-un Lee,
Han Zheng,
Senrui Chen,
Bibek Pokharel,
Alireza Seif,
Liang Jiang
Abstract:
Accurate modeling of mid-circuit measurements (MCMs) is essential for dynamic-circuit operations such as syndrome extraction, measurement-based reset, and the separation of state-preparation and measurement (SPAM) error. Unlike terminal measurement, a noisy MCM both produces a classical outcome and alters the incoming quantum state, thereby influencing subsequent circuit operations. This makes con…
▽ More
Accurate modeling of mid-circuit measurements (MCMs) is essential for dynamic-circuit operations such as syndrome extraction, measurement-based reset, and the separation of state-preparation and measurement (SPAM) error. Unlike terminal measurement, a noisy MCM both produces a classical outcome and alters the incoming quantum state, thereby influencing subsequent circuit operations. This makes conventional confusion-matrix or fidelity-level characterization insufficient. Here we introduce an efficient, self-consistent protocol for learning a single-qubit Z-twirled MCM instrument, retaining the readout-backaction correlations and excitation-decay asymmetry that are erased in Pauli-error descriptions. Remarkably, readout bit strings from only three repeated MCMs on a maximally mixed input determine all learnable parameters of the reduced instrument, up to a single unidentifiable gauge degree of freedom. Physicality constraints convert this non-identifiability into narrow, gauge-aware error intervals. Implemented on IBM superconducting processors, the learned instrument improves Pauli-observable prediction by ${\sim}100\times$ over a conventional confusion-matrix model and reveals a $T_1$-decay dominated backaction. Our protocol provides a compact characterization layer for SPAM error separation, reset optimization, and noise-aware quantum error correction.
△ Less
Submitted 29 May, 2026;
originally announced June 2026.
-
OSDTW: Optimal Shared Depth and Task Weighting for Long-Tailed Recognition
Authors:
Chang Chu,
Qingyue Zhang,
Shao-Lun Huang,
Junxiong Zheng
Abstract:
Long-tailed recognition suffers from a persistent head--tail trade-off: improving tail performance often degrades head accuracy and can increase training instability. Despite strong empirical results from re-weighting, decoupled training, and multi-expert methods, key design choices about representation sharing between head and tail classes and supervision weighting across class groups remain larg…
▽ More
Long-tailed recognition suffers from a persistent head--tail trade-off: improving tail performance often degrades head accuracy and can increase training instability. Despite strong empirical results from re-weighting, decoupled training, and multi-expert methods, key design choices about representation sharing between head and tail classes and supervision weighting across class groups remain largely heuristic. In this work, we propose OSDTW, a principled task-decomposition framework that partitions the original single-label recognition problem into a head task and a tail task, implemented with a shared encoder and task-specific decoders. To handle the mutual exclusivity and statistical dependence between the two label groups, we introduce a factorized model and show that the resulting Kullback--Leibler divergence-based generalization error can be written as the sum of task-wise terms up to an additive constant, yielding a well-defined task-wise objective. We further develop a three-stage training pipeline: independent task training to estimate task-wise optima and the Fisher information matrix, weighted joint training to learn a shared encoder, and branch assembly to construct the final decoupled model. Under a block-diagonal Fisher approximation, we derive a computable second-order expansion of the expected generalization error, decomposing it into encoder variance, encoder bias, and decoder variance. This bias--variance decomposition provides a computable proxy to select the shared depth and task weights, enabling efficient hyper-parameter search. Experiments on standard long-tailed benchmarks demonstrate the effectiveness of the proposed approach over strong baselines.
△ Less
Submitted 24 May, 2026;
originally announced May 2026.
-
MedExpMem: Adapting Experience Memory for Differential Diagnosis
Authors:
Qianhan Feng,
Zhongzhen Huang,
Yakun Zhu,
Yannian Gu,
Winnie Chiu Wing Chu,
Xiaofan Zhang,
Qi Dou
Abstract:
Experienced physicians develop diagnostic expertise through clinical practice, acquiring not only disease knowledge but also the ability to differentiate confusable conditions. Current medical vision-language models (VLMs) lack this capability -- their parameters encode static knowledge that does not evolve across diagnostic encounters. We propose MedExpMem, an experience memory framework enabling…
▽ More
Experienced physicians develop diagnostic expertise through clinical practice, acquiring not only disease knowledge but also the ability to differentiate confusable conditions. Current medical vision-language models (VLMs) lack this capability -- their parameters encode static knowledge that does not evolve across diagnostic encounters. We propose MedExpMem, an experience memory framework enabling VLM-based diagnostic agents to accumulate differential diagnosis expertise. Unlike retrieval-augmented generation, which retrieves encyclopedic disease descriptions, MedExpMem memorizes discriminative experience derived from the agent's own diagnostic failures and organizes them as pairwise differential notes encoding key discriminators, actionable decision rules and reasoning error patterns. The framework adopts a two-phase construction process mirroring physician learning: initial practice exposes knowledge gaps, and reflective re-diagnosis refines understanding. When encountering new cases, the agent retrieves experience memory to guide differential reasoning. We evaluate MedExpMem on a radiology benchmark spanning 11 subspecialties. Results demonstrate consistent accuracy improvements, maximum 7.0%, across diverse models and scales. Analytical experiments validate experience quality and robustness, demonstrating MedExpMem as a competitive method addresses medical adaptation needs beyond the reach of parameteric learning.
△ Less
Submitted 19 May, 2026;
originally announced May 2026.
-
MRecover: A Conditional Generative Model for Recovering Motion-Corrupted MR images Using AI Generated Contrast
Authors:
Jinghang Li,
Tales Santini,
Courtney Clark,
Bruno de Almeida,
Cong Chu,
Salem Alkhateeb,
Andrea Sajewski,
Jacob Berardinelli,
Hecheng Jin,
Tobias Campos,
Jeremy J. Berardo,
Joseph Mettenburg,
Ariel Gildengers,
Howard J. Aizenstein,
Minjie Wu,
Tamer S. Ibrahim
Abstract:
Hippocampal subfield segmentation requires high-resolution T2w turbo spin echo (TSE) MRI, yet this sequence is susceptible to motion artifacts, leading to substantial data loss. We developed a conditional generative model (MRecover) that synthesizes routinely acquired T1w images to create TSE images with autoregressive slice conditioning for volumetric consistency. Trained on 7T MRI data (n=577),…
▽ More
Hippocampal subfield segmentation requires high-resolution T2w turbo spin echo (TSE) MRI, yet this sequence is susceptible to motion artifacts, leading to substantial data loss. We developed a conditional generative model (MRecover) that synthesizes routinely acquired T1w images to create TSE images with autoregressive slice conditioning for volumetric consistency. Trained on 7T MRI data (n=577), the model achieved high in-domain fidelity (n=148, SSIM=0.84, FSIM=0.94) and generalized well to out-of-domain 3T data: subfield volumes from synthesized and the as-acquired images closely matched: (n=416, r=0.87-0.97) and yielded 31.8% more analyzable subjects in the motion-affected ADNI3 dataset after quality control (593 vs 450). The synthesized images also achieved larger effect sizes due to increasing the sample size for diagnostic group differences in hippocampal subfield atrophy (whole hippocampus $ε^2$= 0.121-0.100 vs. 0.086-0.062, left-right hemispheres). Project page: https://jinghangli98.github.io/MRecover/
△ Less
Submitted 20 May, 2026;
originally announced May 2026.
-
Kwai Summary Attention Technical Report
Authors:
Chenglong Chu,
Guorui Zhou,
Guowang Zhang,
Han Li,
Hao Peng,
Hongtao Cheng,
Hui Wang,
Jian Liang,
Jiangxia Cao,
Kun Gai,
Lingzhi Zhou,
Lu Ren,
Qi Zhang,
Ruiming Tang,
Ruitao Wang,
Xinchen Luo,
Yi Su,
Zhiyuan Liang,
Ziqi Wang,
Boyang Ding,
Chengru Song,
Dunju Zang,
Jiao Ou,
Jiaxin Deng,
Jijun Shi
, et al. (13 additional authors not shown)
Abstract:
Long-context ability, has become one of the most important iteration direction of next-generation Large Language Models, particularly in semantic understanding/reasoning, code agentic intelligence and recommendation system. However, the standard softmax attention exhibits quadratic time complexity with respect to sequence length. As the sequence length increases, this incurs substantial overhead i…
▽ More
Long-context ability, has become one of the most important iteration direction of next-generation Large Language Models, particularly in semantic understanding/reasoning, code agentic intelligence and recommendation system. However, the standard softmax attention exhibits quadratic time complexity with respect to sequence length. As the sequence length increases, this incurs substantial overhead in long-context settings, leading the training and inference costs of extremely long sequences deteriorate rapidly. Existing solutions mitigate this issue through two technique routings: i) Reducing the KV cache per layer, such as from the head-level compression GQA, and the embedding dimension-level compression MLA, but the KV cache remains linearly dependent on the sequence length at a 1:1 ratio. ii) Interleaving with KV Cache friendly architecture, such as local attention SWA, linear kernel GDN, but often involve trade-offs among KV Cache and long-context modeling effectiveness. Besides the two technique routings, we argue that there exists an intermediate path not well explored: {Maintaining a linear relationship between the KV cache and sequence length, but performing semantic-level compression through a specific ratio $k$}. This $O(n/k)$ path does not pursue a ``minimum KV cache'', but rather trades acceptable memory costs for complete, referential, and interpretable retention of long distant dependency. Motivated by this, we propose Kwai Summary Attention (KSA), a novel attention mechanism that reduces sequence modeling cost by compressing historical contexts into learnable summary tokens.
△ Less
Submitted 5 July, 2026; v1 submitted 27 April, 2026;
originally announced April 2026.
-
Understanding the Prompt Sensitivity
Authors:
Yang Liu,
Chenhui Chu
Abstract:
Prompt sensitivity, which refers to how strongly the output of a large language model (LLM) depends on the exact wording of its input prompt, raises concerns among users about the LLM's stability and reliability. In this work, we consider LLMs as multivariate functions and perform a first-order Taylor expansion, thereby analyzing the relationship between meaning-preserving prompts, their gradients…
▽ More
Prompt sensitivity, which refers to how strongly the output of a large language model (LLM) depends on the exact wording of its input prompt, raises concerns among users about the LLM's stability and reliability. In this work, we consider LLMs as multivariate functions and perform a first-order Taylor expansion, thereby analyzing the relationship between meaning-preserving prompts, their gradients, and the log probabilities of the model's next token. We derive an upper bound on the difference between log probabilities using the Cauchy-Schwarz inequality. We show that LLMs do not internally cluster similar inputs like smaller neural networks do, but instead disperse them. This dispersing behavior leads to an excessively high upper bound on the difference of log probabilities between two meaning-preserving prompts, making it difficult to effectively reduce to 0. In our analysis, we also show which types of meaning-preserving prompt variants are more likely to introduce prompt sensitivity risks in LLMs. In addition, we demonstrate that the upper bound is strongly correlated with an existing prompt sensitivity metric, PromptSensiScore. Moreover, by analyzing the logit variance, we find that prompt templates typically exert a greater influence on logits than the questions themselves. Overall, our results provide a general interpretation for why current LLMs can be highly sensitive to prompts with the same meaning, offering crucial evidence for understanding the prompt sensitivity of LLMs. Code for experiments is available at https://github.com/ku-nlp/Understanding_the_Prompt_Sensitivity.
△ Less
Submitted 20 April, 2026;
originally announced April 2026.
-
Rethinking Meeting Effectiveness: A Benchmark and Framework for Temporal Fine-grained Automatic Meeting Effectiveness Evaluation
Authors:
Yihang Li,
Chenhui Chu
Abstract:
Evaluating meeting effectiveness is crucial for improving organizational productivity. Current approaches rely on post-hoc surveys that yield a single coarse-grained score for an entire meeting. The reliance on manual assessment is inherently limited in scalability, cost, and reproducibility. Moreover, a single score fails to capture the dynamic nature of collaborative discussions. We propose a ne…
▽ More
Evaluating meeting effectiveness is crucial for improving organizational productivity. Current approaches rely on post-hoc surveys that yield a single coarse-grained score for an entire meeting. The reliance on manual assessment is inherently limited in scalability, cost, and reproducibility. Moreover, a single score fails to capture the dynamic nature of collaborative discussions. We propose a new paradigm for evaluating meeting effectiveness centered on novel criteria and temporal fine-grained approach. We define effectiveness as the rate of objective achievement over time and assess it for individual topical segments within a meeting. To support this task, we introduce the AMI Meeting Effectiveness (AMI-ME) dataset, a new meta-evaluation dataset containing 2,459 human-annotated segments from 130 AMI Corpus meetings. We also develop an automatic effectiveness evaluation framework that uses a Large Language Model (LLM) as a judge to score each segment's effectiveness relative to the overall meeting objectives. Through substantial experiments, we establish a comprehensive benchmark for this new task and evaluate the framework's generalizability across distinct meeting types, ranging from business scenarios to unstructured discussions. Furthermore, we benchmark end-to-end performance starting from raw speech to measure the capabilities of a complete system. Our results validate the framework's effectiveness and provide strong baselines to facilitate future research in meeting analysis and multi-party dialogue. Our dataset and code will be publicly available. The AMI-ME dataset and the Automatic Evaluation Framework are available at: this URL.
△ Less
Submitted 4 June, 2026; v1 submitted 19 April, 2026;
originally announced April 2026.
-
Retrieval-Augmented Large Language Models for Evidence-Informed Guidance on Cannabidiol Use in Older Adults
Authors:
Ali Abedi,
Charlene H. Chu,
Shehroz S. Khan
Abstract:
Older adults commonly experience chronic conditions such as pain and sleep disturbances and may consider cannabidiol for symptom management. Safe use requires appropriate dosing, careful titration, and awareness of drug interactions, yet stigma and limited health literacy often limit understanding. Conversational artificial intelligence systems based on large language models and retrieval-augmente…
▽ More
Older adults commonly experience chronic conditions such as pain and sleep disturbances and may consider cannabidiol for symptom management. Safe use requires appropriate dosing, careful titration, and awareness of drug interactions, yet stigma and limited health literacy often limit understanding. Conversational artificial intelligence systems based on large language models and retrieval-augmented generation may support cannabidiol education, but their safety and reliability remain insufficiently evaluated. This study developed a retrieval-augmented large language model framework that combines structured prompt engineering with curated cannabidiol evidence to generate context-aware guidance for older adults, including those with cognitive impairment. We also proposed an automated, annotation-free evaluation framework to benchmark leading standalone and retrieval-augmented models in the absence of standardized benchmarks. Sixty-four diverse user scenarios were generated by varying symptoms, preferences, cognitive status, demographics, comorbidities, medications, cannabis history, and caregiver support. Multiple state-of-the-art models were evaluated, including a novel ensemble retrieval architecture that integrates multiple retrieval systems. Across three automated evaluation strategies, retrieval-augmented models consistently produced more cautious and guideline-aligned recommendations than standalone models, with the ensemble approach performing best. These findings demonstrate that structured retrieval improves the reliability and safety of AI-driven cannabidiol education and provide a reproducible framework for evaluating AI tools used in sensitive health contexts.
△ Less
Submitted 15 January, 2026;
originally announced April 2026.
-
ImagenWorld: Stress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks
Authors:
Samin Mahdizadeh Sani,
Max Ku,
Nima Jamali,
Matina Mahdizadeh Sani,
Paria Khoshtab,
Wei-Chieh Sun,
Parnian Fazel,
Zhi Rui Tam,
Thomas Chong,
Edisy Kin Wai Chan,
Donald Wai Tong Tsang,
Chiao-Wei Hsu,
Ting Wai Lam,
Ho Yin Sam Ng,
Chiafeng Chu,
Chak-Wing Mak,
Keming Wu,
Hiu Tung Wong,
Yik Chun Ho,
Chi Ruan,
Zhuofeng Li,
I-Sheng Fang,
Shih-Ying Yeh,
Ho Kei Cheng,
Ping Nie
, et al. (1 additional authors not shown)
Abstract:
Advances in diffusion, autoregressive, and hybrid models have enabled high-quality image synthesis for tasks such as text-to-image, editing, and reference-guided composition. Yet, existing benchmarks remain limited, either focus on isolated tasks, cover only narrow domains, or provide opaque scores without explaining failure modes. We introduce \textbf{ImagenWorld}, a benchmark of 3.6K condition s…
▽ More
Advances in diffusion, autoregressive, and hybrid models have enabled high-quality image synthesis for tasks such as text-to-image, editing, and reference-guided composition. Yet, existing benchmarks remain limited, either focus on isolated tasks, cover only narrow domains, or provide opaque scores without explaining failure modes. We introduce \textbf{ImagenWorld}, a benchmark of 3.6K condition sets spanning six core tasks (generation and editing, with single or multiple references) and six topical domains (artworks, photorealistic images, information graphics, textual graphics, computer graphics, and screenshots). The benchmark is supported by 20K fine-grained human annotations and an explainable evaluation schema that tags localized object-level and segment-level errors, complementing automated VLM-based metrics. Our large-scale evaluation of 14 models yields several insights: (1) models typically struggle more in editing tasks than in generation tasks, especially in local edits. (2) models excel in artistic and photorealistic settings but struggle with symbolic and text-heavy domains such as screenshots and information graphics. (3) closed-source systems lead overall, while targeted data curation (e.g., Qwen-Image) narrows the gap in text-heavy cases. (4) modern VLM-based metrics achieve Kendall accuracies up to 0.79, approximating human ranking, but fall short of fine-grained, explainable error attribution. ImagenWorld provides both a rigorous benchmark and a diagnostic tool to advance robust image generation.
△ Less
Submitted 29 March, 2026;
originally announced March 2026.
-
Self-organized pattern synchronization modulated by stochasticity in coupled plankton ecosystems
Authors:
Ju Kang,
Yiyuan Niu,
Yuanzhi Li,
Quan-Xing Liu,
Chengjin Chu
Abstract:
Spatial patterning and synchronization are pervasive features of plankton communities, yet the mechanisms that allow such patterns to persist coherently under environmental noise remain unresolved. In vertically structured aquatic ecosystems, plankton populations are often organized into distinct layers, raising the question of how interactions between layers shape both spatial self-organization a…
▽ More
Spatial patterning and synchronization are pervasive features of plankton communities, yet the mechanisms that allow such patterns to persist coherently under environmental noise remain unresolved. In vertically structured aquatic ecosystems, plankton populations are often organized into distinct layers, raising the question of how interactions between layers shape both spatial self-organization and robustness. Here, we develop a spatiotemporal ecosystem model of a two-layer plankton community to examine the role of passive diffusive coupling under stochastic environmental fluctuations. We show that interlayer diffusion induces a sharp transition from independent, layer-specific Turing patterns to fully synchronized spatial patterns once the coupling strength exceeds a critical threshold. Importantly, the same coupling mechanism markedly enhances the stability of spatial patterns against environmental noise, extending their persistence far beyond that of non-coupled layers. Moreover, we uncover a trophic hierarchy in noise sensitivity, with zooplankton exhibiting substantially greater vulnerability than phytoplankton. Together, these results identify passive diffusive coupling as a unifying mechanism that simultaneously promotes spatial synchronization and robustness, providing a mechanistic explanation for the persistence of coherent plankton patterns in fluctuating aquatic environments.
△ Less
Submitted 25 March, 2026;
originally announced March 2026.
-
PFM-VEPAR: Prompting Foundation Models for RGB-Event Camera based Pedestrian Attribute Recognition
Authors:
Minghe Xu,
Rouying Wu,
ChiaWei Chu,
Xiao Wang,
Yu Li
Abstract:
Event-based pedestrian attribute recognition (PAR) leverages motion cues to enhance RGB cameras in low-light and motion-blur scenarios, enabling more accurate inference of attributes like age and emotion. However, existing two-stream multimodal fusion methods introduce significant computational overhead and neglect the valuable guidance from contextual samples. To address these limitations, this p…
▽ More
Event-based pedestrian attribute recognition (PAR) leverages motion cues to enhance RGB cameras in low-light and motion-blur scenarios, enabling more accurate inference of attributes like age and emotion. However, existing two-stream multimodal fusion methods introduce significant computational overhead and neglect the valuable guidance from contextual samples. To address these limitations, this paper proposes an Event Prompter. Discarding the computationally expensive auxiliary backbone, this module directly applies extremely lightweight and efficient Discrete Cosine Transform (DCT) and Inverse DCT (IDCT) operations to the event data. This design extracts frequency-domain event features at a minimal computational cost, thereby effectively augmenting the RGB branch. Furthermore, an external memory bank designed to provide rich prior knowledge, combined with modern Hopfield networks, enables associative memory-augmented representation learning. This mechanism effectively mines and leverages global relational knowledge across different samples. Finally, a cross-attention mechanism fuses the RGB and event modalities, followed by feed-forward networks for attribute prediction. Extensive experiments on multiple benchmark datasets fully validate the effectiveness of the proposed RGB-Event PAR framework. The source code of this paper will be released on https://github.com/Event-AHU/OpenPAR
△ Less
Submitted 19 March, 2026;
originally announced March 2026.
-
Multiscale Switch for Semi-Supervised and Contrastive Learning in Medical Ultrasound Image Segmentation
Authors:
Jingguo Qu,
Xinyang Han,
Yao Pu,
Man-Lik Chui,
Simon Takadiyi Gunda,
Ziman Chen,
Jing Qin,
Ann Dorothy King,
Winnie Chiu-Wing Chu,
Jing Cai,
Michael Tin-Cheung Ying
Abstract:
Medical ultrasound image segmentation faces significant challenges due to limited labeled data and characteristic imaging artifacts including speckle noise and low-contrast boundaries. While semi-supervised learning (SSL) approaches have emerged to address data scarcity, existing methods suffer from suboptimal unlabeled data utilization and lack robust feature representation mechanisms. In this pa…
▽ More
Medical ultrasound image segmentation faces significant challenges due to limited labeled data and characteristic imaging artifacts including speckle noise and low-contrast boundaries. While semi-supervised learning (SSL) approaches have emerged to address data scarcity, existing methods suffer from suboptimal unlabeled data utilization and lack robust feature representation mechanisms. In this paper, we propose Switch, a novel SSL framework with two key innovations: (1) Multiscale Switch (MSS) strategy that employs hierarchical patch mixing to achieve uniform spatial coverage; (2) Frequency Domain Switch (FDS) with contrastive learning that performs amplitude switching in Fourier space for robust feature representations. Our framework integrates these components within a teacher-student architecture to effectively leverage both labeled and unlabeled data. Comprehensive evaluation across six diverse ultrasound datasets (lymph nodes, breast lesions, thyroid nodules, and prostate) demonstrates consistent superiority over state-of-the-art methods. At 5\% labeling ratio, Switch achieves remarkable improvements: 80.04\% Dice on LN-INT, 85.52\% Dice on DDTI, and 83.48\% Dice on Prostate datasets, with our semi-supervised approach even exceeding fully supervised baselines. The method maintains parameter efficiency (1.8M parameters) while delivering superior performance, validating its effectiveness for resource-constrained medical imaging applications. The source code is publicly available at https://github.com/jinggqu/Switch
△ Less
Submitted 19 March, 2026;
originally announced March 2026.
-
Pressure-induced Superconductivity in AgSbTe2
Authors:
Sudaice Kazibwe,
Bishnu Karki,
Wencheng Lu,
Zhongxin Liang,
Minghong Sui,
Melissa Gooch,
Zhifeng Ren,
Pavan Hosur,
Timothy A. Strobel,
Ching-Wu Chu,
Liangzi Deng
Abstract:
AgSbTe2 is a well-known thermoelectric material with a high Seebeck coefficient and intrinsically low thermal conductivity, but its behavior under pressure remains largely unexplored. Here we report a systematic investigation of the structural, electronic, and transport properties of non-stoichiometric AgSbTe2 under high pressure. At ambient pressure, the material can be described as having a cubi…
▽ More
AgSbTe2 is a well-known thermoelectric material with a high Seebeck coefficient and intrinsically low thermal conductivity, but its behavior under pressure remains largely unexplored. Here we report a systematic investigation of the structural, electronic, and transport properties of non-stoichiometric AgSbTe2 under high pressure. At ambient pressure, the material can be described as having a cubic crystal structure that remains stable up to 21.7 GPa beyond which it loses long-range structural order, while its crystal system fully recovers upon decompression. Remarkably, superconductivity emerges at a very low pressure of 0.38 GPa with an onset superconducting critical temperature (Tc) of 3.2 K. Tc increases with increasing pressure, reaching 6.9 K at 31.9 GPa, and peaks at 7.4 K during decompression. Magnetic-field-dependent transport measurements and electronic structure calculations reveal an evolution of the superconducting state driven by an enhanced electronic density of states at the Fermi level under compression. Our findings uncover pressure-induced superconductivity in AgSbTe2 and demonstrate that pressure can effectively tune the electronic ground state of thermoelectric materials, extending their functionality beyond thermoelectric energy conversion.
△ Less
Submitted 18 March, 2026;
originally announced March 2026.
-
When and Why Does Unsupervised RL Succeed in Mathematical Reasoning? A Manifold Envelopment Perspective
Authors:
Zelin Zhang,
Fei Cheng,
Chenhui Chu
Abstract:
Although outcome-based reinforcement learning (RL) significantly advances the mathematical reasoning capabilities of Large Language Models (LLMs), its reliance on computationally expensive ground-truth annotations imposes a severe scalability bottleneck. Unsupervised RL guided by intrinsic rewards offers a scalable alternative, yet it suffers from opaque training dynamics and catastrophic instabil…
▽ More
Although outcome-based reinforcement learning (RL) significantly advances the mathematical reasoning capabilities of Large Language Models (LLMs), its reliance on computationally expensive ground-truth annotations imposes a severe scalability bottleneck. Unsupervised RL guided by intrinsic rewards offers a scalable alternative, yet it suffers from opaque training dynamics and catastrophic instability, such as policy collapse and reward hacking. In this paper, we first design and evaluate a suite of intrinsic rewards that explicitly enforce concise and certain generation. Second, to discover the boundaries of this approach, we test base models across a spectrum of intrinsic reasoning capabilities, revealing how a model's foundational logical prior dictates its success or failure. Finally, to demystify why certain configurations stabilize while others collapse, we introduce a novel geometric diagnostic lens, showing that successful cases are enveloped by manifolds. Ultimately, our work goes beyond merely demonstrating that enforcing concise and certain responses successfully boosts mathematical reasoning; we reveal when this unsupervised approach breaks down and geometrically diagnose why.
△ Less
Submitted 17 March, 2026;
originally announced March 2026.
-
Adaptive Theory of Mind for LLM-based Multi-Agent Coordination
Authors:
Chunjiang Mu,
Ya Zeng,
Qiaosheng Zhang,
Kun Shao,
Chen Chu,
Hao Guo,
Danyang Jia,
Zhen Wang,
Shuyue Hu
Abstract:
Theory of Mind (ToM) refers to the ability to reason about others' mental states, and higher-order ToM involves considering that others also possess their own ToM. Equipping large language model (LLM)-driven agents with ToM has long been considered to improve their coordination in multiagent collaborative tasks. However, we find that misaligned ToM orders-mismatches in the depth of ToM reasoning b…
▽ More
Theory of Mind (ToM) refers to the ability to reason about others' mental states, and higher-order ToM involves considering that others also possess their own ToM. Equipping large language model (LLM)-driven agents with ToM has long been considered to improve their coordination in multiagent collaborative tasks. However, we find that misaligned ToM orders-mismatches in the depth of ToM reasoning between agents-can lead to insufficient or excessive reasoning about others, thereby impairing their coordination. To address this issue, we design an adaptive ToM (A-ToM) agent, which can align in ToM orders with its partner. Based on prior interactions, the agent estimates the partner's likely ToM order and leverages this estimation to predict the partner's action, thereby facilitating behavioral coordination. We conduct empirical evaluations on four multi-agent coordination tasks: a repeated matrix game, two grid navigation tasks and an Overcooked task. The results validate our findings on ToM alignment and demonstrate the effectiveness of our A-ToM agent. Furthermore, we discuss the generalizability of our A-ToM to non-LLM-based agents, as well as what would diminish the importance of ToM alignment.
△ Less
Submitted 17 March, 2026;
originally announced March 2026.
-
Ambient-pressure 151-K superconductivity in HgBa2Ca2Cu3O8+δ via pressure quench
Authors:
Liangzi Deng,
Thacien Habamahoro,
Artin Safezoddeh,
Bishnu Karki,
Sudaice Kazibwe,
Daniel J. Schulze,
Zheng Wu,
Matthew Julian,
Rohit P. Prasankumar,
Hua Zhou,
Jesse S. Smith,
Pavan R. Hosur,
Ching-Wu Chu
Abstract:
Superconductivity has been a vigorously researched topic since its discovery in 1911. Raising the superconducting transition temperature (Tc) has been the main driving force behind such long-sustained efforts due to its potential for impacting humanity and the fundamental knowledge gained from understanding this macroscopic coherent quantum state at high temperatures. The successful development of…
▽ More
Superconductivity has been a vigorously researched topic since its discovery in 1911. Raising the superconducting transition temperature (Tc) has been the main driving force behind such long-sustained efforts due to its potential for impacting humanity and the fundamental knowledge gained from understanding this macroscopic coherent quantum state at high temperatures. The successful development of high-Tc superconductivity will make possible extraordinarily efficient generation, delivery, and utilization of energy, and could also enable the development of controlled fusion while impacting other burgeoning fields like quantum computation and quantum electronics. However, progress has been hindered by a longstanding plateau in the record ambient-pressure Tc, unchanged since 1993. Subsequent significant advancements in Tc have been achieved only under high pressures, preventing the realization of superconductivity's full potential. To directly address this challenge, we developed a pressure-quench protocol (PQP) to stabilize pressure-induced/-enhanced superconducting states at ambient pressure. Here we achieve a record ambient-pressure Tc of 151 K in the cuprate HgBa2Ca2Cu3O8+δ via PQP. The experimental results are further supported by synchrotron X-ray diffraction measurements and phonon and electronic structure calculations. This breakthrough opens new avenues for stabilizing and exploring ambient-pressure high-Tc superconducting states and other quantum states that have been previously only accessible under pressure, paving the way for deeper understanding and practical applications of high-Tc superconductivity and beyond.
△ Less
Submitted 12 March, 2026;
originally announced March 2026.
-
Hawking Radiation from Tunneling in Black Hole Quantum Mechanics
Authors:
Chong-Sun Chu
Abstract:
It was proposed in \cite{Chu:2024qil} that a quantum black hole can be described by a quantum space configuration of a fuzzy sphere together with a half-filled Fermi sea. In this paper we propose that the tunneling of the fuzzy sphere system to a small one describes the quantum decay of black hole by Hawking radiation. Since the Fermi sea shrinks and the quantum mechanical Hamiltonian conserves fe…
▽ More
It was proposed in \cite{Chu:2024qil} that a quantum black hole can be described by a quantum space configuration of a fuzzy sphere together with a half-filled Fermi sea. In this paper we propose that the tunneling of the fuzzy sphere system to a small one describes the quantum decay of black hole by Hawking radiation. Since the Fermi sea shrinks and the quantum mechanical Hamiltonian conserves fermion number, the amplitude of transition naively vanishes unless the tunneling path provides exact number of zero modes to soak up the excess fermi states. We show that a monopole on fuzzy sphere does exactly that. This fixes the tunneling path. The resulting tunneling rate reproduces Page's result for the semi-classical decay rate of black hole. The quantum states released by the monopole corresponds to gravitational Hawking radiation. At the level of probability, the Hawking radiation is found to be given by a Boltzmann distribution at the Hawking temperature. One can go beyond the probabilistic description by determining the full wave function of the multi-partite Hawking quanta. This is possible with a real time formulation of the tunneling process. Unitarity is manifest in our quantum mechanics.
△ Less
Submitted 30 March, 2026; v1 submitted 12 March, 2026;
originally announced March 2026.
-
Gate-tunable anisotropic Josephson diode effect in topological Dirac semimetal Cd$_3$As$_2$ nanowires
Authors:
Yan-Liang Hou,
An-Qi Wang,
Na Li,
Chun-Guang Chu,
Alexander Brinkman,
Zhi-Min Liao,
Chuan Li
Abstract:
The intrinsic Josephson diode effect (JDE) has recently attracted considerable attention due to its sensitivity to broken symmetries in Josephson junctions, offering a powerful probe for uncovering hidden symmetry-breaking mechanisms in materials. The presence of higher-harmonic components in the current-phase relation, together with spin-orbital coupling, makes topological materials ideal platfor…
▽ More
The intrinsic Josephson diode effect (JDE) has recently attracted considerable attention due to its sensitivity to broken symmetries in Josephson junctions, offering a powerful probe for uncovering hidden symmetry-breaking mechanisms in materials. The presence of higher-harmonic components in the current-phase relation, together with spin-orbital coupling, makes topological materials ideal platforms to explore this effect. In this work, we present a systematic study of the JDE in type-I topological Dirac semimetal Cd$_3$As$_2$ nanowire-based Josephson junctions. We observe a pronounced gate-tunable and highly anisotropic diode response under different magnetic-field orientations. By developing a comprehensive phenomenological model, we capture the angular dependence of the diode effect and, through temperature-dependent measurements, disentangle the respective contributions from bulk and topological surface states. Notably, anomalies in the temperature dependence of the diode efficiency reveal the coexistence of multiple transport channels, highlighting the Josephson diode effect as a sensitive probe of hidden topological superconducting states.
△ Less
Submitted 10 March, 2026;
originally announced March 2026.
-
UniPAR: A Unified Framework for Pedestrian Attribute Recognition
Authors:
Minghe Xu,
Rouying Wu,
Jiarui Xu,
Minhao Sun,
Zikang Yan,
Xiao Wang,
ChiaWei Chu,
Yu Li
Abstract:
Pedestrian Attribute Recognition is a foundational computer vision task that provides essential support for downstream applications, including person retrieval in video surveillance and intelligent retail analytics. However, existing research is frequently constrained by the ``one-model-per-dataset" paradigm and struggles to handle significant discrepancies across domains in terms of modalities, a…
▽ More
Pedestrian Attribute Recognition is a foundational computer vision task that provides essential support for downstream applications, including person retrieval in video surveillance and intelligent retail analytics. However, existing research is frequently constrained by the ``one-model-per-dataset" paradigm and struggles to handle significant discrepancies across domains in terms of modalities, attribute definitions, and environmental scenarios. To address these challenges, we propose UniPAR, a unified Transformer-based framework for PAR. By incorporating a unified data scheduling strategy and a dynamic classification head, UniPAR enables a single model to simultaneously process diverse datasets from heterogeneous modalities, including RGB images, video sequences, and event streams. We also introduce an innovative phased fusion encoder that explicitly aligns visual features with textual attribute queries through a late deep fusion strategy. Experimental results on the widely used benchmark datasets, including MSP60K, DukeMTMC, and EventPAR, demonstrate that UniPAR achieves performance comparable to specialized SOTA methods. Furthermore, multi-dataset joint training significantly enhances the model's cross-domain generalization and recognition robustness in extreme environments characterized by low light and motion blur. The source code of this paper will be released on https://github.com/Event-AHU/OpenPAR
△ Less
Submitted 5 March, 2026;
originally announced March 2026.
-
Synthesis and Structural Analysis of an Emissive Colloidal Argyrodite Nanocrystal: Canfieldite Ag8SnS6
Authors:
Francisco Yarur Villanueva,
Victor Quezada Novoa,
Pascal Rusch,
Stefano Toso,
Maxwell W. Terban,
Yurii P. Ivanov,
Joaquin Carlos Chu,
Maxine J. Kirshenbaum,
Ehsan Nikbin,
Maria J. Gendron Romero,
Mirko Prato,
Giorgio Divitini,
Jane Y. Howe,
Mark W. B. Wilson,
Liberato Manna
Abstract:
We resolve a phase identification controversy in the Ag-Sn-S material system by unraveling the polymorphic structure of nanocrystals within the argyrodite material family. Argyrodites are a class of superionic materials used in their bulk form for applications in solid-state batteries and thermoelectrics, where their advantageous properties relate to their polymorphism. However, despite their well…
▽ More
We resolve a phase identification controversy in the Ag-Sn-S material system by unraveling the polymorphic structure of nanocrystals within the argyrodite material family. Argyrodites are a class of superionic materials used in their bulk form for applications in solid-state batteries and thermoelectrics, where their advantageous properties relate to their polymorphism. However, despite their well-studied bulk applications, the limited exploration at the nanoscale has left considerable potential for the discovery of emerging properties due to size effects. Further, phase identification presents a prominent challenge to the study of polymorphs in superionic conductors and related mate-rials. In this work, we synthesize canfieldite-like (Ag8SnS6) nanocrystals to understand their formation and structural behavior at the nanoscale. We observe the emergence of emissive, meta-stable, cluster-like species. Then, high-resolution transmission electron microscopy reveals indistinguishable polymorphs of canfieldite due to identical heavy-atom frameworks. However, using synchrotron X-ray total scattering for pair distribution function analysis, we uncover structural distortions, showing a pseudo-orthorhombic configuration that likely gives rise to the red emission. Further, we investigate the optical properties and structure of Ag8SnS6 nanocrystals upon the addition of Zn2+, the cation of interest in the canfieldite vs. pirquitasite (Ag2ZnSnS4) phase identification controversy. We show that Zn2+ is incorporated in the canfieldite-like structure through the replacement of Ag+, boosting the emission. Our results solve a standing phase identification challenge and uncover fundamental insights for the synthesis and structure of canfieldite nanocrystals, laying the ground for the exploration of other argyrodite materials with emerging properties at the nanoscale.
△ Less
Submitted 27 February, 2026;
originally announced February 2026.
-
High-pressure stabilization of Mg2IrH7: Structural proximity to high-Tc superconductivity
Authors:
Shubham Sinha,
Wencheng Lu,
Mads F. Hansen,
Michael J. Hutcheon,
Trevor W. Bontke,
Lewis J. Conway,
Kapildeb Dolui,
Chris J. Pickard,
Christoph Heil,
Piotr A. Guńka,
Stella Chariton,
Vitali Prakapenka,
Liangzi Deng,
Ching-Wu Chu,
Matthew N. Julian,
Rohit P. Prasankumar,
Timothy A. Strobel
Abstract:
Mg$_2$IrH$_6$ is a metastable complex metal hydride with a predicted superconducting transition temperature as high as 170 K at ambient pressure. Following the synthesis of isomorphic, insulating Mg$_2$IrH$_5$ at low pressure, higher-pressure studies were conducted to investigate the phase behavior and compound formation in this system. X-ray diffraction and Raman spectroscopic measurements indica…
▽ More
Mg$_2$IrH$_6$ is a metastable complex metal hydride with a predicted superconducting transition temperature as high as 170 K at ambient pressure. Following the synthesis of isomorphic, insulating Mg$_2$IrH$_5$ at low pressure, higher-pressure studies were conducted to investigate the phase behavior and compound formation in this system. X-ray diffraction and Raman spectroscopic measurements indicate that cubic Mg$_2$IrH$_7$ is stabilized above ca. 40 GPa and coexists with a related hexagonal hydride with likely composition near Mg$_2$IrH$_5$. Electrical transport measurements show that the cubic Mg$_2$IrH$_7$ is insulating, in agreement with ab initio predictions, and persists during room-temperature decompression until $\sim$20 GPa before reverting back to the cubic Mg$_2$IrH$_5$. The experimental results confirm ground-state structure predictions in the Mg-Ir-H system, and the formation of two nearly identical phases with surrounding compositions opens new opportunities to access superconducting Mg$_2$IrH$_6$ through non-equilibrium processing pathways.
△ Less
Submitted 26 February, 2026;
originally announced February 2026.
-
Diamond-to-graphite transformation under hypersonic impact
Authors:
Abhijit Biswas,
Aniket Mote,
Rajib Sahu,
Marcelo Lopes Pereira Junior,
Shuo Yang,
Sudaice Kazibwe,
Jishnu Murukeshan,
Raphael Benjamim de Oliveira,
Guilherme da Silva Lopes Fabris,
Shreyasi Chattopadhyay,
Gelu Costin,
Jianhua Li,
Robert Vajtai,
Ching-Wu Chu,
Lizhong Lang,
Yu Zou,
Liangzi Deng,
Tobin Filleter,
Douglas Soares Galvão,
Christian Kübel,
Thomas E Lacy Jr,
Pulickel M. Ajayan
Abstract:
Diamond to graphite transformation is a complex kinetically driven process which has been studied under various conditions for its fundamental importance. We report the transformation of diamond embedded ceramic matrix composites during hypersonic impact. Diamond particles embedded in cubic boron nitride matrix provide a superhard composite that was subjected to high impact collisions of metal pro…
▽ More
Diamond to graphite transformation is a complex kinetically driven process which has been studied under various conditions for its fundamental importance. We report the transformation of diamond embedded ceramic matrix composites during hypersonic impact. Diamond particles embedded in cubic boron nitride matrix provide a superhard composite that was subjected to high impact collisions of metal projectiles travelling at speeds reaching Mach 8.45. Our observations suggest that the energy absorption and fracture of the composite is primarily enabled via the phase change of diamond into graphite. Characterization of the impact-fractured composite shows transformed diamond particles and provides details of the shock-induced phase transformation and the nature of diamond-graphite interfaces formed during rapid phase change. The study provides new understanding of phase transformation of diamond under extreme conditions.
△ Less
Submitted 10 August, 2026; v1 submitted 13 February, 2026;
originally announced February 2026.
-
Kelix Technical Report
Authors:
Boyang Ding,
Chenglong Chu,
Dunju Zang,
Han Li,
Jiangxia Cao,
Kun Gai,
Muhao Wei,
Ruiming Tang,
Shiyao Wang,
Siyang Mao,
Xinchen Luo,
Yahui Liu,
Zhixin Ling,
Zhuoran Yang,
Ziming Li,
Chengru Song,
Guorui Zhou,
Guowang Zhang,
Hao Peng,
Hao Wang,
Jiaxin Deng,
Jin Ouyang,
Jinghao Zhang,
Lejian Ren,
Qianqian Wang
, et al. (6 additional authors not shown)
Abstract:
Autoregressive large language models (LLMs) scale well by expressing diverse tasks as sequences of discrete natural-language tokens and training with next-token prediction, which unifies comprehension and generation under self-supervision. Extending this paradigm to multimodal data requires a shared, discrete representation across modalities. However, most vision-language models (VLMs) still rely…
▽ More
Autoregressive large language models (LLMs) scale well by expressing diverse tasks as sequences of discrete natural-language tokens and training with next-token prediction, which unifies comprehension and generation under self-supervision. Extending this paradigm to multimodal data requires a shared, discrete representation across modalities. However, most vision-language models (VLMs) still rely on a hybrid interface: discrete text tokens paired with continuous Vision Transformer (ViT) features. Because supervision is largely text-driven, these models are often biased toward understanding and cannot fully leverage large-scale self-supervised learning on non-text data. Recent work has explored discrete visual tokenization to enable fully autoregressive multimodal modeling, showing promising progress toward unified understanding and generation. Yet existing discrete vision tokens frequently lose information due to limited code capacity, resulting in noticeably weaker understanding than continuous-feature VLMs. We present Kelix, a fully discrete autoregressive unified model that closes the understanding gap between discrete and continuous visual representations.
△ Less
Submitted 12 February, 2026; v1 submitted 10 February, 2026;
originally announced February 2026.
-
Anomaly Induced Current in Boundary Lifshitz Field Theory
Authors:
Chong-Sun Chu,
Himanshu Parihar
Abstract:
We study quantum transport phenomena induced by anisotropic Lifshitz scale anomaly in a boundary Lifshitz field theory (BLFT) coupled to an external electromagnetic background. In this context, we obtain the anisotropic scale anomaly in Lifshitz field theories coupled to a background $U(1)$ gauge field and subsequently compute the anomaly induced near boundary current in a BLFT. Focusing on 5D BLF…
▽ More
We study quantum transport phenomena induced by anisotropic Lifshitz scale anomaly in a boundary Lifshitz field theory (BLFT) coupled to an external electromagnetic background. In this context, we obtain the anisotropic scale anomaly in Lifshitz field theories coupled to a background $U(1)$ gauge field and subsequently compute the anomaly induced near boundary current in a BLFT. Focusing on 5D BLFTs, we find that the temporal and spatial components of the induced current exhibit distinct power law dependencies on the distance from the boundary, reflecting the intrinsic time-space anisotropy of the theory. We further derive this anomalous current holographically from the bulk dual of BLFT and find that the temporal component is independent of the boundary conditions while the spatial component depends explicitly on them. The distance dependence is in exact agreement with the dual field theory result.
△ Less
Submitted 6 February, 2026;
originally announced February 2026.
-
A longitudinal geospatial multimodal dataset of post-discharge frailty, physiology, mobility, and neighborhoods
Authors:
Ali Abedi,
Charlene H. Chu,
Shehroz S. Khan
Abstract:
Frailty in older adults is associated with increased vulnerability to functional decline, reduced mobility, social isolation, and challenges during the transition from hospital to community living. These factors are associated with rehospitalization and may adversely influence recovery. Neighborhood environments can further shape recovery trajectories by affecting mobility opportunities, social en…
▽ More
Frailty in older adults is associated with increased vulnerability to functional decline, reduced mobility, social isolation, and challenges during the transition from hospital to community living. These factors are associated with rehospitalization and may adversely influence recovery. Neighborhood environments can further shape recovery trajectories by affecting mobility opportunities, social engagement, and access to community resources. Multimodal sensing technologies combined with data-driven analytical approaches offer the potential to continuously monitor these multidimensional factors in real-world settings. This Data Descriptor presents GEOFRAIL, a longitudinal geospatial multimodal dataset collected from community-dwelling frail older adults following hospital discharge. The dataset is organized into interconnected tables capturing participant demographics, features derived from multimodal sensors, biweekly clinical assessments of frailty, physical function, and social isolation, and temporal location records linked to neighborhood amenities, crime rates, and census-based socioeconomic indicators. Data were collected over an eight-week post-discharge period using standardized pipelines with privacy-preserving spatial aggregation. Technical validation demonstrates internal consistency across geospatial, sensor-derived, and clinical measures and reports baseline performance of machine learning models for characterizing recovery trajectories.
△ Less
Submitted 20 January, 2026;
originally announced February 2026.
-
Efficient learning of logical noise from syndrome data
Authors:
Han Zheng,
Chia-Tung Chu,
Senrui Chen,
Argyris Giannisis Manes,
Su-un Lee,
Sisi Zhou,
Liang Jiang
Abstract:
Characterizing errors in quantum circuits is essential for device calibration, yet detecting rare error events requires a large number of samples. This challenge is particularly severe in calibrating fault-tolerant, error-corrected circuits, where logical error probabilities are suppressed to higher order relative to physical noise and are therefore difficult to calibrate through direct logical me…
▽ More
Characterizing errors in quantum circuits is essential for device calibration, yet detecting rare error events requires a large number of samples. This challenge is particularly severe in calibrating fault-tolerant, error-corrected circuits, where logical error probabilities are suppressed to higher order relative to physical noise and are therefore difficult to calibrate through direct logical measurements. Recently, Wagner et al. [PRL 130, 200601 (2023)] showed that, for phenomenological Pauli noise models, the logical channel can instead be inferred from syndrome measurement data generated during error correction. Here, we extend this framework to realistic circuit-level noise models. From a unified code-theoretic perspective and spacetime code formalism, we derive necessary and sufficient conditions for learning the logical channel from syndrome data alone and explicitly characterize the learnable degrees of freedom of circuit-level Pauli faults. Using Fourier analysis and compressed sensing, we develop efficient estimators with provable guarantees on sample complexity and computational cost. We further present an end-to-end protocol and demonstrate its performance on several syndrome-extraction circuits, achieving orders-of-magnitude sample-complexity savings over direct logical benchmarking. Our results establish syndrome-based learning as a practical approach to characterizing the logical channel in fault-tolerant quantum devices.
△ Less
Submitted 10 February, 2026; v1 submitted 29 January, 2026;
originally announced January 2026.