-
When Readability and Source Retention Diverge: An Evaluability Gap in AI Translation
Authors:
Chenchen Mao,
Hanjing Shi,
Haiyan Jia,
Emily Wegrzyn,
Dominic DiFranzo
Abstract:
Readable AI output can leave an evaluability gap: even when the source is shown, an overall-quality judgment may not reflect what an output preserves. We investigated how source-text condition and output rendering relate to perceived translation quality, and how output and system appraisals relate to trust and stated disclosure willingness in a plain-text interface. A focal 2 * 2 comparison (N=306…
▽ More
Readable AI output can leave an evaluability gap: even when the source is shown, an overall-quality judgment may not reflect what an output preserves. We investigated how source-text condition and output rendering relate to perceived translation quality, and how output and system appraisals relate to trust and stated disclosure willingness in a plain-text interface. A focal 2 * 2 comparison (N=306) using TransLingo examined simple generated narratives and complex literary-philosophical prose alongside LLM-generated readability-oriented outputs and researcher-revised fidelity-oriented outputs. A descriptive stimulus audit indicated greater source retention in fidelity-oriented outputs in both source-text conditions. Factorial analyses showed a significant rendering-by-source-text-condition interaction in perceived quality. Participants rated fidelity-oriented outputs higher than readability-oriented outputs for the simple narratives, whereas no reliable rendering difference emerged for the complex prose. A corresponding source-condition-dependent pattern was observed for perceived intelligence, agency-oriented anthropomorphic attribution, and task-performance trust. A separate theory-ordered appraisal-structure SEM characterized concurrent associations among perceived quality, perceived intelligence, agency-oriented anthropomorphic attribution, task-performance trust, and stated disclosure willingness across six domains, with task-performance trust as the proximal correlate of stated willingness. The observed rating pattern distinguishes source access from source evaluability: for the complex stimuli, displaying the source did not ensure that one overall-quality rating reflected differences in retained content. It also separates support for evaluating translation output from data-handling support for decisions about what personal text to entrust to a system.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
Execution-grounded evaluation reveals hidden failures in language-model calculations for environmental science
Authors:
Maohao Ran,
Chendong Ma,
Yanting Zhang,
Dailing Jiang,
Yusen Huang,
Meng Gao,
Jun Song
Abstract:
Large language models are increasingly used for quantitative work in the environmental sciences, yet existing evaluations score only final answers, leaving calculation process unobserved. Here we introduce AtmosCoder-Bench, an execution-grounded benchmark that makes the calculation process visible. Built through a transferable semi-automated pipeline (436 problems, 3,910 variants, 7,029 graded qua…
▽ More
Large language models are increasingly used for quantitative work in the environmental sciences, yet existing evaluations score only final answers, leaving calculation process unobserved. Here we introduce AtmosCoder-Bench, an execution-grounded benchmark that makes the calculation process visible. Built through a transferable semi-automated pipeline (436 problems, 3,910 variants, 7,029 graded quantities), every problem is validated to be unambiguous and human-solvable, with uniquely verifiable answers. We find that (i) multiple-choice formats inflate measured accuracy by at least 12 percentage points; (ii) many failures arise not from missing knowledge but from models failing to apply known formulas and constraints consistently throughout multi-step computation; and (iii) even frontier models remain weak when task-specific conditions invalidate familiar methods, often reverting to canonical solution patterns rather than adapting methods to the relevant physical regime, leaving expert oversight essential.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
Environment-Invariant Subspace Learning for Generalizable Deepfake Detection
Authors:
Shenghao Chen,
Hao Jia,
Chen Li,
Chunjie Ma,
Zan Gao,
Shengyong Chen
Abstract:
Cross-distribution generalization remains a critical bottleneck in deepfake detection. While recent efforts leverage the semantic priors of large-scale visual foundation models (VFMs), a noteworthy yet underexplored challenge remains: the susceptibility of these semantic priors to environmental interference from factors such as lighting and style. Crucially, this interference establishes spurious…
▽ More
Cross-distribution generalization remains a critical bottleneck in deepfake detection. While recent efforts leverage the semantic priors of large-scale visual foundation models (VFMs), a noteworthy yet underexplored challenge remains: the susceptibility of these semantic priors to environmental interference from factors such as lighting and style. Crucially, this interference establishes spurious correlations between forgery cues and environmental patterns that severely limit generalization. To address this fundamental challenge, we propose an innovative Environment-Invariant Subspace Learning (EISL) framework. The core contribution of EISL is that it aims to disentangle features into orthogonal forgery-relevant invariant factors and environment-related residual factors via a learnable low-rank projection. To facilitate robust feature disentanglement, we also design an Environmental Intervention module that generates diverse and challenging intervention pairs, simulating out-of-distribution environmental shifts to guide the model toward discovering truly invariant forgery representations. Experiments across cross-dataset, cross-generator, whole-face synthesis, and corruption settings show consistent gains and competitive or leading performance against strong detectors, demonstrating improved robustness to unseen forgery types and environmental variations. This work provides a new perspective and a valuable exploration for understanding and tackling the generalization barriers of VFMs in deepfake detection.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation
Authors:
Haoran Wang,
Chaofan Ma,
Ran Yi,
Lizhuang Ma
Abstract:
Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., "subject composition"), which are ill-suited to this combinatorial setting and lead to fragmented coverage, uncontrolled complexity, and little diagnostic value. Recognizing that diverse multi-reference tasks share a common set of atomi…
▽ More
Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., "subject composition"), which are ill-suited to this combinatorial setting and lead to fragmented coverage, uncontrolled complexity, and little diagnostic value. Recognizing that diverse multi-reference tasks share a common set of atomic operations, we adopt a capability-oriented perspective and formalize four operators: Anchor ($f$), Disentangle ($g$), Apply ($\oplus$), and Compose ($C$). Any multi-reference prompt can then be represented as a compositional formula over these operators, whose structural complexity is quantified by the number of operator slots. Building on this formulation, we construct TRACE-Bench, comprising approximately 1,600 evaluation cases across slot counts 1--8, built from 631 formula templates and around 4,000 reference images spanning diverse artistic styles and real-world subjects. The formula structure directly drives an operator-aligned evaluation protocol for per-capability scoring and a diagnostic tree analysis for recursive failure localization. Evaluating 9 leading models reveals insights invisible to holistic scoring: the primary bottleneck lies in disentanglement ($g$) and attribute binding ($\oplus$) rather than scene-level composition ($C$), with even the best model scoring only 0.74 on attribute fidelity. Project page: https://amuseum-whr.github.io/TraceBench
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
FTA-Mem: Fact-Time-Affect Anchored Memory for Low-Density Long-Term Dialogue
Authors:
Chang Liu,
Shuyi Zhang,
Changsheng Ma,
Yongfeng Tao,
Minqiang Yang,
Bin Hu
Abstract:
Long-term emotional-support agents require memory mechanisms for personalized understanding across sessions. However, emotional-support dialogue is often low-density: turns are incomplete, evidence is scattered, and user states evolve over time. Existing memory methods usually rely on fixed units, such as turn-level notes or session summaries, which may lose details or introduce redundant noise. W…
▽ More
Long-term emotional-support agents require memory mechanisms for personalized understanding across sessions. However, emotional-support dialogue is often low-density: turns are incomplete, evidence is scattered, and user states evolve over time. Existing memory methods usually rely on fixed units, such as turn-level notes or session summaries, which may lose details or introduce redundant noise. We propose FTA-Mem, a structured memory framework for low-density long-term dialogue. FTA-Mem uses Boundary-preserving Window Segmentation (BWS) to form coherent situation fragments, and constructs Fact-Time-Affect Memory Units (FTA Units) that jointly encode factual content, temporal grounding, and affective context. Retrieved units are then synthesized into structured context for answer generation. Experiments on ES-MemEval and LoCoMo show that FTA-Mem improves overall long-term memory question answering across benchmarks with different information-density characteristics. On ES-MemEval, FTA-Mem achieves 0.3871 F1 and 0.6668 BERTScore. Further analysis shows that situation-level FTA construction better balances evidence preservation and construction cost than coarse session-level or overly fine-grained turn-pair construction, providing an effective granularity trade-off for long-term dialogue memory.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
First measurements of the branching fractions of $J/ψ$ and $ψ(3686) \to Σ^{0} \barΣ^{0}η$
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko
, et al. (750 additional authors not shown)
Abstract:
Based on $(10087 \pm 44) \times 10^6$ $J/ψ$ and $(2712 \pm 14) \times 10^6$ $ψ(3686)$ events collected with the BESIII detector at the BEPCII collider, the hadronic decays $J/ψ\to Σ^{0} \barΣ^{0} η$ and $ψ(3686) \to Σ^{0} \barΣ^{0} η$ are observed for the first time. The corresponding branching fractions are measured to be…
▽ More
Based on $(10087 \pm 44) \times 10^6$ $J/ψ$ and $(2712 \pm 14) \times 10^6$ $ψ(3686)$ events collected with the BESIII detector at the BEPCII collider, the hadronic decays $J/ψ\to Σ^{0} \barΣ^{0} η$ and $ψ(3686) \to Σ^{0} \barΣ^{0} η$ are observed for the first time. The corresponding branching fractions are measured to be $\mathcal{B}(J/ψ\to Σ^{0} \barΣ^{0}η)= (7.5 \pm 0.3 \pm 0.8) \times 10^{-5}$ and $\mathcal{B}(ψ(3686) \to Σ^{0} \barΣ^{0}η)= (1.3\pm 0.1 \pm 0.1) \times 10^{-5}$, respectively, where the first uncertainties are statistical, and the second systematic. The ratio $\text{Q} \approx \frac{\mathcal{B}(ψ(3686) \to Σ^{0} \barΣ^{0} η)}{\mathcal{B}(J/ψ\to Σ^{0} \barΣ^{0} η)}$ is determined to be $(17.3 \pm 1.5 \pm 1.7)\%$, which is con sistent with the 12\%-rule within 3.0$σ$.~No significant intermediate states or threshold enhancements are observed in the $Σ^0$($\barΣ^{0}$)$η$ and $Σ^0$$\barΣ^{0}$ invariant mass spectra.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Measurement of Branching Fraction and Transition Magnetic Moment of the Hyperon Dalitz Decay $Σ^0 \rightarrow Λe^+e^-$
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
R. Aliberti,
A. Amoroso,
Q. An,
Y. Bai,
O. Bakina,
Y. Ban,
H. -R. Bao,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko,
R. A. Briere,
A. Brueggemann,
H. Cai
, et al. (683 additional authors not shown)
Abstract:
Based on a data sample of 10 billion $J/ψ$ events collected with the BESIII detector operating at the BEPCII collider, the Dalitz decay $Σ^0 \rightarrow Λe^+e^-$ is studied experimentally for the first time. The $Σ^0$ hyperons are produced through the process $J/ψ\rightarrow Σ^0\barΣ^0$ and analyzed using a double-tag method. The absolute branching fraction is measured to be…
▽ More
Based on a data sample of 10 billion $J/ψ$ events collected with the BESIII detector operating at the BEPCII collider, the Dalitz decay $Σ^0 \rightarrow Λe^+e^-$ is studied experimentally for the first time. The $Σ^0$ hyperons are produced through the process $J/ψ\rightarrow Σ^0\barΣ^0$ and analyzed using a double-tag method. The absolute branching fraction is measured to be $\mathcal{B}(Σ^0 \rightarrow Λe^+e^-) = (6.34 \pm 0.25_{\rm stat.} \pm 0.23_{\rm syst.}) \times 10^{-3}$. This result shows a $2σ$ discrepancy from the theoretical calculation quoted in the PDG, where the uncertainties are statistical and systematic, respectively. In addition to the branching fraction, the transition magnetic moment $μ$ is determined to be $(1.74 \pm 0.03_{\rm stat.} \pm 0.09_{\rm syst.})\,μ_N$, where $μ_N=e/(2m_p)$ represents the nucleon magnetic moment, providing valuable insight into the intrinsic structure of the $Σ^0$ hyperon.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Quantum Information in SYK Model
Authors:
Chen-Te Ma,
Jeff Murugan,
Masaki Tezuka
Abstract:
We investigate the bulk-boundary correspondence in the SYK model from a quantum information perspective. The SYK model describes a system of Majorana fermions with random all-to-all interactions, whose disorder average-typically taken over a Gaussian ensemble-admits a dual description in terms of JT gravity in the large-$N$, low-energy limit. This framework provides a minimal setting for exploring…
▽ More
We investigate the bulk-boundary correspondence in the SYK model from a quantum information perspective. The SYK model describes a system of Majorana fermions with random all-to-all interactions, whose disorder average-typically taken over a Gaussian ensemble-admits a dual description in terms of JT gravity in the large-$N$, low-energy limit. This framework provides a minimal setting for exploring holography and emergent spacetime in nearly AdS$_2$. We probe the holographic principle through diagnostics of quantum chaos and entanglement. In the early-time regime, the SYK model saturates the universal bound on the Lyapunov exponent, signaling maximal chaos consistent with semiclassical black hole dynamics. In the late-time regime, its spectral statistics are governed by random matrix theory, reflecting universal features of strongly chaotic quantum systems. These dynamical properties establish a concrete link between boundary quantum chaos and bulk semiclassical gravity. In parallel, we analyze quantum entanglement and the structure of operator algebras to investigate transitions in the associated von Neumann algebras and their implications for emergent geometry. To explore the robustness of these phenomena, we consider deformations of the SYK model through modified matter couplings and alternative random distributions. Our results clarify how quantum information-theoretic structures encode bulk gravitational dynamics and provide insight into the mechanism of spacetime emergence.
△ Less
Submitted 16 August, 2026;
originally announced August 2026.
-
Ergodic Stochastic Optimal Control Problems
Authors:
Chenglin Ma,
Huaizhong Zhao
Abstract:
In this article, we introduce a novel approach to solving the ergodic stochastic optimal control problem whose dynamics is driven by a controlled stochastic differential equation. The coefficients of this stochastic differential equation are non-autonomous but periodic in time. We first prove that the infinite horizon average stochastic optimal control problem is ergodic, i.e., the value function…
▽ More
In this article, we introduce a novel approach to solving the ergodic stochastic optimal control problem whose dynamics is driven by a controlled stochastic differential equation. The coefficients of this stochastic differential equation are non-autonomous but periodic in time. We first prove that the infinite horizon average stochastic optimal control problem is ergodic, i.e., the value function equals a constant $ρ$ which is independent of the initial conditions. Based on the result of ergodicity, we construct an auxiliary function $w(t,x)$ that is well-defined and periodic in time. We can prove that the pair $(w,ρ)$ satisfies the dynamic programming principle and is a viscosity solution of the associated Hamilton-Jacobi-Bellman equation. Finally, we apply our results to the study of ergodic backward stochastic differential equations. Our method is even new in the homogeneous case.
△ Less
Submitted 15 August, 2026;
originally announced August 2026.
-
An Almost-Covering Threshold for Golomb-Ruler Difference Packings
Authors:
Chaohang Ma,
Xiangjie Yi
Abstract:
For a fixed integer $t\geq 3$, consider families of $t$-mark Golomb rulers whose positive-difference sets are pairwise disjoint and contained in $[1,U]$. Let $P_t(U)$ be the largest number of integers covered by such a family. We determine the threshold for asymptotically complete coverage: \[
P_t(U)=U-o(U) \quad\Longleftrightarrow\quad 3\leq t\leq 5. \] The cases $t=3,4$ follow from the known e…
▽ More
For a fixed integer $t\geq 3$, consider families of $t$-mark Golomb rulers whose positive-difference sets are pairwise disjoint and contained in $[1,U]$. Let $P_t(U)$ be the largest number of integers covered by such a family. We determine the threshold for asymptotically complete coverage: \[
P_t(U)=U-o(U) \quad\Longleftrightarrow\quad 3\leq t\leq 5. \] The cases $t=3,4$ follow from the known existence spectra for perfect difference families. For $t=5$, Wild's product construction, in the form recorded by Mathon and applied to perfect families of orders $121$ and $161$, gives a multiplicative semigroup of exact-covering scales; an elementary density lemma on its logarithms then supplies a scale $(1-o(1))U$ below every sufficiently large $U$.
For the converse, we give a self-contained one-frequency Fourier obstruction. If $x_0\in(π,3π/2)$ is the first positive solution of $\tan x=x$ and \[
γ_0=-\frac{2\sin x_0}{x_0}=0.4344672564\ldots, \] then, for every fixed $t\geq 6$, \[
\liminf_{U\to\infty}\left(1-\frac{P_t(U)}{U}\right)
\geq \frac{(t-1)γ_0-2}{2(t-2)}. \] In particular, the forced gap for six-mark rulers is at least $2.1542035\%$. We also prove a discrete small-difference bound which yields a stronger obstruction for every $t\geq14$ and forces a gap of \[
\frac12-\frac1{\sqrt t}-\frac7{8t}+O(t^{-3/2}) \] as $t\to\infty$.
△ Less
Submitted 13 August, 2026;
originally announced August 2026.
-
Modeling Bond-Dependent Kitaev-like interaction in 2D Edge-Sharing Tetrahedral Magnets: FeX (X=Te, Se)
Authors:
Mengdong Li,
Can Huangb,
Bingjie Liu,
Zhixin Liu,
Yanfei Pan,
Jiyu Fan,
Chunlan Ma,
Daning Shi,
Yan Zhua
Abstract:
Bond-dependent magnetic interactions, exemplified by the Kitaev model, are known to arise from the interplay between spin-orbit coupling(SOC) and specific coordination geometries, yet their existence has so far been predominantly associated with edge-sharing octahedral systems. Whether analogous interactions survive in edge- sharing tetrahedral environments - relevant to iron-based superconductor…
▽ More
Bond-dependent magnetic interactions, exemplified by the Kitaev model, are known to arise from the interplay between spin-orbit coupling(SOC) and specific coordination geometries, yet their existence has so far been predominantly associated with edge-sharing octahedral systems. Whether analogous interactions survive in edge- sharing tetrahedral environments - relevant to iron-based superconductor parent compounds - remains an open question. Here, we construct a Kitaev-like model for monolayer FeTe and FeSe and demonstrate the presence of a previously unrecognized bond-dependent Ising-type interaction, induced jointly by chalcogen-mediated SOC and the tetrahedral crystal-field geometry. By mapping the first-principles calculations derived magnetic anisotropy energy mapping across representative linear magnetic orders, we disentangle the bond-dependent contributions from single-ion anisotropy. We reveal that the Kitaev-like interaction dominates the magnetic anisotropy in FeTe, whereas in FeSe, it fiercely competes with a single-ion anisotropy of opposite sign. The resulting noncollinear local anisotropy axes generate intrinsic single-site spin frustration, providing a microscopic mechanism for magnetic disorder beyond isotropic exchange models. Our results establish edge-sharing tetrahedral magnets as a new setting for bond-dependent interactions and extend the scope of Kitaev physics beyond octahedral coordination.
△ Less
Submitted 13 August, 2026;
originally announced August 2026.
-
H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
Authors:
Dingyi Rong,
Yue Shi,
Chaofan Ma,
Jiezhang Cao,
Zongrui Wang,
Zeyu Zhang,
Yao Mu,
Guangtao Zhai,
Ning Liu
Abstract:
Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich behavioral experiences, but transferring them across embodiments remains challenging due to differences between human hands and robotic end-effectors. Recent advances in video world models…
▽ More
Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich behavioral experiences, but transferring them across embodiments remains challenging due to differences between human hands and robotic end-effectors. Recent advances in video world models offer a promising pathway to synthesize robot-centric manipulation videos from human observations, while their cross-embodiment transfer capability remains largely unexplored. Therefore, we introduce H2R-Bench, a benchmark for evaluating cross-embodiment human-to-robot manipulation video generation, where models transform egocentric human demonstrations into robot manipulation videos under specified embodiments. Each benchmark instance contains a human demonstration video, target embodiment constraints, and source-grounded annotations covering task goals, action events, functional contacts, and object responses. H2R-Bench evaluates generated videos through five dimensions, including goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality. We benchmark eleven state-of-the-art video generation models across six manipulation families and two robot embodiments. Our evaluation reveals that current video world models remain limited in human-to-robot manipulation transfer: even leading models often fail in embodiment consistency, functional interaction, and task execution. H2R-Bench provides a systematic diagnostic framework for evaluating whether video world models can bridge the human-to-robot embodiment gap and convert human manipulation observations into robot-centric training resources.
△ Less
Submitted 13 August, 2026;
originally announced August 2026.
-
High-precision measurement of the space-like $η^\prime$ transition form factor
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
M. S. Anderson,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone
, et al. (758 additional authors not shown)
Abstract:
Using a data sample corresponding to an integrated luminosity of $20.3\ \text{fb}^{-1}$, collected with the BESIII detector at a center-of-mass energy of $3.773\ \text{GeV}$ at the BEPCII collider, we report a precision measurement of the product $Q^2|F(Q^2)|$, where $F(Q^2)$ is the single-virtual space-like transition form factor of the $η'$ meson and $Q^2$ is the squared momentum transfer of the…
▽ More
Using a data sample corresponding to an integrated luminosity of $20.3\ \text{fb}^{-1}$, collected with the BESIII detector at a center-of-mass energy of $3.773\ \text{GeV}$ at the BEPCII collider, we report a precision measurement of the product $Q^2|F(Q^2)|$, where $F(Q^2)$ is the single-virtual space-like transition form factor of the $η'$ meson and $Q^2$ is the squared momentum transfer of the tagged virtual photon. The transition form factor is extracted from the differential Born cross section of the two-photon fusion processes $e^+e^- \to e^+e^-γγ^* \to e^+e^-η^\prime$ using a single-tag technique, where only one scattered lepton is detected. The measurement covers $Q^2 \in [0.1, 6.0]$ GeV$^2$, achieving unprecedented precision, better than $3.0\%$ for $Q^2 < 1.5$ GeV$^2$, and providing the first direct determination at $Q^2 < 0.3$ GeV$^2$.
△ Less
Submitted 12 August, 2026;
originally announced August 2026.
-
SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning
Authors:
Caoyuan Ma,
Wenpu Liu,
Weichu Xie,
Tian Gu,
Shilei Zhao,
Lingxi Min,
Shuai Dong,
Yuqi Xu,
Ji Zhao,
Ziyue Wang,
Wenzheng Chang,
Taiqiang Wu,
Yongfu Zhu,
Wenqi Shao,
Yinqiang Zheng
Abstract:
Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their language backbones. We propose SafeCap, a reinforcement-learning framework that aligns LVLMs through learned self-captioning. SafeCap trains a policy model to first generate a safety-relevant image caption and then produce a final answer; the captio…
▽ More
Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their language backbones. We propose SafeCap, a reinforcement-learning framework that aligns LVLMs through learned self-captioning. SafeCap trains a policy model to first generate a safety-relevant image caption and then produce a final answer; the caption is further optimized by whether it enables a frozen LLM to reach a safety-aligned decision. This caption-mediated objective encourages the policy to expose visual cues relevant to safe response generation rather than relying solely on direct refusal supervision. Across five multimodal safety benchmarks and six vision-utility benchmarks, SafeCap substantially improves aggregate safety performance under its intended DirectCap protocol, with gains of 3.7-19.0 points in safety average across four model settings while maintaining comparable or improved vision utility. Under controlled comparisons on matched backbones and data, SafeCap outperforms safety SFT, DPO, and SafeGRPO, demonstrating the effectiveness of caption-mediated reinforcement learning for multimodal safety alignment.
△ Less
Submitted 11 August, 2026;
originally announced August 2026.
-
FormStruct-Bench:A Hierarchical and Diagnostic Benchmark for Table-Form Document Structure Recognition
Authors:
Lujie Ban,
Jiangtao Zhu,
Yuanheng Yu,
Jiasheng Shi,
Chenhao Ma
Abstract:
Transforming table-form documents into machine-processable records requires recovering not only their visible content but also the multilevel structure that organizes it. However, existing benchmarks evaluate either holistic document outputs or conventional table grids, and their aggregate scores provide little insight into where structural failures occur. We introduce FormStruct-Bench, a hierarch…
▽ More
Transforming table-form documents into machine-processable records requires recovering not only their visible content but also the multilevel structure that organizes it. However, existing benchmarks evaluate either holistic document outputs or conventional table grids, and their aggregate scores provide little insight into where structural failures occur. We introduce FormStruct-Bench, a hierarchical and diagnostic benchmark that evaluates table-form document structure recognition at both the document level and progressively finer component levels, allowing aggregate performance to be traced back to specific structural failure modes. To construct auditable ground truth at scale, we annotate 70 reusable templates and expand them into 7,000 verified instances through a provenance-preserving Director--Artist--Verifier pipeline; all 1,100 instances in the template-disjoint test set additionally receive human review. Our evaluation protocol uses five primary metrics and three structure-specific diagnostics across page, schema, and component levels, together with slices over difficulty, structural constraints, and visual degradation. Across 14 API-hosted and locally deployable systems plus two SFT variants, the best document-level score reaches 83.85%, whereas the best reported fine-grained structural score remains below 18%. These results reveal a pronounced gap between reading document content and recovering the hierarchy and regional organization required for reliable table-form understanding.
△ Less
Submitted 10 August, 2026;
originally announced August 2026.
-
Outer Limits: An Experimental Approach to Controlled Content Manipulation within the Reddit Interface
Authors:
Chenchen Mao,
Hanjing Shi,
Haiyan Jia,
Daniel Unhuryan,
Eric Baumer,
Dominic DiFranzo
Abstract:
Independent researchers often lack access to intervention capabilities for controlled experiments on live social media platforms. We present Outer Limits, a browser-based system for controlled content experiments within the existing Old Reddit interface, rather than in a reconstructed simulation. The system renders content locally, records study events, and contains configured voting and commentin…
▽ More
Independent researchers often lack access to intervention capabilities for controlled experiments on live social media platforms. We present Outer Limits, a browser-based system for controlled content experiments within the existing Old Reddit interface, rather than in a reconstructed simulation. The system renders content locally, records study events, and contains configured voting and commenting actions so that neither constructed content nor experimental write interactions reach Reddit. In a 219-participant perceptual-fidelity study, ART ANOVAs found no significant Post Type, Participant Awareness, or interaction effects. Exploratory TOSTs met the d = plus-minus 0.50 equivalence criterion for the marginal contrasts and for Post Type within the forewarned subgroup. We also illustrate the system with a factorial study varying post frame, comment frame, and comment stance. Outer Limits combines three properties that the approaches considered here provide separately: precise control over experimental content, an existing platform interface, and containment of experimental content and interactions from the host community.
△ Less
Submitted 10 August, 2026;
originally announced August 2026.
-
Realization Variance of Gravitational Wave Background Anisotropies from Shot Noise for Pulsar Timing Arrays
Authors:
Meng-Xiang Lin,
Adam Lidz,
Chung-Pei Ma
Abstract:
Shot-noise anisotropies in the nHz gravitational wave background (GWB) are a promising target for pulsar timing arrays (PTAs). If the nHz GWB is sourced by merging supermassive black hole binaries (SMBHBs), as current evidence suggests, the shot-noise signal is expected to be large, potentially of order unity at observing frequencies of $f \sim 1 \, \mathrm{yr}^{-1}$. In this regime, the signal is…
▽ More
Shot-noise anisotropies in the nHz gravitational wave background (GWB) are a promising target for pulsar timing arrays (PTAs). If the nHz GWB is sourced by merging supermassive black hole binaries (SMBHBs), as current evidence suggests, the shot-noise signal is expected to be large, potentially of order unity at observing frequencies of $f \sim 1 \, \mathrm{yr}^{-1}$. In this regime, the signal is dominated by rare bright binaries, and Poisson fluctuations in the discrete SMBHB population produce significant spatial anisotropies. Here, we use Monte Carlo simulations to model the realization-to-realization scatter in the shot-noise, sampling from empirically calibrated models of the SMBHB source populations. We find that the probability distribution of shot-noise amplitudes is broad, spanning a factor of $\sim 50$ (95\% interval) at fixed frequency, with a long tail towards high amplitudes. The most probable and median amplitudes lie significantly below the ensemble means by factors of $\sim 2-3$, implying that the shot-noise in typical realizations is smaller than the mean. The ensemble-averaged shot-noise also differs from simple estimates based on moments of the strain, $\langle h^4 \rangle/\langle h^2 \rangle^2$, because the average of a ratio is not equal to the ratio of the averages (i.e., $\langle X/Y \rangle \ne \langle X \rangle/\langle Y \rangle$). This difference is a factor of $\sim 3$ at $f = 0.1 \, \rm{yr}^{-1}$, growing to larger than two orders of magnitude by $f \sim 1 \, \rm{yr}^{-1}$, where the GWB is dominated by low abundance, high-strain sources. Shot-noise nevertheless provides a powerful diagnostic for understanding the GWB and SMBHB populations; interpreting PTA measurements, however, requires modeling its full probability distribution.
△ Less
Submitted 10 August, 2026;
originally announced August 2026.
-
Pulsar Timing Array Sensitivity to Anisotropy: Empirical Sensitivity Curves, Scaling Relations, and the Multi-Resolution Pixel Basis
Authors:
Taha T. Moursy,
Nihan S. Pol,
Gabriella Agazie,
Nikita Agarwal,
Akash Anumarlapudi,
Anne M. Archibald,
Zaven Arzoumanian,
Anjana Ashok,
Jeremy G. Baier,
Paul T. Baker,
Bence Bécsy,
Laura Blecha,
Adam Brazier,
Paul R. Brook,
Sarah Burke-Spolaor,
Rand Burnette,
Robin Case,
J. Andrew Casey-Clyde,
Maria Charisi,
Shami Chatterjee,
Tyler Cohen,
James M. Cordes,
Neil J. Cornish,
Fronefield Crawford,
H. Thankful Cromartie
, et al. (93 additional authors not shown)
Abstract:
We quantify pulsar timing array (PTA) sensitivity to anisotropy in the gravitational wave background using the cross-correlation based Fisher information matrix in the pixel and spherical harmonic bases. We use a set of simulations to empirically determine scaling relations of a PTA's sensitivity to anisotropy with the number of pulsars $N_\mathrm{psr}$ in the array, the error $δt$ on the times of…
▽ More
We quantify pulsar timing array (PTA) sensitivity to anisotropy in the gravitational wave background using the cross-correlation based Fisher information matrix in the pixel and spherical harmonic bases. We use a set of simulations to empirically determine scaling relations of a PTA's sensitivity to anisotropy with the number of pulsars $N_\mathrm{psr}$ in the array, the error $δt$ on the times of arrival, the frequency $f_\mathrm{GW}$ of the gravitational waves, and the angular scale $ΔΩ$ of the anisotropy. The sensitivity scales approximately as $N_\mathrm{psr}^{0.8}$, $δt^{-0.08}$, and $ΔΩ^{1.6}-ΔΩ^{2.1}$ (depending on the ranges of $\ell$ and $m$ under consideration). In addition, we use realistic simulations to project the NANOGrav PTA sensitivity to a 30-year baseline and quantify the growth in sensitivity at several timeslices. Except at the lowest frequencies, we find negligible effect on sensitivity through increasing the observation duration only. Finally, we introduce a multi-resolution pixel basis motivated by the large dependence of the sensitivity on sky location, and demonstrate the operation of the basis through a set of injections and recoveries.
△ Less
Submitted 10 August, 2026;
originally announced August 2026.
-
PressureMesh: 3D Human Mesh Estimation from Multi-Device Pressure Images
Authors:
Changhai Ma,
Ziyu Wu,
Yunkang Zhang,
Fangting Xie,
Mengting Niu,
Heyu Ding,
Quan Wan,
Jiayue Yuan,
Boyan Liu,
Yi Ke,
Xiaohui Cai
Abstract:
Human pose monitoring is crucial in fields such as rehabilitation assessment and human-computer interaction. Due to its privacy-preserving nature, pressure-based human pose monitoring has become a primary approach for unobtrusive sensing. However, existing methods are generally limited to a single device, which restricts the effective monitoring range. To address this limitation, we propose MDP-Ne…
▽ More
Human pose monitoring is crucial in fields such as rehabilitation assessment and human-computer interaction. Due to its privacy-preserving nature, pressure-based human pose monitoring has become a primary approach for unobtrusive sensing. However, existing methods are generally limited to a single device, which restricts the effective monitoring range. To address this limitation, we propose MDP-Net, an end-to-end network capable of directly estimating human meshes from temporal pressure data across multiple devices. We introduce a multimodal fusion mechanism inspired by the Mixture of Experts (MoE) framework to achieve effective complementarity and enhancement of cross-device pressure information. To support the training and evaluation of MDP-Net, we constructed MDP, a high-quality multi-device temporal pressure dataset that includes various pose labels such as 2D/3D joints and human meshes. Experimental results demonstrate that MDP-Net achieves a joint position error of 12.6 cm on the MDP dataset. These results prove that fusing multi-device pressure information is an effective and promising new solution for daily human pose monitoring.
△ Less
Submitted 10 August, 2026;
originally announced August 2026.
-
PreGress: Ranking-Native Pre-training and Prompting for Graph Node Ranking
Authors:
Lujie Ban,
Jiasheng shi,
Yingli Zhou,
Kaiwen Xue,
Daiyin Wang,
Xubin Li,
Shuanghua Li,
Chenhao Ma
Abstract:
Node ranking is a fundamental problem in graph information retrieval, measuring the relative importance of nodes and supporting a wide range of applications such as influence analysis, recommendation, and graph-based retrieval augmented generation. However, exact computation of graph-based ranking measures is often computationally prohibitive at scale. Existing GNN-based ranking methods provide sc…
▽ More
Node ranking is a fundamental problem in graph information retrieval, measuring the relative importance of nodes and supporting a wide range of applications such as influence analysis, recommendation, and graph-based retrieval augmented generation. However, exact computation of graph-based ranking measures is often computationally prohibitive at scale. Existing GNN-based ranking methods provide scalable approximations, but they are typically tailored to individual ranking criteria and require retraining for each downstream task, which limits their transferability and efficiency. Recent graph pre-training approaches aim to enable knowledge transfer across tasks, yet their learning objectives are largely misaligned with node ranking, resulting in suboptimal adaptability to ranking-oriented applications. To address these limitations, we propose PreGress, the first ranking-native pre-training and prompting framework for supporting a wide range of node ranking tasks. PreGress performs multi-task pre-training using our carefully designed objectives, including degree centrality prediction and attribute reconstruction, to jointly capture structural and attribute information. To support heterogeneous ranking criteria, we design lightweight, task-specific prompt modules that adapt a frozen ranking backbone to downstream tasks without full retraining. Experiments on six public graphs and two real-world query-to-item benchmarks---Yelp2018 and MovieLens-100K---together with a controlled five-criterion graph-access study demonstrate strong ranking quality with low task-specific state overhead.
△ Less
Submitted 9 August, 2026;
originally announced August 2026.
-
Chemical Potential and Analytic Continuation for Non-Hermitian Lattice Fermions
Authors:
Chen-Te Ma,
Hui Zhang
Abstract:
We introduce a chemical potential for non-Hermitian lattice fermions and show that, for even flavors with degenerate masses and paired chemical potentials $(μ,-μ)$ or $(iμ,iμ)$, the Hybrid Monte Carlo algorithm is free of the sign problem. For one-dimensional free fermions, we demonstrate that the sign problem is a numerical rather than physical obstruction and derive the exact propagator, which i…
▽ More
We introduce a chemical potential for non-Hermitian lattice fermions and show that, for even flavors with degenerate masses and paired chemical potentials $(μ,-μ)$ or $(iμ,iμ)$, the Hybrid Monte Carlo algorithm is free of the sign problem. For one-dimensional free fermions, we demonstrate that the sign problem is a numerical rather than physical obstruction and derive the exact propagator, which is analytic at finite lattice spacing away from its poles but becomes non-analytic in the continuum limit. Finally, we use AI-assisted fitting to perform analytic continuation from imaginary to real chemical potentials.
△ Less
Submitted 9 August, 2026;
originally announced August 2026.
-
Search for the charged lepton flavour violating decay $η'\to eμ$
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
M. S. Anderson,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone
, et al. (744 additional authors not shown)
Abstract:
Based on $(8998\pm40)\times10^6$ $J/ψ$ events collected in $e^+e^-$ collisions at $\sqrt{s} = 3.097$ GeV with the BESIII detector, we present a search for the charged lepton flavour violating decay $η'\to eμ$ with $J/ψ\toγη'$. No significant signal is observed, and an upper limit on its decay branching fraction is set to be $6.3\times10^{-7}$ at the 90% confidence level, improving the previous bes…
▽ More
Based on $(8998\pm40)\times10^6$ $J/ψ$ events collected in $e^+e^-$ collisions at $\sqrt{s} = 3.097$ GeV with the BESIII detector, we present a search for the charged lepton flavour violating decay $η'\to eμ$ with $J/ψ\toγη'$. No significant signal is observed, and an upper limit on its decay branching fraction is set to be $6.3\times10^{-7}$ at the 90% confidence level, improving the previous best result by nearly three orders of magnitude.
△ Less
Submitted 6 August, 2026;
originally announced August 2026.
-
SITA: Semantic Interest Tokens for Target-Aware Compression in Long-Sequence Recommendation
Authors:
Rui Zhou,
Bo Chen,
Qinglin Jia,
Jiezhou Ji,
Chaoyi Ma,
Ruiming Tang,
Hao Wang,
Enhong Chen
Abstract:
As user behavior histories continue to grow on modern Internet platforms, effectively modeling long behavior sequences has become crucial for predicting user interests in candidate items. Existing methods have evolved along two directions. One line dynamically retrieves target-relevant behaviors from long histories, enabling target-aware modeling but requiring target-dependent computation during i…
▽ More
As user behavior histories continue to grow on modern Internet platforms, effectively modeling long behavior sequences has become crucial for predicting user interests in candidate items. Existing methods have evolved along two directions. One line dynamically retrieves target-relevant behaviors from long histories, enabling target-aware modeling but requiring target-dependent computation during inference. The other line compresses entire behavior sequences into compact user representations, achieving high efficiency and scalability but sacrificing target-specific adaptation due to target-independent encoding. The key challenge is therefore to enable target-aware modeling while preserving the efficiency and scalability of compressed user representations. To address this challenge, we propose \textbf{SITA}, a target-aware compression framework for long-sequence recommendation. SITA enables target-aware compression by organizing compressed interests into semantic structures through semantic identifiers learned via parallel semantic quantization. Conditioned on the semantic identifier of the target item, SITA adaptively aggregates the corresponding structured interests to construct the target-specific user representation. Extensive experiments on public datasets and a large-scale industrial dataset demonstrate that SITA consistently outperforms representative baselines while maintaining strong scalability, highlighting its strong potential for real-world recommender systems.
△ Less
Submitted 4 August, 2026;
originally announced August 2026.
-
Effects of light-cluster degrees of freedom on collective flows in heavy-ion collisions at FOPI energies
Authors:
Xin Li,
Si-Pei Wang,
Rui Wang,
Zhen Zhang,
Jie Pu,
Chun-Wang Ma,
Lie-Wen Chen
Abstract:
Within a lattice Boltzmann-Uehling-Uhlenbeck transport model coupled to a kinetic approach for light-cluster formation, we investigate the impact of explicit light-cluster degrees of freedom on collective flows in Au+Au collisions at FOPI energies with beam energies $E_{\rm beam}$= $120$--$1500 A$ MeV by using a density-, momentum-, and isospin-dependent N$5$LO Skyrme pseudopotential. We first ben…
▽ More
Within a lattice Boltzmann-Uehling-Uhlenbeck transport model coupled to a kinetic approach for light-cluster formation, we investigate the impact of explicit light-cluster degrees of freedom on collective flows in Au+Au collisions at FOPI energies with beam energies $E_{\rm beam}$= $120$--$1500 A$ MeV by using a density-, momentum-, and isospin-dependent N$5$LO Skyrme pseudopotential. We first benchmark the kinetic approach by comparing the calculated light-cluster yields with FOPI data in central Au+Au collisions. We then analyze the collective flows of protons and light nuclei (deuterons, tritons, $^{3}\mathrm{He}$, and $^{4}\mathrm{He}$) in mid-central collisions. For protons, calculations with and without dynamical light-cluster degrees of freedom are compared to quantify the influence of dynamical cluster formation on proton directed ($v_1$), elliptic ($v_2$), triangular ($v_3$), and quadrangular ($v_4$) flows. We find that the dynamical light-cluster effect appreciably modifies proton $v_1$--$v_4$ flows at $E_{\rm beam}=120$--$150 A$ MeV, remains visible at $E_{\rm beam}=250$--$400 A$ MeV, and gradually weakens at $E_{\rm beam}\gtrsim 600 A$ MeV. For light nuclei, the kinetic approach captures the overall beam-energy dependence of the FOPI flow data, with better agreement for $E_{\rm beam}\geq 400 A$ MeV. We further examine the nucleon-number scaling of $v_2/A$ in both model calculations and experimental data, finding that the kinetic light-cluster formation approach qualitatively reproduces the observed scaling behavior. These results highlight the importance of a dynamical treatment of light-cluster formation for interpreting collective flows in heavy-ion collisions below about $600 A$ MeV, although the clustering effects on proton flows are minor at higher collision energies.
△ Less
Submitted 3 August, 2026;
originally announced August 2026.
-
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction
Authors:
Ziqiang Cui,
Han Shi,
Bowei He,
Yu Pan,
Peiyang Liu,
Shengyin Sun,
Yankai Chen,
Haoli Bai,
Yichun Yin,
Xue Liu,
Chen Ma
Abstract:
Multi-Token Prediction (MTP) has emerged as an effective paradigm that augments a shared Large Language Model backbone with auxiliary heads, training the model to predict several future tokens in parallel to enrich its supervision signal and accelerate inference. However, existing training frameworks adopt a rigid, fixed-length prediction horizon, disregarding the highly non-uniform information de…
▽ More
Multi-Token Prediction (MTP) has emerged as an effective paradigm that augments a shared Large Language Model backbone with auxiliary heads, training the model to predict several future tokens in parallel to enrich its supervision signal and accelerate inference. However, existing training frameworks adopt a rigid, fixed-length prediction horizon, disregarding the highly non-uniform information density of natural language and code. Forcing the auxiliary heads to predict across high-entropy semantic boundaries injects noisy, conflicting training signals; because these heads share the backbone's latent representations, the resulting gradients backpropagate and interfere with the model's core capabilities. We propose AdaMTP, an adaptive training paradigm that dynamically aligns the prediction horizon with the intrinsic predictability of the sequence. At its core, an entropy-based segmentation algorithm leverages the base model to detect sudden surges in uncertainty as semantic boundaries, partitioning sequences into variable-length groups. Each token is assigned an adaptive prediction depth, and a dynamically masked MTP objective suppresses the loss for predictions that cross these boundaries, attenuating the noisy gradients that degrade the backbone. Across mathematical reasoning, code generation, and general benchmarks on three backbones (Llama-3.1-8B, Qwen-2.5-7B, Gemma-3-12B), AdaMTP consistently outperforms standard MTP in both task performance and inference speedup.
△ Less
Submitted 1 August, 2026;
originally announced August 2026.
-
Bridging the English-Arabic Medical Knowledge Gap: Targeted Low-Rank Adaptation via Causal Layer Selection
Authors:
Chaimae Abouzahir,
Musa Khan,
Hala Ali-Hassan,
Congbo Ma,
Khaled Saleh,
Yousra Sadqi,
Jihad Mallat,
Walid Al-Eisawi,
Nizar Habash,
Farah E. Shamout
Abstract:
Large Language Models (LLMs) perform strongly in English medical tasks but degrade substantially in Arabic, a gap widely attributed to limited training data. We systematically investigate this assumption via tuned lens probing and causal activation patching, and find that Arabic medical knowledge is present in intermediate model representations but fails to surface at the output. This mechanistic…
▽ More
Large Language Models (LLMs) perform strongly in English medical tasks but degrade substantially in Arabic, a gap widely attributed to limited training data. We systematically investigate this assumption via tuned lens probing and causal activation patching, and find that Arabic medical knowledge is present in intermediate model representations but fails to surface at the output. This mechanistic insight motivates a targeted adaptation strategy: rather than fine-tuning the full network, we propose Targeted Low-Rank Adaptation (TLoRA), restricted to the layer window where cross-lingual representations diverge, upstream of the output layers where the failure manifests. We evaluate TLoRA on multiple-choice medical QA, where our approach outperforms full-network LoRA, zero-shot, and few-shot baselines. We further evaluate it on short-answer generation and multi-turn clinical dialogue, where it performs competitively without the need for task-specific finetuning. We additionally introduce AraClinicDialog, a clinician-constructed Arabic medical dialogue benchmark in MSA with validated variants across four Arabic dialects. Together, these contributions demonstrate that mechanistic diagnosis can serve as a practical guide for targeted adaptation in underrepresented-language medical LLMs.
△ Less
Submitted 31 July, 2026;
originally announced August 2026.
-
MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation
Authors:
Dawei Wang,
Di Zhao,
Xinyuan Liu,
Marci Chi Ma,
Xiaoyang Liu,
Chengming Zhou,
Gary Ushaw,
Richard Davison
Abstract:
Credit assignment is a fundamental challenge in cooperative multi-agent reinforcement learning, particularly in embodied AI settings characterized by limited and delayed feedback as well as dynamically changing numbers of active agents. We propose MARS-RA, a framework that reformulates credit assignment as a rank aggregation problem using contribution-based pairwise comparisons among agents genera…
▽ More
Credit assignment is a fundamental challenge in cooperative multi-agent reinforcement learning, particularly in embodied AI settings characterized by limited and delayed feedback as well as dynamically changing numbers of active agents. We propose MARS-RA, a framework that reformulates credit assignment as a rank aggregation problem using contribution-based pairwise comparisons among agents generated by large multimodal models. This shift from absolute to relative estimation ensures robustness against noise and dynamic agent participation, converting comparison results into contribution scores for potential-based reward shaping. We provide theoretical justification for the convergence and robustness of the proposed framework, and show that Shapley values can be used as an interpretive reference. Experimental results on challenging tasks of different types indicate that MARS-RA can guide agents toward effective cooperation.
△ Less
Submitted 30 July, 2026;
originally announced July 2026.
-
Semantic-Aligned Structural Abstraction for Multimodal Sentiment Analysis
Authors:
Wei Chen,
Junkai Li,
Tongguan Wang,
Hui Liu,
Feiyue Xue,
Chuanxiang Ma,
Ying Sha
Abstract:
Multimodal Sentiment Analysis (MSA) aims to interpret complex human emotions by integrating natural language with non-verbal modalities. Non-verbal modalities share a structural isomorphism with natural language, as both can be viewed as feature sequences evolving over time. This isomorphism enables the transformation of non-verbal modalities into text-like tokens for unified semantic reasoning. L…
▽ More
Multimodal Sentiment Analysis (MSA) aims to interpret complex human emotions by integrating natural language with non-verbal modalities. Non-verbal modalities share a structural isomorphism with natural language, as both can be viewed as feature sequences evolving over time. This isomorphism enables the transformation of non-verbal modalities into text-like tokens for unified semantic reasoning. Large Language Models (LLMs), designed to understand and generate sequential data, can thus be utilized to interpret complex affective sequences. However, existing LLM-based methods primarily capture low-level superficial features, failing to model affective semantics arising from structural variations and contextual interactions. To address this limitation, we propose \textbf{SentiLLM}, a unified framework that leverages \textit{Semantic-Aligned Structural Abstraction} to distill continuous raw signals into compact, semantically meaningful tokens. Specifically, we introduce a \textit{Dual-Stream Salience-Context Calibration Mechanism}, which disentangles non-verbal feature sequences into a focus stream and an ambient stream. The focus stream captures salient sentiment shifts (e.g., facial expressions) guided by textual priors, while the ambient stream characterizes stable background states. Through calibrating these dynamic sentiment shifts against background states, SentiLLM effectively projects non-verbal modalities into a unified semantic space, making them naturally understandable for LLMs. Serving as a plug-and-play module, SentiLLM significantly improves discriminative performance with only a small number of trainable parameters. Our method achieves superior performance on four datasets, MOSI, MOSEI, CH-SIMS, and CH-SIMS v2, demonstrating the effectiveness of the structural abstraction paradigm in MSA. Our code is available at: \href{https://github.com/especiallyW/SentiLLM}.
△ Less
Submitted 30 July, 2026;
originally announced July 2026.
-
Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility
Authors:
Yansen Zhang,
Yilu Liu,
Tianyu Liu,
Jiamin Chen,
Xiaokun Zhang,
Kai Xie,
Xue Liu,
Yiyan Qi,
Chen Ma
Abstract:
Large language models increasingly support scientific and algorithmic discovery through inference-time search over evaluated candidates. Existing adaptive discovery controllers assign credit based only on score progress, even though prompt length, retries, and guidance calls cause search actions to incur different token costs. We prove that cost-blind credit can forfeit all but a vanishing fractio…
▽ More
Large language models increasingly support scientific and algorithmic discovery through inference-time search over evaluated candidates. Existing adaptive discovery controllers assign credit based only on score progress, even though prompt length, retries, and guidance calls cause search actions to incur different token costs. We prove that cost-blind credit can forfeit all but a vanishing fraction of attainable quality as frontiers multiply and costs diverge. Under a fixed search-side token budget, the controller must decide which frontier is improving and whether its gain justifies the realized cost before the budget is exhausted. We introduce \textbf{CostAda}, a cost-calibrated adaptive controller built around \emph{cost-calibrated frontier utility}. The utility values frontier progress relative to realized action cost and conditions that credit on the remaining budget. CostAda uses this signal to control local exploration intensity, frontier allocation, and budgeted tactic intervention. Cost and remaining budget therefore shape the search rather than serving only as accounting variables or a stopping rule. CostAda reaches the strongest baseline's full-budget quality with at most half the budget on twelve of sixteen benchmark--backbone pairs while achieving the strongest mean final quality on all eight benchmarks under GLM-5 and GPT-5.4.
△ Less
Submitted 5 August, 2026; v1 submitted 29 July, 2026;
originally announced July 2026.
-
When Does Legacy Data Start to Help? Emergent Transfer in Cross-Configuration Robot Learning
Authors:
Tao Wang,
Hudson Hou,
Yingdong Hu,
Yufeng Liu,
Qinghai Li,
Yingjie Jiang,
Yingzhi Wang,
Cheng Ma,
Richard Wang,
Yang Gao
Abstract:
Robotic hardware evolves over time, but demonstration data is often tied to a specific sensor and actuator configuration. This raises a practical and underexplored question: when does legacy data begin to benefit an upgraded robot? We study this question on a wheeled humanoid platform across two hardware generations, where both the camera and gripper are changed while the overall morphology remain…
▽ More
Robotic hardware evolves over time, but demonstration data is often tied to a specific sensor and actuator configuration. This raises a practical and underexplored question: when does legacy data begin to benefit an upgraded robot? We study this question on a wheeled humanoid platform across two hardware generations, where both the camera and gripper are changed while the overall morphology remains fixed. Contrary to the common assumption that more cross-configuration data is always helpful, we observe a grokking-like transition: legacy data remains ineffective until the upgraded configuration acquires a minimum level of task competence, after which co-training gains rise sharply before diminishing near saturation. We hypothesize that this task-dependent transition is governed by a transfer threshold and characterize the resulting three-phase pattern. Across real-robot manipulation tasks, we observe all three phases: no measurable benefit at low competence ($10.0\% \rightarrow 10.0\%$), a sharp gain after crossing the threshold ($23.3\% \rightarrow 86.7\%$ on flower insertion), and diminishing returns at high competence ($85.0\% \rightarrow 93.3\%$ on pen insertion). We provide a theoretical account based on gradient alignment and residual policy uncertainty, and derive a phase-aware rule for deciding when to collect more new-hardware data and when to reuse legacy demonstrations. We further validate this three-phase pattern on a mobile dual-arm watering task, with results consistent with our predictions.
△ Less
Submitted 28 July, 2026;
originally announced July 2026.
-
Observational Evidence for Anisotropic Metal Excess around Galaxies
Authors:
Cheqiu Lyu,
Enci Wang,
Zeyu Chen,
Chengyu Ma,
Yangyao Chen,
Haoran Yu,
Xu Kong
Abstract:
The exchange of matter and energy between galaxies and their surroundings drives the cosmic baryon cycle, yet mapping metal transport remains an observational challenge. While simulations predict that galactic winds escape anisotropically along minor axes, evidence for chemical enrichment in neighboring galaxies is limited. We analyze 1,433 galaxy pairs from the Dark Energy Spectroscopic Instrumen…
▽ More
The exchange of matter and energy between galaxies and their surroundings drives the cosmic baryon cycle, yet mapping metal transport remains an observational challenge. While simulations predict that galactic winds escape anisotropically along minor axes, evidence for chemical enrichment in neighboring galaxies is limited. We analyze 1,433 galaxy pairs from the Dark Energy Spectroscopic Instrument survey and detect a gas-phase metallicity excess of 14.6% $\pm$ 3.7% to 24.2% $\pm$ 2.6% in neighbors aligned with the minor axis of massive primary at projected separations of 15--60 kpc. This signal, qualitatively consistent with IllustrisTNG simulation, varies from a marginal detection (>92% confidence) at 15--30 kpc to a significant signal (>98% confidence) at 30--60 kpc. In this work, we show that this anisotropic metallicity excess is consistent with a scenario of enrichment via galactic outflows, providing empirical constraints on feedback models and complementing other environmental processes.
△ Less
Submitted 27 July, 2026;
originally announced July 2026.
-
Artificial Intelligence and Innovation Ecosystem: Evolutionary Developments, Challenges, and Future Directions
Authors:
Zhimin Zhang,
Chengzhen Ma,
Jia Chai,
Rongxin Zhan,
Huansheng Ning,
Lingfeng Mao,
Dan Zhang,
Suiping Jiang
Abstract:
The development of the Innovative Ecosystem (IE) presents a new paradigm for economic integration, collaborative advancement, and shared achievements. The rise of Artificial Intelligence (AI) has significantly accelerated the global processes of digitization, informatization, and intelligence. Exploring how AI can leverage inherent characteristics to influence the development trajectory of IE is a…
▽ More
The development of the Innovative Ecosystem (IE) presents a new paradigm for economic integration, collaborative advancement, and shared achievements. The rise of Artificial Intelligence (AI) has significantly accelerated the global processes of digitization, informatization, and intelligence. Exploring how AI can leverage inherent characteristics to influence the development trajectory of IE is a topic that warrants further investigation. Given AI's increasing prominence and role within IE, the paper analyzes this new form, examining both AI's unique contributions to IE and its potential challenges. Firstly, the paper synthesizes the conceptual frameworks surrounding IE, decomposing them into manifestations in physical, social, and thinking spaces. Furthermore, the concept of Artificial Intelligence IE (AIIE) is introduced from a spatial perspective, with an exploration of the characteristics AI contributes to IE. Subsequently, the paper employs an evolutionary perspective to analyze the roles provided by AI during different development periods of AIIE. The paper then verifies the feasibility, effectiveness, and rationality of the AIIE's definition and analyzes AIIE development from an evolutionary perspective using enterprise development examples. Finally, acknowledging AI's inherent limitations, the paper examines potential challenges facing AIIE in the future from four perspectives, aiming to identify new research avenues for the further development of AIIE.
△ Less
Submitted 27 July, 2026;
originally announced July 2026.
-
Precision Measurement of Decay Dynamics in $D^{0(+)}\to π^{-(0)}\ell^+ν_\ell$
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
M. S. Anderson,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone
, et al. (752 additional authors not shown)
Abstract:
The branching fractions of $D^0\to π^-e^+ν_e$, $D^0\to π^-μ^+ν_μ$, $D^+\to π^0e^+ν_e$, and $D^+\to π^0μ^+ν_μ$ are precisely measured, using 20.3 fb$^{-1}$ of $e^+e^-$ collision data collected at the center-of-mass energy of 3.773 GeV with the BESIII detector. The ratios of the decay widths between muon and positron channels are examined in full, across several four-momentum transfer ranges of…
▽ More
The branching fractions of $D^0\to π^-e^+ν_e$, $D^0\to π^-μ^+ν_μ$, $D^+\to π^0e^+ν_e$, and $D^+\to π^0μ^+ν_μ$ are precisely measured, using 20.3 fb$^{-1}$ of $e^+e^-$ collision data collected at the center-of-mass energy of 3.773 GeV with the BESIII detector. The ratios of the decay widths between muon and positron channels are examined in full, across several four-momentum transfer ranges of $\ell^+ν_{\ell}$. No lepton flavor universality violation is found in the current data. From a simultaneous fit to the precisely measured partial decay rates and the first measured forward-backward asymmetries of these four decays, the product of the hadronic transition form factor, $f^{D\toπ}_+(0)$, and the modulus of the $c\to d$ quark mixing element, $|V_{cd}|$, is measured with unprecedented precision to be $f^{D\toπ}_+(0)|V_{cd}|=0.1425\pm0.0005_{\rm stat.}\pm0.0003_{\rm syst.}$. Taking the value of $|V_{cd}|$ from the standard model global fit and $f^{D\toπ}_+(0)$ derived by the lattice quantum chromodynamics calculation as input, we obtain $f^{D\toπ}_+(0)=0.1425\pm0.0005_{\rm stat.}\pm0.0003_{\rm syst.}$ and $|V_{cd}|=0.2262\pm0.0008_{\rm stat.}\pm0.0005_{\rm syst.}\pm0.0018_{\rm LQCD.}$, respectively. The precision of each result is a factor of 2-3 better than the previous best measurements. Additionally, the real and imaginary parts of the scalar current contribution in the $c\to d \ell^+ν_{\ell}$ transition are measured for the first time to be Re $(C_S^μ)=$ $0.022 \pm 0.023_{\rm stat.}\pm 0.003_{\rm syst.}$ and $|\mathrm{Im} (C_S^μ)|=0.000 \pm 0.038_{\rm stat.}\pm 0.012_{\rm syst.}$.
△ Less
Submitted 26 July, 2026;
originally announced July 2026.
-
Precision measurements of semleptonic decays $D^0 \to π^-\ell^+ν_\ell$ and $D^+ \to π^0\ell^+ν_\ell$ ($\ell =e,μ$)
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
M. S. Anderson,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone
, et al. (752 additional authors not shown)
Abstract:
The branching fractions of $D^0\to π^-e^+ν_e$, $D^0\to π^-μ^+ν_μ$, $D^+\to π^0e^+ν_e$, and $D^+\to π^0μ^+ν_μ$ are measured to be $(2.950\pm0.017_{\rm stat.}\pm 0.017_{\rm syst.})\times10^{-3}$, $(2.817\pm0.037_{\rm stat.}\pm 0.019_{\rm syst.})\times10^{-3}$, $(3.622\pm0.034_{\rm stat.}\pm 0.018_{\rm syst.})\times10^{-3}$, and $(3.507\pm0.043_{\rm stat.}\pm 0.026_{\rm syst.})\times10^{-3}$ using…
▽ More
The branching fractions of $D^0\to π^-e^+ν_e$, $D^0\to π^-μ^+ν_μ$, $D^+\to π^0e^+ν_e$, and $D^+\to π^0μ^+ν_μ$ are measured to be $(2.950\pm0.017_{\rm stat.}\pm 0.017_{\rm syst.})\times10^{-3}$, $(2.817\pm0.037_{\rm stat.}\pm 0.019_{\rm syst.})\times10^{-3}$, $(3.622\pm0.034_{\rm stat.}\pm 0.018_{\rm syst.})\times10^{-3}$, and $(3.507\pm0.043_{\rm stat.}\pm 0.026_{\rm syst.})\times10^{-3}$ using $e^+e^-$ collision data with an integrated luminosity of 20.3 fb$^{-1}$ collected at the center-of-mass energy of 3.773 GeV with the BESIII detector. The partial decay rates of these four decays are measured with the best precision to date and their forward-backward asymmetries are determined for the first time. By performing a simultaneous fit to these results, the product of the hadronic transition form factor $f^{D\toπ}_+(0)$ and the modulus of the $c\to d$ Cabibbo-Kobayashi-Maskawa matrix element $|V_{cd}|$ is given by $f^{D\toπ}_+(0)|V_{cd}|=0.1425\pm0.0005_{\rm stat.}\pm0.0003_{\rm syst.}$. Taking the $|V_{cd}|$ provided by the standard model global fit and the $f^{D\toπ}_+(0)$ calculated from the lattice quantum chromodynamics as input, we obtain $f^{D\toπ}_+(0)=0.6339\pm0.0024_{\rm stat.}\pm0.0014_{\rm syst.}$ and $|V_{cd}|=0.2262\pm0.0008_{\rm stat.}\pm0.0005_{\rm syst.}\pm0.0018_{\rm LQCD.}$, respectively. The reported results have the best precision to date. We also search for the scalar current contribution in the $c\to d \ell^+ν_{\ell}$ transition and determine Re$(C_S^μ)=$ $0.022 \pm 0.023_{\rm stat.}\pm 0.003_{\rm syst.}$ and $|{\rm Im}(C_S^μ)|=0.000 \pm $ $0.038_{\rm stat.} \pm 0.012_{\rm syst.}$. In addition, the lepton flavor universality is tested with the ratios of the decay rates between semimuonic and semielectronic decays in full and several $\ell^+ν_\ell$ four-momentum transfer ranges.
△ Less
Submitted 26 July, 2026;
originally announced July 2026.
-
Recovery of latent inner products from an anisotropic Gaussian random geometric graph
Authors:
Cheng Mao,
Vidya Muthukumar
Abstract:
We study the problem of recovering latent inner products from a random geometric graph with anisotropic Gaussian latent points. More precisely, for an i.i.d. sample $x_1, \dots, x_n \sim N(0,Σ)$ where $Σ\in \mathbb{R}^{d \times d}$, an edge $(i,j)$ is present in the graph if and only if $\langle x_i, x_j \rangle \ge ζ$ for a threshold $ζ$. We assume the threshold $ζ$ to be chosen such that the ave…
▽ More
We study the problem of recovering latent inner products from a random geometric graph with anisotropic Gaussian latent points. More precisely, for an i.i.d. sample $x_1, \dots, x_n \sim N(0,Σ)$ where $Σ\in \mathbb{R}^{d \times d}$, an edge $(i,j)$ is present in the graph if and only if $\langle x_i, x_j \rangle \ge ζ$ for a threshold $ζ$. We assume the threshold $ζ$ to be chosen such that the average edge density of the graph is of constant order. To address the undesired degree fluctuations amplified by the anisotropy of the latent points, we consider the doubly centered adjacency matrix of the graph, and estimate the latent inner products using a rank-$d$ spectral approximation of the doubly centered matrix. The estimator obtains a mean squared error with a rate involving the stable rank of the covariance matrix $Σ$. Notably, the rate of estimation matches the state of the art for the isotropic case $Σ= I_d$, and permits an ill-conditioned covariance matrix with a diverging condition number. The analysis of the spectral method proceeds via the entrywise Hermite expansion of the doubly centered adjacency matrix with respect to the latent inner products. Instead of the standard trace method, it uses a decoupling argument recently introduced by Kaushik, Romberg, and Muthukumar (2025) to control nonlinear error terms.
△ Less
Submitted 26 July, 2026;
originally announced July 2026.
-
Mitigation of Measurement-Induced State Transitions via a Fast-Load and Fast-Clear Readout
Authors:
Wei-En Lin,
Li-Chieh Hsiao,
Chen-Hsun Ma,
Erh-Hsiang Yeh,
Wei-Lun Peng,
Hsi-Sheng Goan,
Cen-Shawn Wu,
Yueh-Nan Chen,
Yung-Fu Chen,
Chung-Ting Ke,
Chii-Dong Chen
Abstract:
High-fidelity and rapid qubit readout is essential for superconducting quantum processors, typically realized through the quantum non-demolition (QND) dispersive interaction within a qubit-resonator architecture. However, the achievable readout speed and fidelity are fundamentally limited by measurement-induced state transitions (MIST). For a transmon qubit, MIST is highly sensitive to the offset…
▽ More
High-fidelity and rapid qubit readout is essential for superconducting quantum processors, typically realized through the quantum non-demolition (QND) dispersive interaction within a qubit-resonator architecture. However, the achievable readout speed and fidelity are fundamentally limited by measurement-induced state transitions (MIST). For a transmon qubit, MIST is highly sensitive to the offset charge $n_g$ due to the charge dispersion of its higher-lying energy levels. In this work, we systematically investigate $n_g$-dependent MIST dynamics governed by the diabaticity and symmetry of pulse shaping within a charge-sensitive transmon architecture. We engineer fast-load and fast-clear pulses that effectively suppress resonator photon overshoots, thereby demonstrating a highly practical strategy to mitigate MIST without requiring complex waveforms or real-time feedback. Utilizing active gate-voltage control and rapid feedback, the measurement-induced transition probability is precisely mapped against $n_g$ and the steady-state resonator photon number, exhibiting strong agreement with numerical Floquet branch analysis. Ultimately, we evaluate the $n_g$-averaged total error probabilities for both readout and post-readout stages, verifying that a straightforward three-step pulse scheme consistently minimizes overall readout errors. Within the framework of large-scale superconducting quantum processors, this practical, hardware-free approach inherently offers a better trade-off between the readout signal-to-noise ratio and QND preservation.
△ Less
Submitted 26 July, 2026;
originally announced July 2026.
-
Measurement of Born Cross Section for $e^+e^-\to K_S^0\barΞ^+Σ^-+\rm{c.c.}$ at $\sqrt{s} = 3.51-4.95$ GeV
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko
, et al. (737 additional authors not shown)
Abstract:
Using $e^+e^-$ collision data collected with the BESIII detector at the BEPCII collider corresponding to a total integrated luminosity of 44~fb$^{-1}$, we present the first measurement of the Born cross sections for the process $e^+e^-\to K_S^0\barΞ^+Σ^-+\rm{c.c.}$ at 56 center-of-mass energies from 3.510 to 4.951~GeV. By fitting the dressed cross sections of $e^+e^-\to K_S^0\barΞ^+Σ^-+\rm{c.c.}$…
▽ More
Using $e^+e^-$ collision data collected with the BESIII detector at the BEPCII collider corresponding to a total integrated luminosity of 44~fb$^{-1}$, we present the first measurement of the Born cross sections for the process $e^+e^-\to K_S^0\barΞ^+Σ^-+\rm{c.c.}$ at 56 center-of-mass energies from 3.510 to 4.951~GeV. By fitting the dressed cross sections of $e^+e^-\to K_S^0\barΞ^+Σ^-+\rm{c.c.}$ with the assumption of a power-law function plus a charmonium(-like) resonance, i.e. $ψ(3770)$, $ψ(4040)$, $ψ(4160)$, $Y(4230)$, $Y(4360)$, $ψ(4415)$, {\it Y}(4500), $Y(4660)$, and {\it Y}(4710), no significant signal of any charmonium(-like) state decaying into the $K_S^0\barΞ^+Σ^-+\rm{c.c.}$ is observed. Upper limits on the product of the electronic width and branching fraction at the 90\% confidence level are given for each resonance. Combining this result with the previous measurement of the isospin-symmetric process $e^+e^-\to K^{-} \barΞ^{+} Σ^{0} + \rm{c.c.}$, the ratio of the Born cross sections, $R=σ^{B}(e^+e^-\to K_S^0\barΞ^+Σ^-+\rm{c.c.})/$$σ^{B}(e^+e^-\to K^-\barΞ^+Σ^0+\rm{c.c.})$, is found to be approximately 1.
△ Less
Submitted 24 July, 2026;
originally announced July 2026.
-
Explicit block-encodings for biharmonic boundary-value problems
Authors:
Chuwen Ma,
Zihao Tang
Abstract:
The biharmonic equation is a prototypical fourth-order partial differential equation whose high-dimensional discretization suffers from rapidly growing degrees of freedom and severe ill-conditioning. We develop QSVT--VTAA quantum linear-system algorithms by constructing explicit block-encodings tailored to periodic, simply supported, and Dirichlet--Neumann boundary conditions. For periodic and sim…
▽ More
The biharmonic equation is a prototypical fourth-order partial differential equation whose high-dimensional discretization suffers from rapidly growing degrees of freedom and severe ill-conditioning. We develop QSVT--VTAA quantum linear-system algorithms by constructing explicit block-encodings tailored to periodic, simply supported, and Dirichlet--Neumann boundary conditions. For periodic and simply supported problems, Fourier and sine-transform diagonalizations yield augmented Poisson systems with the condition-number scaling of a second-order operator. For Dirichlet--Neumann problems, we introduce a second-order boundary-corrected finite-difference discretization, establish mesh-independent stability, and construct an explicit block-encoding of the resulting nonsymmetric matrix. We also formulate a coupled-Laplace system with additional boundary unknowns and characterize its complexity in terms of the condition number of the complete augmented matrix. The analysis covers discretization error, block-encoding normalization, gate complexity, and solution extraction under an amplitude-input and quantum-state-output model. Numerical experiments validate the proposed discretizations and the corresponding linear solves.
△ Less
Submitted 24 July, 2026;
originally announced July 2026.
-
Towards Trustworthy and Cost-Efficient Data Integration: From Naïve RAG to Agentic RAG
Authors:
Chuangtao Ma,
Arijit Khan
Abstract:
Large language models (LLMs) and AI agents have demonstrated strong potential for data integration in zero-shot and few-shot settings. However, they continue to face significant accuracy and cost challenges in enterprise environments due to a persistent knowledge gap. This paper envisions trustworthy, scalable, and cost-efficient integration through knowledge-grounded LLMs and agents operating wit…
▽ More
Large language models (LLMs) and AI agents have demonstrated strong potential for data integration in zero-shot and few-shot settings. However, they continue to face significant accuracy and cost challenges in enterprise environments due to a persistent knowledge gap. This paper envisions trustworthy, scalable, and cost-efficient integration through knowledge-grounded LLMs and agents operating within a retrieval-augmented generation (RAG) workflow. Here, trustworthiness refers to evidence-grounded, verifiable reasoning, where integration decisions are transparently supported by retrieved knowledge, robust against hallucination, and consistent across tasks. We trace the evolution from classic RAG to GraphRAG and KG-RAG (knowledge graph-based RAG), highlighting how these paradigms bridge parametric and contextual knowledge. Building on this trajectory, we explore the shift toward Agentic RAG, where autonomous multi-agent systems adaptively plan, retrieve, refine, and reason for complex integration tasks. We examine optimization strategies for cost-efficient integration, addressing computational bottlenecks in large-scale enterprise settings. Finally, we outline open challenges and future directions toward building reliable, explainable, and scalable knowledge-grounded integration systems.
△ Less
Submitted 24 July, 2026;
originally announced July 2026.
-
Enhanced Curie temperature and room-temperature 50-nm skyrmions achieved in hexagonal ferromagnet Mn5Ge3+x synthesized via a high-pressure method
Authors:
Yongsen Zhang,
Wei Liu,
Meng Shi,
Shuisen Zhang,
Sheng Qiu,
Yaodong Wu,
Jialiang Jiang,
Huanhuan Zhang,
Hui Han,
Kang Wang,
Dingfu Shao,
Zhenfa Zi,
Chao Ma,
Haifeng Du,
Mingliang Tian,
Shouguo Wang,
Jin Tang
Abstract:
The development of new high-temperature ultrasmall-size skyrmion materials holds immense significance for the promising applications of topological spintronic devices. In this study, we demonstrate that a high-pressure synthesis technique can significantly elevate the Curie temperature of Mn5Ge3+x crystals, from 294 K to 350 K. This enhancement is attributed to the combined effects of lattice cont…
▽ More
The development of new high-temperature ultrasmall-size skyrmion materials holds immense significance for the promising applications of topological spintronic devices. In this study, we demonstrate that a high-pressure synthesis technique can significantly elevate the Curie temperature of Mn5Ge3+x crystals, from 294 K to 350 K. This enhancement is attributed to the combined effects of lattice contraction and increased Ge content, the conclusion supported by Density Functional Theory calculations. Additionally, our real-space magnetic imaging reveals the stability of dipolar skyrmions with diameters of approximately 50 nm at room temperature. Our micromagnetic simulations closely replicate the diverse experimental topological magnetic textures observed. Furthermore, magnetotransport measurements indicate the potential for the electrical distinction between various topological magnetic textures in skyrmion-based devices. We also report deterministic manipulations on single dipolar skyrmions in confined nanostructures by using in-plane currents. The observation, electrical manipulation, and electrical detection of room-temperature ultrasmall topological magnetic textures underscore the potential of Mn5Ge3+x as a promising platform for spintronic device applications.
△ Less
Submitted 23 July, 2026;
originally announced July 2026.
-
MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement
Authors:
Daiqing Wu,
Dongbao Yang,
Jiashu Yao,
Hongrui Zhang,
Can Ma,
Yu Zhou,
Sicheng Zhao
Abstract:
Affective Image Content Analysis (AICA) aims to recognize and understand emotions elicited by visual content, representing an indispensable step toward Artificial General Intelligence (AGI). However, despite the rapid progress of Multimodal Large Language Models (MLLMs), systematic evaluation of their visual emotional intelligence remains largely absent from recent model releases. We attribute thi…
▽ More
Affective Image Content Analysis (AICA) aims to recognize and understand emotions elicited by visual content, representing an indispensable step toward Artificial General Intelligence (AGI). However, despite the rapid progress of Multimodal Large Language Models (MLLMs), systematic evaluation of their visual emotional intelligence remains largely absent from recent model releases. We attribute this gap to a structural mismatch between conventional AICA paradigms and the open-ended, instruction-driven nature of MLLMs, where further analysis reveals four major limitations: omission of plausible responses, limited emotion taxonomies, neglect of contextual factors, and labor-intensive annotation. To overcome these barriers, we introduce Emotion Statement Judgement (ESJ), a statement-verification formulation that preserves the expressiveness of the input space while constraining outputs to discriminative judgements. We further develop INSETS, a labor-efficient pipeline that instantiates ESJ at scale by constructing INSETS-462k and supporting MVEI, a rigorously refined benchmark spanning sentiment polarity, emotion interpretation, scene context, and perception subjectivity. Beyond evaluation, we build EmObserver, an emotion-oriented MLLM optimized on ESJ through an elaborate multi-stage recipe. Extensive evaluation of broad-spectrum MLLMs on MVEI reveals fine-grained insights into current artificial visual emotional intelligence, while experiments on multiple AICA benchmarks demonstrate the accuracy, generalization, and reasoning faithfulness of EmObserver. Collectively, these results establish ESJ as a practical formulation, MVEI as a comprehensive benchmark, and EmObserver as an advanced baseline for advancing MLLM-oriented visual emotional intelligence. Code will be released at: https://github.com/wdqqdw/EmObserver.
△ Less
Submitted 23 July, 2026;
originally announced July 2026.
-
Probabilistic Residual Learning for Online Recommendations
Authors:
Wenyuan Wang,
Yusong Zhao,
Zihao Xu,
Hengyi Wang,
Qi Xu,
Zhigang Hua,
Yan Xie,
Yi Wang,
Zihao Zhao,
Bo Long,
Chengzhi Mao,
Shuang Yang,
Hengguan Huang,
Hao Wang
Abstract:
Modern recommender systems are typically based on deep learning (DL) models, where a dense encoder learns representations of users and items. As a result, these systems often suffer from the black-box nature and computational complexity of the underlying models, making it difficult to systematically enhance their recommendation capabilities. To address this problem, we propose Probabilistic Residu…
▽ More
Modern recommender systems are typically based on deep learning (DL) models, where a dense encoder learns representations of users and items. As a result, these systems often suffer from the black-box nature and computational complexity of the underlying models, making it difficult to systematically enhance their recommendation capabilities. To address this problem, we propose Probabilistic Residual Learning (PRL), a causal Bayesian recommendation model that models the residual between ground-truth and base predictions, enabling targeted refinement of existing systems. Specifically, PRL (1) probabilistically groups users for localized residual modeling, (2) models domain-level confounders that influence user and item representations, and (3) aggregates cluster-specific residual predictions over the confounders using do-calculus. Experiments demonstrate that our plug-and-play PRL is compatible with various base deep learning recommender systems, improving their performance while automatically discovering meaningful user clusters.
△ Less
Submitted 22 July, 2026;
originally announced July 2026.
-
CANN Bench: Benchmarking Agent Generated Kernels against Real NPU and Algorithmic Limits
Authors:
Xue-Jian Gao,
Deng Pan,
Yueming Su,
Jiasheng Li,
Bin Du,
Fengming Zhu,
Chengdi Ma,
Junyi Fan,
Qichen Liao,
Chengqiu Hu,
Xinxian Chen,
Lingchao Zheng,
Jun Li,
Jiwei Yang,
Yuwei Fan
Abstract:
AI agents are now capable of writing, compiling, and iteratively optimizing low-level operator kernels on different hardware platforms. Existing benchmarks, however, focus almost exclusively on CUDA and Triton, leaving hardware ecosystems with less-exposed programming models without a common evaluation baseline. We present CANN Bench, an open benchmark for AI-generated operator code on Huawei's As…
▽ More
AI agents are now capable of writing, compiling, and iteratively optimizing low-level operator kernels on different hardware platforms. Existing benchmarks, however, focus almost exclusively on CUDA and Triton, leaving hardware ecosystems with less-exposed programming models without a common evaluation baseline. We present CANN Bench, an open benchmark for AI-generated operator code on Huawei's Ascend NPU. The current release covers 53 operators and 1060 test cases organized into four difficulty tiers -- from simple elementwise primitives to MoE dispatch and FlashAttention kernels -- spanning FP16, BF16, FP32, and INT8 precision formats. Evaluation adopts a \textbf{three-dimensional weighted composite score} that treats compilation, functional correctness, and performance as independent axes, providing a principled reward signal for kernel-generation agents. Performance is graded against an out-of-the-box PyTorch-on-Ascend baseline and an analytical per-case Hardware-Anchored Performance (HAP) limit on real NPU hardware, ensuring scores reflect genuine optimization headroom rather than measurement artifacts. The evaluation harness is designed to resist reward hacking from the ground up. CANN Bench is versioned within the official CANN repository and is designed for long-term community co-construction, providing the Ascend ecosystem with a quantitative, reproducible, and sustainably maintained yardstick for AI operator-authoring capability.
△ Less
Submitted 7 July, 2026;
originally announced July 2026.
-
First Measurement of the Relative Phase between Proton Psionic Form Factors
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
Y. Bai,
O. Bakina,
Y. Ban,
H. -R. Bao,
X. L. Bao,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko
, et al. (732 additional authors not shown)
Abstract:
The relative phase between the time-like form factors of the proton is a crucial observable for a complete understanding of its internal structure, yet it has remained unmeasured due to the formidable experimental challenge of determining the final-state polarization or having available polarized beams. With a novel technique that measures polarization via secondary scattering on spectrometer mate…
▽ More
The relative phase between the time-like form factors of the proton is a crucial observable for a complete understanding of its internal structure, yet it has remained unmeasured due to the formidable experimental challenge of determining the final-state polarization or having available polarized beams. With a novel technique that measures polarization via secondary scattering on spectrometer material, we use $10.09\times10^{9}$ $J/ψ$ events collected at BESIII to analyze the reaction $e^+e^-\rightarrow J/ψ\rightarrow p\bar{p}$. This allows the first determination of the sine of the relative phase between the proton psionic form factors, $\sinΔΦ=-0.20\pm0.34_{\textrm{stat}}\pm0.11_{\textrm{syst}}$. This result provides the first direct insight into the complex dynamics of proton formation, and offers valuable new information to constrain theoretical models of nucleon structure.
△ Less
Submitted 22 July, 2026;
originally announced July 2026.
-
Proof of principle for nucleon polarization measurement at BESIII
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
Y. Bai,
O. Bakina,
Y. Ban,
H. -R. Bao,
X. L. Bao,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko
, et al. (732 additional authors not shown)
Abstract:
A novel technique for measuring the spin polarization of final-state nucleons in a general-purpose spectrometer is validated. Using $10.09\times10^{9}$ $J/ψ$ events at BESIII, the asymmetry of polarized proton scattering on detector support material is measured, and is consistent with the expected value. This proves that a general-purpose spectrometer can be utilized as a large-acceptance polarime…
▽ More
A novel technique for measuring the spin polarization of final-state nucleons in a general-purpose spectrometer is validated. Using $10.09\times10^{9}$ $J/ψ$ events at BESIII, the asymmetry of polarized proton scattering on detector support material is measured, and is consistent with the expected value. This proves that a general-purpose spectrometer can be utilized as a large-acceptance polarimeter, providing the spin polarization in addition to the conventional four-momentum information of the final-state particles. With this technique, physics capabilities are enhanced for existing and future facilities in particle and nuclear physics.
△ Less
Submitted 22 July, 2026;
originally announced July 2026.
-
A reduction scheme for general-order Ising-like Hamiltonians in quantum heuristic solvers
Authors:
Chengsi Mao,
Pavel Mosharev,
Yao Wang,
Man-Hong Yung
Abstract:
The Ising model is ubiquitous in various optimization problems but notoriously difficult to solve due to combinatorial explosion. In view of this, Hamiltonian reduction is a useful preprocessing technique for reducing the effective problem size before applying heuristic solvers. However, existing reduction techniques mainly target second-order Ising models, whereas many pseudo-Boolean formulations…
▽ More
The Ising model is ubiquitous in various optimization problems but notoriously difficult to solve due to combinatorial explosion. In view of this, Hamiltonian reduction is a useful preprocessing technique for reducing the effective problem size before applying heuristic solvers. However, existing reduction techniques mainly target second-order Ising models, whereas many pseudo-Boolean formulations naturally contain higher-order interactions. In this work, we generalize the concept of non-separable groups to arbitrary-order Ising-like models and develop a Hamiltonian reduction framework that iteratively detects and merges constrained spin groups into single variables. We benchmark the reduction on synthetic hypergraphs and higher-order network datasets, and evaluate its integration with downstream order-reduction and solver workflows. Our results establish a foundation for Hamiltonian reduction in higher-order Ising-like optimization problems.
△ Less
Submitted 22 July, 2026;
originally announced July 2026.
-
Point-Selection Fine-Tuning Framework for Robust Point Cloud Classification
Authors:
Da Li,
Chang Ma,
Dongfu Yin
Abstract:
Noisy and corrupted points can substantially degrade point cloud recognition performance, especially under challenging corruption settings. In particular, full fine-tuning of 3D pre-trained models may amplify the influence of outliers and overwrite robustness priors learned during pre-training, while naive parameter-efficient adaptation remains sensitive to corrupted tokens. To address this issue,…
▽ More
Noisy and corrupted points can substantially degrade point cloud recognition performance, especially under challenging corruption settings. In particular, full fine-tuning of 3D pre-trained models may amplify the influence of outliers and overwrite robustness priors learned during pre-training, while naive parameter-efficient adaptation remains sensitive to corrupted tokens. To address this issue, we propose PSFT, a point-selection fine-tuning framework that improves robustness while remaining parameter-efficient. PSFT first estimates point-wise influence from pre-pooling features and adaptively retains minimally influential points to suppress outliers. Based on the selected subset, a prompt generation branch predicts layer-wise prompt tokens and injects them into a frozen backbone for lightweight downstream adaptation. To further mitigate residual noise after selection, we append a lightweight feature filter with bottleneck MLP transformation and Beta-gated residual blending to refine patch-token representations before prediction. Extensive experiments show that PSFT consistently reduces corruption error on ModelNet-C and ModelNet40-C across all tested 3D pre-trained backbones, while achieving the strongest ScanObjectNN-C results with ULIP-2 and Uni3D-B among the evaluated tuning strategies. Our implementation can be found at https://github.com/CVChMA/PSFT/tree/master.
△ Less
Submitted 21 July, 2026;
originally announced July 2026.
-
Harness Engineering for LLM-Driven GPU Kernel Generation
Authors:
Yue Shui,
Chenyu Ma,
Hangfei Xu,
Shengzhao Wen,
Yanpeng Wang
Abstract:
Large language models (LLMs) can assist GPU kernel generation, but their practical effectiveness depends on whether generated code can be reliably constrained, validated, profiled, and selected. This paper presents a harness-centered system for LLM-driven GPU kernel optimization in the MLSys 2026 FlashInfer AI Kernel Generation Contest on NVIDIA Blackwell B200 GPUs. The system separates an evaluat…
▽ More
Large language models (LLMs) can assist GPU kernel generation, but their practical effectiveness depends on whether generated code can be reliably constrained, validated, profiled, and selected. This paper presents a harness-centered system for LLM-driven GPU kernel optimization in the MLSys 2026 FlashInfer AI Kernel Generation Contest on NVIDIA Blackwell B200 GPUs. The system separates an evaluation harness from a profile-backed optimization controller: the harness enforces compilation, correctness, official-aligned timing, and artifact archival, while the controller turns profiler and workload evidence into bounded candidate-generation decisions. Human-authored skills capture operator constraints, references, profiling procedures, and promotion rules, while Codex and Claude Code agents generate candidate kernels inside those constraints. Across five operator definitions, the retained official-aligned artifacts achieved mean-latency speedups over supplied FlashInfer baselines of 1.62x, 18.05x, 29.68x, 1.12x, and 13.70x. The Agent-Assisted kernels outperform the Full-Agent artifacts across the evaluated definitions, indicating that expert-provided optimization directions, high-quality references, and workload context remain critical for reliable AI-driven kernel optimization.
△ Less
Submitted 20 July, 2026;
originally announced July 2026.
-
Ah-SCDFT:A general approach for superconductivity with an-harmonic corrections
Authors:
Xiaozheng Fan,
Panshi Jing,
Chuanguang Zhang,
Junshuai Wang,
Chunlan Ma,
Shijing Gong,
Chuanxi Zhao,
Tianxing Wang,
Yipeng An
Abstract:
First-principles studies of superconductivity often neglect anharmonic effects (AHE), despite their crucial role in achieving quantitative accuracy in many materials. To bridge this gap, we introduce a general computational approach, termed anharmonic superconducting density functional theory (ah-SCDFT) which systematically incorporates anharmonic corrections into standard SCDFT. This approach all…
▽ More
First-principles studies of superconductivity often neglect anharmonic effects (AHE), despite their crucial role in achieving quantitative accuracy in many materials. To bridge this gap, we introduce a general computational approach, termed anharmonic superconducting density functional theory (ah-SCDFT) which systematically incorporates anharmonic corrections into standard SCDFT. This approach allows for high-fidelity predictions of superconducting properties with only a modest increase in computational cost for a limited number of superconducting calculation convergence steps. We demonstrate the effectiveness and reliability of ah-SCDFT by applying it to the prototypical superconductor MgB2, accurately reproducing its superconducting behavior under both ambient conditions and applied pressure in excellent agreement with experiment. Our results establish ah-SCDFT as a powerful, efficient, and broadly applicable approach for quantitatively reliable studies of superconductivity and a promising tool for the prediction of new superconducting materials.
△ Less
Submitted 20 July, 2026;
originally announced July 2026.
-
A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix
Authors:
Changzheng Ma
Abstract:
Multi-head Latent Attention (MLA) ships two implementations in Megatron-Core: an explicit form used for training and an absorbed form -- which slashes collective communication by gathering only the compressed latent -- that is fully implemented but hard-asserted out of training (the forward opens with "assert not (self.training and self.cache_mla_latents)"), allowed only in inference decode. The l…
▽ More
Multi-head Latent Attention (MLA) ships two implementations in Megatron-Core: an explicit form used for training and an absorbed form -- which slashes collective communication by gathering only the compressed latent -- that is fully implemented but hard-asserted out of training (the forward opens with "assert not (self.training and self.cache_mla_latents)"), allowed only in inference decode. The library documents no reason. We show the restriction is well-founded and quantify why: ported to training, the absorbed form is a memory trap -- its intermediates live in n_h x d_kv dimensions per token, larger than the per-head K/V they replace -- inflating activation memory by 20-34%, up to 9.2 GB at DeepSeek-V3 scale (n_h=128, seq=16384, SP=8, eager kernel; the gap widens to 19.2 GB under a fused kernel), enough to change device-fit. This measurement, validated on two axes (linear in seq and n_h) and cross-verified on NVIDIA A100, explains the otherwise-undocumented restriction and leaves practitioners with no low-communication MLA training path. We then provide one. LAGA (Latent All-Gather Attention) keeps the absorbed form's latent-gather communication but rejects the absorb reformulation, instead reconstructing per-head K/V locally from the gathered latent. On 8x Ascend 910B at real DeepSeek-V3 dimensions, LAGA cuts collective communication 1.98x, matches explicit memory within 0.5%, is bit-identical to explicit at SP=1 and equivalent to within 1e-3 at SP=2-8, and under a fused attention kernel improves attention-block throughput 1.04-1.06x single-node and 1.07-1.24x cross-node -- leading at all sequence lengths in the cross-node regime MLA is deployed for.
△ Less
Submitted 20 July, 2026;
originally announced July 2026.