-
Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models
Authors:
Zongyang Qiu,
Yihan Wu,
Kaixuan Fan,
Bo Li,
Hui Xiong
Abstract:
Unified multimodal models (UMMs) are motivated by the hope that understanding and generation reinforce each other but controlled ablations repeatedly find that adding a generation objective leaves understanding flat. Joint-training studies cannot settle the disagreement: with overlapping supervision, a gain cannot be attributed to the architecture rather than the data. To further investigate the r…
▽ More
Unified multimodal models (UMMs) are motivated by the hope that understanding and generation reinforce each other but controlled ablations repeatedly find that adding a generation objective leaves understanding flat. Joint-training studies cannot settle the disagreement: with overlapping supervision, a gain cannot be attributed to the architecture rather than the data. To further investigate the relationship between the two directions in UMMs, we separate them by construction. A novel visual entity, a rendered 3D asset paired with a pseudo-word screened for absence from the frozen model's behavior, is bound through exactly one task direction, and the untrained direction is then measured. We find that the channel is real in both directions, but the directions differ in kind: generation training installs a name the model can only match among candidates; understanding training installs one it can also produce. What governs cross-task usability is where the binding enters the shared computation. An alignment probe predicts export across 36 configurations (Spearman $ρ= +0.68$). That objective's alignment term, maximized in closed form over activations with every weight frozen, makes a concept drawable when injected at layer 7 of 28 and is indistinguishable from the base model from layer 14 on, while the weight-based version of the same edit peaks at layers 10-14. In an observational series of four models, this window appears only where the understanding pathway is a semantic vision encoder, suggesting that unified weights are not enough: the two directions must share a semantic format at the entry point. Exploiting the rule, a mid-stack alignment objective acquires the concept for a $0.1\%$ relative loss of the model's general text-to-image ability, against $41\%$ for the standard generative route. Our code is at https://github.com/Zane-ZYQiu/entry-point-umm.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
High-precision measurement of the space-like $η^\prime$ transition form factor
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
M. S. Anderson,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone
, et al. (758 additional authors not shown)
Abstract:
Using a data sample corresponding to an integrated luminosity of $20.3\ \text{fb}^{-1}$, collected with the BESIII detector at a center-of-mass energy of $3.773\ \text{GeV}$ at the BEPCII collider, we report a precision measurement of the product $Q^2|F(Q^2)|$, where $F(Q^2)$ is the single-virtual space-like transition form factor of the $η'$ meson and $Q^2$ is the squared momentum transfer of the…
▽ More
Using a data sample corresponding to an integrated luminosity of $20.3\ \text{fb}^{-1}$, collected with the BESIII detector at a center-of-mass energy of $3.773\ \text{GeV}$ at the BEPCII collider, we report a precision measurement of the product $Q^2|F(Q^2)|$, where $F(Q^2)$ is the single-virtual space-like transition form factor of the $η'$ meson and $Q^2$ is the squared momentum transfer of the tagged virtual photon. The transition form factor is extracted from the differential Born cross section of the two-photon fusion processes $e^+e^- \to e^+e^-γγ^* \to e^+e^-η^\prime$ using a single-tag technique, where only one scattered lepton is detected. The measurement covers $Q^2 \in [0.1, 6.0]$ GeV$^2$, achieving unprecedented precision, better than $3.0\%$ for $Q^2 < 1.5$ GeV$^2$, and providing the first direct determination at $Q^2 < 0.3$ GeV$^2$.
△ Less
Submitted 12 August, 2026;
originally announced August 2026.
-
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA
Authors:
Mind Lab,
:,
Vin Bo,
Asher Cai,
Jingwei Cao,
Song Cao,
Vic Cao,
Amelia Chen,
Andrew Chen,
Kaijie Chen,
Cleon Cheng,
Steven Chiang,
Kaixuan Fan,
Hera Feng,
Huan Feng,
Arthur Fu,
Jun Gao,
Pyke Han,
Nolan Ho,
Ori Hong,
Hailee Hou,
Piers Hua,
Charles Huang,
Miles Jiang,
Nora Jiang
, et al. (52 additional authors not shown)
Abstract:
Macaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments and continuing to learn after deployment. It is organized around two system goals. Adaptation is pursued through recursive improvement of versioned model-harness pairs, where experience from one configuration is evaluated under an external contract and used to construct its success…
▽ More
Macaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments and continuing to learn after deployment. It is organized around two system goals. Adaptation is pursued through recursive improvement of versioned model-harness pairs, where experience from one configuration is evaluated under an external contract and used to construct its successor. Collaboration is pursued via the Mixture-of-LoRA (MoL) architecture that freezes a base model, composes specialist LoRA adapters, and selects one LoRA per user turn. The flagship Macaron-V1-Venti combines a 744B GLM-5.2 base with four LoRAs for chat, agent, coding, and GenUI; the Qwen3.6-based Macaron-V1-Tall (50B) uses the same design for local deployment. This report presents Macaron-V1 as a co-designed system spanning architecture, algorithms, and infrastructure. The MoL architecture supports continual learning through extensible LoRA specialists. The algorithm combines Model-Harness Co-design and recursive self-improvement loop, including the UI4A component-native GenUI harness, a stateful action substrate, versioned HCP contract, and the agentic RL framework MindForge. The supporting infrastructure includes the post-training platform MinT, the long-context RL method LongStraw, and stability techniques for sparse MoE and DSA base models. We evaluate Macaron-V1 on Personal Intelligence, GenUI, and general capability benchmarks against frontier baselines. Our results validate the current system, while compounding gains from continual learning and collective intelligence remain open questions.
△ Less
Submitted 10 August, 2026;
originally announced August 2026.
-
Estimating the sensitivity of the IceCube Upgrade to probe the interior of the Earth using atmospheric neutrino oscillations
Authors:
The IceCube Collaboration,
R. Abbasi,
M. Ackermann,
J. Adams,
S. K. Agarwalla,
J. A. Aguilar,
M. Ahlers,
J. M. Alameddine,
S. Ali,
N. M. Amin,
K. Andeen,
C. Arg{ü}elles,
S. Athanasiadou,
S. N. Axani,
R. Babu,
X. Bai,
A. Balagopal V.,
S. W. Barwick,
V. Basu,
R. Bay,
J. J. Beatty,
J. Becker Tjus,
P. Behrens,
J. Beise,
C. Bellenghi
, et al. (399 additional authors not shown)
Abstract:
The IceCube Upgrade is a densely instrumented central region of the IceCube Neutrino Observatory, deployed during the 2025-26 polar season. It will reduce the detector's energy threshold and improve overall reconstruction capabilities for multi-GeV atmospheric neutrinos, which in turn enhance their sensitivity to Earth matter effects as they traverse through the deep Earth. In this study, we descr…
▽ More
The IceCube Upgrade is a densely instrumented central region of the IceCube Neutrino Observatory, deployed during the 2025-26 polar season. It will reduce the detector's energy threshold and improve overall reconstruction capabilities for multi-GeV atmospheric neutrinos, which in turn enhance their sensitivity to Earth matter effects as they traverse through the deep Earth. In this study, we describe the potential of the IceCube Upgrade to observe Earth matter effects on atmospheric neutrinos and estimate the detector's sensitivity to probe key features of the Preliminary Reference Earth Model by utilizing these observations. We highlight the IceCube Upgrade's capability to estimate the mass of the Earth and verify the non-homogeneous distribution of matter density within the Earth. We also estimate the IceCube Upgrade sensitivity to measure the correlated densities of the Earth layers while incorporating constraints from the mass and moment of inertia of the Earth. Neutrino-based results would be independent and complementary to the seismic and gravitational measurements.
△ Less
Submitted 6 August, 2026;
originally announced August 2026.
-
Configurable and Hierarchical Allreduce
Authors:
Valentino Guerrini,
Ke Fan,
Sidharth Kumar
Abstract:
MPI_Allreduce is among the most performance-critical collectives in large-scale scientific computing and distributed machine learning, yet the small- and medium-message regime remains challenging: latency, synchronization depth, and strong hardware hierarchy between intra- and inter-domain communication all compound per-invocation cost. We present CHIARA, a configurable hierarchical Allreduce that…
▽ More
MPI_Allreduce is among the most performance-critical collectives in large-scale scientific computing and distributed machine learning, yet the small- and medium-message regime remains challenging: latency, synchronization depth, and strong hardware hierarchy between intra- and inter-domain communication all compound per-invocation cost. We present CHIARA, a configurable hierarchical Allreduce that encodes hardware hierarchy through a logical batch-lane topology and executes a staged schedule in which only a bounded portion of the reduction vector is active at a time. Inter-batch communication is distributed across multiple ranks via a rotating-root lane primitive, avoiding centralized leaders. Tool further enables a semi-composed Rabenseifner-style Allreduce by preserving a lane-aligned intermediate layout across the Reduce-Scatter/Allgather boundary, eliminating redundant intra-domain reorganization. We evaluate Tool on Polaris, Aurora, and Fugaku, achieving speedups of up to 1.94x, 13.43x, and 13.48x over vendor MPI_Allreduce, and up to 2.2x end-to-end speedup in a parallel k-means application.
△ Less
Submitted 3 August, 2026;
originally announced August 2026.
-
LimICE: Integrating LLM into ICE Framework for Efficient Loop Invariant Inference
Authors:
Kai Fan,
ShiWen Yu,
GuangSheng Fan,
HaoAng Chi,
WanWei Liu,
Ji Wang
Abstract:
Loop invariant synthesis is a fundamental problem in program verification, yet the inherent undecidability makes it highly challenging. Recent studies have increasingly employed various machine learning techniques to generate loop invariants. However, most of these methods adopt a monolithic approach. Due to the inability to strictly constrain the learning process, learning-based methods struggle…
▽ More
Loop invariant synthesis is a fundamental problem in program verification, yet the inherent undecidability makes it highly challenging. Recent studies have increasingly employed various machine learning techniques to generate loop invariants. However, most of these methods adopt a monolithic approach. Due to the inability to strictly constrain the learning process, learning-based methods struggle to simultaneously consider all necessary conditions and generate complete invariants when tackling complex problems. In fact, a loop invariant is often an ordered sequence of lemmas, rather than a single invariant formula. This motivates us to propose Incremental ICE, a novel learning framework for incremental synthesis. Our framework integrates the incremental philosophy of IC3 into the general invariant learning framework ICE. By defining a lemma-specific learning objective and introducing a counterexample filtering mechanism, we can achieve sound incremental learning. Under this framework, we instantiate a loop invariant synthesis tool, LimICE, which leverages LLMs to generate the ordered sequence of lemmas and incorporates ICE-DT as a fallback mechanism to complement the lemma sequence. Experiments on 367 linear benchmarks and 50 nonlinear benchmarks demonstrate the effectiveness of the proposed approach. LimICE solves 349 (out of 367) linear problems on an average of 15.2 seconds and 47 (out of 50) nonlinear problems on an average of 8.8 seconds. Compared to the state-of-the-art LLM-based baseline, our approach solves 12-24% more instances while running 36-63% faster across linear and nonlinear benchmarks. LimICE also consistently outperforms strong non-LLM baselines and solves at least 86 and 27 additional instances on the linear and nonlinear benchmarks, respectively.
△ Less
Submitted 29 July, 2026;
originally announced July 2026.
-
High-energy neutrino emission from the Milky Way
Authors:
R. Abbasi,
M. Ackermann,
J. Adams,
J. A. Aguilar,
M. Ahlers,
J. M. Alameddine,
S. Ali,
N. M. Amin,
K. Andeen,
C. Argüelles,
S. Athanasiadou,
S. N. Axani,
R. Babu,
X. Bai,
A. Balagopal V.,
S. W. Barwick,
V. Basu,
R. Bay,
J. J. Beatty,
J. Becker Tjus,
P. Behrens,
J. Beise,
C. Bellenghi,
S. Benkel,
S. BenZvi
, et al. (398 additional authors not shown)
Abstract:
The Milky Way hosts astrophysical objects that accelerate cosmic rays to energies beyond the reach of terrestrial particle accelerators. It remains a longstanding goal to locate the sites of these powerful Galactic engines and understand how cosmic rays propagate through the Galaxy, leading to the production of high-energy neutrinos. In this paper, we combine event morphologies characteristic of a…
▽ More
The Milky Way hosts astrophysical objects that accelerate cosmic rays to energies beyond the reach of terrestrial particle accelerators. It remains a longstanding goal to locate the sites of these powerful Galactic engines and understand how cosmic rays propagate through the Galaxy, leading to the production of high-energy neutrinos. In this paper, we combine event morphologies characteristic of all three neutrino flavours and apply recent improvements in ice modelling, calibration and reconstruction to 12 years of IceCube data. With a predefined, global analysis we establish high-energy neutrino emission from the Galactic plane at 5.7 $σ$ significance. A further study shows that the inner region of the Galaxy is a prominent neutrino source, with 217 shower events with visible energy above 5 TeV compared with an expected background of 154.4 $\pm$ 4.1. These results herald a new era of Galactic multi-messenger astronomy, creating new opportunities to study cosmic-ray propagation and probe neutrino properties over kiloparsec distances.
△ Less
Submitted 28 July, 2026;
originally announced July 2026.
-
Precision Measurement of Decay Dynamics in $D^{0(+)}\to π^{-(0)}\ell^+ν_\ell$
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
M. S. Anderson,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone
, et al. (752 additional authors not shown)
Abstract:
The branching fractions of $D^0\to π^-e^+ν_e$, $D^0\to π^-μ^+ν_μ$, $D^+\to π^0e^+ν_e$, and $D^+\to π^0μ^+ν_μ$ are precisely measured, using 20.3 fb$^{-1}$ of $e^+e^-$ collision data collected at the center-of-mass energy of 3.773 GeV with the BESIII detector. The ratios of the decay widths between muon and positron channels are examined in full, across several four-momentum transfer ranges of…
▽ More
The branching fractions of $D^0\to π^-e^+ν_e$, $D^0\to π^-μ^+ν_μ$, $D^+\to π^0e^+ν_e$, and $D^+\to π^0μ^+ν_μ$ are precisely measured, using 20.3 fb$^{-1}$ of $e^+e^-$ collision data collected at the center-of-mass energy of 3.773 GeV with the BESIII detector. The ratios of the decay widths between muon and positron channels are examined in full, across several four-momentum transfer ranges of $\ell^+ν_{\ell}$. No lepton flavor universality violation is found in the current data. From a simultaneous fit to the precisely measured partial decay rates and the first measured forward-backward asymmetries of these four decays, the product of the hadronic transition form factor, $f^{D\toπ}_+(0)$, and the modulus of the $c\to d$ quark mixing element, $|V_{cd}|$, is measured with unprecedented precision to be $f^{D\toπ}_+(0)|V_{cd}|=0.1425\pm0.0005_{\rm stat.}\pm0.0003_{\rm syst.}$. Taking the value of $|V_{cd}|$ from the standard model global fit and $f^{D\toπ}_+(0)$ derived by the lattice quantum chromodynamics calculation as input, we obtain $f^{D\toπ}_+(0)=0.1425\pm0.0005_{\rm stat.}\pm0.0003_{\rm syst.}$ and $|V_{cd}|=0.2262\pm0.0008_{\rm stat.}\pm0.0005_{\rm syst.}\pm0.0018_{\rm LQCD.}$, respectively. The precision of each result is a factor of 2-3 better than the previous best measurements. Additionally, the real and imaginary parts of the scalar current contribution in the $c\to d \ell^+ν_{\ell}$ transition are measured for the first time to be Re $(C_S^μ)=$ $0.022 \pm 0.023_{\rm stat.}\pm 0.003_{\rm syst.}$ and $|\mathrm{Im} (C_S^μ)|=0.000 \pm 0.038_{\rm stat.}\pm 0.012_{\rm syst.}$.
△ Less
Submitted 26 July, 2026;
originally announced July 2026.
-
Precision measurements of semleptonic decays $D^0 \to π^-\ell^+ν_\ell$ and $D^+ \to π^0\ell^+ν_\ell$ ($\ell =e,μ$)
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
M. S. Anderson,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone
, et al. (752 additional authors not shown)
Abstract:
The branching fractions of $D^0\to π^-e^+ν_e$, $D^0\to π^-μ^+ν_μ$, $D^+\to π^0e^+ν_e$, and $D^+\to π^0μ^+ν_μ$ are measured to be $(2.950\pm0.017_{\rm stat.}\pm 0.017_{\rm syst.})\times10^{-3}$, $(2.817\pm0.037_{\rm stat.}\pm 0.019_{\rm syst.})\times10^{-3}$, $(3.622\pm0.034_{\rm stat.}\pm 0.018_{\rm syst.})\times10^{-3}$, and $(3.507\pm0.043_{\rm stat.}\pm 0.026_{\rm syst.})\times10^{-3}$ using…
▽ More
The branching fractions of $D^0\to π^-e^+ν_e$, $D^0\to π^-μ^+ν_μ$, $D^+\to π^0e^+ν_e$, and $D^+\to π^0μ^+ν_μ$ are measured to be $(2.950\pm0.017_{\rm stat.}\pm 0.017_{\rm syst.})\times10^{-3}$, $(2.817\pm0.037_{\rm stat.}\pm 0.019_{\rm syst.})\times10^{-3}$, $(3.622\pm0.034_{\rm stat.}\pm 0.018_{\rm syst.})\times10^{-3}$, and $(3.507\pm0.043_{\rm stat.}\pm 0.026_{\rm syst.})\times10^{-3}$ using $e^+e^-$ collision data with an integrated luminosity of 20.3 fb$^{-1}$ collected at the center-of-mass energy of 3.773 GeV with the BESIII detector. The partial decay rates of these four decays are measured with the best precision to date and their forward-backward asymmetries are determined for the first time. By performing a simultaneous fit to these results, the product of the hadronic transition form factor $f^{D\toπ}_+(0)$ and the modulus of the $c\to d$ Cabibbo-Kobayashi-Maskawa matrix element $|V_{cd}|$ is given by $f^{D\toπ}_+(0)|V_{cd}|=0.1425\pm0.0005_{\rm stat.}\pm0.0003_{\rm syst.}$. Taking the $|V_{cd}|$ provided by the standard model global fit and the $f^{D\toπ}_+(0)$ calculated from the lattice quantum chromodynamics as input, we obtain $f^{D\toπ}_+(0)=0.6339\pm0.0024_{\rm stat.}\pm0.0014_{\rm syst.}$ and $|V_{cd}|=0.2262\pm0.0008_{\rm stat.}\pm0.0005_{\rm syst.}\pm0.0018_{\rm LQCD.}$, respectively. The reported results have the best precision to date. We also search for the scalar current contribution in the $c\to d \ell^+ν_{\ell}$ transition and determine Re$(C_S^μ)=$ $0.022 \pm 0.023_{\rm stat.}\pm 0.003_{\rm syst.}$ and $|{\rm Im}(C_S^μ)|=0.000 \pm $ $0.038_{\rm stat.} \pm 0.012_{\rm syst.}$. In addition, the lepton flavor universality is tested with the ratios of the decay rates between semimuonic and semielectronic decays in full and several $\ell^+ν_\ell$ four-momentum transfer ranges.
△ Less
Submitted 26 July, 2026;
originally announced July 2026.
-
Observation of the $χ_{cJ}$ decays into $pK^{-}\barΛη+\mathrm{c.c.}$
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
M. S. Anderson,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone
, et al. (759 additional authors not shown)
Abstract:
By analyzing $(2712.4 \pm 14.3) \times 10^{6}$ $ψ(3686)$ events collected with the BESIII detector operating at the BEPCII collider, the decays $χ_{cJ} \to pK^{-}\barΛη+ \mathrm{c.c.}$ ($J=0,1,2$) are observed for the first time, with statistical significances exceeding $5σ$ for all three $χ_{cJ}$ states. The measured branching fractions are…
▽ More
By analyzing $(2712.4 \pm 14.3) \times 10^{6}$ $ψ(3686)$ events collected with the BESIII detector operating at the BEPCII collider, the decays $χ_{cJ} \to pK^{-}\barΛη+ \mathrm{c.c.}$ ($J=0,1,2$) are observed for the first time, with statistical significances exceeding $5σ$ for all three $χ_{cJ}$ states. The measured branching fractions are $\mathcal{B}(χ_{c0} \to pK^{-}\barΛη+ \mathrm{c.c.}) = (5.3 \pm 0.7 \pm 0.5) \times 10^{-5}$, $\mathcal{B}(χ_{c1} \to pK^{-}\barΛη+ \mathrm{c.c.}) = (9.8 \pm 0.6 \pm 0.6) \times 10^{-5}$, and $\mathcal{B}(χ_{c2} \to pK^{-}\barΛη+ \mathrm{c.c.}) = (9.3 \pm 0.6 \pm 0.6) \times 10^{-5}$, where the first uncertainties are statistical and the second are systematic. Structures consistent with the known hyperon resonances $Λ(1520)$ and $\barΛ(1690)$ are seen in the $pK^{-}$ and $\barΛη$ invariant mass spectra, respectively. The reported branching fractions include both resonant and non-resonant contributions. These results provide new experimental information on hadronic decays of $P$-wave charmonium states and contribute to the understanding of baryon production and hadronization dynamics in the nonperturbative QCD regime.
△ Less
Submitted 6 July, 2026;
originally announced July 2026.
-
$p$-adic multiple $L$-functions and twisted multiple Bernoulli numbers
Authors:
Ku-Yu Fan
Abstract:
We compute the special values ($p$MLFVs) of the $p$-adic multiple $L$-functions introduced by Furusho, Komori, Matsumoto, and Tsumura at tuples of positive integers. Furusho and Jarossay show that the special values can be expressed as an infinite sum of cyclotomic multiple harmonic values (CMHVs) with coefficients given by cyclotomic multiple Bernoulli numbers (CMBNs). We provide an explicit form…
▽ More
We compute the special values ($p$MLFVs) of the $p$-adic multiple $L$-functions introduced by Furusho, Komori, Matsumoto, and Tsumura at tuples of positive integers. Furusho and Jarossay show that the special values can be expressed as an infinite sum of cyclotomic multiple harmonic values (CMHVs) with coefficients given by cyclotomic multiple Bernoulli numbers (CMBNs). We provide an explicit formula for CMBNs in terms of twisted multiple Bernoulli numbers (TMBNs), which are special values of generalized Euler-Zagier-Lerch type complex multiple zeta functions at tuples of non-positive integers. As a result, we obtain that these $p$MLFVs can be expressed as infinite sums of CMHVs, with coefficients given by the special values of the complex functions at tuples of non-positive integers.
△ Less
Submitted 4 July, 2026;
originally announced July 2026.
-
High-Energy Neutrino Tomography of the Earth's Interior with IceCube
Authors:
The IceCube Collaboration,
R. Abbasi,
M. Ackermann,
J. Adams,
J. A. Aguilar,
M. Ahlers,
J. M. Alameddine,
S. Ali,
N. M. Amin,
K. Andeen,
C. Argüelles,
S. Athanasiadou,
S. N. Axani,
R. Babu,
X. Bai,
A. Balagopal V.,
S. W. Barwick,
V. Basu,
R. Bay,
J. J. Beatty,
J. Becker Tjus,
P. Behrens,
J. Beise,
C. Bellenghi,
S. Benkel
, et al. (395 additional authors not shown)
Abstract:
The Earth's interior reflects its geological evolution, from accretion to present-day dynamics. Its structure drives the geodynamo in the outer core, generating the magnetic field that shields the surface from charged cosmic radiation. The primary observables of the Earth's interior are its radial density distribution and derived quantities such as its mass and moment of inertia. These have tradit…
▽ More
The Earth's interior reflects its geological evolution, from accretion to present-day dynamics. Its structure drives the geodynamo in the outer core, generating the magnetic field that shields the surface from charged cosmic radiation. The primary observables of the Earth's interior are its radial density distribution and derived quantities such as its mass and moment of inertia. These have traditionally been inferred from gravity and seismic wave propagation, which probe the macroscopic response of matter to gravitational and elastic forces. Here we instead constrain the Earth's density profile using high-energy neutrinos observed by the IceCube Neutrino Observatory at the South Pole. We analyze 10.7 years of predominantly muon-neutrino data spanning 500 GeV--100 TeV, including atmospheric neutrinos produced by cosmic-ray interactions in the Earth's atmosphere and the diffuse astrophysical neutrino flux. Neutrino attenuation depends on both the traversed column density and neutrino energy. By measuring the zenith- and energy-dependent flux suppression, we infer the Earth's radial density profile by fitting a concentric uniform-density shell model that incorporates neutrino fluxes, interaction cross sections, detector response, and glacial-ice systematic uncertainties. From the resulting density posteriors, we derive the Earth's mass and polar moment of inertia as measured by neutrinos. These are the most precise weak-interaction measurements of these quantities to date and are consistent with the Preliminary Reference Earth Model and independent gravitational determinations. Our results demonstrate that neutrinos provide a novel probe of planetary interiors via a distinct physical interaction, complementing gravity and seismology. With improved detectors and precision, neutrinos will further contribute to a multifaceted understanding of the Earth's structure.
△ Less
Submitted 7 July, 2026; v1 submitted 2 July, 2026;
originally announced July 2026.
-
Tiles and weak tiles in $\mathbb{Z}_{pq}$
Authors:
Mamateli Kadir,
Kaibo Fan
Abstract:
This paper investigates the relationship between tiles and weak tiles in the context of finite cyclic group $\mathbb{Z}_{pq}$. We prove that weak tiles and translational tiles are equivalent in this group. Our proof employs Fourier analysis, Delsarte parameters, and the Coven-Meyerowitz conditions.
This paper investigates the relationship between tiles and weak tiles in the context of finite cyclic group $\mathbb{Z}_{pq}$. We prove that weak tiles and translational tiles are equivalent in this group. Our proof employs Fourier analysis, Delsarte parameters, and the Coven-Meyerowitz conditions.
△ Less
Submitted 4 July, 2026; v1 submitted 2 July, 2026;
originally announced July 2026.
-
WavePID: Low-energy flavor identification using single-PMT time series in IceCube
Authors:
The IceCube Collaboration,
R. Abbasi,
M. Ackermann,
J. Adams,
J. A. Aguilar,
M. Ahlers,
J. M. Alameddine,
S. Ali,
N. M. Amin,
K. Andeen,
C. Argüelles,
S. Athanasiadou,
S. N. Axani,
R. Babu,
X. Bai,
A. Balagopal V.,
S. W. Barwick,
V. Basu,
R. Bay,
J. J. Beatty,
J. Becker Tjus,
P. Behrens,
J. Beise,
C. Bellenghi,
S. Benkel
, et al. (395 additional authors not shown)
Abstract:
The IceCube Neutrino Observatory, a cubic-kilometer detector at the South Pole, identifies neutrino flavor through event morphology. Sparse photon detection makes this classification particularly challenging in the 5--100~GeV regime, the energy range relevant for oscillation measurements and searches for physics beyond the Standard Model. We introduce WavePID, a template-based log-likelihood-ratio…
▽ More
The IceCube Neutrino Observatory, a cubic-kilometer detector at the South Pole, identifies neutrino flavor through event morphology. Sparse photon detection makes this classification particularly challenging in the 5--100~GeV regime, the energy range relevant for oscillation measurements and searches for physics beyond the Standard Model. We introduce WavePID, a template-based log-likelihood-ratio classifier that exploits nanosecond-scale timing on individual detector modules through three observables: the distance to the reconstructed vertex, the early-charge fraction, and the module-to-module time difference. Evaluated on a cascade-enriched sample selected by a state-of-the-art graph neural network, WavePID improves both cascade purity and classification performance over the neural network alone. This demonstrates that per-module pulse timing carries flavor-identification information complementary to morphology-based classifiers, opening a new physics-motivated observable for low-energy neutrino reconstruction. Geant4 simulations associate this signal with differences in Cherenkov emission geometry between muon tracks and electromagnetic showers. These results motivate exploiting nanosecond-scale pulse timing in future low-energy classifiers and in detector designs with improved per-module timing in next-generation neutrino telescopes.
△ Less
Submitted 20 August, 2026; v1 submitted 2 July, 2026;
originally announced July 2026.
-
Study of the $e^+e^-\to π^+π^-D_s^+D_s^-$ process from $\sqrt{s}$ = 4.42 to 4.95 GeV at BESIII
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
M. S. Anderson,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko
, et al. (762 additional authors not shown)
Abstract:
Based on $8.5~{\rm fb}^{-1}$ of $e^+e^-$ collision data collected at center-of-mass energies between 4.42 and 4.95 GeV with the BESIII detector at the BEPCII storage ring, we investigate the process $e^+e^-\to π^+π^-D_s^+D_s^-$. With no significant signal observed, upper limits on the Born cross sections of $e^+e^-\to π^+π^-D_s^+D_s^-$ at each energy value are determined at the 90% confidence leve…
▽ More
Based on $8.5~{\rm fb}^{-1}$ of $e^+e^-$ collision data collected at center-of-mass energies between 4.42 and 4.95 GeV with the BESIII detector at the BEPCII storage ring, we investigate the process $e^+e^-\to π^+π^-D_s^+D_s^-$. With no significant signal observed, upper limits on the Born cross sections of $e^+e^-\to π^+π^-D_s^+D_s^-$ at each energy value are determined at the 90% confidence level. Additionally, a search for intermediate charmonium-like resonances is performed in the $M(D_s^+D_s^-)$ invariant-mass spectrum, but no significant resonant structures are observed with the current statistics.
△ Less
Submitted 27 June, 2026;
originally announced June 2026.
-
EGG: An Expert-Guided Agent Framework for Kernel Generation
Authors:
Yaochen Han,
Ke Fan,
Hongxu Jiang,
Wanqi Xu,
Weiyu Xie,
Runhua Zhang,
Chenhui Zhu,
Yixiang Zhang
Abstract:
High-performance GPU kernels are critical for reducing the exponentially growing computational costs of large language models (LLMs), but their development heavily relies on manual tuning by domain experts. While recent advances in LLM-based approaches show promise for automating kernel generation, they still struggle to achieve both correctness and high performance. This limitation primarily aris…
▽ More
High-performance GPU kernels are critical for reducing the exponentially growing computational costs of large language models (LLMs), but their development heavily relies on manual tuning by domain experts. While recent advances in LLM-based approaches show promise for automating kernel generation, they still struggle to achieve both correctness and high performance. This limitation primarily arises from the lack of domain-specific optimization guidance, hindering effective exploration of the optimization space. We propose EGG, an Expert-Guided Agent Framework for Kernel Generation, which incorporates expert optimization principles to guide LLMs' decisions. Inspired by expert workflows, we decompose kernel generation into two hierarchical stages: 1) algorithmic structure design, which establishes a high-quality computational structure foundation; 2) hardware-specific tuning, which performs targeted adjustments through parallel mapping, tensor tiling, and memory optimization. This staged decomposition defines explicit optimization objectives, structuring the design space to achieve progressive refinement. To this end, a stage-aware multi-agent collaboration mechanism is designed for inter and intra-stage context management, ensuring stable optimization trajectories. Experiments on KernelBench and real-world workloads show that EGG achieves a 2.13x average speedup over PyTorch, outperforming existing agent-based and RL-based approaches.
△ Less
Submitted 25 June, 2026;
originally announced June 2026.
-
Magnetic-Free Quantum Interference and Universal Josephson Diode Effect Driven by a Supercurrent Gauge Field
Authors:
Haowei Ye,
Wenxue He,
Kaixuan Fan,
Yingpeng Zhang,
Shijin Li,
Yu Pan,
Dechao Geng,
Fan Yang,
Kenji Watanabe,
Takashi Taniguchi,
Hechen Ren
Abstract:
The Josephson effect, a hallmark of superconducting phase coherence, drives modern quantum technologies. However, Josephson-based quantum interference has hitherto been tethered to magnetic fields, despite phase coherence being a quintessential, intrinsic trait of superconductivity. Moreover, the Josephson diode effect (JDE) is typically viewed as an anomalous phenomenon indicative of broken symme…
▽ More
The Josephson effect, a hallmark of superconducting phase coherence, drives modern quantum technologies. However, Josephson-based quantum interference has hitherto been tethered to magnetic fields, despite phase coherence being a quintessential, intrinsic trait of superconductivity. Moreover, the Josephson diode effect (JDE) is typically viewed as an anomalous phenomenon indicative of broken symmetries in exotic phases of matter. Here, in planar Josephson junctions made with $\mathrm{Bi}_2\mathrm{O}_2\mathrm{Se}$ and bilayer graphene, we demonstrate that the JDE is a missing universal property of the Josephson effect. Simultaneously, we present an all-electric technology that replaces magnetic flux for controlling and measuring supercurrent interference. Central to our approach is a supercurrent gauge field (SGF), generated and amplified through high-kinetic-inductance superconductors and novel device architectures. By establishing the physical equivalence between the SGF and a magnetic field, we eliminate the reliance on external fields in quantum interference and reveal a universal, field-free JDE mechanism with broad implications for detecting broken-symmetry states. Finally, we show that the SGF offers capabilities beyond those of a conventional magnetic field by experimentally demonstrating a magnetic-free, phase-sensitive technique to construct and characterize finite-momentum superconductivity, opening new frontiers for exploring novel phases of matter and superconducting quantum architectures.
△ Less
Submitted 29 June, 2026; v1 submitted 22 June, 2026;
originally announced June 2026.
-
Learning Earthquake Wave Arrival Time Picking from Labels with Inaccuracies
Authors:
Sen Li,
Xu Yang,
S. Mostafa Mousavi,
Anye Cao,
Keting Fan,
Yaoqi Liu,
Changbin Wang,
Qiang Niu
Abstract:
Inaccurately labeled training data, or "label noise", poses a significant threat to the integrity of supervised machine learning models. This corruption directly degrades performance by teaching the model erroneous mappings between features and labels, which leads to poor generalization and reduced accuracy on properly labeled validation and test data. Current seismological applications mainly rel…
▽ More
Inaccurately labeled training data, or "label noise", poses a significant threat to the integrity of supervised machine learning models. This corruption directly degrades performance by teaching the model erroneous mappings between features and labels, which leads to poor generalization and reduced accuracy on properly labeled validation and test data. Current seismological applications mainly rely on large-scale training sets or data augmentation to reduce the label-noise impact, which can be labor-intensive and costly. Here, we introduce a Label Noise-Contrastive Robust Learning (LaNCoR) approach that can effectively handle noisy labels in seismic signal processing tasks, without requiring large-scale training datasets. In this approach, the input waveform feature and label representation distributions are aligned in the feature space to correct mislabeling and reduce its impact on the training process. We present LaNCoR's performance on the task of P-phase arrival-time picking of real microseismic data using two baseline models and training approaches. Our results indicate that LaNCoR can improve performance by up to 28.8% across performance metrics. This approach holds great promise for model training in seismology and geosciences.
△ Less
Submitted 13 June, 2026;
originally announced June 2026.
-
IceCube Real-time Searches for High-energy Neutrinos Coincident with LIGO/Virgo/KAGRA Gravitational-Wave Alerts in O4a
Authors:
The IceCube Collaboration,
R. Abbasi,
M. Ackermann,
J. Adams,
J. A. Aguilar,
M. Ahlers,
J. M. Alameddine,
S. Ali,
N. M. Amin,
K. Andeen,
C. Argüelles,
Y. Ashida,
S. Athanasiadou,
S. N. Axani,
R. Babu,
X. Bai,
A. Balagopal V.,
S. W. Barwick,
V. Basu,
R. Bay,
J. J. Beatty,
J. Becker Tjus,
P. Behrens,
J. Beise,
C. Bellenghi
, et al. (396 additional authors not shown)
Abstract:
Gravitational-wave events from mergers of compact objects are a predicted source of high-energy neutrinos. Using data from the IceCube Neutrino Observatory, we search for neutrinos coincident with 85 significant and 945 low-significance gravitational-wave candidate events from compact binary coalescences published in real-time by the LIGO-Virgo-KAGRA collaboration during the first part of its four…
▽ More
Gravitational-wave events from mergers of compact objects are a predicted source of high-energy neutrinos. Using data from the IceCube Neutrino Observatory, we search for neutrinos coincident with 85 significant and 945 low-significance gravitational-wave candidate events from compact binary coalescences published in real-time by the LIGO-Virgo-KAGRA collaboration during the first part of its fourth observing run (O4a) and its preceding engineering run, within a time window of $\pm500$ seconds centered on the merger time. We report improvements to the online pipelines, including automatic sending of notices, which has decreased the IceCube real-time response time to gravitational-wave events. In addition, we search for long-duration neutrino emission (up to two weeks after the merger) from three candidate events: two neutron star-black hole mergers, and one low-significance gravitational-wave event with a possible subthreshold gamma-ray counterpart. We use two methods, both of which have been previously used to search for neutrino emission associated with gravitational-wave transients: an unbinned maximum likelihood analysis on significant alerts and a Bayesian analysis accounting for astrophysical priors on both significant and low-significance alerts. We find no statistically significant emission from any of the individual gravitational-wave events analyzed, and set upper limits on the time-integrated flux and energy emitted in high energy neutrinos assuming isotropic emission from each event.
△ Less
Submitted 11 June, 2026;
originally announced June 2026.
-
On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters
Authors:
Mind Lab,
:,
Vin Bo,
Song Cao,
Vic Cao,
Andrew Chen,
Kaijie Chen,
Cleon Cheng,
Steven Chiang,
Kaixuan Fan,
Hera Feng,
Huan Feng,
Arthur Fu,
Jun Gao,
Hongquan Gu,
Aaron Guan,
Nolan Ho,
Mutian Hong,
Hailee Hou,
Peixuan Hua,
Charles Huang,
Miles Jiang,
Nora Jiang,
Yuyi Jiang,
Qiuyu Jin
, et al. (42 additional authors not shown)
Abstract:
Parameter-efficient fine-tuning (PEFT) is usually treated as a cheaper alternative to full fine-tuning. We study a broader role: small trainable adapters as persistent local state on top of strong shared foundation models. In this framing, the base model provides shared competence while adapters carry instance-specific behavior such as preferences, skills, tool habits, and memory-like updates. We…
▽ More
Parameter-efficient fine-tuning (PEFT) is usually treated as a cheaper alternative to full fine-tuning. We study a broader role: small trainable adapters as persistent local state on top of strong shared foundation models. In this framing, the base model provides shared competence while adapters carry instance-specific behavior such as preferences, skills, tool habits, and memory-like updates. We organize the problem around three scaling axes: Scale Up, where stronger shared priors make small local updates more useful; Scale Down, where we study how small adapters can be while remaining reliable; and Scale Out, where many persistent adapted instances coexist. MinT provides one infrastructure example for managing adapter identity, revision, provenance, evaluation, and serving residency. Together, the results suggest that PEFT can be a compact substrate for persistent personal models rather than only a budget substitute for full fine-tuning.
△ Less
Submitted 2 June, 2026; v1 submitted 1 June, 2026;
originally announced June 2026.
-
Hallucination-Aware Diffusion Sampling for Inverse Problems via Robust Prior Updates
Authors:
Pengfei Jin,
Yiqi Tian,
Kailong Fan,
Bingjie Qi,
Quanzheng Li
Abstract:
Diffusion-based inverse problem solvers can produce realistic reconstructions, but realism alone does not ensure that the recovered details are supported by the measurement. We study this failure as measurement-conditioned hallucination: visually meaningful content that is either implausible or inconsistent with the measured instance. Our analysis separates Bayes-rule-based diffusion inverse solve…
▽ More
Diffusion-based inverse problem solvers can produce realistic reconstructions, but realism alone does not ensure that the recovered details are supported by the measurement. We study this failure as measurement-conditioned hallucination: visually meaningful content that is either implausible or inconsistent with the measured instance. Our analysis separates Bayes-rule-based diffusion inverse solvers into a prior update and a measurement-conditioning step, showing that hallucinated content can enter through the prior-side proposal before the measurement correction is applied. Motivated by this view, we propose Robust Prior Update (RPU), a solver-level module that probes the local stability of the diffusion prior update, re-anchors the resulting displacement at the current iterate, and leaves the measurement update unchanged. We instantiate RPU in DPS and evaluate it on FFHQ and ImageNet inverse problems using automatic metrics and human faithfulness studies. On FFHQ, RPU improves PSNR and LPIPS over DPS across box inpainting, Gaussian deblurring, and motion deblurring. In human judgments, RPU receives 91.9% of blind non-tie majority preferences and 91.1% of ground-truth-assisted non-tie preferences on FFHQ box inpainting, while the ImageNet Gaussian reader study is tie-heavy but favors RPU among non-tie cases. These results support a targeted claim: robustifying the prior update can improve instance faithfulness in diffusion inverse solvers, especially when the prior shapes weakly constrained content.
△ Less
Submitted 1 June, 2026;
originally announced June 2026.
-
IceCube Second Track Data Release IceTracks-DR2: Data from 2008-2022 for Neutrino Source Searches
Authors:
R. Abbasi,
M. Ackermann,
J. Adams,
J. A. Aguilar,
M. Ahlers,
J. M. Alameddine,
S. Ali,
N. M. Amin,
K. Andeen,
C. Argüelles,
Y. Ashida,
S. Athanasiadou,
S. N. Axani,
R. Babu,
X. Bai,
A. Balagopal V.,
S. W. Barwick,
V. Basu,
R. Bay,
J. J. Beatty,
J. Becker Tjus,
P. Behrens,
J. Beise,
C. Bellenghi,
S. Benkel
, et al. (390 additional authors not shown)
Abstract:
We present IceCube's latest release of muon track data for neutrino point-source searches, extending the previously published 10-year dataset to cover 14 years of observations (April 6, 2008 - May 23, 2022). This release features an updated event selection and improved detector calibration for data recorded after June 1, 2010. The release also includes binned instrument response functions and effe…
▽ More
We present IceCube's latest release of muon track data for neutrino point-source searches, extending the previously published 10-year dataset to cover 14 years of observations (April 6, 2008 - May 23, 2022). This release features an updated event selection and improved detector calibration for data recorded after June 1, 2010. The release also includes binned instrument response functions and effective areas, enabling the community to perform sensitive searches for steady and transient neutrino sources. We report on key science results obtained with this dataset using internal IceCube analysis tools and compare them to those derived from analyses based on the binned response functions included in this public release. To facilitate reproducible research, we provide benchmark results obtained using this data release and publicly available software. This release represents IceCube's most sensitive and comprehensive publicly available all-sky muon track dataset to date and should be preferred over previous releases.
△ Less
Submitted 18 May, 2026;
originally announced May 2026.
-
Observation of $η_c(1S)\to Σ^0\bar Σ^0$ and search for $h_c(1P)\to Σ^0\bar Σ^0$ via $ψ(3686)$ transitions
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
M. S. Anderson,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone
, et al. (758 additional authors not shown)
Abstract:
Using $(2712.4 \pm 14.3) \times 10^6~ψ(3686)$ events collected by the BESIII detector operating at the BEPCII collider, the hadronic decay $η_{c}\toΣ^{0}\bar{Σ^{0}}$ is observed for the first time via the radiative transition from $ψ(3686)$. It is found that the branching fraction has a significant dependence on the interference pattern between $η_c(1S)$ and non-$η_c(1S)$ processes. They are deter…
▽ More
Using $(2712.4 \pm 14.3) \times 10^6~ψ(3686)$ events collected by the BESIII detector operating at the BEPCII collider, the hadronic decay $η_{c}\toΣ^{0}\bar{Σ^{0}}$ is observed for the first time via the radiative transition from $ψ(3686)$. It is found that the branching fraction has a significant dependence on the interference pattern between $η_c(1S)$ and non-$η_c(1S)$ processes. They are determined to be $\displaystyle\mathcal{B}(η_c(1S) \to Σ^{0}\bar{Σ^{0}}) = (2.59 \pm 0.14(stat) \pm 0.44(syst)) \times 10^{-3}$ and $(1.18 \pm 0.12(stat) \pm 0.21(syst)) \times 10^{-3}$, for the destructive and constructive interference scenarios, respectively. No significant signal is observed for the decay $h_{c}\toΣ^{0}\bar{Σ^{0}}$ in the hadronic transition $ψ(3686)\toπ^0h_{c}$, and an upper limit on its branching fraction is set to be $1.02\times 10^{-4}$ at the 90\% confidence level.
△ Less
Submitted 16 May, 2026;
originally announced May 2026.
-
3D Skew-Normal Splatting
Authors:
Xiangru Wu,
Ke Fan,
Yanwei Fu
Abstract:
3D Gaussian Splatting (3DGS) has emerged as a leading representation for real-time novel view synthesis and has been widely adopted in various downstream applications. The core strength of 3DGS lies in its efficient kernel-based scene representation, where Gaussian primitives provide favorable mathematical and computational properties. However, under a finite primitive budget, the symmetric shape…
▽ More
3D Gaussian Splatting (3DGS) has emerged as a leading representation for real-time novel view synthesis and has been widely adopted in various downstream applications. The core strength of 3DGS lies in its efficient kernel-based scene representation, where Gaussian primitives provide favorable mathematical and computational properties. However, under a finite primitive budget, the symmetric shape of each primitive directly affects representation compactness, especially near asymmetric structures such as object boundaries and one-sided surfaces. Recent works have explored more complex kernel distributions; however, they either remain within the elliptical family or rely on hard truncation, which limits continuous shape control and introduces distributional discontinuities. In this paper, we propose Skew-Normal Splatting (SNS), which adopts the Azzalini Skew-Normal distribution as the fundamental primitive. By introducing a learnable and bounded skewness parameter, SNS can continuously interpolate between symmetric Gaussians and Half-Gaussian-like shapes, enabling flexible modeling of both sharp boundaries and interior regions. Moreover, SNS preserves analytical tractability under affine transformations and marginalization. This property allows seamless integration into existing Gaussian Splatting rasterization pipelines. Furthermore, to address the strong coupling between scale, rotation, and skewness parameters, we introduce a decoupled parameterization and a block-wise optimization strategy to enhance training stability and accuracy. Extensive experiments on standard novel-view synthesis benchmarks show that SNS consistently improves reconstruction quality over Gaussian and recent non-Gaussian kernels, with clearer benefits on sharp boundaries and thin or one-sided structures.
△ Less
Submitted 15 May, 2026; v1 submitted 14 May, 2026;
originally announced May 2026.
-
MinT: Managed Infrastructure for Training and Serving Millions of LLMs
Authors:
Mind Lab,
:,
Song Cao,
Vic Cao,
Andrew Chen,
Kaijie Chen,
Cleon Cheng,
Steven Chiang,
Kaixuan Fan,
Hera Feng,
Huan Feng,
Arthur Fu,
Jun Gao,
Hongquan Gu,
Aaron Guan,
Nolan Ho,
Mutian Hong,
Hailee Hou,
Peixuan Hua,
Charles Huang,
Miles Jiang,
Nora Jiang,
Yuyi Jiang,
Qiuyu Jin,
Fancy Kong
, et al. (38 additional authors not shown)
Abstract:
We present MindLab Toolkit (MinT), a managed infrastructure system for Low-Rank Adaptation (LoRA) post-training and online serving. MinT targets a setting where many trained policies are produced over a small number of expensive base-model deployments. Instead of materializing each policy as a merged full checkpoint, MinT keeps the base model resident and moves exported LoRA adapter revisions thro…
▽ More
We present MindLab Toolkit (MinT), a managed infrastructure system for Low-Rank Adaptation (LoRA) post-training and online serving. MinT targets a setting where many trained policies are produced over a small number of expensive base-model deployments. Instead of materializing each policy as a merged full checkpoint, MinT keeps the base model resident and moves exported LoRA adapter revisions through rollout, update, export, evaluation, serving, and rollback, hiding distributed training, serving, scheduling, and data movement behind a service interface. MinT scales this path along three axes. Scale Up extends LoRA RL to frontier-scale dense and MoE architectures, including MLA and DSA attention paths, with training and serving validated beyond 1T total parameters. Scale Down moves only the exported LoRA adapter, which can be under 1% of base-model size in rank-1 settings; adapter-only handoff reduces the measured step by 18.3x on a 4B dense model and 2.85x on a 30B MoE, while concurrent multi-policy GRPO shortens wall time by 1.77x and 1.45x without raising peak memory. Scale Out separates durable policy addressability from CPU/GPU working sets: a tensor-parallel deployment supports 10^6-scale addressable catalogs (measured single-engine sweeps through 100K) and thousand-adapter active waves at cluster scale, with cold loading treated as scheduled service work and packed MoE LoRA tensors improving live engine loading by 8.5-8.7x. MinT thus manages million-scale LoRA policy catalogs while training and serving selected adapter revisions over shared 1T-class base models.
△ Less
Submitted 26 May, 2026; v1 submitted 13 May, 2026;
originally announced May 2026.
-
$δ$-mem: Efficient Online Memory for Large Language Models
Authors:
Jingdi Lei,
Di Zhang,
Junxian Li,
Weida Wang,
Kaixuan Fan,
Xiang Liu,
Qihan Liu,
Xiaoteng Ma,
Baian Chen,
Soujanya Poria
Abstract:
Large language models increasingly need to accumulate and reuse historical information in long-term assistants and agent systems. Simply expanding the context window is costly and often fails to ensure effective context utilization. We propose $δ$-mem, a lightweight memory mechanism that augments a frozen full-attention backbone with a compact online state of associative memory. $δ$-mem compresses…
▽ More
Large language models increasingly need to accumulate and reuse historical information in long-term assistants and agent systems. Simply expanding the context window is costly and often fails to ensure effective context utilization. We propose $δ$-mem, a lightweight memory mechanism that augments a frozen full-attention backbone with a compact online state of associative memory. $δ$-mem compresses past information into a fixed-size state matrix updated by delta-rule learning, and uses its readout to generate low-rank corrections to the backbone's attention computation during generation. With only an $8\times8$ online memory state, $δ$-mem improves the average score to $1.10\times$ that of the frozen backbone and $1.15\times$ that of the strongest non-$δ$-mem memory baseline. It achieves larger gains on memory-heavy benchmarks, reaching $1.31\times$ on MemoryAgentBench and $1.20\times$ on LoCoMo, while largely preserving general capabilities. These results show that effective memory can be realized through a compact online state directly coupled with attention computation, without full fine-tuning, backbone replacement, or explicit context extension.
△ Less
Submitted 12 May, 2026;
originally announced May 2026.
-
Sensitivity Projections for Low-Mass Dark Matter Annihilation with the IceCube Upgrade
Authors:
R. Abbasi,
M. Ackermann,
J. Adams,
J. A. Aguilar,
M. Ahlers,
J. M. Alameddine,
S. Ali,
N. M. Amin,
K. Andeen,
C. Argüelles,
Y. Ashida,
S. Athanasiadou,
S. N. Axani,
R. Babu,
X. Bai,
A. Balagopal V.,
S. W. Barwick,
V. Basu,
R. Bay,
J. J. Beatty,
J. Becker Tjus,
P. Behrens,
J. Beise,
C. Bellenghi,
S. Benkel
, et al. (390 additional authors not shown)
Abstract:
The IceCube Upgrade, an extension designed to enhance the IceCube Neutrino Observatory's detection of neutrinos with energies between 1 GeV and 500 GeV, will markedly improve IceCube's sensitivity to low-mass dark matter scenarios. In this study, we present sensitivity projections for the IceCube Upgrade to neutrino fluxes arising from dark matter annihilation. In particular, we consider dark matt…
▽ More
The IceCube Upgrade, an extension designed to enhance the IceCube Neutrino Observatory's detection of neutrinos with energies between 1 GeV and 500 GeV, will markedly improve IceCube's sensitivity to low-mass dark matter scenarios. In this study, we present sensitivity projections for the IceCube Upgrade to neutrino fluxes arising from dark matter annihilation. In particular, we consider dark matter with masses between 3 GeV to 500 GeV from both the core of the Sun and the Galactic Center. These projections indicate that the IceCube Upgrade will enable stringent limits on dark matter in this parameter space, achieving leading sensitivities to some dark matter models with only three years of data taking.
△ Less
Submitted 7 May, 2026;
originally announced May 2026.
-
Defusing the Trigger: Tail-Risk-Informed Attention Rebalancing for LLM Backdoor Mitigation
Authors:
Kaisheng Fan,
Yishu Gao,
Xunzhu Tang,
Tegawendé F. Bissyandé,
Weizhe Zhang
Abstract:
Backdoored large language models (LLMs) exhibit attacker-specified behavior at inference time while retaining normal performance on benign inputs. Existing mitigations often require parameter updates and trusted clean data, or rely on auxiliary generation and repeated model execution, complicating deployment. Across diverse backdoor mechanisms, successful activations exhibit stronger tail concentr…
▽ More
Backdoored large language models (LLMs) exhibit attacker-specified behavior at inference time while retaining normal performance on benign inputs. Existing mitigations often require parameter updates and trusted clean data, or rely on auxiliary generation and repeated model execution, complicating deployment. Across diverse backdoor mechanisms, successful activations exhibit stronger tail concentration in attention over semantic-content tokens than benign inputs and unsuccessful trigger activations. This pattern provides a sample-internal control signal to selectively regulate suspicious attention dynamics. We propose TIARA, a tail-risk-informed attention rebalancing approach for inference-time LLM backdoor mitigation. TIARA filters structural attention sinks, aggregates sparse high-concentration events across rows and heads, and converts the risk signal into selective content-domain power smoothing and adaptive attention-mass reallocation. A constrained reconstruction writes valid attention distributions back before value aggregation. TIARA requires no parameter updates, auxiliary generation, additional target-model passes, or deployment-time clean reference sets. We evaluate TIARA across four backdoor paradigms on dense, reasoning-oriented, and sparse mixture-of-experts LLMs. Across three model families, TIARA reduces average macro ASR to 11.5%, outperforming the strongest no-update inference-time baseline by 7.2 percentage points while limiting clean-task degradation to at most 3.8 percentage points. Under standardized profiling, TIARA adds 12.9% end-to-end latency over a matched eager-attention baseline; the current unfused path is 23.3% slower than fused SDPA. Overall, TIARA establishes sample-conditional attention rebalancing as a practical inference-time control layer for mitigating attention-concentrated LLM backdoors.
△ Less
Submitted 29 July, 2026; v1 submitted 27 April, 2026;
originally announced April 2026.
-
A theory of ROC analysis of rule-out and rule-in diagnostics with applications to mammography data
Authors:
Michelle Mastrianni,
Kwok Lung Fan,
Yee Lam Elim Thompson,
Jessie J. J. Gommers,
Ioannis Sechopoulos,
Fredrik Strand,
Weijie Chen,
Gary Levine,
Mukul Sherekar,
Frank W. Samuelson
Abstract:
Multiple diagnostic tests are frequently used to determine the presence of a disease condition in patients. In this paper, we use bivariate copulas to examine the properties of receiver operating characteristic (ROC) curves formed when two correlated diagnostic tests are used together to rule-out ("believe the negative") and rule-in ("believe the positive") patients for disease. We use this theory…
▽ More
Multiple diagnostic tests are frequently used to determine the presence of a disease condition in patients. In this paper, we use bivariate copulas to examine the properties of receiver operating characteristic (ROC) curves formed when two correlated diagnostic tests are used together to rule-out ("believe the negative") and rule-in ("believe the positive") patients for disease. We use this theory to analyze three mammography data sets where AI devices are applied to reduce radiologists' workload or improve diagnostic performance. Our analysis shows with generality that increasing the radiologist-AI correlation for diseased cases enhances the area under the ROC curve (AUC) of a radiologist-AI rule-out curve, whereas decreasing correlation for non-diseased cases has a similar effect. The opposite trends hold for rule-in scenarios. Applications to clinical mammography data show that projected empirical radiologist performance under a rule-out or rule-in scenario is consistent with the theory.
△ Less
Submitted 25 April, 2026;
originally announced April 2026.
-
Neural posterior estimation of the neutrino direction in IceCube using transformer-encoded normalizing flows on the sphere
Authors:
R. Abbasi,
M. Ackermann,
J. Adams,
J. A. Aguilar,
M. Ahlers,
J. M. Alameddine,
S. Ali,
N. M. Amin,
K. Andeen,
C. Argüelles,
Y. Ashida,
S. Athanasiadou,
S. N. Axani,
R. Babu,
X. Bai,
A. Balagopal V.,
S. W. Barwick,
V. Basu,
R. Bay,
J. J. Beatty,
J. Becker Tjus,
P. Behrens,
J. Beise,
C. Bellenghi,
S. Benkel
, et al. (389 additional authors not shown)
Abstract:
IceCube is a cubic-kilometer-scale neutrino detector located at the geographic South Pole. A precise directional reconstruction of IceCube neutrinos is vital for associations with astronomical objects. In this context, we discuss neural posterior estimation of the neutrino direction via a transformer encoder that maps to a normalizing flow on the 2-sphere. It achieves a new state-of-the-art angula…
▽ More
IceCube is a cubic-kilometer-scale neutrino detector located at the geographic South Pole. A precise directional reconstruction of IceCube neutrinos is vital for associations with astronomical objects. In this context, we discuss neural posterior estimation of the neutrino direction via a transformer encoder that maps to a normalizing flow on the 2-sphere. It achieves a new state-of-the-art angular resolution for the two main event morphologies in IceCube - tracks and showers - while being significantly faster than traditional B-spline-based likelihood reconstructions. All-sky scans can be performed within seconds rather than hours, and take constant computation time, regardless of whether the posterior extent is arc-minutes or spans the whole sky. We utilize a combination of $C^2$-smooth rational-quadratic splines, scale transformations and rotations to define a novel spherical normalizing-flow distribution whose parameters are predicted as a whole as the output of the transformer encoder. We test several structural choices diverting from the vanilla transformer architecture. In particular, we find dual residual streams, nonlinear QKV projection and a separate class token with its own cross-attention processing to boost test-time performance. The angular resolution for both showers and tracks improves substantially over the whole trained energy range from 100 GeV to 100 PeV. At 100 TeV deposited energy, for example, the median angular resolution improves by a factor of $1.3$ for throughgoing tracks, by a factor of $1.7$ for showers and by a factor of $2.5$ for starting tracks compared to state-of-the art likelihood reconstructions based on B-splines. While previous machine-learning (ML) efforts have managed to obtain competitive shower resolutions, this is the first time an ML-based method outperforms likelihood-based muon reconstructions above 100 GeV.
△ Less
Submitted 21 April, 2026;
originally announced April 2026.
-
Observation of field-odd and field-free superconducting diode effects in $\mathrm{Mo}_2\mathrm{C}$ nanoflakes
Authors:
Wei Gao,
Kaixuan Fan,
Menghan Li,
Jinhao Cheng,
Peng Zhu,
Qing Zhang,
Shuaishuai Ding,
Wenping Hu,
Fan Yang,
Dechao Geng,
Hechen Ren
Abstract:
The superconducting diode effect (SDE) enables nonreciprocal supercurrent flow, holding immense potential for ultra-low-power quantum electronics. Intrinsic SDE typically requires materials with inherent symmetry breakings. Here, we report the discovery of SDE in chemical vapor deposition-grown molybdenum carbide ($\mathrm{Mo}_2\mathrm{C}$) nanoflakes, a material traditionally considered centrosym…
▽ More
The superconducting diode effect (SDE) enables nonreciprocal supercurrent flow, holding immense potential for ultra-low-power quantum electronics. Intrinsic SDE typically requires materials with inherent symmetry breakings. Here, we report the discovery of SDE in chemical vapor deposition-grown molybdenum carbide ($\mathrm{Mo}_2\mathrm{C}$) nanoflakes, a material traditionally considered centrosymmetric. Strikingly, this system uniquely hosts both field-odd and field-free SDEs. Transport measurements reveal a field-odd SDE with tunable efficiency exceeding 40% at 4 K under a perpendicular in-plane magnetic field. In a separate sample, a robust field-free SDE persists under zero-field and field-coolings. Out-of-plane field sweeps confirm the intrinsic nature of these phenomena. We propose that domain-boundary supercurrents or charge density wave-like orders drive this unexpected combination of symmetry breakings. Our findings establish air-stable $\mathrm{Mo}_2\mathrm{C}$ as an ideal platform for nonreciprocal superconducting electronics operating at liquid-helium temperatures, expanding the search for SDE into nominally centrosymmetric superconductors.
△ Less
Submitted 4 May, 2026; v1 submitted 21 April, 2026;
originally announced April 2026.
-
OpenGame: Open Agentic Coding for Games
Authors:
Yilei Jiang,
Jinyuan Hu,
Qianyin Xiao,
Yaozhi Zheng,
Ruize Ma,
Kaituo Feng,
Jiaming Han,
Tianshuo Peng,
Kaixuan Fan,
Manyuan Zhang,
Xiangyu Yue
Abstract:
Game development sits at the intersection of creative design and intricate software engineering, demanding the joint orchestration of game engines, real-time loops, and tightly coupled state across many files. While Large Language Models (LLMs) and code agents now solve isolated programming tasks with ease, they consistently stumble when asked to produce a fully playable game from a high-level des…
▽ More
Game development sits at the intersection of creative design and intricate software engineering, demanding the joint orchestration of game engines, real-time loops, and tightly coupled state across many files. While Large Language Models (LLMs) and code agents now solve isolated programming tasks with ease, they consistently stumble when asked to produce a fully playable game from a high-level design, collapsing under cross-file inconsistencies, broken scene wiring, and logical incoherence. We bridge this gap with OpenGame, the first open-source agentic framework explicitly designed for end-to-end web game creation. At its core lies Game Skill, a reusable, evolving capability composed of a Template Skill that grows a library of project skeletons from experience and a Debug Skill that maintains a living protocol of verified fixes - together enabling the agent to scaffold stable architectures and systematically repair integration errors rather than patch isolated syntax bugs. Powering this framework is GameCoder-27B, a code LLM specialized for game engine mastery through a three-stage pipeline of continual pre-training, supervised fine-tuning, and execution-grounded reinforcement learning. Since verifying interactive playability is fundamentally harder than checking static code, we further introduce OpenGame-Bench, an evaluation pipeline that scores agentic game generation along Build Health, Visual Usability, and Intent Alignment via headless browser execution and VLM judging. Across 150 diverse game prompts, OpenGame establishes a new state-of-the-art. We hope OpenGame pushes code agents beyond discrete software engineering problems and toward building complex, interactive real-world applications. Our framework will be fully open-sourced.
△ Less
Submitted 20 April, 2026;
originally announced April 2026.
-
OFlow: Injecting Object-Aware Temporal Flow Matching for Robust Robotic Manipulation
Authors:
Kuanning Wang,
Ke Fan,
Chenhao Qiu,
Zeyu Shangguan,
Yuqian Fu,
Yanwei Fu,
Daniel Seita,
Xiangyang Xue
Abstract:
Robust robotic manipulation requires not only predicting how the scene evolves over time, but also recognizing task-relevant objects in complex scenes. However, existing VLA models face two limitations. They typically act only on the current frame, while future prediction and object-aware reasoning are often learned in separate latent spaces. We propose OFlow (injecting Object-Aware Temporal Flow…
▽ More
Robust robotic manipulation requires not only predicting how the scene evolves over time, but also recognizing task-relevant objects in complex scenes. However, existing VLA models face two limitations. They typically act only on the current frame, while future prediction and object-aware reasoning are often learned in separate latent spaces. We propose OFlow (injecting Object-Aware Temporal Flow Matching into VLAs), a framework that addresses both limitations by unifying temporal foresight and object-aware reasoning in a shared semantic latent space. Our method forecasts future latents with temporal flow matching, factorizes them into object-aware representations that emphasize physically relevant cues while filtering task-irrelevant variation, and conditions continuous action generation on these predictions. By integrating OFlow into VLA pipelines, our method enables more reliable control under distribution shifts. Extensive experiments across LIBERO, LIBERO-Plus, MetaWorld, and SimplerEnv benchmarks and real-world tasks demonstrate that object-aware foresight consistently enhances robustness and success.
△ Less
Submitted 20 April, 2026;
originally announced April 2026.
-
GGD-SLAM: Monocular 3DGS SLAM Powered by Generalizable Motion Model for Dynamic Environments
Authors:
Yi Liu,
Haoxuan Xu,
Hongbo Duan,
Keyu Fan,
Zhengyang Zhang,
Peiyu Zhuang,
Pengting Luo,
Houde Liu
Abstract:
Visual SLAM algorithms achieve significant improvements through the exploration of 3D Gaussian Splatting (3DGS) representations, particularly in generating high-fidelity dense maps. However, they depend on a static environment assumption and experience significant performance degradation in dynamic environments. This paper presents GGD-SLAM, a framework that employs a generalizable motion model to…
▽ More
Visual SLAM algorithms achieve significant improvements through the exploration of 3D Gaussian Splatting (3DGS) representations, particularly in generating high-fidelity dense maps. However, they depend on a static environment assumption and experience significant performance degradation in dynamic environments. This paper presents GGD-SLAM, a framework that employs a generalizable motion model to address the challenges of localization and dense mapping in dynamic environments - without predefined semantic annotations or depth input. Specifically, the proposed system employs a First-In-First-Out (FIFO) queue to manage incoming frames, facilitating dynamic semantic feature extraction through a sequential attention mechanism. This is integrated with a dynamic feature enhancer to separate static and dynamic components. Additionally, to minimize dynamic distractors' impact on the static components, we devise a method to fill occluded areas via static information sampling and design a distractor-adaptive Structure Similarity Index Measure (SSIM) loss tailored for dynamic environments, significantly enhancing the system's resilience. Experiments conducted on real-world dynamic datasets demonstrate that the proposed system achieves state-of-the-art performance in camera pose estimation and dense reconstruction in dynamic scenes.
△ Less
Submitted 14 April, 2026;
originally announced April 2026.
-
LoViF 2026 Challenge on Human-oriented Semantic Image Quality Assessment: Methods and Results
Authors:
Xin Li,
Daoli Xu,
Wei Luo,
Guoqiang Xiang,
Haoran Li,
Chengyu Zhuang,
Zhibo Chen,
Jian Guan,
Weiping Li,
Weixia Zhang,
Wei Sun,
Zhihua Wang,
Dandan Zhu,
Chengguang Zhu,
Ayush Gupta,
Rachit Agarwal,
Shouvik Das,
Biplab Ch Das,
Amartya Ghosh,
Kanglong Fan,
Wen Wen,
Shuyan Zhai,
Tianwu Zhi,
Aoxiang Zhang,
Jianzhao Liu
, et al. (5 additional authors not shown)
Abstract:
This paper reviews the LoViF 2026 Challenge on Human-oriented Semantic Image Quality Assessment. This challenge aims to raise a new direction, i.e., how to evaluate the loss of semantic information from the human perspective, intending to promote the development of some new directions, like semantic coding, processing, and semantic-oriented optimization, etc. Unlike existing datasets of quality as…
▽ More
This paper reviews the LoViF 2026 Challenge on Human-oriented Semantic Image Quality Assessment. This challenge aims to raise a new direction, i.e., how to evaluate the loss of semantic information from the human perspective, intending to promote the development of some new directions, like semantic coding, processing, and semantic-oriented optimization, etc. Unlike existing datasets of quality assessment, we form a dataset of human-oriented semantic quality assessment, termed the SeIQA dataset. This dataset is divided into three parts for this competition: (i) training data: 510 pairs of degraded images and their corresponding ground truth references; (ii) validation data: 80 pairs of degraded images and their corresponding ground-truth references; (iii) testing data: 160 pairs of degraded images and their corresponding ground-truth references. The primary objective of this challenge is to establish a new and powerful benchmark for human-oriented semantic image quality assessment. There are a total of 58 teams registered in this competition, and 6 teams submitted valid solutions and fact sheets for the final testing phase. These submissions achieved state-of-the-art (SOTA) performance on the SeIQA dataset.
△ Less
Submitted 3 August, 2026; v1 submitted 13 April, 2026;
originally announced April 2026.
-
Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation
Authors:
Zunhai Su,
Hengyuan Zhang,
Wei Wu,
Yifan Zhang,
Yaxiu Liu,
He Xiao,
Qingyao Yang,
Yuxuan Sun,
Rui Yang,
Chao Zhang,
Jing Xiong,
Hui Shen,
Keyu Fan,
Weihao Ye,
Chaofan Tao,
Taiqiang Wu,
Zhongwei Wan,
Tiantian Zhang,
Bowen Yan,
Zhen Li,
Yiming Zhang,
Congkai Xie,
Yulei Qian,
Yuchen Xie,
Yik-Chung Wu
, et al. (2 additional authors not shown)
Abstract:
As the foundational architecture of modern machine learning, Transformers have driven remarkable progress across diverse AI domains. Despite their transformative impact, a persistent challenge across various Transformers is Attention Sink (AS), in which a disproportionate amount of attention is focused on a small subset of specific yet uninformative tokens. AS complicates interpretability, signifi…
▽ More
As the foundational architecture of modern machine learning, Transformers have driven remarkable progress across diverse AI domains. Despite their transformative impact, a persistent challenge across various Transformers is Attention Sink (AS), in which a disproportionate amount of attention is focused on a small subset of specific yet uninformative tokens. AS complicates interpretability, significantly affecting the training and inference dynamics, and exacerbates issues such as hallucinations. In recent years, substantial research has been dedicated to understanding and harnessing AS. However, a comprehensive survey that systematically consolidates AS-related research and offers guidance for future advancements remains lacking. To address this gap, we present the first survey on AS, structured around three key dimensions that define the current research landscape: Fundamental Utilization, Mechanistic Interpretation, and Strategic Mitigation. Our work makes a pivotal contribution by highlighting the key concepts and main trends in the field, guiding researchers through the evolution of AS-related studies. We envision this survey as a valuable resource, empowering researchers to effectively manage AS within the current Transformer paradigm, while simultaneously inspiring innovative advancements for the next generation of Transformers. The paper list of this work is available at https://github.com/ZunhaiSu/Awesome-Attention-Sink.
△ Less
Submitted 5 June, 2026; v1 submitted 11 April, 2026;
originally announced April 2026.
-
Next-Scale Generative Reranking: A Tree-based Generative Rerank Method at Meituan
Authors:
Shuli Wang,
Changhao Li,
Ke Fan,
Senjie Kou Junwei Yin,
Chi Wang,
Yinhua Zhu,
Haitao Wang,
Xingxing Wang
Abstract:
In modern multi-stage recommendation systems, reranking plays a critical role by modeling contextual information. Due to inherent challenges such as the combinatorial space complexity, an increasing number of methods adopt the generative paradigm: the generator produces the optimal list during inference, while an evaluator guides the generator's optimization during the training phase. However, the…
▽ More
In modern multi-stage recommendation systems, reranking plays a critical role by modeling contextual information. Due to inherent challenges such as the combinatorial space complexity, an increasing number of methods adopt the generative paradigm: the generator produces the optimal list during inference, while an evaluator guides the generator's optimization during the training phase. However, these methods still face two problems. Firstly, these generators fail to produce optimal generation results due to the lack of both local and global perspectives, regardless of whether the generation strategy is autoregressive or non-autoregressive. Secondly, the goal inconsistency problem between the generator and the evaluator during training complicates the guidance signal and leading to suboptimal performance. To address these issues, we propose the \textbf{N}ext-\textbf{S}cale \textbf{G}eneration \textbf{R}eranking (NSGR), a tree-based generative framework. Specifically, we introduce a next-scale generator (NSG) that progressively expands a recommendation list from user interests in a coarse-to-fine manner, balancing global and local perspectives. Furthermore, we design a multi-scale neighbor loss, which leverages a tree-based multi-scale evaluator (MSE) to provide scale-specific guidance to the NSG at each scale. Extensive experiments on public and industrial datasets validate the effectiveness of NSGR. And NSGR has been successfully deployed on the Meituan food delivery platform.
△ Less
Submitted 6 April, 2026;
originally announced April 2026.
-
Impact of gate voltage on switching field of perpendicular magnetic tunnel junctions with a synthetic antiferromagnetic free layer
Authors:
K. Fan,
S. V. Beek,
G. Talmelli,
V. Kateel,
D. Giuliano,
B. Vermeulen,
K. Cai,
B. Sorée,
J. D. Boeck,
R. Carpenter,
S. Rao,
S. Couet,
V. D. Nguyen,
G. S. Kar
Abstract:
We present micromagnetic simulations and experiments on voltage-assisted field switching in perpendicular magnetic tunnel junctions (MTJs) with a synthetic antiferromagnetic (SAF) free layer, where the magnetic state of one sublayer is detected via tunneling magnetoresistance (TMR). Simulations reveal that local modulation of perpendicular magnetic anisotropy (PMA) in one SAF sublayer leads to dis…
▽ More
We present micromagnetic simulations and experiments on voltage-assisted field switching in perpendicular magnetic tunnel junctions (MTJs) with a synthetic antiferromagnetic (SAF) free layer, where the magnetic state of one sublayer is detected via tunneling magnetoresistance (TMR). Simulations reveal that local modulation of perpendicular magnetic anisotropy (PMA) in one SAF sublayer leads to distinct switching characteristics. The switching field varies linearly with the anisotropy field, indicating voltage-controlled magnetic anisotropy (VCMA)-dominated dynamics similar to single free-layer devices. We then experimentally study the magnetic switching field of MTJ devices with SAF free layers under applied gate voltage. By varying the MgO tunnel barrier thickness to systematically modulate the resistance-area (RA) product, we enable quantitative separation of spin-transfer torque (STT), VCMA, and Joule heating contributions. Our findings indicate that VCMA dominates in devices with a high RA product, while low-RA devices exhibit nonlinear switching behavior due to enhanced contributions from STT and Joule heating. Furthermore, the effective fields derived from STT, VCMA, and Joule heating contributions under various gate voltages show minimal dependence on device critical dimensions, indicating favorable scaling behavior. This work presents a unified framework analyzing the roles of STT, VCMA, and Joule heating in SAF-based voltage-gated spin-orbit torque (SOT) MRAM, offering key insights for the optimization of performance, energy efficiency, and scalability in SOT-MRAM technologies.
△ Less
Submitted 1 April, 2026;
originally announced April 2026.
-
Gen-Searcher: Reinforcing Agentic Search for Image Generation
Authors:
Kaituo Feng,
Manyuan Zhang,
Shuang Chen,
Yunlong Lin,
Kaixuan Fan,
Yilei Jiang,
Hongyu Li,
Dian Zheng,
Chenyang Wang,
Xiangyu Yue
Abstract:
Recent image generation models have shown strong capabilities in generating high-fidelity and photorealistic images. However, they are fundamentally constrained by frozen internal knowledge, thus often failing on real-world scenarios that are knowledge-intensive or require up-to-date information. In this paper, we present Gen-Searcher, as the first attempt to train a search-augmented image generat…
▽ More
Recent image generation models have shown strong capabilities in generating high-fidelity and photorealistic images. However, they are fundamentally constrained by frozen internal knowledge, thus often failing on real-world scenarios that are knowledge-intensive or require up-to-date information. In this paper, we present Gen-Searcher, as the first attempt to train a search-augmented image generation agent, which performs multi-hop reasoning and search to collect the textual knowledge and reference images needed for grounded generation. To achieve this, we construct a tailored data pipeline and curate two high-quality datasets, Gen-Searcher-SFT-10k and Gen-Searcher-RL-6k, containing diverse search-intensive prompts and corresponding ground-truth synthesis images. We further introduce KnowGen, a comprehensive benchmark that explicitly requires search-grounded external knowledge for image generation and evaluates models from multiple dimensions. Based on these resources, we train Gen-Searcher with SFT followed by agentic reinforcement learning with dual reward feedback, which combines text-based and image-based rewards to provide more stable and informative learning signals for GRPO training. Experiments show that Gen-Searcher brings substantial gains, improving Qwen-Image by around 16 points on KnowGen and 15 points on WISE. We hope this work can serve as an open foundation for search agents in image generation, and we fully open-source our data, models, and code.
△ Less
Submitted 22 May, 2026; v1 submitted 30 March, 2026;
originally announced March 2026.
-
Amplitude analysis and branching fraction measurement of the decay $D^0 \to K^+K^-π^0π^0$
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
M. S. Anderson,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone
, et al. (749 additional authors not shown)
Abstract:
An amplitude analysis of the singly Cabibbo-suppressed decay $D^0 \to K^+ K^- π^0 π^0$ is performed, for the first time, to determine the relative magnitudes and phases of different intermediate processes. The analysis uses $e^+e^-$ collision data collected with the BESIII detector at the center-of-mass energy 3.773~GeV corresponding to an integrated luminosity of 20.3 $\rm fb^{-1}$. The absolute…
▽ More
An amplitude analysis of the singly Cabibbo-suppressed decay $D^0 \to K^+ K^- π^0 π^0$ is performed, for the first time, to determine the relative magnitudes and phases of different intermediate processes. The analysis uses $e^+e^-$ collision data collected with the BESIII detector at the center-of-mass energy 3.773~GeV corresponding to an integrated luminosity of 20.3 $\rm fb^{-1}$. The absolute branching fraction of $D^0 \to K^+ K^- π^0 π^0$ is measured to be \BF. The dominant intermediate process is $D^0 \to K^{*}(892)^+K^{*}(892)^-$, with a branching fraction of $(2.79 \pm 0.13_{\rm{stat.}} \pm 0.11_{\rm{syst.}}) \times 10^{-3}$. Amplitude analysis reveals that the $D^0 \to K^{*}(892)^+K^{*}(892)^-$ decay is S-wave dominant. The longitudinal polarization fraction of $D^0 \to K^{*}(892)^+ K^{*}(892)^-$ is measured to be $0.468\pm0.046_{\rm{stat.}}\pm0.011_{\rm{syst.}}$.
△ Less
Submitted 30 March, 2026; v1 submitted 26 March, 2026;
originally announced March 2026.
-
$p$-adic multiple zeta values of integer indices
Authors:
Ku-Yu Fan
Abstract:
This paper concerns the $p$-adic multiple zeta values of integer indices that may contain zero or negative components. We introduce the admissibility and regularizability conditions for integer indices. We define the $p$-adic multiple zeta values associated with admissible integer indices to be finite rational linear combinations of $p$-adic multiple zeta values associated with admissible positive…
▽ More
This paper concerns the $p$-adic multiple zeta values of integer indices that may contain zero or negative components. We introduce the admissibility and regularizability conditions for integer indices. We define the $p$-adic multiple zeta values associated with admissible integer indices to be finite rational linear combinations of $p$-adic multiple zeta values associated with admissible positive integer indices. We prove that the double shuffle relations, that is, the shuffle and stuffle product formulas, both hold for the values.
△ Less
Submitted 24 March, 2026;
originally announced March 2026.
-
Mini-review of charmonium weak decays at BESIII
Authors:
Xuze Li,
Kaixin Fan,
Zhengyun You,
Yu Zhang,
Minggang Zhao
Abstract:
The weak decays of charmonium states such as $J/ψ$ and $ψ(2S)$ are instrumental in probing both nonperturbative QCD dynamics and the flavor structure of the Standard Model (SM). The extreme rarity of charmonium weak decays renders them highly sensitive to physics beyond the SM, particularly in channels that are heavily suppressed in the SM, such as flavor-changing neutral-current (FCNC) decays. Th…
▽ More
The weak decays of charmonium states such as $J/ψ$ and $ψ(2S)$ are instrumental in probing both nonperturbative QCD dynamics and the flavor structure of the Standard Model (SM). The extreme rarity of charmonium weak decays renders them highly sensitive to physics beyond the SM, particularly in channels that are heavily suppressed in the SM, such as flavor-changing neutral-current (FCNC) decays. This review highlights the critical role of the BESIII experiment, which leverages an unprecedented data sample of over $10^{10}$ $J/ψ$ and $2.7\times10^{9}$ $ψ(2S)$ events to achieve leading sensitivity in searches for charmonium weak decays. We present the latest and most stringent upper limits established by BESIII on various semileptonic, nonleptonic, and FCNC charmonium weak decay channels.
△ Less
Submitted 16 May, 2026; v1 submitted 23 March, 2026;
originally announced March 2026.
-
ME-IQA: Memory-Enhanced Image Quality Assessment via Re-Ranking
Authors:
Kanglong Fan,
Tianhe Wu,
Wen Wen,
Jianzhao Liu,
Le Yang,
Yabin Zhang,
Yiting Liao,
Junlin Li,
Li Zhang
Abstract:
Reasoning-induced vision-language models (VLMs) advance image quality assessment (IQA) with textual reasoning, yet their scalar scores often lack sensitivity and collapse to a few values, so-called discrete collapse. We introduce ME-IQA, a plug-and-play, test-time memory-enhanced re-ranking framework. It (i) builds a memory bank and retrieves semantically and perceptually aligned neighbors using r…
▽ More
Reasoning-induced vision-language models (VLMs) advance image quality assessment (IQA) with textual reasoning, yet their scalar scores often lack sensitivity and collapse to a few values, so-called discrete collapse. We introduce ME-IQA, a plug-and-play, test-time memory-enhanced re-ranking framework. It (i) builds a memory bank and retrieves semantically and perceptually aligned neighbors using reasoning summaries, (ii) reframes the VLM as a probabilistic comparator to obtain pairwise preference probabilities and fuse this ordinal evidence with the initial score under Thurstone's Case V model, and (iii) performs gated reflection and consolidates memory to improve future decisions. This yields denser, distortion-sensitive predictions and mitigates discrete collapse. Experiments across multiple IQA benchmarks show consistent gains over strong reasoning-induced VLM baselines, existing non-reasoning IQA methods, and test-time scaling alternatives.
△ Less
Submitted 16 July, 2026; v1 submitted 21 March, 2026;
originally announced March 2026.
-
Controllable Text-to-Motion Generation via Modular Body-Part Phase Control
Authors:
Minyue Dai,
Ke Fan,
Anyi Rao,
Jingbo Wang,
Bo Dai
Abstract:
Text-to-motion (T2M) generation is becoming a practical tool for animation and interactive avatars. However, modifying specific body parts while maintaining overall motion coherence remains challenging. Existing methods typically rely on cumbersome, high-dimensional joint constraints (e.g., trajectories), which hinder user-friendly, iterative refinement. To address this, we propose Modular Body-Pa…
▽ More
Text-to-motion (T2M) generation is becoming a practical tool for animation and interactive avatars. However, modifying specific body parts while maintaining overall motion coherence remains challenging. Existing methods typically rely on cumbersome, high-dimensional joint constraints (e.g., trajectories), which hinder user-friendly, iterative refinement. To address this, we propose Modular Body-Part Phase Control, a plug-and-play framework enabling structured, localized editing via a compact, scalar-based phase interface. By modeling body-part latent motion channels as sinusoidal phase signals characterized by amplitude, frequency, phase shift, and offset, we extract interpretable codes that capture part-specific dynamics. A modular Phase ControlNet branch then injects this signal via residual feature modulation, seamlessly decoupling control from the generative backbone. Experiments on both diffusion- and flow-based models demonstrate that our approach provides predictable and fine-grained control over motion magnitude, speed, and timing. It preserves global motion coherence and offers a practical paradigm for controllable T2M generation. Project page: https://jixiii.github.io/bp-phase-project-page/
△ Less
Submitted 20 March, 2026;
originally announced March 2026.
-
SegviGen: Repurposing 3D Generative Model for Part Segmentation
Authors:
Lin Li,
Haoran Feng,
Zehuan Huang,
Haohua Chen,
Wenbo Nie,
Shaohua Hou,
Keqing Fan,
Pan Hu,
Sheng Wang,
Buyu Li,
Lu Sheng
Abstract:
We introduce SegviGen, a framework that repurposes native 3D generative models for 3D part segmentation. Existing pipelines either lift strong 2D priors into 3D via distillation or multi-view mask aggregation, often suffering from cross-view inconsistency and blurred boundaries, or explore native 3D discriminative segmentation, which typically requires large-scale annotated 3D data and substantial…
▽ More
We introduce SegviGen, a framework that repurposes native 3D generative models for 3D part segmentation. Existing pipelines either lift strong 2D priors into 3D via distillation or multi-view mask aggregation, often suffering from cross-view inconsistency and blurred boundaries, or explore native 3D discriminative segmentation, which typically requires large-scale annotated 3D data and substantial training resources. In contrast, SegviGen leverages the structured priors encoded in pretrained 3D generative model to induce segmentation through distinctive part colorization, establishing a novel and efficient framework for part segmentation. Specifically, SegviGen encodes a 3D asset and predicts part-indicative colors on active voxels of a geometry-aligned reconstruction. It supports interactive part segmentation, full segmentation, and full segmentation with 2D guidance in a unified framework. Extensive experiments show that SegviGen improves over the prior state of the art by 40% on interactive part segmentation and by 15% on full segmentation, while using only 0.32% of the labeled training data. It demonstrates that pretrained 3D generative priors transfer effectively to 3D part segmentation, enabling strong performance with limited supervision. See our project page at https://fenghora.github.io/SegviGen-Page/.
△ Less
Submitted 10 May, 2026; v1 submitted 17 March, 2026;
originally announced March 2026.
-
OCRA: Object-Centric Learning with 3D and Tactile Priors for Human-to-Robot Action Transfer
Authors:
Kuanning Wang,
Ke Fan,
Yuqian Fu,
Siyu Lin,
Hu Luo,
Daniel Seita,
Yanwei Fu,
Yu-Gang Jiang,
Xiangyang Xue
Abstract:
We present OCRA, an Object-Centric framework for video-based human-to-Robot Action transfer that learns directly from human demonstration videos to enable robust manipulation. Object-centric learning emphasizes task-relevant objects and their interactions while filtering out irrelevant background, providing a natural and scalable way to teach robots. OCRA leverages multi-view RGB videos, the state…
▽ More
We present OCRA, an Object-Centric framework for video-based human-to-Robot Action transfer that learns directly from human demonstration videos to enable robust manipulation. Object-centric learning emphasizes task-relevant objects and their interactions while filtering out irrelevant background, providing a natural and scalable way to teach robots. OCRA leverages multi-view RGB videos, the state-of-the-art 3D foundation model VGGT, and advanced detection and segmentation models to reconstruct object-centric 3D point clouds, capturing rich interactions between objects. To handle properties not easily perceived by vision alone, we incorporate tactile priors via a large-scale dataset of over one million tactile images. These 3D and tactile priors are fused through a multimodal module (ResFiLM) and fed into a Diffusion Policy to generate robust manipulation actions. Extensive experiments on both vision-only and visuo-tactile tasks show that OCRA significantly outperforms existing baselines and ablations, demonstrating its effectiveness for learning from human demonstration videos.
△ Less
Submitted 15 March, 2026;
originally announced March 2026.
-
MAPLE: Elevating Medical Reasoning from Statistical Consensus to Process-Led Alignment
Authors:
Kailong Fan,
Anqi Pu,
Yichen Wu,
Wanhua Li,
Yicong Li,
Hanspeter Pfister,
Huafeng Liu,
Xiang Li,
Quanzheng Li,
Ning Guo
Abstract:
Recent advances in medical large language models have explored Test-Time Reinforcement Learning (TTRL) to enhance reasoning. However, standard TTRL often relies on majority voting (MV) as a heuristic supervision signal, which can be unreliable in complex medical scenarios where the most frequent reasoning path is not necessarily the clinically correct one. In this work, we propose a novel and unif…
▽ More
Recent advances in medical large language models have explored Test-Time Reinforcement Learning (TTRL) to enhance reasoning. However, standard TTRL often relies on majority voting (MV) as a heuristic supervision signal, which can be unreliable in complex medical scenarios where the most frequent reasoning path is not necessarily the clinically correct one. In this work, we propose a novel and unified training paradigm that integrates medical process reward models with TTRL to bridge the gap between test-time scaling (TTS) and parametric model optimization. Specifically, we advance the TTRL framework by replacing the conventional MV with a fine-grained, expert-aligned supervision paradigm using Med-RPM. This integration ensures that reinforcement learning is guided by medical correctness rather than mere consensus, effectively distilling search-based intelligence into the model's parametric memory. Extensive evaluations on four different benchmarks have demonstrated that our developed method consistently and significantly outperforms current TTRL and standalone PRM selection. Our findings establish that transitioning from stochastic heuristics to structured, step-wise rewards is essential for developing reliable and scalable medical AI systems
△ Less
Submitted 9 March, 2026;
originally announced March 2026.
-
Search for the charmonium weak decay $ψ(2S)\to D_s^-π^+ + c.c.$ and $ψ(2S)\to D_s^-ρ^+ + c.c.$
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
Y. Bai,
O. Bakina,
Y. Ban,
H. -R. Bao,
X. L. Bao,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko
, et al. (723 additional authors not shown)
Abstract:
We search for the weak decays $ψ(2S)\to D_s^-π^+ + c.c.$ and $ψ(2S)\to D_s^-ρ^+ + c.c.$ for the first time. The search is based on $(2712.4\pm14.3)\times 10^6$ events containing the charmonium state $ψ(2S)$ collected at the center-of-mass energy $\sqrt{s}=3.686\ \rm{GeV}$ with the BESIII detector. This search offers a unique opportunity to test the Standard Model and search for new physics. Since…
▽ More
We search for the weak decays $ψ(2S)\to D_s^-π^+ + c.c.$ and $ψ(2S)\to D_s^-ρ^+ + c.c.$ for the first time. The search is based on $(2712.4\pm14.3)\times 10^6$ events containing the charmonium state $ψ(2S)$ collected at the center-of-mass energy $\sqrt{s}=3.686\ \rm{GeV}$ with the BESIII detector. This search offers a unique opportunity to test the Standard Model and search for new physics. Since no signal excess above the background is observed, the upper limits on the branching fractions at the 90\% confidence level are set to be $1.4\times 10^{-6}$ and $7.0\times 10^{-6}$ for $ψ(2S)\to D_s^-π^+ + c.c.$ and $ψ(2S)\to D_s^-ρ^+ + c.c.$, respectively.
△ Less
Submitted 4 June, 2026; v1 submitted 2 March, 2026;
originally announced March 2026.
-
DeepAFL: Deep Analytic Federated Learning
Authors:
Jianheng Tang,
Yajiang Huang,
Kejia Fan,
Feijiang Han,
Jiaxu Li,
Jinfeng Xu,
Run He,
Anfeng Liu,
Houbing Herbert Song,
Huiping Zhuang,
Yunhuai Liu
Abstract:
Federated Learning (FL) is a popular distributed learning paradigm to break down data silo. Traditional FL approaches largely rely on gradient-based updates, facing significant issues about heterogeneity, scalability, convergence, and overhead, etc. Recently, some analytic-learning-based work has attempted to handle these issues by eliminating gradient-based updates via analytical (i.e., closed-fo…
▽ More
Federated Learning (FL) is a popular distributed learning paradigm to break down data silo. Traditional FL approaches largely rely on gradient-based updates, facing significant issues about heterogeneity, scalability, convergence, and overhead, etc. Recently, some analytic-learning-based work has attempted to handle these issues by eliminating gradient-based updates via analytical (i.e., closed-form) solutions. Despite achieving superior invariance to data heterogeneity, these approaches are fundamentally limited by their single-layer linear model with a frozen pre-trained backbone. As a result, they can only achieve suboptimal performance due to their lack of representation learning capabilities. In this paper, to enable representable analytic models while preserving the ideal invariance to data heterogeneity for FL, we propose our Deep Analytic Federated Learning approach, named DeepAFL. Drawing inspiration from the great success of ResNet in gradient-based learning, we design gradient-free residual blocks in our DeepAFL with analytical solutions. We introduce an efficient layer-wise protocol for training our deep analytic models layer by layer in FL through least squares. Both theoretical analyses and empirical evaluations validate our DeepAFL's superior performance with its dual advantages in heterogeneity invariance and representation learning, outperforming state-of-the-art baselines by up to 5.68%-8.42% across three benchmark datasets.
△ Less
Submitted 28 February, 2026;
originally announced March 2026.
-
Available Energy and Ground States of Convective Hydrodynamic and Hydromagnetic Instabilities
Authors:
Kaixuan Fan,
Yao Zhou
Abstract:
We propose a method for predicting the nonlinear saturation level of convective instabilities in neutral and magnetized fluids. The method combines Gardner's restacking algorithm, which computes the available energy and ground states of collisionless plasmas in phase space, and Lagrangian relaxation, where fluid elements find lower-energy equilibria while preserving local invariants. For the incom…
▽ More
We propose a method for predicting the nonlinear saturation level of convective instabilities in neutral and magnetized fluids. The method combines Gardner's restacking algorithm, which computes the available energy and ground states of collisionless plasmas in phase space, and Lagrangian relaxation, where fluid elements find lower-energy equilibria while preserving local invariants. For the incompressible Rayleigh-Taylor instability, the problem is formally equivalent to Gardner's and the restacking algorithm directly applies in configuration space. To treat compressibility, we follow restacking with Lagrangian relaxation to obtain the ground state, and the results show excellent agreement with direct numerical simulations. Successful extension to the $m=0$ interchange instability in a Z-pinch demonstrates the method's potential as a general framework for estimating the nonlinear extent of convective instabilities, which can facilitate the design and operation of fusion reactors.
△ Less
Submitted 27 February, 2026;
originally announced February 2026.