-
SleuthTalk: Supporting Historical Photo Identification with Private Workspaces for Collective Sensemaking and Deliberation
Authors:
Liling Yuan,
Vikram Mohanty,
Kurt Luther
Abstract:
Identifying individuals in historical photographs is a critical task across fields such as history, journalism, genealogy, and archival research. While AI-based facial recognition can efficiently generate candidate matches, it often produces ambiguous results that require deeper analysis and contextual interpretation. Existing platforms lack robust support for collaborative deliberation, especiall…
▽ More
Identifying individuals in historical photographs is a critical task across fields such as history, journalism, genealogy, and archival research. While AI-based facial recognition can efficiently generate candidate matches, it often produces ambiguous results that require deeper analysis and contextual interpretation. Existing platforms lack robust support for collaborative deliberation, especially in uncertain or high-stakes cases. We present SleuthTalk, a private collaborative workspace integrated into Civil War Photo Sleuth, designed to scaffold structured comparison, discussion, and group decision-making. SleuthTalk enables users to curate custom shortlists, annotate facial features, and build consensus through structured feedback. In a mixed-methods evaluation with experienced historical photo researchers, SleuthTalk enhanced self-reported confidence, surfaced diverse perspectives, and supported transparent, reflective identifications.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
First measurements of the branching fractions of $J/ψ$ and $ψ(3686) \to Σ^{0} \barΣ^{0}η$
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko
, et al. (750 additional authors not shown)
Abstract:
Based on $(10087 \pm 44) \times 10^6$ $J/ψ$ and $(2712 \pm 14) \times 10^6$ $ψ(3686)$ events collected with the BESIII detector at the BEPCII collider, the hadronic decays $J/ψ\to Σ^{0} \barΣ^{0} η$ and $ψ(3686) \to Σ^{0} \barΣ^{0} η$ are observed for the first time. The corresponding branching fractions are measured to be…
▽ More
Based on $(10087 \pm 44) \times 10^6$ $J/ψ$ and $(2712 \pm 14) \times 10^6$ $ψ(3686)$ events collected with the BESIII detector at the BEPCII collider, the hadronic decays $J/ψ\to Σ^{0} \barΣ^{0} η$ and $ψ(3686) \to Σ^{0} \barΣ^{0} η$ are observed for the first time. The corresponding branching fractions are measured to be $\mathcal{B}(J/ψ\to Σ^{0} \barΣ^{0}η)= (7.5 \pm 0.3 \pm 0.8) \times 10^{-5}$ and $\mathcal{B}(ψ(3686) \to Σ^{0} \barΣ^{0}η)= (1.3\pm 0.1 \pm 0.1) \times 10^{-5}$, respectively, where the first uncertainties are statistical, and the second systematic. The ratio $\text{Q} \approx \frac{\mathcal{B}(ψ(3686) \to Σ^{0} \barΣ^{0} η)}{\mathcal{B}(J/ψ\to Σ^{0} \barΣ^{0} η)}$ is determined to be $(17.3 \pm 1.5 \pm 1.7)\%$, which is con sistent with the 12\%-rule within 3.0$σ$.~No significant intermediate states or threshold enhancements are observed in the $Σ^0$($\barΣ^{0}$)$η$ and $Σ^0$$\barΣ^{0}$ invariant mass spectra.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Measurement of Branching Fraction and Transition Magnetic Moment of the Hyperon Dalitz Decay $Σ^0 \rightarrow Λe^+e^-$
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
R. Aliberti,
A. Amoroso,
Q. An,
Y. Bai,
O. Bakina,
Y. Ban,
H. -R. Bao,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko,
R. A. Briere,
A. Brueggemann,
H. Cai
, et al. (683 additional authors not shown)
Abstract:
Based on a data sample of 10 billion $J/ψ$ events collected with the BESIII detector operating at the BEPCII collider, the Dalitz decay $Σ^0 \rightarrow Λe^+e^-$ is studied experimentally for the first time. The $Σ^0$ hyperons are produced through the process $J/ψ\rightarrow Σ^0\barΣ^0$ and analyzed using a double-tag method. The absolute branching fraction is measured to be…
▽ More
Based on a data sample of 10 billion $J/ψ$ events collected with the BESIII detector operating at the BEPCII collider, the Dalitz decay $Σ^0 \rightarrow Λe^+e^-$ is studied experimentally for the first time. The $Σ^0$ hyperons are produced through the process $J/ψ\rightarrow Σ^0\barΣ^0$ and analyzed using a double-tag method. The absolute branching fraction is measured to be $\mathcal{B}(Σ^0 \rightarrow Λe^+e^-) = (6.34 \pm 0.25_{\rm stat.} \pm 0.23_{\rm syst.}) \times 10^{-3}$. This result shows a $2σ$ discrepancy from the theoretical calculation quoted in the PDG, where the uncertainties are statistical and systematic, respectively. In addition to the branching fraction, the transition magnetic moment $μ$ is determined to be $(1.74 \pm 0.03_{\rm stat.} \pm 0.09_{\rm syst.})\,μ_N$, where $μ_N=e/(2m_p)$ represents the nucleon magnetic moment, providing valuable insight into the intrinsic structure of the $Σ^0$ hyperon.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture
Authors:
GigaBrain Team,
Angen Ye,
Axiang Sun,
Can Jin,
Chenxi Cheng,
Chong Shi,
Dengke Shang,
Dingqian Zhang,
Guan Huang,
Guangqiang Wang,
Guangqing Ding,
Guo Li,
Hangcong Li,
Hengyu Zhong,
Hongtao Lu,
Jianbo Qin,
Jiming Mao,
Jing Zhu,
Jindi Lv,
Jingzhi Cui,
Junjie Xie,
Junyi Bao,
Kai Liu,
Lei Yuan,
Limin Long
, et al. (34 additional authors not shown)
Abstract:
Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit from more effective architectural design, scale to substantially larger and more heterogeneous data regimes, and achieve broader generalizatio…
▽ More
Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit from more effective architectural design, scale to substantially larger and more heterogeneous data regimes, and achieve broader generalization across tasks and embodiments. To this end, we present GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments. Specifically, GigaBrain-0.7 unifies understanding, prediction, and action through a three-system architecture, scales pretraining to over 37,000 hours of heterogeneous embodied data, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation. Compared with the preceding GigaBrain-0 series and prior state-of-the-art models including $π_{0.5}$, GigaBrain-0.7 achieves substantial improvements in foundation zero-shot capabilities, language-conditioned instruction following, and post-training task success rates. In particular, on our in-house Maker H01 platform and mainstream robot embodiments, GigaBrain-0.7 demonstrates strong task adaptability and completion ability across both home and industrial scenarios. All training code and pretrained model weights will be released.
△ Less
Submitted 16 August, 2026;
originally announced August 2026.
-
Search for Magnetic and Spin-Independent Inelastic Dark Matter with XENONnT
Authors:
E. Aprile,
J. Aalbers,
K. Abe,
M. Abu Rmilah,
M. Adrover,
S. Ahmed Maouloud,
L. Althueser,
B. Andrieu,
E. Angelino,
D. Antón Martin,
S. R. Armbruster,
F. Arneodo,
L. Baudis,
M. Bazyk,
V. Beligotti,
L. Bellagamba,
R. Biondi,
A. Bismark,
K. Boese,
R. M. Braun,
G. Bruni,
R. Budnik,
C. Cai,
C. Capelli,
J. M. R. Cardoso
, et al. (151 additional authors not shown)
Abstract:
We present a search for Magnetic and Spin-Independent inelastic Dark Matter using 2.1 tonne-years of data from the XENONnT experiment. We consider both single- and double-site event topologies, targeting the unique signature of an initial nuclear recoil followed by a delayed de-excitation photon. To suppress backgrounds, we introduce novel directional and kinematic selections based on the inferred…
▽ More
We present a search for Magnetic and Spin-Independent inelastic Dark Matter using 2.1 tonne-years of data from the XENONnT experiment. We consider both single- and double-site event topologies, targeting the unique signature of an initial nuclear recoil followed by a delayed de-excitation photon. To suppress backgrounds, we introduce novel directional and kinematic selections based on the inferred speed and direction of the excited dark matter particle between the scatter and decay sites. We find that the collected data are consistent with background expectations, and report 90% C.L. upper limits for both models across the GeV/c2-TeV/c2 mass range.
△ Less
Submitted 15 August, 2026;
originally announced August 2026.
-
High-precision measurement of the space-like $η^\prime$ transition form factor
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
M. S. Anderson,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone
, et al. (758 additional authors not shown)
Abstract:
Using a data sample corresponding to an integrated luminosity of $20.3\ \text{fb}^{-1}$, collected with the BESIII detector at a center-of-mass energy of $3.773\ \text{GeV}$ at the BEPCII collider, we report a precision measurement of the product $Q^2|F(Q^2)|$, where $F(Q^2)$ is the single-virtual space-like transition form factor of the $η'$ meson and $Q^2$ is the squared momentum transfer of the…
▽ More
Using a data sample corresponding to an integrated luminosity of $20.3\ \text{fb}^{-1}$, collected with the BESIII detector at a center-of-mass energy of $3.773\ \text{GeV}$ at the BEPCII collider, we report a precision measurement of the product $Q^2|F(Q^2)|$, where $F(Q^2)$ is the single-virtual space-like transition form factor of the $η'$ meson and $Q^2$ is the squared momentum transfer of the tagged virtual photon. The transition form factor is extracted from the differential Born cross section of the two-photon fusion processes $e^+e^- \to e^+e^-γγ^* \to e^+e^-η^\prime$ using a single-tag technique, where only one scattered lepton is detected. The measurement covers $Q^2 \in [0.1, 6.0]$ GeV$^2$, achieving unprecedented precision, better than $3.0\%$ for $Q^2 < 1.5$ GeV$^2$, and providing the first direct determination at $Q^2 < 0.3$ GeV$^2$.
△ Less
Submitted 12 August, 2026;
originally announced August 2026.
-
Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling
Authors:
Xinmu Ge,
Zizhuo Zhang,
Yu Huang,
Jianing Zhu,
Lin Yuan,
Wanli Gu,
Weichang Wu,
Weiran Huang,
Xiaolu Zhang,
Bo Han,
Jun Zhou,
Jiangchao Yao
Abstract:
On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning. It is commonly believed to enable the student model to distill knowledge from a stronger teacher model, thereby expanding capabilities beyond the pre-OPD base model. In this study, we examine this view through the lens of test-time scaling by varying the sampling budget K and evaluating per…
▽ More
On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning. It is commonly believed to enable the student model to distill knowledge from a stronger teacher model, thereby expanding capabilities beyond the pre-OPD base model. In this study, we examine this view through the lens of test-time scaling by varying the sampling budget K and evaluating performance with pass@K and avg@K. Specifically, across several OPD variants, we observe that OPD-trained models maintain superior avg@K performance across sampling budgets, while the advantage in pass@K gradually shifts to the pre-OPD base models as K increases. These results suggest that OPD primarily improves sampling efficiency rather than consistently expanding the student's reasoning capability boundary. The pass@K dynamics throughout OPD training further reveal a progressive shift toward stronger small-K performance at the expense of the large-K capability boundary. Furthermore, a problem-level solvability analysis using pass@1024 as the criterion reveals an asymmetry: OPD causes more previously solvable problems to become unsolvable than previously unsolvable problems to become solvable. Together, these findings suggest that, from the perspective of capability expansion, OPD behaves more like an "illusory distillation": its apparent gains arise primarily from improved sampling efficiency rather than from acquiring genuinely new reasoning capabilities from the teacher.
△ Less
Submitted 12 August, 2026;
originally announced August 2026.
-
SHE: Trajectory-driven Safety Harness Evolution for LLM Agents
Authors:
Wanying Qu,
Qinghua Mao,
Yu Li,
Jiyao Liu,
Xin Zhang,
Dadi Guo,
Yanxu Zhu,
Qingyu Liu,
Leitao Yuan,
Xi Lin,
Shanfeng Zhu,
Yanwei Fu,
Jing Shao,
Xia Hu,
Dongrui Liu
Abstract:
The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control. Existing safety mechanisms often treat the harness as a fixed deployment artifact, limiting their ability to evolve with emerging risks. Moreover, coupled functions across harness components obscure safety responsibil…
▽ More
The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control. Existing safety mechanisms often treat the harness as a fixed deployment artifact, limiting their ability to evolve with emerging risks. Moreover, coupled functions across harness components obscure safety responsibility attribution, making localized evolution difficult. We propose Safety Harness Evolution (SHE), a framework that learns evolving safe boundaries from rollout trajectories. SHE decomposes the harness into four artifacts with explicit safety responsibilities, including the System Prompt, Rule Bank, Safety Memory, and Tool Policy, defining clear functional boundaries for localized evolution. Based on this decomposition, SHE introduces an attribution-guided evolution loop that converts trajectory failures into structured diagnoses, learns artifact-specific boundary refinements, and selects evolved harnesses through safety-utility validation. Experiments on Agent-SafetyBench demonstrate that SHE effectively enhances safety through harness evolution, achieving a 3.1x ASR reduction compared with static SafeHarness, while also improving benign utility. The evolved harness further generalizes to unseen risks on the held-out AgentHarm benchmark and transfers across agent models without additional evolution.
△ Less
Submitted 10 August, 2026;
originally announced August 2026.
-
Evaluating for the long term: Learnings from industry
Authors:
Leif Sigerson,
Tom Cunningham,
Winston Chou,
Sana Pandey,
Jonathan Stray,
Lo-Hua Yuan,
Eytan Bakshy,
Timothy Chan,
Molly Davies,
Maria Dimakopoulou,
Simon Ejdemyr,
Kenneth Hung,
Nathan Kallus,
Madhav Kumar,
Thu Le,
A. Demetri Pananos,
Lee Richardson,
Brennan Schaffner,
Rose Tan,
Martin Tingley,
Nadia Tomova,
Panagiotis Toulis,
Wenjing Zheng,
Zander Arnao,
Dean Eckles
Abstract:
Online platforms prioritize long-term business outcomes, yet typical experiments are far too short to measure these outcomes directly. Our goal in this paper is to collect and share industry knowledge on how to make decisions from short-term experiments that are better aligned with long-term outcomes. Based on a daylong workshop with 26 experts from 15 online platforms and 4 universities, we formu…
▽ More
Online platforms prioritize long-term business outcomes, yet typical experiments are far too short to measure these outcomes directly. Our goal in this paper is to collect and share industry knowledge on how to make decisions from short-term experiments that are better aligned with long-term outcomes. Based on a daylong workshop with 26 experts from 15 online platforms and 4 universities, we formulate a series of propositions that reflect current industry knowledge.
Participants largely agreed that reversals of sign from short-run to long-run treatment effects are rare, with reversals concentrating in specific cases such as treatments involving content quality signals, hyper-monetization, and pricing. Although the magnitude of treatment effects can shift over time, a "univariate autosurrogate", corresponding to the short-run treatment effect on the long-run metric of interest, is often hard to beat. A recurring theme was the importance of surrogates that are not only (or even primarily) unbiased for true long-run outcomes, but that improve decision-making. Thus, participants generally agreed that simple, interpretable surrogates were generally preferable to elaborate but hard-to-explain surrogate indices. Participants also agreed that, due to concerns about confounding and transportability, experimentally-learned surrogates are generally preferable to observationally-learned surrogates. However, the drawback is that learning good surrogates from experiments typically requires a large, representative portfolio of long-run experiments that few platforms possess.
We conclude that there is no substitute for a well-run long-term experiment, whether for learning surrogates or validating them, and we highlight open challenges including evolving treatments, persistent treatments not fully mediated by short-term proxies, and mismatch between experimental samples and the target population.
△ Less
Submitted 8 August, 2026;
originally announced August 2026.
-
CO Structures with Narrow Lines in Nearby Quiescent Regions
Authors:
Ruilin Xia,
Yang Su,
Shiyu Zhang,
Xuepeng Chen,
Ji Yang,
Yan Gong,
Yuehui Ma,
Yan Sun,
Min Fang,
Fujun Du,
Shaobo Zhang,
Xin Zhou,
Lixia Yuan,
Qing-Zeng Yan,
Li Sun,
Jiancheng Feng
Abstract:
Using CO data from Phase I of the Milky Way Imaging Scroll Painting (MWISP) survey, we present a systematic study of molecular structures with narrow lines. We identify 57 CO structures, most of which exhibit low densities and subsonic/transonic turbulence. Among them, structures with large projected areas and diffuse, sheet-like geometries are identified as veil clouds. The low LSR velocities and…
▽ More
Using CO data from Phase I of the Milky Way Imaging Scroll Painting (MWISP) survey, we present a systematic study of molecular structures with narrow lines. We identify 57 CO structures, most of which exhibit low densities and subsonic/transonic turbulence. Among them, structures with large projected areas and diffuse, sheet-like geometries are identified as veil clouds. The low LSR velocities and the concentration of these CO structures toward both the Galactic center (e.g., Ophiuchus, Aquila) and anticenter (e.g., Cepheus, Taurus) regions suggest a local origin for the sample, as supported by distance measurements of about 200--300pc for a subset with relatively large angular extents. These nearby structures likely arise from large-scale compression driven by past supernova activity within the Local Bubble. The observed low-velocity-dispersion emission may trace quiescent regions where turbulence has decayed due to a lack of sustained energy injection. For diffuse veil clouds with an assumed magnetic field of ~10uG, ion-neutral friction may provide an additional mechanism for turbulent dissipation on sub-parsec scales corresponding to their thickness of 0.1--0.3pc. Tracing the atomic-to-molecular transition, veil clouds provide a unique window into the diffuse, quiescent precursor state of dense gas. They likely represent a widespread but previously overlooked component of the Galactic molecular gas reservoir, with significant implications for cloud formation and evolution, the total mass budget and spatial distribution of molecular gas, and the initial conditions of star formation as a related consequence.
△ Less
Submitted 7 August, 2026;
originally announced August 2026.
-
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
Authors:
Tao Feng,
Fangxu Yu,
Haozhen Zhang,
Zhongjie Dai,
Liangqi Yuan,
Zijie Lei,
Weizhi Zhang,
Kunlun Zhu,
Haodong Yue,
Keyang Xuan,
Ge Liu,
Jiaxuan You
Abstract:
No single large language model (LLM) is optimal across all queries and budget constraints, making model routing essential for cost-effective deployment. Existing routers adopt diverse formulations and implementations, making fair comparison and extension difficult. We present a unified formulation of LLM routing as a sequential decision process characterized by five components: context encoders, m…
▽ More
No single large language model (LLM) is optimal across all queries and budget constraints, making model routing essential for cost-effective deployment. Existing routers adopt diverse formulations and implementations, making fair comparison and extension difficult. We present a unified formulation of LLM routing as a sequential decision process characterized by five components: context encoders, model encoders, scoring functions, decision rules, and learning signals, covering single-turn, multi-turn, and personalized routing. Based on this formulation, we develop an automated pipeline for constructing routing supervision and evaluating routers jointly on response quality and inference cost. The resulting benchmark, xRouteBench, spans generic LLM, memory-augmented, vision, time-series, and personalized routing tasks. We further introduce LLMRouter, an open-source modular infrastructure with more than 16 representative routers. Our empirical study shows that learned routers outperform the strongest fixed-model baseline by 14.6% relatively, lightweight routers become more competitive under tight cost constraints, and user-conditioned routing consistently improves personalization.
△ Less
Submitted 7 August, 2026;
originally announced August 2026.
-
Search for the charged lepton flavour violating decay $η'\to eμ$
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
M. S. Anderson,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone
, et al. (744 additional authors not shown)
Abstract:
Based on $(8998\pm40)\times10^6$ $J/ψ$ events collected in $e^+e^-$ collisions at $\sqrt{s} = 3.097$ GeV with the BESIII detector, we present a search for the charged lepton flavour violating decay $η'\to eμ$ with $J/ψ\toγη'$. No significant signal is observed, and an upper limit on its decay branching fraction is set to be $6.3\times10^{-7}$ at the 90% confidence level, improving the previous bes…
▽ More
Based on $(8998\pm40)\times10^6$ $J/ψ$ events collected in $e^+e^-$ collisions at $\sqrt{s} = 3.097$ GeV with the BESIII detector, we present a search for the charged lepton flavour violating decay $η'\to eμ$ with $J/ψ\toγη'$. No significant signal is observed, and an upper limit on its decay branching fraction is set to be $6.3\times10^{-7}$ at the 90% confidence level, improving the previous best result by nearly three orders of magnitude.
△ Less
Submitted 6 August, 2026;
originally announced August 2026.
-
UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models
Authors:
Haiyang Zhou,
Wangbo Yu,
Chaoran Feng,
Xunyu Zhou,
Yonghong Tian,
Li Yuan
Abstract:
The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experiences. However, producing photorealistic and geometrically consistent views with precise camera control remains challenging when input coverage is extremely limited. Recons…
▽ More
The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experiences. However, producing photorealistic and geometrically consistent views with precise camera control remains challenging when input coverage is extremely limited. Reconstruction-based approaches such as NeRF and 3D Gaussian Splatting (3DGS) deteriorate severely under sparse inputs and fail to explicitly handle occlusions. Generative methods ease data requirements but still struggle with large-baseline view synthesis due to inaccurate or implicit geometric guidance. To overcome these limitations, we introduce UniWorld-View, a unified framework for controllable large-baseline novel view synthesis from monocular inputs. UniWorld-View integrates explicit 3D guidance with generative diffusion modeling to enable precise camera control and geometrically consistent view generation. The geometric guidance is obtained through an occlusion-aware point cloud rendering strategy that resolves visibility ambiguities and provides accurate priors for diffusion-based synthesis. By coupling this rendering strategy with powerful video diffusion backbones, UniWorld-View achieves high-fidelity novel view generation even under extreme camera motions and wide-baseline changes, and can further provide multi-view videos for downstream dynamic 3DGS reconstruction. Experiments on the WorldScore benchmark and zero-shot NVS benchmarks demonstrate the effectiveness of UniWorld-View in controllability, geometric consistency, and visual fidelity.
△ Less
Submitted 5 August, 2026;
originally announced August 2026.
-
UniWorld-Design: From Pixel Generation to Layer-Native Design
Authors:
Zongjian Li,
Zhiyuan Yan,
Chenxu Bai,
Chen Chen,
Haoxiang Sun,
Shaodong Wang,
Feize Wu,
Shenghai Yuan,
Bin Lin,
Zheyuan Liu,
Yuwei Niu,
Li Yuan
Abstract:
We introduce UniWorld-Design, a framework that redefines image generation from flat pixel synthesis to structured visual composition, with semantic RGBA layers as the atomic units of generation, understanding, and editing. Our key insight is that pixels define how an image is rendered, whereas layers define how an image is created, understood, and edited. Just as human designers create and manipul…
▽ More
We introduce UniWorld-Design, a framework that redefines image generation from flat pixel synthesis to structured visual composition, with semantic RGBA layers as the atomic units of generation, understanding, and editing. Our key insight is that pixels define how an image is rendered, whereas layers define how an image is created, understood, and edited. Just as human designers create and manipulate visual content through layers rather than raw pixels, UniWorld-Design equips multimodal generative models with a layer-native design space. UniWorld-Design comprises two models. The Text-to-RGBA (T2RGBA) model generates standalone RGBA assets directly from text. The Image-to-Layer (I2L) model conditions on a finished image, a global instruction and per-layer prompts, and jointly produces ordered, complete semantic RGBA layers. Its instruction interface supports top-level decomposition, recursive decomposition and targeted extraction, making layering an instruction-addressable operation for agentic editing. Because I2L learns complete semantic objects rather than visible-pixel partitions, its layers stay usable when moved or removed. On the Crello benchmark, I2L reduces per-layer RGB L1 error by 37% and achieves a 34% relative improvement in Alpha Soft IoU over Qwen-Image-Layered. Separately, T2RGBA achieves the highest CLIP Score, outperforming LayerDiffuse and OmniAlpha.
△ Less
Submitted 4 August, 2026;
originally announced August 2026.
-
LiveEvalBench: Toward Open-World Evaluation for Web Generation
Authors:
Yiyao Wang,
Zhen Wen,
Yinghao Tang,
Yixiao Fu,
Lin Yuan,
Xiaolau Zhang,
Jun Zhou,
Wei Chen
Abstract:
Large language models are increasingly capable of synthesizing executable frontend projects, yet existing benchmarks still treat web generation as a static evaluation problem. We argue that frontend artifacts demand a different paradigm: they are interactive rather than static, admit diverse yet equally valid implementations, and evolve faster than rigid pipelines can accommodate. To address these…
▽ More
Large language models are increasingly capable of synthesizing executable frontend projects, yet existing benchmarks still treat web generation as a static evaluation problem. We argue that frontend artifacts demand a different paradigm: they are interactive rather than static, admit diverse yet equally valid implementations, and evolve faster than rigid pipelines can accommodate. To address these gaps, we present LiveEvalBench, an automated framework that reformulates web-generation evaluation as an agentic, adaptive, and extensible process. LiveEvalBench instantiates evaluation as a collaborative review workflow, in which a Build Engineer, a Code Engineer, and a UI Tester collectively gather evidence across the full lifecycle of a frontend project, from deployment and code inspection to browser-based interaction. To handle implementation diversity, an adaptive protocol couples shared rubrics for cross-model comparability with implementation-grounded criteria tailored to each artifact. The framework further supports incremental integration of new evaluator roles and assessment dimensions without pipeline redesign. Experiments across diverse real-world web-generation scenarios show that LiveEvalBench aligns closely with human expert judgment and provides fine-grained insights into frontier models' web generation capabilities. Code is available at https://github.com/wyysteelhead/LiveEvalBench
△ Less
Submitted 4 August, 2026;
originally announced August 2026.
-
Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts
Authors:
Ziang Wu,
Peng Jin,
Qishen Yin,
Munan Ning,
Hao Li,
Peizhen Zhang,
Li Yuan
Abstract:
Vision-language MoE batches contain different numbers of image and text tokens. Image resolution, image count, tiling, and prompt length all change this token mix. We call the standard token-level Switch auxiliary loss Std-Aux. Std-Aux balances only the mixed load, so large image and text load errors can cancel at one mix. On our main model, the same trained router shows more than a fivefold chang…
▽ More
Vision-language MoE batches contain different numbers of image and text tokens. Image resolution, image count, tiling, and prompt length all change this token mix. We call the standard token-level Switch auxiliary loss Std-Aux. Std-Aux balances only the mixed load, so large image and text load errors can cancel at one mix. On our main model, the same trained router shows more than a fivefold change in load imbalance across image resolutions. We hold the image and text load profiles fixed and derive the exact load curve as the token mix varies. The image-text load gap controls sensitivity to the token mix. Physical preprocessing can also change the conditional profiles. The fixed-profile law excludes such changes. To design a remedy, we examine the router input structure. Image and text occupy distinct regions, while visual tokens group strongly by source image. The modality boundary motivates separate image and text terms. The image boundary motivates one equal-weight routing instance per image. ReBA, or Relax Within, Balance Across, implements both choices. Across four split backbones, ReBA lowers load on every reported benchmark input while keeping mean task accuracy comparable to Std-Aux. ReBA also lowers average load over the tested range and worst physical load under resolution and tiling shifts. Code is available at https://github.com/ZiangWu-77/ReBA.
△ Less
Submitted 1 August, 2026;
originally announced August 2026.
-
On the Robustness of Propagating Bound States in the Continuum
Authors:
Lijun Yuan,
Ya Yan Lu
Abstract:
Bound states in the continuum (BICs) are localized eigenmodes with their frequencies in the radiation continuum of scattering states. The existence of a BIC implies the loss of uniqueness for scattering problems with given incident waves. Perturbed wave systems close to the ideal ones with a BIC exhibit strong resonance effects that are essential to numerous practical applications. A question of f…
▽ More
Bound states in the continuum (BICs) are localized eigenmodes with their frequencies in the radiation continuum of scattering states. The existence of a BIC implies the loss of uniqueness for scattering problems with given incident waves. Perturbed wave systems close to the ideal ones with a BIC exhibit strong resonance effects that are essential to numerous practical applications. A question of fundamental importance is whether a BIC is robust, i.e., whether it can continue its existence when the structure is slightly perturbed. In an earlier work [Yuan and Lu, Optics Letters, Vol.~42, pp.~4490-4493, 2017], for a class of BICs governed by the two-dimensional (2D) Helmholtz equation, which are not trivially protected by symmetry, we uncovered the conditions that ensure robustness and formally constructed the BIC in perturbed systems using a perturbation method. In this paper, we present a rigorous theory on the robustness of BICs in 2D dielectric structures with a single periodic direction. Specifically, we analyze the solvability and provide estimates for each order in the perturbation series, and prove the convergence of the series.
△ Less
Submitted 31 July, 2026;
originally announced August 2026.
-
Divergence Decoding: Training-Free Capability Fusion
Authors:
Yimi Wang,
Hao Li,
Shuo Yang,
He Cao,
Dechen Zhang,
Ziang Wu,
Zhiyuan Yan,
Fanyang Mo,
Li Yuan
Abstract:
While large language models excel in reasoning, these generalists often lack knowledge for specialized scientific domains. Conversely, domain models~(specialists), while knowledgeable, suffer from specialization side-effects including diminished logic and reduced robustness.To address this dilemma, we introduce Divergence Decoding, a training-free framework for capability fusion. It reconstructs t…
▽ More
While large language models excel in reasoning, these generalists often lack knowledge for specialized scientific domains. Conversely, domain models~(specialists), while knowledgeable, suffer from specialization side-effects including diminished logic and reduced robustness.To address this dilemma, we introduce Divergence Decoding, a training-free framework for capability fusion. It reconstructs the "draft-and-verify" skeleton of speculative decoding into an adaptive routing mechanism. The core is using Jensen-Shannon divergence to monitor the distributional disagreement between the two models at each token. When the specialist exhibits significant divergence, our method identifies it as a potential reasoning risk and instantaneously routes control to the generalist. This allows the dynamic injection of general reasoning while preserving domain expertise, achieving inference-time policy composition of the generalist and the specialist.We evaluate Divergence Decoding across diverse model families (Qwen and Llama series) on challenging scientific benchmarks (GPQA, ChemBench, and ChemCoTBench). Experimental results demonstrate that Divergence Decoding outperforms both the domain-specialized and general-purpose models, effectively surpassing the performance of most single-model baseline. This suggests that Divergence Decoding provides a general, training-free paradigm for fusing diverse LLM capabilities through adaptive inference-time collaboration.
△ Less
Submitted 28 July, 2026;
originally announced July 2026.
-
The pair correlation function of the Sine$_6$ process
Authors:
Shengqi Qiu,
Yahui Qu,
Lingfan Yuan,
Benedek Valkó,
Spencer Venancio
Abstract:
We derive an explicit formula for the pair correlation function of the Sine$_6$ process in terms of Bessel functions of the first kind. This provides the first single-variable special function representation of the pair correlation function for the bulk limit of a beta-ensemble beyond the classical values of $β=1,2,$ and $4$.
We derive an explicit formula for the pair correlation function of the Sine$_6$ process in terms of Bessel functions of the first kind. This provides the first single-variable special function representation of the pair correlation function for the bulk limit of a beta-ensemble beyond the classical values of $β=1,2,$ and $4$.
△ Less
Submitted 28 July, 2026;
originally announced July 2026.
-
Search for the $\boldsymbol{B^0 \to K^0_{\rm S} τ^+ τ^-}$ decay
Authors:
Belle,
Belle II Collaborations,
:,
M. Abumusabh,
I. Adachi,
A. Aggarwal,
Y. Ahn,
H. Aihara,
M. Akdag,
N. Akopov,
S. Alghamdi,
M. Alhakami,
A. Aloisio,
N. Althubiti,
K. Amos,
M. Angelsmark,
N. Anh Ky,
C. Antonioli,
D. M. Asner,
H. Atmacan,
T. Aushev,
V. Aushev,
R. Ayad,
V. Babu,
H. Bae
, et al. (410 additional authors not shown)
Abstract:
We present the first search for $B^0 \to K^0_{\rm S} τ^+τ^-$ decays. We look for signal decays in $B^0\bar B^0$ events produced in asymmetric-energy electron-positron collisions. This work uses samples from the Belle and Belle~II detectors, comprising 1.16 billion $Υ(4S)$ events. In $Υ(4S)\to B^0\bar{B}^0$ decays, the non-signal $\bar{B}^0$ meson is fully reconstructed in a hadronic channel. For t…
▽ More
We present the first search for $B^0 \to K^0_{\rm S} τ^+τ^-$ decays. We look for signal decays in $B^0\bar B^0$ events produced in asymmetric-energy electron-positron collisions. This work uses samples from the Belle and Belle~II detectors, comprising 1.16 billion $Υ(4S)$ events. In $Υ(4S)\to B^0\bar{B}^0$ decays, the non-signal $\bar{B}^0$ meson is fully reconstructed in a hadronic channel. For the signal $B^0$ meson, $τ$-lepton decays into final states with a single charged particle are selected. A multivariate classifier is used to combine several discriminating inputs into a single fit observable. We observe no evidence for the signal and set an upper limit on the branching fraction $\mathcal{B}(B^0\to K^0_{\rm S} τ^+τ^-) < 8.3 \times 10^{-4}$ at the 90\% confidence level. Combining this with the recent measurement of the isospin-partner decay $B^+\to K^+τ^+τ^-$, we determine an upper limit $\mathcal{B}(B\to Kτ^+τ^-) < 5.4\times10^{-4}$ at the 90\% confidence level.
△ Less
Submitted 28 July, 2026;
originally announced July 2026.
-
MSVS-VAE: Multi-Scale Anchored VecSet for High-Fidelity 3D Reconstruction
Authors:
Dehao Hao,
Kaiyi Zhang,
Tanghui Jia,
Xiangjun Gao,
Dongyu Yan,
Weikai Chen,
Zeyu Hu,
Lingting Zhu,
Yingda Yin,
Runze Zhang,
Li Yuan,
Xin Wang,
Long Quan
Abstract:
High-fidelity 3D generative modeling increasingly relies on the latent diffusion paradigm, where the reconstruction quality of the underlying 3D VAE becomes a primary bottleneck. Existing approaches largely follow two paradigms: sparse voxel-based representations achieve strong reconstruction quality but incur significant memory and computational overhead, while set-based representations are compa…
▽ More
High-fidelity 3D generative modeling increasingly relies on the latent diffusion paradigm, where the reconstruction quality of the underlying 3D VAE becomes a primary bottleneck. Existing approaches largely follow two paradigms: sparse voxel-based representations achieve strong reconstruction quality but incur significant memory and computational overhead, while set-based representations are compact and continuous yet typically lag in fidelity due to latent sparsity and excessive global smoothness. We propose MSVS-VAE, a hierarchical set-based VAE that closes this fidelity gap without sacrificing compactness. Our key idea is to progressively densify anchored VecSet latents via hierarchical point-shuffle upsampling, increasing spatial capacity for fine-grained geometry modeling. To efficiently decode from the densified hierarchy, we replace global cross-attention with AVS-Conv, a geometry-aware local aggregation operator operating within local neighborhoods rather than the exhaustive latent set. We further introduce multi-scale query decoding to fuse coarse-to-fine latent features, where coarse scales provide stable global context, and fine scales refine localized geometry, reducing artifacts from overly local receptive fields. Extensive experiments on Objaverse, ABO, and in-the-wild benchmarks demonstrate that MSVS-VAE consistently outperforms prior set-based and voxel-based VAEs, delivering approximately 10x faster decoding than prior set-based methods and approximately 10x higher compactness than voxel-based baselines.
△ Less
Submitted 27 July, 2026;
originally announced July 2026.
-
Precision Measurement of Decay Dynamics in $D^{0(+)}\to π^{-(0)}\ell^+ν_\ell$
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
M. S. Anderson,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone
, et al. (752 additional authors not shown)
Abstract:
The branching fractions of $D^0\to π^-e^+ν_e$, $D^0\to π^-μ^+ν_μ$, $D^+\to π^0e^+ν_e$, and $D^+\to π^0μ^+ν_μ$ are precisely measured, using 20.3 fb$^{-1}$ of $e^+e^-$ collision data collected at the center-of-mass energy of 3.773 GeV with the BESIII detector. The ratios of the decay widths between muon and positron channels are examined in full, across several four-momentum transfer ranges of…
▽ More
The branching fractions of $D^0\to π^-e^+ν_e$, $D^0\to π^-μ^+ν_μ$, $D^+\to π^0e^+ν_e$, and $D^+\to π^0μ^+ν_μ$ are precisely measured, using 20.3 fb$^{-1}$ of $e^+e^-$ collision data collected at the center-of-mass energy of 3.773 GeV with the BESIII detector. The ratios of the decay widths between muon and positron channels are examined in full, across several four-momentum transfer ranges of $\ell^+ν_{\ell}$. No lepton flavor universality violation is found in the current data. From a simultaneous fit to the precisely measured partial decay rates and the first measured forward-backward asymmetries of these four decays, the product of the hadronic transition form factor, $f^{D\toπ}_+(0)$, and the modulus of the $c\to d$ quark mixing element, $|V_{cd}|$, is measured with unprecedented precision to be $f^{D\toπ}_+(0)|V_{cd}|=0.1425\pm0.0005_{\rm stat.}\pm0.0003_{\rm syst.}$. Taking the value of $|V_{cd}|$ from the standard model global fit and $f^{D\toπ}_+(0)$ derived by the lattice quantum chromodynamics calculation as input, we obtain $f^{D\toπ}_+(0)=0.1425\pm0.0005_{\rm stat.}\pm0.0003_{\rm syst.}$ and $|V_{cd}|=0.2262\pm0.0008_{\rm stat.}\pm0.0005_{\rm syst.}\pm0.0018_{\rm LQCD.}$, respectively. The precision of each result is a factor of 2-3 better than the previous best measurements. Additionally, the real and imaginary parts of the scalar current contribution in the $c\to d \ell^+ν_{\ell}$ transition are measured for the first time to be Re $(C_S^μ)=$ $0.022 \pm 0.023_{\rm stat.}\pm 0.003_{\rm syst.}$ and $|\mathrm{Im} (C_S^μ)|=0.000 \pm 0.038_{\rm stat.}\pm 0.012_{\rm syst.}$.
△ Less
Submitted 26 July, 2026;
originally announced July 2026.
-
Precision measurements of semleptonic decays $D^0 \to π^-\ell^+ν_\ell$ and $D^+ \to π^0\ell^+ν_\ell$ ($\ell =e,μ$)
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
M. S. Anderson,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone
, et al. (752 additional authors not shown)
Abstract:
The branching fractions of $D^0\to π^-e^+ν_e$, $D^0\to π^-μ^+ν_μ$, $D^+\to π^0e^+ν_e$, and $D^+\to π^0μ^+ν_μ$ are measured to be $(2.950\pm0.017_{\rm stat.}\pm 0.017_{\rm syst.})\times10^{-3}$, $(2.817\pm0.037_{\rm stat.}\pm 0.019_{\rm syst.})\times10^{-3}$, $(3.622\pm0.034_{\rm stat.}\pm 0.018_{\rm syst.})\times10^{-3}$, and $(3.507\pm0.043_{\rm stat.}\pm 0.026_{\rm syst.})\times10^{-3}$ using…
▽ More
The branching fractions of $D^0\to π^-e^+ν_e$, $D^0\to π^-μ^+ν_μ$, $D^+\to π^0e^+ν_e$, and $D^+\to π^0μ^+ν_μ$ are measured to be $(2.950\pm0.017_{\rm stat.}\pm 0.017_{\rm syst.})\times10^{-3}$, $(2.817\pm0.037_{\rm stat.}\pm 0.019_{\rm syst.})\times10^{-3}$, $(3.622\pm0.034_{\rm stat.}\pm 0.018_{\rm syst.})\times10^{-3}$, and $(3.507\pm0.043_{\rm stat.}\pm 0.026_{\rm syst.})\times10^{-3}$ using $e^+e^-$ collision data with an integrated luminosity of 20.3 fb$^{-1}$ collected at the center-of-mass energy of 3.773 GeV with the BESIII detector. The partial decay rates of these four decays are measured with the best precision to date and their forward-backward asymmetries are determined for the first time. By performing a simultaneous fit to these results, the product of the hadronic transition form factor $f^{D\toπ}_+(0)$ and the modulus of the $c\to d$ Cabibbo-Kobayashi-Maskawa matrix element $|V_{cd}|$ is given by $f^{D\toπ}_+(0)|V_{cd}|=0.1425\pm0.0005_{\rm stat.}\pm0.0003_{\rm syst.}$. Taking the $|V_{cd}|$ provided by the standard model global fit and the $f^{D\toπ}_+(0)$ calculated from the lattice quantum chromodynamics as input, we obtain $f^{D\toπ}_+(0)=0.6339\pm0.0024_{\rm stat.}\pm0.0014_{\rm syst.}$ and $|V_{cd}|=0.2262\pm0.0008_{\rm stat.}\pm0.0005_{\rm syst.}\pm0.0018_{\rm LQCD.}$, respectively. The reported results have the best precision to date. We also search for the scalar current contribution in the $c\to d \ell^+ν_{\ell}$ transition and determine Re$(C_S^μ)=$ $0.022 \pm 0.023_{\rm stat.}\pm 0.003_{\rm syst.}$ and $|{\rm Im}(C_S^μ)|=0.000 \pm $ $0.038_{\rm stat.} \pm 0.012_{\rm syst.}$. In addition, the lepton flavor universality is tested with the ratios of the decay rates between semimuonic and semielectronic decays in full and several $\ell^+ν_\ell$ four-momentum transfer ranges.
△ Less
Submitted 26 July, 2026;
originally announced July 2026.
-
Measurement of Born Cross Section for $e^+e^-\to K_S^0\barΞ^+Σ^-+\rm{c.c.}$ at $\sqrt{s} = 3.51-4.95$ GeV
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko
, et al. (737 additional authors not shown)
Abstract:
Using $e^+e^-$ collision data collected with the BESIII detector at the BEPCII collider corresponding to a total integrated luminosity of 44~fb$^{-1}$, we present the first measurement of the Born cross sections for the process $e^+e^-\to K_S^0\barΞ^+Σ^-+\rm{c.c.}$ at 56 center-of-mass energies from 3.510 to 4.951~GeV. By fitting the dressed cross sections of $e^+e^-\to K_S^0\barΞ^+Σ^-+\rm{c.c.}$…
▽ More
Using $e^+e^-$ collision data collected with the BESIII detector at the BEPCII collider corresponding to a total integrated luminosity of 44~fb$^{-1}$, we present the first measurement of the Born cross sections for the process $e^+e^-\to K_S^0\barΞ^+Σ^-+\rm{c.c.}$ at 56 center-of-mass energies from 3.510 to 4.951~GeV. By fitting the dressed cross sections of $e^+e^-\to K_S^0\barΞ^+Σ^-+\rm{c.c.}$ with the assumption of a power-law function plus a charmonium(-like) resonance, i.e. $ψ(3770)$, $ψ(4040)$, $ψ(4160)$, $Y(4230)$, $Y(4360)$, $ψ(4415)$, {\it Y}(4500), $Y(4660)$, and {\it Y}(4710), no significant signal of any charmonium(-like) state decaying into the $K_S^0\barΞ^+Σ^-+\rm{c.c.}$ is observed. Upper limits on the product of the electronic width and branching fraction at the 90\% confidence level are given for each resonance. Combining this result with the previous measurement of the isospin-symmetric process $e^+e^-\to K^{-} \barΞ^{+} Σ^{0} + \rm{c.c.}$, the ratio of the Born cross sections, $R=σ^{B}(e^+e^-\to K_S^0\barΞ^+Σ^-+\rm{c.c.})/$$σ^{B}(e^+e^-\to K^-\barΞ^+Σ^0+\rm{c.c.})$, is found to be approximately 1.
△ Less
Submitted 24 July, 2026;
originally announced July 2026.
-
CRAFT: Exploring Wearable Creative AI on Smart Glasses for Fiction Writing in Real-World Contexts
Authors:
Runze Cai,
Yuxuan Huang,
Lin-Ping Yuan,
Kexin Xiang,
David Hsu,
Collier Nogues,
Jussi Holopainen,
Shengdong Zhao
Abstract:
Creative writing increasingly integrates AI assistance, yet current tools miss in-situ moments when writers draw inspiration from real-world experiences. We envision Context-aware Reality-Fiction Transformation (CRAFT), an approach for AI glasses that translates daily experiences into fiction narratives. We explored its desirability, feasibility, and potential viability through three studies. Inte…
▽ More
Creative writing increasingly integrates AI assistance, yet current tools miss in-situ moments when writers draw inspiration from real-world experiences. We envision Context-aware Reality-Fiction Transformation (CRAFT), an approach for AI glasses that translates daily experiences into fiction narratives. We explored its desirability, feasibility, and potential viability through three studies. Interviews with nine writers yielded desires and three design goals: 1) augmenting in-situ perception to bridge reality-fiction gaps, 2) promoting authenticity grounded in real-world experiences while maintaining fictionalization, and 3) preserving creative agency, enjoyment, and life-art boundaries. Co-design workshops with 16 writers and researchers operationalized these goals into concrete interaction mechanisms using a technology probe. We then conducted supported field trials with eight writers across 24 sessions using a refined probe, revealing writer-perceived benefits (e.g., enriched fictional ideas from serendipitous encounters), emergent practices (e.g., micro-creation), and design considerations for future sustained use. We contribute design explorations for the CRAFT approach, offering design implications and empirical insights on ubiquitous human-AI creative collaboration in everyday life.
△ Less
Submitted 23 July, 2026;
originally announced July 2026.
-
First Measurement of the Relative Phase between Proton Psionic Form Factors
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
Y. Bai,
O. Bakina,
Y. Ban,
H. -R. Bao,
X. L. Bao,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko
, et al. (732 additional authors not shown)
Abstract:
The relative phase between the time-like form factors of the proton is a crucial observable for a complete understanding of its internal structure, yet it has remained unmeasured due to the formidable experimental challenge of determining the final-state polarization or having available polarized beams. With a novel technique that measures polarization via secondary scattering on spectrometer mate…
▽ More
The relative phase between the time-like form factors of the proton is a crucial observable for a complete understanding of its internal structure, yet it has remained unmeasured due to the formidable experimental challenge of determining the final-state polarization or having available polarized beams. With a novel technique that measures polarization via secondary scattering on spectrometer material, we use $10.09\times10^{9}$ $J/ψ$ events collected at BESIII to analyze the reaction $e^+e^-\rightarrow J/ψ\rightarrow p\bar{p}$. This allows the first determination of the sine of the relative phase between the proton psionic form factors, $\sinΔΦ=-0.20\pm0.34_{\textrm{stat}}\pm0.11_{\textrm{syst}}$. This result provides the first direct insight into the complex dynamics of proton formation, and offers valuable new information to constrain theoretical models of nucleon structure.
△ Less
Submitted 22 July, 2026;
originally announced July 2026.
-
Proof of principle for nucleon polarization measurement at BESIII
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
Y. Bai,
O. Bakina,
Y. Ban,
H. -R. Bao,
X. L. Bao,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko
, et al. (732 additional authors not shown)
Abstract:
A novel technique for measuring the spin polarization of final-state nucleons in a general-purpose spectrometer is validated. Using $10.09\times10^{9}$ $J/ψ$ events at BESIII, the asymmetry of polarized proton scattering on detector support material is measured, and is consistent with the expected value. This proves that a general-purpose spectrometer can be utilized as a large-acceptance polarime…
▽ More
A novel technique for measuring the spin polarization of final-state nucleons in a general-purpose spectrometer is validated. Using $10.09\times10^{9}$ $J/ψ$ events at BESIII, the asymmetry of polarized proton scattering on detector support material is measured, and is consistent with the expected value. This proves that a general-purpose spectrometer can be utilized as a large-acceptance polarimeter, providing the spin polarization in addition to the conventional four-momentum information of the final-state particles. With this technique, physics capabilities are enhanced for existing and future facilities in particle and nuclear physics.
△ Less
Submitted 22 July, 2026;
originally announced July 2026.
-
Hypothesis-and-Refinement Learning of Organic Structures from Multimodal Spectroscopic Data
Authors:
Chengchun Liu,
Zhiyuan Yan,
Li Yuan,
Hao Li,
Boxuan Zhao,
Yonghong Tian,
Bartosz A. Grzybowski,
Fanyang Mo
Abstract:
Determining molecular structures from spectroscopic data remains fundamentally challenging because the inverse problem is intrinsically underdetermined: individual spectra are sparse, low-dimensional, and encode only partial structural evidence relative to the vast space of possible molecules. We address this challenge by formulating automated structure elucidation as a scalable hypothesis-refinem…
▽ More
Determining molecular structures from spectroscopic data remains fundamentally challenging because the inverse problem is intrinsically underdetermined: individual spectra are sparse, low-dimensional, and encode only partial structural evidence relative to the vast space of possible molecules. We address this challenge by formulating automated structure elucidation as a scalable hypothesis-refinement paradigm that tightly integrates spectral evidence with large-scale molecular priors. To supply structure-resolving NMR signals for multimodal learning, we construct \textbf{QM9SPIN}, a DFT-derived dataset comprising diverse 1D and 2D spectra, including J-coupling, DEPT experiments, and explicit spin--spin interactions. On this foundation, we introduce \textbf{SpectroMol}, a spectrum-to-structure model that proposes chemically valid molecular hypotheses conditioned on multimodal spectral inputs. Complementarily, we develop \textbf{MS-Mol2Mol}, a high-resolution mass-constrained molecular generator that integrates molecular formula, exact mass, and degree of unsaturation within a conditional generative prior trained on 400 million molecules, ensuring global compositional consistency and chemically realistic refinement. The integrated system achieves 93.8\% top-1 accuracy on the simulated benchmark, adapts effectively from simulated to experimental spectra with limited experimental fine-tuning, and further improves experimental predictions through mass-guided refinement, establishing a scalable route toward automated, data-driven organic structure elucidation.
△ Less
Submitted 22 July, 2026;
originally announced July 2026.
-
Dark matter searches with a 13 meV threshold superconducting sensor array
Authors:
Christopher Albert,
Lanqing Yuan,
Jacob Harris,
Ritoban Basu Thakur,
Andrew Bear,
Karl K. Berggren,
Christopher Cappiello,
Christopher Curwen,
Peter Day,
Byeong H. Eom,
Arjun Ghosh,
William Ho,
Nikita Klimovich,
Henry G. LeDuc,
Karthik Ramanathan,
Alejandro Simon
Abstract:
Many well-motivated dark matter models predict meV-scale energy deposits in interactions with terrestrial experiments, but this regime is challenging to probe due to a lack of mature single-quantum detectors. Here we report results from QUALIPHIDE (QUAntum LImited PHotons In the Dark Experiment), a cryogenic dark matter search using a $41$-pixel array of energy-resolving microwave kinetic inductan…
▽ More
Many well-motivated dark matter models predict meV-scale energy deposits in interactions with terrestrial experiments, but this regime is challenging to probe due to a lack of mature single-quantum detectors. Here we report results from QUALIPHIDE (QUAntum LImited PHotons In the Dark Experiment), a cryogenic dark matter search using a $41$-pixel array of energy-resolving microwave kinetic inductance detectors with a $13$ meV threshold, simultaneously used to look for both conversion photons from THz wavelength hidden photon dark matter and phonons from particle-like light dark matter interactions. The experimental design, with on- and off-focus pixels for the hidden photon search, allows for a data-driven background model, giving the experiment discovery potential. A blind analysis of $22$ hours of data shows no significant excess, setting the strongest constraints on the hidden photon kinetic mixing parameter $χ$ over the mass range of $13$-$90$ meV/$c^2$, reaching $1.5\times10^{-12}$ at $50$ meV/$c^2$. These data also yield among the first terrestrial limits on dark matter scattering off nuclei and electrons, down to $5$ MeV/$c^2$ and $20$ keV/$c^2$, respectively. The low threshold also enables future study of the low-energy excess limiting cryogenic detectors and, as we project, will allow for a terahertz-scale QCD axion search with a magnetic field.
△ Less
Submitted 22 July, 2026; v1 submitted 21 July, 2026;
originally announced July 2026.
-
Exact generalized Turán number of vertex-disjoint paths of length two
Authors:
Qi Wu,
Long-Tu Yuan
Abstract:
We determine the generalized Turán number of vertex-disjoint paths of length two and characterize all corresponding extremal graphs. Our proof combines the Lovász form of the Kruskal--Katona theorem with a discrete convexity argument.
We determine the generalized Turán number of vertex-disjoint paths of length two and characterize all corresponding extremal graphs. Our proof combines the Lovász form of the Kruskal--Katona theorem with a discrete convexity argument.
△ Less
Submitted 20 July, 2026;
originally announced July 2026.
-
Biodegradable, Millimeter-Scale Light-Emitting Sensors for Distributed Environmental Monitoring-Functional Pixie Dust
Authors:
Zhiming Hu,
Danzhen Zhang,
Janghun Ko,
Haohui Zhang,
Jiale Chen,
Chanho Park,
Jiatong Zhang,
Qiuna Zhuang,
Shiwei Xu,
Xiaoran Yang,
Dain Son,
Taehoon Kim,
Uikang Joo,
Zhaojian Xu,
Hyunsoo Kim,
Richard Chai,
Gwangmin Bae,
Wooyoul Maeng,
Qiong Wang,
Sangmin Lim,
Liangsong Zeng,
Un-Seong Baik,
Kaiqing Zhang,
Liming Yuan,
Yonggang Huang
, et al. (2 additional authors not shown)
Abstract:
Methods for large-area, precise monitoring across natural environments are of growing interest due to pressing needs for sustainable management of rapidly increasing anthropogenic activities. Established approaches involve sparse spatial sampling and/or sequential measurements, while emerging techniques exploit miniaturized electronics or passive optical methods. Various constraints in scalability…
▽ More
Methods for large-area, precise monitoring across natural environments are of growing interest due to pressing needs for sustainable management of rapidly increasing anthropogenic activities. Established approaches involve sparse spatial sampling and/or sequential measurements, while emerging techniques exploit miniaturized electronics or passive optical methods. Various constraints in scalability, costs, robustness, operational range and other factors create a need for alternatives. Here, we introduce a concept that overcomes many of these limitations through the combined use of chemically induced light emission and chemically responsive optical filter elements in millimeter-scale systems that we refer to as functional pixie dust (fPD) sensors, designed specifically for monitoring natural water systems during nighttime to eliminate background optical interference and to enhance remote analysis. These floating devices act as Lagrangian tracers to follow surface flows and to simultaneously measure the concentrations of key chemical species along their trajectories. Optimized designs exploit environmentally compatible constituent materials that are also degradable through natural processes to benign end products, thereby eliminating the need for recovery. Spatially and spectrally resolved ratiometric measurement schemes ensure robust operation and ability to address practical requirements in range, operational lifetime, time response and sensitivity. Demonstrations include distributed measurements of pH, Hg2+, and NO2-, each of relevance to industrial discharge, toxic metal contamination, and nitrogen-rich runoff, adapted for static concentration gradients, flow-driven transport conditions, and outdoor aquatic settings. The results establish a framework for environmental sensing using degradable, self-powered microsystems capable of scalable deployment and remote readout.
△ Less
Submitted 20 July, 2026;
originally announced July 2026.
-
DAUPNet: Domain-Aware Uncertainty Modeling for Reliable Prototype Discrimination in Cross-Domain Few-Shot Semantic Segmentation
Authors:
Lei Yuan,
Zhongxu Hu,
Jingyi Wen,
Pengxing Yi
Abstract:
Cross-domain few-shot semantic segmentation (CD-FSS) has predominantly been formulated as learning domain-invariant representations or improving support-query correspondence. Nevertheless, large domain shifts still make prototype matching unreliable: inconsistent hierarchical responses corrupt the support representation, deterministic prototypes cannot express boundary and appearance ambiguity, an…
▽ More
Cross-domain few-shot semantic segmentation (CD-FSS) has predominantly been formulated as learning domain-invariant representations or improving support-query correspondence. Nevertheless, large domain shifts still make prototype matching unreliable: inconsistent hierarchical responses corrupt the support representation, deterministic prototypes cannot express boundary and appearance ambiguity, and treating prototypes with different reliability equally during optimization weakens foreground-background separation. We therefore propose DAUPNet, a unified framework that reformulates cross-domain prototype matching as uncertainty-aware prototype discrimination. DAUPNet first harmonizes hierarchical support-query features to provide stable evidence, then represents foreground and background prototypes probabilistically, and finally uses their estimated uncertainty to regulate contrastive optimization. On four standard target domains, DAUPNet achieves 72.6% and 76.7% average mIoU in the 1-shot and 5-shot settings, respectively, including substantial gains on the two medical domains. These results demonstrate that modeling prototype uncertainty and incorporating it into optimization provides a robust and interpretable approach to CD-FSS under severe domain shift. The code is available at https://github.com/madness-Lei/DAUPNet
△ Less
Submitted 14 July, 2026;
originally announced July 2026.
-
Observation of $η_{c} \to p\bar{p}η$ via $ψ(3686) \to γp\bar{p}η$
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
M. S. Anderson,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone
, et al. (745 additional authors not shown)
Abstract:
The decay $η_c\to p\bar{p}η$ is observed for the first time with a significance of exceeding $10σ$. It is found by analyzing $(2712.4 \pm 14.3)\times10^{6}$ $ψ(3686)$ events accumulated at the BESIII detector. The measured branching fraction of $η_c\to p\bar{p}η$ via $ψ(3686) \to γp \bar{p} η$ is significantly influenced by the interference between the resonant $η_c$ decay and the non-resonant pro…
▽ More
The decay $η_c\to p\bar{p}η$ is observed for the first time with a significance of exceeding $10σ$. It is found by analyzing $(2712.4 \pm 14.3)\times10^{6}$ $ψ(3686)$ events accumulated at the BESIII detector. The measured branching fraction of $η_c\to p\bar{p}η$ via $ψ(3686) \to γp \bar{p} η$ is significantly influenced by the interference between the resonant $η_c$ decay and the non-resonant process $ψ(3686) \to γp \bar{p} η$ and is measured in both constructive- and destructive-interference scenarios. The joint branching fraction of $ψ(3686)\to γη_c$, $η_c\to p\bar{p}η$ is measured to be $(3.2 \pm 0.1 \pm 0.9)\times10^{-6}$ or $(8.7 \pm 0.3 \pm 2.1)\times10^{-6}$ for constructive- or destructive-interference solutions, respectively, where the first uncertainties are statistical and the second systematic. The branching fraction of $η_c\to p\bar{p}η$ is determined to be $\mathcal{B}(η_c\to p\bar{p}η)=(0.90 \pm 0.04 \pm 0.21 \pm 0.13)\times10^{-3}$ or $(2.42 \pm 0.07 \pm 0.48 \pm 0.34)\times10^{-3}$ for the two solutions, respectively, where the third uncertainties are due to the uncertainty in the branching fraction of $ψ(3686)\to γη_c$.
△ Less
Submitted 13 July, 2026;
originally announced July 2026.
-
Efficient Test-Time Optimization for Multi-Agent Proof Autoformalization
Authors:
Tian-Shuo Liu,
Shiyuan Zhang,
Zijie Geng,
Haoyu Liu,
Runjie Xu,
Pengyuan Wang,
Lei Yuan,
Yang Yu
Abstract:
Full-proof autoformalization bridges extensive mathematical proofs in natural language with formally validated reasoning, offering a pathway to elevate the ceiling of verifiable mathematical reasoning. Unlike statement-level formalization, proof autoformalization is a long-horizon challenge requiring coordination of claims, contexts, and dependencies across many proof steps, yet has only recently…
▽ More
Full-proof autoformalization bridges extensive mathematical proofs in natural language with formally validated reasoning, offering a pathway to elevate the ceiling of verifiable mathematical reasoning. Unlike statement-level formalization, proof autoformalization is a long-horizon challenge requiring coordination of claims, contexts, and dependencies across many proof steps, yet has only recently come under focused study. Current approaches either rely on costly model training or apply excessive, unguided repair at inference time. To this end, we introduce ToMap, a multi-agent framework that structures proof autoformalization as a Decomposer-Formalizer-Prover pipeline with efficient test-time optimization guided by formal verification and semantic rubrics for proof quality. Rather than distributing test-time compute across all agents, we perform bottleneck analysis and identify the Decomposer as the critical bottleneck: the quality of its atomic, self-contained proof units directly determines whether downstream agents can successfully formalize and prove each step. ToMap therefore treats the Formalizer and Prover as downstream executors and efficiently focuses test-time compute on Decomposer refinement. This refinement follows a loop inspired by GEPA, evolving prompts over candidate decompositions and using formal verification progress together with semantic proof rubrics to define a Pareto frontier that guides the next decomposition update. Experiments on ProofFlowBench show that ToMap improves over the best previous method by 19.0% when evaluated by both syntactic correctness and semantic faithfulness, while requiring lower test-time cost. Scaling analysis shows that most gains emerge within a few iterations of decomposition evolution, guiding test-time budget selection.
△ Less
Submitted 13 July, 2026;
originally announced July 2026.
-
Spatula: Exploring On-Demand In-Situ Interfaces and Interaction for Attribute Control
Authors:
Boyu Li,
Linjie Qiu,
Lin-Ping Yuan,
Duotun Wang,
Yue Jiang,
Zeyu Wang,
Hongbo Fu
Abstract:
Controlling attributes is a critical step toward achieving the final creative outcome, yet current approaches fall short in supporting users in the iterative refinement of generative content. We propose Spatula, a proof-of-concept system that generates on-demand, in-situ attribute control interfaces and interactions for creating motion graphics. Building on a technical probe that automatically ana…
▽ More
Controlling attributes is a critical step toward achieving the final creative outcome, yet current approaches fall short in supporting users in the iterative refinement of generative content. We propose Spatula, a proof-of-concept system that generates on-demand, in-situ attribute control interfaces and interactions for creating motion graphics. Building on a technical probe that automatically analyzes animation context and generates corresponding attributes and UI, we frame attribute control as an explorable landscape and explore the attribute control space along four key dimensions: Discoverability, Resolution, Scope, and Expandability. Findings from a user study (N=12) show that our system provides intuitive and convenient interactions while supporting diverse needs for fine-grained parameter control. Furthermore, our applications demonstrate that the plug-and-play design generalizes to other domains, such as web design and 3D modeling.
△ Less
Submitted 25 July, 2026; v1 submitted 11 July, 2026;
originally announced July 2026.
-
Dynamical Nonrelativistic Spin Splitting via THz Nonlinear Phononics
Authors:
Linding Yuan,
ChanJu You,
Ankit Disa,
James M. Rondinelli
Abstract:
Nonrelativistic spin splitting (NRSS) in collinear antiferromagnets offers a route to high-frequency spintronics immune to stray fields, but its dynamic control has remained elusive. We demonstrate, using density functional theory (DFT) and nonlinear phononics, that THz laser pulses can achieve ultrafast, reversible control of NRSS on picosecond timescales in antiferromagnets. We derive two symmet…
▽ More
Nonrelativistic spin splitting (NRSS) in collinear antiferromagnets offers a route to high-frequency spintronics immune to stray fields, but its dynamic control has remained elusive. We demonstrate, using density functional theory (DFT) and nonlinear phononics, that THz laser pulses can achieve ultrafast, reversible control of NRSS on picosecond timescales in antiferromagnets. We derive two symmetry criteria, accounting for phonon and magnetic wavevector compatibility and order-parameter parity, to identify which Raman-active phonon modes can activate or amplify NRSS. Applying these rules to NiO and LaFeO$_3$, we show that resonant driving of an infrared-active mode at 11.08 THz transiently converts spin-degenerate NiO into an NRSS state via biquadratic anharmonic coupling, generating a time-averaged spin splitting of $\sim$40 meV. In LaFeO$_3$, selective excitation amplifies the existing NRSS by about 100%. In both cases, the induced spin splitting is accompanied by a transient SOC-induced net moment detectable via the magneto-optical Kerr effect. This framework establishes nonlinear phononics as a general route for ultrafast manipulation of spin-split antiferromagnetic phases well beyond the reach of static strain.
△ Less
Submitted 10 July, 2026;
originally announced July 2026.
-
Search for axion-like particles decaying to two photons at Belle II
Authors:
Belle II Collaboration,
M. Abumusabh,
I. Adachi,
A. Aggarwal,
H. Ahmed,
Y. Ahn,
H. Aihara,
M. Akdag,
N. Akopov,
S. Alghamdi,
M. Alhakami,
A. Aloisio,
N. Althubiti,
K. Amos,
M. Angelsmark,
N. Anh Ky,
C. Antonioli,
K. Arai,
D. M. Asner,
H. Atmacan,
T. Aushev,
V. Aushev,
R. Ayad,
V. Babu,
H. Bae
, et al. (428 additional authors not shown)
Abstract:
Axion-like particles (ALPs) are predicted in many extensions of the Standard Model and provide a well-motivated portal between visible and hidden sectors through their coupling to photons. We search for ALPs produced in the process $e^{+}e^{-}\toγa$, $a\toγγ$, using a data sample corresponding to an integrated luminosity of $408~\mathrm{fb}^{-1}$ recorded by the Belle~II detector at the SuperKEKB…
▽ More
Axion-like particles (ALPs) are predicted in many extensions of the Standard Model and provide a well-motivated portal between visible and hidden sectors through their coupling to photons. We search for ALPs produced in the process $e^{+}e^{-}\toγa$, $a\toγγ$, using a data sample corresponding to an integrated luminosity of $408~\mathrm{fb}^{-1}$ recorded by the Belle~II detector at the SuperKEKB $e^{+}e^{-}$ collider. Events containing three photons are used to reconstruct the ALP as a narrow peak in the di-photon invariant mass spectrum over the range $0.17 < m_{a} < 9.80~\mathrm{GeV}/c^{2}$. No significant excess above background is observed. We set 95\% confidence level upper limits on the production cross section and on the ALP-photon coupling $g_{aγγ}$, reaching sensitivities at the level of $10^{-4}~\mathrm{GeV}^{-1}$. The limits are the most restrictive to date over nearly the entire mass range $0.17 < m_{a} < 5.00~\mathrm{GeV}/c^{2}$, and improve upon previous results by up to a factor 9.
△ Less
Submitted 8 July, 2026;
originally announced July 2026.
-
Search for an isoscalar partner of the $Z_c(3900)$ in $e^+e^-\toπ^+π^-ηJ/ψ$
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
R. Aliberti,
A. Amoroso,
Q. An,
Y. Bai,
O. Bakina,
Y. Ban,
H. -R. Bao,
X. L. Bao,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko,
R. A. Briere,
A. Brueggemann
, et al. (683 additional authors not shown)
Abstract:
Using a data sample corresponding to an integrated luminosity of 20.3 fb$^{-1}$ collected at center-of-mass energies from 4.18 to 4.95 GeV with the BESIII detector, we observe the process $e^+e^-\toπ^+π^-ηJ/ψ$ with a statistical significance of $6.0 σ$, including systematic uncertainties. The isoscalar partner of the $Z_c(3900)$, denoted $X(3900)$, is searched for in the $ηJ/ψ$ final state, and no…
▽ More
Using a data sample corresponding to an integrated luminosity of 20.3 fb$^{-1}$ collected at center-of-mass energies from 4.18 to 4.95 GeV with the BESIII detector, we observe the process $e^+e^-\toπ^+π^-ηJ/ψ$ with a statistical significance of $6.0 σ$, including systematic uncertainties. The isoscalar partner of the $Z_c(3900)$, denoted $X(3900)$, is searched for in the $ηJ/ψ$ final state, and no significant signal is observed. The upper limits on the product of the Born cross section $σ^{\rm Born}[e^{+}e^{-}\toπ^{+}π^{-} X(3900)$] and the branching fraction $\mathcal{B}[X(3900)\toηJ/ψ]$ are given with various assumptions for the mass and width of the $X(3900)$.
△ Less
Submitted 8 July, 2026;
originally announced July 2026.
-
Observation and branching fraction measurements of $J/ψ\to p \bar p K^0_S K^0_S$ and $ψ(3686) \to p \bar p K^0_S K^0_S$
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
M. S. Anderson,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone
, et al. (745 additional authors not shown)
Abstract:
Using $(10.087\pm0.044)\times10^9$ $J/ψ$ events and $(2.712\pm0.014)\times10^9$ $ψ(3686)$ events collected by the BESIII detector operating at the BEPCII collider, we report the first observation of the hadronic decays of $J/ψ\to p \bar p K^0_S K^0_S$ and $ψ(3686) \to p \bar p K^0_S K^0_S$, both with statistical significance greater than $10σ$. Their branching fractions are determined to be…
▽ More
Using $(10.087\pm0.044)\times10^9$ $J/ψ$ events and $(2.712\pm0.014)\times10^9$ $ψ(3686)$ events collected by the BESIII detector operating at the BEPCII collider, we report the first observation of the hadronic decays of $J/ψ\to p \bar p K^0_S K^0_S$ and $ψ(3686) \to p \bar p K^0_S K^0_S$, both with statistical significance greater than $10σ$. Their branching fractions are determined to be $\mathcal{B}(J/ψ\to p \bar p K^0_S K^0_S)=(1.60 \pm 0.02 \pm 0.09)\times10^{-5}$ and $\mathcal{B}(ψ(3686) \to p \bar p K^0_S K^0_S)=(3.93 \pm 0.24 \pm 0.34)\times10^{-6}$. The ratio of their branching fractions is $\mathcal{B}(ψ(3686) \to p \bar p K^0_S K^0_S)/\mathcal{B}(J/ψ\to p \bar p K^0_S K^0_S)=(24.6 \pm 1.5 \pm 2.1)\%$, which deviates from theoretical expectation by 4.6$σ$. Here the first uncertainties are statistical and the second systematic. We have also examined the $p\bar p$ invariant mass distributions in these decays, and no significant enhancement around the $p \bar p$ near threshold is found.
△ Less
Submitted 7 July, 2026;
originally announced July 2026.
-
Observation of the $χ_{cJ}$ decays into $pK^{-}\barΛη+\mathrm{c.c.}$
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
M. S. Anderson,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone
, et al. (759 additional authors not shown)
Abstract:
By analyzing $(2712.4 \pm 14.3) \times 10^{6}$ $ψ(3686)$ events collected with the BESIII detector operating at the BEPCII collider, the decays $χ_{cJ} \to pK^{-}\barΛη+ \mathrm{c.c.}$ ($J=0,1,2$) are observed for the first time, with statistical significances exceeding $5σ$ for all three $χ_{cJ}$ states. The measured branching fractions are…
▽ More
By analyzing $(2712.4 \pm 14.3) \times 10^{6}$ $ψ(3686)$ events collected with the BESIII detector operating at the BEPCII collider, the decays $χ_{cJ} \to pK^{-}\barΛη+ \mathrm{c.c.}$ ($J=0,1,2$) are observed for the first time, with statistical significances exceeding $5σ$ for all three $χ_{cJ}$ states. The measured branching fractions are $\mathcal{B}(χ_{c0} \to pK^{-}\barΛη+ \mathrm{c.c.}) = (5.3 \pm 0.7 \pm 0.5) \times 10^{-5}$, $\mathcal{B}(χ_{c1} \to pK^{-}\barΛη+ \mathrm{c.c.}) = (9.8 \pm 0.6 \pm 0.6) \times 10^{-5}$, and $\mathcal{B}(χ_{c2} \to pK^{-}\barΛη+ \mathrm{c.c.}) = (9.3 \pm 0.6 \pm 0.6) \times 10^{-5}$, where the first uncertainties are statistical and the second are systematic. Structures consistent with the known hyperon resonances $Λ(1520)$ and $\barΛ(1690)$ are seen in the $pK^{-}$ and $\barΛη$ invariant mass spectra, respectively. The reported branching fractions include both resonant and non-resonant contributions. These results provide new experimental information on hadronic decays of $P$-wave charmonium states and contribute to the understanding of baryon production and hadronization dynamics in the nonperturbative QCD regime.
△ Less
Submitted 6 July, 2026;
originally announced July 2026.
-
First measurement of the masses of the $Υ_1(1D)$ and $Υ_3(1D)$ states and the energy dependence of the cross sections for $e^+e^-\toΥ_J(1D)η$ and $e^+e^-\toΥ_J(1D)π^+π^-$
Authors:
Belle,
Belle II Collaborations,
:,
M. Abumusabh,
I. Adachi,
K. Adamczyk,
A. Aggarwal,
H. Ahmed,
Y. Ahn,
M. Akdag,
N. Akopov,
S. Alghamdi,
M. Alhakami,
N. Althubiti,
K. Amos,
M. Angelsmark,
N. Anh Ky,
C. Antonioli,
K. Arai,
H. Atmacan,
T. Aushev,
V. Aushev,
R. Ayad,
V. Babu,
H. Bae
, et al. (376 additional authors not shown)
Abstract:
We study the processes $e^+e^-\toΥ_J(1D)η$ and $e^+e^-\toΥ_J(1D)π^+π^-$ at center-of-mass energies $\sqrt{s}$=(10.73 -- 11.02) GeV using a $142.5\,\mathrm{fb}^{-1}$ data sample, including 122~fb$^{-1}$ near the $Υ$(10860) peak ($\sqrt{s}$ = 10.866 GeV), collected with the Belle detector at the KEKB asymmetric-energy $e^+e^-$ collider. From the peak sample, the products of Born cross section times…
▽ More
We study the processes $e^+e^-\toΥ_J(1D)η$ and $e^+e^-\toΥ_J(1D)π^+π^-$ at center-of-mass energies $\sqrt{s}$=(10.73 -- 11.02) GeV using a $142.5\,\mathrm{fb}^{-1}$ data sample, including 122~fb$^{-1}$ near the $Υ$(10860) peak ($\sqrt{s}$ = 10.866 GeV), collected with the Belle detector at the KEKB asymmetric-energy $e^+e^-$ collider. From the peak sample, the products of Born cross section times branching fraction are obtained for $σ_{\rm Born}(e^+e^-\toΥ_J(1D)η)$ or $σ_{\rm Born}(e^+e^-\toΥ_J(1D)π^+π^-)$ and ${\cal B}(Υ_J(1D)\toχ_{b1}γ)$ or ${\cal B}(Υ_J(1D)\toχ_{b2}γ)$ for each $Υ_J(1D)$ state. The corresponding branching fractions for $Υ(10860)$ decays are also obtained. The significances of the $Υ_1(1D)$, $Υ_2(1D)$, and $Υ_3(1D)$ signals are 4.8$σ$, ${>}10σ$, and 3.0$σ$, respectively, including systematic uncertainties. The mass for $Υ_2(1D)$ is measured to be $(10167.0\pm 1.0\pm 0.2)$ MeV/$c^2$, where the first and second uncertainties are statistical and systematic. The mass splittings $Δm_{12}=m(Υ_2(1D))-m(Υ_1(1D))$ and $Δm_{23}=m(Υ_3(1D))-m(Υ_2(1D))$ are $(11.8\pm1.5\pm0.4)$ MeV/$c^2$ and $(7.6\pm2.4\pm0.6)$ MeV/$c^2$, respectively.~We determine the energy dependence of the cross sections for $e^+e^-\toΥ_J(1D)η$ and $e^+e^-\toΥ_J(1D)π^+π^-$ for the $Υ_1(1D)$, $Υ_2(1D)$, and $Υ_3(1D)$ states, combined.
△ Less
Submitted 5 July, 2026;
originally announced July 2026.
-
SpecGradFilter: A Spectral Gradient Filtering Framework for Taming Federated Heterogeneity
Authors:
Liyang Yuan,
Yibo Yang,
Dandan Guo,
Peter Richtarik,
Zhouchen Lin
Abstract:
Federated Learning (FL) is fundamentally challenged by statistical heterogeneity, where non-identically distributed (non-IID) data induces client drift that severely hampers global convergence. While existing approaches attempt to mitigate this drift through spatial-domain gradient correction or regularization, they overlook the intrinsic spectral structure of optimization signals. In this work, w…
▽ More
Federated Learning (FL) is fundamentally challenged by statistical heterogeneity, where non-identically distributed (non-IID) data induces client drift that severely hampers global convergence. While existing approaches attempt to mitigate this drift through spatial-domain gradient correction or regularization, they overlook the intrinsic spectral structure of optimization signals. In this work, we revisit client drift from a novel frequency-domain perspective and uncover a critical Spectral Bias of Drift: inter-client gradient divergence is predominantly concentrated in low-frequency components which encode client-specific distributional shifts, while high-frequency components representing fine-grained features remain relatively consistent. Motivated by this, we propose SpecGradFilter, a unified Spectral Gradient Filtering Framework that tames heterogeneity by suppressing discordant low-frequency signals. Crucially, we demonstrate that SpecGradFilter is a generalizable principle, effective not only via precise FFT-based truncation but also through spatial approximations like Gaussian detrending. Extensive experiments on benchmarks such as CIFAR-10/100 and Tiny-ImageNet demonstrate that SpecGradFilter significantly performs better performance in highly Non-IID settings with negligible communication overhead, establishing a new paradigm for robust federated optimization.
△ Less
Submitted 5 July, 2026;
originally announced July 2026.
-
FedFFT: Taming Client Drift in Federated SAM via Spectral Perturbation Filtering
Authors:
Liyang Yuan,
Yibo Yang,
Dandan Guo
Abstract:
Federated Learning (FL) enables decentralized training without data sharing, but suffers from statistical heterogeneity across clients, leading to client drift, poor generalization, and sharp minima compared to centralized training. Sharpness-Aware Minimization (SAM) has emerged as a promising approach to improve generalization, yet its application in federated learning still suffers from divergen…
▽ More
Federated Learning (FL) enables decentralized training without data sharing, but suffers from statistical heterogeneity across clients, leading to client drift, poor generalization, and sharp minima compared to centralized training. Sharpness-Aware Minimization (SAM) has emerged as a promising approach to improve generalization, yet its application in federated learning still suffers from divergence problems, since perturbations are computed locally and reflect client-specific loss geometries. To better understand this issue, we provide experimental evidence from a new perspective, the frequency domain, for SAM perturbations in federated settings, revealing that inter-client perturbation inconsistencies are predominantly concentrated in the low-frequency spectrum. Motivated by this insight, we propose Federated learning with Frequency-domain Filtering of SAM perturbations (FedFFT). It is a lightweight and plug-and-play method that filters out low-frequency components of SAM perturbations without requiring additional communication, thereby suppressing inconsistent components in client updates while preserving consistent learning signals. Extensive experiments across multiple benchmarks and diverse backbones demonstrate that FedFFT consistently outperforms SAM-based FL methods, particularly under severe non-IID distributions. These results highlight the effectiveness, scalability, and general applicability of our frequency-domain perspective for sharpness-aware federated optimization.
△ Less
Submitted 5 July, 2026;
originally announced July 2026.
-
GEAR: Guided End-to-End AutoRegression for Image Synthesis
Authors:
Bin Lin,
Zheyuan Liu,
Chenguo Lin,
Sixiang Chen,
Yunyang Ge,
Yunlong Lin,
Jianwei Zhang,
Miles Yang,
Zhao Zhong,
Liefeng Bo,
Li Yuan
Abstract:
Visual generative models are typically trained in two stages. A tokenizer is first trained for reconstruction and then frozen, after which a generator is trained on its discrete indices or continuous latents. This decoupling leaves the tokenizer unaware of what the generator finds easy to model. We present GEAR (Guided End-to-end AutoRegression), which trains a vector-quantized (VQ) tokenizer and…
▽ More
Visual generative models are typically trained in two stages. A tokenizer is first trained for reconstruction and then frozen, after which a generator is trained on its discrete indices or continuous latents. This decoupling leaves the tokenizer unaware of what the generator finds easy to model. We present GEAR (Guided End-to-end AutoRegression), which trains a vector-quantized (VQ) tokenizer and an autoregressive (AR) generator jointly and end-to-end, guided by representation alignment. The key obstacle is that the VQ index fed to the AR model is non-differentiable, so gradients cannot reach the tokenizer, and a straight-through estimator collapses. GEAR resolves this with a dual read-out of the codebook assignment. A hard, one-hot branch trains the AR with next-token prediction, while a differentiable soft branch carries a representation-alignment loss that flows back to guide only the tokenizer. The AR model thereby steers its tokenizer toward an index distribution it can predict more easily. This shifts the alignment burden from the tokenizer to the AR: the tokenizer's own features become less DINOv2-like while the AR's become more so, the opposite of diffusion-side recipes that make the latent itself semantic. GEAR speeds up ImageNet gFID convergence by up to 10x relative to the strong LlamaGen-REPA baseline, learns markedly better patch-level and spatially-coherent features, and generalizes across quantizers (VQVAE, LFQ, IBQ) and to text-to-image generation.
△ Less
Submitted 30 June, 2026;
originally announced June 2026.
-
FARS: A Fully Automated Research System Deployed at Scale
Authors:
Qiong Tang,
Tianxiang Sun,
Xiangkun Hu,
Xiangyang Liu,
Yiran Chen,
Yunfan Shao,
Bobo Li,
Changze Lv,
Cheng Xu,
Chengsong Huang,
Chunyang Li,
Dizhan Xue,
Hao Bai,
Haodong Duan,
Hengquan Guo,
Hongyang He,
Hongyi Chen,
Hui Shen,
Jiahao Yuan,
Jiankai Sun,
Jikang Cheng,
Jinfeng Xu,
Jingqi Tong,
Jingye Chen,
Jinxiu Liu
, et al. (32 additional authors not shown)
Abstract:
Recent automated research systems show that language-model agents can generate hypotheses, run experiments, and write complete manuscripts, but most evidence still comes from selected examples, human-framed topics, or a few pre-defined research tasks. We present FARS (Fully Automated Research System), a fully automated AI-for-AI research system designed to operate across research topics at scale.…
▽ More
Recent automated research systems show that language-model agents can generate hypotheses, run experiments, and write complete manuscripts, but most evidence still comes from selected examples, human-framed topics, or a few pre-defined research tasks. We present FARS (Fully Automated Research System), a fully automated AI-for-AI research system designed to operate across research topics at scale. FARS autonomously generates and advances projects through ideation, planning, experimentation, and writing, using stage-specific agents coordinated through a shared workspace that records proposals, code, logs, results, and manuscripts. In its first public deployment, FARS produced 166 complete research papers spanning 67 fine-grained AI/ML topics while preserving intermediate artifacts as an auditable corpus rather than a curated set of successes. We evaluate this corpus with 282 structured reviews from volunteer reviewers covering 140 papers, including overall ratings, sub-scores, integrity checks, and LLM-use disclosure. The reviews indicate that FARS can produce review-worthy and occasionally strong AI/ML research artifacts in a large-scale public deployment, while also exposing recurring failure modes in narrow experimental scope, methodological limitations, and integrity issues.
△ Less
Submitted 13 July, 2026; v1 submitted 30 June, 2026;
originally announced June 2026.
-
Truth or Sophistry? LoFa: A Benchmark for LLM Robustness Against Logical Fallacies
Authors:
Xudong Shen,
Li Yuan,
Ye Chen,
Xin Wu,
Yi Cai,
Zhiyong Wu
Abstract:
Large Language Models (LLMs) exhibit strong semantic capabilities, yet their resilience to manipulative linguistic patterns such as logical fallacies remains underexplored. Prior work has primarily examined whether LLMs can identify or classify fallacies, leaving their robustness against fallacious persuasion insufficiently studied. To address this gap, we introduce LoFa (Logical Fallacy), a compr…
▽ More
Large Language Models (LLMs) exhibit strong semantic capabilities, yet their resilience to manipulative linguistic patterns such as logical fallacies remains underexplored. Prior work has primarily examined whether LLMs can identify or classify fallacies, leaving their robustness against fallacious persuasion insufficiently studied. To address this gap, we introduce LoFa (Logical Fallacy), a comprehensive benchmark for evaluating LLM robustness against fallacies. LoFa is constructed through a multi-agent pipeline that pairs factual questions with fallacious arguments, and is accompanied by a multi-round debate framework for assessing model resilience under sustained adversarial persuasion. To disentangle fallacy robustness from a model's inherent knowledge limitations, we further propose Logical Fallacy Resistance at k (LFR@k), a metric that quantifies resistance to fallacious attacks. Experiments show that LLMs exhibit varying levels of robustness across different fallacy types, revealing distinct vulnerability profiles among models.
△ Less
Submitted 29 June, 2026;
originally announced June 2026.
-
SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation
Authors:
Juncheng Ma,
Yuxuan Du,
Yanan Sun,
Zhening Xing,
Changlin Li,
Zhenyu Tang,
Bo Li,
Peng-Tao Jiang,
Li Yuan,
Daquan Zhou,
Yonghong Tian
Abstract:
Diffusion Transformers (DiTs) have significantly advanced audio-driven portrait animation, but their high computational cost leads to substantial inference latency. Although training-free diffusion caching accelerates inference significant, existing methods are primarily developed for text-conditioned generation and overlook the spatial and modality imbalances inherent in audio-driven portrait ani…
▽ More
Diffusion Transformers (DiTs) have significantly advanced audio-driven portrait animation, but their high computational cost leads to substantial inference latency. Although training-free diffusion caching accelerates inference significant, existing methods are primarily developed for text-conditioned generation and overlook the spatial and modality imbalances inherent in audio-driven portrait animation. In this paper, we propose SyncCache, a training-free caching acceleration method tailored for DiT-based portrait animation that explicitly exploits asymmetric dynamics. Specifically, high-frequency dynamics driven by audio conditions and concentrated in human regions are more challenging and critical to cache and reuse than the low-frequency visual background in portrait animation. First, we introduce Spatially-Asymmetric Probing to prioritize error sensitivity in dynamic human region. Second, through Modality-Decoupled Caching, we bypass heavy DiT block by reusing stable inter-block residuals, while continuously recomputing lightweight audio blocks to preserve precise lip synchronization. Furthermore, we introduce a cache ratio to control cache capacity and formulate memory-adaptive cache selection as an offline dynamic programming problem without online overhead. Extensive experiments demonstrate that SyncCache achieves superior speed-quality trade-offs, delivering up to 4.12x acceleration on HunyuanVideo-Avatar and 3.75x on Wan-S2V with near-lossless visual fidelity and precise audio alignment.
△ Less
Submitted 29 June, 2026;
originally announced June 2026.
-
Study of the $e^+e^-\to π^+π^-D_s^+D_s^-$ process from $\sqrt{s}$ = 4.42 to 4.95 GeV at BESIII
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
M. S. Anderson,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko
, et al. (762 additional authors not shown)
Abstract:
Based on $8.5~{\rm fb}^{-1}$ of $e^+e^-$ collision data collected at center-of-mass energies between 4.42 and 4.95 GeV with the BESIII detector at the BEPCII storage ring, we investigate the process $e^+e^-\to π^+π^-D_s^+D_s^-$. With no significant signal observed, upper limits on the Born cross sections of $e^+e^-\to π^+π^-D_s^+D_s^-$ at each energy value are determined at the 90% confidence leve…
▽ More
Based on $8.5~{\rm fb}^{-1}$ of $e^+e^-$ collision data collected at center-of-mass energies between 4.42 and 4.95 GeV with the BESIII detector at the BEPCII storage ring, we investigate the process $e^+e^-\to π^+π^-D_s^+D_s^-$. With no significant signal observed, upper limits on the Born cross sections of $e^+e^-\to π^+π^-D_s^+D_s^-$ at each energy value are determined at the 90% confidence level. Additionally, a search for intermediate charmonium-like resonances is performed in the $M(D_s^+D_s^-)$ invariant-mass spectrum, but no significant resonant structures are observed with the current statistics.
△ Less
Submitted 27 June, 2026;
originally announced June 2026.
-
Charged-lepton identification at Belle~II
Authors:
Belle II Collaboration,
M. Abumusabh,
I. Adachi,
A. Aggarwal,
H. Ahmed,
Y. Ahn,
H. Aihara,
M. Akdag,
N. Akopov,
S. Alghamdi,
M. Alhakami,
N. Althubiti,
K. Amos,
M. Angelsmark,
N. Anh Ky,
C. Antonioli,
K. Arai,
H. Atmacan,
V. Aushev,
R. Ayad,
V. Babu,
H. Bae,
N. K. Baghel,
P. Bambade,
Sw. Banerjee
, et al. (387 additional authors not shown)
Abstract:
Effective particle identification capabilities are a strategic priority for the physics program of the Belle~II experiment. We describe the algorithms used at Belle~II for identifying electrons and muons and separating them from charged hadrons. We present the performance obtained by the experiment during Run 1, which consists of 428 fb$^{-1}$ of data collected at the energy-asymmetric $e^+e^-$ co…
▽ More
Effective particle identification capabilities are a strategic priority for the physics program of the Belle~II experiment. We describe the algorithms used at Belle~II for identifying electrons and muons and separating them from charged hadrons. We present the performance obtained by the experiment during Run 1, which consists of 428 fb$^{-1}$ of data collected at the energy-asymmetric $e^+e^-$ collider SuperKEKB between 2019 and 2022 at center-of-mass energies near the mass of the $Υ(4S)$.
△ Less
Submitted 24 June, 2026;
originally announced June 2026.
-
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models
Authors:
Haoxiang Sun,
Tao Wang,
Li Yuan,
Jian Zhao,
Jiancheng Lv
Abstract:
Multimodal Large Language Models (MLLMs) have recently made remarkable progress in unifying vision-language understanding and reasoning, especially following the introduction of models such as OpenAI's O-series and DeepSeek's R-series, which have driven a paradigm shift toward perception-centric intelligence. However, there remains a lack of systematic surveys that examine perception from a truly…
▽ More
Multimodal Large Language Models (MLLMs) have recently made remarkable progress in unifying vision-language understanding and reasoning, especially following the introduction of models such as OpenAI's O-series and DeepSeek's R-series, which have driven a paradigm shift toward perception-centric intelligence. However, there remains a lack of systematic surveys that examine perception from a truly unified vision-language perspective -- one that treats vision and language as an inseparable modality. Existing reviews are often fragmented, focusing separately on either vision or language, and thus rarely capture the cross-modal evolution of perception as an integrated capability. To bridge this gap, we present the first systematic survey of unified vision-language perception in MLLMs. Specifically, we (1) formalize MLLM perception as an intrinsic, unified vision-language capability analogous to human innate perception, (2) introduce a five-stage taxonomy tracing the paradigm evolution of MLLM perception and survey representative methods and milestones at each phase, and (3) identify open challenges and outline promising research directions toward truly general, unified multimodal intelligence. We hope our study will provide both a foundational understanding and an actionable roadmap to foster further innovation on the path toward artificial general intelligence (AGI).
△ Less
Submitted 24 June, 2026;
originally announced June 2026.