-
On Geometric Models of String Algebras: Uniqueness of Surfaces and Existence of Red Punctures
Authors:
Zheng Xin,
Lingchun Zhang
Abstract:
A geometric model for string algebras was recently established in \cite{BC24}. Building upon this framework, we characterize the class of string algebras whose geometric models are unique up to equivalence of labelled tiled surfaces.
Moreover, we provide a necessary and sufficient condition for all geometric models of a string algebra to be entirely free of red punctures, and further give a comb…
▽ More
A geometric model for string algebras was recently established in \cite{BC24}. Building upon this framework, we characterize the class of string algebras whose geometric models are unique up to equivalence of labelled tiled surfaces.
Moreover, we provide a necessary and sufficient condition for all geometric models of a string algebra to be entirely free of red punctures, and further give a combinatorial description of string algebras with a common red puncture across all geometric models. In addition, we derive a sufficient condition for a string algebra guaranteeing the presence of red punctures in all geometric models.
△ Less
Submitted 14 August, 2026;
originally announced August 2026.
-
Error-Aware Reverse Auction Mechanism for Large Language Model Routing
Authors:
Haolong Chen,
Zhengyuan Xin,
Liang Zhang,
Lei Xue,
Guangxu Zhu
Abstract:
Routing each query to a cost-effective large language model (LLM) is critical for balancing quality and cost, yet most routers rely on a centralized task center to predict model performance, creating an information-risk mismatch and a scalability bottleneck as the model pool grows. We propose a market-based routing paradigm that shifts ex-ante prediction to LLM providers via a reverse auction, whe…
▽ More
Routing each query to a cost-effective large language model (LLM) is critical for balancing quality and cost, yet most routers rely on a centralized task center to predict model performance, creating an information-risk mismatch and a scalability bottleneck as the model pool grows. We propose a market-based routing paradigm that shifts ex-ante prediction to LLM providers via a reverse auction, where providers bid with self-predicted success probabilities and execution costs. To account for inherently noisy provider predictions and center evaluations, we introduce the \textit{\textbf{E}rror-\textbf{A}ware \textbf{R}everse \textbf{A}uction \textbf{M}echanism} (EA-RAM), which explicitly models this inherent Dual Error. We prove that EA-RAM is Bayesian incentive compatible and individually rational under the Dual Error, establish sufficient conditions for center rationality, and derive an explicit welfare-loss bound. We further identify robustness effects: opposite-signed errors can cancel, vanishing-tail link functions (e.g., logistic) stabilize clear-cut cases via saturation, and extra noise smooths belief maps, reducing the gains from marginal manipulation. Experiments on simulations and real-world benchmarks show that EA-RAM is robust to the Dual Error and achieves a better cost--performance Pareto frontier than centralized baselines, with additional gains when providers contribute local information, validating its practical effectiveness.
△ Less
Submitted 12 August, 2026;
originally announced August 2026.
-
OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation
Authors:
Kaiyu Li,
Zepeng Xin,
Zixuan Jiang,
Jing Fu,
Lanxuan Xue,
Lingyu Zhang,
Xiangyong Cao
Abstract:
Open-vocabulary Earth observation (EO) aims to localize geospatial concepts specified in natural language rather than a fixed label set. Existing benchmarks, however, usually cover narrow category vocabularies or limited query forms. To fill this gap, we introduce OVEarth-Bench, which extends existing evaluation in two directions: category breadth, through broad hierarchical category coverage with…
▽ More
Open-vocabulary Earth observation (EO) aims to localize geospatial concepts specified in natural language rather than a fixed label set. Existing benchmarks, however, usually cover narrow category vocabularies or limited query forms. To fill this gap, we introduce OVEarth-Bench, which extends existing evaluation in two directions: category breadth, through broad hierarchical category coverage with positive and negative expressions, and query diversity, through vocabulary, referring, and reasoning queries. The benchmark supports mask and box localization under a unified zero-shot protocol. We evaluate a broad set of general and EO-specific methods. The evaluation reveals that: (1) the performance of current methods remains limited, while broader category coverage yields more stable model rankings; (2) MLLM-based methods achieve the strongest overall performance; and (3) EO-specific methods generally underperform general models and rarely match the strongest methods. These findings provide guidance for future open-vocabulary EO method design and highlight the importance of developing more realistic, diverse, high-quality, and large-scale benchmarks for reliable evaluation. Our data and evaluation package are released at https://earth-insights.github.io/OVEarth-bench.
△ Less
Submitted 2 August, 2026; v1 submitted 29 July, 2026;
originally announced July 2026.
-
EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization
Authors:
Xinbang Dai,
Zheyu Xin,
Huikang Hu,
Lin Ren,
Rihui Jin,
Guohui Xiao,
Guilin Qi,
Kuicai Dong,
Zhaocheng Du,
Yuyang Zhang
Abstract:
Large Reasoning Models (LRMs) often suffer from overthinking due to redundant verification steps. Existing approaches for mitigating overthinking, such as fast-slow thinking switching and reasoning trajectory compression, fail to make a fine-grained distinction between beneficial and redundant steps within the LRM's reasoning process, and may thus impair reasoning capability in their pursuit of ef…
▽ More
Large Reasoning Models (LRMs) often suffer from overthinking due to redundant verification steps. Existing approaches for mitigating overthinking, such as fast-slow thinking switching and reasoning trajectory compression, fail to make a fine-grained distinction between beneficial and redundant steps within the LRM's reasoning process, and may thus impair reasoning capability in their pursuit of efficiency. To simultaneously improve reasoning efficiency and capability, we propose EvoThink, a framework that reduces redundant verification and encourages the exploration of new reasoning paths. EvoThink comprises two key components: Self-Pruning Training (SPT), an unsupervised method that iteratively prunes redundant reasoning steps and self-trains on the concise trajectories; and Aha-Moment Preference Optimization (AMPO), which, inspired by genetic algorithms, identifies valuable failed reasoning attempts, synthesizes from-wrong-to-right aha-moment data, and optimizes the model to internalize this reasoning pattern. Extensive evaluations across mathematical reasoning and code generation benchmarks demonstrate that EvoThink not only substantially reduces inference-time token usage but also improves the reasoning capability of LRMs.
△ Less
Submitted 22 July, 2026;
originally announced July 2026.
-
Robust Multimodal Dynamic Object Segmentation
Authors:
Zhe Xin,
Hanzhi Chang,
Penghui Huang,
Yinian Mao,
Guoquan Huang
Abstract:
Dynamic object segmentation plays a critical role in many visual applications such as static scene reconstruction from dynamic videos. However, existing optical flow-based methods fail to ensure consistent static/dynamic segmentation along object boundaries, while 3D reconstruction-based approaches are highly sensitive to reconstruction errors. To address these limitations, we present a dynamic ob…
▽ More
Dynamic object segmentation plays a critical role in many visual applications such as static scene reconstruction from dynamic videos. However, existing optical flow-based methods fail to ensure consistent static/dynamic segmentation along object boundaries, while 3D reconstruction-based approaches are highly sensitive to reconstruction errors. To address these limitations, we present a dynamic object segmentation framework that can generate both precise and complete dynamic masks by integrating multimodal cues including 2D point tracks, 3D reconstruction, and semantic information. We design a network combining Transformer architectures with feature clustering aggregation modules to perform static/dynamic classification of multimodal feature trajectories. It enables the model to adaptively determine which type of feature should dominate based on the characteristics of each scene, while also mitigating the impact of feature degradation. Additionally, we introduce a novel point-query-based SAM post-processing method capable of handling multiple objects within a single mask. Extensive experiments demonstrate that our approach achieves state-of-the-art performance in both dynamic object segmentation and static scene reconstruction tasks.
△ Less
Submitted 20 July, 2026;
originally announced July 2026.
-
Enhanced stability and asymptotic limits to the non-isentropic compressible fluid-particle interaction model with thermal effects
Authors:
Fucai Li,
Jinkai Ni,
Zhouping Xin
Abstract:
In Einstein's seminal work [Ann. Physik, 17 (1905), 549-560], he pointed out that the temperature of a fluid influences the motion of suspended particles dramatically. To describe the effect of the temperature in this physical process more precisely, Boudin et al. [ESAIM Proc., 28 (2009), 195-210] introduced a new fluid-particle interaction model containing of the non-isentropic compressible Euler…
▽ More
In Einstein's seminal work [Ann. Physik, 17 (1905), 549-560], he pointed out that the temperature of a fluid influences the motion of suspended particles dramatically. To describe the effect of the temperature in this physical process more precisely, Boudin et al. [ESAIM Proc., 28 (2009), 195-210] introduced a new fluid-particle interaction model containing of the non-isentropic compressible Euler equations for the fluid and a nonlinear Vlasov-Fokker-Planck type equation for the particles. By adding some viscous and heat conductive terms to the fluid part of this model, Mu and Wang [Calc. Var. Partial Differential Equations, 59 (2020), Paper no. 110] established the global existence of classical solutions near an equilibrium state.
In this paper, through establishing the uniform a priori estimates with respect to the viscosity and heat conductivity coefficients and taking the combined zero viscosity and heat conductivity limits, we show that the model introduced by Boudin et al. still admits a global classical solution and enjoys optimal decay rates thereby improving Mu and Wang's results and confirming Einstein's predications. Our work indicates that the presence of particles indeed emanates new dissipation effects on the non-isentropic compressible fluid-particle model via the differences between the macroscopic velocity of the particles and the fluid velocity, and the macroscopic temperature of the particles and the fluid temperature, which is significantly different from the case of pure non-isentropic compressible Euler equations. To achieve these goals, we have developed new ideas and techniques to surmount substantial obstacles caused by the absence of viscosity and heat conductivity, and the nonlinear interactions between the fluid and particles.
△ Less
Submitted 19 July, 2026;
originally announced July 2026.
-
An MLIR-Based Compilation Method for Large Language Models
Authors:
Pengchao Hu,
Zhibin Xin,
Yifan Chen,
Yangyang Zhou,
Liang Wang,
Xin Zhang
Abstract:
Large Language Models (LLMs) have become the dominant workload on modern AI accelerators, yet deploying them on specialized hardware still faces two core challenges: how to import a trained model into a compiler-friendly intermediate representation, and how to efficiently schedule the autoregressive inference loop under limited on-chip memory. This paper presents an MLIR (Multi-Level Intermediate…
▽ More
Large Language Models (LLMs) have become the dominant workload on modern AI accelerators, yet deploying them on specialized hardware still faces two core challenges: how to import a trained model into a compiler-friendly intermediate representation, and how to efficiently schedule the autoregressive inference loop under limited on-chip memory. This paper presents an MLIR (Multi-Level Intermediate Representation) based compilation method for large language models, illustrated using two dialects of operators, TopOp and TpuOp. TopOp serves as a high-level graph dialect that is independent of both the source framework and the target chip, and is responsible for expressing model semantics; TpuOp serves as the target hardware dialect, carrying chip-related decisions such as quantization, layer groups, and memory layout. A model is first represented as TopOp, then lowered layer by layer to TpuOp, and finally a deployable binary is generated. In addition, each Transformer layer is split into three stages for static compilation: prefill, prefill_kv (prefill with historical key-value cache), and decode, so as to accommodate the different computational characteristics of prompt-parallel processing and per-token generation. The method has been implemented in the TPU-MLIR compiler {https://github.com/sophgo/tpu-mlir} and the LLM-TPU deployment project {https://github.com/sophgo/LLM-TPU}, supporting a variety of generative models including the Qwen, Llama, InternVL, and MiniCPM-V series, as well as multiple quantization and deployment forms such as GPTQ, AWQ, and AutoRound.
△ Less
Submitted 25 July, 2026; v1 submitted 17 July, 2026;
originally announced July 2026.
-
Global dimension of a string algebra
Authors:
Zheng Xin,
Lingchun Zhang
Abstract:
In this paper, we characterize the global dimension of a string algebra by using combinatorial methods. Moreover, we establish a necessary and sufficient condition for when the global dimension of a string algebra is infinite.
In this paper, we characterize the global dimension of a string algebra by using combinatorial methods. Moreover, we establish a necessary and sufficient condition for when the global dimension of a string algebra is infinite.
△ Less
Submitted 20 July, 2026; v1 submitted 13 July, 2026;
originally announced July 2026.
-
A multi-architecture study of specificity refinement and false-positive mechanism analysis in prostate MRI
Authors:
Yongbo Shu,
Kewen Chen,
Yifeng Yuan,
Zirui Xin,
Luo Lei,
Yang Yang,
Xi Chen,
Aijing Luo
Abstract:
Objectives: To characterize residual false positives in prostate MRI detection, and to evaluate a lightweight post-hoc refinement head for case-level specificity. Materials and Methods: This retrospective study used PI-CAI (5-fold cross-validation) and Prostate158 (n=158; external). A context-aware evidence head and an 89,216-parameter refinement head were trained on a frozen detection backbone; t…
▽ More
Objectives: To characterize residual false positives in prostate MRI detection, and to evaluate a lightweight post-hoc refinement head for case-level specificity. Materials and Methods: This retrospective study used PI-CAI (5-fold cross-validation) and Prostate158 (n=158; external). A context-aware evidence head and an 89,216-parameter refinement head were trained on a frozen detection backbone; the evidence head was also trained on four further backbones (bare nnU-Net, bare U-Net, bare Mamba, MIGF-Mamba). For each false-positive region, T2-weighted, apparent-diffusion-coefficient, and high-b-value contrast ratios versus peri-lesional rings were compared against ground-truth lesions and contralateral benign regions. Results: False positives were closer to true cancers than to benign tissue in evidence and raw T2-weighted and apparent-diffusion-coefficient contrast, reproducing 35/35 across five architectures (Cohen's d 1.10; FP/benign evidence ratio 2.38x) and 105/105 across modality-perturbation scenarios. On PI-CAI fold-0, refinement raised case-level specificity from 0.469 to 0.549 (+17.2%) at preserved sensitivity (0.943); 5-fold cross-validation showed fold-conditional behavior (9/15 observations positive; range -22% to +28%). On Prostate158, both models saturated (McNemar pooled p=0.69), while the false-positive contrast-matching finding replicated. Conclusion: Residual false positives are contrast-matched to cancer (sharing raw imaging features rather than histologically confirmed mimicry), reproducing across five architectures -- a data-level imaging property, not model-specific artifacts; post-hoc refinement adds practical specificity in-domain but is fold-conditional.
△ Less
Submitted 29 June, 2026;
originally announced June 2026.
-
TacVerse: A Multi-Sensor Dataset and Benchmark for Cross-Sensor Vision-Based Tactile Perception
Authors:
Lan Wei,
Gurmeher Khurana,
Sirine Bhouri,
Wenhao Hong,
Zeyuan Xin,
Qingzheng Cong,
Wen Fan,
Yanzheng Xiang,
Dandan Zhang
Abstract:
Vision-based tactile sensors (VBTSs) enable robots to infer contact geometry and force-related cues by imaging deformation through an internal camera, yet generalisation across sensor designs remains poorly understood. We present TacVerse, a multi-sensor dataset and benchmark for cross-sensor vision-based tactile perception. The dataset contains 106,800 tactile images from seven VBTSs and supports…
▽ More
Vision-based tactile sensors (VBTSs) enable robots to infer contact geometry and force-related cues by imaging deformation through an internal camera, yet generalisation across sensor designs remains poorly understood. We present TacVerse, a multi-sensor dataset and benchmark for cross-sensor vision-based tactile perception. The dataset contains 106,800 tactile images from seven VBTSs and supports three downstream tasks: shape classification, grating classification, and force regression. Experiments are conducted under three settings: within-sensor training, zero-shot cross-sensor transfer, and few-shot adaptation. Strong within-sensor performance across all tasks indicates that the collected tactile observations are informative for the target objectives. Direct cross-sensor transfer, however, leads to substantial degradation. Shape classification is comparatively robust, whereas grating classification and force regression are more sensitive to sensor shift. Few-shot adaptation for force regression consistently improves performance on unseen target sensors but does not fully close the gap to within-sensor upper bounds. A representation study further shows that MAE (Masked Autoencoder) pretraining provides the most consistent gains across tasks and sensors. TacVerse provides a controlled testbed for studying sensor shift, data-efficient adaptation, and self-supervised learning in tactile perception.
△ Less
Submitted 24 June, 2026;
originally announced June 2026.
-
Act As a Real Researcher: A Suite of Benchmarks Evaluating Frontier LLMs and Agentic Harnesses in Research Lifecycle
Authors:
Jiayu Wang,
Weijiang Lv,
Bowen Fu,
Jing Fu,
Jiayi Song,
Lingyu Zhang,
Lanxuan Xue,
Luodi Chen,
Zepeng Xin,
Kaiyu Li,
Xiangyong Cao
Abstract:
As foundation models advance and agent scaffolding becomes increasingly sophisticated, agents have demonstrated remarkable proficiency in complex, long-horizon coding tasks and even autonomous experiment execution. Despite their evolution from research assistants into autonomous research agents, these systems still exhibit significant limitations in field sensitivity, research ethics, and nuanced…
▽ More
As foundation models advance and agent scaffolding becomes increasingly sophisticated, agents have demonstrated remarkable proficiency in complex, long-horizon coding tasks and even autonomous experiment execution. Despite their evolution from research assistants into autonomous research agents, these systems still exhibit significant limitations in field sensitivity, research ethics, and nuanced scientific judgment. Consequently, frontier agents remain unable to fully replace human researchers. To bridge this gap, we conceptualize the AARR (Act As a Real Researcher) benchmark series. Unlike existing benchmarks that primarily assess macro-level execution capabilities, AARR focuses on whether agents can emulate the professionalism, thoroughness, and nuanced reasoning that characterize human researchers in granular research scenarios. In this work, we propose AARRI-Bench (Act As a Real Research Intern), the first benchmark in this series. We conduct extensive experiments across frontier models and agentic systems, revealing that even the best-performing configuration (Mini-SWE-Agent with Claude Opus 4.7) achieves only 68.3\% success rate, frequently overlooking subtle yet critical details that are obvious to real human researchers. Our results indicate that developing researcher-like AI requires further exploration of research behavior, rather than merely complex scaffolding. Our data is released at https://github.com/AARR-bench/AARRI-bench.
△ Less
Submitted 5 June, 2026;
originally announced June 2026.
-
Weierstrass Positional Encoding for Vision Transformers
Authors:
Zhihang Xin,
Rui Wang,
Xitong Hu,
Xiaojun Wu
Abstract:
Vision Transformers have achieved remarkable success in computer vision, but their common use of learnable one-dimensional positional encodings weakens the inherent two-dimensional spatial structure of images after patch flattening. Existing positional encodings often lack geometric constraints and do not preserve a monotonic relationship between Euclidean spatial distances and sequential index di…
▽ More
Vision Transformers have achieved remarkable success in computer vision, but their common use of learnable one-dimensional positional encodings weakens the inherent two-dimensional spatial structure of images after patch flattening. Existing positional encodings often lack geometric constraints and do not preserve a monotonic relationship between Euclidean spatial distances and sequential index distances, limiting ViTs' ability to exploit spatial proximity priors. Motivated by the usefulness of periodicity in positional encoding, we propose Weierstrass elliptic Positional Encoding (WePE), a mathematically grounded method for encoding two-dimensional coordinates in the complex domain. WePE maps normalized 2D patch coordinates onto the complex plane and constructs compact four-dimensional positional features using the Weierstrass elliptic function and its derivative. The double periodicity provides a principled representation of 2D positions, and its intrinsic lattice structure naturally matches the regular geometry of image patch grids. Its nonlinear geometric properties help model spatial distance relationships more faithfully, while the algebraic addition formula enables relative positional information between arbitrary patch pairs to be derived directly from their absolute encodings. WePE is plug-and-play and resolution-agnostic, allowing seamless integration into existing ViTs. Extensive experiments show that WePE brings consistent performance gains in most settings. With precomputed lookup tables, these improvements introduce no noticeable computational or memory overhead. Additional analyses and ablation studies further validate the effectiveness of the proposed method.
△ Less
Submitted 20 May, 2026;
originally announced May 2026.
-
Stage-adaptive Token Selection for Efficient Omni-modal LLMs
Authors:
Zijie Xin,
Jie Yang,
Ruixiang Zhao,
Tianyi Wang,
Fengyun Rao,
Jing Lyu,
Xirong Li
Abstract:
Omni-modal large language models (om-LLMs) achieve unified audio-visual understanding by encoding video and audio into temporally aligned token sequences interleaved at the window level. However, processing these dense non-textual tokens throughout the LLM incurs substantial computational overhead. Although training-free token selection can reduce this cost, existing methods either focus on visual…
▽ More
Omni-modal large language models (om-LLMs) achieve unified audio-visual understanding by encoding video and audio into temporally aligned token sequences interleaved at the window level. However, processing these dense non-textual tokens throughout the LLM incurs substantial computational overhead. Although training-free token selection can reduce this cost, existing methods either focus on visual-only inputs or prune om-LLM tokens only before the LLM with fixed per-modality ratios, failing to capture how cross-modal token importance evolves across layers. To address this limitation, we first analyze the layer-wise token dependency of om-LLMs. We find that visual and audio dependencies follow a block-wise pattern and gradually weaken with depth, indicating that many late-layer non-textual tokens become redundant after cross-modal fusion. Motivated by this observation, we propose SEATS, a training-free, stage-adaptive token selection method for efficient om-LLM inference. Before the LLM, SEATS removes spatiotemporal redundancy via attention-weighted diversity selection. Inside the LLM, it progressively prunes tokens across blocks and dynamically allocates the retention budget from temporal windows to modalities using query relevance scores. In late layers, it removes all remaining non-textual tokens once cross-modal fusion is complete. Experiments on Qwen2.5-Omni and Qwen3-Omni demonstrate that SEATS effectively improves inference efficiency. Retaining only 10% of visual and audio tokens, it achieves a 9.3x FLOPs reduction and a 4.8x prefill speedup while preserving 96.3% of the original performance.
△ Less
Submitted 19 May, 2026;
originally announced May 2026.
-
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
Authors:
Ruixiang Zhao,
Jie Yang,
Zijie Xin,
Tianyi Wang,
Fengyun Rao,
Jing LYU,
Xirong Li
Abstract:
Omni-proactive streaming video understanding, i.e., autonomously deciding when to speak and what to say from continuous audio-visual streams, is an emerging capability of omni-modal large language models. Existing benchmarks fall short in three key aspects: they rely primarily on visual signals, adopt polling or fixed-timestamp protocols instead of true proactive evaluation, and cover only a limit…
▽ More
Omni-proactive streaming video understanding, i.e., autonomously deciding when to speak and what to say from continuous audio-visual streams, is an emerging capability of omni-modal large language models. Existing benchmarks fall short in three key aspects: they rely primarily on visual signals, adopt polling or fixed-timestamp protocols instead of true proactive evaluation, and cover only a limited range of tasks, preventing reliable assessment and differentiation of omni-proactive streaming models. We present OmniPro, the first benchmark to jointly evaluate omni-modal perception, proactive responding, and diverse video understanding tasks. It comprises 2,700 human-verified samples spanning 9 sub-tasks and 3 cognitive levels, covering 6 basic video understanding capabilities. Notably, 84% of samples require audio signals (speech or non-speech), and each sample is annotated with modality-isolation labels to enable fine-grained multimodal analysis. We further introduce a dual-mode evaluation protocol: Probe mode assesses content understanding by querying the model before and after each ground-truth trigger, while Online mode evaluates full proactive ability by requiring models to autonomously decide when to respond in streaming input. Evaluating 11 representative models reveals three key findings: (1) audio provides consistent gains but with highly variable utilization across models, (2) performance degrades significantly over time, indicating limited long-horizon robustness, and (3) non-speech audio perception remains the weakest dimension.
△ Less
Submitted 18 May, 2026;
originally announced May 2026.
-
DBES: A Systematic Benchmark and Metric Suite for Evaluating Expert Specialization in Large-Scale MoEs
Authors:
Jing Wang,
Hongxuan Lu,
Jazze Young,
Shu Wang,
Zhimin Xin
Abstract:
Expert specialization in Mixture-of-Experts (MoE) models remains poorly understood, with traditional evaluations conflating architectural load-balancing with functional specialization. We introduce DBES, a comprehensive diagnostic framework combining a multi-domain benchmark with five theoretically grounded metrics: Routing Specialization, Normalized Effective Rank, Domain Isolation, Routing Stiff…
▽ More
Expert specialization in Mixture-of-Experts (MoE) models remains poorly understood, with traditional evaluations conflating architectural load-balancing with functional specialization. We introduce DBES, a comprehensive diagnostic framework combining a multi-domain benchmark with five theoretically grounded metrics: Routing Specialization, Normalized Effective Rank, Domain Isolation, Routing Stiffness Score, and N-gram Expertise measures.
Critical findings demonstrate distinct specialization paradigms across models: Qwen-series exhibit modular specialization with high domain isolation, while DeepSeek and GLM employ distributed collaboration. However, we emphasize that specialization is a diagnostic dimension, necessary but not sufficient for downstream performance. Most crucially, interventional evidence validates the actionability of these metrics: by using DBES to identify high-specialization expert paths during domain-specific post-training, we achieved 66% to 94.48% improvement in specialized domains with only 15% of original training resources, demonstrating that these diagnostic tools can be converted into concrete optimization operators. This work provides the first systematic methodology for evaluating expert specialization independently of accuracy metrics, offering crucial insights for the design and post-training optimization of next-generation MoE systems.
△ Less
Submitted 18 May, 2026;
originally announced May 2026.
-
A Proof-of-Concept Study of Multitask Learning for Cranial Synthetic CT Generation Across Heterogeneous MRI Field Strengths
Authors:
Zhuoyao Xin,
Yiren Zhang,
Christopher Wu,
Dong Liu,
Chunming Gu,
Elena Greco,
Erik H. Middlebrooks,
Jun Hua,
Jia Guo
Abstract:
Accurate synthesis of computed tomography (CT) images from magnetic resonance imaging (MRI) is clinically valuable for cranial applications such as attenuation correction, radiotherapy planning, and image-guided interventions. However, heterogeneity across MRI field strengths and acquisition protocols limits the generalizability of existing methods. In this study, we formulate cranial CT synthesis…
▽ More
Accurate synthesis of computed tomography (CT) images from magnetic resonance imaging (MRI) is clinically valuable for cranial applications such as attenuation correction, radiotherapy planning, and image-guided interventions. However, heterogeneity across MRI field strengths and acquisition protocols limits the generalizability of existing methods. In this study, we formulate cranial CT synthesis as a modular, structurally coupled problem and propose a deep learning framework to improve robustness across heterogeneous MRI conditions. The model is designed to adapt to variations in field strength and imaging protocols while preserving anatomical consistency. Experiments on multi-site datasets demonstrate improved performance and generalization compared with conventional approaches. The proposed method enables reliable CT synthesis across heterogeneous MRI settings, supporting broader clinical translation.
△ Less
Submitted 30 April, 2026;
originally announced May 2026.
-
The Second Challenge on Cross-Domain Few-Shot Object Detection at NTIRE 2026: Methods and Results
Authors:
Xingyu Qiu,
Yuqian Fu,
Jiawei Geng,
Bin Ren,
Jiancheng Pan,
Zongwei Wu,
Hao Tang,
Yanwei Fu,
Radu Timofte,
Nicu Sebe,
Mohamed Elhoseiny,
Lingyi Hong,
Mingxi Cheng,
Xingqi He,
Runze Li,
Xingdong Sheng,
Wenqiang Zhang,
Jiacong Liu,
Shu Luo,
Yikai Qin,
Yaze Zhao,
Yongwei Jiang,
Yixiong Zou,
Zhe Zhang,
Yang Yang
, et al. (49 additional authors not shown)
Abstract:
Cross-domain few-shot object detection (CD-FSOD) remains a challenging problem for existing object detectors and few-shot learning approaches, particularly when generalizing across distinct domains. As part of NTIRE 2026, we hosted the second CD-FSOD Challenge to systematically evaluate and promote progress in detecting objects in unseen target domains under limited annotation conditions. The chal…
▽ More
Cross-domain few-shot object detection (CD-FSOD) remains a challenging problem for existing object detectors and few-shot learning approaches, particularly when generalizing across distinct domains. As part of NTIRE 2026, we hosted the second CD-FSOD Challenge to systematically evaluate and promote progress in detecting objects in unseen target domains under limited annotation conditions. The challenge received strong community interest, with 128 registered participants and a total of 696 submissions. Among them, 31 teams actively participated, and 19 teams submitted valid final results. Participants explored a wide range of strategies, introducing innovative methods that push the performance frontier under both open-source and closed-source tracks. This report presents a detailed overview of the NTIRE 2026 CD-FSOD Challenge, including a summary of the submitted approaches and an analysis of the final results across all participating teams. Challenge Codes: https://github.com/ohMargin/NTIRE2026_CDFSOD.
△ Less
Submitted 13 April, 2026;
originally announced April 2026.
-
Backbone-Conditional Behavior of Modality Gating in Multi-Modal Prostate MRI Segmentation: A 5-Fold Cross-Validation and Gate Mechanism Analysis
Authors:
Yongbo Shu,
Wenzhao Xie,
Shanhu Yao,
Zirui Xin,
Luo Lei,
Kewen Chen,
Aijing Luo
Abstract:
Robust segmentation of clinically significant prostate cancer (csPCa) on multi-parametric MRI must tolerate frequent degradation of its most informative diffusion sequences. Multi-modal fusion commonly employs learned modality gating under the assumption that gates implement per-sample modality quality routing -- rarely tested directly. We ask how gating behaves across backbone architectures. We s…
▽ More
Robust segmentation of clinically significant prostate cancer (csPCa) on multi-parametric MRI must tolerate frequent degradation of its most informative diffusion sequences. Multi-modal fusion commonly employs learned modality gating under the assumption that gates implement per-sample modality quality routing -- rarely tested directly. We ask how gating behaves across backbone architectures. We systematically analyze modality-isolated gated fusion (MIGF) for csPCa segmentation on two backbones (nnU-Net and Mamba) using PI-CAI (n=1500), with cross-cohort validation on Prostate158 (n=158): a factorial ablation over gating, modality dropout, and deep supervision under 5-fold cross-validation (180 trained models), plus a gate-weight and counterfactual analysis of 30 trained gating models. Modality gating is backbone-conditional. On nnU-Net, adding gating reduces the ranking score (marginal effect -0.037; gating configurations p<0.05), whereas on Mamba the gating-plus-dropout configuration improves it (+0.024, p=0.037). Gate-weight analysis explains this: nnU-Net gates collapse into a near-static modality prior (across-case SD 0.0033), while Mamba gates retain sample-dependent variation (0.0365, ~11x larger, non-overlapping); replacing per-sample gates with their training-set mean leaves nnU-Net unchanged but degrades Mamba. Modality dropout is the only component beneficial on both backbones. Under cross-cohort shift, convolutional backbones collapse to case-level specificity near zero, whereas Mamba retains it (MIGF-Mamba highest, 0.31). Learned modality gates do not universally perform per-sample quality routing; their effective behavior is conditional on the backbone's inherent modality awareness. Among tested configurations, MIGF-Mamba is the most cross-cohort robust, and training-time modality dropout is the only component beneficial across both backbones.
△ Less
Submitted 24 June, 2026; v1 submitted 12 April, 2026;
originally announced April 2026.
-
Fundus-R1: Training a Fundus-Reading MLLM with Knowledge-Aware Reasoning on Public Data
Authors:
Yuchuan Deng,
Qijie Wei,
Kaiheng Qian,
Jiazhen Liu,
Zijie Xin,
Bangxiang Lan,
Jingyu Liu,
Jianfeng Dong,
Xirong Li
Abstract:
Fundus imaging such as CFP, OCT and UWF is crucial for the early detection of retinal anomalies and diseases. Fundus image understanding, due to its knowledge-intensive nature, poses a challenging vision-language task. An emerging approach to addressing the task is to post-train a generic multimodal large language model (MLLM), either by supervised finetuning (SFT) or by reinforcement learning wit…
▽ More
Fundus imaging such as CFP, OCT and UWF is crucial for the early detection of retinal anomalies and diseases. Fundus image understanding, due to its knowledge-intensive nature, poses a challenging vision-language task. An emerging approach to addressing the task is to post-train a generic multimodal large language model (MLLM), either by supervised finetuning (SFT) or by reinforcement learning with verifiable rewards (RLVR), on a considerable amount of in-house samples paired with high-quality clinical reports. However, these valuable samples are not publicly accessible, which not only hinders reproducibility but also practically limits research to few players. To overcome the barrier, we make a novel attempt to train a reasoning-enhanced fundus-reading MLLM, which we term Fundus-R1, using exclusively public datasets, wherein over 94\% of the data are annotated with only image-level labels. Our technical contributions are two-fold. First, we propose a RAG-based method for composing image-specific, knowledge-aware reasoning traces. Such auto-generated traces link visual findings identified by a generic MLLM to the image labels in terms of ophthalmic knowledge. Second, we enhance RLVR with a process reward that encourages self-consistency of the generated reasoning trace in each rollout. Extensive experiments on three fundus-reading benchmarks, i.e., FunBench, Omni-Fundus and GMAI-Fundus, show that Fundus-R1 clearly outperforms multiple baselines, including its generic counterpart (Qwen2.5-VL) and a stronger edition post-trained without using the generated traces. This work paves the way for training powerful fundus-reading MLLMs with publicly available data.
△ Less
Submitted 9 April, 2026;
originally announced April 2026.
-
AgentVLN: Towards Agentic Vision-and-Language Navigation
Authors:
Zihao Xin,
Wentong Li,
Yixuan Jiang,
Ziyuan Huang,
Bin Wang,
Piji Li,
Jianke Zhu,
Jie Qin,
Shengjun Huang
Abstract:
Vision-and-Language Navigation (VLN) requires an embodied agent to ground complex natural-language instructions into long-horizon navigation in unseen environments. While Vision-Language Models (VLMs) offer strong 2D semantic understanding, current VLN systems remain constrained by limited spatial perception, 2D-3D representation mismatch, and monocular scale ambiguity. In this paper, we propose A…
▽ More
Vision-and-Language Navigation (VLN) requires an embodied agent to ground complex natural-language instructions into long-horizon navigation in unseen environments. While Vision-Language Models (VLMs) offer strong 2D semantic understanding, current VLN systems remain constrained by limited spatial perception, 2D-3D representation mismatch, and monocular scale ambiguity. In this paper, we propose AgentVLN, a novel and efficient embodied navigation framework that can be deployed on edge computing platforms. We formulate VLN as a Partially Observable Semi-Markov Decision Process (POSMDP) and introduce a VLM-as-Brain paradigm that decouples high-level semantic reasoning from perception and planning via a plug-and-play skill library. To resolve multi-level representation inconsistency, we design a cross-space representation mapping that projects perception-layer 3D topological waypoints into the image plane, yielding pixel-aligned visual prompts for the VLM. Building on this bridge, we integrate a context-aware self-correction and active exploration strategy to recover from occlusions and suppress error accumulation over long trajectories. To further address the spatial ambiguity of instructions in unstructured environments, we propose a Query-Driven Perceptual Chain-of-Thought (QD-PCoT) scheme, enabling the agent with the metacognitive ability to actively seek geometric depth information. Finally, we construct AgentVLN-Instruct, a large-scale instruction-tuning dataset with dynamic stage routing conditioned on target visibility. Extensive experiments show that AgentVLN consistently outperforms prior state-of-the-art methods (SOTA) on long-horizon VLN benchmarks, offering a practical paradigm for lightweight deployment of next-generation embodied navigation models. Code: https://github.com/Allenxinn/AgentVLN.
△ Less
Submitted 18 March, 2026;
originally announced March 2026.
-
DecoVLN: Decoupling Observation, Reasoning, and Correction for Vision-and-Language Navigation
Authors:
Zihao Xin,
Wentong Li,
Yixuan Jiang,
Bin Wang,
Runmin Cong,
Jie Qin,
Shengjun Huang
Abstract:
Vision-and-Language Navigation (VLN) requires agents to follow long-horizon instructions and navigate complex 3D environments. However, existing approaches face two major challenges: constructing an effective long-term memory bank and overcoming the compounding errors problem. To address these issues, we propose DecoVLN, an effective framework designed for robust streaming perception and closed-lo…
▽ More
Vision-and-Language Navigation (VLN) requires agents to follow long-horizon instructions and navigate complex 3D environments. However, existing approaches face two major challenges: constructing an effective long-term memory bank and overcoming the compounding errors problem. To address these issues, we propose DecoVLN, an effective framework designed for robust streaming perception and closed-loop control in long-horizon navigation. First, we formulate long-term memory construction as an optimization problem and introduce adaptive refinement mechanism that selects frames from a historical candidate pool by iteratively optimizing a unified scoring function. This function jointly balances three key criteria: semantic relevance to the instruction, visual diversity from the selected memory, and temporal coverage of the historical trajectory. Second, to alleviate compounding errors, we introduce a state-action pair-level corrective finetuning strategy. By leveraging geodesic distance between states to precisely quantify deviation from the expert trajectory, the agent collects high-quality state-action pairs in the trusted region while filtering out the polluted data with low relevance. This improves both the efficiency and stability of error correction. Extensive experiments demonstrate the effectiveness of DecoVLN, and we have deployed it in real-world environments.
△ Less
Submitted 26 March, 2026; v1 submitted 13 March, 2026;
originally announced March 2026.
-
Automated Tensor-Relational Decomposition for Large-Scale Sparse Tensor Computation
Authors:
Yuxin Tang,
Zhiyuan Xin,
Zhimin Ding,
Xinyu Yao,
Daniel Bourgeois,
Tirthak Patel,
Chris Jermaine
Abstract:
A \emph{tensor-relational} computation is a relational computation where individual tuples carry vectors, matrices, or higher-dimensional arrays. An advantage of tensor-relational computation is that the overall computation can be executed on top of a relational system, inheriting the system's ability to automatically handle very large inputs with high levels of sparsity while high-performance ker…
▽ More
A \emph{tensor-relational} computation is a relational computation where individual tuples carry vectors, matrices, or higher-dimensional arrays. An advantage of tensor-relational computation is that the overall computation can be executed on top of a relational system, inheriting the system's ability to automatically handle very large inputs with high levels of sparsity while high-performance kernels (such as optimized matrix-matrix multiplication codes) can be used to perform most of the underlying mathematical operations. In this paper, we introduce upper-case-lower-case \texttt{EinSum}, which is a tensor-relational version of the classical Einstein Summation Notation. We study how to automatically rewrite a computation in Einstein Notation into upper-case-lower-case \texttt{EinSum} so that computationally intensive components are executed using efficient numerical kernels, while sparsity is managed relationally.
△ Less
Submitted 9 March, 2026;
originally announced March 2026.
-
SAVE: Speech-Aware Video Representation Learning for Video-Text Retrieval
Authors:
Ruixiang Zhao,
Zhihao Xu,
Bangxiang Lan,
Zijie Xin,
Jingyu Liu,
Xirong Li
Abstract:
For video-text retrieval, the use of CLIP has been a de facto choice. Since CLIP provides only image and text encoders, this consensus has led to a biased paradigm that entirely ignores the sound track of videos. While several attempts have been made to reintroduce audio -- typically by incorporating an audio encoder and fusing its output with visual features -- these methods face two challenges:…
▽ More
For video-text retrieval, the use of CLIP has been a de facto choice. Since CLIP provides only image and text encoders, this consensus has led to a biased paradigm that entirely ignores the sound track of videos. While several attempts have been made to reintroduce audio -- typically by incorporating an audio encoder and fusing its output with visual features -- these methods face two challenges: ineffective representation of speech content and suboptimal vision-audio fusion. To address these issues jointly, we propose SAVE, a Speech Aware Video rEpresentation learning method. SAVE improves upon AVIGATE, a SOTA audiovisual method, with a dedicated speech branch for more effective speech embedding. Furthermore, we introduce soft-ALBEF for early vision-audio alignment that facilitates fusion. Extensive experiments on five benchmarks show that SAVE compares favorably against the SOTA, outperforming AVIGATE by +4.1% on MSRVTT-9k, +1.9% on MSRVTT-7k, +2.5% on VATEX, +9.8% on Charades, and +2.1% on LSMDC, in light of the SumR metric.
△ Less
Submitted 10 March, 2026; v1 submitted 9 March, 2026;
originally announced March 2026.
-
Design and characterization of W-band and D-band calibration sources for the AliCPT-1 experiment
Authors:
Xu-Fang Li,
Cong-Zhan Liu,
Ai-Mei Zhang,
Zheng-Wei Li,
Xue-Feng Lu,
Zhong-Xue Xin,
Guo-Feng Wang,
Yong-Ping Li,
Yong-Jie Zhang,
Shi-Bo Shu,
Yi-Fei Zhang,
Ya-Qiong Li,
Zhi Chang,
Dai-Kang Yan
Abstract:
Ali Cosmic Microwave Background Polarization Telescope (AliCPT-1) is the first Chinese cosmic microwave background experiment aiming to make sensitive polarization maps of the potential B-mode signal from inflationary gravitational waves. The telescope was deployed on the Tibet Ali site at 5250 m above sea level in early 2025. Before and after each observation season, the instrument performance mu…
▽ More
Ali Cosmic Microwave Background Polarization Telescope (AliCPT-1) is the first Chinese cosmic microwave background experiment aiming to make sensitive polarization maps of the potential B-mode signal from inflationary gravitational waves. The telescope was deployed on the Tibet Ali site at 5250 m above sea level in early 2025. Before and after each observation season, the instrument performance must be carefully calibrated, including the far field beam performance, far sidelobe, spectral response, polarization angle, and cross-polar beam response. To characterize these optical performances, several calibrators have been developed. We developed a W-band source and a D-band source for the AliCPT-1 telescope's beam characterizations. We present the design and performance of the two calibration sources.
△ Less
Submitted 12 February, 2026;
originally announced February 2026.
-
PINN-Based Kolmogorov-Arnold Networks with RAR-D Adaptive Sampling for Solving Elliptic Interface Problems
Authors:
Zijuan Xin,
Chenyao Wang,
Feng Shi,
Yizhong Sun
Abstract:
Physics-Informed Neural Networks (PINNs) have become a popular and powerful framework for solving partial differential equations (PDEs), leveraging neural networks to approximate solutions while embedding PDE constraints, boundary conditions, and interface jump conditions directly into the loss function. However, most existing PINN approaches are based on multilayer perceptrons (MLPs), which may r…
▽ More
Physics-Informed Neural Networks (PINNs) have become a popular and powerful framework for solving partial differential equations (PDEs), leveraging neural networks to approximate solutions while embedding PDE constraints, boundary conditions, and interface jump conditions directly into the loss function. However, most existing PINN approaches are based on multilayer perceptrons (MLPs), which may require large network sizes and extensive training to achieve high accuracy, especially for complex interface problems. In this work, we propose a novel PINN architecture based on Kolmogorov-Arnold Networks (KANs), which offer greater flexibility in choosing activation functions and can represent functions with fewer parameters. Specifically, we introduce a dual KANs structure that couples two KANs across subdomains and explicitly enforces interface conditions. To further boost training efficiency and convergence, we integrate the RAR-D adaptive sampling strategy to dynamically refine training points. Numerical experiments on the elliptic interface problems yield more uniform error distributions across the computational domain, which demonstrates that our PINN-based KANs achieve superior accuracy with significantly smaller network sizes and faster convergence compared to standard PINNs.
△ Less
Submitted 2 February, 2026;
originally announced February 2026.
-
SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing Images
Authors:
Zepeng Xin,
Kaiyu Li,
Luodi Chen,
Wanchen Li,
Yuchen Xiao,
Hui Qiao,
Weizhan Zhang,
Deyu Meng,
Xiangyong Cao
Abstract:
Effectively grounding complex language to pixels in remote sensing (RS) images is a critical challenge for applications like disaster response and environmental monitoring. Current models can parse simple, single-target commands but fail when presented with complex geospatial scenarios, e.g., segmenting objects at various granularities, executing multi-target instructions, and interpreting implici…
▽ More
Effectively grounding complex language to pixels in remote sensing (RS) images is a critical challenge for applications like disaster response and environmental monitoring. Current models can parse simple, single-target commands but fail when presented with complex geospatial scenarios, e.g., segmenting objects at various granularities, executing multi-target instructions, and interpreting implicit user intent. To drive progress against these failures, we present LaSeRS, the first large-scale dataset built for comprehensive training and evaluation across four critical dimensions of language-guided segmentation: hierarchical granularity, target multiplicity, reasoning requirements, and linguistic variability. By capturing these dimensions, LaSeRS moves beyond simple commands, providing a benchmark for complex geospatial reasoning. This addresses a critical gap: existing datasets oversimplify, leading to sensitivity-prone real-world models. We also propose SegEarth-R2, an MLLM architecture designed for comprehensive language-guided segmentation in RS, which directly confronts these challenges. The model's effectiveness stems from two key improvements: (1) a spatial attention supervision mechanism specifically handles the localization of small objects and their components, and (2) a flexible and efficient segmentation query mechanism that handles both single-target and multi-target scenarios. Experimental results demonstrate that our SegEarth-R2 achieves outstanding performance on LaSeRS and other benchmarks, establishing a powerful baseline for the next generation of geospatial segmentation. All data and code will be released at https://github.com/earth-insights/SegEarth-R2.
△ Less
Submitted 22 December, 2025;
originally announced December 2025.
-
Model of incompressible turbulent flows via a kinetic theory
Authors:
Ziyang Xin,
Zhaoli Guo,
Hudong Chen
Abstract:
Kinetic theory offers a promising alternative to conventional turbulence modelling by providing a mesoscopic perspective that naturally captures non-equilibrium physics such as non-Newtonian effects. In this work, we present an extension and theoretical analysis of the recent kinetic model for incompressible turbulent flows developed by Chen et al. (Atmos. 14(7), 1109, 2023), constructed for unbou…
▽ More
Kinetic theory offers a promising alternative to conventional turbulence modelling by providing a mesoscopic perspective that naturally captures non-equilibrium physics such as non-Newtonian effects. In this work, we present an extension and theoretical analysis of the recent kinetic model for incompressible turbulent flows developed by Chen et al. (Atmos. 14(7), 1109, 2023), constructed for unbounded flows. The first extension is to reselect a relaxation time such that the turbulent transport coefficients are obtained more consistently and better align with well-established turbulence theory. The Chapman-Enskog (CE) analysis of the kinetic model reproduces the traditional linear eddy viscosity and gradient diffusion models for Reynolds stress and turbulent kinetic energy flux at the first order, and yields nonlinear eddy viscosity and closure models at the second order. Particularly, a previously unreported CE solution for turbulent kinetic energy flux is obtained. The second extension is to enable the model for wall-bounded turbulent flows with preserved near-wall asymptotic behaviours. This involves developing a low-Reynolds number kinetic model incorporating wall damping effects and viscous diffusion, with boundary conditions enabling both viscous sublayer resolution and wall function application. Comprehensive validation against experimental and DNS data for turbulent plane Couette flow demonstrates excellent agreement in predicting mean velocity profiles, skin friction coefficients, and Reynolds stress distributions. It reveals that an averaged turbulent flow behaves similarly to a rarefied gas flow at a finite Knudsen number, capturing non-Newtonian effects inaccessible to linear eddy viscosity models. This kinetic model provides a physics-based foundation for turbulence modelling with reduced empirical dependence.
△ Less
Submitted 12 March, 2026; v1 submitted 6 December, 2025;
originally announced December 2025.
-
TS-PEFT: Unveiling Token-Level Redundancy in Parameter-Efficient Fine-Tuning
Authors:
Dabiao Ma,
Ziming Dai,
Zhimin Xin,
Shu Wang,
Jian Yang,
Haojun Fei
Abstract:
Current Parameter-Efficient Fine-Tuning (PEFT) methods typically operate under an implicit assumption: Once a target module is selected, every token passing through it contributes equally to the downstream task and requires a parameter update. In this paper, we challenge this convention by revealing a pervasive token-level redundancy in the fine-tuning of large models (LMs). We propose TS-PEFT, a…
▽ More
Current Parameter-Efficient Fine-Tuning (PEFT) methods typically operate under an implicit assumption: Once a target module is selected, every token passing through it contributes equally to the downstream task and requires a parameter update. In this paper, we challenge this convention by revealing a pervasive token-level redundancy in the fine-tuning of large models (LMs). We propose TS-PEFT, a theoretical framework utilizing proximal optimization that acts as a dynamic probe to identify token-level redundancy during the fine-tuning process. Extensive experiments demonstrate that indiscriminately updating all tokens is not only computationally superfluous but often introduces optimization noise. Surprisingly, by discarding 30%-70% of token updates, TS-PEFT consistently matches or exceeds the performance of dense baselines such as LoRA, DoRA. Our in-depth analysis shows that the learned token-level sparsity is a superior indicator of module importance compared to traditional weight criteria, providing a novel data-driven perspective on the intrinsic adaptation mechanism of LMs.
△ Less
Submitted 29 January, 2026; v1 submitted 20 November, 2025;
originally announced November 2025.
-
First measurement of reactor neutrino oscillations at JUNO
Authors:
Angel Abusleme,
Thomas Adam,
Kai Adamowicz,
David Adey,
Shakeel Ahmad,
Rizwan Ahmed,
Timo Ahola,
Sebastiano Aiello,
Fengpeng An,
Guangpeng An,
Costas Andreopoulos,
Giuseppe Andronico,
João Pedro Athayde Marcondes de André,
Nikolay Anfimov,
Vito Antonelli,
Tatiana Antoshkina,
Burin Asavapibhop,
Didier Auguste,
Margherita Buizza Avanzini,
Andrej Babic,
Jingzhi Bai,
Weidong Bai,
Nikita Balashov,
Roberto Barbera,
Andrea Barresi
, et al. (1114 additional authors not shown)
Abstract:
Neutrino oscillations, a quantum effect manifesting at macroscopic scales, are governed by lepton flavor mixing angles and neutrino mass-squared differences that are fundamental parameters of particle physics, representing phenomena beyond the Standard Model. Precision measurements of these parameters are essential for testing the completeness of the three-flavor framework, determining the mass or…
▽ More
Neutrino oscillations, a quantum effect manifesting at macroscopic scales, are governed by lepton flavor mixing angles and neutrino mass-squared differences that are fundamental parameters of particle physics, representing phenomena beyond the Standard Model. Precision measurements of these parameters are essential for testing the completeness of the three-flavor framework, determining the mass ordering of neutrinos, and probing possible new physics. The Jiangmen Underground Neutrino Observatory (JUNO) is a 20 kton liquid-scintillator detector located 52.5 km from multiple reactor cores, designed to resolve the interference pattern of reactor neutrinos with sub-percent precision. Here we report, using the first 59.1 days of data collected since detector completion in August 2025, the first simultaneous high-precision determination of two neutrino oscillation parameters, $\sin^2 θ_{12} = 0.3092\,\pm\,0.0087$ and $Δm^2_{21} = (7.50\,\pm\,0.12)\times10^{-5}\;{\rm eV}^2$ for the normal mass ordering scenario, improving the precision by a factor of 1.6 relative to the combination of all previous measurements. These results advance the basic understanding of neutrinos, validate the detector's design, and confirm JUNO's readiness for its primary goal of resolving the neutrino mass ordering with a larger dataset. The rapid achievement with a short exposure highlights JUNO's potential to push the frontiers of precision neutrino physics and paves the way for its broad scientific program.
△ Less
Submitted 18 November, 2025;
originally announced November 2025.
-
Initial performance results of the JUNO detector
Authors:
Angel Abusleme,
Thomas Adam,
Kai Adamowicz,
David Adey,
Shakeel Ahmad,
Rizwan Ahmed,
Timo Ahola,
Sebastiano Aiello,
Fengpeng An,
Guangpeng An,
Costas Andreopoulos,
Giuseppe Andronico,
João Pedro Athayde Marcondes de André,
Nikolay Anfimov,
Vito Antonelli,
Tatiana Antoshkina,
Burin Asavapibhop,
Didier Auguste,
Margherita Buizza Avanzini,
Andrej Babic,
Jingzhi Bai,
Weidong Bai,
Nikita Balashov,
Roberto Barbera,
Andrea Barresi
, et al. (1114 additional authors not shown)
Abstract:
The Jiangmen Underground Neutrino Observatory (JUNO) started physics data taking on 26 August 2025. JUNO consists of a 20-kton liquid scintillator central detector, surrounded by a 35 kton water pool serving as a Cherenkov veto, and almost 1000 m$^2$ of plastic scintillator veto on top. The detector is located in a shallow underground laboratory with an overburden of 1800 m.w.e. This paper present…
▽ More
The Jiangmen Underground Neutrino Observatory (JUNO) started physics data taking on 26 August 2025. JUNO consists of a 20-kton liquid scintillator central detector, surrounded by a 35 kton water pool serving as a Cherenkov veto, and almost 1000 m$^2$ of plastic scintillator veto on top. The detector is located in a shallow underground laboratory with an overburden of 1800 m.w.e. This paper presents the performance results of the detector, extensively studied during the commissioning of the water phase, the subsequent liquid scintillator filling phase, and the first physics runs. The liquid scintillator achieved an attenuation length of 20.6 m at 430 nm, while the high coverage PMT system and scintillator together yielded about 1785 photoelectrons per MeV of energy deposit at the detector centre, measured using the 2.223 MeV $γ$ from neutron captures on hydrogen with an Am-C calibration source. The reconstructed energy resolution is 3.4% for two 0.511 MeV $γ$ at the detector centre and 2.9% for the 0.93 MeV quenched Po-214 alpha decays from natural radioactive sources. The energy nonlinearity is calibrated to better than 1%. Intrinsic contaminations of U-238 and Th-232 in the liquid scintillator are below 10$^{-16}$ g/g, assuming secular equilibrium. The water Cherenkov detector achieves a muon detection efficiency better than 99.9% for muons traversing the liquid scintillator volume. During the initial science runs, the data acquisition duty cycle exceeded 97.8%, demonstrating the excellent stability and readiness of JUNO for high-precision neutrino physics.
△ Less
Submitted 18 November, 2025;
originally announced November 2025.
-
Prospects for geoneutrino detection with JUNO
Authors:
Thomas Adam,
Shakeel Ahmad,
Rizwan Ahmed,
Fengpeng An,
João Pedro Athayde Marcondes de André,
Costas Andreopoulos,
Giuseppe Andronico,
Nikolay Anfimov,
Vito Antonelli,
Tatiana Antoshkina,
Didier Auguste,
Marcel Büchner,
Weidong Bai,
Nikita Balashov,
Andrea Barresi,
Davide Basilico,
Eric Baussan,
Marco Beretta,
Antonio Bergnoli,
Nikita Bessonov,
Daniel Bick,
Lukas Bieger,
Svetlana Biktemerova,
Thilo Birkenfeld,
Simon Blyth
, et al. (605 additional authors not shown)
Abstract:
Geoneutrinos, which are antineutrinos emitted during the decay of long-lived radioactive elements inside Earth, serve as a unique tool for studying the composition and heat budget of our planet. The Jiangmen Underground Neutrino Observatory (JUNO) experiment in China, which has recently completed construction, is expected to collect a sample comparable in size to the entire existing world geoneutr…
▽ More
Geoneutrinos, which are antineutrinos emitted during the decay of long-lived radioactive elements inside Earth, serve as a unique tool for studying the composition and heat budget of our planet. The Jiangmen Underground Neutrino Observatory (JUNO) experiment in China, which has recently completed construction, is expected to collect a sample comparable in size to the entire existing world geoneutrino dataset in less than a year. This paper presents an updated estimation of sensitivity to geoneutrinos of JUNO using the best knowledge available to date about the experimental site, the surrounding nuclear reactors, the detector response uncertainties, and the constraints expected from the TAO satellite detector. To facilitate comparison with present and future geological models, our results cover a wide range of predicted signal strengths. Despite the significant background from reactor antineutrinos, the experiment will measure the total geoneutrino flux with a precision comparable to that of existing experiments within its first few years, ultimately achieving a world-leading precision of about 8% over ten years. The large statistics of JUNO will also allow separation of the Uranium-238 and Thorium-232 contributions with unprecedented precision, providing crucial constraints on models of formation and composition of Earth. Observation of the mantle signal above the lithospheric flux will be possible but challenging. For models with the highest predicted mantle concentrations of heat-producing elements, a 3-sigma detection over six years requires knowledge of the lithospheric flux to within 15%. Together with complementary measurements from other locations, the geoneutrino results of JUNO will offer cutting-edge, high-precision insights into the interior of Earth, of fundamental importance to both the geoscience and neutrino physics communities.
△ Less
Submitted 10 November, 2025;
originally announced November 2025.
-
Mock Observations for the CSST Mission: Main Surveys--the Stray Light
Authors:
Xian Jing-Tian,
Lin Lin,
Fang Yue-Dong,
Zhang Xin,
Xu You-Hua,
Meng Xian-Min,
Tian Hao,
Zhang Tian-Yi,
Ban Zhang,
Li Guo-Liang,
Xu Shu-Yan,
Wang Wei
Abstract:
Stray light significantly influences the detection capabilities of astronomical telescopes. The actual stray-light level during observations depends not only on the telescope's inherent stray-light suppression capability but also on its operational orbit conditions. Accurate estimation of stray-light levels is crucial for assessing image quality and performing realistic scientific simulations. To…
▽ More
Stray light significantly influences the detection capabilities of astronomical telescopes. The actual stray-light level during observations depends not only on the telescope's inherent stray-light suppression capability but also on its operational orbit conditions. Accurate estimation of stray-light levels is crucial for assessing image quality and performing realistic scientific simulations. To rapidly estimate stray-light levels under realistic, complex operational conditions, we developed an analytical model tailored to the China Space Station Telescope (CSST). Our model simulates stray-light backgrounds generated by off-field sources such as moonlight, starlight, and earthshine, incorporating the effects of zodiacal light, as well as scattering and ghost images induced by bright in-field stars. The proposed method allows quick and accurate evaluation of stray-light conditions, facilitating both image simulation and observational scheduling.
△ Less
Submitted 10 November, 2025;
originally announced November 2025.
-
Design, waterproofing, and mass production of the 3-inch PMT frontend system of JUNO
Authors:
Jilei Xu,
Miao He,
Cédric Cerna,
Yongbo Huang,
Thomas Adam,
Shakeel Ahmad,
Rizwan Ahmed,
Fengpeng An,
Costas Andreopoulos,
Giuseppe Andronico,
João Pedro Athayde Marcondes de André,
Nikolay Anfimov,
Vito Antonelli,
Tatiana Antoshkina,
Didier Auguste,
Weidong Bai,
Nikita Balashov,
Andrea Barresi,
Davide Basilico,
Eric Baussan,
Marco Beretta,
Antonio Bergnoli,
Nikita Bessonov,
Daniel Bick,
Lukas Bieger
, et al. (609 additional authors not shown)
Abstract:
Over 25,600 3-inch photomultiplier tubes (PMTs) have been instrumented for the central detector of the Jiangmen Underground Neutrino Observatory. Each PMT is equipped with a high-voltage divider and a frontend cable with waterproof sealing. Groups of sixteen PMTs are connected to the underwater frontend readout electronics via specialized multi-channel waterproof connectors. This paper outlines th…
▽ More
Over 25,600 3-inch photomultiplier tubes (PMTs) have been instrumented for the central detector of the Jiangmen Underground Neutrino Observatory. Each PMT is equipped with a high-voltage divider and a frontend cable with waterproof sealing. Groups of sixteen PMTs are connected to the underwater frontend readout electronics via specialized multi-channel waterproof connectors. This paper outlines the design and mass production processes for the high-voltage divider, the cable and connector, as well as the waterproof potting of the PMT bases. The results of the acceptance tests of all the integrated PMTs are also presented.
△ Less
Submitted 22 January, 2026; v1 submitted 7 October, 2025;
originally announced October 2025.
-
Efficient E(3)-equivariant framework for universal charge density prediction
Authors:
Xiwen Li,
Zaizhou Xin,
Hongyu Yu,
Yang Zhong,
Xingao Gong,
Hongjun Xiang
Abstract:
Electronic structure is ubiquitously obtained via density functional theory (DFT), where the charge density plays a central role. This work presents EdenGNN (Equivariant Density Graph Neural Network), a machine learning (ML) charge density model for electronic structure. Current universal ML charge density models are hampered by prohibitive computational costs. Furthermore, despite being trained o…
▽ More
Electronic structure is ubiquitously obtained via density functional theory (DFT), where the charge density plays a central role. This work presents EdenGNN (Equivariant Density Graph Neural Network), a machine learning (ML) charge density model for electronic structure. Current universal ML charge density models are hampered by prohibitive computational costs. Furthermore, despite being trained on projector augmented-wave (PAW) based DFT datasets, they predict only the pseudo charge density, which is insufficient to reconstruct the electronic structure. In contrast, EdenGNN overcomes these limitations. It additionally predicts the augmentation occupancies, enabling electronic structure calculations with PAW accuracy. Critically, by employing a basis-expansion formulation with fully trainable radial basis functions and a $Δ$-learning strategy to capture charge transfer, it is over an order of magnitude faster. Trained on the Materials Project database, our universal model, EdenGNN-Uni, accurately predicts the band structures for the majority of materials across a vast chemical space. These findings establish the ML charge density model as a scalable \textit{ab initio} method for large-scale electronic structure calculations and high-throughput screening.
△ Less
Submitted 13 March, 2026; v1 submitted 1 October, 2025;
originally announced October 2025.
-
Integrated Silicon Photonic Multichannel Optical Hybrid for Broadband Parallel Coherent Reception
Authors:
Tong Lin,
Yan Fan,
Jiao Zhang,
Wenqi Yu,
Zhengyu Guo,
Liu Li,
Zhigang Xin,
Mingzheng Lei,
Ziyang Xiong,
Haoran Wang,
Hao Deng,
Min Zhu,
Shihua Chen,
Junpeng Lu,
Zhenhua Ni
Abstract:
We design and demonstrate a monolithically integrated silicon photonic multichannel optical hybrid for versatile broadband coherent reception, addressing the critical limitations of current wavelength multiplexed systems in scalability and power efficiency. The device combines a phase-compensated 90-degree optical hybrid with four robust three-stage Mach-Zehnder interferometer lattice filters, ena…
▽ More
We design and demonstrate a monolithically integrated silicon photonic multichannel optical hybrid for versatile broadband coherent reception, addressing the critical limitations of current wavelength multiplexed systems in scalability and power efficiency. The device combines a phase-compensated 90-degree optical hybrid with four robust three-stage Mach-Zehnder interferometer lattice filters, enabling 34-port functionality (two inputs and 32 outputs) for simultaneous analog and digital signal processing. Leveraging multimode interferometer designs,the chip achieves a broadband response with sub-dB passband uniformity across eight 200 GHz-spaced wavelength channels, while maintaining phase errors below 4 degrees over a 13.5 nm (1539-1552.5 nm) bandwidth with only 2.5 mW thermal tuning power.Experimentally, we validate its parallel-processing capability through RF channelizer reception (showing an average spurious-free dynamic range of 80.8 dB*Hz2/3 and image rejection ratio of 33.26 dB) and coherent optical communication (achieving 1.024 Tb/s data rate for 32-QAM signals with bit error rates far below the 20% SD-FEC threshold). The scheme enhances system performance with fully passive wavelength multiplexing integration, supporting high-fidelity uniformity and projecting scalability to 1.468 Tb/s. This work promises advancements in high-performance optoelectronic devices for next-generation AI-driven data centers and 5G-XG networks.
△ Less
Submitted 30 September, 2025;
originally announced September 2025.
-
The Digital Landscape of God: Narrative, Visuals and Viewer Engagement of Religious Videos on YouTube
Authors:
Rongyi Chen,
Ziyan Xin,
Qing Xiao,
Ruiwei Xiao,
Jingjia Xiao,
Bingbing Zhang,
Hong Shen,
Zhicong Lu
Abstract:
The digital transformation of religious practice has reshaped how billions of people engage with spiritual content, with video-sharing platforms becoming central to contemporary religious communication. Yet HCI research lacks systematic understanding of how narrative and visual elements create meaningful spiritual experiences and foster viewer engagement. We present a mixed-methods study of religi…
▽ More
The digital transformation of religious practice has reshaped how billions of people engage with spiritual content, with video-sharing platforms becoming central to contemporary religious communication. Yet HCI research lacks systematic understanding of how narrative and visual elements create meaningful spiritual experiences and foster viewer engagement. We present a mixed-methods study of religious videos on YouTube across major religions, developing taxonomies of narrative frameworks, visual elements, and viewer interaction. Using LLM-assisted analysis, we studied relationships between content characteristics and viewer responses. Religious videos predominantly adopt lecture-style formats with authority-based persuasion strategies, using salvation narratives for guidance. All prefer bright lighting, with Buddhism favoring warm tones and prominent symbols, Judaism preferring indoor settings, and Hinduism emphasizing sacred objects. We identified differentiated patterns of emotional sharing among religious viewers while revealing significant correlations between content characteristics and engagement, particularly regarding AI-generated content. We provide evidence-based guidance for creating inclusive and engaging spiritual media.
△ Less
Submitted 10 April, 2026; v1 submitted 13 September, 2025;
originally announced September 2025.
-
Beyond flattening: a geometrically principled positional encoding for vision transformers with Weierstrass elliptic functions
Authors:
Zhihang Xin,
Xitong Hu,
Rui Wang
Abstract:
Vision Transformers have demonstrated remarkable success in computer vision tasks, yet their reliance on learnable one-dimensional positional embeddings fundamentally disrupts the inherent two-dimensional spatial structure of images through patch flattening procedures. Traditional positional encoding approaches lack geometric constraints and fail to establish monotonic correspondence between Eucli…
▽ More
Vision Transformers have demonstrated remarkable success in computer vision tasks, yet their reliance on learnable one-dimensional positional embeddings fundamentally disrupts the inherent two-dimensional spatial structure of images through patch flattening procedures. Traditional positional encoding approaches lack geometric constraints and fail to establish monotonic correspondence between Euclidean spatial distances and sequential index distances, thereby limiting the model's capacity to leverage spatial proximity priors effectively. We propose Weierstrass Elliptic Function Positional Encoding (WEF-PE), a mathematically principled approach that directly addresses two-dimensional coordinates through natural complex domain representation, where the doubly periodic properties of elliptic functions align remarkably with translational invariance patterns commonly observed in visual data. Our method exploits the non-linear geometric nature of elliptic functions to encode spatial distance relationships naturally, while the algebraic addition formula enables direct derivation of relative positional information between arbitrary patch pairs from their absolute encodings. Comprehensive experiments demonstrate that WEF-PE achieves superior performance across diverse scenarios, including 63.78\% accuracy on CIFAR-100 from-scratch training with ViT-Tiny architecture, 93.28\% on CIFAR-100 fine-tuning with ViT-Base, and consistent improvements on VTAB-1k benchmark tasks. Theoretical analysis confirms the distance-decay property through rigorous mathematical proof, while attention visualization reveals enhanced geometric inductive bias and more coherent semantic focus compared to conventional approaches.The source code implementing the methods described in this paper is publicly available on GitHub.
△ Less
Submitted 26 August, 2025;
originally announced August 2025.
-
STM3: Mixture of Multiscale Mamba for Long-Term Spatio-Temporal Time-Series Prediction
Authors:
Haolong Chen,
Liang Zhang,
Zhengyuan Xin,
Guangxu Zhu
Abstract:
Recently, spatio-temporal time-series prediction has developed rapidly, yet existing deep learning methods struggle with learning complex long-term spatio-temporal dependencies efficiently. The long-term spatio-temporal dependency learning brings two new challenges: 1) The long-term temporal sequence naturally includes multiscale information, which is hard to extract efficiently; 2) The multiscale…
▽ More
Recently, spatio-temporal time-series prediction has developed rapidly, yet existing deep learning methods struggle with learning complex long-term spatio-temporal dependencies efficiently. The long-term spatio-temporal dependency learning brings two new challenges: 1) The long-term temporal sequence naturally includes multiscale information, which is hard to extract efficiently; 2) The multiscale temporal information from different nodes is highly correlated and hard to model. To address these challenges, we propose Spatio-Temporal Mixture of Multiscale Mamba (STM3). STM3 integrates a Multiscale Mamba architecture within a novel Disentangled Mixture-of-Experts (DMoE) framework to capture diverse multiscale information efficiently, while utilizing an adaptive graph causal network to model complex spatial dependencies. To ensure robust representation learning, we introduce a stable routing strategy and a causal contrastive learning strategy, which work in tandem with hierarchical information aggregation to guarantee scale distinguishability. We theoretically prove that STM3 achieves superior routing smoothness and guarantees pattern disentanglement for each expert. Extensive experiments on 10 real-world benchmarks across domains demonstrate STM3's superior performance, achieving state-of-the-art results in long-term spatio-temporal time-series prediction. Notably, on the PEMSD8 dataset, it achieves significant improvements, surpassing the second-best model by 7.1% in MAE, 8.5% in RMSE, and 15.9% in MAPE. Code is available at https://github.com/IfReasonable/STM3_KDD26.
△ Less
Submitted 22 May, 2026; v1 submitted 17 August, 2025;
originally announced August 2025.
-
Learning Partially-Decorrelated Common Spaces for Ad-hoc Video Search
Authors:
Fan Hu,
Zijie Xin,
Xirong Li
Abstract:
Ad-hoc Video Search (AVS) involves using a textual query to search for multiple relevant videos in a large collection of unlabeled short videos. The main challenge of AVS is the visual diversity of relevant videos. A simple query such as "Find shots of a man and a woman dancing together indoors" can span a multitude of environments, from brightly lit halls and shadowy bars to dance scenes in black…
▽ More
Ad-hoc Video Search (AVS) involves using a textual query to search for multiple relevant videos in a large collection of unlabeled short videos. The main challenge of AVS is the visual diversity of relevant videos. A simple query such as "Find shots of a man and a woman dancing together indoors" can span a multitude of environments, from brightly lit halls and shadowy bars to dance scenes in black-and-white animations. It is therefore essential to retrieve relevant videos as comprehensively as possible. Current solutions for the AVS task primarily fuse multiple features into one or more common spaces, yet overlook the need for diverse spaces. To fully exploit the expressive capability of individual features, we propose LPD, short for Learning Partially Decorrelated common spaces. LPD incorporates two key innovations: feature-specific common space construction and the de-correlation loss. Specifically, LPD learns a separate common space for each video and text feature, and employs de-correlation loss to diversify the ordering of negative samples across different spaces. To enhance the consistency of multi-space convergence, we designed an entropy-based fair multi-space triplet ranking loss. Extensive experiments on the TRECVID AVS benchmarks (2016-2023) justify the effectiveness of LPD. Moreover, diversity visualizations of LPD's spaces highlight its ability to enhance result diversity.
△ Less
Submitted 4 August, 2025;
originally announced August 2025.
-
Few-Shot Object Detection via Spatial-Channel State Space Model
Authors:
Zhimeng Xin,
Tianxu Wu,
Yixiong Zou,
Shiming Chen,
Dingjie Fu,
Xinge You
Abstract:
Due to the limited training samples in few-shot object detection (FSOD), we observe that current methods may struggle to accurately extract effective features from each channel. Specifically, this issue manifests in two aspects: i) channels with high weights may not necessarily be effective, and ii) channels with low weights may still hold significant value. To handle this problem, we consider uti…
▽ More
Due to the limited training samples in few-shot object detection (FSOD), we observe that current methods may struggle to accurately extract effective features from each channel. Specifically, this issue manifests in two aspects: i) channels with high weights may not necessarily be effective, and ii) channels with low weights may still hold significant value. To handle this problem, we consider utilizing the inter-channel correlation to facilitate the novel model's adaptation process to novel conditions, ensuring the model can correctly highlight effective channels and rectify those incorrect ones. Since the channel sequence is also 1-dimensional, its similarity with the temporal sequence inspires us to take Mamba for modeling the correlation in the channel sequence. Based on this concept, we propose a Spatial-Channel State Space Modeling (SCSM) module for spatial-channel state modeling, which highlights the effective patterns and rectifies those ineffective ones in feature channels. In SCSM, we design the Spatial Feature Modeling (SFM) module to balance the learning of spatial relationships and channel relationships, and then introduce the Channel State Modeling (CSM) module based on Mamba to learn correlation in channels. Extensive experiments on the VOC and COCO datasets show that the SCSM module enables the novel detector to improve the quality of focused feature representation in channels and achieve state-of-the-art performance.
△ Less
Submitted 21 July, 2025;
originally announced July 2025.
-
PDFMathTranslate: Scientific Document Translation Preserving Layouts
Authors:
Rongxin Ouyang,
Chang Chu,
Zhikuang Xin,
Xiangyao Ma
Abstract:
Language barriers in scientific documents hinder the diffusion and development of science and technologies. However, prior efforts in translating such documents largely overlooked the information in layouts. To bridge the gap, we introduce PDFMathTranslate, the world's first open-source software for translating scientific documents while preserving layouts. Leveraging the most recent advances in l…
▽ More
Language barriers in scientific documents hinder the diffusion and development of science and technologies. However, prior efforts in translating such documents largely overlooked the information in layouts. To bridge the gap, we introduce PDFMathTranslate, the world's first open-source software for translating scientific documents while preserving layouts. Leveraging the most recent advances in large language models and precise layout detection, we contribute to the community with key improvements in precision, flexibility, and efficiency. The work has been open-sourced at https://github.com/byaidu/pdfmathtranslate with more than 222k downloads.
△ Less
Submitted 22 September, 2025; v1 submitted 2 July, 2025;
originally announced July 2025.
-
MNN-AECS: Energy Optimization for LLM Decoding on Mobile Devices via Adaptive Core Selection
Authors:
Zhengxiang Huang,
Chaoyue Niu,
Zhaode Wang,
Jiarui Xue,
Hanming Zhang,
Yugang Wang,
Zewei Xin,
Xiaotang Jiang,
Chengfei Lv,
Fan Wu,
Guihai Chen
Abstract:
As the demand for on-device Large Language Model (LLM) inference grows, energy efficiency has become a major concern, especially for battery-limited mobile devices. Our analysis shows that the memory-bound LLM decode phase dominates energy use, and yet most existing works focus on accelerating the prefill phase, neglecting energy concerns. We introduce Adaptive Energy-Centric Core Selection (AECS)…
▽ More
As the demand for on-device Large Language Model (LLM) inference grows, energy efficiency has become a major concern, especially for battery-limited mobile devices. Our analysis shows that the memory-bound LLM decode phase dominates energy use, and yet most existing works focus on accelerating the prefill phase, neglecting energy concerns. We introduce Adaptive Energy-Centric Core Selection (AECS) and integrate it into MNN to create the energy-efficient version, MNN-AECS, the first engine-level system solution without requiring root access or OS modifications for energy-efficient LLM decoding. MNN-AECS is designed to reduce LLM decoding energy while keeping decode speed within an acceptable slowdown threshold by dynamically selecting low-power CPU cores. MNN-AECS is evaluated across 5 Android and 2 iOS devices on 5 popular LLMs of various sizes. Compared to original MNN, MNN-AECS cuts down energy use by 23% without slowdown averaged over all 7 devices and 4 datasets. Against other engines, including llama.cpp, executorch, mllm, and MediaPipe, MNN-AECS delivers 39% to 78% energy saving and 12% to 363% speedup on average.
△ Less
Submitted 24 June, 2025;
originally announced June 2025.
-
A Novel ViDAR Device With Visual Inertial Encoder Odometry and Reinforcement Learning-Based Active SLAM Method
Authors:
Zhanhua Xin,
Zhihao Wang,
Shenghao Zhang,
Wanchao Chi,
Yan Meng,
Shihan Kong,
Yan Xiong,
Chong Zhang,
Yuzhen Liu,
Junzhi Yu
Abstract:
In the field of multi-sensor fusion for simultaneous localization and mapping (SLAM), monocular cameras and IMUs are widely used to build simple and effective visual-inertial systems. However, limited research has explored the integration of motor-encoder devices to enhance SLAM performance. By incorporating such devices, it is possible to significantly improve active capability and field of view…
▽ More
In the field of multi-sensor fusion for simultaneous localization and mapping (SLAM), monocular cameras and IMUs are widely used to build simple and effective visual-inertial systems. However, limited research has explored the integration of motor-encoder devices to enhance SLAM performance. By incorporating such devices, it is possible to significantly improve active capability and field of view (FOV) with minimal additional cost and structural complexity. This paper proposes a novel visual-inertial-encoder tightly coupled odometry (VIEO) based on a ViDAR (Video Detection and Ranging) device. A ViDAR calibration method is introduced to ensure accurate initialization for VIEO. In addition, a platform motion decoupled active SLAM method based on deep reinforcement learning (DRL) is proposed. Experimental data demonstrate that the proposed ViDAR and the VIEO algorithm significantly increase cross-frame co-visibility relationships compared to its corresponding visual-inertial odometry (VIO) algorithm, improving state estimation accuracy. Additionally, the DRL-based active SLAM algorithm, with the ability to decouple from platform motion, can increase the diversity weight of the feature points and further enhance the VIEO algorithm's performance. The proposed methodology sheds fresh insights into both the updated platform design and decoupled approach of active SLAM systems in complex environments.
△ Less
Submitted 16 June, 2025;
originally announced June 2025.
-
Table-r1: Self-supervised and Reinforcement Learning for Program-based Table Reasoning in Small Language Models
Authors:
Rihui Jin,
Zheyu Xin,
Xing Xie,
Zuoyi Li,
Guilin Qi,
Yongrui Chen,
Xinbang Dai,
Tongtong Wu,
Gholamreza Haffari
Abstract:
Table reasoning (TR) requires structured reasoning over semi-structured tabular data and remains challenging, particularly for small language models (SLMs, e.g., LLaMA-8B) due to their limited capacity compared to large LMs (LLMs, e.g., GPT-4o). To narrow this gap, we explore program-based TR (P-TR), which circumvents key limitations of text-based TR (T-TR), notably in numerical reasoning, by gene…
▽ More
Table reasoning (TR) requires structured reasoning over semi-structured tabular data and remains challenging, particularly for small language models (SLMs, e.g., LLaMA-8B) due to their limited capacity compared to large LMs (LLMs, e.g., GPT-4o). To narrow this gap, we explore program-based TR (P-TR), which circumvents key limitations of text-based TR (T-TR), notably in numerical reasoning, by generating executable programs. However, applying P-TR to SLMs introduces two challenges: (i) vulnerability to heterogeneity in table layouts, and (ii) inconsistency in reasoning due to limited code generation capability. We propose Table-r1, a two-stage P-TR method designed for SLMs. Stage 1 introduces an innovative self-supervised learning task, Layout Transformation Inference, to improve tabular layout generalization from a programmatic view. Stage 2 adopts a mix-paradigm variant of Group Relative Policy Optimization, enhancing P-TR consistency while allowing dynamic fallback to T-TR when needed. Experiments on four TR benchmarks demonstrate that Table-r1 outperforms all SLM-based methods, achieving at least a 15% accuracy improvement over the base model (LLaMA-8B) across all datasets and reaching performance competitive with LLMs.
△ Less
Submitted 6 June, 2025;
originally announced June 2025.
-
Large-Scale Gaussian Splatting SLAM
Authors:
Zhe Xin,
Chenyang Wu,
Penghui Huang,
Yanyong Zhang,
Yinian Mao,
Guoquan Huang
Abstract:
The recently developed Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have shown encouraging and impressive results for visual SLAM. However, most representative methods require RGBD sensors and are only available for indoor environments. The robustness of reconstruction in large-scale outdoor scenarios remains unexplored. This paper introduces a large-scale 3DGS-based visual SLAM…
▽ More
The recently developed Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have shown encouraging and impressive results for visual SLAM. However, most representative methods require RGBD sensors and are only available for indoor environments. The robustness of reconstruction in large-scale outdoor scenarios remains unexplored. This paper introduces a large-scale 3DGS-based visual SLAM with stereo cameras, termed LSG-SLAM. The proposed LSG-SLAM employs a multi-modality strategy to estimate prior poses under large view changes. In tracking, we introduce feature-alignment warping constraints to alleviate the adverse effects of appearance similarity in rendering losses. For the scalability of large-scale scenarios, we introduce continuous Gaussian Splatting submaps to tackle unbounded scenes with limited memory. Loops are detected between GS submaps by place recognition and the relative pose between looped keyframes is optimized utilizing rendering and feature warping losses. After the global optimization of camera poses and Gaussian points, a structure refinement module enhances the reconstruction quality. With extensive evaluations on the EuRoc and KITTI datasets, LSG-SLAM achieves superior performance over existing Neural, 3DGS-based, and even traditional approaches. Project page: https://lsg-slam.github.io.
△ Less
Submitted 14 May, 2025;
originally announced May 2025.
-
Global well-posedness of the Cauchy problem for the modified Whitham equations
Authors:
Han Cui,
Yuexun Wang,
Zhouping Xin
Abstract:
This paper aims to show global existence and modified scattering for the solutions of the Cauchy problem to the modified Whitham equations for small, smooth and localized initial data. The main difficulties come from slow decay and non-homogeneity of the Fourier multiplier $(\sqrt{\tanh ξ/ξ})ξ$, which will be overcome by introducing an interaction multiplier theorem and estimating the weighted nor…
▽ More
This paper aims to show global existence and modified scattering for the solutions of the Cauchy problem to the modified Whitham equations for small, smooth and localized initial data. The main difficulties come from slow decay and non-homogeneity of the Fourier multiplier $(\sqrt{\tanh ξ/ξ})ξ$, which will be overcome by introducing an interaction multiplier theorem and estimating the weighted norms in the frequency space. When estimating the weighted norms, due to loss of derivatives, the energy estimate will be performed in the frequency space, and the absence of time resonance will be effectively utilized by extracting some good terms arising from integration by parts in time before the energy estimate.
△ Less
Submitted 14 May, 2025;
originally announced May 2025.
-
Development of 6-inch 80-170 GHz broadband silicon plated horn antenna arrays for primordial gravitational wave search
Authors:
Yuanhang He,
Shibo Shu,
Yaqiong Li,
Xuefeng Lu,
Ye Chai,
Xiang Li,
Zhi Chang,
He Gao,
Yudong Gu,
Xufang Li,
Zhengwei Li,
Zhouhui Liu,
Guofeng Wang,
Zhongxue Xin,
Daikang Yan,
Aimei Zhang,
Yifei Zhang,
Yongjie Zhang,
Wenhua Shi,
Juexian Cao,
Congzhan Liu
Abstract:
Searching for primordial gravitational wave in cosmic microwave background (CMB) polarization signal is one of the key topics in modern cosmology. Cutting-edge CMB telescopes requires thousands of pixels to maximize mapping speed. Using modular design, the telescope focal plane is simplified as several detector modules. Each module has hundreds of pixels including antenna arrays, detector arrays,…
▽ More
Searching for primordial gravitational wave in cosmic microwave background (CMB) polarization signal is one of the key topics in modern cosmology. Cutting-edge CMB telescopes requires thousands of pixels to maximize mapping speed. Using modular design, the telescope focal plane is simplified as several detector modules. Each module has hundreds of pixels including antenna arrays, detector arrays, and readout arrays. The antenna arrays, as the beam defining component, determine the overall optical response of the detector module. In this article, we present the developments of 6-inch broadband antenna arrays from 80GHz to 170GHz for the future IHEP focal plane module. The arrays are fabricated from 42 6-inch silicon wafers including 456 antennas, 7% more pixels than usual design. The overall in-band cross polarization is smaller than -20 dB and the in-band beam asymmetry is smaller than 10%, fulfilling the requirements for primordial gravitational wave search.
△ Less
Submitted 20 April, 2025;
originally announced April 2025.
-
SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model
Authors:
Kaiyu Li,
Zepeng Xin,
Li Pang,
Chao Pang,
Yupeng Deng,
Jing Yao,
Guisong Xia,
Deyu Meng,
Zhi Wang,
Xiangyong Cao
Abstract:
Remote sensing has become critical for understanding environmental dynamics, urban planning, and disaster management. However, traditional remote sensing workflows often rely on explicit segmentation or detection methods, which struggle to handle complex, implicit queries that require reasoning over spatial context, domain knowledge, and implicit user intent. Motivated by this, we introduce a new…
▽ More
Remote sensing has become critical for understanding environmental dynamics, urban planning, and disaster management. However, traditional remote sensing workflows often rely on explicit segmentation or detection methods, which struggle to handle complex, implicit queries that require reasoning over spatial context, domain knowledge, and implicit user intent. Motivated by this, we introduce a new task, \ie, geospatial pixel reasoning, which allows implicit querying and reasoning and generates the mask of the target region. To advance this task, we construct and release the first large-scale benchmark dataset called EarthReason, which comprises 5,434 manually annotated image masks with over 30,000 implicit question-answer pairs. Moreover, we propose SegEarth-R1, a simple yet effective language-guided segmentation baseline that integrates a hierarchical visual encoder, a large language model (LLM) for instruction parsing, and a tailored mask generator for spatial correlation. The design of SegEarth-R1 incorporates domain-specific adaptations, including aggressive visual token compression to handle ultra-high-resolution remote sensing images, a description projection module to fuse language and multi-scale features, and a streamlined mask prediction pipeline that directly queries description embeddings. Extensive experiments demonstrate that SegEarth-R1 achieves state-of-the-art performance on both reasoning and referring segmentation tasks, significantly outperforming traditional and LLM-based segmentation methods. Our data and code will be released at https://github.com/earth-insights/SegEarth-R1.
△ Less
Submitted 13 April, 2025;
originally announced April 2025.
-
Opportunity-Cost-Driven Reward Mechanisms for Crowd-Sourced Computing Platforms
Authors:
Shuhao Zheng,
Ziyue Xin,
Zonglun Li,
Xue Liu
Abstract:
This paper introduces a game-theoretic model tailored for reward distribution on crowd-sourced computing platforms. It explores a repeated game framework where miners, as computation providers, decide their computation power contribution in each round, guided by the platform's designed reward distribution mechanism. The reward for each miner in every round is based on the platform's randomized tas…
▽ More
This paper introduces a game-theoretic model tailored for reward distribution on crowd-sourced computing platforms. It explores a repeated game framework where miners, as computation providers, decide their computation power contribution in each round, guided by the platform's designed reward distribution mechanism. The reward for each miner in every round is based on the platform's randomized task payments and the miners' computation transcripts. Specifically, it defines Opportunity-Cost-Driven Incentive Compatibility (OCD-IC) and Dynamic OCD-IC (DOCD-IC) for scenarios where strategic miners might allocate some computation power to more profitable activities, such as Bitcoin mining. The platform must also achieve Budget Balance (BB), aiming for a non-negative total income over the long term. This paper demonstrates that traditional Pay-Per-Share (PPS) reward schemes require assumptions about task demand and miners' opportunity costs to ensure OCD-IC and BB, yet they fail to satisfy DOCD-IC. The paper then introduces Pay-Per-Share with Subsidy (PPSS), a new reward mechanism that allows the platform to provide subsidies to miners, thus eliminating the need for assumptions on opportunity cost to achieve OCD-IC, DOCD-IC, and long-term BB.
△ Less
Submitted 10 April, 2025;
originally announced April 2025.
-
Multi-Object Sketch Animation by Scene Decomposition and Motion Planning
Authors:
Jingyu Liu,
Zijie Xin,
Yuhan Fu,
Ruixiang Zhao,
Bangxiang Lan,
Xirong Li
Abstract:
Sketch animation, which brings static sketches to life by generating dynamic video sequences, has found widespread applications in GIF design, cartoon production, and daily entertainment. While current methods for sketch animation perform well in single-object sketch animation, they struggle in multi-object scenarios. By analyzing their failures, we identify two major challenges of transitioning f…
▽ More
Sketch animation, which brings static sketches to life by generating dynamic video sequences, has found widespread applications in GIF design, cartoon production, and daily entertainment. While current methods for sketch animation perform well in single-object sketch animation, they struggle in multi-object scenarios. By analyzing their failures, we identify two major challenges of transitioning from single-object to multi-object sketch animation: object-aware motion modeling and complex motion optimization. For multi-object sketch animation, we propose MoSketch based on iterative optimization through Score Distillation Sampling (SDS) and thus animating a multi-object sketch in a training-data free manner. To tackle the two challenges in a divide-and-conquer strategy, MoSketch has four novel modules, i.e., LLM-based scene decomposition, LLM-based motion planning, multi-grained motion refinement, and compositional SDS. Extensive qualitative and quantitative experiments demonstrate the superiority of our method over existing sketch animation approaches. MoSketch takes a pioneering step towards multi-object sketch animation, opening new avenues for future research and applications.
△ Less
Submitted 2 August, 2025; v1 submitted 25 March, 2025;
originally announced March 2025.