-
StreamSoccer: Event-Driven Memory for Streaming Soccer Commentary
Authors:
Chenxi Shao,
Bozhong Wang,
Jiaxin Huang,
Zhao Liu,
Sunwei Zhu,
Tianxin Hang,
Gaoqi He,
Yang Li,
Changbo Wang
Abstract:
Streaming video understanding requires models to causally update state as video arrives and organize growing history into semantic units that can evolve, persist, and be recalled under bounded computation and memory. This challenge is pronounced in live soccer commentary, where a system must describe completed events, summarize recent play, recall earlier events, or remain silent using only inform…
▽ More
Streaming video understanding requires models to causally update state as video arrives and organize growing history into semantic units that can evolve, persist, and be recalled under bounded computation and memory. This challenge is pronounced in live soccer commentary, where a system must describe completed events, summarize recent play, recall earlier events, or remain silent using only information available before each utterance. We present StreamSoccer, an event-driven system that uses event memory as its intermediate representation. A fixed-budget active memory integrates the stream; completed event states are retained locally and consolidated into retrievable historical records. A unified generator uses current, recent, and historical context to produce three commentary modes, while a rule-assisted scheduler selects a mode or silence. Unlike streaming video-language models organized around frames, visual tokens, or caches, and soccer-commentary methods based on predefined clips or output timestamps, StreamSoccer explicitly models event lifecycles. We construct a three-track streaming soccer commentary dataset and a layered evaluation protocol. At common reference anchors, StreamSoccer obtains CIDEr scores of 38.62, 23.96, and 17.39 for current-event, recent-window, and historical-memory commentary, ranking first on the current-event and historical-memory tracks and second on recent-window. Controlled ablations show that local completed events improve all tracks and that the full system performs best on all three. Across 174 raw-video runs on 58 matches, per-minute RTF p95 ranges from 0.10 to 0.22 without sustained growth with match history. These results indicate that event memory supports streaming soccer commentary across temporal scopes while controlling long-history computation.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
Parallel single-pixel imaging based on modulation region expansion and overlapping reconstruction
Authors:
Yinran Shen,
Xuri Yao,
Shijian Li,
Chao Shen,
Yuhao Wang,
Chongwu Shao,
Qing Zhao
Abstract:
Parallel single-pixel imaging (PSPI) enhances the data acquisition efficiency of single-pixel imaging, but its reconstruction quality depends on a cumbersome and noise-sensitive calibration process. To address this challenge, a PSPI strategy was introduced that leverages modulation region expansion and overlapping reconstruction. This method results in the calibration of modulation of the subregio…
▽ More
Parallel single-pixel imaging (PSPI) enhances the data acquisition efficiency of single-pixel imaging, but its reconstruction quality depends on a cumbersome and noise-sensitive calibration process. To address this challenge, a PSPI strategy was introduced that leverages modulation region expansion and overlapping reconstruction. This method results in the calibration of modulation of the subregion for each detector, enabling robust operations with undersampled data. It compensates for misalignment via modulation region expansion and overlapping reconstruction, achieving seamless and high-quality imaging that surpasses conventional PSPI in simulations and experiments. Furthermore, this strategy exhibits remarkable robustness, maintaining high imaging quality under extremely nonideal conditions, such as large deflection angles between the array detector and the modulator. This work provides a simple, efficient, and robust framework that simplifies the PSPI workflow and offers broad applicability in high-resolution, high-speed computational imaging.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
When GNNs Fail: Quantifying and Overcoming Temporal Correlation Volatility in Time Series
Authors:
Chen Shao,
Yue Wang,
Zhenyi Zhu,
Zhanbo Huang,
Tobias Käfer,
Zonghan Wu,
Danai Koutra
Abstract:
Modeling multivariate time series by representing them as graphs, where individual series act as nodes and pairwise temporal corre- lations serve as edges, has gained significant traction. Recent advances in Graph Neural Networks (GNNs) have demonstrated strong perfor- mance by assuming a static graph topology and aggregating information from neighboring series. In this work, we investigate the re…
▽ More
Modeling multivariate time series by representing them as graphs, where individual series act as nodes and pairwise temporal corre- lations serve as edges, has gained significant traction. Recent advances in Graph Neural Networks (GNNs) have demonstrated strong perfor- mance by assuming a static graph topology and aggregating information from neighboring series. In this work, we investigate the representa- tional power of GNNs for forecasting under both static and dynamic settings (i.e., when pairwise correlations evolve drastically over time) and identify critical limitations in current architectures. To formalize this, we first propose Temporal Correlation Volatility (TCV), a model- agnostic metric designed to quantify the distributional evolution of these latent structures. We establish a clear connection between TCV and performance degradation, demonstrating that many popular models, including Transformers, generalize poorly in high-TCV settings and are often outperformed by simple structure-agnostic baselines. To address these limitations, we propose Graph Layer for Inference in Dynamic En- vironments (GLIDE), a novel GNN layer enhanced by two theoretically grounded design mechanisms: (D1) Path-based Message Passing, which captures path-based neighborhoods and (D2) Static and Dynamic Propagation Separation, which identifies optimal dynamics via local static approximation. These components significantly improve learning under dynamic topology while preserving robustness in static scenarios. Ex- tensive experiments on synthetic and real-world benchmarks show that GLIDE improves average performance by up to 45.6% across static and dynamic settings, with the largest gain reaching 85.7%. The source code is available at https://github.com/ChenS676/GLIDE.
△ Less
Submitted 7 August, 2026;
originally announced August 2026.
-
Real-Time Megapixel Kilohertz Neuromorphic Shack-Hartmann Wavefront Sensor
Authors:
Yuhan Bao,
Chenxin Shao,
Kaiwei Wang
Abstract:
Conventional frame-based Shack--Hartmann wavefront sensors (SHWFS) are limited by dynamic range and the intrinsic trade-off between spatial and temporal resolution, while high-bandwidth acquisition poses additional challenges for real-time wavefront reconstruction. This work presents a real-time, megapixel, kilohertz neuromorphic SHWFS to overcome these limitations. In static optical metrology, th…
▽ More
Conventional frame-based Shack--Hartmann wavefront sensors (SHWFS) are limited by dynamic range and the intrinsic trade-off between spatial and temporal resolution, while high-bandwidth acquisition poses additional challenges for real-time wavefront reconstruction. This work presents a real-time, megapixel, kilohertz neuromorphic SHWFS to overcome these limitations. In static optical metrology, the proposed pipeline achieves one-shot wavefront acquisition under extreme illumination non-uniformity, reaching a dynamic range of 260 dB at a 20 Hz acquisition frequency. Owing to this high dynamic range and the concomitant high-intensity resolution, wavefront reconstruction errors in dim and bright sub-apertures are reduced by 59\% and 70\%, respectively, relative to conventional frame-based SHWFS. For dynamic wavefront sensing, the system provides kilohertz-rate centroid tracking over a megapixel field of view with microsecond-scale latency. Centroid localization errors are 0.18 pixels during optical alignment supervision and 0.26 pixels in high-speed turbulence observation, verifying the accuracy and reliability of the system across dynamic scenarios. The per-sub-aperture processing throughput reaches 420,737 Hz on a standard CPU, demonstrating high-speed real-time computation without specialized hardware acceleration. Together, these results establish a unified neuromorphic SHWFS framework for high-fidelity one-shot static wavefront reconstruction and real-time high-bandwidth dynamic wavefront sensing.
△ Less
Submitted 28 July, 2026;
originally announced July 2026.
-
RESTOR: Automated Test Oracle Generation for RESTful APIs via Reinforcement Learning
Authors:
Xun Zhou,
Zhen Dong,
Mingyu Ren,
Qiang Li,
JunJie Li,
Sifan Wang,
Xiaolong Yu,
Chaofeng Sha,
Xin Peng
Abstract:
Modern REST API testing faces a critical challenge in defining reliable test oracles, particularly in agile industrial environments where formal specifications (e.g., OpenAPI) are frequently missing or outdated, and historical execution logs are unavailable for newly deployed endpoints. In this paper, we present Restor (Reinforcement Enhanced Single-Traffic Oracle generator for REST APIs), a frame…
▽ More
Modern REST API testing faces a critical challenge in defining reliable test oracles, particularly in agile industrial environments where formal specifications (e.g., OpenAPI) are frequently missing or outdated, and historical execution logs are unavailable for newly deployed endpoints. In this paper, we present Restor (Reinforcement Enhanced Single-Traffic Oracle generator for REST APIs), a framework that generates executable test assertions from a single observed request-response pair in a black-box setting. Unlike existing approaches that rely on rule-based templates or massive training logs, Restor utilizes a novel data augmentation pipeline to fine-tune a lightweight Large Language Model (LLM) via Group Relative Policy Optimization (GRPO). This training process enables the model to internalize testing "common sense" by optimizing a reward function that jointly encourages: (i) the selection of stable, semantically meaningful fields for validation and the avoidance of dynamic noise (e.g., timestamps or trace IDs); (ii) the generation of robust assertions that withstand logic variations. We evaluate Restor on an industrial dataset comprising over 2,300 API traces across 246 real-world services. Comprehensive experiments demonstrate that Restor significantly outperforms prompt-engineered baselines and generalist models, achieving a superior $F_1$ score of 85.42% in key field identification and increasing the proportion of semantically accurate assertions. Furthermore, deployment in a production CI/CD workflow at ByteDance confirms its practical value: the system raised the adoption rate of automatically generated test cases from 74.1% to over 96%, substantially reducing manual Quality Assurance (QA) effort while ensuring high execution stability.
△ Less
Submitted 26 July, 2026;
originally announced July 2026.
-
Reasoning or Memorization: Can LLMs Understand and Generate Chinese Xiehouyu Riddles?
Authors:
Hai Hu,
Siyuan Song,
Chongtian Shao,
Kejia Zhang,
Tianjian Zhu,
Xiaojing Zhao
Abstract:
In this paper, we push the boundary of LLM reasoning by testing them in a Chinese language game, xiehouyu, with novel xiehouyu created by linguists that had not existed before to avoid data contamination. We use multiple-choice questions (MCQ), free-form explanation generation, and new xiehouyu creation to evaluate LLMs' ability to understand and create xiehouyu. In MCQ, we use the delta of accura…
▽ More
In this paper, we push the boundary of LLM reasoning by testing them in a Chinese language game, xiehouyu, with novel xiehouyu created by linguists that had not existed before to avoid data contamination. We use multiple-choice questions (MCQ), free-form explanation generation, and new xiehouyu creation to evaluate LLMs' ability to understand and create xiehouyu. In MCQ, we use the delta of accuracy ($Δ_{acc}$) between existing but low-frequency xiehouyu and novel ones as an index for memorization. $Δ_{acc}$ for native speakers is very low, suggesting similar processing mechanisms. However, we found that frontier Chinese models have on average a $Δ_{acc}$ of 23.6\%, while English-centric models tested have a mean $Δ_{acc}$ of 5.1\%, suggesting that frontier Chinese models are likely trained with much larger Chinese data, thus memorizing more low-frequency xiehouyu. For novel xiehouyu, Gemini 3.1 Pro demonstrated remarkable ability with acc 92.6, which is 24\% higher than human accuracy. In xiehouyu creation, those created by LLMs receive much worse ratings than those by humans. These results suggest that claims about the reasoning abilities of LLMs may need careful re-examination considering the data contamination issue, and that LLMs' creativity in language-related tasks may still be behind human experts, at least in Chinese xiehouyu.
△ Less
Submitted 25 July, 2026;
originally announced July 2026.
-
Constraints on Primordial Black Hole Dressed by Dark Matter Halo from Microlensing Effect of Fast Radio Bursts
Authors:
Hong-Rui Tao,
Huan Zhou,
Cheng-Gang Shao,
Xiao-Long Gong,
Zheng-Xiang Li
Abstract:
Primordial black holes (PBHs) are not only considered as a candidate for dark matter, but also as potential sources of gravitational waves from binary black hole mergers by the LIGO-Virgo-KAGRA and as seeds for the supermassive black holes observed by the James-Webb Space Telescope, thereby remaining intense interest in cosmology and astrophysics. Fast radio bursts (FRBs) are bright millisecond-du…
▽ More
Primordial black holes (PBHs) are not only considered as a candidate for dark matter, but also as potential sources of gravitational waves from binary black hole mergers by the LIGO-Virgo-KAGRA and as seeds for the supermassive black holes observed by the James-Webb Space Telescope, thereby remaining intense interest in cosmology and astrophysics. Fast radio bursts (FRBs) are bright millisecond-duration radio transients whose physical origin remains elusive, which have rapidly developed into one of the most active and rapidly evolving fields in astronomy. The microlensing effect of FRBs offers a clean and powerful probe of PBHs, especially in the mass range above stellar-mass window. In this work, we derive a complete transformation that converts any upper limit on the abundance of PBHs originally derived for `bare' PBHs with monochromatic mass distribution, into the corresponding constraint on `dressed' PBHs with arbitrary extended mass distributions. Based on this framework, we estimate the future constraints on the dressed PBH abundance \(f_{\mathrm{PBH}}\) from FRB observations assuming an expected sample of \(10^5\) FRBs accumulated over the next decade well within the projected detection capabilities of SKA. Our results indicate that including halo enhancement tightens the upper limits on \(f_{\mathrm{PBH}}\) by approximately one order of magnitude, with the most stringent constraint reaching \(\sim10^{-4}\) for the typical mass range from stellar-mass to intermediate-mass black holes.
△ Less
Submitted 23 July, 2026; v1 submitted 19 July, 2026;
originally announced July 2026.
-
A Synthetic 3D Gear Dataset for Manufacturing Quality Inspection (MFGNet-Gear)
Authors:
Ruo-Syuan Mei,
Chenhui Shao
Abstract:
Quality control in smart manufacturing increasingly relies on data-driven methods, particularly deep learning, to automate the inspection of manufactured parts. Recent advances in three-dimensional (3D) metrology have enabled fine-scale assessment of dimensional accuracy, surface quality, and shape conformity. However, deep learning methods for point-cloud-based inspection require large volumes of…
▽ More
Quality control in smart manufacturing increasingly relies on data-driven methods, particularly deep learning, to automate the inspection of manufactured parts. Recent advances in three-dimensional (3D) metrology have enabled fine-scale assessment of dimensional accuracy, surface quality, and shape conformity. However, deep learning methods for point-cloud-based inspection require large volumes of labeled data covering part designs and defect types, which are costly and time-consuming to obtain. Moreover, defective parts are intrinsically rare in mass production, and the resulting class imbalance can degrade model performance and make rare defect types difficult to detect. Synthetic data generation (SDG) offers a promising approach to address these challenges by producing large, balanced, and fully annotated datasets. Yet, applying SDG to precision components requires representing part geometry and defect morphology parametrically, so that design and quality can be co-varied. This article describes MFGNet-Gear, a publicly available synthetic 3D dataset comprising 24,000 paired polygon meshes and point clouds across 12 gear designs and 4 quality classes, with 500 instances per design-quality combination. Gear geometries are generated with parametric computer-aided design software, with dimensional parameters perturbed by $\pm$0.0254 mm and defect parameters sampled from distributions representing defect morphologies. For each mesh, 100,000 points are uniformly sampled using Open3D and stored as N $\times$ 3 coordinate text files. Metadata labels identify the gear design and quality class, supporting part design classification, geometric defect detection, representation learning, and dataset benchmarking. MFGNet-Gear provides an open-source dataset for deep learning-based 3D metrology, with a reproducible generation pipeline extensible to additional part designs.
△ Less
Submitted 11 July, 2026;
originally announced July 2026.
-
Geometry-Optimized Complex-Domain error-diffusion encoding for Fourier Single-Pixel Imaging
Authors:
Chongwu Shao,
Yue Cao,
Wei Zhang,
Xiaopeng-Jin,
Yingran Shen,
Shijian Li,
Xu-Ri Yao
Abstract:
This work proposes a geometry-optimized complex-domain error-diffusion encoding method for Fourier single-pixel imaging. Instead of independently binarizing multiple grayscale phase-shifting patterns, the proposed method directly represents each complex-valued Fourier basis pattern using K (K >= 3) weighted binary patterns while diffusing the residual error in the complex domain. A geometric inter…
▽ More
This work proposes a geometry-optimized complex-domain error-diffusion encoding method for Fourier single-pixel imaging. Instead of independently binarizing multiple grayscale phase-shifting patterns, the proposed method directly represents each complex-valued Fourier basis pattern using K (K >= 3) weighted binary patterns while diffusing the residual error in the complex domain. A geometric interpretation is further established, revealing that the encoding process can be viewed as approximating the Fourier-basis unit circle by a regular polygon in the complex plane. Based on this geometric interpretation, practical optimization strategies are developed for K = 3, K = 4, and K = 7. Both numerical simulations and real-object experiments demonstrate consistently superior reconstruction quality compared with conventional phase-shifting dithering.
△ Less
Submitted 13 July, 2026;
originally announced July 2026.
-
The First ChineseBabyLM Challenge: training data-efficient and cognitively plausible language models for Chinese
Authors:
Siyuan Song,
Zhiheng Qian,
Yunhao Zhang,
Linyang He,
Xiaozhe Ji,
Yingxin Lin,
Hongao Zhu,
Chongtian Shao,
Chuhan Lang,
Luan Li,
Rui Wang,
Renfen Hu,
Shaonan Wang,
Hai Hu
Abstract:
This paper presents the first ChineseBabyLM Challenge, organized as part of NLPCC 2026. The challenge asked participants to train language models from scratch using no more than 102M Chinese words. The models were evaluated on three tracks: natural language understanding, cognitive alignment, and Hanzi knowledge. There were no restrictions on tokenizers, model architectures, or the number of train…
▽ More
This paper presents the first ChineseBabyLM Challenge, organized as part of NLPCC 2026. The challenge asked participants to train language models from scratch using no more than 102M Chinese words. The models were evaluated on three tracks: natural language understanding, cognitive alignment, and Hanzi knowledge. There were no restrictions on tokenizers, model architectures, or the number of training epochs. Eighteen teams submitted 28 distinct models, generating 74 result files. The overall-winning team used a DeBERTa-v2 architecture and introduced an auxiliary pinyin-prediction objective during pretraining. Several submissions also explored curriculum-learning strategies and architectural innovations. Overall, the challenge provides a benchmark for advancing data-efficient and cognitively plausible approaches to Chinese language modeling.
△ Less
Submitted 17 August, 2026; v1 submitted 12 July, 2026;
originally announced July 2026.
-
Late-time cosmological constraints on three holographic dark energy models with DESI DR2 BAO and Type Ia supernovae
Authors:
Jiayuan Huang,
Tonghua Liu,
Chenggang Shao
Abstract:
We constrain three holographic-inspired dark energy models, namely holographic dark energy (HDE), agegraphic dark energy (ADE), and Ricci dark energy (RDE), using late-time observations from cosmic chronometers, Type Ia supernovae (SNe Ia), DESI DR2 baryon acoustic oscillations (BAO), and {redshift-space distortion (RSD) growth measurements}. Five data combinations are considered: $H(z)+$Pantheon+…
▽ More
We constrain three holographic-inspired dark energy models, namely holographic dark energy (HDE), agegraphic dark energy (ADE), and Ricci dark energy (RDE), using late-time observations from cosmic chronometers, Type Ia supernovae (SNe Ia), DESI DR2 baryon acoustic oscillations (BAO), and {redshift-space distortion (RSD) growth measurements}. Five data combinations are considered: $H(z)+$Pantheon+, $H(z)+$DESI DR2+Pantheon+, $H(z)+$DESI DR2+DES-Dovekie, $H(z)+$DESI DR2+DESY5, and {$H(z)+$DESI DR2+DES-Dovekie+RSD}. We perform Bayesian Markov chain Monte Carlo parameter estimation and compare the models with AIC and BIC. In the BAO-included combinations, HDE gives $H_0\simeq67.3$--$68.0~\mathrm{~km~s^{-1}~Mpc^{-1}}$, $Ω_{m0}\simeq0.270$--$0.272$, and $c\simeq1$, indicating an expansion history close to the de Sitter boundary rather than a robust phantom regime. ADE yields a stable agegraphic parameter $n\simeq2.78$--$2.81$, while RDE gives $γ\simeq0.53$--$0.55$ and persistently favors a low matter density, $Ω_{m0}\simeq0.215$--$0.219$. {Treating $r_d$ as a free parameter reveals a strong negative correlation between $H_0$ and $r_d$, and the RSD-included combination provides a growth-level consistency check through $fσ_8(z)$ without constituting a full perturbative stability analysis.} None of the three models significantly alleviates the Hubble tension. Overall, HDE shows the most balanced phenomenological behavior among the three models, although current late-time data do not decisively prefer it over $Λ\text{CDM}$.
△ Less
Submitted 2 July, 2026;
originally announced July 2026.
-
Clock-noise subtraction in geometric time-delay interferometry for space-based gravitational-wave parameter estimation
Authors:
Rui Luo,
Pan-Pan Wang,
Zi-Jiang Yang,
Wei-Liang Qian,
Cheng-Gang Shao
Abstract:
Millihertz gravitational-wave observations with space-based interferometers require time-delay interferometry (TDI) observables whose residual instrumental noise is sufficiently controlled for both detection and parameter inference. Although TDI suppresses laser phase noise in unequal and time-dependent arms, clock jitter from onboard ultra-stable oscillators can remain above the secondary-noise f…
▽ More
Millihertz gravitational-wave observations with space-based interferometers require time-delay interferometry (TDI) observables whose residual instrumental noise is sufficiently controlled for both detection and parameter inference. Although TDI suppresses laser phase noise in unequal and time-dependent arms, clock jitter from onboard ultra-stable oscillators can remain above the secondary-noise floor and bias the effective noise weighting used in data analysis. We formulate a clock-noise subtraction scheme directly in the geometric-TDI framework. The construction introduces generalized clock-noise observables for the four space-time link structures that arise when both delay and time-advance operators are allowed. This makes the clock-noise residual algebraically parallel to the laser-noise residual and yields explicit subtraction terms for arbitrary two-path geometric TDI observables. We illustrate the method with representative first- and second-generation geometric TDI combinations, and test it with time-domain simulations using LISA-like orbits and noise levels. For a modified second-generation U-type observable, the subtraction suppresses the clock-noise residual below the signal region, restores the expected sensitivity to a monochromatic source, and improves the Fisher and Markov-chain Monte Carlo parameter constraints on the source amplitude, frequency and phase. These results show that clock-noise calibration is a necessary component of precision data analysis for future space-based gravitational-wave detectors.
△ Less
Submitted 10 July, 2026;
originally announced July 2026.
-
DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation
Authors:
Xin Cheng,
Xingkai Yu,
Chenze Shao,
Jiashi Li,
Yunfan Xiong,
Yi Qian,
Jiaqi Zhu,
Shirong Ma,
Xiaokang Zhang,
Jiasheng Ye,
Qinyu Chen,
Chengqi Deng,
Jiping Yu,
Damai Dai,
Zhengyan Zhang,
Yixuan Wei,
Yixuan Tan,
Wenkai Yang,
Runxin Xu,
Yu Wu,
Zhean Xu,
Xuanyu Wang,
Muyang Chen,
Rui Tian,
Xiao Bi
, et al. (8 additional authors not shown)
Abstract:
Speculative decoding accelerates Large Language Model (LLM) inference by decoupling draft generation from target verification. While recent parallel drafters efficiently propose long token sequences in a single forward pass, they suffer from rapid acceptance decay due to a lack of inter-token dependencies. Furthermore, indiscriminately verifying these extended blocks wastes critical batch capacity…
▽ More
Speculative decoding accelerates Large Language Model (LLM) inference by decoupling draft generation from target verification. While recent parallel drafters efficiently propose long token sequences in a single forward pass, they suffer from rapid acceptance decay due to a lack of inter-token dependencies. Furthermore, indiscriminately verifying these extended blocks wastes critical batch capacity on tokens with high rejection risks, severely degrading throughput in high-concurrency serving systems. We introduce DSpark, a speculative decoding framework that unifies high-throughput parallel generation with adaptive, load-aware verification. To maintain draft quality, DSpark utilizes a semi-autoregressive architecture, coupling a parallel backbone with a lightweight sequential module, to introduce intra-block dependency modeling and mitigate suffix decay. To optimize system efficiency, DSpark employs confidence-scheduled verification, dynamically tailoring the verification length for each request based on estimated prefix survival probabilities and engine-specific throughput profiles. On offline benchmarks across diverse domains, DSpark substantially improves the accepted length over state-of-the-art autoregressive and parallel drafters. When deployed within the DeepSeek-V4 serving system under live user traffic, DSpark successfully mitigates verification waste. Compared to the established production baseline (MTP-1), DSpark accelerates per-user generation speeds by 60 to 85 percent at matched throughput levels. More importantly, by preventing severe throughput degradation under strict interactivity constraints, it enables performance tiers that were previously unattainable, shifting the Pareto frontier of our serving system.
△ Less
Submitted 6 July, 2026;
originally announced July 2026.
-
Low-ancilla block encodings via Hamiltonian simulation
Authors:
Yuxin Zhang,
Changpeng Shao
Abstract:
Block encodings are a central primitive in quantum algorithms, but standard constructions typically require logarithmic ancilla overhead and complicated controlled operations. Recent lower bounds further show that such ancilla overhead is unavoidable for exact constructions in broad circuit models. We show that this barrier can be bypassed in the approximate setting. Specifically, we present a sim…
▽ More
Block encodings are a central primitive in quantum algorithms, but standard constructions typically require logarithmic ancilla overhead and complicated controlled operations. Recent lower bounds further show that such ancilla overhead is unavoidable for exact constructions in broad circuit models. We show that this barrier can be bypassed in the approximate setting. Specifically, we present a simple single-ancilla construction that converts Hamiltonian evolution into a block encoding of the underlying Hamiltonian, via generalized quantum signal processing. For operators given by Hermitian decompositions $A=\sum_{j=1}^L α_j H_j$, we instantiate this block-encoding construction in two ways, which differ in how the required Hamiltonian evolution is implemented. Using higher-order Trotterization, we obtain an $\varepsilon$-approximate block encoding of $A$ with only one ancilla qubit and circuit depth $\widetilde O\big(L(α/\varepsilon)^{o(1)}\big),$ where $α=\sum_j α_j$. Using multiproduct formulas, we obtain circuit depth $\widetilde O(L)$, at the cost of $O(\log\log(1/\varepsilon))$ ancilla qubits. Our constructions provide alternatives to the standard LCU framework, with a focus on reducing the number of ancilla qubits while maintaining (near-)optimal circuit depth.
△ Less
Submitted 2 July, 2026;
originally announced July 2026.
-
Surrogate Fidelity: When Can Open LLMs Explain Closed Ones?
Authors:
Philippe Chlenski,
Zachariah Carmichael,
Ayush Warikoo,
Chia-Tse Shao,
Yingxiao Ye,
Aobo Yang,
Vivek Miglani,
Nehal Bandi
Abstract:
Mechanistic interpretability (MI) requires full access to model internals, yet the APIs for most widely deployed language models at best expose log-probabilities over output tokens. This creates a surrogate problem: when do measurements made on open models allow us to make claims about a closed model? We evaluate surrogate fidelity at the prediction, attribution, and representation levels. For bin…
▽ More
Mechanistic interpretability (MI) requires full access to model internals, yet the APIs for most widely deployed language models at best expose log-probabilities over output tokens. This creates a surrogate problem: when do measurements made on open models allow us to make claims about a closed model? We evaluate surrogate fidelity at the prediction, attribution, and representation levels. For binary classification tasks, log-odds provide an API-compatible scalar readout of the model's representation space, and leave-one-out attributions provide insight into model behavior. Across eleven models spanning four families (Llama, Qwen, GPT, and Gemini), we find that prediction fidelity substantially overstates attribution fidelity: models that agree on what the answer is often disagree on why. We document an access-validity inversion: white-box signals like attention patterns and perturbation magnitudes are highly stable across models but only weakly predictive of causal attributions, which black-box input ablations capture by design. Mechanistic insight does not automatically transfer to closed targets, and prediction-level agreement is insufficient to warrant such transfer. Code and results are available at https://github.com/facebookresearch/surrogate.
△ Less
Submitted 30 June, 2026;
originally announced June 2026.
-
Vector alignment in matrix Lie groups
Authors:
Congzhou M Sha
Abstract:
Two observers of the same physical system may differ in gauge by a group element acting on their common vector representations, and recovering that element from finite, noisy paired observations is useful in both theory and experiment. The Kabsch and Horn algorithms solve this problem for rotated frames in $\mathbb R^3$ (i.e. $SO(3)$); our earlier Lie algebra method extends it to the Lorentz group…
▽ More
Two observers of the same physical system may differ in gauge by a group element acting on their common vector representations, and recovering that element from finite, noisy paired observations is useful in both theory and experiment. The Kabsch and Horn algorithms solve this problem for rotated frames in $\mathbb R^3$ (i.e. $SO(3)$); our earlier Lie algebra method extends it to the Lorentz group $SO(3,1)_+$. Here we report explicit formulae for the Lie algebra method on the classical matrix Lie groups ($GL(n)$, $SL(n)$, $SO(n)$, $U(n)$, $SO(p,q)$, $Sp(n)$, $Spin(n)$, and $SE(n)$) over the real and complex fields. The four steps (pseudoinverse, matrix logarithm, projection onto the Lie algebra, matrix exponential) are exact for noiseless data. Only the projection is group-dependent, and we show it yields the unique least squares-optimal element of the Lie algebra whenever its image lies in $\mathfrak g$ and its residual is orthogonal to $\mathfrak g$. For noisy data the method is optimal only to leading order, so we add a quasi-Newton correction whose accuracy interpolates between the uncorrected method and direct least squares optimization. The projections, their optimality, and the identity underlying the correction are formally proven in Lean 4.31.0 with Mathlib, and numerical experiments are benchmarked in Julia.
△ Less
Submitted 19 July, 2026; v1 submitted 29 June, 2026;
originally announced June 2026.
-
CellDETR: A Detection-Guided Framework for Scalable Cell Representation Learning from Histopathology Images
Authors:
Shikang Zhang,
Guojun Li,
Yicong Mao,
Chulin Sha
Abstract:
Recent advances in pathology foundation models have substantially improved patch and slide level representation learning from whole-slide images (WSIs).However, cell-level representations learning remain underexplored, limiting cell resolved interpretability, biological discovery, and clinical translation. We propose CellDETR, a detection-guided framework built on Deformable DETR for scalable cell…
▽ More
Recent advances in pathology foundation models have substantially improved patch and slide level representation learning from whole-slide images (WSIs).However, cell-level representations learning remain underexplored, limiting cell resolved interpretability, biological discovery, and clinical translation. We propose CellDETR, a detection-guided framework built on Deformable DETR for scalable cell representation learning from WSIs. By introducing location feature decoupling and box-constrained attention mechanism, CellDETR enables automated extraction of cell-level embeddings, and outperform existing state-of-the-art methods in supervised cell classification on PanNuke data. In addition, by incorporating contrastive learning design, we build a CellDETR-based pretraining model for scalable cell representation learning from unlabeled WSIs, which improves downstream cell classification performance. Furthermore, we show that after pretraining with Xenium spatial transcriptomics-derived cell annotations, CellDETR achieves accurate cross-dataset cell classification, demonstrating the transferability and biological relevance of the learned cell embeddings. Together, CellDETR provides a scalable route toward general cell-level representation learning framework for interpretable computational patholog
△ Less
Submitted 28 June, 2026;
originally announced June 2026.
-
Gravitational-wave response functions for space-borne detectors based on multiple geometric time-delay interferometry links
Authors:
Rui Luo,
Pan-Pan Wang,
Lin-Lin Yang,
Xin-Lei Zhao,
Cheng-Gang Shao
Abstract:
The primary challenge for space-borne gravitational wave (GW) detectors lies in extracting the weak GW signal from instrumental noise that exceeds the signal level by many orders of magnitude. Time-delay interferometry (TDI) addresses this by suppressing the dominant laser phase noise through recombination of time-delayed measurement data. The detector's response to a GW signal is represented in t…
▽ More
The primary challenge for space-borne gravitational wave (GW) detectors lies in extracting the weak GW signal from instrumental noise that exceeds the signal level by many orders of magnitude. Time-delay interferometry (TDI) addresses this by suppressing the dominant laser phase noise through recombination of time-delayed measurement data. The detector's response to a GW signal is represented in the frequency domain by a response function. Currently, the GW signal response is first expressed in terms of the Doppler frequency shift in a single detection arm, and this formulation is then incorporated into specific TDI combinations to derive the corresponding response function. This paper introduces a generalized formulation for TDI combinations based on multiple geometric links. By extending the representation of the laser Doppler frequency shift to include various geometric configurations, such as round-trip and non-round-trip links, we reformulate 45 second-generation TDI combinations. For several of these, the new formulation significantly streamlines their mathematical expressions and enhances physical clarity. Our results demonstrate that the proposed link-mapping rules not only enable efficient construction of response functions for these TDI combinations but also reduce computational complexity. This approach provides a reliable theoretical and algorithmic foundation for data processing in future space-borne GW missions.
△ Less
Submitted 26 June, 2026;
originally announced June 2026.
-
A multi-task spatiotemporal deep neural network for predicting penetration depth and morphology in laser welding
Authors:
Sen Li,
Haichao Cui,
Chendong Shao,
Yaqi Wang,
Xinhua Tang
Abstract:
In laser penetration welding, the assessment of penetration state and weld seam morphology plays a crucial role in determining the weld quality. This paper presents a comprehensive introduction of the innovative muti-task deep learning model that has the capability to predict penetration state, depth, and weld seam morphology with high accuracy. The monitoring platform relies on weld pool images c…
▽ More
In laser penetration welding, the assessment of penetration state and weld seam morphology plays a crucial role in determining the weld quality. This paper presents a comprehensive introduction of the innovative muti-task deep learning model that has the capability to predict penetration state, depth, and weld seam morphology with high accuracy. The monitoring platform relies on weld pool images captured during the laser welding process using a complementary metal-oxide-semiconductor camera. The proposed model integrates spatiotemporal features extracted from top weld pool images along with welding parameters, establishing a deep learning framework based on convolutional neural networks and state space models for more efficient extraction and processing of spatial-temporal information. Furthermore, a reliable method for constructing the dataset is proposed to enhance both robustness and generalization capability of the developed model. Validation results on the test set demonstrate that prediction accuracy for penetration state can reach 99.35%, while prediction error for penetration depth is 1.79 millimeter, and accuracy of reconstructing the weld cross-section is 95.65%. This study provides new insights and methodologies for in-situ quality control strategies in laser penetration welding systems.
△ Less
Submitted 24 June, 2026;
originally announced June 2026.
-
A cross-process welding penetration status prediction algorithm based on unsupervised domain adaptation in laser and TIG welding
Authors:
Sen Li,
Haichao Cui,
Chendong Shao,
Yaqi Wang,
Xinhua Tang
Abstract:
Supervised deep learning has been widely used for weld penetration state classification; however, its performance often degrades significantly under domain shift, such as when transferring models between welding processes with distinct physical mechanisms:for instance, from arc-dominated tungsten inert gas (TIG) welding to keyhole-based laser welding. To overcome this limitation, we propose an uns…
▽ More
Supervised deep learning has been widely used for weld penetration state classification; however, its performance often degrades significantly under domain shift, such as when transferring models between welding processes with distinct physical mechanisms:for instance, from arc-dominated tungsten inert gas (TIG) welding to keyhole-based laser welding. To overcome this limitation, we propose an unsupervised domain adaptation (UDA) framework integrated with a gradual source domain expansion (GSDE) strategy. Evaluated on dedicated TIG and laser welding datasets, our approach achieves high accuracy in both same-process and cross-process transfer tasks. Specifically, it attains average accuracies of 90.65% on TIGFH and 90.72% on LSPS in same-process settings, surpassing a supervised baseline by 35.83% and 38.87%, respectively. More notably, in cross-process scenarios, it reaches 80.48% for TIG to Laser and 81.13% for Laser to TIG, improving upon the baseline by 43.39% and 43.40%. UMAP visualizations verify that the model learns domain-invariant features while maintaining discriminative class boundaries. This method considerably lowers the relabeling cost for new welding processes and enhances the versatility of intelligent monitoring across different welding systems.
△ Less
Submitted 24 June, 2026;
originally announced June 2026.
-
A welding penetration prediction model for laser welding process based on self-supervised learning using physics-informed neural networks
Authors:
Sen Li,
Xiaoying Liu,
Xiaojian Xu,
Chendong Shao,
Yaqi Wang,
Ling Lan,
Xinhua Tang,
Haichao Cui
Abstract:
The laser welding full-penetration is of critical importance, as it constitutes one of the fundamental factors in achieving defect-free welded joints. Accurate prediction of the penetration state is therefore essential for ensuring weld quality. To this end, this paper introduces SimPhysNet, a novel algorithm that achieves high classification accuracy in laser welding penetration prediction using…
▽ More
The laser welding full-penetration is of critical importance, as it constitutes one of the fundamental factors in achieving defect-free welded joints. Accurate prediction of the penetration state is therefore essential for ensuring weld quality. To this end, this paper introduces SimPhysNet, a novel algorithm that achieves high classification accuracy in laser welding penetration prediction using only a limited number of labelled images. This approach effectively overcomes the limitations of supervised learning classification algorithms, which are hindered in industrial applications by their dependence on extensive, high-quality labelled data. The core of SimPhysNet is a unique self-supervised learning paradigm that embeds physical priors into a contrastive learning framework. By incorporating a physics-informed neural network (PINN), the model is guided to extract physically meaningful features of the molten pool and keyhole from a large set of unlabelled data, while three image augmentation tasks further enhance its generalization capabilities. Subsequently, a few-shot learning strategy, based on prototypical networks, enables robust classification by constructing class representations from a minimal set of labelled images. Experimental results demonstrate that SimPhysNet achieves a classification accuracy of 96.06% using only 200 labelled images (approximately 5% of the total labelled dataset), which is comparable to the performance of conventional supervised learning algorithms that utilize the entire labelled dataset. This work presents a new, efficient, and highly accurate method, providing the way for the intelligent automation of laser welding.
△ Less
Submitted 24 June, 2026;
originally announced June 2026.
-
Physics-Guided Spatiotemporal State Space Modeling for Lookahead Molten Pool Segmentation in Laser Wire-Feed Welding
Authors:
Sen Li,
Haichao Cui,
Changhao Yin,
Chendong Shao,
Yaqi Wang,
Xinhua Tang,
Fenggui Lu
Abstract:
Real-time weld-pool perception is critical for closed-loop control in laser wire-feed welding, where sensing, computation, and actuator response introduce unavoidable delay. This paper presents a physics-guided spatiotemporal state space network for lookahead weld-pool segmentation. The model uses historical coaxial grayscale images, welding process parameters, and aligned wire-state electrical si…
▽ More
Real-time weld-pool perception is critical for closed-loop control in laser wire-feed welding, where sensing, computation, and actuator response introduce unavoidable delay. This paper presents a physics-guided spatiotemporal state space network for lookahead weld-pool segmentation. The model uses historical coaxial grayscale images, welding process parameters, and aligned wire-state electrical signals to predict the future semantic layout of three physically meaningful regions: keyhole, wire, and molten pool. It combines a visual encoder, process- and sensor-conditioned feature normalization, patch-level temporal state space modeling, horizon-conditioned latent prediction, dense future feature prediction, and a motion-aware mask decoder. Auxiliary signed-distance-function supervision, temporal consistency, feature distillation, and fine-grained keyhole losses further constrain the predicted geometry and local motion. Experiments on a 43-sequence laser welding dataset show that the proposed WeldMamba reaches 74.63\% mIoU at a 500 ms lookahead. Ablation studies further show that temporal history, patch-level state space modeling, and keyhole motion awareness are the main contributors to robust future segmentation.
△ Less
Submitted 22 June, 2026;
originally announced June 2026.
-
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
Authors:
DeepSeek-AI,
Anyi Xu,
Bangcai Lin,
Bing Xue,
Bingxuan Wang,
Bingzheng Xu,
Bochao Wu,
Bowei Zhang,
Chaofan Lin,
Chen Dong,
Chenchen Ling,
Chengda Lu,
Chenggang Zhao,
Chengqi Deng,
Chengyu Hou,
Chenhao Xu,
Chenze Shao,
Chong Ruan,
Conner Sun,
Damai Dai,
Daya Guo,
Dejian Yang,
Deli Chen,
Donghao Li,
Dongjie Ji
, et al. (294 additional authors not shown)
Abstract:
We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) -- both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: (1) a hybrid attention arc…
▽ More
We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) -- both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: (1) a hybrid attention architecture that combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to improve long-context efficiency; (2) Manifold-Constrained Hyper-Connections (mHC) that enhance conventional residual connections; (3) and the Muon optimizer for faster convergence and greater training stability. We pre-train both models on more than 32T diverse and high-quality tokens, followed by a comprehensive post-training pipeline that unlocks and further enhances their capabilities. DeepSeek-V4-Pro-Max, the maximum reasoning effort mode of DeepSeek-V4-Pro, redefines the state-of-the-art for open models, outperforming its predecessors in core tasks. Meanwhile, DeepSeek-V4 series are highly efficient in long-context scenarios. In the one-million-token context setting, DeepSeek-V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache compared with DeepSeek-V3.2. This enables us to routinely support one-million-token contexts, thereby making long-horizon tasks and further test-time scaling more feasible. The model checkpoints are available at https://huggingface.co/collections/deepseek-ai/deepseek-v4.
△ Less
Submitted 26 April, 2026;
originally announced June 2026.
-
SRL: Combining SLIP Model and Reinforcement Learning for Agile Robotic Jumping
Authors:
Xiaowen Hu,
Linqi Ye,
Yudi Zhu,
Chenyue Shao,
Rankun Li,
Qingdu Li,
Yan Peng
Abstract:
Robotic jumping is pivotal in applications such as search and rescue and logistics, where crossing obstacles and enhancing mobility efficiency are critical. The Spring-Loaded Inverted Pendulum (SLIP) model leverages simplified spring-mass dynamics that naturally encode biologically plausible hopping motions, yet its performance degrades on irregular terrain due to idealized assumptions regarding c…
▽ More
Robotic jumping is pivotal in applications such as search and rescue and logistics, where crossing obstacles and enhancing mobility efficiency are critical. The Spring-Loaded Inverted Pendulum (SLIP) model leverages simplified spring-mass dynamics that naturally encode biologically plausible hopping motions, yet its performance degrades on irregular terrain due to idealized assumptions regarding contact and joint dynamics. Meanwhile, Reinforcement Learning (RL) can adapt to diverse and complex environments but often requires extensive data from unguided exploration. The complementary strengths of SLIP's physically grounded baseline and RL's adaptive capabilities motivate a hybrid framework that overcomes these individual limitations. We therefore propose Spring-loaded Reinforcement Learning (SRL), which integrates SLIP-based feedforward control signals with RL-driven real-time feedback, enabling continuous optimization of robotic jumping. Experimental results demonstrate that SRL can achieve more stable jumps with much less training time than the baseline method, maintaining an average position tracking error below 0.1 m and velocity tracking errors within +/-3% of the target values. Through bipedal and quadrupedal simulations of ground and stair jumping, as well as sim-to-sim and sim-to-real validations, SRL exhibits robust adaptability to various task requirements and environmental complexities, underscoring its potential for real-world deployment.
△ Less
Submitted 16 June, 2026;
originally announced June 2026.
-
Doppler-shifted X-ray Spectroscopy of Nonradiative Electron Capture in Relativistic Collisions of Xe54+ Ions with Kr and Xe Atoms
Authors:
Bian Yang,
Deyang Yu,
Konstantin N. Lyashchenko,
Caojie Shao,
Zhongwen Wu,
Mingwu Zhang,
Oleg Yu. Andreev,
Junliang Liu,
Zhangyong Song,
Yingli Xue,
Wei Wang,
Fangfang Ruan,
Yehong Wu,
Rongchun Lu,
Chenzhong Dong,
Xiaohong Cai
Abstract:
We present an angular-resolved Doppler spectroscopy study of nonradiative electron capture in relativistic collisions of bare Xe54+ ions with Kr and Xe gas targets at the HIRFL-CSR storage ring. The energy spectra and angular distributions of X-rays emitted from fast-moving down-charged projectiles were measured at five observation angles of 35°, 60°, 90°, 120°, and 145° and three collision energi…
▽ More
We present an angular-resolved Doppler spectroscopy study of nonradiative electron capture in relativistic collisions of bare Xe54+ ions with Kr and Xe gas targets at the HIRFL-CSR storage ring. The energy spectra and angular distributions of X-rays emitted from fast-moving down-charged projectiles were measured at five observation angles of 35°, 60°, 90°, 120°, and 145° and three collision energies of 95, 146, and 197 MeV/u by employing the effect of Doppler shift. The transition intensities of Xe53+ ions with small energy differences were precisely determined. In symmetric Xe54+ \to Xe collisions, the transition intensities of Xe53+ and Xe52+ ions were identified when X-rays emitted by projectiles overlapped with K X-rays arising from target ionization. The anisotropy parameters of the K{α_1}(+M2) transition were derived from the angular emission patterns of the corresponding spectral lines. The relative populations of the L, M, and N-shell excited levels of Xe53+ and Xe52+ were further deduced from the intensity ratios of I(Ly-β)/I(Ly-α), I(Ly-γ)/I(Ly-α), and I(Kα)/I(Ly-α). The energy dependence of the population of excited projectile levels was obtained for both targets. Furthermore, the experimental results were compared with theoretical calculations of nonradiative single- and double-electron capture based on the relativistic eikonal approximation and the independent-electron approximation. These findings provide valuable insights into the magnetic-sublevel population and n-resolved state-selective population of excited states produced in relativistic collisions of highly charged heavy ions with multi-electron atoms.
△ Less
Submitted 14 June, 2026;
originally announced June 2026.
-
Data-driven modeling of Galactic diffuse emission with multi-wavelength observations
Authors:
Xi Liu,
Xiaodong Li,
Sujie Lin,
Yihan Liu,
Chengyu Shao,
Lili Yang,
Le Zhang
Abstract:
We present a data-driven investigation of Galactic diffuse emission. Using multi-frequency Planck radio/microwave maps (30-857 GHz) and Fermi-LAT gamma-ray data (50 MeV-814 GeV), we construct a nonlinear mapping between radio emission and gamma-ray intensity through supervised machine learning. Our models achieve high predictive accuracy (R^2 > 0.90 in the 0.1-10 GeV range), demonstrating that mul…
▽ More
We present a data-driven investigation of Galactic diffuse emission. Using multi-frequency Planck radio/microwave maps (30-857 GHz) and Fermi-LAT gamma-ray data (50 MeV-814 GeV), we construct a nonlinear mapping between radio emission and gamma-ray intensity through supervised machine learning. Our models achieve high predictive accuracy (R^2 > 0.90 in the 0.1-10 GeV range), demonstrating that multi-frequency radio observations encode sufficient information to reconstruct both spatial morphology and spectral properties of diffuse gamma-ray emission. By analyzing model performance across different frequency bands and spatial regions, we identify high-frequency radio bands as the dominant predictor, providing direct empirical support for the hadronic origin of Galactic 0.1-10 GeV gamma rays, while low-frequency radio bands for the leptonic origin above 10 GeV. Residual maps reveal coherent large-scale structures, including Loop I and III, highlighting regions where standard interstellar emission models are incomplete or biased. Compared with the GALPROP model, our machine learning approach yields a higher R^2=0.95 and lower mean absolute relative error (14.7%) in the inner Galactic disk and the Galactic center region. Our results illustrate that machine learning serves as a physically interpretable tool for multi-messenger astrophysics, providing a data-driven baseline for separating non-standard emission components and deriving new constraints on cosmic-ray propagation and interstellar medium structure.
△ Less
Submitted 10 June, 2026;
originally announced June 2026.
-
From Affect to Complex Behavior: Advancing Multimodal Human-Centered AI at the 10th ABAW Workshop & Competition
Authors:
Dimitrios Kollias,
Panagiotis Tzirakis,
Alan Cowen,
Stefanos Zafeiriou,
Irene Kotsia,
Eric Granger,
Marco Pedersoli,
Simon Bacon,
Jens Madsen,
Soufiane Belharbi,
Muhammad Haseeb Aslam,
Chunchang Shao,
Guanyu Hu
Abstract:
The 10th Affective & Behavior Analysis in-the-Wild (ABAW) Workshop and Competition, held at CVPR 2026, continues to advance research on modelling, analysis, understanding of human affect and behavior in real-world, unconstrained environments. The workshop maintains its dual structure, comprising both a competition and a paper track. The ABAW Competition introduces a diverse set of challenges targe…
▽ More
The 10th Affective & Behavior Analysis in-the-Wild (ABAW) Workshop and Competition, held at CVPR 2026, continues to advance research on modelling, analysis, understanding of human affect and behavior in real-world, unconstrained environments. The workshop maintains its dual structure, comprising both a competition and a paper track. The ABAW Competition introduces a diverse set of challenges targeting key aspects of affective and behavioral understanding, including continuous affect (valence-arousal) estimation, discrete affect (expression and action unit) recognition, as well as more complex behavior analysis tasks, such as emotional mimicry intensity estimation, ambivalence/hesitancy recognition and fine-grained violence detection. These challenges are built upon large-scale in-the-wild datasets, providing comprehensive benchmarks for state-of-the-art approaches. In parallel, the paper track presents a wide range of contributions spanning pose, motion & behavior estimation, affect modelling & multimodal learning, benchmarks, datasets & evaluation protocols, fairness, robustness & deployment. Overall, the 10th ABAW Workshop and Competition continues to serve as a key platform for benchmarking, collaboration and innovation, shaping the development of next-generation multimodal, human-centered AI systems.
△ Less
Submitted 24 May, 2026;
originally announced May 2026.
-
Reliability-Constrained Blind Beam Alignment for Backscatter-MIMO mounted Target in Cluttered Multipath Channels
Authors:
Xuehui Dong,
Kai Wan,
Gui Zhou,
Chen Shao,
Miyu Feng,
Robert Caiming Qiu
Abstract:
Practical ISAC is constrained by static clutter and NLoS multipath, which obscure target-coupled echoes and induce spurious peaks for beam alignment. Existing receiver-side methods largely model targets as passive scatterers, limiting the structural separability of target echoes from the environment. This paper establishes a structural correspondence between these limitations and target-side Backs…
▽ More
Practical ISAC is constrained by static clutter and NLoS multipath, which obscure target-coupled echoes and induce spurious peaks for beam alignment. Existing receiver-side methods largely model targets as passive scatterers, limiting the structural separability of target echoes from the environment. This paper establishes a structural correspondence between these limitations and target-side Backscatter-MIMO responses: reflection modulation enables waveform-domain separation from unmodulated clutter, while retro-directional passive beamforming concentrates the tagged echo toward the BS-facing direction and suppresses NLoS-induced false-peak locking. To operationalize this correspondence, dual-end spatial locking is required to overcome cascaded backscatter loss and provide beam-domain angular information. We propose a downlink-triggered blind dual-end alignment protocol that jointly selects the BS and Backscatter-MIMO codeword indices from the tagged echo observed at the BS, without pilots, CSI feedback, or target synchronization. We further derive a clutter-aware remodulation waveform robust to fractional timing offsets and construct adjustable-width BS/Backscatter-MIMO codebooks via quadratic phase spoiling. For reliability characterization, we derive closed-form expressions for the coherence-averaged end-to-end success probability. The analysis shows that beam narrowing is not universally beneficial: in NLoS-dominated regimes, enlarging the array aperture may degrade alignment reliability. The optimal beamwidth is instead governed by cross-phase competition between discovery and alignment, yielding a nontrivial feasible region with an analytically characterized boundary. Simulations validate the analysis and demonstrate improved reliability-gated locked-link performance under strong clutter, severe NLoS multipath, and finite coherence time.
△ Less
Submitted 26 May, 2026;
originally announced May 2026.
-
LiveFigure: Generating Editable Scientific Illustration with VLM Agents
Authors:
Chenyang Shao,
Jiahe Liu,
Fengli Xu,
Yong Li
Abstract:
Scientific illustrations are essential for depicting conceptual designs, methodologies, and experimental workflows in research, playing a pivotal role in communicating complex academic insights. However, creating high-quality scientific illustrations remains a labor-intensive task for human scientists. While recent generative image models have advanced prompt-based editing, the synthesis of fully…
▽ More
Scientific illustrations are essential for depicting conceptual designs, methodologies, and experimental workflows in research, playing a pivotal role in communicating complex academic insights. However, creating high-quality scientific illustrations remains a labor-intensive task for human scientists. While recent generative image models have advanced prompt-based editing, the synthesis of fully editable figures remains a fundamental challenge. Valid editability involves structured transformations of graphical elements, scales, attributes, and text, rather than simple pixel-level changes. Existing models generate raster outputs that do not support manual correction or layout adjustment, limiting their utility in scientific publishing, where editable vector figures are typically required for submission. To address this challenge, we introduce LiveFigure, an agentic framework driven by VLM agents that imitates the multi-step drawing workflow of human researchers. It first plans figure blueprints by drawing inspiration from high-quality references in previous works, then generates executable scripts that produce figures via the PowerPoint interface based on skills and experience, and finally refines the outputs with targeted visual diagnostics, producing fully vectorized, editable figures that meet publication standards. Extensive experiments demonstrate that LiveFigure generates inherently editable figures, achieving 80% publication-readiness in only 17 manual edits, far surpassing the 24% rate of the strongest baseline, NanoBanana. Human preference studies further validate this advantage, with LiveFigure securing a 60% win rate against NanoBanana. Our code is available at https://github.com/tsinghua-fib-lab/LiveFigure.git.
△ Less
Submitted 22 May, 2026;
originally announced May 2026.
-
Evidence for Intermediate-Mass Black Holes From Microlensing Signatures in CHIME/FRB catalog 2
Authors:
Huan Zhou,
Zhengxiang Li,
Cheng-Gang Shao,
Xi-Jing Wang,
Kai Liao,
He Gao,
Zong-Hong Zhu
Abstract:
Intermediate-mass black holes (IMBHs) are the missing link in the cosmic hierarchy of black holes, bridging the gap between stellar-mass black holes and supermassive ones. They also serve as unique laboratories for testing strong-field gravity and are prime targets for future multi-messenger observations. However, IMBHs are a population that has remained notoriously difficult to detect. The microl…
▽ More
Intermediate-mass black holes (IMBHs) are the missing link in the cosmic hierarchy of black holes, bridging the gap between stellar-mass black holes and supermassive ones. They also serve as unique laboratories for testing strong-field gravity and are prime targets for future multi-messenger observations. However, IMBHs are a population that has remained notoriously difficult to detect. The microlensing effect of fast radio bursts (FRBs) can serve as a clean and powerful method to probe IMBHs. In this work, we develop a pipeline to search for microlensed FRBs based on their dynamic spectra and apply it to the CHIME/FRB Catalog 2. Two microlensing signatures have been identified in two separate sources, i.e. FRB~20190131D and FRB~20211115A. The inferred lens masses for these two signatures are $\sim[280-467]~M_{\odot}$ and $\sim[539-609]~M_{\odot}$, respectively. Here we interpret them as evidence for IMBHs. If there are no intervening structures-such as galaxies or clusters-along the line of sights for these two sources, the two identified IMBHs might be isolated and of primordial origins. In that case, we obtain primordial black holes (PBHs) within these two mass ranges would constitute $\sim4\%$ of dark matter. Moreover, if these two candidates are not genuine lensing signatures, the abundance of intermediate-mass PBHs with masses $>300,M_{\odot}$ is constrained to be $\sim13\%$ at $95\%$ confidence level. Therefore, more comprehensive observational information for FRBs, together with a deeper understanding of whether the intrinsic emission mechanisms of FRBs can produce lensing-like signals, will be crucial for establishing this effect as a powerful tool for probing (primordial) IMBHs.
△ Less
Submitted 20 August, 2026; v1 submitted 19 May, 2026;
originally announced May 2026.
-
One Model to Translate Them All: Universal Any-to-Any Translation for Heterogeneous Collaborative Perception
Authors:
Yang Li,
Weize Li,
Quan Yuan,
Congzhang Shao,
Guiyang Luo,
Yunqi Ba,
Xuanhan Zhu,
Xinyuan Ding,
Xiaoyuan Fu,
Jinglin Li
Abstract:
By sharing intermediate features, collaborative perception extends each agent's sensing beyond standalone limits, but real-world feature modality heterogeneity remains a key barrier to effective fusion. Most existing methods, including direct adaption and protocol-based transformation, typically rely on training adapters for newly emerging feature modalities and often require additional retraining…
▽ More
By sharing intermediate features, collaborative perception extends each agent's sensing beyond standalone limits, but real-world feature modality heterogeneity remains a key barrier to effective fusion. Most existing methods, including direct adaption and protocol-based transformation, typically rely on training adapters for newly emerging feature modalities and often require additional retraining or fine-tuning. Such repeated training is costly and is often infeasible across manufacturers due to model and data privacy constraints, limiting real-world scalability. To address this issue, we propose UniTrans, a universal any-to-any feature modality translation model that instantiates translators on the fly for arbitrary modalities.
UniTrans pretrains a bank of translator expert parameters and learns their combination coefficients as a function of source-to-target modality mapping. The mapping is measured in a modality-intrinsic latent space, where an intrinsic encoder extracts modality-specific yet scene-invariant codes from single-frame intermediate features, enabling UniTrans to instantiate translators in a zero-shot manner.
Experiments on OPV2V-H and DAIR-V2X demonstrate that UniTrans consistently outperforms state-of-the-art methods in both simulated and real-world settings, enabling efficient any-to-any translation through a universal model. The code is available at https://github.com/CheeryLeeyy/UniTrans.
△ Less
Submitted 18 May, 2026;
originally announced May 2026.
-
Polarization Birefringence and Waveform Systematics in GW231123
Authors:
Tonghua Liu,
Chenggang Shao,
Kai Liao
Abstract:
GW231123 is a short, massive binary-black-hole event whose source properties show strong waveform dependence. We use this event to test gravitational-wave polarization birefringence, modeled as a frequency-dependent rotation of the tensor-polarization basis. Instead of sampling a distance-normalized coefficient directly, we sample the band-differential rotation…
▽ More
GW231123 is a short, massive binary-black-hole event whose source properties show strong waveform dependence. We use this event to test gravitational-wave polarization birefringence, modeled as a frequency-dependent rotation of the tensor-polarization basis. Instead of sampling a distance-normalized coefficient directly, we sample the band-differential rotation $δ_{\rm br}=Δ(448\,\mathrm{Hz})-Δ(20\,\mathrm{Hz})$ with prior $[-π,π]$, and report the derived coefficient $β_{\rm br}^{\rm derived}$ for comparison with standard propagation parametrizations. We analyze three waveform families: IMRPhenomXPHM (XPHM), IMRPhenomXO4a (XO4a), and NRSur7dq4. The derived posteriors are consistent with the general relativity value, giving $90\%$ upper limits $|β_{\rm br}^{\rm derived}|_{90}=0.378,\,0.097,\,0.273$ for XPHM, XO4a, and NRSur7dq4, respectively. The directly sampled $δ_{\rm br}$ posterior remains broad, with $|δ_{\rm br}|_{90}\simeq2.8\,\mathrm{rad}$, so the accumulated rotation across the analysis band is weakly constrained. The Bayes factors are waveform dependent: $\ln\mathcal{B}_{\rm br/GR}=-1.26\pm0.30$, $+3.64\pm0.28$, and $-0.86\pm0.29$, respectively. We therefore find no waveform-robust evidence for parity-violating propagation. The positive XO4a result is better interpreted as a waveform-dependent birefringence-like response associated with the mass-ratio--distance--spin degeneracy of this short high-mass event.
△ Less
Submitted 14 May, 2026;
originally announced May 2026.
-
A putative, computationally stable structure of homotrimeric BP180/collagen XVII
Authors:
Congzhou M Sha
Abstract:
Background: BP180, also known as collagen XVII and BPAG2 (bullous pemphigoid antigen 2), is a 180-kDa transmembrane protein within the hemidesmosomal plaque complex, and which is known to be a major antigen in bullous pemphigoid, gestational pemphigoid, cicatricial (mucous membrane) pemphigoid, and linear IgA bullous disease.
Objective: At present, the 3D structure of BP180 is not known. The goa…
▽ More
Background: BP180, also known as collagen XVII and BPAG2 (bullous pemphigoid antigen 2), is a 180-kDa transmembrane protein within the hemidesmosomal plaque complex, and which is known to be a major antigen in bullous pemphigoid, gestational pemphigoid, cicatricial (mucous membrane) pemphigoid, and linear IgA bullous disease.
Objective: At present, the 3D structure of BP180 is not known. The goal is to predict a reasonable structure for BP180 through machine learning and molecular dynamics.
Methods: In this work, we use the recent Boltz-2 model to predict a putative structure for the intracellular, transmembrane, and proximal extracellular domains, including the NC16A antigenic region and a portion of its first extracellular collagenous domain, Col-15. We computationally embed BP180 in a simple phospholipid bilayer, demonstrate that the putative structure is stable using molecular dynamics, and analyze its allosteric properties.
Results: The structures presented satisfy symmetry and secondary structure properties which are expected from homology modelling. Over three 500 ns trajectories, there is minor instability of the predicted globular head domain, but the homotrimer otherwise stays mostly folded. The putative NC16A domain is stiff, whereas the truncated Col-15 domain is highly flexible. There does not appear to be a nearby stable conformation distinct from the initial state.
Conclusion: The structure presented is a useful starting point for targeting BP180 pharmacologically, for further experimental characterization of BP180, and for generating hypotheses regarding the relevant epitopes contributing to bullous disease. Diffusion models such as Boltz-2 and AlphaFold3 are useful, but their results must be evaluated carefully.
△ Less
Submitted 9 May, 2026;
originally announced May 2026.
-
From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models
Authors:
Yihan Lin,
Haoyang Li,
Yang Li,
Haitao Shen,
Yihan Zhao,
Chao Shao,
Jing Zhang
Abstract:
Latent actions serve as an intermediate representation that enables consistent modeling of vision-language-action (VLA) models across heterogeneous datasets. However, approaches to supervising VLAs with latent actions are fragmented and lack a systematic comparison. This work structures the study of latent action supervision from two perspectives: (i) regularizing the trajectory via image-based la…
▽ More
Latent actions serve as an intermediate representation that enables consistent modeling of vision-language-action (VLA) models across heterogeneous datasets. However, approaches to supervising VLAs with latent actions are fragmented and lack a systematic comparison. This work structures the study of latent action supervision from two perspectives: (i) regularizing the trajectory via image-based latent actions, and (ii) unifying the target space with action-based latent actions. Under a unified VLA baseline, we instantiate and compare four representative integration strategies. Our results reveal a formulation-task correspondence: image-based latent actions benefit long-horizon reasoning and scene-level generalization, whereas action-based latent actions excel at complex motor coordination. Furthermore, we find that directly supervising the VLM with discrete latent action tokens yields the most effective performance. Finally, our experiments offer initial insights into the benefits of latent action supervision in mixed-data, suggesting a promising direction for VLA training. Code is available at https://github.com/RUCKBReasoning/From_Pixels_to_Tokens.
△ Less
Submitted 6 May, 2026;
originally announced May 2026.
-
LLM-ADAM: A Generalizable LLM Agent Framework for Pre-Print Anomaly Detection in Additive Manufacturing
Authors:
Ahmadreza Eslaminia,
Chuhan Cai,
Cameron Smith,
Ruo-Syuan Mei,
Shichen Li,
Rajiv Malhotra,
Klara Nahrstedt,
Chenhui Shao
Abstract:
Additive manufacturing (AM) continues to transform modern manufacturing by enabling flexible, on-demand production of complex geometries across diverse industries. Fused filament fabrication (FFF) has extended AM to laboratories, classrooms, and small production environments, but this accessibility shifts process-planning responsibility to users who may lack manufacturing expertise. A syntacticall…
▽ More
Additive manufacturing (AM) continues to transform modern manufacturing by enabling flexible, on-demand production of complex geometries across diverse industries. Fused filament fabrication (FFF) has extended AM to laboratories, classrooms, and small production environments, but this accessibility shifts process-planning responsibility to users who may lack manufacturing expertise. A syntactically valid slicer profile can still encode thermally or geometrically harmful settings, and subtle G-code edits can alter extrusion, cooling, or adhesion before a print begins. Pre-print G-code screening catches accidental or adversarial machine-program errors before material or machine time is wasted. This paper proposes LLM-ADAM as a generalizable LLM framework for pre-print anomaly detection in AM. The framework decomposes the task into three roles: Extractor-LLM maps a G-code file to a structured process-parameter schema; Reference-LLM converts printer and material documentation into aligned operating ranges; and Judge-LLM interprets a deterministic deviation table and G-code evidence to decide whether a part is non-defective or belongs to an anomaly class. Printers, materials, and LLM backbones are interchangeable test conditions, not fixed assumptions. We evaluate the framework on an N=200 FFF G-code corpus spanning two desktop printer families, two materials, and five classes including non-defective, under-extrusion, over-extrusion, warping, and stringing. The best framework configuration reaches 87.5% accuracy, compared with 59.5% for the strongest engineered single-LLM baseline. The results show that structured decomposition, rather than backbone strength alone, is the dominant source of improvement, with defect classes identified at or near ceiling for leading configurations while residual errors concentrate on conservative false alarms for non-defective samples.
△ Less
Submitted 4 May, 2026;
originally announced May 2026.
-
Effect of reaction temperature on nascent carbonaceous particles from toluene shock-tube pyrolysis: Insights from FTIR and Raman spectroscopy
Authors:
Meysam K. Rezaeian,
Can Shao,
Jürgen Herzler,
Mustapha Fikri,
Greg J. Smallwood,
Christof Schulz
Abstract:
The transition from gaseous precursors to nascent solid particles and their subsequent structural maturation were investigated in single-pulse shock-tube experiments using ex situ Fourier-transform infrared (FTIR) and Raman spectroscopy of sampled products. A mixture of 2% toluene in argon was pyrolyzed at around 2.0 bar with temperature plateau times of 2.0 ms over the 1450-1800 K reaction temper…
▽ More
The transition from gaseous precursors to nascent solid particles and their subsequent structural maturation were investigated in single-pulse shock-tube experiments using ex situ Fourier-transform infrared (FTIR) and Raman spectroscopy of sampled products. A mixture of 2% toluene in argon was pyrolyzed at around 2.0 bar with temperature plateau times of 2.0 ms over the 1450-1800 K reaction temperature range. In situ laser extinction measurements indicate the onset of particle formation at 1570 K. At this temperature, Raman spectra exhibit emerging D and G bands, and transmission electron microscopy (TEM) reveals the disappearance of poorly defined structures, identifying 1570 K as the phase-transition reaction temperature. Approaching this reaction temperature, Raman spectra show a rapid disappearance of sp hybridized triple carbon bonds. At 1670 K reaction temperature, a maximum in primary particle diameter and a decrease in structural disorder inferred from Raman spectroscopy are observed, defining the ordering threshold. Deconvolution of the FTIR spectra enables separation of in ring double carbon bond stretching vibrations from isolated and ring-conjugated side-chain double carbon bond modes. The in-ring double carbon band is used to normalize aliphatic and aromatic C-H vibrations. FTIR analysis reveals ring-edge structures associated with electron-localization sites, including bay regions, five-membered ring defects, and benzylic positions, indicating a radical-rich environment below the phase-transition temperature. Between the phase-transition and ordering-threshold temperatures, K-regions and armchair structures associated with electron delocalization and thermal stability increase. The emergence of these electronic and structural characteristics highlights the critical role of radicals in soot inception and early structural ordering.
△ Less
Submitted 29 April, 2026;
originally announced April 2026.
-
RCSB PDB AI Help Desk: retrieval-augmented generation for protein structure deposition support
Authors:
Vivek Reddy Chithari,
Jasmine Y. Young,
Irina Persikova,
Yuhe Liang,
Gregg V. Crichlow,
Justin W. Flatt,
Sutapa Ghosh,
Brian P. Hudson,
Ezra Peisach,
Monica Sekharan,
Chenghua Shao,
Stephen K. Burley
Abstract:
Motivation: Structural Biologists have contributed more than 245,000 experimentally determined three-dimensional structures of biological macromolecules to the Protein Data Bank (PDB). Incoming data are validated and biocurated by ~20 expert biocurators across the wwPDB. RCSB PDB biocurators who process more than 40% of global depositions face increasing challenges in maintaining efficient Help De…
▽ More
Motivation: Structural Biologists have contributed more than 245,000 experimentally determined three-dimensional structures of biological macromolecules to the Protein Data Bank (PDB). Incoming data are validated and biocurated by ~20 expert biocurators across the wwPDB. RCSB PDB biocurators who process more than 40% of global depositions face increasing challenges in maintaining efficient Help Desk operations, with approximately 19,000 messages in approximately 8,000 entries received from depositors in 2025.
Results: We developed an AI-powered Help Desk using Retrieval-Augmented Generation (RAG) built on LangChain with a pgvector store (PostgreSQL) and GPT-4.1-mini. The system employs pymupdf4llm for Markdown-preserving PDF extraction, two-stage document chunking, Maximal Marginal Relevance retrieval, a topical guardrail that filters off-topic queries, and a specialized system prompt that prevents exposure of internal terminology. A dual-LLM architecture uses separate model configurations for question condensing and response generation. Deployed in production on Kubernetes with PostgreSQL (pgvector), it provides around-the-clock depositor assistance with citation-backed, streaming responses.
Availability and implementation: Freely available at https://rcsb-deposit-help.rcsb.org.
△ Less
Submitted 13 April, 2026;
originally announced April 2026.
-
The antiferromagnetic Chern insulator phase in the Kane-Mele-Hubbard model
Authors:
Bao-Qing Wang,
Can Shao,
Takami Tohyama,
Hong-Gang Luo,
Hantao Lu
Abstract:
The emergence of the antiferromagnetic (AFM) Chern insulator (AFCI) phase in the Kane-Mele-Hubbard (KMH) model with a finite sublattice potential is investigated. The AFCI, characterized by AFM correlations coexisting with quantized Hall conductance, has long raised the question of whether it can exist in the KMH model that respects time-reversal symmetry (TRS). Using exact diagonalization, we ana…
▽ More
The emergence of the antiferromagnetic (AFM) Chern insulator (AFCI) phase in the Kane-Mele-Hubbard (KMH) model with a finite sublattice potential is investigated. The AFCI, characterized by AFM correlations coexisting with quantized Hall conductance, has long raised the question of whether it can exist in the KMH model that respects time-reversal symmetry (TRS). Using exact diagonalization, we analyze the excitation gap, anisotropic AFM correlations along the $z$ axis and in the $xy$ plane, and the fidelity susceptibility under twisted boundary conditions, all of which provide consistent evidence for the AFCI phase. In particular, our numerical evaluation on the (spin) Chern number reveals a breakdown of adiabatic continuity in the twist-angle space, indicating an instability toward TRS breaking driven by Hubbard-induced AFM perturbations. A modified computational scheme is further proposed, which yields a robust quantized Chern number $C=1$ within this phase.
△ Less
Submitted 24 April, 2026;
originally announced April 2026.
-
Astrophysically Realistic Secondary Spins Trigger Chaos in Schwarzschild Spacetime and Discernible Gravitational Wave Signatures
Authors:
Dan-Dan Yuan,
Jia-Geng Jiao,
Yu-Qi Lei,
Jun-Xi Shi,
Jing-Qi Lai,
Caiying Shao,
Yu Tian
Abstract:
Chaos in extreme-mass-ratio inspirals is often thought to require unrealistically large secondary spins, making its astrophysical relevance uncertain. However, we find that chaos persists across the astrophysically realistic spin range for a spinning secondary orbiting a Schwarzschild black hole. This nonintegrable dynamics leaves clear signatures in the emitted gravitational waves. Nearby regular…
▽ More
Chaos in extreme-mass-ratio inspirals is often thought to require unrealistically large secondary spins, making its astrophysical relevance uncertain. However, we find that chaos persists across the astrophysically realistic spin range for a spinning secondary orbiting a Schwarzschild black hole. This nonintegrable dynamics leaves clear signatures in the emitted gravitational waves. Nearby regular and chaotic trajectories can remain similar in the time domain and retain broadly aligned dominant spectral peaks, yet chaotic signals develop a much less discrete frequency-domain structure with dense inter-peak power. Furthermore, we introduce a local spectral-flatness measure and find it to be several hundred times larger for the chaotic signal than for the neighboring regular signals. Finally, a change in the secondary spin by as little as \(1\%\) of its maximal physically allowed value can drive the system from regular to chaotic motion and produce distinctive detector-level waveforms.
△ Less
Submitted 22 April, 2026;
originally announced April 2026.
-
Targeted Exploration via Unified Entropy Control for Reinforcement Learning
Authors:
Chen Wang,
Lai Wei,
Yanzhi Zhang,
Chenyang Shao,
Zedong Dan,
Weiran Huang,
Ge Lan,
Yue Wang
Abstract:
Recent advances in reinforcement learning (RL) have improved the reasoning capabilities of large language models (LLMs) and vision-language models (VLMs). However, the widely used Group Relative Policy Optimization (GRPO) consistently suffers from entropy collapse, causing the policy to converge prematurely and lose diversity. Existing exploration methods introduce additional bias or variance duri…
▽ More
Recent advances in reinforcement learning (RL) have improved the reasoning capabilities of large language models (LLMs) and vision-language models (VLMs). However, the widely used Group Relative Policy Optimization (GRPO) consistently suffers from entropy collapse, causing the policy to converge prematurely and lose diversity. Existing exploration methods introduce additional bias or variance during exploration, making it difficult to maintain optimization stability. We propose Unified Entropy Control for Reinforcement Learning (UEC-RL), a framework that provides targeted mechanisms for exploration and stabilization. UEC-RL activates more exploration on difficult prompts to search for potential and valuable reasoning trajectories. In parallel, a stabilizer prevents entropy from growing uncontrollably, thereby keeping training stable as the model consolidates reliable behaviors. Together, these components expand the search space when needed while maintaining robust optimization throughout training. Experiments on both LLM and VLM reasoning tasks show consistent gains over RL baselines on both Pass@1 and Pass@$k$. On Geometry3K, UEC-RL achieves a 37.9\% relative improvement over GRPO, indicating that it sustains effective exploration without compromising convergence and underscoring UEC-RL as a key for scaling RL-based reasoning in large models. Our code is available at https://github.com/597358816/UEC-RL.
△ Less
Submitted 17 April, 2026; v1 submitted 16 April, 2026;
originally announced April 2026.
-
Adaptive Unknown Fault Detection and Few-Shot Continual Learning for Condition Monitoring in Ultrasonic Metal Welding
Authors:
Ahmadreza Eslaminia,
Kuan-Chieh Lu,
Klara Nahrstedt,
Chenhui Shao
Abstract:
Ultrasonic metal welding (UMW) is widely used in industrial applications but is sensitive to tool wear, surface contamination, and material variability, which can lead to unexpected process faults and unsatisfactory weld quality. Conventional monitoring systems typically rely on supervised learning models that assume all fault types are known in advance, limiting their ability to handle previously…
▽ More
Ultrasonic metal welding (UMW) is widely used in industrial applications but is sensitive to tool wear, surface contamination, and material variability, which can lead to unexpected process faults and unsatisfactory weld quality. Conventional monitoring systems typically rely on supervised learning models that assume all fault types are known in advance, limiting their ability to handle previously unseen process faults. To address this challenge, this paper proposes an adaptive condition monitoring approach that enables unknown fault detection and few-shot continual learning for UMW. Unknown faults are detected by analyzing hidden-layer representations of a multilayer perceptron and leveraging a statistical thresholding strategy. Once detected, the samples from unknown fault types are incorporated into the existing model through a continual learning procedure that selectively updates only the final layers of the network, which enables the model to recognize new fault types while preserving knowledge of existing classes. To accelerate the labeling process, cosine similarity transformation combined with a clustering algorithm groups similar unknown samples, thereby reducing manual labeling effort. Experimental results using a multi-sensor UMW dataset demonstrate that the proposed method achieves 96% accuracy in detecting unseen fault conditions while maintaining reliable classification of known classes. After incorporating a new fault type using only five labeled samples, the updated model achieves 98% testing classification accuracy. These results demonstrate that the proposed approach enables adaptive monitoring with minimal retraining cost and time. The proposed approach provides a scalable solution for continual learning in condition monitoring where new process conditions may constantly emerge over time and is extensible to other manufacturing processes.
△ Less
Submitted 15 April, 2026;
originally announced April 2026.
-
Removing Motion Artifact in MRI by Using a Perceptual Loss Driven Deep Learning Framework
Authors:
Ziheng Guo,
Danqun Zheng,
Shuai Li,
Chengwei Chen,
Boyang Pan,
Xuezhou Li,
Ziqin Yu,
Langdi Zhong,
Chenwei Shao,
Yun Bian,
Nan-Jie Gong
Abstract:
Purpose: Deep learning-based MRI artifact correction methods often demonstrate poor generalization to clinical data. This limitation largely stems from the inability of deep learning models in reliably distinguishing motion artifacts from true anatomical structures, due to insufficient awareness of artifact characteristics. To address this challenge, we proposed PERCEPT-Net, a deep learning framew…
▽ More
Purpose: Deep learning-based MRI artifact correction methods often demonstrate poor generalization to clinical data. This limitation largely stems from the inability of deep learning models in reliably distinguishing motion artifacts from true anatomical structures, due to insufficient awareness of artifact characteristics. To address this challenge, we proposed PERCEPT-Net, a deep learning framework that enhances structure preserving and suppresses artifact through dedicated perceptual supervision.Method: PERCEPT-Net is built on a residual U-Net backbone and incorporates three auxiliary components. The first multi-scale recovery module is designed to preserve both global anatomical context and fine structural details, while the second dual attention mechanisms further improve performance by prioritizing clinically relevant features. At the core of the framework is the third Motion Perceptual Loss (MPL), an artifact-aware perceptual supervision strategy that learns generalized representations of MRI motion artifacts, enabling the model to effectively suppress them while maintaining anatomical fidelity. The model is trained on a hybrid dataset comprising both real and simulated paired volumes, and its performance is validated on a prospective test set using a combination of quantitative metrics and qualitative assessments by experienced radiologists.Result: PERCEPT-Net outperformed state-of-the-art methods on clinical data. Ablation studies identified the Motion Perceptual Loss as the primary contributor to this performance, yielding significant improvements in structural consistency and tissue contrast, as reflected by higher SSIM and PSNR values. These findings were further corroborated by radiologist evaluations, which demonstrated significantly higher diagnostic confidence in the corrected volumes.
△ Less
Submitted 20 April, 2026; v1 submitted 11 April, 2026;
originally announced April 2026.
-
Worst-case Harrow-Hassidim-Lloyd algorithm with average-case correct quantum Fourier transform
Authors:
Changpeng Shao
Abstract:
In [\href{https://quantum-journal.org/papers/q-2022-12-07-872/}{Quantum 6, 872, 2022}], Linden and de Wolf proposed a lightweight protocol for verifying average-case correctness of the quantum Fourier transform (QFT). They showed that good average-case QFT performance is sufficient for good worst-case performance in several quantum information-processing tasks. In this work, we study whether such…
▽ More
In [\href{https://quantum-journal.org/papers/q-2022-12-07-872/}{Quantum 6, 872, 2022}], Linden and de Wolf proposed a lightweight protocol for verifying average-case correctness of the quantum Fourier transform (QFT). They showed that good average-case QFT performance is sufficient for good worst-case performance in several quantum information-processing tasks. In this work, we study whether such average-case guarantees are also sufficient when the QFT is used coherently inside the Harrow--Hassidim--Lloyd algorithm. We show that the original average-case condition is not quite strong enough for this purpose, due to possible relative phase errors between different eigenspaces. To address this, we introduce a strengthened Linden--de Wolf-type verification condition that controls the relevant coherences, and prove that it guarantees worst-case correctness of the HHL algorithm in several natural settings.
△ Less
Submitted 2 July, 2026; v1 submitted 11 April, 2026;
originally announced April 2026.
-
AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery
Authors:
Yu Li,
Chenyang Shao,
Xinyang Liu,
Ruotong Zhao,
Peijie Liu,
Hongyuan Su,
Zhibin Chen,
Qinglong Yang,
Anjie Xu,
Yi Fang,
Qingbin Zeng,
Tianxing Li,
Jingbo Xu,
Fengli Xu,
Yong Li,
Tie-Yan Liu
Abstract:
Artificial intelligence research increasingly depends on prolonged cycles of reproduction, debugging, and iterative refinement to achieve State-Of-The-Art (SOTA) performance, creating a growing need for systems that can accelerate the full pipeline of empirical model optimization. In this work, we introduce AutoSOTA, an end-to-end automated research system that advances the latest SOTA models publ…
▽ More
Artificial intelligence research increasingly depends on prolonged cycles of reproduction, debugging, and iterative refinement to achieve State-Of-The-Art (SOTA) performance, creating a growing need for systems that can accelerate the full pipeline of empirical model optimization. In this work, we introduce AutoSOTA, an end-to-end automated research system that advances the latest SOTA models published in top-tier AI papers to reproducible and empirically improved new SOTA models. We formulate this problem through three tightly coupled stages: resource preparation and goal setting; experiment evaluation; and reflection and ideation. To tackle this problem, AutoSOTA adopts a multi-agent architecture with eight specialized agents that collaboratively ground papers to code and dependencies, initialize and repair execution environments, track long-horizon experiments, generate and schedule optimization ideas, and supervise validity to avoid spurious gains. We evaluate AutoSOTA on recent research papers collected from eight top-tier AI conferences under filters for code availability and execution cost. Across these papers, AutoSOTA achieves strong end-to-end performance in both automated replication and subsequent optimization. Specifically, it successfully discovers 105 new SOTA models that surpass the original reported methods, averaging approximately five hours per paper. Case studies spanning LLM, NLP, computer vision, time series, and optimization further show that the system can move beyond routine hyperparameter tuning to identify architectural innovation, algorithmic redesigns, and workflow-level improvements. These results suggest that end-to-end research automation can serve not only as a performance optimizer, but also as a new form of research infrastructure that reduces repetitive experimental burden and helps redirect human attention toward higher-level scientific creativity.
△ Less
Submitted 25 May, 2026; v1 submitted 7 April, 2026;
originally announced April 2026.
-
DQC1-completeness of normalized trace estimation for functions of log-local Hamiltonians
Authors:
Zhengfeng Ji,
Tongyang Li,
Changpeng Shao,
Xinzhao Wang,
Yuxin Zhang
Abstract:
We study the computational complexity of estimating the normalized trace $2^{-n}Tr[f(A)]$ for a log-local Hamiltonian $A$ acting on $n$ qubits. This problem arises naturally in the DQC1 model, yet its complexity is only understood for a limited class of functions $f(x)$.
We show that if $f(x)$ is a continuous function with approximate degree $Ω({\rm poly}(n))$, then estimating $2^{-n}Tr[f(A)]$ u…
▽ More
We study the computational complexity of estimating the normalized trace $2^{-n}Tr[f(A)]$ for a log-local Hamiltonian $A$ acting on $n$ qubits. This problem arises naturally in the DQC1 model, yet its complexity is only understood for a limited class of functions $f(x)$.
We show that if $f(x)$ is a continuous function with approximate degree $Ω({\rm poly}(n))$, then estimating $2^{-n}Tr[f(A)]$ up to constant additive error is DQC1-complete, under a technical condition on the polynomial approximation error of $f(x)$. This condition holds for a broad class of functions, including exponentials, trigonometric functions, logarithms, and inverse-type functions. We further prove that when $A$ is sparse, the classical query complexity of this problem is exponential in the approximate degree, assuming a conjectured lower bound for a trace variant of the $k$-Forrelation problem in the DQC1 query model. Together, these results identify the approximate degree as the key parameter governing the complexity of normalized trace estimation: it characterizes both the quantum complexity (via efficient DQC1 algorithms) and, conditionally, the classical hardness, yielding an exponential quantum-classical separation. Our proof develops a unified framework that cleanly combines circuit-to-Hamiltonian constructions, periodic Jacobi operators, and tools from polynomial approximation theory, including the Chebyshev equioscillation theorem.
△ Less
Submitted 1 April, 2026;
originally announced April 2026.
-
Testing General Relativity on Galactic Scales via DESI-BAO and Strong Lensing: Circumventing Assumptions on the Hubble Constant, Sound Horizon, and Dark Energy
Authors:
Hengyu Wu,
Tonghua Liu,
Chenggang Shao
Abstract:
We present a cosmological model-independent framework for testing general relativity (GR) on galactic scales by combining baryon acoustic oscillation (BAO) angular scale measurements with 120 galaxy-scale strong gravitational lensing systems. Using artificial neural networks (ANNs) and cubic spline reconstruction, we reconstruct the BAO angular scale from SDSS, BOSS, eBOSS, and DESI Data Release 2…
▽ More
We present a cosmological model-independent framework for testing general relativity (GR) on galactic scales by combining baryon acoustic oscillation (BAO) angular scale measurements with 120 galaxy-scale strong gravitational lensing systems. Using artificial neural networks (ANNs) and cubic spline reconstruction, we reconstruct the BAO angular scale from SDSS, BOSS, eBOSS, and DESI Data Release 2 (DR2), and infer the angular diameter distances to lenses and sources. Crucially, All the quantities used in the GR test are derived from observations and are independent of cosmological parameters such as the Hubble constant, the sound horizon, or the dark energy equation of state, minimizing potential biases from model-dependent distance priors. These distances are then incorporated into the strong lensing likelihood to constrain the parameterized post-Newtonian (PPN) parameter $γ_{\rm PPN}$ under two lens mass models: a constant-density-slope model ($P_1$) and a redshift-evolving model ($P_2$). For the $P_1$ model, the ANN reconstruction yields $γ_{\rm PPN} = 1.102^{+0.148}_{-0.125}$, consistent with GR at $1σ$ confidence level, while the cubic spline gives $γ_{\rm PPN} = 1.150^{+0.139}_{-0.118}$, consistent with GR at $2σ$ confidence level. For the $P_2$ model, the ANN reconstruction gives $γ_{\rm PPN} = 1.315^{+0.181}_{-0.155}$, compatible with GR at $2σ$, while the spline gives $γ_{\rm PPN} = 1.485^{+0.193}_{-0.168}$, showing mild tension at $\sim2.5σ$. The constraints exhibit a clear dependence on the adopted lens mass model, underscoring the critical role of lens modeling. No significant correlation is observed between $γ_{\rm PPN}$ and the Einstein radius. Overall, current galaxy-scale observations are consistent with GR, providing no evidence for deviations from Einstein's theory on kiloparsec scales.
△ Less
Submitted 22 March, 2026;
originally announced March 2026.
-
A Unified Hierarchical Multi-Task Multi-Fidelity Framework for Data-Efficient Surrogate Modeling in Manufacturing
Authors:
Manan Mehta,
Zhiqiao Dong,
Yuhang Yang,
Chenhui Shao
Abstract:
Surrogate modeling is an essential data-driven technique for quantifying relationships between input variables and system responses in manufacturing and engineering systems. Two major challenges limit its effectiveness: (1) large data requirements for learning complex nonlinear relationships, and (2) heterogeneous data collected from sources with varying fidelity levels. Multi-task learning (MTL)…
▽ More
Surrogate modeling is an essential data-driven technique for quantifying relationships between input variables and system responses in manufacturing and engineering systems. Two major challenges limit its effectiveness: (1) large data requirements for learning complex nonlinear relationships, and (2) heterogeneous data collected from sources with varying fidelity levels. Multi-task learning (MTL) addresses the first challenge by enabling information sharing across related processes, while multi-fidelity modeling addresses the second by accounting for fidelity-dependent uncertainty. However, existing approaches typically address these challenges separately, and no unified framework simultaneously leverages inter-task similarity and fidelity-dependent data characteristics. This paper develops a novel hierarchical multi-task multi-fidelity (H-MT-MF) framework for Gaussian process-based surrogate modeling. The proposed framework decomposes each task's response into a task-specific global trend and a residual local variability component that is jointly learned across tasks using a hierarchical Bayesian formulation. The framework accommodates an arbitrary number of tasks, design points, and fidelity levels while providing predictive uncertainty quantification. We demonstrate the effectiveness of the proposed method using a 1D synthetic example and a real-world engine surface shape prediction case study. Compared to (1) a state-of-the-art MTL model that does not account for fidelity information and (2) a stochastic kriging model that learns tasks independently, the proposed approach improves prediction accuracy by up to 19% and 23%, respectively. The H-MT-MF framework provides a general and extensible solution for surrogate modeling in manufacturing systems characterized by heterogeneous data sources.
△ Less
Submitted 10 March, 2026;
originally announced March 2026.
-
PanoAffordanceNet: Towards Holistic Affordance Grounding in 360° Indoor Environments
Authors:
Guoliang Zhu,
Wanjun Jia,
Caoyang Shao,
Yuheng Zhang,
Zhiyong Li,
Kailun Yang
Abstract:
Global perception is essential for embodied agents in 360° spaces, yet current affordance grounding remains largely object-centric and restricted to perspective views. To bridge this gap, we introduce a novel task: Holistic Affordance Grounding in 360° Indoor Environments. This task faces unique challenges, including severe geometric distortions from Equirectangular Projection (ERP), semantic disp…
▽ More
Global perception is essential for embodied agents in 360° spaces, yet current affordance grounding remains largely object-centric and restricted to perspective views. To bridge this gap, we introduce a novel task: Holistic Affordance Grounding in 360° Indoor Environments. This task faces unique challenges, including severe geometric distortions from Equirectangular Projection (ERP), semantic dispersion, and cross-scale alignment difficulties. We propose PanoAffordanceNet, an end-to-end framework featuring a Distortion-Aware Spectral Modulator (DASM) for latitude-dependent calibration and an Omni-Spherical Densification Head (OSDH) to restore topological continuity from sparse activations. By integrating multi-level constraints comprising pixel-wise, distributional, and region-text contrastive objectives, our framework effectively suppresses semantic drift under low supervision. Furthermore, we construct 360-AGD, the first high-quality panoramic affordance grounding dataset. Extensive experiments demonstrate that PanoAffordanceNet significantly outperforms existing methods, establishing a solid baseline for scene-level perception in embodied intelligence. The source code and benchmark dataset will be made publicly available at https://github.com/GL-ZHU925/PanoAffordanceNet.
△ Less
Submitted 16 July, 2026; v1 submitted 10 March, 2026;
originally announced March 2026.
-
Can Oscillatory and Persistent Nonlinearities Be Bridged in Black Hole Ringdown?
Authors:
Jun-Xi Shi,
Zhen-Tao He,
Jiageng Jiao,
Jing-Qi Lai,
Caiying Shao,
Yu Tian,
Hongbao Zhang
Abstract:
Quadratic quasinormal modes (QQNMs) and Christodoulou memory effect are key nonlinear phenomena in gravitational wave physics. QQNMs characterize the near zone nonlinear response of a perturbed black hole, whereas the memory effect is a nonlinear remnant imprinted at null infinity by outgoing radiation. This naturally raises the question of whether and in what sense the two can be bridged. We show…
▽ More
Quadratic quasinormal modes (QQNMs) and Christodoulou memory effect are key nonlinear phenomena in gravitational wave physics. QQNMs characterize the near zone nonlinear response of a perturbed black hole, whereas the memory effect is a nonlinear remnant imprinted at null infinity by outgoing radiation. This naturally raises the question of whether and in what sense the two can be bridged. We show that they are related through bridge coefficients which depend primarily on remnant black hole parameters during ringdown. Future space-based gravitational wave detectors can probe this relation. These results provide a new avenue for testing gravity and a fresh perspective on the nonlinear regime of general relativity.
△ Less
Submitted 19 May, 2026; v1 submitted 8 March, 2026;
originally announced March 2026.
-
The Impact of Dark Matter on Gravitational Wave Detection by Space-based Interferometers
Authors:
Yuezhe Chen,
Pan-Pan Wang,
Bo Wang,
Rui Luo,
Cheng-Gang Shao
Abstract:
The existence of dark matter is supported by multiple astrophysical observations, yet its particle nature remains unknown. The development of gravitational wave astronomy, especially with future space-based detectors such as LISA, provides new opportunities to study the interactions between dark matter and compact-object systems. This review summarizes the main dark matter candidates and their mac…
▽ More
The existence of dark matter is supported by multiple astrophysical observations, yet its particle nature remains unknown. The development of gravitational wave astronomy, especially with future space-based detectors such as LISA, provides new opportunities to study the interactions between dark matter and compact-object systems. This review summarizes the main dark matter candidates and their macroscopic distributions, and highlights three mechanisms through which dark matter can affect gravitational wave observations: (1) modifications to compact-object orbits and the dynamics of systems such as extreme mass-ratio inspirals, including dark matter spikes, dynamical friction, and potential perturbations; (2) gravitational lensing effects induced by the spatial distribution of dark matter, altering waveform amplitudes and phases; and (3) direct couplings between ultralight dark matter fields and detectors. As low-frequency gravitational wave detection techniques are proposed and continue to develop, these effects may offer a novel avenue for probing the properties of dark matter, and combining precise waveform modeling with multi-messenger observations could reveal insights into its microscopic structure.
△ Less
Submitted 7 March, 2026;
originally announced March 2026.