-
A Real-Time Tsetlin Machine-based Non-intrusive Load Monitoring System on MCUs
Authors:
Tianhang Tan,
Han Wu,
Tousif Rahman,
Shengyu Duan,
Alex Yakovlev,
Rishad Shafik
Abstract:
Non-Intrusive Load Monitoring (NILM) systems estimate individual appliance energy consumption from a single aggregate meter, without requiring separate sensors for each device. By installing a single meter that measures a building's total electricity consumption, NILM algorithms can determine the active status of each appliance. However, traditional NILM systems use computationally intensive optim…
▽ More
Non-Intrusive Load Monitoring (NILM) systems estimate individual appliance energy consumption from a single aggregate meter, without requiring separate sensors for each device. By installing a single meter that measures a building's total electricity consumption, NILM algorithms can determine the active status of each appliance. However, traditional NILM systems use computationally intensive optimization algorithms to process offline data, limiting their capability for on-device deployment, where sensitive household data must be processed locally. This paper proposes a Tsetlin Machine (TM)-based NILM framework, targeting real-time applications on resource-constrained microcontrollers (MCUs), enabling privacy-preserving edge deployment. The problem is reformulated as a classification task, and the proposed approach achieves an average precision of 90% and recall of 96% for two-appliance classification, and 77% precision and 80% recall for four appliances on the REDD dataset. The trained model occupies only 18 KB of flash memory and achieves an inference latency of 0.43 ms on an ESP32, demonstrating its suitability for embedded NILM applications on MCUs.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
Beam-Tracing-Based Quantitative Reconstruction of Density Fluctuations in QUEST Using Doppler Backscattering
Authors:
T. Kinoshita,
T. Tokuzawa,
V. H. Hall-Chen,
Y. T. Tan,
T. Ido,
H. Idei,
R. Ikezoe,
K. Hanada,
M. Hasegawa,
T. Onchi,
Y. Peng
Abstract:
A three-channel X-/Ku-band Doppler backscattering (DBS) system has been developed and installed on QUEST for turbulence and electric-field measurements. In spherical tokamaks, the large magnetic-field pitch angle increases the geometric mismatch between the probing beam wave vector and the local magnetic-field vector, reducing the effective perpendicular projection and resulting in a systematic un…
▽ More
A three-channel X-/Ku-band Doppler backscattering (DBS) system has been developed and installed on QUEST for turbulence and electric-field measurements. In spherical tokamaks, the large magnetic-field pitch angle increases the geometric mismatch between the probing beam wave vector and the local magnetic-field vector, reducing the effective perpendicular projection and resulting in a systematic underestimation of the measured scattering intensity. In addition, in QUEST, where low plasma density requires a low-frequency probe beam, beam propagation effects become increasingly significant, further complicating the interpretation of the measured DBS power in terms of local density fluctuation amplitude. To address these issues, a quantitative correction methodology based on the synthetic DBS code SCOTTY was established. All relevant diagnostic response effects were evaluated using SCOTTY along ray trajectories, yielding a correction factor for reconstructing the local turbulence amplitude from the measured scattering signal. The correction factor exhibits strong spatial and frequency dependence, varying by up to an order of magnitude between the plasma core and edge regions, highlighting the necessity of frequency-dependent corrections. By applying the derived correction factor to experimental measurements, quantitative density fluctuation amplitudes were reconstructed from the detected scattering signals. Evaluation of the fluctuation amplitude indicates enhanced turbulence activity in the plasma edge region, where a finite negative radial electric field is inferred. This work demonstrates the first quantitative turbulence evaluation using low-frequency X-/Ku-band DBS measurements in QUEST and establishes a framework for quantitative DBS analysis in spherical tokamaks.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
UC-VLM: Consistency-Driven Learning for AI-Generated Image Detection with Vision-Language Large Models
Authors:
Lei Tan,
Shuwei Li,
Mohan Kankanhalli,
Robby T. Tan
Abstract:
Vision-Language Large Models (VLLMs) are promising for AI-generated image (AIGI) detection because they can produce both a prediction and a natural-language output. However, most existing VLLM-based detectors primarily fine-tune the language side while giving limited attention to low-level visual forensic cues. They also often depend on manually crafted prompts or human-annotated rationales, which…
▽ More
Vision-Language Large Models (VLLMs) are promising for AI-generated image (AIGI) detection because they can produce both a prediction and a natural-language output. However, most existing VLLM-based detectors primarily fine-tune the language side while giving limited attention to low-level visual forensic cues. They also often depend on manually crafted prompts or human-annotated rationales, which limits scalability.We present UC-VLM, a unified multi-stage framework for AIGI detection that relies solely on binary supervision. UC-VLM first identifies effective instruction variants automatically. It then reuses the same binary label within a multi-stage training framework: (i) a visual discrimination objective that strengthens sensitivity to non-semantic forensic cues, and (ii) a label-conditioned generation objective that uses the binary label to supervise textual outputs. This design turns weak binary supervision into a shared supervision signal for both the visual pathway and the language output. Our key novelty is a unified multi-stage binary-supervised framework that consistently reuses the same authenticity labels for visual adaptation and label-conditioned text generation, while leveraging automatically optimized instructions to reduce prompt sensitivity without requiring human-written rationales or hand-crafted prompts.Experiments show that UC-VLM achieves 96.1% average accuracy on GenImage, exceeding the strongest prior result by 4.6%, and obtains 69.6% / 77.9% accuracy on Chameleon under ProGAN / SDV1.4 training, surpassing the best baseline by 11.2% / 15.3%, respectively.
△ Less
Submitted 15 August, 2026;
originally announced August 2026.
-
Skyrmion Fractional Chern Insulator: An Intrinsically Multiband Route to Fractionalization in Rhombohedral Graphene
Authors:
Julian May-Mann,
Tixuan Tan,
Patrick J. Ledwith,
Zhengyan Darius Shi,
Trithep Devakul
Abstract:
We propose an unconventional microscopic origin for the fractional quantum anomalous Hall (FQAH) effect in rhombohedral graphene moiré superlattices: skyrmion fractionalization. We view the state at filling $ν<1$ as a metal of skyrmion vacancies, charge $+e$ objects formed by removing layer-pseudospin skyrmions from the interaction-generated skyrmion lattice Chern insulator at $ν=1$. These vacanci…
▽ More
We propose an unconventional microscopic origin for the fractional quantum anomalous Hall (FQAH) effect in rhombohedral graphene moiré superlattices: skyrmion fractionalization. We view the state at filling $ν<1$ as a metal of skyrmion vacancies, charge $+e$ objects formed by removing layer-pseudospin skyrmions from the interaction-generated skyrmion lattice Chern insulator at $ν=1$. These vacancies are intrinsically multiband degrees of freedom, absent in single Chern band-projected studies. Building on a recently proposed ideal limit, we first develop an effective field theory showing that skyrmion vacancies can themselves fractionalize, thereby inducing charge fractionalization. Focusing on $ν=\frac{2}{3}$, we then construct explicit variational trial wavefunctions for the resulting skyrmion fractional Chern insulator and provide numerical evidence, together with general arguments, showing that this process is energetically favored. Our results establish a realistic route to the FQAH that does not rely on a partially filled Chern band, but instead arises from fractionalization of collective pseudospin textures.
△ Less
Submitted 14 August, 2026;
originally announced August 2026.
-
Revisiting Overestimation Bias Problem of Q-learning: Settling Large Discrete Action Space via Action Intersection
Authors:
Pu Li,
Tao Tan,
Hong Xie,
Xiaoyu Shi,
Mingsheng Shang
Abstract:
This paper considers the overestimation bias problem of Q-learning in the setting of a large action space, for the purpose of relieving the bottleneck of existing methods. We find that the large action space increases the randomness in Q-value estimation. The randomness makes two paradigms that drive the major literature on the overestimation problem have their own bottlenecks: the coupling paradi…
▽ More
This paper considers the overestimation bias problem of Q-learning in the setting of a large action space, for the purpose of relieving the bottleneck of existing methods. We find that the large action space increases the randomness in Q-value estimation. The randomness makes two paradigms that drive the major literature on the overestimation problem have their own bottlenecks: the coupling paradigm, i.e., the optimal action and its Q-value are estimated with the same Q-function, always has a positive bias. This is because randomness leads to some actions having abnormally high estimated values than their true values, and the coupling methods prefer these actions. The decoupling paradigm, i.e., the optimal action and its Q-value are estimated with two independent Q-functions, always has a negative bias. This is because randomness increases the estimation gap between the two independent Q-tables for the same action. This paper shows that action intersection can be a simple yet powerful strategy to relieve these bottlenecks. The action intersection strategy enables semi-decoupling via two designs: (1) it allows two Q-functions to share a certain fraction of trajectory data; (2) if a data sample is shared, each Q-function is updated using the coupling paradigm; otherwise, using the decoupling paradigm. Two properties make the action intersection strategy powerful: (1) attaining a large bias range, i.e., varying the data sharing fraction, the estimation bias varies from underestimating to overestimating; (2) fine granularity: the action intersection size can be made arbitrarily finer to enable finer control. We consider two experiment settings, i.e., tabular and deep RL, deep RL experiments show that our method outperforms several SOTA baselines drastically; tabular experiments reveal why our method can achieve superior performance.
△ Less
Submitted 13 August, 2026;
originally announced August 2026.
-
Population-Scalable Multi-Agent World Modeling
Authors:
Renjie Zhao,
Yuxiang Wu,
Mingyu Zhang,
Jiaxin Li,
Sisi Li,
Yimin Sheng,
Tianxi Tan,
Zhenkai Zhang,
Jianyi Zhu,
Yong-Lu Li
Abstract:
World models have recently achieved impressive progress in visual prediction and interactive generation, but extending them to multi-agent environments introduces a fundamental scalability challenge. Existing methods generally assume a fixed number of agents during training and inference, which ties the model to a pre-determined agent population and limits inference-time scalability. Our key insig…
▽ More
World models have recently achieved impressive progress in visual prediction and interactive generation, but extending them to multi-agent environments introduces a fundamental scalability challenge. Existing methods generally assume a fixed number of agents during training and inference, which ties the model to a pre-determined agent population and limits inference-time scalability. Our key insight is that cross-view consistency should arise from a shared world state whose evolution does not assume a predefined number of agents, while agent-specific observations should be generated by querying this state through a unified rendering interface. Based on this insight, we propose Khora, a scalable multi-agent world model that supports inference-time expansion to arbitrary numbers of agents without retraining. Our framework decouples world-state evolution from visual rendering and introduces a population-agnostic rendering mechanism for incorporating other agent information. This design maintains cross-view consistency through the shared world state rather than through dense interactions among observation streams inside the expensive video generator, enabling approximately linear practical scaling with the number of queried views. Qualitative experiments demonstrate that our approach generalizes to unseen numbers of agents while maintaining visual quality and multi-agent consistency. We further implement a real-time interactive system to demonstrate scalable open-world simulation.
△ Less
Submitted 17 August, 2026; v1 submitted 9 August, 2026;
originally announced August 2026.
-
Who Gets Access? Global Region and Academic Status Bias in AI-Generated Academic Gatekeeping Scenarios
Authors:
Nouar AlDahoul,
Hezerul Abdul Karim,
Myles Joshua Toledo Tan
Abstract:
Equitable access to scientific knowledge often depends on informal gatekeeping decisions, particularly when resources such as paywalled articles, datasets, or professional materials such as curriculum vitae (CV) must be shared selectively. We introduce a controlled simulation framework in which large language model (LLM)-based professors must grant access to only one requestor. Across prompts, req…
▽ More
Equitable access to scientific knowledge often depends on informal gatekeeping decisions, particularly when resources such as paywalled articles, datasets, or professional materials such as curriculum vitae (CV) must be shared selectively. We introduce a controlled simulation framework in which large language model (LLM)-based professors must grant access to only one requestor. Across prompts, requesters vary systematically by global region (Global North vs. Global South) and academic seniority (undergraduate student, PhD candidate, postdoctoral researcher, and tenured professor), while all other factors remain constant. Across varying evaluation scenarios, LLMs exhibit contrasting academic status biases, with some prioritizing PhD candidates, while others favor tenured professors. However, when global regions differ, a distinct divergence emerges based on model architecture: while many frontier LLMs systematically favor requesters from the Global South due to pro-equity bias that results from equity-focused safety alignment, open-weight and small models frequently flip this preference to favor the Global North, reflecting the global region bias and unaligned geographic distribution of their baseline pre-training data. Our findings highlight how normative assumptions embedded in model behavior can shape gatekeeping decisions, underscoring the importance of auditing AI systems for fairness and value alignment.
△ Less
Submitted 27 June, 2026;
originally announced August 2026.
-
BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation
Authors:
Peiyan Li,
Yuze Zhu,
Yixiang Chen,
Qisen Ma,
Yuan Xu,
Jiabing Yang,
He Guan,
Yan Huang,
Hongtao Wu,
Xiao Ma,
Tao Kong,
Liang Wang,
Tieniu Tan
Abstract:
Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existing 3D VLA methods remain data-hungry, exhibit limited generalization under distribution shifts, and lack explicit memory of past observations. These limitations hinder their application to data-scarce, open-world, and me…
▽ More
Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existing 3D VLA methods remain data-hungry, exhibit limited generalization under distribution shifts, and lack explicit memory of past observations. These limitations hinder their application to data-scarce, open-world, and memory-dependent manipulation scenarios. Our previous work, BridgeVLA, improves data efficiency and generalization by preserving the input--output alignment of a pre-trained VLM during 3D action learning: raw point clouds are projected into multi-view images, and intermediate heatmaps are predicted before generating robot actions. In this work, we develop BridgeVLA++ by equipping BridgeVLA with a unified spatio-temporal memory architecture that models persistent spatial context and temporal interaction history. The resulting memory-augmented framework can reason over observation histories while preserving BridgeVLA's data efficiency and generalization capabilities. Extensive experiments show that our framework achieves strong performance on spatial manipulation tasks while exhibiting robust generalization. BridgeVLA++ further achieves state-of-the-art performance on two challenging memory-dependent manipulation benchmarks without sacrificing the data efficiency and generalization of the original BridgeVLA. In addition, BridgeVLA++ performs effectively in bimanual manipulation settings and is validated on an additional real-world robotic platform, demonstrating its scalability across tasks, environments, and robotic platforms. These results establish BridgeVLA++ as a unified 3D vision-language-action framework that simultaneously supports data-efficient learning, robust generalization, and effective memory-aware robot manipulation. Project website: https://bridgevla-plus.github.io/.
△ Less
Submitted 5 August, 2026;
originally announced August 2026.
-
GROW: Group-Relative Advantage-Weighted On-Policy Reinforcement Learning of Autoregressive-Diffusion Text-to-Speech model
Authors:
Guanrou Yang,
Tian Tan,
Qian Chen,
Ziyang Ma,
Yakun Song,
Zhikang Niu,
Qi Chen,
Wenming Tu,
Haitao Li,
Shan Yang,
Xie Chen
Abstract:
Reinforcement learning for flow-matching text-to-speech is complicated by deterministic ODE sampling: trajectory-level policy-gradient methods typically convert the ODE into an SDE and track per-step likelihood ratios, introducing stochastic perturbations and substantial overhead. We propose GROW, a group-relative advantage-weighted on-policy RL method that acts directly on the standard flow-match…
▽ More
Reinforcement learning for flow-matching text-to-speech is complicated by deterministic ODE sampling: trajectory-level policy-gradient methods typically convert the ODE into an SDE and track per-step likelihood ratios, introducing stochastic perturbations and substantial overhead. We propose GROW, a group-relative advantage-weighted on-policy RL method that acts directly on the standard flow-matching objective. For each prompt, GROW samples a group of on-policy utterances, separately standardizes intelligibility and speaker-similarity rewards within the group, and combines them to reweight flow-matching regression. A Wasserstein-2 velocity penalty anchors the updated model to a frozen pretrained reference. A group-mean reward baseline is introduced to convert reward weighting into advantage weighting. For strong pretrained TTS models with concentrated rewards, positive exponential weighting is dominated by reward-agnostic self-imitation, whereas a zero-mean signed advantage preserves effective within-group credit assignment. Instantiated on DiTAR and evaluated on LibriSpeech and Seed-TTS EN/ZH, GROW reduces average WER from 2.016 to 1.558 and raises speaker similarity from 0.676 to 0.715 while keeping UTMOS. With 10-NFE training rollouts and 32-NFE evaluation, GROW retains comparable performance while training 2.9x faster than 32-NFE DiTAR-GRPO. We will open-source complete GROW codes, faithful DiTAR reproduction, and all model checkpoints.
△ Less
Submitted 4 August, 2026;
originally announced August 2026.
-
WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
Authors:
Yuxue Yang,
Shuyao Shang,
Jiahe Wang,
Zitong Zhou,
Liang Tan,
Junhan Zeng,
Ruizhi Li,
Junyan Li,
Yu Liu,
Xiao Yang,
Yong Li,
Jun Zhu,
Hongsheng Li,
Tieniu Tan,
Lue Fan,
Zhaoxiang Zhang
Abstract:
Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated videos to the inherent reactivity of the worlds they depict: the ability to infer from the scene state how the world should react and to generate plausible consequences not explicitly described in the input. Yet existin…
▽ More
Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated videos to the inherent reactivity of the worlds they depict: the ability to infer from the scene state how the world should react and to generate plausible consequences not explicitly described in the input. Yet existing benchmarks mainly assess visual quality or explicit instruction fulfillment by checking whether requested actions and interaction outcomes are realized, leaving inherent reactivity underexamined. We introduce WorldExam, a hierarchical diagnostic benchmark spanning four levels: Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity. It comprises 1,474 cases across eight dedicated tasks and supports unified evaluation of camera-, action-, and language-driven model paradigms. The World Reactivity level evaluates scene-conditioned reactions and goal-directed behaviors beyond what is explicitly specified in the input. Evaluation of 20 representative models reveals a clear capability split. Camera-driven models excel at camera control, but their interfaces do not support dynamic interaction; action-driven models control subjects more precisely but often leave the world unresponsive; and language-driven models perform better on interaction but follow complex controls less faithfully. No model combines broad task coverage with consistently strong performance, showing that high visual quality and explicit instruction fulfillment do not guarantee inherent reactivity.
△ Less
Submitted 3 August, 2026;
originally announced August 2026.
-
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?
Authors:
Dong Yan,
Jian Liang,
Dapeng Hu,
Ran He,
Nicholas Jing Yuan,
Qi Zhang,
Tieniu Tan
Abstract:
Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. However, existing studies predominantly adopt independent evaluation. Consequently, the behavior of self-evolving agents in realistic streaming settings, where agents adapt to diverse and complex task streams, remains poorly understood. To address this gap, we introduce AgentStream, a…
▽ More
Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. However, existing studies predominantly adopt independent evaluation. Consequently, the behavior of self-evolving agents in realistic streaming settings, where agents adapt to diverse and complex task streams, remains poorly understood. To address this gap, we introduce AgentStream, a unified framework that evaluates self-evolving agents spanning diverse evolution components by organizing agentic benchmarks into a configurable task stream and instantiating the \texttt{Isolated}, \texttt{Sequential}, and \texttt{Interleaved} streaming scenarios at test time, which progressively vary the scope and domain composition of the stream. Over these scenarios, we combinatorially evaluate five representative self-evolving methods across three frontier foundation models, disentangling how model capability, method architecture, and streaming scenario jointly shape self-evolution. Our results show that self-evolution reliability varies across streaming scenarios, the benefit of self-evolution is gated by model capability and non-monotonic in model strength, and no single method dominates across models and scenarios. These findings offer concrete guidance for selecting self-evolving methods across models and streaming scenarios. Overall, we advocate that self-evolving agents should be evaluated under realistic task streams rather than isolated single-task settings.
△ Less
Submitted 31 July, 2026;
originally announced August 2026.
-
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL
Authors:
Ruiming Liang,
Yi Zhong,
Yizhen Yuan,
Yinan Zheng,
Tianyi Tan,
Tianyue Wang,
Haiyun Guo,
Jinqiao Wang,
Xianyuan Zhan
Abstract:
Modern large language models (LLMs) are expected not just to answer correctly, but to adapt their behavior to different human values and use cases. As a result, multi-reward reinforcement learning (RL) has become an increasingly important problem for LLMs, where each reward captures a different aspect of desired behavior. However, optimizing with multiple rewards suffers from a more severe alignme…
▽ More
Modern large language models (LLMs) are expected not just to answer correctly, but to adapt their behavior to different human values and use cases. As a result, multi-reward reinforcement learning (RL) has become an increasingly important problem for LLMs, where each reward captures a different aspect of desired behavior. However, optimizing with multiple rewards suffers from a more severe alignment tax issue, where different optimization objectives can trade off or even conflict with each other, leading to unstable and inefficient post-training. In this work, we propose PRISM, a new multi-reward RL framework built upon the idea of policy-space decomposition and composition. Instead of compositing different rewards, PRISM optimizes a set of standalone positive policies and a global negative policy. This alleviates the potential conflict during multi-reward policy optimization, while enabling controllability during inference by flexible policy composition. Experiments on scientific reasoning, tool-use reasoning, and helpfulness-safety alignment show that PRISM consistently outperforms existing multi-reward RL baselines, with extra controllability for inference-time preference control.
△ Less
Submitted 31 July, 2026;
originally announced July 2026.
-
PhiZero: A World Model Built Around Physical Language
Authors:
Shuyao Shang,
Yuqi Wang,
Ruopeng Gao,
Xu Chen,
Tieniu Tan,
Lue Fan,
Zhaoxiang Zhang
Abstract:
We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional visual predictors. Motivated by humans' ability to abstract predictive structure from visual experienc…
▽ More
We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional visual predictors. Motivated by humans' ability to abstract predictive structure from visual experience and organize it in natural language for explicit reasoning, we learn physical language from in-the-wild videos through self-supervision and use it to explicitly reason about how the physical world evolves. Accordingly, PhiZero adopts a reason-then-render paradigm: it first infers future world evolution as a physical-language sequence and then renders the inferred transitions into videos. Extensive experiments across generation and understanding benchmarks validate the ability of PhiZero to model physically coherent world evolution. We further show its potential for realistic and interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.
△ Less
Submitted 30 July, 2026;
originally announced July 2026.
-
DESI DR2 Results IV: Alcock-Paczyński Measurements from the Lyman Alpha Forest and Cosmological Constraints
Authors:
DESI Collaboration,
A. G. Adame,
J. Aguilar,
S. Ahlen,
O. Alves,
A. Anand,
U. Andrade,
E. Armengaud,
S. Avila,
A. Aviles,
P. Bansal,
A. Bault,
J. R. Bermejo-Climent,
F. Beutler,
D. Bianchi,
C. Blake,
S. Blasby,
M. Bonici,
S. Brieden,
A. Brodzeller,
D. Brooks,
A. Carnero Rosell,
K. Carrion,
L. Casas,
F. J. Castander
, et al. (130 additional authors not shown)
Abstract:
We present Alcock-Paczyński (AP) measurements from the full shape of Lyman-$α$ (Ly$α$) forest correlation functions measured from the second data release (DR2) of the Dark Energy Spectroscopic Instrument (DESI). Our measurements include information from the Ly$α$ forest auto-correlation and its cross-correlation with quasars. We constrain the AP effect with $1\%$ precision at an effective redshift…
▽ More
We present Alcock-Paczyński (AP) measurements from the full shape of Lyman-$α$ (Ly$α$) forest correlation functions measured from the second data release (DR2) of the Dark Energy Spectroscopic Instrument (DESI). Our measurements include information from the Ly$α$ forest auto-correlation and its cross-correlation with quasars. We constrain the AP effect with $1\%$ precision at an effective redshift $z_\mathrm{eff}=2.33$, which is twice as tight as the Baryon Acoustic Oscillation (BAO) constraint from the same data. When using the joint Ly$α$ AP and BAO results, we measure the ratios $D_\text{H}(z_\mathrm{eff})/r_\text{d}=8.600 \pm 0.066$ and $D_\text{M}(z_\mathrm{eff})/r_\text{d}=39.32 \pm 0.33$, where $D_\text{M}$ is the transverse comoving distance, $D_\text{H}$ is the Hubble distance, and $r_\text{d}$ is the sound horizon at the drag epoch. Assuming $Λ$CDM, Ly$α$ forest measurements combined with a nucleosynthesis prior produce a constraint on the Hubble constant $H_0=66.5\pm1.3\,\mathrm{km\,s^{-1}\,Mpc^{-1}}$. The Ly$α$ AP result corresponds to a matter fraction constraint $Ω_\text{m}=0.325\pm0.018$ in $Λ$CDM, which is $1.4σ$ higher than DESI BAO. This impacts the DESI results relative to the Cosmic Microwave Background (CMB), slightly reducing their discrepancy from $2.4σ$ to $2.2σ$. We present updated constraints on extended models using the joint DESI DR2 BAO and Ly$α$ forest full shape data, together with external data sets. When considering a time-evolving dark energy equation of state parametrized by $w_0$ and $w_a$, we find it is preferred over $Λ$CDM at $2.7σ$ for the combination of DESI and CMB data, and at $3.2σ$ when also including supernovae. With the new Ly$α$ AP measurement, DESI provides its most precise anchor for the expansion history at $z > 1$ in the matter-dominated Universe.
△ Less
Submitted 4 August, 2026; v1 submitted 29 July, 2026;
originally announced July 2026.
-
Observable Estimation in the Absence of Classical Verification
Authors:
Samantha V. Barron,
Bradley Mitchell,
Vinay Tripathi,
Francesco Grieco,
Ilan Rosen,
Francesca Pietracaprina,
Davide Materia,
Alireza Seif,
Darvin Wanisch,
Ramón L. Panadés-Barrueta,
Ewout van den Berg,
Jay-U Chung,
Andrew Eddins,
Sam Ferracin,
Guillermo García-Pérez,
John Goold,
Luke C. G. Govia,
Holger Haas,
Ian Hincks,
Jesse C. Hoke,
Zoë Holmes,
Su-un Lee,
Youngseok Kim,
Swarnadeep Majumder,
Sabrina Maniscalco
, et al. (23 additional authors not shown)
Abstract:
The predictive success of quantum mechanics underpins many areas of modern science, even as the exact simulation of large, interacting quantum systems remains beyond the reach of classical computation. This success has been enabled by the remarkable advancement of scalable numerical approximation methods, which often demonstrate practical accuracy despite the absence of formal guarantees. As quant…
▽ More
The predictive success of quantum mechanics underpins many areas of modern science, even as the exact simulation of large, interacting quantum systems remains beyond the reach of classical computation. This success has been enabled by the remarkable advancement of scalable numerical approximation methods, which often demonstrate practical accuracy despite the absence of formal guarantees. As quantum simulation pushes into regimes where these approximations struggle, a fundamental challenge arises: How can quantum outcomes be trusted when reliable classical benchmarks are unavailable? Here, we establish a framework for the independent validation of quantum estimates in this setting and present evidence that they provide the most credible result among several considered methods, in the absence of an immediately accessible ground-truth solution. We apply our framework to the semi-scrambling dynamics of a physical model that strains several leading classical simulation methods yet remains experimentally accessible, in part through our introduction of the \textit{operator Loschmidt echo}. We systematically design a series of experiments using quantum heuristics that, taken together, test the underlying assumptions and provide strong confidence in the observable estimates obtained from the quantum computer. We then show how this framework can be extended to place accuracy bounds on quantum estimates via careful characterization and manipulation of the device noise, transforming the problem of validating the observable estimation to validating the noise model. These results establish a route towards trusted quantum computation for scientific discovery, independent of classical verification.
△ Less
Submitted 28 July, 2026;
originally announced July 2026.
-
Screening with Product Mismatch
Authors:
Teck Yong Tan
Abstract:
A monopolist sells a product line whose variants are horizontally differentiated from the buyers' perspective but ordered by production cost. Buyers privately know their ideal product, and willingness to pay may be correlated with horizontal need. The seller screens buyers through product mismatch, and what she must screen determines whether mismatch creates or reduces information rent. When buyer…
▽ More
A monopolist sells a product line whose variants are horizontally differentiated from the buyers' perspective but ordered by production cost. Buyers privately know their ideal product, and willingness to pay may be correlated with horizontal need. The seller screens buyers through product mismatch, and what she must screen determines whether mismatch creates or reduces information rent. When buyers differ only in horizontal need, mismatch creates rent: the seller induces less mismatch, assigning served buyers products closer to their ideals than under the first best. When willingness to pay is correlated with horizontal need, mismatch instead reduces rent: the seller induces more mismatch, sells the basic product to buyers whose efficient products are advanced variants while excluding buyers better matched to it, and stronger horizontal differentiation can expand coverage and raise profit. Because mismatch is type-specific, optimal allocations are determined by individual rationality rather than by incentive compatibility alone.
△ Less
Submitted 23 July, 2026;
originally announced July 2026.
-
Profiling and Endogenous Valuation
Authors:
Anh Nguyen,
Teck Yong Tan
Abstract:
We study a monopolist facing a buyer whose valuation is determined by pre-trade investment. Before setting price, the seller observes a signal about the buyer's private investment cost (buyer profiling). Information that helps the seller extract surplus can also undermine the buyer's incentive to create it. We characterize the buyer-seller payoffs attainable across all possible profiling. On the P…
▽ More
We study a monopolist facing a buyer whose valuation is determined by pre-trade investment. Before setting price, the seller observes a signal about the buyer's private investment cost (buyer profiling). Information that helps the seller extract surplus can also undermine the buyer's incentive to create it. We characterize the buyer-seller payoffs attainable across all possible profiling. On the Pareto frontier, if investment increases, hold-up risk always raises the payoff that the seller captures faster than the surplus that the investment creates. Protecting buyer welfare therefore requires discouraging investment, even though investment is socially efficient.
△ Less
Submitted 21 July, 2026;
originally announced July 2026.
-
Fenced Citation-Context Retrieval for Case Law: Temporal Leakage and Degree Control Across Two Jurisdictions
Authors:
Yao Liu,
Tien-Ping Tan,
Zhilan Liu
Abstract:
Prior case retrieval (PCR) aims to identify the precedent cases relevant to the facts of a query case. Incoming citation context, the text with which later cases characterize a case when citing it, is a powerful relevance signal, yet it is typically evaluated without a temporal constraint, so the retriever is credited with citations made after the query. We introduce a temporally fenced retriever…
▽ More
Prior case retrieval (PCR) aims to identify the precedent cases relevant to the facts of a query case. Incoming citation context, the text with which later cases characterize a case when citing it, is a powerful relevance signal, yet it is typically evaluated without a temporal constraint, so the retriever is credited with citations made after the query. We introduce a temporally fenced retriever with no learned parameters that augments BM25 with incoming citation context restricted to citations predating the query, together with a temporal-admission decomposition that quantifies the phantom fraction: the share of a citation-context gain attributable to citations not known to predate the query. Experiments span two jurisdictions, U.S. federal (CLERC) and European (ECtHR-PCR) case law. On ECtHR-PCR, without any training, the fenced retriever outperforms a strong degree-controlled baseline across the full recall ladder, and a temporal-admission decomposition attributes 14.9% (validation) of an unfenced citation-context gain over BM25 to citations not known to predate the query. Citation-context retrieval must therefore be temporally fenced and degree-controlled before its reported gains can be interpreted.
△ Less
Submitted 2 August, 2026; v1 submitted 19 July, 2026;
originally announced July 2026.
-
Defect assignment of the clock site in $^{229}\text{Th:CaF}_2$
Authors:
Daniel A. Rehn,
Harry W. T. Morgan,
Harris E. Mason,
H. B. Tran Tan,
Ricky Elwell,
Igor M. Savukov,
Michael J. Martin,
Andrei Derevianko,
Eric R. Hudson
Abstract:
The performance of solid-state $^{229}\text{Th}$ nuclear clocks depends sensitively on the microscopic environment of the thorium nucleus in the host crystal. Here we reassess the dominant quadrupole-split thorium site in $^{229}\text{Th:CaF}_2$, which has been assigned to a thorium dimer in recent spectroscopic work. Thermodynamic estimates, density functional theory calculations, and electric-fi…
▽ More
The performance of solid-state $^{229}\text{Th}$ nuclear clocks depends sensitively on the microscopic environment of the thorium nucleus in the host crystal. Here we reassess the dominant quadrupole-split thorium site in $^{229}\text{Th:CaF}_2$, which has been assigned to a thorium dimer in recent spectroscopic work. Thermodynamic estimates, density functional theory calculations, and electric-field-gradient comparisons instead favor an isolated $\text{Th}^{4+}$ substitution on a $\text{Ca}^{2+}$ site charge-compensated by two nearby fluorine interstitials in a relaxed $90^\circ$ motif. The same calculation identifies a higher-energy mixed-shell interstitial motif as a plausible minor site. The clock-active quadrupole-split site is therefore controlled by local fluoride compensation rather than unavoidable thorium aggregation. This defect assignment also has implications for achievable linewidths and provides a microscopic basis for reducing broadening in solid-state nuclear clocks.
△ Less
Submitted 13 July, 2026;
originally announced July 2026.
-
Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO
Authors:
Xin Zhang,
Haochen Wang,
Yikang Zhou,
Jason Li,
Robby T. Tan
Abstract:
This paper introduces Actor as Its Own Critic, a unified reinforcement learning framework, Cycle Group Relative Policy Optimization (CycleGRPO), that jointly optimizes region understanding and localization for Multimodal Large Language Models (MLLMs). Unlike existing separate pipelines, we leverage the inherent duality between the two tasks to construct a self-evaluating reinforcement learning par…
▽ More
This paper introduces Actor as Its Own Critic, a unified reinforcement learning framework, Cycle Group Relative Policy Optimization (CycleGRPO), that jointly optimizes region understanding and localization for Multimodal Large Language Models (MLLMs). Unlike existing separate pipelines, we leverage the inherent duality between the two tasks to construct a self-evaluating reinforcement learning paradigm: "region $\to$ text $\to$ region''. Specifically, a single MLLM first acts as the actor to generate region captions, then immediately transitions to a critic to ground its generated text back in the spatial domain. Therefore, CycleGRPO requires only region inputs, e.g., masks or bounding boxes, entirely bypassing the need for textual ground truths. A quality-aware token-level cycle-consistency reward is employed to assess the semantic discriminability of text captions via their physical localization accuracy. Empirically, built upon SAMTok, our CycleGRPO framework successfully bootstraps both capabilities simultaneously. Without any task-specific fine-tuning, the framework yields consistent performance gains across a wide range of benchmarks, including region captioning, region VQA, grounded dialogue, and referring segmentation. Overall, CycleGRPO offers a straightforward and scalable way to advance pixel-level capabilities in MLLMs. Code and models are released at https://github.com/devinxzhang/CycleGRPO.
△ Less
Submitted 13 July, 2026;
originally announced July 2026.
-
Comparison of DBS measurements of turbulence spectra by vertical-displacement and poloidal-angle scans using the Scotty synthetic diagnostic
Authors:
Y. T. Tan,
V. Hall-Chen,
T. Rhodes
Abstract:
Doppler backscattering (DBS) measures electron density fluctuations. The measured wavenumber is typically varied by changing the probe beam poloidal launch angle. As most DBS systems are unable to steer during a shot, the shot is repeated and the poloidal angle is changed intershot. An alternative method is to keep the DBS launch angle fixed and move the plasma up and down instead, enabling a rang…
▽ More
Doppler backscattering (DBS) measures electron density fluctuations. The measured wavenumber is typically varied by changing the probe beam poloidal launch angle. As most DBS systems are unable to steer during a shot, the shot is repeated and the poloidal angle is changed intershot. An alternative method is to keep the DBS launch angle fixed and move the plasma up and down instead, enabling a range of wavenumbers to be measured within a single shot. We call this the bouncing ball method. We use the Scotty synthetic diagnostic (Hall-Chen, 2022) to evaluate the similarities and differences between these two approaches. Both approaches are capable of measuring a similar range of fluctuation wavenumbers as well as radial locations. However, the vertical-displacement scan has a larger range of measured poloidal locations than the poloidal-angle scan. Using the same synthetic turbulence spectrum as input, we show that the two approaches are expected to have different backscattered powers due to different instrumentation functions. When mismatch attenuation is accounted for via a synthetic diagnostic, the vertical-displacement scan can provide comparable radial and wavenumber coverage while reducing reliance on shot-to-shot repeatability. These results establish vertical-displacement scan as a practical route to single-shot DBS wavenumber spectra.
△ Less
Submitted 10 July, 2026;
originally announced July 2026.
-
Ab initio calculations of $^{229}$Th band-to-band internal conversion rate in $^{229}$ThO$_2$
Authors:
Udeshika C. Perera,
H. B. Tran Tan,
H. W. T. Morgan,
Eric Hudson,
Daniel A. Rehn,
Andrei Derevianko
Abstract:
We present an ab initio calculation of the band-to-band internal-conversion rate of the $\hbarω_{\rm nuc} \approx 8.35$ eV isomeric transition in $^{229}$ThO$_2$. Because the nuclear transition energy exceeds the electronic band gap of ThO$_2$, the isomer can decay nonradiatively by resonantly promoting a valence electron into the conduction band. We formulate this process as a Brillouin-zone sum…
▽ More
We present an ab initio calculation of the band-to-band internal-conversion rate of the $\hbarω_{\rm nuc} \approx 8.35$ eV isomeric transition in $^{229}$ThO$_2$. Because the nuclear transition energy exceeds the electronic band gap of ThO$_2$, the isomer can decay nonradiatively by resonantly promoting a valence electron into the conduction band. We formulate this process as a Brillouin-zone sum over vertical interband transitions weighted by local Th-centered hyperfine matrix elements, which are evaluated directly from all-electron full-potential linearized augmented-plane-wave Bloch spinors. A finite nuclear magnetization model is included to regularize the short-range hyperfine interaction and to account for the Bohr-Weisskopf effect. After applying scissor shifts to span the experimentally reported ThO$_2$ band gaps, we find calculated internal-conversion lifetimes in the range of $1-16~μ{\rm s}$. The lifetime increases strongly as the band gap approaches $ω_{\rm nuc}$ because the resonant interband phase space at the nuclear transition energy is reduced. For the larger reported ThO$_2$ gaps, the calculated lifetime is comparable to the measured conversion-electron Mössbauer lifetime [Nature 648, 300 (2025)]. Our analysis implies that choosing solid-state hosts with band-gap values slightly lower than $ω_{\rm nuc}$ can optimize solid-state nuclear clock performance with internal-conversion electron readout.
△ Less
Submitted 9 July, 2026;
originally announced July 2026.
-
Executable verification through formalized expert reasoning in astronomical spectroscopy
Authors:
Haosong Wang,
Ting Tan,
Ji Yao,
Jiajun Zhang,
Qian Zheng,
Christophe Yeche,
Jean-Paul Kneib,
Huanyuan Shan
Abstract:
Artificial intelligence has reshaped scientific prediction, but scientific verification remains a human bottleneck. Automated systems can map observations to labels, parameters or hypotheses, yet scientific conclusions require evidence, must satisfy physical consistency, and need explicit testing of alternatives before a decision is made. Here we introduce FORMA (Formalized Observational Reasoning…
▽ More
Artificial intelligence has reshaped scientific prediction, but scientific verification remains a human bottleneck. Automated systems can map observations to labels, parameters or hypotheses, yet scientific conclusions require evidence, must satisfy physical consistency, and need explicit testing of alternatives before a decision is made. Here we introduce FORMA (Formalized Observational Reasoning with Auditable Decisions), an executable verification protocol that reconstructs expert reasoning into a workflow: it extracts evidence, generates hypotheses under physical constraints, tests alternatives, and performs auditable consistency checks. Unlike prediction or post-hoc interpretability, executable verification records and tests the evidential path leading to a decision. Astronomical spectroscopy provides a natural testbed, because ambiguous survey spectra are still adjudicated by expert visual inspection. Applied to the Dark Energy Spectroscopic Instrument (DESI) visual inspection catalogue, FORMA combines template-fitting candidate redshifts, spectral evidence extraction and physical audit into an auditable credibility score. A medium-or-higher credibility threshold identifies $331$ definite predictions with $95.5\%$ binary agreement with expert-adjudicated classes, while increasing credibility is associated with improved redshift consistency and higher classification reliability. These results show that automated inference can be coupled to explicit verification, allowing candidate outputs to be evaluated before they enter scientific use.
△ Less
Submitted 7 July, 2026;
originally announced July 2026.
-
USE: A Unified Self-Ensembling Framework for Test-Time Prompt Tuning
Authors:
Siru Jiang,
Jian Liang,
Ran He,
Tieniu Tan
Abstract:
Test-time adaptation (TTA) has emerged as a popular paradigm for improving the performance of vision-language models (e.g., CLIP) on downstream tasks. Among existing CLIP-based TTA methods, Test-Time Prompt Tuning (TPT) is a pioneering work that optimizes textual prompts using multiple test-time augmentations and remains a strong baseline to date. In this work, we revisit TPT and reveal that its o…
▽ More
Test-time adaptation (TTA) has emerged as a popular paradigm for improving the performance of vision-language models (e.g., CLIP) on downstream tasks. Among existing CLIP-based TTA methods, Test-Time Prompt Tuning (TPT) is a pioneering work that optimizes textual prompts using multiple test-time augmentations and remains a strong baseline to date. In this work, we revisit TPT and reveal that its optimization can be interpreted as implicitly learning from self-generated pseudo labels. Building on this perspective, we propose a unified self-ensembling framework (USE) that ensures consistency between the optimization and inference stages. During optimization, we introduce a simple yet effective self-ensembling (SE) strategy that emphasizes the test image itself over its augmented views adaptively to obtain more reliable pseudo labels. To fully exploit the potential of augmentations, we further apply the same strategy at inference time, unifying the objectives of both stages. Notably, SE can also act as a lightweight optimization-free TTA method. Extensive experiments across multiple datasets demonstrate that SE and USE outperform their counterparts, respectively. Furthermore, SE yields consistent performance gains when integrated with existing TTA methods. The code is available at https://github.com/sirujiang/USE.
△ Less
Submitted 4 July, 2026;
originally announced July 2026.
-
Token-Based Affordance Grounding with Large Vision-Language Models
Authors:
Seung Il Lee,
Qinqian Lei,
Daguang Xu,
Dong Yang,
Robby T. Tan,
Yixin Chen,
Bo Wang
Abstract:
Affordance grounding aims to localize image regions that support a specific action, serving as a core capability for physical intelligence and embodied perception. Previous studies have primarily relied on weakly supervised learning with action labels from exocentric images. However, these methods often struggle with visually ambiguous exocentric images containing co-occurring actions; moreover, t…
▽ More
Affordance grounding aims to localize image regions that support a specific action, serving as a core capability for physical intelligence and embodied perception. Previous studies have primarily relied on weakly supervised learning with action labels from exocentric images. However, these methods often struggle with visually ambiguous exocentric images containing co-occurring actions; moreover, they fail to distinguish semantically similar actions because existing methods typically rely on brief action phrases that lack rich semantic details for action-specific localization. Although large vision-language models (LVLMs) encode rich action semantics and their action-conditioned textual outputs implicitly contain spatial cues, they do not directly provide action-specific spatial localization. To address these problems, we propose TokAG, a zero-shot affordance grounding framework that exploits the token-level semantic-spatial signals in LVLMs to localize action-relevant regions without external supervision. We observe that attention maps associated with different LVLM output tokens vary significantly, with many attending to irrelevant regions such as the background. Thus, we introduce a spatial-aware token-selection mechanism to systematically evaluate each output token and select the one whose attention maps exhibit dominant activation over the target object, instead of relying on arbitrary attention maps. By extracting these object-focused attention maps, we transform the LVLM's implicit semantic signals into zero-shot affordance heatmaps. Our zero-shot framework consistently outperforms prior weakly supervised approaches across multiple benchmarks, improving NSS by 10.7% on the unseen split of AGD20K and by 29.7% on HICO-IIF. The code and models will be made publicly available.
△ Less
Submitted 3 July, 2026;
originally announced July 2026.
-
ClinRAG-GRAPH: Clinical-prior Retrieval-Augmented Graph Model with Domain Adversarial Learning for Breast pCR Prediction
Authors:
Yaofei Duan,
Yuhao Huang,
Tianyu Zhang,
Yuan Gao,
Luyi Han,
Xin Wang,
Xinyu Xie,
Xinglong Liang,
Chunyao Lu,
Muzhen He,
Patrick Pang,
Yue Sun,
Ning Mao,
Tao Tan,
Ritse Mann
Abstract:
Neoadjuvant chemotherapy (NAC) response prediction is clinically important for treatment stratification in breast cancer. However, robust pre-treatment pathological complete response (pCR) prediction remains challenging due to insufficient cross-modal modeling, multicenter imaging heterogeneity, and weak evidence-grounded interpretability. We propose ClinRAG-GRAPH, a Clinically informed Retrieval-…
▽ More
Neoadjuvant chemotherapy (NAC) response prediction is clinically important for treatment stratification in breast cancer. However, robust pre-treatment pathological complete response (pCR) prediction remains challenging due to insufficient cross-modal modeling, multicenter imaging heterogeneity, and weak evidence-grounded interpretability. We propose ClinRAG-GRAPH, a Clinically informed Retrieval-Augmented Generation Graph framework, for pre-treatment pCR prediction from DCE-MRI, structured clinical variables, and biopsy-derived pathological biomarkers. ClinRAG-GRAPH constructs an intra-patient clinical-prior graph and applies a prior-guided relation-aware graph convolutional network for structured multimodal representation learning. To improve cross-center robustness, we introduce a dual-branch domain-adversarial learning strategy to suppress protocol-related MRI bias while preserving pCR-relevant features. To enhance interpretability, we further incorporate large language model (LLM)-driven subgraph RAG module that retrieves clinically analogous historical cases and integrates retrieved evidence for pCR inference. We assemble a large-scale multicenter NAC breast cancer cohort for extensive validation, drawing from two public sources and three in-house centers.Results show that ClinRAG-GRAPH achieves AUCs of 0.815 on the internal test set and 0.774/0.712 on two external test sets, demonstrating robust pre-treatment pCR prediction across centers. The code is available at the anonymized https://github.com/miccai26-1181/ClinRAG-GRAPH.
△ Less
Submitted 1 July, 2026;
originally announced July 2026.
-
DiSTILL: A Hybrid Cloud-HPC Workflow System for Reproducible Spatial Transcriptomics Analysis
Authors:
Myles Joshua Toledo Tan,
Vasco Gerardo Hinostroza Fuentes,
Nikhil Yerra,
Maria Kapetanaki,
Parisa Rashidi,
Kejun Huang,
Panayiotis V. Benos
Abstract:
Spatial transcriptomics workflows increasingly combine large annotated data objects, notebook-based analyses, and resource-intensive statistical models that must be executed on high-performance computing (HPC) systems. In practice, these workflows are often difficult to reproduce because configuration, validation, stage execution, and artifact handling are fragmented across $\textit{ad hoc}$ scrip…
▽ More
Spatial transcriptomics workflows increasingly combine large annotated data objects, notebook-based analyses, and resource-intensive statistical models that must be executed on high-performance computing (HPC) systems. In practice, these workflows are often difficult to reproduce because configuration, validation, stage execution, and artifact handling are fragmented across $\textit{ad hoc}$ scripts and manually edited notebooks. We present $\textit{DiSTILL}$ (Disease Diagnosis from Spatial Transcriptomics via Interpretable Latent Learning), a hybrid cloud$-$HPC workflow system for reproducible spatial transcriptomics (ST) analysis. DiSTILL combines an application programming interface (API) backend built with $\texttt{FastAPI}$, a web frontend, a dataset and preset registry, and a Python pipeline generator that materializes run-specific execution bundles and $\texttt{SLURM}$ submission scripts. The system supports local, Secure Shell (SSH)-mediated, and pull-based poller execution modes, enabling HPC submission in environments where persistent API-initiated automation is restricted. We describe the system through the lens of an inflammatory bowel disease (IBD) ST workflow that operationalizes the analytical pipeline of Tan $\textit{et al.}$ into an auditable application layer. Accordingly, the contribution of this paper is a workflow systems contribution centered on reproducible execution, queue-based orchestration, configuration semantics, and deployment across a split cloud$-$HPC architecture. The broader application goal of DiSTILL is to support user-supplied datasets that satisfy the schema assumptions of the wrapped analytical pipeline.
△ Less
Submitted 28 June, 2026;
originally announced June 2026.
-
On the Vulnerability of Parameter-Level Defenses to Model Merging
Authors:
Kuangpu Guo,
Qingyan Zheng,
Jian Liang,
Yongcan Yu,
Zilei Wang,
Ran He,
Tieniu Tan
Abstract:
The training-free integration of expert models via model merging has exposed significant security risks, enabling free-riders to combine specialized models without authorization. Recent works propose parameter-level defenses that employ linear parameter transformations to neutralize this threat. In this paper, we systematically analyze such defenses and reveal that their protected task vectors are…
▽ More
The training-free integration of expert models via model merging has exposed significant security risks, enabling free-riders to combine specialized models without authorization. Recent works propose parameter-level defenses that employ linear parameter transformations to neutralize this threat. In this paper, we systematically analyze such defenses and reveal that their protected task vectors are inherently small in magnitude. Consequently, the protected weights remain overwhelmingly dominated by the pretrained model. Based on this observation, we designate the pretrained model as a static reference anchor and propose the Anchor-Guided Attack (AGA) to circumvent existing safeguards. Specifically, AGA aligns the protected model with this anchor to recover the transformation matrix analytically. Extensive evaluations validate that AGA consistently bypasses both individual and composite defenses under realistic defense-agnostic scenarios. Furthermore, we provide Anchor-Repulsive Fine-tuning (ARF), a defense method to mitigate the anchor dominance leveraged by AGA. Empirical results confirm that ARF effectively defeats the proposed attack. Our code is available at https://github.com/krumpguo/secure-merge-attack.
△ Less
Submitted 29 June, 2026;
originally announced June 2026.
-
ASTEP confirmation of a pair of long-period Jupiter-sized planets with extremely low densities transiting TOI-791
Authors:
Georgina Dransfield,
Antoine C. Petit,
Amaury H. M. J. Triaud,
Tristan Guillot,
François-Xavier Schmider,
Lyu Abe,
Abdelkrim Agabi,
Khalid Barkaoui,
Thomas A. Baycroft,
Philippe Bendjoya,
Rafael Brahm,
Karen A. Collins,
Billy Edwards,
Phil Evans,
Alix V. Freckelton,
Nolan Grieves,
Steve B. Howell,
Franco Mallia,
Djamel Mekarnia,
Angelica Psaridi,
Daniel Sebastian,
Keivan G. Stassun,
Chris Stockdale,
Amalie Stokholm,
Olga Suarez
, et al. (23 additional authors not shown)
Abstract:
Gas giant planets with periods $20~<~P~<~300~\rm days$ orbiting Sun-like stars are a relatively uncommon outcome of planetary formation, and key questions about the nature and formation of this sub-population remain unanswered. Theoretical models for the location of their formation (in- or ex-situ) and for their subsequent migration predict different outcomes in terms of planet masses and eccentri…
▽ More
Gas giant planets with periods $20~<~P~<~300~\rm days$ orbiting Sun-like stars are a relatively uncommon outcome of planetary formation, and key questions about the nature and formation of this sub-population remain unanswered. Theoretical models for the location of their formation (in- or ex-situ) and for their subsequent migration predict different outcomes in terms of planet masses and eccentricities, indicating that observations have a key role to play in disentangling their histories. In this work we present the discovery and confirmation of a pair of long-period Jupiter-sized planets transiting an F7 star: TOI-791 b is a $0.993\pm0.033\rm~R_{Jup}$ planet on a $139.29931_{-0.00012}^{+0.00011}~\rm day$ orbit, and TOI-791 c, a $1.155\pm0.040\rm ~R_{Jup}$ planet on a $232.01570_{-0.00071}^{+0.00067}~\rm day$ orbit. The two planets are within 0.07% of a second-order 5:3 period commensurability leading to transit timing variations (TTVs) of up to 50 minutes. We confirm their planetary nature using ground-based photometry, including multiple full detections of the $>11~\rm hr$ transits of both TOI-791 b and c from Antarctica with ASTEP, making these the longest-duration transits ever observed in their entirety from the ground. Our detailed analysis of the TTV signal allows us to measure dynamical masses for both planets, which yield densities of $ρ_{\rm b}=0.038\pm0.008 \rm ~g~cm^{-3}$ and $ρ_{\rm c}=0.047\pm0.006 \rm ~g~cm^{-3}$, indicating that TOI-791~b and c are two of the lowest density giant planets ever detected. While these measurements are robust, further follow-up is needed to fully characterise the TTV signal and the architecture of the system.
△ Less
Submitted 29 June, 2026;
originally announced June 2026.
-
CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales
Authors:
Xinlong Chen,
Jiafu Tang,
Yue Ding,
Yizhuo Jia,
Bozhou Li,
Bohan Zeng,
Yang Shi,
Shihao Li,
Yiyan Ji,
Qiang Liu,
Weihong Lin,
Yuanxing Zhang,
Pengfei Wan,
Liang Wang,
Tieniu Tan
Abstract:
Accurate and comprehensive video captions with consistent subject references are critical for downstream understanding and generation tasks. However, few existing benchmarks can objectively and comprehensively evaluate these properties across diverse durations and scenarios, thereby hindering the advancement of video captioning models. To bridge this gap, we propose CapRiCorn-1K, a comprehensive b…
▽ More
Accurate and comprehensive video captions with consistent subject references are critical for downstream understanding and generation tasks. However, few existing benchmarks can objectively and comprehensively evaluate these properties across diverse durations and scenarios, thereby hindering the advancement of video captioning models. To bridge this gap, we propose CapRiCorn-1K, a comprehensive benchmark designed to evaluate both video captioning quality and subject referential consistency across long temporal horizons and diverse video domains. To accommodate varied evaluation needs, our benchmark supports both audiovisual and visual-only settings. Extensive experiments on CapRiCorn-1K reveal that current models generally struggle to generate accurate and comprehensive captions while maintaining consistent subject references. Moreover, as video duration increases, both the overall caption quality and subject referential consistency decline. Notably, our evaluation metrics exhibit strong correlations with the performance of downstream understanding and generation tasks conditioned on the generated captions, further validating their effectiveness. The project is available at https://github.com/xlchen0205/CapRiCorn-1K .
△ Less
Submitted 20 June, 2026;
originally announced June 2026.
-
Cavity-enhanced superconductivity in the two-dimensional limit of NbSe2
Authors:
Hanxiang Zhang,
Zexin Feng,
I-Te Lu,
Zhiwei Li,
Songhao Guo,
Qiuyu Shang,
Thomas Tan,
Xiaodan Lyu,
Xiangming Shen,
Dening Luan,
Mingcheng Panmai,
Kenji Watanabe,
Takashi Taniguchi,
Ranjan Singh,
Angel Rubio,
Weibo Gao
Abstract:
Vacuum electromagnetic fluctuations have emerged as a means of controlling collective quantum phases without external driving. Cavity-induced modification of superconductivity has been widely predicted. What sets the size of the effect, and which microscopic channel carries it, remain open. Here we couple few-layer NbSe2 to a terahertz complementary split-ring resonator (CSRR) and show that the en…
▽ More
Vacuum electromagnetic fluctuations have emerged as a means of controlling collective quantum phases without external driving. Cavity-induced modification of superconductivity has been widely predicted. What sets the size of the effect, and which microscopic channel carries it, remain open. Here we couple few-layer NbSe2 to a terahertz complementary split-ring resonator (CSRR) and show that the enhancement grows sharply on approaching the two-dimensional limit. In bilayer NbSe2 the superconducting transition temperature rises by 10%, from 3.02 K to 3.41 K, on a cavity resonant at 0.92 THz - roughly four times the shift measured in a ten-layer device at the same resonance. Within a single device the shift maps onto the simulated cavity field profile, falling from 0.39 K at the field maximum to zero outside the resonator, with the lower critical field following the same spatial ordering; because all regions are measured on one continuous flake in a single cooldown, sample-to-sample variation is excluded by construction. The frequency dependence is non-monotonic, with suppression below resonance and maximal enhancement near 0.96 THz. Quantum electrodynamical density functional theory calculations show that cavity coupling redistributes spectral weight in the Eliashberg function, weakening the total electron-phonon coupling while hardening the logarithmic average phonon frequency; competition between the two reproduces a sign change in Tc. These results identify dimensionality, local field amplitude and detuning as the control parameters of cavity-enhanced superconductivity, and point to electron-phonon reweighting as its microscopic origin.
△ Less
Submitted 19 August, 2026; v1 submitted 17 June, 2026;
originally announced June 2026.
-
ttda704 at SemEval-2026 Task 4: Modeling Narrative Structures via Pseudonymization and Multi-View Sentence Alignment
Authors:
Tai Tran Tan,
An Dinh Thien
Abstract:
We present our approach to SemEval 2026 Task 4: Narrative Story Similarity and Narrative Representation Learning. Our solution uses contrastive learning with fine-tuned sentence transformers to capture narrative similarity across abstract themes, course of action, and outcomes. We develop two pipelines: (Track A) a single-view method that encodes full narratives with smart layer freezing to reduce…
▽ More
We present our approach to SemEval 2026 Task 4: Narrative Story Similarity and Narrative Representation Learning. Our solution uses contrastive learning with fine-tuned sentence transformers to capture narrative similarity across abstract themes, course of action, and outcomes. We develop two pipelines: (Track A) a single-view method that encodes full narratives with smart layer freezing to reduce overfitting, and (Track B) a multi-view method that models theme, plot, and outcome with view-specific projection heads and self-supervised alignment. Both pipelines build on sentence-transformers models and are trained with contrastive loss on synthetic data. The code is available at the following GitHub repository: https://github.com/dinhthienan33/SemEval2026-Task4-ttda704.
△ Less
Submitted 14 June, 2026;
originally announced June 2026.
-
ttda704 at SemEval-2026 Task 6: Structured Chain-of-Thought Prompting for Political Evasion Detection
Authors:
Tai Tran Tan,
An Dinh Thien
Abstract:
This paper describes our system for SemEval-2026 Task 6, which addresses the classification of political evasion strategies in English question-answer pairs extracted from U.S. presidential interviews. We systematically compare two distinct paradigms: (1) Parameter-Efficient Fine-Tuning of Qwen3 models (4B-32B) using QLoRA, enhanced with tiered upsampling and weighted cross-entropy loss to address…
▽ More
This paper describes our system for SemEval-2026 Task 6, which addresses the classification of political evasion strategies in English question-answer pairs extracted from U.S. presidential interviews. We systematically compare two distinct paradigms: (1) Parameter-Efficient Fine-Tuning of Qwen3 models (4B-32B) using QLoRA, enhanced with tiered upsampling and weighted cross-entropy loss to address severe class imbalance, and (2) structured Chain-of-Thought (CoT) prompting of reasoning-capable API models, namely DeepSeek-V3.2 and Grok-4-Fast. Our evaluation demonstrates that structured CoT prompting of reasoning-enabled models substantially outperforms our baseline parameter-efficient fine-tuning implementation in absolute Macro F1. Our best system, Grok-4-Fast with extended reasoning and few-shot hierarchical CoT prompting, achieves a Macro F1 of 0.5147 on Subtask 2 (9-class evasion) and 0.7979 on Subtask 1 (3-class clarity), ranking 8th out of 33 teams on Subtask 2 and 13th out of 41 teams on Subtask 1 on the official leaderboard. Furthermore, our ablation studies reveal key insights into effective prompt design for evasion detection: presenting labels within a hierarchical taxonomy helps structure model reasoning, while few-shot exemplars provide task calibration. However, the strongest prompt variants are not statistically distinguishable in Macro F1, and explicitly enabling extended reasoning modes yields substantial performance gains by facilitating the multi-step pragmatic analysis required to detect evasive intent.
△ Less
Submitted 14 June, 2026;
originally announced June 2026.
-
Charge as a Construct-Validity Factor in Chinese Legal Case Retrieval: A Cross-Benchmark Audit
Authors:
Yao Liu,
Tien-Ping Tan,
Zhilan Liu
Abstract:
Chinese Legal Case Retrieval (LCR) benchmarks grade a reference judgment relevant when its legal characterization matches the query, and strong systems now reach NDCG@10 of 0.85-0.88. Most of the BM25-to-best-trained gap is recoverable with no retrieval model: ranking candidates only by shared primary charge, broken by BM25, closes 99.2% of it on LeCaRDv2 -- with no detectable difference from the…
▽ More
Chinese Legal Case Retrieval (LCR) benchmarks grade a reference judgment relevant when its legal characterization matches the query, and strong systems now reach NDCG@10 of 0.85-0.88. Most of the BM25-to-best-trained gap is recoverable with no retrieval model: ranking candidates only by shared primary charge, broken by BM25, closes 99.2% of it on LeCaRDv2 -- with no detectable difference from the best-trained system. This reflects benchmark design: LeCaRDv2 defines top relevance via the crime's key constitutive elements, which encode the charge, so same-charge cases are relevant by construction (relevance lift 4.49; charge-to-relevance macro-AUC 0.871). Holding charge fixed, the trained reranker's advantage over BM25 collapses to a small within-charge residual (+0.026 NDCG@10, cluster-bootstrap CI excluding zero, about a quarter), the only non-definitional positive. The effect is not uniform: the same rule recovers 84.3% on LeCaRDv1 and is out of spec on CAIL2022, with the charge-to-relevance signal weakening in step (macro-AUC 0.871/0.759/0.728); a predicted-charge cascade reproduces 76.6% on LeCaRDv2 but does not transfer. The construct is also cashable at first stage: an exploratory zero-training charge-pool channel lifts LeCaRDv2 recall (R@100 +0.025, wrong-charge controls hurt), reported as a positive control for the confound, not a retrieval method or novelty claim. Charge is thus a high-leverage construct-validity factor at the benchmark level -- not auniform explanation of NDCG@10, and not evidence that any system relies on charge. We package established construct-validity and partial-input checks as a reusable charge-controlled protocol (CCE); on all three benchmarks its triggers come back null or descriptive, behaving as designed. We release the scripts, schema, and protocol so future benchmarks can be screened before their NDCG@10 is read as legal-reasoning ability.
△ Less
Submitted 14 June, 2026; v1 submitted 11 June, 2026;
originally announced June 2026.
-
EgoTactile: Learning Grasp Pressure for Everyday Objects from Egocentric Video
Authors:
Yuan Zeng,
Yujia Shi,
Tiao Tan,
Xingting Li,
Yaqi Qin,
Zongqing Lu,
Wenming Yang,
Jing-Hao Xue,
Qingmin Liao
Abstract:
Estimating full-hand grasp pressure from egocentric video is critical for immersive VR and robotic manipulation, yet dense tactile sensing often relies on intrusive hardware. Existing vision-based methods predominantly rely on planar surfaces or fingertip contacts, failing to generalize to complex 3D object interactions. Therefore, we introduce EgoTactile, a benchmark pairing egocentric video with…
▽ More
Estimating full-hand grasp pressure from egocentric video is critical for immersive VR and robotic manipulation, yet dense tactile sensing often relies on intrusive hardware. Existing vision-based methods predominantly rely on planar surfaces or fingertip contacts, failing to generalize to complex 3D object interactions. Therefore, we introduce EgoTactile, a benchmark pairing egocentric video with full-hand pressure supervision for diverse everyday objects, incorporating a bare-hand transfer subset to enable generalization to natural scenarios. Leveraging this benchmark, we first establish EgoPressureFormer as a discriminative baseline. Beyond this, to explicitly address the uncertainty in partial observations, we propose EgoPressureDiff, a conditional diffusion framework that adapts a large-scale pre-trained video diffusion backbone. By combining rich world knowledge priors with a Physically-Informed Feature Rectification layer to inject semantic constraints, our approach effectively infers plausible contact patterns and resolves visual-physical ambiguities. Extensive experiments demonstrate that our method achieves superior performance on the benchmark and robust transferability to in-the-wild scenarios. Our project page is available at https://egotactile.github.io/.
△ Less
Submitted 8 June, 2026;
originally announced June 2026.
-
MMAE: A Massive Multitask Audio Editing Benchmark
Authors:
Ziyang Ma,
Ruiqi Yan,
Ruiyang Xu,
Jie Fang,
Zhikang Niu,
Yi-Wen Chao,
Wenming Tu,
Tianrui Wang,
Auden,
Qi Chen,
Wenxi Chen,
Jiaying Chi,
Yanru Huo,
Zixuan Jiang,
Xiquan Li,
Yalin Li,
Junxi Liu,
Minghao Liu,
Binghao Qiang,
Yijia Shan,
Zheshu Song,
Tian Tan,
Zixiang Wang,
Zeyu Xie,
Zhifei Xie
, et al. (13 additional authors not shown)
Abstract:
We introduce MMAE, a Massive Multitask Audio Editing benchmark, serving as the first comprehensive evaluation testbed designed for general-purpose instruction-based audio editing. Spurred by the shift toward intelligent creation, interactive editing has rapidly expanded from visual domains, pioneered by models like Nano-banana 2 for images and Gemini-Omni for video, into audio. However, the curren…
▽ More
We introduce MMAE, a Massive Multitask Audio Editing benchmark, serving as the first comprehensive evaluation testbed designed for general-purpose instruction-based audio editing. Spurred by the shift toward intelligent creation, interactive editing has rapidly expanded from visual domains, pioneered by models like Nano-banana 2 for images and Gemini-Omni for video, into audio. However, the current evaluation infrastructure lags severely, remaining highly fragmented and restricted to specific subdomains or basic operations. Unlike existing benchmarks that are limited in scope, MMAE extends to a broad spectrum of real-world scenarios, encompassing 7 distinct audio modalities, including sound, speech, music, and their mixtures. Furthermore, we establish a comprehensive taxonomy spanning 6 levels of task complexity, from basic modifications to multi-hop reasoning and multi-round editing, 2 levels of granularity, and 8 distinct operation types. Meticulously curated through human-agent collaboration, MMAE comprises 2,000 high-fidelity samples paired with a pioneering rubric-based evaluation framework. By decomposing free-form tasks into 17,741 verifiable criteria, this robust rubric-based paradigm enables a precise, multi-dimensional assessment of both instruction following and context consistency. Our extensive evaluation of leading models reveals that current systems remain far from achieving reliable edits. Strikingly, the Exact Match Rate (EMR) consistently falls below 5% and plummets to an absolute 0% in complex, mixed-modality tasks, exposing critical bottlenecks in precise execution and structural robustness. We hope MMAE will serve as a catalyst for future advances in the intelligent creation community, providing a clear diagnostic roadmap and establishing a standardized, long-lasting evaluation paradigm for next-generation audio editing systems.
△ Less
Submitted 5 June, 2026;
originally announced June 2026.
-
Beyond Generative Decoding: Discriminative Hidden-State Readout from a Native Omni-Modal LLM for Multimodal Sentiment Analysis
Authors:
Bin Wen,
Tien-Ping Tan
Abstract:
Multimodal sentiment analysis (MSA) infers human affect from language, acoustic, and visual signals. Recent methods increasingly adapt large multimodal models (LMMs) via generative readout: prompting the model to emit a sentiment score as a text string. While convenient, this ties continuous regression to discrete autoregressive decoding, incurring unmeasured costs. We revisit this readout mechani…
▽ More
Multimodal sentiment analysis (MSA) infers human affect from language, acoustic, and visual signals. Recent methods increasingly adapt large multimodal models (LMMs) via generative readout: prompting the model to emit a sentiment score as a text string. While convenient, this ties continuous regression to discrete autoregressive decoding, incurring unmeasured costs. We revisit this readout mechanism and propose a discriminative formulation built on the Thinker module of a native omni-modal LLM (Qwen2.5-Omni-7B). Instead of text decoding, we map the final-layer hidden state of the last non-padding token to a continuous score via a lightweight regression head in a single forward pass. Using 4-bit quantization and low-rank adaptation (QLoRA), the entire 7B pipeline -- including video and audio processing -- trains on a single consumer GPU (RTX 5090, 32 GB) with 10-21 GB peak memory and 1.14% trainable parameters. Through a controlled comparison fixing the backbone, data, and LoRA configuration, we isolate the impact of the readout. On CMU-MOSI and CMU-MOSEI, our discriminative readout reaches state-of-the-art accuracy without task-specific feature engineering (MOSI: MAE 0.551, Corr 0.888; MOSEI: MAE 0.506, Corr 0.790) and exhibits strong multi-seed stability. In contrast, the generative readout -- even after equivalent supervised training -- more than doubles the mean absolute error, yields unparsable or out-of-range outputs (2.8% zero-shot), and suffers from higher latency. Modality ablations reveal a text-dominant regime on CMU-MOSI. Our findings indicate that how an LMM is read out is as consequential as how it is trained, demonstrating that a discriminative readout offers a more accurate, efficient, and reliable alternative for continuous MSA.
△ Less
Submitted 4 June, 2026;
originally announced June 2026.
-
TOI-3664 b, TOI-4034 b & TOI-6564 b: Three new hot Jupiters around stars approaching the terminal age main sequence
Authors:
Matthew P. Battley,
Marina Lafarga,
Edward Gillen,
Monika Lendl,
Solène Ulmer-Moll,
Cynthia S. K. Ho,
Emilio Marfil,
Sergio Sousa,
Yolanda Frensch,
Dimitri Veras,
François Bouchy,
Yann Carteret,
Ian J. M. Crossfield,
Tyler Fairnington,
Mathilde Houelle,
Dan Huber,
Marziye Jafariyazani,
Léna Parc,
Don Radford,
TG Tan,
Sara Tavella,
Rob Wittenmyer,
Duncan Wright,
George Zhou
Abstract:
Studying the evolution of hot Jupiters requires a sample of well-characterised systems across all evolutionary states. We present three new gas giant exoplanets around stars approaching the end of the main sequence, a comparatively unexplored epoch of hot Jupiter evolution. These planets were discovered by TESS before being vetted and confirmed through dedicated spectroscopic follow-up programmes…
▽ More
Studying the evolution of hot Jupiters requires a sample of well-characterised systems across all evolutionary states. We present three new gas giant exoplanets around stars approaching the end of the main sequence, a comparatively unexplored epoch of hot Jupiter evolution. These planets were discovered by TESS before being vetted and confirmed through dedicated spectroscopic follow-up programmes by CARMENES, CORALIE and MINERVA-Australis. TOI-3664 b has a period of 3.30 days, a radius of 1.22 +/- 0.03 RJup and a mass of 0.36 +/- 0.12 MJup. TOI-4034 b is a short-period hot Jupiter with a period of 1.80 days, a radius of 1.58 +/- 0.02 RJup and a mass of 0.87 +/- 0.16 MJup. Meanwhile TOI-6564 b has a period of 3.99 days, radius of 1.46 +/- 0.02 RJup and mass of 0.70 +/- 0.07 MJup. All three planets have radii larger than Jupiter but sub-Jupiter masses, in line with slight inflation as their hosts increase in luminosity towards the end of the main sequence. These exoplanets' low densities and hosts' advanced evolutionary states make them interesting planets with which to study the later stages of hot Jupiter evolution. Careful analysis was undertaken to determine the ages of each system, considering astrometry, gyrochronology, stellar isochrones and lithium abundance, yielding ages of 9.0 +2.4/-2.1 Gyr, 5.7 +/- 0.5 Gyr and 4.0 +/- 1.0 Gyr for TOI-3664, TOI-4034 and TOI-6564 respectively, yet each system has a similar evolutionary state because of their differing stellar masses (0.98 +/- 0.03, 1.19 +0.13/-0.03 and 1.18 +0.16/-0.03 M*). These three planets add more steps to the "age-ladder" of exoplanetary evolution, building towards the community's goal of understanding how planets evolve over time.
△ Less
Submitted 16 July, 2026; v1 submitted 3 June, 2026;
originally announced June 2026.
-
QO-Bench: Diagnosing Query-Operator-Preserving Retrieval over Typed Event Tuples
Authors:
Mengao Zhang,
Xiang Yang,
Chang Liu,
Tianhui Tan,
Ke-wei Huang
Abstract:
Many real-world questions over business, legal, and scientific corpora are natural-language versions of database-style queries over records latent in text. Existing retrieval-augmented generation (RAG) systems are optimized primarily for semantic relevance, but retrieving plausible passages does not guarantee correct query execution. We introduce QO-Bench, a diagnostic benchmark for query-operator…
▽ More
Many real-world questions over business, legal, and scientific corpora are natural-language versions of database-style queries over records latent in text. Existing retrieval-augmented generation (RAG) systems are optimized primarily for semantic relevance, but retrieving plausible passages does not guarantee correct query execution. We introduce QO-Bench, a diagnostic benchmark for query-operator question answering over typed event tuples. The benchmark covers 22,984 news articles and 614 corporate events across 18 query templates, evaluated on 785 questions. Each gold answer is deterministically computed from typed event tuples and scored by recall, with answers matched to the gold tuples by exact match rather than an LLM judge. This design enables operator-level diagnosis such as joins and intersection. We evaluate RAG, ReAct RAG, GraphRAG, and information-extraction-to-SQL under matched conditions, with a long-context oracle ceiling to isolate retrieval failure. A two-axis framework -- index-time preservation versus query-time execution -- predicts where each paradigm fails, and the results bear it out: systems retrieve relevant text but discard the typed values operators need, and the deployable paradigm ranking inverts across operators, with similarity retrieval leading on filter/project and extraction-to-SQL on intersection and counting. Even given the gold evidence, a long-context oracle stays far from saturated, so operator execution -- not retrieval alone -- is a core bottleneck that a stronger answer model does not remove. QO-Bench reframes the goal from passage relevance to query-operator-preserving retrieval.
△ Less
Submitted 3 June, 2026;
originally announced June 2026.
-
AnyAudio-Judge: A Dynamic Rubric-Based Benchmark and Evaluator for Audio Instruction Following
Authors:
Haitao Li,
Tian Tan,
Yuguang Yang,
Shan Yang,
Xie Chen
Abstract:
The rapid advancement of instruction-guided audio generation has highlighted the critical need for robust alignment evaluation. Current automated evaluation methods heavily rely on holistic scoring from general-purpose large language models, which struggle to decouple complex instructions, lack interpretability, and fail to capture fine-grained attribute mismatches. To address this, we introduce a…
▽ More
The rapid advancement of instruction-guided audio generation has highlighted the critical need for robust alignment evaluation. Current automated evaluation methods heavily rely on holistic scoring from general-purpose large language models, which struggle to decouple complex instructions, lack interpretability, and fail to capture fine-grained attribute mismatches. To address this, we introduce a novel dynamic rubric-based evaluation paradigm that adaptively decomposes complex audio captions into a variable number of independent, verifiable binary rubric items. To rigorously benchmark this capability, we propose the AnyAudio-Judge Bench, a comprehensive, bilingual benchmark comprising 7,920 meticulously curated samples across four diverse audio domains (speech, sound, music, and mixed), featuring deliberately constructed hard negatives. Furthermore, we construct a large-scale corpus of 105K samples with explicit Chain-of-Thought (CoT) rationales to train our dedicated evaluator, the AnyAudio-Judge model. By employing a training pipeline that combines Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO), our model successfully aligns its reasoning paths with the rubric-based scoring mechanism. Extensive experiments demonstrate that AnyAudio-Judge not only significantly enhances zero-shot alignment detection compared to state-of-the-art baselines, but also provides precise and interpretable reward signals that substantially improve instruction alignment in downstream reinforcement learning for audio generation.
△ Less
Submitted 2 June, 2026;
originally announced June 2026.
-
Visualizing orbital magnetism in electron doped rhombohedral multilayer graphene
Authors:
Owen I. Sheekey,
Trevor B. Arp,
Benjamin A. Foutty,
Ruoxi Zhang,
Tixuan Tan,
Ludwig F. W. Holleis,
Yi Guo,
Sandesh S. Kalantre,
Canxun Zhang,
Mark Zakharyan,
David Gong,
Aidan Keough,
Youngjoon Choi,
Ysun Choi,
Siyuan Xu,
Tian Xie,
Ben Hodder Alexander,
Marisa Hocking,
Qingrui Cao,
Martin E. Huber,
Takashi Taniguchi,
Kenji Watanabe,
Chenhao Jin,
Etienne Lantagne-Hurtubise,
Aaron Sharpe
, et al. (2 additional authors not shown)
Abstract:
Electron doped rhombohedral multilayer graphene at high displacement field features an exceptionally flat band minimum with near-ideal quantum geometry. Experiments in this regime observe the formation of a 'quarter metal,' in which the electron liquid condenses into a single spin- and valley flavor. Remarkably, recent experiments have found a zero resistance state in the same region of the densit…
▽ More
Electron doped rhombohedral multilayer graphene at high displacement field features an exceptionally flat band minimum with near-ideal quantum geometry. Experiments in this regime observe the formation of a 'quarter metal,' in which the electron liquid condenses into a single spin- and valley flavor. Remarkably, recent experiments have found a zero resistance state in the same region of the density- and displacement-field-tuned parameter space, attributed to the formation of a chiral superconductor from an orbitally ferromagnetic normal state. Here, we use nanoSQUID-on-tip magnetometry to map the orbital magnetization of electron-doped rhombohedral graphene devices ranging in thickness between 3 and 15 layers. Magnetization within the quarter metal phases peaks at finite density, consistent with concentration of the Berry curvature in a finite-momentum 'ring of fire'. Correlating transport and local magnetometry data in a superconducting tetralayer sample reveals a finite orbital ferromagnetic moment, providing direct evidence of valley polarization in the superconducting ground state. We further show that widely observed stochastic switching of the resistivity in both metallic and superconducting regimes arises from a density-tuned sign change in the valley-resolved total magnetic moment. This leads to the formation of metastable magnetic domains under typical gate control sequences and can also be harnessed for electric-field controlled switching of the magnetization across the entire device. Finally, high resolution measurements of the magnetization across a superconducting transition allow us to put an upper bound on the 'condensation magnetization' of 0.1 Bohr magneton per carrier, placing a strong quantitative restriction on theoretical models for ferromagnetic superconductivity.
△ Less
Submitted 4 August, 2026; v1 submitted 28 May, 2026;
originally announced May 2026.
-
SAFE-Diff: Scale-Aware Attention and Feature-Dispersive Diffusion with Uncertainty Estimation for Contrast-Enhanced Breast MRI Synthesis
Authors:
Tianyu Zhang,
Xinglong Liang,
Jarek van Dijk,
Luyi Han,
Chunyao Lu,
Antonio Portaluri,
Xinghe Xie,
Yaofei Duan,
Nika Rasoolzadeh,
Xin Wang,
Yuan Gao,
Muzhen He,
Yue Sun,
Jonas Teuwen,
Tao Tan,
Ritse Mann
Abstract:
Synthesizing high fidelity contrast enhanced MRI is clinically valuable for safer and more efficient breast cancer screening, yet remains challenging due to complex lesion textures and heterogeneous enhancement patterns.
Synthesizing high fidelity contrast enhanced MRI is clinically valuable for safer and more efficient breast cancer screening, yet remains challenging due to complex lesion textures and heterogeneous enhancement patterns.
△ Less
Submitted 26 May, 2026; v1 submitted 25 May, 2026;
originally announced May 2026.
-
GlowGS: Generative Semantic Feature Learning for 3D Gaussian Splatting in Nighttime Glow Scenes
Authors:
Beibei Lin,
Xiao Cao,
Jingyuan Guo,
Robby T. Tan
Abstract:
Existing 3DGS methods effectively render high-quality novel views in clear-day scenes. However, they struggle with night scenes, particularly in glow regions, due to the lack of structural features such as textures and edges, which are key cues for splatting-based reconstruction. To address this problem, we leverage a diffusion model and a Vision Foundation Model (VFM) to compensate for missing st…
▽ More
Existing 3DGS methods effectively render high-quality novel views in clear-day scenes. However, they struggle with night scenes, particularly in glow regions, due to the lack of structural features such as textures and edges, which are key cues for splatting-based reconstruction. To address this problem, we leverage a diffusion model and a Vision Foundation Model (VFM) to compensate for missing structural cues. Our method consists of two key novel ideas: semantic feature generation and novel-view semantic learning. First, semantic feature generation produces high-quality semantic features as implicit structural cues for novel views. Specifically, a diffusion model synthesizes novel views with unknown camera poses from training views, while a VFM evaluates their quality. Once high-quality novel views are identified, the VFM extracts robust features to construct the semantic feature bank. Second, novel-view semantic learning enables 3DGS to optimize rendered novel views without requiring ground truth. It achieves this by extracting semantic features from a rendered novel view, searching the feature bank for the most similar features, and minimizing their distance. This process enforces implicit structural constraints, ensuring semantically coherent, artifact-free rendered views. Extensive experiments demonstrate the effectiveness of our GlowGS in generating semantically accurate 3D views, showing significant improvements over existing methods.
△ Less
Submitted 22 May, 2026;
originally announced May 2026.
-
RADAR: Defending RAG Dynamically against Retrieval Corruption
Authors:
Ziyuan Chen,
Yueming Lyu,
Yi Liu,
Weixiang Han,
Jing Dong,
Caifeng Shan,
Tieniu Tan
Abstract:
While RAG systems are increasingly deployed in dynamic web search, temporal volatility amplifies their vulnerability to adversarial attacks. Existing static-oriented defenses struggle to handle evolving threats and incur prohibitive storage costs in dynamic settings. We propose RADAR, a framework that models reliable context selection as a graph-based energy minimization problem, solved exactly vi…
▽ More
While RAG systems are increasingly deployed in dynamic web search, temporal volatility amplifies their vulnerability to adversarial attacks. Existing static-oriented defenses struggle to handle evolving threats and incur prohibitive storage costs in dynamic settings. We propose RADAR, a framework that models reliable context selection as a graph-based energy minimization problem, solved exactly via Max-Flow Min-Cut. By incorporating a Bayesian memory node, RADAR recursively updates a belief state instead of archiving raw historical documents, effectively balancing stability against attacks with adaptability to genuine knowledge shifts. Experiments on a novel dynamic dataset show that RADAR achieves superior robustness and response quality with minimal storage overhead compared to the baselines.
△ Less
Submitted 21 May, 2026;
originally announced May 2026.
-
White-Balance First, Adjust Later: Cross-Camera Color Constancy via Vision-Language Evaluation
Authors:
Shuwei Li,
Lei Tan,
Robby T. Tan
Abstract:
Color constancy aims to keep object colors consistent under varying illumination. Cross-camera generalization in color constancy remains challenging because learning-based models often overfit to the color response characteristics of the training camera, resulting in degraded performance on images captured by other cameras. We propose VLM-CC, a feedback-guided framework that formulates color const…
▽ More
Color constancy aims to keep object colors consistent under varying illumination. Cross-camera generalization in color constancy remains challenging because learning-based models often overfit to the color response characteristics of the training camera, resulting in degraded performance on images captured by other cameras. We propose VLM-CC, a feedback-guided framework that formulates color constancy as an iterative refinement process. Instead of directly estimating the illuminant from raw input, VLM-CC performs iterative correction driven by vision-language model (VLM)-based evaluation. At each iteration, the image is white-balanced using the current estimate and converted to pseudo-sRGB. A lightweight LoRA-tuned VLM then assesses the corrected image, identifying the dominant residual color cast and providing qualitative feedback. This feedback is mapped to a residual illumination direction (red, green, or blue) and used to update the illuminant estimate until convergence. Our key idea is to reframe color constancy as an iterative perceptual feedback problem, leveraging VLM evaluation instead of direct RGB regression. By replacing direct RGB estimation with VLM-guided perceptual feedback, VLM-CC achieves state-of-the-art robustness in cross-camera color constancy across multiple datasets. Code will be available at https://github.com/NothingIknow/VLM-CC.
△ Less
Submitted 19 May, 2026;
originally announced May 2026.
-
From Static Risk to Dynamic Trajectories: Toward World-Model-Inspired Clinical Prediction
Authors:
Pujun Feng,
Xiaoyu Guo,
Seyed Ehsan Saffari,
Min Hun Lee,
Siew-Kei Lam,
Erik Cambria,
Xibin Sun,
Yangtao Zhou,
Tong Yang,
Xiaoyu Zhang,
Tao Tan,
Yue Sun,
Bin Cui
Abstract:
Clinical decision-making is a feedback system where risk estimates influence treatment, which in turn changes disease trajectories, and both shape clinicians' measurement practices. Static prediction often fails clinically: models trained on observational care logs conflate disease biology with clinician behavior, particularly under treatment confounder feedback and irregular or informative observ…
▽ More
Clinical decision-making is a feedback system where risk estimates influence treatment, which in turn changes disease trajectories, and both shape clinicians' measurement practices. Static prediction often fails clinically: models trained on observational care logs conflate disease biology with clinician behavior, particularly under treatment confounder feedback and irregular or informative observation. This Review focuses on intervention-aware disease trajectory modeling in clinical AI--methods estimating patient-specific longitudinal disease evolution and assessing trajectory changes under alternative treatments. We organize the field around six linked components: three decision tasks (factual forecasting, counterfactual estimation, policy evaluation) and three data-generating mechanisms (disease evolution, treatment assignment, observation process) that determine identifiability. We present the first unified framework bridging forecasting, counterfactual trajectories, and policy evaluation across discrete/continuous time, explicitly addressing treatment assignment, time-varying confounding, and observation bias. We synthesize key method families (multistate/joint models, temporal point-process, deep sequence architectures, longitudinal causal inference), map them to relevant components, and align evaluation with claim strength via overlap diagnostics, uncertainty quantification, off-policy robustness, and target-trial validation. This synthesis advances benchmark prediction to decision-grade clinical evidence, enabling treatment-sensitive individualized futures, pre-deployment policy stress-testing, and safer closed-loop learning health systems that adapt/abstain when evidence is insufficient.
△ Less
Submitted 16 May, 2026;
originally announced May 2026.
-
Dispersion-Engineered Terahertz Silicon Interconnects Enabling Terabit-Scale Data Links
Authors:
Bodhan Chakraborty,
Wenhao Wang,
Nikhil Navaratna,
Thomas Caiwei Tan,
Pascal Szriftgiser,
Hadjer Nihel Khelil,
Guillaume Ducournau,
Ranjan Singh
Abstract:
The rapid growth of artificial intelligence (AI) and data-centric computing is driving exabyte-scale data transfer, pushing conventional interconnect technologies toward fundamental bandwidth and energy limits. Although optical interconnects provide high-capacity and long-reach communication, their complexity and energy overhead limit scalability in short-reach chiplet-based and on-chip systems. T…
▽ More
The rapid growth of artificial intelligence (AI) and data-centric computing is driving exabyte-scale data transfer, pushing conventional interconnect technologies toward fundamental bandwidth and energy limits. Although optical interconnects provide high-capacity and long-reach communication, their complexity and energy overhead limit scalability in short-reach chiplet-based and on-chip systems. Terahertz (THz) silicon interconnects offer a promising alternative by bridging electronics and photonics in compact, complementary metal-oxide-semiconductor (CMOS)-compatible platforms capable of high bandwidth and low latency. However, practical THz interconnects require simultaneous multi-band operation, dual-polarization support, low propagation loss, low group-velocity dispersion (GVD), and terabit-per-second throughput, while avoiding Bragg-induced stopbands and dispersion penalties at high frequencies. Here, we demonstrate a CMOS-compatible, centimetre-scale, multi-band on-chip THz data link achieving an aggregate throughput of 1.004 Tbps. The performance is enabled by suppressing Bragg-induced stopbands using dispersion-engineered, effective-medium-supported unclad silicon waveguides, resulting in flat transmission and low-ripple group delay across multiple THz bands. The waveguide platform operates from 220 to 500 GHz and supports both transverse-electric (TE) and transverse-magnetic (TM) polarizations with low path loss, low bending loss, and low GVD. Fourteen channels in a straight waveguide and twelve channels in a 90$^\circ$ bend achieve aggregate data rates of 1.004 Tbps and 0.895 Tbps, respectively, with GVD as low as 0.15 ps$^2$/mm over the full operating band. These results establish a scalable and energy-efficient THz interconnect platform for high-density on-chip and chip-to-chip communication fabrics targeting next-generation AI systems and emerging 6G technologies.
△ Less
Submitted 16 May, 2026;
originally announced May 2026.
-
DexJoCo: A Benchmark and Toolkit for Task-Oriented Dexterous Manipulation on MuJoCo
Authors:
Hanwen Wang,
Weizhi Zhao,
Xiangyu Wang,
Siyuan Huang,
He Lin,
Boyuan Zheng,
Rongtao Xu,
Gang Wang,
Yao Mu,
He Wang,
Lue Fan,
Hongsheng Li,
Zhaoxiang Zhang,
Tieniu Tan
Abstract:
Achieving human-level manipulation requires dexterous robotic hands capable of complex object interactions. Advancing such capabilities further demands standardized benchmarks for systematic evaluation. However, existing dexterous benchmarks lack tasks that reflect the unique manipulation capabilities of dexterous hands over parallel grippers, as well as comprehensive evaluation pipelines. In this…
▽ More
Achieving human-level manipulation requires dexterous robotic hands capable of complex object interactions. Advancing such capabilities further demands standardized benchmarks for systematic evaluation. However, existing dexterous benchmarks lack tasks that reflect the unique manipulation capabilities of dexterous hands over parallel grippers, as well as comprehensive evaluation pipelines. In this paper, we present DexJoCo, a benchmark and toolkit for task-oriented dexterous manipulation, comprising 11 functionally grounded tasks that evaluate tool-use, bimanual coordination, long-horizon execution, and reasoning. We develop a low-cost data collection system and collect 1.1K trajectories across these tasks, with support for domain randomization to assess robustness. We benchmark modern models under diverse settings, including visual and dynamics randomization, multi-task training, and action-head adaptation. Through extensive empirical analysis, we identify several important insights and common limitations of current policies in dexterous manipulation, highlighting key challenges for future research in dexterous hand robot learning. Project page available at: https://dexjoco.github.io
△ Less
Submitted 15 May, 2026;
originally announced May 2026.
-
Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents
Authors:
Minzheng Wang,
Run Luo,
Yanbo Wang,
Zichen Liu,
Yuqiao Tan,
Tao Tan,
Xu Nan,
Yinhe Zheng,
Wenji Mao
Abstract:
While Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for closed-ended tasks, extending it to open-ended social language games via self-play reveals a critical issue: evolution impasse. Due to the vast strategy space, language agents frequently converge to homogenized behaviors, leading to deterministic match outcomes that eliminate the gradient signals necessary for pol…
▽ More
While Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for closed-ended tasks, extending it to open-ended social language games via self-play reveals a critical issue: evolution impasse. Due to the vast strategy space, language agents frequently converge to homogenized behaviors, leading to deterministic match outcomes that eliminate the gradient signals necessary for policy evolution. To tackle this issue, we propose Dual-scale Evolutionary Policy Training (DEPT) for social language games. DEPT introduces a time-scaled evolutionary perception mechanism that detects impasse by quantifying dual-scale value baseline divergence alongside match entropy. Upon perceiving the collapse, it then activates asymmetric advantage reshaping to dynamically modulate the optimization landscape for intervention. Thus, our method effectively restores gradient signals and enforces sustained strategic exploration. Extensive experiments on multiple social language games demonstrate that DEPT outperforms strong baselines, avoiding policy degeneration and driving the continuous evolution of social language agents.
△ Less
Submitted 9 May, 2026;
originally announced May 2026.
-
Reducing Linguistic Hallucination in LM-Based Speech Enhancement via Noise-Invariant Acoustic-Semantic Distillation
Authors:
Zheng Wang,
Xiaobin Rong,
Hang Su,
Tianyi Tan,
Junnan Wu,
Lichun Fan,
Zhenbo Luo,
Jian Luan,
Jing Lu
Abstract:
Language model (LM)-based speech enhancement (SE) can generate natural-sounding speech, but under severe noise it often suffers from unreliable conditioning, leading to perceptually plausible yet linguistically incorrect outputs. To address this issue, we propose L3-SE, a noise-invariant acoustic-semantic distillation framework for reducing linguistic hallucination in LM-based SE. The proposed met…
▽ More
Language model (LM)-based speech enhancement (SE) can generate natural-sounding speech, but under severe noise it often suffers from unreliable conditioning, leading to perceptually plausible yet linguistically incorrect outputs. To address this issue, we propose L3-SE, a noise-invariant acoustic-semantic distillation framework for reducing linguistic hallucination in LM-based SE. The proposed method learns a noise-invariant conditioning encoder from noisy speech by jointly distilling two complementary clean-speech targets: an acoustic target for reconstruction fidelity and a semantic target for linguistic consistency. The resulting noise-invariant acoustic-semantic representations are used to condition a decoder-only autoregressive language model, which predicts clean acoustic tokens that are decoded into enhanced speech. To support high-quality generation, we further employ a high-fidelity codec built on learnable weighted WavLM layer representations as the discrete acoustic interface. By improving the reliability of conditioning under adverse conditions, the proposed framework substantially reduces hallucination and improves content faithfulness. Experiments show that the proposed method consistently outperforms prior LM-based speech enhancement baselines on linguistic consistency metrics, with especially clear gains under low-SNR and reverberant conditions, while maintaining competitive perceptual quality. Audio samples are available at https://max1wz.github.io/L3-SE-Demo-Page/. The complete source code will be released after the manuscript is accepted.
△ Less
Submitted 8 May, 2026;
originally announced May 2026.