-
Point Spread Function Engineering Using Implicit Neural Representations
Authors:
Suet Ying Chan,
Mitchell Gilmore,
Qilin Deng,
Guorong Hu,
Joseph Greene,
Ruipeng Guo,
Lei Tian
Abstract:
Point spread function (PSF) engineering through pupil plane modulation is a technique used in microscopy to achieve specific imaging properties, such as depth encoding or extended depth of field. Existing PSF design methods often rely on extensive domain knowledge and task-specific basis functions, making it difficult to generalize across different applications. We treat the PSF engineering task a…
▽ More
Point spread function (PSF) engineering through pupil plane modulation is a technique used in microscopy to achieve specific imaging properties, such as depth encoding or extended depth of field. Existing PSF design methods often rely on extensive domain knowledge and task-specific basis functions, making it difficult to generalize across different applications. We treat the PSF engineering task as a phase retrieval problem and propose a neural field pupil design method that optimizes a phase profile for any arbitrary, user-defined 3D PSF distribution. This provides a flexible framework for 3D PSF engineering for various applications with implicit regularization that proves robust to initialization compared to pixel-wise optimization methods
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
RAGE-Vis:A Relation-Aware Generative Editing Interface for Natural Language-Based Chart Editing
Authors:
Ziyao Kang,
Yiping Sun,
Linxuan Tian,
Henghuan Qu,
Wei Zeng,
Jiazhi Xia
Abstract:
Natural language offers an easy way for users to express chart editing intents, which are often composite and cross-component (e.g., adjusting style, extending categories, highlighting values). However, existing methods typically map instructions to a single operation or widget, limiting their ability to handle high-level requests and often producing locally plausible but globally inconsistent res…
▽ More
Natural language offers an easy way for users to express chart editing intents, which are often composite and cross-component (e.g., adjusting style, extending categories, highlighting values). However, existing methods typically map instructions to a single operation or widget, limiting their ability to handle high-level requests and often producing locally plausible but globally inconsistent results due to a lack of awareness of relationships between chart components. To address these challenges, we introduce RAGE-Vis, a Relation-Aware Generative Editing interface for natural language-based chart editing. The system supports bitmap chart images as input and converts them into an editable parameterized intermediate representation. Instead of mapping instructions to a single edit or widget, RAGE-Vis parses composite intents, identifies targets and scopes, and generates hierarchical editing panels for underspecified requests, enabling users to adjust both global settings and local parameters. Furthermore, RAGE-Vis identifies potentially affected fields based on visual encoding relations, structural relationships, and expressive consistency relations, and organizes them into actionable widgets to support cross-component coordinated controls. Through two case studies, we demonstrate the applicability of RAGE-Vis in complex editing tasks, including style adjustment, data extension, order rearrangement, legend layout, and color mapping. A user study further shows that participants can effectively handle underspecified requests, explore candidate alternatives, and maintain cross-component consistency with RAGE-Vis.
△ Less
Submitted 11 August, 2026;
originally announced August 2026.
-
TAMS: Task-Aware Multi-View Adaptive Streaming for Wireless Telerobotic Manipulation
Authors:
Zexin Deng,
Zhenhui Yuan,
Lu Tian,
Subhash Lakshminarayana,
Longhao Zou
Abstract:
Wireless telerobotic manipulation relies on timely multi-view video feedback, but the available uplink bandwidth is often limited and dynamic. This paper presents Task-Aware Multi-View Adaptive Streaming (TAMS), a system that allocates video bitrate according to the current manipulation phase. TAMS infers task phase from lightweight robot-side signals and prioritizes the camera view most relevant…
▽ More
Wireless telerobotic manipulation relies on timely multi-view video feedback, but the available uplink bandwidth is often limited and dynamic. This paper presents Task-Aware Multi-View Adaptive Streaming (TAMS), a system that allocates video bitrate according to the current manipulation phase. TAMS infers task phase from lightweight robot-side signals and prioritizes the camera view most relevant to the operator while preserving baseline visibility for secondary views. Experiments on a six-degree-of-freedom (6-DoF) teleoperation testbed under three constrained network conditions show that TAMS improves primary view Structural Similarity Index (SSIM), reduces task completion time, and increases trial success rate compared with equal and static allocation baselines. Under the most constrained bandwidth condition, TAMS reduces mean completion time from 68.9 s to 43.9 s relative to equal allocation and increases trial success rate from 48% to 71%. Code is available at: https://github.com/Dzxx623/TAMS.
△ Less
Submitted 10 August, 2026;
originally announced August 2026.
-
DREAM Technical Report
Authors:
Bin Zhang,
Bowen Zheng,
Chao Yi,
Chengyu Lai,
Dian Chen,
Dimin Wang,
Gaoyang Guo,
Jialin Zhu,
Jian Wu,
Jing Yu,
Jiuning Lin,
Lingqing Zhang,
Lingyun Zheng,
Mao Zhang,
Mingming Pan,
Ruiquan Lan,
Shuai Zhong,
Wen Chen,
Wendong Zhang,
Xiaodong Zhu,
Xuan Chen,
Xunke Xi,
Yifan Lu,
Yiheng Wang,
Yue Zeng
, et al. (52 additional authors not shown)
Abstract:
Industrial recommender systems commonly use cascaded retrieval, ranking, and re-ranking pipelines. Although efficient, these pipelines fragment information and objectives across modules, rely on rigid rules, and have limited awareness of real-time intent, leaving session-level shifts among browsing, comparison, and purchase insufficiently addressed. We present DREAM (Developing Recommender Engine…
▽ More
Industrial recommender systems commonly use cascaded retrieval, ranking, and re-ranking pipelines. Although efficient, these pipelines fragment information and objectives across modules, rely on rigid rules, and have limited awareness of real-time intent, leaving session-level shifts among browsing, comparison, and purchase insufficiently addressed. We present DREAM (Developing Recommender Engine with Agentic Methods), an autonomous optimization control architecture that adds a perception-aware, orchestrable, and auditable policy layer atop existing pipelines without replacing them. DREAM has two core components. First, a three-tier Intent Engine fuses on-device signals into structured L0/L1/L2 intent representations; its edge-cloud trigger chain reduces reporting volume to approximately 8.7%. Second, a Meta Engine uses a MetaModel for layered M1-to-M2-to-M3 reasoning: intent summarization, strategy planning informed by Strategy Memory, and parameter translation. It dispatches the resulting parameters through a unified outlet with safety guardrails. A Reward Dual Loop continuously optimizes both components by combining offline simulation for strategy-space exploration with online feedback for outcome calibration, forming a cycle of generation, execution, evaluation, and experience accumulation. Large-scale A/B tests on Taobao's homepage feed show that re-ranking control alone improves IPV by 2.06%, Core IPV by 2.39%, and GMV by 0.88%. Extending control to fine ranking raises these gains to 2.71%, 3.06%, and 1.31%, respectively, while consistently improving PV by more than 1%. These gains require neither replacement of pipeline models nor compromise of serving stability, supporting agentic meta-control as a viable paradigm for industrial recommendation.
△ Less
Submitted 13 August, 2026; v1 submitted 10 August, 2026;
originally announced August 2026.
-
ConceptADapt: Concept-guided Adaptive Feature Reconstruction with Dynamic Attention for Few-Shot Industrial Anomaly Detection
Authors:
Yufei Li,
Yicheng Ruan,
Long Tian,
Dongsheng Wang,
Liang Bao
Abstract:
Few-shot industrial anomaly detection (FS-IAD) focuses on detecting and localizing visual defects in industrial inspection during the cold-start phase, where only a limited number of normal training samples are available per category. Recent advances in this field predominantly leverage visual features from foundation-model and have achieved promising performance. Despite the strong representation…
▽ More
Few-shot industrial anomaly detection (FS-IAD) focuses on detecting and localizing visual defects in industrial inspection during the cold-start phase, where only a limited number of normal training samples are available per category. Recent advances in this field predominantly leverage visual features from foundation-model and have achieved promising performance. Despite the strong representational power of foundation-model features, the model generalization remains fragile due to the extreme scarcity of normal training data.To address this pivotal issue, we propose ConceptADapt, a concept-guided adaptive feature reconstruction model with dynamic attention. Specifically, our model pre-learns a set of fixed normal concepts from the limited support features and leverages them to mine relationships with query features, thereby recalibrating their statistics for improved anomaly detection at test time. To mitigate the prevalent feature shortcut problem, which is particularly severe under low-data regimes, we further develop a dynamic attention mechanism integrated with sparse autoencoders to learn robust normal concepts during training. Moreover, to enable fast adaptation during inference, our model remains lightweight by incorporating LoRA into the attention module, which introduces only minimal updating parameters.Extensive experiments on three widely adopted FS-IAD benchmarks, including MVTec-AD, VisA, and MPDD, demonstrate that our model consistently outperforms state-of-the-art (SOTA) approaches across both detection and localization tasks, achieving significant improvements under various shot settings.
△ Less
Submitted 6 August, 2026;
originally announced August 2026.
-
$d$-spacing distributions as a probe of nematoelastic response in iron-based superconductors
Authors:
Wenting Zhang,
Ruixian Liu,
Tingjun Zhang,
Weiliang Yao,
Xüe Fu,
Hanqing Xie,
Ziye Mo,
Ting Guo,
Kuo-Feng Tseng,
Thomas Keller,
Jitae T. Park,
Fankang Li,
Masaaki Matsuda,
Avishek Maity,
Long Tian,
Pengcheng Dai,
Xingye Lu
Abstract:
Electronic nematicity in iron-based superconductors (FeSCs) couples bilinearly to orthorhombic strain, allowing nematic correlations to appear in the lattice response. Here we use neutron Larmor diffraction to measure the temperature-dependent distribution of relative $d$ spacings in electron-doped Ba(Fe$_{1-x}$Co$_x$)$_2$As$_2$, hole-doped Ba$_{0.83}$K$_{0.17}$Fe$_2$As$_2$, FeSe, and Fe$_{1.07}$T…
▽ More
Electronic nematicity in iron-based superconductors (FeSCs) couples bilinearly to orthorhombic strain, allowing nematic correlations to appear in the lattice response. Here we use neutron Larmor diffraction to measure the temperature-dependent distribution of relative $d$ spacings in electron-doped Ba(Fe$_{1-x}$Co$_x$)$_2$As$_2$, hole-doped Ba$_{0.83}$K$_{0.17}$Fe$_2$As$_2$, FeSe, and Fe$_{1.07}$Te. In Ba(Fe$_{1-x}$Co$_x$)$_2$As$_2$ crystals without intentionally applied uniaxial stress, the in-plane distribution width, $\varepsilon_{\rm FWHM}$, increases on cooling in the tetragonal phase and can be described phenomenologically by a Curie--Weiss-like form. The fitted scale $T^*$ decreases with Co doping and evolves similarly to the nematic phase diagram inferred from elastoresistance, although the two experiments probe different response functions. Related broadening in Ba$_{0.83}$K$_{0.17}$Fe$_2$As$_2$ and FeSe supports extending this interpretation beyond electron-doped BaFe$_2$As$_2$. By contrast, Fe$_{1.07}$Te shows no extended Curie--Weiss-like regime without applied stress, whereas uniaxial pressure produces a strongly anisotropic broadening that can contain contributions from both the field-biased lattice response and inhomogeneous loading. A mean-field model with bilinear nematoelastic coupling and spatially varying symmetry-breaking stress explains the Curie--Weiss-like broadening in terms of the renormalized orthorhombic compliance. Neutron Larmor diffraction therefore provides a bulk-sensitive probe of nematic-related lattice broadening that complements electronic and elastic measurements.
△ Less
Submitted 2 August, 2026;
originally announced August 2026.
-
CoRenew: A large language model agent-based policy simulation platform for multifamily residential redevelopment
Authors:
Yudi Zhang,
Yuming Lin,
Li Tian,
Yu Wang,
Jianghao Yu
Abstract:
The difficulty of collective action remains a central challenge in the design of policies for multifamily residential redevelopment. Stakeholders continually adjust their decisions in response to evolving negotiation contexts and the reactions of others, meaning that when a policy intervenes and which stakeholders it targets can substantially reshape collective outcomes. Assessing these adaptive r…
▽ More
The difficulty of collective action remains a central challenge in the design of policies for multifamily residential redevelopment. Stakeholders continually adjust their decisions in response to evolving negotiation contexts and the reactions of others, meaning that when a policy intervenes and which stakeholders it targets can substantially reshape collective outcomes. Assessing these adaptive responses ex ante remains difficult because existing simulation models often rely on predefined behavioral rules. Here, we present CoRenew, an open-source platform that uses LLM-based agents to simulate negotiations among multiple stakeholders and evaluate the effects of alternative policy combinations. Integrating open source geographic and demographic data, the platform can generate synthetic residents, simulate negotiation dynamics under alternative policy settings and compares policy performance across competing objectives. It supports both numerical and semantic policy inputs and includes built-in tools for visualization and result export. We validate its behavioral realism against survey responses from 324 residents and a nine-month observed negotiation process from a real redevelopment case. With its modular and adaptable architecture, CoRenew can be used to assess policies across different institutional and cultural contexts.
△ Less
Submitted 28 July, 2026;
originally announced July 2026.
-
From a Word-Level Dictionary to Sentence-Level Semantics: Multilingual Grievance Labelling with Contextual Models
Authors:
Lin Tian,
Marian-Andrei Rizoiu
Abstract:
Grievance is one of the warning signs analysts look for when assessing threats of violence. It is increasingly measured at scale from online text, most often with word-level lexicons like the Grievance Dictionary that score by matching weighted terms. Such matching is a fast and transparent proxy, but it cannot resolve whether a term is asserted, quoted, negated, or condemned. These lexicons are a…
▽ More
Grievance is one of the warning signs analysts look for when assessing threats of violence. It is increasingly measured at scale from online text, most often with word-level lexicons like the Grievance Dictionary that score by matching weighted terms. Such matching is a fast and transparent proxy, but it cannot resolve whether a term is asserted, quoted, negated, or condemned. These lexicons are also often evaluated on pools enriched with the very examples they retrieve, so a high score partly reflects agreement with the lexicon's own selection rule. Examining a five-language, 2{,}000-item evaluation pool, we find its halves separated almost perfectly by the lexicon itself: every item labeled ``random'' is in fact lexicon-negative, so the lexicon's apparent macro-AUROC of 0.686 collapses to a 0.500 floor fixed by construction. We keep the dictionary's 22-construct ontology but replace term matching with context-reading models, evaluated on a non-circular benchmark that separates unconditional-random, lexicon-positive, and lexicon-negative strata across five languages. Reading the full post rather than the target sentence alone helps most where the lexicon is silent, raising average precision on lexicon-negative text from 0.14 to 0.20, with the largest gains on quoted, implicit, and cross-sentence grievance. Together, these results show that grievance is measured more faithfully by reading the surrounding context, and more honestly when tested on text the lexicon did not select. We release our code and benchmark at https://github.com/behavioral-ds/multilingual_grievance.
△ Less
Submitted 23 July, 2026;
originally announced July 2026.
-
Lipschitzian SLLNs for random functions
Authors:
Lai Tian,
Johannes O. Royset
Abstract:
We prove strong laws of large numbers for locally Lipschitz functions in the Lipschitz pseudometric. Our results hold under either a topological or a model-theoretic condition, with the latter encompassing functions jointly definable in o-minimal structures but extending substantially beyond this class. Applications include uniform convergence of limiting and Clarke subdifferentials and finite-sam…
▽ More
We prove strong laws of large numbers for locally Lipschitz functions in the Lipschitz pseudometric. Our results hold under either a topological or a model-theoretic condition, with the latter encompassing functions jointly definable in o-minimal structures but extending substantially beyond this class. Applications include uniform convergence of limiting and Clarke subdifferentials and finite-sample identification of solutions. Consequently, we identify broad classes of functions for which the failure phenomena revealed by our previous negative results [Tian and Royset, arXiv:2511.16568, 2025] do not occur.
△ Less
Submitted 22 July, 2026;
originally announced July 2026.
-
Boson peak and medium-range elastic heterogeneity in calcium silicate hydrate probed by terahertz spectroscopy and low-temperature calorimetry
Authors:
Xiangyu Li,
Ying Chen,
Jipeng Luo,
Ya Chen,
Linhao Wang,
Lidan Tian,
Gan Ding,
Zeyu Lu,
Zhangli Hu,
Biqin Dong,
Yue Li,
Zongjin Li
Abstract:
The boson peak (BP), a universal vibrational anomaly of disordered solids, has been predicted but not systematically characterized in calcium silicate hydrate (C-S-H), the binding phase of hardened cement. Building on a preliminary terahertz survey, we characterize the BP across five Ca/Si ratios (0.5-1.7) using terahertz time-domain spectroscopy (THz-TDS) and low-temperature calorimetry, two prob…
▽ More
The boson peak (BP), a universal vibrational anomaly of disordered solids, has been predicted but not systematically characterized in calcium silicate hydrate (C-S-H), the binding phase of hardened cement. Building on a preliminary terahertz survey, we characterize the BP across five Ca/Si ratios (0.5-1.7) using terahertz time-domain spectroscopy (THz-TDS) and low-temperature calorimetry, two probes of vibrational dynamics that complement the static picture of conventional structural methods. After Bruggeman correction for crystalline impurities, both probes locate the BP near 1 THz; they agree on frequency but diverge in intensity. The terahertz integrated spectral weight and the calorimetric Cp/T3 peak both fall monotonically with Ca/Si, whereas the apparent terahertz peak height is maximal at Ca/Si = 1.0, where damping is low and oscillator strength still substantial. This decoupling marks a structural crossover between silicate-chain depolymerization and interlayer calcium filling. From the BP we obtain a medium-range dynamical correlation length of order 1 nm (0.3-2 nm) and a coherent-potential elastic-heterogeneity parameter that decreases from gamma = 0.98 to 0.48 as Ca/Si rises; the Debye-normalized BP frequency (nu_BP/nu_D = 0.15-0.17) places C-S-H within the range reported for silicate glasses. Because gamma governs the distribution of energy barriers for local structural rearrangements, it provides a quantitative, composition-resolved descriptor relevant to the intrinsic creep and thermal transport of C-S-H, linking nanoscale vibrational dynamics to the macroscopic durability of concrete. The dual-probe boson-peak approach is transferable to other amorphous solids, including the supplementary cementitious materials of low-carbon cements.
△ Less
Submitted 16 July, 2026;
originally announced July 2026.
-
Average Cause-Specific Hazard: A Censoring-Invariant Measure of Event Burden Under Competing Risks
Authors:
Khondoker Nazmoon Nabi,
Xiang Meng,
Lu Tian,
Jean M Connors,
Deb Schrag,
Hajime Uno
Abstract:
Competing events are common in clinical and epidemiologic studies, including semi-competing risks in which a terminal event such as death may follow a nonfatal event but also competes with it beforehand. Standard summaries include the cumulative incidence function (CIF) and the incidence rate (IR), defined as the number of observed events divided by observed event-free person-time. With competing…
▽ More
Competing events are common in clinical and epidemiologic studies, including semi-competing risks in which a terminal event such as death may follow a nonfatal event but also competes with it beforehand. Standard summaries include the cumulative incidence function (CIF) and the incidence rate (IR), defined as the number of observed events divided by observed event-free person-time. With competing events, the naive IR generally depends on the censoring-time distribution unless intensities are constant. We propose the Average Cause-Specific Hazard (ACSH), a survival-weighted rate per event-free person-time that preserves the interpretation of an incidence rate and is defined purely from the event-time distribution, without involving the censoring-time distribution. We develop nonparametric estimation and inference for ACSH and, for two-sample comparisons, introduce ACSH differences and ratios that provide interpretable contrasts without requiring a strong model assumption between two groups. Simulation studies examine the finite-sample performance, and an analysis of the CANVAS trial illustrates the proposed methods.
△ Less
Submitted 14 July, 2026;
originally announced July 2026.
-
The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy
Authors:
Chunzheng Zhu,
Lei Tian,
Bohan Tan,
Ziqi Zhou,
Yuxuan Sun,
Yijun Wang,
Chengchao Lv,
Yilin Wen,
Yijun He,
Jinghao Lin,
Yihang Chen,
Chee Wei Tan,
Qianshan Wei,
Lei Zhao,
Bin Pu,
Kenli Li,
Yuan Xue,
Jianxin Lin
Abstract:
The growing ability of large language models and vision-language models to jointly interpret and reason over images and text is reshaping medical imaging AI, moving it from task-specific predictors toward autonomous agents that perceive, reason, plan, remember, and act in clinical environments. This survey departs from the capability-first perspective of existing literature and instead begins from…
▽ More
The growing ability of large language models and vision-language models to jointly interpret and reason over images and text is reshaping medical imaging AI, moving it from task-specific predictors toward autonomous agents that perceive, reason, plan, remember, and act in clinical environments. This survey departs from the capability-first perspective of existing literature and instead begins from clinical deployment, asking what tasks, contamination-resistant benchmarks, and interactive training environments are required before medical agents can be trusted in practice. Medical agents are formalized as sequential decision-making systems under partial observability, together with a three-level autonomy taxonomy spanning assisted, cooperative, and fully autonomous operation. The field is organized along a unified scaling spine consisting of framework scaling, capability scaling, and environment scaling. Within this framework, clinical environment scaling, the integration of tools, data, and clinical gyms, is identified as the most actionable yet underexplored direction for agents operating in PACS, EHR, and FHIR ecosystems. Clinical self-evolution, where agents improve through interaction with their environments rather than parameter scaling alone, is further positioned as a key research frontier, drawing insights from self-improving agents, agent gyms, and test-time compute scaling. Applications across radiology, pathology, ophthalmology, and hospital workflows are examined together with deployment challenges including hallucination, cascading failures, and fairness. By consolidating more than 300 references, with particular emphasis on advances from 2025 to 2026, this survey provides a roadmap toward trustworthy, self-improving medical imaging systems for real clinical practice.
△ Less
Submitted 12 August, 2026; v1 submitted 13 July, 2026;
originally announced July 2026.
-
Hidden Decoding at Scale: Latent Computation Scaling for Large Language Models
Authors:
Aiwei Liu,
Cheng Shi,
Chuhan Wu,
Ci Lei,
Di Lu,
Donald He,
Fan Zhang,
Fanhao Kong,
Feifei Zhang,
Guan Wang,
Haicheng Wang,
Haoyu Liu,
Houjin Yu,
Jiachen Ding,
Jiayi Feng,
Jie Zhou,
Jijun Chi,
Jindi Shi,
Jing Lei,
Junjie Zhang,
Laiyi Li,
Le Tian,
Linhao Zhang,
Miao Fan,
Sijun Zhang
, et al. (23 additional authors not shown)
Abstract:
Scaling Large Language Models (LLMs) has been driven mainly by enlarging the Transformer backbone, but for an already-strong model this requires another round of costly pretraining. We study whether an existing backbone can keep improving by allocating more computation to each token while leaving the Transformer backbone fixed. Depth-recurrent (looped) Transformers pursue this goal but are hard to…
▽ More
Scaling Large Language Models (LLMs) has been driven mainly by enlarging the Transformer backbone, but for an already-strong model this requires another round of costly pretraining. We study whether an existing backbone can keep improving by allocating more computation to each token while leaving the Transformer backbone fixed. Depth-recurrent (looped) Transformers pursue this goal but are hard to scale, because looped computation does not fit naturally with the pipeline parallelism used to train the largest models. We add computation along the sequence-length dimension, where the extra computation is simply a longer input and stays compatible with standard large-model training. We propose Hidden Decoding, a sequence-length scaling method applied during continued pretraining (CPT). It expands each token into n streams with independent embedding tables and keeps the intermediate streams' key-value cache as context, so each token performs more internal computation without adding or widening Transformer layers. To keep this affordable at scale, we introduce Stream-Factorized Attention, in which most layers attend only within each stream and only a few layers mix across streams, reducing the attention cost from quadratic to roughly linear in n. Experiments support two scaling results. At frontier scale, we train WeLM-HD4-80B and WeLM-HD4-617B at n=4 and improve their matched non-HD baselines, making Hidden Decoding the first demonstrated sequence-length scaling method at the 100B+ MoE scale. Across expansion factors, the gains grow as n increases, showing that sequence-length expansion is a practical fixed-backbone scaling path for frontier-scale LLMs.
△ Less
Submitted 9 July, 2026;
originally announced July 2026.
-
A Large-Scale Dataset and a New Method for RemoteSensing Traffic Object Segmentation
Authors:
Zhigang Yang,
Huiguang Yao,
Linmao Tian,
Qiang Li,
Qi Wang
Abstract:
Remote sensing imagery plays a crucial role in evaluating regional transportation capacity. However, existing segmentation datasets often lack diversity in object categories and scenes, limiting the ability of models to comprehensively evaluate trans portation capacity in real-world scenes. To alleviate this gap, we construct a large-scale and diverse dataset for transportation object segmentation…
▽ More
Remote sensing imagery plays a crucial role in evaluating regional transportation capacity. However, existing segmentation datasets often lack diversity in object categories and scenes, limiting the ability of models to comprehensively evaluate trans portation capacity in real-world scenes. To alleviate this gap, we construct a large-scale and diverse dataset for transportation object segmentation, named as NWPU-Traffic. This dataset encompass four traffic object categories (car, airplane, ship, and train) and a wide range of scenes from 49 cities across 7 countries, with instance-level annotations to ensure precise segmentation of individual objects, which bridges critical shortcomings in resolution and scene diversity in existing datasets. Leveraging this dataset, we establish a benchmark with several popular segmentation networks. Furthermore, we propose a novel segmentation method that leverages spatial-channel preserving feature interaction and an adaptive feature decoder, enabling robust segmentation across varying scales and complex environments. Extensive experiments and ablation studies validate the effectiveness of our approach. The dataset and code are publicly available at https://github.com/CVer-Yang/NWPU-Traffic.
△ Less
Submitted 4 July, 2026;
originally announced July 2026.
-
Non-synchronism in Global Usage of Research Methods in Library and Information Science from 1990 to 2019
Authors:
Chengzhi Zhang,
Liang Tian
Abstract:
The global development of Library and Information Science (LIS) is influenced by various factors such as the economy, society, culture, discipline, tradition, and more. Consequently, the research methods of LIS vary greatly among countries. To better understand these differences, we conducted a study of 5,281 research papers from 81 countries published in internationally representative journals ov…
▽ More
The global development of Library and Information Science (LIS) is influenced by various factors such as the economy, society, culture, discipline, tradition, and more. Consequently, the research methods of LIS vary greatly among countries. To better understand these differences, we conducted a study of 5,281 research papers from 81 countries published in internationally representative journals over the past thirty years. We manually annotated the research methods used in some articles through content analysis, and subsequently developed and trained a deep learning model for automatic classification of research methods. Using this method, we conducted a comparative analysis of the usage of research methods in different countries. Our findings reveal that there are differences in the research methods used across countries, with each country having its unique research profile and distribution of research methods. Even when investigating the same topic, research methods can differ between countries. Our study also uncovers that there are differences between the national and international distribution of research methods, these differences have decreased over the past 30 years. By highlighting the characteristics of discipline development in various countries from the perspective of research methods, our study can help guide discipline development at the national level. This study provides insights into the usage trends of research methods across different countries and highlights the unique characteristics of discipline development in each country. This information can be valuable in promoting collaboration and understanding between countries and in guiding discipline development at the national level.
△ Less
Submitted 2 July, 2026;
originally announced July 2026.
-
Gender Differences in Research Topic and Method Selection in Library and Information Science: Perspectives from Three Top Journals
Authors:
Chengzhi Zhang,
Siqi Wei,
Yi Zhao,
Liang Tian
Abstract:
Research in the social sciences has shown that there are gender differences in the selection of research methods, with women often opting for qualitative methods while men prefer quantitative methods. However, it is important to consider that research methods are generally chosen based on the research topic. To figure out the influence of gender on research method selection, a study was conducted…
▽ More
Research in the social sciences has shown that there are gender differences in the selection of research methods, with women often opting for qualitative methods while men prefer quantitative methods. However, it is important to consider that research methods are generally chosen based on the research topic. To figure out the influence of gender on research method selection, a study was conducted in the field of Library and Information Science, using a more fine-grained method classification system and an automatic classification model called CogFT, which is based on full-text cognition. The findings showed that women tend to use Interview while men prefer Theoretical approach, across a range of topics. The study offers insights into the specific research design processes that contribute to gender differences in method selection and suggests ways to promoting gender inclusivity and equality in academia by considering research method use and guidance.
△ Less
Submitted 2 July, 2026;
originally announced July 2026.
-
When is vaccine prioritization worth optimizing?
Authors:
Mi Feng,
Zhaohua Lin,
Changsong Zhou,
Liang Tian
Abstract:
Optimizing vaccine prioritization is often treated as the default policy response when vaccine supply is limited. Yet optimized prioritization carries administrative, ethical and communication costs, motivating an upstream question: whether differences among vaccine allocations can alter epidemic outcomes enough to make optimization epidemiologically necessary. We show that optimization is not alw…
▽ More
Optimizing vaccine prioritization is often treated as the default policy response when vaccine supply is limited. Yet optimized prioritization carries administrative, ethical and communication costs, motivating an upstream question: whether differences among vaccine allocations can alter epidemic outcomes enough to make optimization epidemiologically necessary. We show that optimization is not always worth pursuing: in some regimes, vaccination markedly reduces epidemic burden, but many feasible allocation rules perform almost equally well, making the necessity of optimization low. We quantify this necessity as the range of epidemic outcomes generated by different allocations under fixed supply and show that it is governed by competition between vaccinating high-contact groups to slow transmission and vaccinating groups that benefit most directly: necessity is low when these protection routes are balanced and high when one dominates. Increasing transmission intensity changes this balance and drives a transition in the optimal allocation from transmission-focused prioritization toward direct protection. Different prevention objectives exhibit distinct transition thresholds, creating regimes in which optimizing one objective substantially compromises another, thereby revealing when the choice of prevention target matters most. This framework reframes vaccine prioritization as a prior decision problem, identifying when optimization is warranted, when simpler rules suffice, and when prevention goals conflict.
△ Less
Submitted 1 July, 2026;
originally announced July 2026.
-
Usage frequency and application variety of research methods in library and information science: Continuous investigation from 1991 to 2021
Authors:
Chengzhi Zhang,
Liang Tian,
Heting Chu
Abstract:
The present study analyzed over 26,000 research articles published between 1991 and 2021 in twenty-one major LIS (Library and Information Science) journals, using the machine learning (ML) approach to categorize the research methods used by LIS scholars. The findings of this study are significant. Firstly, there has been a shift in the research strategy from conceptual research (e.g., "Theoretical…
▽ More
The present study analyzed over 26,000 research articles published between 1991 and 2021 in twenty-one major LIS (Library and Information Science) journals, using the machine learning (ML) approach to categorize the research methods used by LIS scholars. The findings of this study are significant. Firstly, there has been a shift in the research strategy from conceptual research (e.g., "Theoretical approach") to empirical research (e.g., "Interview") in LIS investigations over the past 31 years. Secondly, the research topics explored by LIS scholars during this period have moved from system-centered issues (e.g., "Information retrieval/models and algorithms") to user-centered topics (e.g., "Information services "). Thirdly, the study revealed dynamic and revealing relationships between the 18 research topics identified in the study and the 16 research methods commonly adopted in the LIS field. These dynamic relationships can be visualized by year and longitudinally via an interactive map created in this study.
△ Less
Submitted 29 June, 2026;
originally announced June 2026.
-
A Mapping Sheath with Thermally Drawn Multi-Electrode Basket for Cardiac Electrophysiological Recording and Ablation Catheter Delivery
Authors:
Qindong Zheng,
Anil Demircali,
Jinshi Zhao,
Xiaotong Guo,
Libaihe Tian,
Oliver Jones,
Jamie Kay,
Shengzhe Li,
Alex Ranne,
Elaine Lim,
Huiyi Wu,
Simos Koutsoftidis,
Mohamed Abdelaziz,
Emmanuel Drakakis,
Prapa Kanagaratnam,
Nick Linton,
Burak Temelkuran
Abstract:
Cardiac arrhythmias, particularly atrial fibrillation, represent a major cardiovascular health burden and underscore the need for efficient and integrated strategies for electrical mapping and targeted therapy. Cardiac electrophysiology procedures depend on accurate identification of arrhythmogenic substrates followed by timely catheter ablation, but conventional diagnostic and therapeutic devices…
▽ More
Cardiac arrhythmias, particularly atrial fibrillation, represent a major cardiovascular health burden and underscore the need for efficient and integrated strategies for electrical mapping and targeted therapy. Cardiac electrophysiology procedures depend on accurate identification of arrhythmogenic substrates followed by timely catheter ablation, but conventional diagnostic and therapeutic devices remain separate, often requiring repeated catheter exchanges and multiple access routes. Here, we report an adaptable strategy for functionalizing hollow-core sheaths with EP mapping capabilities, integrating multielectrode recording and ablation catheter delivery within a single compact platform. The device leverages thermal drawing to enable complex geometric fabrication, miniaturization, rapid prototyping, and scalable manufacturing of ultrathin electrode splines arranged circumferentially at the distal end to form an adjustable basket. The mapping sheath exhibited mechanical and electrophysiological properties suitable for intracardiac navigation and electrogram recording in bench-top evaluations, an in vitro left atrial phantom study, and ex vivo Langendorff-perfused porcine heart testing. In vivo porcine studies further demonstrated translational feasibility through vascular introduction, fluoroscopic visualization, intracardiac deployment, tissue contact, electrogram acquisition, and reconstruction of voltage and activation maps. These results support the development of intracardiac platforms with an adapted manufacturing approach, potentially guiding advances in agile cardiac mapping and ablation.
△ Less
Submitted 28 June, 2026;
originally announced June 2026.
-
Light-driven active phase separation and droplet division
Authors:
Zi Lin,
Thomas Beneyton,
Suzanne Lafon,
Edison Rafael Jimenez Granda,
Liangfei Tian,
Alexandre Baron,
Jean-Christophe Baret,
Nicolas Martin
Abstract:
Phase separation organizes matter across scales, yet how it operates under sustained energy input remains poorly understood. Experimental approaches to driven phase separation have largely relied on chemically fueled systems, in which reaction fluxes are intrinsically coupled to fuel consumption and reaction-network complexity. Here we show that continuous molecular switching alone is sufficient t…
▽ More
Phase separation organizes matter across scales, yet how it operates under sustained energy input remains poorly understood. Experimental approaches to driven phase separation have largely relied on chemically fueled systems, in which reaction fluxes are intrinsically coupled to fuel consumption and reaction-network complexity. Here we show that continuous molecular switching alone is sufficient to generate active phase behavior in a minimal two-phase system. Using light-responsive DNA-azobenzene coacervates confined in microfluidic droplets, we modulate intermolecular interactions with spatiotemporal precision and quantitatively track phase separation dynamics under illumination. Light-driven azobenzene isomerization controls both thermodynamics and kinetics, setting phase boundaries and regulating dissolution and nucleation rates. Under single-wavelength illumination that couples forward and backward isomerization into a dynamic photostationary state, coarsening is arrested and micron-sized coacervates are stabilized. When the two photoisomerization pathways are driven independently, spatially unbalanced reaction fluxes generate sustained interfacial instabilities, including surface undulations, budding, and division. These behaviors arise from a physical coupling between reaction kinetics and phase separation, without chemical fuels or biochemical regulation. Our results show that non-equilibrium phase behavior is governed by how opposing reaction fluxes are imposed, establishing reversible molecular switching as a minimal route to active materials from equilibrium building blocks.
△ Less
Submitted 25 June, 2026;
originally announced June 2026.
-
Degenerate Stochastic Delay Modified Equation: Approximation of Stochastic Variance Reduced Gradient
Authors:
Ting Zhang,
Lulu Tian,
Yue Ding
Abstract:
According to the property of stochastic variance reduced gradient (SVRG) algorithms, we construct a class of degenerate stochastic delay modified equations (SDMEs). Using the Lindeberg principle and the Markov property, we approximate the SVRG algorithms by the corresponding SDMEs. We obtain order 1 weak approximation in the smooth Wasserstein distance and the theory is validated through numerical…
▽ More
According to the property of stochastic variance reduced gradient (SVRG) algorithms, we construct a class of degenerate stochastic delay modified equations (SDMEs). Using the Lindeberg principle and the Markov property, we approximate the SVRG algorithms by the corresponding SDMEs. We obtain order 1 weak approximation in the smooth Wasserstein distance and the theory is validated through numerical experiments.
△ Less
Submitted 20 June, 2026;
originally announced June 2026.
-
SAMA: Semantic Anchor-aligned Augmentation for Unified Low-Resource Multimodal Information Extraction
Authors:
Quanjiang Guo,
Chong Mu,
Jiazhou Pan,
Ming Jia,
Ling Tian,
Hui Gao,
Zhao Kang
Abstract:
Multimodal Information Extraction (MIE)-covering tasks such as Multimodal Named Entity Recognition (MNER), Relation Extraction (MRE), and Event Extraction (MEE)-is essential for understanding multimedia content but remains constrained by severe data scarcity. Although data augmentation is a promising remedy, existing approaches are impeded by coarse cross-modal alignment and fragmented, task-speci…
▽ More
Multimodal Information Extraction (MIE)-covering tasks such as Multimodal Named Entity Recognition (MNER), Relation Extraction (MRE), and Event Extraction (MEE)-is essential for understanding multimedia content but remains constrained by severe data scarcity. Although data augmentation is a promising remedy, existing approaches are impeded by coarse cross-modal alignment and fragmented, task-specific designs that fail to exploit shared semantic knowledge. To overcome these limitations, we introduce Semantic Anchor-aligned Multimodal Augmentation (SAMA), a unified framework for generating high-fidelity, task-aware synthetic data. SAMA constructs structured semantic anchors from ground-truth labels to guide a Collaborative Multi-Experts Multimodal Large Language Model (CME-MLLM), which integrates a Universal Adapter for shared semantics with Task-Specific Adapters to produce diverse yet constraint-compliant textual samples. For image synthesis, SAMA employs an Anchor-Preserving Diffusion mechanism that uses anchor-weighted prompts and latent conditioning to maintain critical semantic anchors while diversifying visual contexts. To eliminate the need for manual verification, SAMA further introduces a Dual-Constraint Filtering module that selects synthetic samples based on both cross-modal consistency and anchor fidelity. Extensive experiments across benchmark datasets for MNER, MRE, and MEE demonstrate that SAMA consistently outperforms state-of-the-art augmentation baselines under both fully supervised and low-resource settings, underscoring its versatility, robustness, and effectiveness.
△ Less
Submitted 17 June, 2026;
originally announced June 2026.
-
Intrinsic Nonreciprocity in Electron-Phonon Interaction Driven Thermoelectric Diodes
Authors:
Hao-Kun Ke,
Lie-Run Tian,
Pei-Hao Fu,
Jun-Feng Liu,
Jun Wang,
H. Xu
Abstract:
We study an electron-phonon interaction driven thermoelectric diode. The nonreciprocity in this diode arises from the asymmetry between the probabilities of phonon emission and absorption in the electron-phonon interaction, as well as the structural reflection asymmetry. We reveal the intrinsic nature of this nonreciprocity, as the forward and backward electron transport remains asymmetric even wh…
▽ More
We study an electron-phonon interaction driven thermoelectric diode. The nonreciprocity in this diode arises from the asymmetry between the probabilities of phonon emission and absorption in the electron-phonon interaction, as well as the structural reflection asymmetry. We reveal the intrinsic nature of this nonreciprocity, as the forward and backward electron transport remains asymmetric even when the applied temperature difference is not reversed. This intrinsic nonreciprocity gives rise to two novel transport phenomena. One is a novel thermoelectric effect which is driven by the temperature difference between the leads and the central device region, rather than the conventional temperature difference between the two leads. The second, and more significant, phenomenon is the suppression of electronic backscattering in the load resistor. This suppression decreases the resistance of the load resistor, which leads to the breakdown of Ohm's addition law. Under suitable conditions, the presence of electron-phonon interaction can yield a larger thermoelectric current compared to the case without it. This intrinsic nonreciprocity opens up a new pathway for low-power electronics besides topology and superconductivity, and for nonreciprocal thermoelectric devices.
△ Less
Submitted 10 June, 2026; v1 submitted 10 June, 2026;
originally announced June 2026.
-
Long Live Fine-Tuning: Task-Specific Transformers Outperform Zero-Shot LLMs for Misinformation Response Classification on Reddit
Authors:
JooYoung Lee,
Lin Tian,
Angela Brillantes,
Adriana-Simona Mihăiţă,
Marian-Andrei Rizoiu
Abstract:
As large language models (LLMs) become default tools for online information verification, an implicit assumption follows them: that scale and general capability are sufficient for nuanced classification of misinformation discourse. We test this assumption directly on 900 Reddit comments spanning three PolitiFact-verified misinformation claims (environment, health, immigration), labelled as belief…
▽ More
As large language models (LLMs) become default tools for online information verification, an implicit assumption follows them: that scale and general capability are sufficient for nuanced classification of misinformation discourse. We test this assumption directly on 900 Reddit comments spanning three PolitiFact-verified misinformation claims (environment, health, immigration), labelled as belief (propagates the claim), fact-check (corrects it), or other. We compare nine models across three paradigms -- BART-MNLI, three Llama variants, three commercial frontier LLMs (Claude Haiku 4.5, Gemini Flash Lite 2.5, Claude Sonnet 4.6), and fine-tuned DistilBERT and RoBERTa -- under universal and topic-specific label schemas.
The assumption does not hold. Fine-tuned RoBERTa reaches 0.62 macro-$F_1$ against a best zero-shot result of 0.50 (Claude Haiku 4.5), at a fraction of the per-query cost; the supervised advantage is concentrated on the belief class, the implicit, affective category every zero-shot model under-detects. Scaling does not help: Llama-3-8B matches Llama-3-70B, and Claude Sonnet 4.6 underperforms the smaller Haiku under generic labels, collapsing belief detection to 0.17 and refusing outright on a subset of comments flagged as sensitive. This is a safety-alignment artefact, not a capacity limit. Label schema and topic jointly shape zero-shot performance, with the same model varying by more than 0.13 macro-$F_1$ across topics under matched labels. In a verification context, where missing belief is the costlier error, task-specific fine-tuning remains the more reliable choice despite the proliferation of large generative models.
△ Less
Submitted 2 June, 2026;
originally announced June 2026.
-
Towards Bulk Locality: A Systematic Construction of Contact Interactions from Chord Diagrams
Authors:
Hao Dai,
Yi-Li Wang,
Yu-Ge Chen,
Li-Guo Qin,
Li-Jun Tian
Abstract:
Chord diagrams encode boundary correlators in the double-scaled holographic Sachdev-Ye-Kitaev model, but currently capture only a limited class of bulk interactions that yield pure power-law correlators. In this article, we investigate a general construction based on Fock-space flux models with arbitrary periodic lattice size, clarifies how lattice dimensions control probe configurations and bulk…
▽ More
Chord diagrams encode boundary correlators in the double-scaled holographic Sachdev-Ye-Kitaev model, but currently capture only a limited class of bulk interactions that yield pure power-law correlators. In this article, we investigate a general construction based on Fock-space flux models with arbitrary periodic lattice size, clarifies how lattice dimensions control probe configurations and bulk contact vertices. Developing a systematic matching scheme and using the chord path integral formalism, we compute three- to six-point contact correlators and reproduce a broad class of AdS$_2$ scalar contact Witten diagrams, including those with logarithmic singularities. The results demonstrate that chord diagrams, in full generality, provide a microscopic description of bulk contact interactions and thereby establish a principled framework for reconstructing bulk locality from boundary data.
△ Less
Submitted 29 May, 2026;
originally announced May 2026.
-
The Empirical Content of Revealed Preference in High Dimensions
Authors:
Ian Crawford,
Longye Tian
Abstract:
We examine how the empirical content of revealed preference theory depends on the dimensionality of the choice environment. While higher-dimensional choice problems may appear more demanding, we show that revealed preference restrictions become less informative. Using Selten's Area measure, we establish that for any fixed number of observations, the empirical content of GARP converges to zero expo…
▽ More
We examine how the empirical content of revealed preference theory depends on the dimensionality of the choice environment. While higher-dimensional choice problems may appear more demanding, we show that revealed preference restrictions become less informative. Using Selten's Area measure, we establish that for any fixed number of observations, the empirical content of GARP converges to zero exponentially fast in the number of goods. We provide complementary proofs based on revealed preference graphs and the Afriat inequalities, and show in simulations calibrated to scanner data that the effect is quantitatively large. We also evaluate potential responses in observational and experimental settings and find that, while these can slow the rate, they do not eliminate this loss of empirical content.
△ Less
Submitted 28 May, 2026;
originally announced May 2026.
-
StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration
Authors:
Linrui Tian,
Qi Wang,
Bang Zhang
Abstract:
Real-time streaming joint audio-video generation for character animation requires a generator to speak the requested transcript, maintain visual identity across chunks, and run within a strict playback budget. These requirements are difficult to satisfy simultaneously: chunk-wise autoregressive generation can accumulate transcript-audio misalignment and visual drift, while the few-step distillatio…
▽ More
Real-time streaming joint audio-video generation for character animation requires a generator to speak the requested transcript, maintain visual identity across chunks, and run within a strict playback budget. These requirements are difficult to satisfy simultaneously: chunk-wise autoregressive generation can accumulate transcript-audio misalignment and visual drift, while the few-step distillation needed for low latency often degrades spatial diversity and temporal quality. We present StreamChar, a streaming framework that separates long-horizon orchestration from short-window audio-video denoising. An LLM-based orchestrator uses the transcript and historical context to produce frame-aligned audio conditions, and a joint audio-video DiT performs local bidirectional denoising with reference and motion-frame conditioning. For efficient deployment, we use a two-stage distillation pipeline that first compresses the sampler and then fine-tunes the student under online chunk rollouts. A progress-aware pointer aligns partial transcripts with generated audio during rollout training, and a sink-chunk memory provides a persistent visual anchor for reducing long-horizon drift. Experiments on short-clip and long-horizon protocols show that StreamChar runs in real time on a single H100 GPU and provides a favorable system-level trade-off among transcript fidelity, audio-visual synchronization, visual quality, and streaming stability compared with recent joint and audio-driven baselines.
△ Less
Submitted 25 May, 2026;
originally announced May 2026.
-
Hypergraph as Language
Authors:
Mengqi Lei,
Guohuan Xie,
Shihui Ying,
Shaoyi Du,
Jun-Hai Yong,
Chuan Shi,
Ling Tian,
Siqi Li,
Yue Gao
Abstract:
Large language models (LLMs) have recently shown strong potential in modeling relational structures. However, existing approaches remain fundamentally graph-centric: they focus on processing pairwise graph structures into tokens that LLMs can understand. In contrast, many real-world relational patterns do not naturally conform to the pairwise-edge assumption, and are better modeled as high-order a…
▽ More
Large language models (LLMs) have recently shown strong potential in modeling relational structures. However, existing approaches remain fundamentally graph-centric: they focus on processing pairwise graph structures into tokens that LLMs can understand. In contrast, many real-world relational patterns do not naturally conform to the pairwise-edge assumption, and are better modeled as high-order associations in hypergraphs. For hypergraph structures, existing methods often fail to preserve the native semantics that multiple objects are jointly connected by the same high-order relation, limiting their ability to exploit complex structures. To address this limitation, we put forth the "Hypergraph as Language" perspective and propose Hyper-Align, a hypergraph-native alignment framework for large language models. Hyper-Align compiles the query-object-centered hypergraph context into hypergraph tokens directly consumable by a base LLM. Specifically, we introduce Hypergraph Incidence Detail Template with Overview (HIDT-O), which serializes high-order association structures into a fixed-shape hybrid template combining local incidence details and overview-level summaries. We then design a Hypergraph Incidence Projector (HIP), which maps native high-order incidence structures into the LLM token space through explicit semantic-structural decoupling and bidirectional message passing between vertices and hyperedges. We further define a concrete Hypergraph-as-Language input protocol, which jointly feeds hypergraph tokens and textual prompts into a frozen base LLM, supporting both vertex-level and hyperedge-level tasks under a unified question-answering paradigm. To systematically evaluate different methods in hypergraph structural modeling, we introduce HyperAlign-Bench. Extensive experiments show that Hyper-Align significantly outperforms existing methods across in-domain and zero-shot evaluations.
△ Less
Submitted 15 August, 2026; v1 submitted 20 May, 2026;
originally announced May 2026.
-
KIO-planner: Attention-Guided Single-Stage Motion Planning with Dual Mapping for UAV Navigation
Authors:
Dexing Yao,
Haochen Li,
Junhao Wei,
Yifu Zhao,
Yanxiao Li,
Jiahui Xu,
Jinxuan Hu,
Lele Tian,
Baili Lu,
Zikun Li,
Xu Yang,
Sio-Kei Im,
Dingcheng Yang,
Yapeng Wang
Abstract:
Autonomous UAV flight in confined, wall-dense environments requires low-latency and reliable motion planning under strict safety constraints. Traditional optimization-based planners suffer from mapping latency and easily fall into local minima when navigating through dense structural obstacles. Meanwhile, existing end-to-end learning methods struggle to extract fine-grained geometric features from…
▽ More
Autonomous UAV flight in confined, wall-dense environments requires low-latency and reliable motion planning under strict safety constraints. Traditional optimization-based planners suffer from mapping latency and easily fall into local minima when navigating through dense structural obstacles. Meanwhile, existing end-to-end learning methods struggle to extract fine-grained geometric features from raw depth images and lack hard kinodynamic constraints, leading to unpredictable collisions near walls. To address these issues, we propose KIO-planner, an attention-guided single-stage trajectory planning framework. First, we integrate a Convolutional Block Attention Module (CBAM) into the perception backbone to adaptively focus on critical structural edges and traversable space. Second, we introduce a novel Dual Mapping mechanism--comprising physical bounds activation and a deterministic Geometric Safety Shield in the depth-pixel space--to enforce kinodynamic feasibility and collision-free flight without global map fusion. Extensive high-fidelity simulated experiments demonstrate that KIO-planner enables highly agile navigation at speeds up to 3.0 m/s. Compared to the state-of-the-art baseline, KIO-planner achieves lower inference latency (approximately 24 ms) and generates significantly smoother trajectories, reducing control cost by 28.4%. Most notably, our Dual Mapping substantially increases the worst-case safety margin, measured by minimum distance to obstacles, from 0.48 m to 0.76 m, ensuring fast, smooth, and safer navigation in highly constrained environments.
△ Less
Submitted 19 May, 2026;
originally announced May 2026.
-
Learning Interpretable Point-Based Clinical Risk Scores via Direct Optimization
Authors:
Ying Cui,
Albert M Li,
Vivek Charu,
Yeon-Mi Hwang,
Tina Hernandez-Boussard,
Lu Tian
Abstract:
Many clinical risk scores are deployed as additive rules with nonnegative integer points assigned to relevant binary predictive features. These integer weights not only make the score easier to use in practice but also promote sparsity in the resulting prediction model. Such risk scores are often derived by first fitting a regression model and then rounding the estimated coefficients to the neares…
▽ More
Many clinical risk scores are deployed as additive rules with nonnegative integer points assigned to relevant binary predictive features. These integer weights not only make the score easier to use in practice but also promote sparsity in the resulting prediction model. Such risk scores are often derived by first fitting a regression model and then rounding the estimated coefficients to the nearest integer after appropriate scaling. This approach is computationally fast but does not guarantee optimality of the resulting score. Alternatively, one may search over all possible integer weights to directly optimize a value function by posing the problem as an integer programming task. However, the associated computational burden can be substantial, especially when the value function is nonconcave or even discontinuous. In this paper, we develop new machine learning algorithms that employ a flexible greedy optimization strategy to learn such additive scoring directly under explicit and sensible optimality objectives. We apply the proposed method to a large electronic health record (EHR) cohort in Epic Cosmos to construct an integer-weighted comorbidity score for measuring the risk of post-discharge mortality. We also conduct a simulation study to examine the finite-sample operating characteristics.
△ Less
Submitted 18 May, 2026;
originally announced May 2026.
-
TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing
Authors:
Pei Yang,
Wanyi Chen,
Tongyun Yang,
Pengbin Feng,
Jiarong Xing,
Wentao Guo,
Yuhang Yao,
Yuhang Han,
Hanchen Li,
Xu Wang,
Zeyu Wang,
Jie Xiao,
Anjie Yang,
Liang Tian,
Lynn Ai,
Eric Yang,
Tianyu Shi
Abstract:
LLM routing matters most in long-horizon applications such as coding agents, deep research systems, and computer-use agents, where a single user request triggers many model calls. Routing each call to the cheapest sufficient model can cut costs without sacrificing quality, yet existing router benchmarks evaluate routers only on one-shot prompts. They never expose the router-visible prefix at an in…
▽ More
LLM routing matters most in long-horizon applications such as coding agents, deep research systems, and computer-use agents, where a single user request triggers many model calls. Routing each call to the cheapest sufficient model can cut costs without sacrificing quality, yet existing router benchmarks evaluate routers only on one-shot prompts. They never expose the router-visible prefix at an intermediate agent step, never test whether a cheaper replacement preserves downstream task success, and often rely on online LLM judges at evaluation time. We introduce TwinRouterBench, a step-level routing benchmark with two tracks. The static track provides 970 router-visible prefixes from 520 instances across SWE-bench, BFCL, mtRAG, QMSum, and PinchBench, each paired with an execution-verified target tier estimated under a released downgrade-and-cascade protocol; scoring is deterministic arithmetic over tier labels, trajectory membership, and token costs, with no online evaluator-side LLM judge. The dynamic track supplies a harness that runs routers on the full 500-case SWE-bench Verified suite; in this paper we report a 100-case held-out evaluation disjoint from the static SWE supervision split. At each LLM call the router selects a concrete model from a locked pool, and success is measured by official task resolution and realized API spend. The two tracks support fast offline iteration followed by end-to-end validation under live agent execution. Code and data are available at https://github.com/CommonstackAI/TwinRouterBench.
△ Less
Submitted 21 May, 2026; v1 submitted 14 May, 2026;
originally announced May 2026.
-
"I'm Not Mad, Just Focused'': Understanding Human Emotions in Human-Robot Collaboration
Authors:
Seung Chan Hong,
Dana Kulić,
Leimin Tian
Abstract:
Human-robot collaboration (HRC) can benefit from robots' abilities to interpret human emotional states. However, current emotion recognition (ER) models in HRC often fall short, particularly due to their reliance on acted datasets and single-modality inputs like facial expressions. We propose a novel vision language model (VLM)-based ER system that leverages contextual understanding to improve emo…
▽ More
Human-robot collaboration (HRC) can benefit from robots' abilities to interpret human emotional states. However, current emotion recognition (ER) models in HRC often fall short, particularly due to their reliance on acted datasets and single-modality inputs like facial expressions. We propose a novel vision language model (VLM)-based ER system that leverages contextual understanding to improve emotion interpretation in HRC. We first evaluate the VLM-ER system by assessing its semantic and sentiment similarity with human annotations on an existing HRC dataset. Then, in a user study with a service robot in a collaborative delivery task, we evaluate the effects of modulating the robot's behaviour based on the user's emotional state inferred by the VLM-ER system. The results show that the proposed VLM-ER system achieves higher semantic similarity and positive sentiment alignment with human annotations compared to a baseline convolutional neural network-based system. Further, participants in the user study preferred emotion-adaptive robot behaviour facilitated by the VLM-ER system.
△ Less
Submitted 16 May, 2026;
originally announced May 2026.
-
DeepFilters: Scattering-Aware Pupil Engineering with Learned Digital Filter Reconstruction for Extended Depth of Field Microscopy
Authors:
Joseph L. Greene,
Suet YIng Chan,
Qilin Deng,
Jeffrey Alido,
Alexandra Lion,
Guorong Hu,
Ruipeng Guo,
Tongyu Li,
Kivilcim Kiliç,
Ian Davison,
Lei Tian
Abstract:
Extended depth of field microscopy encodes axial information into a single acquisition through engineered point spread functions, but conventional and deep optics approaches are subject to degradation in scattering tissue. We introduce DeepFilters, a scattering-aware deep optics framework that jointly optimizes a parameterized pupil filter and a digital-filter-based reconstruction network through…
▽ More
Extended depth of field microscopy encodes axial information into a single acquisition through engineered point spread functions, but conventional and deep optics approaches are subject to degradation in scattering tissue. We introduce DeepFilters, a scattering-aware deep optics framework that jointly optimizes a parameterized pupil filter and a digital-filter-based reconstruction network through a calibrated differentiable forward model to achieve broad generalization without retraining. Incorporating empirical scattering kernels, physics-guided regularization, and a hybrid genetic-gradient initialization strategy, DeepFilters extends the PSF from 16 micron to >400 micron in clear media and enables signal recovery beyond 120 micron deep in biological tissues, validated across fixed brain slices and sea urchin embryos.
△ Less
Submitted 13 May, 2026;
originally announced May 2026.
-
Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime
Authors:
Tianshu Zhu,
Wenyu Zhang,
Xiaoying Zuo,
Lun Tian,
Haotian Zhao,
Yucheng Zeng,
Jingnan Gu,
Daxiang Dong,
Jianmin Wu,
Dawei Yin,
Dou Shen
Abstract:
Agentic reinforcement learning (RL) for software engineering spends much of its compute on stateful trajectories whose grouped binary rewards are highly skewed and weakly contrastive. We frame this as pass-rate control and show that the binary reward-side signal is strongest near a 50% rollout pass rate under four criteria: reward entropy, group-filtering survival, leave-one-out (RLOO) advantage e…
▽ More
Agentic reinforcement learning (RL) for software engineering spends much of its compute on stateful trajectories whose grouped binary rewards are highly skewed and weakly contrastive. We frame this as pass-rate control and show that the binary reward-side signal is strongest near a 50% rollout pass rate under four criteria: reward entropy, group-filtering survival, leave-one-out (RLOO) advantage energy under Group Relative Policy Optimization (GRPO), and success-failure pair count. We propose Prefix Sampling (PS), which replays self-generated trajectory prefixes to steer skewed groups toward this regime: successful prefixes give mostly failing groups a head start, while failing prefixes handicap mostly passing groups. Replayed states are reconstructed through the existing rollout path, and replayed tokens are masked from the loss so optimization applies only to current-policy continuations. On SWE-bench Verified, PS reaches the baseline high-score regime within evaluation variability while delivering 2.01x and 1.55x end-to-end wall-clock speedups on Qwen3-14B and Qwen3-32B; the 14B peak improves from 0.274 to 0.295. AIME 2025 experiments on 4B and 8B show the same pass-rate-control pattern, and 4B ablations attribute gains to replay, bidirectional coverage, and adaptive control.
△ Less
Submitted 15 May, 2026; v1 submitted 6 May, 2026;
originally announced May 2026.
-
ARIS: Agentic and Relationship Intelligence System for Social Robots
Authors:
Stavya Datta,
Fucai Ke,
Leimin Tian,
Hamid Rezatofighi
Abstract:
Foundational models have advanced social robotics, enabling richer perception and communicative interaction with users. However, current systems still struggle with multi-turn engagement, social-relationship reasoning, and contextually grounded dialogue at scale. We present ARIS (Agentic and Relationship Intelligence System), an agentic AI framework that unifies multimodal reasoning, a graph-based…
▽ More
Foundational models have advanced social robotics, enabling richer perception and communicative interaction with users. However, current systems still struggle with multi-turn engagement, social-relationship reasoning, and contextually grounded dialogue at scale. We present ARIS (Agentic and Relationship Intelligence System), an agentic AI framework that unifies multimodal reasoning, a graph-based Social World Model, and retrieval-augmented generation (RAG) within a single modular architecture for social robots. We evaluate ARIS with the Pepper robot in a robot-mediated dyadic conversational setting, comparing it against a large language model baseline. A user study (N=23) shows that ARIS yields significantly higher perceived intelligence, animacy, anthropomorphism, and likeability. Our contributions are threefold: (1)~a Social World Model that explicitly maps and updates social relationships between users through a knowledge graph, enabling social reasoning and re-identification across encounters; (2)~an efficient RAG-based conversational pipeline that maintains bounded latency as dialogue histories grow to thousands of exchanges while preserving response relevance; and (3)~system integration and empirical validation of these components within a modular agentic architecture that coordinates speech, vision, and physical action through structured APIs. The implementation of ARIS will be released as open source upon publication.
△ Less
Submitted 1 May, 2026;
originally announced May 2026.
-
AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning
Authors:
Haotian Zhao,
Songlin Zhou,
Yuxin Zhang,
Stephen S. -T. Yau,
Wenyu Zhang,
Lun Tian,
Tianshu Zhu,
Yifeng Huang,
Yucheng Zeng,
Jingnan Gu,
Daxiang Dong,
Jianmin Wu
Abstract:
Reinforcement learning (RL) has substantially improved the ability of large language model (LLM) agents to interact with environments and solve multi-turn tasks. However, effective agentic RL remains challenging: sparse outcome-only rewards provide limited guidance for assigning credit to individual steps within long interaction trajectories. Existing approaches often introduce dense intermediate…
▽ More
Reinforcement learning (RL) has substantially improved the ability of large language model (LLM) agents to interact with environments and solve multi-turn tasks. However, effective agentic RL remains challenging: sparse outcome-only rewards provide limited guidance for assigning credit to individual steps within long interaction trajectories. Existing approaches often introduce dense intermediate supervision, such as process reward models or auxiliary self-supervised signals, which increases supervision and tuning complexity and may limit generalization across tasks and domains. We present AEM, a supervision-free credit assignment method that adaptively modulates entropy dynamics during RL training to improve the exploration-exploitation trade-off. Since in agentic RL the environment is typically affected by a complete response, rather than an individual token, our analysis lifts entropy dynamics from the token level to the response level, aligning uncertainty estimation with the effective action granularity of LLM agents and reducing sensitivity to token-level sampling noise. We further show that entropy drift under natural-gradient updates is governed by the interaction between the sampled-response advantage and its relative surprisal. Motivated by this result, AEM derives a practical response-level uncertainty proxy and uses it to rescale advantages, leveraging the evolving balance between positive and negative samples to naturally transition from exploration to exploitation. Extensive experiments on ALFWorld, WebShop, and SWE-bench-Verified with models ranging from 1.5B to 32B demonstrate that AEM consistently improves strong RL baselines, including a +1.4\% gain when integrated into a state-of-the-art software-engineering RL training framework.
△ Less
Submitted 8 May, 2026; v1 submitted 1 May, 2026;
originally announced May 2026.
-
Effect of neutron-proton asymmetry on the $^3$H clustering in Boron isotopes
Authors:
J. L. Jin,
Q. Zhao,
P. J. Li,
M. Kimura,
D. Beaumel,
B. Zhou,
J. L. Tian
Abstract:
To investigate the influence of neutron-proton asymmetry on the formation of asymmetric clusters, we perform a systematic comparative study of $^{3}$H and $α$ cluster preformation in the Boron isotopic chain ($^{11-14}$B). Within the framework of Antisymmetrized Molecular Dynamics (AMD), we compute the nuclear wave functions and subsequently extract the reduced width amplitudes (RWA) and spectrosc…
▽ More
To investigate the influence of neutron-proton asymmetry on the formation of asymmetric clusters, we perform a systematic comparative study of $^{3}$H and $α$ cluster preformation in the Boron isotopic chain ($^{11-14}$B). Within the framework of Antisymmetrized Molecular Dynamics (AMD), we compute the nuclear wave functions and subsequently extract the reduced width amplitudes (RWA) and spectroscopic factors (SF). The results show that the $α$ cluster SF exhibits a monotonic decrease with increasing neutron number, consistent with the established suppression effect of the neutron skin. In contrast, the $^{3}$H cluster SF displays a non-monotonic behavior, peaking at $^{12}$B. This distinct trend indicates that the formation of the asymmetric $^{3}$H cluster is subject to a competition between suppression from the neutron skin and an enhancement driven by the neutron-proton asymmetry of the parent nucleus. We successfully isolate this enhancement effect by analyzing the ratio of the SFs, SF($^{3}$H)/SF($α$). This approach not only quantifies the enhancement but also proposes the SF ratio as a robust experimental observable for probing insights into asymmetric clustering phenomena.
△ Less
Submitted 27 April, 2026;
originally announced April 2026.
-
One-Step Diffusion with Inverse Residual Fields for Unsupervised Industrial Anomaly Detection
Authors:
Boan Zhang,
Wen Li,
Guanhua Yu,
Xiyang Liu,
Wenchao Chen,
Long Tian
Abstract:
Diffusion models have achieved outstanding performance in unsupervised industrial anomaly detection (uIAD) by learning a manifold of normal data under the common assumption that off-manifold anomalies are harder to generate, resulting in larger reconstruction errors in data space or lower probability densities in the tractable latent space. However, their iterative denoising and noising nature lea…
▽ More
Diffusion models have achieved outstanding performance in unsupervised industrial anomaly detection (uIAD) by learning a manifold of normal data under the common assumption that off-manifold anomalies are harder to generate, resulting in larger reconstruction errors in data space or lower probability densities in the tractable latent space. However, their iterative denoising and noising nature leads to slow inference. In this paper, we propose OSD-IRF, a novel one-step diffusion with inverse residual fields, to address this limitation for uIAD task. We first train a deep diffusion probabilistic model (DDPM) on normal data without any conditioning. Then, for a test sample, we predict its inverse residual fields (IRF) based on the noise estimated by the well-trained parametric noise function of the DDPM. Finally, uIAD is performed by evaluating the probability density of the IRF under a Gaussian distribution and comparing it with a threshold. Our key observation is that anomalies become distinguishable in this IRF space, a finding that has seldom been reported in prior works. Moreover, OSD-IRF requires only single step diffusion for uIAD, thanks to the property that IRF holds for any neighboring time step in the denoising process. Extensive experiments on three widely used uIAD benchmarks show that our model achieves SOTA or competitive performance across six metrics, along with roughly a 2X inference speedup without distillation.
△ Less
Submitted 20 April, 2026;
originally announced April 2026.
-
Elder-Sim: A Psychometrically Validated Platform for Personality-Stable Elderly Digital Twins
Authors:
Jiaqing Wang,
Zhongfang Yang,
Xingyuan Zhu,
Zong'an Huang,
Hao Wang,
Li Tian,
Ying Cao,
Xiaomin Qu,
Xiang Qi,
Bei Wu,
Zheng Zhu
Abstract:
Background: LLMs enable patient-facing conversational agents, creating a pathway toward digital twins that capture older adults' lived experiences and behavioral responses across time. A central barrier is personality drift -- inconsistent trait expression across repeated interactions -- which undermines reliability of generated trajectories and intervention-response simulation in geriatric care.…
▽ More
Background: LLMs enable patient-facing conversational agents, creating a pathway toward digital twins that capture older adults' lived experiences and behavioral responses across time. A central barrier is personality drift -- inconsistent trait expression across repeated interactions -- which undermines reliability of generated trajectories and intervention-response simulation in geriatric care.
Objective: To develop ELDER-SIM, a multi-role elderly-care conversational platform for building personality-stable digital twin agents, and to propose a psychometric validation framework for quantifying personality consistency in LLM-based agents.
Methods: ELDER-SIM was implemented via n8n workflow orchestration with local LLM inference (Ollama/vLLM), integrating (1) Big Five (OCEAN) trait specifications, (2) a Cognitive Conceptualization Diagram (CCD) grounded in Beck's CBT framework, and (3) a MySQL-based long-term memory module. Ablation studies across four conditions -- Baseline, +Memory, +CCD, and +LoRA (fine-tuned on 19,717 instruction pairs from CHARLS) -- were evaluated via Cronbach's $α$, ICC, and role discrimination accuracy.
Results: Reliability was acceptable to excellent across conditions (Cronbach's $α$: 0.70--0.94; ICC: 0.85--0.96). Role discrimination improved from 83.3% (Baseline) to 88.9% (+Memory), 94.4% (+CCD), and 97.2% (+LoRA). CCD produced the largest consistency gain (mean $α$ 0.702$\to$0.892), while LoRA achieved the highest overall consistency ($α$ 0.940; ICC 0.958).
Conclusions: ELDER-SIM provides a psychometrically validated approach for constructing personality-consistent elderly digital twin agents. Structured cognitive modeling and domain adaptation reduce personality drift, supporting reliable longitudinal simulation for elderly mental health care and reproducible in silico evaluation before clinical deployment.
△ Less
Submitted 16 March, 2026;
originally announced April 2026.
-
POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management
Authors:
Yikun Liu,
Yuan Liu,
Haicheng Wang,
Zhongyin Zhao,
Le Tian,
Xiao Zhou,
Jiangchao Yao,
Yanfeng Wang,
Weidi Xie
Abstract:
Large Multimodal Models (LMMs) excel at visual perception but struggle with real-time, knowledge-intensive queries due to their reliance on static parametric knowledge. While multimodal search agents offer a promising solution, developing them from vanilla LMMs presents two major challenges: the lack of open recipes to cultivate search agency from scratch, and the severe context explosion (attenti…
▽ More
Large Multimodal Models (LMMs) excel at visual perception but struggle with real-time, knowledge-intensive queries due to their reliance on static parametric knowledge. While multimodal search agents offer a promising solution, developing them from vanilla LMMs presents two major challenges: the lack of open recipes to cultivate search agency from scratch, and the severe context explosion (attention dilution) that occurs during long-horizon multi-turn investigations. In this paper, we address both challenges by developing a native multimodal search agent. First, we introduce an open training recipe featuring Agentic Seeding, a formative training stage that bootstraps tool-use and planning abilities directly from a non-agentic foundation. Second, to overcome the context bottleneck, we propose V-Fold, an adaptive memory management mechanism. V-Fold retains recent interactions as high-fidelity text while folding stale historical context into the visual space via rendering, exploiting the model's cross-modal alignment to preserve raw evidence without text token redundancy. Combining these innovations, we present POINTS-Seeker-8B, which achieves state-of-the-art performance among models of comparable scale across six multimodal search benchmarks.
△ Less
Submitted 29 July, 2026; v1 submitted 15 April, 2026;
originally announced April 2026.
-
Bayesian-Enhanced Galerkin-Based Reduced Order Modelling for Unsteady Compressible Flows
Authors:
Bijie Yang,
Chengyuan Liu,
Lu Tian,
Yuping Qian,
Mingyang Yang
Abstract:
This work proposes a statistically enhanced framework to address the instability and limited predictive capability of conventional Galerkin-Proper Orthogonal Decomposition (Galerkin-POD) models. The method reformulates the correction of the Galerkin-projected ODE system as a statistical inverse problem, in which the coefficients are inferred through Bayesian inference. By accounting for model unce…
▽ More
This work proposes a statistically enhanced framework to address the instability and limited predictive capability of conventional Galerkin-Proper Orthogonal Decomposition (Galerkin-POD) models. The method reformulates the correction of the Galerkin-projected ODE system as a statistical inverse problem, in which the coefficients are inferred through Bayesian inference. By accounting for model uncertainty arising from POD mode truncation and data uncertainty introduced by data noise and numerical postprocessing, the framework systematically updates the ODE system coefficients using an analytical, sampling-free solution based on Gaussian likelihood and inverse-Gamma priors. The approach is first validated using a self-sustained oscillating flow over a dimpled surface at a moderate Reynolds number (Re=3000), demonstrating stable and accurate reproduction of the temporal dynamics and phase trajectories of coherent structures when compared with direct numerical simulation (DNS). It is then applied to a centrifugal compressor featuring strong tip-leakage vortex breakdown and impeller-diffuser interactions at Re=100000, where the model successfully captures dominant unsteady structures and frequency characteristics despite limited mode retention. Overall, the results show that Bayesian inference substantially enhances the robustness, stability, and predictive fidelity of Galerkin-POD models for compressible flow systems. The proposed methodology combines the physical interpretability of Galerkin projection with the statistical rigour of Bayesian inference, offering a general, computationally efficient, and uncertainty-aware reduced-order modelling framework for complex fluid dynamic applications.
△ Less
Submitted 14 April, 2026;
originally announced April 2026.
-
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs
Authors:
Haicheng Wang,
Yuan Liu,
Yikun Liu,
Zhemeng Yu,
Zhongyin Zhao,
Yangxiu You,
Zilin Yu,
Le Tian,
Xiao Zhou,
Jie Zhou,
Weidi Xie,
Yanfeng Wang
Abstract:
Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable capabilities in cross-modal understanding and generation. However, the rapid growth of visual token sequences--especially in long-video and streaming scenarios--poses a major challenge to their scalability and real-world deployment. Thus, we introduce POINTS-Long, a native dual-mode MLLM featuring dynamic visual token s…
▽ More
Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable capabilities in cross-modal understanding and generation. However, the rapid growth of visual token sequences--especially in long-video and streaming scenarios--poses a major challenge to their scalability and real-world deployment. Thus, we introduce POINTS-Long, a native dual-mode MLLM featuring dynamic visual token scaling inspired by the human visual system. The model supports two complementary perception modes: focus mode and standby mode, enabling users to dynamically trade off efficiency and accuracy during inference. On fine-grained visual tasks, the focus mode retains the optimal performance, while on long-form general visual understanding, the standby mode retains 97.7-99.7% of the original accuracy using only 1/40-1/10th of the visual tokens. Moreover, POINTS-Long natively supports streaming visual understanding via a dynamically detachable KV-cache design, allowing efficient maintenance of ultra-long visual memory. Our work provides new insights into the design of future MLLMs and lays the foundation for adaptive and efficient long-form visual understanding.
△ Less
Submitted 13 April, 2026;
originally announced April 2026.
-
Cost-optimal Sequential Testing via Doubly Robust Q-learning
Authors:
Doudou Zhou,
Yiran Zhang,
Dian Jin,
Yingye Zheng,
Lu Tian,
Tianxi Cai
Abstract:
Clinical decision-making often involves selecting tests that are costly, invasive, or time-consuming, motivating individualized, sequential strategies for what to measure and when to stop ascertaining. We study the problem of learning cost-optimal sequential decision policies from retrospective data, where test availability depends on prior results, inducing informative missingness. Under a sequen…
▽ More
Clinical decision-making often involves selecting tests that are costly, invasive, or time-consuming, motivating individualized, sequential strategies for what to measure and when to stop ascertaining. We study the problem of learning cost-optimal sequential decision policies from retrospective data, where test availability depends on prior results, inducing informative missingness. Under a sequential missing-at-random mechanism, we develop a doubly robust Q-learning framework for estimating optimal policies. The method introduces path-specific inverse probability weights that account for heterogeneous test trajectories and satisfy a normalization property conditional on the observed history. By combining these weights with auxiliary contrast models, we construct orthogonal pseudo-outcomes that enable unbiased policy learning when either the acquisition model or the contrast model is correctly specified. We establish oracle inequalities for the stage-wise contrast estimators, along with convergence rates, regret bounds, and misclassification rates for the learned policy. Simulations demonstrate improved cost-adjusted performance over weighted and complete-case baselines, and an application to a prostate cancer cohort study illustrates how the method reduces testing cost without compromising predictive accuracy.
△ Less
Submitted 14 April, 2026; v1 submitted 13 April, 2026;
originally announced April 2026.
-
Electrochemical stability and lithium insertion at the Li|Li3OCl solid electrolyte interface
Authors:
Deobrat Singh,
Li-Yun Tian,
Moyses Araujo,
Raquel Lizarraga
Abstract:
Solid-state lithium batteries have attracted considerable attention due to their potential to provide improved safety and higher energy density compared with conventional liquid electrolyte batteries. However, the stability of the interface between Li metal anodes and solid electrolytes remains a critical issue that strongly influences battery performance. In this work, first-principles density fu…
▽ More
Solid-state lithium batteries have attracted considerable attention due to their potential to provide improved safety and higher energy density compared with conventional liquid electrolyte batteries. However, the stability of the interface between Li metal anodes and solid electrolytes remains a critical issue that strongly influences battery performance. In this work, first-principles density functional theory calculations are performed to investigate the interfacial properties of a solid-state battery system composed of Li metal anode and Li3OCl solid electrolyte. The structural stability, electronic structure, and electrochemical behavior of the Li|Li3OCl interface are systematically analyzed. Several interface orientations are constructed and compared in order to identify the most energetically favorable configuration. The electronic properties and interfacial charge redistribution are further examined to understand the nature of the interaction between Li metal and the Li3OCl electrolyte. Our results indicate that the Li|Li3OCl interface exhibits stable structural and electronic characteristics, with localized charge redistribution occurring near the interface region. The electrochemical stability against the insertion of an additional Li atom is also evaluated, showing that Li incorporation is energetically unfavorable in most layers of the electrolyte. These results suggest that the Li3OCl electrolyte maintains good electrochemical stability in contact with Li metal. The present study provides atomic-scale insight into the interfacial behavior of Li|Li3OCl and highlights the potential of Li3OCl as a promising solid electrolyte for solid-state lithium batteries.
△ Less
Submitted 12 April, 2026;
originally announced April 2026.
-
AndroTMem: From Interaction Trajectories to Anchored Memory in Long-Horizon GUI Agents
Authors:
Yibo Shi,
Jungang Li,
Linghao Zhang,
Zihao Dongfang,
Biao Wu,
Sicheng Tao,
Yibo Yan,
Chenxi Qin,
Weiting Liu,
Zhixin Lin,
Hanqian Li,
Yu Huang,
Song Dai,
Yonghua Hei,
Yue Ding,
Xiang Li,
Shikang Wang,
Chengdong Xu,
Jingqi Liu,
Xueying Ma,
Zhiwen Zheng,
Xiaofei Zhang,
Bincheng Wang,
Nichen Yang,
Jie Wu
, et al. (3 additional authors not shown)
Abstract:
Long-horizon GUI agents are a key step toward real-world deployment, yet effective interaction memory under prevailing paradigms remains under-explored. Replaying full interaction sequences is redundant and amplifies noise, while summaries often erase dependency-critical information and traceability. We present AndroTMem, a diagnostic framework for anchored memory in long-horizon Android GUI agent…
▽ More
Long-horizon GUI agents are a key step toward real-world deployment, yet effective interaction memory under prevailing paradigms remains under-explored. Replaying full interaction sequences is redundant and amplifies noise, while summaries often erase dependency-critical information and traceability. We present AndroTMem, a diagnostic framework for anchored memory in long-horizon Android GUI agents. Its core benchmark, AndroTMem-Bench, comprises 1,069 tasks with 34,473 interaction steps (avg. 32.1 per task, max. 65). We evaluate agents with TCR (Task Complete Rate), focusing on tasks whose completion requires carrying forward critical intermediate state; AndroTMem-Bench is designed to enforce strong step-to-step causal dependencies, making sparse yet essential intermediate states decisive for downstream actions and centering interaction memory in evaluation. Across open- and closed-source GUI agents, we observe a consistent pattern: as interaction sequences grow longer, performance drops are driven mainly by within-task memory failures, not isolated perception errors or local action mistakes. Guided by this diagnosis, we propose Anchored State Memory (ASM), which represents interaction sequences as a compact set of causally linked intermediate-state anchors to enable subgoal-targeted retrieval and attribution-aware decision making. Across multiple settings and 12 evaluated GUI agents, ASM consistently outperforms full-sequence replay and summary-based baselines, improving TCR by 5%-30.16% and AMS by 4.93%-24.66%, indicating that anchored, structured memory effectively mitigates the interaction-memory bottleneck in long-horizon GUI tasks. The code, benchmark, and related resources are publicly available at [https://github.com/CVC2233/AndroTMem](https://github.com/CVC2233/AndroTMem).
△ Less
Submitted 18 March, 2026;
originally announced March 2026.
-
HRI-SA: A Multimodal Dataset for Online Assessment of Human Situational Awareness during Remote Human-Robot Teaming
Authors:
Hashini Senaratne,
Richard Attfield,
Samith Widhanapathirana,
David Howard,
Cecile Paris,
Dana Kulic,
Leimin Tian
Abstract:
Maintaining situational awareness (SA) is critical in human-robot teams. Yet, under high workload and dynamic conditions, operators often experience SA gaps. Automated detection of SA gaps could provide timely assistance for operators. However, conventional SA measures either disrupt task flow or cannot capture real-time fluctuations, limiting their operational utility. To the best of our knowledg…
▽ More
Maintaining situational awareness (SA) is critical in human-robot teams. Yet, under high workload and dynamic conditions, operators often experience SA gaps. Automated detection of SA gaps could provide timely assistance for operators. However, conventional SA measures either disrupt task flow or cannot capture real-time fluctuations, limiting their operational utility. To the best of our knowledge, no publicly available dataset currently supports the systematic evaluation of online human SA assessment in human-robot teaming. To advance the development of online SA assessment tools, we introduce HRI-SA, a multimodal dataset from 30 participants in a realistic search-and-rescue human-robot teaming context, incorporating eye movements, pupil diameter, biosignals, user interactions, and robot data. The experimental protocol included predefined events requiring timely operator assistance, with ground truth SA latency of two types (perceptual and comprehension) systematically obtained by measuring the time between assistance need onset and resolution. We illustrate the utility of this dataset by evaluating standard machine learning models for detecting perceptual SA latencies using generic eye-tracking features and contextual features. Results show that eye-tracking features alone effectively classified perceptual SA latency (recall=88.91%, F1=67.63%) using leave-one-group-out cross-validation, with performance improved through contextual data fusion (recall=91.51%, F1=80.38%). This paper contributes the first public dataset supporting the systematic evaluation of SA throughout a human-robot teaming mission, while also demonstrating the potential of generic eye-tracking features for continuous perceptual SA latency detection in remote human-robot teaming.
△ Less
Submitted 18 March, 2026;
originally announced March 2026.
-
SWE-Hub: A Unified Production System for Scalable, Executable Software Engineering Tasks
Authors:
Yucheng Zeng,
Shupeng Li,
Daxiang Dong,
Ruijie Xu,
Zimo Chen,
Liwei Zheng,
Yuxuan Li,
Zhe Zhou,
Haotian Zhao,
Lun Tian,
Heng Xiao,
Tianshu Zhu,
Longkun Hao,
Jianmin Wu
Abstract:
Progress in software-engineering agents is increasingly constrained by the scarcity of executable, scalable, and realistic data for training and evaluation. This scarcity stems from three fundamental challenges in existing pipelines: environments are brittle and difficult to reproduce across languages; synthesizing realistic, system-level bugs at scale is computationally expensive; and existing da…
▽ More
Progress in software-engineering agents is increasingly constrained by the scarcity of executable, scalable, and realistic data for training and evaluation. This scarcity stems from three fundamental challenges in existing pipelines: environments are brittle and difficult to reproduce across languages; synthesizing realistic, system-level bugs at scale is computationally expensive; and existing data predominantly consists of short-horizon repairs, failing to capture long-horizon competencies like architectural consistency. We introduce \textbf{SWE-Hub}, an end-to-end system that operationalizes the data factory abstraction by unifying environment automation, scalable synthesis, and diverse task generation into a coherent production stack. At its foundation, the \textbf{Env Agent} establishes a shared execution substrate by automatically converting raw repository snapshots into reproducible, multi-language container environments with standardized interfaces. Built upon this substrate, \textbf{SWE-Scale} engine addresses the need for high-throughput generation, combining cross-language code analysis with cluster-scale validation to synthesize massive volumes of localized bug-fix instances. \textbf{Bug Agent} generates high-fidelity repair tasks by synthesizing system-level regressions involving cross-module dependencies, paired with user-like issue reports that describe observable symptoms rather than root causes. Finally, \textbf{SWE-Architect} expands the task scope from repair to creation by translating natural-language requirements into repository-scale build-a-repo tasks. By integrating these components, SWE-Hub establishes a unified production pipeline capable of continuously delivering executable tasks across the entire software engineering lifecycle.
△ Less
Submitted 28 February, 2026;
originally announced March 2026.
-
Digital Quantum Simulation of the Holstein-Primakoff Transformation on Noisy Qubits
Authors:
Kelvin Yip,
Alessandro Monteros,
Sahel Ashhab,
Lin Tian
Abstract:
Quantum simulation of many-body systems offers a powerful approach to exploring collective quantum dynamics beyond classical computational reach. Although spin and fermionic models have been extensively simulated on digital quantum computers, the simulation of bosonic systems on programmable quantum processors is often hindered by the intrinsically large Hilbert space of bosonic modes. In this wor…
▽ More
Quantum simulation of many-body systems offers a powerful approach to exploring collective quantum dynamics beyond classical computational reach. Although spin and fermionic models have been extensively simulated on digital quantum computers, the simulation of bosonic systems on programmable quantum processors is often hindered by the intrinsically large Hilbert space of bosonic modes. In this work, we study the digital quantum simulation of bosonic modes using the Holstein-Primakoff (HP) transformation and implement this protocol on a cloud-based superconducting quantum processor. Two representative models are realized on quantum hardware: (i) the driven harmonic oscillator and (ii) the Jaynes-Cummings model. Using data obtained from the quantum simulations, we systematically examine the interplay between algorithmic and hardware-induced errors to identify optimal simulation parameters. The dominant algorithmic errors arise from the finite number of qubits used in the HP mapping and the finite number of Trotter steps in the time evolution, while hardware errors mainly originate from gate infidelity, decoherence, and readout errors. This study advances the digital quantum simulation of many-body systems involving bosonic degrees of freedom on currently available cloud quantum processors and provides a framework that can be extended to more complex spin-boson and multimode cavity models.
△ Less
Submitted 19 February, 2026;
originally announced February 2026.
-
Efficient and Debiased Learning of Average Hazard Under Non-Proportional Hazards
Authors:
Xiang Meng,
Lu Tian,
Kenneth Kehl,
Hajime Uno
Abstract:
The hazard ratio from the Cox proportional hazards model is a ubiquitous summary of treatment effect. However, when hazards are non-proportional, the hazard ratio can lose a stable causal interpretation and become study-dependent because it effectively averages time-varying effects with weights determined by follow-up and censoring. We consider the average hazard (AH) as an alternative causal esti…
▽ More
The hazard ratio from the Cox proportional hazards model is a ubiquitous summary of treatment effect. However, when hazards are non-proportional, the hazard ratio can lose a stable causal interpretation and become study-dependent because it effectively averages time-varying effects with weights determined by follow-up and censoring. We consider the average hazard (AH) as an alternative causal estimand: a population-level person-time event rate that remains well-defined and interpretable without assuming proportional hazards. Although AH can be estimated nonparametrically and regression-style adjustments have been proposed, existing approaches do not provide a general framework for flexible, high-dimensional nuisance estimation with valid sqrt{n} inference. We address this gap by developing a semiparametric, doubly robust framework for covariate-adjusted AH. We establish pathwise differentiability of AH in the nonparametric model, derive its efficient influence function, and construct cross-fitted, debiased estimators that leverage machine learning for nuisance estimation while retaining asymptotically normal, sqrt{n}-consistent inference under mild product-rate conditions. Simulations demonstrate that the proposed estimator achieves small bias and near-nominal confidence-interval coverage across proportional and non-proportional hazards settings, including crossing-hazards regimes where Cox-based summaries can be unstable. We illustrate practical utility in comparative effectiveness research by comparing immunotherapy regimens for advanced melanoma using SEER-Medicare linked data.
△ Less
Submitted 13 February, 2026;
originally announced February 2026.
-
GenFaceUI: Meta-Design of Generative Personalized Facial Expression Interfaces for Intelligent Agents
Authors:
Yate Ge,
Lin Tian,
Yi Dai,
Shuhan Pan,
Yiwen Zhang,
Qi Wang,
Weiwei Guo,
Xiaohua Sun
Abstract:
This work investigates generative facial expression interfaces for intelligent agents from a meta-design perspective. We propose the Generative Personalized Facial Expression Interface (GPFEI) framework, which organizes rule-bounded spaces, character identity, and context--expression mapping to address challenges of control, coherence, and alignment in run-time facial expression generation. To ope…
▽ More
This work investigates generative facial expression interfaces for intelligent agents from a meta-design perspective. We propose the Generative Personalized Facial Expression Interface (GPFEI) framework, which organizes rule-bounded spaces, character identity, and context--expression mapping to address challenges of control, coherence, and alignment in run-time facial expression generation. To operationalize this framework, we developed GenFaceUI, a proof-of-concept tool that enables designers to create templates, apply semantic tags, define rules, and iteratively test outcomes. We evaluated the tool through a qualitative study with twelve designers. The results show perceived gains in controllability and consistency, while revealing needs for structured visual mechanisms and lightweight explanations. These findings provide a conceptual framework, a proof-of-concept tool, and empirical insights that highlight both opportunities and challenges for advancing generative facial expression interfaces within a broader meta-design paradigm.
△ Less
Submitted 11 February, 2026;
originally announced February 2026.