Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 494 results for author: Wen, Q

.
  1. arXiv:2608.23501  [pdf, ps, other

    cs.SE

    An Interactive Agent for Requirement-Driven Candidate Sourcing

    Authors: Yuanpeng He, Fangjing Li, Xiangyu Ru, Kexin Sun, Kun Yang, Lijian Li, Chi-Man Pun, Qingsong Wen, Wenpin Jiao, Mingkai Guo, Yirong Feng, Daiheng Gao, Zhi Jin

    Abstract: Finding people from a natural-language description (``ML engineers transitioning to research roles in biotech'') is increasingly delegated to LLM agents and framed as information retrieval. We argue that it is fundamentally a requirements engineering task: such a request is an under-determined requirement with implicit constraints, many valid answers, and no acceptance criterion, so useful answers… ▽ More

    Submitted 24 August, 2026; originally announced August 2026.

    Comments: 12 pages

  2. arXiv:2608.16386  [pdf, ps, other

    cs.CL cs.LG

    Mint-Agent: Introducing Finance-Native Agentic Foundation Models

    Authors: Mint-Agent Team, Kun Wang, Gavin Zhang, Yaze Geng, Lei Tang, Yaoyang Yi, Zonghan Wu, Yifan Hu, Qingsong Wen, Yilei Shao

    Abstract: Financial agents must do more than recall domain knowledge: they must be both reliable, executing precise operations over grounded evidence, and executive, sustaining long-horizon research whose conclusions remain auditable. We present Mint-Agent, a family of finance-native agentic models designed around these two scales of financial intelligence. Mint-Agent is built upon three pillars: data, harn… ▽ More

    Submitted 21 August, 2026; v1 submitted 17 August, 2026; originally announced August 2026.

  3. arXiv:2608.16353  [pdf, ps, other

    cs.CL cs.AI

    HalluTracer: Hallucination Detection via Depth-Averaging Truth Signals

    Authors: Zhihao Guo, Zonghan Wu, Huan Huo, DaYong Ye, Junwei Zhang, Weiran Yao, Zhiwei Liu, Qingsong Wen, Yilei Shao

    Abstract: Even well-aligned large language models confidently generate factually incorrect text, making hallucination a persistent reliability risk in high-stakes deployments. These models nonetheless carry linearly separable truthfulness signals in their internal representations. Existing white-box detectors, however, collapse this evidence to isolated components or a single depth, discarding discriminativ… ▽ More

    Submitted 17 August, 2026; originally announced August 2026.

  4. arXiv:2608.15844  [pdf, ps, other

    cs.CL

    MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model Simulations

    Authors: Sky Ng, Brihi Joshi, Ishan Gupta, Shirley Huang, Zonglin Di, Yun Shen, Qianfeng Wen, Yifan Simon Liu, Ruoqi Gao, Yilan, Fan, Zhiwei Zhang, Muhammad Ahmed Mohsin, Yucheng Lu, Xiaoyi Liu, Heming Liu, Qianyu Zhu, Hanwen Xing, Zhengyang Shan, My Chiffon Nguyen, Guanghui Min, Jianheng, Hou, Yunze, Xiao , et al. (25 additional authors not shown)

    Abstract: Long-horizon, multi-agent language model (LM) simulations are widely proposed for studying social behavior, yet instruments to measure whether persona-conditioned agents maintain identity fidelity under sustained pressure are lacking. We present MicroVerse, a behavioral-science instrument that measures identity drift in generative agents. Agents carry an immutable "soul file" (core values, moral b… ▽ More

    Submitted 16 August, 2026; originally announced August 2026.

  5. arXiv:2608.15838  [pdf, ps, other

    cs.HC

    PersonaEval: Persona-Based User Simulation for Evaluating Interactive Applications

    Authors: Yifan Simon Liu, Qianfeng Wen, Yilan Fan, Shirley Huang, Ruoqi Gao, Jianheng Hou, Muhammad Ahmed Mohsin, Zonglin Di, Brihi Joshi, Xincheng Tan, Yucheng Lu, Xiaoyi Liu, Heming Liu, Hanwen Xing, Guanghui Min, Zhengyang Shan, My Chiffon Nguyen, Ishan Gupta, Yunze Xiao, Hannah Collison, Jintao Huang, Jiatong Li, Sankalp Jajee, Yunhan Zhao, Bing Hu , et al. (18 additional authors not shown)

    Abstract: Real user studies are important for understanding how people interact with systems under test or already deployed. In practice, however, they are often costly, time-consuming, and difficult to scale. To address these challenges, we introduce PersonaEval, a persona-based user simulation framework that approximates real-user behavior across diverse interactive settings. PersonaEval connects simulate… ▽ More

    Submitted 16 August, 2026; originally announced August 2026.

  6. arXiv:2608.14270  [pdf, ps, other

    cs.AI

    TimeSage-EV: A Live Benchmark for Agentic Time Series Analysis in Evolving Environments

    Authors: Qingren Yao, Yaxuan Kong, Yuqi Nie, Yichen Li, Stefan Zohren, Anna Vettoruzzo, Qingsong Wen, Ming Jin, Joaquin Vanschoren

    Abstract: Time series analysis in high-stakes domains relies on recurring data releases, where new observations can alter the evidence base and the validity of later conclusions. Existing time series QA benchmarks mostly rely on fixed snapshots, leaving temporal validity and cutoff-aware evidence use unevaluated. We introduce TimeSage-EV, a live benchmark for agentic time series analysis in evolving environ… ▽ More

    Submitted 14 August, 2026; originally announced August 2026.

  7. arXiv:2608.14106  [pdf, ps, other

    cs.LG cs.AI cs.CE stat.AP stat.ML

    Forecast Collapse in Time-Series Foundation Models

    Authors: Shu Wan, Miles Ma, Hank Zhu, Guangqi Liu, Stephen Wang, Qingsong Wen, Huan Liu

    Abstract: When forecasting hourly returns for 1,000 US equities, we observe an unexpected phenomenon: predictions become nearly flat and show poor stock ranking, as measured by cross-sectional correlation. We call this forecast collapse. Surprisingly, the phenomenon largely disappears when forecasting trading volume under the same setting. We investigate forecast collapse across time-series foundation model… ▽ More

    Submitted 14 August, 2026; originally announced August 2026.

    Comments: 27 pages, 3 figures, 5 tables. Dataset: https://huggingface.co/datasets/abel-lab/finance1k

  8. An Adaptive Longitudinal Platooning Design Based On Concurrent Learning

    Authors: Qiuhao Wen, Di Liu, Jiwei Wang, Simone Baldi

    Abstract: This work proposes a new adaptive longitudinal platooning strategy in the framework of concurrent learning. Adaptive refers to vehicles facing uncertainty in powertrain parameters via on-line estimation; concurrent learning refers to using both current and past data in the estimation. The proposed platooning strategy advances existing ones since convergence to the true powertrain parameters is gua… ▽ More

    Submitted 7 August, 2026; originally announced August 2026.

    Journal ref: IEEE Control Systems Letters, vol. 8, pp. 303-308, 2024

  9. Correct Online Estimation of the Powertrain Time Constants in Adaptive Vehicular Platooning

    Authors: Qiuhao Wen, Simone Baldi, Jiwei Wang, Wenwu Yu, Di Liu

    Abstract: In longitudinal platooning, some key sources of uncertainty are the powertrain time constants of the vehicles. Because such time constants appear in the input matrix of the platooning dynamics, their correct estimation is either impractical with methods requiring persistence of excitation, or impossible with methods requiring the input matrix to be known. This work proposes a novel adaptive longit… ▽ More

    Submitted 7 August, 2026; originally announced August 2026.

    Journal ref: IEEE Transactions on Systems, Man, and Cybernetics: Systems, vol. 56, no. 1, pp. 279-291, Jan. 2026

  10. arXiv:2608.05774  [pdf, ps, other

    cs.CV cs.LG

    SR-JEPA: Learning Predictive Latent State in 3D Scenes

    Authors: Zihan Zhou, Qifu Wen, Xi Zeng

    Abstract: Joint-embedding predictive architectures learn by predicting latent representations of missing observations, yet many masked JEPAs are evaluated primarily through the encoders they produce. We ask what a trained predictive pathway itself infers when an entire entity is absent from a native 3D scene. We introduce SR-JEPA, a point-native JEPA for scene-scale point clouds whose original frozen predic… ▽ More

    Submitted 6 August, 2026; originally announced August 2026.

    Comments: 17 pages, 5 figures, 9 tables

    ACM Class: I.2.10; I.2.6; I.4.8

  11. arXiv:2608.04205  [pdf, ps, other

    cs.AI

    MatrAIx: Simulating the World with 8.3 Billion Persona Agents

    Authors: Xiaomin Li, Yuexing Hao, Jianheng Hou, Jintao Huang, Qianfeng Wen, Shirley Huang, Yifan Liu, Xiaoyi Liu, Yilan Fan, Yijun Wang, Koutian Wu, Ruoqi Gao, Muhammad Ahmed Mohsin, Jing Tang, Brihi Joshi, Heming Liu, Zheyuan Deng, Zonglin Di, Sankalp Jajee, Jiuyao Lu, Zhiwei Zhang, Saksham Kapoor, Ishan Gupta, Yunhan Zhao, Chanwoo Park , et al. (68 additional authors not shown)

    Abstract: Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First,… ▽ More

    Submitted 4 August, 2026; originally announced August 2026.

    Comments: Project website: https://matraix.ai

  12. arXiv:2608.02578  [pdf, ps, other

    cs.RO cs.AI cs.LG

    CoWAM: Coordination Contracts for Selective Policy Intervention with WAMs

    Authors: Shuaijun Liu, Qifu Wen, Shuyang Hao, Qi Luo, Chenglong Zhang, Feiyang You, Chengyu Wu, Ningxin Su

    Abstract: World Action Models (WAMs) augment robot policies with action-conditioned predicted futures, but a plausible future alone does not justify changing the action that a bimanual policy would execute. We present CoWAM, a selective intervention layer that expresses synchronization, role compatibility, and collision convergence as coordination contracts. Each contract combines typed admissibility checks… ▽ More

    Submitted 3 August, 2026; originally announced August 2026.

  13. arXiv:2607.17347  [pdf, ps, other

    cs.IR

    Adapting Embedding Models for Agent Capability Retrieval

    Authors: Tingwei Chen, Yunxiao Shi, Zhengdong Chu, Qingsong Wen, Min Xu

    Abstract: Open agent marketplaces list native agents, tool bundles, and reusable skill packages in the same search interface, yet practitioners still have little guidance on how to retrieve across this mixed catalog. We study whether off-the-shelf retrieval models, trained for general text retrieval, can be adapted to match user queries to executable agent capabilities, and whether the learned signal transf… ▽ More

    Submitted 19 July, 2026; originally announced July 2026.

    Comments: Accepted for oral presentation at the AgentSearch Workshop, SIGIR 2026

  14. arXiv:2607.13678  [pdf, ps, other

    eess.SP

    M3F-UAV: A Missing-Modality Multimodal Foundation Model for Low-Altitude Wireless Sensing

    Authors: Pengxuan Gao, Kai Ying, Botao Wu, Jianhua Mo, Qingsong Wen

    Abstract: Low-altitude unmanned aerial vehicles (UAVs) are emerging as key platforms for wireless intelligence tasks. However, practical low-altitude wireless systems usually operate in complex urban environments, where visual occlusion, sparse geometric observations, multipath propagation, and sensor failures may degrade the reliability of single-modality models. To address these challenges, this paper pro… ▽ More

    Submitted 15 July, 2026; originally announced July 2026.

  15. arXiv:2607.05820  [pdf, ps, other

    physics.flu-dyn physics.bio-ph q-bio.TO

    Continuum modeling of fluidic and elastic flow during growth-driven wound closure in partial-EMT cell monolayers

    Authors: Chaozhen Wei, Han Jiang, Yifan Gu, Nonthakorn Olaranont, Pengbo Wang, Qi Wen, Yubing Sun, Min Wu

    Abstract: Large-scale circular gap closure occurs over a time scale on which cell growth and proliferation become important. Growth is the main driver of the closing process, while cell dynamics such as elongation and intercalation reflect elastic and fluidic contributions to tissue deformation. We develop a novel fluidized growth-elasticity framework as a nonlinear analogue of a Maxwell fluid with growth.… ▽ More

    Submitted 14 August, 2026; v1 submitted 7 July, 2026; originally announced July 2026.

    Comments: 36 pages, 14 figues

    MSC Class: 35Q74; 92Cxx; 74L15; 74Bxx

  16. arXiv:2606.30466  [pdf, ps, other

    hep-th

    Holography and Kinematic Space for Gravitational Sub-regions in AdS

    Authors: Debarshi Basu, Qiang Wen

    Abstract: It is well-known in integral geometry that a maximally symmetric Riemannian manifold, such as a static slice of vacuum AdS spacetime, can be perfectly covered by the geodesics in the Kinematic space, which we call the partial-entanglement-entropy (PEE) threads. In this context, the area of a codimension-one surface in the manifold can be computed by counting its intersections with the PEE threads,… ▽ More

    Submitted 29 July, 2026; v1 submitted 29 June, 2026; originally announced June 2026.

    Comments: 25 pages, 12 figures, comments are welcome; v2 references added.;v3,minor revision on the presentation

  17. arXiv:2606.30203  [pdf

    physics.optics

    Comb-enabled spectral-domain image transport through perturbation-prone multimode fibers

    Authors: Maohan Li, Zijian Wang, Bowen Sun, Zhuoren Wan, Xiangze Ma, Xiuxiu Zhang, Yuan Chen, Mei Yang, Qi Wen, Zhaoyang Wen, Ming Yan, Heping Zeng

    Abstract: Multimode fibers (MMFs) offer a compact platform for imaging, sensing, and information transport, but their practical deployment is hindered by sensitivity to fiber perturbations, which alter modal coupling and invalidate conventional speckle-based calibrations. Here, we demonstrate perturbation-resilient image transport through MMFs by combining image-to-spectrum encoding with dual-comb spectrosc… ▽ More

    Submitted 29 June, 2026; originally announced June 2026.

    Comments: 20 pages,4 figures

  18. arXiv:2606.28356  [pdf, ps, other

    cs.IR cs.AI

    SafeGEO: Understanding Generative Engine Optimization Risks in Recommendation Agents

    Authors: Qianfeng Wen, Yifan Simon Liu, Xin Liu, Difan Jiao, Blair Yang, Junda Wu, Zhenwei Tang

    Abstract: Generative Engine Optimization (GEO) lets content owners rewrite web content to increase their visibility in generative systems. In recommendation agents, this creates a risk that seller-controlled sources make flawed products appear better supported than they are. We study this risk by asking whether recommendation agents preserve utility-aligned decisions when seller-controlled sources are rewri… ▽ More

    Submitted 8 June, 2026; originally announced June 2026.

    Comments: 41 pages,23 figures

  19. arXiv:2606.24113  [pdf, ps, other

    cs.LG

    FedUP: One-Shot Federated Unlearning via Centroid-Guided Plug-in Filters

    Authors: Feihong Nan, Zhengyi Zhong, Pan Wang, Weidong Bao, Xiongtao Zhang, Quan Wen, Ji Wang

    Abstract: Federated unlearning (FU) is critical for complying with legal mandates like the right to be forgotten in decentralized systems, yet current methods face a persistent dilemma between non-target knowledge loss and high request latency. To resolve these issues, we propose FedUP, a one-shot federated unlearning framework utilizing lightweight pluggable filters that act as a "knowledge funnel" to scre… ▽ More

    Submitted 24 June, 2026; v1 submitted 22 June, 2026; originally announced June 2026.

    Comments: Accepted by IJCAI 2026

  20. arXiv:2606.21399  [pdf, ps, other

    cs.AI

    Calibration Is Not Control: Why LLM-Agent Oversight Needs Intervention

    Authors: Chubin Zhang, Zhenglin Wan, Xingrui Yu, Jingxuan Wu, Qi Wen, Pengfei Zhou, Wangbo Zhao, Ivor Tsang

    Abstract: Runtime oversight for LLM agents is commonly framed as scalar risk prediction: estimate failure likelihood, confidence, or uncertainty, then intervene once the score crosses a threshold. We argue that this framing targets the wrong object for control. The relevant question is not how likely the agent is to fail if it continues, but whether an available intervention would improve the outcome. Two t… ▽ More

    Submitted 19 June, 2026; originally announced June 2026.

    Comments: 29 pages

  21. From Memorization to Creation: Evaluating the Cognitive Depth of LLM-Generated Educational Questions

    Authors: Xiaolong Wang, Zhe Zhao, Song Lai, Chaoli Zhang, Zijie Geng, Yu Tong, Ye Wei, Qingsong Wen

    Abstract: While LLMs show promise in automating educational content creation, their ability to generate questions that stimulate higher-order thinking remains understudied. This work evaluates six widely-used LLMs through a Bloom's Taxonomy lens, focusing on their capacity to transcend rote memorization and achieve cognitive leaps. Using a hybrid human--AI evaluation protocol, we generate and analyze 20{,}7… ▽ More

    Submitted 5 May, 2026; originally announced June 2026.

    Comments: Accepted by KDD 2026

    Journal ref: KDD 2026

  22. arXiv:2606.15128  [pdf

    cond-mat.mes-hall

    Oxidation-induced ultrafast spin-to-orbital conversion at heavy-metal interfaces

    Authors: Xiaoxue Zeng, Tianyi Zhang, Yaokai Niu, Qiye Wen, Zhiyong Zhong, Zhi-Min Liao, Peng Yan, Xiufeng Han, Lichuan Jin

    Abstract: Oxidation engineering provides a route to control orbital degrees of freedom, yet its role in spin-to-orbital conversion remains largely unexplored. Here, we report an efficient spin-to-orbital conversion mechanism driven by interfacial oxidation at heavy-metal interfaces. In W/Co/SiO2 heterostructures, terahertz emission exhibits a time delay that scales linearly with the W thickness, identifying… ▽ More

    Submitted 13 June, 2026; originally announced June 2026.

    Comments: 18 pages, 4 figures

  23. arXiv:2606.06293  [pdf, ps, other

    cs.LG stat.ML

    PAC-Bayesian Adversarially Robust Generalization for Message Passing Graph Neural Networks: A Sensitivity Analysis

    Authors: Ziling Liang, Xinping Yi, Qingsong Wen, Shi Jin

    Abstract: Whilst the vulnerability of graph neural networks (GNNs) to adversarial attacks poses a critical threat to graph representation learning, the understanding of the robust generalization behavior remains a fundamental challenge in the adversarial setting. Recently, PAC-Bayesian margin-based generalization analysis substantially advances this line of research by providing a flexible and data-dependen… ▽ More

    Submitted 4 June, 2026; originally announced June 2026.

  24. arXiv:2606.01498  [pdf, ps, other

    cs.CL cs.AI

    TimeSage-MT: A Multi-Turn Benchmark for Evaluating Agentic Time Series Reasoning

    Authors: Yaxuan Kong, Qingren Yao, Yuqi Nie, Yichen Li, Yilei Shao, Stefan Zohren, Anna Vettoruzzo, Joaquin Vanschoren, Ming Jin, Qingsong Wen

    Abstract: Time series data inform critical decisions across many real-world domains. While large language model (LLM) agents can analyze data through natural language and tools, it remains unclear whether they can conduct reliable time series analysis across multi-turn conversations. Existing benchmarks focus on single-step tasks such as forecasting and anomaly detection, overlooking practical workflows whe… ▽ More

    Submitted 31 May, 2026; originally announced June 2026.

  25. arXiv:2605.29440  [pdf, ps, other

    cs.CL cs.AI cs.IR

    SkillBrew: Multi-Objective Curation of Skill Banks for LLM Agents

    Authors: Wentao Hu, Zhendong Chu, Yiming Zhang, Junda Wu, Ming Jin, Xiangyu Zhao, Yilei Shao, Yanfeng Wang, Qingsong Wen

    Abstract: Retrieval-augmented LLM agents increasingly rely on curated skill banks: collections of reusable textual principles that guide decision making on complex tasks. Existing approaches typically expand these banks in an append-only fashion, continuously adding new skills without removing redundant, outdated, or harmful ones, resulting in inefficient and poorly curated repositories. In this paper, we f… ▽ More

    Submitted 28 May, 2026; originally announced May 2026.

    Comments: 16 pages. Preprint. Under review

  26. arXiv:2605.25373  [pdf, ps, other

    cs.CV

    Physics-Aware 3D Gaussian Editing for Driving Scene Generation

    Authors: Feng Zhou, Jian Zhang, Yuhang Sun, He Wang, Qiong Wen, Debao Kong, Tieru Wu, Rui Ma

    Abstract: 3D Gaussian Splatting (3DGS) has shown great potential in autonomous driving simulation and data generation, enabling photorealistic reconstruction and flexible scene manipulation. However, existing 3DGS scene editing methods have limited support for road geometry editing (e.g., inserting speed humps or sunken roads), and generally do not couple such edits with plausible vehicle-road interaction d… ▽ More

    Submitted 24 May, 2026; originally announced May 2026.

  27. arXiv:2605.24156  [pdf, ps, other

    math.ST

    Long Memory in Intrinsically Dynamic Factor Models

    Authors: Qin Wen, Clifford M. Hurvich

    Abstract: We study the generalized dynamic factor model in a long-memory setting. Unlike most recent work, which assumes a finite-dimensional factor space and short memory, our framework allows the factor space to be infinite-dimensional and the common components to exhibit long memory. We employ the two-sided estimation method of Forni, Hallin, Lippi and Reichlin (2000, Review of Economics and Statistics)… ▽ More

    Submitted 22 May, 2026; originally announced May 2026.

    Comments: 92 pages, 11 figures

    MSC Class: 62P20 ACM Class: G.3

  28. arXiv:2605.23204  [pdf, ps, other

    cs.AI

    AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery

    Authors: Guiyao Tie, Jiawen Shi, Dingjie Song, Yixiao Huang, Ziji Sheng, Xueyang Zhou, Daizong Liu, Pan Zhou, Yongchao Chen, Ran Xu, Lifang He, Qingsong Wen, Manling Li, Cong Lu, Shuai Li, Pengtao Xie, Yixuan Yuan, Rui Meng, Lei Xing, Lichao Sun, Caiming Xiong, Philip S. Yu, Jianfeng Gao

    Abstract: Scientific research is being reshaped by AI systems that move beyond isolated assistance toward longer-horizon workflows spanning literature grounding, hypothesis generation, experimentation, validation, reporting, and revision. This shift marks a transition from task-level AI for science to workflow-level research automation. Yet current systems remain fragmented, differing in autonomy, domain sc… ▽ More

    Submitted 21 May, 2026; originally announced May 2026.

    Comments: 49 pages, 12 figures, 10 tables

  29. arXiv:2605.19249  [pdf, ps, other

    cs.LG

    Beyond Extrapolation: Knowledge Utilization Paradigm with Bidirectional Inspiration for Time Series Forecasting

    Authors: Liu Chong, Yingjie Zhou, Hao Li, Pengyang Wang, Qingsong Wen, Ce Zhu

    Abstract: Time-series forecasting is critical in various scenarios, such as energy, transportation, and public health. However, most existing forecasters rely primarily on one-way inference, \textit{i.e.}, mapping \textbf{history} to \textbf{target}, and overlook the structural information provided by a revised natural chain (``\textbf{history} (model input) -- \textbf{target} (ground-truth output) -- \text… ▽ More

    Submitted 18 May, 2026; originally announced May 2026.

    Comments: Accepted to ICML 2026. 18 pages, 6 figures

  30. arXiv:2605.17340  [pdf, ps, other

    cs.LG

    Olivia: Harmonizing Time Series Foundation Models with Power Spectral Density

    Authors: Jingru Fei, Kun Yi, Alex Xing Wang, Qingsong Wen, Xiangxiang Zhu, Wei Fan

    Abstract: Time series foundation models rely on large-scale pretraining over diverse datasets across domains, yet their heterogeneity in temporal patterns could hinder the effectiveness of training and learning transferable time series representations. Inspired a fundamental concept, normalized power spectral density (PSD) in signal processing, we assume harmonizing datasets via PSDs in the spectral domain… ▽ More

    Submitted 18 May, 2026; v1 submitted 17 May, 2026; originally announced May 2026.

    Comments: Accepted by ICML 2026

  31. arXiv:2605.16679  [pdf, ps, other

    cs.CL cs.AI

    CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?

    Authors: Haolin Chen, Deon Metelski, Leon Qi, Tao Xia, Joonyul Lee, Steve Brown, Kevin Riley, Frank Wang, T. Y. Alvin Liu, Hank Capps MD, Zeyu Tang, Xiangchen Song, Lingjing Kong, Fan Feng, Tianyi Zeng, Zhiwei Liu, Zixian Ma, Hang Jiang, Fangli Geng, Yuan Yuan, Chenyu You, Qingsong Wen, Hua Wei, Yanjie Fu, Yue Zhao , et al. (8 additional authors not shown)

    Abstract: End-to-end automation of realistic healthcare operations stresses three capabilities underrepresented in current benchmarks: policy density, decisions must be grounded in a large library of medical, insurance, and operational rules; Multi-role composition: a single task requires the agent to play multiple roles with handoffs; and multilateral interaction: intermediate workflow steps are multi-turn… ▽ More

    Submitted 19 May, 2026; v1 submitted 15 May, 2026; originally announced May 2026.

    Comments: Website: https://actava.ai/benchmarks Code: https://github.com/actava-ai/chi-bench Dataset: https://huggingface.co/datasets/actava/chi-bench

  32. arXiv:2605.12899  [pdf, ps, other

    stat.ML cs.LG

    Robust Sequential Experimental Design for A/B Testing

    Authors: Qianglin Wen, Xiangkun Wu, Chengchun Shi, Ting Li, Niansheng Tang, Yingying Zhang, Hongtu Zhu

    Abstract: Experimental design has emerged as a powerful approach for improving the sample efficiency of A/B testing, yet existing designs rely critically on correctly specified models. We study robust sequential experimental design under model misspecification and develop a unified framework that covers both contextual bandit and dynamic settings. Theoretically, we prove that our design bounds the worst-cas… ▽ More

    Submitted 12 May, 2026; originally announced May 2026.

  33. arXiv:2605.11400  [pdf, ps, other

    cs.MM

    UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning

    Authors: Hayes Bai, Yinyi Luo, Wenwen Wang, Qingsong Wen, Jindong Wang

    Abstract: Unified multimodal models (UMMs) aim to integrate understanding and generation within a single architecture. However, it remains underexplored how to effectively coordinate these two capabilities for more effective and efficient reasoning. Existing coordination approaches either perform coupling during training, without explicit inference-time coordination, or impose a fixed coordination pattern f… ▽ More

    Submitted 11 May, 2026; originally announced May 2026.

  34. arXiv:2605.10870  [pdf, ps, other

    cs.AI

    Remember the Decision, Not the Description: A Rate-Distortion Framework for Agent Memory

    Authors: Mingxi Zou, Zhihan Guo, Langzhang Liang, Zhuo Wang, Qifan Wang, Qingsong Wen, Irwin King, Lizhen Qu, Zenglin Xu

    Abstract: Long-horizon language agents must operate under limited runtime memory, yet existing memory mechanisms often organize experience around descriptive criteria such as relevance, salience, or summary quality. For an agent, however, memory is valuable not because it faithfully describes the past, but because it preserves the distinctions between histories that must remain separated under a fixed budge… ▽ More

    Submitted 11 May, 2026; originally announced May 2026.

  35. arXiv:2605.08935  [pdf, ps, other

    cs.AI cs.LG

    PnP-Corrector: A Universal Correction Framework for Coupled Spatiotemporal Forecasting

    Authors: Hao Wu, Fan Xu, Yuxu Lu, Penghao Zhao, Fan Zhang, Hao Jia, Yuxuan Liang, Ruijian Gou, Qingsong Wen, Xian Wu, Xiaomeng Huang, Yuan Gao

    Abstract: Coupled spatiotemporal forecasting is important for predicting the future evolution of multiple interacting dynamical systems, such as in climate models. However, existing methods are severely constrained by the persistent bottleneck of compounding errors. In coupled systems, errors from each subsystem simulator propagate and amplify one another, a phenomenon we term Reciprocal Error Amplification… ▽ More

    Submitted 2 June, 2026; v1 submitted 9 May, 2026; originally announced May 2026.

  36. arXiv:2605.06310  [pdf, ps, other

    cs.LG

    Perceive, Route and Modulate: Dynamic Pattern Recalibration for Time Series Forecasting

    Authors: Siru Zhong, Zhao Meng, Haohuan Fu, Haoyang Li, Qingsong Wen, Yuxuan Liang

    Abstract: Local temporal patterns in real-world time series continuously shift, rendering globally shared transformations suboptimal. Current deep forecasting models, despite their scale and complexity, rely on fixed weight matrices applied uniformly to all temporal tokens. This creates a static pattern response: models settle into a compromised average, unable to adapt to changing local dynamics. We introd… ▽ More

    Submitted 7 May, 2026; originally announced May 2026.

    Comments: 22 pages, 6 figures. Preprint

  37. arXiv:2605.01520  [pdf, ps, other

    cs.CV cs.CL

    MIRL: Mutual Information-Guided Reinforcement Learning for Vision-Language Models

    Authors: Yin Zhang, Jiaxuan Zhao, Zonghan Wu, Zengxiang Li, Junfeng Fang, Kun Wang, Qingsong Wen, Yilei Shao

    Abstract: Vision-Language Models (VLMs) frequently suffer from visual perception errors and hallucinations that compromise answer accuracy in complex reasoning tasks. Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising solution by optimizing policies using answer correctness signals. Despite their effectiveness, prevailing RLVR methods face two critical limitations. First, much of the s… ▽ More

    Submitted 2 May, 2026; originally announced May 2026.

  38. arXiv:2604.27646  [pdf, ps, other

    q-bio.CB

    Benchmarking virtual cell models for in-the-wild perturbation response

    Authors: Xinjie Mao, Songming Zhang, Qianhong Wen, Xiangyu Wen, Kedu Jin, Hao Wu, Shuizhou Chen, Yuqiang Li, Lei Bai, Qi Liu, Ning Ding, Siqi Sun, Zhangyang Gao

    Abstract: Virtual cell (VC) models aim to predict cellular responses to any perturbations in silico and have emerged as a promising approach for drug discovery and precision medicine. Yet, a clear gap still remains: while models routinely reported impressive results on standard benchmarks, it is unclear whether their predictions are truly meaningful in practice. This is mainly due to limitations in current… ▽ More

    Submitted 30 April, 2026; originally announced April 2026.

  39. arXiv:2604.25721  [pdf, ps, other

    cs.HC

    Designing and Evaluating Next-Generation Learning Interfaces: Linking AI, HCI, and the Learning Sciences

    Authors: Meng Xia, Yan Chen, Qiao Jin, Yang Shi, Paul Denny, Tiffany Barnes, Qingsong Wen, Vincent Aleven

    Abstract: This workshop addresses this gap by bringing together researchers and practitioners from AI, HCI, and the learning sciences to explore how interactive systems can better support learning. We focus on the design and evaluation of human-AI collaborative learning interfaces that are technically robust, human-centered, and pedagogically grounded. By fostering interdisciplinary dialogue, the workshop a… ▽ More

    Submitted 28 April, 2026; originally announced April 2026.

  40. arXiv:2604.24041  [pdf, ps, other

    cs.LG cs.AI

    End-to-End Learning for Partially-Observed Time Series with PyPOTS

    Authors: Wenjie Du, Yiyuan Yang, Tianxiang Zhan, Qingsong Wen

    Abstract: Partially-observed time series (POTS) is ubiquitous in real-world applications, yet most existing toolchains separate missing-value handling from downstream learning, which limits reproducibility and overall performance. This tutorial introduces PyPOTS, an open-source Python ecosystem for end-to-end data mining and machine learning on POTS. We present practical workflows spanning missingness simul… ▽ More

    Submitted 27 April, 2026; originally announced April 2026.

    Comments: Accepted by KDD 2026

  41. Bayesian Active Learning with Gaussian Processes Guided by LLM Relevance Scoring for Dense Passage Retrieval

    Authors: Junyoung Kim, Anton Korikov, Jiazhou Liang, Justin Cui, Yifan Simon Liu, Qianfeng Wen, Mark Zhao, Scott Sanner

    Abstract: While Large Language Models (LLMs) exhibit exceptional zero-shot relevance modeling, their high computational cost necessitates framing passage retrieval as a budget-constrained global optimization problem. Existing approaches passively rely on first-stage dense retrievers, which leads to two limitations: (1) failing to retrieve relevant passages in semantically distinct clusters, and (2) failing… ▽ More

    Submitted 20 April, 2026; originally announced April 2026.

    Comments: ACL 2026 Findings

    Journal ref: Findings of the Association for Computational Linguistics: ACL 2026

  42. arXiv:2604.08061  [pdf, ps, other

    physics.flu-dyn

    Effects of Soret diffusion on the intrinsic instability of premixed hydrogen/air flames

    Authors: Qizhe Wen, Yan Wang, Linlin Yang, Youhi Morii, Thorsten Zirwes, Shengkai Wang, Zheng Chen

    Abstract: Hydrogen flames exhibit multiple intrinsic instabilities. The low molar masses of H and H2 lead to significant Soret diffusion near the flame front; however, its influence on hydrogen flame instabilities remains to be quantified. This study investigates the effect of Soret diffusion on instability evolution dynamics via one-dimensional counterflow analysis and two-dimensional, high-fidelity direct… ▽ More

    Submitted 5 July, 2026; v1 submitted 9 April, 2026; originally announced April 2026.

  43. arXiv:2604.01591  [pdf, ps, other

    cs.AI

    ThinkTwice: Jointly Optimizing Large Language Models for Reasoning and Self-Refinement

    Authors: Difan Jiao, Qianfeng Wen, Blair Yang, Zhenwei Tang, Ashton Anderson

    Abstract: We introduce ThinkTwice, a simple two-phase framework that jointly optimizes LLMs to solve reasoning problems and refine the answers, based on Group Relative Policy Optimization (GRPO). In each pair of training steps, ThinkTwice first optimizes the model on solving reasoning problems, then optimizes it on refining its own solutions to the same problems, using the same binary correctness reward in… ▽ More

    Submitted 6 April, 2026; v1 submitted 2 April, 2026; originally announced April 2026.

    Comments: 27 pages,7 figures,5 tables

  44. arXiv:2603.28503  [pdf, ps, other

    cs.CV

    Bridging the Geometry Mismatch: Frequency-Aware Anisotropic Serialization for Thin-Structure SSMs

    Authors: Jin Bai, Huiyao Zhang, Qi Wen, Ningyang Li, Shengyang Li, Atta ur Rahman, Xiaolin Tian

    Abstract: The segmentation of thin linear structures is inherently topology allowbreak-critical, where minor local errors can sever long-range connectivity. While recent State-Space Models (SSMs) offer efficient long-range modeling, their isotropic serialization (e.g., raster scanning) creates a geometry mismatch for anisotropic targets, causing state propagation across rather than along the structure traje… ▽ More

    Submitted 30 March, 2026; originally announced March 2026.

  45. arXiv:2603.28249  [pdf, ps, other

    physics.flu-dyn

    Effects of gravity on lean hydrogen/air flame instability: From linear scaling law to nonlinear morphology evolution

    Authors: Qizhe Wen, Yan Wang, Linlin Yang, Yiqing Wang, Thorsten Zirwes, Shengkai Wang, Zheng Chen

    Abstract: The instability characteristics of lean hydrogen/air flames have attracted considerable research attention, yet the effect of gravity remains insufficiently understood. In this study, time-resolved two-dimensional simulations with detailed chemistry and transport are conducted to investigate the influence of gravity-induced Rayleigh-Taylor (RT) instability on the linear growth rate of disturbances… ▽ More

    Submitted 30 March, 2026; originally announced March 2026.

  46. arXiv:2603.24961  [pdf, ps, other

    cs.AI cs.CL cs.CV

    Can MLLMs Read Students' Minds? Unpacking Multimodal Error Analysis in Handwritten Math

    Authors: Dingjie Song, Tianlong Xu, Yi-Fan Zhang, Hang Li, Zhiling Yan, Xing Fan, Haoyang Li, Lichao Sun, Qingsong Wen

    Abstract: Assessing student handwritten scratchwork is crucial for personalized educational feedback but presents unique challenges due to diverse handwriting, complex layouts, and varied problem-solving approaches. Existing educational NLP primarily focuses on textual responses and neglects the complexity and multimodality inherent in authentic handwritten scratchwork. Current multimodal large language mod… ▽ More

    Submitted 25 March, 2026; originally announced March 2026.

    Comments: Accepted by the 27th International Conference on Artificial Intelligence in Education (AIED'26)

    ACM Class: I.2.7; K.3.1

  47. arXiv:2603.23268  [pdf, ps, other

    cs.LG cs.AI

    SafeSeek: Universal Attribution of Safety Circuits in Language Models

    Authors: Miao Yu, Siyuan Fu, Moayad Aloqaily, Zhenhong Zhou, Safa Otoum, Xing fan, Kun Wang, Yufei Guo, Qingsong Wen

    Abstract: Mechanistic interpretability reveals that safety-critical behaviors (e.g., alignment, jailbreak, backdoor) in Large Language Models (LLMs) are grounded in specialized functional components. However, existing safety attribution methods struggle with generalization and reliability due to their reliance on heuristic, domain-specific metrics and search algorithms. To address this, we propose \ourmetho… ▽ More

    Submitted 24 March, 2026; originally announced March 2026.

  48. arXiv:2603.20510  [pdf, ps, other

    cs.AI

    Grounded Chess Reasoning in Language Models via Master Distillation

    Authors: Zhenwei Tang, Qianfeng Wen, Seth Grief-Albert, Yahya Elgabra, Blair Yang, Honghua Dong, Ashton Anderson

    Abstract: Language models often lack grounded reasoning capabilities in specialized domains where training data is scarce but bespoke systems excel. We introduce a general framework for distilling expert system reasoning into natural language chain-of-thought explanations, enabling compact models to acquire domain expertise and the ability to generate faithful, grounded explanations. Rather than distilling… ▽ More

    Submitted 23 June, 2026; v1 submitted 20 March, 2026; originally announced March 2026.

  49. arXiv:2603.19899  [pdf, ps, other

    stat.ML cs.LG stat.AP

    Deep Autocorrelation Modeling for Time-Series Forecasting: Progress and Prospects

    Authors: Hao Wang, Licheng Pan, Qingsong Wen, Jialin Yu, Zhichao Chen, Chunyuan Zheng, Xiaoxi Li, Zhixuan Chu, Chao Xu, Mingming Gong, Haoxuan Li, Yuan Lu, Zhouchen Lin, Philip Torr, Yan Liu

    Abstract: Autocorrelation is a defining characteristic of time-series data, where each observation is statistically dependent on its predecessors. In the context of deep time-series forecasting, autocorrelation arises in both the input history and the label sequences, presenting two central research challenges: (1) designing neural architectures that model autocorrelation in history sequences, and (2) devis… ▽ More

    Submitted 20 March, 2026; originally announced March 2026.

  50. arXiv:2603.14783  [pdf, ps, other

    cs.LG

    Orthogonal Subspace Clustering: Enhancing High-Dimensional Data Analysis through Adaptive Dimensionality Reduction and Efficient Clustering

    Authors: Qing-Yuan Wen, Da-Qing Zhang

    Abstract: This paper presents Orthogonal Subspace Clustering (OSC), an innovative method for high-dimensional data clustering. We first establish a theoretical theorem proving that high-dimensional data can be decomposed into orthogonal subspaces in a statistical sense, whose form exactly matches the paradigm of Q-type factor analysis. This theorem lays a solid mathematical foundation for dimensionality red… ▽ More

    Submitted 15 March, 2026; originally announced March 2026.