Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 69 results for author: Xuan, W

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.11528  [pdf, ps, other

    cs.CL

    Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment

    Authors: Haokai Zhao, Yunze Xiao, Weihao Xuan, Flora Salim, Benjamin Tag, Aditya Joshi

    Abstract: Group alignment adapts a language model to a demographic group to produce responses that reflect the group's opinions, values, and preferences. Sycophancy, a well-documented by-product of alignment, causes the model to over-agree with the user regardless of factual and objective information. However, existing group alignment methods and evaluations focus only on how closely the model matches the g… ▽ More

    Submitted 11 August, 2026; originally announced August 2026.

    Comments: 9 pages main text, 23 pages in total, under review

  2. arXiv:2608.11116  [pdf, ps, other

    cs.AR

    You Only Charge Once 2.0 : A End-to-End Analog Computing-in-Memory Architecture with Reconfigurable Switched Capacitors

    Authors: Zihao Xuan, Yewen Li, Jia Chen, Wei Xuan, Xiao Huo, Fengbin Tu

    Abstract: Analog Computing-in-Memory (ACiM) accelerates deep neural networks by keeping weights inside memory arrays and executing dot products in the analog domain. However, modern ACiM accelerators are often limited by the "ADC wall": analog-to-digital converters consume a large fraction of energy and area, while bit-sliced execution repeatedly invokes these converters. Existing designs reduce this cost w… ▽ More

    Submitted 11 August, 2026; originally announced August 2026.

    Comments: 13 pages, 19 figures

  3. arXiv:2607.28087  [pdf, ps, other

    cs.AI

    Diversifying Personalized Research Ideation against AI-Induced Homogenization

    Authors: Rui Xu, Yunke Wang, Linwei Tao, Wenjie Xuan, Yong Luo

    Abstract: AI-assisted research ideation has emerged as a promising paradigm for accelerating scientific discovery, with systems now capable of generating research directions conditioned on papers, topics, or lightweight researcher contexts. Yet current systems largely optimize individual suggestions in isolation. This leaves two blind spots. First, coarse researcher representations may elicit mainstream dir… ▽ More

    Submitted 30 July, 2026; originally announced July 2026.

  4. arXiv:2607.25933  [pdf, ps, other

    cs.CL cs.AI

    Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases

    Authors: Rui Yang, Weihao Xuan, Yi Lin, Zhuhan Bao, Jonathan Chong Kai Liew, Matthew Yu Heng Wong, Nicolás Lescano, Nikita R. Paripati, Emily Ling-Lin Pai, Jiarui Liu, Heli Qi, Heng-Jui Chang, Benny Kai Guo Loo, Huitao Li, Kunyu Yu, Yufan Wang, Chuan Hong, Shijian Lu, Douglas Teodoro, Naoto Yokoya, Ross Koppel, Mona Diab, Hua Xu, David W. Bates, Nan Liu , et al. (1 additional authors not shown)

    Abstract: Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive disclosure of multimodal information, dynamic updating of diagnostic hypotheses, and continuous refinement of clinical reasoning. However, existing evaluations of multimodal large language models (MLLMs) typically rely on sin… ▽ More

    Submitted 28 July, 2026; originally announced July 2026.

  5. arXiv:2607.22746  [pdf, ps, other

    cs.CV cs.AI eess.IV

    Advancing All-Weather Building Damage Mapping to the Instance Level: Outcomes and Insights from the 2026 Bright Challenge

    Authors: Hongruixuan Chen, He Huang, Haifeng Wang, Jian Song, Junjue Wang, Weihao Xuan, Hamish Mitchell, Jiepan Li, Wei He, Liangpei Zhang, Zijie Wang, Chen Zhong, Jiazhen Zhao, Lei Hu, Ting Hu, Hongyan Zhang, Gregory Angelides, Miriam Cha, Clifford Broni-Bediako, Junshi Xia, Taylor Perron, Naoto Yokoya

    Abstract: Rapid post-disaster response requires timely, building-level information on whether structures remain intact, are damaged, or are destroyed. Post-event optical imagery, however, may be unavailable because of cloud, smoke, or darkness. The Bright Challenge evaluated all-weather building damage mapping from a submeter-resolution pre-event optical image and a post-event SAR image. Participants were r… ▽ More

    Submitted 23 July, 2026; originally announced July 2026.

  6. RRAM-DP: Device-Calibrated Differential Privacy for In-Memory Edge Learning

    Authors: Kwunhang Wong, Jichang Yang, Karl M. H. Lai, Hegan Chen, Songqi Wang, Wei Xuan, Ning Lin, Han Wang, Xiaojuan Qi, Zhongrui Wang

    Abstract: Edge Artificial Intelligence of Things (AIoT) systems often collect sensitive data in situ, raising serious privacy concerns. Resistive-switching random-access memory (RRAM) is an attractive substrate for efficient AIoT thanks to its multi-bit storage and compute-in-memory (CiM) capabilities, while its inherently stochastic write behavior provides a natural source of randomness that can be leverag… ▽ More

    Submitted 31 July, 2026; v1 submitted 20 July, 2026; originally announced July 2026.

    Comments: International Conference on Computer-Aided Design 2026

  7. arXiv:2607.13924  [pdf, ps, other

    cs.HC

    ExpressionCueLens: A Cross-Cultural Analysis of Human-AI Companion Conversations on Social Media

    Authors: Lynnette Hui Xian Ng, Yunze Xiao, Lionel Z. Wang, Weihao Xuan, Mona Diab

    Abstract: LLM-based AI companion agents are increasingly being perceived not only as tools but also as social companions. On social media, people recount conversations where these agents comfort, negotiate and assert boundaries, reflecting a growing attribution of human-like qualities. To profile how agency is perceived in human-AI (HAI) interactions, we introduce the ExpressionCueLens framework, which orga… ▽ More

    Submitted 18 July, 2026; v1 submitted 15 July, 2026; originally announced July 2026.

    Comments: Accepted at Journal of Ambient Intelligence and Humanized Computing

  8. arXiv:2607.00491  [pdf, ps, other

    cs.CV cs.AI cs.CL

    MindEdit-Bench: Benchmarking Object-Level Counterfactual Spatial Reasoning in VLMs from In-the-Wild Photos

    Authors: Leyuan Yu, Xiao Tang, Minghao Liu, Xinyuan Li, Xiaokai Bai, Sheng Zhou, Qunshu Lin, Weihao Xuan, Naoto Yokoya

    Abstract: Benchmarks for vision-language models (VLMs) mostly test observational spatial reasoning: models describe relations already visible in the input. Existing what-if tasks typically vary the observer while keeping the scene fixed. Can VLMs instead predict the consequences of hypothetically moving or rotating an object? We introduce MindEdit-Bench, a benchmark of six spatial reasoning tasks built from… ▽ More

    Submitted 1 July, 2026; originally announced July 2026.

    Comments: 18 pages, 7 figures. Dataset available at https://huggingface.co/datasets/ZODAOfficial/MindEdit-Bench

  9. arXiv:2606.23830  [pdf, ps, other

    cs.LG cs.AI

    Deciphering Fingerprints of 3D Molecular Surfaces for Accurate Epitope Prediction

    Authors: Fang Wu, Weihao Xuan, Jure Leskovec, Yejin Choi, Li Erran Li

    Abstract: Molecular surfaces encode the geometric and physicochemical patterns that determine antibody-antigen recognition, central to epitope prediction. However, existing methods rely on sequences or backbone structures and struggle to capture discontinuous, surface-driven epitopes. This study presents SurfBind, a surface-centric learning framework for epitope prediction that operates directly on molecula… ▽ More

    Submitted 22 June, 2026; originally announced June 2026.

    Journal ref: KDD 2026 AI4Science

  10. arXiv:2606.15345  [pdf, ps, other

    cs.CL cs.IR

    Beyond Monolingual Deep Research: Evaluating Agents and Retrievers with Cross-Lingual BrowseComp-Plus

    Authors: Yuheng Lu, Qingcheng Zeng, Heli Qi, Puxuan Yu, Fuheng Zhao, Rui Yang, Hitomi Yanaka, Naoto Yokoya, Weihao Xuan

    Abstract: Deep research agents are increasingly evaluated on their ability to search for evidence, reason over retrieved sources, and produce grounded answers. Existing browsing benchmarks, however, largely assume that the user's query and the supporting evidence are written in the same language, leaving open whether agentic search systems can operate when relevant evidence appears in another language. We i… ▽ More

    Submitted 17 June, 2026; v1 submitted 13 June, 2026; originally announced June 2026.

    Comments: Preprint

  11. arXiv:2606.11232  [pdf, ps, other

    cs.CL cs.AI

    Every Act Has Its Price: Compressed Moral Composition in Frontier LLMs

    Authors: Weijia Zhang, Ruiqi Chen, Yunze Xiao, Weihao Xuan

    Abstract: Existing LLM moral benchmarks usually ask which isolated moral act, value, or foundation a model prefers. This is useful but incomplete. Realistic judgments often require a model to combine several moral signals within the same option. We introduce **Moral Trolley Arena**, a two-stage blind ELO benchmark for measuring how LLMs compose moral evidence. The single-scene arena first calibrates individ… ▽ More

    Submitted 28 May, 2026; originally announced June 2026.

  12. arXiv:2606.05104  [pdf, ps, other

    cs.AI

    Knowledge Index of Noah's Ark

    Authors: Sheng Jin, Minghao Liu, Yunze Xiao, Zeqi Zhou, Heli Qi, Yifan Yao, Meishu Song, Kaijing Ma, Xuan Zhang, Sicong Jiang, Yizhe Li, Ningshan Ma, Jie Wei, Ziniu Li, Minglai Yang, Bangya Liu, Yiming Liang, Xiao Fang, Qingcheng Zeng, Jiarui Liu, Rui Yang, Shen Yan, Wenhao Huang, Jiaheng Liu, Zihan Wang , et al. (2 additional authors not shown)

    Abstract: Knowledge benchmarks for LLMs face three issues: scaling-driven designs that do not operationalize disciplinary representativeness; flat-payment annotation that permits lazy consensus; and unaudited ranking instability under bounded test budgets. We introduce KINA, an 899-item benchmark across 261 fine-grained disciplines, with two formal results. First, we cast representativeness as a coverage-st… ▽ More

    Submitted 4 June, 2026; v1 submitted 3 June, 2026; originally announced June 2026.

  13. arXiv:2605.16865  [pdf, ps, other

    cs.CL

    MixSD: Mixed Contextual Self-Distillation for Knowledge Injection

    Authors: Jiarui Liu, Lechen Zhang, Yongjin Yang, Yinghui He, Yingheng Wang, Weihao Xuan, Zhijing Jin, Mona Diab

    Abstract: Supervised fine-tuning (SFT) is widely used to inject new knowledge into language models, but it often degrades pretrained capabilities such as reasoning and general-domain performance. We argue this forgetting arises because fine-tuning targets from humans or external systems diverge from the model's autoregressive distribution, forcing the optimizer to imitate low-probability token sequences. To… ▽ More

    Submitted 17 June, 2026; v1 submitted 16 May, 2026; originally announced May 2026.

  14. arXiv:2605.11931  [pdf, ps, other

    cs.CV

    Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training

    Authors: Qihuang Zhong, Liang Ding, Wenjie Xuan, Juhua Liu, Bo Du, Dacheng Tao

    Abstract: Post-training with explicit reasoning traces is common to improve the reasoning capabilities of Multimodal Large Language Models (MLLMs). However, acquiring high-quality reasoning traces is often costly and time-consuming. Hence, the self-improvement paradigm has emerged, enabling MLLMs to self-generate reasoning traces for training without external supervision. Despite its effectiveness, we revea… ▽ More

    Submitted 12 May, 2026; originally announced May 2026.

    Comments: Accepted by ICML 2026

  15. arXiv:2605.11633  [pdf, ps, other

    cs.AI

    Can LLM Agents Respond to Disasters? Benchmarking Heterogeneous Geospatial Reasoning in Emergency Operations

    Authors: Junjue Wang, Weihao Xuan, Heli Qi, Pengyu Dai, Kunyi Liu, Hongruixuan Chen, Zhuo Zheng, Junshi Xia, Stefano Ermon, Naoto Yokoya

    Abstract: Operational disaster response goes beyond damage assessment, requiring responders to integrate multi-sensor signals, reason over road networks, populations and key facilities, plan evacuations, and produce actionable reports. However, prior work largely isolates remote-sensing perception or evaluates generic tool use, leaving the end-to-end workflows of emergency operations underexplored. In this… ▽ More

    Submitted 12 May, 2026; originally announced May 2026.

    Comments: DORA stress-tests LLM agents on real-world disaster operations that demand comprehensive orchestration of 108 specialized tools over heterogeneous geospatial data

    ACM Class: I.4.9

  16. arXiv:2605.06226  [pdf, ps, other

    cs.AI q-bio.GN

    A Versatile AI Agent for Rare Disease Diagnosis and Risk Gene Prioritization

    Authors: Tianyu Liu, Wangjie Zheng, Rui Yang, Benny Kai Guo Loo, Hui Zhang, Jeffries Lauran, Jianlei Gu, Botao Yu, Weihao Xuan, Kexin Huang, Nan Liu, James Zou, Yonghui Jiang, Hua Xu, Hongyu Zhao

    Abstract: Accurate and timely diagnosis is essential for effective treatment, particularly in the context of rare diseases. However, current diagnostic workflows often lead to prolonged assessment times and low accuracy. To address these limitations, we introduce Hygieia, a multi-modal AI agent system designed to support precision disease diagnosis by integrating diverse data sources, including phenotypic f… ▽ More

    Submitted 9 May, 2026; v1 submitted 7 May, 2026; originally announced May 2026.

    Comments: 32 pages, 6 figures

  17. arXiv:2605.02937  [pdf, ps, other

    cs.LG cs.AI cs.CE

    Proteo-R1: Reasoning Foundation Models for De Novo Protein Design

    Authors: Fang Wu, Weihao Xuan, Heli Qi, Hanqun Cao, Heng-Jui Chang, Zeqi Zhou, Haokai Zhao, Ma Jian, Carl Ma, Yu-Chi Cheng, Kuan Pang, Xiangru Tang, Zehong Wang, Guanlue Li, Hanchen Wang, Kejun Ying, Pan Lu, Chiho Im, Seungju Han, Peng Xia, Tinson Xu, Yinxi Li, Deyao Zhu, Pheng-Ann Heng, Naoto Yokoya , et al. (4 additional authors not shown)

    Abstract: Deep learning in de novo protein design has achieved atomic-level fidelity. However, existing models remain largely non-deliberative: they directly synthesize molecular geometries without explicitly reasoning about which residues or interactions are functionally essential. As a result, design decisions are entangled with continuous sampling dynamics, limiting interpretability, controllability, and… ▽ More

    Submitted 10 August, 2026; v1 submitted 1 May, 2026; originally announced May 2026.

    Journal ref: ICML 2026

  18. arXiv:2604.25317  [pdf, ps, other

    cs.AR

    FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory Architecture

    Authors: Zihao Xuan, Jia Chen, Yewen Li, Wei Xuan, Hegan Chen, Xiao Huo, Fengbin Tu

    Abstract: In this paper, we propose FusionCIM, an operator-fusion-driven compute-in-memory (CIM) accelerator architecture for efficient and scalable LLM inference, with three key innovations: (1) a hybrid CIM pipeline architecture that maps QKT computation on inner-product-based CIM (IP-CIM) and PV aggregation on outer-product-based CIM (OP-CIM) for efficient matrix multiplications fusion; (2) a QO-stationa… ▽ More

    Submitted 28 April, 2026; originally announced April 2026.

    Comments: 7 Pages, 10 figures

  19. arXiv:2604.24698  [pdf, ps, other

    cs.CL

    The Chameleon's Limit: Investigating Persona Collapse and Homogenization in Large Language Models

    Authors: Yunze Xiao, Vivienne J. Zhang, Chenghao Yang, Ningshan Ma, Weihao Xuan, Jen-tse Huang

    Abstract: Applications based on large language models (LLMs), such as multi-agent simulations, require population diversity among agents. We identify a pervasive failure mode we term \emph{Persona Collapse}: agents each assigned a distinct profile nonetheless converge into a narrow behavioral mode, producing a homogeneous simulated population. To quantify persona collapse, we propose a framework that measur… ▽ More

    Submitted 27 April, 2026; originally announced April 2026.

  20. arXiv:2604.17632  [pdf, ps, other

    cs.IR

    Code-Switching Information Retrieval: Benchmarks, Analysis, and the Limits of Current Retrievers

    Authors: Qingcheng Zeng, Yuheng Lu, Zeqi Zhou, Heli Qi, Puxuan Yu, Fuheng Zhao, Hitomi Yanaka, Weihao Xuan, Naoto Yokoya

    Abstract: Code-switching is a pervasive linguistic phenomenon in global communication, yet modern information retrieval systems remain predominantly designed for, and evaluated within, monolingual contexts. To bridge this critical disconnect, we present a holistic study dedicated to code-switching IR. We introduce CSR-L (Code-Switching Retrieval benchmark-Lite), constructing a dataset via human annotation t… ▽ More

    Submitted 19 April, 2026; originally announced April 2026.

    Comments: Finding of ACL 2026

  21. arXiv:2604.07607  [pdf, ps, other

    cs.RO cs.CV

    EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World

    Authors: Ryan Punamiya, Simar Kareer, Zeyi Liu, Josh Citron, Ri-Zhao Qiu, Xiongyi Cai, Alexey Gavryushin, Jiaqi Chen, Davide Liconti, Lawrence Y. Zhu, Patcharapong Aphiwetsa, Baoyu Li, Aniketh Cheluva, Pranav Kuppili, Yangcen Liu, Dhruv Patel, Aidan Gao, Hye-Young Chung, Ryan Co, Renee Zbizika, Jeff Liu, Xiaomeng Xu, Haoyu Xiong, Geng Chen, Sebastiano Oliani , et al. (15 additional authors not shown)

    Abstract: Robot learning increasingly depends on large and diverse data, yet robot data collection remains expensive and difficult to scale. Egocentric human data offer a promising alternative by capturing rich manipulation behavior across everyday environments. However, existing human datasets are often limited in scope, difficult to extend, and fragmented across institutions. We introduce EgoVerse, a coll… ▽ More

    Submitted 7 July, 2026; v1 submitted 8 April, 2026; originally announced April 2026.

  22. arXiv:2604.06409  [pdf, ps, other

    cs.CR cs.AI cs.CL

    Say Something Else: Rethinking Contextual Privacy as Information Sufficiency

    Authors: Yunze Xiao, Wenkai Li, Xiaoyuan Wu, Ningshan Ma, Yueqi Song, Weihao Xuan

    Abstract: LLM agents increasingly draft messages on behalf of users, yet users routinely overshare sensitive information and disagree on what counts as private. Existing systems support only suppression (omitting sensitive information) and generalization (replacing information with an abstraction), and are typically evaluated on single isolated messages, leaving both the strategy space and evaluation settin… ▽ More

    Submitted 7 April, 2026; originally announced April 2026.

  23. arXiv:2603.22148  [pdf, ps, other

    cs.CV

    OpenEarth-Agent: From Tool Calling to Tool Creation for Open-Environment Earth Observation

    Authors: Sijie Zhao, Feng Liu, Xueliang Zhang, Hao Chen, Xinyu Gu, Zhe Jiang, Fenghua Ling, Ben Fei, Wenlong Zhang, Junjue Wang, Weihao Xuan, Pengfeng Xiao, Naoto Yokoya, Lei Bai

    Abstract: Earth Observation (EO) is essential for perceiving dynamic land surface changes, yet deploying autonomous EO in open environments is hindered by the immense diversity of multi-source data and heterogeneous tasks. While remote sensing agents have emerged to streamline EO workflows, existing tool-calling agents are confined to closed environments. They rely on pre-defined tools and are restricted to… ▽ More

    Submitted 23 March, 2026; originally announced March 2026.

    Comments: 15 pages, 4 figures

  24. arXiv:2602.20521  [pdf, ps, other

    cs.CR

    Towards Secure and Efficient DNN Accelerators via Hardware-Software Co-Design

    Authors: Wei Xuan, Zihao Xuan, Rongliang Fu, Ning Lin, Kwunhang Wong, Zikang Yuan, Lang Feng, Zhongrui Wang, Tsung-Yi Ho, Yuzhong Jiao, Luhong Liang

    Abstract: The rapid deployment of deep neural network (DNN) accelerators in safety-critical domains such as autonomous vehicles, healthcare systems, and financial infrastructure necessitates robust mechanisms to safeguard data confidentiality and computational integrity. Existing security solutions for DNN accelerators, however, suffer from excessive hardware resource demands and frequent off-chip memory ac… ▽ More

    Submitted 23 February, 2026; originally announced February 2026.

  25. arXiv:2602.19063  [pdf, ps, other

    cs.CV

    Direction-aware 3D Large Multimodal Models

    Authors: Quan Liu, Weihao Xuan, Junjue Wang, Naoto Yokoya, Ling Shao, Shijian Lu

    Abstract: 3D large multimodal models (3D LMMs) rely heavily on ego poses for enabling directional question-answering and spatial reasoning. However, most existing point cloud benchmarks contain rich directional queries but lack the corresponding ego poses, making them inherently ill-posed in 3D large multimodal modelling. In this work, we redefine a new and rigorous paradigm that enables direction-aware 3D… ▽ More

    Submitted 22 February, 2026; originally announced February 2026.

    Comments: In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026

  26. arXiv:2602.02559  [pdf, ps, other

    cs.AI cs.CV cs.LG cs.MA

    Experience-Driven Multi-Agent Systems Are Training-free Context-aware Earth Observers

    Authors: Pengyu Dai, Weihao Xuan, Junjue Wang, Hongruixuan Chen, Jian Song, Yafei Ou, Naoto Yokoya

    Abstract: Recent advances have enabled large language model (LLM) agents to solve complex tasks by orchestrating external tools. However, these agents often struggle in specialized, tool-intensive domains that demand long-horizon execution, tight coordination across modalities, and strict adherence to implicit tool constraints. Earth Observation (EO) tasks exemplify this challenge due to the multi-modal and… ▽ More

    Submitted 30 January, 2026; originally announced February 2026.

    Comments: 21 pages, 6 figures

  27. arXiv:2601.18027  [pdf, ps, other

    cs.AI cs.CL

    Sentipolis: Emotion-Aware Agents for Social Simulations

    Authors: Chiyuan Fu, Lyuhao Chen, Yunze Xiao, Weihao Xuan, Carlos Busso, Mona Diab

    Abstract: LLM agents are increasingly used for social simulation, yet emotion is often treated as a transient cue, causing emotional amnesia and weak long-horizon continuity. We present Sentipolis, a framework for emotionally stateful agents that integrates continuous Pleasure-Arousal-Dominance (PAD) representation, dual-speed emotion dynamics, and emotion--memory coupling. Across thousands of interactions… ▽ More

    Submitted 20 April, 2026; v1 submitted 25 January, 2026; originally announced January 2026.

  28. arXiv:2601.07264  [pdf, ps, other

    cs.CL

    The Confidence Dichotomy: Analyzing and Mitigating Miscalibration in Tool-Use Agents

    Authors: Weihao Xuan, Qingcheng Zeng, Heli Qi, Yunze Xiao, Junjue Wang, Naoto Yokoya

    Abstract: Autonomous agents based on large language models (LLMs) are rapidly evolving to handle multi-turn tasks, but ensuring their trustworthiness remains a critical challenge. A fundamental pillar of this trustworthiness is calibration, which refers to an agent's ability to express confidence that reliably reflects its actual performance. While calibration is well-established for static models, its dyna… ▽ More

    Submitted 12 January, 2026; originally announced January 2026.

  29. arXiv:2601.05473  [pdf, ps, other

    cs.CL cs.HC

    Towards Valid Student Simulation with Large Language Models

    Authors: Zhihao Yuan, Yunze Xiao, Ming Li, Weihao Xuan, Richard Tong, Mona Diab, Tom Mitchell

    Abstract: This paper presents a conceptual and methodological framework for large language model (LLM) based student simulation in educational settings. The authors identify a core failure mode, termed the "competence paradox" in which broadly capable LLMs are asked to emulate partially knowledgeable learners, leading to unrealistic error patterns and learning dynamics. To address this, the paper reframes s… ▽ More

    Submitted 8 January, 2026; originally announced January 2026.

  30. arXiv:2601.02186  [pdf

    cs.CL

    Toward Global Large Language Models in Medicine

    Authors: Rui Yang, Huitao Li, Weihao Xuan, Heli Qi, Xin Li, Kunyu Yu, Yingjian Chen, Rongrong Wang, Jacques Behmoaras, Tianxi Cai, Bibhas Chakraborty, Qingyu Chen, Lionel Tim-Ee Cheng, Marie-Louise Damwanza, Chido Dzinotyiwei, Aosong Feng, Chuan Hong, Yusuke Iwasawa, Yuhe Ke, Linah Kitala, Taehoon Ko, Jisan Lee, Irene Li, Jonathan Chong Kai Liew, Hongfang Liu , et al. (25 additional authors not shown)

    Abstract: Despite continuous advances in medical technology, the global distribution of health care resources remains uneven. The development of large language models (LLMs) has transformed the landscape of medicine and holds promise for improving health care quality and expanding access to medical information globally. However, existing LLMs are primarily trained on high-resource languages, limiting their… ▽ More

    Submitted 5 January, 2026; originally announced January 2026.

    Comments: 182 pages, 65 figures

  31. arXiv:2512.07276  [pdf, ps, other

    cs.CV

    Geo3DVQA: Evaluating Vision-Language Models for 3D Geospatial Reasoning from Aerial Imagery

    Authors: Mai Tsujimoto, Junjue Wang, Weihao Xuan, Naoto Yokoya

    Abstract: Three-dimensional geospatial analysis is critical for applications in urban planning, climate adaptation, and environmental assessment. However, current methodologies depend on costly, specialized sensors, such as LiDAR and multispectral sensors, which restrict global accessibility. Additionally, existing sensor-based and rule-driven methods struggle with tasks requiring the integration of multipl… ▽ More

    Submitted 21 December, 2025; v1 submitted 8 December, 2025; originally announced December 2025.

    Comments: Accepted at WACV 2026. 32 pages long including the appendix. Revision details are provided in the supplements

  32. arXiv:2511.17652  [pdf, ps, other

    q-bio.QM cs.CV

    TeamPath: Building MultiModal Pathology Experts with Reasoning AI Copilots

    Authors: Tianyu Liu, Weihao Xuan, Hao Wu, Peter Humphrey, Marcello DiStasio, Mohamed Kahila, Alfonso Garcia Tan, Heli Qi, Rui Yang, Simeng Han, Tinglin Huang, Fang Wu, Chen Liu, Qingyu Chen, Nan Liu, Irene Li, Hua Xu, Hongyu Zhao

    Abstract: Advances in AI have introduced several strong models in computational pathology to usher it into the era of multi-modal diagnosis, analysis, and interpretation. However, the current pathology-specific visual language models still lack capacities in making the diagnosis with rigorous reasoning paths as well as handling divergent tasks, and thus, challenges of building AI Copilots for real scenarios… ▽ More

    Submitted 6 April, 2026; v1 submitted 20 November, 2025; originally announced November 2025.

    Comments: 45 pages, 6 figures

  33. arXiv:2511.09228  [pdf, ps, other

    cs.CV cs.CL

    Taming Object Hallucinations with Verified Atomic Confidence Estimation

    Authors: Jiarui Liu, Weihao Xuan, Zhijing Jin, Mona Diab

    Abstract: Multimodal Large Language Models (MLLMs) often suffer from hallucinations, particularly errors in object existence, attributes, or relations, which undermine their reliability. We introduce TACO (Verified Atomic Confidence Estimation), a simple framework that mitigates hallucinations through self-verification and confidence calibration without relying on external vision experts. TACO decomposes re… ▽ More

    Submitted 12 November, 2025; originally announced November 2025.

  34. arXiv:2511.05901  [pdf

    cs.CL cs.AI

    Retrieval-Augmented Generation in Medicine: A Scoping Review of Technical Implementations, Clinical Applications, and Ethical Considerations

    Authors: Rui Yang, Matthew Yu Heng Wong, Huitao Li, Xin Li, Wentao Zhu, Jingchi Liao, Kunyu Yu, Jonathan Chong Kai Liew, Weihao Xuan, Yingjian Chen, Yuhe Ke, Jasmine Chiat Ling Ong, Douglas Teodoro, Chuan Hong, Daniel Shi Wei Ting, Nan Liu

    Abstract: The rapid growth of medical knowledge and increasing complexity of clinical practice pose challenges. In this context, large language models (LLMs) have demonstrated value; however, inherent limitations remain. Retrieval-augmented generation (RAG) technologies show potential to enhance their clinical applicability. This study reviewed RAG applications in medicine. We found that research primarily… ▽ More

    Submitted 13 November, 2025; v1 submitted 8 November, 2025; originally announced November 2025.

  35. arXiv:2509.25454  [pdf, ps, other

    cs.AI cs.CL

    DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search

    Authors: Fang Wu, Weihao Xuan, Heli Qi, Ximing Lu, Aaron Tu, Li Erran Li, Yejin Choi

    Abstract: Although RLVR has become an essential component for developing advanced reasoning skills in language models, contemporary studies have documented training plateaus after thousands of optimization steps, i.e., notable decreases in performance gains despite increased computational investment. This limitation stems from the sparse exploration patterns inherent in current RLVR practices, where models… ▽ More

    Submitted 6 April, 2026; v1 submitted 29 September, 2025; originally announced September 2025.

  36. arXiv:2509.23102  [pdf, ps, other

    cs.AI cs.CL

    Multiplayer Nash Preference Optimization

    Authors: Fang Wu, Xu Huang, Weihao Xuan, Zhiwei Zhang, Yijia Xiao, Guancheng Wan, Xiaomin Li, Bing Hu, Peng Xia, Jure Leskovec, Yejin Choi

    Abstract: Reinforcement learning from human feedback (RLHF) has emerged as the standard paradigm for aligning large language models with human preferences. However, reward-based methods grounded in the Bradley-Terry assumption struggle to capture the nontransitivity and heterogeneity of real-world preferences. To address this, recent studies have reframed alignment as a two-player Nash game, giving rise to… ▽ More

    Submitted 10 August, 2026; v1 submitted 27 September, 2025; originally announced September 2025.

    Journal ref: ICLR 2026 Oral

  37. arXiv:2509.21882  [pdf, ps, other

    cs.LG cs.AI

    Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards

    Authors: Fang Wu, Aaron Tu, Weihao Xuan, Heli Qi, Xu Huang, Qingcheng Zeng, Shayan Talaei, Yijia Xiao, Peng Xia, Xiangru Tang, Yuchen Zhuang, Yinxi Li, Bing Hu, Hanqun Cao, Wenqi Shi, Rui Yang, Nan Liu, Huaxiu Yao, Ge Liu, Li Erran Li, Amin Saberi, Naoto Yokoya, Jure Leskovec, Yejin Choi

    Abstract: Reinforcement learning with verifiable rewards (RLVR) is a practical, scalable way to improve large language models on math, code, and other structured tasks. However, we argue that many headline RLVR gains are not yet well validated because reports often conflate policy improvement with three confounds: (i) budget mismatch between RLVR and baseline evaluations, (ii) attempt inflation and calibrat… ▽ More

    Submitted 25 May, 2026; v1 submitted 26 September, 2025; originally announced September 2025.

  38. arXiv:2508.19057  [pdf, ps, other

    cs.DS

    DTC: Real-Time and Accurate Distributed Triangle Counting in Fully Dynamic Graph Streams

    Authors: Wei Xuan, Yan Liang, Huawei Cao, Ning Lin, Xiaochun Ye, Dongrui Fan

    Abstract: Triangle counting is a fundamental problem in graph mining, essential for analyzing graph streams with arbitrary edge orders. However, exact counting becomes impractical due to the massive size of real-world graph streams. To address this, approximate algorithms have been developed, but existing distributed streaming algorithms lack adaptability and struggle with edge deletions. In this article, w… ▽ More

    Submitted 25 January, 2026; v1 submitted 26 August, 2025; originally announced August 2025.

    Comments: Accepted by International Symposium on Reliable Distributed Systems (SRDS) 2024

  39. arXiv:2508.18924  [pdf, ps, other

    cs.AR

    SeDA: Secure and Efficient DNN Accelerators with Hardware/Software Synergy

    Authors: Wei Xuan, Zhongrui Wang, Lang Feng, Ning Lin, Zihao Xuan, Rongliang Fu, Tsung-Yi Ho, Yuzhong Jiao, Luhong Liang

    Abstract: Ensuring the confidentiality and integrity of DNN accelerators is paramount across various scenarios spanning autonomous driving, healthcare, and finance. However, current security approaches typically require extensive hardware resources, and incur significant off-chip memory access overheads. This paper introduces SeDA, which utilizes 1) a bandwidth-aware encryption mechanism to improve hardware… ▽ More

    Submitted 26 August, 2025; originally announced August 2025.

    Comments: Accepted by Design Automation Conference (DAC), 2025

  40. arXiv:2508.04026  [pdf, ps, other

    cs.HC

    VeriWeb: Verifiable Long-Chain Web Benchmark for Agentic Information-Seeking

    Authors: Shunyu Liu, Minghao Liu, Huichi Zhou, Zhenyu Cui, Yang Zhou, Yuhao Zhou, Jialiang Gao, Heng Zhou, Yunhao Yang, Wendong Fan, puzhen zhang, Ge Zhang, Jiajun Shi, Weihao Xuan, Jiaxing Huang, Shuang Luo, Fang Wu, Heli Qi, Qingcheng Zeng, Junjie Wang, Aosong Feng, Jindi Lv, Sicong Jiang, Ziqi Ren, Wangchunshu Zhou , et al. (9 additional authors not shown)

    Abstract: Recent advances have showcased the extraordinary capabilities of Large Language Model (LLM) agents in tackling web-based information-seeking tasks. However, existing efforts mainly focus on single-fact retrieval and rely on outcome-only verification, thereby limiting their scalability in realistic knowledge-intensive scenarios that involve long-horizon web tasks requiring large-scale retrieval and… ▽ More

    Submitted 27 February, 2026; v1 submitted 5 August, 2025; originally announced August 2025.

  41. arXiv:2507.14843  [pdf, ps, other

    cs.LG cs.AI cs.CL

    The Invisible Leash: Why RLVR May or May Not Escape Its Origin

    Authors: Fang Wu, Weihao Xuan, Ximing Lu, Mingjie Liu, Yi Dong, Zaid Harchaoui, Yejin Choi

    Abstract: Recent advances highlight Reinforcement Learning with Verifiable Rewards (RLVR) as a promising method for enhancing LLMs' capabilities. However, it remains unclear whether the current practice of RLVR truly expands a model's reasoning boundary or mainly amplifies high-reward outputs that the base model already knows, thereby improving precision. This study presents an empirical investigation that… ▽ More

    Submitted 3 February, 2026; v1 submitted 20 July, 2025; originally announced July 2025.

  42. arXiv:2506.20983  [pdf, ps, other

    cs.CV

    Rethink Sparse Signals for Pose-guided Text-to-image Generation

    Authors: Wenjie Xuan, Jing Zhang, Juhua Liu, Bo Du, Dacheng Tao

    Abstract: Recent works favored dense signals (e.g., depth, DensePose), as an alternative to sparse signals (e.g., OpenPose), to provide detailed spatial guidance for pose-guided text-to-image generation. However, dense representations raised new challenges, including editing difficulties and potential inconsistencies with textual prompts. This fact motivates us to revisit sparse signals for pose guidance, o… ▽ More

    Submitted 25 June, 2025; originally announced June 2025.

    Comments: accepted by ICCV 2025

  43. arXiv:2505.21089  [pdf, ps, other

    cs.CV

    DisasterM3: A Remote Sensing Vision-Language Dataset for Disaster Damage Assessment and Response

    Authors: Junjue Wang, Weihao Xuan, Heli Qi, Zhihao Liu, Kunyi Liu, Yuhan Wu, Hongruixuan Chen, Jian Song, Junshi Xia, Zhuo Zheng, Naoto Yokoya

    Abstract: Large vision-language models (VLMs) have made great achievements in Earth vision. However, complex disaster scenes with diverse disaster types, geographic regions, and satellite sensors have posed new challenges for VLM applications. To fill this gap, we curate a remote sensing vision-language dataset (DisasterM3) for global-scale disaster assessment and response. DisasterM3 includes 26,988 bi-tem… ▽ More

    Submitted 20 October, 2025; v1 submitted 27 May, 2025; originally announced May 2025.

    Comments: A multi-hazard, multi-sensor, and multi-task vision-language dataset for global-scale disaster assessment and response

    ACM Class: I.4.9

  44. arXiv:2505.21076  [pdf, ps, other

    cs.CV

    DynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City Understanding

    Authors: Weihao Xuan, Junjue Wang, Heli Qi, Zihang Chen, Zhuo Zheng, Yanfei Zhong, Junshi Xia, Naoto Yokoya

    Abstract: Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in visual understanding, but their application to long-term Earth observation analysis remains limited, primarily focusing on single-temporal or bi-temporal imagery. To address this gap, we introduce DVL-Suite, a comprehensive framework for analyzing long-term urban dynamics through remote sensing imagery. Our suite… ▽ More

    Submitted 26 October, 2025; v1 submitted 27 May, 2025; originally announced May 2025.

    Comments: NeurIPS 2025

  45. arXiv:2505.20236  [pdf, ps, other

    cs.CV

    Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models

    Authors: Weihao Xuan, Qingcheng Zeng, Heli Qi, Junjue Wang, Naoto Yokoya

    Abstract: Uncertainty quantification is essential for assessing the reliability and trustworthiness of modern AI systems. Among existing approaches, verbalized uncertainty, where models express their confidence through natural language, has emerged as a lightweight and interpretable solution in large language models (LLMs). However, its effectiveness in vision-language models (VLMs) remains insufficiently s… ▽ More

    Submitted 26 May, 2025; originally announced May 2025.

  46. arXiv:2505.18497  [pdf, ps, other

    cs.CL

    The Pragmatic Mind of Machines: Tracing the Emergence of Pragmatic Competence in Large Language Models

    Authors: Kefan Yu, Qingcheng Zeng, Weihao Xuan, Wanxin Li, Jingyi Wu, Rob Voigt

    Abstract: Current large language models (LLMs) have demonstrated emerging capabilities in social intelligence tasks, including implicature resolution and theory-of-mind reasoning, both of which require substantial pragmatic understanding. However, how LLMs acquire this pragmatic competence throughout the training process remains poorly understood. In this work, we introduce ALTPRAG, a dataset grounded in th… ▽ More

    Submitted 11 January, 2026; v1 submitted 24 May, 2025; originally announced May 2025.

  47. arXiv:2504.06564  [pdf, ps, other

    cs.CL

    Thinking Out Loud: Do Reasoning Models Know When They're Right?

    Authors: Qingcheng Zeng, Weihao Xuan, Leyang Cui, Rob Voigt

    Abstract: Large reasoning models (LRMs) have recently demonstrated impressive capabilities in complex reasoning tasks by leveraging increased test-time computation and exhibiting behaviors reminiscent of human-like self-reflection. While LRMs show a clear capacity for valuable self-reflection, how this ability interacts with other model behaviors remains underexplored. We investigate this connection by anal… ▽ More

    Submitted 19 October, 2025; v1 submitted 8 April, 2025; originally announced April 2025.

    Comments: EMNLP 2025

  48. arXiv:2503.10497  [pdf, other

    cs.CL

    MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation

    Authors: Weihao Xuan, Rui Yang, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing, Junjue Wang, Fan Gao, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen, Douglas Teodoro, Nan Liu , et al. (7 additional authors not shown)

    Abstract: Existing large language model (LLM) evaluation benchmarks primarily focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-linguistic reasoning abilities. This dual limitation makes it challenging to comprehensively assess LLMs' performance in the multilingual setting. To fill this gap, we introduce MMLU-ProX, a comprehensive benchmark covering 29… ▽ More

    Submitted 26 May, 2025; v1 submitted 13 March, 2025; originally announced March 2025.

  49. Predicting and Understanding College Student Mental Health with Interpretable Machine Learning

    Authors: Meghna Roy Chowdhury, Wei Xuan, Shreyas Sen, Yixue Zhao, Yi Ding

    Abstract: Mental health issues among college students have reached critical levels, significantly impacting academic performance and overall wellbeing. Predicting and understanding mental health status among college students is challenging due to three main factors: the necessity for large-scale longitudinal datasets, the prevalence of black-box machine learning models lacking transparency, and the tendency… ▽ More

    Submitted 10 June, 2025; v1 submitted 10 March, 2025; originally announced March 2025.

    Comments: 12 pages, 10 figures, ACM/IEEE International Conference on Connected Health: Applications, Systems and Engineering Technologies (CHASE '25), June 24--26, 2025, New York, NY, USA

  50. arXiv:2503.07637  [pdf, other

    cs.LG

    Is Pre-training Applicable to the Decoder for Dense Prediction?

    Authors: Chao Ning, Wanshui Gan, Weihao Xuan, Naoto Yokoya

    Abstract: Pre-trained encoders are widely employed in dense prediction tasks for their capability to effectively extract visual features from images. The decoder subsequently processes these features to generate pixel-level predictions. However, due to structural differences and variations in input data, only encoders benefit from pre-learned representations from vision benchmarks such as image classificati… ▽ More

    Submitted 15 March, 2025; v1 submitted 5 March, 2025; originally announced March 2025.