Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 195 results for author: Yao, B

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.10621  [pdf, ps, other

    cs.LG

    ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions

    Authors: Xinzhe Huang, Biwu Yao, Kedong Xiu, Mengnan Zhao, Di Wang, Puning Zhao, Tianhang Zheng

    Abstract: Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. Existing guardrails typically formulate safety assessment as a deterministic classification task, mapping a discrete token sequence to a discrete safety label. However, this paradigm has two limitations: First, safety assessment is inherently an uncertain problem, particularly during… ▽ More

    Submitted 11 August, 2026; originally announced August 2026.

  2. arXiv:2608.01856  [pdf, ps, other

    cs.AI

    EchoChange: A Diffusion Language Model with Dual Pass Remasking for Factual Remote Sensing Disaster Change Captioning

    Authors: Dongwei Sun, Bowen Yao, Yujie Zhang, Pei Liu, Jing Yao, Xiangyong Cao

    Abstract: Bi-temporal remote-sensing disaster change captioning often needs to identify sparse and spatially localized changes across large pre- and post-event scenes and then translate them into coherent, factual descriptions. However, existing change captioning methods always follow an autoregressive decoding paradigm to generate the change description and thus an early misinterpretation of the changed ob… ▽ More

    Submitted 14 August, 2026; v1 submitted 3 August, 2026; originally announced August 2026.

  3. arXiv:2608.00793  [pdf, ps, other

    cs.RO

    DynamicWAM: Dual-Path Motion Conditioning for World-Action Models in Dynamic Manipulation

    Authors: Yunfan Lou, Hewen Gao, Xiyu Zhu, Zhuoran Qiao, Xuan Han, Yifan Yang, Yifan Ye, Boxian Yao, Zhibo Pang

    Abstract: Dynamic manipulation requires robots to infer target motion and respond promptly, yet existing World-Action Models (WAMs) typically condition only on the current frame and execute large backbones synchronously, limiting motion awareness and responsive control in dynamic scenes. We propose DynamicWAM, a compact WAM for dynamic object manipulation with dual-path motion conditioning. DynamicWAM intro… ▽ More

    Submitted 6 August, 2026; v1 submitted 1 August, 2026; originally announced August 2026.

    Comments: 18 pages, 9 figures. Project page: https://dynamicwam.github.io/

  4. arXiv:2607.17977  [pdf, ps, other

    cs.RO

    RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

    Authors: Kehan Li, Bohan Hou, Minghao Zhu, Tianyi Zhang, Zesen Cheng, Zhikai Wang, Sicong Leng, Xin Li, Xiao Lin, Biying Yao, Minghua Zeng, Jiangpin Liu, Ronghao Dang, Jiayan Guo, Siteng Huang, Haoyu Zhao, Heng Ping, Yaxi Zhao, Tong Zhao, Kexiang Wang, Tong Lu, Shengke Xue, Jiahao Tang, Yulei Wang, Zejing Wang , et al. (6 additional authors not shown)

    Abstract: We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, RynnBrain 1.1 supports embodied perception, spatial reasoning, localization, and planning. Compared with RynnBrain 1.0, it further introduces contact-point prediction across the model family and native 3D grounding for the… ▽ More

    Submitted 31 July, 2026; v1 submitted 20 July, 2026; originally announced July 2026.

    Comments: KL,BH,MZ,TZ,ZC,ZW,SL,XL,XL,BY,MZ,JL,RD contribute equally. Project Lead: Kehan Li and Xin Li project: https://alibaba-damo-academy.github.io/RynnBrain github: https://github.com/alibaba-damo-academy/RynnBrain huggingface: https://huggingface.co/collections/Alibaba-DAMO-Academy/rynnbrain-11 modelscope: https://modelscope.cn/collections/DAMO_Academy/RynnBrain-11

  5. arXiv:2606.29648  [pdf, ps, other

    cs.CL cs.AI cs.LG cs.MA

    Hybrid Retriever Evolution for Multimodal Document Reasoning Agents

    Authors: Bohan Yao, Shruthan Radhakrishna, Vikas Yadav

    Abstract: Different retrievers, including lexical, semantic, and multimodal approaches, provide highly complementary strengths for multimodal document understanding, yet most systems combine them through fixed pipelines that cannot adapt to the demands of individual reasoning steps. In this work, we ask whether retrieval orchestration itself can be learned as part of the reasoning process. We introduce a fa… ▽ More

    Submitted 28 June, 2026; originally announced June 2026.

    Comments: 17 pages, 3 figures

  6. arXiv:2606.19303  [pdf, ps, other

    cs.LG

    P-K-GCN: Physics-augmented Koopman-enhanced Graph Convolutional Network for Deep Spatiotemporal Super-resolution

    Authors: Xizhuo, Zhang, Zekai Wang, Fei Liu, Bing Yao

    Abstract: High-fidelity simulation of spatiotemporal dynamics is computationally prohibitive, necessitating efficient super-resolution techniques to reconstruct high-resolution data from coarse-grained inputs. Traditional data-driven methods often lack physical constraints, and simple physics-informed learning struggles with irregular spatial geometries and intricately evolving temporal dynamics. To tackle… ▽ More

    Submitted 17 June, 2026; originally announced June 2026.

  7. arXiv:2606.06399  [pdf, ps, other

    cs.CL

    CollabSim: A CSCW-Grounded Methodology for Investigating Collaborative Competence of LLM Agents through Controlled Multi-Agent Experiments

    Authors: Jiaju Chen, Bo Sun, Yuxuan Lu, Yun Wang, Dakuo Wang, Bingsheng Yao

    Abstract: Multi-agent systems (MAS) built on large language models have shown growing promise, with their effectiveness resting on agents' ability to coordinate through text-based channels much as human teams do. Yet recent study suggests that MAS often falter not because agents lack individual task-solving ability, but because they lack collaborative competence: the capacity to establish common ground, mai… ▽ More

    Submitted 5 June, 2026; v1 submitted 4 June, 2026; originally announced June 2026.

    MSC Class: 68T50

  8. arXiv:2606.06388  [pdf, ps, other

    cs.AI cs.CL

    Humans' ALMANAC: A Human Collaboration Dataset of Action-Level Mental Model Annotations for Agent Collaboration

    Authors: Jiaju Chen, Yuxuan Lu, Jiayi Su, Chaoran Chen, Songlin Xiao, Zheng Zhang, Yun Wang, Yunyao Li, Jian Zhao, Tongshuang Wu, Toby Jia-Jun Li, Dakuo Wang, Bingsheng Yao

    Abstract: Recent advances in LLM agents have enabled complex cognitive capabilities, such as multi-step reasoning, planning, and tool use, that increasingly position these agents as human collaborators. Effective collaboration, however, requires collaborators to continuously maintain and align mental models of their own reasoning,partners' intentions, and shared goals during the collaborative process. Today… ▽ More

    Submitted 5 June, 2026; v1 submitted 4 June, 2026; originally announced June 2026.

    MSC Class: 68T50

  9. arXiv:2606.02892  [pdf, ps, other

    cs.LG

    Multi-Modal Machine Learning for Breast Cancer Recurrence Prediction

    Authors: Jiahao Shao, Xudong Wang, Anam Nawaz Khan, Christopher Brett, Xueping Li, Bing Yao

    Abstract: Breast cancer recurrence, a leading cause of long-term mortality among survivors, requires timely and accurate risk assessment to guide follow-up care and treatment planning. Traditional predictive models, often limited to either structured or unstructured data alone, struggle to capture the full clinical context. This study examines the impact of integrating multi-modal clinical data, including t… ▽ More

    Submitted 1 June, 2026; originally announced June 2026.

    Comments: 33 pages, 10 figures

  10. arXiv:2605.03882  [pdf, ps, other

    cs.HC cs.AI cs.CY

    Deco: Extending Personal Physical Objects into Pervasive AI Companion through a Dual-Embodiment Framework

    Authors: Zhihan Jiang, Mengyuan Millie Wu, Ruishi Zou, Shiyu Xu, Xun Qian, Emma Macmanus, Steven Liao, Ping Zhang, Bingsheng Yao, Tingyu Cheng, James L. David, Nabila El-Bassel, Lena Mamykina, Frances R. Levin, Ryan Sultan, Dakuo Wang, Xuhai Xu

    Abstract: Individuals frequently form deep attachments to physical objects (e.g., plush toys) that usually cannot sense or respond to their emotions. While AI companions offer responsiveness and personalization, they exist independently of these physical objects and lack an ongoing connection to them. To bridge this gap, we conducted a formative study (N=9) to explore how digital agents could inherit and ex… ▽ More

    Submitted 5 May, 2026; originally announced May 2026.

    Comments: 27 pages, 7 figures

  11. arXiv:2604.21312  [pdf, ps, other

    cs.CV cs.AI

    The First Challenge on Remote Sensing Infrared Image Super-Resolution at NTIRE 2026: Benchmark Results and Method Overview

    Authors: Kai Liu, Haoyang Yue, Zeli Lin, Zheng Chen, Jingkai Wang, Jue Gong, Jiatong Li, Xianglong Yan, Libo Zhu, Jianze Li, Ziqing Zhang, Zihan Zhou, Xiaoyang Liu, Radu Timofte, Yulun Zhang, Junye Chen, Zhenming Yan, Yucong Hong, Ruize Han, Song Wang, Li Pang, Heng Zhao, Xinqiao Wu, Deyu Meng, Xiangyong Cao , et al. (43 additional authors not shown)

    Abstract: This paper presents the NTIRE 2026 Remote Sensing Infrared Image Super-Resolution (x4) Challenge, one of the associated challenges of NTIRE 2026. The challenge aims to recover high-resolution (HR) infrared images from low-resolution (LR) inputs generated through bicubic downsampling with a x4 scaling factor. The objective is to develop effective models or solutions that achieve state-of-the-art pe… ▽ More

    Submitted 23 April, 2026; originally announced April 2026.

    Comments: Github Repo: https://github.com/Kai-Liu001/NTIRE2026_infraredSR

  12. arXiv:2604.17789  [pdf, ps, other

    cs.CV cs.AI cs.CL

    DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization

    Authors: Haokun Lin, Xinle Jia, Haobo Xu, Bingchen Yao, Xianglong Guo, Yichen Wu, Zhichao Lu, Ying Wei, Qingfu Zhang, Zhenan Sun

    Abstract: The MXFP4 microscaling format, which partitions tensors into blocks of 32 elements sharing an E8M0 scaling factor, has emerged as a promising substrate for efficient LLM inference, backed by native hardware support on NVIDIA Blackwell Tensor Cores. However, activation outliers pose a unique challenge under this format: a single outlier inflates the shared block scale, compressing the effective dyn… ▽ More

    Submitted 21 April, 2026; v1 submitted 20 April, 2026; originally announced April 2026.

    Comments: Technical Report

  13. arXiv:2604.14958  [pdf, ps, other

    cs.CV

    Frequency-Enhanced Dual-Subspace Networks for Few-Shot Fine-Grained Image Classification

    Authors: Meijia Wang, Guochao Wang, Haozhen Chu, Bin Yao, Weichuan Zhang, Yuan Wang, Junpo Yang

    Abstract: Few-shot fine-grained image classification aims to recognize subcategories with high visual similarity using only a limited number of annotated samples. Existing metric learning-based methods typically rely solely on spatial domain features. Confined to this single perspective, models inevitably suffer from inherent texture biases, entangling essential structural details with high-frequency backgr… ▽ More

    Submitted 16 April, 2026; originally announced April 2026.

  14. arXiv:2604.14558  [pdf, ps, other

    cs.CV

    The Fourth Challenge on Image Super-Resolution ($\times$4) at NTIRE 2026: Benchmark Results and Method Overview

    Authors: Zheng Chen, Kai Liu, Jingkai Wang, Xianglong Yan, Jianze Li, Ziqing Zhang, Jue Gong, Jiatong Li, Lei Sun, Xiaoyang Liu, Radu Timofte, Yulun Zhang, Jihye Park, Yoonjin Im, Hyungju Chun, Hyunhee Park, MinKyu Park, Zheng Xie, Xiangyu Kong, Weijun Yuan, Zhan Li, Qiurong Song, Luen Zhu, Fengkai Zhang, Xinzhe Zhu , et al. (128 additional authors not shown)

    Abstract: This paper presents the NTIRE 2026 image super-resolution ($\times$4) challenge, one of the associated competitions of the NTIRE 2026 Workshop at CVPR 2026. The challenge aims to reconstruct high-resolution (HR) images from low-resolution (LR) inputs generated through bicubic downsampling with a $\times$4 scaling factor. The objective is to develop effective super-resolution solutions and analyze… ▽ More

    Submitted 15 April, 2026; originally announced April 2026.

    Comments: NTIRE 2026 webpage: https://cvlai.net/ntire/2026. Code: https://github.com/zhengchen1999/NTIRE2026_ImageSR_x4

  15. arXiv:2604.13328  [pdf, ps, other

    cs.LG

    Multi-Task LLM with LoRA Fine-Tuning for Automated Cancer Staging and Biomarker Extraction

    Authors: Jiahao Shao, Anam Nawaz Khan, Christopher Brett, Tom Berg, Xueping Li, Bing Yao

    Abstract: Pathology reports serve as the definitive record for breast cancer staging, yet their unstructured format impedes large-scale data curation. While Large Language Models (LLMs) offer semantic reasoning, their deployment is often limited by high computational costs and hallucination risks. This study introduces a parameter-efficient, multi-task framework for automating the extraction of Tumor-Node-M… ▽ More

    Submitted 14 April, 2026; originally announced April 2026.

    Comments: 11 pages, 3 figures and 4 tables in the main manuscript. Additional content, figures and tables are in supplementary material section. 17 pages in total

  16. arXiv:2604.10112  [pdf, ps, other

    cs.CV

    Dual-Branch Remote Sensing Infrared Image Super-Resolution

    Authors: Xining Ge, Gengjia Chang, Weijun Yuan, Zhan Li, Zhanglu Chen, Boyang Yao, Yihang Chen, Yifan Deng, Shuhong Liu

    Abstract: Remote sensing infrared image super-resolution aims to recover sharper thermal observations from low-resolution inputs while preserving target contours, scene layout, and radiometric stability. Unlike visible-image super-resolution, thermal imagery is weakly textured and more sensitive to unstable local sharpening, which makes complementary local and global modeling especially important. This pape… ▽ More

    Submitted 27 April, 2026; v1 submitted 11 April, 2026; originally announced April 2026.

  17. arXiv:2604.09643  [pdf, ps, other

    cs.CV

    PA-SFM: Tracker-free differentiable acoustic radiation for freehand 3D photoacoustic imaging

    Authors: Shuang Li, Jian Gao, Chulhong Kim, Seongwook Choi, Qian Chen, Yibing Wang, Shuang Wu, Yu Zhang, Tingting Huang, Yucheng Zhou, Boxin Yao, Yao Yao, Changhui Li

    Abstract: Three-dimensional (3D) handheld photoacoustic tomography typically relies on bulky and expensive external positioning sensors to correct motion artifacts, which severely limits its clinical flexibility and accessibility. To address this challenge, we present PA-SFM, a tracker-free framework that leverages exclusively single-modality photoacoustic data for both sensor pose recovery and high-fidelit… ▽ More

    Submitted 23 March, 2026; originally announced April 2026.

  18. arXiv:2604.07765  [pdf, ps, other

    cs.CV

    RemoteAgent: Bridging Vague Human Intents and Earth Observation with RL-based Agentic MLLMs

    Authors: Liang Yao, Shengxiang Xu, Fan Liu, Chuanyi Zhang, Bishun Yao, Rui Min, Yongjun Li, Chaoqian Ouyang, Shimin Di, Min-Ling Zhang

    Abstract: Earth Observation (EO) systems are essentially designed to support domain experts who often express their requirements through vague natural language rather than precise, machine-friendly instructions. Depending on the specific application scenario, these vague queries can demand vastly different levels of visual precision. Consequently, a practical EO AI system must bridge the gap between ambiguo… ▽ More

    Submitted 12 April, 2026; v1 submitted 8 April, 2026; originally announced April 2026.

  19. arXiv:2603.14602  [pdf, ps, other

    cs.CL cs.AI cs.LG

    PA3: Policy-Aware Agent Alignment through Chain-of-Thought

    Authors: Shubhashis Roy Dipta, Daniel Bis, Kun Zhou, Lichao Wang, Benjamin Z. Yao, Chenlei Guo, Ruhi Sarikaya

    Abstract: Conversational assistants powered by large language models (LLMs) excel at tool-use tasks but struggle with adhering to complex, business-specific rules. While models can reason over business rules provided in context, including all policies for every query introduces high latency and wastes compute. Furthermore, these lengthy prompts lead to long contexts, harming overall performance due to the "… ▽ More

    Submitted 21 March, 2026; v1 submitted 15 March, 2026; originally announced March 2026.

  20. arXiv:2602.23167  [pdf, ps, other

    cs.CR cs.LG

    SettleFL: Trustless and Scalable Reward Settlement Protocol for Federated Learning on Permissionless Blockchains (Extended version)

    Authors: Shuang Liang, Yang Hua, Linshan Jiang, Peishen Yan, Tao Song, Bin Yao, Haibing Guan

    Abstract: In open Federated Learning (FL) environments where no central authority exists, ensuring collaboration fairness relies on decentralized reward settlement, yet the prohibitive cost of permissionless blockchains directly clashes with the high-frequency, iterative nature of model training. Existing solutions either compromise decentralization or suffer from scalability bottlenecks due to linear on-ch… ▽ More

    Submitted 26 February, 2026; originally announced February 2026.

  21. arXiv:2602.09295  [pdf, ps, other

    cs.LG cs.SD

    Positive-Unlabelled Active Learning to Curate a Dataset for Orca Resident Interpretation

    Authors: Bret Nestor, Bohan Yao, Jasmine Moore, Jasper Kanes

    Abstract: This work presents the largest curation of Southern Resident Killer Whale (SRKW) acoustic data to date, also containing other marine mammals in their environment. We systematically search all available public archival hydrophone data within the SRKW habitat (over 30 years of audio data). The search consists of a weakly-supervised, positive-unlabelled, active learning strategy to identify all insta… ▽ More

    Submitted 13 April, 2026; v1 submitted 9 February, 2026; originally announced February 2026.

  22. arXiv:2602.05987  [pdf, ps, other

    cs.HC

    From Human-Human Collaboration to Human-Agent Collaboration: A Vision, Design Philosophy, and an Empirical Framework for Achieving Successful Partnerships Between Humans and LLM Agents

    Authors: Bingsheng Yao, Chaoran Chen, April Yi Wang, Sherry Tongshuang Wu, Toby Jia-jun Li, Dakuo Wang

    Abstract: The emergence of Large Language Model (LLM) agents enables us to build agent-based intelligent systems that move beyond the role of a "tool" to become genuine collaborators with humans, thereby realizing a novel human-agent collaboration paradigm. Our vision is that LLM agents should resemble remote human collaborators, which allows HCI researchers to ground the future exploration in decades of re… ▽ More

    Submitted 5 February, 2026; originally announced February 2026.

  23. arXiv:2601.19266  [pdf, ps, other

    cs.CV

    A Multi-View Consistency Framework with Semi-Supervised Domain Adaptation

    Authors: Yuting Hong, Li Dong, Xiaojie Qiu, Hui Xiao, Baochen Yao, Siming Zheng, Chengbin Peng

    Abstract: Semi-Supervised Domain Adaptation (SSDA) leverages knowledge from a fully labeled source domain to classify data in a partially labeled target domain. Due to the limited number of labeled samples in the target domain, there can be intrinsic similarity of classes in the feature space, which may result in biased predictions, even when the model is trained on a balanced dataset. To overcome this limi… ▽ More

    Submitted 27 January, 2026; originally announced January 2026.

    Comments: 11 pages, 7 figures

  24. arXiv:2601.09050  [pdf, ps, other

    cs.CL

    SITA: Learning Speaker-Invariant and Tone-Aware Speech Representations for Low-Resource Tonal Languages

    Authors: Tianyi Xu, Xuan Ouyang, Binwei Yao, Shoua Xiong, Sara Misurelli, Maichou Lor, Junjie Hu

    Abstract: Tonal low-resource languages are widely spoken yet remain underserved by modern speech technology. A key challenge is learning representations that are robust to nuisance variation such as gender while remaining tone-aware for different lexical meanings. To address this, we propose SITA, a lightweight adaptation recipe that enforces Speaker-Invariance and Tone-Awareness for pretrained wav2vec-styl… ▽ More

    Submitted 13 January, 2026; originally announced January 2026.

    Comments: 8 pages (excluding references, limitations, ethics, acknowledgement, and appendix); 4 figures in the main paper; appendix included

    ACM Class: I.2.7

  25. arXiv:2601.03570  [pdf, ps, other

    cs.CL

    How Do Large Language Models Learn Concepts During Continual Pre-Training?

    Authors: Barry Menglong Yao, Sha Li, Yunzhi Yao, Minqian Liu, Zaishuo Xia, Qifan Wang, Lifu Huang

    Abstract: Human beings primarily understand the world through concepts (e.g., dog), abstract mental representations that structure perception, reasoning, and learning. However, how large language models (LLMs) acquire, retain, and forget such concepts during continual pretraining remains poorly understood. In this work, we study how individual concepts are acquired and forgotten, as well as how multiple con… ▽ More

    Submitted 17 August, 2026; v1 submitted 6 January, 2026; originally announced January 2026.

    Comments: 19 pages, 27 figures

  26. arXiv:2601.00268  [pdf, ps, other

    cs.CL cs.AI

    Beyond Perfect APIs: A Comprehensive Evaluation of LLM Agents Under Real-World API Complexity

    Authors: Doyoung Kim, Zhiwei Ren, Jie Hao, Zhongkai Sun, Lichao Wang, Xiyao Ma, Zack Ye, Xu Han, Jun Yin, Heng Ji, Wei Shen, Xing Fan, Benjamin Yao, Chenlei Guo

    Abstract: We introduce WildAGTEval, a benchmark designed to evaluate large language model (LLM) agents' function-calling capabilities under realistic API complexity. Unlike prior work that assumes an idealized API system and disregards real-world factors such as noisy API outputs, WildAGTEval accounts for two dimensions of real-world complexity: 1. API specification, which includes detailed documentation an… ▽ More

    Submitted 1 January, 2026; originally announced January 2026.

    Comments: 26 pages

  27. arXiv:2512.24824  [pdf, ps, other

    cs.DB

    LMG Index: A Robust and Efficient Learned Index Framework for Multi-Dimensional Performance Balance

    Authors: Yuzhen Chen, Bin Yao

    Abstract: Index structures are fundamental for efficient query processing on large-scale datasets. Learned indexes model the indexing process as a prediction problem to overcome the inherent trade-offs of traditional indexes. However, most existing learned indexes optimize only for limited objectives like query latency or space usage, neglecting other practical evaluation dimensions such as update efficienc… ▽ More

    Submitted 29 March, 2026; v1 submitted 31 December, 2025; originally announced December 2025.

  28. arXiv:2512.14086  [pdf, ps, other

    cs.LG math.NA

    Derivative-Informed Fourier Neural Operator: Universal Approximation and Applications to PDE-Constrained Optimization

    Authors: Boyuan Yao, Dingcheng Luo, Lianghao Cao, Nikola Kovachki, Thomas O'Leary-Roseberry, Omar Ghattas

    Abstract: We present approximation theories and efficient training methods for derivative-informed Fourier neural operators (DIFNOs) with applications to PDE-constrained optimization. A DIFNO is an FNO trained by minimizing its prediction error jointly on output and Fréchet derivative samples of a high-fidelity operator (e.g., a parametric PDE solution operator). As a result, a DIFNO can closely emulate not… ▽ More

    Submitted 15 March, 2026; v1 submitted 15 December, 2025; originally announced December 2025.

    MSC Class: 41A65; 68T07 (Primary) 65K10; 65J22 (Secondary)

  29. arXiv:2512.04448   

    cs.DL cs.CY

    Has ACL Lost Its Crown? A Decade-Long Quantitative Analysis of Scale and Impact Across Leading AI Conferences

    Authors: Jianglin Ma, Ben Yao, Xiang Li, Yazhou Zhang

    Abstract: The recent surge of language models (LMs) has rapidly expanded NLP/AI research, driving an exponential rise in submissions and acceptances at major conferences. Yet this growth has been shadowed by escalating concerns over conference quality, such as plagiarism, reviewer inexperience, and collusive bidding. However, existing studies rely largely on qualitative accounts, for example expert intervie… ▽ More

    Submitted 10 August, 2026; v1 submitted 3 December, 2025; originally announced December 2025.

    Comments: The authors have identified substantive issues in the current version that require substantial revisions to the analysis and interpretation. We are therefore withdrawing this version to avoid potential misunderstanding or inappropriate citation of its current findings

  30. arXiv:2510.25110  [pdf, ps, other

    cs.CL

    DEBATE: A Large-Scale Benchmark for Evaluating Opinion Dynamics in Role-Playing LLM Agents

    Authors: Yun-Shiuan Chuang, Ruixuan Tu, Chengtao Dai, You Li, Smit Vasani, Binwei Yao, Michael Henry Tessler, Sijia Yang, Dhavan Shah, Robert Hawkins, Junjie Hu, Timothy T. Rogers

    Abstract: Accurately modeling opinion change through social interactions is crucial for understanding and mitigating polarization, misinformation, and societal conflict. Recent work simulates opinion dynamics with role-playing LLM agents (RPLAs), but multi-agent simulations often display unnatural group behavior, such as premature convergence, and lack empirical benchmarks for assessing alignment with real… ▽ More

    Submitted 28 May, 2026; v1 submitted 28 October, 2025; originally announced October 2025.

  31. arXiv:2510.14205  [pdf, ps, other

    cs.CL cs.AI

    DPRF: A Generalizable Dynamic Persona Refinement Framework for Optimizing Behavior Alignment Between Personalized LLM Role-Playing Agents and Humans

    Authors: Bingsheng Yao, Bo Sun, Yuanzhe Dong, Yuxuan Lu, Dakuo Wang

    Abstract: The emerging large language model role-playing agents (LLM RPAs) aim to simulate individual human behaviors, but the persona fidelity is often undermined by manually-created profiles (e.g., cherry-picked information and personality characteristics) without validating the alignment with the target individuals. To address this limitation, our work introduces the Dynamic Persona Refinement Framework… ▽ More

    Submitted 28 October, 2025; v1 submitted 15 October, 2025; originally announced October 2025.

    Comments: In Submission

  32. arXiv:2510.13601  [pdf, ps, other

    cs.LG

    Physics-augmented Multi-task Gaussian Process for Modeling Spatiotemporal Dynamics

    Authors: Xizhuo Zhang, Bing Yao

    Abstract: Recent advances in sensing and imaging technologies have enabled the collection of high-dimensional spatiotemporal data across complex geometric domains. However, effective modeling of such data remains challenging due to irregular spatial structures, rapid temporal dynamics, and the need to jointly predict multiple interrelated physical variables. This paper presents a physics-augmented multi-tas… ▽ More

    Submitted 15 October, 2025; originally announced October 2025.

    Comments: 13 pages, 5 figures

  33. arXiv:2510.05746  [pdf, ps, other

    cs.AI cs.CL cs.LG

    ARM: Discovering Agentic Reasoning Modules for Generalizable Multi-Agent Systems

    Authors: Bohan Yao, Shiva Krishna Reddy Malay, Vikas Yadav

    Abstract: Large Language Model (LLM)-powered Multi-agent systems (MAS) have achieved state-of-the-art results on various complex reasoning tasks. Recent works have proposed techniques to automate the design of MASes, eliminating the need for manual engineering. However, these techniques perform poorly, often achieving similar or inferior performance to simple baselines. Furthermore, they require computation… ▽ More

    Submitted 19 May, 2026; v1 submitted 7 October, 2025; originally announced October 2025.

    Comments: 29 pages, 2 figures

  34. arXiv:2509.23509  [pdf, ps, other

    cs.HC

    Exploring Collaboration Breakdowns Between Provider Teams and Patients in Post-Surgery Care

    Authors: Bingsheng Yao, Menglin Zhao, Zhan Zhang, Pengqi Wang, Emma G Chester, Changchang Yin, Tianshi Li, Varun Mishra, Lace Padilla, Odysseas Chatzipanagiotou, Timothy Pawlik, Ping Zhang, Weidan Cao, Dakuo Wang

    Abstract: Post-surgery care involves ongoing collaboration between provider teams and patients, which starts from post-surgery hospitalization through home recovery after discharge. While prior HCI research has primarily examined patients' challenges at home, less is known about how provider teams coordinate discharge preparation and care handoffs, and how breakdowns in communication and care pathways may a… ▽ More

    Submitted 16 February, 2026; v1 submitted 27 September, 2025; originally announced September 2025.

    Comments: Accepted at CHI 2026

  35. arXiv:2509.23055  [pdf, ps, other

    cs.CL

    Peacemaker or Troublemaker: How Sycophancy Shapes Multi-Agent Debate

    Authors: Binwei Yao, Chao Shang, Wanyu Du, Jianfeng He, Ruixue Lian, Yi Zhang, Hang Su, Sandesh Swamy, Yanjun Qi

    Abstract: Large language models (LLMs) often display sycophancy, a tendency toward excessive agreeability. This behavior poses significant challenges for multi-agent debating systems (MADS) that rely on productive disagreement to refine arguments and foster innovative thinking. LLMs' inherent sycophancy can collapse debates into premature consensus, potentially undermining the benefits of multi-agent debate… ▽ More

    Submitted 26 September, 2025; originally announced September 2025.

  36. arXiv:2509.21501  [pdf, ps, other

    cs.HC cs.CL

    LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?

    Authors: Lu Sun, Shihan Fu, Bingsheng Yao, Yuxuan Lu, Wenbo Li, Hansu Gu, Jiri Gesi, Jing Huang, Chen Luo, Dakuo Wang

    Abstract: Agentic AI is emerging, capable of executing tasks through natural language, such as Copilot for coding or Amazon Rufus for shopping. Evaluating these systems is challenging, as their rapid evolution outpaces traditional human evaluation. Researchers have proposed LLM Agents to simulate participants as digital twins, but it remains unclear to what extent a digital twin can represent a specific cus… ▽ More

    Submitted 25 September, 2025; originally announced September 2025.

  37. arXiv:2509.18008  [pdf, ps, other

    cs.HC cs.AI cs.CL

    Through the Lens of Human-Human Collaboration: A Configurable Research Platform for Exploring Human-Agent Collaboration

    Authors: Bingsheng Yao, Jiaju Chen, Chaoran Chen, April Wang, Toby Jia-jun Li, Dakuo Wang

    Abstract: Intelligent systems have traditionally been designed as tools rather than collaborators, often lacking critical characteristics that collaboration partnerships require. Recent advances in large language model (LLM) agents open new opportunities for human-LLM-agent collaboration by enabling natural communication and various social and cognitive behaviors. Yet it remains unclear whether principles o… ▽ More

    Submitted 15 February, 2026; v1 submitted 22 September, 2025; originally announced September 2025.

    Comments: Accepted at CHI 2026

  38. arXiv:2509.10723  [pdf, ps, other

    cs.HC cs.AI

    Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight

    Authors: Jingyu Tang, Chaoran Chen, Jiawen Li, Zhiping Zhang, Bingcan Guo, Ibrahim Khalilov, Simret Araya Gebreegziabher, Bingsheng Yao, Dakuo Wang, Yanfang Ye, Tianshi Li, Ziang Xiao, Yaxing Yao, Toby Jia-Jun Li

    Abstract: The dark patterns, deceptive interface designs manipulating user behaviors, have been extensively studied for their effects on human decision-making and autonomy. Yet, with the rising prominence of LLM-powered GUI agents that automate tasks from high-level intents, understanding how dark patterns affect agents is increasingly important. We present a two-phase empirical study examining how agents,… ▽ More

    Submitted 12 September, 2025; originally announced September 2025.

  39. arXiv:2508.18445  [pdf, ps, other

    cs.CV

    VQualA 2025 Challenge on Face Image Quality Assessment: Methods and Results

    Authors: Sizhuo Ma, Wei-Ting Chen, Qiang Gao, Jian Wang, Chris Wei Zhou, Wei Sun, Weixia Zhang, Linhan Cao, Jun Jia, Xiangyang Zhu, Dandan Zhu, Xiongkuo Min, Guangtao Zhai, Baoying Chen, Xiongwei Xiao, Jishen Zeng, Wei Wu, Tiexuan Lou, Yuchen Tan, Chunyi Song, Zhiwei Xu, MohammadAli Hamidi, Hadi Amirpour, Mingyin Bai, Jiawang Du , et al. (34 additional authors not shown)

    Abstract: Face images play a crucial role in numerous applications; however, real-world conditions frequently introduce degradations such as noise, blur, and compression artifacts, affecting overall image quality and hindering subsequent tasks. To address this challenge, we organized the VQualA 2025 Challenge on Face Image Quality Assessment (FIQA) as part of the ICCV 2025 Workshops. Participants created li… ▽ More

    Submitted 25 August, 2025; originally announced August 2025.

    Comments: ICCV 2025 VQualA workshop FIQA track

  40. arXiv:2508.17886  [pdf, ps, other

    cs.DB

    PGTuner: An Efficient Framework for Automatic and Transferable Configuration Tuning of Proximity Graphs

    Authors: Hao Duan, Yitong Song, Bin Yao, Anqi Liang

    Abstract: Approximate Nearest Neighbor Search (ANNS) plays a crucial role in many key areas. Proximity graphs (PGs) are the leading method for ANNS, offering the best balance between query efficiency and accuracy. However, their performance heavily depends on various construction and query parameters, which are difficult to optimize due to their complex inter-dependencies. Given that users often prioritize… ▽ More

    Submitted 25 August, 2025; originally announced August 2025.

  41. arXiv:2508.17828  [pdf, ps, other

    cs.DB

    TRIM: Accelerating High-Dimensional Vector Similarity Search with Enhanced Triangle-Inequality-Based Pruning

    Authors: Yitong Song, Pengcheng Zhang, Chao Gao, Bin Yao, Kai Wang, Zongyuan Wu, Lin Qu

    Abstract: High-dimensional vector similarity search (HVSS) is critical for many data processing and AI applications. However, traditional HVSS methods often require extensive data access for distance calculations, leading to inefficiencies. Triangle-inequality-based lower bound pruning is a widely used technique to reduce the number of data access in low-dimensional spaces but becomes less effective in high… ▽ More

    Submitted 25 August, 2025; originally announced August 2025.

  42. arXiv:2508.16517  [pdf, ps, other

    cs.SE cs.PL

    ARSP: Automated Repair of Verilog Designs via Semantic Partitioning

    Authors: Bingkun Yao, Ning Wang, Xiangfeng Liu, Yuxin Du, Yuchen Hu, Hong Gao, Zhe Jiang, Nan Guan

    Abstract: Debugging functional Verilog bugs consumes a significant portion of front-end design time. While Large Language Models (LLMs) have demonstrated great potential in mitigating this effort, existing LLM-based automated debugging methods underperform on industrial-scale modules. A major reason for this is bug signal dilution in long contexts, where a few bug-relevant tokens are overwhelmed by hundreds… ▽ More

    Submitted 22 August, 2025; originally announced August 2025.

  43. arXiv:2508.15189  [pdf, ps, other

    cs.AI cs.CV eess.IV

    SurgWound-Bench: A Benchmark for Surgical Wound Diagnosis

    Authors: Jiahao Xu, Changchang Yin, Odysseas Chatzipanagiotou, Diamantis Tsilimigras, Kevin Clear, Bingsheng Yao, Dakuo Wang, Timothy Pawlik, Ping Zhang

    Abstract: Surgical site infection (SSI) is one of the most common and costly healthcare-associated infections and and surgical wound care remains a significant clinical challenge in preventing SSIs and improving patient outcomes. While recent studies have explored the use of deep learning for preliminary surgical wound screening, progress has been hindered by concerns over data privacy and the high costs as… ▽ More

    Submitted 20 August, 2025; originally announced August 2025.

  44. arXiv:2508.07959  [pdf, ps, other

    cs.CL

    Large Language Models for Subjective Language Understanding: A Survey

    Authors: Changhao Song, Yazhou Zhang, Hui Gao, Ben Yao, Peng Zhang

    Abstract: Subjective language understanding refers to a broad set of natural language processing tasks where the goal is to interpret or generate content that conveys personal feelings, opinions, or figurative meanings rather than objective facts. With the advent of large language models (LLMs) such as ChatGPT, LLaMA, and others, there has been a paradigm shift in how we approach these inherently nuanced ta… ▽ More

    Submitted 11 August, 2025; originally announced August 2025.

  45. arXiv:2507.21028  [pdf, ps, other

    cs.CL

    Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation

    Authors: Jiaju Chen, Yuxuan Lu, Xiaojie Wang, Huimin Zeng, Jing Huang, Jiri Gesi, Ying Xu, Bingsheng Yao, Dakuo Wang

    Abstract: Nearly all human work is collaborative; thus, the evaluation of real-world NLP applications often requires multiple dimensions that align with diverse human perspectives. As real human evaluator resources are often scarce and costly, the emerging "LLM-as-a-judge" paradigm sheds light on a promising approach to leverage LLM agents to believably simulate human evaluators. Yet, to date, existing LLM-… ▽ More

    Submitted 28 July, 2025; originally announced July 2025.

    MSC Class: 68T50

  46. arXiv:2507.18973  [pdf, ps, other

    cs.CL cs.AI cs.LG

    A Toolbox, Not a Hammer -- Multi-TAG: Scaling Math Reasoning with Multi-Tool Aggregation

    Authors: Bohan Yao, Vikas Yadav

    Abstract: Augmenting large language models (LLMs) with external tools is a promising avenue for developing high-performance mathematical reasoning systems. Prior tool-augmented approaches typically finetune an LLM to select and invoke a single tool at each reasoning step and show promising results on simpler math reasoning benchmarks such as GSM8K. However, these approaches struggle with more complex math p… ▽ More

    Submitted 21 August, 2025; v1 submitted 25 July, 2025; originally announced July 2025.

    Comments: Published at EMNLP Findings 2025; 21 pages, 3 figures

  47. arXiv:2507.07221  [pdf, ps, other

    cs.RO

    Self-Wearing Adaptive Garments via Soft Robotic Unfurling

    Authors: Nam Gyun Kim, William E. Heap, Yimeng Qin, Elvy B. Yao, Jee-Hwan Ryu, Allison M. Okamura

    Abstract: Robotic dressing assistance has the potential to improve the quality of life for individuals with limited mobility. Existing solutions predominantly rely on rigid robotic manipulators, which have challenges in handling deformable garments and ensuring safe physical interaction with the human body. Prior robotic dressing methods require excessive operation times, complex control strategies, and con… ▽ More

    Submitted 9 July, 2025; originally announced July 2025.

  48. arXiv:2507.03855  [pdf, ps, other

    cs.LG stat.ML

    Transformer with Koopman-Enhanced Graph Convolutional Network for Spatiotemporal Dynamics Forecasting

    Authors: Zekai Wang, Bing Yao

    Abstract: Spatiotemporal dynamics forecasting is inherently challenging, particularly in systems defined over irregular geometric domains, due to the need to jointly capture complex spatial correlations and nonlinear temporal dynamics. To tackle these challenges, we propose TK-GCN, a two-stage framework that integrates geometry-aware spatial encoding with long-range temporal modeling. In the first stage, a… ▽ More

    Submitted 4 July, 2025; originally announced July 2025.

  49. arXiv:2507.00093  [pdf, ps, other

    cs.DM cs.AI cs.DS math.ST

    $σ$-Maximal Ancestral Graphs

    Authors: Binghua Yao, Joris M. Mooij

    Abstract: Maximal Ancestral Graphs (MAGs) provide an abstract representation of Directed Acyclic Graphs (DAGs) with latent (selection) variables. These graphical objects encode information about ancestral relations and d-separations of the DAGs they represent. This abstract representation has been used amongst others to prove the soundness and completeness of the FCI algorithm for causal discovery, and to d… ▽ More

    Submitted 30 June, 2025; originally announced July 2025.

    Comments: It has beee accepted by the 41st Conference on Uncertainty in Artificial Intelligence (UAI)

    Journal ref: Proceedings of the Forty-first Conference on Uncertainty in Artificial Intelligence, PMLR 286:4775-4805, 2025

  50. arXiv:2506.14100  [pdf, ps, other

    cs.RO eess.SY

    A Hierarchical Test Platform for Vision Language Model (VLM)-Integrated Real-World Autonomous Driving

    Authors: Yupeng Zhou, Can Cui, Juntong Peng, Zichong Yang, Juanwu Lu, Jitesh H Panchal, Bin Yao, Ziran Wang

    Abstract: Vision-Language Models (VLMs) have demonstrated notable promise in autonomous driving by offering the potential for multimodal reasoning through pretraining on extensive image-text pairs. However, adapting these models from broad web-scale data to the safety-critical context of driving presents a significant challenge, commonly referred to as domain shift. Existing simulation-based and dataset-dri… ▽ More

    Submitted 16 June, 2025; originally announced June 2025.