Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 3,279 results for author: Xue, X

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.21341  [pdf, ps, other

    cs.SE

    Natural-Language Workflows Are Not Software Yet: Artifact-Driven Compilation for Reliable Agent Execution

    Authors: Xiangzhe Xu, Hanxi Guo, Guangyu Shen, Siyuan Cheng, Xiangyu Zhang

    Abstract: Natural-language workflows offer a software-like interface for agents: domain experts can write reusable procedures, and agents can execute them as instructions. This promise is not yet reliable. Workflow descriptions often leave data dependencies implicit, so the executor must infer which prior results a step should use; agents can also fail to follow long or branching instructions under context… ▽ More

    Submitted 21 August, 2026; originally announced August 2026.

    Comments: The first two authors contributed equally

  2. arXiv:2608.21101  [pdf, ps, other

    cs.CR cs.AI

    ClawSentry: A Progressive Multi-Tier Security Monitor for Safeguarding Autonomous LLM Agents

    Authors: Kai Wang, Zeming Wei, BiaoJie Zeng, Chang Jin, An Wang, Xiaokun Luan, Zhixiao Lin, Jingjing Qu, Xia Hu, Xingcheng Xu

    Abstract: As large language model (LLM) agents move from conversation to executing code, reading local files, and orchestrating external tools, a single agent hijacked by a malicious third-party skill can cause data exfiltration, privilege escalation, or cascading compromise. We argue that agentic risk is progressive: it can enter at four loci of the agent control loop--skill admission, invocation-time inte… ▽ More

    Submitted 21 August, 2026; originally announced August 2026.

    Comments: 35 pages, 14 figures. Code: https://github.com/Elroyper/ClawSentry

  3. arXiv:2608.20334  [pdf, ps, other

    cs.CV

    Exploring the Performance Frontier of Compact Unified Image Generation Models

    Authors: Taihang Hu, Zhao Wang, Zuan Gao, Tao Liu, Hao Yan, Zhengze Xu, Yuhang Yu, Yongchao Du, Xingjian Wang, Jun Zheng, Qinye Zhou, Yaqi Cai, Zhengrui Chen, Chao Lin, Yefeng Shen, Yuan Wang, Zhengtao Wu, Ge Wu, Xiaoli Xu, Denghui Yang, Huayu Zhang, Mingzhou Zhang, Mengting Chen

    Abstract: We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi-image editing. Our goal is to explore how far a relatively small visual generator can be pushed through systematic training engineering under a constrained computational budget. Swift-Image adopts an efficient 6B single-stream DiT and a progressive training pipeline that evolves from broad… ▽ More

    Submitted 21 August, 2026; v1 submitted 20 August, 2026; originally announced August 2026.

    Comments: 28 pages, 11 figures

  4. arXiv:2608.19492  [pdf, ps, other

    cs.LG cs.RO

    Beyond Multimodal Alignment: Certifying Physical Language through Response Substitution and Ordered Execution

    Authors: Kaizhen Tan, Xin Xu, Siru Tao, Yixiao Li, Hanzhe Hong, Yang Feng, Heqing Du

    Abstract: World models increasingly treat compact multimodal representations as interfaces between perception and physical interaction, yet existing probes do not establish whether different sensors carry the same executable meaning or whether that meaning survives a new action composition. We introduce an operational capability hierarchy and the Disjoint-Bridge Operator-Substitution Certificate (DBOSC), wh… ▽ More

    Submitted 19 August, 2026; originally announced August 2026.

  5. arXiv:2608.19181  [pdf, ps, other

    cs.LG cs.AI cs.CL

    Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning

    Authors: Zhu Zhang, Jixun Wang, Xiaoang Xu, Xiaorong Wang, Zihan Zhou, Zhiyuan Wang, Shuo Wang, Chaojun Xiao, Yuezhi Zhou

    Abstract: On-policy distillation (OPD) trains a student on its own responses using dense token-level guidance from a stronger teacher. In long-context tasks, however, token-level teacher support can favor locally plausible responses that omit evidence distributed across the input or violate global task constraints. Task-specific verifiers, in contrast, evaluate task completion at the response level and may… ▽ More

    Submitted 19 August, 2026; originally announced August 2026.

    Comments: 20 pages, 5 figures

  6. arXiv:2608.18076  [pdf, ps, other

    cs.CV cs.AI

    From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation

    Authors: Xingjian Wang, Zhao Wang, Taihang Hu, Jun Zheng, Qing Jin, Qinye Zhou, Zhengtao Wu, Yongchao Du, Zuan Gao, Chao Lin, Yefeng Shen, Xiaoli Xu, Zhengze Xu, Hao Yan, Yuhang Yu, Mingzhou Zhang, Mengting Chen

    Abstract: Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. A central challenge is not only how to curate each task-specific corpus, but also how to organize heterogeneous supervision according to the dependencies among generative capabilities. We present a \textbf… ▽ More

    Submitted 18 August, 2026; originally announced August 2026.

    Comments: 19 pages, 10 figures

  7. arXiv:2608.17719  [pdf, ps, other

    cs.SE cs.AI cs.CL

    What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations

    Authors: Xiaonan Xu, Wenjing Wu

    Abstract: Context: Software systems that depend on commercial large language model APIs must migrate to successor versions when vendors deprecate older models. Migration decisions typically rely on aggregate benchmark scores, which compress heterogeneous item-level behaviour into a single net figure. Objective: We measure what that compression conceals. Method: On three pairwise upgrades in the GPT-5.4 to G… ▽ More

    Submitted 18 August, 2026; originally announced August 2026.

    Comments: 25 pages, 1 figure, 10 tables (including 8 appendix tables)

    ACM Class: D.2.5; I.2.7

  8. arXiv:2608.16551  [pdf, ps, other

    cs.CR

    What to Remember, What to Reveal: Privacy-Aware Memory for Conversational Agents

    Authors: Wenjie Wang, Wenhe Si, Xinyue Xu, Yue Xu

    Abstract: Long-term memory enables personalized conversational agents to retain user information across sessions. However, existing memory architectures primarily optimize for utility while neglecting the risks of unnecessarily storing and reusing private attributes such as personally identifiable information (PII). Addressing privacy risks in personalized memory is challenging because simply removing sensi… ▽ More

    Submitted 17 August, 2026; originally announced August 2026.

  9. arXiv:2608.16465  [pdf, ps, other

    cs.AI

    JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills

    Authors: Xiaoyu Wen, Jiajia Li, Zhida He, Peng Yu, Chenxu Wang, Han Qi, Ziyuan Zhou, Cheng Jin, Ying Wen, Xingcheng Xu, Shuyue Hu, Tianhang Zheng, Chaochao Lu, Qiaosheng Zhang

    Abstract: Automated red-teaming has produced a growing collection of attack strategies, yet they typically remain scattered across prompts and workflows, making them difficult to systematically integrate, reuse, and improve at scale. We introduce \textsc{JailbreakSkill}, a skill-centric framework for scaling automated red-teaming through reusable and continuously evolving attack capabilities. \textsc{Jailbr… ▽ More

    Submitted 17 August, 2026; originally announced August 2026.

  10. arXiv:2608.16419  [pdf, ps, other

    cs.LG cs.AI q-bio.QM

    PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data

    Authors: Zhenchao Tang, Xiaogang Xu, Tianxu Lv, Jiahui Guan, Jiale Zhou, Haohuai He, Zhi Song, Hanbo Huang, Jiehui Huang, Jiafei Wu, Zhe Liu

    Abstract: Large language models can describe mechanisms, yet scalable post-training still depends on costly, manually curated biological reasoning traces. Here we show that cellular perturbation atlases can instead become reinforcement-learning environments, where measured gene responses provide computable rewards for biological reasoning. We introduce PertMind, which combines trusted-trajectory supervised… ▽ More

    Submitted 17 August, 2026; originally announced August 2026.

    Comments: Project page: https://shapsider.github.io/PertMind/ Code: https://github.com/shapsider/PertMind Model: https://huggingface.co/tzcfly/PertMind

    ACM Class: I.2.6; I.2.7

  11. arXiv:2608.16162  [pdf, ps, other

    cs.SD

    ACE-Cap: Active Evidence Acquisition via Agentic Co-Evolution for Long-Paragraph Fine-Grained Audio Captioning

    Authors: Fengji Ma, Yan Rong, Xu Li, Xuenan Xu, Chen Zhang, Li Liu

    Abstract: Long-paragraph fine-grained audio captioning requires models to recover diverse acoustic facts while avoiding omissions and unsupported details. However, prevailing captioners remain passive one-shot generators: once a detail is overlooked, they cannot identify the evidence gap, query the audio for targeted information, or decide when sufficient evidence has been collected. We formulate this task… ▽ More

    Submitted 20 August, 2026; v1 submitted 17 August, 2026; originally announced August 2026.

  12. arXiv:2608.15659  [pdf, ps, other

    cs.CV cs.GR

    WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations

    Authors: Xiaojie Xu, Zhengyuan Lin, Runyi Li, Yihao Liu, Kaipeng Zhang, Yongtao Ge

    Abstract: Learning to generate or reconstruct explorable worlds requires video paired with more than RGB: camera motion, scene geometry, temporal correspondence and, for interactive models, control signals. Real capture can provide some of these signals, but dense geometry and long-range correspondence usually rely on estimation or specialised instrumentation. Rendering provides these quantities directly, y… ▽ More

    Submitted 19 August, 2026; v1 submitted 16 August, 2026; originally announced August 2026.

    Comments: update code and data links

  13. arXiv:2608.15438  [pdf, ps, other

    cs.DB cs.IR cs.LG

    NeuRoute: Logit-Guided Neural Routing for Billion-Scale Vector Search with Sub-Hour Index Construction

    Authors: Xingqiao Wang, Zi Wang, Xiaowei Xu

    Abstract: Building approximate nearest neighbor (ANN) indexes at billion scale is often dominated by expensive global clustering or graph construction, making time-to-index a first-order systems concern. We present NeuRoute, a learned hashing index that turns short binary codes into an effective routing primitive for large-scale vector search. NeuRoute trains a lightweight neural network encoder with a sele… ▽ More

    Submitted 15 August, 2026; originally announced August 2026.

    Comments: 34 pages, 9 figures

  14. arXiv:2608.14957  [pdf, ps, other

    nucl-ex cs.DB nucl-th

    Best Reaction Target To Determine Proton Distribution Radii of Atomic Nuclei

    Authors: Jun-Yao Xu, Bao-Hua Sun, Isao Tanihata, Satoru Terashima, Jian-Wei Zhao, Ji-Chao Zhang, Ge Guo, Shi-Tao Wang, Lei Shen, Jun Su, Xiao-Dong Xu, Andrej Prochazka, Guang-Shuai Li, Xiu-Lin Wei, Chang-Jian Wang, Feng Wang, Meng Wang, Jing Wang, Liu-Chun He, Chuan-Ye Liu, Wen-Jian Lin, Wei-Ping Lin, Zhong Liu, Pei-Pei Ren, Yu Zhang , et al. (7 additional authors not shown)

    Abstract: We found that a heavy target such as Pb is most suitable for determining the proton distribution radii of unstable nuclei through charge-changing cross-section ($σ_\text{cc}$) measurements. As a heavy ion probe, low-$Z$ targets are routinely used to determine nucleon distribution radii of unstable isotopes. This approach has recently been extended to study proton distribution radii from… ▽ More

    Submitted 14 August, 2026; originally announced August 2026.

  15. arXiv:2608.14797  [pdf, ps, other

    cs.CL

    Beyond Tokens: A Survey on Decoding Methods for Large Language and Vision-Language Models

    Authors: Haoran Wang, Xiongxiao Xu, Philip S. Yu, Kai Shu

    Abstract: Large language models (LLMs) and large vision-language models (LVLMs) have demonstrated impressive generative capabilities, yet ensuring their outputs align with user intent is still challenging. While most existing approaches address this issue at the training stage, inference-time approaches like decoding methods offer a more efficient and scalable solution. Decoding methods control model genera… ▽ More

    Submitted 14 August, 2026; originally announced August 2026.

    Comments: ACM SIGKDD Explorations Newsletter, Volume 28, Issue 1

  16. arXiv:2608.14783  [pdf, ps, other

    cs.CV cs.GR

    MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling

    Authors: Manwen Liao, Xinyu Lian, Jian Mao, Kaixu Chen, Li Luo, Jinghao Yan, Wanshui Gan, Qiao Yu, Weitian Zhang, Chunhua Shen, Guang Chen, Bo Dai, Xudong Xu, Zhaoyang Lyu

    Abstract: Part-aware 3D object generation is essential for graphics applications such as controllable modeling, editing, and articulation, where objects are represented as coherent assemblies of semantic parts. However, existing part-aware generation methods, do not scale well to highly complex objects. As the number of parts increases, generating detailed geometry becomes prohibitively expensive in token l… ▽ More

    Submitted 14 August, 2026; originally announced August 2026.

    Comments: 12 pages, 6 pages appendix, 13 figures, technical report

  17. arXiv:2608.14546  [pdf, ps, other

    cs.CV

    CPI-Bench: A Comprehensive, Practical and Intelligent Benchmark for Real-World Image Editing

    Authors: Qinye Zhou, Jun Zheng, Yongchao Du, Yuan Wang, Zhengrui Chen, Zuan Gao, Taihang Hu, Chao Lin, Yefeng Shen, Xingjian Wang, Zhao Wang, Zhengtao Wu, Xiaoli Xu, Zhengze Xu, Hao Yan, Denghui Yang, Yuhang Yu, Huayu Zhang, Mingzhou Zhang, Mengting Chen

    Abstract: With the rapid advancement of image editing models and their widespread application across various domains, there is an increasingly urgent need to deploy these model capabilities directly into real-world scenarios. However, existing benchmarks remain confined to simple single-image tasks, suffering from limited coverage dimensions and an inability to effectively differentiate performance among di… ▽ More

    Submitted 18 August, 2026; v1 submitted 14 August, 2026; originally announced August 2026.

    Comments: 13 pages, benchmark report

  18. arXiv:2608.13505  [pdf, ps, other

    cs.LG cs.CL cs.CV

    Intern-S2-Preview: Scientific Agentic Foundation Model

    Authors: Lei Bai, Jiaqi Cao, Chiyu Chen, Guanzhou Chen, Kai Chen, Guangran Cheng, Erfei Cui, Xuanlang Dai, Shengyuan Ding, Shangheng Du, Yanhui Duan, Yue Fan, Youqing Fang, Quan Gan, Yuanyuan Gao, Jiaye Ge, Lixin Gu, Yuzhe Gu, Qipeng Guo, Junjun He, Xin Hong, Ming Hu, Zhouqi Hua, Haian Huang, Junhao Huang , et al. (100 additional authors not shown)

    Abstract: Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. We present Intern-S2-Preview, a series of scientific agentic foundation models designed to support multimodal scientific understanding, reasoning, generation, and long-horizon tas… ▽ More

    Submitted 13 August, 2026; originally announced August 2026.

    Comments: 35 pages, 12 figures

  19. arXiv:2608.13492  [pdf, ps, other

    cs.AI

    AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1)

    Authors: AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Mingliang Zhai, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao

    Abstract: This report presents an improved version of AlayaWorld. While the backbone architecture, chunk-wise autoregressive generation scheme, and training data remain unchanged from the previous release, we substantially revise how conditioning signals are represented and integrated into the model. The new design is guided by a simple principle: conditioning signals should match the generated content as c… ▽ More

    Submitted 13 August, 2026; originally announced August 2026.

    Comments: Authors are listed alphabetically by the first name and their role. See the contribution section for details

  20. arXiv:2608.13120  [pdf, ps, other

    cs.AI

    SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback

    Authors: Qianxi Yan, Chunrong Chen, Jiuzhou Zhao, Min Zhang, Yongzhou Xu, Xiaochuan Xu

    Abstract: Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through which they might improve from the interaction failures they actually cause. Recent work does close this loop, but derives its feedback from single-turn question-answering evaluation. The consequence is a sharp asymmetry: once the first round has patched the gaps… ▽ More

    Submitted 13 August, 2026; originally announced August 2026.

  21. arXiv:2608.13108  [pdf, ps, other

    cs.AI

    Robust Dempster-Shafer Evidence Fusion with Chaos-Conflict Measurement and Historical-Experience Weighting

    Authors: Huiyu Li, Weibo Liu, Xinru Xu, Dongchen Gao, Meng Zhang, Junhua Hu

    Abstract: Multi-source evidence fusion under Dempster-Shafer theory faces two persistent challenges: existing conflict measures assess inter-evidence inconsistency and intra-evidence uncertainty independently, yielding incomplete evaluations, and current fusion methods evaluate evidence sources exclusively through instantaneous comparisns without exploiting their long-term reliability across diverse decisio… ▽ More

    Submitted 13 August, 2026; originally announced August 2026.

  22. arXiv:2608.13010  [pdf, ps, other

    cs.CL cs.CR cs.IR

    RAGSieve: Self-Referenced Local Contrast for Knowledge-Poison Detection in Retrieval-Augmented Generation

    Authors: Xinlong Xu, Yoshua Y. Li

    Abstract: Retrieval-augmented generation treats an external corpus as inference evidence, allowing injected documents to promote attacker-chosen claims. Existing detectors depend on trusted references, specific attack artifacts, or global thresholds sensitive to corpus topology. We present RAGSieve, a self-referenced detection framework that constructs its reference from the inspected system. RAGSieve-Query… ▽ More

    Submitted 13 August, 2026; originally announced August 2026.

  23. arXiv:2608.13006  [pdf, ps, other

    cs.CL cs.IR

    EviReform: Evidence-Guided Query Reformulation for Multi-Hop Graph Retrieval

    Authors: Xinlong Xu, Yoshua Y. Li

    Abstract: Multi-hop retrieval must recover passages that provide sufficient evidence together. An initial passage often resolves an entity or relation implicit in the question, making the missing evidence easier to describe only after retrieval begins. Graph retrieval improves access to related evidence through stored corpus structure, but its retrieval signal is commonly derived from the original question.… ▽ More

    Submitted 13 August, 2026; originally announced August 2026.

  24. arXiv:2608.12788  [pdf, ps, other

    cs.AI

    ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs

    Authors: Jiale Cui, Yueyao Yuan, Kaixi Zhong, Xiaogang Xu, Jiafei Wu, Zhe Liu

    Abstract: The rapid advancement of Auto-Research has surfaced a fundamental evaluation challenge: how can we measure the alignment, logical coherence, and evolutionary completeness of its research trajectory with human research behavior? We propose Auto-Research's Alignment and Completeness, ARAC-Bench: a Researcher-Mimicking Evaluation framework that shifts the objective from matching final answers to repr… ▽ More

    Submitted 12 August, 2026; originally announced August 2026.

    Comments: 26 pages, 3 figures

  25. arXiv:2608.12778  [pdf, ps, other

    cs.IR

    DrEM: Dual-Side Robust Ensemble Ranking from Noisy User Preference Predictions in Video Recommendation

    Authors: Canwei Huang, Tiantian He, Xiaoxiao Xu, Jun Zhang, Ziran Deng, Weike Pan, Chunjie Chen, Kaiqiao Zhan

    Abstract: Industrial video recommendation systems typically adopt a multi-stage architecture. At the ensemble ranking stage, multi-dimensional user preference predictions (pxtrs) from an upstream multi-task model are fused into a unified ranking score to reflect user satisfaction. Since users' true satisfaction is difficult to observe directly, ensemble ranking models commonly use pxtrs both as input featur… ▽ More

    Submitted 12 August, 2026; originally announced August 2026.

  26. arXiv:2608.12724  [pdf, ps, other

    cs.LG

    MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning

    Authors: Zirui Cheng, Xun Xu, Tiankai Chen, Fady Rezk, Bowen Zheng, Xiaodong Shi, Shijie Li, Kangkang Lu, Bharadwaj Veeravalli, Nancy F. Chen

    Abstract: Few-shot in-context learning (ICL) with multi-modal large language models (MLLMs) enables task adaptation without parameter updates, but its performance is highly sensitive to the quality and coverage of the selected demonstrations. While unlabeled multi-modal data is abundant, it remains elusive how to exploit them for ICL. We propose MAG (MAnifold-Guided semi-supervised in-context demonstra- tio… ▽ More

    Submitted 12 August, 2026; originally announced August 2026.

  27. arXiv:2608.12574  [pdf, ps, other

    cs.AI cs.FL

    Trie Automata for Constrained Decoding over Large Finite Sets

    Authors: Xingzi Xu, Karim Bouyarmane

    Abstract: Large language models increasingly need to generate structured outputs that conform to predefined schemas, with one common constraint being selection from a finite set of valid strings. Current constrained decoding systems handle this through general-purpose grammar compilation, which becomes prohibitively slow as the number of valid values grows into the thousands, a cardinality wall. We introduc… ▽ More

    Submitted 12 August, 2026; originally announced August 2026.

  28. arXiv:2608.12036  [pdf, ps, other

    cs.AI cs.CL cs.HC cs.LG cs.MA

    Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

    Authors: Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhang, Julian McAuley, Tat Seng Chua, Huajun Chen

    Abstract: AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly automated, mechanistic exploration remains largely manual, widening the gap between what models can do and our ability to understand and control them. To bridge this gap, we introd… ▽ More

    Submitted 19 August, 2026; v1 submitted 12 August, 2026; originally announced August 2026.

    Comments: Work in progress

  29. arXiv:2608.12034  [pdf, ps, other

    eess.AS cs.SD

    The SLT 2026 SmartGlasses Challenge: Benchmarking Egocentric Multi-Talker Speech Recognition and Understanding with Audio-Language Models

    Authors: Dehui Gao, Zhixian Zhao, Zhennan Lin, Yujie Liao, Yuhang Dai, Yike Zhu, Longshuai Xiao, Hui Bu, Xin Xu, Xie Chen, Shuai Wang, Liumeng Xue, Zhonghua Fu, Jun Du, Eng-Siong Chng, Jun Zhou, Lei Xie

    Abstract: Recent advances in large language models (LLMs) and multimodal LLMs (MLLMs) have created new opportunities for wearable speech interfaces, with smart glasses providing an egocentric platform for continuous audio sensing and assistance. However, speech recognition and understanding in this setting remain challenging because of dynamic acoustic conditions, speaker overlap, and the spatial ambiguity… ▽ More

    Submitted 12 August, 2026; originally announced August 2026.

    Comments: 7 pages, 7 figures

  30. arXiv:2608.11820  [pdf, ps, other

    cs.CV

    TD-VAD: Breaking Visual Dependence in Video Anomaly Detection with Text-Driven Learning

    Authors: Shuangqing Zhang, Lei-Lei Ma, Zhao Wang, Wen Dong, Xinyi Xu, Guo-Sen Xie, Caifeng Shan, Fang Zhao

    Abstract: Visual data is typically a prerequisite for training existing video anomaly detection (VAD) methods. However, obtaining sufficient annotated anomaly data for training is challenging and not scalable due to the rarity of anomaly data and the wide variety of abnormal events. In this work, we advocate that the effectiveness of treating texts as video sequences for the VAD model and propose a novel Te… ▽ More

    Submitted 12 August, 2026; originally announced August 2026.

    Comments: Accepted to ICML2026

  31. arXiv:2608.11386  [pdf, ps, other

    cs.SE

    The Devil Is in the Interface: Evaluating How Tool Architecture Shapes Coding Agent Behavior

    Authors: Xiangzhe Xu, Hamidreza Saghir, Qianhui Wu, Marc-Alexandre Côté, Tong Wang, Kiran Lakkaraju, Kexin Pei, Xiangyu Zhang

    Abstract: As large language models continue to improve, agentic systems are becoming increasingly important, and tools are a key design dimension because they determine how agents access information and take action in their environments. Prior work on agent tooling has primarily focused on expanding what agents can do, but has paid less systematic attention to how those capabilities are organized and expose… ▽ More

    Submitted 11 August, 2026; originally announced August 2026.

  32. arXiv:2608.10615  [pdf, ps, other

    cs.CL

    Simplex Relaxation for Discrete Diffusion

    Authors: Jinya Sakurai, Patrick Pynadath, Satoshi Hayakawa, Jaehong Yoon, Xulei Yang, Nancy F. Chen, Xun Xu

    Abstract: Discrete diffusion models for categorical generation are defined by a corruption kernel, which determines the intermediate state space and the associated reverse prediction problem. We study uniform discrete diffusion and ask whether its training objective and reverse transitions can be enriched without changing the underlying categorical corruption process. We introduce Simplax, an exact Dirichle… ▽ More

    Submitted 11 August, 2026; originally announced August 2026.

  33. arXiv:2608.10420  [pdf, ps, other

    cs.AI

    Reasoning Shortcuts and Value Symmetries: What Symmetry Permits, Architecture Realizes, and Optimization Selects

    Authors: Xin Xu

    Abstract: Reasoning shortcuts are solutions of a neurosymbolic system's rules that produce correct predictions through unintended concepts. A recent framework of Takemura, Inoue, and Nishino analyzes them through an automorphism group of value relabelings and asks, as its central open question, when rules pin concepts down. We first show that the framework's key definition, one shared permutation applied at… ▽ More

    Submitted 10 August, 2026; originally announced August 2026.

    Comments: 55 pages, 2 figures, 8 tables

  34. arXiv:2608.10173  [pdf, ps, other

    cs.CV

    DoseBridge: Denoising Diffusion Bridge Model for Dose Prediction in Lung Intensity-Modulated Proton Therapy

    Authors: Zerun Zhang, Xiaoda Cong, Xiangkun Xu, Peter Y. Chen, Xuanfeng Ding

    Abstract: Most radiotherapy dose-prediction models use only CT images and anatomical structures, although intensity-modulated proton therapy (IMPT) dose also depends strongly on beam geometry and available clinical datasets are often small. We present DoseBridge, a denoising diffusion bridge model that uses the patient CT as a structured bridge endpoint and encodes plan-specific beam geometry in a spatially… ▽ More

    Submitted 10 August, 2026; originally announced August 2026.

  35. arXiv:2608.09766  [pdf, ps, other

    cs.CL cs.AI

    Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness

    Authors: Pinzhen Chen, Koel Dutta Chowdhury, Xiaoya Xu, David Tan, Doreen Osmelak, Ona de Gibert, Ariun-Erdene Tumurchuluun, Ashok Urlana, Fedor Sizov, Hale Sirin, Jesujoba Alabi, Karrar Talib Abed, Mateusz Klimaszewski, Nikolay Bogoychev, Niyati Bafna, Patricia Schmidtova, Preksha Manjunath Shanbhag, Sherrie Shen, Vilem Zouhar, Vivek Iyer, Yasser Hamidullah, Yusser Al Ghussin, Zheng Zhao

    Abstract: Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale and cultural considerations. We therefore advocate for source-contrastive evaluation and instantiate it with Cultivar, a localised subset of FLORES, which enables locale-s… ▽ More

    Submitted 10 August, 2026; originally announced August 2026.

  36. arXiv:2608.09742  [pdf, ps, other

    cs.LG cs.AI cs.DC

    Rethinking Factor Sharing in Federated LoRA: A Rank-Aware Adaptive Approach

    Authors: Xinyi Xu, Bingnan Xiao, Shuang Qin, Gang Feng, Tony Q. S. Quek

    Abstract: Low-rank adaptation (LoRA) represents large language model (LLM) updates with two compact matrix factors, i.e., $A$ and $B$, providing an efficient way to fine-tune large models in federated learning paradigm. Inspired by the asymmetric roles of the LoRA factors, we study whether $A$ should be shared across clients while $B$ remains client-specific (Share-A/Local-B), or whether $B$ should instead… ▽ More

    Submitted 10 August, 2026; originally announced August 2026.

    Comments: 20 pages

  37. arXiv:2608.09184  [pdf, ps, other

    cs.AI

    Agentic Router: An Execution-Grounded Continual Learning Approach With Memory

    Authors: Yuxuan Chen, Rongpeng Li, Zhifeng Zhao, Yuntao Liu, Xing Xu, Honggang Zhang

    Abstract: Large language model (LLM) agents provide a promising interface for command-line-based network operations, but a plausible command may still fail or introduce operational risk after execution. Existing approaches mainly focus on command generation or final configuration correctness, and do not use execution-grounded experience to jointly improve candidate coverage and action selection. We propose… ▽ More

    Submitted 10 August, 2026; originally announced August 2026.

  38. arXiv:2608.09120  [pdf, ps, other

    cs.CR cs.LO cs.PL

    Renaming or Tightness: Enforcing Disjunctive Information Flow Policies

    Authors: Xin Xu, Siru Tao, Kaizhen Tan

    Abstract: A disjunctive policy allows a value to depend on at most one of two secrets and never on both: an analyst may consult one client's file or the other's, a share of a split secret may be released but not its sibling. Such policies are not lattice-shaped, and Hunt and Sands introduced the quantale of information to give them a semantics, leaving the enforcement layer open. We build the flow-sensitive… ▽ More

    Submitted 10 August, 2026; originally announced August 2026.

    Comments: 13 pages, 3 figures, 1 table. Under review

  39. arXiv:2608.08982  [pdf, ps, other

    cs.LG

    Twin Rollouts: Noise-Coupled Counterfactual Branching in Interactive Video World Models

    Authors: Yu Ma, Hongli Shi, Xinran Xu

    Abstract: Interactive video world models generate rollouts autoregressively under an action stream, yet they are trained and evaluated almost exclusively on factual prediction. We study counterfactual generation inside the rollout: given a trajectory the model has itself generated, what would have happened had the actions differed from step t* onward? We formalize noise-coupled twin rollouts --- a factual a… ▽ More

    Submitted 9 August, 2026; originally announced August 2026.

  40. arXiv:2608.08765  [pdf, ps, other

    cs.PL quant-ph

    Bona: Automatic Management of Dirty Ancilla Borrowing in Quantum Circuits

    Authors: Xiaoquan Xu, Chenke Liu, Boning Meng, Zihao Shen, Li Zhou

    Abstract: The management of ancilla qubits has become a critical technique for reducing quantum circuit width. Dirty ancillas, which may be borrowed from any temporarily idle qubit regardless of their initial states, offer substantial flexibility for width optimization, but their use has so far required manual and error-prone handling. We formalize the dirty-qubit borrowing problem and establish a fundament… ▽ More

    Submitted 9 August, 2026; originally announced August 2026.

  41. arXiv:2608.08542  [pdf, ps, other

    cs.LG

    When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs

    Authors: Yu Ma, Hongli Shi, Jing Li, Xinran Xu, Weiwei Hou

    Abstract: Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, code, or domain specialists into a safety-aligned base using task arithmetic, TIES, or DARE. This convenience is known to carry a safety cost, but almost all of that evidence rests on static refusal tests: fixed harmful prompts scored for compliance… ▽ More

    Submitted 9 August, 2026; originally announced August 2026.

  42. arXiv:2608.07395  [pdf, ps, other

    cs.SE cs.AI

    PACE: Primitive-Aware Code Evolution for Automated Algorithm Design

    Authors: Zhuoliang Xie, Ruihao Zheng, Xiang Xu, Genghui Li, Zhengkun Wang

    Abstract: Large Language Model (LLM)-based automated algorithm design typically evolves algorithms as complete, indivisible programs. While this whole-program perspective simplifies the search space, it fundamentally couples the useful local logic to its host program. Consequently, valuable code snippets vanish when the overall program is discarded, making it highly difficult to assess the contribution of i… ▽ More

    Submitted 7 August, 2026; originally announced August 2026.

  43. arXiv:2608.06900  [pdf, ps, other

    cs.SD eess.AS

    MMAG: A Multi-Control Mixed Audio Generation Benchmark

    Authors: Zihao Zheng, Xuenan Xu, Jiahao Mei, Yixuan Li, Minghao Lv, Wen Wu, Chao Zhang, Mengyue Wu

    Abstract: Recent audio generation systems have progressed from single-modality synthesis to generating complex acoustic scenes containing speech, music, and sound effects. Therefore, evaluating these models requires assessing multiple interacting capabilities, including semantic fidelity, speaker consistency, and temporal control, yet existing benchmarks focus on isolated domains or coarse-grained descripti… ▽ More

    Submitted 7 August, 2026; originally announced August 2026.

    Comments: 15 pages, 6 figures

    MSC Class: 68Txx ACM Class: I.2

  44. arXiv:2608.06712  [pdf, ps, other

    cs.CV

    Suppress and Diversify: Refining Robust Pathways for Corruption Robustness

    Authors: Jiangang Yang, Wenhui Shi, Xiaoran Xu, Wenyue Chong, Luqing Luo, Jing Xing, Jian Liu

    Abstract: Model robustness against natural image corruptions is essential for safety-critical applications. While existing methods primarily focus on implicit representation learning, we provide the first systematic exploration of computational pathways to explicitly characterize internal robustness. We identify a progressive decay of robust features across network layers and establish a functional dependen… ▽ More

    Submitted 6 August, 2026; originally announced August 2026.

    Comments: Accepted at ICML 2026

  45. arXiv:2608.06085  [pdf, ps, other

    cs.AI

    Signal or Spurious Cue? A Randomized Audit of Survey-Country Metadata in LLM Social Inference

    Authors: Yifan Lyu, Xinran Li, Jiaqi Qiao, Xiujuan Xu

    Abstract: Survey-country metadata can improve an LLM's forecast of an individual response when informative, yet the same cue may redirect the forecast when assigned at random. A within-record audit tests whether disclosing a random label's uniform, record-independent origin reduces its country-directed uptake, and whether verified survey country lowers held-out Brier loss. Independent population anchors and… ▽ More

    Submitted 6 August, 2026; originally announced August 2026.

    Comments: 7 pages, 2 figures

  46. arXiv:2608.05970  [pdf, ps, other

    cs.RO cs.AI

    SkillMemo: Expert-guided Skill Memory Framework for Compositional Embodied Manipulation

    Authors: Changyuan Wang, Chubin Zhang, Zhenyu Wu, Runhao Li, Angyuan Ma, Ke Chao, Yinan Liang, Xiuwei Xu, Ziwei Wang, Yansong Tang, Jiwen Lu

    Abstract: Embodied visuomotor models, including Diffusion Policy (DP) and Vision-Language-Action (VLA) models, have demonstrated promising performance on robotic manipulation benchmarks. However, their potential remains fundamentally constrained by the scarcity of large-scale embodied trajectory datasets, leading to insufficient compositional generalization in out-of-distribution (OOD) scenarios with limite… ▽ More

    Submitted 6 August, 2026; originally announced August 2026.

  47. arXiv:2608.05757  [pdf, ps, other

    cs.CV

    Beyond Relevance: Bayesian Evidence Acquisition for Agentic Whole-Slide Image Reasoning

    Authors: Bryan Wong, Xun Xu, Huazhu Fu, Nancy F. Chen, Mun Yong Yi

    Abstract: Whole-slide image (WSI) reasoning requires an agent to sequentially acquire visual evidence before answering a diagnostic question. Existing training-free agentic frameworks formulate this process as iterative patch retrieval based on semantic relevance to the question. However, semantic relevance does not necessarily imply diagnostic informativeness in computational pathology, where competing dia… ▽ More

    Submitted 6 August, 2026; originally announced August 2026.

  48. arXiv:2608.05716  [pdf, ps, other

    cs.AI

    BlockPython: A Process-Aware Agent-Supported Platform for the Transition from Block-Based to Python Programming

    Authors: Jesse Yusuf Chan, Haoming Wang, Mingwei Xu, Xianlong Xu

    Abstract: The transition from block-based to text-based programming requires learners to convert visible program structures into abstract textual expressions, which may create a cognitive gap between understanding computational concepts and expressing them in Python syntax. To support this transition, we designed and implemented BlockPython. The platform centers on bidirectional translation between blocks a… ▽ More

    Submitted 6 August, 2026; originally announced August 2026.

    Comments: AIED 2026 Interactive Event Track

  49. arXiv:2608.05604  [pdf, ps, other

    cs.CL cs.AI

    SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries

    Authors: Xingyu Tan, Xiaoyang Wang, Qing Liu, Xiwei Xu, Xin Yuan, Liming Zhu, Wenjie Zhang

    Abstract: Large Language Models (LLMs) increasingly act as agents whose procedural knowledge is stored in reusable skill packages and loaded at inference time. As skill libraries grow, a central challenge is to expose the smallest sufficient executable context under a limited context budget. Existing systems struggle to reuse routines below the whole-skill level, preserve procedural contracts during compres… ▽ More

    Submitted 6 August, 2026; originally announced August 2026.

  50. Large Language Models and Social Media Information Integrity: Opportunities, Challenges, and Research Directions

    Authors: Junjie Xiong, Zhengyuan Jiang, Xiaoran Xu, Chi Zhang, Changjia Zhu, Ning Wang, Mingkui Wei, Zhuo Lu, Yao Liu, Lingyao Li

    Abstract: Large Language Models (LLMs) have emerged as powerful tools that impact information integrity on social media platforms. This comprehensive review examines the dual role of LLMs in both facilitating and mitigating various information integrity challenges, including misinformation, disinformation, fake news, social bots, and privacy concerns. \textcolor{black}{We conduct a comprehensive review of t… ▽ More

    Submitted 4 August, 2026; originally announced August 2026.

    Comments: It has been accepted by Computing Surveys. Preview From: htong@illinois.edu Congratulations! Your manuscript, "Large Language Models and Social Media Information Integrity: Opportunities, Challenges, and Research Directions," has been accepted for publication in ACM Computing Surveys.Your paper will be returned to your Author Center. Dr. Hanghang Tong Editor-in-Chief ACM Computing Surveys

    Journal ref: Just accpeted by ACM Computing Surveys 2026