Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 2,024 results for author: Xue, M

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.21060  [pdf, ps, other

    cs.AI cs.CV

    CellPath-Bench: A Multidimensional Benchmark for Whole-Slide Cellular Representations in Pathology Foundation Models

    Authors: Bokai Zhao, Yiyang Zhang, Hanqing Chao, Yawei Ma, Long Bai, Tai Ma, Minfeng Xu, Ming Song, Tianzi Jiang

    Abstract: Pathology foundation models (PFMs) are increasingly used as general-purpose backbones, yet existing benchmarks cannot systematically diagnose their whole-slide cellular representation capabilities, including the decodability of cell-type information and the transferability of such information across tissue sections, datasets, and anatomical organs. We introduce CellPath-Bench, a cellular-resolutio… ▽ More

    Submitted 21 August, 2026; originally announced August 2026.

  2. arXiv:2608.21034  [pdf, ps, other

    cs.DC

    AI Infrastructure in Space: How Far Can We Go?

    Authors: Qing Li, Qiyang Zhang, Daliang Xu, Tianze Huang, Dingge Zhang, Yihao Zhao, Xiaolong Huang, Jinfeng Wen, Xiameng Hu, Tao Qi, Mengwei Xu, Shangguang Wang, Xuanzhe Liu

    Abstract: Satellites are becoming programmable computing platforms capable of running increasingly demanding AI workloads. This shift raises a systems problem: how can AI services remain deployable, manageable, and recoverable after launch when compute capacity, connectivity, energy, and thermal headroom vary over orbital time? This paper develops a systems vision for AI infrastructure in space. We define i… ▽ More

    Submitted 21 August, 2026; originally announced August 2026.

    Comments: 20 pages, 4 figures

  3. arXiv:2608.17988  [pdf, ps, other

    cs.CV

    GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation

    Authors: Ming Qian, Zijian Wang, Minchao Sun, Jincheng Xiong, Hang Zhang, Mu Xu, Chi Wang, Baoquan Chen

    Abstract: Many scalable latent 3D generators operate on structured tensors, whereas pre-optimized 3D Gaussian Splatting (3DGS) reconstructions are unordered, spatially irregular, and vary widely in primitive count. We present GS-Voxel, a fitting-free structured latent framework, and evaluate it for large-scale aerial 3D Gaussian scene generation. GS-Voxel deterministically converts a compatible pre-optimize… ▽ More

    Submitted 18 August, 2026; originally announced August 2026.

  4. arXiv:2608.17396  [pdf, ps, other

    cs.SE

    SNIPTEST: Fuzzing Multi-Level Code Slices for Validating Vulnerabilities

    Authors: Aniruddhan Murali, Nobble Saji Mathews, Mahmoud Alfadel, Meng Xu, Meiyappan Nagappan

    Abstract: Modern software systems are increasingly complex, and static analysis tools are commonly used to identify potentially vulnerable code by issuing warnings. However, these warnings often require manual inspection to confirm whether the reported issues are real, making the process time-consuming and error-prone. Directed fuzzing has emerged as a powerful automated technique to validate the warnings.… ▽ More

    Submitted 18 August, 2026; originally announced August 2026.

  5. arXiv:2608.15830  [pdf, ps, other

    cs.CV

    MITE-Net: SWaP-Optimized 4K Video Tiny Target Perception for Embodied Edge SAR

    Authors: Mingshuo Xu, Mu Hua, Jigen Peng, Qi Wang, Shigang Yue

    Abstract: Real-time tiny target perception in high-resolution imagery is critical for embodied Search-and-Rescue (SAR) missions. However, strict Size, Weight, and Power (SWaP) constraints on edge devices like UAVs create a bottleneck: traditional image downsampling causes severe feature loss, while slice-based processing incurs prohibitive latency. To address this gap, this paper introduces a comprehensive… ▽ More

    Submitted 16 August, 2026; originally announced August 2026.

    Comments: Under double blind review

  6. arXiv:2608.15402  [pdf, ps, other

    cs.LG

    Towards a theory of inference-time alignment with unknown rewards

    Authors: Steve Hanneke, Hongao Wang, Mingyue Xu

    Abstract: Generative model alignment has received broad interest, and significant progress has been made in supervised fine-tuning and inference-time computation. Yet, alignment has remained poorly understood from a statistical learning perspective. We formulate inference-time alignment as a weak-to-strong learning problem, where a reference policy (weak learner) is assumed to be fairly good and the goal is… ▽ More

    Submitted 15 August, 2026; originally announced August 2026.

  7. arXiv:2608.15118  [pdf, ps, other

    cs.DC

    Collective Communication for Distributed LLM Systems: Planning, Runtime Adaptation, and Computation Coordination

    Authors: Xuebin Song, Menghao Zhang, Yuezheng Liu, Jinyi Xia, Shucan Yang, Xiaohe Hu, Chunming Hu, Mingwei Xu

    Abstract: Distributed large language model (LLM) systems increasingly rely on collective communication primitives such as AllReduce (AR), ReduceScatter (RS), AllGather (AG), and AlltoAll (A2A). In modern LLM training and serving clusters, heterogeneous GPU interconnects, multi-NIC networking, mixed parallelism strategies, low-latency inference requests, and high-throughput training pipelines have motivated… ▽ More

    Submitted 15 August, 2026; originally announced August 2026.

  8. arXiv:2608.14533  [pdf, ps, other

    cs.CR

    Finding Vulnerabilities via LLM-Augmented Semantics-Aware Type-Checking

    Authors: Ruizhe Wang, Meng Xu, N. Asokan

    Abstract: Vulnerability detection via static analysis traditionally relies on security experts encoding insecure coding patterns into algorithmic rules. However, this approach often focuses on syntactic patterns and overlooks deeper semantic information in the code, such as the meanings of variable and function names. As software systems grow more complex, modeling vulnerabilities using only syntactic rules… ▽ More

    Submitted 14 August, 2026; originally announced August 2026.

  9. arXiv:2608.13878  [pdf

    cs.RO eess.SY

    Knowledge-Data-Dual-Driven Reinforcement Learning for Autonomous Vehicle Control in Mixed Traffic

    Authors: Jie Fang, Wei Zheng, Mengyun Xu, Eui-Jin Kim

    Abstract: In mixed traffic, decision-making for autonomous vehicles (AVs) confronts three interrelated challenges. First, physics-based priors incorporated into reinforcement learning (RL) models fail to capture latent interactive vehicle intentions and diverse driver behaviors, limiting the proactive reasoning capabilities. Second, abrupt maneuvers by surrounding vehicles cause non-stationarity, leaving lo… ▽ More

    Submitted 13 August, 2026; originally announced August 2026.

    Comments: 16 pages, 17 figures

  10. Print&Fold: Printing and Folding Shape-accurate 3D Models

    Authors: Archit Kumar, Zachary Grimm, Mingsheng Xu, Shlok Rathi, Martin Nisser

    Abstract: This paper introduces Print&Fold, a tool to allow FDM 3D printing of complex models with less time and material while preserving shape accuracy. Key to this work is a folding algorithm that planarizes foldable faces internal to the 3D model. While folding techniques typically discretize a target model's surface, thereby fabricating low fidelity counterparts, our method preserves the surface featur… ▽ More

    Submitted 13 August, 2026; originally announced August 2026.

    Comments: Proceedings of ACM UIST 2026

  11. AirForesight: Current-to-Future Spatial Map Imagination with Cross-Space Planning Consistency for UAV-VLN

    Authors: Yutong Liu, Xiaojie Li, Mingzhu Xu, Jianlong Wu

    Abstract: Unmanned Aerial Vehicle Vision-Language Navigation (UAV-VLN) requires agents to follow language instructions, infer spatial structure from sparse multi-view observations, and execute feasible 3D motion in complex outdoor environments. Despite recent progress with large language models, most existing methods still map vision-language inputs directly to actions, providing limited explicit scene grou… ▽ More

    Submitted 13 August, 2026; originally announced August 2026.

    Comments: Accepted by ACM Multimedia 2026

  12. arXiv:2608.10827  [pdf, ps, other

    cs.CV cs.AI

    MIRA: Medical Image Reflection for Agentic Diagnosis

    Authors: Shengzhi Wang, Jun Yang, Kai Wu, Xiaozhong Ji, Yiwen Ye, Ziyang Chen, Mingliang Xiong, Wen Fang, Mingqing Liu, Mengyuan Xu, Miaoxuan Shan, Caiyan Liu, Bin He, Qingwen Liu

    Abstract: Medical visual agents can use tools to inspect images and retrieve external knowledge, but indiscriminate tool use may introduce noisy or misleading evidence. Reliable diagnosis therefore requires not only acquiring additional observations, but also verifying whether tool actions are necessary and whether the resulting evidence supports the current hypothesis. We introduce MIRA (Medical Image Refl… ▽ More

    Submitted 11 August, 2026; originally announced August 2026.

  13. arXiv:2608.09226  [pdf, ps, other

    cs.CV cs.AI

    RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation

    Authors: Yuhan Li, Fangao Zeng, Sicong Kang, Mengfei Xu, Hao Zhou, Wei Li, Pipei Huang, Bingbing Ni

    Abstract: Efficient text-to-image generation requires both reinforcement-learning (RL)-based reward alignment and few-step distillation, yet these procedures are typically performed sequentially, increasing training cost and risking the loss of reward gains during compression. We instead take an RL-native perspective: diffusion RL already generates reward-scored finite-step trajectories, whose intermediate… ▽ More

    Submitted 10 August, 2026; originally announced August 2026.

  14. arXiv:2608.08416  [pdf, ps, other

    cs.LG

    Optimal Learning Under Tsybakov Noise

    Authors: Steve Hanneke, Hongao Wang, Mingyue Xu

    Abstract: Probably Approximately Correct (PAC) learning [Val84] is a fundamental learning model that has been extensively investigated. In this model, $\mathcal{H} \subseteq \{0,1\}^{\mathcal{X}}$ is a concept class, and $h^*\in\mathcal{H}$ is the target concept to be learned. Having access to i.i.d. labeled examples from a distribution $\mathcal{D}$ over $\mathcal{X}\times\{0,1\}$, which admits $h^*$ as th… ▽ More

    Submitted 8 August, 2026; originally announced August 2026.

  15. arXiv:2608.08264  [pdf, ps, other

    cs.AI

    OBLIVION: Workflow-Level Operational Skill Unlearning for Deployed Agents

    Authors: Zhengyang Shan, Xu Qian, Jiayun Xin, Kun Li, Yue Zhang, Minghui Xu

    Abstract: Large language model agents are becoming operational interfaces to files, memories, registries, and external tools. This deployment shift creates a new skill revocation problem: after a skill is removed from an explicit registry, an agent may still reconstruct it from residual carriers such as archives, transcripts, schemas, or memory entries. We study this problem as operational skill unlearning,… ▽ More

    Submitted 8 August, 2026; originally announced August 2026.

  16. arXiv:2608.05810  [pdf, ps, other

    cs.AI cs.CL

    When Self-Evolution Backfires: Pre-Commit Gating against Skill Contamination in LLM Agents

    Authors: Linfang Shang, Ming Xu, Yiding Sun, Tianle Xia, Lingxiang Hu, Lan Xu, Ning Zheng

    Abstract: Self-evolving agents accumulate capability by distilling reusable skills from their execution trajectories, but we find this process is not monotonic: past a critical pool size, newly added skills degrade performance instead of improving it. We formalize this capability-contamination phase transition and trace it to a structural cause: once a defective skill enters the decision context, it becomes… ▽ More

    Submitted 6 August, 2026; originally announced August 2026.

  17. arXiv:2608.05716  [pdf, ps, other

    cs.AI

    BlockPython: A Process-Aware Agent-Supported Platform for the Transition from Block-Based to Python Programming

    Authors: Jesse Yusuf Chan, Haoming Wang, Mingwei Xu, Xianlong Xu

    Abstract: The transition from block-based to text-based programming requires learners to convert visible program structures into abstract textual expressions, which may create a cognitive gap between understanding computational concepts and expressing them in Python syntax. To support this transition, we designed and implemented BlockPython. The platform centers on bidirectional translation between blocks a… ▽ More

    Submitted 6 August, 2026; originally announced August 2026.

    Comments: AIED 2026 Interactive Event Track

  18. arXiv:2608.05393  [pdf, ps, other

    cs.CV

    Adapting Vision Foundation Models with Cascaded Semantics

    Authors: Xi Xiao, Xingjian Li, Cheng Han, Tianyang Wang, Lin Zhao, Yunbei Zhang, Guosheng Hu, Runmin Jiang, Xi Li, Xiao Wang, Min Xu

    Abstract: Prompt tuning, a leading parameter-efficient adaptation paradigm in NLP, has recently been extended to computer vision. Visual prompt tuning (VPT) adapts pre-trained vision transformers (ViTs) by updating a small set of additional prompt parameters. However, existing visual prompts are randomly initialized and do not exploit prior knowledge, such as instructions in NLP. We address this gap by inje… ▽ More

    Submitted 5 August, 2026; originally announced August 2026.

    Comments: Accepted by Transactions on Machine Learning Research (TMLR), 2026. Project page: https://xixiaouab.github.io/Cascaded-Semantics/

  19. arXiv:2608.04339  [pdf, ps, other

    cs.CL cs.AI cs.CE stat.ML

    Equitable System-Prompt Selection via Constrained Mixed-Strategy GroupDRO

    Authors: Mengyu Xu, Qiaoxin Yang, Zhihan Liu, Ruiyao Xu, Zachary Liu, Kezhen Chen, Chongyang Gao

    Abstract: Large language models are increasingly used for information seeking, yet semantically equivalent questions phrased in different ways can receive answers of considerably different quality. System prompts are widely employed to steer response behavior, but they are typically optimized for average-case quality, so some question phrasings may still receive incomplete or low-quality answers. To address… ▽ More

    Submitted 4 August, 2026; originally announced August 2026.

  20. arXiv:2608.04243  [pdf, ps, other

    cs.LG

    Attention-based representations for multi-task computation

    Authors: Daniel Hsu, Mingyue Xu

    Abstract: Multi-head attention layers produce vector representations that support multiple downstream tasks. We establish bounds on the number of heads required in two simple and concrete multi-task scenarios. In the first scenario, a vector representation is sought so that linear predictors can compute both the smallest and largest numbers in a given list. In this case, it is known two attention heads with… ▽ More

    Submitted 4 August, 2026; originally announced August 2026.

  21. arXiv:2608.03700  [pdf, ps, other

    cs.CR cs.CL cs.CY

    When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills

    Authors: Yongli Xiang, Zhifang Zhang, Bojun Yang, Ziming Hong, Lei Feng, Miao Xu, Tongliang Liu

    Abstract: Persona skills distill personal interaction histories into portable and executable artifacts for downstream agents. While enabling flexible personalization, this process concentrates fragmented personal signals, amplifies their impact through reuse, and challenges defenses designed for individual records or retrieval-based memory. To systematically investigate the safety of the persona-skill pipel… ▽ More

    Submitted 4 August, 2026; originally announced August 2026.

    Comments: Project page: https://yonglixiang.github.io/AntiSkillBench

  22. arXiv:2608.03682  [pdf, ps, other

    cs.AI cs.RO

    PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud

    Authors: Chenghua Wang, Daliang Xu, Dongqi Cai, Duojin Sun, Hao Zhang, Haoze Qian, Huaiyuan Zhang, Jinshuo Cui, Junbo Cui, Kezhao Zhao, Longxi Gao, Mengwei Xu, Rongjie Yi, Ruixin Liu, Shangguang Wang, Tam Sikyuen, Tianyue Zhang, Weikai Xie, Xuanzhe Liu, Yingying Qin, Yiwen Lu, Yuan Yao, Yuezhi Zu, Yunhan Guo, Yuxin Zheng , et al. (1 additional authors not shown)

    Abstract: Physical AI policies require inference throughout their lifecycle, including model evaluation, cloud reinforcement learning rollout, edge GPU serving, and onboard deployment. Although these settings share the same checkpoint and action semantics, they often rely on separate inference programs. To unify them, we build PhyAI, a Physical AI inference engine with a single runtime that keeps architectu… ▽ More

    Submitted 14 August, 2026; v1 submitted 4 August, 2026; originally announced August 2026.

    Comments: 25 pages, 9 figures

  23. arXiv:2608.03655  [pdf, ps, other

    cs.CL cs.AI

    Decoupling Generation and Selection for Budget-Constrained Faithful Summarization

    Authors: Zeyu Wang, Guanghua Wang, Meng Xu

    Abstract: Abstractive summarization models remain vulnerable to factual inconsistency, redundancy, and weak length control. We propose a modular generation-and-selection framework for sentence-budget-constrained summarization. A pretrained generator produces multiple candidate summaries, which are decomposed into sentence-level candidates. A combinatorial selector then constructs the final summary by balanc… ▽ More

    Submitted 4 August, 2026; originally announced August 2026.

  24. arXiv:2608.03391  [pdf, ps, other

    cs.LG

    TimeRLM: Recursive Language Models Enable Precise Anomaly Localization in Long-Context Time-Series

    Authors: Nicolas Zumarraga, Lorenzo Steno, Ning Wang, Max Rosenblattl, Thomas Kaar, Maxwell A. Xu, Kevin O'Sullivan, Markus Kreft, Elgar Fleisch, Paul Schmiedmayer, Patrick Langer, Robert Jakob

    Abstract: Precise anomaly localization over long-context time series is a crucial task in monitoring applications across clinical care, industrial operations, financial services, and logistics, where brief evidence may hide inside long spans of high-frequency data. Time-Series Language Models (TSLMs) are able to ingest time series data and verbalize findings on anomalies in natural language; however, recent… ▽ More

    Submitted 4 August, 2026; originally announced August 2026.

    Comments: Open source code and datasets: https://github.com/OpenTSLM/TimeRLM

  25. arXiv:2608.03296  [pdf, ps, other

    cs.RO cs.CV

    PLS-Calib: A Partial Least Squares Framework for Event Camera and Odometry Calibration under Ground Motion Constraints

    Authors: Guangyu Li, Xiao Li, Yujie Wu, Changshuo Wang, Prayag Tiwari, Jiang Cai, Fangwen Yu, Mingkun Xu

    Abstract: Accurate extrinsic rotation calibration between sensors is fundamental to the performance of robotic perception systems. However, most existing calibration techniques rely on full 6-DoF motion to excite all degrees of freedom, which is often infeasible for ground-constrained robots with limited motion capabilities. Recent approaches designed for such restricted settings, such as Canonical Correlat… ▽ More

    Submitted 4 August, 2026; originally announced August 2026.

    Comments: 8 pages, 10 figures, 4 tables. Accepted at the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026)

  26. arXiv:2608.03247  [pdf, ps, other

    cs.CV cs.CL

    CIGTSurv: Clinical Information Guided Tri-modal Survival Prediction with Local Prototype Association and Global Feature Alignment

    Authors: Jing Dai, Qibin Zhang, Weiwei Zhou, Mingde Xu, Jingsong Liu, Jingdong Zhang, Hongming Xu

    Abstract: Multimodal learning has significantly advanced survival prediction by integrating pathology images with genomic data. However, clinical information, despite its critical role in reflecting a patient' s overall health, remains underutilized due to its discrete, sparse, and low-dimensional nature. Furthermore, the inherent heterogeneity across these modalities pose significant challenges in modeling… ▽ More

    Submitted 4 August, 2026; originally announced August 2026.

    Comments: Accepted at MICCAI 2026

  27. arXiv:2608.02831  [pdf, ps, other

    cs.SD cs.CL

    Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning

    Authors: Fangxu Yu, Tao Feng, Dehai Min, Zinan Lin, Weijia Xu, Michael Xu, Philip S. Yu, Ge Liu, Tianyi Zhou

    Abstract: Audio reasoning is essential for machine understanding of the acoustic world. Reinforcement learning with verifiable rewards can elicit such reasoning, yet existing reward designs are complementary in their limitations: outcome-based rewards supervise only the final answer and let the model reach it without attending to the audio, whereas process-based rewards score the reasoning itself but rely o… ▽ More

    Submitted 3 August, 2026; originally announced August 2026.

  28. arXiv:2608.01910  [pdf, ps, other

    cs.CV

    PNEC-Mamba: Prototype-Guided Positive-Negative Evidence Calibration for Hyperspectral Image Classification

    Authors: Mingzhen Xu, Can Xu, Di Wang, Haonan Guo, Bo Du

    Abstract: In real-world hyperspectral scenes, pixel representations are often ambiguous due to factors such as spectral similarity, mixed pixels, and local context interference, which may simultaneously encode discriminative evidence and interfering information. Existing methods mainly focus on learning more powerful representations or modeling broader contexts, but rarely investigate whether the learned re… ▽ More

    Submitted 3 August, 2026; originally announced August 2026.

  29. arXiv:2608.01907  [pdf, ps, other

    cs.GR

    SNAP-tFDP: Massively Scalable Graph Layouts via Sparse Negative Sampling

    Authors: Xin Chen, Shuowei Hou, Yifan Wang, Mingliang Xue, Zezheng Feng, Oliver Deussen, Weidong Huang, Yunhai Wang

    Abstract: Force-Directed Placement (FDP) is a widely used approach for network visualization, yet scaling it to massive graphs while preserving clear community structures remains a major computational and visual challenge. Existing approximation methods often rely on auxiliary data structures (e.g., spatial trees), which introduce substantial memory overhead; furthermore, traditional power-function-based fo… ▽ More

    Submitted 3 August, 2026; originally announced August 2026.

    Comments: Accepted by IEEE VIS 2026

  30. arXiv:2608.01566  [pdf, ps, other

    cs.IT

    Linear network codes for vector-linear network function computation over three-layer networks

    Authors: Min Xu, Gennian Ge

    Abstract: We study vector-linear function computation over three-layer networks with a fixed target function and a fixed source-access pattern. We develop a support-constrained row-space framework that represents a linear computing code by a global row space. This space must contain the target row space and be generated by rows satisfying the local support-constraints of the network. We prove that this repr… ▽ More

    Submitted 2 August, 2026; originally announced August 2026.

  31. arXiv:2608.01035  [pdf, ps, other

    cs.RO cs.AI cs.CV

    WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA

    Authors: Zhihao Zhu, Hanlin Shang, Mingwang Xu, Feipeng Cai, Zhuolin He, Yaoyi Li, Jianhua Han, Hang Xu, Siyu Zhu

    Abstract: Vision-Language-Action (VLA) models have emerged as a prominent paradigm for end-to-end autonomous driving; however, their efficient deployment is severely constrained by high computational latency and exposure bias arising from sequential autoregressive decoding. Conversely, while specialized diffusion policies enable low-latency, parallel execution, training them from scratch typically yields na… ▽ More

    Submitted 18 August, 2026; v1 submitted 2 August, 2026; originally announced August 2026.

  32. arXiv:2608.00814  [pdf, ps, other

    cs.CL

    OoO-Spec: Out-of-Order Semantic Speculation for Fast Tool Calling

    Authors: Zhiheng Zhang, Mujie Xu, Feiyu Sun, Zhixin Zhang

    Abstract: LLMs generate tool calls token by token, even though the function choice and argument values can often be predicted in parallel from the request and tool schema. ToolSpec reduces this cost by drafting schema tokens and retrieving earlier calls, but cannot propose request-specific values absent from either source. We present OoO-Spec, which computes these missing semantics out of order. At request… ▽ More

    Submitted 1 August, 2026; originally announced August 2026.

    Comments: 10 pages, 4 figures, 6 tables; supplementary material included

  33. arXiv:2607.29071  [pdf, ps, other

    cs.LG cs.AI

    Federated Foundation Models Fine-Tuning with Heterogeneous Compressed Clients

    Authors: Shengkun Zhu, Jinshan Zeng, Zhihua Allen-Zhao, Mayi Xu, Quanqing Xu, Wei Ren, Qiang Yang, Yang Liu

    Abstract: Federated learning of foundation models faces a fundamental resource-asymmetry challenge: the institutions holding the most valuable domain-specific data cannot host billion-parameter models. Existing heterogeneous federated approaches attempt to bridge this gap through parameter-efficient tuning, model pruning, or knowledge distillation, yet each trades away a critical property, whether full-mode… ▽ More

    Submitted 31 July, 2026; originally announced July 2026.

  34. arXiv:2607.29039  [pdf, ps, other

    cs.CV

    ReMoE: Report-Guided Mixture-of-Experts for Multimodal OCT/OCTA Anomaly Detection

    Authors: Zihan Nie, Qincheng Qiao, Muhao Xu, Wei Feng, Xinguo Hou, Weiye Song, Zongyuan Ge

    Abstract: Multimodal medical anomaly detection identifies samples deviating from normal patterns, where scarce abnormal cases make normality modeling from normal data practical. In retinal Optical Coherence Tomography (OCT) and OCT Angiography (OCTA) anomaly detection, existing unsupervised methods rely on visual feature distributions, reconstruction residuals, or encoder-decoder discrepancies, making anoma… ▽ More

    Submitted 31 July, 2026; originally announced July 2026.

  35. arXiv:2607.29002  [pdf, ps, other

    cs.AI

    MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping Agents

    Authors: Zeying Hao, Hao Guo, Mengtao Xu, Yimin Hu, Yuheng Song, Zesheng Zhou, Jinsong Lan, Xiaoyong Zhu

    Abstract: Online shoppers increasingly turn to AI shopping assistants, using images and multi-turn dialogue to express and refine product needs that are difficult to articulate in text alone. However, existing benchmarks largely rely on text-only or synthetic requests, underrepresenting complex real-world shopping requirements jointly expressed through images and language. We introduce MMShopBench, the firs… ▽ More

    Submitted 31 July, 2026; originally announced July 2026.

    Comments: 16 pages, 6 figures, including appendix

  36. arXiv:2607.28993  [pdf, ps, other

    cs.RO cs.CV

    ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts

    Authors: Mingxin Wang, Bin Hu, Bin Qian, Kaitao Jiang, Haoning Wu, Feng Yan, Bowen Jing, Ruiyang Hao, Enyi Wang, Kangning Niu, Yandan Yang, Mu Xu, Yan Wang, Houde Liu, Tianlun Li

    Abstract: World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future supervision can entangle action-relevant state transitions with task-irrelevant visual content, limiting robustness under visual distribution shifts. We identify Training-Distribution Hallucination, a recurring phenomenon i… ▽ More

    Submitted 30 July, 2026; originally announced July 2026.

    Comments: 9 pages, 5 figures

  37. arXiv:2607.28154  [pdf, ps, other

    cs.CV cs.AI

    OPLD: On-Policy Latent Distillation for Multimodal Reasoning

    Authors: Shoutai Zhu, Tianyang Xu, Bin Sun, Mingyuan Xu, Yu Liu, Qinzhen Guo

    Abstract: Interleaved multimodal Chain-of-Thought (CoT) improves visual reasoning by incorporating auxiliary visual evidence into intermediate reasoning. However, existing approaches remain constrained by externally defined reasoning traces and visual operations, limiting their ability to develop flexible and abstract visual thinking. Reasoning with latent has recently offered a promising direction by inter… ▽ More

    Submitted 30 July, 2026; originally announced July 2026.

  38. arXiv:2607.28128  [pdf, ps, other

    cs.CL cs.AI cs.CY

    Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models

    Authors: Shuyi Fan, Boyuan Deng, Mengyu Xu, Jiale Liu, Hongyang Zhang, Qiaoxin Yang, Chongyang Gao

    Abstract: LLM tutoring poses a measurement problem: can a general-purpose helpfulness rubric distinguish direct answer-giving from pedagogical guidance? We audit this signal in a pre-registered study. Within each of three tutor bases, we compare conversational and pedagogical policies instantiated with the same underlying model and paired with one fixed weak simulated student. Deterministic detectors measur… ▽ More

    Submitted 31 July, 2026; v1 submitted 30 July, 2026; originally announced July 2026.

    Comments: 24 pages, 4 figures, 6 tables

    MSC Class: 68T50 (Primary) 68T05; 97U50; 62G10 (Secondary) ACM Class: K.3.1; I.2.7

  39. arXiv:2607.27138  [pdf, ps, other

    cs.RO cs.AI cs.CV

    DLAM: Distributional Latent Actions with Temporal Constraints

    Authors: Zuojin Tang, Feifan Luo, Haoyun Liu, Botai Yuan, Dekang Qi, Ronghan Chen, Yandan Yang, Tong Lin, Xinyuan Chang, Mu Xu, Bin Liu, De Ma, Zhiheng Ma

    Abstract: Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant observations of physical change. Latent action models can extract such priors, but reconstruction-trained codes may predict future observations without the structure required for joint generation with robot actions. Existing structured methods add temporal constrain… ▽ More

    Submitted 29 July, 2026; originally announced July 2026.

  40. arXiv:2607.26246  [pdf, ps, other

    cs.LG

    Weak-to-Strong On-Policy Distillation

    Authors: Fangxu Yu, Weijia Xu, Michael Xu, Tianyi Zhou, Zinan Lin

    Abstract: On-policy distillation (OPD), which aligns a student with the teacher's token-level distribution on the student's own rollouts, is an effective paradigm for transferring capabilities across LLMs. Prevailing approaches assume a teacher at least as capable as the student: they either distill a larger model into a smaller one, which fails at the frontier where no larger teacher exists, or consolidate… ▽ More

    Submitted 2 August, 2026; v1 submitted 28 July, 2026; originally announced July 2026.

    Comments: Technical Report

  41. arXiv:2607.24453  [pdf, ps, other

    cs.CV cs.AI cs.LG

    ESRVS: Extreme Semi-Supervised Retinal Vessel Segmentation with a Single Annotated Image

    Authors: Mingzhi Xu, Yizhe Zhang

    Abstract: Learning from minimal human supervision is a long-standing goal in medical image analysis, where dense expert annotations are costly. We study retinal vessel segmentation in an extreme semi-supervised setting with one annotated image and a pool of unlabeled images. We propose ESRVS, which selects a representative reference image for manual annotation and transfers vessel cues using target-domain-a… ▽ More

    Submitted 27 July, 2026; originally announced July 2026.

    Comments: 15 pages, 2 figures.Accepted by PRCV 2026

  42. arXiv:2607.23108  [pdf, ps, other

    cs.RO

    The Curse of Precision: A Data Scaling Law for High-Precision Robotic Manipulation

    Authors: Cuijie Xu, Yuanfan Xu, Min Xue, Jianjie Lin, Jian Wang, Xudong Zhang, Yu Wang, Jincheng Yu

    Abstract: While scaling laws for imitation learning have primarily focused on generalization in open-world settings, the relationship between data and precision in closed-world tasks like robotic assembly remains largely unexplored. This paper systematically investigates this relationship and introduces a novel scaling law. We find that to achieve a fixed success rate, the required number of demonstrations… ▽ More

    Submitted 25 July, 2026; originally announced July 2026.

    Comments: 8 pages, 4 figures. Accepted to the 2026 IEEE International Conference on Robotics and Automation (ICRA 2026)

  43. arXiv:2607.23015  [pdf, ps, other

    cs.CR

    Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks

    Authors: Ying JinCheng, Minghui Xu, Yinhao Xiao, Xiuzhen Cheng, Wencheng Yang

    Abstract: Large language models (LLMs) are safety-aligned before deployment to reduce harmful content generation. Yet neuron-level pruning attacks show that refusal can depend on a small set of removable units: disabling them can remove safety behavior while leaving much of the model usable. To address this problem, we introduce Mask2Shield (M2S), a masked-forward alignment method that trains a model under… ▽ More

    Submitted 24 July, 2026; originally announced July 2026.

  44. arXiv:2607.22633  [pdf, ps, other

    cs.AI

    Evolving from Lessons: Skill-Augmented Table Graph Reasoning for Operation-wise Table Question Answering

    Authors: Guixin Su, Qiankun Pi, Mayi Xu, Wenli Li, Ming Zhong, Yuanyuan Zhu, Jiawei Jiang, Tieyun Qian

    Abstract: Table Question Answering (TableQA) aims to reason over tables to answer user queries. Existing research treats all questions uniformly and evaluates solely through overall accuracy, obscuring a critical reality that LLMs excel at simple lookups yet struggle with complex operations like aggregation and arithmetic. To reveal this disparity, we introduce a novel \emph{Operation-wise TableQA} task wit… ▽ More

    Submitted 20 June, 2026; originally announced July 2026.

  45. arXiv:2607.22036  [pdf, ps, other

    cs.NE

    On the Runtime Analysis of Reinforcement Learning Hyper-Heuristics

    Authors: Pietro S. Oliveto, Zhenyu Wang, Peizhou Wu, Mengqing Xu

    Abstract: Selection Hyper-heuristics (HHs) automate algorithmic design by selecting from a set of low-level heuristics which one to apply at each stage of the optimisation process. Several impressive results have been recently rigorously proven regarding the performance of selection hyper-heuristics (HHs) for standard benchmark functions. However, the learning mechanisms employed by these HHs are considerab… ▽ More

    Submitted 24 July, 2026; originally announced July 2026.

    Comments: Published in PPSN 2026

  46. arXiv:2607.21151  [pdf, ps, other

    cs.AI

    V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure

    Authors: Zhetong Zhang, Honghao Fu, Miao Xu, Yiwei Wang, Yujun Cai

    Abstract: As Video Large Language Models are increasingly deployed in real-world applications, ensuring their safety alignment has become critical. Counterintuitively, we find that harmful videos paired with benign queries achieve higher attack success rates than the same videos paired with explicitly harmful queries. To understand the underlying mechanism of this vulnerability, we present V-DEAL, a three-l… ▽ More

    Submitted 26 July, 2026; v1 submitted 23 July, 2026; originally announced July 2026.

  47. arXiv:2607.20857  [pdf, ps, other

    cs.LG cs.AI

    Multilevel Graph Wavelet Compressed Sensing with Scale-Aware Neural Recovery

    Authors: Amirhossein Nouranizadeh, Sarang Rajendra Patil, Alan John Varghese, Varsha Narayanan, Amit Chakraborty, Mengjia Xu

    Abstract: Scientific machine learning methods such as neural operators and physics-informed neural networks have advanced engineering applications and inverse problems, but their training typically requires large volumes of simulated data. This makes data preparation and model training expensive. We propose Graph Wavelet Compressed Sensing (GWCS), a learning-based framework for offline compression of graph… ▽ More

    Submitted 8 August, 2026; v1 submitted 22 July, 2026; originally announced July 2026.

  48. arXiv:2607.20767  [pdf, ps, other

    cs.CL

    Rushes: A Human Preference Dataset for Pluralistic Alignment

    Authors: Michael Xu, Jorge Leandro, Sudha Rao, Weijia Xu, Nebojsa Jojic, Gabriel DesGarennes, Chris Quirk, Bill Dolan

    Abstract: We introduce Rushes, a dataset and benchmark for studying revealed human engagement preferences in interactive narrative environments. Rushes is collected through a game interface where users interact with AI-generated branching narratives and select one choice from a small, explicit candidate set at each decision point. Each interaction logs the full candidate set, the user's choice, and the evol… ▽ More

    Submitted 22 July, 2026; originally announced July 2026.

  49. arXiv:2607.19191  [pdf, ps, other

    cs.CV cs.AI cs.LG

    ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

    Authors: Fan Jiang, Zhaoxu Sun, Mengchao Wang, Ziyu Zhu, Chiyu Wang, Yunpeng Zhang, Wenlin Liu, Yun Wang, Xue Zheng, Rui Sun, Junfeng Ni, Hongyu Pan, Zhongxu Sun, Fei Yu, Zengye Ge, Mengmeng Du, Nianfei Fan, Mingchao Sun, Yu Liu, Yongchang, Yanqing Zhu, Jiahang Wang, Ning Ying, Yuze Xuan, Di Yang , et al. (16 additional authors not shown)

    Abstract: We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn controllable world dynamics. WorldExplorer performs agent-driven collection guided by training feedback, while a unified pipeline applies 14 deterministic quality ch… ▽ More

    Submitted 21 July, 2026; originally announced July 2026.

  50. arXiv:2607.18142  [pdf, ps, other

    cs.CV cs.AI cs.CL cs.MA

    O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning

    Authors: Mei Yuan, Qi Long, Qifeng Wu, Zhenyang Li, Yizhou Zhao, Lei Wang, Yang Liu, Min Xu

    Abstract: Industrial Video Anomaly Detection (IVAD) aims to identify anomalous objects and events in an industrial process, which is crucial for modern manufacturing and quality control systems. Existing VLM-based anomaly reasoning methods are capable of detecting open-ended anomalies in general domains. However, their performance declines in industrial settings characterized by intricate object transformat… ▽ More

    Submitted 20 July, 2026; originally announced July 2026.

    Comments: Accepted to ECCV 2026