Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 325 results for author: Lv, X

.
  1. arXiv:2608.20312  [pdf, ps, other

    cs.CV

    Inter-X++: A Comprehensive Benchmark for Multimodal Human-Human Interaction Analysis

    Authors: Liang Xu, Chengqun Yang, Zili Lin, Xintao Lv, Yichao Yan, Xin Jin, Zhibo Chen, Xiaokang Yang, Wenjun Zeng

    Abstract: The capability to perceive and synthesize human-human interactions is fundamental to developing intelligent digital human systems. However, existing datasets and modeling approaches are fundamentally constrained by low-fidelity kinematics, the omission of dexterous hand gestures and a severe lack of rich multimodal annotations. Furthermore, fragmented interaction representations and inconsistent e… ▽ More

    Submitted 20 August, 2026; originally announced August 2026.

    Comments: 24 pages, 10 figures

  2. arXiv:2608.16837  [pdf, ps, other

    cs.RO cs.AI

    HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL

    Authors: Langzhe Gu, Chengkai Hou, Meng Li, Xinhua Wang, Jiaming Liu, Xinyuan Lv, Bowei Zhang, Shuanghao Bai, Guangrun Li, Jingyang He, Gaole Dai, Ziluo Ding, Zhiyuan Xu, Kuan Cheng, Jian Tang, Zhengping Che, Shanghang Zhang

    Abstract: Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-language-action (VLA) foundation models are not readily applicable to humanoid whole-body loco-manipulation. The high dimensionality and interdependence of humanoid motions make it challenging for conventional single-stage VLA architectures to coordinate locomotion, waist posture, and… ▽ More

    Submitted 17 August, 2026; originally announced August 2026.

    Comments: Project page: https://grange007.github.io/HAF

  3. arXiv:2608.10905  [pdf, ps, other

    cs.LG

    ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation

    Authors: Ximo Zhu, Ruiqi Liu, Rong Wang, Ping Wu, Xiang Zheng, Wenzhuo Xu, Xubin Yao, Zhiyuan Yan, Bo Li, Jun Gao, Xiaolei Lv

    Abstract: On-policy distillation (OPD) applies token-level teacher supervision to student-generated trajectories, but this supervision is not always reliable. Existing methods use local confidence or teacher-student agreement to weight, filter, or truncate the sampled trajectory. These signals do not directly determine whether the teacher can continue a student prefix to a correct answer, and trajectory-lev… ▽ More

    Submitted 11 August, 2026; originally announced August 2026.

  4. arXiv:2608.01316  [pdf, ps, other

    quant-ph cs.AR

    Compiler Framework for 3D Neutral-Atom Quantum Computers

    Authors: Chen Huang, Zhemin Zhang, Zhao Zhang, Xudong Lv, Zhiding Liang

    Abstract: Neutral-atom quantum computers can now arrange atoms in three-dimensional tweezer arrays, yet every existing compiler assumes a flat geometry. We present Piqasso, a compiler that exploits the vertical axis by stacking storage, entanglement, and readout into distinct layers. Its pipeline pairs an analytical placement respecting axial-clearance optics with a router that brings gate partners together… ▽ More

    Submitted 2 August, 2026; originally announced August 2026.

    Comments: 13 pages, 9 figures

  5. arXiv:2607.29025  [pdf, ps, other

    cs.CV

    Evaluation-Verification Reward for Consistent Multi-Reference Image Editing

    Authors: Yingmao Miao, Pengfei Zhang, Xiaochen Lv, Meng Yu, Lei Sun, Xiangxiang Chu, Chao Shen, Chenhao Lin

    Abstract: While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency across references and ensuring overall visual harmony. Reinforcement learning has proven highly effective for text-to-image generation and single-image editing, but its extension to multi-reference editing is hindered by the absence of suitable rew… ▽ More

    Submitted 5 August, 2026; v1 submitted 31 July, 2026; originally announced July 2026.

  6. arXiv:2607.28568  [pdf, ps, other

    cs.CL

    Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering

    Authors: Junlin Yang, Che Jiang, Yu Fu, Tianwei Luo, Can Ren, Weizhi Wang, Kaikai Zhao, Hongyi Liu, Yuxin Zuo, Yuru Wang, Yuchen Fan, Kai Tian, Zhenzhao Yuan, Xiaojian Lin, Li Sheng, Rushi Qiang, Guoli Jia, Xingtai Lv, Ermo Hua, Dianqiao Lei, Youbang Sun, Ning Ding, Bowen Zhou, Kaiyan Zhang

    Abstract: Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability. We introduce OpenMLE, an open full-stack system for RSI research in MLE, spanning verifiable task environments with execution feedback (OpenMLE-Gym), operator learning (OpenMLE-RL), and lon… ▽ More

    Submitted 30 July, 2026; originally announced July 2026.

  7. arXiv:2607.26500  [pdf, ps, other

    cs.IR

    Multi-Decoder OneRec: Controllable Generative Retrieval for Multi-Objective Industrial Recommendation

    Authors: You Wang, Zhao Liu, Guoping Tang, Yiqing Yang, Shuo Su, Jing Liu, Naifu Zhou, Xiaoyou Zhou, Wei Jiang, Jian Liang, Xiao Lv, Ruiming Tang, Liyin Hong, Wenwu Ou

    Abstract: Industrial recommender systems build candidate pools by assigning explicit quotas to objective-specific retrieval routes. This design offers quota control but increasingly fragments modeling, training, and serving as the route set grows. Semantic-ID-based generative retrieval provides a unified alternative, yet a single decoder entangles objective policies and limits candidate complementarity. We… ▽ More

    Submitted 29 July, 2026; originally announced July 2026.

    Comments: 9 pages, 4 figures, 11 tables, 2 algorithms

  8. arXiv:2607.24241  [pdf, ps, other

    cs.CV cs.AI

    FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

    Authors: Shengyi Wang, Niantong Li, Guangzheng Hu, Hong Qi, Fei Ding, Weixu Qiao, Jinlin Wang, Xiaotong Lv, Peng Han, Zimeng Li, Fanshu Ding, Yushu Wang, Han Wu, Jingjing Chen, Chongxiao Wang, Yanhao Wu, Chenglong Huang, Xiaoqian Zhu, Jie Tian, Hua Li, Jingjing Fan, Mingshuang Tang, Zhong Li, Hengxia Qiang, Weibin Chen , et al. (5 additional authors not shown)

    Abstract: Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than… ▽ More

    Submitted 29 July, 2026; v1 submitted 27 July, 2026; originally announced July 2026.

  9. arXiv:2607.23938  [pdf, ps, other

    eess.AS

    Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm

    Authors: Bajian Xiang, Cheng Wen, Han Zhao, Hao Wang, Haoxu Wang, Jiawei Jin, Jiayan Cui, Jie Chen, Mengxi Nie, Tianyu Zhao, Weiqin Li, Xiang Lv, Xiangang Li, Yang Xiang, Yang Zhou

    Abstract: In this report, we present Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system that jointly advances content consistency, speaker similarity, prosodic naturalness, audio quality, controllability, multilingual coverage, efficiency, and robustness. It combines a 12.5~Hz low-frame-rate speech tokenizer for reduced inference latency with a five-stage progressive training paradigm for coo… ▽ More

    Submitted 26 July, 2026; originally announced July 2026.

    Comments: 19 pages

  10. arXiv:2607.21400  [pdf, ps, other

    cs.RO cs.AI

    VoLN: Vision-Only Long-Horizon Navigation---Paradigm, Benchmark, and Method

    Authors: Jiabin Lou, Haopeng Wang, Yuanshuai Wang, Xinyu Liu, Xuxin Lv, Yuxin Guo, Lei Huang, Rongye Shi, Wenjun Wu

    Abstract: Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions. However, route-level instructions commonly encode spatial priors, such as orientation, distance, and layout, that are not explicitly available from onboard sensing at deployment in open, GPS-denied environments. Benchmark performance under such interfaces therefore jointly reflects visual navigati… ▽ More

    Submitted 23 July, 2026; originally announced July 2026.

    Comments: 10 pages, 7 figures, 2 tables. Project page: https://admire-ljb.github.io/VoLN-UAV/

  11. arXiv:2607.18204  [pdf, ps, other

    quant-ph

    QuantiSpect: A Structure-Aware Lightweight 3D CNN Pre-Decoder for Scalable Surface Code Quantum Error Correction

    Authors: Pan Gao, Xu-Sheng Xu, Ji-Ze Han, Jing-Wei Wen, Ling Qian, Xudong Lv, Run-Qing Zhang, Xiao-Xiao Hu, Gui-Lu Long

    Abstract: Real-time decoding is a critical bottleneck for large-scale fault-tolerant quantum computing. AI-based neural pre-decoders locally correct most physical errors before passing residual syndromes to a global decoder, enabling sub-microsecond latencies. However, existing architectures carry significant overhead from dense 3D convolutions. We present QuantiSpect, a lightweight 3D convolutional neural… ▽ More

    Submitted 5 August, 2026; v1 submitted 20 July, 2026; originally announced July 2026.

    Comments: 22 pages, 10 figures

  12. RECAP: Feedback-Driven Streaming Semantic User Profiles for Short-Video Recommendation

    Authors: Ziyi Zhao, Xiaoyou Zhou, Xiao Lv, Yangyang Li, Chubo He, Zhao Liu, Jiayao Shen, Yuqi Liu, He Li, Chengyi Zhang, Jian Liang, Ming Li, Chongming Gao, Fuli Feng, Ruiming Tang, Han Li

    Abstract: Language-based user profiles convert long behavioral histories into explicit semantic representations for recommendation. However, most profile generators are optimized in an open loop: they may summarize past behavior fluently, but are not directly trained to improve future recommendation. We study this problem in real-world short-video recommendation, where user behaviors continuously arrive as… ▽ More

    Submitted 17 July, 2026; originally announced July 2026.

    Comments: 9 pages, 5 figures, 2 tables. Accepted at RecSys 2026

  13. SpeechGuard: Online Defense against Backdoor Attacks on Speech Recognition Models

    Authors: Jinwen Xin, Xixiang Lv

    Abstract: Backdoor attacks pose a critical threat to neural network models, allowing attackers to implant a backdoor during the training phase by manipulating a small portion of the training data. In security-sensitive applications such as voice interaction for autonomous driving, the presence of backdoor attacks introduces substantial security risks. This study focuses on implementing backdoor defense meas… ▽ More

    Submitted 17 July, 2026; originally announced July 2026.

    Comments: 8 pages

    Journal ref: 2024 International Joint Conference on Neural Networks (IJCNN), pp. 1-8, 2024

  14. arXiv:2606.28758  [pdf, ps, other

    cs.CV cs.AI

    X-Mind: Efficient Visual Chain-of-Thought via Predictive World Model for End-to-End Driving

    Authors: Bohao Zhao, Chengrui Wei, Guangfeng Jiang, Ruixin Liu, Xuejie Lv, Liu Liang, Sutao Deng, Xiuyang Fan, Pengkun Zheng, Jinyun Zhou, Rui Guo, Hanpeng Liu, Yutong Zheng, Yi Guo, Xinlong Zheng, Qingyu Luo, Zhuangzhuang Ding, Yu Zhang, Hang Zhang, Xianming Liu

    Abstract: Predicting future states is essential for autonomous agents, yet current Vision-Language-Action (VLA) models fundamentally lack this capability, relying instead on reactive perception-action mapping. While integrating Predictive World Models (PWMs) addresses this gap, existing approaches either incur prohibitive cascaded latency or act as shallow terminal tasks that fail to deeply embed forward-lo… ▽ More

    Submitted 27 June, 2026; originally announced June 2026.

  15. arXiv:2606.26859  [pdf, ps, other

    cs.AI cs.CL cs.IR

    AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems

    Authors: Changxin Lao, Fei Pan, Guozhuang Ma, Han Li, Huihuang Lin, Jijun Shi, Kangzhi Zhao, Kun Gai, Mo Zhou, Qinqin Zhou, Quan Chen, Ruochen Yang, Shifu Bie, Shijie Yi, Shuang Yang, Shuo Yang, Wenhao Li, Wentao Xie, Xiao Lv, Xuming Wang, Yijun Wang, Yiming Chen, Yusheng Huang, Zhongyuan Wang, Zibo Zhao , et al. (37 additional authors not shown)

    Abstract: Recommendation algorithm iteration is moving from an artisanal, engineer-bound process toward an industrialized research loop, but this transition remains blocked by a structural execution bottleneck: the idea-to-launch cycle still depends on human engineers to generate hypotheses, modify production code, launch A/B experiments, and attribute online results. Innovation therefore scales linearly wi… ▽ More

    Submitted 26 June, 2026; v1 submitted 25 June, 2026; originally announced June 2026.

    Comments: Authors are listed alphabetically by their first name

  16. arXiv:2606.24206  [pdf, ps, other

    cs.CV cs.AI

    Inclusive Interactive Collisions for Multi-View Consistent Compositional 3D Generation

    Authors: Chang Liu, Mingwen Shao, Xiang Lv, Xinyuan Chen, Lingzhuang Meng, Qiao Zhang, Zhengyi Gong, Jinghao Hu

    Abstract: Recent breakthroughs in 3D generation have advanced notably with the development of text-to-image diffusion model. However, existing methods remain two practical challenges: (1) They primarily generate single 3D object, but struggle to generate multi-object compositional 3D assets due to the lack of the modeling for Gaussian primitives in reasonable interactions. (2) They often suffer from cross-v… ▽ More

    Submitted 23 June, 2026; originally announced June 2026.

  17. arXiv:2606.23190  [pdf, ps, other

    eess.AS cs.SD

    FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech

    Authors: Haoxu Wang, Biao Tian, Weiqin Li, Xiang Lv, Han Zhao, Xiangang Li

    Abstract: Existing Reinforcement Learning (RL) research for Text-to-Speech (TTS) focuses on large language models (LLMs), leaving Flow-Matching (FM) under-explored. We present FlowTTS-GRPO, an online RL framework for FM-based TTS. By converting ordinary differential equation (ODE) trajectories into stochastic differential equation (SDE) paths, our method enables direct fine-tuning of open-source FM models w… ▽ More

    Submitted 8 July, 2026; v1 submitted 22 June, 2026; originally announced June 2026.

    Comments: Accepted by Interspeech 2026

  18. arXiv:2606.18375  [pdf, ps, other

    cs.RO

    PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation

    Authors: Yuhang Huang, Xuan Lv, Junyan Xu, Zhiyuan Yu, Jiazhao Zhang, Ruizhen Hu, Wancheng Feng, Shilong Zou, Hewen Xiao, Ziqiao Zhou, Kaiyun Huang, Zhiyu Peng, Juzhan Xu, Hang Zhao, Chenyang Zhu, Renjiao Yi, Yifei Huang, Douhui Wu, Yan Zhang, Kexu Cheng, Chunhe Song, Yunzhi Xue, Xiuhong Zhang, Leitao Guo, Yunji Chen , et al. (3 additional authors not shown)

    Abstract: World foundation models (WFMs) are powerful simulators, yet they predominantly operate in a single-view setting and lack the multi-view 3D consistency required for robotic manipulation. While robotic systems rely on multiple cameras (egocentric, eye-to-hand, and wrist-mounted) for policy learning, current multi-view world models simply concatenate view tokens without explicit geometric reasoning.… ▽ More

    Submitted 23 June, 2026; v1 submitted 16 June, 2026; originally announced June 2026.

  19. arXiv:2606.15724  [pdf, ps, other

    math.FA math.CV

    Carleson measures, tent embeddings, and Volterra-type integral operators on the unit ball

    Authors: Xiaofen Lv, Xiaomin Tang, Jani A. Virtanen

    Abstract: In this paper, we establish a sharp comparison between Carleson-cube and Bergman-metric-ball conditions on the open unit ball $\B$ and combine it with a Berezin-type characterization to prove embedding theorems for Besov spaces and Bergman spaces on $\B$ into logarithmic tent spaces in the Bergman metric. As applications, we characterize the boundedness, compactness, and essential norms of the Vol… ▽ More

    Submitted 14 June, 2026; originally announced June 2026.

    Comments: 23 pages

    MSC Class: Primary 47B38; Secondary 32A37

  20. arXiv:2606.12281  [pdf, ps, other

    cs.MA cs.AI cs.LG

    CCKS: Consensus-based Communication and Knowledge Sharing

    Authors: Jinyuan Zu, Xiaowei Lv, Yongcai Wang, Deying Li, Yunjun Han, Wenping Chen, Fengyi Zhang, Naiqi Wu

    Abstract: In Decentralized Training and Decentralized Execution (DTDE) for cooperative Multi-Agent Reinforcement Learning (MARL), action-advising-based knowledge sharing promotes interpretable and scalable cooperation among agents. However, current action advising approaches often adhere too much to the teacher's guidance without evaluating teacher-student compatibility, which causes excessive advising, sub… ▽ More

    Submitted 10 June, 2026; originally announced June 2026.

  21. arXiv:2606.07324  [pdf, ps, other

    physics.atom-ph

    Anomalous Autler-Townes Splitting in Resonant Multiphoton Ionization Driven by Bright Squeezed Vacuum

    Authors: Xu Zhang, Liding Li, Yutong Deng, Xinyou Lv, Yang Li, Marcelo F. Ciappina, Peixiang Lu, Yueming Zhou

    Abstract: Bright squeezed vacuum (BSV) light has a vanishing mean optical electric field yet can strongly enhance strong-field nonlinear responses beyond the conventional semiclassical paradigm. Here we examine this scenario in the light-matter strong-coupling regime by investigating resonant multiphoton ionization of atoms driven by BSV, using a fully quantum treatment of both the electron and the field. O… ▽ More

    Submitted 5 June, 2026; originally announced June 2026.

    Comments: 8 pages, 5 figures

  22. arXiv:2606.06260  [pdf, ps, other

    cs.IR cs.AI cs.CL

    OneReason Technical Report

    Authors: OneRec Team, Biao Yang, Boyang Ding, Chenglong Chu, Dunju Zang, Fei Pan, Han Li, Hao Jiang, Honghui Bao, Huanjie Wang, Jian Liang, Jiangxia Cao, Jiao Ou, Jiaxin Deng, Jinghao Zhang, Kun Gai, Lu Ren, Peiru Du, Pengfei Zheng, Rongzhou Zhang, Ruiming Tang, Shiyao Wang, Siyang Mao, Siyuan Lou, Teng Shi , et al. (59 additional authors not shown)

    Abstract: Generative recommendation models in the OneRec family have been widely deployed in many real-world services, such as short-video, live-streaming, advertising, and e-commerce. However, these generative models can only benefit from the scaling advantage, while their reasoning ability is hard to activate, since we cannot construct meaningful Chain-of-Thought (CoT) sequences consisting of itemic token… ▽ More

    Submitted 4 June, 2026; originally announced June 2026.

    Comments: Work in progress

  23. arXiv:2606.05575  [pdf, ps, other

    cs.SD eess.AS

    SB-RF: Schrödinger Bridge Rectified Flow for One-Step Robust Speech Enhancement

    Authors: Caixia Lu, Xueyang Lv, Penglong Hu, Jiaming Xu

    Abstract: Generative models have shown promising results for speech enhancement (SE), but they often rely on multi-step inference, limiting low-latency deployment. We propose SB-RF, a one-step generative framework that integrates Rectified Flow (RF) with Schrödinger Bridge (SB) theory. During training, SB-RF samples intermediate states from an SB time marginal and trains a conditional velocity field with th… ▽ More

    Submitted 2 July, 2026; v1 submitted 3 June, 2026; originally announced June 2026.

  24. arXiv:2606.04584  [pdf, ps, other

    cs.SD

    SHB-AE: Spherical harmonic beamforming based Ambisonics encoding and upscaling method for smartphone microphone array

    Authors: Yuhuan You, Yufan Qian, Tianshu Qu, Bin Wang, Xueyang Lv

    Abstract: With the rapid development of virtual reality (VR) and augmented reality (AR), spatial audio recording and reproduction have gained increasing research interest. Higher Order Ambisonics (HOA) stands out for its adaptability to various playback devices and its ability to integrate head orientation. However, current HOA recordings often rely on bulky spherical microphone arrays (SMA), and portable d… ▽ More

    Submitted 3 June, 2026; originally announced June 2026.

    Comments: Accepted for presentation at AES Europe 2025 Convention (AES 158th Convention), Warsaw, Poland, May 22-24, 2025

  25. arXiv:2606.04570  [pdf, ps, other

    cs.SD

    Flow-HOA: Generative Joint Optimization for Ambisonics Encoding via Flow Matching

    Authors: Yuhuan You, Yufan Qian, Tianshu Qu, Bin Wang, Xueyang Lv

    Abstract: Higher-Order Ambisonics (HOA) encoding from sparse, irregular microphone arrays remains a critical challenge for consumer spatial audio capture in immersive communication and XR. We propose Flow-HOA, a generative framework that jointly optimizes a multi-dimensional objective encompassing time-domain, spectral, and spatial fidelity while producing a deployable, time-invariant bank of Finite Impulse… ▽ More

    Submitted 3 June, 2026; originally announced June 2026.

    Comments: Accepted for presentation at AES Europe 2026 Convention (AES 160th Convention), Copenhagen, Denmark, May 28-30, 2026

  26. Beyond Independent Manipulation: Individual Fairness-aware Strategic Classification with Peer Imitation

    Authors: Xinpeng Lv, Chunyuan Zheng, Yunxin Mao, Renzhe Xu, Jinxuan Yang, Yuanlong Chen, Wangrong Huang, Shaowu Yang, Wenjing Yang, Xinwang Liu, Peng Cui, Haotian Wang

    Abstract: Strategic classification (SC) investigates scenarios where agents manipulate their features to obtain favorable decisions from predictive models. Existing fairness-aware SC approaches primarily focus on group fairness and typically assume that agents respond independently. However, when individual fairness is required, ensuring similar individuals receive similar outcomes, agents' manipulation bec… ▽ More

    Submitted 25 June, 2026; v1 submitted 30 May, 2026; originally announced June 2026.

    Comments: Accepted by SIGKDD2026

  27. Partial Fairness Awareness: Belief-Guided Strategic Mechanism for Strategic Agents

    Authors: Xinpeng Lv, Chunyuan Zheng, Yunxin Mao, Renzhe Xu, Hao Zou, Shanzhi Gu, Liyang Xu, Huan Chen, Yuanlong Chen, Wenjing Yang, Haotian Wang

    Abstract: Strategic machine learning investigates scenarios where agents manipulate their features to receive favorable decisions from predictive models. To address fairness concerns intrinsic to strategic classification, recent work has introduced group-specific fairness constraints. However, current fairness-aware approaches face a fundamental dilemma in the issue of fairness exposure: making these constr… ▽ More

    Submitted 30 May, 2026; originally announced June 2026.

    Comments: Accepted by AAAI2026

  28. arXiv:2605.21924  [pdf, ps, other

    cs.CV

    Visual-Advantage On-Policy Distillation for Vision-Language Models

    Authors: Ruiqi Liu, Xiaolei Lv, Gengsheng Li, Ximo Zhu, Zhiheng Wang, Zhengbo Zhang, Junkai Chen, Zhiheng Li, Bo Li, Jun Gao, Shu Wu

    Abstract: On-policy knowledge distillation has proven effective for language models, yet its application to vision-language models (VLMs) remains underexplored. We observe that standard on-policy distillation can improve a student's output quality while failing to strengthen its reliance on visual input: on vision-critical tokens, the student's predictions remain largely unchanged whether or not fine-graine… ▽ More

    Submitted 20 May, 2026; originally announced May 2026.

  29. arXiv:2605.19674  [pdf, ps, other

    cs.AI

    Beyond Rational Illusion: Behaviorally Realistic Strategic Classification

    Authors: Xinpeng Lv, Yunxin Mao, Renzhe Xu, Chunyuan Zheng, Yikai Chen, Haoxuan Li, Yang Shi, Jinxuan Yang, Zhouchen Lin, Yuanlong Chen, Yuanxing Zhang, Shaowu Yang, Wenjing Yang, Haotian Wang

    Abstract: Strategic classification(SC) studies the interaction between decision models and agents who strategically manipulate their features for favorable outcomes. Existing SC frameworks typically rely on the idealized assumption that agents are strictly rational. However, evidence from behavioral economics and psychology consistently shows that real-world decision-making is often shaped by cognitive bias… ▽ More

    Submitted 6 June, 2026; v1 submitted 19 May, 2026; originally announced May 2026.

    Comments: Accepted by ICML2026

  30. arXiv:2605.19662  [pdf, ps, other

    cs.AI

    When Tabular Foundation Models Meet Strategic Tabular Data: A Prior Alignment Approach

    Authors: Xinpeng Lv, Yunxin Mao, Renzhe Xu, Chunyuan Zheng, Yikai Chen, Haoxuan Li, Jinxuan Yang, Kun Kuang, Yuanlong Chen, Mingyang Geng, Wanrong Huang, Shixuan Liu, Shaowu Yang, Wenjing Yang, Zhouchen Lin, Haotian Wang

    Abstract: Tabular foundation models based on pretrained prior-data fitted networks~(PFNs) have shown strong generalization on diverse tabular tasks, but they are typically designed for \emph{non-strategic} settings where data distributions are independent of deployed classifiers. In many real-world decision scenarios, however, individuals may strategically modify their features after deployment to obtain fa… ▽ More

    Submitted 6 June, 2026; v1 submitted 19 May, 2026; originally announced May 2026.

    Comments: Accepted by ICML2026

  31. arXiv:2605.18643  [pdf, ps, other

    cs.LG cs.AI cs.CL

    Post-Trained MoE Can Skip Half Experts via Self-Distillation

    Authors: Xingtai Lv, Li Sheng, Kaiyan Zhang, Yichen You, Siyan Gao, Xueheng Luo, Yuxin Zuo, Yuchen Fan, Junlin Yang, Ganqu Cui, Bingning Wang, Fan Yang, Youbang Sun, Ning Ding, Bowen Zhou

    Abstract: Mixture-of-Experts (MoE) scales language models efficiently through sparse expert activation, and its dynamic variant further reduces computation by adjusting the activated experts in an input-dependent manner. Existing dynamic MoE methods usually rely on pre-training from scratch or task-specific adaptation, leaving the practical conversion of fully trained MoE underexplored. Enabling such adapta… ▽ More

    Submitted 7 June, 2026; v1 submitted 18 May, 2026; originally announced May 2026.

  32. arXiv:2605.17863  [pdf, ps, other

    cs.IR

    DADF: A Distribution-Aware Debiasing Framework for Watch-Time Regression in Recommender Systems

    Authors: Yiqing Yang, Xinlong Zhao, Zhao Liu, Xiao Lv, Ruiming Tang, Han Li, Kun Gai

    Abstract: Watch-time predictors in short-video recommender systems can be approximately calibrated by their own scores while still overestimating short observations and underestimating long ones. We study whether this label-space mean shrinkage contains inference-time-predictable residual structure that can be corrected without replacing a mature first-stage model. We propose DADF, a distribution-aware seco… ▽ More

    Submitted 31 July, 2026; v1 submitted 18 May, 2026; originally announced May 2026.

    Comments: 11 pages, 7 figures, 3 tables

  33. arXiv:2605.14478  [pdf, ps, other

    cs.SE cs.AI cs.CL

    When Retrieval Hurts Code Completion: A Diagnostic Study of Stale Repository Context

    Authors: Haojun Weng, Qianqian Yang, Hao Fu, Haobin Pan, Xinwei Lv

    Abstract: Context: Retrieval-augmented code generation relies on cross-file repository context, but retrieved snippets may come from obsolete project states. Objectives: We study whether temporally stale repository snippets act as harmless noise or actively induce current-state-incompatible code. Methods: We conduct a controlled diagnostic study on a curated 17-sample set of production-helper signature… ▽ More

    Submitted 14 May, 2026; originally announced May 2026.

    Comments: 31 pages, 2 tables. Submitted to Information and Software Technology (Elsevier)

    ACM Class: D.2.5; D.2.7; I.2.7

  34. arXiv:2605.07810  [pdf, ps, other

    physics.optics cs.CV

    Pre-training Enables Extraordinary All-optical Image Denoising

    Authors: Xudong Lv, Yuxiang Sun, Shuo Wang, Nanxing Chen, Jun Guan, Jingtian Hu

    Abstract: Optical neural networks are emerging as powerful machine learning and information processing tools because of their potential advantages in speed and energy efficiency. The training methods of these physical models, however, remain underexplored compared to their digital counterparts and are leading to suboptimal performance. This paper reports a pre-training-driven approach that leads to snapshot… ▽ More

    Submitted 8 May, 2026; originally announced May 2026.

  35. arXiv:2605.07492  [pdf, ps, other

    cs.CV

    How Far Is Document Parsing from Solved? PureDocBench: A Source-TraceableBenchmark across Clean, Degraded, and Real-World Settings

    Authors: Zhiheng Li, Zongyang Ma, Jiaxian Chen, Jianing Zhang, Zhaolong Su, Yutong Zhang, Zhiyin Yu, Ruiqi Liu, Xiaolei Lv, Bo Li, Jun Gao, Ziqi Zhang, Chunfeng Yuan, Bing Li, Weiming Hu

    Abstract: The past year has seen over 20 open-source document parsing models, yet thefield still benchmarks almost exclusively on OmniDocBench, a 1,355-pagemanually annotated dataset whose top scores have saturated above 90%. Athree-stage audit pipeline we run on OmniDocBench screens its 21,353evaluator-scored blocks and confirms 2,580 errors (12.08%); combined with overa year of public availability, both a… ▽ More

    Submitted 8 May, 2026; originally announced May 2026.

    Comments: 42 pages, 20 figures, 16 tables

  36. arXiv:2604.28009  [pdf, ps, other

    quant-ph

    Learning quantum disentanglement scheduling from reduced states via modular hybrid policies

    Authors: Y. -X. Xiao, J. -Z. Han, Z. Zheng, Z. -H. Zhang, M. Xue, J. Li, X. Lv

    Abstract: Quantum control with restricted state access is central to near-term quantum devices, where full wave-function information is unavailable. We study this problem through multiqubit disentanglement scheduling from partial observations, where a controller receives only two-qubit reduced density matrices and selects which qubit pair to disentangle at each step. We introduce a modular hybrid quantum--c… ▽ More

    Submitted 30 April, 2026; originally announced April 2026.

    Comments: 13 pages, 11 figures; Preliminary manuscript prepared for summer program applications

  37. arXiv:2604.23783  [pdf, ps, other

    cs.IR cs.AI

    S2G-RAG: Structured Sufficiency and Gap Judging for Iterative Retrieval-Augmented QA

    Authors: Minghan Li, Junjie Zou, Xinxuan Lv, Chao Zhang, Guodong Zhou

    Abstract: Retrieval-Augmented Generation (RAG) grounds language models in external evidence, but multi-hop question answering remains difficult because iterative pipelines must control what to retrieve next and when the available evidence is adequate. In practice, systems may answer from incomplete evidence chains, or they may accumulate redundant or distractor-heavy text that interferes with later retrieva… ▽ More

    Submitted 26 April, 2026; originally announced April 2026.

    Comments: Accepted to ACL 2026 Main Conference

  38. arXiv:2604.22321  [pdf

    physics.optics

    Gate- and Optically Controlled Nonlinear Optical Response in Graphene via Non-Perturbative Ultrafast Carrier Dynamics

    Authors: Xiaolong Lv, Yu Zhang, Yuxuan Wei, Chuanshan Tian

    Abstract: While the Dirac band structure of graphene has established it as a leading platform for ultrafast optoelectronics, its non-perturbative nonlinear response under intense excitation remains poorly understood. Here, we report ultrafast spectral modulation of nonlinear optical signals in graphene. By utilizing a robust suspended-graphene platform that allows for both wide-range electrostatic gating an… ▽ More

    Submitted 24 April, 2026; originally announced April 2026.

  39. arXiv:2604.07993  [pdf, ps, other

    cs.RO

    HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation

    Authors: Shuanghao Bai, Meng Li, Xinyuan Lv, Jiawei Wang, Xinhua Wang, Fei Liao, Chengkai Hou, Langzhe Gu, Wanqi Zhou, Kun Wu, Ziluo Ding, Zhiyuan Xu, Lei Sun, Shanghang Zhang, Zhengping Che, Jian Tang, Badong Chen

    Abstract: Humans achieve complex manipulation through coordinated whole-body control, whereas most Vision-Language-Action (VLA) models treat robot body parts largely independently, making high-DoF humanoid control challenging and often unstable. We present HEX, a state-centric framework for coordinated manipulation on full-sized bipedal humanoid robots. HEX introduces a humanoid-aligned universal state repr… ▽ More

    Submitted 19 May, 2026; v1 submitted 9 April, 2026; originally announced April 2026.

    Comments: Project page: https://hex-humanoid.github.io/

  40. arXiv:2604.06950  [pdf, ps, other

    cs.CV

    Making MLLMs Blind: Adversarial Smuggling Attacks in MLLM Content Moderation

    Authors: Zhiheng Li, Zongyang Ma, Yuntong Pan, Ziqi Zhang, Xiaolei Lv, Bo Li, Jun Gao, Jianing Zhang, Chunfeng Yuan, Bing Li, Weiming Hu

    Abstract: Multimodal Large Language Models (MLLMs) are increasingly being deployed as automated content moderators. Within this landscape, we uncover a critical threat: Adversarial Smuggling Attacks. Unlike adversarial perturbations (for misclassification) and adversarial jailbreaks (for harmful output generation), adversarial smuggling exploits the Human-AI capability gap. It encodes harmful content into h… ▽ More

    Submitted 8 April, 2026; v1 submitted 8 April, 2026; originally announced April 2026.

    Comments: Accepted to ACL 2026. 19 pages, 6 figures

  41. arXiv:2603.29211  [pdf, ps, other

    cs.AI cs.CL cs.CV

    Xuanwu: Evolving General Multimodal Models into an Industrial-Grade Foundation for Content Ecosystems

    Authors: Zhiqian Zhang, Xu Zhao, Xiaoqing Xu, Guangdong Liang, Weijia Wang, Xiaolei Lv, Bo Li, Jun Gao

    Abstract: In recent years, multimodal large models have continued to improve on general benchmarks. However, in real-world content moderation and adversarial settings, mainstream models still suffer from degraded generalization and catastrophic forgetting because of limited fine-grained visual perception and insufficient modeling of long-tail noise. In this paper, we present Xuanwu VL-2B as a case study of… ▽ More

    Submitted 30 March, 2026; originally announced March 2026.

    Comments: 41 pages, 10 figures

  42. arXiv:2603.28421  [pdf, ps, other

    quant-ph cs.AI

    Learning Unified Control of Intrinsic Nonlinear Spin Dynamics in Atomic Qudits for Magnetometry

    Authors: C. Z. Cao, J. Z. Han, M. Xiong, M. Deng, L. Wang, X. Lv, M. Xue

    Abstract: Generating and preserving metrologically useful quantum states is a central challenge in quantum-enhanced metrology. In low-field atomic magnetometry with multilevel atoms, the nonlinear Zeeman (NLZ) effect is both a resource and a limitation. It can generate internal spin squeezing within a single atomic qudit, but under fixed readout it also rotates and distorts the measurement-relevant quadratu… ▽ More

    Submitted 28 April, 2026; v1 submitted 30 March, 2026; originally announced March 2026.

    Comments: (6+3+2.5) pages, (4+2) figures, 1 table

  43. arXiv:2603.25080  [pdf, ps, other

    physics.ins-det astro-ph.IM

    A Telescope System for Charge and Position Measurement of High Energy Nuclei

    Authors: Dexing Miao, Zhiyu Xiang, Giovanni Ambrosi, Mattia Barbanera, Baasansuren Batsukh, Mengke Cai, Xudong Cai, Yuan-Hann Chang, Shanzhen Chen, Hsin-Yi Chou, Xingzhu Cui, Mingyi Dong, Matteo Duranti, Ke Gong, Mingjie Feng, Valerio Formato, Daojin Hong, Maria Ionica, Xiaojie Jiang, Yaozu Jiang, Liangchenglong Jin, Shengjie Jin, Vladimir Koutsenko, Tiange Li, Zuhao Li , et al. (21 additional authors not shown)

    Abstract: A high-granularity telescope system with a large sensitive area and low material budget has been developed for high-energy heavy ion beam tests. The telescope consists of nine layers of silicon microstrip detectors (SSDs), whose performance was validated through a heavy ion beam test at the CERN SPS. A hybrid machine learning algorithm is proposed to address the challenges of nuclear charge measur… ▽ More

    Submitted 26 March, 2026; originally announced March 2026.

  44. arXiv:2603.17794  [pdf, ps, other

    hep-th cond-mat.soft hep-ph nucl-th physics.flu-dyn

    Hydrodynamics of dilation and spin currents

    Authors: Zhong-Hua Zhang, Xi-Hu Lv, Xu-Guang Huang

    Abstract: We formulate a relativistic hydrodynamic theory for fluids with spin and intrinsic dilation charges. Using an entropy-current analysis, we derive constitutive relations featuring a bulk viscosity and a dilation conductivity governing the relaxation and diffusion of dilation charge. Linear mode analysis reveals a gapped dilation excitation and the freeze-out of long-wavelength sound modes, similar… ▽ More

    Submitted 18 March, 2026; originally announced March 2026.

    Comments: 7 pages, 1 figure

  45. arXiv:2603.12201  [pdf, ps, other

    cs.CL cs.LG

    IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse

    Authors: Yushi Bai, Qian Dong, Ting Jiang, Xin Lv, Zhengxiao Du, Aohan Zeng, Jie Tang, Juanzi Li

    Abstract: Long-context agentic workflows have emerged as a defining use case for large language models, making attention efficiency critical for both inference speed and serving cost. Sparse attention addresses this challenge effectively, and DeepSeek Sparse Attention (DSA) is a representative production-grade solution: a lightweight lightning indexer selects the top-k most relevant tokens per query, reduci… ▽ More

    Submitted 12 March, 2026; originally announced March 2026.

  46. arXiv:2603.08660  [pdf, ps, other

    cs.LG cs.CL

    How Far Can Unsupervised RLVR Scale LLM Training?

    Authors: Bingxiang He, Yuxin Zuo, Zeyuan Liu, Shangziqi Zhao, Zixuan Fu, Junlin Yang, Cheng Qian, Kaiyan Zhang, Yuchen Fan, Ganqu Cui, Xiusi Chen, Youbang Sun, Xingtai Lv, Xuekai Zhu, Li Sheng, Ran Li, Huan-ang Gao, Yuchen Zhang, Bowen Zhou, Zhiyuan Liu, Ning Ding

    Abstract: Unsupervised reinforcement learning with verifiable rewards (URLVR) offers a pathway to scale LLM training beyond the supervision bottleneck by deriving rewards without ground truth labels. Recent works leverage model intrinsic signals, showing promising early gains, yet their potential and limitations remain unclear. In this work, we revisit URLVR and provide a comprehensive analysis spanning tax… ▽ More

    Submitted 9 March, 2026; originally announced March 2026.

    Comments: Accepted to the ICLR 2026

  47. arXiv:2603.01955  [pdf, ps, other

    astro-ph.GA

    Large scale mapping of [CI] and the [CI]-to-CO transition in $ρ$ Ophiuchus molecular cloud

    Authors: Jifeng Xia, Ningyu Tang, Thomas G. Bisbas, Chen Wang, Gan Luo, Sihan Jiao, Xin Lv, Xuejian Jiang, Donghui Quan, Jinzeng Li, Paul F. Goldsmith, Gary A. Fuller, Di Li

    Abstract: Atomic carbon ([CI]) is a key species in the carbon chemistry of the interstellar medium (ISM). Using the Submillimeter Wave Astronomy Satellite (SWAS), we conducted a [CI]($^3$P$_1$--$^3$P$_0$) 492 GHz survey covering approximately 4 deg$^2$ of the L1688 and L1689 regions in the $ρ$ Oph molecular cloud, achieving a spatial resolution of 4.25$\hbox{$^{\prime}$}$. The derived [CI] column densities,… ▽ More

    Submitted 2 March, 2026; originally announced March 2026.

    Comments: 16 pages, 12 figures. Accepted for publication in Science China Physics, Mechanics & Astronomy

  48. arXiv:2603.01853  [pdf, ps, other

    cs.CL

    Let the Agent Search: Autonomous Exploration Beats Rigid Workflows in Temporal Question Answering

    Authors: Xufei Lv, Jiahui Yang, Haoyuan Sun, Xialin Su, Zhiliang Tian, Yifu Gao, Linbo Qiao, Houde Liu

    Abstract: Temporal Knowledge Graph Question Answering (TKGQA) is challenging because it requires multi-hop reasoning under complex temporal constraints. Recent LLM-based approaches have improved semantic modeling for this task, but many still rely on fixed reasoning workflows or costly post-training, which can limit adaptability and make error recovery difficult. We show that enabling an off-the-shelf Large… ▽ More

    Submitted 25 March, 2026; v1 submitted 2 March, 2026; originally announced March 2026.

    Comments: Revised version with three added authors and additional experiments

  49. arXiv:2602.20616  [pdf, ps, other

    cs.CV cs.LG

    Knowing the Unknown: Interpretable Open-World Object Detection via Concept Decomposition Model

    Authors: Xueqiang Lv, Shizhou Zhang, Yinghui Xing, Di Xu, Peng Wang, Yanning Zhang

    Abstract: Open-world object detection (OWOD) requires incrementally detecting known categories while reliably identifying unknown objects. Existing methods primarily focus on improving unknown recall, yet overlook interpretability, often leading to known-unknown confusion and reduced prediction reliability. This paper aims to make the entire OWOD framework interpretable, enabling the detector to truly "know… ▽ More

    Submitted 24 February, 2026; originally announced February 2026.

  50. arXiv:2602.15763  [pdf, ps, other

    cs.LG cs.CL

    GLM-5: from Vibe Coding to Agentic Engineering

    Authors: GLM-5-Team, :, Aohan Zeng, Xin Lv, Zhenyu Hou, Zhengxiao Du, Qinkai Zheng, Bin Chen, Da Yin, Chendi Ge, Chenghua Huang, Chengxing Xie, Chenzheng Zhu, Congfeng Yin, Cunxiang Wang, Gengzheng Pan, Hao Zeng, Haoke Zhang, Haoran Wang, Huilong Chen, Jiajie Zhang, Jian Jiao, Jiaqi Guo, Jingsen Wang, Jingzhao Du , et al. (162 additional authors not shown)

    Abstract: We present GLM-5, a next-generation foundation model designed to transition the paradigm of vibe coding to agentic engineering. Building upon the agentic, reasoning, and coding (ARC) capabilities of its predecessor, GLM-5 adopts DSA to significantly reduce training and inference costs while maintaining long-context fidelity. To advance model alignment and autonomy, we implement a new asynchronous… ▽ More

    Submitted 24 February, 2026; v1 submitted 17 February, 2026; originally announced February 2026.