Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 286 results for author: Cui, X

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.16491  [pdf, ps, other

    cs.DB cs.IR

    FROG: Efficient Range-Filtering Approximate Nearest Neighbor Search on GPUs

    Authors: Xiaokun Cui, Pengbo Liu, Jiadong Xie, Yingfan Liu, Hui Li, Jeffrey Xu Yu, Jiangtao Cui

    Abstract: Range-filtering approximate nearest neighbor search (RFANNS) is a fundamental operation in modern vector databases. Given a query vector $q$ and a numerical range predicate, RFANNS returns the $k$-approximate nearest neighbors ($k$-ANN) of the query $q$ among the objects whose attributes satisfy the range predicate. However, existing RFANNS methods are not well suited to high-throughput GPU execut… ▽ More

    Submitted 17 August, 2026; originally announced August 2026.

  2. arXiv:2608.15419  [pdf, ps, other

    cs.CV

    ArtLang: Structured Language-to-Kinematics Grounding for Articulated 3D Actuation

    Authors: Sylvia Yuan, Dan Wang, Ravi Ramamoorthi, Xinrui Cui

    Abstract: Articulated-object reconstructions recover explicit geometry and kinematics, but their parts often remain semantically anonymous and must be controlled through part indices and numerical joint parameters. We present ArtLang, a framework for open-vocabulary language control of persistent reconstructed articulated assets. ArtLang represents an asset as a semantic-kinematic articulation graph and aug… ▽ More

    Submitted 15 August, 2026; originally announced August 2026.

  3. arXiv:2608.13499  [pdf, ps, other

    cs.DC

    OpScale: Operator-level Provisioning and Autoscaling for LLM Serving

    Authors: Xingqi Cui, Chieh-Jan Mike Liang, Ziang Tang, Jiarong Xing, Haoran Qiu

    Abstract: Achieving cost efficiency while meeting strict user-facing SLOs (e.g., time-to-first-token) remains a fundamental challenge for cloud GPU clusters serving large language models (LLMs). Autoscaling is the key mechanism for cluster resource management, yet a basic system design question is open for serving LLMs: what should be the unit of scaling? Existing approaches primarily treat the entire model… ▽ More

    Submitted 13 August, 2026; originally announced August 2026.

    Comments: 22 pages, 29 figures

  4. arXiv:2608.07570  [pdf, ps, other

    cs.CV cs.AI

    COMEX: A Composition-Grounded Benchmark and Learning Framework for Explainable Aesthetic Image Cropping

    Authors: Rui Yang, Wei Zhou, Dingyong Gou, Xiaohui Cui, Cong Li, Yinyin Gong, Yipo Huang, Jiliang Zhao

    Abstract: Explainable aesthetic image cropping requires not only localizing a visually pleasing crop but also explaining why it is preferred. Existing crop-and-explain methods largely treat explanation as post-hoc text generation and overlook composition, a key aesthetic factor that links crop decisions with interpretable reasoning. In this paper, we reformulate explainable aesthetic image cropping as a str… ▽ More

    Submitted 3 August, 2026; originally announced August 2026.

  5. arXiv:2608.03554  [pdf, ps, other

    cs.CR cs.DC

    ReputationChain: Robust Trust Updating for Blockchain-Enabled Supply Chains

    Authors: Adnan Iftekhar, Chengliang Zheng, Xiaohui Cui, Mir Hassan

    Abstract: Blockchain can preserve supply-chain records, but ledger integrity alone does not show whether a participant should be trusted in a future risk-sensitive transaction. Existing reputation systems mainly address product evidence, global feedback aggregation, or review authenticity, while giving less attention to repeated bilateral inflation, identity multiplicity, and unfair decay for honest partici… ▽ More

    Submitted 4 August, 2026; originally announced August 2026.

  6. arXiv:2607.28638  [pdf, ps, other

    cs.CL cs.LG

    Learning Stateful Predictive Knowledge From Experience

    Authors: Yan Song, Xidong Feng, Bo Liu, Xinyu Cui, Haotian Fu, Zichen Liu, Mengyue Yang, Cheng Deng, Jian Zhao, Jun Wang

    Abstract: As large language model (LLM) agents increasingly learn from experience, they primarily rely on trajectory-level reflection to extract insights. Viewed through the lens of predictive knowledge, we argue that this approach operates on episodic hindsight rather than predictive foresight, yielding brittle, path-dependent heuristics. To address this, we propose Stateful Knowledge Learning (SKL). SKL s… ▽ More

    Submitted 19 May, 2026; originally announced July 2026.

  7. arXiv:2607.12376  [pdf

    cs.CV cs.AI

    Demonstration of the common dual-channel feature decoupling characteristic of front-door mediation causal inference methods in whole-slice image classification

    Authors: Zhirui Zhang, Tianhang Nan, Yong Ding, Zhuolun Song, Dayu Hu, Xiaoyu Cui

    Abstract: Causal inference using front door intervention and multi-instance learning (MIL) has advanced the analysis of Whole Slide Images (WSI) in digital pathology. These methods adjust feature distributions of subtle evidence sub-images to correctly associate them with WSI-level diagnoses. We propose and prove 2 hypotheses for evaluating such methods: 1) Causal inference MIL introduces an independent cla… ▽ More

    Submitted 14 July, 2026; originally announced July 2026.

    Comments: The manuscript is being submitted for publication to a journal

  8. arXiv:2607.08493  [pdf, ps, other

    cs.LG cs.CL

    Ensemble Diversity Optimization for Subjective Supervision

    Authors: Xia Cui, Ziyi Huang, N. R. Abeynayake

    Abstract: Subjective NLP tasks often exhibit systematic annotator disagreement, requiring models that represent uncertainty rather than collapse it. We introduce Ensemble Diversity Optimization (EDO), a prediction-space framework that jointly optimizes ensemble weights, effective cardinality, and calibration through a unified differentiable objective. EDO learns ensemble composition and size end-to-end via… ▽ More

    Submitted 9 July, 2026; originally announced July 2026.

  9. arXiv:2607.05861  [pdf, ps, other

    cs.CL cs.LG

    Mitigating Factual Hallucination in Large Reasoning Models via Mixed-Mode Advantage Regularization

    Authors: Kaishen Wang, Tong Zheng, Xuehao Cui, Ruibo Chen, Tianyi Xiong, Heng Huang

    Abstract: Large reasoning models (LRMs) improve language model capabilities by generating explicit thinking traces before final answers. In factuality-oriented question answering (QA), such thinking often improves overall performance by helping the model recover relevant knowledge and refine its answers. However, we find that this benefit is not uniform at the instance level: explicit thinking can also over… ▽ More

    Submitted 7 July, 2026; originally announced July 2026.

    Comments: 19 pages, 3 figures, 8 tables

  10. arXiv:2607.04814  [pdf, ps, other

    cs.CL cs.AI

    Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition

    Authors: Andrei Florian, Cynthia Jayne Amol, Hope Kerubo Ombaba, Xiaoyu Cui, Boniface Mwau, Biatus Maina Kamau, Lilian Diana Awuor Wanzare, Christiane Fellbaum, Happy Buzaaba

    Abstract: Extending automatic speech recognition (ASR) to low-resource African languages is constrained by the prohibitive demands of data collection at scale. A promising direction is to leverage the linguistic relatedness between a low-resource target language and languages previously seen by a model to reduce the volume of target-language data needed for effective adaptation. Although this approach has p… ▽ More

    Submitted 27 July, 2026; v1 submitted 6 July, 2026; originally announced July 2026.

  11. arXiv:2607.01810  [pdf, ps, other

    cs.SE cs.AI

    Decoupling Code Complexity from Newcomer Participation: A Causal Study of AI Coding Agent Adoption in OSS

    Authors: Weiwei Xu, Xuanning Cui, Hengzhi Ye, Minghui Zhou

    Abstract: Open-source projects depend on a steady inflow of newcomers. A growing concern is that AI coding agents (tools such as Cursor and Claude Code that write code from natural-language instructions) will crowd them out, by absorbing the simple tasks that beginners start with and by making code harder to read. We give this concern a causal answer. Using GitHub code search we identify 1,888 projects that… ▽ More

    Submitted 2 July, 2026; originally announced July 2026.

  12. arXiv:2606.27067  [pdf

    cs.HC

    Floor Raiser or Ceiling Limiter? Differential Storytelling Outcomes with a Child-Centric GenAI System Across Individual Differences

    Authors: Min Fan, Wanqing Ma, Xinyue Cui, Xiaolu Dai, Shengyu Huang

    Abstract: Generative AI (GenAI) holds promise for democratizing creative literacy, yet whether it benefits all children equally remains unclear. Using a child-centric GenAI storytelling system for children aged 7-12, we conducted a mixed-methods within-subjects experiment (N = 40, Grades 2-6) comparing GenAI-assisted and traditional storyboard conditions. Three findings emerged. First, the GenAI-assisted co… ▽ More

    Submitted 25 June, 2026; originally announced June 2026.

  13. arXiv:2606.25224  [pdf, ps, other

    cs.RO

    Spatio-Temporal Retrieval-based Priors for Adaptive Computational Teaching in Driving

    Authors: Deepak Edakkattil Gopinath, Xiongyi Cui, Jonathan DeCastro, Avinash Balachandran, Guy Rosman

    Abstract: Learning-based automated coaching systems for complex motor tasks such as high-performance driving remain limited in the ability to be adaptive by their reliance only on local, context-dependent reasoning, failing to account for the long-term temporal nature of student learning and the cumulative impact of repeated teacher-student interactions. In this paper, we propose an imitation learning based… ▽ More

    Submitted 23 June, 2026; originally announced June 2026.

    Comments: 20 pages, 8 figures

  14. arXiv:2606.18680  [pdf, ps, other

    cs.RO

    High-Degree-of-Freedom Lightweight Bioinspired Leg for Enhanced Mobility in Small Robots

    Authors: Haoqi Han, Yifei Yu, Jiaming Zhang, Xinru Cui, Linxi Feng, Hesheng Wang

    Abstract: In microrobotics, enhancing locomotion capabilities by increasing the degrees of freedom (DoF) of leg mechanisms under severe spatial constraints remains a significant challenge. Inspired by insect locomotion, this paper presents a novel micro-scale parallel leg mechanism with four degrees of freedom, and systematically analyzes its mechanical design, electrical system, and kinematics. The design… ▽ More

    Submitted 17 June, 2026; originally announced June 2026.

    Journal ref: 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026)

  15. arXiv:2606.18661  [pdf, ps, other

    cs.CV cs.AI

    LandslideAgent with Multimodal LandslideBench: A Domain-Rule-Augmented Agent for Autonomous Landslide Identification and Analysis

    Authors: Chengfu Liu, Dongyang Hou, Junwu Xiang, Cheng Yang, Xuezhi Cui, Zeyuan Wang, Liangtian Liu, Zelang Miao

    Abstract: Intelligent landslide hazard interpretation is critical for disaster prevention, yet current paradigms struggle to simultaneously extract visual features and high-level geoscientific semantics, while general-purpose vision-language models (VLMs) suffer from perceptual limitations and domain hallucinations in complex geological scenarios. To address these challenges, we propose an instruction-drive… ▽ More

    Submitted 17 June, 2026; originally announced June 2026.

  16. arXiv:2606.11251  [pdf, ps, other

    cs.LG

    Mechanical Field Networks: Structured Neural Dynamics for Multivariate Systems

    Authors: Xingji Cui

    Abstract: Many multivariate dynamical systems are observed only through trajectories, leaving the mechanisms governing their joint dynamics hidden. Existing approaches can impose interpretable dynamics or learn flexible state transitions, yet the resulting interaction structure is typically either specified in advance or left implicit within the learned dynamics. We introduce MF-Net, a recurrent dynamical m… ▽ More

    Submitted 8 June, 2026; originally announced June 2026.

  17. arXiv:2606.10718  [pdf, ps, other

    cs.LG cs.AI

    Transformer Based Model for Spatiotemporal Feature Learning in EEG Emotion Recognition

    Authors: Xinglong Cui, Dian Gu

    Abstract: Electroencephalography (EEG) is a widely adopted technique for monitoring brain activity, offering valuable insights into neurological states due to its high temporal resolution and cost-effectiveness. To enhance the analysis of complex EEG data, we propose EEG-TransNet, an architecture designed to capture temporal, regional, and synchronous features of EEG signals. EEG-TransNet introduces three k… ▽ More

    Submitted 9 June, 2026; originally announced June 2026.

  18. arXiv:2606.07915  [pdf, ps, other

    cs.AI

    EditSR: Enhancing Neural Symbolic Regression via Edit-based Rectification

    Authors: Da Li, Xinxin Li, Xingyu Cui, Jin Xu, Juan Zhang, Junping Yin

    Abstract: Neural symbolic regression models improve inference efficiency by shifting structural search to pretraining, but their one-pass autoregressive decoding is prone to error accumulation, which may lead to generating structurally incorrect expressions, especially in complex expression generation scenarios. Existing rectification strategies can alleviate this issue, but they often depend on restarting… ▽ More

    Submitted 5 June, 2026; originally announced June 2026.

  19. arXiv:2606.07538  [pdf, ps, other

    cs.IR cs.AI

    Bidirectional Semantic Complementary Tool Retrieval for Remote Sensing Agents

    Authors: Zeyuan Wang, Dongyang Hou, Cheng Yang, Xuezhi Cui, Linrui Xu, Bo Yu, Gaozhi Zhou, Ziyu Li, Liangtian Liu, Kai Ouyang, Wang Guo, Lili Zhu, Chao Tao

    Abstract: Large language model (LLM)-based agents provide a novel paradigm for the automated processing of remote sensing(RS) data. Their success in complex RS tasks rely on extensive specialized tool libraries. However, tool documentation often exceeds the context window limits of LLMs, making precise tool retrieval essential for agentic workflows. Existing tool retrieval methods face "semantic asymmetry"… ▽ More

    Submitted 29 April, 2026; originally announced June 2026.

  20. arXiv:2606.05883  [pdf, ps, other

    cs.CV

    Geometry-Aware Dataset Condensation for Diffusion Model Training

    Authors: Xiao Cui, Yulei Qin, Mo Zhu, Wengang Zhou, Hongsheng Li, Houqiang Li

    Abstract: Dataset condensation aims to construct compact datasets from real data via synthesis or selection. However, existing approaches are ill-suited for diffusion model training: synthetic data generation often yields low-fidelity samples unsuitable for authentic modeling, while real subset selection typically fails to preserve the distributional geometry required by diffusion likelihood objectives. To… ▽ More

    Submitted 17 June, 2026; v1 submitted 4 June, 2026; originally announced June 2026.

    Comments: ICML 2026

  21. Rain: RDMA-assisted In-Network Scheduling for Microsecond-scale Workloads

    Authors: Zhihuang Ma, Xingming Cui, Xiaoliang Chen, Zuqing Zhu

    Abstract: Modern data center applications increasingly require microsecond-scale service time with strict tail latency requirements, which can hardly be realized with existing in-network task schedulers due to their inherent limitations. Specifically, software-based schedulers struggle to balance throughput and latency, while switch-based designs either lack global coordination, rely on packet recirculation… ▽ More

    Submitted 2 June, 2026; originally announced June 2026.

    Comments: 21 pages, 11 figures. Published in Proceedings of the ACM on Networking (PACMNET), CoNEXT2

    Journal ref: Proc. ACM Netw. 4, CoNEXT2, Article 22, June 2026, 21 pages

  22. arXiv:2605.26461  [pdf, ps, other

    cs.DC

    Characterization-Guided GPU Fault Resilience in NVIDIA MPS

    Authors: Rixin Liu, Xingqi Cui, Kaijian Wang, Xinheng Ding, Zirui Liu, Yuke Wang, Jiarong Xing

    Abstract: NVIDIA Multi-Process Service (MPS) enables fine-grained GPU sharing by allowing multiple processes to execute concurrently on the same GPU, making it an important mechanism for improving GPU utilization. However, MPS has weak fault resilience: a fault in one process can terminate all co-running processes, limiting its adoption in resilience-critical settings such as multi-tenant GPU clusters. In t… ▽ More

    Submitted 25 May, 2026; originally announced May 2026.

    Comments: 16 pages, 9 figures, 5 tables

  23. arXiv:2605.21028  [pdf, ps, other

    cs.CV cs.AI

    DySink: Dynamic Frame Sinks for Autoregressive Long Video Generation

    Authors: Bo Ye, Xinyu Cui, Jian Zhao, Tong Wei, Min-Ling Zhang

    Abstract: Autoregressive long video generation often adopts bounded-memory streaming for efficiency, typically combining local windows for short-term continuity with static early-frame sinks as long-range anchors. However, this fixed allocation keeps early frames cached even when the current visual state has substantially diverged from them, while discarding potentially more relevant intermediate history. A… ▽ More

    Submitted 31 July, 2026; v1 submitted 20 May, 2026; originally announced May 2026.

  24. arXiv:2605.19301  [pdf, ps, other

    cs.CV

    iGSP:Implicit Gradient Subspace Projection for Efficient Continual Learning of Vision-Language Models

    Authors: Xuezhi Cui, Dongbo Zhou, Wang Guo, Zeyuan Wang, Ziyu Li, Gaozhi Zhou, Xian Li, Ling Zhao, Wentao Yang, Chao Tao, Haifeng Li

    Abstract: Vision-Language Models require efficient adaptation to continually emerging downstream tasks. While Parameter-Efficient Fine-Tuning mitigates catastrophic forgetting, assigning isolated modules per task leads to parameter explosion. Conversely, recent similarity-driven sharing mechanisms falsely equate superficial visual similarity with underlying alignment consistency. This fundamental mismatch t… ▽ More

    Submitted 18 May, 2026; originally announced May 2026.

  25. arXiv:2605.19027  [pdf, ps, other

    cs.CV

    MedFM-Robust: Benchmarking Robustness of Medical Foundation Models

    Authors: Xiangxiang Cui, Tianjin Huang, Yifang Wang, Lijie Hu, Lu Yin

    Abstract: Medical foundation models have achieved remarkable clinical performance, yet their robustness under real-world perturbations remains underexplored. We present a robustness benchmark comprising 40 perturbation types (12 base, 28 medical-specific) across eight imaging modalities, evaluating five VLMs (LLaVA-Med, MedGemma, MedGemma-1.5, Gemini-2.5-flash and GPT-4o-mini) on VQA, visual grounding, and… ▽ More

    Submitted 22 May, 2026; v1 submitted 18 May, 2026; originally announced May 2026.

    Comments: MICCAI2026

  26. arXiv:2605.18601  [pdf, ps, other

    cs.CV

    Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models

    Authors: Shangwen Zhu, Qianyu Peng, Zhao Pu, Zhilei Shu, Xiangrui Ke, Zhaohu Xing, Zizhao Tong, Zeqing Wang, Xinyu Cui, Zian Zheng, Huangji Wang, Jian Zhao, Yeying Jin, Fan Cheng, Ruili Feng

    Abstract: Modern interactive video world models have achieved impressive visual fidelity, yet lack fine-grained multi-entity control and cross-entity, cross-world generalization. We trace this gap to the action interface: standard control protocols (e.g. animation IDs, device inputs, scene-level captions) bind action semantics to specific entities or engines at design time. We propose natural language as th… ▽ More

    Submitted 12 July, 2026; v1 submitted 18 May, 2026; originally announced May 2026.

  27. arXiv:2605.16582  [pdf, ps, other

    cs.CV

    ArtMesh: Part-Aware Articulated Mesh Fields with Motion-Consistent Dynamics

    Authors: Sylvia Yuan, Dan Wang, Ravi Ramamoorthi, Xinrui Cui

    Abstract: We present ArtMesh, a mesh-native method for reconstructing articulated objects explicitly as connected triangle meshes with per-part rigid motion from multi-view images in start and end states. Existing 3D Gaussian Splatting pipelines for articulated reconstruction inherit the unstructured point-based geometry of their splatting base, which provides no surface topology for reasoning about part bo… ▽ More

    Submitted 15 May, 2026; originally announced May 2026.

  28. arXiv:2605.10543  [pdf, ps, other

    cs.CV

    TIE: Time Interval Encoding for Video Generation over Events

    Authors: Zhilei Shu, Shangwen Zhu, Zihang Liang, Xiaofan Li, Qianyu Peng, Xinyu Cui, Bo Ye, Yiming Li, Fan Cheng, Jian Zhao, Yang Cao, Zheng-Jun Zha, Ruili Feng

    Abstract: Director-style prompting, robotic action prediction, and interactive video agents demand temporal grounding over concurrent events -- a regime in which 68% of general clips and over 99% of robotics/gameplay clips contain overlapping events, yet existing multi-event generators rest on a single-active-prompt assumption. However, modern video generators, such as Diffusion Transformers (DiT), represen… ▽ More

    Submitted 25 May, 2026; v1 submitted 11 May, 2026; originally announced May 2026.

  29. arXiv:2604.25646  [pdf, ps, other

    cs.CV cs.RO

    SAMe: A Semantic Anatomy Mapping Engine for Robotic Ultrasound

    Authors: Jing Zhang, Duojie Chen, Wentao Jiang, Zihan Lou, Jianxin Liu, Xinwu Cui, Qinghong Zhao, Bo Du, Christoph F. Dietrich, Dacheng Tao

    Abstract: Robotic ultrasound has advanced local image-driven control, contact regulation, and view optimization, yet current systems lack the anatomical understanding needed to determine what to scan, where to begin, and how to adapt to individual patient anatomy. These gaps make systems still reliant on expert intervention to initiate scanning. Here we present SAMe, a semantic anatomy mapping engine that p… ▽ More

    Submitted 18 May, 2026; v1 submitted 28 April, 2026; originally announced April 2026.

    Comments: Supplementary information included. Code will be released at https://github.com/MiliLab/Echo-SAMe

  30. arXiv:2604.20317  [pdf, ps, other

    cs.CV

    MD-Face: MoE-Enhanced Label-Free Disentangled Representation for Interactive Facial Attribute Editing

    Authors: Xuan Cui, Yunfei Zhao, Bo Liu, Wei Duan, Xingrong Fan

    Abstract: GAN-based facial attribute editing is widely used in virtual avatars and social media but often suffers from attribute entanglement, where modifying one face attribute unintentionally alters others. While supervised disentangled representation learning can address this, it relies heavily on labeled data, incurring high annotation costs. To address these challenges, we propose MD-Face, a label-free… ▽ More

    Submitted 22 April, 2026; originally announced April 2026.

  31. POLAR-PIC: A Holistic Framework for Matrixized PIC with Co-Designed Compute, Layout, and Communication

    Authors: Yizhuo Rao, Xingjian Cui, Shangzhi Pang, Jiabin Xie, Guangnan Feng, Jinhui Wei, Ziyan Zhang, Languang Gao, Zhenyu Wang, Zhiguang Chen, Yutong Lu

    Abstract: Particle-in-Cell (PIC) simulations are fundamental to plasma physics but often suffer from limited scalability due to particle-grid interaction bottlenecks and particle redistribution costs. Specifically, the particle-grid interaction computations have not taken full advantage of the emerging Matrix Processing Units (MPUs), the particle motion introduces irregular memory accesses, and the bulk-syn… ▽ More

    Submitted 21 April, 2026; originally announced April 2026.

    Comments: Accepted for publication at HPDC 2026

    Journal ref: The 35th International Symposium on High-Performance Parallel and Distributed Computing (HPDC '26), July 13--16, 2026, Cleveland, OH, USA

  32. arXiv:2604.17504  [pdf, ps, other

    cs.CV cs.AI

    RS-HyRe-R1: A Hybrid Reward Mechanism to Overcome Perceptual Inertia for Remote Sensing Images Understanding

    Authors: Gaozhi Zhou, Hu He, Peng Shen, Jipeng Zhang, Liujue Zhang, Linrui Xu, Zeyuan Wang, Ziyu Li, Xuezhi Cui, Wang Guo, Haifeng Li

    Abstract: Reinforcement learning (RL) post-training substantially improves remote sensing vision-language models (RS-VLMs). However, when handling complex remote sensing imagery (RSI) requiring exhaustive visual scanning, models tend to rely on localized salient cues for rapid inference. We term this RL-induced bias "perceptual inertia". Driven by reward maximization, models favor quick outcome fitting, lea… ▽ More

    Submitted 19 April, 2026; originally announced April 2026.

  33. arXiv:2604.16918  [pdf, ps, other

    cs.CL cs.LG

    Freshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement Learning

    Authors: Weiyu Ma, Yongcheng Zeng, Yan Song, Xinyu Cui, Jian Zhao, Xuhui Liu, Mohamed Elhoseiny

    Abstract: Reinforcement Learning (RL) has achieved impressive success in post-training Large Language Models (LLMs) and Vision-Language Models (VLMs), with on-policy algorithms such as PPO, GRPO, and REINFORCE++ serving as the dominant paradigm. However, these methods discard all collected trajectories after a single gradient update, resulting in poor sample efficiency, particularly wasteful for agentic tas… ▽ More

    Submitted 18 April, 2026; originally announced April 2026.

  34. arXiv:2604.16848  [pdf, ps, other

    cs.CV cs.AI

    TowerDataset: A Heterogeneous Benchmark for Transmission Corridor Segmentation with a Global-Local Fusion Framework

    Authors: Xu Cui, Xinyan Liu, Chen Yang, Zhaobo Qi, Beichen Zang, Weigang Zhang, Antoni B. Chan

    Abstract: Fine-grained semantic segmentation of transmission-corridor point clouds is fundamental for intelligent power-line inspection. However, current progress is limited by realistic data scarcity and the difficulty of modeling global corridor structure and local geometric details in long, heterogeneous scenes. Existing public datasets usually provide only a few coarse categories or short cropped scenes… ▽ More

    Submitted 18 April, 2026; originally announced April 2026.

  35. arXiv:2604.11230  [pdf, ps, other

    cs.CV

    NTIRE 2026 The 3rd Restore Any Image Model (RAIM) Challenge: AI Flash Portrait (Track 3)

    Authors: Ya-nan Guan, Shaonan Zhang, Hang Guo, Yawen Wang, Xinying Fan, Tianqu Zhuang, Jie Liang, Hui Zeng, Guanyi Qin, Lishen Qu, Tao Dai, Shu-Tao Xia, Lei Zhang, Radu Timofte, Bin Chen, Yuanbo Zhou, Hongwei Wang, Qinquan Gao, Tong Tong, Yanxin Qian, Lizhao You, Jingru Cong, Lei Xiong, Shuyuan Zhu, Zhi-Qiang Zhong , et al. (33 additional authors not shown)

    Abstract: In this paper, we present a comprehensive overview of the NTIRE 2026 3rd Restore Any Image Model (RAIM) challenge, with a specific focus on Track 3: AI Flash Portrait. Despite significant advancements in deep learning for image restoration, existing models still encounter substantial challenges in real-world low-light portrait scenarios. Specifically, they struggle to achieve an optimal balance am… ▽ More

    Submitted 13 April, 2026; originally announced April 2026.

    Comments: Accepted to CVPR 2026 Workshop. Includes supplementary material as ancillary file

  36. arXiv:2604.02837  [pdf, ps, other

    cs.CR cs.AI

    Towards Secure Agent Skills: Architecture, Threat Taxonomy, and Security Analysis

    Authors: Zhiyuan Li, Jingzheng Wu, Xiang Ling, Xing Cui, Tianyue Luo

    Abstract: Agent Skills is an emerging open standard that defines a modular, filesystem-based packaging format enabling LLM-based agents to acquire domain-specific expertise on demand. Despite rapid adoption across multiple agentic platforms and the emergence of large community marketplaces, the security properties of Agent Skills have not been systematically studied. This paper presents the first comprehens… ▽ More

    Submitted 3 April, 2026; originally announced April 2026.

  37. arXiv:2603.22951  [pdf, ps, other

    cs.LG

    Weak-PDE-Net: Discovering Open-Form PDEs via Differentiable Symbolic Networks and Weak Formulation

    Authors: Xinxin Li, Xingyu Cui, Jin Qi, Juan Zhang, Da Li, Junping Yin

    Abstract: Discovering governing Partial Differential Equations (PDEs) from sparse and noisy data is a challenging issue in data-driven scientific computing. Conventional sparse regression methods often suffer from two major limitations: (i) the instability of numerical differentiation under sparse and noisy data, and (ii) the restricted flexibility of a pre-defined candidate library. We propose Weak-PDE-Net… ▽ More

    Submitted 24 March, 2026; originally announced March 2026.

  38. arXiv:2603.21152  [pdf, ps, other

    physics.geo-ph cs.AI

    TRACE: A Multi-Agent System for Autonomous Physical Reasoning for Seismology

    Authors: Feng Liu, Jian Xu, Xin Cui, Xinghao Wang, Zijie Guo, Jiong Wang, S. Mostafa Mousavi, Xinyu Gu, Hao Chen, Ben Fei, Lihua Fang, Fenghua Ling, Zefeng Li, Lei Bai

    Abstract: Inferring physical mechanisms that govern earthquake sequences from geophysical observations remains a challenging task, particularly across tectonically distinct environments where similar seismic patterns can reflect different underlying processes. Current seismological processing and interpretation rely heavily on experts' choice of parameters and the synthesis of various seismological products… ▽ More

    Submitted 25 March, 2026; v1 submitted 22 March, 2026; originally announced March 2026.

    Comments: 25 pages for main text and 164 pages for appendices

  39. arXiv:2603.18625  [pdf, ps, other

    cs.CV

    GenVideoLens: Where LVLMs Fall Short in AI-Generated Video Detection?

    Authors: Yueying Zou, Pei Pei Li, Zekun Li, Xinyu Guo, Xing Cui, Huaibo Huang, Ran He

    Abstract: In recent years, AI-generated videos have become increasingly realistic and sophisticated. Meanwhile, Large Vision-Language Models (LVLMs) have shown strong potential for detecting such content. However, existing evaluation protocols largely treat the task as a binary classification problem and rely on coarse-grained metrics such as overall accuracy, providing limited insight into where LVLMs succ… ▽ More

    Submitted 19 March, 2026; originally announced March 2026.

    Comments: ECCV 2026 submission. 14 pages, 6 figures, 4 tables. Supplementary material included

  40. arXiv:2603.17687  [pdf, ps, other

    cs.LG cs.AI

    Objective Mispricing Detection for Shortlisting Undervalued Football Players via Market Dynamics and News Signals

    Authors: Chinenye Omejieke, Shuyao Chen, Xia Cui

    Abstract: We present a practical, reproducible framework for identifying undervalued football players grounded in objective mispricing. Instead of relying on subjective expert labels, we estimate an expected market value from structured data (historical market dynamics, biographical and contract features, transfer history) and compare it to the observed valuation to define mispricing. We then assess whether… ▽ More

    Submitted 18 March, 2026; originally announced March 2026.

  41. arXiv:2603.14112  [pdf, ps, other

    cs.CV

    Revisiting the Perception-Distortion Trade-off with Spatial-Semantic Guided Super-Resolution

    Authors: Dan Wang, Haiyan Sun, Shan Du, Z. Jane Wang, Zhaochong An, Serge Belongie, Xinrui Cui

    Abstract: Image super-resolution (SR) aims to reconstruct high resolution images with both high perceptual quality and low distortion, but is fundamentally limited by the perception-distortion trade-off. GAN-based SR methods reduce distortion but still struggle with realistic fine-grained textures, whereas diffusion-based approaches synthesize rich details but often deviate from the input, hallucinating str… ▽ More

    Submitted 14 March, 2026; originally announced March 2026.

  42. arXiv:2603.14004  [pdf, ps, other

    cs.CV cs.AI

    U-Face: An Efficient and Generalizable Framework for Unsupervised Facial Attribute Editing via Subspace Learning

    Authors: Bo Liu, Xuan Cui, Run Zeng, Wei Duan, Chongwen Liu, Jinrui Qian, Lianggui Tang, Hongping Gan

    Abstract: Latent space-based facial attribute editing methods have gained popularity in applications such as digital entertainment, virtual avatar creation, and human-computer interaction systems due to their potential for efficient and flexible attribute manipulation, particularly for continuous edits. Among these, unsupervised latent space-based methods, which discover effective semantic vectors without r… ▽ More

    Submitted 14 March, 2026; originally announced March 2026.

  43. arXiv:2603.13792  [pdf, ps, other

    cs.LG cs.AI

    IGU-LoRA: Adaptive Rank Allocation via Integrated Gradients and Uncertainty-Aware Scoring

    Authors: Xuan Cui, Huiyue Li, Run Zeng, Yunfei Zhao, Jinrui Qian, Wei Duan, Bo Liu, Zhanpeng Zhou

    Abstract: As large language models (LLMs) scale to billions of parameters, full-parameter fine-tuning becomes compute- and memory-prohibitive. Parameter-efficient fine-tuning (PEFT) mitigates this issue by updating only a small set of task-specific parameters while keeping the base model frozen. Among PEFT approaches, low-rank adaptation (LoRA) is widely adopted; however, it enforces a uniform rank across l… ▽ More

    Submitted 14 March, 2026; originally announced March 2026.

  44. arXiv:2603.12920  [pdf, ps, other

    cs.CL stat.ML

    HMS-BERT: Hybrid Multi-Task Self-Training for Multilingual and Multi-Label Cyberbullying Detection

    Authors: Zixin Feng, Xinying Cui, Yifan Sun, Zheng Wei, Jiachen Yuan, Jiazhen Hu, Ning Xin, Md Maruf Hasan

    Abstract: Cyberbullying on social media is inherently multilingual and multi-faceted, where abusive behaviors often overlap across multiple categories. Existing methods are commonly limited by monolingual assumptions or single-task formulations, which restrict their effectiveness in realistic multilingual and multi-label scenarios. In this paper, we propose HMS-BERT, a hybrid multi-task self-training framew… ▽ More

    Submitted 13 March, 2026; originally announced March 2026.

  45. arXiv:2603.05410  [pdf, ps, other

    cs.RO

    PhysiFlow: Physics-Aware Humanoid Whole-Body VLA via Multi-Brain Latent Flow Matching and Robust Tracking

    Authors: Weikai Qin, Sichen Wu, Ci Chen, Mengfan Liu, Linxi Feng, Xinru Cui, Haoqi Han, Hesheng Wang

    Abstract: In the domain of humanoid robot control, the fusion of Vision-Language-Action (VLA) with whole-body control is essential for semantically guided execution of real-world tasks. However, existing methods encounter challenges in terms of low VLA inference efficiency or an absence of effective semantic guidance for whole-body control, resulting in instability in dynamic limb-coordinated tasks. To brid… ▽ More

    Submitted 5 March, 2026; originally announced March 2026.

  46. arXiv:2603.04819  [pdf, ps, other

    cs.RO cs.AI cs.LG

    On the Strengths and Weaknesses of Data for Open-set Embodied Assistance

    Authors: Pradyumna Tambwekar, Andrew Silva, Deepak Gopinath, Jonathan DeCastro, Xiongyi Cui, Guy Rosman

    Abstract: Embodied foundation models are increasingly performant in real-world domains such as robotics or autonomous driving. These models are often deployed in interactive or assistive settings, where it is important that these assistive models generalize to new users and new tasks. Diverse interactive data generation offers a promising avenue for providing data-efficient generalization capabilities for i… ▽ More

    Submitted 5 March, 2026; originally announced March 2026.

  47. arXiv:2603.04073  [pdf, ps, other

    cs.RO

    Swimming Under Constraints: A Safe Reinforcement Learning Framework for Quadrupedal Bio-Inspired Propulsion

    Authors: Xinyu Cui, Fei Han, Hang Xu, Yongcheng Zeng, Luoyang Sun, Ruizhi Zhang, Jian Zhao, Haifeng Zhang, Weikun Li, Hao Chen, Jun Wang, Dixia Fan

    Abstract: Bio-inspired aquatic propulsion offers high thrust and maneuverability but is prone to destabilizing forces such as lift fluctuations, which are further amplified by six-degree-of-freedom (6-DoF) fluid coupling. We formulate quadrupedal swimming as a constrained optimization problem that maximizes forward thrust while minimizing destabilizing fluctuations. Our proposed framework, Accelerated Const… ▽ More

    Submitted 4 March, 2026; originally announced March 2026.

  48. arXiv:2603.04057  [pdf, ps, other

    cs.RO cs.AI

    Sim2Sea: Sim-to-Real Policy Transfer for Maritime Vessel Navigation in Congested Waters

    Authors: Xinyu Cui, Xuanfa Jin, Xue Yan, Yongcheng Zeng, Luoyang Sun, Siying Wei, Ruizhi Zhang, Jian Zhao, Haifeng Zhang, Jun Wang

    Abstract: Autonomous navigation in congested maritime environments is a critical capability for a wide range of real-world applications. However, it remains an unresolved challenge due to complex vessel interactions and significant environmental uncertainties. Existing methods often fail in practical deployment due to a substantial sim-to-real gap, which stems from imprecise simulation, inadequate situation… ▽ More

    Submitted 4 March, 2026; originally announced March 2026.

  49. arXiv:2603.02083  [pdf, ps, other

    cs.RO cs.CV

    $π$-StepNFT: Wider Space Needs Finer Steps in Online RL for Flow-based VLAs

    Authors: Siting Wang, Xiaofeng Wang, Zheng Zhu, Minnan Pei, Xinyu Cui, Cheng Deng, Jian Zhao, Guan Huang, Haifeng Zhang, Jun Wang

    Abstract: Flow-based vision-language-action (VLA) models excel in embodied control but suffer from intractable likelihoods during multi-step sampling, hindering online reinforcement learning. We propose \textbf{\textit{$\boldsymbolπ$-StepNFT}} (Step-wise Negative-aware Fine-Tuning), a critic-and-likelihood-free framework that requires only a single forward pass per optimization step and eliminates auxiliary… ▽ More

    Submitted 9 March, 2026; v1 submitted 2 March, 2026; originally announced March 2026.

  50. arXiv:2602.23369  [pdf, ps, other

    cs.IR cs.AI cs.CL

    Reason to Contrast: A Cascaded Multimodal Retrieval Framework

    Authors: Xuanming Cui, Hong-You Chen, Hao Yu, Hao Yuan, Zihao Wang, Shlok Kumar Mishra, Hanchao Yu, Yonghuan Yang, Jun Xiao, Ser-Nam Lim, Jianpeng Cheng, Qi Guo, Xiangjun Fan

    Abstract: Traditional multimodal retrieval systems rely primarily on bi-encoder architectures, where performance is closely tied to embedding dimensionality. Recent work, Think-Then-Embed (TTE), shows that incorporating multimodal reasoning to elicit additional informative tokens before embedding can further improve retrieval. In this paper, we extend this paradigm with TTE-v2, a hybrid multimodal retrieval… ▽ More

    Submitted 20 December, 2025; originally announced February 2026.