Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 1,056 results for author: Pan, Y

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.19906  [pdf, ps, other

    cs.LG

    PETA:Parameter-Efficient Test-Time Adaptation for Virtual Screening

    Authors: Jia-Qi Lin, Yinghua Yao, Chang-Dong Wang, Yew-Soon Ong, Yuangang Pan

    Abstract: Accurately ranking active ligands for a target protein pocket from massive chemical libraries remains a central challenge in virtual screening. DrugCLIP and its recent extensions substantially accelerate this process by encoding protein pockets and molecules into a shared embedding space. Despite this progress, further performance improvements typically require retraining the entire model, incurri… ▽ More

    Submitted 20 August, 2026; originally announced August 2026.

  2. arXiv:2608.19890  [pdf, ps, other

    cs.LG

    Reliable Neural Collapse Approximation for Open-World Test-Time Adaptation

    Authors: Jia-Qi Lin, Yuangang Pan, Chang-Dong Wang, Haizhang Zhang, Ivor W. Tsang, Joey Tianyi Zhou

    Abstract: Test-Time Adaptation (TTA) methods aim to bridge the domain gap between the source and target domains. However, traditional TTA methods become ineffective when the label distribution shift occurs, a challenge commonly referred to as an open-world scenario. In this paper, we introduce a new method named Reliable Neural Collapse approximation (ReNC) for Open-World Test-Time Adaptation (OWTTA). Speci… ▽ More

    Submitted 20 August, 2026; originally announced August 2026.

  3. Image-Guided Pavement Defect Recognition in GPR Data with novel 3D Deep Learning Architecture

    Authors: Yuandong Pan, Linjun Lu, Mudan Wang, Florian Noichl, Fan Xue, Brian Sheil, Lavindra de Silva, André Borrmann, Ioannis Brilakis

    Abstract: Ground Penetrating Radar (GPR) is a widely adopted non-destructive sensing technology for subsurface inspection in civil and transportation engineering. Despite its potential for pavement condition assessment, the large-scale application of GPR in automated inspection has two key challenges: the scarcity of annotated real-world datasets and the lack of deep learning models designed for the unique… ▽ More

    Submitted 19 August, 2026; originally announced August 2026.

  4. arXiv:2608.18503  [pdf, ps, other

    cs.LG

    LLM-Powered Predictive Decision-Making for Sustainable Data Center Operations

    Authors: Hanzhao Wang, Jingxuan Wu, Yumeng Li, Yu Pan, Guanting Chen

    Abstract: The growing demand for AI-driven workloads, particularly from Large Language Models (LLMs), has raised concerns about the significant energy and resource consumption in data centers. This work introduces a novel LLM-based predictive scheduling system designed to enhance operational efficiency while reducing the environmental impact of data centers. Our system utilizes an LLM to predict key metrics… ▽ More

    Submitted 18 August, 2026; originally announced August 2026.

  5. arXiv:2608.17271  [pdf, ps, other

    cs.AI

    ASI-Bench: At the Dawn of Artificial Superintelligence

    Authors: Junwei Zhou, Zhen Sun, Binyu Li, Jiangyu Zhou, Yuexi Pan, Hengyu Wang, Honghe Ren, Xiaohan Jia, Xueyang Zhou, Xiaoyu Cao, Yongchao Chen, Yuanning Feng, Junhao Wu, Cheng Zhang, Sijia Chen, Haoyu Xue, Chengsong You, Huan Wang, Koutian Wu, Peigan Gao, Jiakun Wu, Wenzhe Li, Ergan Shang, Qingyuan Zheng, Jingjing Zhou , et al. (17 additional authors not shown)

    Abstract: Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce… ▽ More

    Submitted 17 August, 2026; originally announced August 2026.

    Comments: 16 pages, 5 figures, 2 tables

    ACM Class: I.2.0

  6. arXiv:2608.17255  [pdf, ps, other

    cs.CV cs.AI

    Learning Where and What to Lift for Bi-planar X-ray-to-CT Reconstruction

    Authors: Yifei Wu, Yicheng Wu, Qiang Ma, Qi Chen, Renyang Gu, Xinyu Liu, Yongsheng Pan, Yong Xia

    Abstract: X-ray imaging can be approximately modeled as the projection of an underlying volumetric attenuation field, with each measurement recording the accumulated attenuation along a corresponding ray path. Reconstructing a CT volume from only a few X-ray views is therefore severely ill-posed, as the projections collapse depth information and leave 3D locations of anatomical regions and their correspondi… ▽ More

    Submitted 17 August, 2026; originally announced August 2026.

  7. arXiv:2608.15877  [pdf, ps, other

    cs.AI

    Dear Algo: A Precision-First Agentic Intent Layer for Unified Search and Recommendation

    Authors: Rui Wang, Jiazhou Wang, Zheng Wei, Chenglin Lu, Fangcheng Sun, Ivy Sun, Jin Sun, Hui Geng, Lillian Zhang, Chao Yang, Lei Chen, Shahin Sefati, Reem Helou, Joe Zhou, Babak Shakibi, Yiyi Pan, Bi Xue, Hong Yan, Shujian Bu

    Abstract: Search and recommendation serve a shared discovery objective but encode intent differently. We study this boundary through Dear Algo on Threads, a deployed product where open-ended requests such as \emph{more NBA news} or \emph{less politics} steer subsequent feed recommendations rather than return a one-shot result list. Its agentic intent layer compiles explicit, inferred, negative, and compound… ▽ More

    Submitted 16 August, 2026; originally announced August 2026.

  8. arXiv:2608.12980  [pdf, ps, other

    cs.CV

    DiCoR: Decoupled Referent Disambiguation and Contour Recalibration for Efficient Referring Remote Sensing Image Segmentation

    Authors: Ziyang Gao, Zhizhuo Jiang, Jingjing Chang, Yixin Yang, Yuwen Pan, Yong-Qiang Mao, Yu Liu, Hai-Bao Chen

    Abstract: Referring remote sensing image segmentation (RRSIS) aims to delineate targets specified by natural language expressions in remote sensing imagery. Existing methods mainly follow joint fusion segmentation (JFS) or decoupled prompt segmentation (DPS). JFS is efficient but often suffers from limited accuracy because referent localization and mask delineation are optimized under a unified objective, w… ▽ More

    Submitted 13 August, 2026; originally announced August 2026.

  9. arXiv:2608.10413  [pdf, ps, other

    cs.CV

    DriveVLA-M0: Failure-Aware Memory Augmentation for Autonomous Driving

    Authors: Zebin Xing, Yupeng Zheng, Qiang Chen, Linbo Wang, Yichen Zhang, Pengxuan Yang, Junli Wang, Deheng Qian, Xiaoqing Ye, Junyu Han, Yifeng Pan, Qichao Zhang, Dongbin Zhao

    Abstract: Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for end-to-end autonomous driving by enabling unified reasoning across perception, language, and planning. However, existing approaches lack mechanisms to exploit past failures or adapt to distribution shifts, causing the model to persistently underperform on similar scenarios where it has previously failed. In this… ▽ More

    Submitted 10 August, 2026; originally announced August 2026.

  10. arXiv:2608.10011  [pdf, ps, other

    q-bio.QM cs.LG eess.SP

    HIPNO: Symmetry-Aware Physics-Informed Neural Operators for Noninvasive Hemodynamic Inference

    Authors: Yunbei Pan, Jiahang Sha, Simon A. Lee, Maxime Cannesson, Wei Wang, Jeffrey N. Chiang

    Abstract: Continuous hemodynamic monitoring guides treatment decisions in surgery and intensive care. However, gold-standard signals are only measured in severe cases due to risks associated with invasive measurement. In this work, we introduce HIPNO (Hemodynamic Inference via Physics-informed Neural Operators) to recover hemodynamic state from ubiquitous, non-invasive signals and expand access to advanced… ▽ More

    Submitted 8 August, 2026; originally announced August 2026.

    Comments: 12 pages, 3 figures

  11. arXiv:2608.09492  [pdf, ps, other

    cs.RO

    Rethink Before You Execute: Adaptive Execution for World Action Models

    Authors: Feng Ye, Yiming Zhao, Yong Yu, Hongxu Zhou, Yong Pan, Yuan Xue, Peng Jia, Chuanmin Jia

    Abstract: World Action Models (WAMs) jointly predict future actions and the evolution of the environment. At each inference, a WAM generates a chunk of actions and the robot executes a fixed prefix before replanning. We argue that this fixed execution horizon is poorly matched to execution dynamics: the chunk reliability varies across task stages, so when to replan depends on the result of accumulated execu… ▽ More

    Submitted 10 August, 2026; originally announced August 2026.

  12. arXiv:2608.08734  [pdf, ps, other

    cs.CV

    IDATA: Scalable Invertible Diffusion for Unrestricted Adversarial Transfer Attack

    Authors: Yi Pan, Jun-Jie Huang, Tianrui Liu, Zihan Chen, Lin Liu, Zhao Wentao

    Abstract: Unrestricted adversarial transfer attacks are important for evaluating the black-box robustness of deep visual models. Diffusion-based attacks have shown promising transferability and visual imperceptibility by optimizing adversarial perturbations along denoising trajectories in latent space. However, existing methods are limited by two challenges: memory-intensive multistep backpropagation and fr… ▽ More

    Submitted 9 August, 2026; originally announced August 2026.

    Comments: Adversarial attack,Invertible diffusion model,Memory-efficient,Low-frequency

  13. arXiv:2608.08621  [pdf, ps, other

    cs.AI

    Business Arena: Benchmarking LLM Agents in a Realistic Marketplace

    Authors: Yijun Pan, Yukun Lian, Kunyu Shi, Junbo Li, Hongwei Xue, Sicong Xie, Guannan Zhang, Xiaoying Xing

    Abstract: Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under uncertainty, adapt to delayed outcomes in a changing market, and satisfy regulatory obligations before trading legally. Frontier LLM agents can increasingly complete complex workflows, yet business-related capabilities are rarely evaluated in existing agent be… ▽ More

    Submitted 9 August, 2026; originally announced August 2026.

  14. arXiv:2608.07110  [pdf, ps, other

    cs.LG cs.CL

    Modular TTT: Rethinking Test-Time Training as Composable Modules

    Authors: Bohao Tang, Zhen Qin, Yuqi Pan, Zheng Li, Pengfei Liu, Ya Zhang

    Abstract: Test-time training (TTT) views sequence modeling as an online learning problem in which fast weights are updated by an internal learning rule. Despite the growing number of TTT variants, existing approaches typically hard-code each variant separately, which makes it difficult to design new TTT methods and to isolate the role of each component. To address this, we propose Modular TTT, a framework t… ▽ More

    Submitted 7 August, 2026; originally announced August 2026.

    Comments: Code: https://github.com/ByteDance-Seed/Modular-TTT

  15. SubtleTalk: Generating Controllable Weakly-correlated Facial Dynamics for 3D Talking Heads via Residual Flow Matching

    Authors: Chenyang Ding, Shuai Tan, Qunfen Lin, Xinwei Jiang, Zijiao Zeng, Ye Pan

    Abstract: Audio-driven 3D facial animation aims to synthesize realistic and temporally coherent motions from speech. Despite notable progress in lip synchronization, weakly correlated dynamics, including eyebrow movements, eye blinks, and head motion, which are essential to photorealistic facial animation, remain difficult to model faithfully and often appear static or unnaturally repetitive. We attribute t… ▽ More

    Submitted 3 August, 2026; originally announced August 2026.

  16. arXiv:2608.05369  [pdf, ps, other

    cs.RO cs.CV

    World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation

    Authors: Yuhao Pan, Haosong Peng, Zhengshen Zhang, Zhengyang Yan, Yalun Dai, Fushuo Huo, Chujie Wang, Tianyu Qi, Xiucheng Wang, Nan Cheng, Wenchao Xu

    Abstract: Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wrist-local interactions may evolve under the global task context. To address this limitation, we present World-to-Wrist VLA (W2-VLA), a VLA model for fine-grained… ▽ More

    Submitted 5 August, 2026; originally announced August 2026.

  17. arXiv:2608.05201  [pdf, ps, other

    cs.CR cs.AI

    ASTELD: A Six-Axis Classification Framework for Autonomous AI Agents - Design, Evaluation, and an OpenClaw Case Study

    Authors: Siyuan Li, Peng Shu, Churan Yu, Peilong Wang, Ruidong Zhang, Bowen Guo, Xinliang Li, Ruiyu Yan, Arif Hassan Zidan, Yi Pan, Wei Ruan, Lifeng Chen, Junhao Chen, Zhaojun Ding, Yiwei Li, Zhengliang Liu, Haixing Dai, Lin Zhao, Yu Bao, Xiang Li, Wei Zhang, Tianming Liu

    Abstract: Autonomous AI agent platforms differ substantially in architecture, security, tool integration, execution, autonomy, and deployment, yet the field lacks a common classification scheme for comparing these design choices. We propose ASTELD, an operational six-axis classification framework for autonomous AI agents: Architecture pattern, Security posture, Tool integration model, Execution paradigm, Le… ▽ More

    Submitted 5 August, 2026; originally announced August 2026.

    Comments: 40 pages, 4 figures, 6 tables. Introduces and empirically evaluates the ASTELD six-axis classification framework across eight autonomous AI agent platforms, with OpenClaw as an in-depth case study

  18. arXiv:2608.04625  [pdf, ps, other

    cs.AI

    A/B Agent: A Self-Evolving Agent for Strategy Iteration in Industrial A/B Testing

    Authors: Zhuohang Jiang, Yuxin Chen, Yongsen Pan, Zheng Hu, Wenqi Fan, Qing Li, Hongyang Wang, Jun Wang, Wenwu Ou

    Abstract: Industrial recommendation strategy iteration heavily relies on large-scale A/B experimentation. Traditional tuning requires experts to repeatedly design strategies, configure experiments, analyze results, and adjust parameters, making the process labor-intensive and time-consuming. Meanwhile, valuable knowledge from historical experiments is often fragmented, making systematic reuse difficult thro… ▽ More

    Submitted 5 August, 2026; originally announced August 2026.

  19. arXiv:2608.04586  [pdf, ps, other

    cs.CL cs.AI

    Breaking the Curse of Multilinguality in Many-to-Many Speech-to-Text Translation via a Resource-Aware Mixture of Speech Encoders

    Authors: Yexing Du, Kaiyuan Liu, Youcheng Pan, Bo Yang, Chengpeng Fu, Yu Wang, Ming Liu

    Abstract: Multimodal large language models (MLLMs) have achieved significant success in speech-to-text translation (S2TT). However, when processing multilingual speech inputs, a single speech encoder shared across all languages suffers from the curse of multilinguality: languages at different resource levels compete for limited representation capacity, leading to strong high-resource performance but substan… ▽ More

    Submitted 5 August, 2026; v1 submitted 5 August, 2026; originally announced August 2026.

  20. arXiv:2608.04223  [pdf, ps, other

    cs.DS cs.DM math.CO

    Bicriteria Approximation Algorithms for Demand Matching

    Authors: Yuchong Pan, Michel X. Goemans

    Abstract: The demand matching problem generalizes both the knapsack problem and the $b$-matching problem. In this problem, each edge of a graph has a demand and a weight, each vertex has a capacity, and the goal is to find a maximum weight subset of edges whose total incident demand at every vertex does not exceed its capacity. We study $(α, β)$-bicriteria approximation algorithms, which return a solution o… ▽ More

    Submitted 4 August, 2026; originally announced August 2026.

  21. arXiv:2608.03740  [pdf, ps, other

    cs.AI

    MissClick: Exploiting Digit-Serialized Coordinates to Attack GUI Grounding Models

    Authors: Yu Ran, Wentao Zhao, Xin Zhang, Yi Pan

    Abstract: Recent GUI visual grounding models generate screen coordinates as sequences of digit tokens that are parsed into numerical values and mapped to executable clicks. The security implications of this coordinate generation process have been largely overlooked. We observe that each coordinate digit is predicted as a categorical token, yet after parsing, changing a hundreds-place digit by one changes th… ▽ More

    Submitted 4 August, 2026; originally announced August 2026.

  22. arXiv:2608.01862  [pdf, ps, other

    cs.AI

    Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge

    Authors: Xiaofeng Shi, Xiaosong Qiu, Wenxin Ma, Qian Kou, Yiming Pan, Longbin Yu, Ying Liu, Haiping Wang, Hua Zhou

    Abstract: Enterprise question answering requires models to acquire proprietary knowledge without discarding general capabilities. We present Wnuan, a three-stage pipeline that constructs task-oriented supervision from documents, performs supervised fine-tuning with general-data replay, and applies reinforcement learning to residual errors. On the 707-question WnuanBench, the primary 32B route raises accepta… ▽ More

    Submitted 3 August, 2026; originally announced August 2026.

    Comments: 22 pages, 10 figures. Includes Supplementary Appendices A--L. Xiaofeng Shi and Xiaosong Qiu contributed equally

  23. arXiv:2608.00434  [pdf, ps, other

    cs.CL cs.AI

    AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction

    Authors: Ziqiang Cui, Han Shi, Bowei He, Yu Pan, Peiyang Liu, Shengyin Sun, Yankai Chen, Haoli Bai, Yichun Yin, Xue Liu, Chen Ma

    Abstract: Multi-Token Prediction (MTP) has emerged as an effective paradigm that augments a shared Large Language Model backbone with auxiliary heads, training the model to predict several future tokens in parallel to enrich its supervision signal and accelerate inference. However, existing training frameworks adopt a rigid, fixed-length prediction horizon, disregarding the highly non-uniform information de… ▽ More

    Submitted 1 August, 2026; originally announced August 2026.

  24. arXiv:2608.00215  [pdf, ps, other

    cs.AI

    Personalizing Large Language Model Agents with Small Policy Models

    Authors: Dian Jin, Zhi Zhang, Huichao Li, Yihe Pan, Rundong Huang, Doudou Zhou

    Abstract: Large language model (LLM) agents can retrieve memory, call tools, ask clarifying questions, and vary response style, yet adapting these execution decisions to an individual user remains difficult. Fine-tuning a separate LLM is costly or impossible for proprietary systems, while prompts and memory primarily expose user information to the agent rather than adapt its execution decisions from feedbac… ▽ More

    Submitted 31 July, 2026; originally announced August 2026.

  25. arXiv:2607.28834  [pdf, ps, other

    cs.CV

    FocusGS: Spatial Delta Layers for Local Repair and Deterministic Editing of Trained 3D Gaussian Assets

    Authors: Yiqun Pan, Yukun Shi

    Abstract: 3D Gaussian Splatting (3DGS) is evolving from one-time reconstruction into deliverable, inspectable, and maintainable visual assets. Existing workflows focus on global reconstruction, training-time density control, or open-ended generative editing, leaving trained assets without precise local maintenance. We propose FocusGS, which unifies local repair and deterministic editing as composite spatial… ▽ More

    Submitted 30 July, 2026; originally announced July 2026.

    Comments: 8 pages, 6 figures, 5 tables. Ancillary demonstration video included

  26. arXiv:2607.28421  [pdf, ps, other

    cs.AI

    When Derived Measurements Mislead: Quantifying and Mitigating LLM Over-Trust with Privileged-Modality Reliability Evidence

    Authors: Zongheng Guo, Tao Chen, Tianli Li, Mingzhe Cui, Yang Jiao, Lei Xie, Yi Pan, Xiao Hu, Manuela Ferrario

    Abstract: Derived measurements increasingly enter large language model (LLM) pipelines as direct facts despite their instance-dependent validity. We define derived-feature over-trust (DFOT) as the failure in which a downstream LLM assigns such a measurement the epistemic status of a direct fact or uses it outside its valid scope. Using physiological sensing as a case study, D1 tests acceptance of a PPG-deri… ▽ More

    Submitted 30 July, 2026; originally announced July 2026.

    Comments: 25 pages, including references and supplementary material; 3 figures and 19 tables. Code: https://github.com/Zongheng-Guo/When-Derived-Measurements-Mislead

    ACM Class: I.2.7; I.2.6; J.3

  27. arXiv:2607.27614  [pdf, ps, other

    cs.CL cs.AI

    DualAnchor: Preserving Language Priors and Improving Lexical Fidelity in Gloss-Free Sign Language Translation

    Authors: Hongbin Zhang, Junhao Liu, Xuefeng Bai, Youcheng Pan, Yang Xiang, Kehai Chen

    Abstract: Recent advances in large language models (LLMs) have led sign language translation (SLT), the task of converting sign-language videos into spoken-language text, to increasingly adopt LLMs as textual backbones. However, despite their strong language modeling capabilities, existing LLM-based SLT methods often undermine rather than exploit this language prior, producing disfluent translations, a fail… ▽ More

    Submitted 29 July, 2026; originally announced July 2026.

  28. arXiv:2607.25857  [pdf, ps, other

    cs.CL cs.CV

    Shieldstral

    Authors: Antonia Calvi, Avinash Sooriyarachchi, Giada Pistilli, Guillaume Lample, Maarten Buyl, Maximilian Augustin, Maximilian Müller, Pierre Stock, Tom Bewley, Wassim Bouaziz, Yimu Pan, Abdelaziz Bounhar, Abhijeet Somani, Aditi Kabra, Adrian Valente, Adrien Petralia, Adrien Sadé, Alan Jeffares, Albert Jiang, Aleksandr Timashov, Alexandre Cahill, Alexandre Gavaudan, Alexandre Laval, Alexandre Sablayrolles, Amélie Héliou , et al. (251 additional authors not shown)

    Abstract: We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7$\times$ its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content moderation as a binary question-answering task. This simple formulation unifies diverse moderation tasks into a single yes/no p… ▽ More

    Submitted 4 August, 2026; v1 submitted 28 July, 2026; originally announced July 2026.

  29. arXiv:2607.24419  [pdf, ps, other

    cs.AI

    Failures Reveal What Metrics Miss: An Evidence-Driven Agent for Recursive Refinement of ECG Classifiers

    Authors: Jinliang Deng, Yiming Niu, Yibo Pan, Zhiqi Shao, Qin Luo, Yongxin Tong

    Abstract: Deep models have substantially advanced 12-lead ECG classification, yet their refinement still relies heavily on human experts to inspect failures and iteratively revise classifier designs. Recent LLM-based agents have demonstrated the potential for automated model design, but when guided only by aggregate performance metrics, they lack insight into why individual cases fail and how the classifier… ▽ More

    Submitted 27 July, 2026; originally announced July 2026.

  30. arXiv:2607.21354  [pdf, ps, other

    cs.AI

    SPORD: A Simulation-Propose-then-OR-Dispose Approach for Supply Chain Planning

    Authors: Jiayin He, Yutong Pan, Sen Yang, Ningxuan Kang, Yongzhi Qi, Jianshen Zhang, Wei Qi, Zuo-Jun Max Shen

    Abstract: For years, supply chain planning at e-commerce firms has operated as a collection of isolated projects. Each planning task from static network planning to dynamic warehouse assortment planning requires analysts to spend weeks building models from scratch, calibrating and persuading executives to act on outputs they cannot verify. Three barriers drive this: bespoke models proliferate because standa… ▽ More

    Submitted 23 July, 2026; originally announced July 2026.

  31. arXiv:2607.20785  [pdf, ps, other

    cs.RO cs.AI

    Robostral Navigate

    Authors: Abdelaziz Bounhar, Abhijeet Somani, Aditi Kabra, Adrian Valente, Adrien Petralia, Adrien Sade, Alan Jeffares, Albert Jiang, Aleksandr Timashov, Alexandre Cahill, Alexandre Gavaudan, Alexandre Laval, Alexandre Sablayrolles, Amelie Heliou, Amos You, Andre Jonasson, Andrew Bai, Andrew Ehrenberg, Andrew Zhao, Angele Lenglemetz, Anmol Agarwal, Antonia Calvi, Arata Suzuki, Arjun Majumdar, Arthur Fournier , et al. (251 additional authors not shown)

    Abstract: Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently. Yet, today's best systems depend on depth sensors, multi-camera rigs, or pre-built maps, limiting the hardware they support and increasing deployment cost. We introduce Robostral Navigate, an 8B vision-language model built around this scalability… ▽ More

    Submitted 31 July, 2026; v1 submitted 22 July, 2026; originally announced July 2026.

  32. arXiv:2607.18979  [pdf, ps, other

    cs.AI

    Fishing Out Free Riders: Shapley-Based Reward Attribution for Parallel Reasoning via Reinforcement Learning

    Authors: Wentao Zhang, Haoyu Zhang, Xinke Jiang, Yuxuan Cheng, Yuhan Pan, Miao Li, Zhipeng Qiao, Tao Feng, Zhen Tao, Dengji Zhao

    Abstract: Large Language Models (LLMs) excel at multi-step reasoning, yet current parallel reasoning approaches often fail to distinguish the contributions of individual reasoning paths. Many paths may be redundant, misleading, or even detrimental, but outcome-level rewards assign uniform reward, leading to ambiguous learning signals and unstable training. We propose Parallel Shapley, a reinforcement learni… ▽ More

    Submitted 21 July, 2026; originally announced July 2026.

    Comments: 19 pages, 4 figures, 8 Tables

  33. arXiv:2607.18794  [pdf, ps, other

    cs.RO

    Beyond Transformers: Linear Attention Policy for Open-Vocabulary Object Goal Navigation

    Authors: Jiahong Zhang, Yifan Lin, Yandong Zhang, Sijun Shen, Kexin Wang, Yuqi Pan, Hongjuan Pei, Wei Wang, Guoqi Li

    Abstract: Open-Vocabulary Object Goal Navigation (OVON) requires agents to operate under partial observability, making effective internal state updates critical for navigation performance. This update is implemented by the policy network, where recent approaches adopt Transformer-based backbones with self-attention over a context window to integrate temporal information. However, our controlled experiments… ▽ More

    Submitted 21 July, 2026; originally announced July 2026.

    Comments: 12 pages, 7 figures

  34. arXiv:2607.16248  [pdf, ps, other

    cs.LG cs.AI

    High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration

    Authors: Gradwell Dzikanyanga, Yanqi Pan, Weihao Yang, Donglei Wu, Wen Xia, Hao Huang

    Abstract: Long-context large language model inference relies on the KV cache to avoid redundant attention computation, but incurs high memory and bandwidth overheads. Low-bit KV-cache quantization reduces this cost, yet it severely degrade quality; particularly, one-bit quantization reduces accuracy from 84.2% to 47.8% on Llama-3.1-8B under RULER. Rather than common beliefs that absolute error of logits,… ▽ More

    Submitted 26 June, 2026; originally announced July 2026.

  35. arXiv:2607.16189  [pdf, ps, other

    cs.CV

    Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA

    Authors: Ce Zhang, Ziyang Wang, Yulu Pan, Oluwatumininu Oguntola, Pranav Wagh, Qiyu Wu, Hiromi Wakaki, Mohit Bansal, Gedas Bertasius

    Abstract: Grounded long-video question answering (Grounded LVQA) requires answering a question about a long video while localizing the short evidence interval that supports the answer. Recent agentic methods frame this task as multi-turn exploration with a single crop_video(start, end) action, which supports coarse-to-fine narrowing but provides no primitive for fine-to-coarse backtracking. As a result, the… ▽ More

    Submitted 17 July, 2026; originally announced July 2026.

  36. arXiv:2607.15659  [pdf

    cs.RO

    Continuously Stable Structure through Plastic Deformation

    Authors: Junlong Xiao, Yaoqiang Pan, Xuan Zhang, Michael Yu Wang, Chao Chen

    Abstract: Soft robots have seen widespread adoption in interactive tasks due to their inherent compliance and adaptability. However, these advantages often come at the cost of stability, posing challenges in a dynamic environment. This limitation is especially critical in soft grippers, where instability under acceleration or external disturbances can result in grasp failure. In this study, we present a con… ▽ More

    Submitted 20 July, 2026; v1 submitted 17 July, 2026; originally announced July 2026.

    Comments: 24 pages, 8 figures

  37. arXiv:2607.14497  [pdf, ps, other

    cs.CV

    Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation

    Authors: Chi Kit Wong, Ye Pan, Yuanhuiyi Lyu, Xu Zheng, Zidong Cao, Lutao Jiang, Zixin Zhang, Huiyu Zhou, Xuming Hu

    Abstract: Egocentric Visual Question Answering (VQA) has attracted widespread attention as an important task for enabling Multimodal Large Language Models (MLLMs) to interact with the real world. However, existing MLLMs struggle to perform effective spatial reasoning in complex egocentric scenes due to their limited spatial perception capabilities. To this end, we introduce Ego Scene Augmentation (ESA), an… ▽ More

    Submitted 20 August, 2026; v1 submitted 15 July, 2026; originally announced July 2026.

    Comments: 14 pages, 8 figures. Chi Kit Wong and Ye Pan contributed equally. Code: https://github.com/Chikit-WONG/spatialGraph

  38. arXiv:2607.06216  [pdf, ps, other

    cs.CV

    MoWorld: A Flash World Model

    Authors: Team Moxin, Deyi Ji, Tianrun Chen, Xin Zhang, Jiale Yang, Qi Zhu, An Zhao, Zihao Xie, Han Wang, Xuanyi Liu, Yixiang Zhou, Pei Liu, Yi Tan, Cheng Chen, Dayi Zhu, Mingyu Wei, Hanjie Xu, Jun Liao, Siqi Li, Lingyu Lu, Hongye Fang, Hongming Tan, Youjiang Zhu, Taiyu Zhang, Zejian Li , et al. (15 additional authors not shown)

    Abstract: The future of World Models depends not only on scaling model capability, but also on scaling practicality and inference efficiency. High-frame-rate inference enables responsive perception, planning, and control in real-world autonomous systems. To this end, we present MoWorld, a cost-effective yet high-performance Flash World Model with an end-to-end framework spanning data generation, pre-trainin… ▽ More

    Submitted 3 August, 2026; v1 submitted 7 July, 2026; originally announced July 2026.

    Comments: Project Page: https://moxin-tech.github.io/moworld/

  39. arXiv:2607.06186  [pdf, ps, other

    cs.RO

    Calf-Integrated Arms for Bimanual Quadruped Loco-Manipulation

    Authors: Yan Pan, Yuanchuan Ren, Chipui Chan, Jingcheng Sun, Chengxu Zhou

    Abstract: Most quadruped loco-manipulation designs trade manipulation capability against stance. A trunk-mounted arm sits high and usually carries a single arm; using the legs as manipulators lifts the manipulating leg off the ground; and even leg-mounted grippers reach two-handed tasks only by rearing onto the hind legs. This paper integrates a manipulator with a prismatic slider, two revolute joints, and… ▽ More

    Submitted 11 July, 2026; v1 submitted 7 July, 2026; originally announced July 2026.

    Comments: 6 pages, 6 figures

  40. arXiv:2607.05772  [pdf, ps, other

    cs.SE

    Detecting Vulnerability-Inducing Commits via Multi-Stage Reasoning with LLM-Based Agents

    Authors: Liyou Chen, Hailong Sun, Xiang Gao, Yue Pan

    Abstract: Detecting vulnerability-inducing commits (VICs) at submission time is critical for improving the security and reliability of software systems. However, this task is highly challenging because it requires reasoning about the semantic impact of code changes from heterogeneous information sources, including code diffs, commit messages, and the surrounding contextual code. Existing approaches often st… ▽ More

    Submitted 6 July, 2026; originally announced July 2026.

  41. arXiv:2607.04718  [pdf, ps, other

    cs.AI

    FORGE: Research-Trajectory Hijacking Attacks on Deep Research Agents

    Authors: Yue Pan, Ziheng Zhang, Junxiang Lei, Changhao Jia, Qingyi Si, Hongcheng Guo

    Abstract: Deep research agents decompose open-ended queries into subtasks, retrieve web evidence over multiple rounds, and synthesize long-form reports. This workflow creates a planning-layer poisoning surface: adversarial documents that enter the retrieval pool can steer follow-up questions and turn a local injection into report-level contamination. We present FORGE (Fabricated Orchestrated Reasoning chain… ▽ More

    Submitted 6 July, 2026; originally announced July 2026.

    Comments: 20 pages,8 figures,Code available at https://github.com/yvepan/FORGE

  42. arXiv:2607.01804  [pdf, ps, other

    cs.RO

    VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon

    Authors: Yi Pan, Miao Pan, Qi Lu, Jiaming Huang, Man Zhang, Siteng Huang, Xin Li, Jie Zhang, Yongliang Shen, Xuhong Zhang, Wenqi Zhang

    Abstract: Vision-Language-Action (VLA) foundation models have recently achieved strong progress in embodied intelligence. To reduce policy-call frequency while preserving temporal coherence, most generative policies adopt an action chunk mechanism, executing multiple future actions in an open-loop manner under a fixed action horizon. However, this "predict-then-blindly-execute" paradigm sacrifices closed-lo… ▽ More

    Submitted 2 July, 2026; originally announced July 2026.

    Comments: 22 pages, 14 figures

  43. arXiv:2606.30560  [pdf, ps, other

    cs.LG cs.AI cs.PF

    TraceLab: Characterizing Coding Agent Workloads for LLM Serving

    Authors: Kan Zhu, Mathew Jacob, Chenxi Ma, Yi Pan, Stephanie Wang, Arvind Krishnamurthy, Baris Kasikci

    Abstract: Coding agents are rapidly becoming a major application of agentic LLMs, but serving them efficiently remains challenging. Progress on this challenge requires understanding real workload patterns, yet the data needed for such analysis is largely absent. Existing public traces and benchmarks do not capture real, day-to-day coding-agent usage across multiple agents and model families for serving-syst… ▽ More

    Submitted 30 June, 2026; v1 submitted 29 June, 2026; originally announced June 2026.

  44. arXiv:2606.29934  [pdf, ps, other

    cs.RO

    RoamFlow: Reinforcement-Aligned One-Step Action MeanFlow Policy for Image-Goal Navigation

    Authors: Zixuan Zhang, Yuqi Chen, Junjie Gao, Siyuan Song, Yongzhou Pan, Beichen Wang, Mir Feroskhan

    Abstract: Image-goal navigation is a key challenge in embodied robotics, where an agent must reach a target specified solely by a goal image. While existing reinforcement learning approaches map perceptual observations directly to actions, they struggle to model long-horizon dependencies, often leading to suboptimal trajectories. To address this limitation, we propose RoamFlow, a generative navigation frame… ▽ More

    Submitted 29 June, 2026; originally announced June 2026.

  45. arXiv:2606.29917  [pdf, ps, other

    cs.RO

    Flying to Image-Specified Objects: 3D Quadrotor Navigation via Cross-Graph Memory and Viewpoint Planning

    Authors: Junjie Gao, Yuqi Chen, Yongzhou Pan, Yaosheng Deng, Jiaping Xiao, Mir Feroskhan

    Abstract: Instance-Specific Image-Goal Navigation (InstanceImageNav) requires a robot to navigate toward the exact object instance depicted in a query image. Extending this task to quadrotors is challenging due to continuous 3D control, limited field of view (FOV), and safety constraints, which make successful navigation highly dependent on selecting informative viewpoints. We propose a hierarchical navigat… ▽ More

    Submitted 29 June, 2026; originally announced June 2026.

  46. arXiv:2606.29837  [pdf, ps, other

    cs.CV

    Robust Trajectory Distillation: Hybrid Reweighting Meets Teacher-Inspired Targets

    Authors: Kaifeng Chen, Lechao Cheng, Jiyang Li, Shengeng Tang, Fan Zhang, Yantao Pan, Yaxiong Wang, Tuanrui Hui, Zhun Zhong

    Abstract: Dataset distillation (DD) condenses large corpora into compact, information-rich subsets for efficient training and reuse. However, under noisy supervision, DD risks condensing corrupted associations together with useful signals, degrading robustness. Conventional noisy-label remedies (sample selection, loss weighting, label correction) tightly couple noise estimation with model optimization, ofte… ▽ More

    Submitted 29 June, 2026; originally announced June 2026.

  47. arXiv:2606.27457  [pdf, ps, other

    cs.PF cs.CL

    Cluster, Route, Escalate: Cascaded Framework for Cost-Aware LLM Serving

    Authors: Yasmin Moslem, Magdalena Kacmajor, Vasudevan Nedumpozhimana, Ammar Abbas, Solmaz Panahi, David Lynch, Zhuangzhuang Nie, Alexandros Agapitos, Aleksandar Milenovic, Hongmeng Song, Yucheng Shi, Yue Pan, Patricia Buffini, John D. Kelleher

    Abstract: Efficient deployment of large language models (LLMs) in production forces a trade-off between accuracy and cost. Operators often default to a single model that is either expensive for easy queries or insufficient for hard ones. To address this challenge, we propose a two-stage cascaded solution. Stage 1 clusters incoming queries and assigns each cluster to its most cost-effective model. The cost b… ▽ More

    Submitted 25 June, 2026; originally announced June 2026.

  48. arXiv:2606.27285  [pdf, ps, other

    cs.LG cs.IT math.CA math.DS

    Recovering Governing Equations from Solution Data: Identifiability Bounds for Linear and Nonlinear ODEs

    Authors: Yang Pan, Helmut Bölcskei

    Abstract: Learning governing equations from observed solution data is a fundamental challenge in scientific machine learning, yet the theoretical conditions under which a ground-truth ODE can be uniquely and stably identified from multiple solution observations remain largely undeveloped, and no quantitative analysis of the sample complexity of such learning tasks exists in the literature. To address this g… ▽ More

    Submitted 26 June, 2026; v1 submitted 25 June, 2026; originally announced June 2026.

  49. arXiv:2606.27095  [pdf, ps, other

    cs.LG cs.AI

    Data-Free Reservoir Features for Efficient Long-Horizon Cold-Start Continual Learning

    Authors: Augustinas Jučas, Yangchen Pan

    Abstract: Cold-start exemplar-free class-incremental learning requires learning a growing set of classes without replay, external pretraining, or a large initial task. Existing cold-start methods typically either train the backbone throughout the stream and compensate for semantic drift, or freeze a backbone after the first task, producing features biased toward the initial classes. These choices also creat… ▽ More

    Submitted 25 June, 2026; originally announced June 2026.

  50. arXiv:2606.26588  [pdf, ps, other

    cs.RO

    Inference-Time Robot Behavior Steering through Physically-Aware Reconfiguration of Task-Structure

    Authors: Yiyuan Pan, Hanjiang Hu, Shangtao Li, Xusheng Luo, Changliu Liu

    Abstract: A central challenge in deploying learned robot policies is inference-time behavior steering: redirecting a policy at test time to satisfy user preferences not anticipated during training, without retraining. Existing methods fail in two modes: end-to-end methods require fine-tuning or expert-level guidance, while neuro-symbolic methods rely on predefined symbols whose edits can result in logically… ▽ More

    Submitted 25 June, 2026; originally announced June 2026.