Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 296 results for author: Jiao, P

.
  1. arXiv:2608.20909  [pdf, ps, other

    cs.LG cs.RO

    Decoupling Policy Extraction for Offline Reinforcement Learning

    Authors: Xuyao Lin, Yixiang Shan, Jinru Duan, Tao Yang, Xinyu Zhao, Runyu Lei, Yiming Zhao, Jiaxin Fan, Zongbao Feng, Peng Jia

    Abstract: Offline RL methods commonly jointly train the actor and critic, where the critic is used to guide the actor toward higher-value actions. This coupled learning process is well motivated in online RL, where an improved actor collects new data that can further update the actor and the critic. However, training data remains fixed in offline RL, making actor-side policy improvement unable to generate n… ▽ More

    Submitted 21 August, 2026; originally announced August 2026.

  2. arXiv:2608.16274  [pdf

    cs.IR cs.AI

    Decoupled Temporal Encoding for Generative Recommendation

    Authors: Pengfei Jia, Jingjian Wang, Jingmao Li, Ge Zhang, Feng Shi

    Abstract: Positional encoding is a fundamental component of Transformer-based generative recommendation models, where user histories are modeled as autoregressive item sequences. Most positional encoding methods are inherited from natural language processing and mainly represent discrete item order. However, recommendation sequences go beyond ordered lists, as timestamps and temporal effects also shape item… ▽ More

    Submitted 17 August, 2026; originally announced August 2026.

    Comments: accepted by CIKM '26

  3. arXiv:2608.15110  [pdf, ps, other

    cs.CV cs.AI

    CETalk: Continuous Valence-Arousal Control for Audio-Driven 3D Talking Head Generation

    Authors: Peng Jia, Li Dai, Zhen Xiao, Xueliang Liu, Jia Li

    Abstract: Emotional 3D talking head generation aims to synthesize expressive facial animations with accurate lip synchronization. However, existing methods often rely on discrete emotion categories, which fail to capture the continuous evolution of affect. They also overlook the temporal frequency mismatch between audio articulation and emotional expression. In this paper, we propose CETalk, an audio-driven… ▽ More

    Submitted 15 August, 2026; originally announced August 2026.

    Comments: 14 pages, 6 figures, 3 tables

  4. arXiv:2608.09492  [pdf, ps, other

    cs.RO

    Rethink Before You Execute: Adaptive Execution for World Action Models

    Authors: Feng Ye, Yiming Zhao, Yong Yu, Hongxu Zhou, Yong Pan, Yuan Xue, Peng Jia, Chuanmin Jia

    Abstract: World Action Models (WAMs) jointly predict future actions and the evolution of the environment. At each inference, a WAM generates a chunk of actions and the robot executes a fixed prefix before replanning. We argue that this fixed execution horizon is poorly matched to execution dynamics: the chunk reliability varies across task stages, so when to replan depends on the result of accumulated execu… ▽ More

    Submitted 10 August, 2026; originally announced August 2026.

  5. arXiv:2608.06917  [pdf, ps, other

    cs.AI

    ReGraph: Learning to Generate Recipe Graphs from Food Images

    Authors: Guoshan Liu, Bin Zhu, Pengkun Jiao, Jingjing Chen, Chong-Wah Ngo, Yu-Gang Jiang

    Abstract: Recent Large Multimodal Models (LMMs) have achieved impressive performance in recipe generation from food images.However, cooking is a structured transformation process in which ingredients undergo state changes through ordered actions,while free-form recipe language leaves the corresponding entities, intermediate states, and dependencies largely implicit and entangled.A graph representation makes… ▽ More

    Submitted 7 August, 2026; originally announced August 2026.

  6. arXiv:2608.04471  [pdf, ps, other

    cs.LG cs.AI

    Beyond Linear Dynamics: Neural Bilinear Dynamical Models for Time Series Forecasting

    Authors: Mengzhou Gao, Huangqian Yu, Pengfei Jiao

    Abstract: Time series in real-world applications are often generated by nonlinear dynamical systems, making accurate forecasting challenging. Existing approaches that explicitly model system dynamics typically rely on linear assumptions or Koopman-based linearizations, which may inadequately capture complex nonlinear behaviors and lead to error accumulation in long-horizon prediction. To address this limita… ▽ More

    Submitted 5 August, 2026; originally announced August 2026.

  7. arXiv:2607.24662  [pdf, ps, other

    cs.LG cs.SI

    When Can You Correct Distribution Drift in Temporal Graph Generation? A Sharpening--Drift Tension and an Impossibility for Observation-Based Correction

    Authors: Tianpeng Li, Xuan Guo, Wenjun Wang, Wang Zhang, Pengfei Jiao

    Abstract: Generative models of temporal graphs are trained on one stretch of an evolving network and deployed on the next, and they degrade badly in the gap. We show this degradation is derivable, general, and not fixable from observations. The masked flow-matching loss decomposes exactly, with no independence assumption, into an irreducible entropy plus a divergence whose derivative along the training path… ▽ More

    Submitted 27 July, 2026; originally announced July 2026.

  8. arXiv:2607.24017  [pdf, ps, other

    cs.CV cs.AI

    Disentangling Semantic Attention from Structural Bias in the Attention Manifold

    Authors: Pengkun Jiao, Bin Zhu, Jingjing Chen, Yu-gang Jiang

    Abstract: The empirical success of attention mechanism in Multimodal Large Language Models (MLLMs) often obscures its inherent, subtle flaws. Specifically, MLLMs consistently exhibit disproportionate attention toward certain semantically uninformative visual tokens, a phenomenon termed "register" or "Visual Attention Sinks." While existing inference intervention methods attempt to identify these sink tokens… ▽ More

    Submitted 27 July, 2026; originally announced July 2026.

  9. arXiv:2607.22783  [pdf, ps, other

    eess.IV cs.CV

    JPEG AIC2026: A large-scale dataset for fine-grained assessment of image coding

    Authors: Mohsen Jenadeleh, Jon Sneyers, João Ascenso, Thomas Richter, Alexander Karabutov, Panqi Jia, Elena Alshina, Osamu Watanabe, António Pinheiro, Touradj Ebrahimi, Dietmar Saupe

    Abstract: Recent advances in conventional and learning-based image coding have increased the demand for benchmark datasets that support fine-grained assessment of compressed image quality, particularly for learning-based image compression methods. This paper introduces Assessment of Image Coding 2026 (AIC2026), a large-scale dataset for high-fidelity image compression containing 70 source images selected fr… ▽ More

    Submitted 24 July, 2026; originally announced July 2026.

  10. arXiv:2607.19437  [pdf, ps, other

    eess.IV cs.CV cs.MM

    Group-of-Latents: Perceptual Video Compression at Extreme Bitrates via Masked Latent Generative Modeling

    Authors: Shaokang Wang, Jinchang Xu, Peidong Jia, Zhijian Hao, Siyuan Qian, Fei Zhao, Rui Ma, Xiaozhu Ju, Jian Tang, Xiaodong Xie, Shanghang Zhang, Huizhu Jia

    Abstract: Most existing video compression algorithms follow a paradigm of transformation and quantization, optimizing the trade-off between distortion and bitrate. However, extremely low-bitrate compression remains an underexplored frontier where perceptual quality optimization under severely constrained coding resources has not been adequately addressed. In this paper, we propose a unified generative frame… ▽ More

    Submitted 21 July, 2026; originally announced July 2026.

  11. arXiv:2607.14739  [pdf, ps, other

    cs.CV cs.AI

    FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models

    Authors: Wei Li, Peijin Jia, Yuan Ma, Xuefeng Jiang, Titong Jiang, Sheng Sun, Yujian Li, Xin Wen, Han Hong, Zhikang Liu, Bailin Li, Kun Zhan

    Abstract: Vision-Language-Action (VLA) models have achieved impressive results in visuomotor policy learning, yet remain fundamentally reactive, mapping current observations and language to actions without explicit forward prediction of world dynamics. Existing visual foresight methods predict future visual states but lack explicit motion guidance: they show where to go but not how to get there. We argue th… ▽ More

    Submitted 16 July, 2026; originally announced July 2026.

    Comments: 12 pages, 7 figures, 8 tables. Project page: https://liauto-research.github.io/FoMoVLA

  12. arXiv:2606.25319  [pdf, ps, other

    cs.CV

    V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning

    Authors: Haoxiang Sun, Zhihang Yi, Langxuan Deng, Yuhao Zhou, Peiqi Jia, Jian Zhao, Li Yuan, Jiancheng Lv, Tao Wang

    Abstract: Fine-grained visual reasoning requires multimodal large language models (MLLMs) to identify task-relevant visual evidence and ground their reasoning in local image regions. Existing agentic methods typically rely on reinforcement learning with verifiable rewards or supervised fine-tuning on large-scale annotated reasoning traces, leading to costly exploration, hand-designed verification rules, or… ▽ More

    Submitted 23 June, 2026; originally announced June 2026.

  13. arXiv:2606.23685  [pdf, ps, other

    cs.RO

    LaST-HD: Learning Latent Physical Reasoning from Scalable Human Data for Robot Manipulation

    Authors: Jiaming Liu, Yinxi Wang, Chenyang Gu, Siyuan Qian, Xiangju Mi, Hao Chen, Jiawei Chen, Qingpo Wuwu, Xiaoqi Li, Nuowei Han, Yiming Zhang, Xuheng Zhang, Yang Yue, Yeqing Yang, Lei Wang, Peng Jia, Hao Tang, Shanghang Zhang

    Abstract: Human-hand demonstrations provide a direct and scalable source of physical interaction data for robot learning. While manual retargeting is indispensable for establishing kinematic action correspondence across different morphologies, robust transfer requires going beyond geometry to address the underlying alignment of physical dynamics between human and robot manipulation. To address this, we intr… ▽ More

    Submitted 22 June, 2026; originally announced June 2026.

  14. arXiv:2606.23296  [pdf, ps, other

    cs.RO

    IOI: Decoupling Kinematics and Physics for Interactive World Models

    Authors: Chengyu Bai, Peidong Jia, Tiecheng Guo, Yukai Wang, Rui Ma, Fangyuan Zhao, Chunkai Fan, Xiaobao Wei, Jintao Chen, Hao Wang, Ying Li, Xiaozhu Ju, Jian Tang, Shanghang Zhang

    Abstract: Developing generalist embodied agents requires interactive environments providing visually realistic feedback and accurate action-conditioned dynamics. Interactive world models address this by simulating such complex dynamics. However, purely data-driven methods struggle to ensure precise control alignment and physically plausible visual feedback due to a lack of explicit structural constraints. T… ▽ More

    Submitted 22 June, 2026; originally announced June 2026.

  15. arXiv:2606.21088  [pdf, ps, other

    cs.RO

    MV-WAM: Manifold-Aware World Action Model with Value Augmentation

    Authors: Jintao Chen, Peidong Jia, Qingpo Wuwu, Jiaming Liu, Mengfei Du, Chun-Kai Fan, Xiaowei Chi, Hao Chen, Chengyu Bai, Zezhong Qian, Hao Wang, Jiajun Cao, Weishi Mi, Xiaozhu Ju, Jian Tang, Shanghang Zhang

    Abstract: Achieving robust and generalizable manipulation across diverse environments remains a fundamental challenge in embodied robotics. Recent world action models achieve strong in-domain performance, yet their gains do not extend proportionally to out-of-distribution scenarios. We attribute this to a structural mismatch between visual and action modalities, whose intrinsically heterogeneous manifolds c… ▽ More

    Submitted 19 June, 2026; originally announced June 2026.

    Comments: 20 pages, 9 figures, 7 tables

  16. arXiv:2606.20698  [pdf, ps, other

    cs.RO

    SafeDojo: Safe Reinforcement Learning for VLA via Interactive World Model

    Authors: Kai Tang, Peidong Jia, Zhong Chu, Jixian Wu, Rui Ma, Jiajun Cao, Fangyuan Zhao, Sixiang Chen, Yichen Guo, Xiaowei Chi, Chun-Kai Fan, Kevin Zhang, Jinchang Xu, Fubing Yang, Weishi Mi, Xiaozhu Ju, Jian Tang, Shanghang Zhang

    Abstract: Safe control is a prerequisite for real-world embodied intelligence, for which safe reinforcement learning has emerged as a promising paradigm. However, existing safe reinforcement learning methods either require costly real-world exploration or depend on hand-crafted safety functions. Neither scales to vision-language-action models deployed in open-world physical environments. We propose SafeDojo… ▽ More

    Submitted 15 June, 2026; originally announced June 2026.

    Comments: 20 pages, 5 figures, 8 tables

  17. arXiv:2606.14048  [pdf, ps, other

    cs.CV cs.RO

    WAM4D: Fast 4D World Action Model via Spatial Register Tokens

    Authors: Ying Li, Xiaobao Wei, Jiajun Cao, Hao Wang, Xiaowei Chi, Chengyu Bai, Qianpu Sun, Jiajun Li, Xiaojie Zhang, Peidong Jia, Jian Tang, Sirui Han, Shanghang Zhang

    Abstract: World action models (WAMs) have recently shown promise in jointly modeling future observations and executable robot actions. However, most existing WAMs still operate in 2D video or latent spaces, where visually plausible rollouts miss the 3D spatial constraints and occluded contact geometry required for precise manipulation. While geometric foundation models offer strong priors for recovering den… ▽ More

    Submitted 7 July, 2026; v1 submitted 11 June, 2026; originally announced June 2026.

    Comments: 15 pages, 7figures, 9tables

  18. arXiv:2606.12495  [pdf, ps, other

    cs.SD

    Missing-Token Prompted Reliability-Aware Fusion for Robust Polyglot Speaker Identification

    Authors: Peng Jia, Li Dai, Jia Li, Zhenzhen Hu, Ye Zhao, Richang Hong

    Abstract: Accurate and robust multimodal speaker identification is essential for multimedia understanding and biometric authentication. However, real-world polyglot scenarios pose two key challenges: speaker-discriminative representations should generalize across languages, and the model should remain reliable when face information is unavailable. To address these challenges, we propose MRAF, a Missing-Toke… ▽ More

    Submitted 10 June, 2026; originally announced June 2026.

    Comments: 8 pages, 3 figures, 4 tables

  19. arXiv:2606.11074  [pdf, ps, other

    cs.CL cs.AI

    Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models

    Authors: Peiqi Jia, Haonan Jia, Ziqi Miao, Linkang Du, Yuntao Wang, Zhou Su

    Abstract: With the widespread deployment of Multimodal Large Language Models (MLLMs) in social interaction, understanding and controlling their behavior under complex personality conditions is essential. This paper introduces explicit personality conditioning and establishes a systematic evaluation framework encompassing single-personality induction, multi-personality induction, and personality switching. E… ▽ More

    Submitted 9 June, 2026; v1 submitted 9 June, 2026; originally announced June 2026.

    Comments: 16 pages, 4 figures, 10 tables

  20. arXiv:2606.08737  [pdf, ps, other

    cs.RO

    Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation

    Authors: Yunfan Lou, Yifan Ye, Yankai Fu, Jun Cen, Xiaowei Chi, Yaoxu Lyu, Peidong Jia, Sirui Han, Zhihe Lu, Shanghang Zhang

    Abstract: World action models inherit the predictive capability of world models, enabling action generation to be guided by anticipated future observations. However, they rely primarily on vision and often fail in contact-rich manipulation, where critical cues arise from physical interaction. In this paper, we propose Dream-Tac, a unified Tactile-World Action Model that jointly models actions, future visual… ▽ More

    Submitted 7 June, 2026; originally announced June 2026.

    Comments: 16 pages,13 figures

  21. arXiv:2605.30195  [pdf

    cond-mat.mtrl-sci cs.AI cs.LG

    What drives performance in molecular MPNNs? An operator-level factorial benchmark

    Authors: Panyu Jiao, Shuizhou Chen, Yiheng Shen, Yuyang Wang, Runhai Ouyang, Wei Xie

    Abstract: Message-passing neural networks (MPNNs) are widely used for molecular property prediction, but their deployment as monolithic architectures makes it difficult to identify how specific message-passing operators affect performance. We present an operator-level factorial benchmark that decomposes 2D molecular MPNNs into the three families of message-seed initialization, node-edge fusion, and node upd… ▽ More

    Submitted 28 May, 2026; originally announced May 2026.

  22. TGFormer: Towards Temporal Graph Transformer with Auto-Correlation Mechanism

    Authors: Hongjiang Chen, Pengfei Jiao, Ming Du, Xuan Guo, Zhidong Zhao, Di Jin, Xiao Liu

    Abstract: The growing interest in Temporal Graph Neural Networks (TGNNs) stems from their ability to model complex dynamics and deliver superior performance. However, TGNNs encounter fundamental challenges in capturing long-term dependencies and identifying periodic patterns. To address these limitations, we propose TGFormer, a novel Transformer architecture specifically designed for temporal graphs. Our mo… ▽ More

    Submitted 24 May, 2026; originally announced May 2026.

    Journal ref: Pattern Recognition 170 (2026): 112053

  23. arXiv:2605.21482  [pdf, ps, other

    cs.AI

    DeepWeb-Bench: A Deep Research Benchmark Demanding Massive Cross-Source Evidence and Long-Horizon Derivation

    Authors: Sixiong Xie, Zhuofan Shi, Haiyang Shen, Jiuzheng Wang, Siqi Zhong, Mugeng Liu, Chongyang Pan, Peilun Jia, Baoqing Sun, Xiang Jing, Yun Ma

    Abstract: Deep research, in which an agent searches the open web, collects evidence, and derives an answer through extended reasoning, is a prominent use case for frontier language models. Frontier deep research products score high on existing benchmarks, making it difficult to distinguish their capabilities from current evaluation data alone. We introduce DeepWeb-Bench, a deep research benchmark that is su… ▽ More

    Submitted 20 May, 2026; originally announced May 2026.

    Comments: Work in Progress. 27 pages, 10 figures, 4 tables. Project page: https://sixiongxie1001-dot.github.io/deep-research-benchmark2.0

  24. arXiv:2605.19957  [pdf, ps, other

    cs.CV cs.AI cs.RO

    World-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied Tasks

    Authors: Zuyao Lin, Jianhui Zhang, Peidong Jia, Xiaoguang Zhao, Shanghang Zhang, Xingyu Chen

    Abstract: World models are widely explored in embodied intelligence, yet they typically predict distinct evolutions of the world and the ego within a single stream, where the world captures persistent instruction-agnostic scene regularities and the ego captures robot-centric instruction-conditioned dynamics. This world-ego entanglement leads to a degradation in long-horizon embodied scenarios, particularly… ▽ More

    Submitted 19 May, 2026; originally announced May 2026.

  25. arXiv:2605.19822  [pdf, ps, other

    cs.LG cs.AI

    ST-TGExplainer: Disentangling Stability and Transition Patterns for Temporal GNN Interpretability

    Authors: Hongjiang Chen, Xin Zheng, Pengfei Jiao, Huan Liu, Zhidong Zhao, Huaming Wu, Feng Xia, Shirui Pan

    Abstract: Temporal graph neural networks (TGNNs) have gained significant traction for solving real-world temporal graph tasks. However, their interpretability remains limited, as most TGNNs fail to identify which historical interactions most influence a given prediction. Despite promising progress on interpretable TGNNs, existing methods predominantly focus on previously seen historical interactions, which… ▽ More

    Submitted 19 May, 2026; originally announced May 2026.

  26. arXiv:2605.17904  [pdf, ps, other

    cs.CV

    Beyond Euclidean Prototypes: Spectral Disentanglement and Geodesic Matching for Few-Shot Medical Image Segmentation

    Authors: Penghao Jia, Zhiyong Huang, Mingyang Hou, Zhi Yu, Shuai Miao, Jiahong Wang, Yan Yan

    Abstract: Few-Shot Medical Image Segmentation (FSMIS) aims to delineate novel anatomical targets from one or a few annotated support images, addressing the annotation scarcity in medical imaging. Notwithstanding recent advancements, current prototype-based methods are bottlenecked by two coupled limitations: 1) cue entanglement, where a single spatial-domain prototype is forced to summarise organ silhouette… ▽ More

    Submitted 18 May, 2026; originally announced May 2026.

  27. arXiv:2605.15577  [pdf, ps, other

    physics.ins-det hep-ex

    A Novel Segment-Based Tracking Algorithm for HLT under High-Occupancy and Complex Conditions

    Authors: Pengkun Jia, Zhujun Fang, Hang Zhou, Yuhe Huang, Changqing Feng, Jianbei Liu

    Abstract: In the High-Level Trigger (HLT) of both electron-positron and hadron collision experiments, the tracking process for large-volume gaseous detectors typically consumes a latency of hundreds of milliseconds. Upgrades of existing experiments and the development of next-generation facilities demand enhanced HLT tracking performance: handling higher detector occupancy and suppressing latency. To addres… ▽ More

    Submitted 14 May, 2026; originally announced May 2026.

  28. arXiv:2605.13034  [pdf, ps, other

    cs.CV cs.IR

    ViDR: Grounding Multimodal Deep Research Reports in Source Visual Evidence

    Authors: Zhuofan Shi, Peilun Jia, Baoqin Sun, Haiyang Shen, Sixiong Xie, Yun Ma, Xiang Jing

    Abstract: Recent deep research systems have improved the ability of large language models to produce long, grounded reports through iterative retrieval and reasoning. However, most text-centered systems rely mainly on textual evidence, while multimodal systems often retrieve images only weakly or generate charts themselves, leaving source figures underused as evidence. We present ViDR, a multimodal deep res… ▽ More

    Submitted 13 May, 2026; originally announced May 2026.

  29. arXiv:2605.10942  [pdf, ps, other

    cs.RO

    HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models

    Authors: Qiuxuan Feng, Jiale Yu, Jiaming Liu, Yueru Jia, Zhuangzhe Wu, Hao Chen, Zezhong Qian, Shuo Gu, Peng Jia, Siwei Ma, Shanghang Zhang

    Abstract: World Action Models (WAMs) have emerged as a promising paradigm for robot control by modeling physical dynamics. Current WAMs generally follow two paradigms: the "Imagine-then-Execute" approach, which uses video prediction to infer actions via inverse dynamics, and the "Joint Modeling" approach, which jointly models actions and video representations. Based on systematic experiments, we observe a f… ▽ More

    Submitted 11 May, 2026; originally announced May 2026.

  30. arXiv:2605.10530  [pdf, ps, other

    cs.IR

    Personalized Deep Research: A User-Centric Framework, Dataset, and Hybrid Evaluation for Knowledge Discovery

    Authors: Xiaopeng Li, Wenlin Zhang, Yingyi Zhang, Pengyue Jia, Yejing Wang, Yichao Wang, Yong Liu, Huifeng Guo, Xiangyu Zhao

    Abstract: Deep Research agents driven by LLMs have automated the scholarly discovery pipeline, from planning and query formulation to iterative web exploration. Yet they remain constrained by a static, ``one-size-fits-all'' retrieval paradigm. Current systems fail to adaptively adjust the depth and breadth of exploration based on the user's existing expertise or latent interests, frequently resulting in rep… ▽ More

    Submitted 11 May, 2026; originally announced May 2026.

    Comments: Accepted to SIGIR 2026

  31. arXiv:2605.10527  [pdf, ps, other

    cs.IR

    UniRank: Unified List-wise Reranking via Confidence-Ordered Denoising

    Authors: Pengyue Jia, Hailan Yang, Shuchang Liu, Xiaobei Wang, Wanyu Wang, Xiang Li, Yongqi Liu, Kaiqiao Zhan, Kun Gai, Xiangyu Zhao

    Abstract: List-wise reranking arranges a request-specific pool of candidate items into an ordered slate that maximizes user satisfaction. Existing generative rerankers fall into two paradigms: Autoregressive (AR) rerankers construct the slate left to right and capture inter-item dependencies in the exposure list, but they suffer from error propagation because early mistakes affect subsequent slots. Non-auto… ▽ More

    Submitted 11 May, 2026; originally announced May 2026.

  32. arXiv:2605.10179  [pdf, ps, other

    cs.LG cs.AI

    One-Step Graph-Structured Neural Flows for Irregular Multivariate Time Series Classification

    Authors: Mengzhou Gao, Kaiwei Wang, Pengfei Jiao

    Abstract: Neural Flows efficiently model irregular multivariate time series by directly learning ODE solution trajectories with neural networks, bypassing step-by-step numerical solvers. Despite their efficiency, many existing approaches treat variables independently, leaving inter-variable interactions underexplored. Moreover, their one-step mapping makes interaction modeling inherently challenging, as it… ▽ More

    Submitted 11 May, 2026; originally announced May 2026.

  33. arXiv:2605.09956  [pdf, ps, other

    cs.CV cs.AI

    SDTalk: Structured Facial Priors and Dual-Branch Motion Fields for Generalizable Gaussian Talking Head Synthesis

    Authors: Peng Jia, Zhen Xiao, Jia Li, Xueliang Liu, Zhenzhen Hu, Lingyun Yu

    Abstract: High-quality, real-time talking head synthesis remains a fundamental challenge in computer vision. Existing reconstruction- and rendering-based methods typically rely on identity-specific models, limiting cross-identity generalization. To address this issue, we propose SDTalk, a one-shot 3D Gaussian Splatting (3DGS)-based framework that generalizes to unseen identities without personalized trainin… ▽ More

    Submitted 11 May, 2026; originally announced May 2026.

    Comments: 5 pages, 4 figures, 4 tables

  34. arXiv:2605.00702  [pdf, ps, other

    cs.CL

    Learning How and What to Memorize: Cognition-Inspired Two-Stage Optimization for Evolving Memory

    Authors: Derong Xu, Shuochen Liu, Pengfei Luo, Pengyue Jia, Yingyi Zhang, Yi Wen, Yimin Deng, Wenlin Zhang, Enhong Chen, Xiangyu Zhao, Tong Xu

    Abstract: Large language model (LLM) agents require long-term user memory for consistent personalization, but limited context windows hinder tracking evolving preferences over long interactions. Existing memory systems mainly rely on static, hand-crafted update rules; although reinforcement learning (RL)-based agents learn memory updates, sparse outcome rewards provide weak supervision, resulting in unstabl… ▽ More

    Submitted 1 May, 2026; originally announced May 2026.

  35. arXiv:2604.28192  [pdf, ps, other

    cs.RO cs.CV

    LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning

    Authors: Hao Chen, Jiaming Liu, Zhonghao Yan, Nuowei Han, Renrui Zhang, Chenyang Gu, Jialin Gao, Ziyu Guo, Siyuan Qian, Yinxi Wang, Peng Jia, Shanghang Zhang, Pheng-Ann Heng

    Abstract: Robotic foundation models require reasoning over complex visual scenes to execute adaptive actions in dynamic environments. While recent studies on latent-reasoning Vision-Language-Action (VLA) models have demonstrated the capability to capture fine-grained physical dynamics, they remain predominantly confined to static imitation learning, severely limiting their adaptability and generalization. I… ▽ More

    Submitted 7 May, 2026; v1 submitted 30 April, 2026; originally announced April 2026.

  36. arXiv:2604.25291  [pdf, ps, other

    cs.IR

    From Local Indices to Global Identifiers: Generative Reranking for Recommender Systems via Global Action Space

    Authors: Pengyue Jia, Xiaobei Wang, Yingyi Zhang, Shuchang Liu, Yupeng Hou, Hailan Yang, Xu Gao, Xiaopeng Li, Yejing Wang, Julian McAuley, Xiang Li, Lantao Hu, Yongqi Liu, Kaiqiao Zhan, Han Li, Kun Gai, Xiangyu Zhao

    Abstract: In modern recommender systems, list-wise reranking serves as a critical phase within the multi-stage pipeline, finalizing the exposed item sequence and directly impacting user satisfaction by modeling complex intra-list item dependencies. Existing methods typically formulate this task as selecting indices from the local input list. However, this approach suffers from a semantically inconsistent ac… ▽ More

    Submitted 28 April, 2026; originally announced April 2026.

  37. Precision extraction of the deuteron electric polarizability via the Baldin sum rule with full low-energy coverage

    Authors: Zi-Rui Hao, Gong-Tao Fan, Qian-Kun Sun, Hong-Wei Wang, Hang-Hua Xu, Long-Xiang Liu, Yue Zhang, Jiunn-Wei Chen, Yu-Xuan Yang, Sheng Jin, Kai-Jie Chen, Zhen-Wei Wang, Xiang-Fei Wang, Meng-Ke Xu, Zhi-Cai Li, Pu Jiao, Meng-Die Zhou, Shan Ye, Yu-Long Shen, Yin-Ji Chen, Hao Zhang, Jian-Jun He, Wen-Qing Shen, Yu-Gang Ma

    Abstract: The photodisintegration cross sections of the deuteron have been systematically measured over the photon energy range of 2.33-19.65 MeV at the Shanghai Laser Electron Gamma Source (SLEGS). By applying the well-established Baldin sum rule to the newly obtained data, the sum of the electric and magnetic dipole polarizabilities of the deuteron is extracted for the first time based solely on a dense a… ▽ More

    Submitted 27 April, 2026; originally announced April 2026.

  38. arXiv:2604.24186  [pdf, ps, other

    cs.CL cs.AI

    MultiDx: A Multi-Source Knowledge Integration Framework towards Diagnostic Reasoning

    Authors: Yimin Deng, Zhenxi Lin, Yejing Wang, Guoshuai Zhao, Pengyue Jia, Zichuan Fu, Derong Xu, Yefeng Zheng, Xiangyu Zhao, Li Zhu, Xian Wu, Xueming Qian

    Abstract: Diagnostic prediction and clinical reasoning are critical tasks in healthcare applications. While Large Language Models (LLMs) have shown strong capabilities in commonsense reasoning, they still struggle with diagnostic reasoning due to limited domain knowledge. Existing approaches often rely on internal model knowledge or static knowledge bases, resulting in knowledge insufficiency and limited ad… ▽ More

    Submitted 27 April, 2026; originally announced April 2026.

    Comments: ACL 2026 findings

  39. arXiv:2604.17873  [pdf, ps, other

    cs.CV

    Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models

    Authors: Ziyao Tang, Pengkun Jiao, Bin Zhu, Huiyan Qi, Jingjing Chen, Yu-Gang Jiang

    Abstract: Video Large Language Models (Vid-LLMs) have demonstrated remarkable performance in video understanding tasks, yet their robustness under conversational interaction remains largely underexplored. In this paper, we identify spatiotemporal sycophancy, a failure mode in which Vid-LLMs retract initially correct, visually grounded judgments and conform to misleading user feedback under negation-based ga… ▽ More

    Submitted 20 April, 2026; originally announced April 2026.

  40. arXiv:2604.17265  [pdf, ps, other

    cs.IR

    MemSearch-o1: Empowering Large Language Models with Reasoning-Aligned Memory Growth in Agentic Search

    Authors: Sheng Zhang, Junyi Li, Yingyi Zhang, Pengyue Jia, Yichao Wang, Xiaowei Qian, Wenlin Zhang, Maolin Wang, Yong Liu, Xiangyu Zhao

    Abstract: Recent advances in large language models (LLMs) have scaled the potential for reasoning and agentic search, wherein models autonomously plan, retrieve, and reason over external knowledge to answer complex queries. However, the iterative think-search loop accumulates long system memories, leading to memory dilution problem. In addition, existing memory management methods struggle to capture fine-gr… ▽ More

    Submitted 12 May, 2026; v1 submitted 19 April, 2026; originally announced April 2026.

  41. arXiv:2604.08363  [pdf, ps, other

    cs.SD

    CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation

    Authors: Xiaosu Su, Zihan Sun, Peilei Jia, Jun Gao

    Abstract: Voice design from natural language descriptions is emerging as a new task in text-to-speech multimodal generation, aiming to synthesize speech with target timbre and speaking style without relying on reference audio. However, existing methods mainly focus on single-utterance generation, leaving conversational voice design largely unexplored. In this work, we extend voice design to dialogue, enabli… ▽ More

    Submitted 9 April, 2026; originally announced April 2026.

    Comments: 14 pages, 2 figures

  42. arXiv:2604.07146  [pdf, ps, other

    cs.CV

    Learning to Search: A Decision-Based Agent for Knowledge-Based Visual Question Answering

    Authors: Zhuohong Chen, Zhenxian Wu, Yunyao Yu, Hangrui Xu, Zirui Liao, Zhifang Liu, Xiangwen Deng, Pen Jiao, Haoqian Wang

    Abstract: Knowledge-based visual question answering (KB-VQA) requires vision-language models to understand images and use external knowledge, especially for rare entities and long-tail facts. Most existing retrieval-augmented generation (RAG) methods adopt a fixed pipeline that sequentially retrieves information, filters it, and then produces an answer. Such a design makes it difficult to adapt to diverse q… ▽ More

    Submitted 9 April, 2026; v1 submitted 8 April, 2026; originally announced April 2026.

  43. arXiv:2603.24376  [pdf, ps, other

    cs.CV

    GeoRouter: Dynamic Paradigm Routing for Worldwide Image Geolocalization

    Authors: Pengyue Jia, Derong Xu, Yingyi Zhang, Xiaopeng Li, Wenlin Zhang, Yi Wen, Yuanshao Zhu, Xiangyu Zhao

    Abstract: Worldwide image geolocalization aims to predict precise GPS coordinates for images captured anywhere on Earth, which is challenging due to the large visual and geographic diversity. Recent methods mainly follow two paradigms: retrieval-based approaches that match queries against a reference database, and generation-based approaches that directly predict coordinates using Large Vision-Language Mode… ▽ More

    Submitted 25 March, 2026; originally announced March 2026.

  44. arXiv:2603.21155  [pdf, ps, other

    cs.AI

    Can LLMs Fool Graph Learning? Exploring Universal Adversarial Attacks on Text-Attributed Graphs

    Authors: Zihui Chen, Yuling Wang, Pengfei Jiao, Kai Wu, Xiao Wang, Xiang Ao, Dalin Zhang

    Abstract: Text-attributed graphs (TAGs) enhance graph learning by integrating rich textual semantics and topological context for each node. While boosting expressiveness, they also expose new vulnerabilities in graph learning through text-based adversarial surfaces. Recent advances leverage diverse backbones, such as graph neural networks (GNNs) and pre-trained language models (PLMs), to capture both struct… ▽ More

    Submitted 22 March, 2026; originally announced March 2026.

    Comments: Accepted by TheWebConf (WWW) 2026

  45. arXiv:2603.19677  [pdf, ps, other

    cs.LG cs.AI cs.MA

    GoAgent: Group-of-Agents Communication Topology Generation for LLM-based Multi-Agent Systems

    Authors: Hongjiang Chen, Xin Zheng, Yixin Liu, Pengfei Jiao, Shiyuan Li, Huan Liu, Zhidong Zhao, Ziqi Xu, Ibrahim Khalil, Shirui Pan

    Abstract: Large language model (LLM)-based multi-agent systems (MAS) have demonstrated exceptional capabilities in solving complex tasks, yet their effectiveness depends heavily on the underlying communication topology that coordinates agent interactions. Within these systems, successful problem-solving often necessitates task-specific group structures to divide and conquer subtasks. However, most existing… ▽ More

    Submitted 20 March, 2026; originally announced March 2026.

  46. arXiv:2603.16245  [pdf, ps, other

    cs.CV cs.CL

    How to Utilize Complementary Vision-Text Information for 2D Structure Understanding

    Authors: Jiancheng Dong, Pengyue Jia, Derong Xu, Jiawei Cheng, Jingyu Peng, Chao Zhang, Bowen Liu, Xin Sun, Lixin Su, Shuaiqiang Wang, Dawei Yin, Xiangyu Zhao

    Abstract: LLMs typically linearize 2D tables into 1D sequences to fit their autoregressive architecture, which weakens row-column adjacency and other layout cues. In contrast, purely visual encoders can capture spatial cues, yet often struggle to preserve exact cell text. Our analysis reveals that these two modalities provide highly distinct information to LLMs and exhibit strong complementarity. However, d… ▽ More

    Submitted 17 March, 2026; originally announced March 2026.

    Comments: 16 pages, 5 figures

  47. arXiv:2603.15618  [pdf, ps, other

    cs.CV

    Look Before Acting: Enhancing Vision Foundation Representations for Vision-Language-Action Models

    Authors: Yulin Luo, Hao Chen, Zhuangzhe Wu, Bowen Sui, Jiaming Liu, Chenyang Gu, Zhuoyang Liu, Qiuxuan Feng, Jiale Yu, Shuo Gu, Peng Jia, Pheng-Ann Heng, Shanghang Zhang

    Abstract: Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for robotic manipulation, in which reliable action prediction critically depends on accurately interpreting and integrating visual observations conditioned on language instructions. Although recent works have sought to enhance the visual capabilities of VLA models, most approaches treat the LLM backbone as a black bo… ▽ More

    Submitted 17 March, 2026; v1 submitted 16 March, 2026; originally announced March 2026.

  48. arXiv:2603.10363  [pdf

    cond-mat.mes-hall

    Symmetry Breaking and Transition to Robust Excitonic Topological Order in InAs/GaSb Bilayers

    Authors: Xinghao Wang, Wenfeng Zhang, Yujiang Dong, Weiliang Qiao, Peizhe Jia, Rui-Rui Du

    Abstract: Symmetry and topology are fundamental concepts deeply intertwined in various fields of physics, especially in the studies of quantum phases of matter. The critical role that Coulomb interactions play in symmetry breaking during topological transitions is a fundamental problem that has not been fully understood. Utilizing gated indium arsenide-gallium antimonide bilayers, we demonstrate that Coulom… ▽ More

    Submitted 10 March, 2026; originally announced March 2026.

    Comments: 19 pages, 5 figures

  49. arXiv:2603.09250  [pdf, ps, other

    cs.IR

    Evoking User Memory: Personalizing LLM via Recollection-Familiarity Adaptive Retrieval

    Authors: Yingyi Zhang, Junyi Li, Wenlin Zhang, Penyue Jia, Xianneng Li, Yichao Wang, Derong Xu, Yi Wen, Huifeng Guo, Yong Liu, Xiangyu Zhao

    Abstract: Personalized large language models (LLMs) rely on memory retrieval to incorporate user-specific histories, preferences, and contexts. Existing approaches either overload the LLM by feeding all the user's past memory into the prompt, which is costly and unscalable, or simplify retrieval into a one-shot similarity search, which captures only surface matches. Cognitive science, however, shows that hu… ▽ More

    Submitted 10 March, 2026; originally announced March 2026.

    Comments: Accepted by ICLR 2026

  50. arXiv:2603.08063  [pdf, ps, other

    cs.CV

    SkyLink: A Large Vision-Language Model Driven Re-ranking Framework for Cross-View UAV geolocalization

    Authors: Bowen Liu, Pengyue Jia, Wanyu Wang, Derong Xu, Jiawei Cheng, Jiancheng Dong, Xiao Han, Zimo Zhao, Chao Zhang, Bowen Yu, Fangyu Hong, Xiangyu Zhao

    Abstract: Cross-view UAV geolocalization is fundamentally a challenging large-scale image retrieval task, aiming to determine the geographic coordinates of Unmanned Aerial Vehicle (UAV) queries by matching them against an extensive geo-tagged satellite image database. Most existing methods learn separate feature representations for each view and determine the final prediction using naive heuristics to asses… ▽ More

    Submitted 15 May, 2026; v1 submitted 9 March, 2026; originally announced March 2026.