Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 145 results for author: Yin, G

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.05159  [pdf

    cs.AI cs.CL

    Agentic Nesting: A New Methodology for Existing Enterprise Application Integration and Services

    Authors: Xi Wang, Kun Li, Xianyao Ling, Gang Yin, Liang Zhang, Jiang Wu, Wenbo Lei, Jun Xu, Annie Wang, Fu Zhang, Weizhe Wang

    Abstract: Enterprise operations extensively rely on multiple heterogeneous business systems and information applications, which also result in severe data silos and process fragmentation. Enterprises have invested considerable financial and material resources in building these applications, however, effectively leveraging and orchestrating them remains a formidable challenge. Conventional approaches to ente… ▽ More

    Submitted 25 May, 2026; originally announced August 2026.

  2. arXiv:2608.01743  [pdf, ps, other

    cs.LG cs.CL

    Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning

    Authors: Li Wang, Xiaodong Lu, Xiaohan Wang, Jiajun Chai, Wei Lin, Tianhao Peng, Guojun Yin

    Abstract: Reinforcement learning (RL) has become a central paradigm for large language model (LLM) post-training, but optimization toward new objectives can degrade capabilities already present in the base model. KL regularization is widely used to mitigate such forgetting by constraining policy drift toward a reference model. However, standard full-policy KL regularization constrains the entire response di… ▽ More

    Submitted 3 August, 2026; originally announced August 2026.

  3. arXiv:2607.26017  [pdf, ps, other

    cs.CL

    UniMem: Complementary Episodic-to-Parametric Memory for Boundary-Agnostic Task Streams

    Authors: Siyu Xia, Chenheng Zhang, Yanting Wu, Haoxuan Li, Jiajun Chai, Xiaohan Wang, Guojun Yin, Wei Lin, Zhouchen Lin, Haifeng Zhang, Jun Wang

    Abstract: Memory is essential for LLM agents to accumulate task experience and reuse task-specific execution strategies. However, real-world deployment over boundary-agnostic and evolving task streams exposes a fundamental stability-plasticity dilemma. External retrieval-based memory can rapidly absorb new evidence, but it often fails to internalize recurring execution patterns and incurs inference-time ret… ▽ More

    Submitted 28 July, 2026; originally announced July 2026.

  4. arXiv:2607.25110  [pdf, ps, other

    cs.IR cs.LG

    Memory Layer: Train the In-Model Cache for Recommendation Models

    Authors: Liangyuan Na, Gufan Yin, Yixin Bao, Xianjie Chen, Justin Lin, Ziheng huang, Xinyuan Zhang, Wen Zhang, Hao Lin, Xiaoheng Mao, Shuo Tang, Min Yu, Lei Chen, Chao yang, Ziliang Zhao, Mengjiao Zhou, Zheng Qi, Dmitry Barablin, Chuo-Yun Yang, Kaustubh Vartak, Tingting Zhang, Arun Kumar Singh

    Abstract: Early ranking stages in recommendation systems precompute item embeddings and cache them in-model for scoring within strict latency constraints. Because this cache exists only at serving time, outside the training loop, training and serving use different item representations, a structural discrepancy that limits quality and adds operational fragility. We show that co-designing the training and ser… ▽ More

    Submitted 27 July, 2026; originally announced July 2026.

  5. arXiv:2607.19606  [pdf

    physics.geo-ph cs.AI

    Juxtaposition of Shallow Reservoir-Triggered Seismicity and Deep Tectonic Locking in the Qiaojia-Dongchuan Seismic Gap

    Authors: Yuxin Zhou, Huai Zhang, S. Mostafa Mousavi, Guangyao Yin, Pei He, Yicun Guo, Shuang Yi, Yaolin Shi

    Abstract: Identifying the critical state of mature seismic gaps is challenging, especially when anthropogenic stress perturbations, such as reservoir impoundment, superimpose on tectonic loading. Here, utilizing a high-resolution dense array catalog from the Qiaojia-Dongchuan seismic gap (hosting the second-largest hydropower station in the world), we reveal a distinct vertical decoupling mechanism. The sha… ▽ More

    Submitted 21 July, 2026; originally announced July 2026.

  6. arXiv:2607.12281  [pdf, ps, other

    cs.IR cs.LG

    SlimPer: Make Personalization Model Slim and Smart

    Authors: Siqi Wang, Xianjie Chen, Shaofeng Deng, Albert Chen, Romil Shah, Jiawei Huang, Zhaoqin Wang, Zhang Zhang, Yiqun Liu, Meilei Jiang, Anish Dubey, Moyan Mei, Tongxin Wang, Nathan Berrebbi, Misael Manjarres, Armand Sauzay, Shardul Kothapalli, Aryaman Vinchhi, Kevin Johnstone, Juheon Lee, Gufan Yin, Ziheng Huang, Justin Lin, Mert Terzihan, Yilin Qi , et al. (20 additional authors not shown)

    Abstract: Transformer-style architectures are increasingly adopted for industrial recommendation systems, yet they inherit a design premise misaligned with the task: generative models rely on per-token autoregressive prediction, which justifies maintaining large intermediate tensors that scale with sequence length. In contrast, recommendation systems produce a single set of relevance scores for each <user,… ▽ More

    Submitted 13 July, 2026; originally announced July 2026.

  7. arXiv:2606.09249  [pdf, ps, other

    cs.CV

    DECIS: Dual-Evidence Corrective Verification for Interpretable Strabismus Diagnostic Decision-Making

    Authors: Xikai Tang, Yifan Wang, Jiafan Zhuang, Li Luo, Jinming Guo, Xiaoling Xie, Jiacheng Liu, Peiwei Wei, Lihao Zhong, Xiaoli Kang, Jie Cen, Guangqiang Yin, Kunliang Qiu, Ce Zheng, Zhun Fan

    Abstract: Strabismus is a common ocular disorder that requires fine-grained subtype diagnosis for individualized treatment planning. However, existing deep learning methods mainly provide diagnostic predictions without transparent reasoning, while recent large vision-language models (LVLMs), although promising for joint image understanding and report generation, remain highly prone to hallucination in this… ▽ More

    Submitted 20 July, 2026; v1 submitted 8 June, 2026; originally announced June 2026.

  8. arXiv:2606.05784  [pdf, ps, other

    cs.AI

    TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents

    Authors: Chengqi Dong, Chuhuai Yue, Hang He, yandong liu, Fenghe Tang, S Kevin Zhou, Xiaohan Wang, Jiajun Chai, Guojun Yin

    Abstract: We identify and formally characterize credit misassignment as a systematic failure mode of GRPO in tool-augmented multimodal search agents: its uniform broadcast of trajectory-level advantages to all tokens causes valuable tool-use steps in failing trajectories to be penalized no differently from valueless ones. We further empirically quantify the scale of this phenomenon. Over half of failing tra… ▽ More

    Submitted 4 June, 2026; originally announced June 2026.

  9. arXiv:2606.03273  [pdf, ps, other

    cs.CV cs.AI cs.CL

    VistaHop: Benchmarking Long-Horizon Visual DeepSearch

    Authors: Hang He, Chuhuai Yue, Chengqi Dong, Chengcheng Wan, Ting Su, Haiying Sun, Jiajun Chai, Xiaohan Wang, Guojun Yin

    Abstract: Visual DeepSearch tasks require multimodal large language models (MLLMs) to resolve complex visual queries by repeatedly inspecting image regions, grounding reasoning in visual evidence, and connecting fine-grained clues across multiple steps. However, existing benchmarks primarily evaluate single-step visual understanding or isolated visual-query response generation. They have limited difficulty,… ▽ More

    Submitted 29 July, 2026; v1 submitted 2 June, 2026; originally announced June 2026.

  10. arXiv:2605.31490  [pdf, ps, other

    cs.CL

    Are Full Rollouts Necessary for On-Policy Distillation?

    Authors: Yaocheng Zhang, Jiajun Chai, Yuqian Fu, Songjun Tu, Xiaohan Wang, Wei Lin, Guojun Yin, Qichao Zhang, Yuanheng Zhu, Dongbin Zhao

    Abstract: On-policy distillation (OPD) provides dense teacher feedback along student-generated rollouts rather than fixed teacher traces and has emerged as a promising post-training paradigm. However, standard OPD typically generates full rollouts during training, which is computationally expensive and may expose the student to unreliable teacher feedback at late rollout positions, especially during early t… ▽ More

    Submitted 1 June, 2026; v1 submitted 29 May, 2026; originally announced May 2026.

    Comments: 15 pages, 14 figures

  11. arXiv:2605.28184  [pdf, ps, other

    cs.LG

    Joint Training of Multi-Token Prediction in Reinforcement Learning via Optimal Coefficient Calibration

    Authors: Zili Wang, Jiajun Chai, Lin Chen, Xiaohan Wang, Shiming Xiang, Guojun Yin

    Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as the standard paradigm for improving reasoning capability of large language models, while Multi-Token Prediction (MTP) has been a widely adopted module in pretraining. Combining them is a natural approach, yet current RL practices detach MTP gradients because joint training degrades the performance. We revisit this failure from an… ▽ More

    Submitted 27 May, 2026; originally announced May 2026.

  12. arXiv:2605.28069  [pdf, ps, other

    cs.AI

    ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay

    Authors: Zhexin Hu, Li Wang, Xiaohan Wang, Jiajun Chai, Xiaojun Guo, Wei Lin, Guojun Yin

    Abstract: Adaptive context compression is vital for scaling Large Language Models (LLMs) to complex, multi-turn agent tasks. However, rule-based compression methods may discard task-critical nuances, while Reinforcement Learning (RL) approaches usually struggle to balance information retention and token efficiency under the sparse rewards inherent to long-horizon workflows. To bridge this gap, we propose Zi… ▽ More

    Submitted 27 May, 2026; originally announced May 2026.

  13. arXiv:2605.25864  [pdf, ps, other

    cs.LG cs.CL

    When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards

    Authors: Li Wang, Xiaodong Lu, Xiaohan Wang, Yikun Ban, Jiajun Chai, Wei Lin, Tianhao Peng, Guojun Yin

    Abstract: Large Language Models (LLMs) have achieved remarkable advancements in reasoning capabilities empowered by Reinforcement Learning with Verifiable Rewards (RLVR). Nonetheless, RLVR intrinsically relies on ground-truth labels for reward computation, the acquisition of which is often prohibitively expensive in real-world scenarios. While unsupervised RLVR paradigms attempt to circumvent this by traini… ▽ More

    Submitted 25 May, 2026; originally announced May 2026.

  14. arXiv:2605.21083  [pdf, ps, other

    physics.app-ph cs.LG physics.bio-ph physics.med-ph

    AIMBio-Mat: An AI-Native FAIR Platform for Closed-Loop Materials Discovery and Biomedical Translation

    Authors: D. -M. Mei, K. Acharya, C. M. Adhikari, M. Adhikari, S. Aryal, B. V. Benson, K. Bhatta, S. Bhattarai, N. Budhathoki, A. M. Castillo, D. Chakraborty, S. Chhetri, S. Choudhury, T. A. Chowdhury, R. D. Cruz, B. Cui, S. Dhital, K. -M. Dong, R. Gapuz, A. Ghasemi, E. Z. Gnimpieba, B. D. S. Gurung, H. A. Hashim, R. I. Harry, K. -E. Hasin , et al. (29 additional authors not shown)

    Abstract: Materials discovery and biomedical translation increasingly require models that can reason across composition, processing, structure, biological response, manufacturability, safety, and governance constraints. Existing materials and biomedical data ecosystems are powerful but remain poorly coupled for AI-guided discovery. Here we present AIMBio, a conceptual framework for an AI-native, FAIR, and g… ▽ More

    Submitted 20 May, 2026; originally announced May 2026.

    Comments: 35 pages, 4 figures, and 12 tables

  15. arXiv:2605.18529  [pdf, ps, other

    cs.AI

    AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment

    Authors: Zhenlin Wei, Pu Jian, Yingzhuo Deng, Xiaohan Wang, Jiajun Chai, Zhexin Hu, Wei Lin, Shanbin Zhang, Guojun Yin

    Abstract: The alignment of Large Language Models (LLMs) for complex reasoning heavily relies on Reinforcement Learning with Verifiable Rewards (RLVR). However, standard algorithms like GRPO apply sequence-level rewards uniformly to all tokens, creating a severe credit-assignment bottleneck. While on-policy self-distillation attempts to resolve this by conditioning a self-teacher on privileged contexts, dire… ▽ More

    Submitted 18 May, 2026; originally announced May 2026.

  16. arXiv:2605.18500  [pdf, ps, other

    cs.CL

    Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning

    Authors: Li Wang, Xiaohan Wang, Xiaodong Lu, Zipeng Zhang, Jinyang Wu, Jiajun Chai, Wei Lin, Guojun Yin

    Abstract: Large language models (LLMs) have increasingly leveraged tool invocation to enhance their reasoning capabilities. However, existing approaches typically tightly couple tool invocation with immediate execution. Such immediate tool interaction may disrupt the reasoning coherence of LLMs and constrain their expressivity, ultimately degrading reasoning performance. To this end, for the first time, we… ▽ More

    Submitted 18 May, 2026; originally announced May 2026.

  17. arXiv:2605.11874  [pdf, ps, other

    cs.IR

    RecRM-Bench: Benchmarking Multidimensional Reward Modeling for Agentic Recommender Systems

    Authors: Wenwen Zeng, Jinhui Zhang, Hao Chen, Zhaoyu Hu, Yongqi Liang, Jiajun Chai, Dengcan Liu, Zhenfeng Liu, Shurui Yan, Minglong Xue, Xiaohan Wang, Wei Lin, Guojun Yin

    Abstract: The integration of Large Language Model (LLM) agents is transforming recommender systems from simple query-item matching towards deeply personalized and interactive recommendations. Reinforcement Learning (RL) provides an essential framework for the optimization of these agents in recommendation tasks. However, current methodologies remain limited by a reliance on single dimensional outcome-based… ▽ More

    Submitted 12 May, 2026; originally announced May 2026.

  18. arXiv:2605.08308  [pdf, ps, other

    cs.LG cs.AI eess.SP

    Practical Wi-Fi-based Motion Recognition Under Variable Traffic Patterns

    Authors: Guolin Yin, Junqing Zhang, Guanxiong Shen, Simon L. Cotton

    Abstract: Wi-Fi sensing detects human motions and activities by analysing the channel state information (CSI) derived from Wi-Fi transmissions. However, the impact of variable transmission traffic, which dictates the effective sampling rate and interval, is often overlooked. Existing Wi-Fi sensing systems are trained with fixed input size and sampling rate, which suffer from poor sampling rate generalisatio… ▽ More

    Submitted 8 May, 2026; originally announced May 2026.

    Comments: 17 Pages

  19. arXiv:2605.00380  [pdf, ps, other

    cs.LG cs.CL

    ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning

    Authors: Zihan Lin, Xiaohan Wang, Jie Cao, Jiajun Chai, Li Wang, Xiaodong Lu, Wei Lin, Ran He, Guojun Yin

    Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) enhances reasoning of Large Language Models (LLMs) but usually exhibits limited generation diversity due to the over-incentivization of positive rewards. Although methods like Negative Sample Reinforcement (NSR) mitigate this issue by upweighting penalty from negative samples, they may suppress the semantic distributions shared between positive… ▽ More

    Submitted 8 May, 2026; v1 submitted 30 April, 2026; originally announced May 2026.

    Comments: Accepted to ICML 2026. Preprint version. https://github.com/1229095296/ResRL.git

  20. arXiv:2604.17337  [pdf, ps, other

    cs.AI

    AutoSearch: Adaptive Search Depth for Efficient Agentic RAG via Reinforcement Learning

    Authors: Jingbo Sun, Wenyue Chong, Songjun Tu, Qichao Zhang, Yaocheng Zhang, Jiajun Chai, Xiaohan Wang, Wei Lin, Guojun Yin, Dongbin Zhao

    Abstract: Agentic retrieval-augmented generation (RAG) systems enable large language models (LLMs) to solve complex tasks through multi-step interaction with external retrieval tools. However, such multi-step interaction often involves redundant search steps, incurring substantial computational cost and latency. Prior work limits search depth (i.e., the number of search steps) to reduce cost, but this often… ▽ More

    Submitted 19 April, 2026; originally announced April 2026.

  21. arXiv:2604.14054  [pdf, ps, other

    cs.LG cs.CL

    $π$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data

    Authors: Yaocheng Zhang, Yuanheng Zhu, Wenyue Chong, Songjun Tu, Qichao Zhang, Jiajun Chai, Xiaohan Wang, Wei Lin, Guojun Yin, Dongbin Zhao

    Abstract: Deep search agents have emerged as a promising paradigm for addressing complex information-seeking tasks, but their training remains challenging due to sparse rewards, weak credit assignment, and limited labeled data. Self-play offers a scalable route to reduce data dependence, but conventional self-play optimizes students only through sparse outcome rewards, leading to low learning efficiency. In… ▽ More

    Submitted 25 May, 2026; v1 submitted 15 April, 2026; originally announced April 2026.

    Comments: 23 pages, 11 figures

  22. arXiv:2604.10425  [pdf, ps, other

    cs.CV

    DiningBench: A Hierarchical Multi-view Benchmark for Perception and Reasoning in the Dietary Domain

    Authors: Song Jin, Juntian Zhang, Xun Zhang, Zeying Tian, Fei Jiang, Guojun Yin, Wei Lin, Yong Liu, Rui Yan

    Abstract: Recent advancements in Vision-Language Models (VLMs) have revolutionized general visual understanding. However, their application in the food domain remains constrained by benchmarks that rely on coarse-grained categories, single-view imagery, and inaccurate metadata. To bridge this gap, we introduce DiningBench, a hierarchical, multi-view benchmark designed to evaluate VLMs across three levels of… ▽ More

    Submitted 11 April, 2026; originally announced April 2026.

    Comments: ACL 2026 Main

  23. arXiv:2604.00860  [pdf, ps, other

    cs.LG

    Policy Improvement Reinforcement Learning

    Authors: Huaiyang Wang, Xiaojie Li, Xiaohan Wang, Zhixia Zhang, Xiaodong Lu, Zixuan Huang, Jiajun Chai, Guojun Yin, Deqing Wang, Haoyi Zhou, Yaodong Yang, Jianxin Li, Yikun Ban

    Abstract: Reinforcement learning has become a central post-training paradigm for improving LLM and agent capabilities. Yet existing RL post-training methods share a common blind spot: they construct local learning signals from sampled trajectories, rewards, or feedback-conditioned targets, then update the policy without explicitly verifying whether the resulting policy outperforms its predecessor. Optimizin… ▽ More

    Submitted 7 July, 2026; v1 submitted 1 April, 2026; originally announced April 2026.

  24. arXiv:2603.08035  [pdf, ps, other

    cs.AI cs.LG

    CDRRM: Contrast-Driven Rubric Generation for Reliable and Interpretable Reward Modeling

    Authors: Dengcan Liu, Fengkai Yang, Xiaohan Wang, Shurui Yan, Jiajun Chai, Jiahao Li, Yikun Ban, Zhendong Mao, Wei Lin, Guojun Yin

    Abstract: Reward modeling is essential for aligning Large Language Models(LLMs) with human preferences, yet conventional reward models suffer from poor interpretability and heavy reliance on costly expert annotations. While recent rubric-based approaches enhance evaluation transparency, they lack systematic quality control, yielding noisy and redundant criteria, failing to mitigate persistent biases (e.g.,… ▽ More

    Submitted 9 March, 2026; originally announced March 2026.

  25. arXiv:2603.06595  [pdf, ps, other

    cs.CL

    Rethinking Personalization in Large Language Models at the Token Level

    Authors: Chenheng Zhang, Yijun Lu, Lizhe Fang, Chunyuan Zheng, Jiajun Chai, Xiaohan Wang, Guojun Yin, Wei Lin, Yisen Wang, Zhouchen Lin

    Abstract: With large language models (LLMs) now performing strongly across diverse tasks, there is growing demand for them to personalize outputs for individual users. Personalization is typically framed as an additional layer on top of a base NLP task, requiring model responses to meet user-specific needs while still accomplishing the underlying task. From a token-level perspective, different tokens in a r… ▽ More

    Submitted 4 February, 2026; originally announced March 2026.

  26. arXiv:2603.02908  [pdf, ps, other

    cs.AI

    SAE as a Crystal Ball: Interpretable Features Predict Cross-domain Transferability of LLMs without Training

    Authors: Qi Zhang, Yifei Wang, Xiaohan Wang, Jiajun Chai, Guojun Yin, Wei Lin, Yisen Wang

    Abstract: In recent years, pre-trained large language models have achieved remarkable success across diverse tasks. Besides the pivotal role of self-supervised pre-training, their effectiveness in downstream applications also depends critically on the post-training process, which adapts models to task-specific data and objectives. However, this process inevitably introduces model shifts that can influence p… ▽ More

    Submitted 3 March, 2026; originally announced March 2026.

  27. arXiv:2602.17100  [pdf, ps, other

    cs.MA

    AgentConductor: Topology Evolution for Multi-Agent Competition-Level Code Generation

    Authors: Siyu Wang, Ruotian Lu, Zhihao Yang, Yuchao Wang, Yanzhou Zhang, Lei Xu, Qimin Xu, Guojun Yin, Cailian Chen, Xinping Guan

    Abstract: Large language model(LLM)-driven multi-agent systems(MAS) coordinate specialized agents through predefined interaction topologies and have shown promise for complex tasks such as competition-level code generation. Recent studies demonstrate that carefully designed multi-agent workflows and communication graphs can significantly improve code generation performance by leveraging collaborative reason… ▽ More

    Submitted 19 February, 2026; originally announced February 2026.

  28. arXiv:2602.12304  [pdf, ps, other

    cs.SD cs.AI cs.MM eess.AS

    OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model

    Authors: Maomao Li, Zhen Li, Kaipeng Zhang, Guosheng Yin, Zhifeng Li, Dong Xu

    Abstract: Existing mainstream video customization methods focus on generating identity-consistent videos based on given reference images and textual prompts. Benefiting from the rapid advancement of joint audio-video generation, this paper proposes a more compelling new task: sync audio-video customization, which aims to synchronously customize both video identity and audio timbre. Specifically, given a ref… ▽ More

    Submitted 23 July, 2026; v1 submitted 11 February, 2026; originally announced February 2026.

    Comments: code: https://github.com/OmniCustom-project/OmniCustom

  29. arXiv:2602.08499  [pdf, ps, other

    cs.LG cs.AI

    Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards

    Authors: Xiaodong Lu, Xiaohan Wang, Jiajun Chai, Guojun Yin, Wei Lin, Zhijun Chen, Yu Luo, Fuzhen Zhuang, Yikun Ban, Deqing Wang

    Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is an effective paradigm for improving the reasoning capabilities of large language models. However, existing RLVR methods utilize rollouts in an indiscriminate and short-horizon manner: responses of heterogeneous quality within each prompt are treated uniformly, and historical rollouts are discarded after a single use. This leads to noisy supe… ▽ More

    Submitted 24 May, 2026; v1 submitted 9 February, 2026; originally announced February 2026.

  30. arXiv:2602.07306  [pdf, ps, other

    cs.DC cs.LG

    Parallel Track Transformers: Enabling Fast GPU Inference with Reduced Synchronization

    Authors: Chong Wang, Nan Du, Tom Gunter, Tao Lei, Kulin Seth, Senyu Tong, Jianyu Wang, Guoli Yin, Xiyou Zhou, Kelvin Zou, Ruoming Pang

    Abstract: Efficient large-scale inference of transformer-based large language models (LLMs) remains a fundamental systems challenge, frequently requiring multi-GPU parallelism to meet stringent latency and throughput targets. Conventional tensor parallelism decomposes matrix operations across devices but introduces substantial inter-GPU synchronization, leading to communication bottlenecks and degraded scal… ▽ More

    Submitted 6 February, 2026; originally announced February 2026.

  31. arXiv:2601.10721  [pdf, ps, other

    cs.RO

    Collaborative Continuum Robots: A Survey

    Authors: Xinyu Li, Qian Tang, Guoxin Yin, Gang Zheng, Jessica Burgner-Kahrs, Cesare Stefanini, Ke Wu

    Abstract: Continuum robots (CRs), owing to their compact structure, inherent compliance, and flexible deformation, have been widely applied in various fields. By coordinating multiple CRs to form collaborative continuum robots (CCRs), task adaptability, workspace, flexibility, load capacity, and operational stability can be further improved, thus offering significant advantages. In recent years, interest in… ▽ More

    Submitted 22 December, 2025; originally announced January 2026.

  32. arXiv:2601.08521  [pdf, ps, other

    cs.LG

    Your Group-Relative Advantage Is Biased

    Authors: Fengkai Yang, Zherui Chen, Xiaohan Wang, Xiaodong Lu, Jiajun Chai, Guojun Yin, Wei Lin, Shuai Ma, Fuzhen Zhuang, Deqing Wang, Yaodong Yang, Jianxin Li, Yikun Ban

    Abstract: Reinforcement Learning from Verifier Rewards (RLVR) has emerged as a widely used approach for post-training large language models on reasoning tasks, with group-based methods such as GRPO and its variants gaining broad adoption. These methods rely on group-relative advantage estimation to avoid learned critics, yet its theoretical properties remain poorly understood. In this work, we uncover a f… ▽ More

    Submitted 21 January, 2026; v1 submitted 13 January, 2026; originally announced January 2026.

  33. arXiv:2601.07468  [pdf, ps, other

    cs.AI

    Beyond Dialogue Time: Temporal Semantic Memory for Personalized LLM Agents

    Authors: Miao Su, Yucan Guo, Zhongni Hou, Long Bai, Zixuan Li, Yufei Zhang, Guojun Yin, Wei Lin, Xiaolong Jin, Jiafeng Guo, Xueqi Cheng

    Abstract: Memory enables Large Language Model (LLM) agents to perceive, store, and use information from past dialogues, which is essential for personalization. However, existing methods fail to properly model the temporal dimension of memory in two aspects: 1) Temporal inaccuracy: memories are organized by dialogue time rather than their actual occurrence time; 2) Temporal fragmentation: existing methods fo… ▽ More

    Submitted 12 January, 2026; originally announced January 2026.

  34. arXiv:2512.19126  [pdf, ps, other

    cs.CL

    AWPO: Enhancing Tool-Use of Large Language Models through Adaptive Integration of Reasoning Rewards

    Authors: Zihan Lin, Xiaohan Wang, Hexiong Yang, Jiajun Chai, Jie Cao, Guojun Yin, Wei Lin, Ran He

    Abstract: While Reinforcement Learning (RL) shows promise in training tool-use Large Language Models (LLMs) using verifiable outcome rewards, existing methods largely overlook the potential of reasoning rewards based on chain-of-thought quality for better tool utilization. Furthermore, naïvely combining reasoning and outcome rewards may yield suboptimal performance or conflict with the primary optimization… ▽ More

    Submitted 14 January, 2026; v1 submitted 22 December, 2025; originally announced December 2025.

  35. arXiv:2512.16149  [pdf, ps, other

    cs.AI

    ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIs

    Authors: Hao Chen, Zhexin Hu, Jiajun Chai, Haocheng Yang, Hang He, Xiaohan Wang, Wei Lin, Luhang Wang, Guojun Yin, Zhuofeng zhao

    Abstract: Training LLMs to invoke tools and leverage retrieved information necessitates high-quality, diverse data. However, existing pipelines for synthetic data generation often rely on tens of thousands of real API calls to enhance generalization, incurring prohibitive costs while lacking multi-hop reasoning and self-reflection. To address these limitations, we introduce ToolForge, an automated synthesis… ▽ More

    Submitted 17 December, 2025; originally announced December 2025.

    Comments: 13 pages, 9 tables, 6 figures. Code available at https://github.com/Buycar-arb/ToolForge

    ACM Class: I.2.7; I.2.8; H.3.3

  36. arXiv:2512.08980  [pdf, ps, other

    cs.CV cs.AI

    Training Multi-Image Vision Agents via End2End Reinforcement Learning

    Authors: Chengqi Dong, Chuhuai Yue, Hang He, Rongge Mao, Fenghe Tang, S Kevin Zhou, Zekun Xu, Xiaohan Wang, Jiajun Chai, Guojun Yin

    Abstract: Recent VLM-based agents aim to replicate OpenAI O3's "thinking with images" via tool use, yet most open-source methods restrict inputs to a single image, limiting their applicability to real-world multi-image QA tasks. To address this gap, we propose IMAgent, an open-source visual agent trained with end-to-end reinforcement learning for fine-grained single/multi-image reasoning. During inference,… ▽ More

    Submitted 3 April, 2026; v1 submitted 5 December, 2025; originally announced December 2025.

  37. LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services

    Authors: Hang He, Chuhuai Yue, Chengqi Dong, Mingxue Tian, Hao Chen, Zhenfeng Liu, Jiajun Chai, Xiaohan Wang, Yufei Zhang, Qun Liao, Guojun Yin, Wei Lin, Chengcheng Wan, Haiying Sun, Ting Su

    Abstract: Recent advances in large reasoning models LRMs have enabled agentic search systems to perform complex multi-step reasoning across multiple sources. However, most studies focus on general information retrieval and rarely explores vertical domains with unique challenges. In this work, we focus on local life services and introduce LocalSearchBench, which encompass diverse and complex business scenari… ▽ More

    Submitted 31 May, 2026; v1 submitted 8 December, 2025; originally announced December 2025.

    Comments: 12 pages; accepted to KDD 2026

    Journal ref: Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 (KDD 2026), August 9--13, 2026, Jeju Island, Republic of Korea. ACM, New York, NY, USA, 12 pages

  38. arXiv:2512.06392   

    cs.LG cs.AI

    RLAX: Large-Scale, Distributed Reinforcement Learning for Large Language Models on TPUs

    Authors: Runlong Zhou, Lefan Zhang, Shang-Chen Wu, Kelvin Zou, Hanzhi Zhou, Ke Ye, Yihao Feng, Dong Yin, Alex Guillen Garcia, Dmytro Babych, Rohit Chatterjee, Matthew Hopkins, Xiang Kong, Chang Lan, Lezhi Li, Yiping Ma, Daniele Molinari, Senyu Tong, Yanchao Sun, Thomas Voice, Jianyu Wang, Chong Wang, Simon Wang, Floris Weers, Yechen Xu , et al. (7 additional authors not shown)

    Abstract: Reinforcement learning (RL) has emerged as the de-facto paradigm for improving the reasoning capabilities of large language models (LLMs). We have developed RLAX, a scalable RL framework on TPUs. RLAX employs a parameter-server architecture. A master trainer periodically pushes updated model weights to the parameter server while a fleet of inference workers pull the latest weights and generates ne… ▽ More

    Submitted 10 December, 2025; v1 submitted 6 December, 2025; originally announced December 2025.

    Comments: The submission is being withdrawn because internal stakeholders determined that it is not appropriate to publish work on this topic at this time

  39. arXiv:2512.05682  [pdf, ps, other

    cs.RO eess.SY

    Scenario-aware Uncertainty Quantification for Trajectory Prediction with Statistical Guarantees

    Authors: Yiming Shu, Jiahui Xu, Linghuan Kong, Fangni Zhang, Guodong Yin, Chen Sun

    Abstract: Reliable uncertainty quantification in trajectory prediction is crucial for safety-critical autonomous driving systems, yet existing deep learning predictors lack uncertainty-aware frameworks adaptable to heterogeneous real-world scenarios. To bridge this gap, we propose a novel scenario-aware uncertainty quantification framework to provide the predicted trajectories with prediction intervals and… ▽ More

    Submitted 5 December, 2025; originally announced December 2025.

  40. arXiv:2512.03745  [pdf, ps, other

    cs.CV

    Dual-level Modality Debiasing Learning for Unsupervised Visible-Infrared Person Re-Identification

    Authors: Jiaze Li, Yan Lu, Bin Liu, Guojun Yin, Mang Ye

    Abstract: Two-stage learning pipeline has achieved promising results in unsupervised visible-infrared person re-identification (USL-VI-ReID). It first performs single-modality learning and then operates cross-modality learning to tackle the modality discrepancy. Although promising, this pipeline inevitably introduces modality bias: modality-specific cues learned in the single-modality training naturally pro… ▽ More

    Submitted 9 April, 2026; v1 submitted 3 December, 2025; originally announced December 2025.

  41. arXiv:2511.17676  [pdf

    cs.DB cs.AI cs.CL

    LLM and Agent-Driven Data Analysis: A Systematic Approach for Enterprise Applications and System-level Deployment

    Authors: Xi Wang, Xianyao Ling, Kun Li, Gang Yin, Liang Zhang, Jiang Wu, Annie Wang, Weizhe Wang

    Abstract: The rapid progress in Generative AI and Agent technologies is profoundly transforming enterprise data management and analytics. Traditional database applications and system deployment are fundamentally impacted by AI-driven tools, such as Retrieval-Augmented Generation (RAG) and vector database technologies, which provide new pathways for semantic querying over enterprise knowledge bases. In the m… ▽ More

    Submitted 21 November, 2025; originally announced November 2025.

  42. arXiv:2511.07800  [pdf, ps, other

    cs.CL

    From Experience to Strategy: Empowering LLM Agents with Trainable Graph Memory

    Authors: Siyu Xia, Zekun Xu, Jiajun Chai, Wentian Fan, Yan Song, Xiaohan Wang, Guojun Yin, Wei Lin, Haifeng Zhang, Jun Wang

    Abstract: Large Language Models (LLMs) based agents have demonstrated remarkable potential in autonomous task-solving across complex, open-ended environments. A promising approach for improving the reasoning capabilities of LLM agents is to better utilize prior experiences in guiding current decisions. However, LLMs acquire experience either through implicit memory via training, which suffers from catastrop… ▽ More

    Submitted 10 November, 2025; originally announced November 2025.

  43. arXiv:2511.03196  [pdf, ps, other

    cs.LG stat.ML

    Cross-Modal Alignment via Variational Copula Modelling

    Authors: Feng Wu, Tsai Hor Chan, Fuying Wang, Guosheng Yin, Lequan Yu

    Abstract: Various data modalities are common in real-world applications (e.g., electronic health records, medical images and clinical notes in healthcare). It is essential to develop multimodal learning methods to aggregate various information from multiple modalities. The main challenge is how to appropriately align and fuse the representations of different modalities into a joint distribution. Existing me… ▽ More

    Submitted 5 November, 2025; originally announced November 2025.

    Journal ref: published by ICML2025

  44. arXiv:2511.02755  [pdf, ps, other

    cs.CL

    Controlling Performance and Budget of a Centralized Multi-agent LLM System with Reinforcement Learning

    Authors: Bowen Jin, TJ Collins, Donghan Yu, Mert Cemri, Shenao Zhang, Mengyu Li, Jay Tang, Tian Qin, Zhiyang Xu, Jiarui Lu, Guoli Yin, Jiawei Han, Zirui Wang

    Abstract: Large language models (LLMs) exhibit complementary strengths across domains and come with varying inference costs, motivating the design of multi-agent LLM systems where specialized models collaborate efficiently. Existing approaches predominantly rely on decentralized frameworks, which invoke multiple LLMs for every input and thus lead to substantial and uncontrolled inference costs. In this work… ▽ More

    Submitted 4 November, 2025; originally announced November 2025.

    Comments: 14 pages

  45. arXiv:2510.25510  [pdf, ps, other

    cs.AI

    MTIR-SQL: Multi-turn Tool-Integrated Reasoning Reinforcement Learning for Text-to-SQL

    Authors: Zekun Xu, Siyu Xia, Chuhuai Yue, Jiajun Chai, Mingxue Tian, Xiaohan Wang, Wei Lin, Haoxuan Li, Guojun Yin

    Abstract: As large language models (LLMs) are increasingly used in Text-to-SQL tasks, Reinforcement Learning (RL) has become a common method for improving performance. Existing methods primarily rely on static execution feedback, which restricts real-time error correction. However, integrating multi-turn tool invocation along with dynamic feedback could significantly improve adaptability and robustness, ult… ▽ More

    Submitted 29 October, 2025; originally announced October 2025.

  46. arXiv:2510.24285  [pdf, ps, other

    cs.CV cs.AI cs.CL

    ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model

    Authors: Juntian Zhang, Song Jin, Chuanqi Cheng, Yuhan Liu, Yankai Lin, Xun Zhang, Yufei Zhang, Fei Jiang, Guojun Yin, Wei Lin, Rui Yan

    Abstract: The limited capacity for fine-grained visual perception presents a critical bottleneck for Vision-Language Models (VLMs) in real-world applications. Addressing this is challenging due to the scarcity of high-quality data and the limitations of existing methods: supervised fine-tuning (SFT) often compromises general capabilities, while reinforcement fine-tuning (RFT) prioritizes textual reasoning o… ▽ More

    Submitted 28 October, 2025; originally announced October 2025.

  47. arXiv:2510.22651  [pdf, ps, other

    cs.LG cs.AI

    Variational Polya Tree

    Authors: Lu Xu, Tsai Hor Chan, Kwok Fai Lam, Lequan Yu, Guosheng Yin

    Abstract: Density estimation is essential for generative modeling, particularly with the rise of modern neural networks. While existing methods capture complex data distributions, they often lack interpretability and uncertainty quantification. Bayesian nonparametric methods, especially the \polya tree, offer a robust framework that addresses these issues by accurately capturing function behavior over small… ▽ More

    Submitted 26 October, 2025; originally announced October 2025.

  48. arXiv:2510.20736  [pdf, ps, other

    cs.LG

    Amplifying Prominent Representations in Multimodal Learning via Variational Dirichlet Process

    Authors: Tsai Hor Chan, Feng Wu, Yihang Chen, Guosheng Yin, Lequan Yu

    Abstract: Developing effective multimodal fusion approaches has become increasingly essential in many real-world scenarios, such as health care and finance. The key challenge is how to preserve the feature expressiveness in each modality while learning cross-modal interactions. Previous approaches primarily focus on the cross-modal alignment, while over-emphasis on the alignment of marginal distributions of… ▽ More

    Submitted 23 October, 2025; originally announced October 2025.

    Comments: Accepted by NeruIPS 2025

  49. arXiv:2510.19254  [pdf, ps, other

    cs.SE

    Trace: Securing Smart Contract Repository Against Access Control Vulnerability

    Authors: Chong Chen, Jiachi Chen, Lingfeng Bao, David Lo, Yanlin Wang, Zhenyu Shan, Ting Chen, Guangqiang Yin, Jianxing Yu, Zibin Zheng

    Abstract: Smart contract vulnerabilities, particularly improper Access Control that allows unauthorized execution of restricted functions, have caused billions of dollars in losses. GitHub hosts numerous smart contract repositories containing source code, documentation, and configuration files-these serve as intermediate development artifacts that must be compiled and packaged before deployment. Third-party… ▽ More

    Submitted 22 October, 2025; originally announced October 2025.

  50. arXiv:2510.16416  [pdf, ps, other

    cs.CV cs.AI

    SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning

    Authors: Xiaojun Guo, Runyu Zhou, Yifei Wang, Qi Zhang, Chenheng Zhang, Stefanie Jegelka, Xiaohan Wang, Jiajun Chai, Guojun Yin, Wei Lin, Yisen Wang

    Abstract: Vision-language models (VLMs) have shown remarkable abilities by integrating large language models with visual inputs. However, they often fail to utilize visual evidence adequately, either depending on linguistic priors in vision-centric tasks or resorting to textual shortcuts during reasoning. Although reinforcement learning (RL) can align models with desired behaviors, its application to VLMs h… ▽ More

    Submitted 17 May, 2026; v1 submitted 18 October, 2025; originally announced October 2025.