Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 56 results for author: Si, Q

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.07196  [pdf, ps, other

    cs.AI

    EMAS: Stabilizing Multi-Agent System Evolution through Evidence-Guided Revision

    Authors: Chao Fei, Qingyi Si, Kaihua Liang, Yanghua Xiao, Panos Kalnis, Hongcheng Guo

    Abstract: Many methods for automated multi-agent system design optimize prompts and topologies during an initial design stage and then deploy the resulting system unchanged on subsequent samples. Experience from these samples is rarely consolidated into reusable system updates, while accuracy-oriented designs may incur high token costs. We introduce EMAS (Evolving Multi-Agent System), which uses this experi… ▽ More

    Submitted 7 August, 2026; originally announced August 2026.

  2. arXiv:2607.13940  [pdf, ps, other

    cs.AI

    A Self-Evolving Agent for Longitudinal Personal Health Management

    Authors: Haoran Li, Jiebi Deng, Tong Jin, Jinghong Han, Yuxin Wang, Zexin Wang, Qingyi Si, Weikang Gong, Xiahai Zhuang, Jia You, Wei Cheng, Jianfeng Feng, Hongcheng Guo

    Abstract: Personal health management unfolds over repeated encounters, yet most health AI systems treat each request in isolation. We developed HealthClaw, an open-source agent architecture that updates support as a person's routines, preferences, measurements and risks change. It separates shared safety rules and medical knowledge from private longitudinal memory containing profile facts, reusable procedur… ▽ More

    Submitted 15 July, 2026; originally announced July 2026.

    Comments: 20 pages, 4 figures, 6 supplementary tables. Code: https://github.com/HC-Guo/HealthClaw

  3. arXiv:2607.04718  [pdf, ps, other

    cs.AI

    FORGE: Research-Trajectory Hijacking Attacks on Deep Research Agents

    Authors: Yue Pan, Ziheng Zhang, Junxiang Lei, Changhao Jia, Qingyi Si, Hongcheng Guo

    Abstract: Deep research agents decompose open-ended queries into subtasks, retrieve web evidence over multiple rounds, and synthesize long-form reports. This workflow creates a planning-layer poisoning surface: adversarial documents that enter the retrieval pool can steer follow-up questions and turn a local injection into report-level contamination. We present FORGE (Fabricated Orchestrated Reasoning chain… ▽ More

    Submitted 6 July, 2026; originally announced July 2026.

    Comments: 20 pages,8 figures,Code available at https://github.com/yvepan/FORGE

  4. arXiv:2606.30626  [pdf, ps, other

    cs.AI

    DOPD: Dual On-policy Distillation

    Authors: Xinlei Yu, Gen Li, Qingyi Si, Guibin Zhang, Yuqi Xu, Congcong Wang, Shuai Dong, Kaiwen Tuo, Xiangyu Zeng, Kaituo Feng, Qunzhong Wang, Yang Shi, Xiaobin Hu, Xiangyu Yue, Jiaqi Wang, Shuicheng Yan

    Abstract: On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense token-level signals. To furnish high-quality supervision sources and thereby elevate the performance frontier of distillation, an intuitive direction is to infuse privileged information to either teacher or student itself. However, this additional input induces a potential failure… ▽ More

    Submitted 29 June, 2026; originally announced June 2026.

  5. arXiv:2606.14777  [pdf, ps, other

    cs.CV cs.AI

    JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence

    Authors: Dingyu Yao, Junhao Zhou, Chenxu Yang, Chuanyu Qin, Haowen Hou, Zheming Liang, Congcong Wang, Yuhang Cao, Shenglong Ye, Shuai Xie, Shuhuan Gu, Haoyang Huang, Qingyi Si, Nan Duan, Jiaqi Wang

    Abstract: Many moments in the real world do not wait for a user to ask. A fire starts on a security monitor, an expression flickers across a video call, or a product a viewer wants flashes by in a livestream. Yet today's large models remain mostly turn-based by design: they answer only when addressed, and even video-call apps that appear interactive still operate as question-answer systems, reacting only wh… ▽ More

    Submitted 9 June, 2026; originally announced June 2026.

  6. arXiv:2606.08615  [pdf, ps, other

    cs.CV cs.CL

    Harnessing Streaming Video in the Wild

    Authors: Dingyu Yao, Shuhuan Gu, Qingyi Si, Junhao Zhou, Chenxu Yang, Chuanyu Qin, Naibin Gu, Zheng Lin, Weiping Wang, Nan Duan, Jiaqi Wang

    Abstract: Vision-Language Models (VLMs) are increasingly required to process unbounded video streams in applications such as video-call assistants, live commentary, and embodied robots. An ideal streaming system should support proactive interaction, long-horizon memory, and real-time processing, while resting on a VLM backbone capable of handling diverse in-the-wild streaming tasks. However, existing VLMs e… ▽ More

    Submitted 7 June, 2026; originally announced June 2026.

  7. arXiv:2606.05874  [pdf, ps, other

    cs.CL

    Evaluating Stochastic Collapse and Implicit Bias in Multimodal Large Language Models

    Authors: Huiyuan Zheng, Houtao Zhang, Boyang Wang, Qingyi Si, Hongcheng Guo

    Abstract: Current evaluations for Multimodal Large Language Models (MLLMs) overwhelmingly focus on utility-driven objectives, leaving model behavior under logic-neutral scenarios largely underexplored. Stochasticity is essential in scenarios where multiple actions are equally valid, such as recommending travel itineraries or daily schedules where multiple options have similar utility. In such settings, dete… ▽ More

    Submitted 4 June, 2026; originally announced June 2026.

  8. arXiv:2606.03087  [pdf, ps, other

    cs.LG

    Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR

    Authors: Chuanyu Qin, Chenxu Yang, Qingyi Si, Naibin Gu, Peng Fu, Zheng Lin

    Abstract: Reinforcement learning with verifiable rewards (RLVR) improves the ability of large language model, yet headline accuracy gains often conceal a hidden cost: previously solved problems quietly become unsolvable as training proceeds. We frame this phenomenon as \emph{correct-set turnover}, representing the coupled dynamics of solution acquisition and regression over the mastered set. Under this view… ▽ More

    Submitted 1 June, 2026; originally announced June 2026.

  9. arXiv:2606.02569  [pdf, ps, other

    cs.CV cs.AI cs.CL

    AdaCodec: A Predictive Visual Code for Video MLLMs

    Authors: Haowen Hou, Zhen Huang, Zheming Liang, Qingyi Si, Chenglin Li, Shuai Dong, Kele Shao, Ruilin Li, Dianyi Wang, Nan Duan, Jiaqi Wang

    Abstract: Video is temporally redundant: adjacent frames usually share most objects, background, and layout. Yet existing video multimodal large language models (video MLLMs) usually encode each sampled frame as an independent RGB image, causing visual tokens to repeat content already present in earlier frames. This suggests a more direct video interface: send a full reference frame only when the scene cann… ▽ More

    Submitted 1 June, 2026; originally announced June 2026.

    Comments: 23 pages

  10. arXiv:2605.28258  [pdf, ps, other

    cs.SE cs.AI cs.CV cs.HC

    GUI Agents for Continual Game Generation

    Authors: Yixu Huang, Bo Li, Na Li, Zhe Wang, Kaijie Chen, Haonan Ge, Qingyi Si, Yuanzhe Shen, Ruihan Yang, Guangjing Wang, Hongcheng Guo

    Abstract: Generating a game is not the same as making one that can be played. Despite advances in code generation, existing approaches treat game generation as one-shot translation from prompt to artifact, leaving interaction-level failures undetected. We argue that evaluating and improving game generation requires a player, and study two roles for graphical user interface (GUI) agents in this process: (1)… ▽ More

    Submitted 27 May, 2026; originally announced May 2026.

  11. arXiv:2605.18055  [pdf, ps, other

    cs.LG cs.AI

    FLAG: Foundation model representation with Latent diffusion Alignment via Graph for spatial gene expression prediction

    Authors: Qi Si, Penglei Wang, Yushuai Wu, Yifeng Jiao, Xuyang Liu, Xin Guo, Yuan Qi, Yuan Cheng

    Abstract: Predicting spatial gene expression from routine H\&E enables large-scale molecular profiling, yet current models treat this as isolated pointwise tasks, thereby overlooking essential biological structures like gene coordination and spatial distribution. To preserve these relationships, we introduce \textbf{FLAG}, a diffusion-based framework that redefines this task as structured distribution model… ▽ More

    Submitted 18 May, 2026; originally announced May 2026.

    Comments: 9 pages for main text, 3 pages for references, 19 pages for appendix. accepted by ICML 2026

  12. arXiv:2605.00323  [pdf, ps, other

    cs.CV cs.LG

    Online Self-Calibration Against Hallucination in Vision-Language Models

    Authors: Minghui Chen, Chenxu Yang, Hengjie Zhu, Dayan Wu, Zheng Lin, Qingyi Si

    Abstract: Large Vision-Language Models (LVLMs) often suffer from hallucinations, generating descriptions that include visual details absent from the input image. Recent preference alignment methods typically rely on supervision distilled from stronger models such as GPT. However, this offline paradigm introduces a Supervision-Perception Mismatch: the student model is forced to align with fine-grained detail… ▽ More

    Submitted 30 April, 2026; originally announced May 2026.

    Comments: IJCAI 2026

  13. arXiv:2604.27083  [pdf, ps, other

    cs.LG

    Co-Evolving Policy Distillation

    Authors: Naibin Gu, Chenxu Yang, Qingyi Si, Chuanyu Qin, Dingyu Yao, Peng Fu, Zheng Lin, Weiping Wang, Nan Duan, Jiaqi Wang

    Abstract: RLVR and OPD have become standard paradigms for post-training. We provide a unified analysis of these two paradigms in consolidating multiple expert capabilities into a single model, identifying capability loss in different ways: mixed RLVR suffers from inter-capability divergence cost, while the pipeline of first training experts and then performing OPD, though avoiding divergence, fails to fully… ▽ More

    Submitted 29 April, 2026; originally announced April 2026.

    Comments: Work in progress

  14. arXiv:2604.20733  [pdf, ps, other

    cs.LG

    Near-Future Policy Optimization

    Authors: Chuanyu Qin, Chenxu Yang, Qingyi Si, Naibin Gu, Dingyu Yao, Zheng Lin, Peng Fu, Nan Duan, Jiaqi Wang

    Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a core post-training recipe. Introducing suitable off-policy trajectories into on-policy exploration accelerates RLVR convergence and raises the performance ceiling, yet finding a source of such trajectories remains the key challenge. Existing mixed-policy methods either import trajectories from external teachers (high-quality but di… ▽ More

    Submitted 22 April, 2026; originally announced April 2026.

    Comments: Work in progress

  15. arXiv:2604.16893  [pdf, ps, other

    cs.CV cs.LG

    EasyVideoR1: Easier RL for Video Understanding

    Authors: Chuanyu Qin, Chenxu Yang, Qingyi Si, Naibin Gu, Dingyu Yao, Zheng Lin, Peng Fu, Nan Duan, Jiaqi Wang

    Abstract: Reinforcement learning from verifiable rewards (RLVR) has demonstrated remarkable effectiveness in improving the reasoning capabilities of large language models. As models evolve into natively multimodal architectures, extending RLVR to video understanding becomes increasingly important yet remains largely unexplored, due to the diversity of video task types, the computational overhead of repeated… ▽ More

    Submitted 18 April, 2026; originally announced April 2026.

  16. arXiv:2604.03128  [pdf, ps, other

    cs.LG cs.CL

    Self-Distilled RLVR

    Authors: Chenxu Yang, Chuanyu Qin, Qingyi Si, Minghui Chen, Naibin Gu, Dingyu Yao, Zheng Lin, Weiping Wang, Jiaqi Wang, Nan Duan

    Abstract: On-policy distillation (OPD) has become a popular training paradigm in the LLM community. This paradigm selects a larger model as the teacher to provide dense, fine-grained signals for each sampled trajectory, in contrast to reinforcement learning with verifiable rewards (RLVR), which only obtains sparse signals from verifiable outcomes in the environment. Recently, the community has explored on-p… ▽ More

    Submitted 8 April, 2026; v1 submitted 3 April, 2026; originally announced April 2026.

    Comments: Work in progress

  17. arXiv:2603.27240  [pdf, ps, other

    cs.CV cs.AI

    Diagnosing and Repairing Unsafe Channels in Vision-Language Models via Causal Discovery and Dual-Modal Safety Subspace Projection

    Authors: Jinhu Fu, Yihang Lou, Qingyi Si, Shudong Zhang, Yan Bai, Sen Su

    Abstract: Large Vision-Language Models (LVLMs) have achieved impressive performance across multimodal understanding and reasoning tasks, yet their internal safety mechanisms remain opaque and poorly controlled. In this work, we present a comprehensive framework for diagnosing and repairing unsafe channels within LVLMs (CARE). We first perform causal mediation analysis to identify neurons and layers that are… ▽ More

    Submitted 28 March, 2026; originally announced March 2026.

    Comments: Accepted by CVPR 2026 main conference

  18. arXiv:2603.15518  [pdf, ps, other

    cs.CL

    Beyond the Covariance Trap: Unlocking Generalization in Same-Subject Knowledge Editing for Large Language Models

    Authors: Xiyu Liu, Qingyi Si, Zhengxiao Liu, Chenxu Yang, Naibin Gu, Zheng Lin

    Abstract: While locate-then-edit knowledge editing efficiently updates knowledge encoded within Large Language Models (LLMs), a critical generalization failure mode emerges in the practical same-subject knowledge editing scenario: models fail to recall the updated knowledge when following user instructions, despite successfully recalling it in the original edited form. This paper identifies the geometric ro… ▽ More

    Submitted 16 March, 2026; originally announced March 2026.

    Comments: 23 pages, 20 figures

  19. arXiv:2603.15402  [pdf, ps, other

    cs.CL cs.AI

    A Closer Look into LLMs for Table Understanding

    Authors: Jia Wang, Chuanyu Qin, Mingyu Zheng, Qingyi Si, Peize Li, Zheng Lin

    Abstract: Despite the success of Large Language Models (LLMs) in table understanding, their internal mechanisms remain unclear. In this paper, we conduct an empirical study on 16 LLMs, covering general LLMs, specialist tabular LLMs, and Mixture-of-Experts (MoE) models, to explore how LLMs understand tabular data and perform downstream tasks. Our analysis focus on 4 dimensions including the attention dynamic… ▽ More

    Submitted 16 March, 2026; originally announced March 2026.

  20. arXiv:2603.15397  [pdf, ps, other

    cs.CR cs.AI

    SFCoT: Safer Chain-of-Thought via Active Safety Evaluation and Calibration

    Authors: Yu Pan, Wenlong Yu, Tiejun Wu, Xiaohu Ye, Qiannan Si, Guangquan Xu, Bin Wu

    Abstract: Large language models (LLMs) have demonstrated remarkable capabilities in complex reasoning tasks. However, they remain highly susceptible to jailbreak attacks that undermine their safety alignment. Existing defense mechanisms typically rely on post hoc filtering applied only to the final output, leaving intermediate reasoning steps unmonitored and vulnerable to adversarial manipulation. To addres… ▽ More

    Submitted 16 March, 2026; originally announced March 2026.

  21. arXiv:2603.14807  [pdf, ps, other

    cs.CV cs.RO

    HiMemVLN: Enhancing Reliability of Open-Source Zero-Shot Vision-and-Language Navigation with Hierarchical Memory System

    Authors: Kailin Lyu, Kangyi Wu, Pengna Li, Xiuyu Hu, Qingyi Si, Cui Miao, Ning Yang, Zihang Wang, Long Xiao, Lianyu Hu, Jingyuan Sun, Ce Hao

    Abstract: LLM-based agents have demonstrated impressive zero-shot performance in vision-language navigation (VLN) tasks. However, most zero-shot methods primarily rely on closed-source LLMs as navigators, which face challenges related to high token costs and potential data leakage risks. Recent efforts have attempted to address this by using open-source LLMs combined with a spatiotemporal CoT framework, but… ▽ More

    Submitted 26 July, 2026; v1 submitted 16 March, 2026; originally announced March 2026.

    Comments: 9 pages, 7 figures

  22. arXiv:2601.21414  [pdf, ps, other

    cs.AI cs.CL

    System 1&2 Synergy via Dynamic Model Interpolation

    Authors: Chenxu Yang, Qingyi Si, Chong Tian, Xiyu Liu, Dingyu Yao, Chuanyu Qin, Zheng Lin, Weiping Wang, Jiaqi Wang

    Abstract: Training a unified language model that adapts between intuitive System 1 and deliberative System 2 remains challenging due to interference between their cognitive modes. Recent studies have thus pursued making System 2 models more efficient. However, these approaches focused on output control, limiting what models produce. We argue that this paradigm is misaligned: output length is merely a sympto… ▽ More

    Submitted 29 January, 2026; originally announced January 2026.

  23. arXiv:2601.19232  [pdf, ps, other

    cs.LG cs.AI

    Structure-based RNA Design by Step-wise Optimization of Latent Diffusion Model

    Authors: Qi Si, Xuyang Liu, Penglei Wang, Xin Guo, Yuan Qi, Yuan Cheng

    Abstract: RNA inverse folding, designing sequences to form specific 3D structures, is critical for therapeutics, gene regulation, and synthetic biology. Current methods, focused on sequence recovery, struggle to address structural objectives like secondary structure consistency (SS), minimum free energy (MFE), and local distance difference test (LDDT), leading to suboptimal structural accuracy. To tackle th… ▽ More

    Submitted 27 January, 2026; originally announced January 2026.

    Comments: 20 pages (7 pages content + 2 pages references + 11 pages appendix), 11 figures, 8 tables. Source code available at https://github.com/darkflash03/SOLD Accepted to AAAI 2026

  24. arXiv:2601.07408  [pdf, ps, other

    cs.CL cs.LG

    Outcome-Grounded Advantage Reshaping for Fine-Grained Credit Assignment in Mathematical Reasoning

    Authors: Ziheng Li, Liu Kang, Feng Xiao, Luxi Xing, Qingyi Si, Zhuoran Li, Weikang Gong, Deqing Yang, Yanghua Xiao, Hongcheng Guo

    Abstract: Group Relative Policy Optimization (GRPO) has emerged as a promising critic-free reinforcement learning paradigm for reasoning tasks. However, standard GRPO employs a coarse-grained credit assignment mechanism that propagates group-level rewards uniformly to to every token in a sequence, neglecting the varying contribution of individual reasoning steps. We address this limitation by introducing Ou… ▽ More

    Submitted 3 June, 2026; v1 submitted 12 January, 2026; originally announced January 2026.

  25. arXiv:2601.00677  [pdf, ps, other

    cs.LG cs.AI

    IRPM: Intergroup Relative Preference Modeling for Pointwise Generative Reward Models

    Authors: Haonan Song, Qingchen Xie, Huan Zhu, Feng Xiao, Luxi Xing, Liu Kang, Fuzhen Li, Zhiyong Zheng, Feng Jiang, Ziheng Li, Kun Yan, Qingyi Si, Yanghua Xiao, Hongcheng Guo, Fan Yang

    Abstract: Generative Reward Models (GRMs) have demonstrated strong performance in reward modeling, due to their interpretability and potential for refinement through reinforcement learning (RL). However, widely used pairwise GRMs create a computational bottleneck in reinforcement learning from human feedback (RLHF), when calibrating or aggregating preference signals over n candidates, often incurring O(n^2)… ▽ More

    Submitted 30 January, 2026; v1 submitted 2 January, 2026; originally announced January 2026.

    Comments: Comments: Updated title for clarity; improved theoretical derivations; added experiments at additional parameter scales and more ablations; added experimental details in the appendix; updated author list (added five co-authors) to reflect contributions to experiments and writing

  26. arXiv:2512.19530  [pdf, ps, other

    cs.LG cs.AI

    Learning Continuous Solvent Effects from Transient Flow Data: A Graph Neural Network Benchmark on Catechol Rearrangement

    Authors: Hongsheng Xing, Qiuxin Si

    Abstract: Predicting reaction outcomes across continuous solvent composition ranges remains a critical challenge in organic synthesis and process chemistry. Traditional machine learning approaches often treat solvent identity as a discrete categorical variable, which prevents systematic interpolation and extrapolation across the solvent space. This work introduces the \textbf{Catechol Benchmark}, a high-thr… ▽ More

    Submitted 22 December, 2025; originally announced December 2025.

    Comments: 13 pages, 6 figures

    MSC Class: 68T07; 92E20; 62M45 ACM Class: I.2.1; I.2.6; J.2

  27. arXiv:2510.11292  [pdf, ps, other

    cs.LG cs.AI

    LouisKV: Efficient KV Cache Retrieval for Long Input-Output Sequences

    Authors: Wenbo Wu, Qingyi Si, Xiurui Pan, Ye Wang, Jie Zhang

    Abstract: While Key-Value (KV) cache succeeds in reducing redundant computations in auto-regressive models, it introduces significant memory overhead, limiting its practical deployment in long-sequence scenarios. Existing KV retrieval methods mitigate this by dynamically retaining only a subset of KV entries on the GPU. However, they still suffer from notable efficiency and accuracy bottlenecks due to per-t… ▽ More

    Submitted 13 October, 2025; originally announced October 2025.

  28. arXiv:2510.09221  [pdf, ps, other

    cs.RO

    HANDO: Hierarchical Autonomous Navigation and Dexterous Omni-loco-manipulation

    Authors: Jingyuan Sun, Chaoran Wang, Mingyu Zhang, Cui Miao, Hongyu Ji, Zihan Qu, Han Sun, Bing Wang, Qingyi Si

    Abstract: Seamless loco-manipulation in unstructured environments requires robots to leverage autonomous exploration alongside whole-body control for physical interaction. In this work, we introduce HANDO (Hierarchical Autonomous Navigation and Dexterous Omni-loco-manipulation), a two-layer framework designed for legged robots equipped with manipulators to perform human-centered mobile manipulation tasks. T… ▽ More

    Submitted 10 October, 2025; originally announced October 2025.

    Comments: 4 pages, 2 figures, this paper has been accepted for the workshop Perception and Planning for Mobile Manipulation in Changing Environments (PM2CE) at IROS 2025

  29. arXiv:2508.18651  [pdf, ps, other

    cs.CL cs.AI

    Breaking the Trade-Off Between Faithfulness and Expressiveness for Large Language Models

    Authors: Chenxu Yang, Qingyi Si, Zheng Lin

    Abstract: Grounding responses in external knowledge represents an effective strategy for mitigating hallucinations in Large Language Models (LLMs). However, current LLMs struggle to seamlessly integrate knowledge while simultaneously maintaining faithfulness (or fidelity) and expressiveness, capabilities that humans naturally possess. This limitation results in outputs that either lack support from external… ▽ More

    Submitted 25 August, 2025; originally announced August 2025.

  30. arXiv:2508.11661  [pdf, ps, other

    cs.LG cs.CL

    Sparse Attention across Multiple-context KV Cache

    Authors: Ziyi Cao, Qingyi Si, Jingbin Zhang, Bingquan Liu

    Abstract: Large language models face significant cost challenges in long-sequence inference. To address this, reusing historical Key-Value (KV) Cache for improved inference efficiency has become a mainstream approach. Recent advances further enhance throughput by sparse attention mechanisms to select the most relevant KV Cache, thereby reducing sequence length. However, such techniques are limited to single… ▽ More

    Submitted 5 August, 2025; originally announced August 2025.

  31. arXiv:2508.02511  [pdf, ps, other

    cs.AI cs.CL

    Test-time Prompt Intervention

    Authors: Chenxu Yang, Qingyi Si, Mz Dai, Dingyu Yao, Mingyu Zheng, Minghui Chen, Zheng Lin, Weiping Wang

    Abstract: Test-time compute has led to remarkable success in the large language model (LLM) community, particularly for complex tasks, where longer chains of thought (CoTs) are generated to enhance reasoning capabilities. However, growing evidence reveals that such reasoning models often produce CoTs plagued by excessive redundancy, including unnecessary verification steps and repetitive reasoning shifts. T… ▽ More

    Submitted 22 October, 2025; v1 submitted 4 August, 2025; originally announced August 2025.

    Comments: 24 pages, 20 figures, under review

  32. arXiv:2505.18086  [pdf, other

    cs.AI cs.LG

    Stable Reinforcement Learning for Efficient Reasoning

    Authors: Muzhi Dai, Shixuan Liu, Qingyi Si

    Abstract: The success of Deepseek-R1 has drawn the LLM community's attention to reinforcement learning (RL) methods like GRPO. However, such rule-based 0/1 outcome reward methods lack the capability to regulate the intermediate reasoning processes during chain-of-thought (CoT) generation, leading to severe overthinking phenomena. In response, recent studies have designed reward functions to reinforce models… ▽ More

    Submitted 23 May, 2025; originally announced May 2025.

  33. arXiv:2505.07686  [pdf, ps, other

    cs.AI cs.LG

    S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models

    Authors: Muzhi Dai, Chenxu Yang, Qingyi Si

    Abstract: As Test-Time Scaling emerges as an active research focus in the large language model community, advanced post-training methods increasingly emphasize extending chain-of-thought (CoT) generation length, thereby enhancing reasoning capabilities to approach Deepseek R1-like reasoning models. However, recent studies reveal that reasoning models (even Qwen3) consistently exhibit excessive thought redun… ▽ More

    Submitted 17 May, 2025; v1 submitted 12 May, 2025; originally announced May 2025.

  34. arXiv:2504.15895  [pdf, ps, other

    cs.CL cs.AI

    Dynamic Early Exit in Reasoning Models

    Authors: Chenxu Yang, Qingyi Si, Yongjie Duan, Zheliang Zhu, Chenyu Zhu, Qiaowei Li, Minghui Chen, Zheng Lin, Weiping Wang

    Abstract: Recent advances in large reasoning language models (LRLMs) rely on test-time scaling, which extends long chain-of-thought (CoT) generation to solve complex tasks. However, overthinking in long CoT not only slows down the efficiency of problem solving, but also risks accuracy loss due to the extremely detailed or redundant reasoning steps. We propose a simple yet effective method that allows LLMs t… ▽ More

    Submitted 28 September, 2025; v1 submitted 22 April, 2025; originally announced April 2025.

    Comments: 41 pages, 18 figures

  35. arXiv:2503.20484  [pdf, other

    cs.CV cs.AI

    Contrastive Learning Guided Latent Diffusion Model for Image-to-Image Translation

    Authors: Qi Si, Bo Wang, Zhao Zhang

    Abstract: The diffusion model has demonstrated superior performance in synthesizing diverse and high-quality images for text-guided image translation. However, there remains room for improvement in both the formulation of text prompts and the preservation of reference image content. First, variations in target text prompts can significantly influence the quality of the generated images, and it is often chal… ▽ More

    Submitted 26 March, 2025; originally announced March 2025.

    Comments: 11 pages, 13 figures

  36. arXiv:2503.12559  [pdf, ps, other

    cs.CV cs.CL cs.MM

    AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding

    Authors: Xiao Wang, Qingyi Si, Jianlong Wu, Shiyu Zhu, Li Cao, Liqiang Nie

    Abstract: Multimodal Large Language Models (MLLMs) have revolutionized video understanding, yet are still limited by context length when processing long videos. Recent methods compress videos by leveraging visual redundancy uniformly, yielding promising results. Nevertheless, our quantitative analysis shows that redundancy varies significantly across time and model layers, necessitating a more flexible comp… ▽ More

    Submitted 8 June, 2025; v1 submitted 16 March, 2025; originally announced March 2025.

  37. arXiv:2412.20504  [pdf, other

    cs.CV cs.CL cs.MM

    ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding

    Authors: Xiao Wang, Qingyi Si, Jianlong Wu, Shiyu Zhu, Li Cao, Liqiang Nie

    Abstract: Video Large Language Models (VideoLLMs) have made significant strides in video understanding but struggle with long videos due to the limitations of their backbone LLMs. Existing solutions rely on length extrapolation, which is memory-constrained, or visual token compression, which primarily leverages low-level temporal redundancy while overlooking the more effective high-level knowledge redundanc… ▽ More

    Submitted 23 March, 2025; v1 submitted 29 December, 2024; originally announced December 2024.

    Comments: Rewrite the methods section. Add more ablation studies and results in LongVideoBench. Update metadata

  38. arXiv:2412.14880  [pdf, other

    cs.CV

    Multimodal Hypothetical Summary for Retrieval-based Multi-image Question Answering

    Authors: Peize Li, Qingyi Si, Peng Fu, Zheng Lin, Yan Wang

    Abstract: Retrieval-based multi-image question answering (QA) task involves retrieving multiple question-related images and synthesizing these images to generate an answer. Conventional "retrieve-then-answer" pipelines often suffer from cascading errors because the training objective of QA fails to optimize the retrieval stage. To address this issue, we propose a novel method to effectively introduce and re… ▽ More

    Submitted 19 December, 2024; originally announced December 2024.

    Comments: AAAI 2025

  39. arXiv:2411.02457  [pdf, other

    cs.CL cs.AI

    A Multi-Task Role-Playing Agent Capable of Imitating Character Linguistic Styles

    Authors: Siyuan Chen, Qingyi Si, Chenxu Yang, Yunzhi Liang, Zheng Lin, Huan Liu, Weiping Wang

    Abstract: The advent of large language models (LLMs) has significantly propelled the advancement of Role-Playing Agents (RPAs). However, current Role-Playing Agents predominantly focus on mimicking a character's fundamental attributes while neglecting the replication of linguistic style, and they are incapable of effectively replicating characters when performing tasks beyond multi-turn dialogues, which res… ▽ More

    Submitted 3 November, 2024; originally announced November 2024.

  40. arXiv:2408.00300  [pdf, other

    cs.CV cs.MM

    Towards Flexible Evaluation for Generative Visual Question Answering

    Authors: Huishan Ji, Qingyi Si, Zheng Lin, Weiping Wang

    Abstract: Throughout rapid development of multimodal large language models, a crucial ingredient is a fair and accurate evaluation of their multimodal comprehension abilities. Although Visual Question Answering (VQA) could serve as a developed test field, limitations of VQA evaluation, like the inflexible pattern of Exact Match, have hindered MLLMs from demonstrating their real capability and discourage ric… ▽ More

    Submitted 1 August, 2024; originally announced August 2024.

  41. arXiv:2406.08100  [pdf, other

    cs.CL cs.AI

    Multimodal Table Understanding

    Authors: Mingyu Zheng, Xinwei Feng, Qingyi Si, Qiaoqiao She, Zheng Lin, Wenbin Jiang, Weiping Wang

    Abstract: Although great progress has been made by previous table understanding methods including recent approaches based on large language models (LLMs), they rely heavily on the premise that given tables must be converted into a certain text sequence (such as Markdown or HTML) to serve as model input. However, it is difficult to access such high-quality textual table representations in some real-world sce… ▽ More

    Submitted 12 June, 2024; originally announced June 2024.

    Comments: 23 pages, 16 figures, ACL 2024 main conference, camera-ready version

  42. arXiv:2406.04758  [pdf, other

    cs.CL

    Think out Loud: Emotion Deducing Explanation in Dialogues

    Authors: Jiangnan Li, Zheng Lin, Lanrui Wang, Qingyi Si, Yanan Cao, Mo Yu, Peng Fu, Weiping Wang, Jie Zhou

    Abstract: Humans convey emotions through daily dialogues, making emotion understanding a crucial step of affective intelligence. To understand emotions in dialogues, machines are asked to recognize the emotion for an utterance (Emotion Recognition in Dialogues, ERD); based on the emotion, then find causal utterances for the emotion (Emotion Cause Extraction in Dialogues, ECED). The setting of the two tasks… ▽ More

    Submitted 7 June, 2024; originally announced June 2024.

  43. arXiv:2405.16982  [pdf, other

    cs.IR

    Robust kernel-free quadratic surface twin support vector machine with capped $L_1$-norm distance metric

    Authors: Qi Si, Zhi Xia Yang

    Abstract: Twin support vector machine (TSVM) is a very classical and practical classifier for pattern classification. However, the traditional TSVM has two limitations. Firstly, it uses the L_2-norm distance metric that leads to its sensitivity to outliers. Second, it needs to select the appropriate kernel function and the kernel parameters for nonlinear classification. To effectively avoid these two proble… ▽ More

    Submitted 27 May, 2024; originally announced May 2024.

  44. arXiv:2402.02549  [pdf, other

    cs.CL cs.AI cs.LG

    Are Large Language Models Table-based Fact-Checkers?

    Authors: Hanwen Zhang, Qingyi Si, Peng Fu, Zheng Lin, Weiping Wang

    Abstract: Table-based Fact Verification (TFV) aims to extract the entailment relation between statements and structured tables. Existing TFV methods based on small-scaled models suffer from insufficient labeled data and weak zero-shot ability. Recently, the appearance of Large Language Models (LLMs) has gained lots of attraction in research fields. They have shown powerful zero-shot and in-context learning… ▽ More

    Submitted 13 November, 2024; v1 submitted 4 February, 2024; originally announced February 2024.

    Comments: CSCWD 2024

  45. arXiv:2401.16699  [pdf, other

    cs.RO

    Towards Unified Interactive Visual Grounding in The Wild

    Authors: Jie Xu, Hanbo Zhang, Qingyi Si, Yifeng Li, Xuguang Lan, Tao Kong

    Abstract: Interactive visual grounding in Human-Robot Interaction (HRI) is challenging yet practical due to the inevitable ambiguity in natural languages. It requires robots to disambiguate the user input by active information gathering. Previous approaches often rely on predefined templates to ask disambiguation questions, resulting in performance reduction in realistic interactive scenarios. In this paper… ▽ More

    Submitted 18 February, 2024; v1 submitted 29 January, 2024; originally announced January 2024.

    Comments: Accepted to ICRA 2024

  46. arXiv:2401.09442  [pdf, other

    cs.CV cs.AI

    Object Attribute Matters in Visual Question Answering

    Authors: Peize Li, Qingyi Si, Peng Fu, Zheng Lin, Yan Wang

    Abstract: Visual question answering is a multimodal task that requires the joint comprehension of visual and textual information. However, integrating visual and textual semantics solely through attention layers is insufficient to comprehensively understand and align information from both modalities. Intuitively, object attributes can naturally serve as a bridge to unify them, which has been overlooked in p… ▽ More

    Submitted 20 December, 2023; originally announced January 2024.

    Comments: AAAI 2024

  47. arXiv:2310.07328  [pdf, other

    cs.CL cs.AI

    An Empirical Study of Instruction-tuning Large Language Models in Chinese

    Authors: Qingyi Si, Tong Wang, Zheng Lin, Xu Zhang, Yanan Cao, Weiping Wang

    Abstract: The success of ChatGPT validates the potential of large language models (LLMs) in artificial general intelligence (AGI). Subsequently, the release of LLMs has sparked the open-source community's interest in instruction-tuning, which is deemed to accelerate ChatGPT's replication process. However, research on instruction-tuning LLMs in Chinese, the world's most spoken language, is still in its early… ▽ More

    Submitted 20 October, 2023; v1 submitted 11 October, 2023; originally announced October 2023.

    Comments: EMNLP 2023

  48. arXiv:2305.06407  [pdf, other

    cs.CV cs.AI

    Combo of Thinking and Observing for Outside-Knowledge VQA

    Authors: Qingyi Si, Yuchen Mo, Zheng Lin, Huishan Ji, Weiping Wang

    Abstract: Outside-knowledge visual question answering is a challenging task that requires both the acquisition and the use of open-ended real-world knowledge. Some existing solutions draw external knowledge into the cross-modality space which overlooks the much vaster textual knowledge in natural-language space, while others transform the image into a text that further fuses with the textual knowledge into… ▽ More

    Submitted 10 May, 2023; originally announced May 2023.

    Comments: ACL-23, Main Conference

  49. arXiv:2210.14558  [pdf, other

    cs.CV

    Compressing And Debiasing Vision-Language Pre-Trained Models for Visual Question Answering

    Authors: Qingyi Si, Yuanxin Liu, Zheng Lin, Peng Fu, Weiping Wang

    Abstract: Despite the excellent performance of vision-language pre-trained models (VLPs) on conventional VQA task, they still suffer from two problems: First, VLPs tend to rely on language biases in datasets and fail to generalize to out-of-distribution (OOD) data. Second, they are inefficient in terms of memory footprint and computation. Although promising progress has been made in both problems, most exis… ▽ More

    Submitted 11 October, 2023; v1 submitted 26 October, 2022; originally announced October 2022.

    Comments: EMNLP 2023

  50. arXiv:2210.04692  [pdf, other

    cs.CV

    Language Prior Is Not the Only Shortcut: A Benchmark for Shortcut Learning in VQA

    Authors: Qingyi Si, Fandong Meng, Mingyu Zheng, Zheng Lin, Yuanxin Liu, Peng Fu, Yanan Cao, Weiping Wang, Jie Zhou

    Abstract: Visual Question Answering (VQA) models are prone to learn the shortcut solution formed by dataset biases rather than the intended solution. To evaluate the VQA models' reasoning ability beyond shortcut learning, the VQA-CP v2 dataset introduces a distribution shift between the training and test set given a question type. In this way, the model cannot use the training set shortcut (from question ty… ▽ More

    Submitted 10 October, 2022; originally announced October 2022.

    Comments: Fingdings of EMNLP-2022