Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 105 results for author: Yue, L

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.20111  [pdf, ps, other

    cs.RO cs.ET

    Planning-Oriented End-to-End Autonomous Driving: Architectures, Evaluation, and Emerging Paradigms

    Authors: Yanchen Guan, Xingcheng Liu, Bin Rao, Chengyue Wang, Guofa Li, Yunjian Li, Lishengsa Yue, Zhiyong Cui, Chengzhong Xu, Zhenning Li

    Abstract: End-to-end autonomous driving has evolved from camera-to-control regression toward planning-oriented systems that use structured representations, trajectory-level outputs, and increasingly realistic evaluation protocols. This survey reviews this transition across behavior cloning, conditional imitation learning, privileged distillation, BEV and vectorized planning, unified perception-prediction-pl… ▽ More

    Submitted 20 August, 2026; originally announced August 2026.

  2. arXiv:2608.14667  [pdf, ps, other

    cs.AI cs.HC

    Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems

    Authors: Patrick Emami, Sameera Horawalavithana, Truc Nguyen, Gihan Panapitiya, Bruno Jacob, Siddhisanket Raskar, Saumya Sinha, Jared D. Willard, Andrew Glaws, Nithin Somasekharan, Ling Yue, Brian Lu, Shaowu Pan, Jason Eisner

    Abstract: Large language model-based agents are increasingly deployed as collaborators in scientific discovery yet most current work focuses on the autonomous capabilities of "AI Scientists". We argue that this overlooks the social aspects of scientific teamwork, and that studying AI Scientists as human-agent systems (HAS)--where the unit of analysis is the human-agent pair--is both underexplored and underv… ▽ More

    Submitted 2 August, 2026; originally announced August 2026.

    Comments: 15 pages. Accepted at the COLM 2nd Workshop on Language Models for Scientific Discovery

  3. arXiv:2608.04423  [pdf, ps, other

    cs.CV

    Foreseeing the Invisible: Amodal Reconstruction of Leaf Fossil Images

    Authors: Liuxiang Yue, Ailin Zhang, Ziyue Zhao, Yikun Duan

    Abstract: Fossil leaves are rarely preserved whole -- sedimentary rock hides, breaks, and erodes the lamina, yet paleobotany depends on the complete shape and outline of the leaf. We cast the recovery of the missing tissue as amodal reconstruction and present AmodalDINO, a multi-head dense-prediction model that predicts four masks from a single RGB image: visible leaf, amodal complete leaf, amodal main vein… ▽ More

    Submitted 5 August, 2026; originally announced August 2026.

    Comments: 12 pages, 10 figures

  4. arXiv:2607.22745  [pdf, ps, other

    cs.CV cs.AI

    AI-generated Images Challenge Visual Trust in High-risk Scenarios

    Authors: Yi-Zhi Wang, Yichen Xiao, Linan Yue, Weibo Gao, Yichao Du, Pengfei Fang, Shimin Di, Min-Ling Zhang

    Abstract: Rapid advances in image generation are eroding the evidentiary value of visual content in settings where authenticity can affect public safety and personal reputation. Yet existing detection benchmarks rarely examine synthetic images in public- and individual-safety contexts, where misleading visual content may carry substantial risks. Here we introduce SafeIMG, a safety-oriented benchmark spannin… ▽ More

    Submitted 23 July, 2026; originally announced July 2026.

  5. arXiv:2607.21596  [pdf, ps, other

    cs.AI

    FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills

    Authors: Zeyu Ren, Ling Yue, Ran Li, Yishu Wang, Shengxiang Xu, Hanmo Liu, Shaowu Pan, Shimin Di

    Abstract: Large language model agents can adapt to complex tasks by constructing workflows at inference time, but procedures discovered in one episode are usually discarded after execution. Existing skill libraries provide reusable executable routines, but are typically assembled offline and do not grow from the agent's own workflows. We introduce FlowEvo, a training-free framework in which workflows and sk… ▽ More

    Submitted 20 August, 2026; v1 submitted 18 April, 2026; originally announced July 2026.

    Comments: Published as a conference paper at the Conference on Language Modeling (COLM) 2026. 25 pages, 3 figures, 16 tables. Code: https://github.com/DEFENSE-SEU/FlowEvo

    ACM Class: I.2.7; I.2.11; I.2.8

  6. arXiv:2607.17952  [pdf, ps, other

    cs.CL cs.CE

    What Transfers Under Source Shift? Definitions, Examples, and Fine-Tuning for Climate Disclosure Classification

    Authors: Guosheng Li, Fenghui Ren, Bin Liu, Chuan Yu, Kaiying Ji, Lin Yue, Jun Shen, Sasa Qian

    Abstract: Climate disclosure classification is a fundamental task for analysing corporate climate disclosures, yet such disclosures appear in many different sources -- annual reports, press releases, and earnings calls -- that differ in length, purpose, and writing style. Existing evaluations are mostly conducted within a single source, leaving open whether common LLM adaptation strategies remain effective… ▽ More

    Submitted 20 July, 2026; originally announced July 2026.

    Comments: 15 pages, 12 figures. Code: https://github.com/Leoccino/TCFD-SourceShift

  7. arXiv:2606.27731  [pdf, ps, other

    cs.CL cs.AI cs.CE cs.LG

    Enhancing Numerical Prediction in LLMs via Smooth MMD Alignment

    Authors: Zhuo Zuo, Li Yue, Wenhao Zheng, Chenpeng Wang, Xianggen Liu

    Abstract: Despite their strong general capabilities, large language models (LLMs) often remain unreliable when outputs must be numerically precise. A key reason is the training objective: standard cross-entropy treats numeric tokens as unstructured categories and ignores the metric structure of their values. We address this mismatch with Smooth Maximum Mean Discrepancy (SMMD), which builds on the classic MM… ▽ More

    Submitted 26 June, 2026; originally announced June 2026.

  8. arXiv:2606.19849  [pdf, ps, other

    cs.CV

    ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference

    Authors: Yang Tan, Junlong Tong, Linan Yue, Hao Wu, Pengfei Fang, Xiaoyu Shen

    Abstract: Streaming VideoLLMs must continuously process incoming video while maintaining low query latency, making both video-ingestion throughput and query-time responsiveness critical for real-time deployment. Existing methods largely focus on accelerating individual modules, such as visual encoding, token pruning, or KV-cache compression, but provide limited insight into whether the resulting system can… ▽ More

    Submitted 18 June, 2026; originally announced June 2026.

    Comments: 19 pages, 7 figures, 13 tables

    ACM Class: I.2.7; I.2.10; H.5.1

  9. arXiv:2606.15225  [pdf, ps, other

    cs.LG cs.AI cs.IR

    Edu-Theater: A Data-Efficient Agent Framework for Scalable Learner Behavior Simulation through Staging Roll-Call

    Authors: Weibo Gao, Qi Liu, Linan Yue, Zheng Zhang, Yichao Du, Fangzhou Yao, Ao Yu, Zhenya Huang, Shijin Wang

    Abstract: Large-scale learner-task interaction data are crucial for intelligent educational systems but are costly to collect and constrained by privacy and learner engagement. Learner simulators play a critical role in simulating scalable learner behavior without the need for continuous involvement of real learners. However, existing methods are predominantly \textbf{individual-centric}, pairing a simulato… ▽ More

    Submitted 13 June, 2026; originally announced June 2026.

    Comments: LLM Agent, Educational Data Mining, Data Synthesis, Human Simulation

  10. arXiv:2606.12674  [pdf, ps, other

    cs.AI

    Evoflux: Inference-Time Evolution of Executable Tool Workflows for Compact Agents

    Authors: Kushal Raj Bhandari, Ling Yue, Ching-Yun Ko, Dhaval Patel, Shaowu Pan, Pin-Yu Chen, Jianxi Gao

    Abstract: Compact language models (LMs) reduce cost, latency, and deployment risk for tool agents. Yet MCP-style tool use requires more than isolated function calling: an agent must discover tools from live catalogs, satisfy schemas, preserve dependencies across intermediate outputs, and ground final responses in executed evidence. Small planners often generate plausible workflow graphs that fail under tool… ▽ More

    Submitted 10 June, 2026; originally announced June 2026.

    Comments: Code is available at https://github.com/IBM/Evoflux

  11. arXiv:2606.00138  [pdf, ps, other

    cs.AI

    A Multi-AI-agent Framework Enabling End-to-end Finite Element Analysis for Solid Mechanics Problems

    Authors: Titu Ranjan Sarker, Muhammed Jawaad Zulqernine, Ling Yue, Shaowu Pan, Chenxi Wang, Shiyao Lin

    Abstract: Finite element analysis (FEA) is the most important numerical approach for solid mechanics. Challenges of FEA include a steep learning curve for entry-level users and potential false simulations due to incorrect definitions of key simulation components, such as boundary conditions, load cases, and solution variables. Years of engineering experience are usually necessary for real-world problem-solv… ▽ More

    Submitted 28 May, 2026; originally announced June 2026.

  12. arXiv:2605.27431  [pdf, ps, other

    cs.LG cs.AI

    Tackling Multimodal Learning Challenges with Mixture-of-Expert: A Survey

    Authors: Liangwei Nathan Zheng, Wei Emma Zhang, Olaf Maennel, Lin Yue, Weitong Chen

    Abstract: Mixture-of-Experts (MoE) presents a naturally compatible and scalable framework for multimodal learning, demonstrating strong adaptability across diverse modalities and tasks. Despite its growing success, a comprehensive and systematic review on the MoE metho addressing multimodal challenges remains lacking. Existing surveys tend to evaluate either multimodal learning or MoE independently from met… ▽ More

    Submitted 22 May, 2026; originally announced May 2026.

    Comments: This survey paper has just been accepted by IJCAI 2026. Results were released by 30 April 2026. As I could not find a particular place to drop the acceptance email. I have upload the acceptance email alongside the LaTeX files of the paper, named as Acceptance_email.pdf

  13. arXiv:2605.25511  [pdf, ps, other

    cs.CL

    CRPO: Character-centric Group Relative Policy Optimization for Role-aware Reasoning in Role-playing Agents

    Authors: Yihong Tang, Kehai Chen, Liang Yue, Benyou Wang, Min Zhang

    Abstract: Recent advancements in Reinforcement Learning (RL), particularly Group Relative Policy Optimization (GRPO), have significantly enhanced the reasoning capabilities of Large Language Models. However, applying these problem-centric optimization methods to role-playing agents often leads to a loss of character fidelity and style collapse, as they prioritize context-specific utility over persona alignm… ▽ More

    Submitted 25 May, 2026; originally announced May 2026.

  14. arXiv:2605.24456  [pdf, ps, other

    cs.CV

    EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy

    Authors: Jinzhao Li, Yinuo Chen, Dongxu Piao, Panwang Pan, Yifan Yu, Dong Wang, Honglei Yan, Liang Yue, Shaofei Wang, Yixin Chen, Siyuan Huang, Miao Liu

    Abstract: Humans constantly reason about 3D proximity, the relations between their body and surrounding objects, to guide perception and action in daily life. Whether multimodal large language models (MLLMs) can perform such embodied 3D reasoning remains unclear. To this end, we introduce EgoProx, a benchmark for egocentric 3D proximity reasoning. We organize our tasks along a cognitive chain, covering inte… ▽ More

    Submitted 26 May, 2026; v1 submitted 23 May, 2026; originally announced May 2026.

    Comments: Accepted to CVPR 2026

  15. arXiv:2605.10404  [pdf, ps, other

    cs.CV

    Position: Life-Logging Video Streams Make the Privacy-Utility Trade-off Inevitable

    Authors: Tianyuan Zou, Liang Yue, Yang Liu, Ya-Qin Zhang, Sijie Cheng

    Abstract: With the growing prevalence of always-on hardware such as smart glasses, body cameras, and home security systems, life-logging visual sensing is becoming inevitable, forming the backbone of persistent, always-on AI systems. Meanwhile, recent advances in proactive agents and world models signal a fundamental shift from episodic, prompt-driven tools to next-generation AI systems that continuously pe… ▽ More

    Submitted 11 May, 2026; originally announced May 2026.

    Comments: 19 pages, 7 figures

  16. arXiv:2605.09850  [pdf, ps, other

    cs.CV cs.AI

    Probing Routing-Conditional Calibration in Attention-Residual Transformers

    Authors: Wenhao Liang, Lin Yue, Wei Emma Zhang, Miao Xu, Mingyu Guo, Olaf Maennel, Weitong Chen

    Abstract: Post-hoc calibration is usually evaluated as a function of logits or softmax confidence alone, even as routing-augmented architectures increasingly accompany predictions with sample-specific internal routing traces and pair them with claims of calibration-relevant uncertainty. We ask a basic question: do these traces provide stable routing-specific evidence for post-hoc calibration beyond confiden… ▽ More

    Submitted 10 May, 2026; originally announced May 2026.

    Comments: Under reviewing

  17. arXiv:2605.08518  [pdf, ps, other

    cs.AI

    Results and Retrospective Analysis of the CODS 2025 AssetOpsBench Challenge

    Authors: Dhaval Patel, Chathurangi Shyalika, Suryanarayana Reddy Yarrabothula, Ling Yue, Shuxin Lin, Nianjun Zhou, James Rayfield

    Abstract: Competition retrospectives are useful when they explain what a leaderboard measured, how hidden evaluation changed conclusions, and which design patterns were rewarded. We revisit the CODS 2025 \assetopslive{} challenge, a privacy-aware Codabench competition on industrial multi-agent orchestration built on \assetops{}. We combine final rank sheets, a 300-submission server log, 149-team registratio… ▽ More

    Submitted 8 May, 2026; originally announced May 2026.

    Comments: 43 pages, 32 Figures

  18. arXiv:2605.06607  [pdf, ps, other

    physics.flu-dyn cs.AI

    AI CFD Scientist: Toward Open-Ended Computational Fluid Dynamics Discovery with Physics-Aware AI Agents

    Authors: Nithin Somasekharan, Rabi Pathak, Manushri Dhanakoti, Tingwen Zhang, Ling Yue, Andy Zhu, Shaowu Pan

    Abstract: Recent LLM-based agents have closed substantial portions of the scientific discovery loop in software-only machine-learning research, in chemistry, and in biology. Extending the same loop to high-fidelity physical simulators is harder, because solver completion does not imply physical validity and many failure modes appear only in field-level imagery rather than in solver logs. We present AI CFD S… ▽ More

    Submitted 12 May, 2026; v1 submitted 7 May, 2026; originally announced May 2026.

    Comments: 9 main pages and rest in appendix

  19. When AI reviews science: Can we trust the referee?

    Authors: Jialiang Wang, Yuchen Liu, Hang Xu, Kaichun Hu, Shimin Di, Wangze Ni, Linan Yue, Min-Ling Zhang, Kui Ren, Lei Chen

    Abstract: The volume of scientific submissions continues to climb, outpacing the capacity of qualified human referees and stretching editorial timelines. At the same time, modern large language models (LLMs) offer impressive capabilities in summarization, fact checking, and literature triage, making the integration of AI into peer review increasingly attractive -- and, in practice, unavoidable. Yet early de… ▽ More

    Submitted 26 April, 2026; originally announced April 2026.

    Journal ref: The Innovation Informatics 2:100030 (2026)

  20. arXiv:2604.20857  [pdf, ps, other

    cs.IR cs.AI

    DiagramBank: A Quality-Audited Dataset of Scientific Schematic Diagrams with Multi-Level Document Context

    Authors: Ling Yue, Tingwen Zhang, Jiaying Wang, Zhen Xu, Shaowu Pan

    Abstract: Scientific papers use schematic diagrams to communicate methods, workflows, and system structure, yet existing scientific-figure corpora often mix them with plots, screenshots, and photographs and rarely preserve document context. We introduce DiagramBank, a quality-audited dataset of 57,100 schematic diagrams curated from OpenReview-hosted AI/ML venues. Each record links a diagram image to its pa… ▽ More

    Submitted 27 May, 2026; v1 submitted 27 February, 2026; originally announced April 2026.

  21. arXiv:2604.10959  [pdf, ps, other

    cs.DB cs.CY

    Ozone: A Unified Platform for Transportation Research

    Authors: Ou Zheng, Ruyi Feng, Yufeng Yang, Shengxuan Ding, Lishengsa Yue, Ye Li, Yunhan Zheng, Minwei Kong, Dingyi Zhuang, Ao Qu, Zhibin Li, Meng Li, Dongjie Wang, Wangyang Ying

    Abstract: Intelligent Transportation Systems increasingly depend on heterogeneous data from roadside cameras, UAV imagery, LiDAR, and in-vehicle sensors, yet the lack of unified data standards, model interfaces, and evaluation protocols across these sources hampers reproducibility, cross-dataset benchmarking, and cross-region transferability of research findings. Existing trajectory datasets follow incompat… ▽ More

    Submitted 19 May, 2026; v1 submitted 12 April, 2026; originally announced April 2026.

  22. arXiv:2604.04074  [pdf, ps, other

    cs.AI cs.LG

    FactReview: Evidence-Grounded Peer Review with Execution-Based Claim Verification

    Authors: Ling Yue, Chaoqian Ouyang, Hang Xu, Ruijun Huang, Yuchen Liu, Libin Zheng, Wei Liu, Shaowu Pan, Shimin Di, Min-Ling Zhang

    Abstract: Large language model (LLM)-based reviewing systems typically assess manuscripts in isolation, leaving literature- and code-dependent claims difficult to verify. We present FactReview, an audit pipeline that extracts review-relevant claims, grounds them in related work and reference checks, and, when code is available, executes released artifacts under a fixed repair budget. On 26 paper-disjoint te… ▽ More

    Submitted 16 August, 2026; v1 submitted 5 April, 2026; originally announced April 2026.

  23. arXiv:2603.22386  [pdf, ps, other

    cs.AI cs.CL

    From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents

    Authors: Ling Yue, Kushal Raj Bhandari, Ching-Yun Ko, Dhaval Patel, Shuxin Lin, Nianjun Zhou, Jianxi Gao, Pin-Yu Chen, Shaowu Pan

    Abstract: Large language model (LLM)-based systems are becoming increasingly popular for solving tasks by constructing executable workflows that interleave LLM calls, information retrieval, tool use, code execution, memory updates, and verification. This survey reviews recent methods for designing and optimizing such workflows, which we treat as agentic computation graphs (ACGs). We organize the literature… ▽ More

    Submitted 23 March, 2026; originally announced March 2026.

  24. arXiv:2603.17368  [pdf, ps, other

    cs.AI

    Towards Safer Large Reasoning Models by Promoting Safety Decision-Making before Chain-of-Thought Generation

    Authors: Jianan Chen, Zhifang Zhang, Shuo He, Linan Yue, Lei Feng, Minling Zhang

    Abstract: Large reasoning models (LRMs) achieved remarkable performance via chain-of-thought (CoT), but recent studies showed that such enhanced reasoning capabilities are at the expense of significantly degraded safety capabilities. In this paper, we reveal that LRMs' safety degradation occurs only after CoT is enabled, and this degradation is not observed when CoT is disabled. This observation motivates u… ▽ More

    Submitted 3 May, 2026; v1 submitted 18 March, 2026; originally announced March 2026.

  25. arXiv:2603.09290  [pdf, ps, other

    cs.SE cs.CE cs.MA

    ToolRosella: Translating Code Repositories into Standardized Tools for Scientific Agents

    Authors: Shimin Di, Xujie Yuan, Hanghui Guo, Chaoqian Ouyang, Yongxu Liu, Ling Yue, Zhangze Chen, Libin Zheng, Jia Zhu, Shaowu Pan, Jian Yin, Yong Rui, Min-Ling Zhang

    Abstract: Large Language Model (LLM)-based agent systems are increasingly used for scientific tasks, yet their practical capability remains constrained by the narrow scope of manually curated tools they can invoke. Much scientific computational functionality already exists in open-source code repositories, but these resources remain difficult to standardize, operationalize, and invoke reliably for agent use… ▽ More

    Submitted 9 June, 2026; v1 submitted 10 March, 2026; originally announced March 2026.

    Comments: 20 pages

  26. arXiv:2603.09163  [pdf, ps, other

    cs.RO

    SPAN-Nav: Generalized Spatial Awareness for Versatile Vision-Language Navigation

    Authors: Jiahang Liu, Tianyu Xu, Jiawei Chen, Lu Yue, Jiazhao Zhang, Zhiyong Wang, Minghan Li, Qisheng Zhao, Anqi Li, Qi Su, Zhizheng Zhang, He Wang

    Abstract: Recent embodied navigation approaches leveraging Vision-Language Models (VLMs) demonstrate strong generalization in versatile Vision-Language Navigation (VLN). However, reliable path planning in complex environments remains challenging due to insufficient spatial awareness. In this work, we introduce SPAN-Nav, an end-to-end foundation model designed to infuse embodied navigation with universal 3D… ▽ More

    Submitted 9 March, 2026; originally announced March 2026.

  27. arXiv:2602.21612  [pdf, ps, other

    cs.RO

    Jumping Control for a Quadrupedal Wheeled-Legged Robot via NMPC and DE Optimization

    Authors: Xuanqi Zeng, Lingwei Zhang, Linzhu Yue, Zhitao Song, Hongbo Zhang, Tianlin Zhang, Yun-Hui Liu

    Abstract: Quadrupedal wheeled-legged robots combine the advantages of legged and wheeled locomotion to achieve superior mobility, but executing dynamic jumps remains a significant challenge due to the additional degrees of freedom introduced by wheeled legs. This paper develops a mini-sized wheeled-legged robot for agile motion and presents a novel motion control framework that integrates the Nonlinear Mode… ▽ More

    Submitted 25 February, 2026; originally announced February 2026.

    Comments: 8 pages, 12 figures

  28. arXiv:2602.09485  [pdf, ps, other

    cs.AI

    Bridging Efficiency and Transparency: Explainable CoT Compression in Multimodal Large Reasoning Models

    Authors: Yizhi Wang, Linan Yue, Min-Ling Zhang

    Abstract: Long chains of thought (Long CoTs) are widely employed in multimodal reasoning models to tackle complex tasks by capturing detailed visual information. However, these Long CoTs are often excessively lengthy and contain redundant reasoning steps, which can hinder inference efficiency. Compressing these long CoTs is a natural solution, yet existing approaches face two major challenges: (1) they may… ▽ More

    Submitted 10 February, 2026; originally announced February 2026.

  29. arXiv:2601.23032  [pdf, ps, other

    cs.AI

    Guided by Trajectories: Repairing and Rewarding Tool-Use Trajectories for Tool-Integrated Reasoning

    Authors: Siyu Gong, Linan Yue, Weibo Gao, Fangzhou Yao, Shimin Di, Lei Feng, Min-Ling Zhang

    Abstract: Tool-Integrated Reasoning (TIR) enables large language models (LLMs) to solve complex tasks by interacting with external tools, yet existing approaches depend on high-quality synthesized trajectories selected by scoring functions and sparse outcome-based rewards, providing limited and biased supervision for learning TIR. To address these challenges, in this paper, we propose AutoTraj, a two-stage… ▽ More

    Submitted 30 January, 2026; originally announced January 2026.

  30. arXiv:2601.19259  [pdf, ps, other

    cs.CE

    Learning Collective Medication Effects via Multi-level Abstraction for Medication Recommendation

    Authors: Yanda Wang, Weitong Chen, Chao Tan, Ian Nabney, Lin Yue, Genlin Ji

    Abstract: Historical prescriptions and selected candidate drugs relevant to the current visit serve as important references for medication recommendation. However, in the absence of explicit intrinsic principles for semantic composition, existing methods treat synergistic drugs as independent entities and fail to capture their collective therapeutic effects, resulting in a mismatch between medication-level… ▽ More

    Submitted 27 January, 2026; originally announced January 2026.

  31. arXiv:2601.12766  [pdf, ps, other

    cs.CV eess.SY

    Spatial-VLN: Zero-Shot Vision-and-Language Navigation With Explicit Spatial Perception and Exploration

    Authors: Lu Yue, Yue Fan, Shiwei Lian, Yu Zhao, Jiaxin Yu, Liang Xie, Feitian Zhang

    Abstract: Zero-shot Vision-and-Language Navigation (VLN) agents leveraging Large Language Models (LLMs) excel in generalization but suffer from insufficient spatial perception. Focusing on complex continuous environments, we categorize key perceptual bottlenecks into three spatial challenges: door interaction,multi-room navigation, and ambiguous instruction execution, where existing methods consistently suf… ▽ More

    Submitted 19 January, 2026; originally announced January 2026.

  32. arXiv:2512.24609  [pdf

    cs.AI

    Reinforcement Learning-Augmented LLM Agents for Collaborative Decision Making and Performance Optimization

    Authors: Dong Qiu, Duo Xu, Limengxi Yue

    Abstract: Large Language Models (LLMs) perform well in language tasks but often lack collaborative awareness and struggle to optimize global performance in multi-agent settings. We present a reinforcement learning-augmented LLM agent framework that formulates cooperation as a decentralized partially observable Markov decision process (Dec-POMDP) and adopts centralized training with decentralized execution (… ▽ More

    Submitted 30 December, 2025; originally announced December 2025.

    Comments: Accepted by IEEE ICFTIC 2025

  33. arXiv:2512.18956  [pdf, ps, other

    cs.AI cs.LG

    Training Multimodal Large Reasoning Models Needs Better Thoughts: A Three-Stage Framework for Long Chain-of-Thought Synthesis and Selection

    Authors: Yizhi Wang, Linan Yue, Min-Ling Zhang

    Abstract: Large Reasoning Models (LRMs) have demonstrated remarkable performance on complex reasoning tasks through long Chain-of-Thought (CoT) reasoning. Extending these successes to multimodal reasoning remains challenging due to the increased complexity of integrating diverse input modalities and the scarcity of high-quality long CoT training data. Existing multimodal datasets and CoT synthesis methods s… ▽ More

    Submitted 14 February, 2026; v1 submitted 21 December, 2025; originally announced December 2025.

  34. arXiv:2512.18571  [pdf, ps, other

    cs.AI cs.CV

    ESearch-R1: Learning Cost-Aware MLLM Agents for Interactive Embodied Search via Reinforcement Learning

    Authors: Weijie Zhou, Xuangtang Xiong, Ye Tian, Lijun Yue, Xinyu Wu, Wei Li, Chaoyang Zhao, Honghui Dong, Ming Tang, Jinqiao Wang, Zhengyou Zhang

    Abstract: Multimodal Large Language Models (MLLMs) have empowered embodied agents with remarkable capabilities in planning and reasoning. However, when facing ambiguous natural language instructions (e.g., "fetch the tool" in a cluttered room), current agents often fail to balance the high cost of physical exploration against the cognitive cost of human interaction. They typically treat disambiguation as a… ▽ More

    Submitted 20 December, 2025; originally announced December 2025.

  35. arXiv:2511.17532  [pdf, ps, other

    cs.NI cs.AI

    Denoising Refinement Diffusion Models for Simultaneous Generation of Multi-scale Mobile Network Traffic

    Authors: Xiaoqian Qi, Haoye Chai, Sichang Liu, Lei Yue, Raoyuan Pan, Yue Wang, Yong Li

    Abstract: The planning, management, and resource scheduling of cellular mobile networks require joint estimation of mobile traffic across different layers and nodes. Mobile traffic generation can proactively anticipate user demands and capture the dynamics of network load. However, existing methods mainly focus on generating traffic at a single spatiotemporal resolution, making it difficult to jointly model… ▽ More

    Submitted 24 November, 2025; v1 submitted 29 October, 2025; originally announced November 2025.

  36. arXiv:2511.17092  [pdf, ps, other

    cs.CV

    SPAGS: Sparse-View Articulated Object Reconstruction from Single State via Planar Gaussian Splatting

    Authors: Di Wu, Liu Liu, Xueyu Yuan, Wenxiao Chen, Lijun Yue, Liuzhu Chen, Yiming Tang, Meng Wang

    Abstract: Articulated objects are ubiquitous in daily environments, and their 3D reconstruction holds great significance across various fields. However, existing articulated object reconstruction methods typically require costly inputs such as multi-stage and multi-view observations. To address the limitations, we propose a category-agnostic articulated object reconstruction framework via planar Gaussian Sp… ▽ More

    Submitted 26 April, 2026; v1 submitted 21 November, 2025; originally announced November 2025.

    Comments: 10 pages, 7 figures

  37. arXiv:2511.16140  [pdf, ps, other

    cs.CV

    Real-Time 3D Object Detection with Inference-Aligned Learning

    Authors: Chenyu Zhao, Xianwei Zheng, Zimin Xia, Linwei Yue, Nan Xue

    Abstract: Real-time 3D object detection from point clouds is essential for dynamic scene understanding in applications such as augmented reality, robotics and navigation. We introduce a novel Spatial-prioritized and Rank-aware 3D object detection (SR3D) framework for indoor point clouds, to bridge the gap between how detectors are trained and how they are evaluated. This gap stems from the lack of spatial r… ▽ More

    Submitted 20 November, 2025; originally announced November 2025.

    Comments: Accepted by AAAI 2026

  38. arXiv:2511.14446  [pdf, ps, other

    cs.CV cs.AI

    Agentic Video Intelligence: A Flexible Framework for Advanced Video Exploration and Understanding

    Authors: Hong Gao, Yiming Bao, Xuezhen Tu, Yutong Xu, Yue Jin, Yiyang Mu, Bin Zhong, Linan Yue, Min-Ling Zhang

    Abstract: Video understanding requires not only visual recognition but also complex reasoning. While Vision-Language Models (VLMs) demonstrate impressive capabilities, they typically process videos largely in a single-pass manner with limited support for evidence revisit and iterative refinement. While recently emerging agent-based methods enable long-horizon reasoning, they either depend heavily on expensi… ▽ More

    Submitted 18 November, 2025; originally announced November 2025.

  39. arXiv:2510.24120  [pdf, ps, other

    cs.LG

    Graph-Guided Concept Selection for Efficient Retrieval-Augmented Generation

    Authors: Ziyu Liu, Yijing Liu, Jianfei Yuan, Minzhi Yan, Le Yue, Honghui Xiong, Yi Yang

    Abstract: Graph-based RAG constructs a knowledge graph (KG) from text chunks to enhance retrieval in Large Language Model (LLM)-based question answering. It is especially beneficial in domains such as biomedicine, law, and political science, where effective retrieval often involves multi-hop reasoning over proprietary documents. However, these methods demand numerous LLM calls to extract entities and relati… ▽ More

    Submitted 28 October, 2025; originally announced October 2025.

  40. arXiv:2510.17491  [pdf, ps, other

    cs.CL

    Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents

    Authors: Yihong Tang, Kehai Chen, Liang Yue, Jinxin Fan, Caishen Zhou, Xiaoguang Li, Yuyang Zhang, Mingming Zhao, Shixiong Kai, Kaiyang Guo, Xingshan Zeng, Wenjing Cun, Lifeng Shang, Min Zhang

    Abstract: With the rise of large language models (LLMs), LLM agents capable of autonomous reasoning, planning, and executing complex tasks have become a frontier in artificial intelligence. However, how to translate the research on general agents into productivity that drives industry transformations remains a significant challenge. To address this, this paper systematically reviews the technologies, applic… ▽ More

    Submitted 20 October, 2025; originally announced October 2025.

  41. arXiv:2509.20374  [pdf, ps, other

    cs.CL cs.AI

    CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics

    Authors: Nithin Somasekharan, Ling Yue, Yadi Cao, Weichao Li, Patrick Emami, Pochinapeddi Sai Bhargav, Anurag Acharya, Xingyu Xie, Shaowu Pan

    Abstract: Large Language Models (LLMs) have demonstrated strong performance across general NLP tasks, but their utility in automating numerical experiments of complex physical system -- a critical and labor-intensive component -- remains underexplored. As the major workhorse of computational science over the past decades, Computational Fluid Dynamics (CFD) offers a uniquely challenging testbed for evaluatin… ▽ More

    Submitted 26 April, 2026; v1 submitted 19 September, 2025; originally announced September 2025.

    Comments: 40 pages

    Journal ref: Journal of Data-Centric Machine Learning Research, 2026

  42. arXiv:2509.18178  [pdf, ps, other

    cs.AI cs.CE cs.LG

    Foam-Agent 2.0: An End-to-End Composable Multi-Agent Framework for Automating CFD Simulation in OpenFOAM

    Authors: Ling Yue, Nithin Somasekharan, Tingwen Zhang, Yadi Cao, Shaowu Pan

    Abstract: Computational Fluid Dynamics (CFD) is an essential simulation tool in engineering, yet its steep learning curve and complex manual setup create significant barriers. To address these challenges, we introduce Foam-Agent, a multi-agent framework that automates the entire end-to-end OpenFOAM workflow from a single natural language prompt. Our key innovations address critical gaps in existing systems:… ▽ More

    Submitted 30 September, 2025; v1 submitted 17 September, 2025; originally announced September 2025.

  43. arXiv:2509.05941  [pdf, ps, other

    cs.SE cs.LG cs.MA

    Code2MCP: Transforming Code Repositories into MCP Services

    Authors: Chaoqian Ouyang, Ling Yue, Shimin Di, Libin Zheng, Linan Yue, Shaowu Pan, Jian Yin, Min-Ling Zhang

    Abstract: The Model Context Protocol (MCP) aims to create a standard for how Large Language Models use tools. However, most current research focuses on selecting tools from an existing pool. A more fundamental, yet largely overlooked, problem is how to populate this pool by converting the vast number of existing software projects into MCP-compatible services. To bridge this gap, we introduce Code2MCP, an ag… ▽ More

    Submitted 10 February, 2026; v1 submitted 7 September, 2025; originally announced September 2025.

  44. arXiv:2508.08547  [pdf, ps, other

    cs.CV

    Calibration Attention: Learning Reliability-Aware Representations for Vision Transformers

    Authors: Wenhao Liang, Wei Emma Zhang, Lin Yue, Miao Xu, Mingyu Guo, Olaf Maennel, Weitong Chen

    Abstract: Most calibration methods operate at the logit level, implicitly assuming that miscalibration can be corrected without changing the underlying representation. We challenge this assumption and propose \textbf{Calibration Attention (CalAttn)}, a \emph{representation-aware} calibration module for vision transformers that couples instance-wise temperature scaling to transformer token geometry under a p… ▽ More

    Submitted 19 January, 2026; v1 submitted 11 August, 2025; originally announced August 2025.

    Comments: UnderReview

  45. arXiv:2508.02120  [pdf, ps, other

    cs.AI

    Don't Overthink It: A Survey of Efficient R1-style Large Reasoning Models

    Authors: Linan Yue, Yichao Du, Yizhi Wang, Weibo Gao, Fangzhou Yao, Li Wang, Ye Liu, Ziyu Xu, Qi Liu, Shimin Di, Min-Ling Zhang

    Abstract: Recently, Large Reasoning Models (LRMs) have gradually become a research hotspot due to their outstanding performance in handling complex tasks. Among them, DeepSeek R1 has garnered significant attention for its exceptional performance and open-source nature, driving advancements in the research of R1-style LRMs. Unlike traditional Large Language Models (LLMs), these models enhance logical deducti… ▽ More

    Submitted 4 August, 2025; originally announced August 2025.

  46. arXiv:2506.14278  [pdf, ps, other

    cs.RO

    Whole-Body Control Framework for Humanoid Robots with Heavy Limbs: A Model-Based Approach

    Authors: Tianlin Zhang, Linzhu Yue, Hongbo Zhang, Lingwei Zhang, Xuanqi Zeng, Zhitao Song, Yun-Hui Liu

    Abstract: Humanoid robots often face significant balance issues due to the motion of their heavy limbs. These challenges are particularly pronounced when attempting dynamic motion or operating in environments with irregular terrain. To address this challenge, this manuscript proposes a whole-body control framework for humanoid robots with heavy limbs, using a model-based approach that combines a kino-dynami… ▽ More

    Submitted 15 November, 2025; v1 submitted 17 June, 2025; originally announced June 2025.

  47. arXiv:2506.04953  [pdf, ps, other

    cs.CV

    APVR: Hour-Level Long Video Understanding with Adaptive Pivot Visual Information Retrieval

    Authors: Hong Gao, Yiming Bao, Xuezhen Tu, Bin Zhong, Linan Yue, Minling Zhang

    Abstract: Current multimodal large language models (MLLMs) struggle with hour-level video understanding, facing significant challenges not only in modeling the substantial information volume of long videos but also in overcoming the memory wall and resource constraints during both training and inference. Although recent training-free approaches have alleviated resource demands by compressing visual features… ▽ More

    Submitted 15 November, 2025; v1 submitted 5 June, 2025; originally announced June 2025.

    Comments: Accepted by AAAI 2026

  48. arXiv:2506.02689  [pdf, ps, other

    cs.CL

    MASTER: Enhancing Large Language Model via Multi-Agent Simulated Teaching

    Authors: Liang Yue, Yihong Tang, Kehai Chen, Jie Liu, Min Zhang

    Abstract: Instruction fine-tuning is crucial in NLP tasks, enhancing pretrained models' instruction-following capabilities and task-specific performance. However, obtaining high-quality fine-tuning data for large models is challenging due to data collection difficulties and high production costs. To address this, we propose MASTER, a novel data augmentation method that enriches original data through interac… ▽ More

    Submitted 3 June, 2025; v1 submitted 3 June, 2025; originally announced June 2025.

  49. Foam-Agent: A Large Language Model-Based Multi-Agent Framework for Automating Computational Fluid Dynamics Workflows

    Authors: Ling Yue, Nithin Somasekharan, Tingwen Zhang, Yadi Cao, Zhangze Chen, Shimin Di, Shaowu Pan

    Abstract: Computational fluid dynamics (CFD) has been the main workhorse of computational physics, yet its steep learning curve and fragmented, multi-stage workflow create significant barriers to entry. We present Foam-Agent, a multi-agent framework that leverages large language models (LLMs) to automate the end-to-end CFD workflow in OpenFOAM from a single natural-language prompt. Foam-Agent rests on three… ▽ More

    Submitted 13 August, 2026; v1 submitted 8 May, 2025; originally announced May 2025.

    Comments: 34 pages, 9 figures, 9 tables

    MSC Class: 76-04 ACM Class: I.2.11; I.2.1; J.2

    Journal ref: Computer Methods in Applied Mechanics and Engineering, Vol. 461, 119271 (2026)

  50. arXiv:2505.02027  [pdf, ps, other

    cs.LG cs.AI cs.SI

    GraphPrompter: Multi-stage Adaptive Prompt Optimization for Graph In-Context Learning

    Authors: Rui Lv, Zaixi Zhang, Kai Zhang, Qi Liu, Weibo Gao, Jiawei Liu, Jiaxia Yan, Linan Yue, Fangzhou Yao

    Abstract: Graph In-Context Learning, with the ability to adapt pre-trained graph models to novel and diverse downstream graphs without updating any parameters, has gained much attention in the community. The key to graph in-context learning is to perform downstream graphs conditioned on chosen prompt examples. Existing methods randomly select subgraphs or edges as prompts, leading to noisy graph prompts and… ▽ More

    Submitted 4 May, 2025; originally announced May 2025.

    Comments: 14 pages. IEEE International Conference on Data Engineering (ICDE'2025), accepted