Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 261 results for author: Feng, Q

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.12737  [pdf, ps, other

    cs.CV

    Dual-Manifold Geometry Guided Representation Learning: Adaptive Coupling between Kernel and Data Spaces

    Authors: Wencong Zhang, Yue Zhang, Meiyan Huang, Wei Yang, Qianjin Feng

    Abstract: Deep representation learning has primarily focused on how features evolve across network layers, while largely overlooking the structured geometry embedded in network parameters. We introduce a dual-manifold perspective in which each convolutional layer contains two coupled geometric spaces: a Kernel Manifold induced by convolutional filters and a Data Manifold characterized by intermediate featur… ▽ More

    Submitted 12 August, 2026; originally announced August 2026.

  2. arXiv:2608.02471  [pdf

    cs.CV cs.AI

    Action-grounded tissue affordance enables anticipatory auto-framing that lowers surgeon cognitive workload during laparoscopic surgery

    Authors: Jiayu Gu, Yiwei Wang, Jie Zhang, Guojun Cao, Keshen Lyu, Song Zhou, Yimeng Chen, Haorui Wang, Qingmin Feng, Shenchao Shi, Hongkuan Shi, Qiuyu Yu, Qiang Xie, Huan Zhao, Wenbin Chen, Caihua Xiong, Chidan Wan, Jing Samantha Pan, Xiong Cai, Han Ding

    Abstract: In laparoscopy, surgeon gaze tracks where the instruments will act; easing this demand through visual attention modeling requires dense labels of those interaction loci. These encode tacit knowledge: experts converge on consensus loci yet struggle to state the rules. Here we show that such labels can be recovered from completed actions in surgical videos, in which recorded instrument trajectories… ▽ More

    Submitted 21 August, 2026; v1 submitted 3 August, 2026; originally announced August 2026.

    Comments: Preprint. 59 pages, including supplementary information and 8 main figures

    ACM Class: I.2.10; I.4.8; I.5.4; I.2.6

  3. arXiv:2607.27713  [pdf, ps, other

    cs.RO eess.SY

    Write-Safe Flow Field Mapping under Ambiguous Onboard Sensing and Localization Drift

    Authors: Linhao Jin, Qimin Feng, Peter Gunnarson, Qiang Zhong

    Abstract: Mobile robots can infer local flow structure from onboard sensing, but a locally plausible estimate is not always safe to write into a global map. Similar flow structures may produce ambiguous observations, while localization drift causes predicted patches to be written at incorrect locations. Repeated misregistered updates then accumulate into persistent ghost structures. We address this failure… ▽ More

    Submitted 30 July, 2026; originally announced July 2026.

  4. arXiv:2607.24653  [pdf, ps, other

    cs.CL cs.LG

    Kimi K3: Open Frontier Intelligence

    Authors: Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, M. C., Jianfeng Cai, Xinyuan Cai, Peizhou Cao, Yuxuan Cao, Ziwei Chai, Y. Charles, H. S. Che, Guanduo Chen, Guangyu Chen, Guanzheng Chen, Huarong Chen, Jia Chen, Jianlong Chen, Jun Chen, Kexin Chen, Peng Chen, Ruijue Chen, Wentao Chen, Xin Chen, Yang Chen , et al. (377 additional authors not shown)

    Abstract: We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. Together with Stable LatentMoE, which effectively activates 16 of 896 routed experts per token… ▽ More

    Submitted 7 August, 2026; v1 submitted 27 July, 2026; originally announced July 2026.

    Comments: K3 tech report

  5. Nexus: Native Mesh Generation with Diffusion

    Authors: Hanxiao Wang, Ying-Tian Liu, Yuan-Chen Guo, Qi-Yuan Feng, Zi-Xin Zou, Ding Liang, Biao Zhang, Yan-Pei Cao

    Abstract: Generating high-quality triangle meshes is essential for film, gaming, and interactive 3D applications. Mainstream methods rely on mesh serialization and autoregressive processes, which stuggles in effective inference and is sensitive to error accumulation. In this paper, we present Nexus, a diffusion method that achieves holistic mesh generation via decoupled vertex and topology generation. First… ▽ More

    Submitted 15 July, 2026; originally announced July 2026.

  6. arXiv:2607.08269  [pdf, ps, other

    cs.AI

    PolyUQuest: Verifiable Structure-Aware Web RAG over Heterogeneous Graphs

    Authors: Ying Liu, Yi Ye, Quanyu Feng, Mingxi Ye, Mingtao Zhang, Haoyang Li, Chen Jason Zhang, Qing Li

    Abstract: Existing retrieval-augmented generation (RAG) systems treat web pages as flat text, losing the structural and semantic signals encoded in HTML. We present PolyUQuest, a verifiable, structure-aware web RAG framework built on a heterogeneous graph that unifies hyperlink topology between pages, DOM hierarchy within pages, and entity-relation knowledge across pages. A two-tier router dispatches each q… ▽ More

    Submitted 9 July, 2026; originally announced July 2026.

  7. arXiv:2607.02840  [pdf, ps, other

    cs.RO

    TACO: TActile World Model as a Self-COrrector forScalable VLA Post-Training

    Authors: Shengbang Liu, Yueru Jia, Yuyang Yan, Jiaming Liu, Xinran Zhang, Qiuxuan Feng, Yandong Guo, Shiji Zhou, Boxin Shi, Shanghang Zhang

    Abstract: Vision-Language-Action (VLA) models have shown promising generalization in robotic manipulation, but they still struggle with contact-rich tasks, where minor contact perturbations can cause unrecoverable failures that are hard to detect from vision alone. Since these failures are localized rather than task-level semantic errors, tactile-aware corrective post-training offers an efficient way to imp… ▽ More

    Submitted 2 July, 2026; originally announced July 2026.

  8. arXiv:2606.29883  [pdf

    cs.CV

    Building artificial intelligence virtual tissue (AIVT) for tissue state representation, feature prediction, and dynamic simulation

    Authors: Qiqi Lu, Qianjin Feng, Shaoqun Zeng, Shenghua Cheng

    Abstract: Modeling tissue states and their transitions is essential for understanding tissue homeostasis in health and pathological remodeling in disease. However, conventional computational modeling approaches are inadequate to capture the complexity of tissues as spatially organized, multiscale biological systems. Artificial intelligence (AI) has shown a remarkable ability for representing intricate syste… ▽ More

    Submitted 29 June, 2026; originally announced June 2026.

  9. arXiv:2606.28132  [pdf, ps, other

    cs.SE

    CrossLangFuzzer: Differential Testing of Cross-Language JVM Compilers

    Authors: Xiaotian Ma, Qiong Feng, Yongqiang Tian, Wei Song, Peng Liang

    Abstract: Modern JVM software increasingly integrates multiple programming languages, such as Java, Kotlin, Groovy, and Scala, within a single application. Supporting such interoperability requires JVM compilers to perform cross-language compilation while reconciling subtle semantic differences across language boundaries. Errors in this process can lead to critical miscompilations, yet existing compiler tes… ▽ More

    Submitted 26 June, 2026; originally announced June 2026.

  10. arXiv:2606.27720  [pdf, ps, other

    cs.CV

    Scene and Human in One World: Reconstruction in a Feedforward Pass

    Authors: Boao Shi, Qiao Feng, Yiming Huang, Lingjie Liu

    Abstract: Reconstructing humans in dynamic scenes from moving monocular cameras remains challenging due to scale ambiguity, human-scene misalignment, and occlusion interference. Rather than treating human mesh recovery and scene reconstruction as separate tasks, we believe that accurate human-scene reconstruction requires the two tasks to mutually inform each other: parametric human models offer semantic st… ▽ More

    Submitted 26 June, 2026; originally announced June 2026.

  11. arXiv:2606.24330  [pdf, ps, other

    cs.CV

    REDI-Match: Rotation-Equivariant Distillation for Efficient and Robust Dense Matching

    Authors: Yinji Ge, Guixu Zheng, Wulong Guo, Qian Feng, Xu Wu, Kai Zhou, Xinyuan Liu, Fei Xing

    Abstract: Vision Foundation Models (VFMs) have significantly advanced dense feature matching, yet severe in-plane rotation remains a critical challenge. Existing solutions face a fundamental dilemma: data-driven methods require inefficient parameter scaling to implicitly learn rotations, whereas strictly equivariant networks lack the semantic capacity of modern VFMs. Consequently, current frameworks typical… ▽ More

    Submitted 30 June, 2026; v1 submitted 23 June, 2026; originally announced June 2026.

  12. arXiv:2606.22082  [pdf, ps, other

    cs.SE cs.AI

    CodeTeam: An LLM-Powered Multi-Agent Framework for Repository-Level Code Generation

    Authors: Yifei Wang, Ruiyin Li, Peng Liang, Qiong Feng, Zengyang Li, Mojtaba Shahin, Arif Ali Khan

    Abstract: Natural language to repository generation (NL2Repo) requires a system to construct an entire software repository from a natural-language requirements document. Compared with function-level code generation, this task demands longer planning horizons, stable interfaces across files, and iterative debugging of cross-file inconsistencies. To address these challenges, we propose CodeTeam, an LLM-based… ▽ More

    Submitted 20 June, 2026; originally announced June 2026.

    Comments: 36 pages, 5 images, 9 tables, Manuscript submitted to a Journal (2026)

  13. arXiv:2606.15597  [pdf, ps, other

    cs.CV

    Fusion-E2Pulse: A Multimodal Event-RGB Fusion Network for Non-contact Pulse Wave Reconstruction

    Authors: Qian Feng, Hao Guo, Yan Niu, Zhenhuan Xu, Yidi Li

    Abstract: Non-contact pulse wave reconstruction hinges on the precise recovery of waveform morphology, including the dicrotic notch. Conventional Red-Green-Blue (RGB)-based methods, which extract physiological signals from recorded facial videos, are constrained by the integral imaging mechanism of standard cameras, where the exposure process induces a smoothing effect that attenuates subtle vascular pulsat… ▽ More

    Submitted 14 June, 2026; originally announced June 2026.

    Comments: Accepted by MICCAI 2026. The final version will appear in the official MICCAI proceedings published by Springer

  14. arXiv:2606.07148  [pdf, ps, other

    cs.DB

    Efficient $(α,β)$-core Computation and On-the-fly Query at Billion Scale with GPUs

    Authors: Qingshuai Feng, Shunyang Li, Kai Wang, Xuemin Lin, Kongzhang Hao, Long Yuan

    Abstract: In bipartite graphs, $(α,β)$-core is a widely used model for cohesive subgraph mining. Specifically, an $(α,β)$-core is a maximal subgraph in which each vertex in the upper layer has degree at least $α$, and each vertex in the lower layer has degree at least $β$. The state-of-the-art CPU-based solutions incur extensive costs to construct an index structure for all $α$ and $β$ combinations, leading… ▽ More

    Submitted 5 June, 2026; originally announced June 2026.

    Comments: 10 pages, 8 figures

  15. arXiv:2606.06453  [pdf, ps, other

    cs.AI

    Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents

    Authors: Zhuoming Chen, Xinrui Zhong, Qilong Feng, Ranajoy Sadhukhan, Yang Zhou, Michael Qizhe Shieh, Zhihao Jia, Beidi Chen

    Abstract: Sparse attention is becoming increasingly important for serving large language models (LLMs) as generation lengths continue to grow. However, deploying and evaluating new sparse attention algorithms at scale remains highly engineering-intensive, slowing both human researchers and AI agents in exploring the sparse attention design. To address this challenge, we present Vortex, a system that combine… ▽ More

    Submitted 4 June, 2026; originally announced June 2026.

  16. arXiv:2606.03943  [pdf, ps, other

    cs.RO cs.CV cs.LG

    PointAction: 3D Points as Universal Action Representations for Robot Control

    Authors: Mutian Tong, Han Jiang, Qiao Feng, Lingjie Liu, Jiatao Gu

    Abstract: Video-Action Models (VAMs) leverage the broad visual dynamics captured by pre-trained video diffusion models, offering a promising path toward generalizable robot manipulation. However, RGB-only video rollouts are not directly actionable: they leave metric 3D motion, contact geometry, and fine-grained spatial constraints under-specified, making action grounding ambiguous. Meanwhile, scaling action… ▽ More

    Submitted 2 June, 2026; originally announced June 2026.

    Comments: Project page: https://oriontmt.github.io/pointaction/

    ACM Class: I.2.9; I.2.10; I.2.6

  17. arXiv:2605.29237  [pdf, ps, other

    cs.CR

    Evolving Skill-Structured Attack Memory Enhances LLM Jailbreaking

    Authors: Junke Zhang, Jianwei Wang, Sishuo Chen, Yizhang He, Qingshuai Feng, Zhengyi Yang

    Abstract: Jailbreak attacks on large language models (LLMs) aim to induce LLMs to produce content that they are expected to refuse. Automated black-box jailbreak generation is important for safety evaluation, where the attacker observes only model outputs and needs to search for effective adversarial prompts. Existing black-box jailbreak methods either depend on sample-wise heuristic search or leverage atta… ▽ More

    Submitted 14 August, 2026; v1 submitted 27 May, 2026; originally announced May 2026.

    Comments: Under review

  18. arXiv:2605.26638  [pdf, ps, other

    cs.RO

    HyperSim: A Holistic Sim-To-Real Framework For Robust Robotic Manipulation

    Authors: Junyi Dong, Haotian Luo, Ziwei Xu, Shengwei Bian, Heng Zhang, Sitong Mao, Jingyi Guo, Yang Xu, Wenhao Chen, Qiuyu Feng, Yao Mu, Ping Luo, Shunbo Zhou, Xiaodong Wu

    Abstract: Scaling data volume and diversity is critical for generalizing embodied intelligence. While synthetic data generation offers a scalable alternative to expensive physical data acquisition, transferring robotic manipulation policies from simulation to the real world (sim-to-real) remains a formidable challenge due to the domain gap. This paper presents HyperSim, a holistic framework spanning from sy… ▽ More

    Submitted 26 May, 2026; originally announced May 2026.

    Comments: 9 pages, 8 figures

  19. arXiv:2605.26456  [pdf, ps, other

    cs.CV

    Sparse-LiDAR Prompting of Monocular Geometry Foundations: An Empirical Study Toward Long-Range Driving Depth

    Authors: Kai Zheng, Qiang Feng, Xingjian Liu, Wenquan Tan, Yuan Li

    Abstract: Sparse-LiDAR-prompted depth foundation models (PromptDA, Prior Depth Anything, DMD3C) have shown strong results on indoor scenes or within KITTI's standard 80-meter evaluation cap. However, two limitations remain: (i) systematic distance-stratified evaluation in long-range driving regimes (50-150 m) is largely absent; (ii) prior approaches built on disparity-based foundations rely on pre-interpola… ▽ More

    Submitted 25 May, 2026; originally announced May 2026.

    Comments: 6 pages, 3 figures, 2 tables

  20. arXiv:2605.25826  [pdf, ps, other

    math.NA cs.CE cs.LG

    Branched Signature Kernel Solvers for ODEs with rough Single-Trajectory signals

    Authors: Munawar Ali, Qi Feng, Charlie Pyle, George Xu

    Abstract: We develop a branched signature kernel solver for linear and nonlinear ordinary differential equations driven by a \emph{single observed trajectory} of a possibly rough forcing signal--a setting common within earthquake engineering, finance, biology, and structural health monitoring, where only one forcing realization is available, and the solver must respect the underlying physical law without an… ▽ More

    Submitted 13 July, 2026; v1 submitted 25 May, 2026; originally announced May 2026.

    Comments: revised for journal submission

    MSC Class: 60L10; 60L20; 46E22; 60G17; 65C20; 65C30; 60H10; 91B70

  21. arXiv:2605.22872  [pdf, ps, other

    cs.LG cs.AI cs.CV

    MedExpMem: Adapting Experience Memory for Differential Diagnosis

    Authors: Qianhan Feng, Zhongzhen Huang, Yakun Zhu, Yannian Gu, Winnie Chiu Wing Chu, Xiaofan Zhang, Qi Dou

    Abstract: Experienced physicians develop diagnostic expertise through clinical practice, acquiring not only disease knowledge but also the ability to differentiate confusable conditions. Current medical vision-language models (VLMs) lack this capability -- their parameters encode static knowledge that does not evolve across diagnostic encounters. We propose MedExpMem, an experience memory framework enabling… ▽ More

    Submitted 19 May, 2026; originally announced May 2026.

    Comments: MICCAI 2026 Early Accept. Submission Version

  22. arXiv:2605.19412  [pdf, ps, other

    cs.SE cs.PL

    DRReduce: Enhancing Syntax-Guided Program Reduction with Dependency Reconstruction

    Authors: Qiong Feng, Xiaotian Ma, Yongqiang Tian, Wei Song, Peng Liang

    Abstract: Program reduction is a technique for simplifying large, failure-inducing programs into minimal reproducible test cases. Language-specific tools such as CReduce achieve strong performance by leveraging deep semantic knowledge of C/C++, but are tightly coupled to a single language family. Language-agnostic reducers such as Perses address this by applying syntax-guided search across any grammar, yet… ▽ More

    Submitted 19 May, 2026; originally announced May 2026.

    Comments: 22 pages, 6 images, 5 tables, Manuscript submitted to a Journal (2026)

  23. arXiv:2605.13050  [pdf, ps, other

    cs.CL cs.AI

    Context Training with Active Information Seeking

    Authors: Zeyu Huang, Adhiguna Kuncoro, Qixuan Feng, Jiajun Shen, Lucio Dery, Arthur Szlam, Marc'Aurelio Ranzato

    Abstract: Most existing large language models (LLMs) are expensive to adapt after deployment, especially when a task requires newly produced information or niche domain knowledge. Recent work has shown that, by manipulating and optimizing their context, LLMs can be tailored to downstream tasks without updating their weights. However, most existing methods remain closed-loop, relying solely on the model's in… ▽ More

    Submitted 14 May, 2026; v1 submitted 13 May, 2026; originally announced May 2026.

    Comments: Preprint

  24. arXiv:2605.10942  [pdf, ps, other

    cs.RO

    HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models

    Authors: Qiuxuan Feng, Jiale Yu, Jiaming Liu, Yueru Jia, Zhuangzhe Wu, Hao Chen, Zezhong Qian, Shuo Gu, Peng Jia, Siwei Ma, Shanghang Zhang

    Abstract: World Action Models (WAMs) have emerged as a promising paradigm for robot control by modeling physical dynamics. Current WAMs generally follow two paradigms: the "Imagine-then-Execute" approach, which uses video prediction to infer actions via inverse dynamics, and the "Joint Modeling" approach, which jointly models actions and video representations. Based on systematic experiments, we observe a f… ▽ More

    Submitted 11 May, 2026; originally announced May 2026.

  25. arXiv:2605.10365  [pdf, ps, other

    cs.AI

    Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values

    Authors: Haonan Dong, Qiguan Feng, Kehan Jiang, Haoran Ye, Xin Zhang, Guojie Song

    Abstract: Autonomous agents have rapidly matured as task executors and seen widespread deployment via harnesses such as OpenClaw. Safety concerns have rightly drawn growing research attention, and beneath them lie the values silently steering agent behavior. Existing value benchmarks, however, remain confined to LLMs, leaving agent values largely uncharted. From intuitive, empirical, and theoretical vantage… ▽ More

    Submitted 11 May, 2026; originally announced May 2026.

  26. arXiv:2605.01392  [pdf, ps, other

    cs.SE cs.AI

    Using LLMs in Software Design: An Empirical Study of GitHub and A Practitioner Survey

    Authors: Yifei Wang, Ruiyin Li, Peng Liang, Yangxiao Cai, Zengyang Li, Mojtaba Shahin, Arif Ali Khan, Qiong Feng

    Abstract: Recent advancements in Large Language Models (LLMs) have demonstrated significant potential across a wide range of software engineering tasks, including software design, an area traditionally regarded as highly dependent on human expertise and judgment. However, there has been little research focusing on how LLMs are used in software design, nor on the associated benefits and drawbacks. This paper… ▽ More

    Submitted 2 May, 2026; originally announced May 2026.

    Comments: 29 pages, 8 images, 6 tables, Manuscript submitted to a Journal (2026)

  27. arXiv:2604.27654  [pdf, ps, other

    cs.CV

    MSR:Hybrid Field Modeling for CT-MRI Rigid-Deformable Registration of the Cervical Spine with an Annotated Dataset

    Authors: Bohai Zhang, Wenjie Chen, Mu Li, Kaixing Long, Xing Shen, Xinqiang Yao, Jincheng Yang, Jianting Chen, Wei Yang, Qianjin Feng, Lei Cao

    Abstract: Accurate CT-MRI registration of the cervical spine is essential for preoperative planning because this region is anatomically complex,highly variable,and vulnerable to injury of the vertebral arteries and spinal cord. However,cervical CT-MRI registration remains underexplored,particularly for rigid-deformable hybrid modeling,and the lack of high-quality annotated multimodal data further limits pro… ▽ More

    Submitted 30 April, 2026; originally announced April 2026.

  28. arXiv:2604.25080  [pdf, ps, other

    cs.DC

    CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration

    Authors: Sean Nian, Jiahao Fang, Qilong Feng, Zhiyu Wu, Fan Lai

    Abstract: KV cache restoration has emerged as a dominant bottleneck in serving long-context LLM workloads, including multi-turn conversations, retrieval-augmented generation, and agentic pipelines. Existing approaches treat restoration as a per-request tradeoff between recomputation and I/O transfer, recomputing KV states from scratch or offloading them from external storage (e.g., CPU memory or remote mach… ▽ More

    Submitted 27 April, 2026; originally announced April 2026.

    Comments: 11 pages, 10 figures

  29. arXiv:2604.16083  [pdf, ps, other

    cs.CV

    DINOv3 Beats Specialized Detectors: A Simple Foundation Model Baseline for Image Forensics

    Authors: Jieming Yu, Qiuxiao Feng, Zhuohan Wang, Xiaochen Ma

    Abstract: With the rapid advancement of deep generative models, realistic fake images have become increasingly accessible, yet existing localization methods rely on complex designs and still struggle to generalize across manipulation types and imaging conditions. We present a simple but strong baseline based on DINOv3 with LoRA adaptation and a lightweight convolutional decoder. Under the CAT-Net protocol,… ▽ More

    Submitted 17 April, 2026; originally announced April 2026.

    Comments: Technical report

  30. arXiv:2604.06373  [pdf, ps, other

    cs.SE cs.AI

    Beyond Functional Correctness: Design Issues in AI IDE-Generated Large-Scale Projects

    Authors: Syed Mohammad Kashif, Ruiyin Li, Peng Liang, Amjed Tahir, Qiong Feng, Zengyang Li, Mojtaba Shahin

    Abstract: New generation of AI coding tools, including AI-powered IDEs equipped with agentic capabilities, can generate code within the context of the project. These AI IDEs are increasingly perceived as capable of producing project-level code at scale. However, there is limited empirical evidence on the extent to which they can generate large-scale software systems and what design issues such systems may e… ▽ More

    Submitted 7 April, 2026; originally announced April 2026.

    Comments: 40 pages, 19 images, 5 tables, Manuscript submitted to a Journal (2026)

  31. arXiv:2604.01674  [pdf, ps, other

    cs.AI

    Can Heterogeneous Language Models Be Fused?

    Authors: Shilian Chen, Jie Zhou, Qin Chen, Wen Wu, Xin Li, Qi Feng, Liang He

    Abstract: Model merging aims to integrate multiple expert models into a single model that inherits their complementary strengths without incurring the inference-time cost of ensembling. Recent progress has shown that merging can be highly effective when all source models are \emph{homogeneous}, i.e., derived from the same pretrained backbone and therefore share aligned parameter coordinates or compatible ta… ▽ More

    Submitted 15 May, 2026; v1 submitted 2 April, 2026; originally announced April 2026.

  32. arXiv:2603.26699  [pdf, ps, other

    eess.SP cs.CV cs.LG

    EMPD: An Event-based Multimodal Physiological Dataset for Remote Pulse Wave Detection

    Authors: Qian Feng, Pengfei Li, Rongshan Gao, Jiale Xu, Rui Gong, Yidi Li

    Abstract: Remote photoplethysmography (rPPG) based on traditional frame-based cameras often struggles with motion artifacts and limited temporal resolution. To address these limitations, we introduce EMPD (Event-based Multimodal Physiological Dataset), the first benchmark dataset specifically designed for non-contact physiological sensing via event cameras. The dataset leverages a laser-assisted acquisition… ▽ More

    Submitted 16 March, 2026; originally announced March 2026.

    Comments: 12 pages, 4 figures, 2 tables

  33. arXiv:2603.15618  [pdf, ps, other

    cs.CV

    Look Before Acting: Enhancing Vision Foundation Representations for Vision-Language-Action Models

    Authors: Yulin Luo, Hao Chen, Zhuangzhe Wu, Bowen Sui, Jiaming Liu, Chenyang Gu, Zhuoyang Liu, Qiuxuan Feng, Jiale Yu, Shuo Gu, Peng Jia, Pheng-Ann Heng, Shanghang Zhang

    Abstract: Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for robotic manipulation, in which reliable action prediction critically depends on accurately interpreting and integrating visual observations conditioned on language instructions. Although recent works have sought to enhance the visual capabilities of VLA models, most approaches treat the LLM backbone as a black bo… ▽ More

    Submitted 17 March, 2026; v1 submitted 16 March, 2026; originally announced March 2026.

  34. arXiv:2603.03537  [pdf, ps, other

    cs.RO

    Passive Phase-Oriented Impedance Shaping for Rapid Acceleration in Soft Robotic Swimmers

    Authors: Qimin Feng, Orion A. Roberts, Qiang Zhong

    Abstract: Rapid acceleration and burst maneuvers in underwater robots depend less on maintaining precise resonance and more on force--velocity phase alignment during thrust generation. In this work, we investigate constrained-layer damping (CLD) as a passive mechanism for frequency-selective impedance shaping in soft robotic swimmers. Unlike conventional stiffness-tuning approaches, CLD selectively amplifie… ▽ More

    Submitted 16 July, 2026; v1 submitted 3 March, 2026; originally announced March 2026.

    Comments: Accepted to the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026). Camera-ready version

  35. arXiv:2602.23926  [pdf, ps, other

    cs.CV

    Leveraging Geometric Prior Uncertainty and Complementary Constraints for High-Fidelity Neural Indoor Surface Reconstruction

    Authors: Qiyu Feng, Jiwei Shan, Shing Shin Cheng, Hesheng Wang

    Abstract: Neural implicit surface reconstruction with signed distance function has made significant progress, but recovering fine details such as thin structures and complex geometries remains challenging due to unreliable or noisy geometric priors. Existing approaches rely on implicit uncertainty that arises during optimization to filter these priors, which is indirect and inefficient, and masking supervis… ▽ More

    Submitted 27 February, 2026; originally announced February 2026.

    Comments: Accepted by ICRA 2026

  36. arXiv:2602.16444  [pdf, ps, other

    cs.RO cs.AI cs.LG

    RoboGene: Boosting VLA Pre-training via Diversity-Driven Agentic Framework for Real-World Task Generation

    Authors: Yixue Zhang, Kun Wu, Zhi Gao, Zhen Zhao, Pei Ren, Zhiyuan Xu, Fei Liao, Xinhua Wang, Shichao Fan, Di Wu, Qiuxuan Feng, Meng Li, Zhengping Che, Chang Liu, Jian Tang

    Abstract: The pursuit of general-purpose robotic manipulation is hindered by the scarcity of diverse, real-world interaction data. Unlike data collection from web in vision or language, robotic data collection is an active process incurring prohibitive physical costs. Consequently, automated task curation to maximize data value remains a critical yet under-explored challenge. Existing manual methods are uns… ▽ More

    Submitted 18 February, 2026; v1 submitted 18 February, 2026; originally announced February 2026.

  37. arXiv:2602.13964  [pdf, ps, other

    cs.CL

    HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam

    Authors: Weiqi Zhai, Zhihai Wang, Jinghang Wang, Boyu Yang, Xiaogang Li, Xander Xu, Bohan Wang, Peng Wang, Xingzhe Wu, Anfeng Li, Qiyuan Feng, Yuhao Zhou, Taolin Han, Wenjie Luo, Yiyuan Li, Xiang Zheng, Yaxuan Wang, Ruixiang Luo, Guojie Lin, Peiyao Xiao, Chengliang Xu, Ben Wang, Zeyu Wang, Zichao Chen, Jianan Ye , et al. (11 additional authors not shown)

    Abstract: Humanity's Last Exam (HLE) has become a widely used benchmark for evaluating frontier large language models on challenging, multi-domain questions. However, community-led analyses have raised concerns that HLE contains a non-trivial number of noisy items, which can bias evaluation results and distort cross-model comparisons. To address this challenge, we introduce HLE-Verified, a verified and revi… ▽ More

    Submitted 17 August, 2026; v1 submitted 14 February, 2026; originally announced February 2026.

    Comments: 14 pages, 10 figures

  38. arXiv:2602.07092  [pdf, ps, other

    cs.MA cs.AI

    Lemon Agent Technical Report

    Authors: Haipeng Jiang, Kailong Ren, Zimo Yin, Zhetao Sun, Xin Gan, Guangyi Lv, Ming He, Peng Wang, Congli Yin, Hong Pan, Changwen Zhang, Shan Tong, Zhengyu Xu, Zeping Chen, Yubin Huangfu, Yanzhi Xu, Xing Su, Qin Feng, Dong An, Jianping Fan

    Abstract: Recent advanced LLM-powered agent systems have exhibited their remarkable capabilities in tackling complex, long-horizon tasks. Nevertheless, they still suffer from inherent limitations in resource efficiency, context management, and multimodal perception. Based on these observations, Lemon Agent is introduced, a multi-agent orchestrator-worker system built on a newly proposed AgentCortex framewor… ▽ More

    Submitted 6 February, 2026; originally announced February 2026.

  39. arXiv:2602.04244  [pdf, ps, other

    cs.LG

    GraphVec: Cross-Domain Graph Vectorization for Graph-Level Representation Learning

    Authors: Qi Feng, Jicong Fan

    Abstract: Learning universal graph representations across heterogeneous domains is difficult because graph datasets differ in topology, node-attribute semantics, feature dimensions, and even attribute availability. We propose GraphVec, a language-model-free graph vectorization model that maps diverse graphs into transferable fixed-dimensional embeddings for graph-level tasks. Instead of directly using incom… ▽ More

    Submitted 7 May, 2026; v1 submitted 4 February, 2026; originally announced February 2026.

  40. arXiv:2602.02276  [pdf, ps, other

    cs.CL cs.AI cs.LG

    Kimi K2.5: Visual Agentic Intelligence

    Authors: Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, S. H. Cai, Yuan Cao, Ziwei Chai, Y. Charles, H. S. Che, Cheng Chen, Guanduo Chen, Huarong Chen, Jia Chen, Jianlong Chen, Jun Chen, Kefan Chen, Liang Chen, Ruijue Chen, Xinhao Chen, Yanru Chen, Yanxu Chen, Yicun Chen, Yimin Chen, Yingjiang Chen, Yuankun Chen , et al. (312 additional authors not shown)

    Abstract: We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint optimization of text and vision so that two modalities enhance each other. This includes a series of techniques such as joint text-vision pre-training, zero-vision SFT, and joint text-vision reinforcement learning. Building on this multimodal foundation, K2.5… ▽ More

    Submitted 7 August, 2026; v1 submitted 2 February, 2026; originally announced February 2026.

    Comments: Kimi K2.5 tech report

  41. arXiv:2601.19352  [pdf, ps, other

    cs.LG

    GraphSB: Boosting Imbalanced Node Classification on Graphs through Structural Balance

    Authors: Zhixiao Wang, Chaofan Zhu, Qihan Feng, Jian Zhang, Xiaobin Rui, Philip S Yu

    Abstract: Imbalanced node classification is a critical challenge in graph learning, where most existing methods typically utilize Graph Neural Networks (GNNs) to learn node representations. These methods can be broadly categorized into the data-level and the algorithm-level. The former aims to synthesize minority-class nodes to mitigate quantity imbalance, while the latter tries to optimize the learning pro… ▽ More

    Submitted 27 January, 2026; originally announced January 2026.

  42. arXiv:2601.15250  [pdf, ps, other

    cs.CV cs.RO

    FlowSSC: Universal Generative Monocular Semantic Scene Completion via One-Step Latent Diffusion

    Authors: Zichen Xi, Hao-Xiang Chen, Nan Xue, Hongyu Yan, Qi-Yuan Feng, Levent Burak Kara, Joaquim Jorge, Qun-Ce Xu

    Abstract: Semantic Scene Completion (SSC) from monocular RGB images is a fundamental yet challenging task due to the inherent ambiguity of inferring occluded 3D geometry from a single view. While feed-forward methods have made progress, they often struggle to generate plausible details in occluded regions and preserve the fundamental spatial relationships of objects. Such accurate generative reasoning capab… ▽ More

    Submitted 21 January, 2026; originally announced January 2026.

    Comments: Under Review

  43. arXiv:2601.13578  [pdf, ps, other

    cs.LG cs.CV

    FG-OrIU: Towards Better Forgetting via Feature-Gradient Orthogonality for Incremental Unlearning

    Authors: Qian Feng, JiaHang Tu, Mintong Kang, Hanbin Zhao, Chao Zhang, Hui Qian

    Abstract: Incremental unlearning (IU) is critical for pre-trained models to comply with sequential data deletion requests, yet existing methods primarily suppress parameters or confuse knowledge without explicit constraints on both feature and gradient level, resulting in \textit{superficial forgetting} where residual information remains recoverable. This incomplete forgetting risks security breaches and di… ▽ More

    Submitted 19 January, 2026; originally announced January 2026.

    Comments: This paper has been accepted by ICCV 2025. code: \url{https://github.com/RAIAN08/FG-OrIU}

  44. arXiv:2601.06123  [pdf, ps, other

    cs.LG cs.AI

    Latent Space Communication via K-V Cache Alignment

    Authors: Lucio M. Dery, Zohar Yahav, Henry Prior, Qixuan Feng, Jiajun Shen, Arthur Szlam

    Abstract: Solving increasingly complex problems with large language models (LLMs) necessitates a move beyond individual models and towards multi-model systems that can effectively collaborate. While text has traditionally served as the medium for inter-model communication, a richer and more efficient exchange is possible if models can access each other's internal states directly. In this paper, we propose l… ▽ More

    Submitted 3 January, 2026; originally announced January 2026.

    Comments: 15 pages, 6 figures, 4 tables

  45. arXiv:2601.04377  [pdf, ps, other

    cs.CL cs.AI cs.LG

    Disco-RAG: Discourse-Aware Retrieval-Augmented Generation

    Authors: Dongqi Liu, Hang Ding, Qiming Feng, Xurong Xie, Zhucun Xue, Chengjie Wang, Jian Li, Jiangning Zhang, Yabiao Wang

    Abstract: Retrieval-Augmented Generation (RAG) has emerged as an important means of enhancing the performance of large language models (LLMs) in knowledge-intensive tasks. However, most existing RAG strategies treat retrieved passages in a flat and unstructured way, which prevents the model from capturing structural cues and constrains its ability to synthesize knowledge from dispersed evidence across docum… ▽ More

    Submitted 17 April, 2026; v1 submitted 7 January, 2026; originally announced January 2026.

    Comments: ACL 2026 Main & Long Conference Paper

  46. arXiv:2512.24653  [pdf, ps, other

    cs.RO

    RoboMIND 2.0: A Multimodal, Bimanual Mobile Manipulation Dataset for Generalizable Embodied Intelligence

    Authors: Chengkai Hou, Kun Wu, Jiaming Liu, Zhengping Che, Di Wu, Fei Liao, Guangrun Li, Jingyang He, Qiuxuan Feng, Zhao Jin, Chenyang Gu, Zhuoyang Liu, Nuowei Han, Xiangju Mi, Yaoxu Lv, Yankai Fu, Gaole Dai, Langzhe Gu, Tao Li, Yuheng Zhang, Yixue Zhang, Xinhua Wang, Shichao Fan, Meng Li, Zhen Zhao , et al. (8 additional authors not shown)

    Abstract: While data-driven imitation learning has revolutionized robotic manipulation, current approaches remain constrained by the scarcity of large-scale, diverse real-world demonstrations. Consequently, the ability of existing models to generalize across long-horizon bimanual tasks and mobile manipulation in unstructured environments remains limited. To bridge this gap, we present RoboMIND 2.0, a compre… ▽ More

    Submitted 27 February, 2026; v1 submitted 31 December, 2025; originally announced December 2025.

  47. arXiv:2512.22827  [pdf, ps, other

    cs.SE cs.AI

    FasterPy: An LLM-based Code Execution Efficiency Optimization Framework

    Authors: Yue Wu, Minghao Han, Ruiyin Li, Peng Liang, Amjed Tahir, Zengyang Li, Qiong Feng, Mojtaba Shahin

    Abstract: Code often suffers from performance bugs. These bugs necessitate the research and practice of code optimization. Traditional rule-based methods rely on manually designing and maintaining rules for specific performance bugs (e.g., redundant loops, repeated computations), making them labor-intensive and limited in applicability. In recent years, machine learning and deep learning-based methods have… ▽ More

    Submitted 15 June, 2026; v1 submitted 28 December, 2025; originally announced December 2025.

    Comments: 38 pages, 5 images, 14 tables, Manuscript revision submitted to a Journal (2026)

  48. arXiv:2512.20145  [pdf, ps, other

    cs.CL cs.AI cs.CV cs.IR cs.LG

    Retrieval-augmented Prompt Learning for Pre-trained Foundation Models

    Authors: Xiang Chen, Yixin Ou, Quan Feng, Lei Li, Piji Li, Haibo Ye, Sheng-Jun Huang, Shuofei Qiao, Shumin Deng, Huajun Chen, Ningyu Zhang

    Abstract: The pre-trained foundation models (PFMs) have become essential for facilitating large-scale multimodal learning. Researchers have effectively employed the ``pre-train, prompt, and predict'' paradigm through prompt learning to induce improved few-shot performance. However, prompt learning approaches for PFMs still follow a parametric learning paradigm. As such, the stability of generalization in me… ▽ More

    Submitted 23 December, 2025; originally announced December 2025.

    Comments: IEEE/ACM Transactions on Audio, Speech and Language Processing

  49. arXiv:2512.16969  [pdf, ps, other

    cs.AI cs.CL cs.LG

    Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows

    Authors: Wanghan Xu, Yuhao Zhou, Yifan Zhou, Qinglong Cao, Shuo Li, Jia Bu, Bo Liu, Yixin Chen, Xuming He, Xiangyu Zhao, Xiang Zhuang, Fengxiang Wang, Zhiwang Zhou, Qiantai Feng, Wenxuan Huang, Jiaqi Wei, Hao Wu, Yuejin Yang, Guangshuai Wang, Sheng Xu, Ziyan Huang, Xinyao Liu, Jiyao Liu, Cheng Tang, Wei Li , et al. (82 additional authors not shown)

    Abstract: Despite advances in scientific AI, a coherent framework for Scientific General Intelligence (SGI)-the ability to autonomously conceive, investigate, and reason across scientific domains-remains lacking. We present an operational SGI definition grounded in the Practical Inquiry Model (PIM: Deliberation, Conception, Action, Perception) and operationalize it via four scientist-aligned tasks: deep res… ▽ More

    Submitted 18 December, 2025; originally announced December 2025.

  50. arXiv:2512.15171  [pdf, ps, other

    cs.CV

    Cross-modal ultra-scale learning with tri-modalities of renal biopsy images for glomerular multi-disease auxiliary diagnosis

    Authors: Kaixing Long, Danyi Weng, Yun Mi, Zhentai Zhang, Yanmeng Lu, Jian Geng, Zhitao Zhou, Liming Zhong, Qianjin Feng, Wei Yang, Lei Cao

    Abstract: Constructing a multi-modal automatic classification model based on three types of renal biopsy images can assist pathologists in glomerular multi-disease identification. However, the substantial scale difference between transmission electron microscopy (TEM) image features at the nanoscale and optical microscopy (OM) or immunofluorescence microscopy (IM) images at the microscale poses a challenge… ▽ More

    Submitted 17 December, 2025; originally announced December 2025.