Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 174 results for author: Ge, R

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.02806  [pdf

    cs.CR cs.CV

    Fast Object Removal Attacks on Safety-Critical Video-based Perception Systems

    Authors: Mohammad Imtiaz Hasan, M Sabbir Salek, Nathan Jones, Mashrur Chowdhury, Rong Ge

    Abstract: By leveraging data from video-based perception systems, intelligent transportation systems (ITS) support safety-critical applications that improve road safety. However, adversaries may manipulate video frames to compromise downstream perception modules, causing failures in safety-critical functions and increasing risks to vulnerable road users. This paper presents a novel attack model and an end-t… ▽ More

    Submitted 3 August, 2026; originally announced August 2026.

  2. arXiv:2607.26208  [pdf, ps, other

    cs.ET cond-mat.mes-hall quant-ph

    MEDA: Measurement-Efficient Disorder-Aware Majorana Zero Mode Detection in Realistic Devices

    Authors: Nathan Jones, Binayyak Roy, Valentine Mohaugen, Ian Lewis, Toby Cox, Sumanta Tewari, Rong Ge

    Abstract: Fault-tolerant topological quantum computing relies on identifying Majorana zero modes (MZMs), but reliable detection in realistic devices remains challenging. Conventional topological indicators are inherently biased in finite, disordered systems, blurring the distinction between true MZMs and trivial states. Furthermore, attempts to map these indicators to real observables via machine learning r… ▽ More

    Submitted 16 August, 2026; v1 submitted 28 July, 2026; originally announced July 2026.

    Comments: Accepted to 2026 IEEE International Conference on Quantum Computing and Engineering

  3. arXiv:2607.17341  [pdf, ps, other

    cs.CV

    Understanding From Human Perspective: A Multi-agent System for Interactive Egocentric Medical Image Segmentation

    Authors: Rongjun Ge, Dongyang Wang, Heng Zhu, Zhirui Li, Yang Chen, Yuting He

    Abstract: Interactive egocentric medical image segmentation (IEMIS) plays an important role in smart-glasses-assisted medical image review, segmenting the medical targets a clinician refers to from their egocentric view. Once it succeeds, the object-level visual evidence it provides strengthens the review and underpins fine-grained analysis and clinical decision-making. However, the instruction and the vide… ▽ More

    Submitted 19 July, 2026; originally announced July 2026.

  4. arXiv:2607.07039  [pdf

    eess.IV cs.CV physics.med-ph

    From Data Completeness to Data Sufficiency: A Task-Driven Imaging Framework for Intraoperative CBCT under Quality-Time-Dose Trade-offs

    Authors: Yi Jia, Rongjun Ge, Yang Chen, Yan Xi, Wenjun Xia

    Abstract: Mobile C-arm cone-beam computed tomography (CBCT) has been widely used for real-time intraoperative 3D imaging. However, current practice often mechanically applies the fan-beam CT criterion of "180° plus fan angle" in pursuit of "data completeness" in reconstruction. This review argues that, under the single circular trajectory of three-dimensional cone-beam geometry, complete data are mathematic… ▽ More

    Submitted 8 July, 2026; originally announced July 2026.

  5. arXiv:2606.20990  [pdf, ps, other

    cs.RO

    Duet: Dual-Robot Understanding via Efficient Teaching

    Authors: Yiqi Zhao, Ruohai Ge, Celina Shiyu Wang, Junjie Ye, Muchen Xu, Minhao Li, Sergey Zakharov, Basile Van Hoorick, Vitor Campagnolo Guizilini, Leonidas Guibas, Gaurav S. Sukhatme, Jyotirmoy V. Deshmukh, Yue Wang

    Abstract: Dual-robot collaboration enables tasks that exceed the reach and payload of a single robot, such as collaboratively transporting objects across environments and executing coordinated handovers. Data acquisition is the primary bottleneck for training these systems. To this end, we introduce DUET, a dual-robot learning framework for mobile manipulation. For efficient data collection, we create a uni… ▽ More

    Submitted 18 June, 2026; originally announced June 2026.

  6. arXiv:2606.19348  [pdf, ps, other

    cs.CL cs.AI

    DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

    Authors: DeepSeek-AI, Anyi Xu, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, Chenchen Ling, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chengyu Hou, Chenhao Xu, Chenze Shao, Chong Ruan, Conner Sun, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Donghao Li, Dongjie Ji , et al. (294 additional authors not shown)

    Abstract: We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) -- both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: (1) a hybrid attention arc… ▽ More

    Submitted 26 April, 2026; originally announced June 2026.

  7. arXiv:2605.27774  [pdf, ps, other

    cs.LG

    Fine-Tuning Dynamics of In-Context Factual Recall in Transformers

    Authors: Ruomin Huang, Eshaan Nichani, Jason D. Lee, Rong Ge

    Abstract: In-context learning \ -- performing tasks based on examples given in the prompt \ -- is an important capability that has emerged in large language models and has received significant attention in both theory and practice. Existing theoretical work often focuses on settings where the learning uses information purely from the prompt. However, many practical instances of in-context learning require t… ▽ More

    Submitted 26 May, 2026; originally announced May 2026.

  8. arXiv:2605.06230  [pdf, ps, other

    cs.AI cs.DC

    Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence

    Authors: Xinquan Chen, Zhenyun Yin, Shan He, Bin Huang, Shanzhe Lei, Pengcheng Shi, Kun Cai, Bei Chen, Bangwei Liu, Zeyu Kang, Chao Huang, Yang Zhang, Wenjie Li, Ruijun Ge, Yajie Wang, Tianshun Fang, Tianyang Xu, Yiwen Cong, Meng Jin, Gaolei Li, Xuansheng Wu, Linhan Liu, Zijing He, An Li, Yan Teng , et al. (16 additional authors not shown)

    Abstract: As large models evolve from conversational assistants into autonomous agents, challenges increasingly arise from long-horizon decision making, tool use, and real environment interaction. Existing agenticinfrastructure remain fragmented across evaluation, data management, and agent evolution, making it difficult to discover risks systematically and improve models in a continuous closed loop. In thi… ▽ More

    Submitted 7 May, 2026; v1 submitted 7 May, 2026; originally announced May 2026.

    Comments: 50 pages, 21 figures

  9. arXiv:2604.15459  [pdf, ps, other

    eess.IV cs.AI cs.CV

    RelativeFlow: Taming Medical Image Denoising Learning with Noisy Reference

    Authors: Yuxin Liu, Yiqing Dong, Wenxue Yu, Zhan Wu, Rongjun Ge, Yang Chen, Yuting He

    Abstract: Medical image denoising (MID) lacks absolutely clean images for supervision, leading to a noisy reference problem that fundamentally limits denoising performance. Existing simulated-supervised discriminative learning (SimSDL) and simulated-supervised generative learning (SimSGL) treat noisy references as clean targets, causing suboptimal convergence or reference-biased learning, while self-supervi… ▽ More

    Submitted 16 April, 2026; originally announced April 2026.

    Comments: Accepted by CVPR 2026

  10. arXiv:2604.12162  [pdf, ps, other

    cs.CL

    AlphaEval: Evaluating Agents in Production

    Authors: Pengrui Lu, Bingyu Xu, Wenjun Zhang, Shengjia Hua, Xuanjian Gao, Ranxiang Ge, Lyumanshan Ye, Linxuan Wu, Yiran Li, Junfei Fish Yu, Yibo Zhang, Ruixin Li, Manxiang Li, Xiao Han, Xiaocong Zhou, Guangyao Chi, Zisheng Chen, Kaishen Chen, Kun Wang, Qihua Xu, Fengyue Meng, Yuchen Ni, Jiajun Li, Jinxiu Liu, Danfeng Zhang , et al. (2 additional authors not shown)

    Abstract: The rapid deployment of AI agents in commercial settings has outpaced the development of evaluation methodologies that reflect production realities. Existing benchmarks measure agent capabilities through retrospectively curated tasks with well-specified requirements and deterministic metrics -- conditions that diverge fundamentally from production environments where requirements contain implicit c… ▽ More

    Submitted 13 April, 2026; originally announced April 2026.

  11. arXiv:2603.18157  [pdf, ps, other

    cs.DS cs.LG

    Learning-Augmented Algorithms for $k$-median via Online Learning

    Authors: Anish Hebbar, Rong Ge, Amit Kumar, Debmalya Panigrahi

    Abstract: The field of learning-augmented algorithms seeks to use ML techniques on past instances of a problem to inform an algorithm designed for a future instance. In this paper, we introduce a novel model for learning-augmented algorithms inspired by online learning. In this model, we are given a sequence of instances of a problem and the goal of the learning-augmented algorithm is to use prior instances… ▽ More

    Submitted 18 March, 2026; originally announced March 2026.

    Comments: NeurIPS 2025

  12. arXiv:2603.16843  [pdf, ps, other

    cs.AI

    Internalizing Agency from Reflective Experience

    Authors: Rui Ge, Yichao Fu, Yuyang Qian, Junda Su, Yiming Zhao, Peng Zhao, Hao Zhang

    Abstract: Large language models are increasingly deployed as autonomous agents that must plan, act, and recover from mistakes through long-horizon interaction with environments that provide rich feedback. However, prevailing outcome-driven post-training methods (e.g., RL with verifiable rewards) primarily optimize final success signals, leaving rich environment feedback underutilized. Consequently, they oft… ▽ More

    Submitted 17 March, 2026; originally announced March 2026.

    Comments: 17 pages, 5 figures; Submitted to ICML 2026

  13. arXiv:2602.19285  [pdf, ps, other

    cs.CV

    MRI Contrast Enhancement Kinetics World Model

    Authors: Jindi Kong, Yuting He, Cong Xia, Rongjun Ge, Shuo Li

    Abstract: Clinical MRI contrast acquisition suffers from inefficient information yield, which presents as a mismatch between the risky and costly acquisition protocol and the fixed and sparse acquisition sequence. Applying world models to simulate the contrast enhancement kinetics in the human body enables continuous contrast-free dynamics. However, the low temporal resolution in MRI acquisition restricts t… ▽ More

    Submitted 22 February, 2026; originally announced February 2026.

    Comments: Accepted by CVPR 2026

  14. arXiv:2512.20898  [pdf, ps, other

    cs.CV cs.AI

    DGSAN: Dual-Graph Spatiotemporal Attention Network for Pulmonary Nodule Malignancy Prediction

    Authors: Xiao Yu, Zhaojie Fang, Guanyu Zhou, Yin Shen, Huoling Luo, Ye Li, Ahmed Elazab, Xiang Wan, Ruiquan Ge, Changmiao Wang

    Abstract: Lung cancer continues to be the leading cause of cancer-related deaths globally. Early detection and diagnosis of pulmonary nodules are essential for improving patient survival rates. Although previous research has integrated multimodal and multi-temporal information, outperforming single modality and single time point, the fusion methods are limited to inefficient vector concatenation and simple… ▽ More

    Submitted 23 December, 2025; originally announced December 2025.

  15. arXiv:2512.11251  [pdf, ps, other

    cs.LG

    Insight Miner: A Time Series Analysis Dataset for Cross-Domain Alignment with Natural Language

    Authors: Yunkai Zhang, Yawen Zhang, Ming Zheng, Kezhen Chen, Chongyang Gao, Ruian Ge, Siyuan Teng, Amine Jelloul, Jinmeng Rao, Xiaoyuan Guo, Chiang-Wei Fang, Zeyu Zheng, Jie Yang

    Abstract: Time-series data is critical across many scientific and industrial domains, including environmental analysis, agriculture, transportation, and finance. However, mining insights from this data typically requires deep domain expertise, a process that is both time-consuming and labor-intensive. In this paper, we propose \textbf{Insight Miner}, a large-scale multimodal model (LMM) designed to generate… ▽ More

    Submitted 11 December, 2025; originally announced December 2025.

  16. arXiv:2512.08337  [pdf, ps, other

    cs.CV

    DINO-BOLDNet: A DINOv3-Guided Multi-Slice Attention Network for T1-to-BOLD Generation

    Authors: Jianwei Wang, Qing Wang, Menglan Ruan, Rongjun Ge, Chunfeng Yang, Yang Chen, Chunming Xie

    Abstract: Generating BOLD images from T1w images offers a promising solution for recovering missing BOLD information and enabling downstream tasks when BOLD images are corrupted or unavailable. Motivated by this, we propose DINO-BOLDNet, a DINOv3-guided multi-slice attention framework that integrates a frozen self-supervised DINOv3 encoder with a lightweight trainable decoder. The model uses DINOv3 to extra… ▽ More

    Submitted 9 December, 2025; originally announced December 2025.

  17. arXiv:2512.02556  [pdf, ps, other

    cs.CL

    DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

    Authors: DeepSeek-AI, Aixin Liu, Aoxue Mei, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenhao Xu, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Erhang Li, Fangqi Zhou, Fangyun Lin, Fucong Dai, Guangbo Hao , et al. (239 additional authors not shown)

    Abstract: We introduce DeepSeek-V3.2, a model that harmonizes high computational efficiency with superior reasoning and agent performance. The key technical breakthroughs of DeepSeek-V3.2 are as follows: (1) DeepSeek Sparse Attention (DSA): We introduce DSA, an efficient attention mechanism that substantially reduces computational complexity while preserving model performance in long-context scenarios. (2)… ▽ More

    Submitted 2 December, 2025; originally announced December 2025.

  18. arXiv:2512.01274  [pdf, ps, other

    cs.CL cs.AI cs.LG

    SUPERChem: A Multimodal Reasoning Benchmark in Chemistry

    Authors: Zehua Zhao, Zhixian Huang, Junren Li, Siyu Lin, Junting Zhou, Fengqi Cao, Kun Zhou, Rui Ge, Tingting Long, Yuexiang Zhu, Yan Liu, Jie Zheng, Junnian Wei, Rong Zhu, Peng Zou, Wenyu Li, Zekai Cheng, Tian Ding, Yaxuan Wang, Yizhao Yan, Tingru Wei, Haowei Ming, Weijie Mao, Chen Sun, Yiming Liu , et al. (6 additional authors not shown)

    Abstract: Current benchmarks for evaluating the chemical reasoning capabilities of Large Language Models (LLMs) are limited by oversimplified tasks, lack of process-level evaluation, and misalignment with expert-level chemistry skills. To address these issues, we introduce SUPERChem, a benchmark of 500 expert-curated reasoning-intensive chemistry problems, covering diverse subfields and provided in both mul… ▽ More

    Submitted 30 November, 2025; originally announced December 2025.

    Comments: 35 pages, 11 figures, 5 tables

  19. arXiv:2511.21042  [pdf, ps, other

    cs.CV

    LungNoduleAgent: A Collaborative Multi-Agent System for Precision Diagnosis of Lung Nodules

    Authors: Cheng Yang, Hui Jin, Xinlei Yu, Zhipeng Wang, Yaoqun Liu, Fenglei Fan, Dajiang Lei, Gangyong Jia, Changmiao Wang, Ruiquan Ge

    Abstract: Diagnosing lung cancer typically involves physicians identifying lung nodules in Computed tomography (CT) scans and generating diagnostic reports based on their morphological features and medical expertise. Although advancements have been made in using multimodal large language models for analyzing lung CT scans, challenges remain in accurately describing nodule morphology and incorporating medica… ▽ More

    Submitted 25 November, 2025; originally announced November 2025.

    Comments: Accepted by AAAI 2026

  20. arXiv:2511.17765  [pdf, ps, other

    cs.RO cs.LG cs.MA

    LEARN: Learning End-to-End Aerial Resource-Constrained Multi-Robot Navigation

    Authors: Darren Chiu, Zhehui Huang, Ruohai Ge, Gaurav S. Sukhatme

    Abstract: Nano-UAV teams offer great agility yet face severe navigation challenges due to constrained onboard sensing, communication, and computation. Existing approaches rely on high-resolution vision or compute-intensive planners, rendering them infeasible for these platforms. We introduce LEARN, a lightweight, two-stage safety-guided reinforcement learning (RL) framework for multi-UAV navigation in clutt… ▽ More

    Submitted 21 November, 2025; originally announced November 2025.

    Comments: 20 pages, 15 figures

  21. arXiv:2511.08987  [pdf, ps, other

    cs.CV

    WDT-MD: Wavelet Diffusion Transformers for Microaneurysm Detection in Fundus Images

    Authors: Yifei Sun, Yuzhi He, Junhao Jia, Jinhong Wang, Ruiquan Ge, Changmiao Wang, Hongxia Xu

    Abstract: Microaneurysms (MAs), the earliest pathognomonic signs of Diabetic Retinopathy (DR), present as sub-60 $μm$ lesions in fundus images with highly variable photometric and morphological characteristics, rendering manual screening not only labor-intensive but inherently error-prone. While diffusion-based anomaly detection has emerged as a promising approach for automated MA screening, its clinical ap… ▽ More

    Submitted 12 December, 2025; v1 submitted 12 November, 2025; originally announced November 2025.

    Comments: 9 pages, 6 figures, 8 tables, accepted by AAAI 2026

  22. arXiv:2510.25726  [pdf, ps, other

    cs.CL cs.AI

    The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution

    Authors: Junlong Li, Wenshuo Zhao, Jian Zhao, Weihao Zeng, Haoze Wu, Xiaochen Wang, Rui Ge, Yuxuan Cao, Yuzhen Huang, Wei Liu, Junteng Liu, Zhaochen Su, Yiyang Guo, Fan Zhou, Lueyang Zhang, Juan Michelini, Xingyao Wang, Xiang Yue, Shuyan Zhou, Graham Neubig, Junxian He

    Abstract: Real-world language agents must handle complex, multi-step workflows across diverse Apps. For instance, an agent may manage emails by coordinating with calendars and file systems, or monitor a production database to detect anomalies and generate reports following an operating manual. However, existing language agent benchmarks often focus on narrow domains or simplified tasks that lack the diversi… ▽ More

    Submitted 26 February, 2026; v1 submitted 29 October, 2025; originally announced October 2025.

    Comments: ICLR 2026, Website: https://toolathlon.xyz/

  23. arXiv:2510.21857  [pdf, ps, other

    cs.CV cs.AI

    Poisson Flow Consistency Training

    Authors: Anthony Zhang, Mahmut Gokmen, Dennis Hein, Rongjun Ge, Wenjun Xia, Ge Wang, Jin Chen

    Abstract: The Poisson Flow Consistency Model (PFCM) is a consistency-style model based on the robust Poisson Flow Generative Model++ (PFGM++) which has achieved success in unconditional image generation and CT image denoising. Yet the PFCM can only be trained in distillation which limits the potential of the PFCM in many data modalities. The objective of this research was to create a method to train the PFC… ▽ More

    Submitted 22 October, 2025; originally announced October 2025.

    Comments: 5 pages, 3 figures, 1 table

    MSC Class: 68T07 (Primary); 68T45 (Secondary)

  24. arXiv:2510.16322  [pdf, ps, other

    cs.LG stat.ML

    Memorizing Long-tail Data Can Help Generalization Through Composition

    Authors: Mo Zhou, Haoyang Ma, Rong Ge

    Abstract: Deep learning has led researchers to rethink the relationship between memorization and generalization. In many settings, memorization does not hurt generalization due to implicit regularization and may help by memorizing long-tailed examples. In this paper, we consider the synergy between memorization and simple composition -- the ability to make correct prediction on a combination of long-tailed… ▽ More

    Submitted 17 October, 2025; originally announced October 2025.

    Comments: 30 pages

  25. arXiv:2510.00355  [pdf, ps, other

    cs.AI cs.LG

    Hierarchical Reasoning Models: Perspectives and Misconceptions

    Authors: Renee Ge, Qianli Liao, Tomaso Poggio

    Abstract: Transformers have demonstrated remarkable performance in natural language processing and related domains, as they largely focus on sequential, autoregressive next-token prediction tasks. Yet, they struggle in logical reasoning, not necessarily because of a fundamental limitation of these models, but possibly due to the lack of exploration of more creative uses, such as latent space and recurrent r… ▽ More

    Submitted 7 October, 2025; v1 submitted 30 September, 2025; originally announced October 2025.

    Comments: Found errors in some results of v1. Removed them and changed conclusions

  26. arXiv:2508.17478  [pdf, ps, other

    cs.CV

    GraphMMP: A Graph Neural Network Model with Mutual Information and Global Fusion for Multimodal Medical Prognosis

    Authors: Xuhao Shan, Ruiquan Ge, Jikui Liu, Linglong Wu, Chi Zhang, Siqi Liu, Wenjian Qin, Wenwen Min, Ahmed Elazab, Changmiao Wang

    Abstract: In the field of multimodal medical data analysis, leveraging diverse types of data and understanding their hidden relationships continues to be a research focus. The main challenges lie in effectively modeling the complex interactions between heterogeneous data modalities with distinct characteristics while capturing both local and global dependencies across modalities. To address these challenges… ▽ More

    Submitted 24 August, 2025; originally announced August 2025.

  27. arXiv:2508.08566  [pdf

    cs.CV

    Think as Cardiac Sonographers: Marrying SAM with Left Ventricular Indicators Measurements According to Clinical Guidelines

    Authors: Tuo Liu, Qinghan Yang, Yu Zhang, Rongjun Ge, Yang Chen, Guangquan Zhou

    Abstract: Left ventricular (LV) indicator measurements following clinical echocardiog-raphy guidelines are important for diagnosing cardiovascular disease. Alt-hough existing algorithms have explored automated LV quantification, they can struggle to capture generic visual representations due to the normally small training datasets. Therefore, it is necessary to introduce vision founda-tional models (VFM) wi… ▽ More

    Submitted 11 August, 2025; originally announced August 2025.

  28. arXiv:2508.04205  [pdf, ps, other

    cs.CV

    Small Lesions-aware Bidirectional Multimodal Multiscale Fusion Network for Lung Disease Classification

    Authors: Jianxun Yu, Ruiquan Ge, Zhipeng Wang, Cheng Yang, Chenyu Lin, Xianjun Fu, Jikui Liu, Ahmed Elazab, Changmiao Wang

    Abstract: The diagnosis of medical diseases faces challenges such as the misdiagnosis of small lesions. Deep learning, particularly multimodal approaches, has shown great potential in the field of medical disease diagnosis. However, the differences in dimensionality between medical imaging and electronic health record data present challenges for effective alignment and fusion. To address these issues, we pr… ▽ More

    Submitted 6 August, 2025; originally announced August 2025.

  29. arXiv:2508.03357  [pdf, ps, other

    eess.IV cs.CV

    GL-LCM: Global-Local Latent Consistency Models for Fast High-Resolution Bone Suppression in Chest X-Ray Images

    Authors: Yifei Sun, Zhanghao Chen, Hao Zheng, Yuqing Lu, Lixin Duan, Fenglei Fan, Ahmed Elazab, Xiang Wan, Changmiao Wang, Ruiquan Ge

    Abstract: Chest X-Ray (CXR) imaging for pulmonary diagnosis raises significant challenges, primarily because bone structures can obscure critical details necessary for accurate diagnosis. Recent advances in deep learning, particularly with diffusion models, offer significant promise for effectively minimizing the visibility of bone structures in CXR images, thereby improving clarity and diagnostic accuracy.… ▽ More

    Submitted 5 August, 2025; originally announced August 2025.

    Comments: 11 pages, 3 figures, accepted by MICCAI 2025

  30. Extreme Cardiac MRI Analysis under Respiratory Motion: Results of the CMRxMotion Challenge

    Authors: Kang Wang, Chen Qin, Zhang Shi, Haoran Wang, Xiwen Zhang, Chen Chen, Cheng Ouyang, Chengliang Dai, Yuanhan Mo, Chenchen Dai, Xutong Kuang, Ruizhe Li, Xin Chen, Xiuzheng Yue, Song Tian, Alejandro Mora-Rubio, Kumaradevan Punithakumar, Shizhan Gong, Qi Dou, Sina Amirrajab, Yasmina Al Khalil, Cian M. Scannell, Lexiaozi Fan, Huili Yang, Xiaowu Sun , et al. (24 additional authors not shown)

    Abstract: Deep learning models have achieved state-of-the-art performance in automated Cardiac Magnetic Resonance (CMR) analysis. However, the efficacy of these models is highly dependent on the availability of high-quality, artifact-free images. In clinical practice, CMR acquisitions are frequently degraded by respiratory motion, yet the robustness of deep learning models against such artifacts remains an… ▽ More

    Submitted 25 July, 2025; originally announced July 2025.

  31. arXiv:2507.18576  [pdf, ps, other

    cs.AI cs.CL cs.CV

    SafeWork-R1: Coevolving Safety and Intelligence under the AI-45$^{\circ}$ Law

    Authors: Shanghai AI Lab, :, Yicheng Bao, Guanxu Chen, Mingkang Chen, Yunhao Chen, Chiyu Chen, Lingjie Chen, Sirui Chen, Xinquan Chen, Jie Cheng, Yu Cheng, Dengke Deng, Yizhuo Ding, Dan Ding, Xiaoshan Ding, Yi Ding, Zhichen Dong, Lingxiao Du, Yuyu Fan, Xinshun Feng, Yanwei Fu, Yuxuan Gao, Ruijun Ge, Tianle Gu , et al. (93 additional authors not shown)

    Abstract: We introduce SafeWork-R1, a cutting-edge multimodal reasoning model that demonstrates the coevolution of capabilities and safety. It is developed by our proposed SafeLadder framework, which incorporates large-scale, progressive, safety-oriented reinforcement learning post-training, supported by a suite of multi-principled verifiers. Unlike previous alignment methods such as RLHF that simply learn… ▽ More

    Submitted 7 August, 2025; v1 submitted 24 July, 2025; originally announced July 2025.

    Comments: 47 pages, 18 figures, authors are listed in alphabetical order by their last names; v3 modifies minor issues

  32. arXiv:2507.15230  [pdf, ps, other

    cs.DC cs.GR

    GALE: Leveraging Heterogeneous Systems for Efficient Unstructured Mesh Data Analysis

    Authors: Guoxi Liu, Thomas Randall, Rong Ge, Federico Iuricich

    Abstract: Unstructured meshes present challenges in scientific data analysis due to irregular distribution and complex connectivity. Computing and storing connectivity information is a major bottleneck for visualization algorithms, affecting both time and memory performance. Recent task-parallel data structures address this by precomputing connectivity information at runtime while the analysis algorithm exe… ▽ More

    Submitted 30 July, 2025; v1 submitted 21 July, 2025; originally announced July 2025.

    Comments: Accepted at IEEE VIS 2025

  33. arXiv:2506.23121  [pdf, ps, other

    eess.IV cs.AI cs.CV cs.LG

    CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ Segmentation

    Authors: Xinlei Yu, Changmiao Wang, Hui Jin, Ahmed Elazab, Gangyong Jia, Xiang Wan, Changqing Zou, Ruiquan Ge

    Abstract: Multi-organ medical segmentation is a crucial component of medical image processing, essential for doctors to make accurate diagnoses and develop effective treatment plans. Despite significant progress in this field, current multi-organ segmentation models often suffer from inaccurate details, dependence on geometric prompts and loss of spatial information. Addressing these challenges, we introduc… ▽ More

    Submitted 13 July, 2025; v1 submitted 29 June, 2025; originally announced June 2025.

    Comments: Accepted By ACMMM25

  34. arXiv:2506.21843  [pdf, ps, other

    cs.CV

    3D-Telepathy: Reconstructing 3D Objects from EEG Signals

    Authors: Yuxiang Ge, Jionghao Cheng, Ruiquan Ge, Zhaojie Fang, Gangyong Jia, Xiang Wan, Nannan Li, Ahmed Elazab, Changmiao Wang

    Abstract: Reconstructing 3D visual stimuli from Electroencephalography (EEG) data holds significant potential for applications in Brain-Computer Interfaces (BCIs) and aiding individuals with communication disorders. Traditionally, efforts have focused on converting brain activity into 2D images, neglecting the translation of EEG data into 3D objects. This limitation is noteworthy, as the human brain inheren… ▽ More

    Submitted 26 June, 2025; originally announced June 2025.

  35. arXiv:2506.19830  [pdf, ps, other

    cs.LG cs.CL

    Scaling Speculative Decoding with Lookahead Reasoning

    Authors: Yichao Fu, Rui Ge, Zelei Shao, Zhijie Deng, Hao Zhang

    Abstract: Reasoning models excel by generating long chain-of-thoughts, but decoding the resulting thousands of tokens is slow. Token-level speculative decoding (SD) helps, but its benefit is capped, because the chance that an entire $γ$-token guess is correct falls exponentially as $γ$ grows. This means allocating more compute for longer token drafts faces an algorithmic ceiling -- making the speedup modest… ▽ More

    Submitted 24 June, 2025; originally announced June 2025.

  36. arXiv:2503.12066  [pdf

    cs.LG q-bio.NC q-bio.QM

    Dataset Properties Shape the Success of Neuroimaging-Based Patient Stratification: A Benchmarking Analysis Across Clustering Algorithms

    Authors: Yuetong Yu, Ruiyang Ge, Ilker Hacihaliloglu, Alexander Rauscher, Roger Tam, Sophia Frangou

    Abstract: Background: Data driven stratification of patients into biologically informed subtypes holds promise for precision neuropsychiatry, yet neuroimaging-based clustering methods often fail to generalize across cohorts. While algorithmic innovations have focused on model complexity, the role of underlying dataset characteristics remains underexplored. We hypothesized that cluster separation, size imbal… ▽ More

    Submitted 10 June, 2025; v1 submitted 15 March, 2025; originally announced March 2025.

  37. arXiv:2503.03686  [pdf, other

    cs.CL cs.MA

    MAS-GPT: Training LLMs to Build LLM-based Multi-Agent Systems

    Authors: Rui Ye, Shuo Tang, Rui Ge, Yaxin Du, Zhenfei Yin, Siheng Chen, Jing Shao

    Abstract: LLM-based multi-agent systems (MAS) have shown significant potential in tackling diverse tasks. However, to design effective MAS, existing approaches heavily rely on manual configurations or multiple calls of advanced LLMs, resulting in inadaptability and high inference costs. In this paper, we simplify the process of building an MAS by reframing it as a generative language task, where the input i… ▽ More

    Submitted 5 March, 2025; originally announced March 2025.

    Comments: 26 pages, 7 figures

  38. arXiv:2502.05282  [pdf, other

    cs.CV cs.AI

    Homeomorphism Prior for False Positive and Negative Problem in Medical Image Dense Contrastive Representation Learning

    Authors: Yuting He, Boyu Wang, Rongjun Ge, Yang Chen, Guanyu Yang, Shuo Li

    Abstract: Dense contrastive representation learning (DCRL) has greatly improved the learning efficiency for image-dense prediction tasks, showing its great potential to reduce the large costs of medical image collection and dense annotation. However, the properties of medical images make unreliable correspondence discovery, bringing an open problem of large-scale false positive and negative (FP&N) pairs in… ▽ More

    Submitted 7 February, 2025; originally announced February 2025.

    Comments: Accepted by T-PAMI 2025

  39. arXiv:2502.03781  [pdf, ps, other

    cs.CV eess.IV

    Gaze-Assisted Human-Centric Domain Adaptation for Cardiac Ultrasound Image Segmentation

    Authors: Ruiyi Li, Yuting He, Rongjun Ge, Chong Wang, Daoqiang Zhang, Yang Chen, Shuo Li

    Abstract: Domain adaptation (DA) for cardiac ultrasound image segmentation is clinically significant and valuable. However, previous domain adaptation methods are prone to be affected by the incomplete pseudo-label and low-quality target to source images. Human-centric domain adaptation has great advantages of human cognitive guidance to help model adapt to target domain and reduce reliance on labels. Docto… ▽ More

    Submitted 6 February, 2025; originally announced February 2025.

  40. arXiv:2502.00695  [pdf, other

    cs.CV cs.AI

    TMI-CLNet: Triple-Modal Interaction Network for Chronic Liver Disease Prognosis From Imaging, Clinical, and Radiomic Data Fusion

    Authors: Linglong Wu, Xuhao Shan, Ruiquan Ge, Ruoyu Liang, Chi Zhang, Yonghong Li, Ahmed Elazab, Huoling Luo, Yunbi Liu, Changmiao Wang

    Abstract: Chronic liver disease represents a significant health challenge worldwide and accurate prognostic evaluations are essential for personalized treatment plans. Recent evidence suggests that integrating multimodal data, such as computed tomography imaging, radiomic features, and clinical information, can provide more comprehensive prognostic information. However, modalities have an inherent heterogen… ▽ More

    Submitted 2 February, 2025; originally announced February 2025.

    Comments: 6 pages, 3 figures, accepted by IEEE ISBI 2025

  41. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

    Authors: DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei Feng, Chengda Lu , et al. (175 additional authors not shown)

    Abstract: General reasoning represents a long-standing and formidable challenge in artificial intelligence. Recent breakthroughs, exemplified by large language models (LLMs) and chain-of-thought prompting, have achieved considerable success on foundational reasoning tasks. However, this success is heavily contingent upon extensive human-annotated demonstrations, and models' capabilities are still insufficie… ▽ More

    Submitted 3 January, 2026; v1 submitted 22 January, 2025; originally announced January 2025.

    Journal ref: Nature volume 645, pages 633-638 (2025)

  42. arXiv:2501.11276  [pdf, other

    eess.IV cs.CV

    ITCFN: Incomplete Triple-Modal Co-Attention Fusion Network for Mild Cognitive Impairment Conversion Prediction

    Authors: Xiangyang Hu, Xiangyu Shen, Yifei Sun, Xuhao Shan, Wenwen Min, Liyilei Su, Xiaomao Fan, Ahmed Elazab, Ruiquan Ge, Changmiao Wang, Xiaopeng Fan

    Abstract: Alzheimer's disease (AD) is a common neurodegenerative disease among the elderly. Early prediction and timely intervention of its prodromal stage, mild cognitive impairment (MCI), can decrease the risk of advancing to AD. Combining information from various modalities can significantly improve predictive accuracy. However, challenges such as missing data and heterogeneity across modalities complica… ▽ More

    Submitted 20 January, 2025; originally announced January 2025.

    Comments: 5 pages, 1 figure, accepted by IEEE ISBI 2025

  43. arXiv:2412.19437  [pdf, other

    cs.CL cs.AI

    DeepSeek-V3 Technical Report

    Authors: DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Han Bao , et al. (175 additional authors not shown)

    Abstract: We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2. Furthermore, DeepSeek-V3 pioneers an auxiliary-loss-free strategy for loa… ▽ More

    Submitted 18 February, 2025; v1 submitted 26 December, 2024; originally announced December 2024.

  44. BS-LDM: Effective Bone Suppression in High-Resolution Chest X-Ray Images with Conditional Latent Diffusion Models

    Authors: Yifei Sun, Zhanghao Chen, Hao Zheng, Wenming Deng, Jin Liu, Wenwen Min, Ahmed Elazab, Xiang Wan, Changmiao Wang, Ruiquan Ge

    Abstract: Lung diseases represent a significant global health challenge, with Chest X-Ray (CXR) being a key diagnostic tool due to its accessibility and affordability. Nonetheless, the detection of pulmonary lesions is often hindered by overlapping bone structures in CXR images, leading to potential misdiagnoses. To address this issue, we develop an end-to-end framework called BS-LDM, designed to effectivel… ▽ More

    Submitted 6 July, 2025; v1 submitted 20 December, 2024; originally announced December 2024.

    Comments: 12 pages, 8 figures, accepted by IEEE Journal of Biomedical and Health Informatics (JBHI) on July 4, 2025

  45. arXiv:2411.04656  [pdf, other

    cs.CV

    ICH-SCNet: Intracerebral Hemorrhage Segmentation and Prognosis Classification Network Using CLIP-guided SAM mechanism

    Authors: Xinlei Yu, Ahmed Elazab, Ruiquan Ge, Hui Jin, Xinchen Jiang, Gangyong Jia, Qing Wu, Qinglei Shi, Changmiao Wang

    Abstract: Intracerebral hemorrhage (ICH) is the most fatal subtype of stroke and is characterized by a high incidence of disability. Accurate segmentation of the ICH region and prognosis prediction are critically important for developing and refining treatment plans for post-ICH patients. However, existing approaches address these two tasks independently and predominantly focus on imaging data alone, thereb… ▽ More

    Submitted 7 November, 2024; originally announced November 2024.

    Comments: 6 pages, 2 figures, 3 tables, published to BIBM 2024

  46. arXiv:2410.01591  [pdf, ps, other

    eess.IV cs.AI cs.CV

    Imaging foundation model for universal enhancement of non-ideal measurement CT

    Authors: Rongjun Ge, Yuxin Liu, Zhan Wu, Shangwen Yang, Yuan Gao, Chenyu You, Ge Wang, Shuo Li, Yuting He, Yang Chen

    Abstract: Non-ideal measurement computed tomography (NICT) employs suboptimal imaging protocols to expand CT applications. However, the resulting trade-offs degrade image quality, limiting clinical acceptability. Although deep learning methods have been used to enhance NICT images, their reliance on large training datasets and limited generalizability across diverse settings hinder practical use. We propose… ▽ More

    Submitted 22 March, 2026; v1 submitted 2 October, 2024; originally announced October 2024.

    Comments: This paper has been accepted by Nature Communications

  47. arXiv:2409.07136  [pdf, other

    cs.CL cs.AI cs.MA

    Leveraging Unstructured Text Data for Federated Instruction Tuning of Large Language Models

    Authors: Rui Ye, Rui Ge, Yuchi Fengting, Jingyi Chai, Yanfeng Wang, Siheng Chen

    Abstract: Federated instruction tuning enables multiple clients to collaboratively fine-tune a shared large language model (LLM) that can follow humans' instructions without directly sharing raw data. However, existing literature impractically requires that all the clients readily hold instruction-tuning data (i.e., structured instruction-response pairs), which necessitates massive human annotations since c… ▽ More

    Submitted 11 September, 2024; originally announced September 2024.

    Comments: 11 pages, work in progress

  48. arXiv:2409.00726  [pdf, other

    cs.CV cs.AI

    LPUWF-LDM: Enhanced Latent Diffusion Model for Precise Late-phase UWF-FA Generation on Limited Dataset

    Authors: Zhaojie Fang, Xiao Yu, Guanyu Zhou, Ke Zhuang, Yifei Chen, Ruiquan Ge, Changmiao Wang, Gangyong Jia, Qing Wu, Juan Ye, Maimaiti Nuliqiman, Peifang Xu, Ahmed Elazab

    Abstract: Ultra-Wide-Field Fluorescein Angiography (UWF-FA) enables precise identification of ocular diseases using sodium fluorescein, which can be potentially harmful. Existing research has developed methods to generate UWF-FA from Ultra-Wide-Field Scanning Laser Ophthalmoscopy (UWF-SLO) to reduce the adverse reactions associated with injections. However, these methods have been less effective in producin… ▽ More

    Submitted 1 September, 2024; originally announced September 2024.

    Comments: 13 pages, 7 figures

  49. arXiv:2408.05705  [pdf, other

    eess.IV cs.AI cs.CV

    TC-KANRecon: High-Quality and Accelerated MRI Reconstruction via Adaptive KAN Mechanisms and Intelligent Feature Scaling

    Authors: Ruiquan Ge, Xiao Yu, Yifei Chen, Guanyu Zhou, Fan Jia, Shenghao Zhu, Junhao Jia, Chenyan Zhang, Yifei Sun, Dong Zeng, Changmiao Wang, Qiegen Liu, Shanzhou Niu

    Abstract: Magnetic Resonance Imaging (MRI) has become essential in clinical diagnosis due to its high resolution and multiple contrast mechanisms. However, the relatively long acquisition time limits its broader application. To address this issue, this study presents an innovative conditional guided diffusion model, named as TC-KANRecon, which incorporates the Multi-Free U-KAN (MF-UKAN) module and a dynamic… ▽ More

    Submitted 6 January, 2025; v1 submitted 11 August, 2024; originally announced August 2024.

    Comments: 11 pages, 3 figures

  50. arXiv:2406.15968  [pdf, other

    cs.CL cs.LG

    ReCaLL: Membership Inference via Relative Conditional Log-Likelihoods

    Authors: Roy Xie, Junlin Wang, Ruomin Huang, Minxing Zhang, Rong Ge, Jian Pei, Neil Zhenqiang Gong, Bhuwan Dhingra

    Abstract: The rapid scaling of large language models (LLMs) has raised concerns about the transparency and fair use of the data used in their pretraining. Detecting such content is challenging due to the scale of the data and limited exposure of each instance during training. We propose ReCaLL (Relative Conditional Log-Likelihood), a novel membership inference attack (MIA) to detect LLMs' pretraining data b… ▽ More

    Submitted 23 May, 2025; v1 submitted 22 June, 2024; originally announced June 2024.

    Comments: Accepted to EMNLP 2024 Main Conference