Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 282 results for author: Fu, R

.
  1. arXiv:2608.14249  [pdf, ps, other

    cs.SD

    AT-ADD: All-Type Audio Deepfake Detection Challenge Summary

    Authors: Yuankun Xie, Haonan Cheng, Jiayi Zhou, Xiaoxuan Guo, Tao Wang, Changhao Zhang, Jian Liu, Weiqiang Wang, Ruibo Fu, Xiaopeng Wang, Hengyan Huang, Xiaoying Huang, Long Ye, Guangtao Zhai

    Abstract: This paper summarizes the ACM Multimedia 2026 AT-ADD Grand Challenge on all-type audio deepfake detection. AT-ADD contains two tracks: robust speech deepfake detection under realistic acoustic and channel variations, and type-agnostic detection over speech, environmental sound, singing voice, and music. We describe the challenge tasks, dataset and evaluation-set design, official leaderboard result… ▽ More

    Submitted 14 August, 2026; originally announced August 2026.

    Comments: Accepted to ACM MM 2026

  2. arXiv:2608.12751  [pdf, ps, other

    cs.AR cs.AI

    SynAct: A Reasoning-Acting Large Language Model Agent for Adaptive Synthesis Optimization

    Authors: Fangzhou Liu, Peiyi Han, Jiawei Liu, Yuan Pu, Zhuolun He, Rongliang Fu, Tsung-Yi Ho, Bei Yu

    Abstract: Logic synthesis transforms RTL designs into gate-level netlists, where PPA results are highly sensitive to the choice of optimization commands, making synthesis tuning both high-dimensional and expensive. Previous approaches fall into two categories: automated methods, which perform black-box search over fixed action spaces with limited decision-level interpretability, and LLM-based methods, which… ▽ More

    Submitted 13 August, 2026; v1 submitted 12 August, 2026; originally announced August 2026.

    Comments: 12 pages, 8 figures

  3. arXiv:2608.09391  [pdf, ps, other

    cs.CV

    CoInS-Net: A Continuous Position-Aware Network for Joint Medical Image Interpolation and Segmentation

    Authors: Yujia Sun, Ningfeng Que, Peiting Shi, Rongrong Fu, Yingying Yang, Xinhang Li, Yin Dai

    Abstract: Accurate medical image interpolation and anatomical structure segmentation are fundamental for computer-aided diagnosis and treatment planning. Anisotropic medical volumes with sparse through-plane sampling often suffer from structural discontinuity and boundary blur, hindering reliable clinical image analysis. Most existing methods implement interpolation and segmentation independently, which int… ▽ More

    Submitted 10 August, 2026; originally announced August 2026.

  4. arXiv:2608.08553  [pdf, ps, other

    cs.CV cs.LG cs.MM

    MotionCraft: Latent World Modeling with Sparse Attention for Visual Upscaling

    Authors: Rong Fu, Chunlei Meng, Yangchen Zeng, Xiaowen Ma, Yongtai Liu, Wangyu Wu, Shuo Yin, Zijian Zhang, Sicheng Li, Yingrui Ji, Chenhao Wang, Simon Fong

    Abstract: Video super-resolution (VSR) aims to recover high-fidelity high-resolution videos from low-resolution inputs and is central to applications ranging from mobile capture to streaming and archival restoration. Existing approaches trade off among local-detail fidelity, long-range spatio-temporal modeling, perceptual realism, and efficiency: convolutional alignment techniques preserve local structure b… ▽ More

    Submitted 9 August, 2026; originally announced August 2026.

    Comments: 14 pages, 6 figures

  5. arXiv:2608.03738  [pdf, ps, other

    cs.AI

    AgenticECO: An Agentic Framework for ECO on 3D Integrated Circuits

    Authors: Shuo Ren, Yaohui Han, Libo Shen, Zhiqiang Jia, Rongliang Fu, Bei Yu, Tsung-Yi Ho

    Abstract: As Moore's law slows, the industry is turning to three-dimensional integration; yet in merged 3D-IC flows, routed designs expose bond-level defects with no 2D analogue, and post-route engineering change orders (ECO) remain manual, expertise-bound work. Worse, the standard edit-then-fully-reroute practice entangles a repair with router churn, so a signoff number cannot be attributed to the edit tha… ▽ More

    Submitted 4 August, 2026; originally announced August 2026.

    Comments: 20 pages, 12 figures

  6. Odin: Primitive-Level Synchronization for Distributed Point-Based Neural Rendering

    Authors: Zhenxiang Ma, Zeyu He, Yuanzhen Zhou, Zhenyu Yang, Yuchang Zhang, Miao Tao, Rong Fu, Jidong Zhai, Hengjie Li

    Abstract: Point-based neural rendering (PBNR) represents 3D scenes as explicit, trainable primitives and underpins high-quality reconstruction and emerging embodied AI and world-model pipelines. Unlike layer-structured neural networks, PBNR has primitive-indexed dependencies: each view reads and updates only a sparse, view-dependent subset of mutable scene state. As large scenes require distributed training… ▽ More

    Submitted 22 July, 2026; originally announced July 2026.

    Comments: Accepted at HPDC 26 (Poster), further submission in progress

  7. arXiv:2607.17244  [pdf, ps, other

    cs.LG

    DynImmune-BERT: Dynamic Immune Repertoire Modeling with Neural ODE Driven Continuous Transformers

    Authors: Rong Fu, Yongtai Liu, Xiaowen Ma, Haoyu Zhao, Shuo Yin, Yiqing Lyu, Long Zhang, Wangyu Wu

    Abstract: Longitudinal T cell receptor repertoires contain signals of clonal expansion, contraction, disappearance, and reappearance after immune perturbation. Static repertoire language models usually summarize a sample as a bag of sequences, so the sampling interval, sequencing depth, and clone presence pattern are only weakly represented. This paper presents DynImmune-BERT, a continuous time repertoire m… ▽ More

    Submitted 3 August, 2026; v1 submitted 19 July, 2026; originally announced July 2026.

    Comments: 13 pages, 6 figures

  8. arXiv:2607.17029  [pdf, ps, other

    cs.DS

    Degeneracy-Guided List Compression for Greedy Graph Coloring

    Authors: Rong Fu, Yongtai Liu, Xiaowen Ma, Wangyu Wu, Long Zhang, Hongbo Zhang, Yangchen Zeng, Hoi Leong Lee, Hao Zhang

    Abstract: We study degeneracy guided list compression for greedy graph coloring when graph structure is available before colors are sampled. Our exposure calibrated ordering framework assigns each vertex an independent uniform list according to its backward neighborhood in a color independent order. Its certified instantiation, Profiled Structure Aware Asymmetric Palette Sparsification, or P-SAPST, reverses… ▽ More

    Submitted 4 August, 2026; v1 submitted 18 July, 2026; originally announced July 2026.

    Comments: 26 pages, 16 figures

  9. arXiv:2607.12788  [pdf, ps, other

    cs.AR

    CLIP-3D: Closed-Loop Evaluation of Performance and Physical Constraints for 3D ICs

    Authors: Shuo Ren, Libo Shen, Yaohui Han, Leilei Jin, Chenghan Wang, Zhen Zhuang, Rongliang Fu, Bei Yu, Tsung-Yi Ho

    Abstract: 3D integration packs more power into a smaller footprint, so a candidate design's actual throughput depends on its layout: which macro sits on which tier, where the hot spot lands, and how cache geometry maps to access cycles. Architectural simulators like gem5 report IPC under idealized timing. They do not produce the per-block power map, the cache cycle counts, or the 3D layout that decide the r… ▽ More

    Submitted 14 July, 2026; originally announced July 2026.

    Comments: 12 pages, 4 figures

  10. arXiv:2607.09742  [pdf, ps, other

    cs.AR

    Chiplet3D: Pin- and Thermal-Aware 3D Chiplet Floorplanning via Convolution-Embedded MILP

    Authors: Shuo Ren, Libo Shen, Yaohui Han, Rongliang Fu, Junying Huang, Bei Yu, Tsung-Yi Ho

    Abstract: As traditional Moore's Law scaling slows down, 3D-ICs stack multiple active dies vertically to sustain performance scaling. However, this vertical stacking traps heat inside, making temperature a design concern. Although we can fix thermal issues at different design steps, floorplanning is the earliest and most cost-effective stage to solve it. Previous methods handle this by assuming wires connec… ▽ More

    Submitted 3 July, 2026; originally announced July 2026.

    Comments: 8 pages, 6 figures

  11. arXiv:2607.08161  [pdf, ps, other

    cs.CL cs.CV

    SQuaD-SQL: Efficient Text-to-SQL with Small Language Models via LLM-Guided Knowledge Distillation

    Authors: Wangyu Wu, Xiaojian Lin, Rong Fu, Zaiyang Yu, Xuhang Chen, Wenjun Yu, Zhenhong Chen

    Abstract: Text-to-SQL is a fundamental task in natural language processing that enables users to interact with structured databases using natural language. While large language models (LLMs) have demonstrated remarkable performance on this task, their substantial computational requirements hinder deployment in resource-constrained settings. In this paper, we introduce SQuaD-SQL (Small-Qualified and Distille… ▽ More

    Submitted 9 July, 2026; originally announced July 2026.

    Comments: Accepted at IEEE SMC 2026

  12. arXiv:2607.05390  [pdf, ps, other

    cs.RO cs.CV

    Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models

    Authors: Hongyu Li, Wanjia Fu, Xiaoyan Cong, Zekun Li, Binghao Huang, Hanxiao Jiang, Xintong He, Yiqing Liang, Rao Fu, Tao Lu, Srinath Sridhar, Kevin A. Smith, George Konidaris, Yunzhu Li

    Abstract: Predicting object dynamics (i.e., world modeling) is a fundamental challenge for robotic manipulation, and modeling deformable objects presents a particularly difficult case due to their high-dimensional state spaces and complex material properties. While current world models approach this through two distinct paradigms: learning the dynamics over the 2D pixel space or more explicit 3D geometric s… ▽ More

    Submitted 6 July, 2026; originally announced July 2026.

    Comments: Accepted by ECCV 2026

  13. arXiv:2607.04758  [pdf, ps, other

    cs.AI

    AgenticPD: A Stage-Aware Agentic Framework for Physical Design QoR Optimization

    Authors: Shuo Ren, Zijin Cheng, Yaohui Han, Libo Shen, Leilei Jin, Wanting Tian, Rongliang Fu, Chao Wang, Bei Yu, Tsung-Yi Ho

    Abstract: Physical design quality-of-results~(QoR) optimization is hard and expensive. Choices made at one stage can help or hurt later stages. Each evaluation requires a costly EDA run through the full flow. While existing methods still treat optimization as flat parameter tuning or a LLM-based script generation task, we present AgenticPD, a stage-aware agentic framework for physical design QoR optimizatio… ▽ More

    Submitted 7 July, 2026; v1 submitted 6 July, 2026; originally announced July 2026.

    Comments: 7 pages, 6 figures

  14. arXiv:2607.02141  [pdf, ps, other

    cs.AI

    A$^{2}$utoLPBench: An Auto-Generated, Agent-Friendly LP Benchmark via Inverse-KKT Construction

    Authors: Shuo Ren, Yaohui Han, Yifan Shi, Libo Shen, Haodong Lu, Dongfang Wu, Rongliang Fu, Bei Yu, Tsung-Yi Ho

    Abstract: Most LP-from-text benchmarks are static datasets of word problems written and labeled by hand. Once such a dataset is released, its size is fixed, its difficulty is fixed, and every problem can leak into the training data of future LLMs. We present \textbf{A$^{2}$utoLPBench}, a benchmark for testing LLM-driven agents on linear programming problems written in plain text. We first pick a feasible po… ▽ More

    Submitted 2 July, 2026; originally announced July 2026.

    Comments: 25 pages and 4 figures

  15. arXiv:2606.28226  [pdf, ps, other

    cs.CV cs.AI

    Exposure Bias Can Alleviate Itself via Directional and Frequency Rectification in Flow Matching

    Authors: Guanbo Huang, Jingjia Mao, Fanding Huang, Fengkai Liu, Xiangyang Luo, Yaoyuan Liang, Jiasheng Lu, Xiaoe Wang, Pei Liu, Ruiliu Fu, Ruqi Huang, Shao-Lun Huang

    Abstract: Flow Matching (FM) has achieved remarkable generative performance, yet it suffers from exposure bias due to discrepancies between training and inference. Existing mitigation strategies typically rely on static constraints or external heuristics. In this work, we propose that exposure bias itself inherently contains dynamic signals that can guide its own rectification. To leverage this, we introduc… ▽ More

    Submitted 26 June, 2026; originally announced June 2026.

    Comments: arXiv admin note: text overlap with arXiv:2512.04904

  16. arXiv:2606.25472  [pdf

    physics.flu-dyn physics.med-ph

    From Propulsion to Suction: Unraveling Thrust Reversal in Propellers at Intermediate Reynolds Numbers

    Authors: Rong Fu, Siyu Li, Yang Ding

    Abstract: This study investigates propeller hydrodynamics at intermediate Reynolds numbers (Re), crucial for small-scale robotic systems but still uncharted. Experiments on a propeller-driven underwater vehicle and numerical simulations reveal thrust reversal--a phenomenon where clockwise propeller rotation leads to backward motion--in the approximate range 1.3 < Re < 150 under specific conditions. Notably,… ▽ More

    Submitted 24 June, 2026; originally announced June 2026.

    Journal ref: Proc. Natl. Acad. Sci. U.S.A. 122 (40) e2504153122 (2025)

  17. arXiv:2606.14066   

    cs.SE

    FastContext: Training Efficient Repository Explorer for Coding Agents

    Authors: Shaoqiu Zhang, Maoquan Wang, Yuling Shi, Yuhang Wang, Xiaodong Gu, Yongqiang Yao, Tori Gong, Sheng Chen, Rao Fu, Anisha Agarwal, Spandan Grag, Gabriel Ryan, Colin Merkel, Yufan Huang, Shengyu Fu

    Abstract: Large Language Model (LLM) coding agents have achieved strong results on software engineering tasks, yet repository exploration remains a major bottleneck: locating relevant code consumes substantial token budget and pollutes the agent's context with irrelevant snippets. In most agents, the same model explores the repository and solves the task, leaving exploratory reads and searches in the solver… ▽ More

    Submitted 29 June, 2026; v1 submitted 11 June, 2026; originally announced June 2026.

    Comments: The current article involves some product IP issues and needs to be withdrawn and re-approved

  18. arXiv:2606.08436  [pdf, ps, other

    cs.CV

    CACR:Reinforcing Temporal Answer Grounding in Instructional Video via Candidate-Aware Causal Reasoning

    Authors: Muge Qi, Rong Fu, Pengbin Feng, Xianda Li, Yu Cai, Yifu Guo, Shizhe Zhang, Simon James Fong, Lei Ma, Bin Li

    Abstract: The task of temporal answer grounding in instructional video (TAGV), which aims to locate precise video segments that respond to natural language queries, is increasingly important for direct video answer retrieval. This task remains challenging due to the need to comprehend semantically complex questions and to address the significant length mismatch between untrimmed videos and short target mome… ▽ More

    Submitted 11 June, 2026; v1 submitted 6 June, 2026; originally announced June 2026.

  19. arXiv:2606.05405  [pdf, ps, other

    cs.AI cs.CL cs.LG

    Agents' Last Exam

    Authors: Yiyou Sun, Xinyang Han, Weichen Zhang, Yuanbo Pang, Tianyu Wang, Yuhan Cao, Yixiao Huang, Chris Duroiu, Haoyun Zhang, Jeffrey Lin, Weishu Zhang, Tyler Zeng, Ying Yan, Bo Liu, Hanson Wen, Mingyang Xu, Xiaoyuan Liu, Zimeng Chen, Weiyan Shi, Amanda Dsouza, Vincent Sunn Chen, Patrick Bryant, Carl Boettiger, Yamini Rangan, Bradley Rothenberg , et al. (285 additional authors not shown)

    Abstract: Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional domains. We argue that this gap is largely an evaluation problem: widely used benchmarks lack sustained performance measurement on real and economically valuable workflows. This paper introduces Agents' Last Exam (ALE), a… ▽ More

    Submitted 11 June, 2026; v1 submitted 3 June, 2026; originally announced June 2026.

    Comments: Project website: https://agents-last-exam.org Code: https://github.com/rdi-berkeley/agents-last-exam

  20. arXiv:2606.00704  [pdf, ps, other

    cs.CV

    VICR: Visual In-Context Restoration for Real-World Image Super-Resolution

    Authors: Qichang Zhang, Hailong Wang, Baiang Li, Linhao Wang, Rong Fu, Erkang Cheng, Simon James Fong

    Abstract: Real-world image super-resolution (Real-ISR) requires balancing structural fidelity to degraded observations with realistic detail synthesis. However, existing generative Real-ISR methods often rely on entangled conditioning mechanisms, leading to structural drift or semantically inconsistent details. To address this issue, we propose Visual In-Context Restoration (VICR), a Diffusion Transformer (… ▽ More

    Submitted 30 May, 2026; originally announced June 2026.

    Comments: 28 pages, 11 figures, 9 tables

  21. arXiv:2605.29558  [pdf, ps, other

    cs.CV

    TAE: Target-aware enhancer for nighttime UAV tracking

    Authors: Yanyan Chen, Ruigang Fu, Yu Song, Ping Zhong

    Abstract: Severe image degradation under low-light nighttime conditions constitutes a core bottleneck preventing all-day applications for UAV-based single object tracking. Existing image enhancement methods often struggle to distinguish between target and background regions, which can easily lead to amplified background noise or compromise target features. To overcome this limitation, we propose TAE, a targ… ▽ More

    Submitted 28 May, 2026; originally announced May 2026.

    Comments: Accepted at ICIP 2026. Dataset is avaliable at: https://github.com/Fu0511/DarkSOT-Dataset

  22. arXiv:2605.29232  [pdf, ps, other

    cs.IR

    On the Practice of Scaling Search Conversion Rate Prediction

    Authors: James Pak, Jyun-Yu Jiang, Fan Zhang, Sen Wang, Taekmin Kim, Henry Tsai, Vijay Rajaram, Juexin Lin, Mohitdeep Singh, Alessandro Magnani, Johnny Chen, Qian Zhao, Rao Fu, Zhirong Liang, Jordan Gilliland, Winter Jiao

    Abstract: Scaling a Search Conversion Rate (CVR) prediction model, especially in high-traffic environments, presents a challenge: superior model quality needs to be balanced with strict constraints on training cost and serving latency. This paper details an effective approach for scaling modern search CVR prediction models. We begin with an empirical study to understand the scaling performance of search CVR… ▽ More

    Submitted 27 May, 2026; originally announced May 2026.

  23. arXiv:2605.18797  [pdf, ps, other

    cs.LG cs.AI

    Simply Stabilizing the Loop via Fully Looped Transformer

    Authors: Rao Fu, Zixuan Yang, Jiankun Zhang, Jing Ma, Hechang Chen, Yu Li, Yi Chang

    Abstract: Scaling model performance typically requires increasing model size. Looped Transformer offers a compelling alternative by iteratively reusing the same Transformer blocks, trading additional computation for improved performance without increasing parameter count or context length. Because the number of loop iterations can be adjusted at inference, it also provides a natural mechanism for balancing… ▽ More

    Submitted 25 May, 2026; v1 submitted 11 May, 2026; originally announced May 2026.

  24. arXiv:2605.17949  [pdf, ps, other

    cs.CV

    SkyNative: A Native Multimodal Framework for Remote Sensing Visual Evidence Reasoning

    Authors: Xiao Yang, Ronghao Fu, Zhiwen Lin, Zhuoran Duan, Jiashun Zhu, Jiasen Hu, Lang Sun, Weipeng Zhang, Jiaqi Liu, Xu Na, Haoran Liu, Weijie Zhang, Bo Yang

    Abstract: Remote sensing vision-language models commonly rely on pretrained visual encoders to convert images into semantic features before language-model reasoning. While effective for scene-level understanding, this pipeline may prematurely compress local visual evidence, making fine-grained spatial reasoning vulnerable to language priors, especially in ultra-high-resolution remote sensing imagery. We pre… ▽ More

    Submitted 18 May, 2026; originally announced May 2026.

  25. arXiv:2605.14475  [pdf, ps, other

    cs.CV

    GeoVista: Visually Grounded Active Perception for Vision-Language Understanding of Ultra-High-Resolution Remote Sensing Images

    Authors: Jiashun Zhu, Ronghao Fu, Jiasen Hu, Jing Huang, Nachuan Xing, Bo Yang

    Abstract: Interpreting ultra-high-resolution (UHR) remote sensing images requires models to search for sparse and tiny visual evidence across large-scale scenes. Existing remote sensing vision-language models can inspect local regions with zooming and cropping tools, but most exploration strategies follow either a one-shot focus or a single sequential trajectory. Such single-path exploration can lose global… ▽ More

    Submitted 21 June, 2026; v1 submitted 14 May, 2026; originally announced May 2026.

  26. arXiv:2605.14366  [pdf, ps, other

    cs.CL cs.LG

    Reinforcement Learning with Semantic Rewards Enables Low-Resource Language Expansion without Alignment Tax

    Authors: Zeli Su, Ziyin Zhang, Zhou Liu, Xuexian Song, Zhankai Xu, Longfei Zheng, Xiaolu Zhang, Rong Fu, Guixian Xu, Wentao Zhang

    Abstract: Extending large language models (LLMs) to low-resource languages often incurs an "alignment tax": improvements in the target language come at the cost of catastrophic forgetting in general capabilities. We argue that this trade-off arises from the rigidity of supervised fine-tuning (SFT), which enforces token-level surface imitation on narrow and biased data distributions. To address this limitati… ▽ More

    Submitted 14 May, 2026; originally announced May 2026.

    Comments: ACL 2026 Findings

  27. arXiv:2605.10544  [pdf, ps, other

    cs.CL

    Where Does Long-Context Supervision Actually Go? Effective-Context Exposure Balancing

    Authors: Jinchang Zhu, Jindong Li, Chengyu Zou, Rong Fu, Chao Wang, Haowei He, Menglin Yang

    Abstract: Long-context adaptation is often viewed as window scaling, but this misses a token-level supervision mismatch: in packed training with document masking, each target token's effective context remains short. We introduce EXACT, a supervision-allocation objective that assigns extra weight to long effective-context targets by inverse frequency within the long tail. Across seven Qwen/LLaMA CPT configur… ▽ More

    Submitted 11 May, 2026; originally announced May 2026.

  28. arXiv:2605.10504  [pdf, ps, other

    cs.CL

    Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining

    Authors: Jinchang Zhu, Jindong Li, Yuwen Hao, Chengyu Zou, Rong Fu, Menglin Yang

    Abstract: A causal-decoder block is hierarchical: lower layers build the residual basis that upper layers attend over. We identify a failure mode in GPT pretraining: upper layers commit to sharp attention patterns before lower-layer features stabilize. We call this premature upper-layer attention specialization. Temporarily slowing only upper-layer Q/K projections during early training improves final perple… ▽ More

    Submitted 11 May, 2026; originally announced May 2026.

  29. arXiv:2605.01836  [pdf, ps, other

    cs.AR

    PipeRTL: Timing-Aware Pipeline Optimization at IR-Level for RTL Generation

    Authors: Shuo Yin, Fangzhou Liu, Lancheng Zou, Rongliang Fu, Wenqian Zhao, Chen Bai, Tsung-Yi Ho, Yuan Xie, Bei Yu

    Abstract: Modern hardware compilers increasingly rely on rich intermediate representations (IRs) to preserve optimization-relevant semantics before generating RTL code. However, one important optimization is still largely deferred to backend tools: pipeline optimization. In common RTL flows, registers are inserted by frontend heuristics or hardware designers and later adjusted by backend retiming after the… ▽ More

    Submitted 3 May, 2026; originally announced May 2026.

  30. arXiv:2605.00370  [pdf, ps, other

    cs.LG cs.CY cs.MM

    Group Cognition Learning: Making Everything Better Through Governed Two-Stage Agents Collaboration

    Authors: Chunlei Meng, Pengbin Feng, Rong Fu, Hoi Leong Lee, Xiaojing Du, Zhaolu Kang, Zeyu Zhang, Weilin Zhou, Chun Ouyang, Zhongxue Gan

    Abstract: Centralized multimodal learning commonly compresses language, acoustic, and visual signals into a single fused representation for prediction. While effective, this paradigm suffers from two limitations: modality dominance, where optimization gravitates towards the path of least resistance, ignoring weaker but informative modalities, and spurious modality coupling, where models overfit to incidenta… ▽ More

    Submitted 11 May, 2026; v1 submitted 30 April, 2026; originally announced May 2026.

    Comments: This study has been Accepted by ICML 2026. The current version is a manuscript, please refer to the official version released at ICML 2026 for the final published version

  31. MappingEvolve: LLM-Driven Code Evolution for Technology Mapping

    Authors: Rongliang Fu, Yi Liu, Qiang Xu, Tsung-Yi Ho

    Abstract: Technology mapping is a critical yet challenging stage in logic synthesis. While Large Language Models (LLMs) have been applied to generate optimization scripts, their potential for core algorithm enhancement remains untapped. We introduce MappingEvolve, an open-source framework that pioneers the use of LLMs to directly evolve technology mapping code. Our method abstracts the mapping process into… ▽ More

    Submitted 29 April, 2026; originally announced April 2026.

  32. arXiv:2604.23781  [pdf, ps, other

    cs.CV cs.SE

    ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

    Authors: Fanqing Meng, Lingxiao Du, Zijian Wu, Guanzheng Chen, Xiangyan Liu, Jiaqi Liao, Chonghe Jiang, Zhenglin Wan, Jiawei Gu, Pengfei Zhou, Rui Huang, Ziqi Zhao, Shengyuan Ding, Ailing Yu, Bo Peng, Bowei Xia, Hao Sun, Haotian Liang, Ji Xie, Jiajun Chen, Jiajun Song, Liu Yang, Ming Xu, Qionglin Qiu, Runhao Fu , et al. (24 additional authors not shown)

    Abstract: Language-model agents are increasingly used as persistent coworkers that assist users across multiple working days. During such workflows, the surrounding environment may change independently of the agent: new emails arrive, calendar entries shift, knowledge-base records are updated, and evidence appears across images, scanned PDFs, audio, video, and spreadsheets. Existing benchmarks do not adequa… ▽ More

    Submitted 5 May, 2026; v1 submitted 26 April, 2026; originally announced April 2026.

    Comments: github repo: https://github.com/evolvent-ai/ClawMark

  33. arXiv:2604.22209  [pdf, ps, other

    eess.AS cs.AI cs.CL cs.SD

    UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions

    Authors: Chunyu Qiang, Xiaopeng Wang, Kang Yin, Yuzhe Liang, Yuxin Guo, Teng Ma, Ziyu Zhang, Tianrui Wang, Cheng Gong, Yushen Chen, Ruibo Fu, Chen Zhang, Longbiao Wang, Jianwu Dang

    Abstract: Generative audio modeling has largely been fragmented into specialized tasks, text-to-speech (TTS), text-to-music (TTM), and text-to-audio (TTA), each operating under heterogeneous control paradigms. Unifying these modalities remains a fundamental challenge due to the intrinsic dissonance between structured semantic representations (speech/music) and unstructured acoustic textures (sound effects).… ▽ More

    Submitted 24 April, 2026; originally announced April 2026.

    Comments: Accepted to ACL 2026 main conference (oral)

  34. arXiv:2604.20268  [pdf, ps, other

    cs.CV

    Opportunistic Bone-Loss Screening from Routine Knee Radiographs Using a Multi-Task Deep Learning Framework with Sensitivity-Constrained Threshold Optimization

    Authors: Zhaochen Li, Xinghao Yan, Runni Zhou, Xiaoyang Li, Chenjie Zhu, Gege Wang, Yu Shi, Lixin Zhang, Rongrong Fu, Liehao Yan, Yuan Chai

    Abstract: Background: Osteoporosis and osteopenia are often undiagnosed until fragility fractures occur. Dual-energy X-ray absorptiometry (DXA) is the reference standard for bone mineral density (BMD) assessment, but access remains limited. Knee radiographs are obtained at high volume for osteoarthritis evaluation and may offer an opportunity for opportunistic bone-loss screening. Objective: To develop an… ▽ More

    Submitted 22 April, 2026; originally announced April 2026.

  35. arXiv:2604.16548  [pdf, ps, other

    cs.CR cs.AI cs.CL

    A Survey on Long-Term Memory Security in LLM Agents: Attacks, Defenses, and Governance Across the Memory Lifecycle

    Authors: Zehao Lin, Xixuan Hao, Renyu Fu, Shaobo Cui, Kai Chen, Chunyu Li, Zhiyu Li, Feiyu Xiong

    Abstract: The emergence of writable, cross-session persistent memory in LLM agents introduces a qualitatively different threat landscape from conventional input-centric security concerns, characterized by three properties: persistence, statefulness, and propagation. To systematically characterize this landscape, we propose a Memory Lifecycle Framework that organizes attacks, defenses, and their cross-phase… ▽ More

    Submitted 11 June, 2026; v1 submitted 17 April, 2026; originally announced April 2026.

    ACM Class: K.6.5; I.2.0; D.4.6

  36. arXiv:2604.16024  [pdf, ps, other

    cs.MA cs.CV

    AstroVLM: Expert Multi-agent Collaborative Reasoning for Astronomical Imaging Quality Diagnosis

    Authors: Yaohui Han, Tianshuo Wang, Zixi Zhao, Zhengchun Zhu, Shuo Ren, Yiru Wang, Rongliang Fu, Tinghuan Chen, Tsung-Yi Ho

    Abstract: Vision Language Models (VLMs) have been applied to several specific domains and have shown strong problem-solving capabilities. However, astronomical imaging, a quite complex problem involving multidisciplinary knowledge and several subtasks, has not been adequately studied. Due to the complexity of the astronomical imaging process, both world-class astronomical organizations, such as NASA, and ex… ▽ More

    Submitted 17 April, 2026; originally announced April 2026.

  37. arXiv:2604.14538  [pdf, ps, other

    cond-mat.str-el cond-mat.mes-hall

    Discovery of an odd-parity f-wave charge order in a kagome metal

    Authors: Jiangchang Zheng, Caiyun Chen, Ruiqin Fu, Luca Buiarelli, Zihan Lin, Fazhi Yang, Tianhao Guo, Ganesh Pokharel, Andrea Capa Salinas, Sen Zhou, Turan Birol, Stephen D. Wilson, Junzhang Ma, Daniel J. Schultz, Xianxin Wu, Berthold Jäck

    Abstract: The spontaneous breaking of symmetries is a cornerstone of physics, defining the phases of matter from the cosmological scale to the quantum realm. In condensed matter, electronic orders are classified by their behavior under fundamental symmetries like spatial inversion (parity). While even-parity orders, such as conventional superconductivity and charge density waves, are ubiquitous, their odd-p… ▽ More

    Submitted 15 April, 2026; originally announced April 2026.

    Comments: 5 figures

  38. arXiv:2604.08299  [pdf, ps, other

    cs.CL cs.AI

    SeLaR: Selective Latent Reasoning in Large Language Models

    Authors: Renyu Fu, Guibo Luo

    Abstract: Chain-of-Thought (CoT) has become a cornerstone of reasoning in large language models, yet its effectiveness is constrained by the limited expressiveness of discrete token sampling. Recent latent reasoning approaches attempt to alleviate this limitation by replacing discrete tokens with soft embeddings (probability-weighted mixtures of token embeddings) or hidden states, but they commonly suffer f… ▽ More

    Submitted 19 April, 2026; v1 submitted 9 April, 2026; originally announced April 2026.

    Comments: Camera-ready for ACL 2026 (main conference)

  39. arXiv:2604.08184  [pdf, ps, other

    cs.SD cs.AI

    AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan

    Authors: Yuankun Xie, Haonan Cheng, Jiayi Zhou, Xiaoxuan Guo, Tao Wang, Jian Liu, Weiqiang Wang, Ruibo Fu, Xiaopeng Wang, Hengyan Huang, Xiaoying Huang, Long Ye, Guangtao Zhai

    Abstract: The rapid advancement of Audio Large Language Models (ALLMs) has enabled cost-effective, high-fidelity generation and manipulation of both speech and non-speech audio, including sound effects, singing voices, and music. While these capabilities foster creativity and content production, they also introduce significant security and trust challenges, as realistic audio deepfakes can now be generated… ▽ More

    Submitted 9 April, 2026; originally announced April 2026.

    Comments: Accepted to the ACM Multimedia 2026 Grand Challenge

  40. arXiv:2604.03586  [pdf, ps, other

    cs.CL

    MultiPress: A Multi-Agent Framework for Interpretable Multimodal News Classification

    Authors: Tailong Luo, Hao Li, Rong Fu, Xinyue Jiang, Huaxuan Ding, Yiduo Zhang, Zilin Zhao, Simon Fong, Guangyin Jin, Jianyuan Ni

    Abstract: With the growing prevalence of multimodal news content, effective news topic classification demands models capable of jointly understanding and reasoning over heterogeneous data such as text and images. Existing methods often process modalities independently or employ simplistic fusion strategies, limiting their ability to capture complex cross-modal interactions and leverage external knowledge. T… ▽ More

    Submitted 4 April, 2026; originally announced April 2026.

    Comments: Accepted in International Joint Conference on Neural Networks (IJCNN) 2026

  41. arXiv:2604.03212  [pdf, ps, other

    cs.CV

    ProtoFlow: Mitigating Forgetting in Class-Incremental Remote Sensing Segmentation via Low-Curvature Prototype Flow

    Authors: Jiekai Wu, Rong Fu, Chuangqi Li, Zijian Zhang, Guangxin Wu, Hao Zhang, Shiyin Lin, Yang Li, Dongxu Zhang, Amir H. Gandomi, Simon Fong, Pengbin Feng

    Abstract: Remote sensing segmentation in real deployment is inherently continual: new semantic categories emerge, and acquisition conditions shift across seasons, cities, and sensors. Despite recent progress, many incremental approaches still treat training steps as isolated updates, which leaves representation drift and forgetting insufficiently controlled. We present ProtoFlow, a time-aware prototype dyna… ▽ More

    Submitted 20 August, 2026; v1 submitted 3 April, 2026; originally announced April 2026.

  42. arXiv:2603.27813  [pdf, ps, other

    cs.CV

    MuSEAgent: A Multimodal Reasoning Agent with Stateful Experiences

    Authors: Shijian Wang, Jiarui Jin, Runhao Fu, Zexuan Yan, Xingjian Wang, Mengkang Hu, Eric Wang, Xiaoxi Li, Kangning Zhang, Li Yao, Wenxiang Jiao, Xuelian Cheng, Yuan Lu, Zongyuan Ge

    Abstract: Research agents have recently achieved significant progress in information seeking and synthesis across heterogeneous textual and visual sources. In this paper, we introduce MuSEAgent, a multimodal reasoning agent that enhances decision-making by extending the capabilities of research agents to discover and leverage stateful experiences. Rather than relying on trajectory-level retrieval, we propos… ▽ More

    Submitted 29 March, 2026; originally announced March 2026.

  43. arXiv:2603.25688  [pdf, ps, other

    cs.RO

    Intelligent Navigation and Obstacle-Aware Fabrication for Mobile Additive Manufacturing Systems

    Authors: Yifei Li, Ruizhe Fu, Huihang Liu, Guha Manogharan, Feng Ju, Ilya Kovalenko

    Abstract: As the demand for mass customization increases, manufacturing systems must become more flexible and adaptable to produce personalized products efficiently. Additive manufacturing (AM) enhances production adaptability by enabling on-demand fabrication of customized components directly from digital models, but its flexibility remains constrained by fixed equipment layouts. Integrating mobile robots… ▽ More

    Submitted 26 March, 2026; originally announced March 2026.

    Comments: 8 pages, 4 figures, conference

  44. arXiv:2603.23512  [pdf, ps, other

    cs.CL cs.AI cs.IR

    S-Path-RAG: Semantic-Aware Shortest-Path Retrieval Augmented Generation for Multi-Hop Knowledge Graph Question Answering

    Authors: Rong Fu, Yemin Wang, Tianxiang Xu, Yongtai Liu, Weizhi Tang, Wangyu Wu, Xiaowen Ma, Simon Fong

    Abstract: We present S-Path-RAG, a semantic-aware shortest-path Retrieval-Augmented Generation framework designed to improve multi-hop question answering over large knowledge graphs. S-Path-RAG departs from one-shot, text-heavy retrieval by enumerating bounded-length, semantically weighted candidate paths using a hybrid weighted $k$-shortest, beam, and constrained random-walk strategy, learning a differenti… ▽ More

    Submitted 5 March, 2026; originally announced March 2026.

    Journal ref: WWW 2026

  45. arXiv:2603.18634  [pdf, ps, other

    cs.CV cs.LG

    SwiftGS: Episodic Priors for Immediate Satellite Surface Recovery

    Authors: Rong Fu, Jiekai Wu, Xiaowen Ma, Shiyin Lin, Kangan Qian, Chuang Liu, Simon James Fong

    Abstract: Rapid, large-scale 3D reconstruction from multi-date satellite imagery is vital for environmental monitoring, urban planning, and disaster response, yet remains difficult due to illumination changes, sensor heterogeneity, and the cost of per-scene optimization. We introduce SwiftGS, a meta-learned system that reconstructs 3D surfaces in a single forward pass by predicting geometry-radiation-decoup… ▽ More

    Submitted 28 July, 2026; v1 submitted 19 March, 2026; originally announced March 2026.

    Comments: 23 pages, 6 figures

  46. arXiv:2603.18031  [pdf, ps, other

    cs.LG cs.AI

    InfoMamba: An Attention-Free Hybrid Mamba-Transformer Model

    Authors: Youjin Wang, Jiaqiao Zhao, Rong Fu, Run Zhou, Ruizhe Zhang, Jiani Liang, Suisuai Cao, Feng Zhou

    Abstract: Balancing fine-grained local modeling with long-range dependency capture under computational constraints remains a central challenge in sequence modeling. While Transformers provide strong token mixing, they suffer from quadratic complexity, whereas Mamba-style selective state-space models (SSMs) scale linearly but often struggle to capture high-rank and synchronous global interactions. We present… ▽ More

    Submitted 8 March, 2026; originally announced March 2026.

  47. arXiv:2603.16306  [pdf, ps, other

    cs.CV

    DriveFix: Spatio-Temporally Coherent Driving Scene Restoration

    Authors: Heyu Si, Brandon James Denis, Muyang Sun, Dragos Datcu, Yaoru Li, Xin Jin, Ruiju Fu, Yuliia Tatarinova, Federico Landi, Jie Song, Mingli Song, Qi Guo

    Abstract: Recent advancements in 4D scene reconstruction, particularly those leveraging diffusion priors, have shown promise for novel view synthesis in autonomous driving. However, these methods often process frames independently or in a view-by-view manner, leading to a critical lack of spatio-temporal synergy. This results in spatial misalignment across cameras and temporal drift in sequences. We propose… ▽ More

    Submitted 17 March, 2026; originally announced March 2026.

  48. arXiv:2603.14938  [pdf, ps, other

    cs.CV

    FAR-Drive: Frame-AutoRegressive Video Generation in Closed-Loop Autonomous Driving

    Authors: Yaoru Li, Federico Landi, Marco Godi, Xin Jin, Ruiju Fu, Yufei Ma, Muyang Sun, Heyu Si, Qi Guo

    Abstract: Despite rapid progress in autonomous driving, reliable training and evaluation of driving systems remain fundamentally constrained by the lack of scalable and interactive simulation environments. Recent generative video models achieve remarkable visual fidelity, yet most operate in open-loop settings and fail to support fine-grained frame-level interaction between agent actions and environment evo… ▽ More

    Submitted 16 March, 2026; originally announced March 2026.

  49. arXiv:2603.11606  [pdf, ps, other

    cs.CV

    Articulat3D: Reconstructing Articulated Digital Twins From Monocular Videos with Geometric and Motion Constraints

    Authors: Lijun Guo, Haoyu Zhao, Xingyue Zhao, Rong Fu, Linghao Zhuang, Siteng Huang, Zhongyu Li, Hua Zou

    Abstract: Building high-fidelity digital twins of articulated objects from visual data remains a central challenge. Existing approaches depend on multi-view captures of the object in discrete, static states, which severely constrains their real-world scalability. In this paper, we introduce Articulat3D, a novel framework that constructs such digital twins from casually captured monocular videos by jointly e… ▽ More

    Submitted 24 June, 2026; v1 submitted 12 March, 2026; originally announced March 2026.

    Comments: Accepted to ECCV 2026. 26 pages, 12 figures

  50. arXiv:2603.09566  [pdf, ps, other

    cs.CV

    GeoAlignCLIP: Enhancing Fine-Grained Vision-Language Alignment in Remote Sensing via Multi-Granular Consistency Learning

    Authors: Xiao Yang, Ronghao Fu, Zhuoran Duan, Zhiwen Lin, Xueyan Liu, Bo Yang

    Abstract: Vision-language pretraining models have made significant progress in bridging remote sensing imagery with natural language. However, existing approaches often fail to effectively integrate multi-granular visual and textual information, relying primarily on global image-text alignment. This limitation hinders the model's ability to accurately capture fine-grained details in images, thus restricting… ▽ More

    Submitted 10 March, 2026; originally announced March 2026.