Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 502 results for author: Meng, X

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.21252  [pdf, ps, other

    cs.CL cs.AI cs.DB cs.IR

    EnSI-RAG: Entity-Structure-Indexed Retrieval-Augmented Generation for Long-Document Question Answering

    Authors: Xuanyu Meng, Jiashuo Sun, Jash Rajesh Parekh, Jiawei Han

    Abstract: Question answering (QA) over long, connected documents remains challenging because relevant evidence may span multiple entities and their relationships. Existing retrieval-augmented generation (RAG) methods typically index documents as raw chunks and retrieve them through embedding similarity. Their performance degrades when chunk boundaries separate entities from supporting evidence or when a que… ▽ More

    Submitted 21 August, 2026; originally announced August 2026.

    Comments: 21 pages, preprint

  2. arXiv:2608.13611  [pdf, ps, other

    cs.IT

    SN-ASMO: Satellite-Navigation Array Spatial-Manifold Precise Observation Theory A Mathematical Foundation for Observation Formation, Unified U(1) Geometry, Intrinsic Information, and Preservation of the RTK Integer Structure

    Authors: Xianwei Meng

    Abstract: Suppressive interference, high-dynamic motion of the receiving platform, and carrier-phase RTK lead satellite navigation to one fundamental question: when the receiver actively participates in observation formation through its array geometry, channel states, weights, and signal-processing rules, how can the resulting carrier observation continue to represent the same objective propagation process… ▽ More

    Submitted 12 August, 2026; originally announced August 2026.

  3. arXiv:2608.09139  [pdf, ps, other

    cs.CV

    CodecArena: Codec Quality Assessment via Visual Reinforcement Learning

    Authors: Jiaye Fu, Weiqi Li, Qiankun Gao, Yanchen Zhao, Xiandong Meng, Jian Zhang, Siwei Ma, Jiaqi Zhang

    Abstract: Video coding is advancing into the low and ultra-low bitrate regime, driven by end-to-end codecs that replace the hand-crafted pipeline with jointly optimized neural networks and generative codecs that exploit the priors of video generation models. Yet the dominant metrics, LPIPS and DISTS, measure feature and texture similarity rather than content fidelity: a reconstruction that hallucinates a wr… ▽ More

    Submitted 10 August, 2026; originally announced August 2026.

    Comments: The project page is: https://jyfu-vcl.github.io/codecarena

  4. arXiv:2608.08418  [pdf, ps, other

    cs.CV

    Learning Deep Modality-Shared Self-Expressiveness for Image Clustering with Textual Information

    Authors: Xianghan Meng, Wei He, Zhiyuan Huang, Chun-Guang Li

    Abstract: Leveraging textual information for image clustering has emerged as a promising direction, largely owing to the powerful representations learned by Vision-Language Models (VLMs). Existing approaches typically retrieve a textual counterpart for each image and then refine multimodal representations by directly enforcing cross-modal agreement, e.g., maximizing image-text similarity inherited from pret… ▽ More

    Submitted 8 August, 2026; originally announced August 2026.

  5. arXiv:2608.07078  [pdf, ps, other

    cs.DC cs.AI

    Scalable High-Fidelity Macromolecular Docking for GPU-Accelerated Supercomputers

    Authors: Xiangyu Meng, Peng Chen, Mingzhen Li, Jianmin Wang, Sen Wang, Guangming Tan, Weile Jia, Mohamed Wahib, Tao Luo, Xun Wang

    Abstract: Flexible macromolecular docking offers high-fidelity predictions of biomolecular interactions, but remains prohibitively expensive at scale. Among existing approaches, LightDock leverages Glowworm Swarm Optimization (GSO) for accuracy, yet suffers from limited parallelism, irregular computation, and severe load imbalance, preventing efficient execution on GPU supercomputers. We present SparkleDock… ▽ More

    Submitted 21 August, 2026; v1 submitted 7 August, 2026; originally announced August 2026.

    Comments: Accepted in the International Conference for High Performance Computing, Networking, Storage, and Analysis(SC'26)

  6. arXiv:2608.03304  [pdf, ps, other

    cs.CV

    Recurrent Contrastive Learning for Imbalanced Medical Image Classification

    Authors: Zhiyuan Zhu, Xinling Meng, Junxuan Yu, Jiongquan Chen, Qiongying Ni, Tuhang Shao, Yuhao Huang, Luping Zhou, Ruiyang Huang, Yuxue Wang, Rongliang Zhang, Xue Wang, Tianhong Tang, Likun Wang, Junbo Chen, Yong Jiang, Yongping Lu, Xin Yang

    Abstract: Medical image classification often suffers from class imbalance due to the inherent disparities in disease incidence. Existing approaches, such as class resampling and loss reweighting, mainly improve learning within the observed feature distribution, but do not explicitly enlarge the latent support region of tail classes. As a result, tail-class representations remain overly compact and are easil… ▽ More

    Submitted 4 August, 2026; originally announced August 2026.

    Comments: 10 pages, 3 figures

    Journal ref: The 7th MICCAI Workshop on Advances in Simplifying Medical UltraSound.2026

  7. arXiv:2608.02206  [pdf, ps, other

    cs.CV

    CLEAR: Conflict-aware Learning via Evidence-guided Adaptive Routing for Unified Sparse-View 3D Gaussian Super-Resolution

    Authors: Hantang Li, Qiang Zhu, Xiandong Meng, Debin Zhao, Xiaopeng Fan

    Abstract: Sparse-view 3D Gaussian Splatting Super-resolution is highly challenging since the sparse and low-resolution (LR) inputs lack sufficient geometric and high-frequency information for accurate reconstruction. To achieve high-quality reconstruction, existing sparse-view super-resolution methods adhere to two-stage pipeline that performs LR Gaussian reconstruction and then high-resolution (HR) Gaussia… ▽ More

    Submitted 3 August, 2026; originally announced August 2026.

    Comments: 9 pages, 5 figures

  8. arXiv:2607.28126  [pdf, ps, other

    cs.AI

    ConMem: Contribution-Aware Memory for Long-Horizon Manufacturing Inspection Logs

    Authors: Bingchen Liu, Yuanyuan Fang, Lei Liu, Guangyuan Dong, Xing Fu, Yuanyuan Gao, Shuyue Wei, Xin Li, Xiangtian Meng

    Abstract: Long-horizon steel-equipment inspection requires reasoning over heterogeneous records accumulated across repeated inspection cycles. Existing retrieval-augmented generation systems treat historical logs as a static corpus and retain records without estimating their diagnostic value, failing to report early risk. To this end, we propose ConMem, a contribution-aware memory framework for LLM-assisted… ▽ More

    Submitted 9 August, 2026; v1 submitted 30 July, 2026; originally announced July 2026.

  9. arXiv:2607.27952  [pdf, ps, other

    cs.CV cs.AI

    LAST: The Last Query Token Guides Visual Token Pruning for Edge-Cloud Collaborative MLLM Inference

    Authors: Feng Yang, Xinrui Ju, Keyang Zhang, Xiandong Meng, Rongqun Lin, Howard Leung, Shiqi Wang, Haoliang Li, Chris Xing Tian

    Abstract: Multimodal foundation models are reshaping edge-cloud visual intelligence from task-specific feature pipelines into token-based interfaces, where edge devices encode visual inputs into tokens for a general-purpose cloud MLLM. However, dense visual-token sequences increase cloud-side inference costs. Existing pruning methods mainly target centralized inference: vision-driven methods can operate bef… ▽ More

    Submitted 30 July, 2026; originally announced July 2026.

  10. arXiv:2607.25430  [pdf, ps, other

    cs.CR

    SafeStats: Efficient 2PC Protocols for Data Statistic-Related Functions

    Authors: Tanren Liu, Xianjia Meng, Yang Liu, Xin Kang, Chenhui You, Yong Zeng, Zhuo Ma

    Abstract: Statistical analysis on sensitive datasets like medical records and financial transactions is essential for decision-making, but raises significant privacy concerns. While existing secure Two-Party Computation (2PC) makes extensive efforts in designing the common secure primitives (e.g., addition and multiplication) or machine learning-related functions, few pay attention to the statistical functi… ▽ More

    Submitted 28 July, 2026; originally announced July 2026.

    Comments: 18 pages, 6 figures

  11. arXiv:2607.23444  [pdf, ps, other

    cs.CR

    Isolated but Exposed: Persistence-Based Memory Extraction Attack on LLM Agents

    Authors: Xinyu Gao, Wenyu Chen, Xiangtao Meng, Li Wang, Chuanchao Zang, Jianing Wang, Zheng Li, Shanqing Guo

    Abstract: LLM-based agents extend large language models with long-term memory (LTM) that persists privacy-sensitive user data across sessions. Production systems mitigate extraction risks through memory isolation, binding each user's LTM to a unique identifier. This defense has blocked known attacks on shared storage, fostering the assumption that isolated LTM is secure. We identify the tool interface as an… ▽ More

    Submitted 25 July, 2026; originally announced July 2026.

  12. arXiv:2607.17896  [pdf, ps, other

    cs.CV

    Locality-Aware Density Control for Efficient Gaussian-based Image Representation

    Authors: Jiacong Chen, Qingyu Mao, Xiandong Meng, Shuai Liu, Chao Li, Fanyang Meng, Youneng Bao, Yongsheng Liang

    Abstract: 2D Gaussian Splatting is an attractive direction for image representation due to its explicit formulation, fast rasterization, and favorable decoding efficiency. The representation quality of this paradigm depends on the proper allocation of Gaussian capacity to the demanding regions. However, existing methods fail to allocate Gaussian capacity efficiently during optimization: under-reconstructed… ▽ More

    Submitted 20 July, 2026; originally announced July 2026.

    Comments: Accepted by ACMMM 2026

  13. arXiv:2607.17269  [pdf, ps, other

    cs.AI cs.DB

    An Explicit World Model Based on Data-First Ontology: DaoQL Multimodal Storage Validation and Counterfactual Reasoning Evaluation

    Authors: Zhanbo Li, Shifeng Wu, Xiangjin Meng, Wenjie Cai

    Abstract: Large language models encode world models implicitly in neural weights, which exposes four structural risks in high-precision domains such as medicine and finance: hallucination, frozen knowledge, poor explainability, and poor modifiability. This paper proposes data-first ontology: LLMs are treated as reasoning and language engines, while deterministic knowledge is moved into an explicit multimoda… ▽ More

    Submitted 19 July, 2026; originally announced July 2026.

    Comments: 20 pages, 2 figures. Code: https://github.com/zhanbolee/DaoQL-Edu

    ACM Class: I.2.4; H.2.4; I.2.6

  14. arXiv:2607.16257  [pdf, ps, other

    cs.LG cs.AI cs.CL

    From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training

    Authors: Zishang Jiang, Tingyun Li, Jinyi Han, Xinyi Wang, Sihang Jiang, Yizhou Ying, Xiaojun Meng, Jiansheng Wei, Jiaqing Liang, Yanghua Xiao

    Abstract: Reinforcement learning (RL) has become a widely adopted technique for improving large language models (LLMs) on complex tasks. Despite this progress, existing RL methods still face challenges in training agents with longer-horizon interactions. One major bottleneck is distinguishing the contribution of different actions in long-horizon interaction, leading to high optimization variance. To address… ▽ More

    Submitted 28 June, 2026; originally announced July 2026.

    Comments: Accepted at ICML 2026

  15. arXiv:2607.13079  [pdf, ps, other

    cs.AR cs.PL

    ChipVerilog: A Large-Scale OpenCores-Derived Benchmark for LLM-Based Verilog RTL Generation

    Authors: Yan Tan, Jiping Du, Xiangchen Meng, Yangdi Lyu

    Abstract: Large language models have shown strong potential for Verilog RTL generation. However, many existing benchmarks are built from short, self-contained module-level tasks. These tasks are useful for controlled evaluation, but they do not fully capture the code scale, hierarchy, and module interactions found in practical IP and processor-core RTL. We present ChipVerilog, a description-to-Verilog gener… ▽ More

    Submitted 16 August, 2026; v1 submitted 12 July, 2026; originally announced July 2026.

  16. arXiv:2607.12495  [pdf, ps, other

    cs.CV

    RealSkin: Spatio-Spectral Partial Neural Adjoint Maps for Image-to-3D Attribute Transfer

    Authors: Jing Li, Yawei Luo, Xiangze Meng, Ying Li, Tieru Wu, Rui Ma

    Abstract: Creating photorealistic 3D assets requires bridging the appearance gap between real-world observations and synthetic models. A promising approach is to transfer visual attributes from real images onto synthetic 3D surfaces. Traditional methods struggle with resolution mismatch and the inherent discreteness of point correspondences. In contrast, resolution-robust functional maps enable smooth attri… ▽ More

    Submitted 14 July, 2026; originally announced July 2026.

  17. arXiv:2607.12186  [pdf, ps, other

    cs.CV

    Overview of Cross-Component In-loop Filters in Video Coding Standards

    Authors: Zhaoyu Li, Xuewei Meng, Jiaqi Zhang, Cheng Huang, Chuanmin Jia, Siwei Ma, Yun Jiang

    Abstract: In-loop filters have been comprehensively explored during the development of video coding standards due to their remarkable noise-reduction capability. In the early stage of video coding, in-loop filters, such as Deblocking Filter, Sample Adaptive Offset, and Adaptive Loop Filter, were performed separately for each component. Recently, cross-component filters were studied to improve the chroma fid… ▽ More

    Submitted 13 July, 2026; originally announced July 2026.

    Comments: This paper was submitted and accepted by ZTE Communications

  18. arXiv:2607.08221  [pdf, ps, other

    cs.CV

    LUMI: Tokenizer-Agnostic LLM-Based Lossless Image Compression

    Authors: Chris Xing Tian, Chengkai Wu, Ziyu Wang, Rongqun Lin, Kecheng Chen, Xiandong Meng, Haoliang Li, Shiqi Wang, Siwei Ma

    Abstract: Large language model (LLM)-based lossless image compression methods typically represent pixel data through the native text interface of a pretrained model, converting pixel values into token sequences that the LLM processes through its vocabulary head. This design shows that pretrained language models can provide probability estimates for image coding, but it also couples compression to tokenizer… ▽ More

    Submitted 9 July, 2026; originally announced July 2026.

    Comments: Preprint

  19. arXiv:2607.06943  [pdf, ps, other

    cs.CV

    General Incomplete Multimodal Learning via Dynamic Quality Perception

    Authors: Xiangyu Meng, Shicai Wei

    Abstract: Multimodal learning robust to missing modalities is essential for real-world applications. Existing methods mainly focus on inter-modality missing, where entire modalities are absent, while overlooking intra-modality degradation, where modalities are present but severely corrupted. In practice, these two types of missing often coexist, making existing approaches ineffective. To address this limita… ▽ More

    Submitted 7 July, 2026; originally announced July 2026.

    Comments: Accepted by ECCV 2026. Corresponding author: Shicai Wei

  20. arXiv:2607.05737  [pdf, ps, other

    cs.CV

    Optimized Adaptive Loop Filter in Versatile Video Coding

    Authors: Xuewei Meng, Jiaqi Zhang, Chuanmin Jia, Xinfeng Zhang, Shanshe Wang, Siwei Ma

    Abstract: In the Versatile Video Coding~(VVC) standard, adaptive loop filter~(ALF), including Geometry transformation-based Adaptive Loop Filter~(GALF) and Cross Component Adaptive Loop Filter~(CCALF), plays an essential role in reducing compression artifacts. However, it also has high coding complexity and requires many picture buffer accesses in the encoder that will increase external memory access and is… ▽ More

    Submitted 8 July, 2026; v1 submitted 6 July, 2026; originally announced July 2026.

    Comments: This paper was submitted to DCC 2021 and accepted as a poster

  21. arXiv:2607.05471  [pdf, ps, other

    cs.SE cs.AI

    KAT-Coder-V2.5 Technical Report

    Authors: Bo Huang, Fengxiang Li, Hao Xu, Haoyang Huang, Hongyi Fu, Jinhua Hao, Kun Yuan, Minglei Zhang, Pengcheng Xu, Shiyang Liu, Wenhao Zhuang, Yuze Shi, Zongxian Feng, Chao Wang, Cheng He, Chongling Rao, Deyu Cao, Fan Yang, Gang Xiong, Haochen Liu, Jiabao Li, Jian Liang, Jinghui Jia, Jingwen Chang, Jun Du , et al. (28 additional authors not shown)

    Abstract: We present KAT-Coder-V2.5, a coding-focused agentic model trained to act autonomously inside real, executable repositories rather than as a single-turn code generator. Its capability is bottlenecked less by model scale than by the scarcity of reproducible environments, verifiable rewards, and high-value trajectories, which we address with an end-to-end agentic post-training framework. AutoBuilder… ▽ More

    Submitted 6 July, 2026; originally announced July 2026.

    Comments: 24 pages, 5 figures

  22. arXiv:2607.05241   

    cs.RO

    GelNeuro: A Sensing-Computing Integrated Neuromorphic Tactile System for Texture Recognition

    Authors: Luoyang Bian, Xinpan Meng, Zhenghua Ma, Houcheng Li, Long Cheng

    Abstract: Neuromorphic visuo-tactile sensing offers a promising paradigm for low-latency and low-power robotic perception. However, existing systems still rely heavily on a host computer for event readout, preprocessing, or relaying prior to chip inference. This paper presents GelNeuro, a fully integrated sensing-computing visuo-tactile system that directly pairs a GelSight Mini-based optical tactile front… ▽ More

    Submitted 11 July, 2026; v1 submitted 6 July, 2026; originally announced July 2026.

    Comments: The authors withdraw this preprint as the work requires substantial revision and additional validation. The current version is not suitable for public dissemination, and further development of the research is needed

  23. arXiv:2607.03387  [pdf, ps, other

    cs.RO

    Feeling the Unexpected: ResTacVLA for Contact-Rich Manipulation via Residual Tactile Representation

    Authors: Pengwei Zhang, Bin Xie, Xinpan Meng, Xinyu Guo, Ce Hao, Fang Deng, Long Cheng, Tiancai Wang

    Abstract: Tactile perception is indispensable for contact-rich manipulation, yet integrating it into Vision-Language-Action (VLA) models often induces modality collapse, where high-bandwidth visual features overshadow sparse tactile cues. Inspired by Predictive Coding, a neural mechanism where the brain attenuates predictable inputs to prioritize surprising stimuli, we propose ResTacVLA. Rather than treatin… ▽ More

    Submitted 19 July, 2026; v1 submitted 3 July, 2026; originally announced July 2026.

    Comments: 8 pages, 6 figures, 3 tables. Accepted by IROS 2026, Project page: https://awilekong.github.io/ResTacVLA/

  24. arXiv:2607.03348  [pdf, ps, other

    cs.IT eess.SP

    Diffusion-Based Noise-Adaptive Null-Space Channel Estimation for OFDM Systems

    Authors: Heqiang Qi, Yirun Chen, Xiangming Meng, Chunxiao Jiang, Sheng Wu, Linling Kuang

    Abstract: Accurate channel estimation in orthogonal frequency division multiplexing (OFDM) systems remains challenging when demodulation reference signal (DMRS) observations are sparse and noisy, and when DMRS configurations vary across deployment scenarios. This paper proposes DANCE (Diffusion-based Noise-Adaptive Null-space Channel Estimation), a diffusion-based channel estimator for OFDM systems. We form… ▽ More

    Submitted 3 July, 2026; originally announced July 2026.

  25. arXiv:2607.02407  [pdf, ps, other

    cs.AI cs.CV

    Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments

    Authors: Xianhui Meng, Zirui Song, Yuchen Zhang, Li Zhang, Yongxuan Lv, Xiuying Chen, Kun Wang, Yan Luo, Kai Chen, Hangjun Ye, Long Chen, Jun Liu, Xiaoshuai Hao

    Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in 3D indoor synthesis for Manhattan environments. However, existing methods often fail to capture plausible object layout patterns in non-Manhattan settings, primarily because they struggle to model non-orthogonal spatial relationships, leading to high geometric violations and low physical fidelity. To address this challenge,… ▽ More

    Submitted 2 July, 2026; originally announced July 2026.

  26. arXiv:2607.00498  [pdf, ps, other

    cs.CV

    Robust 3D Alignment of Generative Reconstructions via Partial Monocular Observations

    Authors: Yuchen Zhang, Luanyuan Dai, Yiwei Wang, Xiwei Xu, Jianing Zhang, Johnny. r. zhang, Xianhui Meng, Yanbiao Ma, Jiayi Ma, Xiaoshuai Hao

    Abstract: Aligning generative 3D reconstructions with partial monocular observations is a critical but under-explored challenge in computer vision. This task is inherently ill-posed due to severe asymmetries between noisy, sparse monocular inputs and dense generative priors, whose scale ambiguity and geometric hallucinations, combined with the lack of initial overlap, render traditional registration pipelin… ▽ More

    Submitted 1 July, 2026; originally announced July 2026.

  27. arXiv:2606.23879  [pdf

    eess.IV cs.AI

    Promise and challenges of heart chamber segmentation from non-contrast CT scans using contrastive unpaired image translation: a feasibility study

    Authors: Jing Wang, Tong Yu, Hao-En Lu, Zixue Zeng, Joseph K. Leader, Xin Meng, Jianbing Zhu, Jiantao Pu

    Abstract: Purpose: To evaluate the feasibility and challenges of heart chamber segmentation from non-contrast CT scans using contrastive unpaired image translation and deep learning-based segmentation. Approach: We developed ChameleonNet, a framework utilizing the Contrastive Unpaired Translation (CUT) network with decoupled contrastive learning (DCL) loss to synthesize non-contrast CT from contrast CT scan… ▽ More

    Submitted 22 June, 2026; originally announced June 2026.

  28. arXiv:2606.23565  [pdf, ps, other

    cs.RO cs.CV

    HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory

    Authors: Xiaolin Zhou, Liu Liu, Tingyang Xiao, Wei Feng, Fa Fu, Xinrui Meng, Xinjie Wang, Jialiang Han, Boyang Yu, Yun Du, Wei Sui, Zhizhong Su

    Abstract: LLM agents follow a practical execution loop in digital environments: they reason over structured states, invoke tools, inspect feedback, and revise actions. Extending this loop to physical robots is difficult because physical execution is continuous, embodiment-dependent, uncertain, and constrained by safety. Existing embodied-AI systems have advanced manipulation, spatial understanding, navigati… ▽ More

    Submitted 22 June, 2026; originally announced June 2026.

  29. arXiv:2606.23023  [pdf, ps, other

    cs.CV

    Boosting Neural Video Codec via Scale-Driven Online Flow Refinement

    Authors: Tiange Zhang, Rongqun Lin, Haocheng Tang, Xiandong Meng, Weijia Jiang, Zhimeng Huang, Siwei Ma

    Abstract: Although state-of-the-art neural video codecs (NVCs) have achieved remarkable performance, they suffer from limited generalization when encountering complex motion patterns unseen during training. To bridge this domain gap without the expensive cost of online fine-tuning, we propose a Training-Free Scale-Driven Online Flow Refinement (SOFR) method. Serving as a plug-and-play module, SOFR integrate… ▽ More

    Submitted 22 June, 2026; originally announced June 2026.

    Comments: Accepted to ICME 2026 as an oral paper

    ACM Class: I.4.2

  30. arXiv:2606.16575  [pdf, ps, other

    cs.LG math-ph

    RepNN: Tackling spectral bias in deep neural networks via parameter reparameterization

    Authors: Yong Wang, Tao Zhou, Xuhui Meng

    Abstract: Deep neural networks (DNNs) have achieved remarkable success in scientific computing, yet they often suffer from spectral bias in capturing oscillatory and multiscale behaviors. In this study, we investigate this limitation by examining the failure of shallow ReLU neural networks in fitting high-frequency functions. This observation identifies two important factors in resolving rapid oscillations:… ▽ More

    Submitted 18 June, 2026; v1 submitted 15 June, 2026; originally announced June 2026.

  31. arXiv:2606.14173  [pdf, ps, other

    cs.GR physics.optics

    HoloPathTracer: Fast and Accurate Wave Path Tracing for Holography

    Authors: Wenbin Zhou, Xiangyu Meng, Jiankai Xing, Xin Liu, Suyeon Choi, Yifan Peng

    Abstract: Holography offers unique advantages for delivering perceptual realism while preserving compact form factors in VR/AR. Its perceptual quality, however, hinges on encoding rich wavefronts of photorealistic scenes into interference patterns and then incoherently multiplexing the resulting wave fields for perception. Existing CGH paradigms decouple radiance estimation from wave propagation by pre-rend… ▽ More

    Submitted 16 June, 2026; v1 submitted 12 June, 2026; originally announced June 2026.

    Comments: ACM Transactions on Graphics, Vol. 45, No. 4, Article 39, July 2026, 12 pages. Presented at SIGGRAPH 2026

    ACM Class: I.3.7; I.3.3

    Journal ref: ACM Trans. Graph. 45, 4, Article 39 (July 2026), 12 pages

  32. arXiv:2606.12550  [pdf, ps, other

    cs.RO cs.AI

    Foresight: Iterative Reasoning About Clues that Matter for Navigation

    Authors: Arthur Zhang, Carl Qi, Donne Su, Xiangyun Meng, Amy Zhang, Joydeep Biswas

    Abstract: Open-world mapless navigation from sparse language instructions requires resolving underspecified goals and inferring which environmental cues are relevant for reaching the goal. For instance, reaching an out-of-view destination may require interpreting ramps, signs, or detours that reveal where to go or which route to take. Prior works are limited by their reliance on known navigation factors and… ▽ More

    Submitted 10 June, 2026; originally announced June 2026.

    Comments: 22 pages, 10 figures, 3 tables

  33. arXiv:2606.07020  [pdf, ps, other

    cs.CL

    MADE: Beyond Scoring via a Multilingual Agentic Diagnosing Engine for Fine-Grained Evaluation Insights

    Authors: Yilun Liu, Miao Zhang, Shimin Tao, Minggui He, Chunguang Zhao, Chenxin Liu, Li Zhang, Chen Liu, Cheng Qian, Liqun Deng, Xiaojun Meng, Daimeng Wei

    Abstract: Multilingual and multicultural benchmarks now cover dozens of languages and model families, but the resulting score landscapes remain metric-rich and insight-poor, necessitating fine-grained multilingual post-evaluation diagnosis. However, single LLMs and open-ended agents are easily swamped by the long, noisy diagnostic input, and no reusable taxonomy exists for it. To address this, we propose MA… ▽ More

    Submitted 5 June, 2026; originally announced June 2026.

  34. arXiv:2606.06370  [pdf, ps, other

    cs.RO

    Ensuring Interaction Safety in Multitask Exoskeleton Control: A Simulation-Trained Variable Impedance Framework

    Authors: Muyuan Ma, Houcheng Li, Haotian Zhai, Lijun Han, Xinpan Meng, Xiuze Xia, Long Cheng

    Abstract: Wearable exoskeletons can augment human phys ical capabilities during complex activities. However, ensuring adaptation across diverse tasks while guaranteeing interaction safety remains a critical challenge. To address this, a simulation trained variable impedance control approach with stability guarantees is proposed. First, a simulation-based human exoskeleton motion data generation pipeline is… ▽ More

    Submitted 4 June, 2026; originally announced June 2026.

  35. Spatio-Temporal Correlation Guided Geometric Partitioning for Versatile Video Coding

    Authors: Xuewei Meng, Chuanmin Jia, Xinfeng Zhang, Shanshe Wang, Siwei Ma

    Abstract: Geometric partitioning has attracted increasing attention by its remarkable motion field description capability in the hybrid video coding framework. However, the existing geometric partitioning (GEO) scheme in Versatile Video Coding (VVC) causes a non-negligible burden for signaling the side information. Consequently, the coding efficiency is limited. In view of this, we propose a spatio-temporal… ▽ More

    Submitted 1 June, 2026; originally announced June 2026.

    Journal ref: IEEE Transactions on Image Processing, vol. 31, pp. 30-42, 2022

  36. Edge-directed geometric partitioning for versatile video coding

    Authors: Xuewei Meng, Xinfeng Zhang, Chuanmin Jia, Xia Li, Shanshe Wang, Siwei Ma

    Abstract: To improve the coding performance, geometric partition (GEO) was proposed for the upcoming VVC standard. GEO provides 140 partition candidates. The index of optimal GEO mode needs to be signaled explicitly. Considering different structural characteristics of different CUs and the correlation between spatial adjacent blocks and temporal collocated blocks, we propose a GEO mode prediction strategy b… ▽ More

    Submitted 31 May, 2026; originally announced June 2026.

    Comments: This paper has been published in IEEE ICME

    Journal ref: IEEE International Conference on Multimedia and Expo (ICME), 2020, pp. 1-6

  37. Deformable Wiener Filter for Future Video Coding

    Authors: Xuewei Meng, Chuanmin Jia, Xinfeng Zhang, Shanshe Wang, Siwei Ma

    Abstract: In-loop filters have attracted increasing attention due to the remarkable noise-reduction capability in the hybrid video coding framework. However, the existing in-loop filters in Versatile Video Coding (VVC) mainly take advantage of the image local similarity. Although some non-local based in-loop filters can make up for this shortcoming, the widely-used unsupervised parameter estimation method b… ▽ More

    Submitted 31 May, 2026; originally announced June 2026.

    Comments: This paper has been published in IEEE Transactions on Image Processing

    Journal ref: IEEE Transactions on Image Processing, vol. 31, pp. 7222-7236, 2022

  38. arXiv:2605.29607  [pdf, ps, other

    cs.LG

    Cluster-Level Attention-Guided Parallel Decoding for Masked Diffusion Language Models

    Authors: Heqiang Qi, Wei Huang, Mingyuan Bai, Xiangming Meng

    Abstract: Masked diffusion language models (MDLMs) enable parallel decoding by predicting all masked positions at each denoising step, yet existing training-free samplers usually decide which positions to commit at token-level granularity. We revisit this granularity and observe that reliable predictions often emerge as contiguous high-confidence spans, suggesting that the unit of parallel commitment can be… ▽ More

    Submitted 28 May, 2026; originally announced May 2026.

  39. arXiv:2605.29280  [pdf, ps, other

    cs.LG cs.AI cs.IR

    LoopFM: Learning frOm HistOrical RePresentations of Foundation Model for Recommendation

    Authors: Shali Jiang, Hua Zheng, Boyang Liu, Laming Chen, Kenny Lov, Chuanqi Xu, Lisang Ding, Qinghai Zhou, Can Cui, Xiaolong Liu, Xiaoyi Liu, Yasmine Badr, Xin Xu, Jiyan Yang, Ellie Dingqiao Wen, Gerard Jonathan Mugisha Akkerhuis, Chenxiao Guan, Rong Jin, Ruichao Qiu, Xian Chen, Shifu Xu, Zhehui Zhou, Ping Chen, Rui Yang, Haicheng Chen , et al. (18 additional authors not shown)

    Abstract: Knowledge distillation (KD) transfers a single scalar prediction from a large foundation model (FM) to compact vertical models (VMs), suffering from diminishing transfer ratio -- the fraction of FM improvement captured by the VM -- as a single scalar cannot convey the rich intermediate knowledge that larger FMs learn. To address this bottleneck, we propose LoopFM (Learning frOm HistOrical RePresen… ▽ More

    Submitted 2 June, 2026; v1 submitted 27 May, 2026; originally announced May 2026.

    Comments: Shali Jiang, Hua Zheng, Boyang Liu contributed equally to this work

  40. arXiv:2605.25820  [pdf, ps, other

    cs.LG

    Visual-Redundancy-Controlled Parallel Decoding for Diffusion-Based Multimodal Large Language Models

    Authors: Yulin Yuan, Hongshuo Zhao, Xiangming Meng

    Abstract: Diffusion-based multimodal large language models (dMLLMs) decode by iteratively predicting tokens at multiple masked positions in parallel. This turns each decoding step into a position-selection problem: the model must choose not only which predictions are reliable in isolation, but also which positions should be committed together as context for later decoding steps. Existing confidence-based de… ▽ More

    Submitted 9 June, 2026; v1 submitted 25 May, 2026; originally announced May 2026.

    Comments: 18 pages, 5 figures, preprint. Code is available at https://github.com/infiniteYuanyl/VRCD

  41. arXiv:2605.20646  [pdf, ps, other

    cs.SI

    DisImpact: Quantifying the Physi-Social Impact of Natural Disasters Through Social Media

    Authors: Ruichen Yao, Tejna Dasari, Xuanyu Meng, Elliot Cao, Zelin Li, Yifan Liu, Yaokun Liu, Lanyu Shang, Dong Wang

    Abstract: Natural disasters not only cause large-scale physical destruction, but also cascading social consequences that are difficult to quantify with traditional surveys and reports. Social media platforms offer an alternative perspective that captures multimodal, real-time, and user-generated content that can be leveraged for disaster impacts. In this paper, we introduce DisImpact, a two-stage framework… ▽ More

    Submitted 19 May, 2026; originally announced May 2026.

    Comments: Accepted by ICWSM 2026

  42. arXiv:2605.14514  [pdf, ps, other

    cs.CR

    Defenses at Odds: Measuring and Explaining Defense Conflicts in Large Language Models

    Authors: Xiangtao Meng, Wenyu Chen, Chuanchao Zang, Xinyu Gao, Jianing Wang, Li Wang, Zheng Li, Shanqing Guo

    Abstract: Large Language Models (LLMs) deployed in high-stakes applications must simultaneously manage multiple risks, yet existing defenses are almost exclusively evaluated in isolation under a one-shot deployment assumption. In practice, providers patch models incrementally throughout their lifecycle-responding to newly exposed vulnerabilities or targeted data-removal requests without retraining from scra… ▽ More

    Submitted 14 May, 2026; originally announced May 2026.

    Comments: Under Review

  43. arXiv:2605.13476  [pdf, ps, other

    cs.CV

    Neural Video Compression with Domain Transfer

    Authors: Tiange Zhang, Rongqun Lin, Xiandong Meng, Haofeng Wang, Xing Tian, Qi Zhang, Siwei Ma

    Abstract: Content-adaptive compression has always been a key direction in neural video coding (NVC), aiming to mitigate the domain gap between training and testing data. Such gaps often arise from distributional discrepancies between training and inference data, which may cause noticeable performance degradation when the testing content differs from the training distribution. To tackle this challenge, we pr… ▽ More

    Submitted 13 May, 2026; originally announced May 2026.

    Comments: Accepted to ISCAS 2026 as an oral paper

    ACM Class: I.4.2

  44. arXiv:2605.12072  [pdf, ps, other

    cs.CV

    PairDropGS: Paired Dropout-Induced Consistency Regularization for Sparse-View Gaussian Splatting

    Authors: Hantang Li, Qiang Zhu, Xiandong Meng, Xingtao Wang, Debin Zhao, Xiaopeng Fan

    Abstract: Dropout-based sparse-view 3D Gaussian Splatting (3DGS) methods alleviate overfitting by randomly suppressing Gaussian primitives during training. Existing methods mainly focus on designing increasingly sophisticated dropout strategies, while they overlook the resulting inconsistencies among different dropped Gaussian subsets. This oversight often leads to unstable reconstruction and suboptimal Gau… ▽ More

    Submitted 13 May, 2026; v1 submitted 12 May, 2026; originally announced May 2026.

    Comments: 11 pages,8 figures

  45. arXiv:2605.11222  [pdf, ps, other

    cs.LG

    ADMM-Q: An Improved Hessian-based Weight Quantizer for Post-Training Quantization of Large Language Models

    Authors: Ryan Lucas, Mehdi Makni, Xiang Meng, Adam Deng, Rahul Mazumder

    Abstract: Quantization is an effective strategy to reduce the storage and computation footprint of large language models (LLMs). Post-training quantization (PTQ) is a leading approach for compressing LLMs. Popular weight quantization procedures, including GPTQ and RTN, suffer in model utility, especially at aggressive quantization levels (sub-4-bit). We propose ADMM-Q, a novel weight quantization algorithm… ▽ More

    Submitted 11 May, 2026; originally announced May 2026.

  46. arXiv:2605.10114  [pdf, ps, other

    cs.CL

    SkillRAE: Agent Skill-Based Context Compilation for Retrieval-Augmented Execution

    Authors: Xiangcheng Meng, Shu Wang, Yixiang Fang

    Abstract: Large Language Model (LLM)-based agents (e.g., OpenClaw) increasingly rely on reusable skill libraries to solve artifact-rich tasks such as document-centric workflows and data-intensive analysis. As these libraries grow, a few works have attempted to study the Retrieval-Augmented Execution (RAE), which often first retrieves some external skills and other knowledge, then compiles the context using… ▽ More

    Submitted 11 May, 2026; originally announced May 2026.

  47. arXiv:2605.05753  [pdf, ps, other

    cs.CV

    Jointly Learning Structured Representations and Stabilized Affinity for Human Motion Segmentation

    Authors: Xianghan Meng, Zhiyuan Huang, Zhengyu Tong, Chun-Guang Li

    Abstract: Human Motion Segmentation (HMS), which aims to partition a video into non-overlapping segments corresponding to different human motions, has recently attracted increasing research attention. Existing HMS approaches are predominantly based on subspace clustering, which are grounded on the assumption that the distribution of high-dimensional temporal features well aligns with a Union-of-Subspaces (U… ▽ More

    Submitted 7 May, 2026; originally announced May 2026.

    Comments: This manuscript is currently under review by the IEEE Transactions on Circuits and Systems for Video Technology (TCSVT)

  48. arXiv:2605.04649   

    cs.RO

    From Reach to Insert: Tactile-Augmented Precision Assembly under Sub-Millimeter Tolerances

    Authors: Xinpan Meng, Siyao Huang, JingPu Yang, Muyuan Ma, Zhenghua Ma, Lijun Han, Gao Yuan, Houcheng Li, Long Cheng

    Abstract: High-precision assembly frequently involves tight-tolerance insertions, where even slight pose errors can cause jamming or excessive interaction forces, making robust and safe insertion policies difficult to obtain. This paper proposes a tactile-augmented two-stage method that combines Imitation Learning (IL) and Reinforcement Learning (RL) for precision insertion tasks. In the first stage, IL lea… ▽ More

    Submitted 14 August, 2026; v1 submitted 6 May, 2026; originally announced May 2026.

    Comments: The current version is not yet suitable for public dissemination and requires further development and refinement

  49. arXiv:2604.25710  [pdf

    stat.AP cs.LG stat.ME stat.ML

    Adaptive Meta-Learning Stochastic Gradient Hamiltonian Monte Carlo Simulation for Bayesian Updating of Structural Dynamic Models

    Authors: Xianghao Meng, James L. Beck, Yong Huang, Hui Li

    Abstract: In the last few decades, Markov chain Monte Carlo (MCMC) methods have been widely applied to Bayesian updating of structural dynamic models in the field of structural health monitoring. Recently, several MCMC algorithms have been developed that incorporate neural networks to enhance their performance for specific Bayesian model updating problems. However, a common challenge with these approaches l… ▽ More

    Submitted 28 April, 2026; originally announced April 2026.

    Journal ref: Comput Meth Appl Mech Eng; 437: 117753 (2025)

  50. arXiv:2604.24806  [pdf, ps, other

    cs.IR cs.AI cs.DB

    Versioned Late Materialization for Ultra-Long Sequence Training in Recommendation Systems at Scale

    Authors: Liang Guo, Ge Song, Litao Deng, Jianhui Sun, Chufeng Hu, Lu Zhang, Zhen Ma, Shouwei Chen, Weiran Liu, Sarang Masti Sreeshylan, Xiaoxuan Meng, Yanzun Huang

    Abstract: Modern Deep Learning Recommendation Models (DLRMs) follow scaling laws with sequence length, driving the frontier toward ultra-long User Interaction History (UIH). However, the industry-standard "Fat Row" paradigm, which pre-materializes these sequences into every training example, creates a storage and I/O wall where data infrastructure usage exceeds GPU training capacity due to data redundancy t… ▽ More

    Submitted 10 June, 2026; v1 submitted 27 April, 2026; originally announced April 2026.