Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 1,047 results for author: Jia, Y

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.20617  [pdf, ps, other

    cs.AI cs.LG

    Dual-Cache Latent Space Communication between Heterogeneous Language Models

    Authors: Jiyao Liu, Qi Zhang, Yaoyi Jia, Ziwen Kan, Song Wang

    Abstract: Multi-agent LLM systems split work across models, so answering often requires knowledge that sits in another agent's context: a Sharer has encoded information that a Receiver needs to complete its task. They usually communicate by exchanging text, which puts autoregressive decoding on the critical path and reduces the exchange to a discrete message written without sight of the receiver's state. Re… ▽ More

    Submitted 20 August, 2026; originally announced August 2026.

  2. arXiv:2608.20467  [pdf, ps, other

    eess.SY cs.RO

    Learning-Based Measurement-Robust Control Barrier Functions for Obstacle Avoidance under State Estimation Error

    Authors: Nicholas Rober, Yixuan Jia, Jonathan P. How

    Abstract: Safety filters are an effective tool for enforcing constraints in safety-critical systems, but most existing methods assume perfect state information, which is rarely available in practice. Recent work has begun to close this gap by developing filtering mechanisms that are robust to state estimation error, but these methods can still exhibit safety violations or overly conservative behavior as est… ▽ More

    Submitted 20 August, 2026; originally announced August 2026.

    Comments: 8 Pages, 6 figures

  3. arXiv:2608.20320  [pdf, ps, other

    cs.AI cs.CL

    An Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Prediction

    Authors: Narges Ahmadi, Yubo Jiao, Jônatas Augusto Manzolli, Jiangbo Yu, Luis Miranda-Moreno

    Abstract: Travel behavior research increasingly combines digital data collection with predictive modeling, yet these stages are often developed and evaluated separately. This study proposes a three-agent workflow integrating conversational data collection, structured data processing, and behavioral prediction. A chatbot-administered, image-augmented stated-preference survey collected mode choices from stude… ▽ More

    Submitted 20 August, 2026; originally announced August 2026.

  4. arXiv:2608.18713  [pdf, ps, other

    cs.IT

    Joint Power Allocation and Phase-Shift Design for Beyond-Diagonal Stacked Intelligent Metasurfaces-Aided ISAC Systems

    Authors: Yuhui Jiao, Qian Zhang, Xuejun Cheng, Meihui Liu, Jiancheng An, Ju Liu

    Abstract: Stacked intelligent metasurfaces (SIM) provide an efficient architecture for integrated sensing and communication (ISAC) with few radio-frequency (RF) chains. However, diagonal SIM provide only element-wise phase control, so balancing multiuser communication and sensing performance may require additional layers. In this letter, we propose a beyond-diagonal SIM (BD-SIM) architecture for ISAC, enabl… ▽ More

    Submitted 19 August, 2026; originally announced August 2026.

    Journal ref: IEEE Wireless Communications Letters, 2026

  5. arXiv:2608.18606  [pdf, ps, other

    cs.IR

    OneModel: A Unified Foundation for Platform-Scale Multi-Scenario Ranking

    Authors: Yinqi Zhang, Peiyu Hu, Yuntian Tang, Siying Gu, Jiahao Liang, Longxin Kou, Haiqing Hu, Shuman Zhuang, Yubin Xu, Chenggen Sun, Bin Ye, Donghui Xu, Zhaoyu Liu, Jiang Rong, Yuting Jia, Zhaokai Luo, Leilei Ma, Yiying Xie, Yao Hu

    Abstract: Platform-scale recommender systems often span multiple business streams such as organic recommendation, advertising, and merchant services, where user behaviors form a continuous cross-stream trajectory. Maintaining separate ranking systems fragments user representations and increases engineering cost. We propose \textbf{OneModel}, a unified framework for multi-stream final ranking. OneModel maps… ▽ More

    Submitted 19 August, 2026; v1 submitted 19 August, 2026; originally announced August 2026.

  6. arXiv:2608.18458  [pdf, ps, other

    cs.IT

    Joint Beamforming and Phase Shifts Design for RIS-Enabled RSMA-ISAC Systems

    Authors: Xuejun Cheng, Qian Zhang, Yuhui Jiao, Yufei Zhao, Zheng Dong, Ju Liu

    Abstract: This paper investigates the sensing-centric design of reconfigurable intelligent surface (RIS)-enabled rate-splitting multiple access-integrated sensing and communication (RSMA-ISAC) systems. Specifically, we propose a new beam-gain approximation method to enhance the sensing beam gain while satisfying communication quality-of-service (QoS) constraints.Since the joint optimization of the beamformi… ▽ More

    Submitted 18 August, 2026; originally announced August 2026.

    Journal ref: IEEE Wireless Communications Letters, 2026

  7. arXiv:2608.17770  [pdf, ps, other

    cs.CR

    Efficient Fuzzy PSI under One-Sided Assumptions

    Authors: Xinpeng Yang, Meng Hao, Yanxue Jia, Chenkai Weng, Yonggang Wen, Tianwei Zhang

    Abstract: Fuzzy private set intersection (PSI) enables two parties to identify approximately matching elements between their input sets, where two elements are considered a match if their distance is at most a threshold $δ$ under a given metric. Although substantial progress has been made, existing constructions for general Minkowski distances either rely on strong two-sided geometric separation assumptions… ▽ More

    Submitted 18 August, 2026; originally announced August 2026.

    Comments: Accepted to ACM CCS 2026

  8. arXiv:2608.17163  [pdf, ps, other

    cs.LG cs.AI

    Q-Learning With World Models

    Authors: Perry Dong, Yueru Jia, Chelsea Finn, Dorsa Sadigh

    Abstract: Off-policy reinforcement learning (RL) has become increasingly sample-efficient, enabling applications such as RL fine-tuning of Vision-Language-Action models into reliable, high-performing policies. World models offer a further lever for sample efficiency, as they predict state changes rather than actions alone, but their success has largely been confined to supervised policy learning. Prior mode… ▽ More

    Submitted 17 August, 2026; originally announced August 2026.

  9. arXiv:2608.16351  [pdf, ps, other

    cs.RO

    Arm-Aware Guided Dexterous Grasp Generation with Arm-Agnostic Grasp Models

    Authors: Yongyi Jia, Yongpeng Jiang, Kangchen Lv, Yi Ren, Mingrui Yu, Xiang Li

    Abstract: Dexterous grasp generation that considers arm-related constraints is crucial in real-world scenarios involving arm environment collision avoidance, workspace boundary grasps, and consecutive grasping. Existing hand-centric grasp models, which primarily focus on the floating hand's pose, are insufficient for such cases. Conventional arm-aware methods either rely on rejection sampling to discard inf… ▽ More

    Submitted 17 August, 2026; originally announced August 2026.

  10. arXiv:2608.16134  [pdf, ps, other

    cs.LG cs.HC eess.SP q-bio.NC

    Multi-Feature Riemannian Hypergraph for Online Test-Time Adaptation of Motor Imagery Brain-Computer Interface

    Authors: Siqi Li, Zhi Li, Tong Liu, Shuai Zhang, Yanfei Jia, Zhiqiang Yi, Jue Xie, Ni Ji

    Abstract: In clinical motor imagery brain-computer interface (MI-BCI) decoding, cross-day transferability and online operation remain two critical challenges. Hypergraphs can improve transferability by capturing higher-order sample relationships, yet existing hypergraph-based methods for online emotion recognition neglect the cross-day benefits of Riemannian geometry widely adopted in EEG transfer learning.… ▽ More

    Submitted 17 August, 2026; originally announced August 2026.

  11. arXiv:2608.14611  [pdf

    cs.CY

    The 2026 Singapore Consensus on Global AI Safety Research Priorities

    Authors: Stephen Casper, Oskar Galeev, Yoshua Bengio, Mohan Kankanhalli, Lee Wan Sie, Tegan Maharaj, Chris Meserole, Luke Ong, Stuart Russell, Dawn Song, Max Tegmark, Brian Tse, Xue Lan, Andrew Yao, Zhang Ya-Qin, Zhou Bowen, Imane Bello, Kwan Yee Ng, Vanessa Wilfred, Erica Liaw, Lee Chein Inn, Lin Wanxuan, Ng En Qi, Jonathan Lee, José Villalobos , et al. (95 additional authors not shown)

    Abstract: Frontier AI capabilities and autonomy are advancing rapidly. A growing number of real-world incidents make a trusted AI ecosystem essential to embracing AI with confidence. The 2026 Singapore Consensus is an outcome of the second International Scientific Exchange on AI Safety, bringing together over 100 contributors spanning 13 countries from frontier developers, government safety institutes, acad… ▽ More

    Submitted 8 July, 2026; originally announced August 2026.

    Comments: Available at https://aisafetypriorities.org/

  12. arXiv:2608.14401  [pdf, ps, other

    stat.ML cs.LG

    Offline Deep Q* Estimation with Diffusion Models

    Authors: Xiaohong Chen, Yuling Jiao, Lican Kang, Jerry Zhijian Yang, Chen Zhong

    Abstract: In offline RL, estimating the optimal action-value function $Q^*$ can be formulated as solving the optimal Bellman equation based solely on offline observations. A fundamental challenge is that the reward function and transition kernel are unknown, so the optimal Bellman operator is not directly observable from data. To address this issue, we propose a novel framework that decouples operator estim… ▽ More

    Submitted 14 August, 2026; originally announced August 2026.

  13. arXiv:2608.14085  [pdf, ps, other

    cs.CV

    CoDS: Robust Collaborative Perception via Expert-driven Detection and BEV Segmentation

    Authors: Jinlong Wang, Yuang Jia, Junhong Lin, Nannan Li, Wei Gao

    Abstract: Collaborative perception breaks through single-view limitations via multi-agent information exchange. However, multi-source noise such as pose errors and communication delays degrades fusion feature quality, constraining perception performance. Joint training of detection and BEV segmentation provides a natural remedy, where segmented road regions help constrain target distributions and detection… ▽ More

    Submitted 14 August, 2026; originally announced August 2026.

    Comments: 10 pages, 6 figures

    Journal ref: ACMMM 2026

  14. arXiv:2608.13937  [pdf, ps, other

    stat.ML cs.LG

    Deep Vision in Smart Manufacturing: MODERN Framework for Intelligent Quality Monitoring and Diagnosis

    Authors: Yicheng Kang, Yuling Jiao, Xin Geng, Mahesh Nagarajan

    Abstract: Smart manufacturing processes are often installed with a large number of sensors, imaging devices and computers, which not only enable instant communication across various modules of a production system but also aid in intelligent manufacturing management. In this paper, we introduce MODERN, a deep learning framework for quality monitoring and fault isolation, which integrates these enhanced capab… ▽ More

    Submitted 14 August, 2026; originally announced August 2026.

  15. arXiv:2608.13833  [pdf, ps, other

    cs.IR cs.AI

    AdsWorldEngine: A Self-Evolving Conversational Advertising Agent through Orchestrator and Tool Coevolution

    Authors: Simiao Zuo, Chenhui Xu, Yimeng Jia, Qiang Lou, Jian Jiao, Denis Charles

    Abstract: Conversational advertising aims to deliver useful ads within multi-turn assistant interactions. Unlike conventional query-based advertising, where the user's intent is often expressed in a short standalone query, conversational ads must infer latent commercial intent from the current user query, the assistant response, and dialogue history while also deciding whether an ad would be helpful rather… ▽ More

    Submitted 13 August, 2026; originally announced August 2026.

  16. arXiv:2608.11685  [pdf, ps, other

    cs.CV

    EGM-Det: Entropy-Guided Multimodal Adaptive Fusion for UAV RGB-IR Object Detection

    Authors: Cunzheng Fan, Dawei Yan, Guanlin Wang, Xingshuo Yang, Yupeng Jia, Jing Yang, Haokui Zhang

    Abstract: Joint use of RGB and infrared (IR) imagery can improve UAV-view object detection, but most existing methods fuse multimodal features with static or fixed weights and therefore overlook spatially varying modality reliability. We propose EGM-Det, an entropy-guided multimodal adaptive fusion framework for RGB-IR object detection. EGM-Det employs a dual-stream architecture to preserve modality-specifi… ▽ More

    Submitted 12 August, 2026; originally announced August 2026.

    Comments: 14 pages, 7 figures, 6 tables

  17. arXiv:2608.10590  [pdf, ps, other

    cs.CV

    Rethinking Data Efficiency in Industrial Dense Prediction: Pretraining Coherence, Not Inductive Bias, Determines ViTs Low-Data Advantage

    Authors: Haoran Sui, Yaoyuan Jia

    Abstract: Vision Transformers (ViTs) are widely believed to require more labeled data than CNNs for industrial dense prediction. Through controlled experiments on four industrial datasets, we show that the data-efficiency gap stems from pretraining incoherence, which refers to the statistical mismatch between ImageNet-pretrained ViT backbones and COCO-pretrained CNN necks, rather than from inherent self-att… ▽ More

    Submitted 11 August, 2026; originally announced August 2026.

    Comments: 14 pages, 10 figures, 17 tables. This paper targets industrial defect detection via vision transformer and CNN alignment grafting

    ACM Class: I.2.6; I.2.10; I.4.8

  18. arXiv:2608.09292  [pdf, ps, other

    cs.LG cs.CL

    Beyond the Capability Boundary: Zeroth-Order Optimization for Self-Evolving LLM Agents

    Authors: Bingzhen Liu, Xiaomeng Fan, Yuwei Wu, Zhi Gao, Mingyang Gao, Chuanhao Li, Yunde Jia

    Abstract: Self-evolving methods improve the capabilities of LLM agents by sampling trajectories from the underlying LLMs and learning from these trajectories. However, these methods struggle to learn beyond the inherent capability boundary of the agents, since the agents cannot sample correct trajectories on difficult examples for further improvements. In this paper, we propose a zeroth-order self-evolution… ▽ More

    Submitted 10 August, 2026; originally announced August 2026.

  19. arXiv:2608.09273  [pdf, ps, other

    cs.AI cs.SE

    Entropy-based Code Adversarial Translation for Real-world Repository Migration

    Authors: Yushun Tang, Yisen Cao, Zhicheng Chen, Lin Peng, Junkang Mao, Fengyi Song, Yantao Jia

    Abstract: LLMs have demonstrated strong capabilities in code generation and automated program repair, but migrating an entire repository rarely produces a runnable application because long-horizon translation challenges LLM-based agents' ability to maintain repository-level migration objectives. In this work, we propose Entropy-based Code Adversarial Translation (ECAT), a multi-agent framework for automated… ▽ More

    Submitted 11 August, 2026; v1 submitted 10 August, 2026; originally announced August 2026.

  20. arXiv:2608.08204  [pdf, ps, other

    stat.ML cs.LG

    Conditional Diffusion for Nonparametric Instrumental Variable Quantile Regression

    Authors: Xingdong Feng, Xinhong Jiang, Yuling Jiao, Lican Kang, Junwei Liu

    Abstract: This work proposes deep nonparametric Instrumental variable quantile regression (IVQR), a two-stage estimator that combines conditional diffusion modeling with a kernel-smoothed conditional moment formulation. In the first stage, we estimate the joint conditional distribution of the outcome and endogenous covariates given the instrument using a variance-preserving conditional diffusion model. In t… ▽ More

    Submitted 8 August, 2026; originally announced August 2026.

  21. arXiv:2608.07341  [pdf, ps, other

    cs.CL cs.AI

    Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination

    Authors: Ruijie Hou, Yueyang Jiao, Zhao Wang, Yingming Li

    Abstract: Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. \textbf{Contamination mitigation evaluation} intervenes in the decoding process to suppress memorization and restore a contaminated model's genuine capability, but its prevailing metric, the \textbf{G-AP} (\textbf{G}ap of \textbf{A}ggregate \textbf{P}erformance), is flawed. Discr… ▽ More

    Submitted 7 August, 2026; originally announced August 2026.

  22. arXiv:2608.03046  [pdf, ps, other

    cs.CV

    CAPE-T2V: Captioner-Anchored Prompt Enhancement toward Two-Sided Conditioning Alignment in Text-to-Video Generation

    Authors: Yizhuo Jia, Jingyun Hua, Yuanxing Zhang

    Abstract: Text-to-video (T2V) diffusion transformers (DiTs) are trained with detailed video captions, whereas inference often relies on user prompts rewritten by a prompt enhancer (PE). Prior work has improved generation by optimizing the PE, the DiT, or both; some methods have also sought to narrow the training-inference mismatch through shared schemas. Yet even within a shared schema, inference-time PE ou… ▽ More

    Submitted 3 August, 2026; originally announced August 2026.

    Comments: Includes appendix; 11 figures. Project page: https://github.com/yizzz927/CAPE-T2V

  23. arXiv:2608.02162  [pdf, ps, other

    cs.SE cs.AI cs.PL

    Lossless Tensor Compression as Program Synthesis

    Authors: Jieke Shi, Junda He, Wenjia Jiang, Weifeng Sun, Shidong Pan, Zhensu Sun, Chengran Yang, Peixin Zhang, Yifan Jia, Zhou Yang, Thong Hoang, Xiwei Xu, Zhenchang Xing, David Lo

    Abstract: Model checkpoints are growing in both number and size, which makes archival, transfer, and deployment increasingly costly. General-purpose compressors can reduce storage requirements but ignore tensor structure, whereas existing tensor-specific compressors rely on fixed and format-specific pipelines. We present Brevis, which formulates lossless tensor compression as program synthesis. We design a… ▽ More

    Submitted 3 August, 2026; originally announced August 2026.

  24. arXiv:2608.01635  [pdf, ps, other

    cs.CV

    Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning

    Authors: Qianlong Yang, Bowen Ye, Xianda Guo, Yanlun Peng, Wenke Huang, Hongyuan Zhang, Yulei Jia

    Abstract: Despite the progress of multimodal large language models (MLLMs), they continue to exhibit deficiencies in visual perception. Following visual instruction tuning, internal MLLM representations rapidly deviate from their original semantic states during inference, causing severe information degradation. While existing methods attempt to leverage external vision foundation models (VFMs) to align inte… ▽ More

    Submitted 2 August, 2026; originally announced August 2026.

    Comments: This paper has been accepted by ACM MM 2026

  25. arXiv:2608.00965  [pdf, ps, other

    cs.CR cs.AI

    An AI Approach to Verified Production Cryptographic Libraries

    Authors: Chuyue Sun, Su Fong, Zhiyi Kuang, Yizheng Jiao, Nina Narodytska, Haoze Wu, David L. Dill, Clark Barrett

    Abstract: Cryptographic code is critical infrastructure that must be correct, yet formally verifying production libraries remains difficult. Existing language-model proof systems solve isolated obligations with specifications and premises already given, leaving production-library verification unresolved. We present CryptoProver, an AI-based system that synthesizes internal specifications and Verus-checked… ▽ More

    Submitted 1 August, 2026; originally announced August 2026.

  26. arXiv:2607.29593  [pdf, ps, other

    cs.LG math.OC

    Convergence and Regret of the Policy Gradient for Multi-Armed Bandits in Diffusion Environment

    Authors: Yanwei Jia, Du Ouyang

    Abstract: This paper studies the policy gradient update for a multi-arm bandit problem in diffusion environment that is described by a stochastic differential equation (SDE) under the continuous-time reinforcement learning framework by Wang et al. (2020), Jia and Zhou (2022b). With the logit parameterization for the stochastic policy, we show that it converges almost surely to the optimal arm under an arbit… ▽ More

    Submitted 31 July, 2026; originally announced July 2026.

  27. arXiv:2607.28674  [pdf, ps, other

    cs.AI cs.CL cs.LG

    How Hard Does It Think? Analyzing Step-Aware Reasoning Energy in LLM Chain-of-Thought Trajectories

    Authors: Hui Wei, Junda Wu, Sheldon Yu, Sizhe Zhou, Yizhu Jiao, Ming Zhong, Bowen Jin, Tong Yu, Shijia Pan, Jiawei Han, Julian McAuley

    Abstract: Understanding how computational effort is allocated across individual chain-of-thought (CoT) reasoning steps remains an open challenge: existing interpretability methods rely on output-level signals or collapse processing depth into a single trajectory-level scalar, leaving step-wise effort opaque. We propose Step-Aware Reasoning Energy (SARE), a geometric framework that quantifies effort at the g… ▽ More

    Submitted 28 July, 2026; originally announced July 2026.

    Comments: 13 pages, 3 figures

  28. arXiv:2607.28421  [pdf, ps, other

    cs.AI

    When Derived Measurements Mislead: Quantifying and Mitigating LLM Over-Trust with Privileged-Modality Reliability Evidence

    Authors: Zongheng Guo, Tao Chen, Tianli Li, Mingzhe Cui, Yang Jiao, Lei Xie, Yi Pan, Xiao Hu, Manuela Ferrario

    Abstract: Derived measurements increasingly enter large language model (LLM) pipelines as direct facts despite their instance-dependent validity. We define derived-feature over-trust (DFOT) as the failure in which a downstream LLM assigns such a measurement the epistemic status of a direct fact or uses it outside its valid scope. Using physiological sensing as a case study, D1 tests acceptance of a PPG-deri… ▽ More

    Submitted 30 July, 2026; originally announced July 2026.

    Comments: 25 pages, including references and supplementary material; 3 figures and 19 tables. Code: https://github.com/Zongheng-Guo/When-Derived-Measurements-Mislead

    ACM Class: I.2.7; I.2.6; J.3

  29. arXiv:2607.27632  [pdf, ps, other

    cs.LG

    First-order Constrained Trilevel Optimization Over Distributed Networks for Robust Coreset Selection

    Authors: Yang Jiao, Kaixuan Jiao, Kai Yang, Nadjib Aitsaadi, Ilhem Fajjari, Renwei, Li

    Abstract: With the rapid advancement of the Internet of Things (IoT), massive amounts of data are generated across distributed edge networks. Training models on full data incurs significant computational overhead and storage bottlenecks, rendering coreset selection a critical paradigm. Furthermore, given the privacy-sensitive nature of local data and the escalating demand for model robustness in real-world… ▽ More

    Submitted 29 July, 2026; originally announced July 2026.

  30. arXiv:2607.26637  [pdf, ps, other

    cs.CL cs.AI

    Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability

    Authors: Sizhe Zhou, Sheldon Yu, Hui Wei, Junda Wu, Siru Ouyang, Yizhu Jiao, Shijia Pan, Julian McAuley, Yu Zhang, Tong Yu, Jiawei Han

    Abstract: Deployed LLM agents increasingly keep their long-term memory as a filesystem: a directory tree of markdown files that the agent itself reads, writes, and reorganizes through generic file tools. Yet research has largely passed over this medium: prior systems design bespoke memory representations and study retrieval over them, leaving the default's two working assumptions untested: that an agent can… ▽ More

    Submitted 29 July, 2026; originally announced July 2026.

    Comments: 59 pages, 12 figures, 18 tables

  31. arXiv:2607.24665  [pdf, ps, other

    cs.CV cs.GR cs.LG

    MMOE: Modernizing Diffusion Transformers with Efficient Expert Design

    Authors: Yanhao Jia, Jiepeng Wang, Haibin Huang, Chi Zhang, Erik Cambria, Xuelong Li

    Abstract: Modern large language models scale successfully by pairing capacity growth with efficiency, keeping per-token and deployment costs under control as capacity grows. AIGC Foundation Models (AFMs), especially diffusion-transformer backbones, have begun to adopt sparse experts, but recent efforts mostly enlarge total parameter counts and sparsity ratios without importing the efficiency mechanisms that… ▽ More

    Submitted 27 July, 2026; originally announced July 2026.

    Comments: 13 pages, 5 figures

  32. arXiv:2607.23702  [pdf, ps, other

    cs.RO

    Try Once, Then Optimal: De-Redundified Procedure Memory for Cross-Episode Exploration Amortization

    Authors: Haizhou Ge, Haochen Ouyang, Zhixing Chen, Yufei Jia, Yue Li, Lu Shi, Lei Han, Guyue Zhou, Ruqi Huang

    Abstract: Manipulating objects with hidden internal state, such as a latched microwave, forces a robot to probe before it can act. Yet a robot that has solved an instance once re-runs the same probes whenever it encounters that instance again, because existing cross-episode memories target task success and organize reuse around states, not the object or the cost of re-exploring it. We present Instance-Orien… ▽ More

    Submitted 26 July, 2026; originally announced July 2026.

    Comments: 8 pages, 3 figures, 3 tables

    ACM Class: I.2.9; I.2.6; I.2.10

  33. arXiv:2607.22184  [pdf, ps, other

    cs.CR cs.SE

    DeFiScreener: Efficient DeFi Attack Pre-screening in Smart Contracts via Historical Case Matching

    Authors: Rui Cao, Shaojing Fan, Zhimei Sui, Liming Fang, Ziqi Yang, Yingying Jiao, Zhenguang Liu

    Abstract: Blockchain and its killer applications, particularly decentralized finance (DeFi), are gaining widespread adoption, with over 5,200 DeFi projects deployed on mainstream blockchains as of January 2026. At the same time, security risks in DeFi are becoming increasingly serious. However, existing DeFi detection tools usually cover only specific attack types, exhibiting severely limited detection cove… ▽ More

    Submitted 24 July, 2026; originally announced July 2026.

  34. arXiv:2607.21556  [pdf, ps, other

    cs.CV cs.AI

    Visual Contrastive Self-Distillation

    Authors: Yijun Liang, Yunjie Tian, Yijiang Li, Yuqi Jia, Furong Huang, Tianyi Zhou, Di Fu

    Abstract: On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation (OPD), yet it still needs asymmetric information between teacher and student to ensure that the self-teacher provides a stronger learning signal than the student. Existing methods create this asymmetry either through privileged answers or visual evidence. We ask whether both can be… ▽ More

    Submitted 23 July, 2026; originally announced July 2026.

    Comments: 15 pages

  35. arXiv:2607.21417  [pdf, ps, other

    cs.CV

    Towards Privacy-Preserving Federated Prompt Tuning under Data Heterogeneity: A Subspace-Decomposed Expert Approach

    Authors: Yuhua Wang, Xiaodong Li, Yihao Guo, Yuxiang Jia, Qinnan Zhang, Yifan Sun, Hainan Zhang, Yongxin Tong, Zhiming Zheng

    Abstract: Federated prompt tuning (FPT) enables collaborative adaptation of vision--language models (VLMs) using lightweight prompts. Existing methods often address heterogeneity and privacy through a split-prompt design under local differential privacy (DP), combining a shared prompt for global transfer with private prompts for local adaptation. However, a single shared prompt may over-smooth diverse trans… ▽ More

    Submitted 23 July, 2026; originally announced July 2026.

    Comments: Accepted by ACM MM 2026

  36. arXiv:2607.20557  [pdf, ps, other

    cs.LG cs.AI

    Monkey King Bang: A Unified Scientific Multimodal Foundation Model

    Authors: Hesen Chen, Xinyu Su, Xiaomeng Yang, Yuetan Lin, Zixiong Yang, Junyi An, Fenglei Cao, Yifeng Jiao, Yunqi Zhang, Yuan Cheng, Zhiyu Tan, Hao Li, Libo Wu, Yuan Qi

    Abstract: Scientific discovery is increasingly shifting from isolated disciplines to multi-domain reasoning, and AI for science faces a similar transition. Existing systems are either specialised for individual domains or unify scientific data mainly through text tokenisation and prompt-based interfaces, limiting their ability to handle diverse scientific inputs, produce modality-native outputs, and support… ▽ More

    Submitted 17 July, 2026; originally announced July 2026.

  37. arXiv:2607.20467  [pdf, ps, other

    cs.AI cs.LG

    DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding

    Authors: Yanhua Jiao, Tianyi Wu, Xiaoxi Sun, Yulin Li, HuiLing Zhen, Libo Qin, Baotian Hu, Zhuotao Tian, Min Zhang

    Abstract: While parallel decoding is central to the efficiency of Diffusion Large Language Models (dLLMs), current strategies are often hindered by overly conservative confidence thresholds. These thresholds, necessitated by the Joint Probability Dependence Error (JPDE), result in redundant denoising iterations and suboptimal inference speeds. To overcome this, we propose DC-Leap, a training-free framework… ▽ More

    Submitted 19 May, 2026; originally announced July 2026.

  38. arXiv:2607.19333  [pdf, ps, other

    cs.LG cs.AI stat.ML

    Provable diffusion-based posterior sampling for linear inverse problems via DDIM

    Authors: Yuchen Jiao, Na Li, Changxiao Cai, Yuxin Chen, Gen Li

    Abstract: Diffusion-based methods have achieved remarkable empirical success in solving inverse problems. However, many existing posterior samplers either lack rigorous theoretical guarantees or incur substantial computational overhead. We propose a simple and efficient algorithm, called \pddim, for solving linear inverse problems with diffusion priors via a DDIM-type sampler. Our method requires only light… ▽ More

    Submitted 21 July, 2026; originally announced July 2026.

  39. arXiv:2607.17708  [pdf, ps, other

    cs.AI

    LaT: LLM-as-Trainer for Multi-Task Vehicle Routing Solvers

    Authors: Yang Wang, Ya-Hui Jia, Wei-Neng Chen, Yi Mei, Wen Song, Zhiguang Cao

    Abstract: Multi-task neural solvers aim to handle multiple Vehicle Routing Problem (VRP) variants within a unified model, avoiding separate training for each constraint combination. However, VRP variants differ in optimization difficulty, while existing methods lack stage-wise feedback on their training status, making the model biased to some specific variants. Although meta-learning can support adaptive tr… ▽ More

    Submitted 20 July, 2026; originally announced July 2026.

    Comments: 28 pages

  40. arXiv:2607.17062  [pdf, ps, other

    cs.GT

    Equilibrium analysis of three-player General Lotto game with leader-follower framework

    Authors: Yang Jiao, Dunbiao Niu, Yiguang Hong

    Abstract: In this paper, we introduce the General Lotto game with a regulator (R-Lotto), a leader-follower extension of the classical two-player General Lotto game. The model captures regulatory interventions in competitive resource allocation, where a regulator first chooses an intervention parameter to influence the subsequent competition between two resource-constrained followers. The intervention parame… ▽ More

    Submitted 19 July, 2026; originally announced July 2026.

  41. arXiv:2607.16725  [pdf, ps, other

    stat.ML cs.LG

    Semi-Supervised Conditional Generative Learning through Stochastic Interpolation and Sufficient Representations

    Authors: Changyu Liu, Yuling Jiao, Jian Huang

    Abstract: Conditional generative modeling remains a challenging problem in semi-supervised settings where labeled data is scarce but unlabeled samples are abundant. To effectively leverage structural information embedded within the unlabeled dataset and compensate for sparse conditioning signals, we propose a semi-supervised framework combining conditional stochastic interpolation with low-dimensional laten… ▽ More

    Submitted 18 July, 2026; originally announced July 2026.

    Comments: 35 pages, 4 figures and 1 table

    MSC Class: 62G05; 68T07

  42. arXiv:2607.16682  [pdf, ps, other

    cs.LG cs.AI

    Building a Neural Network from Scratch: Implementation, Evaluation, and Optimization

    Authors: Yuanzhe Jia

    Abstract: The widespread adoption of high-level deep learning libraries, while accelerating model development, has increasingly abstracted away the internal mechanics of neural networks, creating a gap between practical usage and fundamental understanding. To address this, the paper presents a self-contained neural network framework implemented entirely from scratch -- without relying on automatic different… ▽ More

    Submitted 18 July, 2026; originally announced July 2026.

  43. arXiv:2607.10789  [pdf, ps, other

    cs.AI

    Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging

    Authors: Siyi Chen, Jiahe Ying, Yixuan Jia, Yuxuan Gu, Enze Ye, Weimin Bai, Zhijun Zeng, Shaochi Ren, Binhong Gao, Yubing Li, Tianhan Zhang, He Sun

    Abstract: Computational imaging, which recovers hidden signals from indirect, noisy measurements, underpins quantitative discovery across scientific disciplines, yet building a correct reconstruction pipeline demands deep domain expertise and remains laborious even for domain scientists. We introduce Imaging-101, a benchmark of 57 expert-verified computational imaging tasks spanning six scientific domains,… ▽ More

    Submitted 12 July, 2026; originally announced July 2026.

  44. arXiv:2607.10237  [pdf, ps, other

    cs.CV

    CoSAG: Compact Semantic Anchor Gaussians via Training-Free Rate-Distortion Coding

    Authors: Yuang Jia, Jinlong Wang, Junhong Lin, Ruiting Dai, Wei Gao

    Abstract: Open-vocabulary 3D scene understanding is commonly achieved by embedding 2D vision-language features such as CLIP into a 3D Gaussian Splatting scene, turning it into a text-queryable semantic field. However, attaching a high-dimensional feature to each of millions of Gaussians inflates a single scene to gigabytes, which makes storage and deployment the real bottleneck of these fields. Existing com… ▽ More

    Submitted 11 July, 2026; originally announced July 2026.

  45. arXiv:2607.07039  [pdf

    eess.IV cs.CV physics.med-ph

    From Data Completeness to Data Sufficiency: A Task-Driven Imaging Framework for Intraoperative CBCT under Quality-Time-Dose Trade-offs

    Authors: Yi Jia, Rongjun Ge, Yang Chen, Yan Xi, Wenjun Xia

    Abstract: Mobile C-arm cone-beam computed tomography (CBCT) has been widely used for real-time intraoperative 3D imaging. However, current practice often mechanically applies the fan-beam CT criterion of "180° plus fan angle" in pursuit of "data completeness" in reconstruction. This review argues that, under the single circular trajectory of three-dimensional cone-beam geometry, complete data are mathematic… ▽ More

    Submitted 8 July, 2026; originally announced July 2026.

  46. arXiv:2607.06700  [pdf, ps, other

    cs.RO

    CILC: Cryptographically-secure Inter-agent Loop Closure Candidate Detection for Multi-Agent Collaborative SLAM

    Authors: Andrew Fishberg, Yixuan Jia, Jonathan P. How

    Abstract: Multi-agent Simultaneous Localization and Mapping (SLAM) and collaborative SLAM (CSLAM) require robots to continuously exchange global descriptors (GDs) to detect inter-agent loop closures (ILCs). While encrypted radios protect this traffic from external eavesdroppers, they offer no protection against a compromised swarm member. We show this threat is concrete by demonstrating how a corrupted agen… ▽ More

    Submitted 7 July, 2026; originally announced July 2026.

  47. arXiv:2607.06564  [pdf, ps, other

    cs.RO cs.CV

    Lift3D-VLA: Lifting VLA Models to 3D Geometry and Dynamics-Aware Manipulation

    Authors: Jiaming Liu, Qingpo Wuwu, Nuowei Han, Hao Chen, Zhuoyang Liu, Fan Fei, Yueru Jia, Chenyang Gu, Yandong Guo, Boxin Shi, Shanghang Zhang

    Abstract: Recently, Vision-Language-Action (VLA) models have demonstrated strong generalization across diverse tasks. However, effective robotic manipulation in physical environments fundamentally requires geometric understanding and spatial reasoning. While some VLA approaches attempt to incorporate 3D information, they are constrained by limited data availability and geometric information loss in current… ▽ More

    Submitted 7 July, 2026; originally announced July 2026.

    Comments: 14 pages, 7 figures. Project website: https://lift3dvla.github.io/

  48. arXiv:2607.04825  [pdf, ps, other

    cs.MA

    Dynamic Airspace Management for UAVs in Evolving Urban Environments: Collaborative Coordination and Human Safety

    Authors: Lin Sun, Yuhang Wang, Fan Meng Hong, Haopeng Chen, Yan Jiao, Yongming Xu

    Abstract: The low-altitude economy is an emerging industry with significant development potential, in which the safety of unmanned aerial vehicle (UAV) operations is a critical challenge. Particularly within complex urban topographies and human-populated environments, UAV airspace management must prioritize collision avoidance and human safety. We propose Pharos, a collaborative multi-UAV airspace managemen… ▽ More

    Submitted 6 July, 2026; originally announced July 2026.

  49. arXiv:2607.04554  [pdf, ps, other

    cs.RO

    HUGS: Guiding Unified Dexterous Grasp Synthesis Across Modes and Scales via Learned Human Priors

    Authors: Mingrui Yu, Yongpeng Jiang, Yongyi Jia, Kangchen Lv, Li Huang, Yi Ren, Xiang Li

    Abstract: Dexterous grasping across diverse object scales requires contact modes ranging from two-finger pinches to bimanual grasps. Existing dexterous grasp synthesis methods reduce the high-dimensional optimization space with manually designed expected contacts and initialization heuristics, which struggle to balance synthesis success rate and diversity. We present HUGS (Human-prior-guided Unified Dextero… ▽ More

    Submitted 5 July, 2026; originally announced July 2026.

    Comments: The first two authors contributed equally. Project website: https://hugs-dex.github.io/

  50. arXiv:2607.04241  [pdf, ps, other

    cs.CV cs.AI cs.LG

    Hierarchical Multi-to-Single-Modal Knowledge Distillation for Disruption Prediction in EAST

    Authors: Qiang Chen, Xiao Wang, Hao Si, Qingquan Yang, Meiwen Chen, Jianhua Yang, Xiaofeng Han, Yunhu Jia, Ran Chen, Liang Wang, Jin Tang, Guosheng Xu

    Abstract: Plasma disruption is a critical threat to tokamak safety. Existing data-driven predictors mainly rely on time-series diagnostic signals, while visible images provide complementary spatial cues including plasma deformation, local brightening, and radiation-structure evolution. Although the image modality improves the model's discriminative capability, it also substantially increases the computation… ▽ More

    Submitted 5 July, 2026; originally announced July 2026.