Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 308 results for author: Dong, M

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.16022  [pdf, ps, other

    cs.SE

    OpenHarmony Bench: Evaluating LLMs and Coding Agents on OpenHarmony App Development

    Authors: Li Li, Han Hu, Tianjian Zhang, Xin Peng, Fangzhu Mao, Qingyu Zhang, Xiaoheng Xie, Zhongmin Tang, Zhihao Lin, Haolin Ruan, Miaomiao Dong, Liuchuan Zhu, Yue Li, Chi Chen, Wenkang Zhong, Mingfei Zhang, Yang Yu, Bo Sun, Chaorui Zhang, Weixi Zhang, Wei Han, Bo Bai, Kui Liu, Gang Fan, Siru Liu , et al. (5 additional authors not shown)

    Abstract: We present OPENHARMONY BENCH, an app-level coding benchmark for evaluating LLM-based coding agents on OpenHarmony ArkTS applications. Unlike function-level benchmarks, it evaluates complete app-level changes: each task requires an agent to modify a buildable ArkTS project so that a requested behavior works end to end, involving UI state, data persistence, build configuration, and platform APIs. Th… ▽ More

    Submitted 16 August, 2026; originally announced August 2026.

  2. arXiv:2608.07088  [pdf, ps, other

    cs.CV cs.AI

    RoRA: Role-Oriented Regional Allocation for Visual Token Pruning in MLLMs

    Authors: Qiyanhui Lu, Han Wu, Rongjian Xu, Tingzhang Luo, Cheng Fan, Xinghao Chen, Minjing Dong, Jufeng Yang, Jianyuan Guo

    Abstract: Multimodal large language models (MLLMs) encode images as long visual token sequences, making prefilling and KV-cache storage expensive. Existing training-free pruning methods select tokens by importance, diversity, or spatial coverage, but treat retained tokens as interchangeable and do not explicitly track which object-related regions are already covered. We present RoRA, a training-free framewo… ▽ More

    Submitted 7 August, 2026; originally announced August 2026.

    Comments: 9 pages, 4 figures, 4 tables. Code is available at https://github.com/LukieLuu/RoRA

  3. arXiv:2608.04756  [pdf, ps, other

    cs.CR cs.AI

    PURPOSE: Poisoning Conflict Resolution in RAG via Proxy-Fact-Grounded Updates

    Authors: Zijian Wang, Yubo Zhu, Muzhi Dong, Yanjun Lou, Yisheng Li, ZiLiang Zhang, Wei Tong, Yuan Zhang, Jingyu Hua, Sheng Zhong

    Abstract: In Retrieval-Augmented Generation (RAG), post-retrieval conflict resolution arbitrates among noisy or contradictory retrieved passages. However, the robustness of this safeguard against knowledge poisoning has not been adequately studied. Existing black-box poisoning methods all assert the target answer in frontal contradiction with what the resolver treats as settled, the very signal these method… ▽ More

    Submitted 5 August, 2026; originally announced August 2026.

  4. arXiv:2608.00643  [pdf, ps, other

    cs.CR cs.NI

    Domain Decoupling Attack: Exploiting the Validation Gap Between Protective DNS and Shared Edge Routing

    Authors: Weizhe Wang, Minhong Dong, Jinhao Li, Yao Zhang, Hao Liu, Qiang Hu, Tao Luo, Guangquan Xu, Bin Wu

    Abstract: Network attackers often conceal malicious communication within legitimate Internet traffic. Existing CDN-based evasion techniques rely on SNI--Host inconsistency, insufficient domain ownership verification, or provider-specific routing rewrites, which limit their applicability in modern CDN environments. We identify a validation gap in DNS-based authorization, where permission derived from an allo… ▽ More

    Submitted 1 August, 2026; originally announced August 2026.

  5. arXiv:2607.28312  [pdf, ps, other

    cs.CV cs.AI

    ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding

    Authors: Mingkang Dong, Muxin Pu, Jie Li, Bohan Guo, Songruo Chen, Bin Ren, Xu Zheng, Chen Zhao, Tianwen Qian, Mohamed Elhoseiny, Yuqian Fu

    Abstract: Streaming video understanding requires models to continuously retain useful visual evidence before future questions are known. Existing approaches primarily manage the growing visual context according to token importance, temporal redundancy, or segment-level relevance, but rarely organize evidence around objects that persist and evolve over time. Thus, in this paper, we introduce ObjectStream, a… ▽ More

    Submitted 1 August, 2026; v1 submitted 30 July, 2026; originally announced July 2026.

    Comments: https://github.com/DMK041218/ObjectStream

  6. arXiv:2607.26899  [pdf

    cs.HC cs.AI cs.CY

    Human diversity fuels collective creativity that large language models cannot simulate or sustain

    Authors: Mengchen Dong, Hiromu Yakura

    Abstract: Diverse human groups produce diverse ideas, the raw material of innovation. Generative AI challenges this engine twice over: everyday AI assistance may homogenize what diverse people create, and AI-simulated diversity may replace the people altogether. We tested both challenges in a preregistered creative metaphor experiment with native (L1) and non-native (L2) English writers, who wrote without A… ▽ More

    Submitted 29 July, 2026; originally announced July 2026.

  7. arXiv:2607.24957  [pdf, ps, other

    cs.CV

    PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

    Authors: Zichao Lin, Yifeng Xie, Bowen Qu, Haiming Wang, Jia Li, Haoning Wu, Yuhao Dong, Zuhao Yang, Jinguo Zhu, Haoyu Lu, Zijia Zhao, Tongtian Yue, Zhangyang Qi, Junwei Yang, Mengfan Dong, Peizhou Cao, Chenzhuang Du, Zaida Zhou, Haotian Yao, Hao Yang, Hongcheng Gao, Lin Sui, Weihong Li, Xinxing Zu, Jia Chen , et al. (8 additional authors not shown)

    Abstract: We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks only cover narrow, fragmented domains shaped by heu… ▽ More

    Submitted 27 July, 2026; originally announced July 2026.

  8. arXiv:2607.24653  [pdf, ps, other

    cs.CL cs.LG

    Kimi K3: Open Frontier Intelligence

    Authors: Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, M. C., Jianfeng Cai, Xinyuan Cai, Peizhou Cao, Yuxuan Cao, Ziwei Chai, Y. Charles, H. S. Che, Guanduo Chen, Guangyu Chen, Guanzheng Chen, Huarong Chen, Jia Chen, Jianlong Chen, Jun Chen, Kexin Chen, Peng Chen, Ruijue Chen, Wentao Chen, Xin Chen, Yang Chen , et al. (377 additional authors not shown)

    Abstract: We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. Together with Stable LatentMoE, which effectively activates 16 of 896 routed experts per token… ▽ More

    Submitted 7 August, 2026; v1 submitted 27 July, 2026; originally announced July 2026.

    Comments: K3 tech report

  9. arXiv:2607.21819  [pdf, ps, other

    cs.HC eess.SY

    Adaptive Driving Style for SAE Level-2 Driving Automation: Minimizing Preference Mismatch

    Authors: Kumar Akash, Zhaobo Zheng, Teruhisa Misu, Vidya Krishnamoorthy, Mia Dong, Yuni Lee, Gaojian Huang

    Abstract: Driving style is a key factor in the comfort and acceptance of automated vehicle (AV) features. In SAE Level-2 automation, where the driver must supervise the system and remain ready to intervene, mismatches between the automation's driving style and the driver's preference can reduce trust and trigger takeovers. This paper proposes an adaptive driving-style control framework that minimizes such p… ▽ More

    Submitted 23 July, 2026; originally announced July 2026.

    Comments: Published in the 2026 American Control Conference (ACC), May 26-29, 2026, New Orleans, Louisiana, USA

  10. arXiv:2607.20489  [pdf, ps, other

    cs.AI cs.DB

    EvoSQL: Memory-Augmented Critic-Generator Co-Evolution for Text-to-SQL

    Authors: Jiawei Zhou, Jianwei Wang, Chenyu Zhou, Chaojian Shi, Ming Dong, Kai Wang

    Abstract: Text-to-SQL has advanced rapidly with large language models, but complex database queries still require reasoning beyond one-shot generation, including multi-step decomposition, execution-based diagnosis, and targeted correction. We present EvoSQL, a co-evolution framework that formulates SQL synthesis as an iterative interaction between a generator and a critic. EvoSQL maintains a contextualized… ▽ More

    Submitted 4 June, 2026; originally announced July 2026.

  11. arXiv:2606.29451  [pdf, ps, other

    cs.CV cs.CR

    The Platonic Defense: Backdoor Defense for Self-Supervised Encoders in the Era of Large Scale Pre-training

    Authors: Tuo Chen, Minjing Dong, Benlei Cui, Jian Liu, Jie Gui

    Abstract: Self-supervised learning (SSL) pretrained models have become a dominant paradigm for visual representation learning, but they are vulnerable to backdoor attacks. Existing defenses struggle to defend against such attacks in a fully black-box setting because they often require access to labels, attack patterns, or training data. To tackle this issue, we propose a new attack-agnostic, model-agnostic,… ▽ More

    Submitted 28 June, 2026; originally announced June 2026.

  12. arXiv:2606.26273  [pdf, ps, other

    cs.LG

    Equivariance and Augmentation for Bayesian Neural Networks

    Authors: Miaowen Dong, Axel Flinth, Jan E. Gerken

    Abstract: Symmetries are important for many deep learning tasks, ranging from applications in the sciences to medical imaging. However, there is an ongoing debate about whether to impose symmetry constraints on the neural network architecture (yielding equivariant neural networks) or learn them from augmented training data. Although equivariant networks are well-studied theoretically, much less is known abo… ▽ More

    Submitted 24 June, 2026; originally announced June 2026.

  13. arXiv:2606.20683  [pdf, ps, other

    cs.AI cs.CL

    From Question Answering to Task Completion: A Survey on Agent System and Harness Design

    Authors: Jianyuan Guo, Zhiwei Hao, Chengcheng Wang, Cheng Fan, Tingzhang Luo, Hongguang Li, Ying Gao, Hefei Mei, Jiankun Peng, Rongjian Xu, Minjing Dong, Han Wu, Mengyu Zheng, Kai Han, Shiqi Wang, Chang Xu, Yunhe Wang

    Abstract: LLM-based agents mark a shift from passive question answering to active task completion: they perceive environments, invoke tools, maintain state, and act over extended horizons. As agent systems have evolved from prompt engineering to workflows and context engineering, harness engineering, and agent-native training with co-evolution, a central question has become increasingly important: where doe… ▽ More

    Submitted 14 June, 2026; originally announced June 2026.

  14. arXiv:2606.15384  [pdf, ps, other

    cs.DB

    CoeusBI: A Comprehensive Interactive Business Intelligence System Powered by LLMs at Baidu [Extended Version]

    Authors: Jinqing Lian, Chaofan Li, Yingxia Shao, Ming Wang, Yang Dong, Xinyi Liu, Wei Zhang, Chaoxian Gui, Tianqi Wan, Ming Dong

    Abstract: The advent of Large Language Models has catalyzed the emergence of interactive Business Intelligence (BI) systems. Although commercial BI products increasingly adopt semantic layers paired with natural language interfaces, they predominantly rely on manual configurations to define metrics and dimensions. Real-world deployments face critical challenges: (a) frequent JOIN operations degrade the accu… ▽ More

    Submitted 13 June, 2026; originally announced June 2026.

    Comments: Accepted by VLDB 2026

    ACM Class: H.2.3

  15. arXiv:2606.04990  [pdf, ps, other

    cs.CR cs.AI

    From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents

    Authors: Yiqi Wang, Jiaqi Zhang, Taotao Cai, Zirui Liu, Qingqiang Sun, Zequn Sun, Zhangkai Wu, Manqing Dong, Mingkai Zheng, Xuefei Yin, Yanming Zhu

    Abstract: Large language model (LLM)-based agents are evolving from passive text generators into autonomous systems capable of planning, tool use, retrieval, memory access, environmental interaction, and multi-agent collaboration. These capabilities expand agent autonomy, but also make agent behavior harder to verify, debug, and audit. Final-answer accuracy alone cannot explain how an output was produced, w… ▽ More

    Submitted 28 June, 2026; v1 submitted 3 June, 2026; originally announced June 2026.

  16. arXiv:2606.03406  [pdf, ps, other

    cs.CV

    SAMatcher: Co-Visibility Modeling with Segment Anything for Robust Feature Matching

    Authors: Xu Pan, Qiyuan Ma, Mingyue Dong, He Chen, Wei Ji, Xianwei Zheng

    Abstract: Reliable correspondence estimation is a fundamental problem in image processing, underpinning applications such as Structure from Motion, visual localization, and image registration. Existing learning-based methods have significantly improved local feature representations, yet most still operate at the pixel or patch level and lack explicit modeling of regions that are jointly visible across views… ▽ More

    Submitted 2 June, 2026; originally announced June 2026.

    Comments: 14 pages

  17. arXiv:2605.26761  [pdf, ps, other

    cs.CV

    Once-For-All: A Train-Once and Select-Anytime Framework for Multimodal Instruction Tuning

    Authors: Mingkang Dong, Hongyi Cai, Xiwen Lei, Jie Li, Tao Zhang, Muxin Pu

    Abstract: Multimodal instruction tuning is the de facto recipe for adapting vision language models (VLMs), yet instruction data are highly redundant, making data selection critical for training efficiency. Existing methods derive selection signals from a specific model or dataset, so whenever the target model or candidate pool changes, the criteria must be recomputed from scratch at substantial cost. To add… ▽ More

    Submitted 4 June, 2026; v1 submitted 26 May, 2026; originally announced May 2026.

    Comments: 15 pages, 6 figures. Mingkang Dong and Hongyi Cai contributed equally to this work. Muxin Pu is the corresponding author

    ACM Class: I.2.6; I.2.10; I.4.0

  18. arXiv:2605.19344  [pdf, ps, other

    cs.CL

    Retrieval-Augmented Linguistic Calibration

    Authors: Yi-Fan Yeh, Linwei Tao, Minjing Dong, Tao Huang, Jialin Yu, Philip Torr, Chang Xu

    Abstract: Linguistic cues such as "I believe" and "probably" offer an intuitive interface for communicating confidence, yet a generalisable, principled calibration framework for linguistic confidence expressions remains underexplored. In particular, co-occurring linguistic cues, contextual variation, and subjective audience interpretation pose unique challenges. We therefore model linguistic confidence as a… ▽ More

    Submitted 29 May, 2026; v1 submitted 19 May, 2026; originally announced May 2026.

  19. arXiv:2605.14950  [pdf, ps, other

    cs.CV cs.RO

    Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model

    Authors: Tao Lin, Yuxin Du, Jiting Liu, Nuobei Zhu, Yunhe Li, Yuqian Fu, Yinxinyu Chen, Hongyi Cai, Zewei Ye, Bing Cheng, Kai Ye, Yiran Mao, Yilei Zhong, MingKang Dong, Junchi Yan, Gen Li, Bo Zhao

    Abstract: Vision-Language-Action models have emerged as a promising paradigm for robotic manipulation by unifying perception, language grounding, and action generation. However, they often struggle in scenarios requiring precise spatial understanding, as current VLA models primarily rely on 2D visual representations that lack depth information and detailed spatial relationships. While recent approaches inco… ▽ More

    Submitted 14 May, 2026; originally announced May 2026.

  20. arXiv:2605.14821  [pdf, ps, other

    cs.CV

    HDRFace: Rethinking Face Restoration with High-Dimensional Representation

    Authors: Zirui Wang, Xianhui Lin, Yi Dong, Bo Wei, Gangjian Zhang, Siteng Ma, Zebiao Zheng, Xing Liu, Hong Gu, Minjing Dong

    Abstract: Face restoration under complex degradations still remains an ill-posed inverse problem due to severe information loss. Although diffusion models benefit from strong generative priors, most methods still condition only on low-quality inputs, making it difficult to recover identity-critical details under heavy degradations. In this work, we propose HDRFace, a High-Dimensional Representation conditio… ▽ More

    Submitted 14 May, 2026; originally announced May 2026.

  21. arXiv:2604.24642  [pdf, ps, other

    cs.CV

    Probing CLIP's Comprehension of 360-Degree Textual and Visual Semantics

    Authors: Hai Wang, Xiaochen Yang, Mingzhi Dong, Jing-Hao Xue

    Abstract: The dream of instantly creating rich 360-degree panoramic worlds from text is rapidly becoming a reality, yet a crucial gap exists in our ability to reliably evaluate their semantic alignment. Contrastive Language-Image Pre-training (CLIP) models, standard AI evaluators, predominantly trained on perspective image-text pairs, face an open question regarding their understanding of the unique charact… ▽ More

    Submitted 27 April, 2026; originally announced April 2026.

    Comments: Project Page: https://littlewhitesea.github.io/360Semantics.github.io/

  22. arXiv:2604.10789  [pdf, ps, other

    cs.CV

    ReplicateAnyScene: Zero-Shot Video-to-3D Composition via Textual-Visual-Spatial Alignment

    Authors: Mingyu Dong, Chong Xia, Mingyuan Jia, Weichen Lyu, Long Xu, Zheng Zhu, Yueqi Duan

    Abstract: Humans exhibit an innate capacity to rapidly perceive and segment objects from video observations, and even mentally assemble them into structured 3D scenes. Replicating such capability, termed compositional 3D reconstruction, is pivotal for the advancement of Spatial Intelligence and Embodied AI. However, existing methods struggle to achieve practical deployment due to the insufficient integratio… ▽ More

    Submitted 12 April, 2026; originally announced April 2026.

    Comments: Project Page: https://xiac20.github.io/ReplicateAnyScene/

  23. arXiv:2604.06912  [pdf, ps, other

    cs.CV cs.AI

    Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models

    Authors: Yuheng Shi, Xiaohuan Pei, Linfeng Wen, Minjing Dong, Chang Xu

    Abstract: MLLMs require high-resolution visual inputs for fine-grained tasks like document understanding and dense scene perception. However, current global resolution scaling paradigms indiscriminately flood the quadratic self-attention mechanism with visually redundant tokens, severely bottlenecking inference throughput while ignoring spatial sparsity and query intent. To overcome this, we propose Q-Zoom,… ▽ More

    Submitted 8 April, 2026; originally announced April 2026.

    Comments: 16 pages, 9 figures

  24. arXiv:2604.02623  [pdf, ps, other

    cs.CR cs.AI

    Poison Once, Exploit Forever: Environment-Injected Memory Poisoning Attacks on Web Agents

    Authors: Wei Zou, Mingwen Dong, Miguel Romero Calvo, Shuaichen Chang, Jiang Guo, Dongkyu Lee, Xing Niu, Xiaofei Ma, Yanjun Qi, Jiarong Jiang

    Abstract: Memory makes LLM-based web agents personalized, powerful, yet exploitable. By storing past interactions to personalize future tasks, agents inadvertently create a persistent attack surface that spans websites and sessions. While existing security research on memory assumes attackers can directly inject into memory storage or exploit shared memory across users, we present a more realistic threat mo… ▽ More

    Submitted 7 April, 2026; v1 submitted 2 April, 2026; originally announced April 2026.

  25. arXiv:2603.22879  [pdf, ps, other

    cs.LG cs.AI

    Confidence Calibration under Ambiguous Ground Truth

    Authors: Linwei Tao, Haoyang Luo, Minjing Dong, Chang Xu

    Abstract: Confidence calibration assumes a unique ground-truth label per input, yet this assumption fails wherever annotators genuinely disagree. Post-hoc calibrators fitted on majority-voted labels, the standard single-label targets used in practice, can appear well-calibrated under conventional evaluation yet remain substantially miscalibrated against the underlying annotator distribution. We show that th… ▽ More

    Submitted 24 March, 2026; originally announced March 2026.

  26. Bridging neuroscience and AI: adaptive, culturally sensitive technologies transforming aphasia rehabilitation

    Authors: Andreea I. Niculescu, Jochen Ehnes, Minghui Dong

    Abstract: Aphasia, a language impairment primarily resulting from stroke or brain injury, profoundly disrupts communication and everyday functioning. Despite advances in speech therapy, barriers such as limited therapist availability and the scarcity of personalized, culturally relevant tools continue to hinder optimal rehabilitation outcomes. This paper reviews recent developments in neurocognitive researc… ▽ More

    Submitted 22 March, 2026; originally announced March 2026.

    Comments: 12 pages, 2 figures, Proceedings of the 20th International Conference on linguistic resources and tools for natural language processing (ConsILR 2025)

    ACM Class: H.5.2; J.3

  27. arXiv:2603.15542  [pdf, ps, other

    cs.CY cs.AI

    InterveneBench: Benchmarking LLMs for Intervention Reasoning and Causal Study Design in Real Social Systems

    Authors: Shaojie Shi, Zhengyu Shi, Lingran Zheng, Xinyu Su, Anna Xie, Bohao Lv, Rui Xu, Zijian Chen, Zhichao Chen, Guolei Liu, Naifu Zhang, Mingjian Dong, Zhuo Quan, Bohao Chen, Teqi Hao, Yuan Qi, Yinghui Xu, Libo Wu

    Abstract: Causal inference in social science relies on end-to-end, intervention-centered research-design reasoning grounded in real-world policy interventions, but current benchmarks fail to evaluate this capability of large language models (LLMs). We present InterveneBench, a benchmark designed to assess such reasoning in realistic social settings. Each instance in InterveneBench is derived from an empiric… ▽ More

    Submitted 16 March, 2026; originally announced March 2026.

    Comments: 35pages,3 figures

  28. arXiv:2603.10444  [pdf, ps, other

    cs.LG cs.AI

    The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training

    Authors: Hengjie Cao, Zhendong Huang, Mengyi Chen, Yifeng Yang, Fang Dong, Anrui Chen, Ruijun Huang, Xin Zhang, Mingzhi Dong, Yujiang Wang, Jinlong Hou, Qin Lv, Robert P. Dick, Yuan Cheng, Tun Lu, Fan Yang, Yixuan Chen, Li Shang

    Abstract: FP4 training promises substantial memory and compute savings for large language models, but remains fragile because blockwise quantization is dictated by extreme activation magnitudes, which inflate dynamic range and compress long-tail signals. We identify a counterintuitive source of this failure: dominant activation outliers are not merely arbitrary sparse events, but are largely induced by a co… ▽ More

    Submitted 12 June, 2026; v1 submitted 11 March, 2026; originally announced March 2026.

  29. arXiv:2603.07992  [pdf, ps, other

    cs.DC

    SI-ChainFL: Shapley-Incentivized Secure Federated Learning for High-Speed Rail Data Sharing

    Authors: Mingjie Zhao, Cheng Dai, Fei Chen, Xin Chen, Kaoru Ota, Mianxiong Dong, Bing Guo

    Abstract: In high-speed rail (HSR) systems, federated learning (FL) enables cross-departmental flow prediction without sharing raw data. However, existing schemes suffer from two key limitations: (1) insufficient incentives, leading to free-riding and model poisoning; and (2) centralized aggregation, which introduces a single point of failure. We propose a secure and efficient framework SI-ChainFL that addr… ▽ More

    Submitted 9 March, 2026; originally announced March 2026.

    Comments: 17 pages, 19 figures

    MSC Class: 68M14 ACM Class: I.2.11

  30. arXiv:2603.01195  [pdf, ps, other

    cs.CV cs.AI

    VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning

    Authors: Mingkang Dong, Hongyi Cai, Jie Li, Sifan Zhou, Bin Ren, Kunyu Peng, Yuqian Fu

    Abstract: The effectiveness of multimodal instruction tuning depends not only on dataset scale, but critically on whether training samples genuinely require visual reasoning. However, existing instruction datasets often contain a substantial portion of visually redundant samples (solvable from text alone), as well as multimodally misaligned supervision that can degrade learning. To address this, we propose… ▽ More

    Submitted 25 June, 2026; v1 submitted 1 March, 2026; originally announced March 2026.

    Comments: Accepted at ECCV 2026. Project Page: https://dmk041218.github.io/VisNec/

    ACM Class: I.2.10; I.2.6; I.2.7

  31. arXiv:2602.19418  [pdf, ps, other

    cs.CV

    PA-Attack: Guiding Gray-Box Attacks on LVLM Vision Encoders with Prototypes and Attention

    Authors: Hefei Mei, Zirui Wang, Chang Xu, Jianyuan Guo, Minjing Dong

    Abstract: Large Vision-Language Models (LVLMs) are foundational to modern multimodal applications, yet their susceptibility to adversarial attacks remains a critical concern. Prior white-box attacks rarely generalize across tasks, and black-box methods depend on expensive transfer, which limits efficiency. The vision encoder, standardized and often shared across LVLMs, provides a stable gray-box pivot with… ▽ More

    Submitted 22 February, 2026; originally announced February 2026.

  32. arXiv:2602.12587  [pdf, ps, other

    cs.LG

    Multi-Head Attention as a Source of Catastrophic Forgetting in MoE Transformers

    Authors: Anrui Chen, Ruijun Huang, Xin Zhang, Fang Dong, Hengjie Cao, Zhendong Huang, Yifeng Yang, Mengyi Chen, Jixian Zhou, Mingzhi Dong, Yujiang Wang, Jinlong Hou, Qin Lv, Robert P. Dick, Yuan Cheng, Tun Lu, Fan Yang, Li Shang

    Abstract: Mixture-of-Experts (MoE) architectures are often considered a natural fit for continual learning because sparse routing should localize updates and reduce interference, yet MoE Transformers still forget substantially even with sparse, well-balanced expert utilization. We attribute this gap to a pre-routing bottleneck: multi-head attention concatenates head-specific signals into a single post-atten… ▽ More

    Submitted 12 February, 2026; originally announced February 2026.

  33. arXiv:2602.12556  [pdf, ps, other

    cs.LG cs.AI

    SD-MoE: Spectral Decomposition for Effective Expert Specialization

    Authors: Ruijun Huang, Fang Dong, Xin Zhang, Hengjie Cao, Zhendong Huang, Anrui Chen, Jixian Zhou, Mengyi Chen, Yifeng Yang, Mingzhi Dong, Yujiang Wang, Jinlong Hou, Qin Lv, Robert P. Dick, Yuan Cheng, Fan Yang, Tun Lu, Chun Zhang, Li Shang

    Abstract: Mixture-of-Experts (MoE) architectures scale Large Language Models via expert specialization induced by conditional computation. In practice, however, expert specialization often fails: some experts become functionally similar, while others functioning as de facto shared experts, limiting the effective capacity and model performance. In this work, we analysis from a spectral perspective on paramet… ▽ More

    Submitted 12 February, 2026; originally announced February 2026.

  34. arXiv:2602.11185  [pdf, ps, other

    cs.LG cs.AI

    Spectra: Rethinking Optimizers for LLMs Under Spectral Anisotropy

    Authors: Zhendong Huang, Hengjie Cao, Fang Dong, Ruijun Huang, Mengyi Chen, Yifeng Yang, Xin Zhang, Anrui Chen, Mingzhi Dong, Yujiang Wang, Jinlong Hou, Qin Lv, Robert P. Dick, Yuan Cheng, Fan Yang, Tun Lu, Li Shang

    Abstract: Gradient signals in LLM training are highly anisotropic: recurrent linguistic structure concentrates energy into a small set of dominant spectral directions, while context specific information resides in a long tail. We show that this spike tail separation persists throughout training, with the spike occupying only about 1.5% of directions yet dominating optimizer statistics. This dominance suppre… ▽ More

    Submitted 30 January, 2026; originally announced February 2026.

  35. arXiv:2602.10504  [pdf, ps, other

    cs.CL

    On the Robustness of Knowledge Editing for Detoxification

    Authors: Ming Dong, Shiyi Tang, Ziyan Peng, Guanyi Chen, Tingting He

    Abstract: Knowledge-Editing-based (KE-based) detoxification has emerged as a promising approach for mitigating harmful behaviours in Large Language Models. Existing evaluations, however, largely rely on automatic toxicity classifiers, implicitly assuming that reduced toxicity scores reflect genuine behavioural suppression. In this work, we propose a robustness-oriented evaluation framework for KE-based deto… ▽ More

    Submitted 10 February, 2026; originally announced February 2026.

  36. arXiv:2602.08762  [pdf, ps, other

    cs.LG cs.CR

    HoGS: Homophily-Oriented Graph Synthesis for Local Differentially Private GNN Training

    Authors: Wen Xu, Zhetao Li, Yong Xiao, Pengpeng Qiao, Mianxiong Dong, Kaoru Ota

    Abstract: Graph neural networks (GNNs) have demonstrated remarkable performance in various graph-based machine learning tasks by effectively modeling high-order interactions between nodes. However, training GNNs without protection may leak sensitive personal information in graph data, including links and node features. Local differential privacy (LDP) is an advanced technique for protecting data privacy in… ▽ More

    Submitted 9 February, 2026; originally announced February 2026.

  37. arXiv:2602.02503  [pdf, ps, other

    eess.SP cs.AI cs.IT

    Joint single-shot ToA and DoA estimation for VAA-based BLE ranging with phase ambiguity: A deep learning-based approach

    Authors: Jincheng Xie, Yili Deng, Jiguang He, Pengyu Wang, Miaomiao Dong, Rui Tang, Zhongyi Huang

    Abstract: Conventional direction-of-arrival (DoA) estimation methods rely on multi-antenna arrays, which are costly to implement on size-constrained Bluetooth Low Energy (BLE) devices. Virtual antenna array (VAA) techniques enable DoA estimation with a single antenna, making angle estimation feasible on such devices. However, BLE only provides a single-shot two-way channel frequency response (CFR) with a bi… ▽ More

    Submitted 21 January, 2026; originally announced February 2026.

  38. arXiv:2602.02276  [pdf, ps, other

    cs.CL cs.AI cs.LG

    Kimi K2.5: Visual Agentic Intelligence

    Authors: Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, S. H. Cai, Yuan Cao, Ziwei Chai, Y. Charles, H. S. Che, Cheng Chen, Guanduo Chen, Huarong Chen, Jia Chen, Jianlong Chen, Jun Chen, Kefan Chen, Liang Chen, Ruijue Chen, Xinhao Chen, Yanru Chen, Yanxu Chen, Yicun Chen, Yimin Chen, Yingjiang Chen, Yuankun Chen , et al. (312 additional authors not shown)

    Abstract: We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint optimization of text and vision so that two modalities enhance each other. This includes a series of techniques such as joint text-vision pre-training, zero-vision SFT, and joint text-vision reinforcement learning. Building on this multimodal foundation, K2.5… ▽ More

    Submitted 7 August, 2026; v1 submitted 2 February, 2026; originally announced February 2026.

    Comments: Kimi K2.5 tech report

  39. arXiv:2602.01308  [pdf, ps, other

    cs.LG cs.AI

    Dispelling the Curse of Singularities in Neural Network Optimizations

    Authors: Hengjie Cao, Mengyi Chen, Yifeng Yang, Fang Dong, Ruijun Huang, Anrui Chen, Jixian Zhou, Mingzhi Dong, Yujiang Wang, Dongsheng Li, Wenyi Fang, Yuanyi Lin, Fan Wu, Li Shang

    Abstract: This work investigates the optimization instability of deep neural networks from a less-explored yet insightful perspective: the emergence and amplification of singularities in the parametric space. Our analysis reveals that parametric singularities inevitably grow with gradient updates and further intensify alignment with representations, leading to increased singularities in the representation s… ▽ More

    Submitted 12 February, 2026; v1 submitted 1 February, 2026; originally announced February 2026.

  40. arXiv:2601.22667  [pdf, ps, other

    cs.SE cs.AI

    From Horizontal Layering to Vertical Integration: A Comparative Study of the AI-Driven Software Development Paradigm

    Authors: Chi Zhang, Zehan Li, Ziqian Zhong, Haibing Ma, Dan Xiao, Chen Lin, Ming Dong

    Abstract: This paper examines the organizational implications of Generative AI adoption in software engineering through a multiple-case comparative study. We contrast two development environments: a traditional enterprise (brownfield) and an AI-native startup (greenfield). Our analysis reveals that transitioning from Horizontal Layering (functional specialization) to Vertical Integration (end-to-end ownersh… ▽ More

    Submitted 30 January, 2026; originally announced January 2026.

  41. arXiv:2601.20911  [pdf, ps, other

    cs.CV cs.AI

    Non-Markov Multi-Round Conversational Image Generation with History-Conditioned MLLMs

    Authors: Haochen Zhang, Animesh Sinha, Felix Juefei-Xu, Haoyu Ma, Kunpeng Li, Zhipeng Fan, Meng Dong, Xiaoliang Dai, Tingbo Hou, Peizhao Zhang, Zecheng He

    Abstract: Conversational image generation requires a model to follow user instructions across multiple rounds of interaction, grounded in interleaved text and images that accumulate as chat history. While recent multimodal large language models (MLLMs) can generate and edit images, most existing multi-turn benchmarks and training recipes are effectively Markov: the next output depends primarily on the most… ▽ More

    Submitted 28 January, 2026; originally announced January 2026.

    Comments: 19 pages, 19 figures, plan for TIP

  42. arXiv:2601.17900  [pdf, ps, other

    cs.CV

    Revisiting 3D Reconstruction Kernels as Low-Pass Filters

    Authors: Shengjun Zhang, Min Chen, Yibo Wei, Mingyu Dong, Yueqi Duan

    Abstract: 3D reconstruction is to recover 3D signals from the sampled discrete 2D pixels, with the goal to converge continuous 3D spaces. In this paper, we revisit 3D reconstruction from the perspective of signal processing, identifying the periodic spectral extension induced by discrete sampling as the fundamental challenge. Previous 3D reconstruction kernels, such as Gaussians, Exponential functions, and… ▽ More

    Submitted 25 January, 2026; originally announced January 2026.

    Comments: 14 pages, 5 figures

  43. arXiv:2601.14129  [pdf, ps, other

    cs.OS cs.DC

    "Range as a Key" is the Key! Fast and Compact Cloud Block Store Index with RASK

    Authors: Haoru Zhao, Mingkai Dong, Erci Xu, Zhongyu Wang, Haibo Chen

    Abstract: In cloud block store, indexing is on the critical path of I/O operations and typically resides in memory. With the scaling of users and the emergence of denser storage media, the index has become a primary memory consumer, causing memory strain. Our extensive analysis of production traces reveals that write requests exhibit a strong tendency to target continuous block ranges in cloud storage syste… ▽ More

    Submitted 20 January, 2026; originally announced January 2026.

  44. arXiv:2601.12719  [pdf, ps, other

    cs.CV

    S2DiT: Sandwich Diffusion Transformer for Mobile Streaming Video Generation

    Authors: Lin Zhao, Yushu Wu, Aleksei Lebedev, Dishani Lahiri, Meng Dong, Arpit Sahni, Michael Vasilkovsky, Hao Chen, Ju Hu, Aliaksandr Siarohin, Sergey Tulyakov, Yanzhi Wang, Anil Kag, Yanyu Li

    Abstract: Diffusion Transformers (DiTs) have recently improved video generation quality. However, their heavy computational cost makes real-time or on-device generation infeasible. In this work, we introduce S2DiT, a Streaming Sandwich Diffusion Transformer designed for efficient, high-fidelity, and streaming video generation on mobile hardware. S2DiT generates more tokens but maintains efficiency with nove… ▽ More

    Submitted 6 March, 2026; v1 submitted 18 January, 2026; originally announced January 2026.

    Comments: https://snap-research.github.io/S2DiT/

  45. arXiv:2601.03267  [pdf, ps, other

    cs.CL cs.AI

    OpenAI GPT-5 System Card

    Authors: Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, Aidan McLaughlin, Aiden Low, AJ Ostrow, Akhila Ananthram, Akshay Nathan, Alan Luo, Alec Helyar, Aleksander Madry, Aleksandr Efremov, Aleksandra Spyra, Alex Baker-Whitcomb, Alex Beutel, Alex Karpenko, Alex Makelov, Alex Neitz, Alex Wei, Alexandra Barr, Alexandre Kirchmeyer, Alexey Ivanov , et al. (461 additional authors not shown)

    Abstract: This is the system card published alongside the OpenAI GPT-5 launch, August 2025. GPT-5 is a unified system with a smart and fast model that answers most questions, a deeper reasoning model for harder problems, and a real-time router that quickly decides which model to use based on conversation type, complexity, tool needs, and explicit intent (for example, if you say 'think hard about this' in… ▽ More

    Submitted 1 May, 2026; v1 submitted 19 December, 2025; originally announced January 2026.

    Comments: May 2026: Added monitorability evals and authors

  46. arXiv:2512.23952  [pdf, ps, other

    cs.DC

    Squeezing Edge Performance: A Sensitivity-Aware Container Management for Heterogeneous Tasks

    Authors: Yongmin Zhang, Pengyu Huang, Mingyi Dong, Jing Yao

    Abstract: Edge computing enables latency-critical applications to process data close to end devices, yet task heterogeneity and limited resources pose significant challenges to efficient orchestration. This paper presents a measurement-driven, container-based resource management framework for intra-node optimization on a single edge server hosting multiple heterogeneous applications. Extensive profiling exp… ▽ More

    Submitted 29 December, 2025; originally announced December 2025.

  47. arXiv:2512.23463  [pdf, ps, other

    cs.CV

    Deterministic Image-to-Image Translation via Denoising Brownian Bridge Models with Dual Approximators

    Authors: Bohan Xiao, Peiyong Wang, Qisheng He, Ming Dong

    Abstract: Image-to-Image (I2I) translation involves converting an image from one domain to another. Deterministic I2I translation, such as in image super-resolution, extends this concept by guaranteeing that each input generates a consistent and predictable output, closely matching the ground truth (GT) with high fidelity. In this paper, we propose a denoising Brownian bridge model with dual approximators (… ▽ More

    Submitted 29 December, 2025; originally announced December 2025.

    Comments: Minor correction to a reference entry

  48. arXiv:2512.22766  [pdf

    eess.IV cs.CV nucl-ex

    SwinCCIR: An end-to-end deep network for Compton camera imaging reconstruction

    Authors: Minghao Dong, Xinyang Luo, Xujian Ouyang, Yongshun Xiao

    Abstract: Compton cameras (CCs) are a kind of gamma cameras which are designed to determine the directions of incident gammas based on the Compton scatter. However, the reconstruction of CCs face problems of severe artifacts and deformation due to the fundamental reconstruction principle of back-projection of Compton cones. Besides, a part of systematic errors originated from the performance of devices are… ▽ More

    Submitted 27 December, 2025; originally announced December 2025.

    Comments: 10 pages, 7 figures

  49. arXiv:2512.15037  [pdf, ps, other

    cs.CR

    RELIC-GNN: Efficient State Registers Identification with Graph Neural Network for Reverse Engineering

    Authors: Weitao Pan, Meng Dong, Zhiliang Qiu, Jianlei Yang, Zhixiong Di, Yiming Gao

    Abstract: Reverse engineering of gate-level netlist is critical for Hardware Trojans detection and Design Piracy counteracting. The primary task of gate-level reverse engineering is to separate the control and data signals from the netlist, which is mainly realized by identifying state registers with topological comparison.However, these methods become inefficient for large scale netlist. In this work, we p… ▽ More

    Submitted 16 December, 2025; originally announced December 2025.

  50. Truthful and Trustworthy IoT AI Agents via Immediate-Penalty Enforcement under Approximate VCG Mechanisms

    Authors: Xun Shao, Ryuuto Shimizu, Zhi Liu, Kaoru Ota, Mianxiong Dong

    Abstract: The deployment of autonomous AI agents in Internet of Things (IoT) energy systems requires decision-making mechanisms that remain robust, efficient, and trustworthy under real-time constraints and imperfect monitoring. While reinforcement learning enables adaptive prosumer behaviors, ensuring economic consistency and preventing strategic manipulation remain open challenges, particularly when sensi… ▽ More

    Submitted 1 December, 2025; v1 submitted 29 November, 2025; originally announced December 2025.

    Journal ref: IEEE Internet of Things Journal, 2026