Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 161 results for author: Duan, S

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.20942  [pdf, ps, other

    cs.CV math-ph

    LHMCF-Net: A Learned Hyperbolic Mean Curvature Flow Network for Medical Images Segmentation

    Authors: Shuangshuang Duan, Chunlei He, Shoujun Huang, Dexing Kong

    Abstract: Motivated by the classical Chan-Vese model and the ability of deep priors to capture complex spatial structures, we develop a segmentation model that leverages learned hyperbolic mean curvature flow (LHMCF) as a mathematical foundation for integrating feature space data fidelity and deep structural priors within a unified high-dimensional framework. The proposed LHMCF model is governed by a second… ▽ More

    Submitted 21 August, 2026; originally announced August 2026.

  2. arXiv:2608.18780  [pdf, ps, other

    cs.LG

    A Real-Time Tsetlin Machine-based Non-intrusive Load Monitoring System on MCUs

    Authors: Tianhang Tan, Han Wu, Tousif Rahman, Shengyu Duan, Alex Yakovlev, Rishad Shafik

    Abstract: Non-Intrusive Load Monitoring (NILM) systems estimate individual appliance energy consumption from a single aggregate meter, without requiring separate sensors for each device. By installing a single meter that measures a building's total electricity consumption, NILM algorithms can determine the active status of each appliance. However, traditional NILM systems use computationally intensive optim… ▽ More

    Submitted 19 August, 2026; originally announced August 2026.

    Comments: Accepted by International Symposium on the Tsetlin Machine (ISTM 2026)

  3. arXiv:2608.15875  [pdf, ps, other

    cs.RO

    GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture

    Authors: GigaBrain Team, Angen Ye, Axiang Sun, Can Jin, Chenxi Cheng, Chong Shi, Dengke Shang, Dingqian Zhang, Guan Huang, Guangqiang Wang, Guangqing Ding, Guo Li, Hangcong Li, Hengyu Zhong, Hongtao Lu, Jianbo Qin, Jiming Mao, Jing Zhu, Jindi Lv, Jingzhi Cui, Junjie Xie, Junyi Bao, Kai Liu, Lei Yuan, Limin Long , et al. (34 additional authors not shown)

    Abstract: Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit from more effective architectural design, scale to substantially larger and more heterogeneous data regimes, and achieve broader generalizatio… ▽ More

    Submitted 16 August, 2026; originally announced August 2026.

    Comments: https://gigaai.cc/blog/gigabrain07

  4. arXiv:2608.01924   

    cs.CE cs.AI

    TransNRank: Towards Accurate Neoantigen Ranking with Transformer

    Authors: Zhiyin An, Yuenan Hou, Shumeng Duan, Yiming Zhou, Yuanting Zheng, Leming Shi

    Abstract: Personalized neoantigen prediction is challenging due to the scarcity of positive samples, the noise of the experimental data, the severe class imbalance trait and the complex of immunogenicity features. Prior arts, such as linear regression and XGBoost fail to model long-range dependencies and contextual relationships within peptide features, therefore the performance of neoantigen positive recal… ▽ More

    Submitted 7 August, 2026; v1 submitted 3 August, 2026; originally announced August 2026.

    Comments: The final authorship has not been determined and this version has not been polished

  5. arXiv:2607.27843  [pdf, ps, other

    cs.CV

    VCP-DCN: Beyond Visual Concealed Property via Depth Collaborative Network for Camouflaged Object Detection

    Authors: Songsong Duan, Xi Yang, Nannan Wang

    Abstract: Camouflaged Object Detection (COD) aims to identify and segment camouflaged objects in complex environments, which are often concealed because their color and texture are similar to the background. Several existing COD methods introduce depth maps to boost detection performance via learning complementary RGB-D features, ignoring modality-specific characteristics of concealed objects in the depth d… ▽ More

    Submitted 30 July, 2026; originally announced July 2026.

    Comments: Accepted by ECCV 2026

  6. arXiv:2607.20116  [pdf, ps, other

    cs.CV

    RIM: A Retrieval-In-Matching Framework for Cross-Domain Global Visual Localization of UAVs

    Authors: Xin Li, Siyuan Duan, Shang Wang, Zhimin Mao, Bingliang Hu, Geng Zhang

    Abstract: Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, acquisition-time and imaging-platform differences between UAV and reference imagery induce substantial cross-domain appearance and viewpoint shifts, challenging robust six-degree-of-freedom (6-DoF) pose estimation. We address these shifts by sampling UAV-vi… ▽ More

    Submitted 22 July, 2026; originally announced July 2026.

    Comments: 56 pages, 10 figures, and 18 tables. Supplementary material is included

  7. arXiv:2606.26942  [pdf, ps, other

    cs.CV

    TraMP-LLaMA: Generative Interpretability with Decoupled Instruction Tuning for Facial Expression Quality Assessment

    Authors: Shuchao Duan, Alan Whone, Hossein Rahmani, Jun Liu, Majid Mirmehdi

    Abstract: Existing facial expression quality assessment (FEQA) methods typically produce only a severity score, without explicitly communicating the observable facial motion evidence that supports the prediction. This limits interpretability and makes it difficult to inspect the basis of model outputs in Parkinson's disease assessment. To address this gap, we propose TraMP-LLaMA, a unified multimodal framew… ▽ More

    Submitted 25 June, 2026; originally announced June 2026.

  8. arXiv:2606.15598  [pdf, ps, other

    cs.AI

    Integrating Reasoning and Generalization in Text-to-SQL via Self-Enhanced Fine-Tuning

    Authors: Feng Lyu, Jinfeng Cen, Sijing Duan, Hao Wu, Shucheng Li, Weixu Zhang, Haolun Wu

    Abstract: Text-to-SQL aims to translate natural language questions into executable SQL queries over structured databases, enabling non-expert users to access data intuitively. While recent advances in large language models (LLMs) have shown promise in this task, existing LLM-based approaches often struggle to strike a balance between strong reasoning capabilities and robust generalization. To address these… ▽ More

    Submitted 14 June, 2026; originally announced June 2026.

    Comments: 14 pages, 13 figures, 7 tables

  9. arXiv:2606.13171  [pdf, ps, other

    cs.CL cs.AI

    NTS-CoT: Mitigating Hallucinations in LLM-based News Timeline Summarization with Chain-of-Thought Reasoning

    Authors: Feng Lyu, Huiqin Yan, Sijing Duan, Hao Wu, Shuang Gu, Xue Qiao, Weixu Zhang, Haolun Wu

    Abstract: The rapid updates of online news make tracking event developments challenging, highlighting the need for timeline summarization (TLS). Hallucinations, where LLM-generated content deviates from source news, still remain a critical issue in LLM-based TLS and are not well studied in existing works. To bridge this gap, we identify two primary types of hallucinations: unfaithful content during news sum… ▽ More

    Submitted 11 June, 2026; originally announced June 2026.

  10. arXiv:2605.24475  [pdf, ps, other

    cs.CV cs.AI cs.MM

    Robust Fuzzy Multi-view Learning under View Conflict

    Authors: Siyuan Duan, Yuan Sun, Dezhong Peng, Yingke Chen, Xi Peng, Peng Hu

    Abstract: Trusted multi-view classification aims to deliver reliable fusion for accurate predictions and has recently attracted substantial attention in both academia and industry. However, existing TMVC methods typically assume strict alignment across different views during both training and testing phases, which is often impractical in real-world scenarios. This limitation motivates us to revisit TMVC and… ▽ More

    Submitted 23 May, 2026; originally announced May 2026.

  11. arXiv:2605.23282  [pdf, ps, other

    eess.IV cs.CV cs.LG

    Discontinuous Galerkin Neural Operator for Pathology Defocus Deblurring

    Authors: Shaoqing Duan, Haofei Song, Xintian Mao, Qingli Li, Yan Wang

    Abstract: Defocus deblurring in pathological microscopy remains challenging due to the spatially varying and locally discontinuous nature of optical blur induced by a position-dependent integral imaging process. Existing deep learning methods, constrained by shift-invariance assumptions and limited interpretability, are not well suited to such heterogeneous blur patterns. Neural operators provide a prin… ▽ More

    Submitted 22 May, 2026; originally announced May 2026.

    Comments: 17 pages, 9 figures. Accepted by ICML 2026

  12. arXiv:2605.09888  [pdf, ps, other

    cs.NI

    Mixed-Criticality Flow Scheduling with Low Delay and Limited Bandwidth in TSN

    Authors: Wenyan Yan, Sijing Duan, Dongsheng Wei

    Abstract: Time-Sensitive Networking (TSN) is a promising Ethernet protocol with time determinism, widely used in time-critical systems such as industrial automation, automotive networks, and avionics. By allocating dedicated time windows for time-sensitive flows, TSN enables deterministic transmission; however, as network traffic grows, multiple flows may contend for the same window, causing large delays. F… ▽ More

    Submitted 10 May, 2026; originally announced May 2026.

    Comments: 7 pages

  13. arXiv:2605.01950  [pdf, ps, other

    cs.LG cs.AI

    TRAP: Tail-aware Ranking Attack for World-Model Planning

    Authors: Siyuan Duan, Ke Zhang, Xizhao Luo

    Abstract: World models enable long-horizon planning by internally generating and evaluating imagined trajectories, making them a promising foundation for generalist agents. However, this imagination-driven decision process also introduces new security risks. Existing backdoor attacks typically aim to manipulate local features, one-step predictions, or instantaneous policy outputs. While such objectives may… ▽ More

    Submitted 3 May, 2026; originally announced May 2026.

  14. arXiv:2604.26752  [pdf, ps, other

    cs.CV

    GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

    Authors: GLM-V Team, :, Wenyi Hong, Xiaotao Gu, Ziyang Pan, Zhen Yang, Yuting Wang, Yue Wang, Yuanchang Yue, Yu Wang, Yanling Wang, Yan Wang, Xijun Liu, Wenmeng Yu, Weihan Wang, Wei Li, Shuaiqi Duan, Sheng Yang, Ruiliang Lv, Mingdao Liu, Lihang Pan, Ke Ning, Junhui Ji, Jinjiang Wang, Jing Chen , et al. (73 additional authors not shown)

    Abstract: We present GLM-5V-Turbo, a step toward native foundation models for multimodal agents. As foundation models are increasingly deployed in real environments, agentic capability depends not only on language reasoning, but also on the ability to perceive, interpret, and act over heterogeneous contexts such as images, videos, webpages, documents, GUIs. GLM-5V-Turbo is built around this objective: multi… ▽ More

    Submitted 12 May, 2026; v1 submitted 29 April, 2026; originally announced April 2026.

  15. arXiv:2604.22335  [pdf, ps, other

    cs.CL

    Context-Fidelity Boosting: Enhancing Faithful Generation through Watermark-Inspired Decoding

    Authors: Weixu Zhang, Fanghua Ye, Qiang Gao, Jian Li, Haolun Wu, Yuxing Tian, Sijing Duan, Nan Du, Xiaolong Li, Xue Liu

    Abstract: Large language models (LLMs) often produce content that contradicts or overlooks information provided in the input context, a phenomenon known as faithfulness hallucination. In this paper, we propose Context-Fidelity Boosting (CFB), a lightweight and general decoding-time framework that reduces such hallucinations by increasing the generation probability of source-supported tokens. Motivated by lo… ▽ More

    Submitted 24 April, 2026; originally announced April 2026.

    Comments: Accepted at ACL 2026

  16. arXiv:2604.15001  [pdf, ps, other

    cs.AI

    COEVO: Co-Evolutionary Framework for Joint Functional Correctness and PPA Optimization in LLM-Based RTL Generation

    Authors: Heng Ping, Peiyu Zhang, Shixuan Li, Wei Yang, Anzhe Cheng, Shukai Duan, Xiaole Zhang, Paul Bogdan

    Abstract: LLM-based RTL code generation methods increasingly target both functional correctness and PPA quality, yet existing approaches universally decouple the two objectives, optimizing PPA only after correctness is fully achieved. Whether through sequential multi-agent pipelines, evolutionary search with binary correctness gates, or hierarchical reward dependencies, partially correct but architecturally… ▽ More

    Submitted 17 April, 2026; v1 submitted 16 April, 2026; originally announced April 2026.

  17. arXiv:2604.02356  [pdf, ps, other

    cs.NI cs.LG

    MLFCIL: A Multi-Level Forgetting Mitigation Framework for Federated Class-Incremental Learning in LEO Satellites

    Authors: Heng Zhang, Xiaohong Deng, Sijing Duan, Wu Ouyang, KM Mahfujul, Yiqin Deng, Zhigang Chen

    Abstract: Low-Earth-orbit (LEO) satellite constellations are increasingly performing on-board computing. However, the continuous emergence of new classes under strict memory and communication constraints poses major challenges for collaborative training. Federated class-incremental learning (FCIL) enables distributed incremental learning without sharing raw data, but faces three LEO-specific challenges: non… ▽ More

    Submitted 14 March, 2026; originally announced April 2026.

    Comments: Submitted to IEEE Internet of Things Journal

  18. arXiv:2603.24186  [pdf, ps, other

    cs.LG cs.AR

    TsetlinWiSARD: On-Chip Training of Weightless Neural Networks using Tsetlin Automata on FPGAs

    Authors: Shengyu Duan, Marcos L. L. Sartori, Rishad Shafik, Alex Yakovlev

    Abstract: Increasing demands for adaptability, privacy, and security at the edge have persistently pushed the frontiers for a new generation of machine learning (ML) algorithms with training and inference capabilities on-chip. Weightless Neural Network (WNN) is such an algorithm that is principled on lookup table based simple neuron structures. As a result, it offers architectural benefits, such as low-late… ▽ More

    Submitted 25 March, 2026; originally announced March 2026.

    Comments: Accepted at the 63rd Design Automation Conference (DAC 2026)

  19. arXiv:2603.22343  [pdf, ps, other

    cs.LG cs.NI

    Cloud-Edge Collaborative Large Models for Robust Photovoltaic Power Forecasting

    Authors: Nan Qiao, Shuning Wang, Sijing Duan, Wenpeng Cui, Yuzhe Chen, Qingchen Yang, Xingyuan Hua, Ju Ren

    Abstract: Photovoltaic (PV) power forecasting in edge-enabled grids requires balancing forecasting accuracy, robustness under weather-driven distribution shifts, and strict latency constraints. Existing models work well under normal conditions but often struggle with rare ramp events and unexpected weather changes. Relying solely on cloud-based large models often leads to significant communication delays, w… ▽ More

    Submitted 25 March, 2026; v1 submitted 21 March, 2026; originally announced March 2026.

  20. arXiv:2603.12483  [pdf, ps, other

    cs.AI cs.LG

    Generating Expressive and Customizable Evals for Timeseries Data Analysis Agents with AgentFuel

    Authors: Aadyaa Maddi, Prakhar Naval, Deepti Mande, Shane Duan, Muckai Girish, Vyas Sekar

    Abstract: Across many domains (e.g., IoT, observability, telecommunications, cybersecurity), there is an emerging adoption of conversational data analysis agents that enable users to "talk to your data" to extract insights. Such data analysis agents operate on timeseries data models; e.g., measurements from sensors or events monitoring user clicks and actions in product analytics. We evaluate 6 popular data… ▽ More

    Submitted 12 March, 2026; originally announced March 2026.

  21. arXiv:2603.10910  [pdf, ps, other

    cs.CL

    GLM-OCR Technical Report

    Authors: Shuaiqi Duan, Yadong Xue, Weihan Wang, Zhe Su, Huan Liu, Sheng Yang, Guobing Gan, Guo Wang, Zihan Wang, Shengdong Yan, Dexin Jin, Yuxuan Zhang, Guohong Wen, Yanfeng Wang, Yutao Zhang, Xiaohan Zhang, Wenyi Hong, Yukuo Cen, Da Yin, Bin Chen, Wenmeng Yu, Xiaotao Gu, Jie Tang

    Abstract: GLM-OCR is an efficient 0.9B-parameter compact multimodal model designed for real-world document understanding. It combines a 0.4B-parameter CogViT visual encoder with a 0.5B-parameter GLM language decoder, achieving a strong balance between computational efficiency and recognition performance. To address the inefficiency of standard autoregressive decoding in deterministic OCR tasks, GLM-OCR intr… ▽ More

    Submitted 16 March, 2026; v1 submitted 11 March, 2026; originally announced March 2026.

  22. arXiv:2603.00680  [pdf, ps, other

    cs.AI

    MemPO: Self-Memory Policy Optimization for Long-Horizon Agents

    Authors: Ruoran Li, Xinghua Zhang, Haiyang Yu, Shitong Duan, Xiang Li, Wenxin Xiang, Chonghua Liao, Xudong Guo, Yongbin Li, Jinli Suo

    Abstract: Long-horizon agents face the challenge of growing context size during interaction with environment, which degrades the performance and stability. Existing methods typically introduce the external memory module and look up the relevant information from the stored memory, which prevents the model itself from proactively managing its memory content and aligning with the agent's overarching task objec… ▽ More

    Submitted 14 June, 2026; v1 submitted 28 February, 2026; originally announced March 2026.

  23. arXiv:2602.17011  [pdf, ps, other

    cs.MM

    CAFE: Channel-Autoregressive Factorized Encoding for Robust Biosignal Spatial Super-Resolution

    Authors: Hongjun Liu, Leyu Zhou, Zijianghao Yang, Rujun Han, Shitong Duan, Kuanjian Tang, Chao Yao

    Abstract: High-density biosignal recordings are critical for neural decoding and clinical monitoring, yet real-world deployments often rely on low-density (LD) montages due to hardware and operational constraints. This motivates spatial super-resolution from LD observations, but heterogeneous dependencies under sparse and noisy measurements often lead to artifact propagation and false non-local correlations… ▽ More

    Submitted 18 February, 2026; originally announced February 2026.

  24. arXiv:2601.17356  [pdf, ps, other

    cs.CR cs.PF

    Obfuscation as an Effective Signal for Prioritizing Cross-Chain Smart Contract Audits: Large-Scale Measurement and Risk Profiling

    Authors: Yao Zhao, Zhang Sheng, Shengchen Duan, Shen Wang, Daoyuan Wu, Zhiyuan Wan

    Abstract: Obfuscation raises the interpretation cost of smart-contract auditing, yet its signals are hard to transfer across chains. We present HOBFNET, a fast surrogate of OBFPROBE, enabling million-scale cross-chain scoring. The model aligns with tool outputs on Ethereum (PCC 0.9158, MAPE 8.20 percent) and achieves 8-9 ms per contract, yielding a 2.3k-5.2k times speedup. Across BSC, Polygon, and Avalanche… ▽ More

    Submitted 30 January, 2026; v1 submitted 24 January, 2026; originally announced January 2026.

  25. arXiv:2601.12137  [pdf, ps, other

    cs.LG cs.CV

    EMoE: Eigenbasis-Guided Routing for Mixture-of-Experts

    Authors: Anzhe Cheng, Shukai Duan, Shixuan Li, Chenzhong Yin, Mingxi Cheng, Shahin Nazarian, Paul Thompson, Paul Bogdan

    Abstract: The relentless scaling of deep learning models has led to unsustainable computational demands, positioning Mixture-of-Experts (MoE) architectures as a promising path towards greater efficiency. However, MoE models are plagued by two fundamental challenges: 1) a load imbalance problem known as the``rich get richer" phenomenon, where a few experts are over-utilized, and 2) an expert homogeneity prob… ▽ More

    Submitted 17 January, 2026; originally announced January 2026.

    Comments: accepted by ICASSP2026

  26. arXiv:2601.11525  [pdf, ps, other

    cs.HC

    PlotGen-Bench: Evaluating VLMs on Generating Visualization Code from Diverse Plots across Multiple Libraries

    Authors: Yi Zhao, Zhen Yang, Shuaiqi Duan, Wenmeng Yu, Zhe Su, Jibing Gong, Jie Tang

    Abstract: Recent advances in vision-language models (VLMs) have expanded their multimodal code generation capabilities, yet their ability to generate executable visualization code from plots, especially for complex 3D, animated, plot-to-plot transformations, or multi-library scenarios, remains underexplored. To address this gap, we introduce PlotGen-Bench, a comprehensive benchmark for evaluating plot-to-co… ▽ More

    Submitted 13 November, 2025; originally announced January 2026.

    Comments: 30 pages, 27 figures

  27. arXiv:2601.06002  [pdf, ps, other

    cs.CL cs.AI

    The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning

    Authors: Qiguang Chen, Yantao Du, Ziniu Li, Jinhao Liu, Songyao Duan, Jiarui Guo, Minghao Liu, Jiaheng Liu, Tong Yang, Ge Zhang, Libo Qin, Wanxiang Che, Wenhao Huang

    Abstract: Large language models (LLMs) often fail to learn effective long chain-of-thought (Long CoT) reasoning from human or non-Long-CoT LLMs imitation. To understand this, we propose that effective and learnable Long CoT trajectories feature stable molecular-like structures in unified view, which are formed by three interaction types: Deep-Reasoning (covalent-like), Self-Reflection (hydrogen-bond-like),… ▽ More

    Submitted 13 January, 2026; v1 submitted 9 January, 2026; originally announced January 2026.

    Comments: Preprint

  28. arXiv:2601.03888  [pdf, ps, other

    cs.SD cs.AI

    IndexTTS 2.5 Technical Report

    Authors: Yunpei Li, Xun Zhou, Jinchao Wang, Lu Wang, Yong Wu, Siyi Zhou, Yiquan Zhou, Yining Wang, Yaogen Yang, Zhetao Hu, Shiyao Duan, Jiacheng Xu, Jingchen Shu, Bin Xia

    Abstract: In prior work, we introduced IndexTTS 2, a zero-shot neural text-to-speech foundation model comprising two core components: a transformer-based Text-to-Semantic (T2S) module and a non-autoregressive Semantic-to-Mel (S2M) module, which together enable faithful emotion replication and establish the first autoregressive duration-controllable generative paradigm. Building upon this, we present IndexTT… ▽ More

    Submitted 11 August, 2026; v1 submitted 7 January, 2026; originally announced January 2026.

    Comments: 11 pages, 4 figures

  29. arXiv:2512.21711  [pdf, ps, other

    cs.CL cs.AI

    Do Latent Tokens Think? A Causal and Adversarial Analysis of Chain-of-Continuous-Thought

    Authors: Yuyi Zhang, Boyu Tang, Tianjie Ju, Sufeng Duan, Gongshen Liu

    Abstract: Latent tokens are gaining attention for enhancing reasoning in large language models (LLMs), yet their internal mechanisms remain unclear. This paper examines the problem from a reliability perspective, uncovering fundamental weaknesses: latent tokens function as uninterpretable placeholders rather than encoding faithful reasoning. While resistant to perturbation, they promote shortcut usage over… ▽ More

    Submitted 25 December, 2025; originally announced December 2025.

    Comments: 13 pages, 5 figures

    ACM Class: I.2.7; K.6.5

  30. arXiv:2512.06899  [pdf, ps, other

    cs.CR

    Patronus: Identifying and Mitigating Transferable Backdoors in Pre-trained Language Models

    Authors: Tianhang Zhao, Haodong Zhao, Wei Du, Pengzhou Cheng, Junxian Li, Sufeng Duan, Haojin Zhu, Gongshen Liu

    Abstract: The ``Pre-train, then fine-tune'' paradigm has revolutionized Natural Language Processing (NLP). In this context, transferable backdoors pose a severe threat to the Pre-trained Language Models (PLMs) supply chain, yet defensive research remains nascent, primarily relying on detecting anomalies in the output feature space. We identify a critical flaw that fine-tuning on downstream tasks inevitably… ▽ More

    Submitted 18 June, 2026; v1 submitted 7 December, 2025; originally announced December 2025.

    Comments: Work in progress

  31. arXiv:2511.10971  [pdf, ps, other

    cs.CV

    ERMoE: Eigen-Reparameterized Mixture-of-Experts for Stable Routing and Interpretable Specialization

    Authors: Anzhe Cheng, Shukai Duan, Shixuan Li, Chenzhong Yin, Mingxi Cheng, Heng Ping, Tamoghna Chattopadhyay, Sophia I Thomopoulos, Shahin Nazarian, Paul Thompson, Paul Bogdan

    Abstract: Mixture-of-Experts (MoE) architectures expand model capacity by sparsely activating experts but face two core challenges: misalignment between router logits and each expert's internal structure leads to unstable routing and expert underutilization, and load imbalances create straggler bottlenecks. Standard solutions, such as auxiliary load-balancing losses, can reduce load disparities but often we… ▽ More

    Submitted 26 March, 2026; v1 submitted 14 November, 2025; originally announced November 2025.

    Comments: Accepted in CVPR2026 Main Track

  32. arXiv:2510.27045  [pdf, ps, other

    cs.CL cs.CY

    Quantitative Intertextuality from the Digital Humanities Perspective: A Survey

    Authors: Siyu Duan

    Abstract: The connection between texts is referred to as intertextuality in literary theory, which served as an important theoretical basis in many digital humanities studies. Over the past decade, advancements in natural language processing have ushered intertextuality studies into the quantitative age. Large-scale intertextuality research based on cutting-edge methods has continuously emerged. This paper… ▽ More

    Submitted 30 October, 2025; originally announced October 2025.

  33. arXiv:2510.15653  [pdf, ps, other

    cs.LG

    Fast and Compact Tsetlin Machine Inference on CPUs Using Instruction-Level Optimization

    Authors: Yefan Zeng, Shengyu Duan, Rishad Shafik, Alex Yakovlev

    Abstract: The Tsetlin Machine (TM) offers high-speed inference on resource-constrained devices such as CPUs. Its logic-driven operations naturally lend themselves to parallel execution on modern CPU architectures. Motivated by this, we propose an efficient software implementation of the TM by leveraging instruction-level bitwise operations for compact model representation and accelerated processing. To furt… ▽ More

    Submitted 17 October, 2025; originally announced October 2025.

  34. arXiv:2510.10965  [pdf, ps, other

    cs.CL cs.AI

    Judge Before Answer: Can MLLM Discern the False Premise in Question?

    Authors: Jidong Li, Lingyong Fang, Haodong Zhao, Sufeng Duan, Gongshen Liu

    Abstract: Multimodal large language models (MLLMs) have witnessed astonishing advancements in recent years. Despite these successes, MLLMs remain vulnerable to flase premise problems. However, existing benchmarks targeting this issue are limited in scope: they often lack fine-grained categorization, exhibit insufficient coverage, and thus fail to provide a rigorous evaluation of the ability of models to rec… ▽ More

    Submitted 12 October, 2025; originally announced October 2025.

  35. arXiv:2510.00991  [pdf, ps, other

    cs.DC

    An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters

    Authors: Mingjun Zhang, Xiaohe Hu, Menghao Zhang, Ziteng Chen, Yanmin Jia, Yan Zhang, Da Liu, Qing Chen, Fangzheng Jiao, Jun Chen, He Liu, Aohan Zeng, Shuaixing Duan, Ruya Gu, Yang Jing, Bowen Han, Wei Chen, Wenqi Xie, Jinlong Hou, Yuan Cheng, Hongzhou Zhang, Bohua Xu, Mingwei Xu, Chunming Hu

    Abstract: Large-scale LLM training requires collective communication libraries to exchange data among distributed GPUs. As a company dedicated to building and operating large-scale GPU training clusters, we encounter several practical limitations of NCCL in production, including 1) SM competition between computation and communication, 2) expensive restart costs under link failures, and 3) insufficient obser… ▽ More

    Submitted 31 May, 2026; v1 submitted 1 October, 2025; originally announced October 2025.

    Comments: 19 pages, 21 figures

  36. arXiv:2509.22551  [pdf, ps, other

    quant-ph cs.AI cs.LG

    ConQuER: Modular Architectures for Control and Bias Mitigation in IQP Quantum Generative Models

    Authors: Xiaocheng Zou, Shijin Duan, Charles Fleming, Gaowen Liu, Ramana Rao Kompella, Shaolei Ren, Xiaolin Xu

    Abstract: Quantum generative models based on instantaneous quantum polynomial (IQP) circuits show great promise in learning complex distributions while maintaining classical trainability. However, current implementations suffer from two key limitations: lack of controllability over generated outputs and severe generation bias towards certain expected patterns. We present a Controllable Quantum Generative Fr… ▽ More

    Submitted 11 October, 2025; v1 submitted 26 September, 2025; originally announced September 2025.

  37. arXiv:2509.19403  [pdf, ps, other

    eess.SP cs.AI cs.LG

    Online Adaptation via Dual-Stage Alignment and Self-Supervision for Fast-Calibration Brain-Computer Interfaces

    Authors: Sheng-Bin Duan, Jian-Long Hao, Tian-Yu Xiang, Xiao-Hu Zhou, Mei-Jiang Gui, Xiao-Liang Xie, Shi-Qi Liu, Zeng-Guang Hou

    Abstract: Individual differences in brain activity hinder the online application of electroencephalogram (EEG)-based brain computer interface (BCI) systems. To overcome this limitation, this study proposes an online adaptation algorithm for unseen subjects via dual-stage alignment and self-supervision. The alignment process begins by applying Euclidean alignment in the EEG data space and then updates batch… ▽ More

    Submitted 23 September, 2025; originally announced September 2025.

  38. arXiv:2509.03937  [pdf, ps, other

    cs.CL cs.AI

    SPFT-SQL: Enhancing Large Language Model for Text-to-SQL Parsing by Self-Play Fine-Tuning

    Authors: Yuhao Zhang, Shaoming Duan, Jinhang Su, Chuanyi Liu, Peiyi Han

    Abstract: Despite the significant advancements of self-play fine-tuning (SPIN), which can transform a weak large language model (LLM) into a strong one through competitive interactions between models of varying capabilities, it still faces challenges in the Text-to-SQL task. SPIN does not generate new information, and the large number of correct SQL queries produced by the opponent model during self-play re… ▽ More

    Submitted 11 October, 2025; v1 submitted 4 September, 2025; originally announced September 2025.

    Comments: EMNLP 2025 Findings

  39. arXiv:2508.13993  [pdf, ps, other

    cs.CL cs.AI

    Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization

    Authors: Shaohua Duan, Pengcheng Huang, Xinze Li, Zhenghao Liu, Xiaoyuan Yi, Yukun Yan, Shuo Wang, Yu Gu, Ge Yu, Maosong Sun

    Abstract: Long-context modeling is critical for a wide range of real-world tasks, including long-context question answering, summarization, and complex reasoning tasks. Recent studies have explored fine-tuning Large Language Models (LLMs) with synthetic data to enhance their long-context capabilities. However, the effectiveness of such approaches is often limited by the low diversity and factual inconsisten… ▽ More

    Submitted 9 April, 2026; v1 submitted 19 August, 2025; originally announced August 2025.

    Comments: 17 pages

  40. arXiv:2508.12769  [pdf, ps, other

    cs.CL cs.AI

    CRED-SQL: Enhancing Real-world Large Scale Database Text-to-SQL Parsing through Cluster Retrieval and Execution Description

    Authors: Shaoming Duan, Zirui Wang, Chuanyi Liu, Zhibin Zhu, Yuhao Zhang, Peiyi Han, Liang Yan, Zewu Peng

    Abstract: Recent advances in large language models (LLMs) have significantly improved the accuracy of Text-to-SQL systems. However, a critical challenge remains: the semantic mismatch between natural language questions (NLQs) and their corresponding SQL queries. This issue is exacerbated in large-scale databases, where semantically similar attributes hinder schema linking and semantic drift during SQL gener… ▽ More

    Submitted 20 August, 2025; v1 submitted 18 August, 2025; originally announced August 2025.

  41. arXiv:2508.08719  [pdf, ps, other

    cs.CL cs.AI cs.CY

    IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization

    Authors: Yuzhuo Bai, Shitong Duan, Muhua Huang, Jing Yao, Zhenghao Liu, Peng Zhang, Tun Lu, Xiaoyuan Yi, Maosong Sun, Xing Xie

    Abstract: Trained on various human-authored corpora, Large Language Models (LLMs) have demonstrated a certain capability of reflecting specific human-like traits (e.g., personality or values) by prompting, benefiting applications like personalized LLMs and social simulations. However, existing methods suffer from the superficial elicitation problem: LLMs can only be steered to mimic shallow and unstable sty… ▽ More

    Submitted 27 November, 2025; v1 submitted 12 August, 2025; originally announced August 2025.

    Comments: This paper is accepted by AAAI 2026

  42. arXiv:2508.01746  [pdf, ps, other

    cs.AI

    Bayes-Entropy Collaborative Driven Agents for Research Hypotheses Generation and Optimization

    Authors: Shiyang Duan, Yuan Tian, Qi Bing, Xiaowei Shao

    Abstract: The exponential growth of scientific knowledge has made the automated generation of scientific hypotheses that combine novelty, feasibility, and research value a core challenge. Existing methods based on large language models fail to systematically model the inherent in hypotheses or incorporate the closed-loop feedback mechanisms crucial for refinement. This paper proposes a multi-agent collabora… ▽ More

    Submitted 3 August, 2025; originally announced August 2025.

    Comments: Corresponding author: Xiaowei Shao. 12 pages, 4 figures

    ACM Class: I.2.4

  43. arXiv:2508.01219  [pdf, ps, other

    cs.CV cs.LG

    Eigen Neural Network: Unlocking Generalizable Vision with Eigenbasis

    Authors: Anzhe Cheng, Chenzhong Yin, Mingxi Cheng, Shukai Duan, Shahin Nazarian, Paul Bogdan

    Abstract: The remarkable success of Deep Neural Networks(DNN) is driven by gradient-based optimization, yet this process is often undermined by its tendency to produce disordered weight structures, which harms feature clarity and degrades learning dynamics. To address this fundamental representational flaw, we introduced the Eigen Neural Network (ENN), a novel architecture that reparameterizes each layer's… ▽ More

    Submitted 2 August, 2025; originally announced August 2025.

  44. arXiv:2508.00308  [pdf, ps, other

    cs.CV

    Exploring Fourier Prior and Event Collaboration for Low-Light Image Enhancement

    Authors: Chunyan She, Fujun Han, Chengyu Fang, Shukai Duan, Lidan Wang

    Abstract: The event camera, benefiting from its high dynamic range and low latency, provides performance gain for low-light image enhancement. Unlike frame-based cameras, it records intensity changes with extremely high temporal resolution, capturing sufficient structure information. Currently, existing event-based methods feed a frame and events directly into a single model without fully exploiting modalit… ▽ More

    Submitted 1 August, 2025; originally announced August 2025.

    Comments: Accepted by ACM MM 2025

  45. arXiv:2507.21503  [pdf, ps, other

    cs.AI cs.CV

    MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions

    Authors: Yanxu Zhu, Shitong Duan, Xiangxu Zhang, Jitao Sang, Peng Zhang, Tun Lu, Xiao Zhou, Jing Yao, Xiaoyuan Yi, Xing Xie

    Abstract: Recently Multimodal Large Language Models (MLLMs) have achieved considerable advancements in vision-language tasks, yet produce potentially harmful or untrustworthy content. Despite substantial work investigating the trustworthiness of language models, MMLMs' capability to act honestly, especially when faced with visually unanswerable questions, remains largely underexplored. This work presents th… ▽ More

    Submitted 13 January, 2026; v1 submitted 29 July, 2025; originally announced July 2025.

    Comments: AAAI2026 Oral

  46. arXiv:2507.19951  [pdf, ps, other

    cs.SE

    PDLogger: Automated Logging Framework for Practical Software Development

    Authors: Shengcheng Duan, Yihua Xu, Sheng Zhang, Shen Wang, Yue Duan

    Abstract: Logging is indispensable for maintaining the reliability and diagnosability of modern software, yet developers still struggle to decide where and how to log effectively. Existing automated logging techniques focus on isolated sub-tasks - predicting a single log position, level, or message - and therefore cannot produce complete, high-quality log statements that reflect real-world practice in which… ▽ More

    Submitted 26 July, 2025; originally announced July 2025.

    Comments: 10 pages, 10 figures

    ACM Class: D.2

  47. arXiv:2507.18433  [pdf, ps, other

    eess.IV cs.CV

    DiagR1: A Vision-Language Model Trained via Reinforcement Learning for Digestive Pathology Diagnosis

    Authors: Minxi Ouyang, Lianghui Zhu, Yaqing Bao, Qiang Huang, Jingli Ouyang, Tian Guan, Xitong Ling, Jiawen Li, Song Duan, Wenbin Dai, Li Zheng, Xuemei Zhang, Yonghong He

    Abstract: Multimodal large models have shown great potential in automating pathology image analysis. However, current multimodal models for gastrointestinal pathology are constrained by both data quality and reasoning transparency: pervasive noise and incomplete annotations in public datasets predispose vision language models to factual hallucinations when generating diagnostic text, while the absence of ex… ▽ More

    Submitted 24 July, 2025; originally announced July 2025.

  48. arXiv:2507.01006  [pdf, ps, other

    cs.CV cs.AI cs.LG

    GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

    Authors: GLM-V Team, :, Wenyi Hong, Wenmeng Yu, Xiaotao Gu, Guo Wang, Guobing Gan, Haomiao Tang, Jiale Cheng, Ji Qi, Junhui Ji, Lihang Pan, Shuaiqi Duan, Weihan Wang, Yan Wang, Yean Cheng, Zehai He, Zhe Su, Zhen Yang, Ziyang Pan, Aohan Zeng, Baoxu Wang, Bin Chen, Boyan Shi, Changyu Pang , et al. (69 additional authors not shown)

    Abstract: We present GLM-4.1V-Thinking, GLM-4.5V, and GLM-4.6V, a family of vision-language models (VLMs) designed to advance general-purpose multimodal understanding and reasoning. In this report, we share our key findings in the development of the reasoning-centric training framework. We first develop a capable vision foundation model with significant potential through large-scale pre-training, which argu… ▽ More

    Submitted 1 January, 2026; v1 submitted 1 July, 2025; originally announced July 2025.

  49. arXiv:2506.20966  [pdf, ps, other

    cs.RO cs.AI

    Parallels Between VLA Model Post-Training and Human Motor Learning: Progress, Challenges, and Trends

    Authors: Tian-Yu Xiang, Ao-Qun Jin, Xiao-Hu Zhou, Mei-Jiang Gui, Xiao-Liang Xie, Shi-Qi Liu, Shuang-Yi Wang, Sheng-Bin Duan, Fu-Chao Xie, Wen-Kai Wang, Si-Cheng Wang, Ling-Yun Li, Tian Tu, Zeng-Guang Hou

    Abstract: Vision-language-action (VLA) models extend vision-language models (VLM) by integrating action generation modules for robotic manipulation. Leveraging the strengths of VLM in vision perception and instruction understanding, VLA models exhibit promising generalization across diverse manipulation tasks. However, applications demanding high precision and accuracy reveal performance gaps without furthe… ▽ More

    Submitted 28 January, 2026; v1 submitted 25 June, 2025; originally announced June 2025.

  50. SpikeSMOKE: Spiking Neural Networks for Monocular 3D Object Detection with Cross-Scale Gated Coding

    Authors: Xuemei Chen, Huamin Wang, Jing Peng, Hangchi Shen, Shukai Duan, Shiping Wen, Tingwen Huang

    Abstract: With the wide application of 3D object detection in some fields such as autonomous driving, its energy consumption is constantly increasing, making the research on low-power consumption alternatives a key research area. The spiking neural networks (SNNs), possessing low-power consumption characteristics, offer a novel solution for this research. Consequently, we apply SNNs to monocular 3D object d… ▽ More

    Submitted 10 March, 2026; v1 submitted 9 June, 2025; originally announced June 2025.