Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 3,435 results for author: Zha, W

.
  1. arXiv:2608.20019  [pdf, ps, other

    cs.AI

    Contrastive Mixed Prompt Learning for Incomplete Multimodal Sentiment Analysis with Unseen Modality Combination

    Authors: Kaixin Xu, NaiJin Liu, Yulin Kang, Tangyue Jin, Zixuan Yu, Wenxi Zhao, Yibei Liu, Qianle Zhang, Yangyang Wu, Mengying Zhu, Meng Xi

    Abstract: Incomplete multimodal sentiment analysis has garnered significant attention in recent years. Existing approaches typically assume that data is missing at random or are designed specifically for certain missing patterns, ignoring the modality combination inconsistency between training and testing phases. However, in real-world scenarios, the testing phase often encounters modal combinations that we… ▽ More

    Submitted 20 August, 2026; originally announced August 2026.

  2. arXiv:2608.18618  [pdf, ps, other

    cs.RO

    LabDex: A Hierarchical Benchmark for Dexterous Manipulation in Laboratories

    Authors: Zhipeng Tang, Sihang Chen, Sha Zhang, Peihao Yang, Yan Liu, Wentao Zhao, Xinrui Liu, Rui Huang, Wensheng Du, Yuting Huang, Jiajun Deng, Lidian Wang, Yuan Zhang, Yanyong Zhang

    Abstract: Autonomous laboratories hold great promise for accelerating scientific discovery. To achieve this vision, robots are supposed to dexterously manipulate diverse labware and instruments and execute long-horizon, state-dependent experimental procedures. Yet existing benchmarks do not jointly capture dexterous hand use, real-world laboratory interactions, and multi-stage experimental procedures, limit… ▽ More

    Submitted 19 August, 2026; originally announced August 2026.

    Comments: 13 pages, 3 figures

  3. arXiv:2608.18589  [pdf, ps, other

    hep-ph nucl-th physics.comp-ph

    Fourier Transforms of Color Glass Condensate Multi-Wilson-Line Correlators via Filon Quadrature

    Authors: Haowu Duan, Si-Wei Dai, Cong Yi, Wenbin Zhao

    Abstract: Calculating cross sections in the Color Glass Condensate effective theory requires Fourier transforms of multi-Wilson-line correlators from transverse coordinate space to transverse momentum space. Under the common assumption of impact-parameter independence, each transform reduces to a set of Hankel transforms whose Bessel-function kernels oscillate rapidly at phenomenologically relevant momenta,… ▽ More

    Submitted 19 August, 2026; originally announced August 2026.

    Comments: 42 pages, 14 figures, 5 tables

  4. arXiv:2608.16798  [pdf, ps, other

    cs.CL cs.AI cs.LG

    ClawGym II: Exploring Black-Box RL on Agent Harness

    Authors: Huatong Song, Fei Bai, Ming Yang, Renyuan Li, Jia Deng, Jujie He, Zhange Zhang, Daixuan Cheng, Yan Xing, Qi Yun, Xuxing Chen, Danyang Li, Feng Chang, Chuan Hao, Ran Tao, Jian Yang, Bryan Dai, Wayne Xin Zhao, Mingjie Tang, Ji-Rong Wen

    Abstract: Agent harnesses have substantially improved performance on long-horizon tasks by coordinating agent interactions with the environment. However, reinforcement learning through complex harnesses remains largely unexplored, as scaling such training to long-horizon agent tasks introduces fundamental challenges. In this work, we present a unified black-box RL framework for stable and scalable optimizat… ▽ More

    Submitted 17 August, 2026; originally announced August 2026.

  5. arXiv:2608.16578  [pdf, ps, other

    cs.AI cs.MA cs.SI

    Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents

    Authors: Batu El, Jinhee Paeng, Fatih Dinc, Shiye Su, Mete Erdogan, Aneesh Pappu, Haotian Ye, Wanjia Zhao, Surya Ganguli, James Zou

    Abstract: AI agents increasingly operate as part of interacting systems rather than in isolation. As agents exchange information and jointly make decisions, their interactions can improve collective reasoning but may also produce herding, polarization, or amplify shared biases. Understanding and predicting these collective dynamics is therefore important for designing effective and aligned multi-agent syste… ▽ More

    Submitted 17 August, 2026; originally announced August 2026.

    Comments: 51 pages, 20 figures, 9 tables

  6. arXiv:2608.14392  [pdf, ps, other

    cs.AI

    Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons

    Authors: Wei Zhao, Zhe Li, Peixin Zhang, Jun Sun

    Abstract: Neuron- and path-level interventions offer the finest-grained route to defending large language models (LLMs) against jailbreak attacks, yet existing methods fall short of this promise, i.e., they often compromise model utility significantly. Specifically, one line of work suppresses toxic neurons to erase harmful semantics, but since such semantics are distributed across the network, blocking eve… ▽ More

    Submitted 14 August, 2026; originally announced August 2026.

  7. arXiv:2608.14391  [pdf, ps, other

    cs.CV cs.AI

    Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination

    Authors: Shuo Liang, Yixing Ma, Pengfei Zhou, Zhenglin Wan, Xingyan Chen, Zihan Mei, Manting Li, Feihan Chen, Zhiwen Wang, Bin Xu, Haotian Zhang, Jiajun Song, Shiya Su, Run Liu, Zhenghang Ni, Yifa Yu, Jintao Hong, Bolong Feng, Yifei Liu, Zirui Zhang, Jingxuan Zhang, Songlin Zhao, Yifan Bai, Kang Tan, Yizhe Liu , et al. (11 additional authors not shown)

    Abstract: Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-world crises, creating substantial risks of misinformation. Existing benchmarks, however, provide limited evidence on detector and generator behavior in such settings, including how detectability varies with generation conditions, how people perceive generated videos, and whether detec… ▽ More

    Submitted 16 August, 2026; v1 submitted 14 August, 2026; originally announced August 2026.

    Comments: 63 pages, 20 figures, 32 tables

  8. arXiv:2608.13267  [pdf, ps, other

    cs.CL cs.AI cs.CV cs.LG

    How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures

    Authors: Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee, Christian Greisinger, Steffen Eger, Pius Onobhayedo, Wei Zhao

    Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when visual evidence is missing or misleading). We introduce SciFigBench, a diagnostic VLM benchmark for scientific figure understanding that jointly evalu… ▽ More

    Submitted 13 August, 2026; originally announced August 2026.

    Comments: 25 pages including appendix. Project website: https://scifigbench.nlp4sci.com

    ACM Class: I.2.7; I.2.10

  9. arXiv:2608.13102  [pdf, ps, other

    cs.CV

    RbFT-Net: Rectify-Before-Fuse Temporal Radar Anchors for 4D Radar-Camera Depth Completion

    Authors: Wentao Zhao, Shouxuan Wu, Yongtao Cen, Tianchen Deng, Yuyang Zhang, Jingchuan Wang

    Abstract: Dense metric depth prediction from cameras and millimeter-wave radar offers a cost-effective sensing solution for autonomous systems. However, radar measurements are inherently sparse and susceptible to clutter, multipath reflections, and projection errors. While aggregating multiple radar frames provides denser metric cues, it also introduces temporal misalignment and dynamic-object interference.… ▽ More

    Submitted 13 August, 2026; originally announced August 2026.

  10. arXiv:2608.12307  [pdf, ps, other

    cs.LG cs.AI cs.CL

    AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses

    Authors: Cheng Qian, Wenting Zhao, Liangwei Yang, Heng Wang, Jielin Qiu, Heng Ji, Silvio Savarese, Huan Wang, Shelby Heinecke

    Abstract: Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's parameters, through teacher forcing, on-policy distillation, and related training-time methods. In this paper, we ask whether such transfer can instead occur at test time. We study strong-to-weak scaffolding: whether a stronger builder model can construct inference-time harnesses t… ▽ More

    Submitted 12 August, 2026; originally announced August 2026.

    Comments: 23 Pages, 12 Figures, 6 Tables

  11. arXiv:2608.12127  [pdf, ps, other

    cs.CV

    SCOPE-Router: Cost-Aware Open-Set VLM Routing for Execution-Oriented Tasks

    Authors: Tao Yu, Yifei Qu, Zhiqing Cui, Pengfei Zhou, Zhongtian Luo, Yujia Yang, Shenghua Chai, Haopeng Jin, Zhenghao Zhang, Xinming Wang, Hongzhu Yi, Wangbo Zhao, Zhenglin Wan, Yan Huang, Yeshani, Jinwen Luo, Yang You

    Abstract: Model routing aims to select the most suitable model from a candidate pool for each query, balancing quality and cost. Existing VLM routing research is limited to traditional VQA evaluation, lacks systematic calibration optimization for open-set scenarios, and employs training objectives that dilute multi-positive signals via softmax normalization without incorporating cost. We address these limit… ▽ More

    Submitted 19 August, 2026; v1 submitted 12 August, 2026; originally announced August 2026.

  12. arXiv:2608.09819  [pdf, ps, other

    cs.LG cs.CL

    Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA

    Authors: Mind Lab, :, Vin Bo, Asher Cai, Jingwei Cao, Song Cao, Vic Cao, Amelia Chen, Andrew Chen, Kaijie Chen, Cleon Cheng, Steven Chiang, Kaixuan Fan, Hera Feng, Huan Feng, Arthur Fu, Jun Gao, Pyke Han, Nolan Ho, Ori Hong, Hailee Hou, Piers Hua, Charles Huang, Miles Jiang, Nora Jiang , et al. (52 additional authors not shown)

    Abstract: Macaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments and continuing to learn after deployment. It is organized around two system goals. Adaptation is pursued through recursive improvement of versioned model-harness pairs, where experience from one configuration is evaluated under an external contract and used to construct its success… ▽ More

    Submitted 10 August, 2026; originally announced August 2026.

    Comments: 49 pages, technical report

  13. arXiv:2608.09568  [pdf, ps, other

    cs.CL

    Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization

    Authors: Wenxiao Zhao, Shu Wang, Ying Nian Wu

    Abstract: Direct Preference Optimization (DPO) aggregates token-level log-probability ratios via uniform summation, implicitly treating all tokens as contributing equally to the preference signal. However, the contribution of individual tokens to the preference signal varies. We introduce token credit, which modulates each token's KL regularization based on its contribution to the preference outcome. We der… ▽ More

    Submitted 10 August, 2026; originally announced August 2026.

    Comments: 16 pages, 2 figures, COLM2026

  14. arXiv:2608.09280  [pdf, ps, other

    cs.CL

    Is the ACL Responsible NLP Checklist a Box-Ticking Exercise? A Large-Scale Analysis of EMNLP 2025

    Authors: Nusrath Jinnath, Wei Zhao

    Abstract: Responsible NLP practice includes a) transparency, b) ethics, and c) societal impacts. The Responsible NLP Checklist aims to push these goals, and promote responsible practice. Recently, ACL released the EMNLP 2025 Checklists to aid transparency on the current research practice, which we focus on. We curate and release the first two datasets of: a) all the checklist responses and justifications fr… ▽ More

    Submitted 10 August, 2026; originally announced August 2026.

  15. arXiv:2608.09096  [pdf, ps, other

    cs.CL

    Evo-Bench: Can Language Models Improve Agent Harness?

    Authors: Lisheng Huang, Chen Yang, Hao Zhou, Huatong Song, Zongchao Chen, Ran Le, Yang Song, Wayne Xin Zhao, Tao Zhang

    Abstract: Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously optimize its own operating harness. However, systematically benchmarking this capability remains challenging, as existing evaluations fail to isolate harness improvements from… ▽ More

    Submitted 10 August, 2026; v1 submitted 9 August, 2026; originally announced August 2026.

  16. arXiv:2608.08802  [pdf, ps, other

    cs.AI

    Improving Generalization Robustness of Multimodal RLVR

    Authors: Pengfei Zhou, Zhiwei Tang, Xiaopeng Peng, Chenrui Zhou, Lama Moukheiber, Yixing Ma, Bin Xu, Jiajun Song, Zhenglin Wan, Wangbo Zhao, Jiasheng Tang, Bohan Zhuang, Fan Wang, Yang You

    Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) makes Multimodal Large Language Models more accurate, but the gains are brittle: simply paraphrasing a question or changing the prompt template can degrade them, which challenges reliable deployment in high-stakes scenarios like medical VQA. We trace this to two issues of the standard RL objective. First, the binary verifier conflates format wi… ▽ More

    Submitted 14 August, 2026; v1 submitted 9 August, 2026; originally announced August 2026.

    Comments: 32 pages, 5 figures

  17. arXiv:2608.08768  [pdf, ps, other

    cs.IR

    BOUND: Brief-Guided Corrective Preference Distillation at Search-Control Boundaries

    Authors: Qingying Niu, Ruiyang Ren, Wayne Xin Zhao, Yaliang Li

    Abstract: Large language model (LLM)-based deep search agents solve tasks through iterative retrieval and reasoning, but locally relevant evidence can cause persistent wrong-anchor drift, constraint drift, or local-topic drift. Existing methods supervise trajectories, outcomes, or steps, but rarely distinguish task-aligned continuations from locally plausible ones that reinforce drift. We propose BOUND, a b… ▽ More

    Submitted 9 August, 2026; originally announced August 2026.

    Comments: 15 pages

  18. arXiv:2608.08596  [pdf, ps, other

    cs.CV

    Goal-oriented Navigation Instruction Generation with Tour Video Priors

    Authors: Fangdi Li, Juncheng Liao, Changxu Cheng, Jiazhi Wang, Senda Chen, Tao Wang, Wuyue Zhao

    Abstract: Navigation Instruction Generation (NIG) aims to produce step-by-step natural language instructions for navigation guidance. Existing studies primarily treat NIG as an auxiliary task for vision-andlanguage navigation (VLN), focusing on data augmentation or multi-task learning. However, generating navigation instructions from compact environmental priors requires meticulous spatial reasoning, especi… ▽ More

    Submitted 9 August, 2026; originally announced August 2026.

  19. arXiv:2608.07918  [pdf, ps, other

    econ.TH

    A Note on Market Segmentation and Bertrand Competition

    Authors: Zhang Xu, Mingsheng Zhang, Wei Zhao

    Abstract: In this note, we show that equilibrium profit is zero in Bertrand competition with a finite number of firms and consumers whose willingness to pay are bounded, under any market segmentation profile.

    Submitted 8 August, 2026; originally announced August 2026.

  20. arXiv:2608.07525  [pdf, ps, other

    cs.CL cs.AI

    Unified Hallucination Fuzzing for Multimodal Large Language Models

    Authors: Pengfei Zhou, Jiajun Song, Zhiwei Tang, Yixing Ma, Xiaopeng Peng, Donghui Si, Yuhang Xu, Huiqi Song, Yiyuan Miao, Yichen Qian, Weihua Chen, Wangbo Zhao, Bohan Zhuang, Jiasheng Tang, Yang You

    Abstract: Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applications. Existing evaluations, predominantly based on static benchmarks, suffer from narrow taxonomical coverage and rapid performance saturation, failing to reflect model robustness in evolving real-world scenarios. To bridge this gap, we present a sys… ▽ More

    Submitted 15 July, 2026; originally announced August 2026.

    Comments: 47 pages, 17 figures

  21. arXiv:2608.07055  [pdf, ps, other

    cs.IR

    Teacher Retains Full Tokens, Student Merges Efficiently: TM20K for E-Commerce Sequence Modeling in Ad Recommendation

    Authors: Xinchun Li, Duoru Zheng, Wenlin Zhao, Haoran Ding, Ziyi Zhou, Jingxuan Tan, Huizhi Yang, Yuchen Jiang, Zhe Chen, Yuchao Zheng, Linlan Chen, Dongjian Wang, Dongyue Wang, Xiaosong Li, Hongyue Mao, Yaocheng Tan

    Abstract: Benefiting from ultra-long behavior sequence modeling, existing recommender systems bring users a better experience via simultaneously considering their long-term and short-term interests. Nevertheless, extended sequence lengths introduce substantial burdens on training efficiency and serving throughput. Prior approaches typically utilize search-based or cluster-based compression on ultra-long seq… ▽ More

    Submitted 13 August, 2026; v1 submitted 7 August, 2026; originally announced August 2026.

    Comments: ByteDance 20K Ultra-long Sequence Modeling for Ad E-Commerce Recommendation

  22. arXiv:2608.06838  [pdf, ps, other

    cs.DC

    StateFlow: Sequence Pipeline Parallelism for Long-Context Modeling with Linear Recurrence

    Authors: Wenxuan Zhao, Yingfa Chen, Xu Han, Wenjing Han, Tianbo Huang, Zhiyu Li, Ao Sun, Jingheng Xu, Lin Gan, Guangwen Yang

    Abstract: Long-context training is increasingly important for large language models, and linear attention and state space models have become popular for improving long-context efficiency. However, efficiently parallelizing long-sequence training for recurrent and hybrid models remains challenging. We present StateFlow, a sequence pipeline parallelism system for models with linear recurrence. StateFlow par… ▽ More

    Submitted 7 August, 2026; originally announced August 2026.

  23. arXiv:2608.06352  [pdf, ps, other

    cs.LG cs.CL

    CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

    Authors: Fanzhe Meng, Guoxin Chen, Jiale Zhao, Shuang Sun, Zhiyu Lin, Wayne Xin Zhao, Ruihua Song, Ji-Rong Wen, Kai Jia

    Abstract: Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable validation establishes feasibility, yet does not reveal how a task behaves relative to a given solver setting. In this paper, we present CalibForge, an autonomous terminal-task synthesis system that uses verified solver behavior to revise candidate… ▽ More

    Submitted 6 August, 2026; originally announced August 2026.

    Comments: Dataset: https://huggingface.co/datasets/AweAI-Team/CalibForge. Repository: https://github.com/AweAI-Team/CalibForge

  24. arXiv:2608.05377  [pdf, ps, other

    nucl-ex hep-ex

    ePIC Early Science Report

    Authors: D. Abbott, N. Abdelrahman, S. Abhijit, I. Abualrob, R. B. Achari, J. Adam, L. Adamczyk, K. Adkins, A. Affolder, K. Agarwal, J. Agarwala, N. Agrawal, C. A. Aidala, W. Akers, A. Al-bataineh, S. N. Alam, M. Alekseev, P. R. Altieri, J. -S. Alvarado Gallenao, S. B. L. Amar, R. Ammendola, I. Amos Cali, G. An, D. Anderson, E. Anderssen , et al. (774 additional authors not shown)

    Abstract: This Early Science Report from the ePIC Collaboration outlines the compelling physics program achievable during the first years of operation of the Electron-Ion Collider (EIC), prior to the establishment of the full design luminosity and energy range. The analyses are based on realistic early-running beam configurations and detailed Geant4 ePIC detector simulations, hit digitization and data recon… ▽ More

    Submitted 5 August, 2026; originally announced August 2026.

    Report number: epic-AN-AC-2026-004

  25. arXiv:2608.04996  [pdf, ps, other

    cs.RO

    DreamWAM: Beyond RGB Future Prediction for World Action Models

    Authors: Shanglin Yuan, Weiheng Zhao, Xin Shi, Haoyi Jiang, Xianda Guo, Liu Liu, Wenyu Liu, Wei Sui, Xinggang Wang

    Abstract: World Action Models (WAMs) learn action-relevant representations by predicting how the observed world will evolve. Most existing WAMs define this future in RGB space, where task-relevant state transitions are entangled with nuisance variations in texture, illumination, background, and viewpoint. We argue that WAMs should explicitly predict action-relevant future state rather than relying on RGB pr… ▽ More

    Submitted 5 August, 2026; originally announced August 2026.

  26. arXiv:2608.04872  [pdf, ps, other

    cs.CL cs.AI cs.LG

    A-SR: Self-Evolving Agentic LLMs for Symbolic Regression via Hierarchical Coordination

    Authors: Wenxiao Zhao, Dong Liu, Kaiyi Xu, Feng Liu, Zhen Zhao, Fei Ben, Shu Wang, Wenhao Li, Ying Nian Wu, Fenghua Ling, Haobo Li, Lei Bai

    Abstract: Symbolic regression aims to discover closed-form equations from data, but existing LLM-guided methods often rely on a unified proposal loop that compresses heterogeneous search failures into a scalar score and a single prompt. We propose A-SR, a self-evolving agentic framework that shifts the control unit from expression edits to role-conditioned evidence views. A-SR coordinates formula discovery… ▽ More

    Submitted 6 August, 2026; v1 submitted 5 August, 2026; originally announced August 2026.

    Comments: 18 pages, 8 figures, including appendix

  27. arXiv:2608.04721  [pdf, ps, other

    cs.RO

    Enabling Urgency-aware Robot Swarm Intralogistics using Smart IoT Tags

    Authors: Youssef Alboraei, Murray Groves, Shane Wen, Wenda Zhao, Senhui Qiu, Mohammud J. Bocus, Robert Piechocki, Sabine Hauert, Kerstin Eder

    Abstract: Warehouse items differ in how urgently they must be moved: perishable goods, pharmaceutical shipments, and just-in-time production materials must be delivered sooner than the rest of the stock. Decentralised robot swarms suit warehouses that cannot justify fixed automation infrastructure, but current swarm controllers treat all items alike or rely on an external scheduler to set priorities, so urg… ▽ More

    Submitted 5 August, 2026; originally announced August 2026.

    Comments: 16 pages, 8 figures, 3 Tables

  28. arXiv:2608.04404  [pdf, ps, other

    cs.CV

    Faster-WAM: Efficient Inference-Time Future Conditioning for Robust World Action Models

    Authors: Weiheng Zhao, Haoyi Jiang, Xin Shi, Liu Liu, Fan Huang, Zhizhong Su, Wei Sui, Xinggang Wang

    Abstract: World Action Models (WAMs) improve robot manipulation by learning how the environment evolves beyond the current observation. However, existing approaches face a fundamental dilemma: Joint-WAMs preserve future-aware representations during inference but incur prohibitive computation costs, while efficient alternatives remove future modeling at inference time and may lose the robustness benefits of… ▽ More

    Submitted 4 August, 2026; originally announced August 2026.

  29. arXiv:2608.03740  [pdf, ps, other

    cs.AI

    MissClick: Exploiting Digit-Serialized Coordinates to Attack GUI Grounding Models

    Authors: Yu Ran, Wentao Zhao, Xin Zhang, Yi Pan

    Abstract: Recent GUI visual grounding models generate screen coordinates as sequences of digit tokens that are parsed into numerical values and mapped to executable clicks. The security implications of this coordinate generation process have been largely overlooked. We observe that each coordinate digit is predicted as a categorical token, yet after parsing, changing a hundreds-place digit by one changes th… ▽ More

    Submitted 4 August, 2026; originally announced August 2026.

  30. arXiv:2608.03457  [pdf, ps, other

    cs.AI

    LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models

    Authors: Fengqi Zhu, Shaoxuan Xu, Jingyang Ou, Zebin You, Yipeng Xing, Huabin Liu, Xiaolu Zhang, Jun Zhou, Zhenzhong Lan, Yankai Lin, Wayne Xin Zhao, Jianguo Li, Chongxuan Li, Ji-Rong Wen

    Abstract: Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling behavior of Mixture-of-Experts (MoE) dLLMs remains poorly understood. We systematically characterize how optimization hyperparameters, compute allocation, and architecture scale for MoE dLLMs, identifying quantitative differences from scaling trends previously reported for AR models. Sp… ▽ More

    Submitted 4 August, 2026; originally announced August 2026.

  31. arXiv:2608.01638  [pdf, ps, other

    cs.CV

    Dynamic Resolution Routing for Efficient Egocentric Grounding

    Authors: Huixin Sun, Wangbo Zhao, Fanyue Wei, Qiuxia Lin, Pengzhan Sun, Angela Yao

    Abstract: Egocentric visual grounding requires high-resolution inputs to localize small objects. However, scaling Multimodal Large Language Models to this domain is constrained by the excessive cost of visual token processing. We identify that current efficient strategies based on token reduction are unreliable for selecting object-centric spatial evidence. To overcome this, we propose SmartRes, a framework… ▽ More

    Submitted 2 August, 2026; originally announced August 2026.

  32. arXiv:2608.01315  [pdf, ps, other

    cs.IR

    Collaborative Memory Augmentation for Generative Recommendation

    Authors: Enze Liu, Zhen Tian, Wayne Xin Zhao

    Abstract: Generative Recommendation (GR) has exhibited great potential by modeling item transitions as a sequence-to-sequence task. Despite the success of GR, existing frameworks primarily focus on modeling individual user sequences within a constrained internal parametric space, failing to explicitly leverage cross-user collaborative signals. To address this issue, we propose \textbf{OMEGA}, a cOllaborativ… ▽ More

    Submitted 2 August, 2026; originally announced August 2026.

    Comments: Accepted by KDD 2026 Research Track

  33. arXiv:2608.01239  [pdf

    eess.SP q-bio.CB

    Smart membrane: high content in situ monitoring barrier on chip with artificial neural network

    Authors: Bo Tang, Victor Krajka, Mengxi Liu, Wei Zhao, Gazal Goekkus, Paul Lukowicz, Lili Zhu, Pu Chen, Stephan Reichl, Andreas Dietzel

    Abstract: Conventional transepithelial electrical resistance (TEER) technique provides only a low-content analysis of cell-layer conditions, necessitating repeated microscopic assessments of morphology and cell-cell contacts outside the incubator for barrier-on-chip systems. This work presents a novel high-content TEER device in the form of a novel nanoporous membrane that facilitates continuous electrical… ▽ More

    Submitted 4 August, 2026; v1 submitted 2 August, 2026; originally announced August 2026.

  34. arXiv:2608.00301  [pdf, ps, other

    cs.LG cs.CL

    Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning

    Authors: Xujun Che, Yuchen Yuan, Weida Zhao, Chenyang Yu

    Abstract: Error-penalized scoring rules ($+1$ for a correct answer, $-λ$ for a wrong one, $0$ for abstaining) are increasingly prescribed against hallucination: a rational agent facing such a rule answers exactly when its correctness probability exceeds Chow's threshold $t^\ast=λ/(1+λ)$. We prove that a KL-anchored gradient learner can do the opposite. When abstention is a discrete action, the reward gradie… ▽ More

    Submitted 31 July, 2026; originally announced August 2026.

  35. arXiv:2607.29222  [pdf, ps, other

    cs.CV

    Is It Time for the Renaissance of Salient Object Detection in the Era of MLLMs?

    Authors: Wenzhuo Zhao, Xiuzhi Li, Zhongkuan Mao, Ronghao Xian, Yao Jiang, Zhao Gao, Keren Fu, Qijun Zhao, Jian Cheng

    Abstract: The zero-shot capabilities of multimodal large language models (MLLMs) are pushing salient object detection (SOD) beyond task-specific supervision. To disentangle MLLMs beyond conventional mask-based evaluation, we decompose SOD into localization and segmentation, and re-engineer datasets with phrases, boxes, and attributes, establishing a diagnostic benchmark for MLLM saliency perception (SaliLLM… ▽ More

    Submitted 31 July, 2026; originally announced July 2026.

    Comments: 10 pages, 4 figures, conference

  36. arXiv:2607.28777  [pdf, ps, other

    cs.CL

    Self-Supervised Skill Optimization

    Authors: Siran Peng, Cuiyu Yang, Tianyu Fu, Tianshuo Zhang, Haoyuan Zhang, Weisong Zhao, Anyang Su, Minghui Wu, Huiying Li, Xiangyu Zhu, Chenxu Zhao, Zhen Lei

    Abstract: Agent skills provide frozen large language model (LLM) agents with reusable procedural guidance, and recent work shows that such skills can be optimized with ground-truth (GT) feedback. Many applications, however, lack GT labels, task scores, rewards, or reliable task-specific evaluators. We therefore introduce Self-Supervised Skill Optimization (SSO), a comparative framework that learns a reusabl… ▽ More

    Submitted 30 July, 2026; originally announced July 2026.

  37. arXiv:2607.27982  [pdf, ps, other

    cs.CV

    ViP-Rig: Visual-Prompted Controllable Rigging

    Authors: Zihan Qin, Mingze Sun, Yifan Mao, Jialei Xu, Jingfeng Guo, Changrong Hu, Wenbo Zhao, Junjun Jiang, Xianming Liu

    Abstract: Rigging is inherently task-dependent because the same mesh may require different skeletons and deformation behaviors across animation tasks. In practice, artists often inspect an initial rig and repeatedly edit its skeletal structure and deformation behavior to meet specific animation requirements. Existing automatic methods primarily generate a plausible rig from geometry, offering limited explic… ▽ More

    Submitted 30 July, 2026; originally announced July 2026.

    Comments: 8 pages, 4 figures. Zihan Qin and Mingze Sun contributed equally. Xianming Liu is the corresponding author

  38. arXiv:2607.27845  [pdf, ps, other

    cs.CL cs.AI

    AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification

    Authors: Haobo Li, Eunseo Jung, Wenxiao Zhao, Feng Liu, Jiong Wang, Kaiyi Xu, Zijie Guo, Zixin Chen, Ben Fei, Fenghua Ling, Lei Bai

    Abstract: Recent advances in large language models (LLMs) have enabled AI systems to assist scientific research and peer review. However, an essential capability for reliable AI-assisted scientific workflows remains underexplored: verifying whether reviewer feedback leads to meaningful and evidence-supported manuscript improvements. We introduce AutoSupervision, which evaluates whether scientific manuscript… ▽ More

    Submitted 30 July, 2026; originally announced July 2026.

  39. arXiv:2607.27830  [pdf, ps, other

    cs.CV

    Thinking Once Is Enough: Intermediate-Layer Evidence Routing for High-Resolution VQA

    Authors: Zhongkuan Mao, Xianjie Liu, Tianyu Meng, Yidong Wang, Wenzhuo Zhao, Ronghao Xian, Yao Jiang, Fei Shen, Junfeng Fang, Yong Dai, Yi Zhang, Keren Fu

    Abstract: High-resolution visual question answering (HR-VQA) is often treated as a problem of insufficient evidence acquisition, where failing multimodal large language models must inspect images again through cropping, re-encoding, or multi-round search. We show that this view is incomplete: in many cases, fine-grained evidence has already survived visual encoding and become identifiable and influential wi… ▽ More

    Submitted 30 July, 2026; originally announced July 2026.

  40. arXiv:2607.27764  [pdf, ps, other

    cs.CV cs.AI

    Private Face Recognition Training Dataset Publication via Identity-Decoupled and Geometry-Preserving Face Distillation

    Authors: Shuhuan Chen, Xiangyu Zhu, Weisong Zhao, Siran Peng, Tianshuo Zhang, Haoyuan Zhang, Haichao Shi, Xiao-Yu Zhang, Zhen Lei

    Abstract: Publishing private face recognition~(FR) training datasets is privacy-sensitive because faces expose identity information. Private FR training dataset publication mitigates this risk by releasing protected proxies as substitutes for private training faces. However, training FR models with such data introduces an identity paradox: \emph{the identity cues that make released faces useful for recognit… ▽ More

    Submitted 30 July, 2026; originally announced July 2026.

  41. arXiv:2607.27603  [pdf, ps, other

    hep-ph nucl-ex nucl-th

    Unbiased Data-Driven Determination of the Nuclear Dipole Amplitude in the Color Glass Condensate

    Authors: Si-Wei Dai, Haowu Duan, Long-Gang Pang, Guang-You Qin, Shu-Yi Wei, Han-Zhong Zhang, Wenbin Zhao

    Abstract: Gluon saturation limits the growth of parton densities at small Bjorken-$x$ and is expected to be most pronounced in heavy nuclei. Yet quantitative extractions of the nuclear gluon dipole amplitude have long relied on parametrized initial conditions, introducing uncontrolled model dependence that obscures genuine nuclear effects. We introduce a physics-informed neural-network framework that embeds… ▽ More

    Submitted 29 July, 2026; originally announced July 2026.

  42. arXiv:2607.26509  [pdf, ps, other

    cs.LG cs.AI

    Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning

    Authors: Gong Gao, Xiao Lai, Ziqi Xie, Guojie Chen, Xianhui Liu, Weidong Zhao

    Abstract: Deep off-policy reinforcement learning algorithms for continuous control typically rely on neural value function approximation to guide policy improvement. However, temporal-difference (TD) learning introduces noisy targets, resulting in non-stationary optimization, while greedy policy updates amplify early-stage estimation errors. The recursive propagation of such errors leads to persistent overe… ▽ More

    Submitted 29 July, 2026; originally announced July 2026.

  43. arXiv:2607.26367  [pdf, ps, other

    cs.AI

    Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?

    Authors: Wanyu Zhao, Wanbing Zhao

    Abstract: An important skill in theoretical physics is to recognize when a new problem can be transformed into a known model. We study this skill as an AI-agent task: can LLM-based agents discover statistical mechanical mappings from a raw partition function to a tractable representation? To probe this question, we introduce StatMechBench-v0, a benchmark of six Ising-type problems covering transfer-matrix m… ▽ More

    Submitted 28 July, 2026; originally announced July 2026.

    Comments: Accepted to the AID-Wild workshop at CAIS 2026

  44. arXiv:2607.26329  [pdf, ps, other

    astro-ph.EP cond-mat.mtrl-sci physics.geo-ph

    Nanoscale Storage of Incompatible Elements at Olivine Grain Boundaries in Natural Basalts

    Authors: Wenhao Zhao, Reid Cooper, Stephen Parman, Austin Akey, Greg Hirth

    Abstract: Grain boundaries are pervasive in polycrystalline olivine, yet their structure and trace-element chemistry remain poorly constrained. We characterize boundaries in undeformed olivine aggregates from Piton de la Fournaise (La Reunion) and Mauna Loa (Hawaii) by correlating electron backscatter diffraction, transmission electron microscopy, and atom probe tomography. The boundaries are crystalline, w… ▽ More

    Submitted 28 July, 2026; originally announced July 2026.

    Comments: 46 pages; 10 main figures and 10 supplementary figures; supplementary text and tables included

  45. arXiv:2607.25216  [pdf, ps, other

    cs.IR cs.AI

    TopoGR: Revealing and Preserving Latent Structure of Semantic ID in Generative Recommendation

    Authors: Ziyu Zheng, Zhengshun Du, Yaming Yang, Bin Tong, Guan Wang, Meng Yan, Ziyu Guan, Wei Zhao

    Abstract: Semantic ID-based generative recommendation tokenizes each item into a sequence of discrete semantic IDs and predicts the next item by generating semantic IDs. However, existing methods typically regard SIDs as independent discrete symbols, while often overlooking the topology of the learned semantic ID space. We identify a structural mismatch between tokenization and generation: the tokenizer lea… ▽ More

    Submitted 27 July, 2026; originally announced July 2026.

    Comments: under review

  46. arXiv:2607.24889  [pdf, ps, other

    cs.LG cs.AI cs.CE

    GAUGE: Grading Agent-Built Financial Models Without a Golden Answer

    Authors: Jiacheng Lu, Sinuo Wang, Wentao Zhao, Rui Sun, Cheng Hua, Tao Song, Hui Cai, Beidi Luan, Zhengze Wu, Lingjing Teng, Yijia He, Jing Li, Daxin Jiang, Zuo Bai, Haibing Guan

    Abstract: Financial models combine public disclosures with analyst assumptions to produce forecasts and valuations. While some components can be checked mechanically, forecasts, discount rates, and target prices often admit multiple reasonable answers. Existing benchmarks nevertheless tend to grade such outputs against a single expert reference. Using independently built analyst models for the same companie… ▽ More

    Submitted 27 July, 2026; originally announced July 2026.

    Comments: 35 pages, including appendices. Code, benchmark materials, and evaluation artifacts will be publicly released

  47. arXiv:2607.24653  [pdf, ps, other

    cs.CL cs.LG

    Kimi K3: Open Frontier Intelligence

    Authors: Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, M. C., Jianfeng Cai, Xinyuan Cai, Peizhou Cao, Yuxuan Cao, Ziwei Chai, Y. Charles, H. S. Che, Guanduo Chen, Guangyu Chen, Guanzheng Chen, Huarong Chen, Jia Chen, Jianlong Chen, Jun Chen, Kexin Chen, Peng Chen, Ruijue Chen, Wentao Chen, Xin Chen, Yang Chen , et al. (377 additional authors not shown)

    Abstract: We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. Together with Stable LatentMoE, which effectively activates 16 of 896 routed experts per token… ▽ More

    Submitted 7 August, 2026; v1 submitted 27 July, 2026; originally announced July 2026.

    Comments: K3 tech report

  48. arXiv:2607.24091  [pdf, ps, other

    gr-qc

    Gravitational Lensing of Gravitational Waves: Towards a Higher-order Geometric-optics Approach

    Authors: Zhao Li, Shaoqi Hou, Wen Zhao

    Abstract: In this work, we study the gravitational lensing of gravitational waves (GWs) by extending the geometric-optics approximation to higher order. With the help of the Newman-Penrose formalism, we reexpress the GW propagation equations as a series of scalar equations and present explicit expressions for the Weyl scalars that describe the GW polarizations. By combining the approaches of solving geodesi… ▽ More

    Submitted 27 July, 2026; originally announced July 2026.

    Comments: 26 pages, 5 figures

  49. arXiv:2607.23958  [pdf, ps, other

    cs.CV

    RODR: Riemannian Orthogonally Decoupled Regularization for Disentangled Manifold Representation

    Authors: Jiayu Zhu, Wenlai Zhao

    Abstract: Point cloud denoising is essentially a geometric recovery task that aims to reconstruct the intrinsic structure of a smooth 2D Riemannian manifold embedded in R^3 from noisy, discrete ambient-space samples. Despite the remarkable progress of modern manifold-aware encoders and generative transport models in geometric representation learning, a fundamental objective-geometry mismatch remains underex… ▽ More

    Submitted 26 July, 2026; originally announced July 2026.

    Comments: 10 pages, 4 figures

    MSC Class: 68T05; 53B20 ACM Class: I.2.6; G.1.6; I.3.5

  50. arXiv:2607.23802  [pdf, ps, other

    cs.AI

    From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

    Authors: Qinsi Wang, Jing Shi, Huazheng Wang, Kun Wan, Yiran Wu, Bo Liu, Qingyun Wu, Hai Helen Li, Yiran Chen, Handong Zhao, Wentian Zhao

    Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-scale optimization. However, its applicability remains largely limited to domains such as mathematics and coding, where correctness can be deterministically verifiable. Open-ended tasks instead often rely on human preferences, reward models, or LLM-b… ▽ More

    Submitted 30 July, 2026; v1 submitted 26 July, 2026; originally announced July 2026.

    Comments: COLM 2026