Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 13,924 results for author: liu, y

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.21355  [pdf, ps, other

    cs.RO

    ViTacPhys: Physical Property-Aware Grasping from Human Visual-Tactile Demonstrations

    Authors: Yiwen Liu, Yujun Zhu, Kui Jia, Zhao Liao, Yangwei You, Shuaijun Wang

    Abstract: Recent vision-based action models have demonstrated strong capabilities in complex manipulation, but they rarely leverage explicit object physical properties to adapt their policies. We introduce ViTacPhys, a visual-tactile framework and data acquisition system that estimates object mass and friction-coefficient classes, together with continuous stiffness, from human manipulation demonstrations. T… ▽ More

    Submitted 21 August, 2026; originally announced August 2026.

    Comments: 11 pages, 7 figures. Project page: https://vitacphys.github.io/ViTacPhys/

  2. arXiv:2608.21243  [pdf, ps, other

    cs.IR cs.AI

    Adapting Knowledge Graphs for Behavior Denoising in Sequential Recommendation

    Authors: Zichun Jin, Zihan Zhou, Yinan Liu, Bin Wang, Xiaochun Yang

    Abstract: Sequential recommendation predicts the next item from a user's interaction history, but not every interaction is equally informative. Real logs combine persistent preferences with temporary needs, exploration, and incidental behavior, so some interactions can distort history representations or provide unreliable supervision. Existing denoising methods judge such interactions mainly from co-occurre… ▽ More

    Submitted 21 August, 2026; originally announced August 2026.

  3. arXiv:2608.21229  [pdf, ps, other

    cs.CV

    Anchoring Instruction Outside Mask: Exact Reference Caching for Efficient In-Context Diffusion Transformers

    Authors: Yangshuai Liu, Zheming Li, Jiaao Li, Kang He, Ziliang Lai, Zhitai Liu, Chengru Song

    Abstract: Omnimodal generation is central to a wide range of content creation and editing applications. In-context conditioning is essential to this paradigm. It allows diffusion transformers to process text instructions and visual references in a shared attention sequence. However, each reference image introduces thousands of tokens. Computation therefore grows rapidly with the number of references. Existi… ▽ More

    Submitted 21 August, 2026; originally announced August 2026.

  4. arXiv:2608.21218  [pdf, ps, other

    cs.AI cs.CL cs.IR

    Enhancing LLMs in Predictive Political QA with Semi-Structured Data

    Authors: Yinan Liu, Zihan Zhou, Zichun Jin, Xinyu Wang, Bin Wang, Xiaochun Yang

    Abstract: Predictive political question answering (QA), such as predicting how a political actor will vote, goes beyond factual lookup. External political resources offer rich historical evidence, but rarely contain the answer itself. Existing LLM augmentation methods, including actor-profile-based simulation and knowledge graph evidence injection, improve political reasoning but largely treat external reso… ▽ More

    Submitted 21 August, 2026; originally announced August 2026.

  5. arXiv:2608.21019  [pdf, ps, other

    cs.CL cs.AI

    Target-Aware Calibration Data Selection for Preserving Uncertainty in Quantized Language Models

    Authors: Zhen Yang, Sizai Hou, Kaiwen Zheng, Yaofang Liu, Liang He, Yixuan Chen, Kangning Cui

    Abstract: Quantization is widely used to deploy large language models, but its effect on uncertainty behavior, such as confidence, margins, and abstention, is rarely treated as a primary objective. We frame calibration-data selection for quantization as a target-dependent uncertainty-preservation problem. Different deployments emphasize different regions of the input distribution, yet prior work mainly opti… ▽ More

    Submitted 21 August, 2026; originally announced August 2026.

    Comments: 20 pages, 5 figures. Accepted to EMNLP Findings 2026

  6. arXiv:2608.20999  [pdf, ps, other

    cs.CV

    Latent Ordinal Evidence, Misaligned Outputs: Inference-Time Ordinal Lens Alignment for Multimodal LLMs

    Authors: Haiming Li, Yingsheng Liu, Jingmin Zhu, Siyuan Yan, Xieji Li, Jiajun Sun, Zhen Yu, Zongyuan Ge

    Abstract: Multimodal LLMs apply the language model interface to visual inputs, where ordinal regression tasks such as age estimation, image quality assessment, and disease grading require autoregressive decisions over ordered class labels. We ask whether MLLMs reliably convert internal ordinal evidence into ordered digit-token outputs. Across four ordinal benchmarks and four MLLM backbones, ordinal labels a… ▽ More

    Submitted 21 August, 2026; originally announced August 2026.

    Comments: EMNLP 2026 (Main Conference)

  7. Beyond Truth Discovery: A Two-Stage Framework to Assess the Severity of False Claim during Disasters

    Authors: Ruichen Yao, Tejna Dasari, Gulshat Baispay, Aizhan Zaurbek, Yifan Liu, Yaokun Liu, Zelin Li, Dong Wang

    Abstract: False information spreads rapidly on social media during disasters and can undermine emergency response efforts, public trust, and crisis communication. Existing research primarily focuses on determining whether social media posts contain false information, but provides limited insight into the specific false claims embedded within posts and the severity of individual false claims. To address the… ▽ More

    Submitted 21 August, 2026; originally announced August 2026.

    Comments: The 2026 ACM Conference on Human-AI Complementarity and Alignment

  8. arXiv:2608.20913  [pdf, ps, other

    cs.CV cs.AI

    Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation Optimization

    Authors: Zhu Xu, Jiaqi Tang, Pokai Chen, Yuxin Peng, Yang Liu

    Abstract: Explainable deepfake detection extends binary classification by requiring models to not only predict authenticity but also provide interpretable justifications. This expanded scope is critical in practice, where users like forensic analysts need insight into the rationale behind the detection. Despite advancements, current approaches suffer from two critical deficiencies: (1)vulnerability to image… ▽ More

    Submitted 21 August, 2026; originally announced August 2026.

  9. arXiv:2608.20820  [pdf, ps, other

    cs.AI

    Certified Multi-Turn Robustness for LLM Safety via Compositional Bounds and Safety Persistence

    Authors: Yang Liu, Bin Chong, Wenkai Yang, Shuai Zhang, Yancheng Chen, Feiyu Han, GuoZhen, Cheng Zhang, Huaibing Xie, Changze Lv, Shihan Dou, Pluto Zhou

    Abstract: Large language models (LLMs) are vulnerable to multi-turn jailbreak attacks that progressively manipulate conversation context. Existing certified robustness methods are limited to single-turn inputs; naive multi-turn composition yields bounds that degrade exponentially in the number of turns. We introduce Multi-Turn Certified Robustness (MTCR), a framework that models conversational safety via St… ▽ More

    Submitted 21 August, 2026; originally announced August 2026.

  10. arXiv:2608.20749  [pdf, ps, other

    cs.CV

    Identity-Preserving Text-to-Video Generation via Agentic Enhancement and Semantic Repair

    Authors: Jiayi Gao, Changcheng Hua, Jiaqi Tang, Yuxin Peng, Yang Liu

    Abstract: Identity-preserving video generation aims to synthesize videos that follow natural-language instructions while maintaining the visual identity of a given subject. Recent commercial video generation models have achieved strong visual quality and motion realism, but they still suffer from identity drift, incomplete instruction following, and missing visual details under complex prompts. Since these… ▽ More

    Submitted 21 August, 2026; originally announced August 2026.

  11. arXiv:2608.20738  [pdf, ps, other

    cs.AI

    Continuous-Time Quantum Walks based Graph Neural Network

    Authors: Yuliang Zhan, Zefeng Gao, Jian Li, Yang Liu, Hao sun

    Abstract: Graph Neural Networks (GNNs) are widely used on graph-structured data, but most suffer from two key weaknesses. First, message passing behaves as a low-pass filter under the homophily assumption, leading to poor performance on heterophilic graphs. Second, stacking layers drives node features toward constants, causing over-smoothing. Existing methods usually address these issues separately, while t… ▽ More

    Submitted 21 August, 2026; originally announced August 2026.

    Journal ref: CIKM 2026

  12. arXiv:2608.20711  [pdf, ps, other

    cs.CL

    AsmEvo: Agentic Assembly-Level Optimization of AMD GPU Kernels with Functional Equivalence Verification

    Authors: Ji Liu, Puyuan Yang, Rongzhang Zheng, Fan Wang, Jinglin Wang, Muhammad A. Awad, Mortis Huang, Andy Chang, Zekai Li, Zeping Li, Zihao An, Yue Liu, Yuchen Yang, Jianghui Wang, Chushi Chen, Ziqiong Liu, Fuwei Yang, Dong Li, Wen Heng Chung, Shengcai Liu, Emad Barsoum

    Abstract: High-performance ML systems increasingly rely on GPU kernels whose editable source is unavailable, generated, or too distant from final machine code to expose remaining optimizations. Existing LLM kernel optimizers and autotuners mainly operate on CUDA, Triton, HIP, or tensor-program source and validate against reference implementations. We study a stricter setting: optimizing an already compiled… ▽ More

    Submitted 20 August, 2026; originally announced August 2026.

  13. arXiv:2608.20611  [pdf, ps, other

    cs.AI

    Difficulty-Aware Semantic-ID Optimization for Generative Recommendation

    Authors: Xin Yu, Stephen Li, Sina Aghaei, Zifan Zhu, Jiamu Bai, Guanjie Huang, Bo Peng, Yiyao Liu, Lingzhou Xue

    Abstract: Semantic-ID-based generative recommendation casts retrieval and ranking as autoregressive generation over hierarchical item identifiers. A common recipe is SFT followed by GRPO, yet vanilla GRPO is poorly matched to this tree-structured task. Under the frozen SFT checkpoint, the exact target is absent from the first 16 candidates of the 50-beam constrained ranking for many prompts, and in harder c… ▽ More

    Submitted 20 August, 2026; originally announced August 2026.

  14. arXiv:2608.20375  [pdf, ps, other

    cs.CL

    GRAFT: Adaptive DLM-Based Draft Tree Construction with Target-Distilled Edge Scoring

    Authors: Xuming Ye, Zeming Ma, Runjie Yu, Yuan Liu, Tianle Li, Shuhan Bai, Jian Zhou, Fei Wu

    Abstract: Tree-based speculative decoding raises the mean accepted tokens of standard speculative decoding by verifying multiple draft paths, and existing tree builders typically construct these paths through parent-conditioned expansion, where each child token is generated conditioned on its parent path. This construction is incompatible with diffusion language model (DLM) drafters such as DFlash, which pr… ▽ More

    Submitted 23 June, 2026; originally announced August 2026.

  15. arXiv:2608.20350  [pdf, ps, other

    cs.CL cs.AI

    How to Train a Real-World Silicon Concierge? Internalizing Complex Business Workflow to Only OneModel

    Authors: Chang Liu, Chaoyang Ning, Dayi Jiang, Enrui Gu, Fang Ran, Hongyan Xue, Huaqing Li, Hui Cai, Jia Liu, Jiang-Ming Yang, Jianshe Li, Jiawei Luo, Jin Zhou, Leshen Zhu, Lihui Chen, Liying Ma, Lyuxin Xue, Mengjian Ji, Ruijia Xu, Wei Ren, Wei Wu, Xiaoling Qu, Xiaoyun Feng, Xin Zhang, Xixie Zhou , et al. (10 additional authors not shown)

    Abstract: Traditional industrial agents rely on modular pipelines, including Router, Retriever, Planner, Executor, Responder, Reviewer, and other components. These systems often fracture into a labyrinth of ad-hoc patches, leading to cascading errors and high latency. We propose OneModel, an applicable paradigm shift from external workflows to internalized knowledge representation. Unlike modular systems th… ▽ More

    Submitted 15 June, 2026; originally announced August 2026.

    Comments: Accepted to the ACL 2026 Industry Track (Oral). To appear in Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Industry Track)

  16. arXiv:2608.20308  [pdf, ps, other

    cs.CV

    DreamHand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery

    Authors: Yufei Liu, Xixi Wang, Hao Li, Ganlong Zhao, Kaitong Cai, Chengkai Jin, Chunxiao Liu, Jianbo Liu, Siyuan Huang, Xingang Pan, Hongsheng Li

    Abstract: Egocentric video offers scalable manipulation data for embodied AI, yet recovering metric 3D hand trajectories remains challenging due to severe object occlusion and frequent out-of-sight gaps. Existing single-frame and windowed temporal regressors fail when hand shortly leaves the frame, while recent video diffusion models (VDMs) rely on heavy, stochastic multi-step sampling as pixel-space render… ▽ More

    Submitted 20 August, 2026; originally announced August 2026.

    Comments: Project Page: https://ggxxii.github.io/dreamhand/

  17. arXiv:2608.20237  [pdf, ps, other

    cs.AI

    Rule-Compliant Visual Spatial Planning for Multimodal Large Language Models

    Authors: Yu Chen, Ting Lei, Yaoyi Li, Jia Cai, Zhecen Wu, Yang Liu

    Abstract: Multimodal large language models (MLLMs) combine linguistic reasoning with visual perception, yet their ability to perform visual spatial planning under explicit or previously unseen rule constraints remains underexplored. This setting requires models to jointly understand spatial layouts, interpret natural-language rules, and plan valid actions accordingly. To address this gap, we introduce RuleM… ▽ More

    Submitted 20 August, 2026; originally announced August 2026.

  18. arXiv:2608.20160  [pdf, ps, other

    cs.CR cs.CY

    Chameleon: Robust Defense Against Tor Website Fingerprinting via Many-to-Many Traffic Morphing

    Authors: Yuwen Cui, Kai Wei, Kehan Shen, Ning Wang, Zhuo Lu, Yao Liu, Guangjing Wang

    Abstract: Website fingerprinting (WF) attacks can infer users' browsing activities from encrypted Tor traffic by exploiting side-channel features. Although many WF defenses have been proposed, we find that most existing defenses create learnable web trace mapping features. We further show that robustness against adversarial training does not necessarily imply robustness against defense-aware autoencoder (DA… ▽ More

    Submitted 20 August, 2026; originally announced August 2026.

  19. arXiv:2608.20127  [pdf, ps, other

    cs.CV

    ID-VTG: Image-Disambiguated Video Temporal Grounding

    Authors: Minghang Zheng, Jingli Wei, Hongyi Yang, Yang Liu

    Abstract: Video Temporal Grounding (VTG) faces significant challenges when natural language queries must distinguish between multiple events involving visually similar entities, particularly when relying on fine-grained visual attributes that are difficult to describe accurately in words alone. To address this, we introduce Image-Disambiguated Video Temporal Grounding (ID-VTG), a task that leverages multimo… ▽ More

    Submitted 20 August, 2026; originally announced August 2026.

    Comments: ACM-MM 2026

  20. arXiv:2608.20019  [pdf, ps, other

    cs.AI

    Contrastive Mixed Prompt Learning for Incomplete Multimodal Sentiment Analysis with Unseen Modality Combination

    Authors: Kaixin Xu, NaiJin Liu, Yulin Kang, Tangyue Jin, Zixuan Yu, Wenxi Zhao, Yibei Liu, Qianle Zhang, Yangyang Wu, Mengying Zhu, Meng Xi

    Abstract: Incomplete multimodal sentiment analysis has garnered significant attention in recent years. Existing approaches typically assume that data is missing at random or are designed specifically for certain missing patterns, ignoring the modality combination inconsistency between training and testing phases. However, in real-world scenarios, the testing phase often encounters modal combinations that we… ▽ More

    Submitted 20 August, 2026; originally announced August 2026.

  21. arXiv:2608.19901  [pdf, ps, other

    cs.CR cs.AI

    MaliciousSkillBench: A Comprehensive Benchmark for Malicious Agent Skill Detection

    Authors: Yue Wang, Yi Liu, Gelei Deng, Ying Zhang, Yuekang Li, Zhenyu Chen, Leo Zhang

    Abstract: Agent Skills extend LLM agents with reusable instruction packages that may also include scripts, resources, and service configuration. This creates a direct distribution channel for malicious behavior, yet existing malicious-Skill datasets are fragmented across sources, artifact formats, evidence regimes, and benign coverage; duplicated and structurally related content further complicates direct a… ▽ More

    Submitted 20 August, 2026; originally announced August 2026.

    Comments: Project page: https://protectskills.github.io/MaliciousSkillBench/

  22. arXiv:2608.19869  [pdf, ps, other

    cs.IT cs.DM

    Sub-optimality of Marton's Inner Bound for the Two-Receiver Broadcast Channel

    Authors: Mian Huang, Yanxiao Liu, Yi Liu

    Abstract: Marton's inner bound, the best-known achievable region for a general discrete memoryless broadcast channel, was proposed by Katalin Marton in 1979, and whether it always achieves the capacity region has remained open since then. In this paper, we establish its strict sub-optimality: we show that the capacity region of some discrete memoryless broadcast channels can be strictly larger than Marton's… ▽ More

    Submitted 20 August, 2026; originally announced August 2026.

  23. arXiv:2608.19842  [pdf, ps, other

    cs.AI

    SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning

    Authors: Dayang Liang, Lang Feng, Bo An, Yunlong Liu

    Abstract: Agentic reinforcement learning (RL) has become a critical stage in the post-training of large language models. Existing critic-free, group-relative methods estimate policy advantages from multiple rollouts, avoiding the substantial memory overhead of conventional proximal policy optimization (PPO) and achieving strong performance on long-horizon interactive tasks. Despite their success, recent stu… ▽ More

    Submitted 20 August, 2026; originally announced August 2026.

  24. arXiv:2608.19750  [pdf, ps, other

    cs.CR

    TGL-APT: Temporal Graph Learning with Graph Distillation for Efficient APT Investigation

    Authors: Jing Chen, Ayong Ye, Yuanhuang Liu, Yuexin Zhang

    Abstract: Advanced Persistent Threat (APT) attacks pose a critical challenge to modern systems, as their stealthy, multi-stage nature renders conventional detection methods ineffective. While provenance graphs provide rich behavioral context for attack investigation, attack-relevant evidence is often sparse and embedded in large volumes of routine system activity, making full-graph learning both computation… ▽ More

    Submitted 20 August, 2026; originally announced August 2026.

  25. arXiv:2608.19738  [pdf, ps, other

    cs.CV cs.AI

    Learning to Beat: Phenotype-Guided Latent Flow with Regional Motion Priors for Biventricular Motion Synthesis

    Authors: Xuan Yang, Xiaohan Yuan, Hao Li, Lingyu Chen, Yanan Liu, Qingya Li, Lei Li

    Abstract: Full-cycle biventricular geometry is essential for characterizing cardiac function. However, dense and temporally consistent 3D+t biventricular meshes are not routinely available, whereas end-diastolic (ED) anatomy can often be obtained reliably. We therefore investigate full-cycle biventricular motion synthesis from a single ED mesh. This task is challenging because cardiac deformation is spatial… ▽ More

    Submitted 20 August, 2026; originally announced August 2026.

    Comments: 14pages, 10 figures

  26. arXiv:2608.19621  [pdf, ps, other

    cs.CL

    Mitigating Identity Essentialism in LLM Agents with Longitudinal Life Trajectories

    Authors: Hexi Wang, Yujia Zhou, Bangde Du, Weihang Su, Xinyuan Cao, Qingyi Pan, Qingyao Ai, Yueyue Wu, Min Zhang, Yiqun Liu

    Abstract: Large language models (LLMs) offer a scalable approach to social simulation, but their credibility depends on how agents are constructed. Existing methods can partially reproduce population-level patterns, yet often fail to capture human-like diversity. Our analysis shows that static-profile agents exhibit stronger demographic separation and within-group compression than humans, a pattern consiste… ▽ More

    Submitted 21 August, 2026; v1 submitted 20 August, 2026; originally announced August 2026.

    Comments: 23 pages, 12 figures

    MSC Class: 68T50 ACM Class: I.2.7

  27. arXiv:2608.19430  [pdf, ps, other

    cs.IR cs.CL cs.CR

    HARP: Hierarchical Adaptive Ranking with Preference-Adaptive Fusion for Query-Based CVE Prioritization

    Authors: Haochen Liu, Zhengzhang Chen, Haoyu Wang, Yanchi Liu, Jundong Li, Haifeng Chen

    Abstract: Vulnerability prioritization is inherently preference dependent, since the same CVE can receive different remediation priority under different operational preference scenarios. Existing scoring systems and ranking methods typically assume a fixed criterion. In practice, organizations already operate under a preference scenario, but this preference is often implicit and difficult to express as a wr… ▽ More

    Submitted 19 August, 2026; originally announced August 2026.

  28. arXiv:2608.19355  [pdf, ps, other

    cs.MM cs.CV

    GRACE: Grounded Reasoning via Adapter Composition and Evidence-Aware Calibration for Educational Visual Question Answering

    Authors: Xinjin Li, Yudi Xia, Xi Zhao, Yiliu Xu, Yining Liu, Cheng Lu, Yujian Long, Yu Ma, Jinghan Cao, Liang Fan, Yeyun Xu

    Abstract: Educational visual question answering, or VQA, requires models to solve curriculum-oriented multiple-choice questions using both language and visual evidence. Compared with conventional open-ended VQA, educational examples often include structured assessment metadata, diagrams or image contexts, and semantically close answer options, creating strong opportunities for question-option shortcuts. We… ▽ More

    Submitted 19 August, 2026; originally announced August 2026.

  29. arXiv:2608.19082  [pdf, ps, other

    stat.ML cs.LG stat.AP

    Learning Random Geometric Graphs Drawn in Probabilistic Metric Spaces

    Authors: Dalia Chakrabarty, Kangrui Wang, Chuqiao Zhang, Ye Liu

    Abstract: We present a new data-driven learning of a Random Geometric Graph (RGG) of a multivariate dataset, where the graph is drawn in a probabilistic metric space. This graph learning works for generic datasets, irrespective of the type of the observables; their probability distributions; or size of the data. We identify a metric of the space that the graph is drawn in, as a probability distribution of a… ▽ More

    Submitted 19 August, 2026; originally announced August 2026.

    MSC Class: 60-XX (Primary) 05C12; 62H20 (Secondary) }

  30. arXiv:2608.18903  [pdf, ps, other

    cs.LG

    A FEM-Based Surrogate Modelling and Optimization Framework for Physics-Constrained Electromagnetic Coil Design

    Authors: Yucheng Liu

    Abstract: This work evaluates surrogate-assisted optimization of a seven-parameter current-excited coil--core benchmark subject to geometric, manufacturing, and separate core and copper mass constraints. A Python--MPh--COMSOL workflow couples a two-dimensional axisymmetric finite-element method (FEM) model to a Matern 5/2 Gaussian-process (GP) probabilistic surrogate. Here, physics-constrained denotes a des… ▽ More

    Submitted 19 August, 2026; originally announced August 2026.

    Comments: 18 pages, 9 figures

  31. arXiv:2608.18839  [pdf, ps, other

    quant-ph cs.CC cs.IT

    Quantum Mixedness Testing with Pauli Measurements

    Authors: Jayadev Acharya, Abhilash Dharmavarapu, Yuhan Liu, Nengkun Yu

    Abstract: We consider a fundamental problem of \emph{mixedness testing}: Given $n$ copies of an $N$-qubit state $ρ$, determine whether $ρ= \mathbb{I}_d/d$ or $\|ρ-\mathbb{I}_d/d\|_1 \geq \varepsilon$ with high probability, where $d = 2^N$. In particular, we focus on performing this task in the practical setting of single-qubit measurements, where measurements are prepared independently on each qubit. We pro… ▽ More

    Submitted 19 August, 2026; originally announced August 2026.

    ACM Class: E.4; F.2.0; G.3

  32. arXiv:2608.18764  [pdf, ps, other

    cs.IR

    GateDiffInt: Gate-Mediated Controllable Diffusion and Multi-Intent LLM Distillation for User Behavior Modeling

    Authors: Jialong Duan, Zichen Zhang, Zirui Tu, Zheng Zhang, Zepeng Li, Qingyao Cui, Qinwen Wang, Yudan Liu, Luo Yang, Yao Hu

    Abstract: Existing ranking models encode intent only implicitly, making it hard to disentangle structured intents of varying strength and temporal scale. Noise and intent in behavior sequences are mutually reinforcing---we call this Noise--Intent Coupling (NIC). Noise dilutes true intents, while the lack of structured intent priors leaves denoising without a clear target. To address NIC, we propose GateDiff… ▽ More

    Submitted 19 August, 2026; v1 submitted 19 August, 2026; originally announced August 2026.

  33. arXiv:2608.18685  [pdf, ps, other

    cs.CV

    DocClaw: A Unified Agentic System for Intelligent Document Processing

    Authors: Siqi Xiang, Zhipeng Xu, Yufei Liu, Junhao Ji, Qing Liu, Zulong Chen, Zhibo Yang, Chunyan Miao, Shijian Lu

    Abstract: Intelligent document processing (IDP) encompasses a broad range of tasks, including optical character recognition (OCR), document question answering (DocQA), and key information extraction (KIE). Despite their distinct objectives, these tasks share a common need to perceive document content, acquire task-relevant information, and progressively refine intermediate results. However, they are typical… ▽ More

    Submitted 19 August, 2026; originally announced August 2026.

  34. arXiv:2608.18665  [pdf, ps, other

    cs.AI

    Candidate-Fate Accounting for Transparent Sensor Diagnostic Pipeline Search

    Authors: Haotao Xie, Yutian Chen, Yangqi Liu, Xiaoyu Jiang

    Abstract: Industrial sensor diagnostics relies on preprocessing, representation, and classification pipelines, making automated pipeline search useful for reducing manual design cost. However, existing automated machine/deep learning (AutoML/AutoDL) reports typically retain only fitted trials, scores, and winners, omitting generated candidates that are invalid, pruned, skipped, cached, or unfitted. This omi… ▽ More

    Submitted 19 August, 2026; originally announced August 2026.

  35. arXiv:2608.18637  [pdf, ps, other

    cs.IR

    PILOT Technical Report

    Authors: Jiuning Lin, Ruiquan Lan, Xiaodong Zhu, Bin Zhang, Chengyu Lai, Chuxin Chen, Dimin Wang, Han Zhu, Hongtao Cheng, Jialin Zhu, Lingqing Zhang, Shuai Zhong, Tao Wang, Weipeng Huang, Yinjiang Cai, Yinnan Song, Yuan Liu, Zhibo Xiao, Zhixin Ma, Zihong Huang

    Abstract: Existing agentic approaches for recommendation system optimization remain fundamentally reactive: they adjust parameters in response to observed metric changes but lack the ability to proactively design controlled experiments, personalize strategies at the user-segment level, or accumulate reusable experimental methodology across tasks. We present PILOT (Proactive Insight Learner for Online Tree-E… ▽ More

    Submitted 19 August, 2026; v1 submitted 19 August, 2026; originally announced August 2026.

    Comments: Technical Report, 42 pages, 10 figures

  36. arXiv:2608.18618  [pdf, ps, other

    cs.RO

    LabDex: A Hierarchical Benchmark for Dexterous Manipulation in Laboratories

    Authors: Zhipeng Tang, Sihang Chen, Sha Zhang, Peihao Yang, Yan Liu, Wentao Zhao, Xinrui Liu, Rui Huang, Wensheng Du, Yuting Huang, Jiajun Deng, Lidian Wang, Yuan Zhang, Yanyong Zhang

    Abstract: Autonomous laboratories hold great promise for accelerating scientific discovery. To achieve this vision, robots are supposed to dexterously manipulate diverse labware and instruments and execute long-horizon, state-dependent experimental procedures. Yet existing benchmarks do not jointly capture dexterous hand use, real-world laboratory interactions, and multi-stage experimental procedures, limit… ▽ More

    Submitted 19 August, 2026; originally announced August 2026.

    Comments: 13 pages, 3 figures

  37. arXiv:2608.18274  [pdf, ps, other

    cs.CR cs.LG

    Model Card for OpenAI Privacy Filter

    Authors: Charles de Bourcy, Sahra Ghalebikesabi, Avi Schwarzschild, Alex Gorbachev, Mihai Maruseac, Annie Chu, Vol Kyrylov, Tong Mu, Ally Bennett, Andy Nguyen, Casey Meehan, Jessica Gan Lee, Shane Bauer, Harold Nguyen, Rodolpho Eckhardt, Yuqi Liu, Charlie Oxborough, Marco Rougeth, Omar Chedid, Caio Costa, Yash Parikh, Yao Li, Congzheng Song, Om Thakkar, Vinnie Monaco

    Abstract: OpenAI Privacy Filter is a compact, bidirectional token-classification model for detecting and redacting personally identifiable information (PII) and secrets in unstructured text. The model is derived from an autoregressively pretrained checkpoint and converted into a bidirectional, banded-attention classifier that labels an input sequence in a single forward pass. A constrained Viterbi decoder p… ▽ More

    Submitted 18 August, 2026; originally announced August 2026.

    Comments: 20 pages, 3 figures, 11 tables

  38. arXiv:2608.18234  [pdf, ps, other

    cs.RO cs.AI cs.LG

    GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction

    Authors: Ziyang Cheng, Tianshu Tang, Jinxin Lan, Xinze Chen, Yuhan Gong, Zhichao Liu, Changzhong Wu, Yahao Mao, Zongyan Deng, Mingxuan Ma, Huasen Xi, Yilong Liu, Yutong Wu, Xiaofeng Wang, Yang Wang, Yun Ye, Guan Huang, Xiaojie Jin, Zheng Zhu, Jiwen Lu

    Abstract: Whole-body motion tracking policies turn a humanoid into a robust control interface: the teleoperator---or an upstream model---only supplies a coarse movement intent, while the low-level policy keeps the robot balanced and physically feasible. Existing trackers deliver this interface only on flat ground: trained in empty scenes, they never learn how contact with terrain and objects reshapes their… ▽ More

    Submitted 18 August, 2026; originally announced August 2026.

    Comments: 20 pages, 8 figures, 4 tables. Technical report. Project page: https://shepherd1226.github.io/gigabrain-wbc-0.5/

  39. arXiv:2608.18187  [pdf, ps, other

    cs.CR

    AdaRare: Telemetry-Guided Joint Profile Control for Greybox Fuzzing

    Authors: Jingchuan Ma, Tongan Liu, Yanhua Liu, Qiaoyun Huang

    Abstract: Greybox fuzzers combine interacting queue, mutation, dictionary, energy, and comparison-solving control surfaces, while prior adaptive systems typically optimize other decision objects or control layers. We present AdaRare, an AFL++ extension that coordinates five internal actuation mechanisms as one bounded in-process profile updated every 5,000 ms. Completed-window, action-induced telemetry feed… ▽ More

    Submitted 18 August, 2026; originally announced August 2026.

    Comments: 20 pages, 4 figures

  40. arXiv:2608.18185  [pdf, ps, other

    cs.LG cs.CV

    H$^2$EDL: Hyper Evidential Deep Learning for Hierarchical Classification

    Authors: Yuanye Liu, Xiahai Zhuang

    Abstract: Fine-grained recognition often involves hierarchical label spaces, where a model may be confident about a coarse semantic concept while remaining uncertain among its descendant classes. Such structured ambiguity requires uncertainty representations that capture both fine-grained classes and intermediate concepts. However, existing tools each capture only half of it: flat evidential classifiers qua… ▽ More

    Submitted 18 August, 2026; originally announced August 2026.

  41. arXiv:2608.18096  [pdf, ps, other

    cs.CL cs.LG

    MAVEN: A Macro-Societal Value Evaluation Framework of Multimodal Content with Compact Aligned Evaluators

    Authors: Zijuan Zhao, Zheren Fu, Hou Xia, Licheng Zhang, Yi Liu, Zhendong Mao

    Abstract: Assessing whether multimodal content aligns with macro-societal values, such as peace, justice, and freedom, has become an increasingly urgent challenge. Existing frameworks are largely confined to safety-oriented taxonomies, text-only psychometric probes, or single-label classification. Therefore, we propose MAVEN, a hierarchical framework for macro-societal value evaluation of multimodal content… ▽ More

    Submitted 8 June, 2026; originally announced August 2026.

    Comments: 18 pages, 6 figures

    ACM Class: I.2.7; I.2.10; K.4.1

  42. arXiv:2608.18034  [pdf, ps, other

    cs.CV

    Deep Academic Survey: Stateful Agentic Closed-Loop Paradigm for Academic Survey Automation

    Authors: Zhikai Xu, Zhucun Xue, Teng Hu, Yabiao Wang, Yong Liu, Jiangning Zhang

    Abstract: Academic surveys play a central role in organizing rapidly expanding scholarly literature, yet their construction requires extensive paper analysis, coherent knowledge organization, fine-grained citation support, and reliable manuscript assembly. Existing Deep Research and automated survey generation systems address parts of this process, but typically do not coordinate paper understanding, litera… ▽ More

    Submitted 18 August, 2026; originally announced August 2026.

    Comments: Project page: https://zhikaixu24.github.io/projects/DAS/ | Code: https://github.com/ZhikaiXu24/DAS | Data: https://huggingface.co/datasets/ZhikaiXu24/DAS-2M

  43. arXiv:2608.17834  [pdf, ps, other

    cs.HC cs.AI

    AdaLens: Interactive Storyline for Monitoring and Steering Long-Running Agentic Data Analysis

    Authors: Yangtian Liu, Yan Miao, Shuhan Liu, Yunfan Zhou, Dae Hyun Kim, Di Weng, Yingcai Wu

    Abstract: Large language models are pushing data science toward increasingly autonomous and agentic workflows, with recent systems already supporting multi-step and long-running analyses. As these workflows become more autonomous, conventional interfaces no longer provide adequate support for two critical requirements: observability for understanding an agent's evolving reasoning and evidence, and steerabil… ▽ More

    Submitted 18 August, 2026; originally announced August 2026.

  44. arXiv:2608.17756  [pdf, ps, other

    cs.AI

    D$^2$ACCI: A Dual-Loop Diagnostic Protocol for Evidence-Preserving Agent Memory

    Authors: Xule Liu, Yijun Liu, Chao Li, Shao Kun

    Abstract: Memory is a key capability of LLM agents. Persistent memory extends this across sessions---enabling recall, revision, and personalization. Yet its multi-stage pipeline (ingestion, retrieval, filtering, generation) makes failures difficult to localize: end-to-end evaluation reveals that an error occurred, but not which stage caused it. Existing evaluations often report aggregate performance without… ▽ More

    Submitted 18 August, 2026; v1 submitted 18 August, 2026; originally announced August 2026.

    Comments: Preprint

  45. arXiv:2608.17671  [pdf, ps, other

    cs.SE cs.AI cs.CR

    Benchmarking Automated Security Patch Backporting: How Far Are We?

    Authors: Jincheng Yang, Yulong Fu, Chengwei Liu, Lyuye Zhang, Fangyuan Zhang, Bingyang Ren, Yang Liu, Hui Li

    Abstract: Automated security patch backporting is critical for mitigating N-day vulnerabilities. Recent tools report success rates above 80% on their respective datasets. However, these evaluations are often confined to homogeneous environments, such as one repository or specific project versions. Consequently, it remains unclear how well these tools generalize beyond their originally targeted scenarios. We… ▽ More

    Submitted 18 August, 2026; originally announced August 2026.

    Comments: 13 pages, 3 figures. Accepted at ASE 2026. Artifact: https://doi.org/10.5281/zenodo.21785770

  46. arXiv:2608.17566  [pdf, ps, other

    cs.CV

    CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing

    Authors: Fuchen Long, Cong Wang, Zitao Gao, Wenhao Zhong, Yu Cheng, Xiaolu Hou, Yan Li, Xiao Cao, Xinlong Sun, Xi Chen, Yu Liu

    Abstract: The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets mainly focus on single editing operations and fall short in supporting compositional instruction-guided video editing. In particular, multiple editing intents must be jointly understood and faithfully executed within the same video. To address this issue, we introduce CoinVE-200K, a… ▽ More

    Submitted 18 August, 2026; originally announced August 2026.

    Comments: Project page: https://coinve200k.github.io; Dataset is available at https://huggingface.co/datasets/FireCRT/CoinVE-200K; see source codes at https://github.com/coinve200k/CoinVE-200K

  47. arXiv:2608.17490  [pdf, ps, other

    cs.CV

    When More Foundation Models Means Less: Diagnosing and Addressing Multi-View Fusion Failure

    Authors: Yibo Liu, Bowen Jiang

    Abstract: Foundation-model hubs turn multi-view fusion into a selection problem: from a large heterogeneous encoder pool, which views should be fused, and how many? We show that downstream performance is non-monotonic in the number of fused encoders; later views can be redundant or task-misaligned, causing accuracy to saturate or decline. We formalise this setting as view-set composition and propose KAGES (… ▽ More

    Submitted 18 August, 2026; originally announced August 2026.

    Comments: 26 pages, 4 figures. Code and results: https://github.com/yibol9768-alt/Quantifying-Representation-Reliability

  48. arXiv:2608.17453  [pdf, ps, other

    cs.RO

    EATR-Stereo: Embodiment-Aware Token Routing of Paired Stereo Evidence for Humanoid Vision-Language-Action Control

    Authors: Songwei Wu, Rui Zhao, Fan Yang, Zhongqiang Nie, Zhiduo Jiang, Wandong Sun, Yuwei Li, Jian Hu, Yang Liu, Hong Liu

    Abstract: Long-horizon humanoid vision--language--action (VLA) control with head-mounted stereo cameras requires visual interfaces that can exploit complementary views while maintaining compatibility with pretrained representations. Existing interfaces often discard complementary stereo evidence or fuse additional observations without preserving the native primary-view pathway and adapting auxiliary informa… ▽ More

    Submitted 20 August, 2026; v1 submitted 18 August, 2026; originally announced August 2026.

    Comments: 8 pages, 5 figures

  49. arXiv:2608.17414  [pdf, ps, other

    cs.CV cs.PL

    REChart: Reasoning-Efficient Chart Editing with Large Reasoning Models

    Authors: Yuanbang Liu, Chenxi Ruan, Yihan Hou, Qiong Luo, Wei Zeng

    Abstract: Chart editing requires inferring and modifying visualization code from a reference chart image based on an editing instruction, challenging fine-grained visual reasoning, instruction following, and executable code synthesis capabilities of MLLMs. Large reasoning models (LRMs) with extended Chain-of-Thought (CoT) reasoning are suitable for tackling such complex multimodal tasks. However, our prelim… ▽ More

    Submitted 18 August, 2026; originally announced August 2026.

  50. arXiv:2608.17365  [pdf, ps, other

    math.CO cs.CC

    A Counting Lemma for Somewhat Restricted 3-APs

    Authors: Amey Bhangale, Subhash Khot, Yang P. Liu, Dor Minzer

    Abstract: For a prime $p\geq 3$, a somewhat restricted $3$-AP in $\mathbb{F}_p^n$ is a triplet $(x,x+a,x+2a)$, where $x\in\mathbb{F}_p^n$ and $a\in \{0,1,2\}^n$. We prove a counting lemma for somewhat restricted $3$-APs in dense sets in $\mathbb{F}_p^n$. More precisely, we prove that for all $α>0$, there exists $β>0$, such that for sufficiently large $n$, if a set $A\subseteq \mathbb{F}_p^n$ has density at… ▽ More

    Submitted 18 August, 2026; originally announced August 2026.

    Comments: 58 pages