-
Inter-X++: A Comprehensive Benchmark for Multimodal Human-Human Interaction Analysis
Authors:
Liang Xu,
Chengqun Yang,
Zili Lin,
Xintao Lv,
Yichao Yan,
Xin Jin,
Zhibo Chen,
Xiaokang Yang,
Wenjun Zeng
Abstract:
The capability to perceive and synthesize human-human interactions is fundamental to developing intelligent digital human systems. However, existing datasets and modeling approaches are fundamentally constrained by low-fidelity kinematics, the omission of dexterous hand gestures and a severe lack of rich multimodal annotations. Furthermore, fragmented interaction representations and inconsistent e…
▽ More
The capability to perceive and synthesize human-human interactions is fundamental to developing intelligent digital human systems. However, existing datasets and modeling approaches are fundamentally constrained by low-fidelity kinematics, the omission of dexterous hand gestures and a severe lack of rich multimodal annotations. Furthermore, fragmented interaction representations and inconsistent evaluation protocols also impede fair and rigorous benchmarking. To systematically address these bottlenecks, we present Inter-X++, a comprehensive and large-scale benchmark designed to empower versatile HHI analysis. Captured via a novel hybrid motion capture system, Inter-X++ provides 11,388 high-fidelity interaction sequences and over 8.1M frames, featuring precise whole-body movements and detailed finger articulations. Meanwhile, we enrich the data foundation with multifaceted annotations, including hierarchical fine-grained textual descriptions, interaction categories, causal interaction orders, the relationship and personality of the subjects, as well as vertex-level contact maps and physically regularized constraints. Leveraging these elaborate annotations, we formulate a unified testing ground comprising four categories of downstream tasks that symmetrically span both generative and perceptive paradigms. To eliminate benchmarking ambiguities, we systematically standardize the interaction representations and evaluation protocols. Finally, we go beyond dataset construction to propose OpenHHI, a single and unified HHI representation and modeling framework that jointly optimizes interaction reconstruction and semantic understanding. Extensive experiments reveal that OpenHHI achieves state-of-the-art performance on both generation and perception tasks. This definitively proves that our unified representation successfully bridges interaction understanding and generation simultaneously.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL
Authors:
Langzhe Gu,
Chengkai Hou,
Meng Li,
Xinhua Wang,
Jiaming Liu,
Xinyuan Lv,
Bowei Zhang,
Shuanghao Bai,
Guangrun Li,
Jingyang He,
Gaole Dai,
Ziluo Ding,
Zhiyuan Xu,
Kuan Cheng,
Jian Tang,
Zhengping Che,
Shanghang Zhang
Abstract:
Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-language-action (VLA) foundation models are not readily applicable to humanoid whole-body loco-manipulation. The high dimensionality and interdependence of humanoid motions make it challenging for conventional single-stage VLA architectures to coordinate locomotion, waist posture, and…
▽ More
Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-language-action (VLA) foundation models are not readily applicable to humanoid whole-body loco-manipulation. The high dimensionality and interdependence of humanoid motions make it challenging for conventional single-stage VLA architectures to coordinate locomotion, waist posture, and dual-arm manipulation effectively. Moreover, policies trained through offline behavior cloning can remain suboptimal during real-world deployment. Although online reinforcement learning can refine policies through real-world interaction, directly tuning large VLA backbones demands excessive computation and may introduce safety risks during real-robot exploration. To address these bottlenecks, we introduce HAF (Humanoid Adaptation Framework), a two-part framework consisting of HAF-VLA and HAF-Steer that transfers off-the-shelf generalist VLA foundation models to humanoid whole-body loco-manipulation. HAF-VLA is a hierarchical action-flow generator built on a pretrained flow-matching VLA. It splits full-body action denoising into three sequential stages with stage embeddings and cross-stage KV caches that retain kinematic dependencies, avoiding incoherent whole-body actions from one-shot generation. On top of the frozen HAF-VLA, HAF-Steer is a latent offline-to-online RL pipeline that leverages flow-matching invertibility and DCT-based dimensionality reduction to restrict RL optimization to a compact noise subspace and train a regularized SAC policy. This avoids updating the large VLA backbone and enables efficient real-world policy refinement. Evaluated on seven real-world humanoid loco-manipulation tasks, HAF surpasses vanilla single-stage VLA baselines and improves whole-body coordination and task performance. Project website: https://grange007.github.io/HAF .
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation
Authors:
Ximo Zhu,
Ruiqi Liu,
Rong Wang,
Ping Wu,
Xiang Zheng,
Wenzhuo Xu,
Xubin Yao,
Zhiyuan Yan,
Bo Li,
Jun Gao,
Xiaolei Lv
Abstract:
On-policy distillation (OPD) applies token-level teacher supervision to student-generated trajectories, but this supervision is not always reliable. Existing methods use local confidence or teacher-student agreement to weight, filter, or truncate the sampled trajectory. These signals do not directly determine whether the teacher can continue a student prefix to a correct answer, and trajectory-lev…
▽ More
On-policy distillation (OPD) applies token-level teacher supervision to student-generated trajectories, but this supervision is not always reliable. Existing methods use local confidence or teacher-student agreement to weight, filter, or truncate the sampled trajectory. These signals do not directly determine whether the teacher can continue a student prefix to a correct answer, and trajectory-level interventions can conflate one rollout's unreliability with low expected training value of its prompt. We define prompt-level teacher continuation reliability $R$ as the teacher's probability of reaching a correct answer from a student prefix, averaged over prefixes and trajectories induced by the current student. Oracle experiments show that high-$R$ prompts yield larger OPD gains and that descending-$R$ training outperforms random and ascending orders on a fixed prompt pool. Because estimating $R$ requires many teacher continuations, we use the maximum ROUGE-5 F1 between one independent student rollout and verifier-correct same-prompt teacher trajectories. Across ten equal-frequency bins of this actual score, mean $R$ rises monotonically, showing that the proxy separates coarse reliability levels. ReOrder-OPD sorts prompts by the proxy, then draws independent on-policy training trajectories for vanilla OPD. It improves every matched aggregate comparison across Qwen3 and Gemma4 mathematics settings and Qwen3 code settings. Gains in all six FiRe-OPD and ExOPD settings show that prompt ordering complements within-trajectory supervision.
△ Less
Submitted 11 August, 2026;
originally announced August 2026.
-
Compiler Framework for 3D Neutral-Atom Quantum Computers
Authors:
Chen Huang,
Zhemin Zhang,
Zhao Zhang,
Xudong Lv,
Zhiding Liang
Abstract:
Neutral-atom quantum computers can now arrange atoms in three-dimensional tweezer arrays, yet every existing compiler assumes a flat geometry. We present Piqasso, a compiler that exploits the vertical axis by stacking storage, entanglement, and readout into distinct layers. Its pipeline pairs an analytical placement respecting axial-clearance optics with a router that brings gate partners together…
▽ More
Neutral-atom quantum computers can now arrange atoms in three-dimensional tweezer arrays, yet every existing compiler assumes a flat geometry. We present Piqasso, a compiler that exploits the vertical axis by stacking storage, entanglement, and readout into distinct layers. Its pipeline pairs an analytical placement respecting axial-clearance optics with a router that brings gate partners together via short vertical hops---bypassing in-plane crossing conflicts through out-of-plane detours---and a multi-AOD scheduler that parallelizes transport across focal planes. On 34 circuits, Piqasso reduces atom transport distance by 2.1$\times$ over a state-of-the-art planar compiler, yielding up to 7.3$\times$ faster execution, 2.2$\times$ higher movement fidelity, and 1.8$\times$ fewer serialized transport rounds, with all gains widening at scale.
△ Less
Submitted 2 August, 2026;
originally announced August 2026.
-
Evaluation-Verification Reward for Consistent Multi-Reference Image Editing
Authors:
Yingmao Miao,
Pengfei Zhang,
Xiaochen Lv,
Meng Yu,
Lei Sun,
Xiangxiang Chu,
Chao Shen,
Chenhao Lin
Abstract:
While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency across references and ensuring overall visual harmony. Reinforcement learning has proven highly effective for text-to-image generation and single-image editing, but its extension to multi-reference editing is hindered by the absence of suitable rew…
▽ More
While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency across references and ensuring overall visual harmony. Reinforcement learning has proven highly effective for text-to-image generation and single-image editing, but its extension to multi-reference editing is hindered by the absence of suitable reward models that capture multi-image relational constraints. Moreover, naively using multimodal large language models(MLLMs) as zero-shot evaluators faces a key tension between hallucination-prone long-form reasoning and the limited deductive power of short-form judgments. We address these issues with a Multi-dimensional Evaluation-Verification Reward(EVR). EVR decomposes evaluation into distinct visual criteria; for each criterion, an MLLM Evaluator generates multiple candidate hypotheses, and a Verifier grounds each claim in concrete visual evidence to accept or reject it, producing reliable and fine-grained reward signals. Together with a scalable data pipeline, our method enables RL fine-tuning of off-the-shelf editors without architectural changes. Extensive experiments show substantial gains over the base Qwen-Image-Edit, improving consistency and harmony to match or surpass NanoBanana.
△ Less
Submitted 5 August, 2026; v1 submitted 31 July, 2026;
originally announced July 2026.
-
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
Authors:
Junlin Yang,
Che Jiang,
Yu Fu,
Tianwei Luo,
Can Ren,
Weizhi Wang,
Kaikai Zhao,
Hongyi Liu,
Yuxin Zuo,
Yuru Wang,
Yuchen Fan,
Kai Tian,
Zhenzhao Yuan,
Xiaojian Lin,
Li Sheng,
Rushi Qiang,
Guoli Jia,
Xingtai Lv,
Ermo Hua,
Dianqiao Lei,
Youbang Sun,
Ning Ding,
Bowen Zhou,
Kaiyan Zhang
Abstract:
Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability. We introduce OpenMLE, an open full-stack system for RSI research in MLE, spanning verifiable task environments with execution feedback (OpenMLE-Gym), operator learning (OpenMLE-RL), and lon…
▽ More
Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability. We introduce OpenMLE, an open full-stack system for RSI research in MLE, spanning verifiable task environments with execution feedback (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo). On this stack we post-train Frontis-MA1 (35B) as a meta-evolution agent for MLE, aligning post-training and inference around four atomic program-evolution operators (Draft, Improve, Debug, Crossover): the same operators are trained via execution-grounded SFT and RL on data deduplicated against all evaluation benchmarks, then composed into long-horizon search, coupling learning and evolution in a single loop. On MLE-Bench Lite under a 12-hour per-task budget on one RTX 4090 capped at 12 GB VRAM, Frontis-MA1 (35B) improves Medal Average from 39.39% to 60.61% over its base model with OpenMLE-Evo, and reaches 71.21% with OpenMLE-Evo-Max (benchmark-independent experience priors and asynchronous search), exceeding GPT-5.5 + Codex and approaching GPT-5.6 Sol and the 2.8T Kimi K3. On held-out NatureBench Lite, both components transfer: with the framework fixed, swapping in the trained model raises Match-SOTA from 50% to 70%; with the model fixed, swapping in OpenMLE-Evo raises it from 20% to 50%. We release the model weights and the full OpenMLE stack to enable reproducible research on executable AI4AI toward RSI. Code: https://github.com/FrontisAI/OpenRSI
△ Less
Submitted 30 July, 2026;
originally announced July 2026.
-
Multi-Decoder OneRec: Controllable Generative Retrieval for Multi-Objective Industrial Recommendation
Authors:
You Wang,
Zhao Liu,
Guoping Tang,
Yiqing Yang,
Shuo Su,
Jing Liu,
Naifu Zhou,
Xiaoyou Zhou,
Wei Jiang,
Jian Liang,
Xiao Lv,
Ruiming Tang,
Liyin Hong,
Wenwu Ou
Abstract:
Industrial recommender systems build candidate pools by assigning explicit quotas to objective-specific retrieval routes. This design offers quota control but increasingly fragments modeling, training, and serving as the route set grows. Semantic-ID-based generative retrieval provides a unified alternative, yet a single decoder entangles objective policies and limits candidate complementarity. We…
▽ More
Industrial recommender systems build candidate pools by assigning explicit quotas to objective-specific retrieval routes. This design offers quota control but increasingly fragments modeling, training, and serving as the route set grows. Semantic-ID-based generative retrieval provides a unified alternative, yet a single decoder entangles objective policies and limits candidate complementarity. We propose Multi-Decoder OneRec, a controllable framework that combines shared representations, isolated objective adaptation, and coordinated decoding. All objectives share a user-context module and the General Decoder, while each objective adds an isolated, parameter-efficient LoRA expert. During training, exposure-sample next-token prediction (NTP) updates the shared base, target-filtered NTP updates the event-based experts, and Kullback-Leibler (KL)-regularized policy optimization updates the Watch-time expert; gradient routing isolates these updates, and the General Decoder supplies a stop-gradient reference. At inference, explicit route quotas allocate the fixed budget and Multi-Decoder Constrained Beam Search reduces cross-route overlap. We publicly release Kwai26, a large-scale multi-objective benchmark with 1.31 billion raw item-level records, 31.85 million Item-ID entries, and 25.03 million items with valid Semantic IDs, together with predefined splits and an evaluation protocol. Under the same 512-item retrieval budget, Multi-Decoder OneRec improves over the single-decoder OneRec baseline by 1.69%-5.62% across four Recall@512 metrics. In a production A/B test, it yields relative gains of 0.37% in app usage time per device, 0.19% in Day-7 retained users, 0.19% in devices with at least one share, and 2.09% in new-content Cold-Start. These results show that generative retrieval can combine shared modeling with objective-specific control and complementary candidate generation.
△ Less
Submitted 29 July, 2026;
originally announced July 2026.
-
FilmBench: A Film-Grade Benchmark for Cinematic Video Generation
Authors:
Shengyi Wang,
Niantong Li,
Guangzheng Hu,
Hong Qi,
Fei Ding,
Weixu Qiao,
Jinlin Wang,
Xiaotong Lv,
Peng Han,
Zimeng Li,
Fanshu Ding,
Yushu Wang,
Han Wu,
Jingjing Chen,
Chongxiao Wang,
Yanhao Wu,
Chenglong Huang,
Xiaoqian Zhu,
Jie Tian,
Hua Li,
Jingjing Fan,
Mingshuang Tang,
Zhong Li,
Hengxia Qiang,
Weibin Chen
, et al. (5 additional authors not shown)
Abstract:
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than…
▽ More
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they assess basic video plausibility rather than film-grade craft. We introduce FilmBench, a text-to-video (T2V) and reference-to-video (R2V) benchmark grounded in the professional Cinematic Language of the film- academy tradition and co-developed with directors and faculty from the Beijing Film Academy and the Hujing Digital Media & Entertainment Group film studio. It rests on three choices. First, prompts are reverse-engineered from clips of award-winning films spanning 20 cinematic genres and chosen by professional directors, so every prompt is anchored to a verified live-action reference; the prompts follow real shot lists, and most script multiple shots (1,056 of the 1,169 prompts are multi-shot), unlike prior single-clip benchmarks. Second, evaluation follows a three-level Cinematic taxonomy of 3 axes, 12 components and 35 (T2V) +3 (R2V-only) sub-metrics. Third, we develop an in-house expert-grade automatic evaluation agent and open-source its core suite of Cinematic Language operators (FilmOps). Benchmarking leading video generation models (9 for T2V, 7 for R2V), the evaluator reproduces the human model ranking at model-level Spearman \r{ho} = 0.95 (T2V) and 0.96 (R2V). Scores fall well below prior web-style benchmarks, with two consistent gaps in dynamic aesthetics and a marked single- to multi-shot performance drop that widens for weaker models.
△ Less
Submitted 29 July, 2026; v1 submitted 27 July, 2026;
originally announced July 2026.
-
Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm
Authors:
Bajian Xiang,
Cheng Wen,
Han Zhao,
Hao Wang,
Haoxu Wang,
Jiawei Jin,
Jiayan Cui,
Jie Chen,
Mengxi Nie,
Tianyu Zhao,
Weiqin Li,
Xiang Lv,
Xiangang Li,
Yang Xiang,
Yang Zhou
Abstract:
In this report, we present Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system that jointly advances content consistency, speaker similarity, prosodic naturalness, audio quality, controllability, multilingual coverage, efficiency, and robustness. It combines a 12.5~Hz low-frame-rate speech tokenizer for reduced inference latency with a five-stage progressive training paradigm for coo…
▽ More
In this report, we present Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system that jointly advances content consistency, speaker similarity, prosodic naturalness, audio quality, controllability, multilingual coverage, efficiency, and robustness. It combines a 12.5~Hz low-frame-rate speech tokenizer for reduced inference latency with a five-stage progressive training paradigm for coordinated language model (LM) and flow-matching model (FM) optimization. The model provides production-level control through free-style natural-language instructions and fine-grained inline tags, while supporting 16 languages, 20 Chinese dialect regions, one-pass long-form synthesis up to 3 minutes, and robust generation from noisy, reverberant, or unclear reference speech. Across SEED-TTS-Eval, CV3-Eval, instruction-following, long-form, and acoustic-robustness evaluations, Qwen-Audio-3.0-TTS achieves state-of-the-art performance on many reported dimensions or the strongest aggregate results. It also ranks first on the independent Artificial Analysis Text-to-Speech Leaderboard. These results establish Qwen-Audio-3.0-TTS as a strong foundation for production-level speech synthesis.
△ Less
Submitted 26 July, 2026;
originally announced July 2026.
-
VoLN: Vision-Only Long-Horizon Navigation---Paradigm, Benchmark, and Method
Authors:
Jiabin Lou,
Haopeng Wang,
Yuanshuai Wang,
Xinyu Liu,
Xuxin Lv,
Yuxin Guo,
Lei Huang,
Rongye Shi,
Wenjun Wu
Abstract:
Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions. However, route-level instructions commonly encode spatial priors, such as orientation, distance, and layout, that are not explicitly available from onboard sensing at deployment in open, GPS-denied environments. Benchmark performance under such interfaces therefore jointly reflects visual navigati…
▽ More
Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions. However, route-level instructions commonly encode spatial priors, such as orientation, distance, and layout, that are not explicitly available from onboard sensing at deployment in open, GPS-denied environments. Benchmark performance under such interfaces therefore jointly reflects visual navigation ability and the use of route structure explicitly supplied by the task description. As a complementary formulation, we propose Vision-Only Long-Horizon Navigation (VoLN), which shifts route-relevant information from externally supplied instructions and global guidance to locally observable in-scene cues. In VoLN, goal views specify the destination, while route-relevant information is available only through locally observable in-scene cues that the agent must detect, interpret, and select online. We instantiate VoLN for aerial navigation through VoLN-UAV, a 7,210-episode benchmark that combines long-horizon goal-directed flight, continuous 3D motion, large viewpoint changes, and context-dependent beacon selection. We further provide VoLN-MLLM as an initial reference baseline. It aligns self-supervised visual features with a structured semantic space and predicts short-horizon waypoint segments from observation history, goal views, retrieved visual--semantic tokens, and proprioception. On the five-environment Test-Unseen split, it obtains success rates of 7.4%, 4.5%, and 1.8% on Easy, Normal, and Hard episodes, respectively. These results provide an initial evaluation of VoLN and reveal substantial remaining challenges in long-horizon evidence integration, cross-view goal matching, and closed-loop stability. Project page: https://admire-ljb.github.io/VoLN-UAV/
△ Less
Submitted 23 July, 2026;
originally announced July 2026.
-
QuantiSpect: A Structure-Aware Lightweight 3D CNN Pre-Decoder for Scalable Surface Code Quantum Error Correction
Authors:
Pan Gao,
Xu-Sheng Xu,
Ji-Ze Han,
Jing-Wei Wen,
Ling Qian,
Xudong Lv,
Run-Qing Zhang,
Xiao-Xiao Hu,
Gui-Lu Long
Abstract:
Real-time decoding is a critical bottleneck for large-scale fault-tolerant quantum computing. AI-based neural pre-decoders locally correct most physical errors before passing residual syndromes to a global decoder, enabling sub-microsecond latencies. However, existing architectures carry significant overhead from dense 3D convolutions. We present QuantiSpect, a lightweight 3D convolutional neural…
▽ More
Real-time decoding is a critical bottleneck for large-scale fault-tolerant quantum computing. AI-based neural pre-decoders locally correct most physical errors before passing residual syndromes to a global decoder, enabling sub-microsecond latencies. However, existing architectures carry significant overhead from dense 3D convolutions. We present QuantiSpect, a lightweight 3D convolutional neural network (CNN) pre-decoder for the rotated surface code, built on the decoding pipeline of Chamberland et al. The key idea is to replace the dense 3D convolutions with three parallel branches in each residual block: a depthwise spatial branch, a depthwise temporal branch, and a grouped spatio-temporal branch, followed by a squeeze-and-excitation channel gate. This reflects the structure of surface code errors, where spatial and temporal syndrome correlations are partially separable. On a unified 4xA100 GPU benchmark, QuantiSpect matches the receptive field of the Accurate baseline at R=13 while using ~2.71x fewer parameters (0.663M vs 1.80M) and ~2.84x fewer per-voxel convolutional MACs. It matches Accurate's circuit-level threshold and accuracy at moderate and large code distances, reduces the logical error rate by up to ~1.85x relative to uncorrelated PyMatching at d=13, p=0.5%, and speeds up the PyMatching decode by up to 3.11x at d=23. We also explored enlarging the receptive field by adding blocks. Even at R=21, the model uses only 1.18M parameters, fewer than both the R=13 Accurate baseline (1.80M) and the R=17 dense model (4.22M), despite its larger receptive field. This expanded variant significantly outperforms the Accurate model, raising the circuit-level threshold to ~0.80% and further reducing the logical error rate. Together, both variants show that a structure-aware factorized design is an effective, parameter-efficient alternative to a dense one for decoding the surface code.
△ Less
Submitted 5 August, 2026; v1 submitted 20 July, 2026;
originally announced July 2026.
-
RECAP: Feedback-Driven Streaming Semantic User Profiles for Short-Video Recommendation
Authors:
Ziyi Zhao,
Xiaoyou Zhou,
Xiao Lv,
Yangyang Li,
Chubo He,
Zhao Liu,
Jiayao Shen,
Yuqi Liu,
He Li,
Chengyi Zhang,
Jian Liang,
Ming Li,
Chongming Gao,
Fuli Feng,
Ruiming Tang,
Han Li
Abstract:
Language-based user profiles convert long behavioral histories into explicit semantic representations for recommendation. However, most profile generators are optimized in an open loop: they may summarize past behavior fluently, but are not directly trained to improve future recommendation. We study this problem in real-world short-video recommendation, where user behaviors continuously arrive as…
▽ More
Language-based user profiles convert long behavioral histories into explicit semantic representations for recommendation. However, most profile generators are optimized in an open loop: they may summarize past behavior fluently, but are not directly trained to improve future recommendation. We study this problem in real-world short-video recommendation, where user behaviors continuously arrive as streams and profiles must be incrementally updated under limited capacity. This requires maintaining a consistent bounded profile state and constructing profile-targeted semantic feedback from industrial implicit behavior logs. We propose RECAP, an offline closed-loop framework for optimizing streaming structured semantic profiles with historical recommendation feedback. RECAP maintains each profile as a bounded structured memory by combining LLM-based semantic updates with deterministic lifecycle and capacity control. RECAP constructs profile-targeted semantic feedback by filtering label-consistent behavior pairs with an LLM judge and training a dual-tower evaluator whose matching score serves as a GRPO reward. Experiments on Kuaishou short-video data show that RECAP improves uAUC by 0.0084 and Recall@2000 by about 4.9% over the base generator. Further analyses confirm the benefits of feedback construction and policy optimization, and show more grounded refinement and user-level abstraction in profile updates. A seven-day online A/B test further shows a statistically significant 0.139% improvement in average application usage time per user.
△ Less
Submitted 17 July, 2026;
originally announced July 2026.
-
SpeechGuard: Online Defense against Backdoor Attacks on Speech Recognition Models
Authors:
Jinwen Xin,
Xixiang Lv
Abstract:
Backdoor attacks pose a critical threat to neural network models, allowing attackers to implant a backdoor during the training phase by manipulating a small portion of the training data. In security-sensitive applications such as voice interaction for autonomous driving, the presence of backdoor attacks introduces substantial security risks. This study focuses on implementing backdoor defense meas…
▽ More
Backdoor attacks pose a critical threat to neural network models, allowing attackers to implant a backdoor during the training phase by manipulating a small portion of the training data. In security-sensitive applications such as voice interaction for autonomous driving, the presence of backdoor attacks introduces substantial security risks. This study focuses on implementing backdoor defense measures for speech recognition models in run-time, taking into account the characteristics of audio signals. We propose SpeechGuard, the first online backdoor defense pipeline designed to identify and purify poisoned audio samples. Specifically, we improve STRIP method to perform adaptive perturbation injection to detect and filter poisoned samples, named as S-STRIP. More importantly, we further consider the purification of poisoned samples. We utilize time-frequency (T-F) masking to suppress the expression of trigger signals and autonomously generate masks based on an autoencoder. The two-stage processing prevents the backdoor in the model from being triggered, and even input speech carrying triggers can be accurately predicted. Extensive experimental demonstrate that SpeechGuard can accurately filter out poisoned samples. Through purification, it can significantly mitigate the backdoor threat while maintaining a certain prediction accuracy.
△ Less
Submitted 17 July, 2026;
originally announced July 2026.
-
X-Mind: Efficient Visual Chain-of-Thought via Predictive World Model for End-to-End Driving
Authors:
Bohao Zhao,
Chengrui Wei,
Guangfeng Jiang,
Ruixin Liu,
Xuejie Lv,
Liu Liang,
Sutao Deng,
Xiuyang Fan,
Pengkun Zheng,
Jinyun Zhou,
Rui Guo,
Hanpeng Liu,
Yutong Zheng,
Yi Guo,
Xinlong Zheng,
Qingyu Luo,
Zhuangzhuang Ding,
Yu Zhang,
Hang Zhang,
Xianming Liu
Abstract:
Predicting future states is essential for autonomous agents, yet current Vision-Language-Action (VLA) models fundamentally lack this capability, relying instead on reactive perception-action mapping. While integrating Predictive World Models (PWMs) addresses this gap, existing approaches either incur prohibitive cascaded latency or act as shallow terminal tasks that fail to deeply embed forward-lo…
▽ More
Predicting future states is essential for autonomous agents, yet current Vision-Language-Action (VLA) models fundamentally lack this capability, relying instead on reactive perception-action mapping. While integrating Predictive World Models (PWMs) addresses this gap, existing approaches either incur prohibitive cascaded latency or act as shallow terminal tasks that fail to deeply embed forward-looking reasoning. To endow VLA models with this reasoning capability, we propose X-Mind. Rather than treating PWMs as an external auxiliary module, this framework internalizes them as the Visual Chain-of-Thought (Visual CoT). By enforcing a world rollout prior to action, the model is constrained to imagine future evolution first, yielding a driving policy that is robustly grounded in environmental dynamics and aware of the future consequences its actions will unfold. The challenge here is efficiency, and we tackle it on two fronts. First, we introduce a compact representation of visual thinking: an abstract sketch that fuses a Bird's-Eye-View (BEV) layout with abstract driving priors (e.g., navigation intents and traffic rules). Rather than rolling out dense future frames, the model reasons over this sketch as a mental canvas; aided by a Deep Compression Autoencoder (DC-AE), a 12-frame future rollout is reduced to merely 96 tokens, alleviating the long-context computational bottleneck. Second, to accelerate generation further, we propose a recurrent block diffusion scheme that unrolls the denoising steps across the layers of the large drive model, folding iterative refinement into the backbone's one forward pass. Trained and validated on large-scale real-world data, X-Mind achieves competitive end-to-end driving performance, which makes it a highly practical, low-latency solution that successfully deploys large-scale cognitive reasoning directly onto resource-constrained vehicle platforms.
△ Less
Submitted 27 June, 2026;
originally announced June 2026.
-
AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems
Authors:
Changxin Lao,
Fei Pan,
Guozhuang Ma,
Han Li,
Huihuang Lin,
Jijun Shi,
Kangzhi Zhao,
Kun Gai,
Mo Zhou,
Qinqin Zhou,
Quan Chen,
Ruochen Yang,
Shifu Bie,
Shijie Yi,
Shuang Yang,
Shuo Yang,
Wenhao Li,
Wentao Xie,
Xiao Lv,
Xuming Wang,
Yijun Wang,
Yiming Chen,
Yusheng Huang,
Zhongyuan Wang,
Zibo Zhao
, et al. (37 additional authors not shown)
Abstract:
Recommendation algorithm iteration is moving from an artisanal, engineer-bound process toward an industrialized research loop, but this transition remains blocked by a structural execution bottleneck: the idea-to-launch cycle still depends on human engineers to generate hypotheses, modify production code, launch A/B experiments, and attribute online results. Innovation therefore scales linearly wi…
▽ More
Recommendation algorithm iteration is moving from an artisanal, engineer-bound process toward an industrialized research loop, but this transition remains blocked by a structural execution bottleneck: the idea-to-launch cycle still depends on human engineers to generate hypotheses, modify production code, launch A/B experiments, and attribute online results. Innovation therefore scales linearly with headcount rather than compounding with evidence, compute, and accumulated experimental knowledge. We present AgentX, a production-deployed multi-agent system that fundamentally restructures this production function. AgentX operates as a self-evolving development engine: it autonomously generates, implements, evaluates, and learns from recommendation experiments at a scale and pace that no manual workflow can sustain.
The system orchestrates four tightly coupled stages in a closed loop. A Brainstorm Agent synthesizes evidence from historical experiments, system architecture, data analysis, and external research into ranked, executable proposals. A Developing Agent translates each proposal into production-ready code through repository-grounded generation and multi-dimensional reliability verification. An Evaluation Agent conducts safe online rollout with guardrail-vetoed A/B judgment, converting both successes and failures into structured knowledge assets. A Harness Evolution layer (SGPO) then distills execution trajectories into semantic-gradient updates that continuously sharpen the agents themselves -- making the system not merely automated, but self-improving.
△ Less
Submitted 26 June, 2026; v1 submitted 25 June, 2026;
originally announced June 2026.
-
Inclusive Interactive Collisions for Multi-View Consistent Compositional 3D Generation
Authors:
Chang Liu,
Mingwen Shao,
Xiang Lv,
Xinyuan Chen,
Lingzhuang Meng,
Qiao Zhang,
Zhengyi Gong,
Jinghao Hu
Abstract:
Recent breakthroughs in 3D generation have advanced notably with the development of text-to-image diffusion model. However, existing methods remain two practical challenges: (1) They primarily generate single 3D object, but struggle to generate multi-object compositional 3D assets due to the lack of the modeling for Gaussian primitives in reasonable interactions. (2) They often suffer from cross-v…
▽ More
Recent breakthroughs in 3D generation have advanced notably with the development of text-to-image diffusion model. However, existing methods remain two practical challenges: (1) They primarily generate single 3D object, but struggle to generate multi-object compositional 3D assets due to the lack of the modeling for Gaussian primitives in reasonable interactions. (2) They often suffer from cross-view inconsistency during 3D optimization, as Score Distillation Sampling inherently performs on each single view, inevitably resulting in cross-view hallucinations. To solve above issues, we propose I2C-3D, a novel optimization-based method to generate multi-view consistent compositional 3D assets with reasonable interactions. Specifically, we propose an Inclusive Interactive Collisions strategy to guide Gaussian primitives appearing in reasonable interaction regions naturally, thereby ensuring objects in the compositional scene interact in a physically plausible and visually coherent way. Additionally, to enhance multi-view consistency, Multi-View Adaptive Score Distillation Sampling is devised to distill multi-view consistency prior and layout prior from pre-trained diffusion model by modulating attention map of instance token and spatial token across viewpoints. Benefiting from above elaborate designs, I2C-3D not only generates high-fidelity multi-view consistent compositional 3D assets but also supports 3D editing flexibly, facilitating complex scene generation. Extensive experiments demonstrate our I2C-3D outperforms existing methods in generation quality and multi-view consistency.
△ Less
Submitted 23 June, 2026;
originally announced June 2026.
-
FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech
Authors:
Haoxu Wang,
Biao Tian,
Weiqin Li,
Xiang Lv,
Han Zhao,
Xiangang Li
Abstract:
Existing Reinforcement Learning (RL) research for Text-to-Speech (TTS) focuses on large language models (LLMs), leaving Flow-Matching (FM) under-explored. We present FlowTTS-GRPO, an online RL framework for FM-based TTS. By converting ordinary differential equation (ODE) trajectories into stochastic differential equation (SDE) paths, our method enables direct fine-tuning of open-source FM models w…
▽ More
Existing Reinforcement Learning (RL) research for Text-to-Speech (TTS) focuses on large language models (LLMs), leaving Flow-Matching (FM) under-explored. We present FlowTTS-GRPO, an online RL framework for FM-based TTS. By converting ordinary differential equation (ODE) trajectories into stochastic differential equation (SDE) paths, our method enables direct fine-tuning of open-source FM models without auxiliary models. We show that a weighted reward combination converges faster than a probabilistic scheme, and identify three practical optimizations: omitting classifier-free guidance (CFG) during training accelerates convergence; synthesizing hard cases improves robustness; and applying RL to the FM component enhances audio-detail metrics. Experiments on CosyVoice 3.0 and F5-TTS demonstrate objective and subjective preference gains in speaker similarity and perceptual quality, with F5-TTS also improving intelligibility.
△ Less
Submitted 8 July, 2026; v1 submitted 22 June, 2026;
originally announced June 2026.
-
PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation
Authors:
Yuhang Huang,
Xuan Lv,
Junyan Xu,
Zhiyuan Yu,
Jiazhao Zhang,
Ruizhen Hu,
Wancheng Feng,
Shilong Zou,
Hewen Xiao,
Ziqiao Zhou,
Kaiyun Huang,
Zhiyu Peng,
Juzhan Xu,
Hang Zhao,
Chenyang Zhu,
Renjiao Yi,
Yifei Huang,
Douhui Wu,
Yan Zhang,
Kexu Cheng,
Chunhe Song,
Yunzhi Xue,
Xiuhong Zhang,
Leitao Guo,
Yunji Chen
, et al. (3 additional authors not shown)
Abstract:
World foundation models (WFMs) are powerful simulators, yet they predominantly operate in a single-view setting and lack the multi-view 3D consistency required for robotic manipulation. While robotic systems rely on multiple cameras (egocentric, eye-to-hand, and wrist-mounted) for policy learning, current multi-view world models simply concatenate view tokens without explicit geometric reasoning.…
▽ More
World foundation models (WFMs) are powerful simulators, yet they predominantly operate in a single-view setting and lack the multi-view 3D consistency required for robotic manipulation. While robotic systems rely on multiple cameras (egocentric, eye-to-hand, and wrist-mounted) for policy learning, current multi-view world models simply concatenate view tokens without explicit geometric reasoning. This causes cross-view object drift, depth inconsistency, and texture misalignment. We trace these failures to two deficiencies: the absence of an explicit inter-view communication mechanism and the lack of a 3D geometric prior. We argue that resolving both simultaneously is necessary and sufficient. To address this, we present PAIWorld, a framework that augments diffusion-transformer world models via three core components: (1) Geometry-Aware Cross-View Attention blocks that establish an explicit pathway across views, (2) Geometric Rotary Position Embedding that encodes camera ray directions and extrinsic poses into the attention mechanism, and (3) Latent 3D-REPA, which distills 3D-aware features from frozen 3D foundation models to ensure 3D consistency. Built upon a DiT-based world foundation model, PAIWorld achieves state-of-the-art multi-view 3D consistency on robotic manipulation benchmarks, ranking 1st on the WorldArena leaderboard and 2nd on the AgiBot-Challenge2026 leaderboard, while enabling downstream applications such as model-based planning, world action models, and multi-view policy post-training.
△ Less
Submitted 23 June, 2026; v1 submitted 16 June, 2026;
originally announced June 2026.
-
Carleson measures, tent embeddings, and Volterra-type integral operators on the unit ball
Authors:
Xiaofen Lv,
Xiaomin Tang,
Jani A. Virtanen
Abstract:
In this paper, we establish a sharp comparison between Carleson-cube and Bergman-metric-ball conditions on the open unit ball $\B$ and combine it with a Berezin-type characterization to prove embedding theorems for Besov spaces and Bergman spaces on $\B$ into logarithmic tent spaces in the Bergman metric. As applications, we characterize the boundedness, compactness, and essential norms of the Vol…
▽ More
In this paper, we establish a sharp comparison between Carleson-cube and Bergman-metric-ball conditions on the open unit ball $\B$ and combine it with a Berezin-type characterization to prove embedding theorems for Besov spaces and Bergman spaces on $\B$ into logarithmic tent spaces in the Bergman metric. As applications, we characterize the boundedness, compactness, and essential norms of the Volterra-type integral operators $T_g$ and $I_g$ acting from the Besov space $B_t(\B)$ to the general function space $F(p,q,s)$.
△ Less
Submitted 14 June, 2026;
originally announced June 2026.
-
CCKS: Consensus-based Communication and Knowledge Sharing
Authors:
Jinyuan Zu,
Xiaowei Lv,
Yongcai Wang,
Deying Li,
Yunjun Han,
Wenping Chen,
Fengyi Zhang,
Naiqi Wu
Abstract:
In Decentralized Training and Decentralized Execution (DTDE) for cooperative Multi-Agent Reinforcement Learning (MARL), action-advising-based knowledge sharing promotes interpretable and scalable cooperation among agents. However, current action advising approaches often adhere too much to the teacher's guidance without evaluating teacher-student compatibility, which causes excessive advising, sub…
▽ More
In Decentralized Training and Decentralized Execution (DTDE) for cooperative Multi-Agent Reinforcement Learning (MARL), action-advising-based knowledge sharing promotes interpretable and scalable cooperation among agents. However, current action advising approaches often adhere too much to the teacher's guidance without evaluating teacher-student compatibility, which causes excessive advising, suboptimal stability, and degraded performance. To overcome these challenges, this paper presents a Consensus-based Communication and Knowledge Sharing (CCKS) framework, which allows agents to adopt recommendations based on consensus-derived constraints and to follow the teacher's instructions more smartly. This mechanism enables agents to balance exploration and learning from experienced teachers, improving overall performance. The key is the consensus model construction, for which we propose to employ contrastive learning to construct consensus models based on local observations in the agents' training phase. In action selection, agents score and choose actions based on consensus and shared knowledge. Designed as a plug-and-play solution, CCKS integrates seamlessly with existing DTDE algorithms. Experiments conducted in the Google Research Football environment and the complex StarCraft II Multi-Agent Challenge demonstrate that the integration with CCKS significantly improves cooperation efficiency, learning speed, and overall performance compared with current DTDE baselines. The code is available at https://github.com/yuanxpy/CCKS.
△ Less
Submitted 10 June, 2026;
originally announced June 2026.
-
Anomalous Autler-Townes Splitting in Resonant Multiphoton Ionization Driven by Bright Squeezed Vacuum
Authors:
Xu Zhang,
Liding Li,
Yutong Deng,
Xinyou Lv,
Yang Li,
Marcelo F. Ciappina,
Peixiang Lu,
Yueming Zhou
Abstract:
Bright squeezed vacuum (BSV) light has a vanishing mean optical electric field yet can strongly enhance strong-field nonlinear responses beyond the conventional semiclassical paradigm. Here we examine this scenario in the light-matter strong-coupling regime by investigating resonant multiphoton ionization of atoms driven by BSV, using a fully quantum treatment of both the electron and the field. O…
▽ More
Bright squeezed vacuum (BSV) light has a vanishing mean optical electric field yet can strongly enhance strong-field nonlinear responses beyond the conventional semiclassical paradigm. Here we examine this scenario in the light-matter strong-coupling regime by investigating resonant multiphoton ionization of atoms driven by BSV, using a fully quantum treatment of both the electron and the field. Our results show that the photoelectron energy spectrum exhibits an anomalous Autler-Townes splitting whose magnitude grows with the Above-threshold-ionization (ATI) order, rather than remaining essentially ATI-order independent as in the case of coherent driving. This behavior reflects a general scaling with the number of absorbed photons and originates from the broad photon-number fluctuations of the driving field together with the resulting electron-field entanglement. We further show that the BSV-induced enhancement of ionization yields evolves with intensity, crossing over from the $g^{(p+1)}$ limit to the $g^{(p)}$ limit as Rabi oscillations become established. These results identify a quantum regime of strong-field ionization governed by the interplay of photon statistics, nonlinear transitions, strong coupling, and nonseparable light-matter dynamics.
△ Less
Submitted 5 June, 2026;
originally announced June 2026.
-
OneReason Technical Report
Authors:
OneRec Team,
Biao Yang,
Boyang Ding,
Chenglong Chu,
Dunju Zang,
Fei Pan,
Han Li,
Hao Jiang,
Honghui Bao,
Huanjie Wang,
Jian Liang,
Jiangxia Cao,
Jiao Ou,
Jiaxin Deng,
Jinghao Zhang,
Kun Gai,
Lu Ren,
Peiru Du,
Pengfei Zheng,
Rongzhou Zhang,
Ruiming Tang,
Shiyao Wang,
Siyang Mao,
Siyuan Lou,
Teng Shi
, et al. (59 additional authors not shown)
Abstract:
Generative recommendation models in the OneRec family have been widely deployed in many real-world services, such as short-video, live-streaming, advertising, and e-commerce. However, these generative models can only benefit from the scaling advantage, while their reasoning ability is hard to activate, since we cannot construct meaningful Chain-of-Thought (CoT) sequences consisting of itemic token…
▽ More
Generative recommendation models in the OneRec family have been widely deployed in many real-world services, such as short-video, live-streaming, advertising, and e-commerce. However, these generative models can only benefit from the scaling advantage, while their reasoning ability is hard to activate, since we cannot construct meaningful Chain-of-Thought (CoT) sequences consisting of itemic tokens only. Inspired by the success of the reasoning-style ``think before answer'' paradigm in the LLM field, we conduct preliminary studies (i.e., OneRec-Think, OpenOneRec) to explore reasoning capability in generative recommendation. Nevertheless, we notice an unexpected phenomenon: the thinking mode does not show advantages over the non-thinking mode. Drawing insights from recent findings on CoT robustness in multi-modal language models, we argue that effective reasoning in recommendation rests on two factors: perception, the ability to ground itemic tokens in their underlying language semantics, and cognition, the ability to reorganize a user's behavior sequence into coherent latent interest points. We therefore propose OneReason, which includes: (1) strong itemic token perception in pre-training, (2) a three-level cognition-enhanced CoT format for recommendation tasks in SFT, and (3) a specialize-then-unify training recipe in RL to enhance the thinking ability.
△ Less
Submitted 4 June, 2026;
originally announced June 2026.
-
SB-RF: Schrödinger Bridge Rectified Flow for One-Step Robust Speech Enhancement
Authors:
Caixia Lu,
Xueyang Lv,
Penglong Hu,
Jiaming Xu
Abstract:
Generative models have shown promising results for speech enhancement (SE), but they often rely on multi-step inference, limiting low-latency deployment. We propose SB-RF, a one-step generative framework that integrates Rectified Flow (RF) with Schrödinger Bridge (SB) theory. During training, SB-RF samples intermediate states from an SB time marginal and trains a conditional velocity field with th…
▽ More
Generative models have shown promising results for speech enhancement (SE), but they often rely on multi-step inference, limiting low-latency deployment. We propose SB-RF, a one-step generative framework that integrates Rectified Flow (RF) with Schrödinger Bridge (SB) theory. During training, SB-RF samples intermediate states from an SB time marginal and trains a conditional velocity field with the RF velocity-matching objective. At inference, SB-RF starts from the noisy observation and applies a single Euler update. Experiments show that SB-RF achieves competitive performance among generative methods on the VoiceBank-DEMAND benchmark. To further assess performance beyond this standard setting, we evaluate SB-RF on a simulated low signal-to-noise ratio test set using an expanded training dataset. Under these conditions, SB-RF achieves superior performance over the compared baselines, supporting its potential for real-world applications.
△ Less
Submitted 2 July, 2026; v1 submitted 3 June, 2026;
originally announced June 2026.
-
SHB-AE: Spherical harmonic beamforming based Ambisonics encoding and upscaling method for smartphone microphone array
Authors:
Yuhuan You,
Yufan Qian,
Tianshu Qu,
Bin Wang,
Xueyang Lv
Abstract:
With the rapid development of virtual reality (VR) and augmented reality (AR), spatial audio recording and reproduction have gained increasing research interest. Higher Order Ambisonics (HOA) stands out for its adaptability to various playback devices and its ability to integrate head orientation. However, current HOA recordings often rely on bulky spherical microphone arrays (SMA), and portable d…
▽ More
With the rapid development of virtual reality (VR) and augmented reality (AR), spatial audio recording and reproduction have gained increasing research interest. Higher Order Ambisonics (HOA) stands out for its adaptability to various playback devices and its ability to integrate head orientation. However, current HOA recordings often rely on bulky spherical microphone arrays (SMA), and portable devices like smartphones are limited by array configuration and number of microphones. We propose SHB-AE, a spherical harmonic beamforming based method for Ambisonics encoding using a smartphone microphone array (SPMA). By designing beamformers for each order of spherical harmonic functions based on the array manifold, the method enables Ambisonics encoding and up-scaling. Validation on a real SPMA and its simulated free-field counterpart in noisy and reverberant conditions showed that the method successfully encodes and up-scales Ambisonics up to the fourth order with just four irregularly arranged microphones.
△ Less
Submitted 3 June, 2026;
originally announced June 2026.
-
Flow-HOA: Generative Joint Optimization for Ambisonics Encoding via Flow Matching
Authors:
Yuhuan You,
Yufan Qian,
Tianshu Qu,
Bin Wang,
Xueyang Lv
Abstract:
Higher-Order Ambisonics (HOA) encoding from sparse, irregular microphone arrays remains a critical challenge for consumer spatial audio capture in immersive communication and XR. We propose Flow-HOA, a generative framework that jointly optimizes a multi-dimensional objective encompassing time-domain, spectral, and spatial fidelity while producing a deployable, time-invariant bank of Finite Impulse…
▽ More
Higher-Order Ambisonics (HOA) encoding from sparse, irregular microphone arrays remains a critical challenge for consumer spatial audio capture in immersive communication and XR. We propose Flow-HOA, a generative framework that jointly optimizes a multi-dimensional objective encompassing time-domain, spectral, and spatial fidelity while producing a deployable, time-invariant bank of Finite Impulse Response (FIR) encoding filters. Using conditional flow matching, the model learns to map a simple prior distribution to the target distribution of FIR filter coefficients. Training is guided by a composite loss that balances time-domain waveform fidelity, multi-resolution spectral consistency, sub-band energy preservation, and spatial directivity constraints. Objective evaluations on synthetically simulated data demonstrate improved performance over strong model-based baselines in both signal fidelity and spatial accuracy metrics. Subjective listening tests on real microphone array recordings further confirm that Flow-HOA yields higher overall sound quality with reduced artifacts, demonstrating generalization from synthetic training data to real-world capture conditions.
△ Less
Submitted 3 June, 2026;
originally announced June 2026.
-
Beyond Independent Manipulation: Individual Fairness-aware Strategic Classification with Peer Imitation
Authors:
Xinpeng Lv,
Chunyuan Zheng,
Yunxin Mao,
Renzhe Xu,
Jinxuan Yang,
Yuanlong Chen,
Wangrong Huang,
Shaowu Yang,
Wenjing Yang,
Xinwang Liu,
Peng Cui,
Haotian Wang
Abstract:
Strategic classification (SC) investigates scenarios where agents manipulate their features to obtain favorable decisions from predictive models. Existing fairness-aware SC approaches primarily focus on group fairness and typically assume that agents respond independently. However, when individual fairness is required, ensuring similar individuals receive similar outcomes, agents' manipulation bec…
▽ More
Strategic classification (SC) investigates scenarios where agents manipulate their features to obtain favorable decisions from predictive models. Existing fairness-aware SC approaches primarily focus on group fairness and typically assume that agents respond independently. However, when individual fairness is required, ensuring similar individuals receive similar outcomes, agents' manipulation becomes interdependent: an agent's preferred manipulation depends on the neighborhoods' outcomes. This induces a mismatch between classical SC formulations and fairness-aware decision settings, where independent models no longer accurately characterize strategic manipulations. To address this issue, we introduce individual fairness-aware strategic classification (IFSC), a framework that models peer-driven manipulation arising from individual fairness, where agents imitate nearby positively decided peers to obtain favorable outcomes. IFSC characterizes strategic manipulation as similarity-based imitation toward visible accepted peers and learns classifiers under the resulting post-manipulation distributions. To account for uncertainty in peer observability, IFSC employs a robust learning process that introduces stochastic perturbations during manipulation simulation. Experiments on synthetic and real-world datasets demonstrate that IFSC improves individual-fairness consistency and mitigates imitation-induced distortions.
△ Less
Submitted 25 June, 2026; v1 submitted 30 May, 2026;
originally announced June 2026.
-
Partial Fairness Awareness: Belief-Guided Strategic Mechanism for Strategic Agents
Authors:
Xinpeng Lv,
Chunyuan Zheng,
Yunxin Mao,
Renzhe Xu,
Hao Zou,
Shanzhi Gu,
Liyang Xu,
Huan Chen,
Yuanlong Chen,
Wenjing Yang,
Haotian Wang
Abstract:
Strategic machine learning investigates scenarios where agents manipulate their features to receive favorable decisions from predictive models. To address fairness concerns intrinsic to strategic classification, recent work has introduced group-specific fairness constraints. However, current fairness-aware approaches face a fundamental dilemma in the issue of fairness exposure: making these constr…
▽ More
Strategic machine learning investigates scenarios where agents manipulate their features to receive favorable decisions from predictive models. To address fairness concerns intrinsic to strategic classification, recent work has introduced group-specific fairness constraints. However, current fairness-aware approaches face a fundamental dilemma in the issue of fairness exposure: making these constraints public enables strategic manipulation and can lead to fairness reversal, while keeping them hidden may reduce social welfare and discourage genuine improvement. To fill this gap, we subsequently propose the problem of partial fairness awareness (PFA), as our theoretical analysis informs that such a dilemma can be mitigated by releasing the candidate set of fairness constraints and concealing the grounding constraint. To be specific, we introduce a belief-guided strategic mechanism, wherein agents iteratively interact with the decision system and maintain a belief distribution over the candidate set of fairness constraints. This belief-guided process enables agents, through iterative interaction and feedback, to update their belief distribution over the candidate set, thereby gradually aligning their belief with the grounding fairness constraint employed by the system. Extensive experiments on real-world and synthetic datasets demonstrate that PFA achieves lower group fairness gaps, higher acceptance of truly qualified individuals, and more stable outcomes compared to fully public or private fairness regimes.
△ Less
Submitted 30 May, 2026;
originally announced June 2026.
-
Visual-Advantage On-Policy Distillation for Vision-Language Models
Authors:
Ruiqi Liu,
Xiaolei Lv,
Gengsheng Li,
Ximo Zhu,
Zhiheng Wang,
Zhengbo Zhang,
Junkai Chen,
Zhiheng Li,
Bo Li,
Jun Gao,
Shu Wu
Abstract:
On-policy knowledge distillation has proven effective for language models, yet its application to vision-language models (VLMs) remains underexplored. We observe that standard on-policy distillation can improve a student's output quality while failing to strengthen its reliance on visual input: on vision-critical tokens, the student's predictions remain largely unchanged whether or not fine-graine…
▽ More
On-policy knowledge distillation has proven effective for language models, yet its application to vision-language models (VLMs) remains underexplored. We observe that standard on-policy distillation can improve a student's output quality while failing to strengthen its reliance on visual input: on vision-critical tokens, the student's predictions remain largely unchanged whether or not fine-grained visual detail is present, even though the teacher's predictions depend heavily on it.To make this difference observable, we introduce visual advantage (VA), the token-level log-probability difference when the teacher scores a student-generated rollout with versus without access to fine-grained visual detail. VA is concentrated in a small minority of tokens, and these high-VA tokens are the ones that actually carry the visual supervision signal. This motivates a distillation objective that treats them differently from language scaffolding, so their contribution is not diluted by the abundant surrounding language tokens.We propose Visual-Advantage On-Policy Distillation (VA-OPD), which uses VA at two granularities: rollout-level reweighting by trajectory-averaged VA, and token-level KL averaged within high-VA and low-VA groups separately. We train on two math datasets (Geometry3K and ViRL39K) and evaluate on eight benchmarks covering both mathematical reasoning and visual understanding, across three teacher sizes (4B, 8B, and 32B) on the Qwen3-VL family. VA-OPD improves over standard on-policy distillation on every benchmark, with the gain growing monotonically along both the teacher-size and data-scale axes, suggesting that these factors compound consistently.
△ Less
Submitted 20 May, 2026;
originally announced May 2026.
-
Beyond Rational Illusion: Behaviorally Realistic Strategic Classification
Authors:
Xinpeng Lv,
Yunxin Mao,
Renzhe Xu,
Chunyuan Zheng,
Yikai Chen,
Haoxuan Li,
Yang Shi,
Jinxuan Yang,
Zhouchen Lin,
Yuanlong Chen,
Yuanxing Zhang,
Shaowu Yang,
Wenjing Yang,
Haotian Wang
Abstract:
Strategic classification(SC) studies the interaction between decision models and agents who strategically manipulate their features for favorable outcomes. Existing SC frameworks typically rely on the idealized assumption that agents are strictly rational. However, evidence from behavioral economics and psychology consistently shows that real-world decision-making is often shaped by cognitive bias…
▽ More
Strategic classification(SC) studies the interaction between decision models and agents who strategically manipulate their features for favorable outcomes. Existing SC frameworks typically rely on the idealized assumption that agents are strictly rational. However, evidence from behavioral economics and psychology consistently shows that real-world decision-making is often shaped by cognitive biases, deviating from pure rationality. To formalize this limitation, we identify and define a new problem setting, termed the behaviorally realistic strategic classification problem, where agents' strategic manipulations deviate from full rationality due to psychological biases. Motivated by the identified limitation, we propose the Prospect-Guided Strategic Framework (Pro-SF) to address the problem, a principled framework grounded in prospect theory to model and learn under behaviorally realistic strategic responses. Specifically, to capture behaviorally realistic strategic manipulations, our framework reformulates the Stackelberg-style interaction between agents and the decision-maker by incorporating three key mechanisms inspired by prospect theory, including the asymmetry between benefits and costs, different subjective reference points, and non-rational probability distortion. Experiments on synthetic and real-world datasets establish Pro-SF as a behaviorally grounded approach to strategic classification, bridging machine learning and behavioral economics for more reliable deployment in the real world.
△ Less
Submitted 6 June, 2026; v1 submitted 19 May, 2026;
originally announced May 2026.
-
When Tabular Foundation Models Meet Strategic Tabular Data: A Prior Alignment Approach
Authors:
Xinpeng Lv,
Yunxin Mao,
Renzhe Xu,
Chunyuan Zheng,
Yikai Chen,
Haoxuan Li,
Jinxuan Yang,
Kun Kuang,
Yuanlong Chen,
Mingyang Geng,
Wanrong Huang,
Shixuan Liu,
Shaowu Yang,
Wenjing Yang,
Zhouchen Lin,
Haotian Wang
Abstract:
Tabular foundation models based on pretrained prior-data fitted networks~(PFNs) have shown strong generalization on diverse tabular tasks, but they are typically designed for \emph{non-strategic} settings where data distributions are independent of deployed classifiers. In many real-world decision scenarios, however, individuals may strategically modify their features after deployment to obtain fa…
▽ More
Tabular foundation models based on pretrained prior-data fitted networks~(PFNs) have shown strong generalization on diverse tabular tasks, but they are typically designed for \emph{non-strategic} settings where data distributions are independent of deployed classifiers. In many real-world decision scenarios, however, individuals may strategically modify their features after deployment to obtain favorable outcomes, inducing a post-deployment distribution shift. This paper studies whether PFN-style tabular foundation models can generalize to such \emph{strategic} tabular data. We show that strategic manipulation creates a mismatch between the non-strategic prior learned during pretraining and the post-manipulation strategic prior, which leads to systematic prediction bias. To address this issue, we propose \textbf{Strategic Prior-data Fitted Network}~\textit{(SPN)}, an inference-time strategy-aware framework that adapts tabular foundation models to strategic environments without retraining. SPN constructs strategic in-context examples to approximate post-manipulation inputs and aligns PFN predictions with the induced strategic distribution. Experiments on real-world and synthetic tabular datasets show that SPN consistently improves robustness and predictive performance under strategic manipulation compared with both tabular foundation models and classical tabular methods.
△ Less
Submitted 6 June, 2026; v1 submitted 19 May, 2026;
originally announced May 2026.
-
Post-Trained MoE Can Skip Half Experts via Self-Distillation
Authors:
Xingtai Lv,
Li Sheng,
Kaiyan Zhang,
Yichen You,
Siyan Gao,
Xueheng Luo,
Yuxin Zuo,
Yuchen Fan,
Junlin Yang,
Ganqu Cui,
Bingning Wang,
Fan Yang,
Youbang Sun,
Ning Ding,
Bowen Zhou
Abstract:
Mixture-of-Experts (MoE) scales language models efficiently through sparse expert activation, and its dynamic variant further reduces computation by adjusting the activated experts in an input-dependent manner. Existing dynamic MoE methods usually rely on pre-training from scratch or task-specific adaptation, leaving the practical conversion of fully trained MoE underexplored. Enabling such adapta…
▽ More
Mixture-of-Experts (MoE) scales language models efficiently through sparse expert activation, and its dynamic variant further reduces computation by adjusting the activated experts in an input-dependent manner. Existing dynamic MoE methods usually rely on pre-training from scratch or task-specific adaptation, leaving the practical conversion of fully trained MoE underexplored. Enabling such adaptation would directly alleviate the inference costs by allowing easy tokens to bypass unnecessary expert during serving. This paper introduces Zero-Expert Self-Distillation Adaptation (ZEDA), a low-cost framework that transforms post-trained static MoE models into efficient dynamic ones. To stabilize this architectural conversion, ZEDA injects parameter-free zero-output experts into each MoE layer and adapts the augmented model through two-stage self-distillation, utilizing the original MoE as a frozen teacher and applying a group-level balancing loss. On Qwen3-30B-A3B and GLM-4.7-Flash across 11 benchmarks spanning math, code, and instruction following, ZEDA eliminates over 50% of expert FLOPs at marginal accuracy loss. It outperforms the strongest dynamic MoE baseline by 6.1 and 4.0 points on the two models, and delivers ~1.20$\times$ end-to-end inference speedup.
△ Less
Submitted 7 June, 2026; v1 submitted 18 May, 2026;
originally announced May 2026.
-
DADF: A Distribution-Aware Debiasing Framework for Watch-Time Regression in Recommender Systems
Authors:
Yiqing Yang,
Xinlong Zhao,
Zhao Liu,
Xiao Lv,
Ruiming Tang,
Han Li,
Kun Gai
Abstract:
Watch-time predictors in short-video recommender systems can be approximately calibrated by their own scores while still overestimating short observations and underestimating long ones. We study whether this label-space mean shrinkage contains inference-time-predictable residual structure that can be corrected without replacing a mature first-stage model. We propose DADF, a distribution-aware seco…
▽ More
Watch-time predictors in short-video recommender systems can be approximately calibrated by their own scores while still overestimating short observations and underestimating long ones. We study whether this label-space mean shrinkage contains inference-time-predictable residual structure that can be corrected without replacing a mature first-stage model. We propose DADF, a distribution-aware second-stage framework that applies multiplicative correction to a frozen watch-time predictor. DADF stabilizes long-tailed correction targets with group-specific transformations, uses video duration to route specialized correction experts, and incorporates auxiliary engagement representations. Duration is used only to index heterogeneous residual distributions, not treated as the cause of the observed pattern. Experiments on KuaiRec and WeChat21 with seven first-stage backbones, together with a large-scale industrial ranking system, show that DADF reduces offline MAE by 4.33% and improves XAUC by 4.01% on average. In production, it reduces MAE by 12.57%. Three online A/B tests across full ranking, rough ranking, and degraded serving improve average time spent per device by 0.649%, 0.235%, and 0.199%, respectively, and all three integrations were subsequently deployed to 100% of traffic. These results show that DADF is a practical, model-agnostic plug-in for correcting predictable conditional residuals while preserving the serving interface of mature first-stage models. Code is available at https://github.com/liuzhao09/DADF.
△ Less
Submitted 31 July, 2026; v1 submitted 18 May, 2026;
originally announced May 2026.
-
When Retrieval Hurts Code Completion: A Diagnostic Study of Stale Repository Context
Authors:
Haojun Weng,
Qianqian Yang,
Hao Fu,
Haobin Pan,
Xinwei Lv
Abstract:
Context: Retrieval-augmented code generation relies on cross-file repository context, but retrieved snippets may come from obsolete project states.
Objectives: We study whether temporally stale repository snippets act as harmless noise or actively induce current-state-incompatible code.
Methods: We conduct a controlled diagnostic study on a curated 17-sample set of production-helper signature…
▽ More
Context: Retrieval-augmented code generation relies on cross-file repository context, but retrieved snippets may come from obsolete project states.
Objectives: We study whether temporally stale repository snippets act as harmless noise or actively induce current-state-incompatible code.
Methods: We conduct a controlled diagnostic study on a curated 17-sample set of production-helper signature changes from five Python repositories. For each sample, we compare current-only, stale-only, no-retrieval, and mixed current/stale retrieval conditions under prompts that hide commit freshness and expected current signatures.
Results: Under neutralized prompts, stale-only retrieval induces stale helper references on 15/17 Qwen2.5-Coder-7B-Instruct samples and 13/17 gpt-4.1-mini samples, corresponding to 88.2 and 76.5 percentage-point increases over current-only retrieval. No retrieval produces zero stale references but only 1/17 passing completions. The two models share 75.0% Jaccard overlap among stale-triggering samples, and mixed conditions show that adding valid current evidence largely rescues stale-only failures.
Conclusion: Temporal validity of retrieved repository context is a distinct diagnostic variable for Code RAG robustness: stale context can actively bias models toward obsolete repository state rather than merely removing useful evidence.
△ Less
Submitted 14 May, 2026;
originally announced May 2026.
-
Pre-training Enables Extraordinary All-optical Image Denoising
Authors:
Xudong Lv,
Yuxiang Sun,
Shuo Wang,
Nanxing Chen,
Jun Guan,
Jingtian Hu
Abstract:
Optical neural networks are emerging as powerful machine learning and information processing tools because of their potential advantages in speed and energy efficiency. The training methods of these physical models, however, remain underexplored compared to their digital counterparts and are leading to suboptimal performance. This paper reports a pre-training-driven approach that leads to snapshot…
▽ More
Optical neural networks are emerging as powerful machine learning and information processing tools because of their potential advantages in speed and energy efficiency. The training methods of these physical models, however, remain underexplored compared to their digital counterparts and are leading to suboptimal performance. This paper reports a pre-training-driven approach that leads to snapshot image denoising with substantially improved quality. We demonstrated effective free-space optical denoising by a diffractive network optimized by a two-step process including (1) pre-training using a massive dataset of 3.45 million diverse but simple images and (2) fine-tuning with the corresponding task-specific datasets. Compared to conventional Fourier-domain filtering and directly trained diffractive networks, such a transfer learning process exhibited prominent advantages for denoising images degraded by severe noise, peak signal-to-noise ratio (PSNR) below 8 dB, while preserving fine image features and improving the PSNR to above 18 dB. Importantly, the same pre-trained optical network could be consistently fine-tuned to process degraded images from highly diverse styles ranging from handwritten digits (MNIST) and chest X-rays (ChestMNIST) to CIFAR-10 images and human faces (CelebA). We further demonstrated the critical role of our optical denoisers in vision-based applications, including face detection, plate recognition, and localization of UAVs in noisy conditions.
△ Less
Submitted 8 May, 2026;
originally announced May 2026.
-
How Far Is Document Parsing from Solved? PureDocBench: A Source-TraceableBenchmark across Clean, Degraded, and Real-World Settings
Authors:
Zhiheng Li,
Zongyang Ma,
Jiaxian Chen,
Jianing Zhang,
Zhaolong Su,
Yutong Zhang,
Zhiyin Yu,
Ruiqi Liu,
Xiaolei Lv,
Bo Li,
Jun Gao,
Ziqi Zhang,
Chunfeng Yuan,
Bing Li,
Weiming Hu
Abstract:
The past year has seen over 20 open-source document parsing models, yet thefield still benchmarks almost exclusively on OmniDocBench, a 1,355-pagemanually annotated dataset whose top scores have saturated above 90%. Athree-stage audit pipeline we run on OmniDocBench screens its 21,353evaluator-scored blocks and confirms 2,580 errors (12.08%); combined with overa year of public availability, both a…
▽ More
The past year has seen over 20 open-source document parsing models, yet thefield still benchmarks almost exclusively on OmniDocBench, a 1,355-pagemanually annotated dataset whose top scores have saturated above 90%. Athree-stage audit pipeline we run on OmniDocBench screens its 21,353evaluator-scored blocks and confirms 2,580 errors (12.08%); combined with overa year of public availability, both annotation quality and contamination riskcall its rankings into question. To address these issues, we presentPureDocBench, a programmatically generated, source-traceable benchmark thatrenders document images from HTML/CSS and produces verifiable annotations fromthe same source, covering 10 domains, 66 subcategories, and 1,475 pages, eachin three versions: clean, digitally degraded, and real-degraded (4,425 imagestotal). Evaluating 40 models spanning pipeline specialists, end-to-endspecialists, and general-purpose VLMs, we find: (i) document parsing is farfrom solved: the best model scores only ~74 out of 100, with a 44.6-point gapbetween the strongest and weakest models; (ii) specialist parsers with <=4Bparameters rival or surpass general VLMs that are 5-100x larger, yet formularecognition remains a shared bottleneck where no model exceeds 67% whenaveraging the formula metric across all three tracks; (iii) general VLMs loseonly 0.99/8.52 Overall points under digital/real degradation versus 4.90/14.21for pipeline specialists, producing ranking reversals that make clean-onlyevaluation misleading for deployment. All data, code, and artifacts arepublicly released.
△ Less
Submitted 8 May, 2026;
originally announced May 2026.
-
Learning quantum disentanglement scheduling from reduced states via modular hybrid policies
Authors:
Y. -X. Xiao,
J. -Z. Han,
Z. Zheng,
Z. -H. Zhang,
M. Xue,
J. Li,
X. Lv
Abstract:
Quantum control with restricted state access is central to near-term quantum devices, where full wave-function information is unavailable. We study this problem through multiqubit disentanglement scheduling from partial observations, where a controller receives only two-qubit reduced density matrices and selects which qubit pair to disentangle at each step. We introduce a modular hybrid quantum--c…
▽ More
Quantum control with restricted state access is central to near-term quantum devices, where full wave-function information is unavailable. We study this problem through multiqubit disentanglement scheduling from partial observations, where a controller receives only two-qubit reduced density matrices and selects which qubit pair to disentangle at each step. We introduce a modular hybrid quantum--classical policy framework consisting of classical preprocessing, a parameterized quantum circuit as a compact nonlinear latent block, and classical postprocessing for pair-selection probabilities. Benchmarking 4-, 5-, and 6-qubit tasks, we find that preprocessing is the dominant factor governing performance under reduced-state observations, while the quantum module provides a conditional compact representation whose utility depends on the input features and model budget. We further identify a performance--efficiency trade-off across policy families and find that increasing circuit width is generally more useful than increasing depth. These results provide practical design principles for hybrid policies in reduced-information quantum control.
△ Less
Submitted 30 April, 2026;
originally announced April 2026.
-
S2G-RAG: Structured Sufficiency and Gap Judging for Iterative Retrieval-Augmented QA
Authors:
Minghan Li,
Junjie Zou,
Xinxuan Lv,
Chao Zhang,
Guodong Zhou
Abstract:
Retrieval-Augmented Generation (RAG) grounds language models in external evidence, but multi-hop question answering remains difficult because iterative pipelines must control what to retrieve next and when the available evidence is adequate. In practice, systems may answer from incomplete evidence chains, or they may accumulate redundant or distractor-heavy text that interferes with later retrieva…
▽ More
Retrieval-Augmented Generation (RAG) grounds language models in external evidence, but multi-hop question answering remains difficult because iterative pipelines must control what to retrieve next and when the available evidence is adequate. In practice, systems may answer from incomplete evidence chains, or they may accumulate redundant or distractor-heavy text that interferes with later retrieval and reasoning. We propose S2G-RAG (Structured Sufficiency and Gap-judging RAG), an iterative framework with an explicit controller, S2G-Judge. At each turn, S2G-Judge predicts whether the current evidence memory supports answering and, if not, outputs structured gap items that describe the missing information. These gap items are then mapped into the next retrieval query, producing stable multi-turn retrieval trajectories. To reduce noise accumulation, S2G-RAG maintains a sentence-level Evidence Context by extracting a compact set of relevant sentences from retrieved documents. Experiments on TriviaQA, HotpotQA, and 2WikiMultiHopQA show that S2G-RAG improves multi-hop QA performance and robustness under multi-turn retrieval. Furthermore, S2G-RAG can be integrated into existing RAG pipelines as a lightweight component, without modifying the search engine or retraining the generator.
△ Less
Submitted 26 April, 2026;
originally announced April 2026.
-
Gate- and Optically Controlled Nonlinear Optical Response in Graphene via Non-Perturbative Ultrafast Carrier Dynamics
Authors:
Xiaolong Lv,
Yu Zhang,
Yuxuan Wei,
Chuanshan Tian
Abstract:
While the Dirac band structure of graphene has established it as a leading platform for ultrafast optoelectronics, its non-perturbative nonlinear response under intense excitation remains poorly understood. Here, we report ultrafast spectral modulation of nonlinear optical signals in graphene. By utilizing a robust suspended-graphene platform that allows for both wide-range electrostatic gating an…
▽ More
While the Dirac band structure of graphene has established it as a leading platform for ultrafast optoelectronics, its non-perturbative nonlinear response under intense excitation remains poorly understood. Here, we report ultrafast spectral modulation of nonlinear optical signals in graphene. By utilizing a robust suspended-graphene platform that allows for both wide-range electrostatic gating and high optical damage thresholds, we observe dramatic frequency shifts (up to 8 THz) in third-harmonic generation (THG) and sum-frequency generation (SFG) driven by pump-induced nonequilibrium carrier dynamics. The magnitude and even the direction of this spectral shift can be reversibly controlled by the Fermi level and excitation conditions. A quasiequilibrium theoretical framework based on hot-carrier dynamics quantitatively reproduces the measured spectral evolution, elucidating the critical interplay between carrier heating and the Fermi level. These findings establish a universal mechanism for carrier-mediated spectral control, providing a practical route toward high-speed, gatetunable nonlinear photonic architectures.
△ Less
Submitted 24 April, 2026;
originally announced April 2026.
-
HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation
Authors:
Shuanghao Bai,
Meng Li,
Xinyuan Lv,
Jiawei Wang,
Xinhua Wang,
Fei Liao,
Chengkai Hou,
Langzhe Gu,
Wanqi Zhou,
Kun Wu,
Ziluo Ding,
Zhiyuan Xu,
Lei Sun,
Shanghang Zhang,
Zhengping Che,
Jian Tang,
Badong Chen
Abstract:
Humans achieve complex manipulation through coordinated whole-body control, whereas most Vision-Language-Action (VLA) models treat robot body parts largely independently, making high-DoF humanoid control challenging and often unstable. We present HEX, a state-centric framework for coordinated manipulation on full-sized bipedal humanoid robots. HEX introduces a humanoid-aligned universal state repr…
▽ More
Humans achieve complex manipulation through coordinated whole-body control, whereas most Vision-Language-Action (VLA) models treat robot body parts largely independently, making high-DoF humanoid control challenging and often unstable. We present HEX, a state-centric framework for coordinated manipulation on full-sized bipedal humanoid robots. HEX introduces a humanoid-aligned universal state representation for scalable learning across heterogeneous embodiments, and incorporates a Mixture-of-Experts Unified Proprioceptive Predictor to model whole-body coordination and temporal motion dynamics from large-scale multi-embodiment trajectory data. To efficiently capture temporal visual context, HEX uses lightweight history tokens to summarize past observations, avoiding repeated encoding of historical images during inference. It further employs a residual-gated fusion mechanism with a flow-matching action head to adaptively integrate visual-language cues with proprioceptive dynamics for action generation. Experiments on real-world humanoid manipulation tasks show that HEX achieves state-of-the-art performance in task success rate and generalization, particularly in fast-reaction and long-horizon scenarios.
△ Less
Submitted 19 May, 2026; v1 submitted 9 April, 2026;
originally announced April 2026.
-
Making MLLMs Blind: Adversarial Smuggling Attacks in MLLM Content Moderation
Authors:
Zhiheng Li,
Zongyang Ma,
Yuntong Pan,
Ziqi Zhang,
Xiaolei Lv,
Bo Li,
Jun Gao,
Jianing Zhang,
Chunfeng Yuan,
Bing Li,
Weiming Hu
Abstract:
Multimodal Large Language Models (MLLMs) are increasingly being deployed as automated content moderators. Within this landscape, we uncover a critical threat: Adversarial Smuggling Attacks. Unlike adversarial perturbations (for misclassification) and adversarial jailbreaks (for harmful output generation), adversarial smuggling exploits the Human-AI capability gap. It encodes harmful content into h…
▽ More
Multimodal Large Language Models (MLLMs) are increasingly being deployed as automated content moderators. Within this landscape, we uncover a critical threat: Adversarial Smuggling Attacks. Unlike adversarial perturbations (for misclassification) and adversarial jailbreaks (for harmful output generation), adversarial smuggling exploits the Human-AI capability gap. It encodes harmful content into human-readable visual formats that remain AI-unreadable, thereby evading automated detection and enabling the dissemination of harmful content. We classify smuggling attacks into two pathways: (1) Perceptual Blindness, disrupting text recognition; and (2) Reasoning Blockade, inhibiting semantic understanding despite successful text recognition. To evaluate this threat, we constructed SmuggleBench, the first comprehensive benchmark comprising 1,700 adversarial smuggling attack instances. Evaluations on SmuggleBench reveal that both proprietary (e.g., GPT-5) and open-source (e.g., Qwen3-VL) state-of-the-art models are vulnerable to this threat, producing Attack Success Rates (ASR) exceeding 90%. By analyzing the vulnerability through the lenses of perception and reasoning, we identify three root causes: the limited capabilities of vision encoders, the robustness gap in OCR, and the scarcity of domain-specific adversarial examples. We conduct a preliminary exploration of mitigation strategies, investigating the potential of test-time scaling (via CoT) and adversarial training (via SFT) to mitigate this threat. Our code is publicly available at https://github.com/zhihengli-casia/smugglebench.
△ Less
Submitted 8 April, 2026; v1 submitted 8 April, 2026;
originally announced April 2026.
-
Xuanwu: Evolving General Multimodal Models into an Industrial-Grade Foundation for Content Ecosystems
Authors:
Zhiqian Zhang,
Xu Zhao,
Xiaoqing Xu,
Guangdong Liang,
Weijia Wang,
Xiaolei Lv,
Bo Li,
Jun Gao
Abstract:
In recent years, multimodal large models have continued to improve on general benchmarks. However, in real-world content moderation and adversarial settings, mainstream models still suffer from degraded generalization and catastrophic forgetting because of limited fine-grained visual perception and insufficient modeling of long-tail noise. In this paper, we present Xuanwu VL-2B as a case study of…
▽ More
In recent years, multimodal large models have continued to improve on general benchmarks. However, in real-world content moderation and adversarial settings, mainstream models still suffer from degraded generalization and catastrophic forgetting because of limited fine-grained visual perception and insufficient modeling of long-tail noise. In this paper, we present Xuanwu VL-2B as a case study of how general multimodal models can be developed into an industrial-grade foundation model for content ecosystems. The model adopts a compact InternViT-300M + MLP + Qwen3 1.7B architecture, balancing fine-grained visual perception, language-semantic alignment, and deployment cost within an approximately 2B-parameter budget. To balance business specialization with the retention of general capabilities, we developed a data iteration and curation mechanism and trained the model through a progressive three-stage pipeline: pre-training, mid-training, and post-training. Ablation studies and offline business evaluations show that Xuanwu VL-2B achieves an average score of 67.90 across seven OpenCompass multimodal metrics (vs. 64.27 for InternVL 3.5 2B), an average recall of 94.38% over seven independent business moderation tasks, and a weighted overall recall of 82.82% on policy-violating text in challenging adversarial OCR scenarios, outperforming Gemini-2.5-Pro (76.72%). These results show that, under a limited parameter budget, Xuanwu VL-2B achieves a practical balance among business alignment, visual perception, general capability retention, and deployment cost.
△ Less
Submitted 30 March, 2026;
originally announced March 2026.
-
Learning Unified Control of Intrinsic Nonlinear Spin Dynamics in Atomic Qudits for Magnetometry
Authors:
C. Z. Cao,
J. Z. Han,
M. Xiong,
M. Deng,
L. Wang,
X. Lv,
M. Xue
Abstract:
Generating and preserving metrologically useful quantum states is a central challenge in quantum-enhanced metrology. In low-field atomic magnetometry with multilevel atoms, the nonlinear Zeeman (NLZ) effect is both a resource and a limitation. It can generate internal spin squeezing within a single atomic qudit, but under fixed readout it also rotates and distorts the measurement-relevant quadratu…
▽ More
Generating and preserving metrologically useful quantum states is a central challenge in quantum-enhanced metrology. In low-field atomic magnetometry with multilevel atoms, the nonlinear Zeeman (NLZ) effect is both a resource and a limitation. It can generate internal spin squeezing within a single atomic qudit, but under fixed readout it also rotates and distorts the measurement-relevant quadrature, limiting the usable metrological gain. The problem is further complicated by the time dependence of both the squeezing axis and the nonlinear evolution itself. Here we show that reinforcement learning can transform NLZ dynamics from a source of readout degradation into a sustained metrological resource. Using only experimentally accessible low-order spin moments, a trained agent identifies a unified control policy for this class of intrinsically nonlinear sensing dynamics. We illustrate the approach in the $f=21/2$ manifold of $^{161}\mathrm{Dy}$, where the learned policy rapidly prepares strongly squeezed internal states and stabilizes more than $4\,\mathrm{dB}$ of fixed-axis spin squeezing under continuous NLZ evolution. Including state-preparation overhead, the learned protocol yields a single-atom magnetic-field sensitivity of $13.9\,\mathrm{pT}/\sqrt{\mathrm{Hz}}$, approximately $3\,\mathrm{dB}$ beyond the standard quantum limit. Our results establish learning-based control as an experimentally feasible route for converting unavoidable intrinsic nonlinear dynamics in multilevel atomic sensors into operational metrological advantage.
△ Less
Submitted 28 April, 2026; v1 submitted 30 March, 2026;
originally announced March 2026.
-
A Telescope System for Charge and Position Measurement of High Energy Nuclei
Authors:
Dexing Miao,
Zhiyu Xiang,
Giovanni Ambrosi,
Mattia Barbanera,
Baasansuren Batsukh,
Mengke Cai,
Xudong Cai,
Yuan-Hann Chang,
Shanzhen Chen,
Hsin-Yi Chou,
Xingzhu Cui,
Mingyi Dong,
Matteo Duranti,
Ke Gong,
Mingjie Feng,
Valerio Formato,
Daojin Hong,
Maria Ionica,
Xiaojie Jiang,
Yaozu Jiang,
Liangchenglong Jin,
Shengjie Jin,
Vladimir Koutsenko,
Tiange Li,
Zuhao Li
, et al. (21 additional authors not shown)
Abstract:
A high-granularity telescope system with a large sensitive area and low material budget has been developed for high-energy heavy ion beam tests. The telescope consists of nine layers of silicon microstrip detectors (SSDs), whose performance was validated through a heavy ion beam test at the CERN SPS. A hybrid machine learning algorithm is proposed to address the challenges of nuclear charge measur…
▽ More
A high-granularity telescope system with a large sensitive area and low material budget has been developed for high-energy heavy ion beam tests. The telescope consists of nine layers of silicon microstrip detectors (SSDs), whose performance was validated through a heavy ion beam test at the CERN SPS. A hybrid machine learning algorithm is proposed to address the challenges of nuclear charge measurement with SSDs. The system achieves a spatial resolution of $\mathcal{O}(1) \,$\SI{}{\micro\metre} and a charge resolution better than 0.16 charge units for nuclei from $Z = 1$ to $Z = 29$, with a sensitive area of $8 \times 8 \, \mathrm{cm}^2$. To the best of our knowledge, this represents the most precise charge and spatial resolution simultaneously achieved by a silicon telescope to date.
△ Less
Submitted 26 March, 2026;
originally announced March 2026.
-
Hydrodynamics of dilation and spin currents
Authors:
Zhong-Hua Zhang,
Xi-Hu Lv,
Xu-Guang Huang
Abstract:
We formulate a relativistic hydrodynamic theory for fluids with spin and intrinsic dilation charges. Using an entropy-current analysis, we derive constitutive relations featuring a bulk viscosity and a dilation conductivity governing the relaxation and diffusion of dilation charge. Linear mode analysis reveals a gapped dilation excitation and the freeze-out of long-wavelength sound modes, similar…
▽ More
We formulate a relativistic hydrodynamic theory for fluids with spin and intrinsic dilation charges. Using an entropy-current analysis, we derive constitutive relations featuring a bulk viscosity and a dilation conductivity governing the relaxation and diffusion of dilation charge. Linear mode analysis reveals a gapped dilation excitation and the freeze-out of long-wavelength sound modes, similar to the superhorizon modes in cosmology. In the nonrelativistic limit, the theory reduces to that of microstretch fluids. Upon coupling to electromagnetic field, we show that the scale anomaly permits additional contributions in the electric current, dilation current, and energy-momentum tensor. Our theory naturally applies to nearly conformal fluids undergoing rapid expansion or contraction.
△ Less
Submitted 18 March, 2026;
originally announced March 2026.
-
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
Authors:
Yushi Bai,
Qian Dong,
Ting Jiang,
Xin Lv,
Zhengxiao Du,
Aohan Zeng,
Jie Tang,
Juanzi Li
Abstract:
Long-context agentic workflows have emerged as a defining use case for large language models, making attention efficiency critical for both inference speed and serving cost. Sparse attention addresses this challenge effectively, and DeepSeek Sparse Attention (DSA) is a representative production-grade solution: a lightweight lightning indexer selects the top-k most relevant tokens per query, reduci…
▽ More
Long-context agentic workflows have emerged as a defining use case for large language models, making attention efficiency critical for both inference speed and serving cost. Sparse attention addresses this challenge effectively, and DeepSeek Sparse Attention (DSA) is a representative production-grade solution: a lightweight lightning indexer selects the top-k most relevant tokens per query, reducing core attention from $O(L^2)$ to $O(Lk)$. However, the indexer itself retains $O(L^2)$ complexity and must run independently at every layer, despite the fact that the resulting top-k selections are highly similar across consecutive layers. We present IndexCache, which exploits this cross-layer redundancy by partitioning layers into a small set of Full layers that run their own indexers and a majority of Shared layers that simply reuse the nearest Full layer's top-k indices. We propose two complementary approaches to determine and optimize this configuration. Training-free IndexCache applies a greedy search algorithm that selects which layers to retain indexers by directly minimizing language modeling loss on a calibration set, requiring no weight updates. Training-aware IndexCache introduces a multi-layer distillation loss that trains each retained indexer against the averaged attention distributions of all layers it serves, enabling even simple interleaved patterns to match full-indexer accuracy. Experimental results on a 30B DSA model show that IndexCache can remove 75% of indexer computations with negligible quality degradation, achieving up to 1.82$\times$ prefill speedup and 1.48$\times$ decode speedup compared to standard DSA. These positive results are further confirmed by our preliminary experiments on the production-scale GLM-5 model (Figure 1).
△ Less
Submitted 12 March, 2026;
originally announced March 2026.
-
How Far Can Unsupervised RLVR Scale LLM Training?
Authors:
Bingxiang He,
Yuxin Zuo,
Zeyuan Liu,
Shangziqi Zhao,
Zixuan Fu,
Junlin Yang,
Cheng Qian,
Kaiyan Zhang,
Yuchen Fan,
Ganqu Cui,
Xiusi Chen,
Youbang Sun,
Xingtai Lv,
Xuekai Zhu,
Li Sheng,
Ran Li,
Huan-ang Gao,
Yuchen Zhang,
Bowen Zhou,
Zhiyuan Liu,
Ning Ding
Abstract:
Unsupervised reinforcement learning with verifiable rewards (URLVR) offers a pathway to scale LLM training beyond the supervision bottleneck by deriving rewards without ground truth labels. Recent works leverage model intrinsic signals, showing promising early gains, yet their potential and limitations remain unclear. In this work, we revisit URLVR and provide a comprehensive analysis spanning tax…
▽ More
Unsupervised reinforcement learning with verifiable rewards (URLVR) offers a pathway to scale LLM training beyond the supervision bottleneck by deriving rewards without ground truth labels. Recent works leverage model intrinsic signals, showing promising early gains, yet their potential and limitations remain unclear. In this work, we revisit URLVR and provide a comprehensive analysis spanning taxonomy, theory and extensive experiments. We first classify URLVR methods into intrinsic versus external based on reward sources, then establish a unified theoretical framework revealing that all intrinsic methods converge toward sharpening the model's initial distribution This sharpening mechanism succeeds when initial confidence aligns with correctness but fails catastrophically when misaligned. Through systematic experiments, we show intrinsic rewards consistently follow a rise-then-fall pattern across methods, with collapse timing determined by model prior rather than engineering choices. Despite these scaling limits, we find intrinsic rewards remain valuable in test-time training on small datasets, and propose Model Collapse Step to measure model prior, serving as a practical indicator for RL trainability. Finally, we explore external reward methods that ground verification in computational asymmetries, showing preliminary evidence they may escape the confidence-correctness ceiling. Our findings chart boundaries for intrinsic URLVR while motivating paths toward scalable alternatives.
△ Less
Submitted 9 March, 2026;
originally announced March 2026.
-
Large scale mapping of [CI] and the [CI]-to-CO transition in $ρ$ Ophiuchus molecular cloud
Authors:
Jifeng Xia,
Ningyu Tang,
Thomas G. Bisbas,
Chen Wang,
Gan Luo,
Sihan Jiao,
Xin Lv,
Xuejian Jiang,
Donghui Quan,
Jinzeng Li,
Paul F. Goldsmith,
Gary A. Fuller,
Di Li
Abstract:
Atomic carbon ([CI]) is a key species in the carbon chemistry of the interstellar medium (ISM). Using the Submillimeter Wave Astronomy Satellite (SWAS), we conducted a [CI]($^3$P$_1$--$^3$P$_0$) 492 GHz survey covering approximately 4 deg$^2$ of the L1688 and L1689 regions in the $ρ$ Oph molecular cloud, achieving a spatial resolution of 4.25$\hbox{$^{\prime}$}$. The derived [CI] column densities,…
▽ More
Atomic carbon ([CI]) is a key species in the carbon chemistry of the interstellar medium (ISM). Using the Submillimeter Wave Astronomy Satellite (SWAS), we conducted a [CI]($^3$P$_1$--$^3$P$_0$) 492 GHz survey covering approximately 4 deg$^2$ of the L1688 and L1689 regions in the $ρ$ Oph molecular cloud, achieving a spatial resolution of 4.25$\hbox{$^{\prime}$}$. The derived [CI] column densities, N([CI]), range from 4.85 $\times$ 10$^{14}$ cm$^{-2}$ to 6.29 $\times$ 10$^{17}$ cm$^{-2}$, corresponding to an abundance ratio N([CI])/N($H_2$) of 2.24$\times$ 10$^{-7}$ to 2.39$\times$ 10$^{-4}$, with a median value of 1.8$\times$ 10$^{-5}$. Combining observations with photodissociation region (PDR) modeling, we find that [CI] abundance varies less than CO in regions with UV intensity G$_0$ $> 16$ and N(H$_2$) $<$ 4.6 $\times$ 10$^{21}$ cm$^{-2}$, suggesting [CI] is a more reliable tracer of molecular hydrogen in low-density, high-radiation environments where the [CI]-to-CO transition occurs. Utilizing [CI] as direct H$_2$ tracer, the CO-dark gas fraction is estimated to be 0.43 , meaning that 43% of the total cloud mass will be missed by conventional calculation based on CO observations but can be calibrated by [CI] emission. The [CI] line widths are systematically broader than those of $^{13}$CO, possibly due to contributions from atomic carbon. These findings provide key insights into Galactic [CI] emission and the carbon cycle evolution in the interstellar medium. Future high-sensitivity [CI] ($^3$P$_1$--$^3$P$_0$) surveys with the Chinese Survey Space Telescope (CSST) will significantly advance our understanding of the carbon cycle evolution.
△ Less
Submitted 2 March, 2026;
originally announced March 2026.
-
Let the Agent Search: Autonomous Exploration Beats Rigid Workflows in Temporal Question Answering
Authors:
Xufei Lv,
Jiahui Yang,
Haoyuan Sun,
Xialin Su,
Zhiliang Tian,
Yifu Gao,
Linbo Qiao,
Houde Liu
Abstract:
Temporal Knowledge Graph Question Answering (TKGQA) is challenging because it requires multi-hop reasoning under complex temporal constraints. Recent LLM-based approaches have improved semantic modeling for this task, but many still rely on fixed reasoning workflows or costly post-training, which can limit adaptability and make error recovery difficult. We show that enabling an off-the-shelf Large…
▽ More
Temporal Knowledge Graph Question Answering (TKGQA) is challenging because it requires multi-hop reasoning under complex temporal constraints. Recent LLM-based approaches have improved semantic modeling for this task, but many still rely on fixed reasoning workflows or costly post-training, which can limit adaptability and make error recovery difficult. We show that enabling an off-the-shelf Large Language Model (LLM) to determine its next action is already effective in a zero-shot setting. Based on this insight, we propose AT2QA, an Autonomous and Training-free Agent for TKG Question Answering. AT2QA empowers the LLM to iteratively interact with the TKG via a generic search tool, inherently enabling autonomous exploration and dynamic self-correction during reasoning. To further elicit the LLM's potential for complex temporal reasoning, we introduce a training-free experience mining mechanism that distills a compact few-shot demonstration library from successful self-generated trajectories. AT2QA also yields a transparent audit trail for every prediction. Experiments on three challenging benchmarks -- MultiTQ, Timeline-CronQuestion, and Timeline-ICEWS-Actor -- show that AT2QA achieves new state-of-the-art performance, surpassing the strongest baselines by 10.7, 4.9, and 11.2 absolute points, respectively. Our code is available at https://github.com/AT2QA-Official-Code/AT2QA-Official-Code
△ Less
Submitted 25 March, 2026; v1 submitted 2 March, 2026;
originally announced March 2026.
-
Knowing the Unknown: Interpretable Open-World Object Detection via Concept Decomposition Model
Authors:
Xueqiang Lv,
Shizhou Zhang,
Yinghui Xing,
Di Xu,
Peng Wang,
Yanning Zhang
Abstract:
Open-world object detection (OWOD) requires incrementally detecting known categories while reliably identifying unknown objects. Existing methods primarily focus on improving unknown recall, yet overlook interpretability, often leading to known-unknown confusion and reduced prediction reliability. This paper aims to make the entire OWOD framework interpretable, enabling the detector to truly "know…
▽ More
Open-world object detection (OWOD) requires incrementally detecting known categories while reliably identifying unknown objects. Existing methods primarily focus on improving unknown recall, yet overlook interpretability, often leading to known-unknown confusion and reduced prediction reliability. This paper aims to make the entire OWOD framework interpretable, enabling the detector to truly "knowing the unknown". To this end, we propose a concept-driven InterPretable OWOD framework(IPOW) by introducing a Concept Decomposition Model (CDM) for OWOD, which explicitly decomposes the coupled RoI features in Faster R-CNN into discriminative, shared, and background concepts. Discriminative concepts identify the most discriminative features to enlarge the distances between known categories, while shared and background concepts, due to their strong generalization ability, can be readily transferred to detect unknown categories. Leveraging the interpretable framework, we identify that known-unknown confusion arises when unknown objects fall into the discriminative space of known classes. To address this, we propose Concept-Guided Rectification (CGR) to further resolve such confusion. Extensive experiments show that IPOW significantly improves unknown recall while mitigating confusion, and provides concept-level interpretability for both known and unknown predictions.
△ Less
Submitted 24 February, 2026;
originally announced February 2026.
-
GLM-5: from Vibe Coding to Agentic Engineering
Authors:
GLM-5-Team,
:,
Aohan Zeng,
Xin Lv,
Zhenyu Hou,
Zhengxiao Du,
Qinkai Zheng,
Bin Chen,
Da Yin,
Chendi Ge,
Chenghua Huang,
Chengxing Xie,
Chenzheng Zhu,
Congfeng Yin,
Cunxiang Wang,
Gengzheng Pan,
Hao Zeng,
Haoke Zhang,
Haoran Wang,
Huilong Chen,
Jiajie Zhang,
Jian Jiao,
Jiaqi Guo,
Jingsen Wang,
Jingzhao Du
, et al. (162 additional authors not shown)
Abstract:
We present GLM-5, a next-generation foundation model designed to transition the paradigm of vibe coding to agentic engineering. Building upon the agentic, reasoning, and coding (ARC) capabilities of its predecessor, GLM-5 adopts DSA to significantly reduce training and inference costs while maintaining long-context fidelity. To advance model alignment and autonomy, we implement a new asynchronous…
▽ More
We present GLM-5, a next-generation foundation model designed to transition the paradigm of vibe coding to agentic engineering. Building upon the agentic, reasoning, and coding (ARC) capabilities of its predecessor, GLM-5 adopts DSA to significantly reduce training and inference costs while maintaining long-context fidelity. To advance model alignment and autonomy, we implement a new asynchronous reinforcement learning infrastructure that drastically improves post-training efficiency by decoupling generation from training. Furthermore, we propose novel asynchronous agent RL algorithms that further improve RL quality, enabling the model to learn from complex, long-horizon interactions more effectively. Through these innovations, GLM-5 achieves state-of-the-art performance on major open benchmarks. Most critically, GLM-5 demonstrates unprecedented capability in real-world coding tasks, surpassing previous baselines in handling end-to-end software engineering challenges. Code, models, and more information are available at https://github.com/zai-org/GLM-5.
△ Less
Submitted 24 February, 2026; v1 submitted 17 February, 2026;
originally announced February 2026.