-
MOSS-VL Technical Report
Authors:
Pengyu Wang,
Chenkun Tan,
Shaojun Zhou,
Qirui Zhou,
Yanxin Chen,
Xingyang He,
Huazheng Zeng,
Jijun Cheng,
Chenghao Wang,
Xiaomeng Qian,
Pengfei Wang,
Zhan Huang,
Shanqing Gao,
Wei Huang,
Longjun Cao,
Wu Ran,
Jie Liu,
Changtai Zhu,
Hongkai Wang,
Yixian Tian,
Chenghao Liu,
Zhen Ye,
Xinghao Wang,
Botian Jiang,
Guoguo Feng
, et al. (7 additional authors not shown)
Abstract:
We present MOSS-VL, an open vision-language model family that treats real-time interaction -- perceiving while it speaks -- as a first-class capability. It is co-designed across the stack: the language decoder attends to vision only through gated cross-attention, so the model can naturally see incoming frames while generating; a synthesized interaction corpus supervises when to speak, when to stay…
▽ More
We present MOSS-VL, an open vision-language model family that treats real-time interaction -- perceiving while it speaks -- as a first-class capability. It is co-designed across the stack: the language decoder attends to vision only through gated cross-attention, so the model can naturally see incoming frames while generating; a synthesized interaction corpus supervises when to speak, when to stay silent, and when to revise; and a staged curriculum concentrates all real-time-specific training in one light final stage over a strong offline foundation. Offline, MOSS-VL-Instruct is competitive at comparable scale and leads temporal-reasoning video sets. Across four streaming benchmarks, MOSS-VL-Realtime posts the best average on three (second on the fourth) among open-source streaming models, sweeping the three subsets that squarely test proactive behavior -- 66.0 vs. 37.5 for the best baseline on OmniMMI Proactive Alerting. With 11.3B parameters but visual tokens outside the decoded sequence, MOSS-VL widens its time-to-first-token advantage over same-backbone Qwen3-VL-8B from 2.8x to 5.1x as visual context grows. We release all five checkpoints, the training curriculum, and the real-time inference code at https://github.com/OpenMOSS/MOSS-VL.
△ Less
Submitted 15 August, 2026;
originally announced August 2026.
-
Integrated Sensing, Communication, and Computing in Multi-Tier Systems: Joint Hybrid Beamforming Design and Computation Resource Allocation
Authors:
Peng Liu,
Zesong Fei,
Xinyi Wang,
Qiao Qi,
Zhaohui Yang,
Meng Hua,
Arumugam Nallanathan
Abstract:
This paper proposes a novel integrated sensing, communication, and computing (ISCC) framework over a cloud-edge-device collaborative architecture, where passive sensing is enabled by reusing uplink offloading signals to extract sensing information directly at the edge without incurring additional transmission overhead. Nevertheless, such signal reuse introduces an inherent tradeoff between communi…
▽ More
This paper proposes a novel integrated sensing, communication, and computing (ISCC) framework over a cloud-edge-device collaborative architecture, where passive sensing is enabled by reusing uplink offloading signals to extract sensing information directly at the edge without incurring additional transmission overhead. Nevertheless, such signal reuse introduces an inherent tradeoff between communication efficiency and sensing coverage. To address this challenge, we adopt a hybrid beamforming architecture under practical hardware constraints. In addition, the integration of sensing tasks creates significant resource contention at the mobile edge computing (MEC) server, where latency-sensitive device tasks and computation-intensive sensing inference tasks compete for limited processing capacity. To alleviate this computation burden, we introduce a split inference mechanism that strategically partitions intelligent sensing tasks between the edge and the cloud. Building upon this framework, we formulate a joint optimization problem to minimize the average computation latency of all device tasks subject to strict sensing performance constraints. To tackle the high non-convexity of the formulated problem, we develop an efficient alternating optimization algorithm. In particular, we design a two-layer framework to jointly determine the optimal DNN splitting point and computation resource allocation and employ a weighted minimum mean square error (WMMSE)-based approach with manifold optimization for hybrid beamforming design. Numerical results demonstrate that the proposed framework achieves a superior tradeoff between sensing accuracy and computation latency compared to the benchmark schemes.
△ Less
Submitted 14 August, 2026;
originally announced August 2026.
-
ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models
Authors:
Yilin Jiang,
Xiaorong Zhu,
Fei Tan,
Zicheng Zhang,
Kaiyi Huang,
Yang Yu,
Zexuan Fei,
Yiming Luo,
Keqian Li,
Hao Hao,
Guangtao Zhai,
Aimin Zhou
Abstract:
Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does not: a usable education-facing model is supposed to be accurate, safe under sensitive prompts, instructionally useful, and aligned with pedagogical goals at the same time. Existing benchmarks evaluate these requirements…
▽ More
Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does not: a usable education-facing model is supposed to be accurate, safe under sensitive prompts, instructionally useful, and aligned with pedagogical goals at the same time. Existing benchmarks evaluate these requirements largely in isolation, so none assesses education-facing suitability as an integrated profile. We introduce ELBench, the first benchmark to evaluate all four requirements (General Capability, Safety and Trustworthiness, Basic Education, and High-Level Cultivation) on the same models under a common protocol, combining curated public sources with newly synthesized safety and cultivation data. We evaluate nine models, seven frontier general-purpose systems and two education-specialized variants, and report three findings. First, module-level profiles are more informative than a single aggregate: the top six models are statistically indistinguishable on overall score, yet their module leaders differ substantially, and safety is anti-correlated with practical teaching (r = -0.83). Second, the Chinese-developed models lead the safety module, the most discriminative in the suite; this advantage is largest on region-specific normative content and narrows, but does not vanish, on universal-harm content. Third, the two education-specialized models lead neither education module, and on High-Level Cultivation all models share a systematic blind spot: on the structured judgment task they converge on the same non-reference option, favoring pedagogical style over fit to the stated goal, so the module scores uniformly low and does not separate models. This raises, but does not resolve, whether domain post-training keeps pace with frontier systems on education tasks.
△ Less
Submitted 11 August, 2026; v1 submitted 10 August, 2026;
originally announced August 2026.
-
Universal Scaling of the Minimum Error Probability in Qualification of Quantum States
Authors:
Zhaoyu Fei,
Yaotian Li,
Weicheng Huang,
Xiaoguang Wang,
Y. M. Du
Abstract:
Qualification of quantum states judges which of two sets of quantum states an unknown state lies in, where the two sets are labeled by two distinct parameter regions. We formulate this problem as a composite quantum hypothesis test and uncover universal scaling laws for the minimum error probability for $N$ copies. Taking polarization-direction qualification and purity qualification as examples, w…
▽ More
Qualification of quantum states judges which of two sets of quantum states an unknown state lies in, where the two sets are labeled by two distinct parameter regions. We formulate this problem as a composite quantum hypothesis test and uncover universal scaling laws for the minimum error probability for $N$ copies. Taking polarization-direction qualification and purity qualification as examples, we show that the $N$-copy permutation symmetry and the geometric symmetries of the parameter regions identify the optimal measurements and the "worst pairwise states". The minimum error probability scales as $N^{-3/2}\exp(-Nξ)$ for disjoint regions and as $(NF)^{-1/2}$ for adjacent regions, where $ξ$ and $F$ are the quantum Chernoff divergence and quantum Fisher information associated with the "worst pairwise states", respectively. With the minimum error probability serving as an order parameter, the transition between the scaling behaviors becomes a second-order phase transition as $N\to\infty$. Our approach determines whether a quantum state belongs to a given set without full state tomography, thereby enabling qualification of large ensembles using finite samples.
△ Less
Submitted 10 August, 2026; v1 submitted 5 August, 2026;
originally announced August 2026.
-
Agent Skills Matter: Inferring Proprietary Skills from Execution Trajectories
Authors:
Jianing Geng,
Ruiqi He,
Zekun Fei,
Biao Yi,
Xuansheng Wu,
Ruijie Wang,
Zheli Liu,
Xia Hu,
Qingkai Zeng
Abstract:
Agent skills package reusable procedures that improve downstream performance. Their lightweight, portable form enables marketplace monetization and private deployment behind cloud-hosted agent interfaces, giving providers incentives to keep high-value skills proprietary. Yet hiding the artifacts does not conceal their behavioral effects, which remain observable in execution trajectories and form a…
▽ More
Agent skills package reusable procedures that improve downstream performance. Their lightweight, portable form enables marketplace monetization and private deployment behind cloud-hosted agent interfaces, giving providers incentives to keep high-value skills proprietary. Yet hiding the artifacts does not conceal their behavioral effects, which remain observable in execution trajectories and form a behavioral side channel. We define this exposure as Skill Leakage: reconstructing proprietary skills from trajectories elicited by benign queries, without reference answers or success labels. We introduce SigLeak, a black-box framework that exploits recurring skill signatures in agent behavior. It constructs diverse, decision-rich diagnostic tasks, contrasts matched skill-enabled and skill-disabled trajectories, and iteratively refines a reconstructed skill from the isolated patterns. Across five scenarios, three model families, and three agent frameworks, SigLeak outperforms or matches three baselines in nearly every setting. It raises the success rate by 6.88 percentage points over the skill-disabled reference on average and achieves the highest overall SkillSim, our metric for coarse- and fine-grained semantic similarity. These results show that benign execution trajectories can expose proprietary procedural knowledge. The code is available at https://anonymous.4open.science/r/SigLeak-D1DB.
△ Less
Submitted 30 July, 2026; v1 submitted 28 July, 2026;
originally announced July 2026.
-
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
Authors:
Jun Zhan,
Chen Yang,
Yitian Gong,
Donghua Yu,
Kuangwei Chen,
Wenbo Zhang,
Kexin Huang,
Qi Luo,
Zhe Xu,
Ying Zhu,
Jin Wang,
Tengyue Zhang,
Qi Chen,
Cheng Chang,
Songlin Wang,
Junqi Dai,
Jiasheng Ye,
Xiaogui Yang,
Tianyi Liang,
Xiangyu Peng,
Zhaoye Fei,
Shimin Li,
Qinyuan Cheng,
Xie Chen,
Xinchi Chen
, et al. (1 additional authors not shown)
Abstract:
Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to their fundamental structural differences. Most existing methods use audio and video VAEs trained separately. As a result, t…
▽ More
Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to their fundamental structural differences. Most existing methods use audio and video VAEs trained separately. As a result, the two latent spaces lack cross-modal alignment, leaving the downstream generative model to learn cross-modal synchronization from scratch. We present OmniVAE, a jointly trained audio-video VAE that learns fine-grained semantic alignment between audio and video latent representations. Beyond reconstruction, OmniVAE uses a segment-level audio-video contrastive objective to capture temporal-semantic correspondence and align the two latent spaces. In parallel, it distills features from pretrained modality-specific semantic encoders into each modality, improving the downstream learnability of both latent spaces. Extensive experiments show that both objectives consistently improve the learnability of the latent spaces, translating into higher generation quality and more accurate cross-modal synchronization in downstream text-to-audio-video generation. These findings underscore the importance of learning unified representations as a foundation for omnimodal modeling.1
△ Less
Submitted 31 July, 2026; v1 submitted 26 July, 2026;
originally announced July 2026.
-
HiMe: Hierarchical Embodied Memory for Long-Horizon Vision-Language-Action Control
Authors:
Li Ji,
Siyin Wang,
Pengfang Qian,
Xiaopeng Yu,
Yihai Tian,
Zhaoye Fei,
Jingjing Gong,
Xipeng Qiu
Abstract:
Current Vision-Language-Action (VLA) models excel at robotic manipulation but often struggle with non-Markovian tasks requiring long-term memory and reasoning due to their reliance on immediate observations. Existing solutions face a ''frequency-competence paradox,'' where stronger reasoning models are too slow for real-time control, while faster models lack sufficient reasoning capabilities. To r…
▽ More
Current Vision-Language-Action (VLA) models excel at robotic manipulation but often struggle with non-Markovian tasks requiring long-term memory and reasoning due to their reliance on immediate observations. Existing solutions face a ''frequency-competence paradox,'' where stronger reasoning models are too slow for real-time control, while faster models lack sufficient reasoning capabilities. To resolve this architectural misalignment, we propose HiMe, a Hierarchical Embodied Memory framework that decouples embodied intelligence into a high-frequency Executor for execution, a Sentry for working memory, and a Planner for long-term strategy. We also introduce a dynamic knowledge system based on cross-modal semantic schemas and active management mechanisms, allowing robots to maintain memory plasticity through ''Add, Update, and Delete'' operations. This hierarchical design effectively balances the conflict between real-time execution and slow thinking planning, significantly improving success rates in long-horizon tasks. Experiments demonstrate that this approach not only outperforms flat memory baselines but also exhibits the novel ability to self-correct its internal knowledge based on human preferences.
△ Less
Submitted 3 July, 2026;
originally announced July 2026.
-
GEAR-Seg: A Grounded Explainable Agent for Reasoning Segmentation and Data Engine
Authors:
Yanan Wang,
Wen Li,
Yibin Ying,
Zhenghao Fei
Abstract:
Reasoning segmentation requires localizing targets based on complex, implicit queries. Current end-to-end models typically entangle perception and deduction into an opaque black box, severely limiting interpretability and scalability. To address this, we propose GEAR-Seg (Grounded Explainable Agent for Reasoning Segmentation), an explicitly decoupled agent that shifts the paradigm by translating v…
▽ More
Reasoning segmentation requires localizing targets based on complex, implicit queries. Current end-to-end models typically entangle perception and deduction into an opaque black box, severely limiting interpretability and scalability. To address this, we propose GEAR-Seg (Grounded Explainable Agent for Reasoning Segmentation), an explicitly decoupled agent that shifts the paradigm by translating visual pixels into dense, attribute-rich text. By decoupling class-agnostic segmentation, semantic description, and Large Language Model (LLM) deduction, GEAR-Seg transforms implicit reasoning into an explicit, trackable logic chain. As a zero-shot inference framework, it achieves highly competitive performance across diverse reasoning and fine-grained referring segmentation benchmarks. Furthermore, GEAR-Seg inherently functions as a highly scalable data engine. Utilizing this engine, we construct GEAR-131K, a massive benchmark (over 38k images, 656k QA-mask pairs) introducing a multifaceted taxonomy tailored for complex real-world manipulation-oriented reasoning. Finally, distillation experiments demonstrate that lightweight models supervised exclusively by our automated pipeline closely match the upper-bound performance of costly human-annotated baselines.
△ Less
Submitted 1 July, 2026;
originally announced July 2026.
-
Evolving Intelligent Complex Systems via Intellicise Networks: Architecture, Technologies, and Pathways
Authors:
Ping Zhang,
Rui Meng,
Xiaodong Xu,
Song Gao,
Zixuan Huang,
Yaheng Wang,
Yinqiu Liu,
Ruichen Zhang,
Yiming Liu,
Kaiwen Yu,
Yaping Sun,
Han Meng,
Haonan Tong,
Huishi Song,
Qianqian Yang,
Shuoyao Wang,
Lexi Xu,
Qinghe Du,
Geng Sun,
Jiawen Kang,
Gang Wu,
Yiqing Zhou,
Haixia Zhang,
Zesong Fei,
Aimin Hao
, et al. (1 additional authors not shown)
Abstract:
Future engineering infrastructures are evolving into large-scale, open, heterogeneous, and wirelessly interconnected complex systems. These systems present significant challenges in optimizing network resource utilization, managing high-dimensional information spaces, and accommodating diverse business requirements. Intellicise networks, characterized by Intent-driven operation, semantic-native ca…
▽ More
Future engineering infrastructures are evolving into large-scale, open, heterogeneous, and wirelessly interconnected complex systems. These systems present significant challenges in optimizing network resource utilization, managing high-dimensional information spaces, and accommodating diverse business requirements. Intellicise networks, characterized by Intent-driven operation, semantic-native capability, and distributed intelligence, offer a promising paradigm for enabling such intelligent complex systems. We provide a systematic exploration of future intelligent complex systems from the perspective of intellicise networks. Specifically, we propose a cross-domain intelligent communication network architecture based on intellicise networks, grounded in information theory, systems theory, game theory, and cybernetics. The architecture comprises a cross-layer organizational framework, multi-functional planes, and novel information flows. The cross-layer framework defines the vertical evolution from perception and cognition to decision, while the control, user, data, computation, intelligence, and security planes deliver horizontal intellicise capabilities. Moreover, data, knowledge, model, and task flows interconnect the various layers and planes, forming a closed-loop process that derives simplicity from high-level intelligene while concurrently pursuing enhanced. Building on this architecture, we review key enabling technologies, tracing their evolution from semantic extraction to intent understanding, from heterogeneous resource integration to self-configuration and self-optimization, from generative artificial intelligence (AI) to agentic AI, and from embodied AI to symbodied AI. Additionally, we present a case study on intellicise networks for embodied agent communications and discuss representative applications and services for intelligent complex systems.
△ Less
Submitted 30 June, 2026;
originally announced July 2026.
-
Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy
Authors:
Junhao Shi,
Zezheng Huai,
Siyin Wang,
Jia Chen,
Yubang Wang,
Zhaoye Fei,
Hechang Chen,
Jingjing Gong,
Xipeng Qiu,
Yu-Gang Jiang
Abstract:
Building persistent embodied agents in unstructured environments demands unified orchestration of heterogeneous tools spanning both cyber (APIs, IoT) and physical (manipulation, navigation) domains, coupled with autonomous recovery from physical failures that inevitably arise over extended operation. Existing systems treat these as separate problems: VLM-based planners lack a unified cyber-physica…
▽ More
Building persistent embodied agents in unstructured environments demands unified orchestration of heterogeneous tools spanning both cyber (APIs, IoT) and physical (manipulation, navigation) domains, coupled with autonomous recovery from physical failures that inevitably arise over extended operation. Existing systems treat these as separate problems: VLM-based planners lack a unified cyber-physical action space, agent frameworks accumulate unbounded context that degrades temporal coherence, and VLA policies execute open-loop without detecting their own failures. We argue that persistent autonomy requires not a monolithic model but a hierarchical asynchronous architecture with explicit separation of planning, memory, and verification. To this end, we present OmniAct, a framework integrating a multimodal semantic planner for skill routing across unified action spaces, an adaptive hierarchical memory with event-boundary-driven compression for sub-linear context growth, and an asynchronous visual preemption engine that closes the semantic loop during physical execution. Across 40 real-world long-horizon tasks on two robotic platforms coordinating four IoT devices, OmniAct achieves consistent improvements in end-to-end success across all complexity levels, maintains near-flat token consumption over under 100k+ accumulated interaction tokens, and elevates mid-scale open-weight models to proprietary-level performance.
△ Less
Submitted 25 June, 2026;
originally announced June 2026.
-
WebCQ: Cooperative Multi-Agent Deep Reinforcement Learning for Scalable Web GUI Testing
Authors:
Yujia Fan,
Sinan Wang,
Zebang Fei,
Yao Qin,
Huaxuan Li,
Yepang Liu
Abstract:
Multi-agent reinforcement learning (MARL)-based techniques have shown promise for GUI testing. However, as the complexity of modern GUI software increases, existing MARL-based approaches (e.g., MARG and Fastbot) struggle to scale due to the inherent limitations of their underlying tabular reinforcement learning algorithms. This limits their applicability to large-scale commercial GUI software, esp…
▽ More
Multi-agent reinforcement learning (MARL)-based techniques have shown promise for GUI testing. However, as the complexity of modern GUI software increases, existing MARL-based approaches (e.g., MARG and Fastbot) struggle to scale due to the inherent limitations of their underlying tabular reinforcement learning algorithms. This limits their applicability to large-scale commercial GUI software, especially web applications with vast state spaces and many interactive elements. To fill this gap, we propose WebCQ, a novel MARL-based approach for scalable web GUI testing. WebCQ incorporates QTRAN for multi-agent coordination and a lightweight synchronization mechanism, allowing it to work under asynchronous web testing scenarios. It extracts semantic and exploration features for each UI event to form an action vector. This vector is concatenated with the current state vector and fed into the policy network, enabling DQN-based decision making within a dynamic action space. We evaluated WebCQ on eight large-scale commercial websites. Under the same time budget and agent count, WebCQ explored 33.3% more states and executed 42.2% more unique actions than MARG, while triggering more failures on six of the eight websites under test. It also demonstrated strong scalability, maintaining higher action throughput during 20-hour experiments, and achieving greater performance improvements as the number of agents increased. These results show that WebCQovercomes key limitations of existing MARL-based approaches, providing a scalable and effective solution for enhancing modern web GUI testing.
△ Less
Submitted 21 June, 2026;
originally announced June 2026.
-
Spectroscopic fingerprints of a ferroaxial charge density wave
Authors:
Jiangchang Zheng,
Zhongyi Zhang,
Fazhi Yang,
Josh Leeman,
Luanjing Li,
Zihan Lin,
Zijian Fei,
Tianhao Guo,
Siyu Heng,
Xin Liang,
Leslie M. Schoop,
Junzhang Ma,
Hoi Chun Po,
Berthold Jäck
Abstract:
Unconventional charge density waves (CDWs) with complex order parameters can host exotic collective modes and non-trivial topologies. They have emerged as a new frontier in the study of quantum matter. Recent experiments on rare-earth tritellurides have reported evidence for a ferroaxial CDW through the detection of characteristic Raman modes. This phase, often regarded as a hidden order, has been…
▽ More
Unconventional charge density waves (CDWs) with complex order parameters can host exotic collective modes and non-trivial topologies. They have emerged as a new frontier in the study of quantum matter. Recent experiments on rare-earth tritellurides have reported evidence for a ferroaxial CDW through the detection of characteristic Raman modes. This phase, often regarded as a hidden order, has been recognized to arise from the coupling between charge and orbital degrees of freedom in these materials. Yet, spectroscopic insight into its underlying electronic structure and the explicit form of its order parameter symmetry has remained elusive. Here, we present results from linearly polarized angle-resolved photoemission spectroscopy (ARPES) and scanning tunneling microscopy (STM) measurements of the CDW phase in LaTe$_3$. Our ARPES measurements reveal a complex landscape of spectral gaps across the reconstructed Fermi surface, while our STM-based quasiparticle interference (QPI) mapping, enhanced through the selective deposition of atomic scattering centers, directly reveals an inter-orbital CDW with mixed $p_x$-$p_z$ orbital character. The detailed analysis of the QPI characteristics in terms of the order parameter symmetry within the orbital subspace of the Fermi surface suggests a mixed CDW phase with substantial ferroaxial component, which breaks all vertical mirror symmetries. More broadly, our work establishes a powerful spectroscopic pathway, based on scattering off individual atoms, for identifying and characterizing hidden, multi-component electronic orders in quantum materials using STM and ARPES measurements.
△ Less
Submitted 20 June, 2026;
originally announced June 2026.
-
Selection Integrity for LLM Graph Memory: An Accumulability Criterion for Information-Flow-Blind Retrieval
Authors:
Zeming Fei,
Hongming Fei,
Xiaoyang Wang,
Yang yang,
Prosanta Gope,
Biplab Sikdar,
Ying Zhang
Abstract:
Agent memory is moving to graphs, and the provenance defenses now being built for it all check one thing: the provenance of the records an agent retrieves. We show that this entire class of defense is blind by construction. A long-term graph memory runs a global selection step over writable graph structure, so structure that an untrusted principal writes changes \emph{which} authenticated facts ar…
▽ More
Agent memory is moving to graphs, and the provenance defenses now being built for it all check one thing: the provenance of the records an agent retrieves. We show that this entire class of defense is blind by construction. A long-term graph memory runs a global selection step over writable graph structure, so structure that an untrusted principal writes changes \emph{which} authenticated facts are selected while the cited evidence stays fully authenticated; faithful information-flow control (IFC), checking the provenance of what the reader uses (all of it authenticated), makes the byte-identical decision to no defense at all, across document-QA substrates and real multi-session agent memory. In the most consequential instance, a no-source structural write silently misdirects $28$ irreversible ledger transfers over $499$ live actions: faithful IFC permits every one, and \authselect\ prevents every one. We then characterize exactly which memories are exposed: a selector admits the channel when its structural term can reallocate an $Ω(1)$ share of top-$k$ membership past a selected fact's margin. Personalized PageRank can, since a sourceless write reroutes conserved random-walk mass; a content-fixed reranker cannot, and Graphiti's node-distance, which leans on structure \emph{more} than PageRank does, stays immune. Reallocatability, not reliance, is the predictor. We prove the immune case in general and the open case under a chokepoint condition we verify. Closing the channel forces any provenance defense to recompute selection on the authenticated subgraph, which is what \authselect\ does, at zero over-block and $2$--$3\%$ latency.
△ Less
Submitted 10 June, 2026;
originally announced June 2026.
-
A Causal Probabilistic Framework for Perception-Informed Closed-Loop Simulation of Autonomous Driving
Authors:
Zhennan Fei,
Rickard Johansson,
Mikael Andersson,
Matthias Eng,
Mattias Eriksson,
Kaveh Kianfar,
Sadegh Rahrovani,
Chris van der Ploeg,
Michael Borth,
Maren Buermann,
Michiel Braat,
Henk Goossens,
Zijian Han,
Majid Khorsand Vakilzadeh,
Gabriel Rodrigues de Campos
Abstract:
Software-in-the-loop (SIL) simulation is a cornerstone for the validation of modern automotive safety functions. However, many current frameworks utilize ideal sensing, which bypasses the functional insufficiencies of perception algorithms, leading to over-optimistic safety assessments. This paper proposes a perception-informed SIL testing methodology that bridges the gap between ground-truth simu…
▽ More
Software-in-the-loop (SIL) simulation is a cornerstone for the validation of modern automotive safety functions. However, many current frameworks utilize ideal sensing, which bypasses the functional insufficiencies of perception algorithms, leading to over-optimistic safety assessments. This paper proposes a perception-informed SIL testing methodology that bridges the gap between ground-truth simulation and real-world perception behavior. We present a framework for incorporating causal probabilistic models into standardized, scenario-based simulation toolchains, applicable to both Advanced Driver Assistance Systems (ADAS) and Autonomous Driving Systems (ADS). Our approach enables the systematic injection of realistic perception errors, such as loss of detection, sizing inaccuracies, and positioning offsets, derived from physical triggering conditions like fog, rain, and object-merging scenarios. By evaluating these ``faults'' within a standardized simulation environment, we demonstrate that perception-informed testing reveals latent operational risks that ideal SIL environments fail to capture, providing a scalable pathway for SOTIF (ISO 21448) validation.
△ Less
Submitted 5 June, 2026;
originally announced June 2026.
-
Distributed MoE-based Uplink Detection for Cell-Free Communication Systems
Authors:
Le Zhao,
Xuesong Pan,
Xinyi Wang,
Zhong Zheng,
Zesong Fei
Abstract:
Cell-free Massive multiple input and multiple output (MIMO) is recognized as a key technology for beyond-5G networks, where distributed access points (APs) jointly serve user equipments (UEs) to address the inherent inter-cell interference issue inherent in cellular systems. While conventional distributed signal detection methods offer a practical balance between performance and fronthaul load, th…
▽ More
Cell-free Massive multiple input and multiple output (MIMO) is recognized as a key technology for beyond-5G networks, where distributed access points (APs) jointly serve user equipments (UEs) to address the inherent inter-cell interference issue inherent in cellular systems. While conventional distributed signal detection methods offer a practical balance between performance and fronthaul load, they are fundamentally limited by linear processing constraints. In this paper, we propose a novel deep learning based uplink detection framework by introducing the distributed mixture of experts detection network (DMoE-DetNet). In this architecture, each AP acts as a local expert employing convolutional neural networks (CNNs) for non-linear feature extraction, and transmits the local minimum mean square error (MMSE) detection results and statistical channel information to the central processing unit (CPU). In the CPU, an attention-based encoder module captures complex spatio-temporal dependencies among users for global feature fusion, with a gating network at the central processor dynamically weighting the contributions from different APs. At last, a linear detector outputs the symbol probability. Simulation results demonstrate that the proposed DMoE-DetNet significantly outperforms conventional linear processing based cell-free signal detection methods in terms of symbol error rate, showcasing the potential of artificial intelligence-enabled communication systems.
△ Less
Submitted 1 June, 2026;
originally announced June 2026.
-
MOSS-Audio Technical Report
Authors:
Chen Yang,
Chufan Yu,
Hanfu Chen,
Jie Zhu,
Jingqi Chen,
Ke Chen,
Wenxuan Wang,
Yang Wang,
Yaozhou Jiang,
Yi Jiang,
Zhengyuan Lin,
Ziqi Chen,
Zhaoye Fei,
Chenghao Liu,
Donghua Yu,
Jun Zhan,
Kang Yu,
Kexin Huang,
Liwei Fan,
Mingshu Chen,
Qinyuan Cheng,
Ruixiao Li,
Shimin Li,
Songlin Wang,
Xingjian Zhao
, et al. (5 additional authors not shown)
Abstract:
MOSS-Audio is a unified audio-language model for speech, environmental sound, and music understanding, supporting audio captioning, time-aware question answering, timestamped transcription, and audio-grounded reasoning. MOSS-Audio couples a dedicated audio encoder with a modality adapter and a large language model: the encoder produces 12.5 Hz temporal representations, the adapter projects them in…
▽ More
MOSS-Audio is a unified audio-language model for speech, environmental sound, and music understanding, supporting audio captioning, time-aware question answering, timestamped transcription, and audio-grounded reasoning. MOSS-Audio couples a dedicated audio encoder with a modality adapter and a large language model: the encoder produces 12.5 Hz temporal representations, the adapter projects them into the decoder space, and the decoder generates autoregressive text outputs. Two design choices are central to the system: DeepStack cross-layer feature injection, which exposes the decoder to acoustic information from multiple encoder depths, and time markers, which provide explicit temporal cues by inserting timestamp markers into the audio-token stream. At the data level, we design an event-preserving audio annotation pipeline that segments raw audio at coherent event boundaries, applies branch-specific annotation to speech, music, and general audio, and merges the results into unified captions for pretraining. The intermediate branch-specific captions are further retained to support the construction of task-oriented SFT data. The model is pretrained on large-scale audio-language data, with time-aware objectives incorporated to support temporal grounding, and then undergoes multi-stage post-training to enhance instruction following and audio-grounded reasoning. We release 4B and 8B variants in both Instruct and Thinking configurations. MOSS-Audio achieves strong performance across general audio understanding, speech captioning, ASR, and timestamped ASR, positioning it as a promising understanding foundation for future voice agents.
△ Less
Submitted 5 June, 2026; v1 submitted 1 June, 2026;
originally announced June 2026.
-
OpenCompass: A Universal Evaluation Platform for Large Language Models
Authors:
Maosong Cao,
Kai Chen,
Haodong Duan,
Yixiao Fang,
Zhiwei Fei,
Tong Gao,
Ge Jiaye,
Mo Li,
Hongwei Liu,
Junnan Liu,
Yuan Liu,
Chengqi Lyu,
Han Lyu,
Ningsheng Ma,
Zerun Ma,
Yu Sun,
Zhiyong Wu,
Linchen Xiao,
Zhuozhi Xiong,
Jun Xu,
Haochen Ye,
Zhaohui Yu,
Yike Yuan,
Songyang Zhang,
Yufeng Zhao
, et al. (5 additional authors not shown)
Abstract:
In recent years, the field of artificial intelligence has undergone a paradigm shift from task-specific small-scale models to general-purpose large language models (LLMs). With the rapid iteration of LLMs, objective, quantitative, and comprehensive evaluation of their capabilities has become a critical link in advancing technological development. Currently, the mainstream static benchmark dataset-…
▽ More
In recent years, the field of artificial intelligence has undergone a paradigm shift from task-specific small-scale models to general-purpose large language models (LLMs). With the rapid iteration of LLMs, objective, quantitative, and comprehensive evaluation of their capabilities has become a critical link in advancing technological development. Currently, the mainstream static benchmark dataset-based evaluation methods face challenges such as the diversity of task types, inconsistent evaluation criteria, and fragmentation of data and processing workflows, making it difficult to efficiently conduct cross-domain and large-scale model evaluation. To address the aforementioned issues, this paper proposes and open-sources OpenCompass, a one-stop, scalable, and high-concurrency-supported general-purpose LLM evaluation platform. Adhering to the design philosophy of modularization and component decoupling, the platform boasts three core advantages: high compatibility, flexibility, and high concurrency. The core architecture of OpenCompass comprises five key components: the Configuration System, Task Partitioning Module, Execution and Scheduling Module, Task Execution Unit, and Result Visualization Module. Its workflow provides rule-based, LLM-as-a-Judge, and cascaded evaluators to adapt to the requirements of different task scenarios. Supporting mainstream benchmark datasets across multiple domains, including knowledge, reasoning, computation, science, language, code, etc., the platform offers a unified and efficient LLM evaluation tool for both academia and industry, facilitating the accurate identification of strengths and weaknesses of LLMs as well as their subsequent optimization.
△ Less
Submitted 7 June, 2026; v1 submitted 18 May, 2026;
originally announced May 2026.
-
Berry-Phase-Induced Chirality in Thermodynamics
Authors:
Zhaoyu Fei,
Yu-Han Ma
Abstract:
Geometric phases are foundational to isolated quantum systems, yet their thermodynamic role in open systems remains unrevealed Developing a dissipative adiabatic perturbation expansion, we discover a Berry-phase-induced chiral work difference that survives decoherence. This chirality evolves from an interferometric thermodynamic Aharonov-Bohm effect in the unitary regime to a fringe-free signal in…
▽ More
Geometric phases are foundational to isolated quantum systems, yet their thermodynamic role in open systems remains unrevealed Developing a dissipative adiabatic perturbation expansion, we discover a Berry-phase-induced chiral work difference that survives decoherence. This chirality evolves from an interferometric thermodynamic Aharonov-Bohm effect in the unitary regime to a fringe-free signal in the dissipative regime. We illustrate this framework in a two-level system and assess its experimental feasibility. Our findings clarify the role of quantum geometry in the geometric formulation of thermodynamics.
△ Less
Submitted 13 May, 2026;
originally announced May 2026.
-
World Action Models: The Next Frontier in Embodied AI
Authors:
Siyin Wang,
Junhao Shi,
Zhaoyang Fu,
Xinzhe He,
Feihong Liu,
Chenchen Yang,
Yikang Zhou,
Zhaoye Fei,
Jingjing Gong,
Jinlan Fu,
Mike Zheng Shou,
Xuanjing Huang,
Xipeng Qiu,
Yu-Gang Jiang
Abstract:
Vision-Language-Action (VLA) models have achieved strong semantic generalization for embodied policy learning, yet they learn reactive observation-to-action mappings without explicitly modeling how the physical world evolves under intervention. A growing body of work addresses this limitation by integrating world models, predictive models of environment dynamics, into the action generation pipelin…
▽ More
Vision-Language-Action (VLA) models have achieved strong semantic generalization for embodied policy learning, yet they learn reactive observation-to-action mappings without explicitly modeling how the physical world evolves under intervention. A growing body of work addresses this limitation by integrating world models, predictive models of environment dynamics, into the action generation pipeline. We term this emerging paradigm World Action Models (WAMs): embodied foundation models that unify predictive state modeling with action generation, targeting a joint distribution over future states and actions rather than actions alone. However, the literature remains fragmented across architectures, learning objectives, and application scenarios, lacking a unified conceptual framework. We formally define WAMs and disambiguate them from related concepts, and trace the foundations and early integration of VLA and world model research that gave rise to this paradigm. We organize existing methods into a structured taxonomy of Cascaded and Joint WAMs, with further subdivision by generation modality, conditioning mechanism, and action decoding strategy. We systematically analyze the data ecosystem fueling WAMs development, spanning robot teleoperation, portable human demonstrations, simulation, and internet-scale egocentric video, and synthesize emerging evaluation protocols organized around visual fidelity, physical commonsense, and action plausibility. Overall, this survey provides the first systematic account of the WAMs landscape, clarifies key architectural paradigms and their trade-offs, and identifies open challenges and future opportunities for this rapidly evolving field.
△ Less
Submitted 12 May, 2026;
originally announced May 2026.
-
SafeTune: Search-based Harmfulness Minimisation for Large Language Models
Authors:
Giordano d'Aloisio,
David Williams,
Giusy Annunziata,
Zhiwei Fei,
Antinisca Di Marco,
Federica Sarro
Abstract:
The widespread adoption of Large Language Models (LLMs) raises concerns about the potential harmfulness of their responses. In this paper, we first investigate the harmfulness of responses from four general-purpose LLMs. Next, we propose SafeTune, a multi-objective search-based approach to mitigate harmfulness while increasing response relevance through hyperparameter tuning and system prompt engi…
▽ More
The widespread adoption of Large Language Models (LLMs) raises concerns about the potential harmfulness of their responses. In this paper, we first investigate the harmfulness of responses from four general-purpose LLMs. Next, we propose SafeTune, a multi-objective search-based approach to mitigate harmfulness while increasing response relevance through hyperparameter tuning and system prompt engineering. Our initial evaluation shows that SafeTune significantly reduces the rate of harmful responses generated by Qwen3.5 0.8B and increases prompt-response relevance (both with a large effect size). Among the parameters we explore, we also find that encouraging greater repetition in responses is most impactful in reducing harmfulness while increasing relevance.
△ Less
Submitted 8 May, 2026;
originally announced May 2026.
-
Towards Intelligent Low-Altitude Wireless Network Deployment: Differentiable Channel Knowledge Map Construction and Trajectory Design
Authors:
Le Zhao,
Zesong Fei,
Wenge Shi,
Xinyi Wang,
Jingxuan Huang,
Jihao Luo,
Yong Zeng
Abstract:
Channel knowledge map (CKM) has emerged as a promising technique to leverage prior propagation knowledge in low-altitude wireless networks (LAWNs), yet state-of-the-art grid-based CKM construction methods struggle to support efficient LAWN deployment due to their lack of differentiability with respect to continuous locations of unmanned aerial vehicles (UAVs). To overcome this limitation, we propo…
▽ More
Channel knowledge map (CKM) has emerged as a promising technique to leverage prior propagation knowledge in low-altitude wireless networks (LAWNs), yet state-of-the-art grid-based CKM construction methods struggle to support efficient LAWN deployment due to their lack of differentiability with respect to continuous locations of unmanned aerial vehicles (UAVs). To overcome this limitation, we propose a differentiable CKM-triggered trajectory optimization framework for LAWNs. Firstly, we propose a location-oriented CKM construction method that directly maps continuous spatial coordinates to channel gain. In particular, a shared convolutional neural network (CNN) is employed to encode high-level environmental features from conditional inputs. These features are then sampled based on location information to form a fused regressor-conditional multilayer perceptron (c-MLP) or conditional Kolmogorov-Arnold network (cKAN)-for channel gain prediction. We further propose a joint power, bandwidth, and trajectory optimization (JPBTO) method for multi-UAV systems, with the constructed differentiable CKM employed to evaluate the communication performance. The formulated non-convex problem is solved via alternating optimization and successive convex approximation. Numerical results show that the proposed framework enables location-aware differentiability of the CKM, while achieving higher accuracy than the methods without environmental features. Furthermore, the proposed CKM-JPBTO achieves a significantly higher minimum throughput than the conventional statistical channel model-based JPBTO.
△ Less
Submitted 7 May, 2026;
originally announced May 2026.
-
Misrouter: Exploiting Routing Mechanisms for Input-Only Attacks on Mixture-of-Experts LLMs
Authors:
Zekun Fei,
Zihao Wang,
Weijie Liu,
Ruiqi He,
Jianing Geng,
Zheli Liu,
XiaoFeng Wang
Abstract:
Mixture-of-Experts (MoE) architectures have emerged as a leading paradigm for scaling large language models through sparse, routing-based computation. However, this design introduces a new attack surface: the routing mechanism that determines which experts process each input. Prior work shows that manipulating routing can bypass safety alignment, but existing attacks require model modification and…
▽ More
Mixture-of-Experts (MoE) architectures have emerged as a leading paradigm for scaling large language models through sparse, routing-based computation. However, this design introduces a new attack surface: the routing mechanism that determines which experts process each input. Prior work shows that manipulating routing can bypass safety alignment, but existing attacks require model modification and thus apply only to locally deployed models. By contrast, real-world LLM services are remotely hosted and accessible only through input queries. This raises a fundamental question: can MoE routing be exploited through input-only attacks to induce stronger unsafe behaviors in real-world services? Our key insight is to optimize attacks in a white-box setting on open-source surrogate MoE models and transfer the resulting adversarial inputs to public API services within the same model family. This setting presents three main challenges: routing can be influenced only indirectly through input perturbations, routing control and output generation are tightly coupled, and even a successful safety bypass may still produce low-quality responses. To address these challenges, we propose Misrouter, an input-only attack framework that jointly targets routing behavior and expert functionality. Misrouter identifies weakly aligned experts that are willing to produce target harmful content by analyzing expert activations under harmful queries paired with unsafe continuations. It then optimizes adversarial inputs to steer routing toward these experts and away from strongly aligned ones. It further biases routing toward highly capable general-purpose experts identified from benign question-answering tasks. Finally, because routing and output objectives can conflict, Misrouter uses a two-phase optimization strategy that first steers routing and then optimizes harmful outputs while preserving routing stability.
△ Less
Submitted 5 May, 2026;
originally announced May 2026.
-
From Concept to Capability: Evaluating 3D Gaussian Splatting for Synthetic Scene Editing in Autonomous Driving
Authors:
Ali Nouri,
Yifei Zhang,
Yifan Zhang,
Tayssir Bouraffa,
Zhennan Fei,
Zijian Han,
Håkan Sivencrona,
Anders Heyden
Abstract:
The perception of an Autonomous Driving System (ADS) critically depends on relevant, comprehensive, and diverse datasets to ensure its safety while operating in the environment. Field data collection lacks completeness with respect to the list of rare but still possible safety-related scenarios needed for the development, verification, and validation of the ADS. 3D Gaussian Splatting (3DGS) has sh…
▽ More
The perception of an Autonomous Driving System (ADS) critically depends on relevant, comprehensive, and diverse datasets to ensure its safety while operating in the environment. Field data collection lacks completeness with respect to the list of rare but still possible safety-related scenarios needed for the development, verification, and validation of the ADS. 3D Gaussian Splatting (3DGS) has shown promising capabilities for the reconstruction and editing of scenes based on data collected by cameras and LiDAR sensors. However, the industrial fidelity evaluation of reconstructions is underexplored, which is crucial when employing such methods in safety-related systems, especially for ADS. This becomes more challenging as ADS operates in a dynamic, uncontrolled environment with limited viewpoints and often partially occluded objects. This paper addresses this gap by proposing and implementing a framework (Fig. 1) to systematically analyze the capabilities and limitations of 3DGS for use in the reconstruction of safety-related scenes. It focuses on the quality of reconstruction for vehicles and pedestrians, which are the two most critical object classes for ADS. Our findings provide industry insights into the fidelity degradation of reconstructions from multiple novel viewpoints, both lateral and longitudinal, enabling the integration of these methods into real-world industrial AD software development and testing pipelines.
△ Less
Submitted 3 May, 2026;
originally announced May 2026.
-
Toward Low-Altitude Embodied Intelligence: A Sensing-Communication-Computation-Control Closed-Loop Perspective
Authors:
Jihao Luo,
Zesong Fei,
Xinyi Wang,
Shuntian Tang,
Zilong Liu,
Yiqing Zhou
Abstract:
The rapid growth of the low-altitude economy drives increasingly autonomous unmanned aerial vehicle (UAV) operations, giving rise to low-altitude embodied intelligence (LAEI), in which sensing, communication, computation, and control (SC$^3$) are tightly integrated to enable closed-loop interaction, ensuring timely, effective, and safe responses in complex or unknown environments. This article sys…
▽ More
The rapid growth of the low-altitude economy drives increasingly autonomous unmanned aerial vehicle (UAV) operations, giving rise to low-altitude embodied intelligence (LAEI), in which sensing, communication, computation, and control (SC$^3$) are tightly integrated to enable closed-loop interaction, ensuring timely, effective, and safe responses in complex or unknown environments. This article systematically explores the LAEI networks, from its fundamental architecture to the diverse scenarios that it can support. We examine key enabling techniques that sustain timely information exchange and effective decision feedback within the $\text{SC}^3$ closed loop. A representative low-altitude UAV mission in an unknown urban area is presented as a case study, where the UAV provides communication services and performs environmental sensing to inform closed-loop control, illustrating how coordinated $\text{SC}^3$ capabilities enable efficient and responsive operation. By identifying major challenges and outlining future research directions, this work serves as a cornerstone for developing next-generation low-altitude intelligent systems.
△ Less
Submitted 27 April, 2026;
originally announced April 2026.
-
SkillMOO: Multi-Objective Optimization of Agent Skills for Software Engineering
Authors:
Jingzhi Gong,
Ruizhen Gu,
Zhiwei Fei,
Yazhuo Cao,
Lukas Twist,
Alina Geiger,
Shuo Han,
Dominik Sobania,
Federica Sarro,
Jie M. Zhang
Abstract:
Agent skills are increasingly used to configure coding agents for software engineering (SE) tasks, yet current practice treats them as static, hand-crafted assets, or evolved on pass rate alone. This is insufficient: a skill can improve task success while substantially raising token cost, or introducing misleading guidance. We argue that SE agent skill bundles can be treated as multi-objective sea…
▽ More
Agent skills are increasingly used to configure coding agents for software engineering (SE) tasks, yet current practice treats them as static, hand-crafted assets, or evolved on pass rate alone. This is insufficient: a skill can improve task success while substantially raising token cost, or introducing misleading guidance. We argue that SE agent skill bundles can be treated as multi-objective search objects and present SkillMOO, a framework that evolves skill bundles through LLM-proposed edits and NSGA-II Pareto selection on pass rate and inference cost. Evaluated across all 16 SkillsBench SE tasks, SkillMOO achieves the top pass rate rank on 11 of 12 non-zero-pass tasks while achieving cost reductions of up to 31.7% over static bundles, with pass rate gains up to 21 percentage points. Analysis of 38 skill edits shows that pruning and substitution dominate successful operations, offering actionable principles for skill bundle design. Thereby, the current practice of deploying skills without cost-aware validation leaves better skill configurations unexplored, motivating a new class of cost-aware, search-based skill engineering.
△ Less
Submitted 5 August, 2026; v1 submitted 10 April, 2026;
originally announced April 2026.
-
MOSS-VoiceGenerator: Create Realistic Voices with Natural Language Descriptions
Authors:
Kexin Huang,
Liwei Fan,
Botian Jiang,
Yaozhou Jiang,
Qian Tu,
Jie Zhu,
Yuqian Zhang,
Yiwei Zhao,
Chenchen Yang,
Zhaoye Fei,
Shimin Li,
Xiaogui Yang,
Qinyuan Cheng,
Xipeng Qiu
Abstract:
Voice design from natural language aims to generate speaker timbres directly from free-form textual descriptions, allowing users to create voices tailored to specific roles, personalities, and emotions. Such controllable voice creation benefits a wide range of downstream applications-including storytelling, game dubbing, role-play agents, and conversational assistants, making it a significant task…
▽ More
Voice design from natural language aims to generate speaker timbres directly from free-form textual descriptions, allowing users to create voices tailored to specific roles, personalities, and emotions. Such controllable voice creation benefits a wide range of downstream applications-including storytelling, game dubbing, role-play agents, and conversational assistants, making it a significant task for modern Text-to-Speech models. However, existing models are largely trained on carefully recorded studio data, which produces speech that is clean and well-articulated, yet lacks the lived-in qualities of real human voices. To address these limitations, we present MOSS-VoiceGenerator, an open-source instruction-driven voice generation model that creates new timbres directly from natural language prompts. Motivated by the hypothesis that exposure to real-world acoustic variation produces more perceptually natural voices, we train on large-scale expressive speech data sourced from cinematic content. Subjective preference studies demonstrate its superiority in overall performance, instruction-following, and naturalness compared to other voice design models.
△ Less
Submitted 30 March, 2026;
originally announced March 2026.
-
Not All Entities are Created Equal: A Dynamic Anonymization Framework for Privacy-Preserving RAG
Authors:
Xinyuan Zhu,
Zekun Fei,
Enye Wang,
Ruiqi He,
Jia Guo,
Ruijie Wang,
Zheli Liu,
Qingkai Zeng
Abstract:
Retrieval-Augmented Generation (RAG) enhances the utility of Large Language Models (LLMs) by retrieving external documents. Since the knowledge databases in RAG are predominantly utilized via cloud services, private data in sensitive domains such as finance and healthcare faces the risk of personal information leakage. Thus, effectively anonymizing knowledge bases is crucial for privacy preservati…
▽ More
Retrieval-Augmented Generation (RAG) enhances the utility of Large Language Models (LLMs) by retrieving external documents. Since the knowledge databases in RAG are predominantly utilized via cloud services, private data in sensitive domains such as finance and healthcare faces the risk of personal information leakage. Thus, effectively anonymizing knowledge bases is crucial for privacy preservation. Existing studies equate the privacy risk of text to the linear superposition of the privacy risks of individual, isolated sensitive entities. The "one-size-fits-all" full processing of all sensitive entities severely degrades utility of LLM. To address this issue, we introduce a dynamic anonymization framework named TRIP-RAG. Based on context-aware entity quantification, this framework evaluates entities from the perspectives of marginal privacy risk, knowledge divergence, and topical relevance. It identifies highly sensitive entities while trading off utility, providing a feasible approach for variable-intensity privacy protection scenarios. Our theoretical analysis and experiments indicate that TRIP-RAG can effectively reduce context inference risks. Extensive experimental results demonstrate that, while maintaining privacy protection comparable to full anonymization, TRIP-RAG's Recall@k decreases by less than 35% compared to the original data, and the generation quality improves by up to 56% over existing baselines.
△ Less
Submitted 28 May, 2026; v1 submitted 27 March, 2026;
originally announced March 2026.
-
Towards Semantic-based Agent Communication Networks: Vision, Technologies, and Challenges
Authors:
Ping Zhang,
Rui Meng,
Xiaodong Xu,
Yaheng Wang,
Zixuan Huang,
Yiming Liu,
Ruichen Zhang,
Yinqiu Liu,
Haonan Tong,
Huishi Song,
Gang Wu,
Zhaoming Lu,
Jiawen Kang,
Geng Sun,
Qinghe Du,
Zhaohui Yang,
Jingxuan Zhang,
Han Meng,
Lexi Xu,
Haitao Zhao,
Zesong Fei,
Yiqing Zhou,
Pei Xiao,
Meixia Tao,
Qinyu Zhang
, et al. (2 additional authors not shown)
Abstract:
The International Telecommunication Union (ITU) identifies "Artificial Intelligence (AI) and Communication" as one of six key usage scenarios for 6G. Agentic AI, characterized by its ca-pabilities in multi-modal environmental sensing, complex task coordination, and continuous self-optimization, is anticipated to drive the evolution toward agent-based communication net-works. Semantic communication…
▽ More
The International Telecommunication Union (ITU) identifies "Artificial Intelligence (AI) and Communication" as one of six key usage scenarios for 6G. Agentic AI, characterized by its ca-pabilities in multi-modal environmental sensing, complex task coordination, and continuous self-optimization, is anticipated to drive the evolution toward agent-based communication net-works. Semantic communication (SemCom), in turn, has emerged as a transformative paradigm that offers task-oriented efficiency, enhanced reliability in complex environments, and dynamic adaptation in resource allocation. However, comprehensive reviews that trace their technologi-cal evolution in the contexts of agent communications remain scarce. Addressing this gap, this paper systematically explores the role of semantics in agent communication networks. We first propose a novel architecture for semantic-based agent communication networks, structured into three layers, four entities, and four stages. Three wireless agent network layers define the logical structure and organization of entity interactions: the intention extraction and understanding layer, the semantic encoding and processing layer, and the distributed autonomy and collabora-tion layer. Across these layers, four AI agent entities, namely embodied agents, communication agents, network agents, and application agents, coexist and perform distinct tasks. Furthermore, four operational stages of semantic-enhanced agentic AI systems, namely perception, memory, reasoning, and action, form a cognitive cycle guiding agent behavior. Based on the proposed architecture, we provide a comprehensive review of the state-of-the-art on how semantics en-hance agent communication networks. Finally, we identify key challenges and present potential solutions to offer directional guidance for future research in this emerging field.
△ Less
Submitted 25 March, 2026;
originally announced March 2026.
-
Rateless DeepJSCC for Broadcast Channels: a Rate-Distortion-Complexity Tradeoff
Authors:
Zijun Qin,
Jingxuan Huang,
Zesong Fei,
Haichuan Ding,
Yulin Shao,
Xianhao Chen
Abstract:
In recent years, numerous data-intensive broadcasting applications have emerged at the wireless edge, calling for a flexible tradeoff between distortion, transmission rate, and processing complexity. While deep learning-based joint source-channel coding (DeepJSCC) has been identified as a potential solution to data-intensive communications, most of these schemes are confined to worst-case solution…
▽ More
In recent years, numerous data-intensive broadcasting applications have emerged at the wireless edge, calling for a flexible tradeoff between distortion, transmission rate, and processing complexity. While deep learning-based joint source-channel coding (DeepJSCC) has been identified as a potential solution to data-intensive communications, most of these schemes are confined to worst-case solutions, lack adaptive complexity, and are inefficient in broadcast settings. To overcome these limitations, this paper introduces nonlinear transform rateless source-channel coding (NTRSCC), a variable-length JSCC framework for broadcast channels based on rateless codes. In particular, we integrate learned source transformations with physical-layer LT codes, develop unequal protection schemes that exploit decoder side information, and devise approximations to enable end-to-end optimization of rateless parameters. Our framework enables heterogeneous receivers to adaptively adjust their received number of rateless symbols and decoding iterations in belief propagation, thereby achieving a controllable tradeoff between distortion, rate, and decoding complexity. Simulation results demonstrate that the proposed method enhances image broadcast quality under stringent communication and processing budgets over heterogeneous edge devices.
△ Less
Submitted 23 March, 2026;
originally announced March 2026.
-
MOSS-TTSD: Text to Spoken Dialogue Generation
Authors:
Yuqian Zhang,
Donghua Yu,
Zhengyuan Lin,
Botian Jiang,
Mingshu Chen,
Yaozhou Jiang,
Yiwei Zhao,
Yiyang Zhang,
Yucheng Yuan,
Hanfu Chen,
Kexin Huang,
Jun Zhan,
Cheng Chang,
Zhaoye Fei,
Shimin Li,
Xiaogui Yang,
Qinyuan Cheng,
Xipeng Qiu
Abstract:
Spoken dialogue generation is crucial for applications like podcasts, dynamic commentary, and entertainment content, but poses significant challenges compared to single-utterance text-to-speech (TTS). Key requirements include accurate turn-taking, cross-turn acoustic consistency, and long-form stability, which current models often fail to address due to a lack of dialogue context modeling. To brid…
▽ More
Spoken dialogue generation is crucial for applications like podcasts, dynamic commentary, and entertainment content, but poses significant challenges compared to single-utterance text-to-speech (TTS). Key requirements include accurate turn-taking, cross-turn acoustic consistency, and long-form stability, which current models often fail to address due to a lack of dialogue context modeling. To bridge this gap, we present MOSS-TTSD, a spoken dialogue synthesis model designed for expressive, multi-party conversational speech across multiple languages. With enhanced long-context modeling, MOSS-TTSD generates long-form spoken conversations from dialogue scripts with explicit speaker tags, supporting up to 60 minutes of single-pass synthesis, multi-party dialogue with up to 5 speakers, and zero-shot voice cloning from a short reference audio clip. The model supports various mainstream languages, including English and Chinese, and is adapted to several long-form scenarios. Additionally, to address limitations of existing evaluation methods, we propose TTSD-eval, an objective evaluation framework based on forced alignment that measures speaker attribution accuracy and speaker similarity without relying on speaker diarization tools. Both objective and subjective evaluation results show that MOSS-TTSD surpasses strong open-source and proprietary baselines in dialogue synthesis.
△ Less
Submitted 20 March, 2026;
originally announced March 2026.
-
MOSS-TTS Technical Report
Authors:
Yitian Gong,
Botian Jiang,
Yiwei Zhao,
Yucheng Yuan,
Kuangwei Chen,
Yaozhou Jiang,
Cheng Chang,
Dong Hong,
Mingshu Chen,
Ruixiao Li,
Yiyang Zhang,
Yang Gao,
Hanfu Chen,
Ke Chen,
Songlin Wang,
Xiaogui Yang,
Yuqian Zhang,
Kexin Huang,
ZhengYuan Lin,
Kang Yu,
Ziqi Chen,
Jin Wang,
Zhaoye Fei,
Qinyuan Cheng,
Shimin Li
, et al. (1 additional authors not shown)
Abstract:
This technical report presents MOSS-TTS, a speech generation foundation model built on a scalable recipe: discrete audio tokens, autoregressive modeling, and large-scale pretraining. Built on MOSS-Audio-Tokenizer, a causal Transformer tokenizer that compresses 24 kHz audio to 12.5 fps with variable-bitrate RVQ and unified semantic-acoustic representations, we release two complementary generators:…
▽ More
This technical report presents MOSS-TTS, a speech generation foundation model built on a scalable recipe: discrete audio tokens, autoregressive modeling, and large-scale pretraining. Built on MOSS-Audio-Tokenizer, a causal Transformer tokenizer that compresses 24 kHz audio to 12.5 fps with variable-bitrate RVQ and unified semantic-acoustic representations, we release two complementary generators: MOSS-TTS, which emphasizes structural simplicity, scalability, and long-context/control-oriented deployment, and MOSS-TTS-Local-Transformer, which introduces a frame-local autoregressive module for higher modeling efficiency, stronger speaker preservation, and a shorter time to first audio. Across multilingual and open-domain settings, MOSS-TTS supports zero-shot voice cloning, token-level duration control, phoneme-/pinyin-level pronunciation control, smooth code-switching, and stable long-form generation. This report summarizes the design, training recipe, and empirical characteristics of the released models.
△ Less
Submitted 20 March, 2026; v1 submitted 18 March, 2026;
originally announced March 2026.
-
Echo: Graph-Enhanced Retrieval and Execution Feedback for Issue Reproduction Test Generation
Authors:
Zhiwei Fei,
Yue Pan,
Federica Sarro,
Jidong Ge,
Marc Liu,
Vincent Ng,
He Ye
Abstract:
Identifying the root cause of a bug remains difficult for many developers because bug reports often lack a bug reproducing test case that reliably triggers the failure. Manually writing such test cases is time-consuming and requires substantial effort to understand the codebase and isolate the failing behavior. To address this challenge, we propose Echo, an agent for generating issue reproducing t…
▽ More
Identifying the root cause of a bug remains difficult for many developers because bug reports often lack a bug reproducing test case that reliably triggers the failure. Manually writing such test cases is time-consuming and requires substantial effort to understand the codebase and isolate the failing behavior. To address this challenge, we propose Echo, an agent for generating issue reproducing test cases, which advances previous work in several ways. During generation, Echo strengthens context retrieval by leveraging a code graph and a novel automatic query-refinement strategy. Echo also improves upon previous tools by automatically executing generated test cases, a first-of-its-kind feature that seamlessly integrates into practical development workflows. In addition, Echo generates potential patches and uses the patched version to validate whether a candidate test meets the fail-to-pass criterion and to provide actionable feedback for refinement. Unlike prior bug-reproduction agents that sample and rank multiple candidate tests, Echo generates a single test per issue, offering a better cost-performance trade-off. Experiments on SWT-Bench Verified show that Echo establishes a new state of the art among open-source approaches, achieving a 66.28% success rate.
△ Less
Submitted 7 March, 2026;
originally announced March 2026.
-
BioLLMAgent: A Hybrid Framework with Enhanced Structural Interpretability for Simulating Human Decision-Making in Computational Psychiatry
Authors:
Zuo Fei,
Kezhi Wang,
Xiaomin Chen,
Yizhou Huang
Abstract:
Computational psychiatry faces a fundamental trade-off: traditional reinforcement learning (RL) models offer interpretability but lack behavioral realism, while large language model (LLM) agents generate realistic behaviors but lack structural interpretability. We introduce BioLLMAgent, a novel hybrid framework that combines validated cognitive models with the generative capabilities of LLMs. The…
▽ More
Computational psychiatry faces a fundamental trade-off: traditional reinforcement learning (RL) models offer interpretability but lack behavioral realism, while large language model (LLM) agents generate realistic behaviors but lack structural interpretability. We introduce BioLLMAgent, a novel hybrid framework that combines validated cognitive models with the generative capabilities of LLMs. The framework comprises three core components: (i) an Internal RL Engine for experience-driven value learning; (ii) an External LLM Shell for high-level cognitive strategies and therapeutic interventions; and (iii) a Decision Fusion Mechanism for integrating components via weighted utility. Comprehensive experiments on the Iowa Gambling Task (IGT) across six clinical and healthy datasets demonstrate that BioLLMAgent accurately reproduces human behavioral patterns while maintaining excellent parameter identifiability (correlations $>0.67$). Furthermore, the framework successfully simulates cognitive behavioral therapy (CBT) principles and reveals, through multi-agent dynamics, that community-wide educational interventions may outperform individual treatments. Validated across reward-punishment learning and temporal discounting tasks, BioLLMAgent provides a structurally interpretable "computational sandbox" for testing mechanistic hypotheses and intervention strategies in psychiatric research.
△ Less
Submitted 5 March, 2026;
originally announced March 2026.
-
SkyReels-V4: Multi-modal Video-Audio Generation, Inpainting and Editing model
Authors:
Guibin Chen,
Dixuan Lin,
Jiangping Yang,
Youqiang Zhang,
Zhengcong Fei,
Debang Li,
Sheng Chen,
Chaofeng Ao,
Nuo Pang,
Yiming Wang,
Yikun Dou,
Zheng Chen,
Mingyuan Fan,
Tuanhui Li,
Mingshan Chang,
Hao Zhang,
Xiaopeng Sun,
Jingtao Xu,
Yuqiang Xie,
Jiahua Wang,
Zhiheng Xu,
Weiming Xiong,
Yuzhe Jin,
Baoxuan Gu,
Binjie Mao
, et al. (26 additional authors not shown)
Abstract:
SkyReels V4 is a unified multi modal video foundation model for joint video audio generation, inpainting, and editing. The model adopts a dual stream Multimodal Diffusion Transformer (MMDiT) architecture, where one branch synthesizes video and the other generates temporally aligned audio, while sharing a powerful text encoder based on the Multimodal Large Language Models (MLLM). SkyReels V4 accept…
▽ More
SkyReels V4 is a unified multi modal video foundation model for joint video audio generation, inpainting, and editing. The model adopts a dual stream Multimodal Diffusion Transformer (MMDiT) architecture, where one branch synthesizes video and the other generates temporally aligned audio, while sharing a powerful text encoder based on the Multimodal Large Language Models (MLLM). SkyReels V4 accepts rich multi modal instructions, including text, images, video clips, masks, and audio references. By combining the MLLMs multi modal instruction following capability with in context learning in the video branch MMDiT, the model can inject fine grained visual guidance under complex conditioning, while the audio branch MMDiT simultaneously leverages audio references to guide sound generation. On the video side, we adopt a channel concatenation formulation that unifies a wide range of inpainting style tasks, such as image to video, video extension, and video editing under a single interface, and naturally extends to vision referenced inpainting and editing via multi modal prompts. SkyReels V4 supports up to 1080p resolution, 32 FPS, and 15 second duration, enabling high fidelity, multi shot, cinema level video generation with synchronized audio. To make such high resolution, long-duration generation computationally feasible, we introduce an efficiency strategy: Joint generation of low resolution full sequences and high-resolution keyframes, followed by dedicated super-resolution and frame interpolation models. To our knowledge, SkyReels V4 is the first video foundation model that simultaneously supports multi-modal input, joint video audio generation, and a unified treatment of generation, inpainting, and editing, while maintaining strong efficiency and quality at cinematic resolutions and durations.
△ Less
Submitted 18 March, 2026; v1 submitted 25 February, 2026;
originally announced February 2026.
-
Passive Imaging with Ambient Noise Under Wave Speed Mismatch: Mathematical Analysis and Wave Speed Estimation
Authors:
Zetao Fei,
Josselin Garnier
Abstract:
It is known that waves generated by ambient noise sources and recorded by passive receivers can be used to image the reflectivities of an unknown medium. However, reconstructing the reflectivity of the medium from partial boundary measurements remains a challenging problem, particularly when the background wave speed is unknown. In this paper, we investigate passive correlation-based imaging in th…
▽ More
It is known that waves generated by ambient noise sources and recorded by passive receivers can be used to image the reflectivities of an unknown medium. However, reconstructing the reflectivity of the medium from partial boundary measurements remains a challenging problem, particularly when the background wave speed is unknown. In this paper, we investigate passive correlation-based imaging in the daylight configuration, where uncontrolled noise sources illuminate the medium and only ambient fields are recorded by a sensor array. We first analyze daylight migration for a point reflector embedded in a homogeneous background. By introducing a searching wave speed into the migration functional, we derive an explicit characterization of the deterministic shift and defocusing effects induced by wave-speed mismatch. We show that the maximum of the envelope of the resulting functional provides a reliable estimator of the true wave speed. We then extend the analysis to a random medium with correlation length smaller than the wavelength. Leveraging the shift formula obtained in the homogeneous case, we introduce a virtual guide star that remains fixed under migration with different searching speeds. This property enables an effective wave-speed estimation strategy based on spatial averaging around the virtual guide star. For both homogeneous and random media, we establish resolution analyses for the proposed wave-speed estimators. Numerical experiments are conducted to validate the theoretical result.
△ Less
Submitted 17 February, 2026;
originally announced February 2026.
-
MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models
Authors:
Yitian Gong,
Kuangwei Chen,
Zhaoye Fei,
Xiaogui Yang,
Ke Chen,
Yang Wang,
Kexin Huang,
Mingshu Chen,
Ruixiao Li,
Qingyuan Cheng,
Shimin Li,
Xipeng Qiu
Abstract:
Discrete audio tokenizers are fundamental to empowering large language models with native audio processing and generation capabilities. Despite recent progress, existing approaches often rely on pretrained encoders, semantic distillation, or heterogeneous CNN-based architectures. These designs introduce fixed inductive biases that limit reconstruction fidelity and hinder effective scaling. In this…
▽ More
Discrete audio tokenizers are fundamental to empowering large language models with native audio processing and generation capabilities. Despite recent progress, existing approaches often rely on pretrained encoders, semantic distillation, or heterogeneous CNN-based architectures. These designs introduce fixed inductive biases that limit reconstruction fidelity and hinder effective scaling. In this paper, we argue that discrete audio tokenization should be learned fully end-to-end using a homogeneous and scalable architecture. To this end, we first propose CAT (Causal Audio Tokenizer with Transformer), a purely Transformer-based architecture that jointly optimizes the encoder, quantizer, and decoder from scratch for high-fidelity reconstruction. Building on the CAT architecture, we develop MOSS-Audio-Tokenizer, a large-scale audio tokenizer featuring 1.6 billion parameters, pre-trained on 3 million hours of diverse, general audio data. We show that this simple, fully end-to-end approach built from homogeneous, causal Transformer blocks scales gracefully and supports high-fidelity reconstruction across diverse audio domains. Across speech, sound, and music, MOSS-Audio-Tokenizer consistently outperforms prior codecs over a wide range of bitrates, while exhibiting predictable improvements with increased scale. Notably, leveraging the discrete tokens from our model, we develop the first purely autoregressive TTS model that surpasses prior non-autoregressive and cascaded systems. Furthermore, MOSS-Audio-Tokenizer enables competitive ASR performance without auxiliary encoders. Our findings position the CAT architecture as a unified, scalable interface for the next generation of native audio foundation models.
△ Less
Submitted 11 February, 2026; v1 submitted 11 February, 2026;
originally announced February 2026.
-
Environment-in-the-Loop: Rethinking Code Migration with LLM-based Agents
Authors:
Xiang Li,
Zhiwei Fei,
Ying Ma,
Jerry Zhang,
Sarro Federica,
He Ye
Abstract:
Modern software systems continuously undergo code upgrades to enhance functionality, security, and performance, and Large Language Models (LLMs) have demonstrated remarkable capabilities in code migration tasks. However, while research on automated code migration which including refactoring, API adaptation, and dependency updates has advanced rapidly, the exploration of the automated environment i…
▽ More
Modern software systems continuously undergo code upgrades to enhance functionality, security, and performance, and Large Language Models (LLMs) have demonstrated remarkable capabilities in code migration tasks. However, while research on automated code migration which including refactoring, API adaptation, and dependency updates has advanced rapidly, the exploration of the automated environment interaction that must accompany it remains relatively scarce. In practice, code and its environment are intricately intertwined. Relying solely on static analysis of the environment leads to an inadequate understanding of the target setting, prolongs feedback cycles, and consequently causes significant rework and project delays, thereby reducing overall efficiency. We contend that successful software evolution demands a holistic perspective that integrates both code and environment migration. To understand the current landscape and challenges, we first provide an overview of the status of automated environment construction. We then propose a novel framework paradigm that tightly integrates automated environment setup with the code migration workflow. Finally, we explore the challenges and future directions for automated environment interaction within the code migration domain. Our findings emphasize that without automated environment interaction, the automation of code migration is only half complete.
△ Less
Submitted 10 February, 2026;
originally announced February 2026.
-
JMigBench: A Benchmark for Evaluating LLMs on Source Code Migration (Java 8 to Java 11)
Authors:
Nishil Amin,
Zhiwei Fei,
Xiang Li,
Justyna Petke,
He Ye
Abstract:
We build a benchmark to evaluate large language models (LLMs) for source code migration tasks, specifically upgrading functions from Java 8 to Java 11. We first collected a dataset of function pairs from open-source repositories, but limitations in data quality led us to construct a refined dataset covering eight categories of deprecated APIs. Using this dataset, the Mistral Codestral model was ev…
▽ More
We build a benchmark to evaluate large language models (LLMs) for source code migration tasks, specifically upgrading functions from Java 8 to Java 11. We first collected a dataset of function pairs from open-source repositories, but limitations in data quality led us to construct a refined dataset covering eight categories of deprecated APIs. Using this dataset, the Mistral Codestral model was evaluated with CodeBLEU and keyword-based metrics to measure lexical and semantic similarity as well as migration correctness. Results show that the evaluated model (Mistral Codestral) can handle trivial one-to-one API substitutions with moderate success, achieving identical migrations in 11.11% of the cases, but it struggles with more complex migrations such as CORBA or JAX-WS. These findings suggest Mistral Codestral can partially reduce developer effort by automating repetitive migration tasks but cannot yet replace humans within the scope of the JMigBench benchmark. The benchmark and analysis provide a foundation for future work on expanding datasets, refining prompting strategies, and improving migration performance across different LLMs.
△ Less
Submitted 10 February, 2026;
originally announced February 2026.
-
MOVA: Towards Scalable and Synchronized Video-Audio Generation
Authors:
SII-OpenMOSS Team,
:,
Donghua Yu,
Mingshu Chen,
Qi Chen,
Qi Luo,
Qianyi Wu,
Qinyuan Cheng,
Ruixiao Li,
Tianyi Liang,
Wenbo Zhang,
Wenming Tu,
Xiangyu Peng,
Yang Gao,
Yanru Huo,
Ying Zhu,
Yinze Luo,
Yiyang Zhang,
Yuerong Song,
Zhe Xu,
Zhiyu Zhang,
Chenchen Yang,
Cheng Chang,
Chushu Zhou,
Hanfu Chen
, et al. (17 additional authors not shown)
Abstract:
Audio is indispensable for real-world video, yet generation models have largely overlooked audio components. Current approaches to producing audio-visual content often rely on cascaded pipelines, which increase cost, accumulate errors, and degrade overall quality. While systems such as Veo 3 and Sora 2 emphasize the value of simultaneous generation, joint multimodal modeling introduces unique chal…
▽ More
Audio is indispensable for real-world video, yet generation models have largely overlooked audio components. Current approaches to producing audio-visual content often rely on cascaded pipelines, which increase cost, accumulate errors, and degrade overall quality. While systems such as Veo 3 and Sora 2 emphasize the value of simultaneous generation, joint multimodal modeling introduces unique challenges in architecture, data, and training. Moreover, the closed-source nature of existing systems limits progress in the field. In this work, we introduce MOVA (MOSS Video and Audio), an open-source model capable of generating high-quality, synchronized audio-visual content, including realistic lip-synced speech, environment-aware sound effects, and content-aligned music. MOVA employs a Mixture-of-Experts (MoE) architecture, with a total of 32B parameters, of which 18B are active during inference. It supports IT2VA (Image-Text to Video-Audio) generation task. By releasing the model weights and code, we aim to advance research and foster a vibrant community of creators. The released codebase features comprehensive support for efficient inference, LoRA fine-tuning, and prompt enhancement.
△ Less
Submitted 10 February, 2026; v1 submitted 9 February, 2026;
originally announced February 2026.
-
SkyReels-V3 Technique Report
Authors:
Debang Li,
Zhengcong Fei,
Tuanhui Li,
Yikun Dou,
Zheng Chen,
Jiangping Yang,
Mingyuan Fan,
Jingtao Xu,
Jiahua Wang,
Baoxuan Gu,
Mingshan Chang,
Wenjing Cai,
Yuqiang Xie,
Binjie Mao,
Youqiang Zhang,
Nuo Pang,
Hao Zhang,
Yuzhe Jin,
Zhiheng Xu,
Dixuan Lin,
Guibin Chen,
Yahui Zhou
Abstract:
Video generation serves as a cornerstone for building world models, where multimodal contextual inference stands as the defining test of capability. In this end, we present SkyReels-V3, a conditional video generation model, built upon a unified multimodal in-context learning framework with diffusion Transformers. SkyReels-V3 model supports three core generative paradigms within a single architectu…
▽ More
Video generation serves as a cornerstone for building world models, where multimodal contextual inference stands as the defining test of capability. In this end, we present SkyReels-V3, a conditional video generation model, built upon a unified multimodal in-context learning framework with diffusion Transformers. SkyReels-V3 model supports three core generative paradigms within a single architecture: reference images-to-video synthesis, video-to-video extension and audio-guided video generation. (i) reference images-to-video model is designed to produce high-fidelity videos with strong subject identity preservation, temporal coherence, and narrative consistency. To enhance reference adherence and compositional stability, we design a comprehensive data processing pipeline that leverages cross frame pairing, image editing, and semantic rewriting, effectively mitigating copy paste artifacts. During training, an image video hybrid strategy combined with multi-resolution joint optimization is employed to improve generalization and robustness across diverse scenarios. (ii) video extension model integrates spatio-temporal consistency modeling with large-scale video understanding, enabling both seamless single-shot continuation and intelligent multi-shot switching with professional cinematographic patterns. (iii) Talking avatar model supports minute-level audio-conditioned video generation by training first-and-last frame insertion patterns and reconstructing key-frame inference paradigms. On the basis of ensuring visual quality, synchronization of audio and videos has been optimized.
Extensive evaluations demonstrate that SkyReels-V3 achieves state-of-the-art or near state-of-the-art performance on key metrics including visual quality, instruction following, and specific aspect metrics, approaching leading closed-source systems. Github: https://github.com/SkyworkAI/SkyReels-V3.
△ Less
Submitted 28 January, 2026; v1 submitted 24 January, 2026;
originally announced January 2026.
-
BeamCKMDiff: Beam-Aware Channel Knowledge Map Construction via Diffusion Transformer
Authors:
Le Zhao,
Yining Wang,
Xinyi Wang,
Zesong Fei,
Yong Zeng
Abstract:
Channel knowledge map (CKM) is emerging as a critical enabler for environment-aware 6G networks, offering a site-specific database to significantly reduce pilot overhead. However, existing CKM construction methods typically rely on sparse sampling measurements and are restricted to either omnidirectional maps or discrete codebooks, hindering the exploitation of beamforming gain. To address these l…
▽ More
Channel knowledge map (CKM) is emerging as a critical enabler for environment-aware 6G networks, offering a site-specific database to significantly reduce pilot overhead. However, existing CKM construction methods typically rely on sparse sampling measurements and are restricted to either omnidirectional maps or discrete codebooks, hindering the exploitation of beamforming gain. To address these limitations, we propose BeamCKMDiff, a generative framework for constructing high-fidelity CKMs conditioned on arbitrary continuous beamforming vectors without site-specific sampling. Specifically, we incorporate a novel adaptive layer normalization (adaLN) mechanism into the noise prediction network of the Diffusion Transformer (DiT). This mechanism injects continuous beam embeddings as {global control parameters}, effectively steering the generative process to capture the complex coupling between beam patterns and environmental geometries. Simulation results demonstrate that BeamCKMDiff significantly outperforms state-of-the-art baselines, achieving superior reconstruction accuracy in capturing main lobes and side lobes.
△ Less
Submitted 15 January, 2026;
originally announced January 2026.
-
WESR: Scaling and Evaluating Word-level Event-Speech Recognition
Authors:
Chenchen Yang,
Kexin Huang,
Liwei Fan,
Qian Tu,
Botian Jiang,
Dong Zhang,
Linqi Yin,
Shimin Li,
Zhaoye Fei,
Qinyuan Cheng,
Xipeng Qiu
Abstract:
Speech conveys not only linguistic information but also rich non-verbal vocal events such as laughing and crying. While semantic transcription is well-studied, the precise localization of non-verbal events remains a critical yet under-explored challenge. Current methods suffer from insufficient task definitions with limited category coverage and ambiguous temporal granularity. They also lack stand…
▽ More
Speech conveys not only linguistic information but also rich non-verbal vocal events such as laughing and crying. While semantic transcription is well-studied, the precise localization of non-verbal events remains a critical yet under-explored challenge. Current methods suffer from insufficient task definitions with limited category coverage and ambiguous temporal granularity. They also lack standardized evaluation frameworks, hindering the development of downstream applications. To bridge this gap, we first develop a refined taxonomy of 21 vocal events, with a new categorization into discrete (standalone) versus continuous (mixed with speech) types. Based on the refined taxonomy, we introduce WESR-Bench, an expert-annotated evaluation set (900+ utterances) with a novel position-aware protocol that disentangles ASR errors from event detection, enabling precise localization measurement for both discrete and continuous events. We also build a strong baseline by constructing a 1,700+ hour corpus, and train specialized models, surpassing both open-source audio-language models and commercial APIs while preserving ASR quality. We anticipate that WESR will serve as a foundational resource for future research in modeling rich, real-world auditory scenes.
△ Less
Submitted 7 January, 2026;
originally announced January 2026.
-
Doppler-Resilient LEO Satellite OFDM Transmission with Affine Frequency Domain Pilot
Authors:
Shuntian Tang,
Xiaomei Wu,
Xinyi Wang,
Le Zhao,
Guang Yang,
Zilong Liu,
Fan Liu,
Zesong Fei
Abstract:
Orthogonal frequency division multiplexing (OFDM) based low Earth orbit (LEO) satellite communication system suffers from severe Doppler shifts, while {the Doppler-resilient affine frequency-division multiplexing (AFDM) transmission suffers from significantly high processing complexity in data detection}. In this paper, we explore the channel estimation gain of affine frequency (AF) domain pilot t…
▽ More
Orthogonal frequency division multiplexing (OFDM) based low Earth orbit (LEO) satellite communication system suffers from severe Doppler shifts, while {the Doppler-resilient affine frequency-division multiplexing (AFDM) transmission suffers from significantly high processing complexity in data detection}. In this paper, we explore the channel estimation gain of affine frequency (AF) domain pilot to enhance the OFDM transmission under high mobility. Specifically, we propose a novel AF domain pilot embedding scheme for satellite-ground downlink OFDM systems for capturing the channel characteristics. By exploiting the autoregressive (AR) property of adjacent channels, a long short-term memory (LSTM) based predictor is designed to replace conventional interpolation operation in OFDM channel estimation. Simulation results show that the proposed transmission scheme significantly outperforms conventional OFDM scheme in terms of bit error rate (BER) under high Doppler scenarios, thus paving a new way for the design of next generation non-terrestrial network (NTN) communication systems.
△ Less
Submitted 13 January, 2026; v1 submitted 5 January, 2026;
originally announced January 2026.
-
MOSS Transcribe Diarize Technical Report
Authors:
MOSI. AI,
:,
Donghua Yu,
Zhengyuan Lin,
Hanfu Chen,
Chen Yang,
Yiyang Zhang,
Jingqi Chen,
Ke Chen,
Liwei Fan,
Yi Jiang,
Jie Zhu,
Muchen Li,
Wenxuan Wang,
Yang Wang,
Zhe Xu,
Botian Jiang,
Yitian Gong,
Yuqian Zhang,
Wenbo Zhang,
Songlin Wang,
Zhiyu Wu,
Zhaoye Fei,
Qinyuan Cheng,
Shimin Li
, et al. (1 additional authors not shown)
Abstract:
Speaker-Attributed, Time-Stamped Transcription (SATS) aims to transcribe what is said and to precisely determine the timing of each speaker, which is particularly valuable for meeting transcription. Existing SATS systems rarely adopt an end-to-end formulation and are further constrained by limited context windows, weak long-range speaker memory, and the inability to output timestamps. To address t…
▽ More
Speaker-Attributed, Time-Stamped Transcription (SATS) aims to transcribe what is said and to precisely determine the timing of each speaker, which is particularly valuable for meeting transcription. Existing SATS systems rarely adopt an end-to-end formulation and are further constrained by limited context windows, weak long-range speaker memory, and the inability to output timestamps. To address these limitations, we present MOSS Transcribe Diarize, a unified multimodal large language model that jointly performs Speaker-Attributed, Time-Stamped Transcription in an end-to-end paradigm. Trained on extensive real wild data and equipped with a 128k context window for up to 90-minute inputs, MOSS Transcribe Diarize scales well and generalizes robustly. Across comprehensive evaluations, it outperforms state-of-the-art commercial systems on multiple public and in-house benchmarks.
△ Less
Submitted 16 July, 2026; v1 submitted 4 January, 2026;
originally announced January 2026.
-
DMAConv: Dual Mask-Adaptive Convolution for Remote Sensing Pansharpening
Authors:
Xianghong Xiao,
Zeyu Xia,
Zhou Fei,
Jinliang Xiao,
Haorui Chen,
Liangjian Deng
Abstract:
Pansharpening aims to fuse a high-resolution panchromatic image with a low-resolution multispectral image. Existing deep learning methods, including recent adaptive convolutions, struggle with regional heterogeneity in remote sensing images and often incur prohibitive computational costs. To address these challenges, we propose Dual Mask-Adaptive Convolution (DMAConv), a novel operator that dynami…
▽ More
Pansharpening aims to fuse a high-resolution panchromatic image with a low-resolution multispectral image. Existing deep learning methods, including recent adaptive convolutions, struggle with regional heterogeneity in remote sensing images and often incur prohibitive computational costs. To address these challenges, we propose Dual Mask-Adaptive Convolution (DMAConv), a novel operator that dynamically allocates computational resources based on feature characteristics. DMAConv first employs a lightweight module to generate soft and hard masks. The hard mask separates features into a compact branch for processing redundant information globally and a focused branch that models complex, heterogeneous regions with greater computational investment. The soft mask then preliminarily modulates the input features for both branches. This dual-branch, mask-adaptive design significantly enhances feature representation while minimizing computational overhead. Extensive experiments demonstrate that our method achieves SOTA on a broad array of quantitative benchmarks, with substantially lower parameter counts and the minimal computational cost among adaptive convolution models.
△ Less
Submitted 2 June, 2026; v1 submitted 9 December, 2025;
originally announced December 2025.
-
TokCom-UEP: Semantic Importance-Matched Unequal Error Protection for Resilient Image Transmission
Authors:
Kaizheng Zhang,
Zuolin Jin,
Zhihang Cheng,
Ming Zeng,
Li Qiao,
Zesong Fei
Abstract:
Based on the provided LaTeX code, here is the metadata for the submission form: Title: TokCom-UEP: Semantic Importance-Matched Unequal Error Protection for Resilient Image Transmission Author(s): Kaizheng Zhang, Zuolin Jin, Zhihang Cheng, Ming Zeng, Li Qiao, Zesong Fei Abstract: Token communication (TokCom), an emerging semantic communication framework powered by Large Multimodal Model (LMM), has…
▽ More
Based on the provided LaTeX code, here is the metadata for the submission form: Title: TokCom-UEP: Semantic Importance-Matched Unequal Error Protection for Resilient Image Transmission Author(s): Kaizheng Zhang, Zuolin Jin, Zhihang Cheng, Ming Zeng, Li Qiao, Zesong Fei Abstract: Token communication (TokCom), an emerging semantic communication framework powered by Large Multimodal Model (LMM), has become a key paradigm for resilient data transmission in 6G networks. A key limitation of existing TokCom designs lies in the assumption of uniform token importance, which leads to the adoption of equal error protection (EEP). However, compressed one-dimensional (1D) token sequences inherently exhibit heterogeneous semantic importance hierarchies, rendering EEP schemes suboptimal. To address this, this paper proposes TokCom-UEP, a novel semantic importance-matched unequal error protection (UEP) framework designed for resilient image transmission. TokCom-UEP integrates rateless UEP coding with the non-uniform semantic importance of tokens by partitioning source tokens into nested expanding windows, assigning higher selection probabilities to windows containing critical tokens to ensure their prioritized recovery. Simulation results demonstrate that TokCom-UEP outperforms EEP schemes in terms of three core semantic restoration metrics and spectral efficiency under low-overhead conditions.
△ Less
Submitted 27 November, 2025;
originally announced November 2025.
-
From Code Foundation Models to Agents and Applications: A Comprehensive Survey and Practical Guide to Code Intelligence
Authors:
Jian Yang,
Xianglong Liu,
Weifeng Lv,
Ken Deng,
Shawn Guo,
Lin Jing,
Yizhi Li,
Shark Liu,
Xianzhen Luo,
Yuyu Luo,
Changzai Pan,
Ensheng Shi,
Yingshui Tan,
Renshuai Tao,
Jiajun Wu,
Xianjie Wu,
Zhenhe Wu,
Daoguang Zan,
Chenchen Zhang,
Wei Zhang,
He Zhu,
Terry Yue Zhuo,
Kerui Cao,
Xianfu Cheng,
Jun Dong
, et al. (46 additional authors not shown)
Abstract:
Large language models (LLMs) have fundamentally transformed automated software development by enabling direct translation of natural language descriptions into functional code, driving commercial adoption through tools like Github Copilot (Microsoft), Cursor (Anysphere), Trae (ByteDance), and Claude Code (Anthropic). While the field has evolved dramatically from rule-based systems to Transformer-b…
▽ More
Large language models (LLMs) have fundamentally transformed automated software development by enabling direct translation of natural language descriptions into functional code, driving commercial adoption through tools like Github Copilot (Microsoft), Cursor (Anysphere), Trae (ByteDance), and Claude Code (Anthropic). While the field has evolved dramatically from rule-based systems to Transformer-based architectures, achieving performance improvements from single-digit to over 95\% success rates on benchmarks like HumanEval. In this work, we provide a comprehensive synthesis and practical guide (a series of analytic and probing experiments) about code LLMs, systematically examining the complete model life cycle from data curation to post-training through advanced prompting paradigms, code pre-training, supervised fine-tuning, reinforcement learning, and autonomous coding agents. We analyze the code capability of the general LLMs (GPT-4, Claude, LLaMA) and code-specialized LLMs (StarCoder, Code LLaMA, DeepSeek-Coder, and QwenCoder), critically examining the techniques, design decisions, and trade-offs. Further, we articulate the research-practice gap between academic research (e.g., benchmarks and tasks) and real-world deployment (e.g., software-related code tasks), including code correctness, security, contextual awareness of large codebases, and integration with development workflows, and map promising research directions to practical needs. Last, we conduct a series of experiments to provide a comprehensive analysis of code pre-training, supervised fine-tuning, and reinforcement learning, covering scaling law, framework selection, hyperparameter sensitivity, model architectures, and dataset comparisons.
△ Less
Submitted 6 December, 2025; v1 submitted 23 November, 2025;
originally announced November 2025.
-
First measurement of reactor neutrino oscillations at JUNO
Authors:
Angel Abusleme,
Thomas Adam,
Kai Adamowicz,
David Adey,
Shakeel Ahmad,
Rizwan Ahmed,
Timo Ahola,
Sebastiano Aiello,
Fengpeng An,
Guangpeng An,
Costas Andreopoulos,
Giuseppe Andronico,
João Pedro Athayde Marcondes de André,
Nikolay Anfimov,
Vito Antonelli,
Tatiana Antoshkina,
Burin Asavapibhop,
Didier Auguste,
Margherita Buizza Avanzini,
Andrej Babic,
Jingzhi Bai,
Weidong Bai,
Nikita Balashov,
Roberto Barbera,
Andrea Barresi
, et al. (1114 additional authors not shown)
Abstract:
Neutrino oscillations, a quantum effect manifesting at macroscopic scales, are governed by lepton flavor mixing angles and neutrino mass-squared differences that are fundamental parameters of particle physics, representing phenomena beyond the Standard Model. Precision measurements of these parameters are essential for testing the completeness of the three-flavor framework, determining the mass or…
▽ More
Neutrino oscillations, a quantum effect manifesting at macroscopic scales, are governed by lepton flavor mixing angles and neutrino mass-squared differences that are fundamental parameters of particle physics, representing phenomena beyond the Standard Model. Precision measurements of these parameters are essential for testing the completeness of the three-flavor framework, determining the mass ordering of neutrinos, and probing possible new physics. The Jiangmen Underground Neutrino Observatory (JUNO) is a 20 kton liquid-scintillator detector located 52.5 km from multiple reactor cores, designed to resolve the interference pattern of reactor neutrinos with sub-percent precision. Here we report, using the first 59.1 days of data collected since detector completion in August 2025, the first simultaneous high-precision determination of two neutrino oscillation parameters, $\sin^2 θ_{12} = 0.3092\,\pm\,0.0087$ and $Δm^2_{21} = (7.50\,\pm\,0.12)\times10^{-5}\;{\rm eV}^2$ for the normal mass ordering scenario, improving the precision by a factor of 1.6 relative to the combination of all previous measurements. These results advance the basic understanding of neutrinos, validate the detector's design, and confirm JUNO's readiness for its primary goal of resolving the neutrino mass ordering with a larger dataset. The rapid achievement with a short exposure highlights JUNO's potential to push the frontiers of precision neutrino physics and paves the way for its broad scientific program.
△ Less
Submitted 18 November, 2025;
originally announced November 2025.
-
Initial performance results of the JUNO detector
Authors:
Angel Abusleme,
Thomas Adam,
Kai Adamowicz,
David Adey,
Shakeel Ahmad,
Rizwan Ahmed,
Timo Ahola,
Sebastiano Aiello,
Fengpeng An,
Guangpeng An,
Costas Andreopoulos,
Giuseppe Andronico,
João Pedro Athayde Marcondes de André,
Nikolay Anfimov,
Vito Antonelli,
Tatiana Antoshkina,
Burin Asavapibhop,
Didier Auguste,
Margherita Buizza Avanzini,
Andrej Babic,
Jingzhi Bai,
Weidong Bai,
Nikita Balashov,
Roberto Barbera,
Andrea Barresi
, et al. (1114 additional authors not shown)
Abstract:
The Jiangmen Underground Neutrino Observatory (JUNO) started physics data taking on 26 August 2025. JUNO consists of a 20-kton liquid scintillator central detector, surrounded by a 35 kton water pool serving as a Cherenkov veto, and almost 1000 m$^2$ of plastic scintillator veto on top. The detector is located in a shallow underground laboratory with an overburden of 1800 m.w.e. This paper present…
▽ More
The Jiangmen Underground Neutrino Observatory (JUNO) started physics data taking on 26 August 2025. JUNO consists of a 20-kton liquid scintillator central detector, surrounded by a 35 kton water pool serving as a Cherenkov veto, and almost 1000 m$^2$ of plastic scintillator veto on top. The detector is located in a shallow underground laboratory with an overburden of 1800 m.w.e. This paper presents the performance results of the detector, extensively studied during the commissioning of the water phase, the subsequent liquid scintillator filling phase, and the first physics runs. The liquid scintillator achieved an attenuation length of 20.6 m at 430 nm, while the high coverage PMT system and scintillator together yielded about 1785 photoelectrons per MeV of energy deposit at the detector centre, measured using the 2.223 MeV $γ$ from neutron captures on hydrogen with an Am-C calibration source. The reconstructed energy resolution is 3.4% for two 0.511 MeV $γ$ at the detector centre and 2.9% for the 0.93 MeV quenched Po-214 alpha decays from natural radioactive sources. The energy nonlinearity is calibrated to better than 1%. Intrinsic contaminations of U-238 and Th-232 in the liquid scintillator are below 10$^{-16}$ g/g, assuming secular equilibrium. The water Cherenkov detector achieves a muon detection efficiency better than 99.9% for muons traversing the liquid scintillator volume. During the initial science runs, the data acquisition duty cycle exceeded 97.8%, demonstrating the excellent stability and readiness of JUNO for high-precision neutrino physics.
△ Less
Submitted 18 November, 2025;
originally announced November 2025.
-
Understanding Code Agent Behaviour: An Empirical Study of Success and Failure Trajectories
Authors:
Oorja Majgaonkar,
Zhiwei Fei,
Xiang Li,
Federica Sarro,
He Ye
Abstract:
The increasing deployment of Large Language Model (LLM) agents for complex software engineering tasks has created a need to understand their problem-solving behaviours beyond simple success metrics. While these agents demonstrate impressive capabilities in automated issue resolution, their decision-making processes remain largely opaque. This paper presents an empirical study of agent trajectories…
▽ More
The increasing deployment of Large Language Model (LLM) agents for complex software engineering tasks has created a need to understand their problem-solving behaviours beyond simple success metrics. While these agents demonstrate impressive capabilities in automated issue resolution, their decision-making processes remain largely opaque. This paper presents an empirical study of agent trajectories, namely the execution traces capturing the steps agents take when attempting to resolve software issues. We analyse trajectories from three state-of-the-art code agents (OpenHands, SWE-agent, and Prometheus) on the SWE-Bench benchmark, examining both successful and failed attempts. Our investigation reveals several key insights into agent behaviour. First, we identify how distinct problem-solving strategies, such as defensive programming and context gathering, enable success in different scenarios. Second, we find that failed trajectories are consistently longer and exhibit higher variance than successful ones, with failure patterns differing significantly between agents. Third, our fault localisation analysis shows that while most trajectories correctly identify problematic files (72-81\% even in failures), success depends more on achieving approximate rather than exact code modifications. These and other findings unveiled by our study, provide a foundation for understanding agent behaviour through trajectory analysis, contributing to the development of more robust and interpretable autonomous software engineering systems.
△ Less
Submitted 31 October, 2025;
originally announced November 2025.