-
BiCRVC: An Efficient Bidirectional Neural Video Compression Framework via Coupled Representation Coding
Authors:
Wei Jiang,
Junru Li,
Kai Zhang,
Li Zhang
Abstract:
Neural video compression (NVC) has achieved strong compression performance, but practical random-access coding still faces two technical challenges: existing bidirectional NVCs (BVCs) usually require costly motion-first decoding, and reliable motion estimation is difficult under long-range bidirectional prediction. To address these issues, we present BiCRVC, an efficient bidirectional neural video…
▽ More
Neural video compression (NVC) has achieved strong compression performance, but practical random-access coding still faces two technical challenges: existing bidirectional NVCs (BVCs) usually require costly motion-first decoding, and reliable motion estimation is difficult under long-range bidirectional prediction. To address these issues, we present BiCRVC, an efficient bidirectional neural video compression framework based on coupled representation coding. Instead of coding motion and frame information with two separate codecs, BiCRVC transforms the motion representation and the current-frame latent into a unified latent representation for entropy coding. This design enables motion and frame information to be decoded from the same bitstream with one unified codec, while still reconstructing motion-aligned contexts for frame decoding. To improve motion accuracy, we introduce multi-candidate motion estimation (MCME), which combines multi-scale motion estimation and parallel accumulated motion estimation to better handle diverse and long-range motions. To reduce motion coding overhead, we further propose bidirectional motion feature propagation (BMFP), which reuses previously decoded motion features at both the encoder and decoder as temporal priors for conditional motion coding. In addition, coupled distortion training and random GOP structure training are used to encourage joint motion-frame coding and improve adaptation to hierarchical random-access structures. Experiments show that BiCRVC achieves better compression performance than state-of-the-art BVCs while providing about 30 times faster 1080p decoding than recent BVCs.
△ Less
Submitted 17 August, 2026; v1 submitted 17 August, 2026;
originally announced August 2026.
-
Luna-TTS Family Technical Report
Authors:
Feng Yin,
Shuai Shi,
Junjie Zheng,
Kechenying Zhou,
Yiqiu Wang,
Chenyang He,
Qiuhua Jiang,
Mengxiao Bi,
Yanmin Qian,
Mingxin Chen,
Xun Gong,
Tianteng Gu,
Bing Han,
Peng Jiang,
Chenda Li,
Haiyang Sun,
Han Wang,
Wei Wang,
Yi Wang,
Leying Zhang,
Wangyou Zhang,
Chushu Zhou
Abstract:
Modern text-to-speech (TTS) is dominated by autoregressive (AR) codec language models, whose left-to-right decoding brings latency that grows with utterance length, error accumulation along the committed prefix, and an artificial generation order imposed on the Residual Vector Quantization (RVQ) token grid. We propose Luna-TTS Family, diffusion-language-model-based TTS systems pretrained on 1 mill…
▽ More
Modern text-to-speech (TTS) is dominated by autoregressive (AR) codec language models, whose left-to-right decoding brings latency that grows with utterance length, error accumulation along the committed prefix, and an artificial generation order imposed on the Residual Vector Quantization (RVQ) token grid. We propose Luna-TTS Family, diffusion-language-model-based TTS systems pretrained on 1 million hours of speech across Chinese, English, Japanese, and Korean. The family is built by progressive adaptation of a pretrained AR text LLM, from causal to bidirectional and finally to block-causal attention, and comprises two variants sharing a single tokenizer, data pipeline, and 0.6B backbone lineage. Luna-TTS is fully non-autoregressive: it generates the entire RVQ token grid in a fixed number of parallel refinement steps, with zero-shot voice cloning and speech editing arising natively as infilling. Luna-TTS Realtime, derived by continual training, is autoregressive over blocks of 32 codec frames (1.28s) while denoising each block in parallel; it supports KV-cached blockwise generation and incremental audio delivery, achieving an end-to-end RTF of 0.0240 and 41.6 ms local first-block latency under the warmed serving protocol. An annealed fine-tuning stage adds explicit control over emotion and non-verbal vocalizations (NVVs), and a reinforcement-learning stage applies GRPO with policy ratios computed over the realized denoising trajectory. On Seed-TTS-Eval, Luna-TTS achieves the best results on all four metrics among compared open-source and commercial systems (0.73 CER / 79.7 SIM on test-zh, 1.49 WER / 76.8 SIM on test-en); on the harder in-the-wild CV3-Eval, it posts the lowest Mandarin and English error rates in our comparison. Against leading commercial systems, it achieves the best results on most objective, model-based, and human-rated metrics for NVV and emotion control.
△ Less
Submitted 11 August, 2026;
originally announced August 2026.
-
PHY-Layer Modeling and Throughput-Driven Adaptation for Batteryless V2X Networks
Authors:
Zhaoyu Liu,
Ruikang Li,
Liu Cao,
YuKun Pan,
Xiangkai Wang,
Lyutianyang Zhang
Abstract:
Passive overlay communication for batteryless devices is an important enabling capability for next-generation vehicle-to-everything (V2X) networks. However, enabling reliable passive payload delivery without occupying additional spectrum remains challenging, since overlay signaling must be embedded into short and time-varying vehicular packets while preserving the decodability of the legacy host t…
▽ More
Passive overlay communication for batteryless devices is an important enabling capability for next-generation vehicle-to-everything (V2X) networks. However, enabling reliable passive payload delivery without occupying additional spectrum remains challenging, since overlay signaling must be embedded into short and time-varying vehicular packets while preserving the decodability of the legacy host transmission. This paper investigates a packetized batteryless V2X overlay architecture in which a dedicated short-range communications (DSRC)-based packet simultaneously carries conventional V2X data and a passive overlay payload. A compact PHY-layer model is developed to characterize the coupled effects of attenuation depth, embedded-bit rate, and legacy modulation and coding scheme (MCS) on host-link and passive-link reliability, as well as packet-level embedding feasibility. We then formulate a sum-throughput maximization problem that jointly accounts for the legacy packet error rate and passive decoding error rate. We further propose a multi-agent reinforcement learning (MARL)-based adaptive parameter-selection method. Simulation results show that the proposed MARL controller achieves stable convergence and improves the average throughput by 15\%, demonstrating the effectiveness of throughput-driven PHY adaptation for batteryless V2X overlay communications.
△ Less
Submitted 3 August, 2026;
originally announced August 2026.
-
Multimodal Wearable-Based Olfactory-Induced Emotion Recognition in Arousal-Valence Dimensions
Authors:
Chen-Yang Xu,
Lan Zhang,
Fei-Yi Fan,
Bin Hu,
Qing-Hao Meng
Abstract:
Olfaction is important for emotion regulation because it acts as a non-intrusive and cognitively lightweight pathway that directly engages the brain s affective circuitry and achieves unobtrusive emotional modulation. This trait is essential for advancing practical affective computing in daily and attention-critical scenarios. However, current olfactory emotion research has two key limitations. Fi…
▽ More
Olfaction is important for emotion regulation because it acts as a non-intrusive and cognitively lightweight pathway that directly engages the brain s affective circuitry and achieves unobtrusive emotional modulation. This trait is essential for advancing practical affective computing in daily and attention-critical scenarios. However, current olfactory emotion research has two key limitations. First, it overemphasises the valence dimension while neglecting arousal. Second, it lacks multimodal datasets that synchronously capture central and peripheral physiological responses to olfactory stimuli. To address these issues, we construct a large-scale multimodal olfactory emotion dataset based on 111 subjects, in which odors are labeled in the 2D arousal-valence space and electroencephalogram (EEG), electrocardiogram (ECG), and photoplethysmography (PPG) signals synchronously recorded. Nevertheless, multimodal signals present challenges such as non-stationarity, differences in latency, and cross-modal heterogeneity. Thus, we propose a spatiotemporal-frequency hybrid fusion network (STF-HFNet), which integrates three core modules. Frequency aggregation processing learns adaptive frequency aggregation in order to model non-stationary dynamics. Reciprocal guided attention enables reciprocal bidirectional calibration for cross-modal temporal alignment without synchronisation priors. Hybrid collaborative fusion combines spatial and channel attention mechanisms to enhance cross-modal complementarity while suppressing redundant information. Extensive experiments show that STF-HFNet achieves state-of-the-art (SOTA) recognition accuracies of 88.34% on the AMIGOS dataset and 92.40% on our self-constructed dataset, and outperform the SOTA methods by 8.27% and 5.07%, respectively.
△ Less
Submitted 23 July, 2026;
originally announced August 2026.
-
Data-Driven Batteryless Channel Sounding for Wi-Fi 8-Inspired Downlink MU-MIMO
Authors:
Muhan Zhang,
Chuqi Zhang,
Qitong Xu,
Zhaoyu Liu,
Liu Cao,
Lyutianyang Zhang,
Ming Gan
Abstract:
Batteryless overlays couple passive throughput to Wi-Fi sounding overhead and channel state information (CSI) aging. This paper investigates channel sounding for ultra-high reliability (UHR) operation in a Wi-Fi 8/IEEE 802.11bn-inspired downlink multi-user multiple-input multiple-output (MU-MIMO) system with a batteryless passive overlay. We optimize the post-sounding transmission interval to maxi…
▽ More
Batteryless overlays couple passive throughput to Wi-Fi sounding overhead and channel state information (CSI) aging. This paper investigates channel sounding for ultra-high reliability (UHR) operation in a Wi-Fi 8/IEEE 802.11bn-inspired downlink multi-user multiple-input multiple-output (MU-MIMO) system with a batteryless passive overlay. We optimize the post-sounding transmission interval to maximize the aggregate throughput of the active Wi-Fi and passive links, while jointly accounting for sounding overhead, CSI aging, modulation and coding scheme (MCS), passive attenuation, and passive data rate. A packet-level cross-layer model evaluates the cycle-average throughput, and a data-driven search identifies the optimal interval under different operating conditions. Simulations demonstrate that passive overlay reshapes the conventional sounding tradeoff: depending on the MCS and passive-link configuration, the additional passive throughput may or may not compensate for the associated Wi-Fi reliability loss, causing the optimal interval to shift. The results provide design guidance for reliable and low-power MU-MIMO WLANs.
△ Less
Submitted 31 July, 2026;
originally announced July 2026.
-
Advancing All-Weather Building Damage Mapping to the Instance Level: Outcomes and Insights from the 2026 Bright Challenge
Authors:
Hongruixuan Chen,
He Huang,
Haifeng Wang,
Jian Song,
Junjue Wang,
Weihao Xuan,
Hamish Mitchell,
Jiepan Li,
Wei He,
Liangpei Zhang,
Zijie Wang,
Chen Zhong,
Jiazhen Zhao,
Lei Hu,
Ting Hu,
Hongyan Zhang,
Gregory Angelides,
Miriam Cha,
Clifford Broni-Bediako,
Junshi Xia,
Taylor Perron,
Naoto Yokoya
Abstract:
Rapid post-disaster response requires timely, building-level information on whether structures remain intact, are damaged, or are destroyed. Post-event optical imagery, however, may be unavailable because of cloud, smoke, or darkness. The Bright Challenge evaluated all-weather building damage mapping from a submeter-resolution pre-event optical image and a post-event SAR image. Participants were r…
▽ More
Rapid post-disaster response requires timely, building-level information on whether structures remain intact, are damaged, or are destroyed. Post-event optical imagery, however, may be unavailable because of cloud, smoke, or darkness. The Bright Challenge evaluated all-weather building damage mapping from a submeter-resolution pre-event optical image and a post-event SAR image. Participants were required to detect and delineate each building and assign exactly one of three mutually exclusive damage labels. The challenge extended the globally distributed \textsc{Bright} dataset with instance-level annotations for about 291,000 buildings across 16 disaster events spanning seven disaster types. The final phase was evaluated exclusively on two 2025 events absent from training: a wildfire event in California and a hurricane in Jamaica. A total of 157 participants made 1,289 submissions, and 46 teams entered the final phase. The two winning solutions achieved test mAPs of 0.182 and 0.181, approximately 8.7 times the public baseline of 0.021, but remained far below the best in-domain holdout score of 0.513. Across teams ranked in both phases, performance declined sharply and the rank order changed substantially. The two leading solutions independently favored modality-specific encoding, staged or late optical--SAR fusion, and an optical-dominant separation of building localization from damage recognition. The winning method additionally used scene-aware threshold adjustment and pseudo-label adaptation. These results identify cross-event generalization and stable severity discrimination as the principal remaining challenges. All data, annotations, baseline code, and winning solutions are publicly available at https://github.com/ChenHongruixuan/BRIGHT.
△ Less
Submitted 23 July, 2026;
originally announced July 2026.
-
MedDiT4SR: Tri-Stream Joint Adaptation of Pre-Trained Diffusion Transformers for Medical Image Super-Resolution
Authors:
Zhi Chen,
Le Zhang
Abstract:
Medical image super-resolution (MedSR) requires recovering fine anatomical structures from degraded observations while avoiding unsupported details introduced by generative priors. Large-scale pre-trained multimodal diffusion transformers provide strong visual priors, but their adaptation to MedSR remains non-trivial. In conventional ControlNet-style adaptation, the low-resolution (LR) image is pr…
▽ More
Medical image super-resolution (MedSR) requires recovering fine anatomical structures from degraded observations while avoiding unsupported details introduced by generative priors. Large-scale pre-trained multimodal diffusion transformers provide strong visual priors, but their adaptation to MedSR remains non-trivial. In conventional ControlNet-style adaptation, the low-resolution (LR) image is processed as an external condition and injected into the denoising stream through one-way connections. Consequently, LR anatomical evidence cannot be jointly updated with the evolving denoising and semantic representations. We propose MedDiT4SR, a tri-stream adaptation framework that integrates the LR, noisy latent, and text representations into the same multimodal diffusion-transformer blocks. To complement global token interaction, we introduce a Super-Resolution Adapter (SR Adapter) that aggregates scale-dependent local tokens and suppresses interpolation-induced redundancy. We further propose a Semantic Alignment Refiner (SA Refiner) that calibrates local LR responses using prompt-conditioned semantic information. Experiments under both in-domain and within-modality cross-dataset settings demonstrate the effectiveness of adapting large-scale pre-trained DiT models to medical image super-resolution across diverse imaging domains.
△ Less
Submitted 22 July, 2026;
originally announced July 2026.
-
IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment
Authors:
Jinjian Wu,
Jiaqi Tang,
Wei Wei,
Yingying Yan,
Jianmin Chen,
Botong Geng,
Lei Zhang,
Qifeng Chen
Abstract:
Image Quality Assessment (IQA) in open-world environments remains challenging due to limited generalization and interpretability. Recent approaches based on multimodal large language models (MLLMs) introduce textual reasoning for quality prediction, yet their judgments rely heavily on semantically biased internal representations, making them insensitive to low-level perceptual degradations. We pro…
▽ More
Image Quality Assessment (IQA) in open-world environments remains challenging due to limited generalization and interpretability. Recent approaches based on multimodal large language models (MLLMs) introduce textual reasoning for quality prediction, yet their judgments rely heavily on semantically biased internal representations, making them insensitive to low-level perceptual degradations. We propose IQA-T1, a tool-based visual evidence reasoning framework that augments MLLM reasoning with explicit perceptual observations. During inference, the model autonomously invokes specialized analysis tools to generate structured visual evidence, such as noise residual maps, gradient statistics, and frequency spectra, which are progressively integrated into the reasoning process. To support this paradigm, we construct Q-Tool, a dataset containing 11k multimodal reasoning chains grounded in tool-generated evidence. Extensive experiments on seven IQA benchmarks show that IQA-T1 achieves the best overall performance across datasets while producing interpretable and evidence-grounded quality assessments. Code and dataset are available at https://github.com/zibuyu-02/IQA-T1.
△ Less
Submitted 14 July, 2026;
originally announced July 2026.
-
DiffRadar: Differentiable Physics-Aware Radar SLAM with Gaussian Fields
Authors:
Gaurav Bagwe,
Xiaoyong Yuan,
Yongji Wu,
Lan Zhang
Abstract:
Radar sensing is increasingly used in mobile systems because it operates reliably under poor lighting, adverse weather, and privacy-sensitive settings where cameras and LiDAR often fail. However, most existing radar SLAM systems estimate motion through scan matching on discretized radar heatmaps, which breaks geometric continuity and fails to capture key radar sensing properties, often leading to…
▽ More
Radar sensing is increasingly used in mobile systems because it operates reliably under poor lighting, adverse weather, and privacy-sensitive settings where cameras and LiDAR often fail. However, most existing radar SLAM systems estimate motion through scan matching on discretized radar heatmaps, which breaks geometric continuity and fails to capture key radar sensing properties, often leading to unstable pose estimation and degraded mapping in regenerate or dynamically changing environments. We present DiffRadar, a real-time radar SLAM system that models radar observations as a differentiable, physics-aware Gaussian field rather than discrete scans. DiffRadar represents the scene as anisotropic Gaussian primitives and renders radar measurements in range-azimuth and Doppler-azimuth spaces through a differentiable radar forward model, enabling joint optimization of robot pose and scene structure directly from radar measurements. We implement DiffRadar on commodity FMCW radar hardware and evaluate it on both the public Radarize benchmark and a controlled stress-test suite that targets common radar SLAM failure modes, including corridor degeneracy, motion regime transitions, dynamic clutter, and long-horizon loop closures. DiffRadar achieves substantial reductions in trajectory error on the benchmark, with especially large gains under feature-poor corridor motion, while more than doubling map consistency and maintaining real-time performance at 70 FPS. These results show that modeling radar observations directly in the signal domain enables substantially more robust and consistent radar-only SLAM for mobile platforms.
△ Less
Submitted 13 July, 2026;
originally announced July 2026.
-
Quantifying Realizable Flexibility Limits in Fast and Ultra-Fast EV Charging Using Real-World Data
Authors:
Cesar Diaz-Londono,
Liu Zhang,
Jorge De La Cruz,
Hamidreza Arasteh,
Anand R.,
Daogui Tang,
Josep M. Guerrero
Abstract:
The rapid growth of electric vehicles (EVs) is increasing the need to accurately quantify their flexibility as a resource for power system operation. However, most existing approaches rely on simplified or power-controllable models that overlook the intrinsic constraints of fast and ultra-fast DC charging. In practice, flexibility is fundamentally shaped by battery management system (BMS) behavior…
▽ More
The rapid growth of electric vehicles (EVs) is increasing the need to accurately quantify their flexibility as a resource for power system operation. However, most existing approaches rely on simplified or power-controllable models that overlook the intrinsic constraints of fast and ultra-fast DC charging. In practice, flexibility is fundamentally shaped by battery management system (BMS) behavior, connection time availability, and battery-protection limits. This paper introduces a trajectory-aware data-driven framework to quantify EV charging flexibility as an energy-bounded and time-constrained process. Based on 252 real charging sessions, 141 representative Power-SoC profiles are reconstructed to capture real-world charging dynamics. Unidirectional flexibility is defined through bounds on the maximum shiftable charging energy, while bidirectional flexibility is quantified as the bounds of the maximum extractable discharge energy under feasibility constraints. Results show that flexibility depends on charging state and connection time. Charging beyond 80% SoC increases duration with limited gains, while higher charger power saturates due to BMS limits. Charging time in the 20%-80% range drops by over 60%, and mean power increases by up to 40%. The maximum extractable bidirectional energy can exceed twice its value depending on the point at which flexibility is activated. These results highlight that EV flexibility is not a controllable resource, but a bounded and time-dependent capability. As such, the proposed framework provides actionable limits that can be directly used by system operators and aggregators for scheduling, peak shaving, and short-duration flexibility services.
△ Less
Submitted 29 June, 2026;
originally announced June 2026.
-
Low Complexity Kolmogorov-Arnold Network-based DPD for Analog RoF Fronthaul
Authors:
Carlos Daniel Fontes da Silva,
Tianyu Jiang,
Lu Zhang,
Vjaceslavs Bobrovs,
Xianbin Yu,
Xiaodan Pang,
Oskars Ozolins,
Edson Porto da Silva
Abstract:
This paper proposes and demonstrates experimentally for the first time a Kolmogorov-Arnold Network (KAN)-based digital predistortion (DPD) model, named envelope time-delay KAN (ETDKAN), for mitigating nonlinear distortions in analog radio-over-fiber (A-RoF) systems. The ETDKAN model incorporates physical constraints of radio-frequency (RF) nonlinear devices and, through KAN symbolization, achieves…
▽ More
This paper proposes and demonstrates experimentally for the first time a Kolmogorov-Arnold Network (KAN)-based digital predistortion (DPD) model, named envelope time-delay KAN (ETDKAN), for mitigating nonlinear distortions in analog radio-over-fiber (A-RoF) systems. The ETDKAN model incorporates physical constraints of radio-frequency (RF) nonlinear devices and, through KAN symbolization, achieves a significant reduction in computational complexity while improving interpretability. The proposed model is numerically implemented and optimized alongside multilayer perceptron (MLP) and memory-polynomial-based DPDs. Results show that the resulting symbolic ETDKAN (symbETDKAN) attains ACLR and EVM performance comparable to neural network-based models, while maintaining a computational complexity close to that of memory polynomials. Experimental validation using an A-RoF system confirms the practical feasibility of the proposed approach, which resulted in a 4-5 dB reduction in ACLR in the analyzed scenario.
△ Less
Submitted 25 June, 2026;
originally announced June 2026.
-
Parallel Dynamic Programming for Conic Linear Quadratic Control
Authors:
Luyao Zhang,
Gabriel Bravo-Palacios,
Brian Plancher,
Sergio Grammatico
Abstract:
Linear Quadratic (LQ) control problems are at the heart of linear control theory and Model Predictive Control (MPC). While performant, standard approaches to solving such problems are inherently serial, limiting real-time scalability despite the parallel computing power available on modern multi-core CPUs. Contributing to addressing this challenge and motivated by ``divide and conquer'' strategies…
▽ More
Linear Quadratic (LQ) control problems are at the heart of linear control theory and Model Predictive Control (MPC). While performant, standard approaches to solving such problems are inherently serial, limiting real-time scalability despite the parallel computing power available on modern multi-core CPUs. Contributing to addressing this challenge and motivated by ``divide and conquer'' strategies, we present a parallel-in-time approach that solves computationally demanding conic optimal control problems through the use of the alternating direction method of multipliers (ADMM). In particular, we formulate the inner primal update of ADMM as an LQ problem and split the reformulated problem along the time horizon. This enables us to derive a variant of the Riccati recursion using dynamic programming to solve each subproblem in parallel. Numerical benchmarks on two real-world applications demonstrate as much as a 5x speedup compared to existing related approaches on multi-core CPU hardware.
△ Less
Submitted 23 June, 2026;
originally announced June 2026.
-
E-MRL: Cross-view Aligned Evidence-driven Multimodal Reinforcement Learning for Reliable 3D Tumor Analysis
Authors:
Sijing Li,
Zhongwei Qiu,
Zhuoya Wang,
Boxiang Yun,
Zhenyu Yi,
Jianwei Xu,
Wenqiao Zhang,
Yingda Xia,
Ling Zhang
Abstract:
While Vision-Language Models (VLMs) show great promise in volumetric medical report generation, they frequently suffer from visual hallucinations and a lack of grounding in 3D CT data. Current Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) strategies typically optimize text fidelity alone, essentially rewarding correct diagnoses derived from language priors rather than genuine visual…
▽ More
While Vision-Language Models (VLMs) show great promise in volumetric medical report generation, they frequently suffer from visual hallucinations and a lack of grounding in 3D CT data. Current Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) strategies typically optimize text fidelity alone, essentially rewarding correct diagnoses derived from language priors rather than genuine visual perception. To address this, we propose cross-view aligned Evidence-driven Multimodal Reinforcement Learning (Evidence-MRL, noted as E-MRL), a reliable RL reasoning framework that formulates the generation process as a Markov Decision Process of "diagnosis-localization-verification". Unlike standard approaches, our model is explicitly trained to identify a "key evidence slice" alongside the global diagnostic report, grounding its findings in verifiable visual evidence. Crucially, we introduce a novel cross-view consistency reward, which validates the semantic alignment between the golden-standard report and a local visual re-query of the selected key slice, providing additional rewards for correctly-localized reasoning. Experiments on large-scale 3D CT tumor datasets demonstrate that E-MRL significantly reduces hallucinations and improves diagnostic accuracy compared to SFT and RL baselines, offering a clinically interpretable solution for visually-grounded and tumor analysis.
△ Less
Submitted 22 June, 2026;
originally announced June 2026.
-
LOLLA: Deep Reinforcement Learning for Closed-Loop Link Adaptation Towards a GPU-Accelerated AI-RAN
Authors:
Rui Wang,
Linchao Zhang,
Qiang Liu,
Kun Yang
Abstract:
Outer-loop link adaptation (OLLA) is widely deployed in 5G NR to track channel variations, yet its reliance on first-order, single-bit feedback degrades performance significantly under high-mobility and fast-varying channels. This paper presents LOLLA (Learned Outer-Loop Link Adaptation), a deep reinforcement learning framework that replaces the conventional OLLA staircase with a learned, continuo…
▽ More
Outer-loop link adaptation (OLLA) is widely deployed in 5G NR to track channel variations, yet its reliance on first-order, single-bit feedback degrades performance significantly under high-mobility and fast-varying channels. This paper presents LOLLA (Learned Outer-Loop Link Adaptation), a deep reinforcement learning framework that replaces the conventional OLLA staircase with a learned, continuous SINR offset conditioned on rich PHY/MAC telemetry inaccessible to OLLA. The offset modulates the SINR-to-MCS lookup table, preserving 3GPP-compliant MCS selection and provably subsuming the conventional OLLA update rule. A Proximal Policy Optimization (PPO) policy trained under a Lagrangian block error rate (BLER) constraint automatically enforces tunable reliability targets from 1% to 15% without manual penalty calibration. The framework is realized as the first closed-loop AI-native control dApp on a GPU-accelerated 5G NR stack, achieving end-to-end control latencies under 500 microseconds. Evaluations under 3GPP TDL channel models demonstrate 15% to 92% throughput gains over OLLA across Doppler frequencies up to 400 Hz, while attaining a Pareto frontier that strictly dominates OLLA across all evaluated reliability targets. The learned policy generalizes to unseen channel models and scales to eight concurrent UEs under shared-resource scheduling. In the uplink formulation, the gNB directly observes decoding outcomes, enabling simulation-to-deployment parity.
△ Less
Submitted 22 June, 2026;
originally announced June 2026.
-
ParaBridge: Bridging Paralinguistic Perception and Dialogue Behavior in Speech Language Models
Authors:
Yuxiang Wang,
Qinke Ni,
Shengbo Cai,
Wan Lin,
Liqiang Zhang,
Zhizheng Wu
Abstract:
Speech carries more information than just words: a child's voice, a fearful tone, or a noisy background should all lead a sufficiently competent spoken-dialogue assistant to different replies. Current Speech Language Models (SLMs) can recognize such paralinguistic cues but often ignore them in open-ended dialogue. We observe that a simple paralinguistic instruction scaffold at the inference stage…
▽ More
Speech carries more information than just words: a child's voice, a fearful tone, or a noisy background should all lead a sufficiently competent spoken-dialogue assistant to different replies. Current Speech Language Models (SLMs) can recognize such paralinguistic cues but often ignore them in open-ended dialogue. We observe that a simple paralinguistic instruction scaffold at the inference stage narrows this perception-behavior gap, suggesting that the relevant cues are already latent in the model. Such scaffolds, however, remain brittle under multi-turn context and competing instructions. Therefore, we propose \textbf{ParaBridge}, an on-policy self-distillation method that turns a brittle inference-time scaffold into stable model behavior. During training, the scaffold serves only as a temporary privileged view; the scaffold-free model rolls out its own response, while the scaffolded view supplies dense, full-vocabulary next-token targets along its trajectory. This supervision teaches when non-lexical cues should affect the reply without the need for curated dialogues, human labels, or external reward models. On Qwen3-Omni-thinking, ParaBridge raises scaffold-free VoxSafeBench SAR from $14.6\%$ to $40.3\%$ and improves EchoMind average rating from $3.27$ to $3.92$. It also preserves general ability, with MMAU-Pro, VoiceBench, and GPQA all within $0.4$ points of the original model. Beyond the training distribution, ParaBridge generalizes to unseen paralinguistic cues, transfers from safety-oriented training to empathy-oriented dialogue, and works on a different SLM backbone.
△ Less
Submitted 9 June, 2026;
originally announced June 2026.
-
A No-Regret Framework for Adaptive Incentive Design
Authors:
Georgios Vasileiou,
Lantian Zhang,
Silun Zhang
Abstract:
Incentive design studies how a central authority can influence strategic agents through payments, subsidies, or taxes, so that individual objectives align with collective welfare. This paper introduces a No-Regret Adaptive Incentive Design (RAID) framework for nonlinear games with continuous action spaces and private agent costs. In this framework, the authority (planner) designs incentives that r…
▽ More
Incentive design studies how a central authority can influence strategic agents through payments, subsidies, or taxes, so that individual objectives align with collective welfare. This paper introduces a No-Regret Adaptive Incentive Design (RAID) framework for nonlinear games with continuous action spaces and private agent costs. In this framework, the authority (planner) designs incentives that regulate the Nash equilibrium toward a socially optimal action profile, while simultaneously learning agents' unknown preferences from repeated strategic responses. We formulate the RAID problem and construct a least-squares estimator whose strong consistency requires only diminishing excitation. Leveraging this weak excitation requirement, we propose a switching incentive policy that alternates between probing (exploration) and estimate-based (exploitation) incentives. The resulting policy achieves an $O(t^{-0.5})$ parameter estimation rate and accumulates $O(t^{0.5}\log t)$ squared social-cost regret, almost surely. We further extend the framework to an endogenous-noise response model, where standard least-squares estimation is biased due to an error-in-variables correlation between the noise and agent responses. We utilize a repeated-sampling estimator and corresponding switching policy that retain the same almost-sure convergence and regret rates. Numerical experiments validate the effectiveness and predicted convergence rates of the method.
△ Less
Submitted 1 June, 2026;
originally announced June 2026.
-
Sample Complexity of Policy Gradient for Log-Growth Control
Authors:
Qiuhua Pan,
Yukai Shen,
Liwei Zhang,
Cailian Chen,
Xinping Guan
Abstract:
We study the sample complexity of policy gradient for log-growth control -- the problem of learning, from observed state transitions, a feedback gain that optimally stabilizes a scalar linear system driven through a multiplicative-noise actuation channel. The objective $J(K) = \mathbb{E}[\log|1+BK|]$ is the top Lyapunov exponent of the closed loop. This problem carries a structural difficulty we c…
▽ More
We study the sample complexity of policy gradient for log-growth control -- the problem of learning, from observed state transitions, a feedback gain that optimally stabilizes a scalar linear system driven through a multiplicative-noise actuation channel. The objective $J(K) = \mathbb{E}[\log|1+BK|]$ is the top Lyapunov exponent of the closed loop. This problem carries a structural difficulty we call the cusp obstruction: the optimal gain $K^*$ always places the noise singularity $b_{\rm sing}(K) = -1/K$ in the interior of the support. At this singular optimum the policy gradient exists only as a Cauchy principal value, not as a Lebesgue integral, and the natural single-sample gradient estimator has infinite variance. Standard first-order stochastic-optimization analysis is thus inapplicable at the optimum, and merely smoothing the objective does not resolve the difficulty. The obstruction, however, has an exploitable symmetry: the Cauchy kernel is an odd function of the displacement from the moving pole, so pairing each observation with its reflection through the pole cancels the divergent part. This one cancellation simultaneously controls the population curvature, the gradient-estimator variance, and the bias incurred when the noise density is estimated. Combining these bounds with a closed-form single-transition gradient oracle, we prove that projected mini-batch policy gradient, initialized in any compact subset of the stabilizing region, attains total sample complexity $\tilde{O}(1/η)$ when the noise density is known and $\tilde{O}(η^{-(2s+1)/(2s)})$ when it must be estimated, for $C^s$ noise densities with $s \geq 2$.
△ Less
Submitted 26 May, 2026;
originally announced May 2026.
-
Partition Tree Search Acceleration for VVC: Survey and Evaluation with VTM Evolution
Authors:
M. E. A. Kherchouche,
F. Galpin,
T. Dumas,
L. Zhang,
D. Menard
Abstract:
The Versatile Video Coding (VVC) standard, introduced in 2020, offers 40-50% bitrate savings for equivalent visual quality of reconstructed videos over its predecessor, High Efficiency Video Coding (HEVC), at the cost of significantly increased encoding complexity. This growth in encoding complexity is mainly due to the addition of the Quad Tree Multi Type Tree (QTMTT) partitioning structure, whic…
▽ More
The Versatile Video Coding (VVC) standard, introduced in 2020, offers 40-50% bitrate savings for equivalent visual quality of reconstructed videos over its predecessor, High Efficiency Video Coding (HEVC), at the cost of significantly increased encoding complexity. This growth in encoding complexity is mainly due to the addition of the Quad Tree Multi Type Tree (QTMTT) partitioning structure, which increases the split combinatorial complexity. This paper presents a critical evaluation of state-of-the-art (SOTA) partitioning acceleration techniques designed to reduce the complexity of the partitioning search in VVC. Particular attention is given to how these methods have evolved alongside successive versions of the VVC Test Model (VTM), which serves as the reference software for benchmarking coding tools. These techniques are analyzed in the context of their adaptation to internal changes in VTM, such as updated heuristics for fast partitioning decisions. The study also highlights the challenges involved in improving the trade-off between encoding complexity and compression efficiency. This challenge becomes more pronounced when evaluating methods across diverse VTM configurations and multiple software versions.
△ Less
Submitted 19 May, 2026;
originally announced May 2026.
-
Adversarial Stress Testing of SPARK Humanoid Safety Filters
Authors:
Saurav Ghosh,
Abdou Sow,
Luke Zhang
Abstract:
Humanoid robots are difficult to deploy safely because they have high-dimensional bodies, many collision constraints, and must operate near people and obstacles. Safety filters help by modifying a nominal control action when it may violate collision-avoidance constraints. Still, nominal benchmark scores do not fully show how these filters behave in harder environments. In this work, we study the r…
▽ More
Humanoid robots are difficult to deploy safely because they have high-dimensional bodies, many collision constraints, and must operate near people and obstacles. Safety filters help by modifying a nominal control action when it may violate collision-avoidance constraints. Still, nominal benchmark scores do not fully show how these filters behave in harder environments. In this work, we study the robustness of SPARK humanoid safety filters through replication and stress testing. We replicate the SPARK benchmark case G1SportMode_D1_WG_SO_v1 in MuJoCo and evaluate RSSA, RSSS, SSA, CBF, PFM, and SMA under controlled random seeds. We also built a post-processing pipeline that converts raw SPARK logs into goal-tracking, minimum-distance, and collision-step metrics. Our results show that some methods track the goal more closely, while others reduce collision steps more effectively. The stress tests further indicate that safety behavior can change under obstacle crowding, noisy distance estimates, and delayed obstacle information. These findings suggest that humanoid autonomy should be evaluated beyond nominal performance, using metrics that expose failure modes before deployment.
△ Less
Submitted 18 May, 2026;
originally announced May 2026.
-
DepthPolyp: Pseudo-Depth Guided Lightweight Segmentation for Real-Time Colonoscopy
Authors:
Zhuoyu Wu,
Wenhui Ou,
Lexi Zhang,
Pei-Sze Tan,
Dongjun Wu,
Junhe Zhao,
Wenqi Fang,
Raphaël C. -W. Phan
Abstract:
Accurate polyp segmentation in colonoscopy is essential for early colorectal cancer detection, yet real-world clinical environments pose persistent challenges such as motion blur, specular reflections, and illumination instability. Most existing methods are optimized on clean benchmark images and suffer noticeable performance degradation when deployed in authentic surgical scenarios. We propose De…
▽ More
Accurate polyp segmentation in colonoscopy is essential for early colorectal cancer detection, yet real-world clinical environments pose persistent challenges such as motion blur, specular reflections, and illumination instability. Most existing methods are optimized on clean benchmark images and suffer noticeable performance degradation when deployed in authentic surgical scenarios. We propose DepthPolyp, a lightweight and robust segmentation framework based on pseudo-depth-guided multi-task learning and efficient feature modulation. The architecture combines hierarchical Ghost factorization for compact feature generation, Interleaved Shuffle Fusion for low-cost cross-scale interaction, and Dynamic Group Gating for adaptive group-wise feature weighting. Extensive experiments demonstrate that DepthPolyp achieves strong cross-dataset generalization when trained on degraded data and evaluated on both clean and noisy target domains, consistently outperforming lightweight baselines and remaining competitive with substantially larger models. In real surgical video evaluation on PolypGen, DepthPolyp achieves better segmentation performance than models up to $20\times$ larger while preserving real-time inference speed. With only 3.57M parameters and 0.86 GMACs, the proposed method runs at over 180 FPS on mobile devices, making it well suited for real-time deployment in resource-constrained clinical environments. Code and pretrained weights are available at: https://github.com/ReaganWu/DepthPolyp/
△ Less
Submitted 15 May, 2026;
originally announced May 2026.
-
Toward World Modeling of Physiological Signals with Chaos-Theoretic Balancing and Latent Dynamics
Authors:
Yunfei Luo,
Xi Chen,
Yuliang Chen,
Lanshuang Zhang,
Md Mofijul Islam,
Siwei Zhao,
Peter Kotanko,
Subhasis Dasgupta,
Andrew Campbell,
Rakesh Malhotra,
Tauhidur Rahman
Abstract:
Physiological time series signals reflect complex, multi-scale dynamical processes of the human body. Existing modeling studies focus on static tasks such as classification, event forecasting, or short-horizon next step prediction, while long-horizon signal-level forecasting and predictive nature of physiological signals remain underexplored. We introduce NormWear-2, a world model that encodes bot…
▽ More
Physiological time series signals reflect complex, multi-scale dynamical processes of the human body. Existing modeling studies focus on static tasks such as classification, event forecasting, or short-horizon next step prediction, while long-horizon signal-level forecasting and predictive nature of physiological signals remain underexplored. We introduce NormWear-2, a world model that encodes both multivariate physiological signals and clinical intervention variables into a shared latent space and models their joint temporal evolution as a dynamical system. Our approach combines inference from prior pre-trained knowledge (intuition) with instant non-parametric latent state transition adaptation (insight), enabling coherent forecasting across multiple temporal scales, conditioned on heterogeneous clinical interventions. During the pretraining phase, we find that chaos-theoretic balancing of dynamical regime diversity yields more robust representations, with a smaller balanced corpus outperforming one twice its size and capturing bifurcation regimes. We evaluate the world model performance across diverse real-world physiological datasets spanning heterogeneous temporal resolutions and intervention regimes, covering daily life, point-of-care, and clinical settings, including fitness planning, hemodialysis, diabetes management, and surgical monitoring. These evaluation datasets comprise records from 8,026 subjects, spanning study durations from 3.2 hours for high-resolution signal data to 2.3 years for longitudinal clinical biomarker tracking. NormWear-2 achieves the best overall forecasting performance across time, frequency, and latent representation domains, with significant improvements over state-of-the-art time series foundation models, while maintaining competitive downstream representation quality, providing a step toward general-purpose world models for physiological signals.
△ Less
Submitted 14 May, 2026;
originally announced May 2026.
-
Revisiting Voltage and Synchronization Stability Analysis in Grid-Following Converter-Integrated Weak Grids: Insights from Non-Minimum-Phase Zeros
Authors:
Fuyilong Ma,
Lidong Zhang,
Wangqianyun Tang,
Waisheng Zheng,
Huanhai Xin,
Linbin Huang,
Lennart Harnefors
Abstract:
The increasing penetration of grid-following (GFL) converter-interfaced generators (CIGs) intensifies concerns over small-signal voltage and synchronization stability. While existing theories treat these two stability issues distinctly, practical wisdom in contrast employs a unified and static metric, short-circuit ratio (SCR), to assess both in weak grids. This paper aims to bridge this theory-pr…
▽ More
The increasing penetration of grid-following (GFL) converter-interfaced generators (CIGs) intensifies concerns over small-signal voltage and synchronization stability. While existing theories treat these two stability issues distinctly, practical wisdom in contrast employs a unified and static metric, short-circuit ratio (SCR), to assess both in weak grids. This paper aims to bridge this theory-practice gap by introducing the insight of non-minimum-phase (NMP) zeros. First, we demonstrate that the two stability issues in weak grids can be characterized within a common NMP-zero-based assessment framework: a zero at the origin corresponds to voltage instability, while low-frequency zeros impose fundamental constraints on synchronization dynamics. The traditional SCR is proven to be a special case of our proposed novel stability metric, NMP-zero (NMP-Z) factor, evaluated at the rated operating point. This establishes the theoretical foundation for the empirical success of SCR. Building on this insight, we then develop a unified stability assessment method for multi-converter systems. The method retains the simplicity of SCR, requiring only the NMP-Z factor together with individual CIG dynamic models and enabling stability margin assessment under various operating points. Our work provides a simple yet theoretically rigorous framework for stability analysis in CIG-integrated weak grids, with all theoretical findings and the proposed method validated through detailed time-domain simulations.
△ Less
Submitted 3 August, 2026; v1 submitted 13 May, 2026;
originally announced May 2026.
-
JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions
Authors:
Leying Zhang,
Bowen Shi,
Haibin Wu,
Bach Viet Do,
Yanmin Qian
Abstract:
The rapid advancement of generative audio models has outpaced the development of robust evaluation methodologies. Existing objective metrics and general multimodal large language models (MLLMs) often struggle with domain generalization, zero-shot capabilities, and instructional flexibility. To address these bottlenecks, we propose JASTIN, a generalizable, instruction-driven audio evaluation framew…
▽ More
The rapid advancement of generative audio models has outpaced the development of robust evaluation methodologies. Existing objective metrics and general multimodal large language models (MLLMs) often struggle with domain generalization, zero-shot capabilities, and instructional flexibility. To address these bottlenecks, we propose JASTIN, a generalizable, instruction-driven audio evaluation framework that formulates audio assessment as a self-instructed reasoning task. JASTIN bridges a frozen high-performance audio encoder with a fine-tuned LLM backbone via a trainable audio adapter. To ensure robust zero-shot generalization, we introduce a comprehensive instruction following data preparation pipeline, incorporating Multi-Source, Multi-Task, Multi-Calibration, and Multi-Description data. Experimental results demonstrate that JASTIN achieves state-of-the-art Pearson and Spearman correlations with human subjective ratings. It consistently outperforms general MLLMs across speech, sound, music, and out-of-domain evaluation tasks without the need for task-specific retraining.
△ Less
Submitted 6 May, 2026;
originally announced May 2026.
-
CMTA: Leveraging Cross-Modal Temporal Artifacts for Generalizable AI-Generated Video Detection
Authors:
Hang Wang,
Chao Shen,
Chenhao Lin,
Minghui Yang,
Lei Zhang,
Cong Wang
Abstract:
The proliferation of advanced AI video synthesis techniques poses an unprecedented challenge to digital video authenticity. Existing AI-generated video (AIGV) detection methods primarily focus on uni-modal or spatiotemporal artifacts, but they overlook the rich cues within the visual-textual cross-modal space, especially the temporal stability of semantic alignment. In this work, we identify a dis…
▽ More
The proliferation of advanced AI video synthesis techniques poses an unprecedented challenge to digital video authenticity. Existing AI-generated video (AIGV) detection methods primarily focus on uni-modal or spatiotemporal artifacts, but they overlook the rich cues within the visual-textual cross-modal space, especially the temporal stability of semantic alignment. In this work, we identify a distinctive fingerprint in AIGVs, termed cross-modal temporal artifact (CMTA). Unlike real videos that exhibit natural temporal fluctuations in cross-modal alignment due to semantic variations, AIGVs display unnaturally stable semantic trajectories governed by given input prompts. To bridge this gap, we propose the CMTA framework, a cross-modal detection approach that captures these unique temporal artifacts through joint cross-modal embedding and multi-grained temporal modeling. Specifically, CMTA leverages BLIP to generate frame-level image captions and utilizes CLIP to extract corresponding visual-textual representations. A coarse-grained temporal modeling branch is then designed to characterize temporal fluctuations in cross-modal alignment with a GRU. In parallel, a fine-grained branch is constructed to capture intricate inter-frame variations from integrated visual-textual features with a Transformer encoder. Extensive experiments on 40 subsets across four large-scale datasets, including GenVideo, EvalCrafter, VideoPhy, and VidProM, validate that our approach sets a new state-of-the-art while exhibiting superior cross-generator generalization. Code and models of CMTA will be released at https://github.com/hwang-cs-ime/CMTA
△ Less
Submitted 1 May, 2026;
originally announced May 2026.
-
EDU-Net: Retinal Pathological Fluid Segmentation in OCT Images with Multiscale Feature Fusion and Boundary Optimization
Authors:
Zijun Lei,
Zikang Xu,
Liang Zhang,
Ge Song,
Hanyu Guo,
Dan Cao,
Yujia Zhou,
Qianjin Feng
Abstract:
Objective: Diabetic macular edema (DME) is the leading cause of severe visual impairment in patients with diabetes. Quantification of retinal fluid, particularly intraretinal fluid (IRF) and subretinal fluid (SRF), plays a critical role in the management of DME. Although optical coherence tomography (OCT) can be used for detection, the variable morphology of fluid accumulation and the blurred boun…
▽ More
Objective: Diabetic macular edema (DME) is the leading cause of severe visual impairment in patients with diabetes. Quantification of retinal fluid, particularly intraretinal fluid (IRF) and subretinal fluid (SRF), plays a critical role in the management of DME. Although optical coherence tomography (OCT) can be used for detection, the variable morphology of fluid accumulation and the blurred boundaries caused by noise interference still limit the accuracy of OCT's automatic segmentation. Methods: Retrospective model development and validation study. This study proposes a novel edge-guided dual-branch encoder-decoder network (EDU-Net) to achieve accurate and efficient automatic segmentation of OCT liquid lesions. The local feature extraction branch is based on the EfficientNet model, which precisely captures tiny lesions by leveraging its lightweight separable convolution and high-resolution feature preservation strategy. The global feature extraction branch is based on the large-kernel efficient convolution (LKEC) module and the downsampling layer design to enhance long-range dependencies and global semantics. EDU-Net applies a multi-category edge-guided attention module to fuse high-frequency boundary detail information to each resolution feature to optimize the boundary segmentation performance. Results: Extensive results on the in-house and public datasets demonstrate that EDU-Net achieves state-of-the-art DSC segmentation performance in terms of efficiency and robustness, especially in the segmentation of IRF lesions. Conclusions: EDU-Net integrates local details with global context and optimizes boundaries, achieving an improvement in the accuracy of automatic segmentation of retinal fluid.
△ Less
Submitted 21 April, 2026;
originally announced April 2026.
-
Inertia Matching Principle: Improving Transient Synchronization Stability in Hybrid Power Systems With VSGs and SGs
Authors:
Changjun He,
Li Zhang,
Qi Liu,
Rui Zou
Abstract:
This paper investigates the transient synchronization stability in power systems hybridized with virtual synchronous generators (VSGs) and synchronous generators (SGs). A relative swing equation model is established to capture the transient synchronization dynamics between the VSG and the SG. Based on this model, both static and dynamic characteristics are systematically analyzed, and a quantitati…
▽ More
This paper investigates the transient synchronization stability in power systems hybridized with virtual synchronous generators (VSGs) and synchronous generators (SGs). A relative swing equation model is established to capture the transient synchronization dynamics between the VSG and the SG. Based on this model, both static and dynamic characteristics are systematically analyzed, and a quantitative stability level index is derived to elucidate the underlying stability mechanism. Then, two fundamental inertia matching principles are identified. First, a new instability mechanism induced by improper inertia matching between the VSG and the SG is revealed. It is identified that increasing the VSG's inertia does not monotonically improve transient stability, as commonly presumed. Instead, an optimal inertia matching constant exists that maximizes stability performance. Second, the influence of the VSG share on the synchronization stability is discovered to be strongly influenced by the matching between the VSG's inertia level and its voltage strength (i.e., output impedance). To achieve reliable and robust synchronization stability, proper coordination between the VSG's inertia and virtual impedance is essential. Finally, a coordinated stabilization strategy based on inertia matching and virtual impedance adjustment is proposed to enhance transient synchronization stability performance while suppressing fault current. Simulations conducted on a two-machine system and the IEEE 39-bus system validate the theoretical findings and demonstrate the effectiveness of the proposed strategy.
△ Less
Submitted 20 April, 2026;
originally announced April 2026.
-
Incentive Design without Hypergradients: A Social-Gradient Method
Authors:
Georgios Vasileiou,
Lantian Zhang,
Silun Zhang
Abstract:
Incentive design problems consider a system planner who steers self-interested agents toward a socially optimal Nash equilibrium by issuing incentives in the presence of information asymmetry, that is, uncertainty about the agents' cost functions. A common approach formulates the problem as a Mathematical Program with Equilibrium Constraints (MPEC) and optimizes incentives using hypergradients-the…
▽ More
Incentive design problems consider a system planner who steers self-interested agents toward a socially optimal Nash equilibrium by issuing incentives in the presence of information asymmetry, that is, uncertainty about the agents' cost functions. A common approach formulates the problem as a Mathematical Program with Equilibrium Constraints (MPEC) and optimizes incentives using hypergradients-the total derivatives of the planner's objective with respect to incentives. However, computing or approximating the hypergradients typically requires full or partial knowledge of equilibrium sensitivities to incentives, which is generally unavailable under information asymmetry. In this paper, we propose a hypergradient-free incentive law, called the social-gradient flow, for incentive design when the planner's social cost depends on the agents' joint actions. We prove that the social cost gradient is always a descent direction for the planner's objective, irrespective of the agent cost landscape. In the idealized setting where equilibrium responses are observable, the social-gradient flow converges to the unique socially optimal incentive. When equilibria are not directly observable, the social-gradient flow emerges as the slow-timescale limit of a two-timescale interaction, in which agents' strategies evolve on a faster timescale. It is established that the joint strategy-incentive dynamics converge to the social optimum for any agent learning rule that asymptotically tracks the equilibrium. Theoretical results are also validated via numerical experiments.
△ Less
Submitted 13 April, 2026;
originally announced April 2026.
-
Unsupervised Equivalent Contrastive Learning for Radio Signal Recognition
Authors:
Shilian Zheng,
Jie Chen,
Luxin Zhang,
Xiaoniu Yang
Abstract:
Robust radio signal recognition is fundamental to spectrum management, electromagnetic space security, and intelligent wireless applications, yet existing deep-learning methods rely heavily on large labeled datasets and struggle to capture the multi-domain characteristics inherent in real-world signals. To address these limitations, we propose an unsupervised equivalent contrastive learning method…
▽ More
Robust radio signal recognition is fundamental to spectrum management, electromagnetic space security, and intelligent wireless applications, yet existing deep-learning methods rely heavily on large labeled datasets and struggle to capture the multi-domain characteristics inherent in real-world signals. To address these limitations, we propose an unsupervised equivalent contrastive learning method that leverages four information-lossless equivalent transformations, spanning the time, instantaneous, frequency, and time-frequency domains, to construct multi-view and semantically consistent representations of each signal. An equivalent contrastive learning strategy then aligns these complementary views to learn discriminative and transferable embeddings without requiring labeled data. Once pre-training is completed, the resulting model can be directly fine-tuned on downstream tasks using only raw signal samples, without reapplying any equivalent transformations, which reduces computational overhead and simplifies deployment. Extensive experiments on four public datasets demonstrate that the proposed method consistently outperforms state-of-the-art contrastive baselines under linear evaluation, few-shot semi-supervised learning, and cross-domain transfer settings. Notably, the learned representations yield substantial gains in few-shot regimes and challenging channel conditions, confirming the effectiveness of multi-domain equivalent modeling in enhancing robustness and generalization. This work establishes a principled pathway for exploiting massive unlabeled radio data and provides a foundation for future self-supervised learning frameworks in wireless systems.
△ Less
Submitted 12 April, 2026;
originally announced April 2026.
-
On the Optimization Landscape of Observer-based Dynamic Linear Quadratic Control
Authors:
Jingliang Duan,
Jie Li,
Yinsong Ma,
Liye Tang,
Guofa Li,
Liping Zhang,
Shengbo Eben Li,
Lin Zhao
Abstract:
Understanding the optimization landscape of linear quadratic regulation (LQR) problems is fundamental to the design of efficient reinforcement learning solutions. Recent work has made significant progress in characterizing the landscape of static output-feedback control and linear quadratic Gaussian (LQG) control. For LQG, much of the analysis leverages the separation principle, which allows the c…
▽ More
Understanding the optimization landscape of linear quadratic regulation (LQR) problems is fundamental to the design of efficient reinforcement learning solutions. Recent work has made significant progress in characterizing the landscape of static output-feedback control and linear quadratic Gaussian (LQG) control. For LQG, much of the analysis leverages the separation principle, which allows the controller and estimator to be designed independently. However, this simplification breaks down when the gradients with respect to the estimator and controller parameters are inherently coupled, leading to a more intricate analysis. This paper investigates the optimization landscape of observer-based dynamic output-feedback control of LQR problems. We derive the optimal observer-controller pair in settings where transient quadratic performance cannot be neglected. Our analysis reveals that, in general, the combination of the standard LQR controller and the observer that minimizes the trace of the accumulated estimation error covariance does not correspond to a stationary point of the overall closed-loop performance objective. Moreover, we derive a pair of discrete-time Sylvester equations with symmetric structure, both involving the same set of matrix elements, that characterize the stationary point of the observer-based dynamic LQR problem. These equations offer analytical insight into the structure of the optimality conditions and provide a foundation for developing numerical policy gradient methods aimed at learning complex controllers that rely on reconstructed state information.
△ Less
Submitted 12 April, 2026;
originally announced April 2026.
-
Stochastic Adaptive Control for Systems with Nonlinear Parameterization: Almost Sure Stability and Tracking
Authors:
Lantian Zhang,
Bo Wahlberg,
Silun Zhang
Abstract:
This paper concerns the adaptive control problem for a class of nonlinear stochastic systems in which the state update is given by a nonlinear function of linear dynamics plus additive stochastic noise. Such systems arise in a wide range of applications, including recurrent neural networks, social dynamics, and signal processing. Despite their importance, adaptive control for these systems remains…
▽ More
This paper concerns the adaptive control problem for a class of nonlinear stochastic systems in which the state update is given by a nonlinear function of linear dynamics plus additive stochastic noise. Such systems arise in a wide range of applications, including recurrent neural networks, social dynamics, and signal processing. Despite their importance, adaptive control for these systems remains relatively unexplored in the literature. This gap is primarily due to the inherently nonconvex dependence of the system dynamics on unknown parameters, which significantly complicates both controller design and analysis. To address these challenges, we propose an online nonlinear weighted least-squares (WLS)-based parameter estimation algorithm and establish the global strong consistency of the resulting parameter estimates. In contrast to most existing results, our consistency analysis does not rely on restrictive assumptions such as persistent excitation conditions of the trajectory data, making it applicable to stochastic adaptive control settings. Building on the proposed estimator, we further develop an adaptive control algorithm with an attenuating excitation signal that can effectively combine adaptive estimation and feedback control. Finally, we are able to show that the resulting closed-loop system is globally stable and that the system trajectory can track, in a long-run average sense, the reference trajectory generated with the true system parameters. The proposed methods and theoretical results are finally validated through simulations in two nonlinear interaction network applications.
△ Less
Submitted 8 April, 2026;
originally announced April 2026.
-
Adaptive Incentive Design with Regret Minimization
Authors:
Georgios Vasileiou,
Lantian Zhang,
Silun Zhang
Abstract:
Incentive design constitutes a foundational paradigm for influencing the behavior of strategic agents, wherein a system planner (principal) publicly commits to an incentive mechanism designed to align individual objectives with collective social welfare. This paper introduces the Regret-Minimizing Adaptive Incentive Design (RAID) problem, which aims to synthesize incentive laws under information a…
▽ More
Incentive design constitutes a foundational paradigm for influencing the behavior of strategic agents, wherein a system planner (principal) publicly commits to an incentive mechanism designed to align individual objectives with collective social welfare. This paper introduces the Regret-Minimizing Adaptive Incentive Design (RAID) problem, which aims to synthesize incentive laws under information asymmetry and achieve asymptotically minimal regret compared to an oracle with full information. To this end, we develop the RAID algorithm, which employs a switching policy alternating between probing (exploration) and estimate-based incentivization (exploitation). The associated type estimator relies only on a weaker excitation condition required for strong consistency in least squares estimation, substantially relaxing the persistence-of-excitation assumptions previously used in adaptive incentive design. In addition, we establish the strong consistency of the proposed type estimator and prove that the incentive obtained asymptotically minimizes the planner's average regret almost surely. Numerical experiments illustrate the convergence rate of the proposed methodology.
△ Less
Submitted 7 April, 2026;
originally announced April 2026.
-
BAAI Cardiac Agent: An intelligent multimodal agent for automated reasoning and diagnosis of cardiovascular diseases from cardiac magnetic resonance imaging
Authors:
Taiping Qu,
Hongkai Zhang,
Lantian Zhang,
Can Zhao,
Nan Zhang,
Hui Wang,
Zhen Zhou,
Mingye Zou,
Kairui Bo,
Pengfei Zhao,
Xingxing Jin,
Zixian Su,
Kun Jiang,
Huan Liu,
Yu Du,
Maozhou Wang,
Ruifang Yan,
Zhongyuan Wang,
Tiejun Huang,
Lei Xu,
Henggui Zhang
Abstract:
Cardiac magnetic resonance (CMR) is a cornerstone for diagnosing cardiovascular disease. However, it remains underutilized due to complex, time-consuming interpretation across multi-sequences, phases, quantitative measures that heavily reliant on specialized expertise. Here, we present BAAI Cardiac Agent, a multimodal intelligent system designed for end-to-end CMR interpretation. The agent integra…
▽ More
Cardiac magnetic resonance (CMR) is a cornerstone for diagnosing cardiovascular disease. However, it remains underutilized due to complex, time-consuming interpretation across multi-sequences, phases, quantitative measures that heavily reliant on specialized expertise. Here, we present BAAI Cardiac Agent, a multimodal intelligent system designed for end-to-end CMR interpretation. The agent integrates specialized cardiac expert models to perform automated segmentation of cardiac structures, functional quantification, tissue characterization and disease diagnosis, and generates structured clinical reports within a unified workflow. Evaluated on CMR datasets from two hospitals (2413 patients) spanning 7-types of major cardiovascular diseases, the agent achieved an area under the receiver-operating-characteristic curve exceeding 0.93 internally and 0.81 externally. In the task of estimating left ventricular function indices, the results generated by this system for core parameters such as ejection fraction, stroke volume, and left ventricular mass are highly consistent with clinical reports, with Pearson correlation coefficients all exceeding 0.90. The agent outperformed state-of-the-art models in segmentation and diagnostic tasks, and generated clinical reports showing high concordance with expert radiologists (six readers across three experience levels). By dynamically orchestrating expert models for coordinated multimodal analysis, this agent framework enables accurate, efficient CMR interpretation and highlights its potentials for complex clinical imaging workflows. Code is available at https://github.com/plantain-herb/Cardiac-Agent.
△ Less
Submitted 5 April, 2026;
originally announced April 2026.
-
Boundary-aware Prototype-driven Adversarial Alignment for Cross-Corpus EEG Emotion Recognition
Authors:
Guangli Li,
Canbiao Wu,
Na Tian,
Li Zhang,
Zhen Liang
Abstract:
Electroencephalography (EEG)-based emotion recognition suffers from severe performance degradation when models are transferred across heterogeneous datasets due to physiological variability, experimental paradigm differences, and device inconsistencies. Existing domain adversarial methods primarily enforce global marginal alignment and often overlook class-conditional mismatch and decision boundar…
▽ More
Electroencephalography (EEG)-based emotion recognition suffers from severe performance degradation when models are transferred across heterogeneous datasets due to physiological variability, experimental paradigm differences, and device inconsistencies. Existing domain adversarial methods primarily enforce global marginal alignment and often overlook class-conditional mismatch and decision boundary distortion, limiting cross-corpus generalization. In this work, we propose a unified Prototype-driven Adversarial Alignment (PAA) framework for cross-corpus EEG emotion recognition. The framework is progressively instantiated in three configurations: PAA-L, which performs prototype-guided local class-conditional alignment; PAA-C, which further incorporates contrastive semantic regularization to enhance intra-class compactness and inter-class separability; and PAA-M, the full boundary-aware configuration that integrates dual relation-aware classifiers within a three-stage adversarial optimization scheme to explicitly refine controversial samples near decision boundaries. By combining prototype-guided subdomain alignment, contrastive discriminative enhancement, and boundary-aware aggregation within a coherent adversarial architecture, the proposed framework reformulates emotion recognition as a relation-driven representation learning problem, reducing sensitivity to label noise and improving cross-domain stability. Extensive experiments on SEED, SEED-IV, and SEED-V demonstrate state-of-the-art performance under four cross-corpus evaluation protocols, with average improvements of 6.72\%, 5.59\%, 6.69\%, and 4.83\%, respectively. Furthermore, the proposed framework generalizes effectively to clinical depression identification scenarios, validating its robustness in real-world heterogeneous settings. The source code is available at \textit{https://github.com/WuCB-BCI/PAA}
△ Less
Submitted 17 March, 2026;
originally announced March 2026.
-
CL-SEC: Cross-Layer Semantic Error Correction Empowered by Language Models
Authors:
Yirun Wang,
Yuyang Du,
Soung Chang Liew,
Yuchen Pan,
Feifan Zhang,
Lihao Zhang
Abstract:
Achieving reliable communication has long been a fundamental challenge in networked systems. Semantic Error Correction (SEC) leverages the semantic understanding capabilities of language models (LMs) to perform application-layer error correction, complementing conventional channel decoding. While promising, existing SEC approaches rely solely on context captured by LMs at the application layer, ig…
▽ More
Achieving reliable communication has long been a fundamental challenge in networked systems. Semantic Error Correction (SEC) leverages the semantic understanding capabilities of language models (LMs) to perform application-layer error correction, complementing conventional channel decoding. While promising, existing SEC approaches rely solely on context captured by LMs at the application layer, ignoring the rich information available at the physical layer. To address this limitation, this paper introduces Cross-Layer SEC (CL-SEC), an LM-empowered error correction framework that integrates cross-layer information from both the physical and application layers to jointly correct corrupted words in text communication. Using a Bayesian combination in product form tailored to this framework, CL-SEC achieves significantly improved performance over methods that process information in isolated layers. CL-SEC shows substantial gains across multiple error-correction metrics, including bit-error rate, word-error rate, and semantic fidelity scores. Importantly, unlike most semantic communication systems that focus solely on recovering the semantic meaning of transmitted messages, CL-SEC aims to reconstruct the original transmitted message verbatim, leveraging the semantic understanding capabilities of LMs for precise reconstruction.
△ Less
Submitted 27 March, 2026;
originally announced March 2026.
-
Agentic AI for SAGIN Resource Management_Semantic Awareness, Orchestration, and Optimization
Authors:
Linghao Zhang,
Haitao Zhao,
Bo Xu,
Hongbo Zhu,
Xianbin Wang
Abstract:
Space-air-ground integrated networks (SAGIN) promise ubiquitous 6G connectivity but face significant resource management challenges due to heterogeneous infrastructure, dynamic topologies, and stringent quality-of-service (QoS) requirements. Conventional model-driven approaches struggle with scalability and adaptability in such complex environments. This paper presents an agentic artificial intell…
▽ More
Space-air-ground integrated networks (SAGIN) promise ubiquitous 6G connectivity but face significant resource management challenges due to heterogeneous infrastructure, dynamic topologies, and stringent quality-of-service (QoS) requirements. Conventional model-driven approaches struggle with scalability and adaptability in such complex environments. This paper presents an agentic artificial intelligence (AI) framework for autonomous SAGIN resource management by embedding large language model (LLM)-based agents into a Monitor-Analyze-Plan- Execute-Knowledge (MAPE-K) control plane. The framework incorporates three specialized agents, namely semantic resource perceivers, intent-driven orchestrators, and adaptive learners, that collaborate through natural language reasoning to bridge the gap between operator intents and network execution. A key innovation is the hierarchical agent-reinforcement learning (RL) collaboration mechanism, wherein LLM-based orchestrators dynamically shape reward functions for RL agents based on semantic network conditions. Validation through UAV-assisted AIGC service orchestration in energy-constrained scenarios demonstrates that LLM-driven reward shaping achieves 14% energy reduction and the lowest average service latency among all compared methods. This agentic paradigm offers a scalable pathway toward adaptive, AI-native 6G networks, capable of autonomously interpreting intents and adapting to dynamic environments.
△ Less
Submitted 17 March, 2026;
originally announced March 2026.
-
NLOS-Aided Joint OTA Synchronization and Off-Grid Imaging for Distributed MIMO Systems
Authors:
Xin Tong,
Lechen Zhang,
Yu Ge,
Dario Tagliaferri,
Henk Wymeersch
Abstract:
Distributed multiple-input multiple-output (MIMO) architectures enable large-scale integrated sensing and communication (ISAC) by providing high spatial resolution and robustness through spatial diversity. However, practical phase-coherent sensing is challenged by phase synchronization errors and modeling mismatch caused by grid discretization. Existing over-the-air (OTA) synchronization methods t…
▽ More
Distributed multiple-input multiple-output (MIMO) architectures enable large-scale integrated sensing and communication (ISAC) by providing high spatial resolution and robustness through spatial diversity. However, practical phase-coherent sensing is challenged by phase synchronization errors and modeling mismatch caused by grid discretization. Existing over-the-air (OTA) synchronization methods typically treat synchronization and sensing tasks separately, which may lead to inaccurate phase alignment when multipath components are used for imaging. In this paper, we propose a non-line-of-sight (NLOS)-aided joint OTA synchronization and off-grid imaging framework for distributed MIMO ISAC systems. First, a line-of-sight (LOS)-assisted coarse synchronization is performed to establish initial phase coherence across distributed links. Subsequently, an iterative refinement stage exploits reconstructed NLOS components obtained from imaging results. By modeling off-grid effects via a first-order Taylor expansion, we transform measurements with nonlinear off-grid offset into an augmented linear model with jointly sparse reflectivity and off-set variables. The imaging problem is reformulated as a structured sparse recovery task and solved using a tailored off-grid approximate message passing (OG-AMP) algorithm. The imaging and synchronization modules are coupled within a closed-loop alternative optimization framework, where improved imaging enables more accurate phase refinement, and vice versa. Numerical results show that the proposed framework achieves accurate synchronization and imaging under phase errors. Compared with conventional approaches, it shows superior robustness and accuracy.
△ Less
Submitted 14 March, 2026;
originally announced March 2026.
-
Integrated Spoofing-Robust Automatic Speaker Verification via a Three-Class Formulation and LLR
Authors:
Kai Tan,
Lin Zhang,
Ruiteng Zhang,
Johan Rohdin,
Leibny Paola García-Perera,
Zexin Cai,
Sanjeev Khudanpur,
Matthew Wiesner,
Nicholas Andrews
Abstract:
Spoofing-robust automatic speaker verification (SASV) aims to integrate automatic speaker verification (ASV) and countermeasure (CM). A popular solution is fusion of independent ASV and CM scores. To better modeling SASV, some frameworks integrate ASV and CM within a single network. However, these solutions are typically bi-encoder based, offer limited interpretability, and cannot be readily adapt…
▽ More
Spoofing-robust automatic speaker verification (SASV) aims to integrate automatic speaker verification (ASV) and countermeasure (CM). A popular solution is fusion of independent ASV and CM scores. To better modeling SASV, some frameworks integrate ASV and CM within a single network. However, these solutions are typically bi-encoder based, offer limited interpretability, and cannot be readily adapted to new evaluation parameters without retraining. Based on this, we propose a unified end-to-end framework via a three-class formulation that enables log-likelihood ratio (LLR) inference from class logits for a more interpretable decision pipeline. Experiments show comparable performance to existing methods on ASVSpoof5 and better results on SpoofCeleb. The visualization and analysis also prove that the three-class reformulation provides more interpretability.
△ Less
Submitted 18 March, 2026; v1 submitted 14 March, 2026;
originally announced March 2026.
-
Can LLMs Help Localize Fake Words in Partially Fake Speech?
Authors:
Lin Zhang,
Thomas Thebaud,
Zexin Cai,
Sanjeev Khudanpur,
Daniel Povey,
Leibny Paola García-Perera,
Matthew Wiesner,
Nicholas Andrews
Abstract:
Large language models (LLMs), trained on large-scale text, have recently attracted significant attention for their strong performance across many tasks. Motivated by this, we investigate whether a text-trained LLM can help localize fake words in partially fake speech, where only specific words within a speech are edited. We build a speech LLM to perform fake word localization via next token predic…
▽ More
Large language models (LLMs), trained on large-scale text, have recently attracted significant attention for their strong performance across many tasks. Motivated by this, we investigate whether a text-trained LLM can help localize fake words in partially fake speech, where only specific words within a speech are edited. We build a speech LLM to perform fake word localization via next token prediction. Experiments and analyses on AV-Deepfake1M and PartialEdit indicates that the model frequently leverages editing-style pattern learned from the training data, particularly word-level polarity substitutions for those two databases we discussed, as cues for localizing fake words. Although such particular patterns provide useful information in an in-domain scenario, how to avoid over-reliance on such particular pattern and improve generalization to unseen editing styles remains an open question.
△ Less
Submitted 11 March, 2026;
originally announced March 2026.
-
Universal Speech Content Factorization
Authors:
Henry Li Xinyuan,
Zexin Cai,
Lin Zhang,
Leibny Paola García-Perera,
Berrak Sisman,
Sanjeev Khudanpur,
Nicholas Andrews,
Matthew Wiesner
Abstract:
We propose Universal Speech Content Factorization (USCF), a simple and invertible linear method for extracting a low-rank speech representation in which speaker timbre is suppressed while phonetic content is preserved. USCF extends Speech Content Factorization, a closed-set voice conversion (VC) method, to an open-set setting by learning a universal speech-to-content mapping via least-squares opti…
▽ More
We propose Universal Speech Content Factorization (USCF), a simple and invertible linear method for extracting a low-rank speech representation in which speaker timbre is suppressed while phonetic content is preserved. USCF extends Speech Content Factorization, a closed-set voice conversion (VC) method, to an open-set setting by learning a universal speech-to-content mapping via least-squares optimization and deriving speaker-specific transformations from only a few seconds of target speech. We show through embedding analysis that USCF effectively removes speaker-dependent variation. As a zero-shot VC system, USCF achieves competitive intelligibility, naturalness, and speaker similarity compared to methods that require substantially more target-speaker data or additional neural training. Finally, we demonstrate that as a training-efficient timbre-disentangled speech feature, USCF features can serve as the acoustic representation for training timbre-prompted text-to-speech models. Speech samples and code are publicly available.
△ Less
Submitted 7 June, 2026; v1 submitted 9 March, 2026;
originally announced March 2026.
-
A Curved Monopole Antenna for HF Radar with Enhanced Gain and Bandwidth
Authors:
Masoud Salmani Arani,
Reza Shahidi,
Lihong Zhang
Abstract:
This paper presents the design and simulation of a new curved monopole antenna optimized for skywave HF radar applications, with a systematic investigation of the effects of curvature and fixed-section length on antenna performance. The proposed design achieves improved impedance matching, broader bandwidth, and enhanced realized gain compared to a conventional quarter-wavelength monopole at 15 MH…
▽ More
This paper presents the design and simulation of a new curved monopole antenna optimized for skywave HF radar applications, with a systematic investigation of the effects of curvature and fixed-section length on antenna performance. The proposed design achieves improved impedance matching, broader bandwidth, and enhanced realized gain compared to a conventional quarter-wavelength monopole at 15 MHz. Parametric analysis shows that fully bending the monopole degrades performance, whereas introducing a straight section and carefully optimizing the curvature enables a 18.5% gain increase and a 400 kHz bandwidth expansion. The single-element design is further extended to a 12-element linear array with 0.45λ spacing (where λ is the wavelength), demonstrating stable embedded-element behavior and improved low-to- moderate elevation gain for skywave over-the-horizon radar operation. At θ = 30°, the proposed array achieves 14.04 dBi compared to 13.11 dBi for the reference array, corresponding to 24% gain enhancement, which is significant in high-power HF radar systems. These results confirm that the proposed curved monopole antenna provides a compact, broadband, and scalable solution for next-generation HF radar arrays.
△ Less
Submitted 8 March, 2026;
originally announced March 2026.
-
Machine Learning for the Internet of Underwater Things: From Fundamentals to Implementation
Authors:
Kenechi Omeke,
Attai Abubakar,
Michael Mollel,
Lei Zhang,
Qammer H. Abbasi,
Muhammad Ali Imran
Abstract:
The Internet of Underwater Things (IoUT) is becoming a critical infrastructure for ocean observation, marine resource management, and climate science. Its development is hindered by severe acoustic attenuation, propagation delays far exceeding those of terrestrial wireless systems, strict energy constraints, and dynamic topologies shaped by ocean currents. Machine learning (ML) has emerged as a ke…
▽ More
The Internet of Underwater Things (IoUT) is becoming a critical infrastructure for ocean observation, marine resource management, and climate science. Its development is hindered by severe acoustic attenuation, propagation delays far exceeding those of terrestrial wireless systems, strict energy constraints, and dynamic topologies shaped by ocean currents. Machine learning (ML) has emerged as a key enabler for addressing these limitations, offering data driven mechanisms that enhance performance across all layers of underwater wireless sensor networks. This tutorial survey synthesises ML methodologies supervised, unsupervised, reinforcement, and deep learning specifically contextualised for underwater communication environments. It outlines the algorithmic principles of each paradigm and examines the conditions under which particular approaches deliver superior performance. A layer wise analysis highlights physical layer gains in localisation and channel estimation, MAC layer adaptations that improve channel utilisation, network layer routing strategies that extend operational lifetime, and transport layer mechanisms capable of reducing packet loss by up to 91 percent. At the application layer, ML enables substantial data compression and object detection accuracies reaching 92 percent. Drawing on 300 studies from 2012 to 2025, the survey documents energy efficiency gains of 7 to 29 times, throughput improvements over traditional protocols, and cross layer optimisation benefits of up to 42 percent. It also identifies persistent barriers, including limited datasets, computational constraints, and the gap between theoretical models and real world deployment. The survey concludes with emerging research directions and a technology roadmap supporting ML adoption in operational underwater networks.
△ Less
Submitted 7 March, 2026;
originally announced March 2026.
-
In-batch Relational Features Enhance Precision in An Unsupervised Medical Anomaly Detection Task
Authors:
P. Bilha Githinji,
Ijaz Gul,
Lian Zhang,
Jinhao Xu,
Peiwu Qin,
Dongmei Yu
Abstract:
Confounding pathology with normal anatomical variation remains a significant challenge in unsupervised medical-image anomaly detection, resulting in numerous false positives. To enhance integration of healthy variation, we augment the latent representation of a CNN autoencoder with contextual similarities within a normal cohort through batch-wise hypergraph estimation and a shared-weights graph co…
▽ More
Confounding pathology with normal anatomical variation remains a significant challenge in unsupervised medical-image anomaly detection, resulting in numerous false positives. To enhance integration of healthy variation, we augment the latent representation of a CNN autoencoder with contextual similarities within a normal cohort through batch-wise hypergraph estimation and a shared-weights graph convolution layer, producing a population-aware embedding. On a heterogeneous brain-tumor dataset of 2D MRI scans, the method improves separability between healthy and pathological samples, achieving an AUC-ROC of 0.90 (95% CI 0.84-0.95, 5.7% absolute gain), and a 16% absolute improvement in average precision (0.78 AP, 95% CI 0.66-0.89), thereby lowering false-positive rates. Moreover, both anomaly detection and downstream tumor versus no-tumor classification performance improve with the size of the mini-batch context captured in the augmented representation, suggesting a tunable lever for integrating healthy variation.
△ Less
Submitted 3 July, 2026; v1 submitted 3 March, 2026;
originally announced March 2026.
-
Entropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech Codec
Authors:
Yanzhou Ren,
Noboru Harada,
Daiki Takeuchi,
Siyu Chen,
Wei Liu,
Xiao Zhang,
Liyuan Zhang,
Takehiro Moriya,
Shoji Makino
Abstract:
Neural audio codec (NAC) is essential for reconstructing high-quality speech signals and generating discrete representations for downstream speech language models. However, ensuring accurate semantic modeling while maintaining high-fidelity reconstruction under ultra-low bitrate constraints remains challenging. We propose an entropy-guided group residual vector quantization (EG-GRVQ) for an ultra-…
▽ More
Neural audio codec (NAC) is essential for reconstructing high-quality speech signals and generating discrete representations for downstream speech language models. However, ensuring accurate semantic modeling while maintaining high-fidelity reconstruction under ultra-low bitrate constraints remains challenging. We propose an entropy-guided group residual vector quantization (EG-GRVQ) for an ultra-low bitrate neural speech codec, which retains a semantic branch for linguistic information and incorporates an entropy-guided grouping strategy in the acoustic branch. Assuming that channel activations follow approximately Gaussian statistics, the variance of each channel can serve as a principled proxy for its information content. Based on this assumption, we partition the encoder output such that each group carries an equal share of the total information. This balanced allocation improves codebook efficiency and reduces redundancy. Trained on LibriTTS and VCTK, our model shows improvements in perceptual quality and intelligibility metrics under ultra-low bitrate conditions, with a focus on codec-level fidelity for communication-oriented scenarios.
△ Less
Submitted 2 March, 2026;
originally announced March 2026.
-
SignVLA: A Gloss-Free Vision-Language-Action Framework for Real-Time Sign Language-Guided Robotic Manipulation
Authors:
Xinyu Tan,
Ningwei Bai,
Harry Gardener,
Zhengyang Zhong,
Luoyu Zhang,
Liuhaichen Yang,
Zhekai Duan,
Monkgogi Galeitsiwe,
Zezhi Tang
Abstract:
We present, to our knowledge, the first sign language-driven Vision-Language-Action (VLA) framework for intuitive and inclusive human-robot interaction. Unlike conventional approaches that rely on gloss annotations as intermediate supervision, the proposed system adopts a gloss-free paradigm and directly maps visual sign gestures to semantic instructions. This design reduces annotation cost and av…
▽ More
We present, to our knowledge, the first sign language-driven Vision-Language-Action (VLA) framework for intuitive and inclusive human-robot interaction. Unlike conventional approaches that rely on gloss annotations as intermediate supervision, the proposed system adopts a gloss-free paradigm and directly maps visual sign gestures to semantic instructions. This design reduces annotation cost and avoids the information loss introduced by gloss representations, enabling more natural and scalable multimodal interaction.
In this work, we focus on a real-time alphabet-level finger-spelling interface that provides a robust and low-latency communication channel for robotic control. Compared with large-scale continuous sign language recognition, alphabet-level interaction offers improved reliability, interpretability, and deployment feasibility in safety-critical embodied environments. The proposed pipeline transforms continuous gesture streams into coherent language commands through geometric normalization, temporal smoothing, and lexical refinement, ensuring stable and consistent interaction.
Furthermore, the framework is designed to support future integration of transformer-based gloss-free sign language models, enabling scalable word-level and sentence-level semantic understanding. Experimental results demonstrate the effectiveness of the proposed system in grounding sign-derived instructions into precise robotic actions under diverse interaction scenarios. These results highlight the potential of the framework to advance accessible, scalable, and multimodal embodied intelligence.
△ Less
Submitted 25 February, 2026;
originally announced February 2026.
-
Secure Semantic Communications via AI Defenses: Fundamentals, Solutions, and Future Directions
Authors:
Lan Zhang,
Chengsi Liang,
Zeming Zhuang,
Yao Sun,
Fang Fang,
Xiaoyong Yuan,
Dusit Niyato
Abstract:
Semantic communication (SemCom) redefines wireless communication from reproducing symbols to transmitting task-relevant semantics. However, this AI-native architecture also introduces new vulnerabilities, as semantic failures may arise from adversarial perturbations to models, corrupted training data, desynchronized priors, or misaligned inference even when lower-layer transmission reliability and…
▽ More
Semantic communication (SemCom) redefines wireless communication from reproducing symbols to transmitting task-relevant semantics. However, this AI-native architecture also introduces new vulnerabilities, as semantic failures may arise from adversarial perturbations to models, corrupted training data, desynchronized priors, or misaligned inference even when lower-layer transmission reliability and cryptographic protection remain intact. This survey provides a defense-centered and system-oriented synthesis of security in SemCom via AI defense. We analyze AI-centric threat models by consolidating existing studies and organizing attack surfaces across model-level, channel-realizable, knowledge-based, and networked inference vectors. Building on this foundation, we present a structured taxonomy of defense strategies organized by where semantic integrity can be compromised in SemCom systems despite correct symbol delivery, spanning semantic encoding, wireless transmission, knowledge integrity, and coordination among multiple agents. These categories correspond to distinct security failure modes, including representation fragility, channel-realizable manipulation, semantic prior poisoning or desynchronization, and adversarial propagation through distributed inference. We also examine security utility operating envelopes that capture tradeoffs among semantic fidelity, robustness, latency, and energy under realistic constraints, survey evaluation frameworks and representative applications, and identify open challenges in cross-layer composition and deployment-time certification. Overall, this survey offers a unified system-level perspective that enables readers to understand major threat and defense mechanisms in AI-native SemCom systems and to leverage emerging security techniques in the design and deployment of robust SemCom architectures for next-generation intelligent networks.
△ Less
Submitted 4 March, 2026; v1 submitted 25 February, 2026;
originally announced February 2026.
-
Exploiting Completeness Perception with Diffusion Transformer for Unified 3D MRI Synthesis
Authors:
Junkai Liu,
Nay Aung,
Theodoros N. Arvanitis,
Joao A. C. Lima,
Steffen E. Petersen,
Le Zhang
Abstract:
Missing data problems, such as missing modalities in multi-modal brain MRI and missing slices in cardiac MRI, pose significant challenges in clinical practice. Existing methods rely on external guidance to supply detailed missing-state information for instructing generative models to synthesize missing MRIs. However, manual indicators are not always available or reliable in real-world scenarios du…
▽ More
Missing data problems, such as missing modalities in multi-modal brain MRI and missing slices in cardiac MRI, pose significant challenges in clinical practice. Existing methods rely on external guidance to supply detailed missing-state information for instructing generative models to synthesize missing MRIs. However, manual indicators are not always available or reliable in real-world scenarios due to the unpredictable nature of clinical environments. Moreover, these explicit masks are not informative enough to provide guidance for improving semantic consistency. In this work, we argue that generative models should infer and recognize missing states in a self-perceptive manner, enabling them to better capture subtle anatomical and pathological variations. Towards this goal, we propose CoPeDiT, a shared completeness-perception framework for 3D MRI synthesis, following a common conditioning strategy with task-specific instantiations for different missing-data scenarios. Specifically, we incorporate dedicated pretext tasks into our tokenizer, CoPeVAE, empowering it to learn completeness-aware discriminative prompt tokens, and design MDiT3D, a specialized diffusion transformer architecture for 3D MRI synthesis that effectively uses the completeness-aware prompt tokens as guidance to enhance semantic consistency in 3D space. Comprehensive evaluations on three large-scale MRI datasets demonstrate that CoPeDiT consistently improves upon state-of-the-art methods across diverse missing patterns, yielding high-fidelity and structurally consistent MRI synthesis. Our code is available at https://github.com/JK-Liu7/CoPeDiT.
△ Less
Submitted 20 August, 2026; v1 submitted 20 February, 2026;
originally announced February 2026.
-
MeDUET: Disentangled Unified Pretraining for 3D Medical Image Synthesis and Analysis
Authors:
Junkai Liu,
Ling Shao,
Le Zhang
Abstract:
Self-supervised learning (SSL) and diffusion models have respectively advanced representation learning and generative modeling for high-dimensional 3D visual data, yet they are often developed as separate paradigms. Their unification remains challenging under multi-source heterogeneity, as anatomical content must be preserved for analysis while acquisition-related style varies across centers and a…
▽ More
Self-supervised learning (SSL) and diffusion models have respectively advanced representation learning and generative modeling for high-dimensional 3D visual data, yet they are often developed as separate paradigms. Their unification remains challenging under multi-source heterogeneity, as anatomical content must be preserved for analysis while acquisition-related style varies across centers and affects synthesis. In this paper, we propose MeDUET, a 3D Medical image Disentangled UnifiEd PreTraining framework in the variational autoencoder latent space. MeDUET formulates unified pretraining as an empirical factor identifiability problem, aiming to learn domain-invariant content factors for anatomy and domain-specific style factors for appearance. To improve factor separation, MeDUET first uses token demixing with a standard adversarial domain regularizer to establish basic content-style specialization, and further introduces Mixed Factor Token Distillation and Swap-invariance Quadruplet Contrast to reduce mixed-region factor leakage and organize factor spaces with factor-wise invariance and discriminability. With these learned factors, MeDUET transfers effectively to both synthesis and analysis, yielding higher fidelity, faster convergence, and better controllability for synthesis, while achieving competitive or superior domain generalization and label efficiency on diverse datasets, tasks, and modalities. Overall, MeDUET shows that multi-source heterogeneity can serve as useful supervision, with disentanglement providing an effective interface for unifying 3D medical image synthesis and analysis. Our code is available at https://github.com/JK-Liu7/MeDUET.
△ Less
Submitted 26 June, 2026; v1 submitted 19 February, 2026;
originally announced February 2026.
-
Covo-Audio Technical Report
Authors:
Wenfu Wang,
Chenxing Li,
Liqiang Zhang,
Yiyang Zhao,
Yuxiang Zou,
Hanzhao Li,
Mingyu Cui,
Hao Zhang,
Kun Wei,
Le Xu,
Zikang Huang,
Jiajun Xu,
Jiliang Hu,
Xiang He,
Zeyu Xie,
Jiawen Kang,
Youjun Chen,
Meng Yu,
Dong Yu,
Rilin Chen,
Linlin Di,
Shulin Feng,
Na Hu,
Yang Liu,
Bang Wang
, et al. (1 additional authors not shown)
Abstract:
In this work, we present Covo-Audio, a 7B-parameter end-to-end LALM that directly processes continuous audio inputs and generates audio outputs within a single unified architecture. Through large-scale curated pretraining and targeted post-training, Covo-Audio achieves state-of-the-art or competitive performance among models of comparable scale across a broad spectrum of tasks, including speech-te…
▽ More
In this work, we present Covo-Audio, a 7B-parameter end-to-end LALM that directly processes continuous audio inputs and generates audio outputs within a single unified architecture. Through large-scale curated pretraining and targeted post-training, Covo-Audio achieves state-of-the-art or competitive performance among models of comparable scale across a broad spectrum of tasks, including speech-text modeling, spoken dialogue, speech understanding, audio understanding, and full-duplex voice interaction. Extensive evaluations demonstrate that the pretrained foundation model exhibits strong speech-text comprehension and semantic reasoning capabilities on multiple benchmarks, outperforming representative open-source models of comparable scale. Furthermore, Covo-Audio-Chat, the dialogue-oriented variant, demonstrates strong spoken conversational abilities, including understanding, contextual reasoning, instruction following, and generating contextually appropriate and empathetic responses, validating its applicability to real-world conversational assistant scenarios. Covo-Audio-Chat-FD, the evolved full-duplex model, achieves substantially superior performance on both spoken dialogue capabilities and full-duplex interaction behaviors, demonstrating its competence in practical robustness. To mitigate the high cost of deploying end-to-end LALMs for natural conversational systems, we propose an intelligence-speaker decoupling strategy that separates dialogue intelligence from voice rendering, enabling flexible voice customization with minimal text-to-speech (TTS) data while preserving dialogue performance. Overall, our results highlight the strong potential of 7B-scale models to integrate sophisticated audio intelligence with high-level semantic reasoning, and suggest a scalable path toward more capable and versatile LALMs.
△ Less
Submitted 16 March, 2026; v1 submitted 10 February, 2026;
originally announced February 2026.
-
NarraScore: Bridging Visual Narrative and Musical Dynamics via Hierarchical Affective Control
Authors:
Yufan Wen,
Zhaocheng Liu,
YeGuo Hua,
Ziyi Guo,
Lihua Zhang,
Chun Yuan,
Jian Wu
Abstract:
Synthesizing coherent soundtracks for long-form videos remains a formidable challenge, currently stalled by three critical impediments: computational scalability, temporal coherence, and, most critically, a pervasive semantic blindness to evolving narrative logic. To bridge these gaps, we propose NarraScore, a hierarchical framework predicated on the core insight that emotion serves as a high-dens…
▽ More
Synthesizing coherent soundtracks for long-form videos remains a formidable challenge, currently stalled by three critical impediments: computational scalability, temporal coherence, and, most critically, a pervasive semantic blindness to evolving narrative logic. To bridge these gaps, we propose NarraScore, a hierarchical framework predicated on the core insight that emotion serves as a high-density compression of narrative logic. Uniquely, we repurpose frozen Vision-Language Models (VLMs) as continuous affective sensors, distilling high-dimensional visual streams into dense, narrative-aware Valence-Arousal trajectories. Mechanistically, NarraScore employs a Dual-Branch Injection strategy to reconcile global structure with local dynamism: a \textit{Global Semantic Anchor} ensures stylistic stability, while a surgical \textit{Token-Level Affective Adapter} modulates local tension via direct element-wise residual injection. This minimalist design bypasses the bottlenecks of dense attention and architectural cloning, effectively mitigating the overfitting risks associated with data scarcity. Experiments demonstrate that NarraScore achieves state-of-the-art consistency and narrative alignment with negligible computational overhead, establishing a fully autonomous paradigm for long-video soundtrack generation.
△ Less
Submitted 11 February, 2026; v1 submitted 9 February, 2026;
originally announced February 2026.
-
Late Breaking Results: Conversion of Neural Networks into Logic Flows for Edge Computing
Authors:
Daniel Stein,
Shaoyi Huang,
Rolf Drechsler,
Bing Li,
Grace Li Zhang
Abstract:
Neural networks have been successfully applied in various resource-constrained edge devices, where usually central processing units (CPUs) instead of graphics processing units exist due to limited power availability. State-of-the-art research still focuses on efficiently executing enormous numbers of multiply-accumulate (MAC) operations. However, CPUs themselves are not good at executing such math…
▽ More
Neural networks have been successfully applied in various resource-constrained edge devices, where usually central processing units (CPUs) instead of graphics processing units exist due to limited power availability. State-of-the-art research still focuses on efficiently executing enormous numbers of multiply-accumulate (MAC) operations. However, CPUs themselves are not good at executing such mathematical operations on a large scale, since they are more suited to execute control flow logic, i.e., computer algorithms. To enhance the computation efficiency of neural networks on CPUs, in this paper, we propose to convert them into logic flows for execution. Specifically, neural networks are first converted into equivalent decision trees, from which decision paths with constant leaves are then selected and compressed into logic flows. Such logic flows consist of if and else structures and a reduced number of MAC operations. Experimental results demonstrate that the latency can be reduced by up to 14.9 % on a simulated RISC-V CPU without any accuracy degradation.
The code is open source at https://github.com/TUDa-HWAI/NN2Logic
△ Less
Submitted 29 January, 2026;
originally announced January 2026.