-
The Enceladian crater production function
Authors:
E. W. Wong,
S. C. Werner,
M. R. Kirchoff,
R. Brasser
Abstract:
Interpreting Enceladus's past and present surface history and interior state remains challenging, owing to uncertain prescription of its impact bombardment history and limited interpretation of its crater statistics. Further progress in understanding its evolutionary history can be achieved with an improved crater chronology model and a thorough assessment of Enceladus's surface. Here we present t…
▽ More
Interpreting Enceladus's past and present surface history and interior state remains challenging, owing to uncertain prescription of its impact bombardment history and limited interpretation of its crater statistics. Further progress in understanding its evolutionary history can be achieved with an improved crater chronology model and a thorough assessment of Enceladus's surface. Here we present the first step in the form of a comprehensive, global crater catalogue with geomorphology survey for Enceladus. From our dataset we build the crater production function (CPF), which is the underlying, unmodified crater size-frequency distribution of the satellite surface, assuming no subsequent modification. We obtained the CPF with a data-driven approach; therefore it makes no assumptions about the impactor source population or planet evolution models or the timing of impact. We fit the CPF with a high-order polynomial as is customary for the Moon and Mars, capturing the slope variations across different crater diameter ranges. The Enceladian CPF generally has a steeper cumulative slope than that of the Moon and Mars for small crater diameters D_cr < 10 km, as well as that of the size-frequency distribution of trans-Neptunian objects. This CPF serves as a critical observational input for an Enceladian crater chronology model, enabling the conversion of crater densities into absolute surface ages. Extending the crater cataloguing and CPF derivation of this study to other Saturnian satellites will determine whether the Enceladian CPF is unique; a shared CPF would indicate a common impactor population, providing observational constraints on the size-frequency distribution of small bodies in the outer Solar System.
△ Less
Submitted 14 August, 2026;
originally announced August 2026.
-
Fashion Outfit Generation via Unified Sequential Composition Models
Authors:
Kaicheng Pang,
Xingxing Zou,
Ruohan Xu,
Waikeung Wong
Abstract:
The task of synthesizing stylistically coherent fashion outfits from massive item libraries, known as fashion outfit generation, remains a non-trivial challenge, primarily due to the non-monotonic and implicit nature of aesthetic compatibility, coupled with the exponentially large combinatorial search space. In this paper, we formalize this task as Constrained Ensemble Generation (CEG) and model i…
▽ More
The task of synthesizing stylistically coherent fashion outfits from massive item libraries, known as fashion outfit generation, remains a non-trivial challenge, primarily due to the non-monotonic and implicit nature of aesthetic compatibility, coupled with the exponentially large combinatorial search space. In this paper, we formalize this task as Constrained Ensemble Generation (CEG) and model it as a finite-horizon deterministic Markov Decision Process. To address CEG in fashion, we propose the Unified Sequential Composition Model (USCM), which jointly models set-level compatibility and latent composition intents. Guided by USCM's learned priors, a Latent Expansion Monte Carlo Tree Search (LE-MCTS) mechanism is proposed to handle item retrieval during composition, balancing local aesthetic synergy with global structural balance. Extensive experiments on the Polyvore Outfits dataset, along with zero-shot evaluations on the iFashion and PolyvoreU datasets, demonstrate that our framework achieves state-of-the-art performance across independent human preference evaluations, automated aesthetic proxies, and structural validity metrics for constrained fashion outfit generation.
△ Less
Submitted 13 August, 2026;
originally announced August 2026.
-
Asymptotic Behavior and Error Bounds for Fisher-KPP Equations on the Real Half-Line
Authors:
Chu Chu,
M. W. Wong
Abstract:
We study the Fisher--KPP equation on the half-line under Dirichlet,Neumann, and Robin boundary conditions. For the autonomous logistic equation, we identify bounded stationary profiles converging to $1$ and obtain exponential far-field comparison estimates. We prove local uniform convergence of nontrivial Neumann solutions to $1$. Assuming local uniform convergence of the Robin solution to its sta…
▽ More
We study the Fisher--KPP equation on the half-line under Dirichlet,Neumann, and Robin boundary conditions. For the autonomous logistic equation, we identify bounded stationary profiles converging to $1$ and obtain exponential far-field comparison estimates. We prove local uniform convergence of nontrivial Neumann solutions to $1$. Assuming local uniform convergence of the Robin solution to its stationary profile, we derive asymptotic Neumann--Robin comparison estimates. We then consider small time-periodic Neumann and Robin boundary forcing. Under exponential stability of the homogeneous linearized semigroup, we construct a locally unique small periodic lifted mild solution and obtain a first-order expansion with a uniform $O(\varepsilon^2)$ remainder in $C_0([0,\infty))$.
△ Less
Submitted 13 August, 2026; v1 submitted 27 July, 2026;
originally announced August 2026.
-
Lapis: Laplacian Spiking Attention via First-Spike Timing and Membrane Leakage
Authors:
Kaiwen Tang,
Jiaqi Zheng,
Zixuan Zhu,
Yiqun Wang,
Zhanglu Yan,
Weng-Fai Wong
Abstract:
Self-attention has become central to spiking vision transformers, yet its query-key scoring is still largely inherited from dense networks. Existing spiking variants either simplify dot product scoring or replace it with discrete operators, but spike timing, the native variable of a spiking network, does not directly define how tokens are related. We propose Lapis, a spiking attention mechanism th…
▽ More
Self-attention has become central to spiking vision transformers, yet its query-key scoring is still largely inherited from dense networks. Existing spiking variants either simplify dot product scoring or replace it with discrete operators, but spike timing, the native variable of a spiking network, does not directly define how tokens are related. We propose Lapis, a spiking attention mechanism that scores each token pair by the L1 distance between its query and key first-spike latency vectors under time-to-first-spike coding, and maps this distance to an affinity through a Laplacian kernel. The kernel's exponential decay matches the impulse response of a leaky integrate-and-fire membrane, so the accumulated latency difference determines the decay of a membrane trace, while row normalization reduces to a bit shift under power-of-two rounding. Scoring therefore needs only subtraction, absolute value, and accumulation, and removes all multiplication between query and key channels. Under a matched backbone and training schedule, Lapis reaches 96.56% top-1 accuracy on CIFAR-10, within 0.53 points of dot-product scoring. On ImageNet-1K, it reduces the estimated arithmetic energy of the attention path by 14.5x relative to dense dot-product attention. The deployed 6-bit model attains 83.25% top-1 accuracy at an estimated arithmetic energy of 3.28mJ per image.
△ Less
Submitted 16 August, 2026; v1 submitted 12 August, 2026;
originally announced August 2026.
-
An AoI-oriented Time-Frequency Distributed Access Mechanism in Wireless Sensor Networks with Spectrum Division
Authors:
Jingwei Liu,
Fang Liu,
Wing Shing Wong,
Yuan-Hsun Lo,
Chung Shue Chen
Abstract:
The increasing adoption of spectrum-division techniques enables concurrent uplink transmissions over multiple orthogonal resources, yet low-overhead access design with effective information freshness remains insufficiently studied for large-scale randomly activated sensor networks. In this paper, we apply the age of information (AoI) to measure information freshness and propose an AoI-efficient de…
▽ More
The increasing adoption of spectrum-division techniques enables concurrent uplink transmissions over multiple orthogonal resources, yet low-overhead access design with effective information freshness remains insufficiently studied for large-scale randomly activated sensor networks. In this paper, we apply the age of information (AoI) to measure information freshness and propose an AoI-efficient deterministic time-frequency distributed access (D-TFDA) mechanism. D-TFDA combines centralized configuration and distributed operation through a periodic token-based time-frequency structure, which provides sensors with collision-free and predictable transmission opportunities without considerable run-time overhead. We develop an analytical framework to characterize the long-term average AoI (AAoI) by exploiting the periodicity of the token assignment pattern and modeling the steady local state of each sensor with a one-dimensional discrete-time Markov chain (DTMC). We further reveal structural properties of the token assignment pattern and identify AoI-equivalent token clusters, which substantially reduce the search space of the AAoI-optimal token allocation problem. Based on this structure, we formulate the reduced problem as a linear programming (LP) problem and develop an AAoI-optimal search algorithm, together with an auction-inspired heuristic algorithm of lower complexity. Simulation results validate the proposed AAoI analysis, demonstrate the effectiveness of the token allocation algorithms, and show that D-TFDA achieves substantially lower AAoI than optimized random access baselines by avoiding collisions and exploiting heterogeneous sensor--resource transmission reliability.
△ Less
Submitted 11 August, 2026;
originally announced August 2026.
-
A Bayesian Weakest-Link Framework for Joint Estimation of Material Strength and Stress Profile
Authors:
Shiyu He,
Samuel W. K. Wong
Abstract:
For structural components whose failure is governed by the weakest-link theory, existing reliability models typically either assume that the underlying mechanical model is known or neglect to exploit the spatial information contained in observed failure locations. In practice, however, idealized mechanical models may systematically deviate from the actual stress due to simplifying or incorrect ass…
▽ More
For structural components whose failure is governed by the weakest-link theory, existing reliability models typically either assume that the underlying mechanical model is known or neglect to exploit the spatial information contained in observed failure locations. In practice, however, idealized mechanical models may systematically deviate from the actual stress due to simplifying or incorrect assumptions. To address this limitation, we propose a hierarchical Bayesian weakest-link model that jointly estimates the latent material strength and stress profile from paired failure load and failure zone observations. In our formulation, the stress profile is estimated via a B-spline basis expansion, and the non-differentiable weakest-link mechanism is approximated by a differentiable Softmin function to account for unobserved material flaws and facilitate Bayesian inference. Simulation studies demonstrate that the proposed framework provides accurate and robust estimation under various experiment configurations. Applied to a real-data analysis of Douglas-fir crossarms, the proposed model identifies systematic deviations from idealized beam theory that cannot be captured by deterministic stress derivations.
△ Less
Submitted 2 August, 2026;
originally announced August 2026.
-
Generalized Query-Oriented Image Semantic Coding Empowered by Large AI Models and Semantic-Aware Hybrid Beamforming
Authors:
Sin-Yu Huang,
Vincent W. S. Wong
Abstract:
Semantic communication is an emerging paradigm that can preserve the meaning of data during transmission. However, human users are often interested in specific semantic content based on their intent, and users' intent is often not considered in current semantic coding design. Moreover, most of the existing semantic models are fine-tuned using specific datasets, which limits their generalization ca…
▽ More
Semantic communication is an emerging paradigm that can preserve the meaning of data during transmission. However, human users are often interested in specific semantic content based on their intent, and users' intent is often not considered in current semantic coding design. Moreover, most of the existing semantic models are fine-tuned using specific datasets, which limits their generalization capability. Furthermore, how to prioritize semantically important features in large-scale multiple-input multiple-output orthogonal frequency-division multiplexing (MIMO-OFDM) systems remains largely unexplored. To address the aforementioned challenges, in this paper, we propose a generalized query-oriented image semantic coding (QO-ISC) framework. In the proposed framework, the transmitter extracts features which are relevant to the user's query and the receiver reconstructs an image based on those features. We use a pretrained large artificial intelligence (AI) model (LAM) to enhance general feature representations. We develop a semantic-aware hybrid beamforming (SA-HBF) algorithm to prioritize semantically important features for large-scale MIMO-OFDM system. When evaluated on unseen object categories within the dataset, simulation results show that our proposed generalized QO-ISC framework achieves better performance than the traditional codec and two state-of-the-art semantic coding schemes.
△ Less
Submitted 30 July, 2026;
originally announced July 2026.
-
Student Perceptions and Preferences Regarding AI-Generated Instructional Videos in Computing Education
Authors:
Esse Ciego,
Shubbhi Taneja,
Wilson Wong,
Amanpreet Kapoor
Abstract:
Students differ in how they prefer to engage with learning resources, with some favoring textual materials and others visual or video-based content. Recent advances in generative AI have led CS education research to focus on text-based AI tools for developing learning resources. However, advances in AI video models and the rapid proliferation of AI video generation tools have made it possible for…
▽ More
Students differ in how they prefer to engage with learning resources, with some favoring textual materials and others visual or video-based content. Recent advances in generative AI have led CS education research to focus on text-based AI tools for developing learning resources. However, advances in AI video models and the rapid proliferation of AI video generation tools have made it possible for instructors to create high-quality personalized educational videos efficiently and cost-effectively. Understanding students' perceptions of AI-generated videos is thus critical for helping CS instructors know when and how to use them purposefully. To address this gap, we conducted a descriptive post-test survey study in which 170 computing students at two U.S. institutions watched three 3-minute AI videos on the Markdown markup language created with Knowlify. Students then completed a survey about their perceptions of the Markdown videos and their broader views on the use of AI-generated videos in education. Students rated the Markdown videos as high-quality, accurate, and usable, with nearly half unable to determine whether the videos were AI-generated. At the same time, students expressed limited comfort with the widespread adoption of AI videos in the classroom. They preferred AI videos for simple, supplemental, and visual use cases, while expressing concerns about lower-quality or inaccurate content, reduced instructor interaction, and diminished educational value.
△ Less
Submitted 30 July, 2026;
originally announced July 2026.
-
Assisting Mission-Critical Traffic Flows with Active Queue Management in Industrial Internet of Things
Authors:
Shuo Wang,
Jonathan Kua,
Jiong Jin,
Yew Wee Wong,
Prem Prakash Jayaraman,
Zhibo Pang
Abstract:
Mission-critical Industrial Internet of Things (IIoT) traffic flows require bounded network latency and jitter guarantees to ensure the safe functioning of critical industrial infrastructure. These flows are typically communicated via commodity network routers with conventional First-In-First-Out (FIFO) buffers. FIFO has proven to be the culprit of the well-known bufferbloat phenomenon, and the de…
▽ More
Mission-critical Industrial Internet of Things (IIoT) traffic flows require bounded network latency and jitter guarantees to ensure the safe functioning of critical industrial infrastructure. These flows are typically communicated via commodity network routers with conventional First-In-First-Out (FIFO) buffers. FIFO has proven to be the culprit of the well-known bufferbloat phenomenon, and the deployment of Active Queue Management (AQM) schemes have demonstrated significant performance improvements for latency-sensitive applications over the Internet in the IT domain. However, the bufferbloat phenomenon and the efficacy of AQM schemes have not been studied in IIoT-based OT domain. In this paper, we propose the use of AQM as a lightweight and non-intrusive mechanism for assisting mission-critical traffic flows in IIoT networks. Our experimental results demonstrated that multi-queue AQM schemes provide substantial flow isolation and capacity sharing benefits, and significantly improve the performance of mission-critical traffic flows under network pressure. We further provide deployment recommendations based on our experimental insights.
△ Less
Submitted 15 July, 2026;
originally announced July 2026.
-
Efficient and Robust Spiking Neural Networks for sEMG-Based Muscle Fatigue Detection
Authors:
Kaiwen Tang,
Jiaqi Dong,
Zhanglu Yan,
Weng-Fai Wong
Abstract:
Detecting muscle fatigue via surface electromyography (sEMG) is essential for applications in sports, rehabilitation, and wearable health monitoring. Accurate and timely detection of fatigue is crucial for preventing injuries, optimizing physical performance, and ensuring user safety during prolonged activity. However, existing deep learning models are often unsuitable for this task due to their h…
▽ More
Detecting muscle fatigue via surface electromyography (sEMG) is essential for applications in sports, rehabilitation, and wearable health monitoring. Accurate and timely detection of fatigue is crucial for preventing injuries, optimizing physical performance, and ensuring user safety during prolonged activity. However, existing deep learning models are often unsuitable for this task due to their high computational cost and dependence on large-scale data. In this work, we propose an energy-efficient framework for muscle fatigue detection based on Spiking Neural Networks (SNNs), which exploit sparse, event-driven computation and temporal modeling. We further introduce a quantization-compatible training scheme (SDH) that combines multiple regularization terms to improve robustness under noisy conditions. Evaluated on two public sEMG datasets against a broad set of baselines and under seven noise conditions including physically motivated perturbations, our quantized SNNs match or exceed strong baselines while remaining more stable under diverse noise and reducing estimated energy consumption by up to 201.77x. These results demonstrate the framework's strong potential for real-time deployment in low-power wearable systems.
△ Less
Submitted 13 July, 2026;
originally announced July 2026.
-
Sharp Nordhaus-Gaddum bounds for throttling
Authors:
Ryan Blair,
Gabriel Elvin,
Veronika Furst,
Leslie Hogben,
Tony W. H. Wong
Abstract:
Throttling is a graph optimization problem, where the throttling number of a graph is the minimum sum or minimum product of the number of vertices in an initial set and the time required to complete a certain graph operation. A Nordhaus-Gaddum bound refers to an upper or lower bound of the sum or product of a graph parameter together with that of its complement. In this paper, we study the Nordhau…
▽ More
Throttling is a graph optimization problem, where the throttling number of a graph is the minimum sum or minimum product of the number of vertices in an initial set and the time required to complete a certain graph operation. A Nordhaus-Gaddum bound refers to an upper or lower bound of the sum or product of a graph parameter together with that of its complement. In this paper, we study the Nordhaus-Gaddum sum and product bounds of the various throttling numbers (sum throttling and product throttling with or without initial cost). Graph operations considered are standard zero forcing, positive semidefinite forcing, power domination, and Cops and Robbers.
△ Less
Submitted 29 June, 2026;
originally announced June 2026.
-
On Zeckendorf-Niven numbers and arithmetic progressions
Authors:
Kelly Lao,
Steven J. Miller,
Nicholas Rosa,
Mark Shiliaev,
Garrett Tresch,
Tony W. H. Wong,
Han Zhang
Abstract:
A positive integer is Zeckendorf-Niven (respectively, Lucas-Niven) if it is divisible by the number of summands in its Zeckendorf decomposition (respectively, Lucas decomposition). We show that there exist infinitely many Zeckendorf-Niven numbers and Lucas-Niven numbers in every arithmetic progression. Furthermore, we provide bounds on the maximum number of consecutive Zeckendorf-Niven terms in ce…
▽ More
A positive integer is Zeckendorf-Niven (respectively, Lucas-Niven) if it is divisible by the number of summands in its Zeckendorf decomposition (respectively, Lucas decomposition). We show that there exist infinitely many Zeckendorf-Niven numbers and Lucas-Niven numbers in every arithmetic progression. Furthermore, we provide bounds on the maximum number of consecutive Zeckendorf-Niven terms in certain arithmetic progressions.
△ Less
Submitted 22 June, 2026;
originally announced June 2026.
-
Wireless Personal Agent: Extending Wireless Intelligence from Networks to Terminals
Authors:
Jiedan Tan,
Fang Liu,
Jingwen Tong,
Shengli Zhang,
Jun Zhang,
Wing Shing Wong
Abstract:
Wireless networks are evolving from connectivity-oriented infrastructures into intelligent and personalized service platforms. Existing wireless intelligence remains centered on network-side optimization, improving objectives such as throughput, latency, and coverage. Nevertheless, besides network performance, wireless intelligence also depends on user-perceived experience via application context,…
▽ More
Wireless networks are evolving from connectivity-oriented infrastructures into intelligent and personalized service platforms. Existing wireless intelligence remains centered on network-side optimization, improving objectives such as throughput, latency, and coverage. Nevertheless, besides network performance, wireless intelligence also depends on user-perceived experience via application context, mobility routine, service cost, privacy preference, and long-term usage behavior. This article proposes WISPA, a Wireless Intelligent Self-evolving Personal Agent framework for automated terminal-side resource management based on large language model (LLM)-based agent. To overcome the resource constraints on terminals, WISPA decouples the latency-sensitive online resource execution from offline LLM agent reflection. In this way, a lightweight online executor makes deterministic resource decisions using interpretable preference parameters; While an offline LLM agent analyzes terminal-side traces, refines user profiles, and updates online preference parameters for subsequent decisions. At last, we demonstrate the practical applicability and benefits of WISPA for terminal-side resource allocations on a campus commute route. Numerical results show that WISPA learns user-specific connection styles and adapts access decisions as preferences change.
△ Less
Submitted 22 June, 2026;
originally announced June 2026.
-
Approaching Shannon Bound with Lossless LLM Weight Compression
Authors:
Hongshi Tan,
Yao Chen,
Gustavo Alonso,
Weng-Fai Wong,
Bingsheng He
Abstract:
Large language models (LLMs) now scale to trillions of parameters, driving weight storage into the terabyte regime and creating an acute mismatch with GPU memory capacity. Although lossless compression is widely effective in other domains, it remains underutilized in LLM systems. Through a comprehensive entropy study across models from 1.5B to 405B parameters and numeric formats ranging from bf16…
▽ More
Large language models (LLMs) now scale to trillions of parameters, driving weight storage into the terabyte regime and creating an acute mismatch with GPU memory capacity. Although lossless compression is widely effective in other domains, it remains underutilized in LLM systems. Through a comprehensive entropy study across models from 1.5B to 405B parameters and numeric formats ranging from bf16 to int4 and AWQ/SQ8, we find that LLM weights contain far less intrinsic randomness than their stored bitwidth implies, their effective entropy is 2-10x lower, indicating that up to a 10x footprint reduction is theoretically achievable without altering any weight values. Leveraging this insight, we introduce a tile-level, on-the-fly lossless decompression framework based on Asymmetric Numeral Systems that aligns decoding with the GEMM tiling pattern of GPU inference. Our design achieves bit-rates within 0.01-0.1 bits of the Shannon limit across a wide range of LLM numerical formats, demonstrating that nearly all statistical redundancy is eliminated. Integrated into the SGLang serving framework with multi-GPU support, our approach increases the maximum batch size of Qwen-14B from 47 to 75, improving throughput by up to 1.2x. On Mixtral-176B, the feasible batch size increases from 20 to 95 (4.8x), yielding up to 1.6x throughput improvement. Compared to state-of-the-art lossless compression approaches NeuZip and DFloat11, our design further improves throughput by up to 11x.
△ Less
Submitted 14 June, 2026;
originally announced June 2026.
-
Atoms in the Semigroup of Non-Negative Integer Matrices
Authors:
Lindsay Dever,
Eva G. Goedhart,
Gregory S. Heilbrunn,
Tony W. H. Wong
Abstract:
In the semigroup $M_2(\mathbb{N}_0)^\bullet$, two-by-two matrices with non-negative integer entries and non-zero determinant, we study the factorization of matrices into atoms, or irreducible matrices. In 2022, Baeth et al. listed some fundamental classes of atoms in $M_2(\mathbb{N}_0)^\bullet$; however, the factorability of most matrices in $M_2(\mathbb{N}_0)^\bullet$ remains unknown. We identify…
▽ More
In the semigroup $M_2(\mathbb{N}_0)^\bullet$, two-by-two matrices with non-negative integer entries and non-zero determinant, we study the factorization of matrices into atoms, or irreducible matrices. In 2022, Baeth et al. listed some fundamental classes of atoms in $M_2(\mathbb{N}_0)^\bullet$; however, the factorability of most matrices in $M_2(\mathbb{N}_0)^\bullet$ remains unknown. We identify two additional classes of atoms: a class of atoms with determinant $p$, $2p$, or $4p$, for $p$ prime, and a class of atoms in which the main diagonal is much "larger" than the off-diagonal (or vice versa). Finally, we show that bisymmetric matrices with relatively prime entries are a divisor-closed subset of $M_2(\mathbb{N}_0)^\bullet$ and use a factor search algorithm to classify bisymmetric atoms of $M_2(\mathbb{N}_0)^\bullet$ with minimum entry up to 4000.
△ Less
Submitted 12 June, 2026;
originally announced June 2026.
-
Understanding Truncated Positional Encodings for Graph Neural Networks
Authors:
James Flora,
Mitchell Black,
Weng-Keen Wong,
Amir Nayyeri
Abstract:
Positional encodings (PEs) enhance the power of graph neural networks (GNNs), both theoretically and empirically. Two of the most popular families of PEs - spectral (e.g., Laplacian eigenspaces, effective resistance) and walk-based (polynomials of the adjacency matrix) - are theoretically equivalent in expressive power, with expressivity between the 1-WL and 3-WL tests. However, this equivalence a…
▽ More
Positional encodings (PEs) enhance the power of graph neural networks (GNNs), both theoretically and empirically. Two of the most popular families of PEs - spectral (e.g., Laplacian eigenspaces, effective resistance) and walk-based (polynomials of the adjacency matrix) - are theoretically equivalent in expressive power, with expressivity between the 1-WL and 3-WL tests. However, this equivalence assumes the GNN uses the "complete" version of these PEs, which requires $O(n^3)$ time and space complexity. Instead, practitioners commonly use truncated variants of these encodings, such as the first $k$ eigenspaces or powers of the adjacency matrix. However, the theoretical properties of these truncated PEs are unknown. In this work, we initiate the study of these truncated PEs. Theoretically, we show that, under truncation, several families of PEs are fundamentally different in expressive power. As a corollary, we show that truncated spectral PEs are no longer stronger than the 1-WL test. We also study a family of spectral PEs, the $k$-harmonic distances, to highlight the differences in expressive power of even closely related truncated PEs. Finally, we experimentally show that a mix of truncated PEs is preferable to any single family on real-world datasets.
△ Less
Submitted 11 June, 2026;
originally announced June 2026.
-
Otters++: A Time-to-first-spike Based Energy Efficient Optical Spiking Transformer
Authors:
Zhanglu Yan,
Jiayi Mao,
Kaiwen Tang,
Fanfan Li,
Gang Pan,
Tao Luo,
Bowen Zhu,
Qianhui Liu,
Weng-Fai Wong
Abstract:
Spiking neural networks (SNNs) are promising for energy-efficient inference, and time-to-first-spike (TTFS) coding is especially attractive because each neuron fires at most once. In practice, however, this benefit is often reduced by the cost of computing a temporal decay term and multiplying it by the synaptic weight. We address this issue by turning a physical hardware "bug," the natural signal…
▽ More
Spiking neural networks (SNNs) are promising for energy-efficient inference, and time-to-first-spike (TTFS) coding is especially attractive because each neuron fires at most once. In practice, however, this benefit is often reduced by the cost of computing a temporal decay term and multiplying it by the synaptic weight. We address this issue by turning a physical hardware "bug," the natural signal decay in optoelectronic devices, into the main computation of TTFS, named Otters++. Specifically, we use the measured decay of a custom In$_2$O$_3$ optoelectronic synapse to directly realize the TTFS temporal term, removing the need for explicit digital decay computation. To scale this idea to Transformer models, we establish a layer-wise functional equivalence between the Otters++ and a quantized neural network (QNN), and develop a hybrid training method that uses device-faithful SNN computation in the forward pass and QNN straight-through gradients through the equivalent QNN path in the backward pass, together with model distillation. This avoids differentiation through discrete first-spike events and reduces the over-sparsity problem in direct TTFS-SNN training. We further make training aware of measured device noise by sampling run-to-run variation, and refine the system-level energy model by accounting for device sharing and multi-hop communication. On GLUE dataset, Otters++ improves the average score to 84.17\% while maintaining a clear energy advantage over prior spiking Transformer baselines. These results show that physically grounded TTFS computing can be efficient, trainable, and robust under realistic hardware effects.
△ Less
Submitted 11 June, 2026;
originally announced June 2026.
-
Do VLMs See What Sensors Feel? A Scalable Expert-Guided Design for Wheelchair Accessibility Assessment from Street View
Authors:
Dongdong Wang,
Alina Hagen,
Isabelle Gatmaitan,
Hao Zhou,
Yiwen Dong,
Shabboo Valipoor,
Vivian W. H. Wong,
Lingyao Li
Abstract:
Assessing built-environment interaction, such as wheelchair accessibility, is difficult because real-world mobility is shaped by distributed, context-dependent, and temporary barriers that are hard to capture at scale. To support scalable assessment, this paper examines whether vision-language models (VLMs) can identify accessibility barriers from Google Street View (GSV) imagery. We propose an ex…
▽ More
Assessing built-environment interaction, such as wheelchair accessibility, is difficult because real-world mobility is shaped by distributed, context-dependent, and temporary barriers that are hard to capture at scale. To support scalable assessment, this paper examines whether vision-language models (VLMs) can identify accessibility barriers from Google Street View (GSV) imagery. We propose an expert-guided retrieval-augmented framework that combines GSV images, ADA-informed guidance, and expert-derived rubrics to evaluate accessibility dimensions. We collect a campus-scale dataset at the University of Florida, linking 407 unique GSV locations with GPS-derived wheelchair dwell behavior as a mobility-friction signal. Results show that VLM ratings are both negatively correlated and distributionally similar with dwell time, indicating partial but consistent alignment with a behavioral proxy for mobility friction. Visual cue analysis shows that certain environmental objects, such as curb ramps and crosswalks, are associated with higher VLM accessibility scores, while alignment remains limited for subtle surface conditions, transient obstructions, and viewpoint-dependent barriers. Overall, our findings show the potential of expert-guided VLMs for scalable accessibility assessment aligning with sensor-derived indicators of real-world wheelchair navigation.
△ Less
Submitted 1 June, 2026;
originally announced June 2026.
-
SePO: Self-Evolving Prompt Agent for System Prompt Optimization
Authors:
Wangcheng Tao,
Han Wu,
Weng-Fai Wong
Abstract:
System prompt optimization improves agent behavior without modifying the underlying model, yielding human-readable, model-agnostic instructions. Existing methods build a prompt agent that refines task agents' system prompts, yet leave the prompt agent's own system prompt hand-engineered and fixed. We propose Self-Evolving Prompt Optimization (SePO), which treats the prompt agent's own system promp…
▽ More
System prompt optimization improves agent behavior without modifying the underlying model, yielding human-readable, model-agnostic instructions. Existing methods build a prompt agent that refines task agents' system prompts, yet leave the prompt agent's own system prompt hand-engineered and fixed. We propose Self-Evolving Prompt Optimization (SePO), which treats the prompt agent's own system prompt as an optimization target alongside task agents' system prompts. SePO adopts a self-referential design. A single prompt agent improves both task agents' system prompts and its own under an open-ended evolutionary search that maintains an archive of candidate prompts as stepping stones. Training proceeds in two stages: pre-training evolves the prompt agent on a multi-task pool, and fine-tuning then applies it to a target task. Across five benchmarks spanning math (AIME'25), abstract reasoning (ARC-AGI-1), graduate-level science (GPQA), code generation (MBPP), and logic puzzles (Sudoku), SePO consistently outperforms Manual-CoT, TextGrad, and MetaSPO, improving the average accuracy by 4.49 points compared to Manual-CoT. The prompt optimization skill from pre-training also generalizes to tasks beyond the pre-training mixture, rather than memorizing per-task prompts.
△ Less
Submitted 3 June, 2026;
originally announced June 2026.
-
Low-rank Distributional Matrix Completion
Authors:
Jiayi Wang,
Raymond K. W. Wong
Abstract:
We study a distributional generalization of the matrix completion problem in which each entry of the target matrix is a probability distribution rather than a scalar. In this setting, only a subset of matrix entries is observed, and even for observed entries, the underlying distributions are not directly accessible; instead, we observe finitely many samples drawn from them. To represent distributi…
▽ More
We study a distributional generalization of the matrix completion problem in which each entry of the target matrix is a probability distribution rather than a scalar. In this setting, only a subset of matrix entries is observed, and even for observed entries, the underlying distributions are not directly accessible; instead, we observe finitely many samples drawn from them. To represent distributional entries, we employ kernel mean embeddings and introduce a notion of Tucker rank for distribution-valued matrices to capture their low-rank structure. The infinite-dimensional nature of kernel embeddings poses significant methodological challenges. To address this, we introduce functional unfolding operators that link the proposed distributional low-rank structure to the classical Tucker rank for finite-dimensional tensors. Based on this framework, we propose a novel estimator for distributional matrix completion. We establish non-asymptotic error bounds that characterize the statistical performance of the estimator. Extensive experiments on synthetic data and a real-world application demonstrate the effectiveness of the proposed method.
△ Less
Submitted 2 June, 2026;
originally announced June 2026.
-
Revisiting Neural Processes via Fourier Transform and Volterra Series
Authors:
Peiman Mohseni,
Nick Duffield,
Raymond K. W. Wong
Abstract:
Modeling unknown latent functions from finite, irregularly sampled measurements is a recurring challenge across science and engineering. Neural processes (NPs), a family of probabilistic functional models, are promising solutions -- especially when endowed with domain-specific symmetries like translation equivariance, which improve sample efficiency and generalization. Yet existing translation-equ…
▽ More
Modeling unknown latent functions from finite, irregularly sampled measurements is a recurring challenge across science and engineering. Neural processes (NPs), a family of probabilistic functional models, are promising solutions -- especially when endowed with domain-specific symmetries like translation equivariance, which improve sample efficiency and generalization. Yet existing translation-equivariant NPs face two limitations: (i) they stack generic components with non-linearities, obscuring the induced function class and limiting interpretability; and (ii) convolutional designs are limited by local receptive fields and the need to embed inputs onto a dense uniform grid, while attention-based alternatives lift these restrictions at quadratic cost in the number of observations. We address both with two contributions. First, using the Volterra expansion, we approximate continuous translation-equivariant operators by sums of higher-order convolutions, yielding analytical transparency while admitting efficient evaluation via first-order convolutions. Second, we introduce set Fourier convolutions (SFConvs), a frequency-domain parameterization that operates directly on irregularly sampled points, achieves approximately global receptive fields, and scales linearly in the number of observations. Building on these ideas, we propose two conditional NPs (CNPs): SFConvCNPs, which stack SFConv blocks with non-linearities, and SFVConvCNPs, which integrate the Volterra formulation. Experiments on synthetic and real-world datasets demonstrate our methods' efficacy against state-of-the-art baselines.
△ Less
Submitted 13 July, 2026; v1 submitted 31 May, 2026;
originally announced June 2026.
-
Density-aware Sample-specific Attack
Authors:
Qiyuan Wang,
Yao Li,
Raymond K. W. Wong
Abstract:
Despite recent progress in backdoor attacks, existing methods remain susceptible to post-training defenses that erase the backdoor through fine-tuning or pruning. We revisit the core objectives of backdoor attacks and derive principled criteria characterizing optimal sample-specific trigger construction under a Bayes-optimal model of the victim's training. Our analysis reveals that both attack suc…
▽ More
Despite recent progress in backdoor attacks, existing methods remain susceptible to post-training defenses that erase the backdoor through fine-tuning or pruning. We revisit the core objectives of backdoor attacks and derive principled criteria characterizing optimal sample-specific trigger construction under a Bayes-optimal model of the victim's training. Our analysis reveals that both attack success and clean-accuracy preservation are simultaneously optimized when triggered samples are steered into low-density regions of the clean data distribution, a distributional condition that controls all moments of the poisoned distribution at once rather than a handful of input-space summary statistics. We introduce a bilevel optimization framework that estimates density ratios via conditional time-score matching and optimizes a mixture-model objective to place triggered samples in these sparse regions. Extensive evaluations on MNIST, CIFAR-10, GTSRB, and TinyImageNet demonstrate that our method achieves above 99\% attack success rate before defense and retains 50--85 percentage points higher post-defense ASR than the strongest baselines under fine-tuning defenses. Against neuron-pruning defenses, the method exhibits complete immunity, with zero neurons identified for removal across all pruning thresholds. These results expose a fundamental gap in current defense paradigms and underscore the need for defenses that operate beyond the support of the clean distribution.
△ Less
Submitted 28 May, 2026; v1 submitted 26 May, 2026;
originally announced May 2026.
-
Magneto-optic phonon resonances in magnetic topological EuCd2As2 via helical Raman spectroscopy
Authors:
Jin Ho Kang,
Liangbo Liang,
Ioannis Petrides,
Subhajit Roychowdhury,
Kai-Chi Chang,
Chandra Shekhar,
Claudia Felser,
Prineha Narang,
Chee Wei Wong
Abstract:
EuCd2As2 materials have two magnetic ordering states: antiferromagnetic (AFM) and ferromagnetic (FM) when their chemical tunability is utilized. While AFM-EuCd2As2 has a nonzero magnetoelectric response due to its symmetry breaking with spin configuration, FM-EuCd2As2 is an ideal candidate for studies of Weyl physics because of its minimum number of Weyl points with opposite chirality. In this art…
▽ More
EuCd2As2 materials have two magnetic ordering states: antiferromagnetic (AFM) and ferromagnetic (FM) when their chemical tunability is utilized. While AFM-EuCd2As2 has a nonzero magnetoelectric response due to its symmetry breaking with spin configuration, FM-EuCd2As2 is an ideal candidate for studies of Weyl physics because of its minimum number of Weyl points with opposite chirality. In this article, we examine cryogenic low-frequency Raman spectroscopy of phonon modes in FM-EuCd2As2 crystals using circular polarization configurations, with support from density functional theory calculations, and investigate in-plane magneto-anisotropy by linear polarization configuration below the Curie temperature (Tc = 26 K). We attribute the anomalous enhancements in Raman intensities below the Curie temperature are due to spin-phonon coupling. Furthermore, we see that A-mode peaks can be distinguished by magneto-helical Raman spectroscopy through the magneto-optic effect and that the degree of circular polarization (DCP) of 12.5 meV peak reaches 60% at 4.2 K and becomes saturated. We also examine AFM-EuCd2As2 below Néel temperature (TN = 9 K) to compare with FM-EuCd2As2, but we hardly observe spin-phonon coupling and find negligible DCP values due to almost zero net magnetization. Our results contribute to the understanding of the phonon dynamics and the interplay between topology and magnetism in FM-EuCd2As2, through helical light and external magnetic fields. This lays the foundation for utilizing state-of-the-art Weyl systems for applications in thermoelectrics, phononic devices, and topological quantum computing.
△ Less
Submitted 26 May, 2026; v1 submitted 25 May, 2026;
originally announced May 2026.
-
Clustering based on Stochastic Dominance with application for risk averters and risk seekers
Authors:
Hua Li,
Xue Jia,
Yilin Kang,
Wing-Keung Wong
Abstract:
Stochastic Dominance (SD) theory provides a rigorous framework for selecting superior assets tailored to the asset allocation needs of investors with varying risk preferences (i.e., risk-averse, risk-seeking, and risk-neutral). However, traditional stock clustering methods typically rely on geometric metrics such as Euclidean distance, which often fail to effectively capture the intrinsic risk dom…
▽ More
Stochastic Dominance (SD) theory provides a rigorous framework for selecting superior assets tailored to the asset allocation needs of investors with varying risk preferences (i.e., risk-averse, risk-seeking, and risk-neutral). However, traditional stock clustering methods typically rely on geometric metrics such as Euclidean distance, which often fail to effectively capture the intrinsic risk dominance relationships among assets. To address this limitation, this paper proposes an innovative clustering analysis framework based on SD test statistics. Methodologically, this study deeply integrates SD theory with machine learning algorithms. Transcending the limitations of traditional reliance on geometric distance, we innovatively utilize test statistics from first-, second-, and third-order SD to construct a "Stochastic Dominance Coefficient Matrix." Building upon this matrix, we modify the classic K-means and Hierarchical Clustering algorithms. Specifically, we derive 12 distinct algorithm variants tailored to different orders of SD relationships. Simultaneously, we construct the SD-SC coefficient and the SD-DBI index as specialized validity indices to evaluate the clustering performance. Empirically, we analyze constituent stock data from a representative developed market (the US NASDAQ Index) and an emerging market (China's CSI 100 Index). The results verify the effectiveness and robustness of the proposed method. Furthermore, we apply the clustering results to the modification of the Single Index Model and the construction of Global Minimum Variance Portfolios (GMVP). The findings demonstrate that the proposed method effectively facilitates customized asset allocation for investors, holding significant theoretical value and practical implications.
△ Less
Submitted 23 May, 2026;
originally announced May 2026.
-
Is Agentic AI Ready for Real-World Hardware Engineering? A Deep Dive with Phoenix-bench
Authors:
Qingyun Zou,
Feng Yu,
Hongshi Tan,
Bingsheng He,
WengFai Wong
Abstract:
We ask whether agentic AI systems built for software engineering transfer to realistic hardware engineering. Existing hardware LLM benchmarks isolate sub-tasks but none jointly requires repository navigation, hierarchy-aware localization, Electronic Design Automation (EDA) executable verification, and maintenance-style patching. We introduce \textbf{Phoenix-bench}, a synchronized corpus of 511 ver…
▽ More
We ask whether agentic AI systems built for software engineering transfer to realistic hardware engineering. Existing hardware LLM benchmarks isolate sub-tasks but none jointly requires repository navigation, hierarchy-aware localization, Electronic Design Automation (EDA) executable verification, and maintenance-style patching. We introduce \textbf{Phoenix-bench}, a synchronized corpus of 511 verified Verilator instances from 114 GitHub repositories, each shipped with the developer patch, design-flow labels, fail-to-pass and pass-to-pass testbenches, and a Docker-pinned EDA environment so resolved-rate differences reflect agent behavior rather than toolchain availability. Using Phoenix-bench we run a uniform evaluation of four commercial agents and eight open-source agentic structures across four LLM backbones, plus two diagnostic interventions (file-level oracle localization and one round of testbench-log feedback). Three findings emerge. (i)~Software and hardware are fundamentally different engineering tasks: the same agent loses 37\% to 58\% from SWE-bench Verified to Phoenix-bench because hardware bugs propagate across parallel instantiated modules through signal flow rather than along a software-style call graph, and software-tuned agents stop at the symptom file instead of tracing back through the instantiation chain. (ii)~Failures concentrate on design control-flow / finite state machine (FSM) bugs, verification testbench bugs, and hard cases that demand cross-hierarchy signal-flow tracking and coordinated multi-file edits. (iii)~Localization granularity matters far more than localization itself: a perfect file-level oracle yields only $+1.4$\% because the agent then breaks files that did not need editing, while a single round of test case feedback lifts resolved rate by $42$\% to $45$\% because the test case tells \emph{where} the bug is and \emph{what} the fix has to look like.
△ Less
Submitted 13 May, 2026;
originally announced May 2026.
-
Analog RF Computing: A New Paradigm for Energy-Efficient Edge AI Over MU-MIMO Systems
Authors:
Wentao Yu,
Vincent W. S. Wong
Abstract:
Modern edge devices increasingly rely on neural networks for intelligent applications. However, conventional digital computing-based edge inference requires substantial memory and energy consumption. In analog radio frequency (RF) computing, a base station (BS) encodes the weights of the neural networks and broadcasts the RF waveforms to the clients. Each client reuses its passive mixer to multipl…
▽ More
Modern edge devices increasingly rely on neural networks for intelligent applications. However, conventional digital computing-based edge inference requires substantial memory and energy consumption. In analog radio frequency (RF) computing, a base station (BS) encodes the weights of the neural networks and broadcasts the RF waveforms to the clients. Each client reuses its passive mixer to multiply the received weight-encoded waveform with a locally generated input-encoded waveform. This enables wireless receivers to perform the matrix-vector multiplications (MVMs) that account for most of the computation burden in edge inference with ultra-low energy consumption. Unlike conventional downlink transmissions which are optimized for communications, analog RF computing requires a computing-centric physical layer that controls both the analog MVM accuracy and the energy consumption for inference. Motivated by this, in this paper, we propose a physical layer design framework for analog RF computing in MU-MIMO wireless systems. We derive tractable models for computing accuracy and energy consumption for inference, formulate a joint BS beamforming and client-side scaling problem subject to computing accuracy, transmit power, and hardware constraints, and develop a low-complexity algorithm to solve the non-convex problem. The proposed design provides client- and layer-specific accuracy control for both uniform- and mixed-precision inference. Simulations under 3GPP specifications show that analog RF computing can significantly reduce client-side energy consumption by nearly two orders of magnitude compared to digital computing, while mixed-precision inference requires even lower energy consumption than uniform-precision inference. Overall, these results establish analog RF computing over wireless networks as a promising paradigm for energy-efficient edge inference.
△ Less
Submitted 13 May, 2026;
originally announced May 2026.
-
HLS-Seek: QoR-Aware Code Generation for High-Level Synthesis via Proxy Comparative Reward Reinforcement Learning
Authors:
Qingyun Zou,
Feng Yu,
Hongshi Tan,
Yao Chen,
Bingsheng He,
WengFai Wong
Abstract:
High-Level Synthesis (HLS) compiles algorithmic C/C++ descriptions into hardware, with Quality of Results (QoR) -- latency and resource utilization -- critically governed by pragma configurations and code structure. Existing LLM-based HLS approaches train for functional correctness but ignore QoR entirely. We observe that reinforcement learning (RL) for HLS does not require absolute synthesis resu…
▽ More
High-Level Synthesis (HLS) compiles algorithmic C/C++ descriptions into hardware, with Quality of Results (QoR) -- latency and resource utilization -- critically governed by pragma configurations and code structure. Existing LLM-based HLS approaches train for functional correctness but ignore QoR entirely. We observe that reinforcement learning (RL) for HLS does not require absolute synthesis results -- only relative comparisons between candidates. Based on this insight, we propose \textbf{HLS-Seek}, a QoR-aware NL-to-HLS framework that replaces expensive synthesis-in-the-loop RL with a comparative proxy reward model achieving 99.53\% Pareto-dominance accuracy. To prevent reward hacking, we introduce \textit{uncertainty-aware Monte Carlo (MC) dropout switching} that selectively invokes real Vitis HLS synthesis for low-confidence candidates and online updates the proxy, creating a self-improving reward system. HLS-Seek achieves 81.5\% syntax correctness pass@1 and 81.4\% Func@5 on HLS-eval with only 7B parameters, surpassing GPT-5.1 and other frontier models while achieving 8.5$\times$ faster training than real-reward RL. On QoR evaluation, HLS-Seek achieves the lowest latency on 16/30 kernels and Pareto-dominates HLS-specific baselines on 9 kernels.
△ Less
Submitted 13 May, 2026;
originally announced May 2026.
-
Reward-Weighted On-Policy Distillation with an Open Property-Equivalence Verifier for NL-to-SVA Generation
Authors:
Qingyun Zou,
Yingze Li,
Tianen Liu,
Bingsheng He,
Weng-Fai Wong
Abstract:
LLM-based generation of SystemVerilog Assertions (SVA) is often reported as nearing saturation, with the strongest specialized model reaching ${\sim}76\%$ accuracy on NL2SVA-Human. We show that this aggregate hides a temporal gap: models that appear strong overall still collapse to a few implication templates on bounded-delay and liveness specifications. The core issue is that the dominant recipe,…
▽ More
LLM-based generation of SystemVerilog Assertions (SVA) is often reported as nearing saturation, with the strongest specialized model reaching ${\sim}76\%$ accuracy on NL2SVA-Human. We show that this aggregate hides a temporal gap: models that appear strong overall still collapse to a few implication templates on bounded-delay and liveness specifications. The core issue is that the dominant recipe, supervised fine-tuning on NL/SVA pairs, optimizes token-level mimicry rather than the \emph{property equivalence} that defines SVA correctness. We introduce \emph{Reward-Weighted On-Policy Distillation} (RWOPD), an on-policy distillation method that samples student rollouts, scores them with an open SymbiYosys+Z3 Property-Equivalence Checker (PEC), and applies a verifier-reward-weighted forward-KL gradient from a frozen 14B teacher on verifier-passable rollouts. This keeps the supervision dense at every response token while grounding both selection and loss weight in property-equivalent behavior. RWOPD distills CodeV-SVA-14B into a Qwen2.5-Coder-7B-Instruct student that sets a new state of the art on NL2SVA-Human and NL2SVA-Machine across pass@1, pass@5, and pass@10, surpassing both specialized prior SOTA models and 671B general-purpose baselines.
△ Less
Submitted 13 May, 2026;
originally announced May 2026.
-
Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images
Authors:
Yuangong Chen,
Wai Keung Wong,
Jiaxing Li,
Ioannis Patras,
Xu Zheng
Abstract:
Multimodal Large Language Models (MLLMs) show strong visual perception, yet remain limited in reasoning about space under changing viewpoints. We study this challenge as Perspective-Conditioned Spatial Reasoning (PCSR) in 360-degree omnidirectional images, where broad scene coverage reduces ambiguity from partial observations without eliminating the need for viewpoint-dependent inference. To asses…
▽ More
Multimodal Large Language Models (MLLMs) show strong visual perception, yet remain limited in reasoning about space under changing viewpoints. We study this challenge as Perspective-Conditioned Spatial Reasoning (PCSR) in 360-degree omnidirectional images, where broad scene coverage reduces ambiguity from partial observations without eliminating the need for viewpoint-dependent inference. To assess this capability, we introduce PCSR-Bench, a diagnostic benchmark of 84,373 question-answer pairs from 2,600 omnidirectional images across 26 indoor environments. PCSR-Bench contains eight tasks spanning foundational perception (e.g., object counting, relative distance, and relative direction) and advanced PCSR, including compositional chains, egocentric rotation, perspective re-anchoring, ego-distortion, and limited-FOV visibility. We evaluate 14 representative MLLMs and observe a substantial perception-reasoning gap: accuracy reaches 57.59% on foundational relative direction, but drops to 13.49% on egocentric rotation, 7.13% on egocentric distortion, and 0.64% on open-ended compositional reasoning. To probe the plasticity of this gap, we conduct an RL-based diagnostic study on a 7B-scale model. Reward shaping improves a matched 7B baseline from 31.10% to 60.06% under a controlled setting, suggesting that PCSR is partial plasticity rather than being fully immutable. Still, the gains are task-selective, sensitive to reward design including both weight allocation and reward formulation, and partially dependent on the evaluation protocol. These results position PCSR as a key bottleneck in current MLLMs and highlight limited but meaningful room for recovery under targeted optimization.
△ Less
Submitted 18 May, 2026; v1 submitted 12 May, 2026;
originally announced May 2026.
-
Graduate Training in Quantum Information Science and Engineering: Lessons, Challenges, and a Roadmap from the NSF Research Traineeship Programs
Authors:
Yohannes Abate,
Victor Acosta,
Alessandro Alabastri,
Mehmet Aydeniz,
Viktoriia E. Babicheva,
Lincoln D. Carr,
I-Tung Chen,
Wandi Ding,
Tara Drake,
Mattias Fitzpatrick,
Kai-Mei C. Fu,
Jay Gupta,
Kaden R. A. Hazzard,
Sophia E. Hayes,
Jin Hu,
Hilary M. Hurst,
Sohrab Ismail-Beigi,
Ehsan Khatami,
Junichiro Kono,
Cheng-Yu Lai,
Xiuling Li,
Yingmei Liu,
Sara Mouradian,
Kater Murch,
Borja Peropadre
, et al. (9 additional authors not shown)
Abstract:
Since 2019, eighteen NSF Research Traineeship (NRT) awards in quantum information science and engineering (QISE) and adjacent fields have been funded, constituting the largest NSF-coordinated investment in graduate QISE training in the United States. Synthesizing lessons from our programs, we work through the central tensions that every QISE graduate program must negotiate: between depth in a home…
▽ More
Since 2019, eighteen NSF Research Traineeship (NRT) awards in quantum information science and engineering (QISE) and adjacent fields have been funded, constituting the largest NSF-coordinated investment in graduate QISE training in the United States. Synthesizing lessons from our programs, we work through the central tensions that every QISE graduate program must negotiate: between depth in a home discipline and breadth across the field, between structured instruction and open-ended experiential and hands-on learning, and between training individual specialists and cultivating teams that collectively cover all areas of QISE. We describe the structural and pedagogical innovations the NRT programs have developed in response, assess what is working and what remains unresolved, and sketch 12 open problems the community will need to address as QISE graduate education scales beyond the well-resourced research universities where it has up till now been mainly concentrated. Eight concrete recommendations follow: (1) adopt the startup model of team-based training as an organizing philosophy; (2) invest immediately in sensing and communication curriculum development; (3) build student agency into program governance, not just activities; (4) establish structural mechanisms for industrial engagement rather than depending on goodwill; (5) design for sustainability from year one; (6) develop graduate-level textbooks spanning all three QISE pillars: computing, sensing, and communications; (7) establish shared outcome assessment instruments across programs; and (8) develop structured mechanisms for faculty professional development in QISE.
△ Less
Submitted 8 May, 2026;
originally announced May 2026.
-
Computer Use at the Edge of the Statistical Precipice
Authors:
Pierluca D'Oro,
Sneha Silwal,
William Wong,
Yuxuan Sun,
Fanyi Xiao,
Manchen Wang,
Eric Gan,
Allen Bolourchi,
Joseph Tighe
Abstract:
Evaluating Computer Use Agents (CUAs) on interactive environments is fraught with methodological pitfalls that the field has yet to systematically address. We show that a 1MB replay script that blindly executes a recorded action sequence without ever observing the screen outperforms frontier models on prominent static benchmarks, and prove that its expected success rate is exactly equal to the sou…
▽ More
Evaluating Computer Use Agents (CUAs) on interactive environments is fraught with methodological pitfalls that the field has yet to systematically address. We show that a 1MB replay script that blindly executes a recorded action sequence without ever observing the screen outperforms frontier models on prominent static benchmarks, and prove that its expected success rate is exactly equal to the source agent's pass@k in deterministic environments. We trace this and other failures to two root causes: non-principled environment design (static, unsandboxed, or unreliably verified environments) and non-principled evaluation methodology (naive aggregation and misuse of pass@k for stateful UI interactions). To address the first, we propose PRISM, five design principles for CUA environments (privileged verification, realistic environments, integrity-checked configurations, sandboxed execution, and multifactorial variability) and instantiate them in DigiWorld, a benchmark of 15 realistic sandboxed mobile applications able to evaluate agents in over 3.2 million verified unique configurations. To address the second, we develop an aggregation framework pairing Wilson score intervals with hierarchical bootstrap, producing confidence intervals that correctly account for the nested structure of CUA benchmarks, as we empirically demonstrate. All together, we show that principled environment design and rigorous evaluation methodology are not optional refinements but prerequisites for meaningful CUA research.
△ Less
Submitted 7 May, 2026;
originally announced May 2026.
-
XtraMAC: An Efficient MAC Architecture for Mixed-Precision LLM Inference on FPGA
Authors:
Feng Yu,
Hongshi Tan,
Yao Chen,
Weng-Fai Wong,
Bingsheng He
Abstract:
The widespread adoption of mixed-precision quantization in large language models (LLMs) has created demand for hardware that can efficiently perform multiply-accumulate (MAC) operations across mixed datatypes and switch datatypes at runtime. Existing FPGA-based MAC solutions fall short due to limitations in fixed-datatype design, inefficient spatial or temporal resource sharing, and poor support f…
▽ More
The widespread adoption of mixed-precision quantization in large language models (LLMs) has created demand for hardware that can efficiently perform multiply-accumulate (MAC) operations across mixed datatypes and switch datatypes at runtime. Existing FPGA-based MAC solutions fall short due to limitations in fixed-datatype design, inefficient spatial or temporal resource sharing, and poor support for mixed-precision execution. These limitations collectively lead to under-utilization of DSP resources, limiting achievable parallelism and throughput. In this work, we present XtraMAC, a novel MAC architecture that unifies integer, floating-point, and mixed-precision operations within a single, datatype-adaptive microarchitecture. XtraMAC decomposes all supported MAC formats into a shared integer mantissa product with lightweight sign and exponent handling, enabling dynamic operand packing and efficient DSP resource sharing with constant latency and initiation interval of one across all datatypes. Evaluated on an AMD Xilinx U55c FPGA, XtraMAC achieves 1.4-2.0x higher compute density, reduces per-operation LUT, FF, and DSP consumption by 27-51%, and delivers up to 1.9x greater energy efficiency and 1.2x speedup on representative mixed-precision LLM workloads. The implementation of XtraMAC is open-sourced at https://github.com/Xtra-Computing/XtraMAC.
△ Less
Submitted 7 May, 2026;
originally announced May 2026.
-
ShiftLIF: Efficient Multi-Level Spiking Neurons with Power-of-Two Quantization
Authors:
Kaiwen Tang,
Di Yu,
Jiaqi Zheng,
Changze Lv,
Qianhui Liu,
Zhanglu Yan,
Weng-Fai Wong
Abstract:
Spiking neural networks (SNNs) are promising for edge sensing due to their event-driven computation and temporal filtering capability. However, standard leaky integrate-and-fire (LIF) neurons communicate only through binary spikes, which severely limit representational capacity. Existing multi-level spiking neurons improve information transmission, but often rely on uniform quantization that misma…
▽ More
Spiking neural networks (SNNs) are promising for edge sensing due to their event-driven computation and temporal filtering capability. However, standard leaky integrate-and-fire (LIF) neurons communicate only through binary spikes, which severely limit representational capacity. Existing multi-level spiking neurons improve information transmission, but often rely on uniform quantization that mismatches membrane-potential distributions or introduces costly synaptic multiplications. In this paper, we propose ShiftLIF, a multi-level spiking neuron that maps membrane potentials to a logarithmically spaced power-of-two spike set. This design provides finer representation in the small-amplitude regime, where membrane potentials are densely concentrated, while enabling multiplier-free synaptic computation through bit-shift and accumulation operations. As a result, ShiftLIF improves spike-level expressiveness without sacrificing the hardware-friendly nature of standard SNN computation. We evaluate ShiftLIF on 10 datasets spanning wireless, acoustic, motion, and visual sensing tasks. Results show that ShiftLIF consistently matches or exceeds the accuracy of existing multi-level spiking neurons while maintaining synaptic energy consumption close to standard binary LIF. These results indicate that ShiftLIF provides a favorable accuracy-efficiency trade-off for cross-modal edge sensing.
△ Less
Submitted 3 May, 2026;
originally announced May 2026.
-
LatentDiff: Scaling Semantic Dataset Comparison to Millions of Images
Authors:
James Flora,
Kowshik Thopalli,
Akshay R. Kulkarni,
Weng-Keen Wong,
Shusen Liu
Abstract:
We present LatentDiff, a scalable framework for semantic dataset comparison that operates directly in the latent space of pretrained vision encoders. By combining sparse autoencoder-based divergence testing with density ratio estimation, LatentDiff identifies interpretable semantic differences between datasets at a fraction of the computational cost of caption-based alternatives. We also introduce…
▽ More
We present LatentDiff, a scalable framework for semantic dataset comparison that operates directly in the latent space of pretrained vision encoders. By combining sparse autoencoder-based divergence testing with density ratio estimation, LatentDiff identifies interpretable semantic differences between datasets at a fraction of the computational cost of caption-based alternatives. We also introduce Noisy-Diff, a benchmark capturing realistic sparse distribution shifts that cause existing methods to struggle. Experiments demonstrate that LatentDiff achieves superior accuracy while remaining robust to settings where an extremely small fraction of images (from 5% to <1% ) differ semantically.
△ Less
Submitted 28 April, 2026;
originally announced May 2026.
-
A Multimodal Clinically Informed Coarse-to-Fine Framework for Longitudinal CT Registration in Proton Therapy
Authors:
Caiwen Jiang,
Yuzhen Ding,
Mi Jia,
Samir H. Patel,
Terence T. Sio,
Jonathan B. Ashman,
Lisa A. McGee,
Jean-Claude M. Rwigema,
William G. Rule,
Sameer R. Keole,
Sujay A. Vora,
William W. Wong,
Nathan Y. Yu,
Michele Y. Halyard,
Steven E. Schild,
Dinggang Shen,
Wei Liu
Abstract:
Proton therapy offers superior organ-at-risk sparing but is highly sensitive to anatomical changes, making accurate deformable image registration (DIR) across longitudinal CT scans essential. Conventional DIR methods are often too slow for emerging online adaptive workflows, while existing deep learning-based approaches are primarily designed for generic benchmarks and underutilize clinically rele…
▽ More
Proton therapy offers superior organ-at-risk sparing but is highly sensitive to anatomical changes, making accurate deformable image registration (DIR) across longitudinal CT scans essential. Conventional DIR methods are often too slow for emerging online adaptive workflows, while existing deep learning-based approaches are primarily designed for generic benchmarks and underutilize clinically relevant information beyond images. To address this gap, we propose a clinically scalable coarse-to-fine deformable registration framework that integrates multimodal information from the proton radiotherapy workflow to accommodate diverse clinical scenarios. The model employs dual CNN-based encoders for hierarchical feature extraction and a transformer-based decoder to progressively refine deformation fields. Beyond CT intensities, clinically critical priors, including target and organ-at-risk contours, dose distributions, and treatment planning text, are incorporated through anatomy- and risk-guided attention, text-conditioned feature modulation, and foreground-aware optimization, enabling anatomically focused and clinically informed deformation estimation. We evaluate the proposed framework on a large-scale proton therapy DIR dataset comprising 1,222 paired planning and repeat CT scans across multiple anatomical regions and disease types. Extensive experiments demonstrate consistent improvements over state-of-the-art methods, enabling fast and robust clinically meaningful registration.
△ Less
Submitted 14 April, 2026;
originally announced April 2026.
-
Closing the Loop in Epitaxy with Machine Learning: Joint Optimization of Growth and Geometry in On-Chip Lasers
Authors:
Mihir R. Athavale,
Stephen A. Church,
Wei Wen Wong,
Andre KY Low,
Hark Hoe Tan,
Kedar Hippalgaonkar,
Patrick Parkinson
Abstract:
Achieving device-to-device reproducibility is a critical bottleneck for scalable photonic integrated circuits, as subtle variations in bottom-up epitaxial growth and fabrication severely limit yield. We present a machine learning workflow for III-V multi-quantum well microring lasers that first optimizes growth and geometry parameters via multi-objective Bayesian optimization, then leverages varia…
▽ More
Achieving device-to-device reproducibility is a critical bottleneck for scalable photonic integrated circuits, as subtle variations in bottom-up epitaxial growth and fabrication severely limit yield. We present a machine learning workflow for III-V multi-quantum well microring lasers that first optimizes growth and geometry parameters via multi-objective Bayesian optimization, then leverages variational autoencoders (VAEs) to attribute residual device-to-device variability to its underlying sources. By explicitly targeting threshold variance alongside absolute performance, we demonstrate 100% lasing yield across all designs. The optimized multi-quantum well microring laser fields achieved a median lasing threshold of $16~μ\mathrm{J}\,\mathrm{cm}^{-2}\,\mathrm{pulse}^{-1}$, a $73\%$ reduction in threshold variance relative to the previously reported best values, and a median emission wavelength of $1333~\mathrm{nm}$, in the telecommunications O-band. Furthermore, to diagnose residual performance dispersion under nominally identical conditions, VAEs were used to isolate the key components of device morphology that impact performance. This analysis successfully decoupled geometric from material disorder, quantitatively linking previously unmeasured morphological variations to population-level threshold fluctuations. This data-driven workflow bridges the gap between fundamental epitaxy and reliable manufacturing, establishing a generalizable blueprint for designing and yield-optimizing complex, non-linear optoelectronic devices.
△ Less
Submitted 9 April, 2026;
originally announced April 2026.
-
Incremental GNN Embedding Computation on Streaming Graphs
Authors:
Qiange Wang,
Haoran Lv,
Yanfeng Zhang,
Weng-Fai Wong,
Bingsheng He
Abstract:
Graph Neural Network (GNN) on streaming graphs has gained increasing popularity. However, its practical deployment remains challenging, as the inference process relies on Runtime Embedding Computation (RTEC) to capture recent graph changes. This process incurs heavyweight multi-hop graph traversal overhead, which significantly undermines computation efficiency. We observe that the intermediate res…
▽ More
Graph Neural Network (GNN) on streaming graphs has gained increasing popularity. However, its practical deployment remains challenging, as the inference process relies on Runtime Embedding Computation (RTEC) to capture recent graph changes. This process incurs heavyweight multi-hop graph traversal overhead, which significantly undermines computation efficiency. We observe that the intermediate results for large portions of the graph remain unchanged during graph evolution, and thus redundant computations can be effectively eliminated through carefully designed incremental methods. In this work, we propose an efficient framework for incrementalizing RTEC on streaming graphs.The key idea is to decouple GNN computation into a set of generalized, fine-grained operators and safely reorder them, transforming the expensive full-neighbor GNN computation into a more efficient form over the affected subgraph. With this design, our framework preserves the semantics and accuracy of the original full-neighbor computation while supporting a wide range of GNN models with complex message-passing patterns. To further scale to graphs with massive historical results, we develop a GPU-CPU co-processing system that offloads embeddings to CPU memory with communication-optimized scheduling. Experiments across diverse graph sizes and GNN models show that our method reduces computation by 64%-99% and achieves 1.7x-145.8x speedups over existing solutions.
△ Less
Submitted 20 March, 2026;
originally announced March 2026.
-
High-dimensional quantum communication with scalable photonic entanglement in time and frequency
Authors:
Kai-Chi Chang,
Murat Can Sarihan,
Nicky Kai Hong Li,
Florian Kanitschar,
Kemal Enes Akyuz,
Yujie Chen,
Dong-Il Lee,
Jin Ho Kang,
Alwaleed Aldhafeeri,
Andrew Mueller,
Matthew D. Shaw,
Boris Korzh,
Maria Spiropulu,
Paul Erker,
Marcus Huber,
Chee Wei Wong
Abstract:
High-dimensional photonic entanglement holds significant promise for advancing quantum communication, computation, and metrology. For example, large-alphabet quantum communication protocols are known to benefit from enhanced noise resilience and information capacity via multi-bit time-bin encoding. Yet, characterizing high-dimensional entangled states is challenging, as full state tomography becom…
▽ More
High-dimensional photonic entanglement holds significant promise for advancing quantum communication, computation, and metrology. For example, large-alphabet quantum communication protocols are known to benefit from enhanced noise resilience and information capacity via multi-bit time-bin encoding. Yet, characterizing high-dimensional entangled states is challenging, as full state tomography becomes prohibitively costly and often requires unrealizable measurements. Here, we demonstrate a scan-free method to characterize high-dimensional entanglement in the time-frequency domain. Our reconstruction achieves a record $5.70\pm0.07$ ebits and a fidelity of $65.4\pm0.4\%$ with the maximally entangled state of local dimension $1021$, certifying the presence of $668$-dimensional entanglement. We further prove the attainability of a secure key rate of $15.6$ kB/s in a composable finite-size, entanglement-based protocol, and show that in continuous operation, the setup can quickly approach asymptotic key rates. Using commercial telecom components and state-of-the-art low-jitter single-photon detectors, our scalable architecture offers a practical path towards high-rate, noise-resilient quantum communication testbeds.
△ Less
Submitted 18 March, 2026;
originally announced March 2026.
-
Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and Accessories
Authors:
Junyao Hu,
Zhongwei Cheng,
Waikeung Wong,
Xingxing Zou
Abstract:
Virtual try-on (VTON) has advanced single-garment visualization, yet real-world fashion centers on full outfits with multiple garments, accessories, fine-grained categories, layering, and diverse styling, remaining beyond current VTON systems. Existing datasets are category-limited and lack outfit diversity. We introduce Garments2Look, the first large-scale multimodal dataset for outfit-level VTON…
▽ More
Virtual try-on (VTON) has advanced single-garment visualization, yet real-world fashion centers on full outfits with multiple garments, accessories, fine-grained categories, layering, and diverse styling, remaining beyond current VTON systems. Existing datasets are category-limited and lack outfit diversity. We introduce Garments2Look, the first large-scale multimodal dataset for outfit-level VTON, comprising 80K many-garments-to-one-look pairs across 40 major categories and 300+ fine-grained subcategories. Each pair includes an outfit with 3-12 reference garment images (Average 4.48), a model image wearing the outfit, and detailed item and try-on textual annotations. To balance authenticity and diversity, we propose a synthesis pipeline. It involves heuristically constructing outfit lists before generating try-on results, with the entire process subjected to strict automated filtering and human validation to ensure data quality. To probe task difficulty, we adapt SOTA VTON methods and general-purpose image editing models to establish baselines. Results show current methods struggle to try on complete outfits seamlessly and to infer correct layering and styling, leading to misalignment and artifacts.
△ Less
Submitted 14 March, 2026;
originally announced March 2026.
-
Induced Numerical Instability: Hidden Costs in Multimodal Large Language Models
Authors:
Wai Tuck Wong,
Jun Sun,
Arunesh Sinha
Abstract:
The use of multimodal large language models has become widespread, and as such the study of these models and their failure points has become of utmost importance. We study a novel mode of failure that causes degradation in performance indirectly by optimizing a loss term that seeks to maximize numerical instability in the inference stage of these models. We apply this loss term as the optimization…
▽ More
The use of multimodal large language models has become widespread, and as such the study of these models and their failure points has become of utmost importance. We study a novel mode of failure that causes degradation in performance indirectly by optimizing a loss term that seeks to maximize numerical instability in the inference stage of these models. We apply this loss term as the optimization target to construct images that, when used on multimodal large language models, cause significant degradation in the output. We validate our hypothesis on state of the art models large vision language models (LLaVa-v1.5-7B, Idefics3-8B, SmolVLM-2B-Instruct) against standard datasets (Flickr30k, MMVet, TextVQA, VQAv2, POPE, COCO) and show that performance degrades significantly, even with a very small change to the input image, compared to baselines. Our results uncover a fundamentally different vector of performance degradation, highlighting a failure mode not captured by adversarial perturbations.
△ Less
Submitted 27 February, 2026;
originally announced March 2026.
-
ASFL: An Adaptive Model Splitting and Resource Allocation Framework for Split Federated Learning
Authors:
Chuiyang Meng,
Ming Tang,
Vincent W. S. Wong
Abstract:
Federated learning (FL) enables multiple clients to collaboratively train a machine learning model without sharing their raw data. However, the limited computation resources of the clients may result in a high delay and energy consumption on training. In this paper, we propose an adaptive split federated learning (ASFL) framework over wireless networks. ASFL exploits the computation resources of t…
▽ More
Federated learning (FL) enables multiple clients to collaboratively train a machine learning model without sharing their raw data. However, the limited computation resources of the clients may result in a high delay and energy consumption on training. In this paper, we propose an adaptive split federated learning (ASFL) framework over wireless networks. ASFL exploits the computation resources of the central server to train part of the model and enables adaptive model splitting as well as resource allocation during training. To optimize the learning performance (i.e., convergence rate) and efficiency (i.e., delay and energy consumption) of ASFL, we theoretically analyze the convergence rate and formulate a joint learning performance and resource allocation optimization problem. Solving this problem is challenging due to the long-term delay and energy consumption constraints as well as the coupling of the model splitting and resource allocation decisions. We propose an online optimization enhanced block coordinate descent (OOE-BCD) algorithm to solve the problem iteratively. Experimental results show that when compared with five baseline schemes, our proposed ASFL framework converges faster and reduces the total delay and energy consumption by up to 75% and 80%, respectively.
△ Less
Submitted 19 February, 2026;
originally announced March 2026.
-
ZorBA: Zeroth-order Federated Fine-tuning of LLMs with Heterogeneous Block Activation
Authors:
Chuiyang Meng,
Ming Tang,
Vincent W. S. Wong
Abstract:
Federated fine-tuning of large language models (LLMs) enables collaborative tuning across distributed clients. However, due to the large size of LLMs, local updates in federated learning (FL) may incur substantial video random-access memory (VRAM) usage. Moreover, frequent model exchange may lead to significant communication overhead. To tackle these challenges, in this paper we propose ZorBA, a z…
▽ More
Federated fine-tuning of large language models (LLMs) enables collaborative tuning across distributed clients. However, due to the large size of LLMs, local updates in federated learning (FL) may incur substantial video random-access memory (VRAM) usage. Moreover, frequent model exchange may lead to significant communication overhead. To tackle these challenges, in this paper we propose ZorBA, a zeroth-order optimization-based federated fine-tuning framework with heterogeneous block activation. ZorBA leverages zeroth-order optimization to eliminate the storage of gradients at the clients by forward passes. ZorBA includes a heterogeneous block activation mechanism in which the central server allocates different subsets of transformer blocks to clients in order to accelerate the convergence rate and reduce the VRAM usage. Furthermore, ZorBA utilizes shared random seeds and the finite differences of gradients in order to reduce the communication overhead. We conduct theoretical analysis to characterize the effect of block activation decisions on the convergence rate and VRAM usage. To jointly enhance the convergence rate and reduce the VRAM usage, we formulate an optimization problem to optimize the block activation decisions. We propose an $ε$-constraint lexicographic algorithm to solve this problem. Experimental results show that ZorBA outperforms three federated fine-tuning baselines in VRAM usage by up to 62.41% and incurs a low communication overhead.
△ Less
Submitted 19 February, 2026;
originally announced March 2026.
-
A Thermodynamic Structure of Asymptotic Inference
Authors:
Willy Wong
Abstract:
A thermodynamic framework for asymptotic inference is developed in which sample size and parameter variance define a state space. Within this description, Shannon information plays the role of entropy, and an integrating factor organizes its variation into a first-law-type balance equation. The framework supports a cyclic inequality analogous to a reversed second law, derived for the estimation of…
▽ More
A thermodynamic framework for asymptotic inference is developed in which sample size and parameter variance define a state space. Within this description, Shannon information plays the role of entropy, and an integrating factor organizes its variation into a first-law-type balance equation. The framework supports a cyclic inequality analogous to a reversed second law, derived for the estimation of the mean. A non-trivial third-law-type result emerges as a lower bound on entropy set by representation noise. Optimal inference paths, global bounds on information gain, and a natural Carnot-like information efficiency follow from this structure, with efficiency fundamentally limited by a noise floor. Finally, de Bruijn's identity and the I-MMSE relation in the Gaussian-limit case appear as coordinate projections of the same underlying thermodynamic structure. This framework suggests that ensemble physics and inferential physics constitute shadow processes evolving in opposite directions within a unified thermodynamic description.
△ Less
Submitted 25 March, 2026; v1 submitted 25 February, 2026;
originally announced February 2026.
-
Agentic AI for Intent-driven Optimization in Cell-free O-RAN
Authors:
Mohammad Hossein Shokouhi,
Vincent W. S. Wong
Abstract:
Agentic artificial intelligence (AI) is emerging as a key enabler for autonomous radio access networks (RANs), where multiple large language model (LLM)-based agents reason and collaborate to achieve operator-defined intents. The open RAN (O-RAN) architecture enables the deployment and coordination of such agents. However, most existing works consider simple intents handled by independent agents,…
▽ More
Agentic artificial intelligence (AI) is emerging as a key enabler for autonomous radio access networks (RANs), where multiple large language model (LLM)-based agents reason and collaborate to achieve operator-defined intents. The open RAN (O-RAN) architecture enables the deployment and coordination of such agents. However, most existing works consider simple intents handled by independent agents, while complex intents that require coordination among agents remain unexplored. In this paper, we propose an agentic AI framework for intent translation and optimization in cell-free O-RAN. A supervisor agent translates the operator intents into an optimization objective and minimum rate requirements. Based on this information, a user weighting agent retrieves relevant prior experience from a memory module to determine the user priority weights for precoding. If the intent includes an energy-saving objective, then an open radio unit (O-RU) management agent will also be activated to determine the set of active O-RUs by using a deep reinforcement learning (DRL) algorithm. A monitoring agent measures and monitors the user data rates and coordinates with other agents to guarantee the minimum rate requirements are satisfied. To enhance scalability, we adopt a parameter-efficient fine-tuning (PEFT) method that enables the same underlying LLM to be used for different agents. Simulation results show that the proposed agentic AI framework reduces the number of active O-RUs by 41.93% when compared with three baseline schemes in energy-saving mode. Using the PEFT method, the proposed framework reduces the memory usage by 92% when compared with deploying separate LLM agents.
△ Less
Submitted 25 February, 2026;
originally announced February 2026.
-
cyclinbayes: Bayesian Causal Discovery with Linear Non-Gaussian Directed Acyclic and Cyclic Graphical Models
Authors:
Robert Lee,
Raymond K. W. Wong,
Yang Ni
Abstract:
We introduce cyclinbayes, an open-source R package for discovering linear causal relationships with both acyclic and cyclic structures. The package employs scalable Bayesian approaches with spike-and-slab priors to learn directed acyclic graphs (DAGs) and directed cyclic graphs (DCGs) under non-Gaussian noise. A central feature of cyclinbayes is comprehensive uncertainty quantification, including…
▽ More
We introduce cyclinbayes, an open-source R package for discovering linear causal relationships with both acyclic and cyclic structures. The package employs scalable Bayesian approaches with spike-and-slab priors to learn directed acyclic graphs (DAGs) and directed cyclic graphs (DCGs) under non-Gaussian noise. A central feature of cyclinbayes is comprehensive uncertainty quantification, including posterior edge inclusion probabilities, posterior probabilities of network motifs, and posterior probabilities over entire graph structures. Our implementation addresses two limitations in existing software: (1) while methods for linear non-Gaussian DAG learning are available in R and Python, they generally lack proper uncertainty quantification, and (2) reliable implementations for linear non-Gaussian DCG remain scarce. The package implements computationally efficient hybrid MCMC algorithms that scale to large datasets. Beyond uncertainty quantification, we propose a new decision-theoretic approach to summarize posterior samples of graphs, yielding principled point estimates based on posterior expected loss such as posterior expected structural Hamming distance and structural intervention distance. The package, a supplementary material, and a tutorial are available on GitHub at https://github.com/roblee01/cyclinbayes.
△ Less
Submitted 24 February, 2026;
originally announced February 2026.
-
Cooperative ISAC for Joint Localization and Velocity Estimation in Cell-Free MIMO Systems
Authors:
Zihuan Wang,
Vincent W. S. Wong,
Robert Schober
Abstract:
In this paper, we explore a cooperative integrated sensing and communication (ISAC) framework that utilizes orthogonal frequency division multiplexing (OFDM) waveforms. Under the control of a central processing unit (CPU), multiple access points (APs) collaboratively perform multistatic sensing while providing communication service in a cell-free multiple-input multiple-output (MIMO) system. Achie…
▽ More
In this paper, we explore a cooperative integrated sensing and communication (ISAC) framework that utilizes orthogonal frequency division multiplexing (OFDM) waveforms. Under the control of a central processing unit (CPU), multiple access points (APs) collaboratively perform multistatic sensing while providing communication service in a cell-free multiple-input multiple-output (MIMO) system. Achieving high sensing accuracy requires the collection of global sensing information at the CPU, which can lead to significant fronthaul signaling overhead due to the feedback of the sensing signals from each AP. To tackle this issue, we propose a collaborative processing scheme in which the APs locally compress and quantize the received sensing signals before forwarding them to the CPU. The CPU then aggregates the information from all APs to estimate the location and velocity of the targets. We develop a distributed vector-quantized variational autoencoder (D-VQVAE) to enable an end-to-end implementation of this scheme. D-VQVAE consists of distributed encoders at the APs to locally encode the received sensing signals, codebooks for quantizing the encoded results, and a decoder at the CPU for location and velocity estimation. It effectively reduces the amount of data transmitted from each AP to the CPU while maintaining a high sensing accuracy. We employ a collaborative learning-assisted scheme to train D-VQVAE in an end-to-end manner. Simulation results show that the proposed D-VQVAE network outperforms the baseline schemes in sensing accuracy and reduces fronthaul signaling overhead by 99% when compared with the centralized sensing approach.
△ Less
Submitted 23 February, 2026;
originally announced February 2026.
-
FLoRG: Federated Fine-tuning with Low-rank Gram Matrices and Procrustes Alignment
Authors:
Chuiyang Meng,
Ming Tang,
Vincent W. S. Wong
Abstract:
Parameter-efficient fine-tuning techniques such as low-rank adaptation (LoRA) enable large language models (LLMs) to adapt to downstream tasks efficiently. Federated learning (FL) further facilitates this process by enabling collaborative fine-tuning across distributed clients without sharing private data. However, the use of two separate low-rank matrices in LoRA for federated fine-tuning introdu…
▽ More
Parameter-efficient fine-tuning techniques such as low-rank adaptation (LoRA) enable large language models (LLMs) to adapt to downstream tasks efficiently. Federated learning (FL) further facilitates this process by enabling collaborative fine-tuning across distributed clients without sharing private data. However, the use of two separate low-rank matrices in LoRA for federated fine-tuning introduces two types of challenges. First, aggregation error can arise from separately aggregating the two low-rank matrices. Second, even if the server aggregates the product of two low-rank matrices, it needs to decompose the aggregated matrix back into low-rank matrices. Since the decomposition is not unique, it can lead to decomposition drift. To tackle the aforementioned challenges, we propose federated low-rank Gram-matrix aggregation (FLoRG), a federated fine-tuning framework which employs a single low-rank matrix for fine-tuning and aggregates its Gram matrix (i.e., the matrix of inner products of its column vectors). FLoRG can eliminate the aggregation error and reduce the communication overhead. It also minimizes the decomposition drift by introducing a Procrustes alignment approach which aligns the decomposed matrix between consecutive fine-tuning rounds for consistent updates. We theoretically analyze the convergence of FLoRG and prove that adopting the Procrustes alignment results in a tighter convergence bound. Experimental results across multiple LLM fine-tuning benchmarks demonstrate that FLoRG outperforms five state-of-the-art baseline schemes by providing higher downstream task accuracy and can reduce the communication overhead by up to 2041$\times$.
△ Less
Submitted 6 March, 2026; v1 submitted 19 February, 2026;
originally announced February 2026.
-
GKP-inspired high-dimensional superdense coding with energy-time entanglement
Authors:
Kai-Chi Chang,
Arjun Mirani,
Murat Can Sarihan,
Xiang Cheng,
Michelle Harasimowicz,
Patrick Hayden,
Chee Wei Wong
Abstract:
Superdense coding, the application of entanglement to boost classical communication capacity, is a cornerstone of quantum communication. In this paper, we propose a high-dimensional superdense coding protocol using energy-time entangled states. These states are biphoton frequency combs, an example of entangled time-frequency Gottesman-Kitaev-Preskill (TFGKP) states or time-frequency grid states. I…
▽ More
Superdense coding, the application of entanglement to boost classical communication capacity, is a cornerstone of quantum communication. In this paper, we propose a high-dimensional superdense coding protocol using energy-time entangled states. These states are biphoton frequency combs, an example of entangled time-frequency Gottesman-Kitaev-Preskill (TFGKP) states or time-frequency grid states. Inspired by GKP codes, our protocol involves discretizing the continuous time and frequency degrees of freedom and encoding information by time-frequency displacements. This approach leverages the inherently large Hilbert space found in quantum frequency combs, with resilience against both temporal and spectral errors. In addition to describing the theoretical structure of the protocol, we propose an experimental implementation using standard telecommunication components, time-resolving single-photon detectors and a frequency beamsplitter. We also analyze the effect of experimental noise and errors on the channel capacity of the protocol. We demonstrate that for realistic experimental parameters, contemporary technologies satisfy the prerequisites for superdense coding with biphoton frequency combs, achieving a transmission rate of approximately 8.91 bits per transmitted photon (equivalent to 481 distinguishable messages with asymptotically vanishing errors). This more than doubles the previously highest transmission rate of 4 bits achieved by the Kwiat-Weinfurter scheme, while also having competitive optical loss. Furthermore, our results beat the rate achievable using a single-photon frequency comb with identical parameters by 4.6 times. Our protocol thus represents an experimentally feasible application of time-frequency grid states to entanglement-assisted communication, contributing to the active fields of continuous-variable and high-dimensional quantum information.
△ Less
Submitted 19 February, 2026; v1 submitted 16 February, 2026;
originally announced February 2026.
-
Kirin: Improving ANN efficiency with SNN Hybridization
Authors:
Chenyu Wang,
Zhanglu Yan,
Zhi Zhou,
Xu Chen,
Weng-Fai Wong
Abstract:
Artificial neural networks (ANNs), particularly large language models (LLMs), demonstrate powerful inference capabilities but consume substantial energy. Conversely, spiking neural networks (SNNs) exhibit exceptional energy efficiency due to their binary and event-driven characteristics, thus motivating the study of ANN-to-SNN conversion. In this process, quantization plays a pivotal role, mapping…
▽ More
Artificial neural networks (ANNs), particularly large language models (LLMs), demonstrate powerful inference capabilities but consume substantial energy. Conversely, spiking neural networks (SNNs) exhibit exceptional energy efficiency due to their binary and event-driven characteristics, thus motivating the study of ANN-to-SNN conversion. In this process, quantization plays a pivotal role, mapping LLMs' floating-point parameters to discrete SNN parameters via the temporal dimension of the time window. However, several challenges remain in the conversion process: (i) converting high bit-width quantization values into binary spikes requires longer time windows, increasing system latency; and (ii) the inherent trade-off between the information loss of single-spike schemes and the energy costs of multi-spike ones in SNN. To address these challenges, we propose Kirin, a integer and spike hybrid based SNN to achieve accuracy lossless ANN-to-SNN conversion with time and energy efficiency. Specifically, we first propose a Spike Matrix Hybridization strategy that encoding low bit-width parameters that leading to small time window size into binary spikes while preserving the rest in integer format, thereby reducing the overall latency of SNN execution. Second, we introduce a silence threshold mechanism to regulate the timing of single-spike firing, ensuring the output is mathematically equivalent to the LLM's output and preserves accuracy. Experimental results demonstrate that Kirin, under a W4A4\&8 quantization setting, achieves near-FP16 accuracy while reducing energy consumption by up to 84.66\% and shortening time steps by 93.75\%.
△ Less
Submitted 9 February, 2026;
originally announced February 2026.
-
Method on Using Shadow Altitude to Remove Geocoronal H$α$
Authors:
Wai-Kiu Ricky Wong,
Renbin Yan,
Zesen Lin
Abstract:
Spectroscopic surveys allow spatially resolved spectroscopy of galaxies to study their interstellar medium (ISM). However, observations of Galactic H$α$ emission are contaminated by geocoronal H$α$ emission. The latter is known to depend on the shadow altitude, a geometric parameter relating the line of sight to Earth's shadow cone. Using fibres on blank skys from the SDSS-IV/MaStar survey, we est…
▽ More
Spectroscopic surveys allow spatially resolved spectroscopy of galaxies to study their interstellar medium (ISM). However, observations of Galactic H$α$ emission are contaminated by geocoronal H$α$ emission. The latter is known to depend on the shadow altitude, a geometric parameter relating the line of sight to Earth's shadow cone. Using fibres on blank skys from the SDSS-IV/MaStar survey, we established an empirical relation between the geocoronal H$α$ emission and the shadow altitude, with a root mean square fractional scatter of 23.52$\%$. This relation can be used to predict geocoronal H$α$ emission so that it can be removed from observed spectra. This removal method is advantageous when the observed targets are extensive in the sky, and it does not require a large velocity separation between the observed target and the local standard of rest. This will enable reliable studies of Galactic H$α$ in intermediate spectral resolution integral field spectroscopic surveys. We also find tentative evidences for the dependences of geocoronal emission on solar activity and the distance between the Earth and the Sun.
△ Less
Submitted 31 January, 2026;
originally announced February 2026.