-
Centroid-Referenced Mahalanobis Matching (CRM): A Scalable, Representation-Based Framework for Causal Inference in Large Observational Studies
Authors:
Keming Hu,
Yingpei He
Abstract:
Matching for causal inference can be computationally expensive at scale and can silently change the target population when overlap is limited. We propose Centroid-Referenced Mahalanobis Matching (CRM), which replaces global pairwise search with stratified sampling in two reference coordinates: each unit's Mahalanobis distance from the treated centroid and its Fisher coordinate along the treated-co…
▽ More
Matching for causal inference can be computationally expensive at scale and can silently change the target population when overlap is limited. We propose Centroid-Referenced Mahalanobis Matching (CRM), which replaces global pairwise search with stratified sampling in two reference coordinates: each unit's Mahalanobis distance from the treated centroid and its Fisher coordinate along the treated-control mean shift. All covariates enter through the treated covariance geometry; CRM is therefore not principal-component preprocessing followed by nearest-neighbor matching. For $n$ units and $p$ pretreatment covariates, its implemented cost is $O(np^2+p^3+n\log n)$, simplifying to $O(np^2+n\log n)$ when $n \ge p$.
We derive an error decomposition separating representation, support, discretization, and stochastic components. A pre-matching shortage fraction $\hatπ$ estimates the population support restriction $π$, which enters a gap bound under bounded treatment-effect heterogeneity. Final retention is reported separately for capacity-driven exclusions. Under representation sufficiency, smoothness, and adequate cell capacity, CRM has a conservative two-dimensional histogram mean-squared-error bound $O(n_T^{-1/2})$; representation sufficiency is an additional assumption, not a consequence of ignorability given the original covariates.
On Criteo, CRM retains at least 99.4% of treated units, has lower MaxSMD than corrected propensity-score matching in 31 of 36 large-scale configurations, and is roughly an order of magnitude faster. Moderate-size simulations favor some pairwise and weighting baselines on balance, locating CRM's contribution in scalability and explicit support diagnostics rather than universal finite-sample dominance.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Writing and erasing skyrmions by single ultrafast laser pulses in monolayer Janus 2D magnets
Authors:
Guangyao Miao,
Yonglong Ga,
Chang Liu,
Pan Chen,
Yichen Jin,
Florian Kronast,
Wenxin Cheng,
Zhaoqing Ding,
Kai Hu,
Zongnan Zhang,
Nikolai Severin,
Chenxi Meng,
Patil Shubhada,
Sergio Valencia,
Meng Meng,
Qinlin Guo,
Xiaoran Liu,
Jiandi Zhang,
Yangmu Li,
Carlos-Andres Palma,
Jürgen P. Rabe,
Hongxin Yang,
Weihua Wang,
Jiandong Guo
Abstract:
Skyrmions in 2D magnets are promising candidates for nonvolatile, low-power, and high-density spintronic memories. However, their experimental realization at the 2D limit remains challenging, owing to the difficulty in engineering the required chiral magnetic interactions. Here, we report the creation and direct imaging of Néel-type skyrmions in Janus 2D chromium chalcogenides using synchrotron X-…
▽ More
Skyrmions in 2D magnets are promising candidates for nonvolatile, low-power, and high-density spintronic memories. However, their experimental realization at the 2D limit remains challenging, owing to the difficulty in engineering the required chiral magnetic interactions. Here, we report the creation and direct imaging of Néel-type skyrmions in Janus 2D chromium chalcogenides using synchrotron X-ray photoemission electron microscopy, and scanning nitrogen-vacancy magnetometry, which exhibit field-free stability, nonvolatility, and size tunability. First-principles calculations and micromagnetic simulations reveal that Janus-surface-induced inversion-symmetry breaking enhances the Dzyaloshinskii-Moriya interaction, providing the microscopic mechanism for skyrmion stabilization and tunability. We further achieve reversible skyrmion writing and erasing using a single ultrafast laser pulse in a magnetic field as low as 300 Oe, demonstrating the excellent manipulability of this 2D magnetic system. These results establish Janus engineering as a route to creating and manipulating nonvolatile skyrmions in atomically thin magnets, with implications for skyrmion-based low-power spintronic devices.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
VoiceChat-TTS: A Low-Latency Continuous Speech Synthesis Model for Interactive Agents
Authors:
Edresson Casanova,
Jaehyeon Kim,
Mariana Graterol Fuenmayor,
Shehzeen Hussain,
Viacheslav Klimkov,
Valentin Mendelev,
Mikyas Desta,
Paarth Neekhara,
Piotr Zelasko,
Chen Chen,
Elena Rastorgueva,
Ke Hu,
Ankita Pasad,
Xuesong Yang,
Aya Alja'fari,
Rajarshi Roy,
Rohan Badlani,
Jason Roche,
Jason Li,
Zhehuai Chen
Abstract:
Spoken dialogue is a natural form of human--computer interaction, yet most speech language models remain limited to turn-based operation and lack real-time adaptability, such as user barge-in. Recent duplex speech-to-speech and speech-to-text models reduce latency by replacing multi-stage pipelines, but often compromise speech quality because accurate ASR, interruption handling, and high-fidelity…
▽ More
Spoken dialogue is a natural form of human--computer interaction, yet most speech language models remain limited to turn-based operation and lack real-time adaptability, such as user barge-in. Recent duplex speech-to-speech and speech-to-text models reduce latency by replacing multi-stage pipelines, but often compromise speech quality because accurate ASR, interruption handling, and high-fidelity synthesis must be optimized jointly. We propose VoiceChat-TTS, a low-latency, continuous, and streamable text-to-speech model for interactive agents. VoiceChat-TTS is driven directly by LLM text-token streams, supports explicit interruption via control tokens, and produces silence when no textual input is available. The model enables always-on, responsive speech generation while preserving modularity and high speech quality, and it supports mid-utterance interruptions without resetting the KV cache.
△ Less
Submitted 13 August, 2026;
originally announced August 2026.
-
Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning
Authors:
Xulin Fan,
Jialu Li,
Mohammad Nur Hossain Khan,
Kexin Hu,
Bashima Islam,
Mark Hasegawa-Johnson,
Nancy L. McElwain
Abstract:
Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challenging due to limited labeled data, low signal-to-noise ratio, and cross-family domain shifts. We present a family-conditioned, multi-tier audio tagger that combines a LoRA-finetuned Whisper encoder with a lightweight, targe…
▽ More
Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challenging due to limited labeled data, low signal-to-noise ratio, and cross-family domain shifts. We present a family-conditioned, multi-tier audio tagger that combines a LoRA-finetuned Whisper encoder with a lightweight, target-speaker-aware Transformer for long-context inference and framewise prediction across tiers. To improve temporal coherence, we incorporate a simple sequence-level smoothing loss, and to enhance robustness across households, we introduce a factorized speaker-token design with a shared tier token and a learned family-specific offset, reducing family bias and promoting generalizable representations. Together, these choices enable efficient and effective infant-centered audio tagging of daylong audio recordings in home environments.
△ Less
Submitted 11 August, 2026;
originally announced August 2026.
-
JSGS: JPEG State-Guided Supervision for 3D Gaussian Splatting from Mixed-Quality Views
Authors:
Jinhua Cui,
Anhong Wang,
Kai Hu,
Donghan Bu,
Peihao Li,
Tammam Tillo,
Hao Jing,
Shiao Xu
Abstract:
Standard 3D Gaussian Splatting (3DGS) assumes that every input image faithfully samples scene radiance. However, mixed-quality JPEG images violate this assumption because compression-induced blocking and ringing artifacts can corrupt updates to Gaussians shared across views. To address this problem, we propose JPEG State-Guided Supervision for 3D Gaussian Splatting from Mixed-Quality Views (JSGS).…
▽ More
Standard 3D Gaussian Splatting (3DGS) assumes that every input image faithfully samples scene radiance. However, mixed-quality JPEG images violate this assumption because compression-induced blocking and ringing artifacts can corrupt updates to Gaussians shared across views. To address this problem, we propose JPEG State-Guided Supervision for 3D Gaussian Splatting from Mixed-Quality Views (JSGS). JSGS uses luminance and chrominance quantization tables stored in each JPEG file to construct a view-specific JPEG observation operator. This operator encodes and decodes each rendered view for domain-matched comparison with the corresponding decoded input image. The luminance quantization table supplies continuous weights within a fixed middle frequency band. A loss in the low frequency band anchors coarse structure, while the weighted middle frequency loss redistributes supervision among the selected DCT coordinates. The resulting block disagreement also guides the Gaussian Controller to regularize small primitives with high opacity in disagreement regions. Across seven scenes and three mixed-quality schedules, JSGS achieves the lowest mean LPIPS and the highest mean SSIM under every schedule while rendering at approximately 150 FPS. Code: https://github.com/Jayden-Cui/JSGS.
△ Less
Submitted 9 August, 2026;
originally announced August 2026.
-
Position: It's Time to Optimize LLMs for Self-Consistency
Authors:
Itamar Pres,
Belinda Z. Li,
Laura Ruis,
Zifan Carl Guo,
Keya Hu,
Mehul Damani,
Isha Puri,
Ekdeep Singh Lubana,
Jacob Andreas
Abstract:
Despite ever-increasing sophistication in language model (LM) pre- and post-training pipelines, many important failures persist: models overcondition on user framing ("sycophancy"), exhibit incomplete logical generalization, and produce confident but incorrect responses. We argue that these failures arise from a modeling assumption permeating all aspects of the pipeline: that behavior can be speci…
▽ More
Despite ever-increasing sophistication in language model (LM) pre- and post-training pipelines, many important failures persist: models overcondition on user framing ("sycophancy"), exhibit incomplete logical generalization, and produce confident but incorrect responses. We argue that these failures arise from a modeling assumption permeating all aspects of the pipeline: that behavior can be specified and evaluated independently on single-output pairs. Many model failures are difficult, if not impossible, to detect without reasoning about relationships between a model's responses across inputs. In this position paper, we propose self-consistency as a framework for understanding these failures. We first observe that a wide variety of techniques designed to improve specific aspects of LM behavior-targeting properties as diverse as adversarial robustness and factual coherence-can be understood as special cases of a common "consistency optimization" procedure and addressed with a standard set of optimization tools. We next outline a set of new model properties that could be achieved by optimizing for consistency, and conclude with a discussion of what it would mean to develop generally consistent LMs, including the capabilities they would enable and the objections they raise.
△ Less
Submitted 31 July, 2026;
originally announced August 2026.
-
HumanCLAW: Can Vision-Language Models Act Through a Body?
Authors:
Li Siyao,
Jiawei Gu,
Shuai Liu,
Kairui Hu,
Zekun Li,
Linjie Li,
Chengcheng Tang,
Po-Chen Wu,
Ivan Shugurov,
Lingni Ma,
Michael Zollhoefer,
Sizhe An,
Abhay Mittal,
Amy Zhao,
Ranjay Krishna,
Manling Li,
Ziwei Liu,
Chuan Guo
Abstract:
Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed to execute it, e.g., losing balance and falling. In this work, we introduce HumanCLAW, an evaluation framework that decou…
▽ More
Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed to execute it, e.g., losing balance and falling. In this work, we introduce HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution. At every step, a harnessed, off-the-shelf VLM issues an atomic skill command, and the command is translated into a sub-second chunk of continuous full-body motion with real physical consequences, including gravity and collisions. The body can therefore act freely in the physical world, while execution-side disturbances, balance and motor errors, are factored out. What remains measurable is the model's action intelligence: its moment-to-moment choice of what the body should execute next. Based on this framework, we build HumanCLAW-Bench: 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes. We test nine state-of-the-art VLMs and find that none solves the benchmark; the best model reaches only a 16.8% success rate. Recognizing the target is not the bottleneck. What current VLMs lack is embodied self-awareness: they lose track of their own body, failing to tell where it is, whether it has reached the goal, or whether it has hit an obstacle.
△ Less
Submitted 3 August, 2026; v1 submitted 29 July, 2026;
originally announced July 2026.
-
Optimal block preconditioners for a mass-conserving mixed stress formulation of Stokes flow
Authors:
Kaibo Hu,
Jongho Park,
Jindong Wang
Abstract:
We present optimal block diagonal and triangular preconditioners for a mass-conserving mixed stress formulation of Stokes flow. The algebraic formulation leads to a double saddle point system with unknowns corresponding to discrete stress, velocity, vorticity, and pressure. MINRES equipped with a block diagonal preconditioner for an augmented Lagrangian formulation of this system is analyzed and s…
▽ More
We present optimal block diagonal and triangular preconditioners for a mass-conserving mixed stress formulation of Stokes flow. The algebraic formulation leads to a double saddle point system with unknowns corresponding to discrete stress, velocity, vorticity, and pressure. MINRES equipped with a block diagonal preconditioner for an augmented Lagrangian formulation of this system is analyzed and shown to be optimal, in the sense that the convergence rate is independent of key parameters such as mesh size and kinematic viscosity. GMRES equipped with a block triangular preconditioner is also analyzed using a field-of-values approach. Finally, we present numerical results for both two- and three-dimensional model problems to validate the parameter robustness of the proposed preconditioners.
△ Less
Submitted 24 July, 2026;
originally announced July 2026.
-
ReLU$^k$ Neural de Rham Complexes
Authors:
Kaibo Hu,
Jindong Wang,
Jinchao Xu
Abstract:
We construct finite-dimensional de Rham subcomplexes generated by fixed-neuron shallow ReLU$^k$ neural networks, a class of spaces known to provide optimal approximation rates. For neurons of the form $s_i(x)=ω_i\cdot x+b_i$, we introduce spaces of neural differential forms: differential $p$-forms whose coefficients are the ReLU$^k$ ridge functions $σ_{k-p}(s_i)$. These spaces are compatible with…
▽ More
We construct finite-dimensional de Rham subcomplexes generated by fixed-neuron shallow ReLU$^k$ neural networks, a class of spaces known to provide optimal approximation rates. For neurons of the form $s_i(x)=ω_i\cdot x+b_i$, we introduce spaces of neural differential forms: differential $p$-forms whose coefficients are the ReLU$^k$ ridge functions $σ_{k-p}(s_i)$. These spaces are compatible with the exterior derivative because differentiating a ReLU power lowers its order by one, and for each fixed neuron, differentiation amounts to exterior multiplication by the fixed one-form $ d s_i$. Under a linear independence assumption on the lowest-order family $\{σ_{k-d}(s_i)\}_{i=1}^n$, the global complex decomposes into independent neuron-wise Koszul complexes. We prove exactness in arbitrary dimension and provide a geometric sufficient condition for the required linear independence. Numerical experiments based on the resulting complex provide evidence of stable discretizations and of convergence rates consistent with the underlying approximation theory, and exhibit no spurious modes in eigenvalue problems considered.
△ Less
Submitted 24 July, 2026;
originally announced July 2026.
-
Apple-$π$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
Authors:
Runmao Yao,
Kairui Hu,
Yukang Cao,
Ruisi Wang,
Shulin Tian,
Ziang Cao,
Weichen Fan,
Ziqi Huang,
Yuhao Dong,
Hao Li,
Zhaoxi Chen,
Zhongang Cai,
Lei Yang,
Ziwei Liu
Abstract:
Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without verifying whether the model arrives there through a faithful, law-grounded reasoning process. We introduce Apple-PI, the first benchmark that anchors video-model evaluation explic…
▽ More
Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without verifying whether the model arrives there through a faithful, law-grounded reasoning process. We introduce Apple-PI, the first benchmark that anchors video-model evaluation explicitly in physical laws. Apple-PI comprises three components. 1) Orchard: a dataset of 400 videos covering ten canonical tasks in classical mechanics. It separates single-law tasks for confounder-free diagnosis from multi-law tasks for probing generalization. 2) Benchmark Protocol: a three-stage protocol based on scientific reasoning, including Perception, Formulation, and Deduction. It uses chain-of-frames prompting on infographic-annotated first frames, treating the generated video as the model's visible reasoning trace. 3) Evaluation Suite: a hybrid evaluation suite that combines MLLM-based subjective scoring with physics-law-grounded objective measures. This enables stage-resolved diagnosis of not only whether a model fails, but where it fails. Benchmarking 11 models shows that current video models remain far from reliable law-grounded world simulators, with the best video model scoring only 0.473. Our stage-, pillar-, and source-resolved analyses further expose a Perception-to-Formulation-to-Deduction bottleneck, weak multi-law state transfer, and a persistent Sim-to-Real gap. These findings position Apple-PI as a diagnostic foundation for guiding future video models toward world models with law-grounded physical intelligence.
△ Less
Submitted 17 July, 2026;
originally announced July 2026.
-
MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers
Authors:
Huanxi Liu,
Kun Hu,
Jiaqi Liao,
Qiang Wang,
Pengfei Qian,
YuanZhao Zhai,
Dawei Feng,
Bo Ding,
Huaimin Wang
Abstract:
As Model Context Protocol (MCP) servers emerge as the core infrastructure for connecting LLMs with external tools, existing benchmarks leverage real-world MCP servers to evaluate LLM agents' tool-using capabilities. However, these benchmarks overlook the continuous evolution of tool interfaces and functionalities within MCP servers, resulting in flawed assessments that fail to capture the agent's…
▽ More
As Model Context Protocol (MCP) servers emerge as the core infrastructure for connecting LLMs with external tools, existing benchmarks leverage real-world MCP servers to evaluate LLM agents' tool-using capabilities. However, these benchmarks overlook the continuous evolution of tool interfaces and functionalities within MCP servers, resulting in flawed assessments that fail to capture the agent's adaptability in changing tool landscapes. To bridge this gap, we introduce \textbf{MCPEvol-Bench}, a novel benchmark for evaluating the task-solving capabilities of LLM agents under dynamic toolset evolution. Inspired by large-scale empirical study, we propose 11 mutation operators to simulate realistic tool evolution within 123 MCP servers. We benchmark 12 state-of-the-art LLMs on multiple versions of MCP servers, revealing that even frontier models struggle to adapt to evolving tools. For instance, GPT-5.4 and Claude-Sonnet-4-6 exhibit performance declines of 13.7\% and 14.4\% in evolved MCP servers, respectively, accompanied by substantial increases in planning and reasoning errors. These findings highlight the vulnerability of LLM-driven workflows, establishing MCPEvol-Bench as a standard for evaluating agent adaptability in dynamic tool environments.
△ Less
Submitted 16 July, 2026;
originally announced July 2026.
-
ProfMalPlus: Agent-Coordinated Detection of Malicious NPM Packages via Static-Dynamic Analysis Synergy
Authors:
Yiheng Huang,
Zhijia Zhao,
Bihuan Chen,
Susheng Wu,
Zhuotong Zhou,
Yiheng Cao,
Kun Hu,
Xin Hu,
Xin Peng
Abstract:
Open source software is vulnerable to supply-chain attacks through transitive dependencies, especially malicious code injected into NPM packages. Existing detectors often inadequately model obfuscated behavior, overlook JavaScript's object-centric features, poorly coordinate static and dynamic analysis, and lose semantic information during behavior abstraction. We propose ProfMalPlus, a malicious…
▽ More
Open source software is vulnerable to supply-chain attacks through transitive dependencies, especially malicious code injected into NPM packages. Existing detectors often inadequately model obfuscated behavior, overlook JavaScript's object-centric features, poorly coordinate static and dynamic analysis, and lose semantic information during behavior abstraction. We propose ProfMalPlus, a malicious NPM package detector combining object-sensitive behavior graphs with coordinated LLM reasoning over annotated code slices. It identifies installation commands and entry files, then constructs graphs capturing sensitive APIs, third-party calls, and unresolved calls. From these graphs, ProfMalPlus extracts security-relevant slices and adds inline static analysis evidence. Local judge agents independently assess each slice. Self-consistency consolidates repeated judgements to reduce LLM variance, while a global judge synthesizes their reports into an entry-level verdict. For undetermined cases, a router selects either third-party enrichment, which adds registry derived module and method semantics, or dynamic augmentation, which executes the package in a sandbox to resolve runtime dependent behavior. The enriched evidence is fed back for reassessment. Finally, a localization agent reports malicious code snippets with explanations. ProfMalPlus achieves a 98.1% F1-score, outperforming state-of-the-art detectors by 3.5% to 52.6%. It also identified 597 previously unknown malicious packages, all confirmed and removed from NPM.
△ Less
Submitted 15 July, 2026;
originally announced July 2026.
-
Q-BridgeNet: A Quantization Network for Cross-Lingual Sign Language Translation
Authors:
Liqian Feng,
Lintao Wang,
Xiaochen Liu,
Anusha Withana,
Ken-Tye Yong,
Dehui Kong,
Zhiyong Wang,
Kun Hu
Abstract:
Most sign language translation (SLT) methods focus on isolated native sign-spoken pairs (e.g., American Sign Language - English). Extending language-specific SLT models to multilingual translation would improve accessibility by enabling communication across diverse sign and spoken language communities. However, existing multilingual SLT approaches still struggle to learn a unified model that minim…
▽ More
Most sign language translation (SLT) methods focus on isolated native sign-spoken pairs (e.g., American Sign Language - English). Extending language-specific SLT models to multilingual translation would improve accessibility by enabling communication across diverse sign and spoken language communities. However, existing multilingual SLT approaches still struggle to learn a unified model that minimizes cross-lingual conflicts while capturing shared cross-lingual semantics and preserving language-specific variations across different sign languages. Therefore, we propose Q-BridgeNet, a unified framework for multilingual SLT that jointly mitigates cross-lingual conflicts across both the sign language and spoken language sides. On the sign language side, Q-BridgeNet learns discrete Q-units via adaptive segmentation and residual vector quantization: a shared base codebook provides language-agnostic semantic primitives, while language-specific residual codebooks refine heterogeneous signing semantics. On the spoken language side, a multilingual LLM is fine-tuned to operate in the Q-unit space, leveraging cross-lingual priors to enable a unified SLT model. Experiments on PHOENIX14T, How2Sign, and CSL-Daily show that Q-BridgeNet effectively mitigates cross-lingual conflicts, achieving state-of-the-art performance on native sign-spoken pairs while also demonstrating strong generalization to non-native pairs. Our source code is publicly available at: https://github.com/FengLiQ/Q-BridgeNet
△ Less
Submitted 13 July, 2026;
originally announced July 2026.
-
A Multi-Frequency Input-Admittance Model of Locomotive Rectifier Considering PWM Sideband Harmonic Coupling in Electrical Railways
Authors:
Xiangyu Meng,
Zhigang Liu,
Guorong Li,
Xunjun Chen,
Siqi Wu,
Keting Hu
Abstract:
Electrical railway harmonic instability issues are common in the high-frequency range. The effective frequency of the traditional converter's small-signal averaging model is below 1/2 switching frequency since the pulse width modulation (PWM) sideband harmonic components are ignored. In this article, the dynamic propagations of perturbation frequency and the generated PWM sideband components are c…
▽ More
Electrical railway harmonic instability issues are common in the high-frequency range. The effective frequency of the traditional converter's small-signal averaging model is below 1/2 switching frequency since the pulse width modulation (PWM) sideband harmonic components are ignored. In this article, the dynamic propagations of perturbation frequency and the generated PWM sideband components are constructed first. Then the locomotive rectifier's multi-frequency input-admittance model is derived appropriately. Afterward, an admittance conversion approach is used to convert the multi-frequency model into the single-input-single-output (SISO) model whereas retaining the sideband frequency couplings. The proposed SISO model is more accurate than the traditional small-signal averaging model in the frequency range higher than 1 / 2 switching frequency. It is found that PWM sideband harmonics dominate the locomotive rectifier's input-admittance characteristic higher than 1 / 2 switching frequency. Finally, based on the proposed model, the influence of different switching frequencies, control bandwidths, and traction network impedance on system harmonic stability is revealed by the hardware-in-the-loop (HIL) results.
△ Less
Submitted 10 July, 2026;
originally announced July 2026.
-
Reinforcement Learning for Evidence-Seeking Diagnostic Reasoning with Large Language Models
Authors:
Shengyi Hua,
Kangzhe Hu,
Conghui He,
Xiaofan Zhang,
Shaoting Zhang
Abstract:
Recent reasoning-centric Large Language Models (LLMs) have made significant strides, yet they predominantly operate on a passive-inference pattern that assumes complete information. In contrast, real-world clinical intelligence is inherently an iterative investigative process requiring strategic evidence acquisition. To bridge this gap, we formalize medical diagnosis as an Iterative Evidence-Seeki…
▽ More
Recent reasoning-centric Large Language Models (LLMs) have made significant strides, yet they predominantly operate on a passive-inference pattern that assumes complete information. In contrast, real-world clinical intelligence is inherently an iterative investigative process requiring strategic evidence acquisition. To bridge this gap, we formalize medical diagnosis as an Iterative Evidence-Seeking Task. We leverage Reinforcement Learning with Verifiable Rewards (RLVR) to elicit intrinsic reasoning within a closed-loop environment, guided by a novel suite of rewards that enforce diagnostic precision and examination consistency. To facilitate this, we introduce the Retrieval-Augmented Generation-based Examination Simulator (RAGES), a high-fidelity clinical oracle that provides realistic, knowledge-grounded follow-up evidence. Empirical results across diverse datasets demonstrate that our framework enables LLMs to transition from passive responders to autonomous assistants. Notably, our model demonstrates comparable performance to larger and reasoning-enhanced baselines, while RAGES proves superior to vanilla LLMs in generating biologically plausible clinical feedback.
△ Less
Submitted 3 July, 2026;
originally announced July 2026.
-
Steal the Patch Size: Adversarially Manipulate Vision-Language Models
Authors:
Kai Hu,
Akash Bharadwaj,
Weichen Yu,
Matt Fredrikson
Abstract:
We present a black-box model-stealing attack that recovers private vision-tokenizer configurations of deployed vision-language models (VLMs), including the visual patch size and input preprocessing pipeline. The key idea is a task-level side channel induced by ViT-style patchification: when a synthetic grid image is aligned with the hidden patch grid, boundary cues are erased at tokenization, caus…
▽ More
We present a black-box model-stealing attack that recovers private vision-tokenizer configurations of deployed vision-language models (VLMs), including the visual patch size and input preprocessing pipeline. The key idea is a task-level side channel induced by ViT-style patchification: when a synthetic grid image is aligned with the hidden patch grid, boundary cues are erased at tokenization, causing periodic accuracy drop. By sweeping the grid cell size and measuring these collapses, we infer the patch size; by introducing padding and a consistency-check test, we further identify whether preprocessing is dynamic- or fixed-resolution and recover the target resize resolution. Across open-source Qwen-VL variants and proprietary models including GPT and Claude, we reliably recover tokenizer-related parameters. Finally, we show that such leakage enables preprocessing-aware transfer attacks and model-targeted adversarial manipulation.
△ Less
Submitted 30 June, 2026;
originally announced July 2026.
-
Energy-Resolved Limits on Orbital X-ray Polarization Modulation in Cygnus X-1
Authors:
Sohee Chun,
Bert Vander Meulen,
Kun Hu,
Henric Krawczynski
Abstract:
Reflection off the companion star and its focused stellar wind is predicted to modulate the X-ray polarization of black hole X-ray binaries at half the orbital period ($P_{\rm orb}/2$), with an energy-dependent amplitude. We test this prediction against all publicly available IXPE observations of Cygnus X-1, comprising 26 one-day bins from 12 observation IDs spanning 2022-2024. Since the normalize…
▽ More
Reflection off the companion star and its focused stellar wind is predicted to modulate the X-ray polarization of black hole X-ray binaries at half the orbital period ($P_{\rm orb}/2$), with an energy-dependent amplitude. We test this prediction against all publicly available IXPE observations of Cygnus X-1, comprising 26 one-day bins from 12 observation IDs spanning 2022-2024. Since the normalized Stokes parameters correlate linearly with the spectral hardness ratio in all three energy bands (2-4, 4-6, and 6-8 keV), we employ a simultaneous harmonic regression that decouples spectral variability from orbital modulation at both $P_{\rm orb}/2$ and $P_{\rm orb}$, complemented by direct fitting of 3D Monte Carlo radiative transfer stellar companion and wind-scattering templates. After removing the spectral hardness trend, neither approach reveals statistically significant orbital modulation: permutation tests yield $p > 0.01$ in all bands, with 99% confidence upper limits of 0.47%, 0.67%, and 1.81% on the $P_{\rm orb}$ amplitude and 0.54%, 0.77%, and 2.13% on the $P_{\rm orb}/2$ amplitude in the 2-4 keV, 4-6 keV, and 6-8 keV bands, respectively. The best-fit stellar companion and wind-scattering amplitude scaling factors in the three bands of $A = $ 0.78$\pm$0.89, 0.96$\pm$0.62, and $-$1.02$\pm$1.11 are consistent with a null result. These non-detections are sensitivity-limited, as the predicted stellar companion and wind-scattering RMS amplitudes in the three bands of $\approx$0.10%, $\approx$0.33%, and $\approx$0.49% are at or below the statistical noise floor of $\sim$0.15%, $\sim$0.31%, and $\sim$0.84%. We quantify the additional exposure required to detect the predicted signal and constrain the wind physics.
△ Less
Submitted 29 June, 2026;
originally announced June 2026.
-
Enhanced Magnon Synchronization in Coupled WGM Optomagnonic Resonators with Phase-Dependent Photon Hopping
Authors:
Le-Ji Xue,
Ying-Jian Zhu,
Jaspal Singh,
Ahmad Zahia,
Kong-Ming Hu,
Jia-Xin Peng,
S. K. Singh
Abstract:
We investigate quantum synchronization in a coupled cavity optomagnonic system which consists of two spatially separated optical whispering-gallery-mode (WGM) resonators and each resonator is also coupled to a yttrium iron garnet (YIG) sphere through the optomagnonic interaction. Phase-dependent single-photon hopping factor couples the two optical resonators and provides an indirect interaction be…
▽ More
We investigate quantum synchronization in a coupled cavity optomagnonic system which consists of two spatially separated optical whispering-gallery-mode (WGM) resonators and each resonator is also coupled to a yttrium iron garnet (YIG) sphere through the optomagnonic interaction. Phase-dependent single-photon hopping factor couples the two optical resonators and provides an indirect interaction between the two distant magnon modes. We then investigate complete synchronization, φ-synchronization, and quantum phase synchronization using the covariance-matrix formalism as well as also studying the effects of the hopping term on the overall synchronization dynamics of two distant magnon modes. It can be seen that the photon-hopping phase provides an efficient way to control the synchronization dynamics and when it is varied from 0 to π, the magnon trajectories gradually evolve from weakly correlated motion to a highly synchronized state, which is also accompanied by a significant reduction in the synchronization error. The influence of the photon-hopping strength and thermal fluctuations is also investigated, where it can be seen that stronger photon hopping enhances all synchronization measures, while thermal noise weakens the coherent correlations responsible for synchronized dynamics. Our results demonstrate that the phase of the hopping factor offers a simple and effective approach for controlling synchronization dynamics in WGM based coupled cavity optomagnonic systems and also provide a useful route towards coherent control of collective magnon dynamics in such quantum optomganonic devices.
△ Less
Submitted 29 June, 2026;
originally announced June 2026.
-
Classification of regular Cayley maps of skew-type three on semidihedral groups
Authors:
Kan Hu,
Tao Qiu
Abstract:
It is well known that every regular Cayley map $M = \CM(G,X,p)$ on a finite group $G$ with respect to an inverse-closed generating set $X$ of $G$ and a specified cyclic permutation $p$ on $X$ corresponds to a skew morphism $\varphi$ on $G$ such that the restriction of $\varphi$ to $X$ is $p$. The skew-type of the map $M$ is defined as the index $[G:\Ker \varphi]$, which equals the number of distin…
▽ More
It is well known that every regular Cayley map $M = \CM(G,X,p)$ on a finite group $G$ with respect to an inverse-closed generating set $X$ of $G$ and a specified cyclic permutation $p$ on $X$ corresponds to a skew morphism $\varphi$ on $G$ such that the restriction of $\varphi$ to $X$ is $p$. The skew-type of the map $M$ is defined as the index $[G:\Ker \varphi]$, which equals the number of distinct values in $\mathbb{Z}_{|\varphi|}$ taken by the associated power function $π$ of the skew morphism $\varphi$. In this paper, we develop a covering theory of skew morphisms and as an application we provide a classification of regular Cayley maps of skew-type three on the semidihedral groups.
△ Less
Submitted 24 June, 2026;
originally announced June 2026.
-
RoBoSR: Structured Scene Representations for Embodied Robotic Reasoning
Authors:
Kewei Hu,
Wanchan Yu,
Fangwen Chen,
Jing Jiajian,
Zimeng Li,
Ying Wei,
Tianhao Liu,
Michael Zhang,
Hanwen Kang
Abstract:
Despite rapid progress, embodied reasoning under real-world variability remains challenging. Existing approaches rely on demonstration-driven sequential biases, limiting flexibility in open-ended and long-horizon tasks that require structured reasoning over evolving states.
We introduce RoBoSR, an intermediate structural representation that formulates manipulation as step-wise state transitions…
▽ More
Despite rapid progress, embodied reasoning under real-world variability remains challenging. Existing approaches rely on demonstration-driven sequential biases, limiting flexibility in open-ended and long-horizon tasks that require structured reasoning over evolving states.
We introduce RoBoSR, an intermediate structural representation that formulates manipulation as step-wise state transitions over semantically grounded, object-centric scene graphs. By modeling object states and their spatial relations at the perception-action interface, RoBoSR disentangles high-level task reasoning from raw inputs and enables structured reasoning over preconditions, effects, and goal states. This representation endows the agent with causal reasoning capability, enforcing subtask dependencies and supporting coherent long-horizon task planning.
To learn such structure-aware reasoning, we construct Manip-Cognition-1.6M, an open-world dataset that jointly supervises scene understanding, instruction interpretation, and subtask planning across diverse tasks.
Across several benchmarks and real-world demonstrations, our method consistently outperforms prompting-based methods and classical TAMP baselines in zero-shot generalization and long-horizon tasks. The results underscore structured intermediate representations as a critical inductive bias for scalable embodied reasoning.
△ Less
Submitted 23 June, 2026;
originally announced June 2026.
-
CODEBLOCK: Learning to Supervise Code at the Right Granularity
Authors:
Zhijie Deng,
Ling Li,
Jinlong Pang,
Kaiqin Hu,
Qi Xuan,
Zhaowei Zhu,
Jiaheng Wei
Abstract:
Supervised fine-tuning of code LLMs typically applies uniform cross-entropy loss to all response tokens, implicitly assuming that every token provides equally useful learning signal. Recent token-level selection methods challenge this assumption in natural-language SFT by supervising only high-value tokens. However, directly transferring token-level masking to code can break syntactically and sema…
▽ More
Supervised fine-tuning of code LLMs typically applies uniform cross-entropy loss to all response tokens, implicitly assuming that every token provides equally useful learning signal. Recent token-level selection methods challenge this assumption in natural-language SFT by supervising only high-value tokens. However, directly transferring token-level masking to code can break syntactically and semantically coherent program units, because code depends on structural completeness and definition-use relations. We therefore propose CodeBlock, a structure-aware sparse supervision framework that selects structure-complete code evidence rather than isolated tokens. CodeBlock first selects high-quality instruction-response pairs, then partitions code responses into syntactically coherent coding items, estimates their utility by aggregating generalized cross-entropy over core logic tokens, and reranks them with data-flow reach and bridge signals to prioritize blocks that propagate or connect important program dependencies. During training, the full response remains available as context, while loss is applied only to selected code items and informative natural-language tokens. Experiments on six code-generation benchmarks show that CodeBlock achieves stronger average pass@1 than full-token SFT and competitive selection baselines, while using only 1.9% of supervised response tokens.
△ Less
Submitted 10 June, 2026;
originally announced June 2026.
-
Steering Generative Reinforcement Learning into Stable Robotic Controller
Authors:
Yixuan Wang,
Shutong Ding,
Ke Hu,
Tianxiang Gui,
Jingya Wang,
Ye Shi
Abstract:
Diffusion and flow-based generative policies provide a powerful policy class for reinforcement learning by inducing rich stochastic exploration through iterative action generation. However, the stochasticity of diffusion policies is not suitable for stable and precise control in high-dimensional robotic systems, where small action variations can accumulate into inconsistent motion and reduced robu…
▽ More
Diffusion and flow-based generative policies provide a powerful policy class for reinforcement learning by inducing rich stochastic exploration through iterative action generation. However, the stochasticity of diffusion policies is not suitable for stable and precise control in high-dimensional robotic systems, where small action variations can accumulate into inconsistent motion and reduced robustness. To address this issue, we propose SteerGenPO, a latent-space reinforcement learning framework that steers a trained generative policy into a robust deterministic robotic controller. The key idea is to replace stochastic latent sampling of the trained generative policy with a learned latent actor that predicts a state-dependent latent input for the generative policies. This separates exploration and control: stochastic generative sampling provides diverse action proposals during policy learning, while deterministic latent steering provides stable and adaptive control at deployment. We evaluate SteerGenPO on six Isaac Lab benchmarks and a Unitree G1 locomotion task. The results show SteerGenPO improves over both classical RL and generative RL baselines, while its deterministic latent steering produces more stable inference-time behaviors and more reliable command responses.
△ Less
Submitted 15 June, 2026;
originally announced June 2026.
-
Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale
Authors:
Ang Li,
Ben Liu,
Bin Han,
Bin Hu,
Bin Jing,
Binbin Hu,
Bing Li,
Cai Chen,
Caizhi Tang,
Changxin Tian,
Chao Huang,
Chao Zhang,
Chen Liang,
Chen Qian,
Chengfu Tang,
Chengyao Wen,
Chilin Fu,
Chunwei Wu,
Cong Zhang,
Cunyin Peng,
Daixin Wang,
Dalong Zhang,
Deng Zhao,
Dingnan Jin,
Dingyuan Zhu
, et al. (193 additional authors not shown)
Abstract:
Efficient and scalable agentic intelligence requires models that can deliver both low-latency responses and strong reasoning capabilities while remaining practical to train, serve, and deploy. In this report, we present Ling-2.6 and Ring-2.6, a family of models designed to address this challenge at scale. Ling-2.6 is optimized for instant response generation and high capability per output token, w…
▽ More
Efficient and scalable agentic intelligence requires models that can deliver both low-latency responses and strong reasoning capabilities while remaining practical to train, serve, and deploy. In this report, we present Ling-2.6 and Ring-2.6, a family of models designed to address this challenge at scale. Ling-2.6 is optimized for instant response generation and high capability per output token, whereas Ring-2.6 is tailored for deeper reasoning and more advanced agentic workflows. Instead of training from scratch, we upgrade the Ling-2.0 base model through architectural migration pre-training and large-scale post-training. This upgrade is guided by a unified co-design of model architecture, optimization objectives, serving systems, and agent training environments, enabling improvements in both model capability and deployment efficiency. At the architectural level, we introduce a hybrid linear attention design that integrates Lightning Attention with MLA, improving the efficiency of long-context training and decoding. To further enhance token efficiency, we optimize capability per output token through Evolutionary Chain-of-Thought, Linguistic Unit Policy Optimization, bidirectional preference alignment, and shortest-correct-response distillation. For agentic capabilities, we propose KPop, a reinforcement learning framework designed to support stable training of Ring-2.6-1T on large-scale environment-grounded data. KPop improves training efficiency through asynchronous scheduling across coding, search, tool use, and workflow execution, enabling scalable learning from complex agent-environment interactions. Together, Ling-2.6 and Ring-2.6 provide a practical pathway toward efficient, scalable, and open agentic systems. We open-source all checkpoints in the 2.6 family to support further research and development in practical agentic intelligence.
△ Less
Submitted 12 June, 2026;
originally announced June 2026.
-
GeoRoPE: Ground-Aware Rotary Adaptation for Remote Sensing Foundation Models
Authors:
Yu Luo,
Kun Hu,
Mengwei He,
Xiaogang Zhu,
Shan Zeng,
Allen Benter,
Wei Xiang,
Patrick Filippi,
Thomas Francis Bishop,
Zhiyong Wang
Abstract:
Remote-sensing foundation models (RSFMs) benefit from pretraining on imagery from multiple sensors and ground sampling distances (GSDs), but such exposure alone does not resolve scale mismatch during downstream adaptation. A fixed token-grid offset can correspond to different ground distances across sensors, making grid-based positional priors physically inconsistent. Meanwhile, heterogeneous spat…
▽ More
Remote-sensing foundation models (RSFMs) benefit from pretraining on imagery from multiple sensors and ground sampling distances (GSDs), but such exposure alone does not resolve scale mismatch during downstream adaptation. A fixed token-grid offset can correspond to different ground distances across sensors, making grid-based positional priors physically inconsistent. Meanwhile, heterogeneous spatial granularity means that compact urban regions and homogeneous landscapes may require different positional sensitivities even under the same GSD. Therefore, we propose {GeoRoPE}, a ground-aware, RoPE-compatible, and parameter-efficient spatial adaptation method for RSFMs. GeoRoPE recalibrates token-level positional interactions from two complementary aspects. First, \textit{Geo-Coordinate Calibration (GCC)} rescales raw token-grid offsets according to the ground distance represented by one token-grid step, producing geo-calibrated relative coordinates across GSDs. Second, \textit{Geo-Frequency Calibration (GFC)} adjusts the native RoPE frequency with a relation-specific factor, enabling position sensitive adaptation to scene-dependent spatial granularity. GeoRoPE is injected into pretrained RSFMs through a lightweight adapter, preserving the frozen spatial prior while adding geo-aware positional corrections. Experiments across multiple RSFMs, sensors, resolutions, and downstream tasks demonstrate that GeoRoPE improves cross-resolution robustness and scale-sensitive representation learning.
△ Less
Submitted 8 June, 2026;
originally announced June 2026.
-
Block coordinate descent for joint delay-energy optimization in multi-hop D2D networks
Authors:
Kai-Xiang Hu,
Jacek Gondzio,
Caixia Kou
Abstract:
In multi-hop device-to-device (D2D) networks, the optimization of network-level metrics is particularly difficult due to the tight coupling between network-layer routing and physical-layer resource allocation. Departing from traditional average-performance metrics, this paper addresses the joint optimization of routing paths, transmission power, and bandwidth allocation. We formulate a generalized…
▽ More
In multi-hop device-to-device (D2D) networks, the optimization of network-level metrics is particularly difficult due to the tight coupling between network-layer routing and physical-layer resource allocation. Departing from traditional average-performance metrics, this paper addresses the joint optimization of routing paths, transmission power, and bandwidth allocation. We formulate a generalized cost function to minimize the maximum transmission time (i.e., the bottleneck delay) alongside the total energy consumption. To tackle the resulting highly non-convex formulation, we propose a novel block coordinate descent (BCD) framework. At the network layer, we develop two adaptive routing algorithms: a matrix-free Frank-Wolfe (MF-FW) algorithm for fast execution in dense topologies, and a low-rank primal-dual interior-point method (LR-PDIPM) that bypasses dense matrix inversions via the Sherman-Morrison formula for high-precision solutions. At the physical layer, we design a parallel dual ascent algorithm leveraging a time-domain perspective transformation to solve the resource allocation subproblem to global optimality. The proposed BCD framework is proven to converge to an ε-neighborhood of a stationary point. Through comprehensive experiments, the proposed BCD framework establishes its superiority in achieving the optimal delay-energy trade-off. Specifically, the LR-PDIPM variant achieves a maximum 9.14-fold reduction in total energy consumption and up to an order of magnitude improvement in energy efficiency, while maintaining a bounded maximum delay gap (up to 3.78-fold) relative to the best baseline. Meanwhile, the warm-start MF-FW variant identifies near-optimal solutions in mere seconds, serving as a highly practical engineering approach.
△ Less
Submitted 7 June, 2026;
originally announced June 2026.
-
GenPO++: Generative Policy Optimization with Jacobian-free Likelihood Ratios
Authors:
Ke Hu,
Shutong Ding,
Panxin Tao,
Jingya Wang,
Ye Shi
Abstract:
Generative policies provide expressive and multimodal action distributions, making them attractive for reinforcement learning (RL) in complex continuous-control tasks. Among them, flow-based policies are especially appealing because they generate actions through deterministic transport maps. However, applying such generative policies to likelihood-based on-policy learning remains limited by the di…
▽ More
Generative policies provide expressive and multimodal action distributions, making them attractive for reinforcement learning (RL) in complex continuous-control tasks. Among them, flow-based policies are especially appealing because they generate actions through deterministic transport maps. However, applying such generative policies to likelihood-based on-policy learning remains limited by the difficulty of evaluating the probability of executed actions. Existing flow RL methods either replace the true action-density ratio with approximate surrogates, which can introduce biased updates, or recover exact likelihoods through dummy-action augmentation, which enlarges the policy space and increases computation. In this work, we propose GenPO++, a reversible generative policy optimization framework that uses history states as auxiliary memory in a high-order reversible ODE solver, yielding exact inversion without changing the original action dimension. The resulting generative policy map has a log-determinant determined only by fixed solver coefficients, enabling exact and Jacobian-free likelihood-ratio computation. This design preserves the expressiveness of generative flow policies while avoiding both action ratio bias and dummy-action overhead. We evaluate GenPO++ on large-scale simulated control, fine-tuning, and real-world robotic manipulation tasks, where it achieves competitive or superior performance over state-of-the-art on-policy RL methods, while improving training stability and computational efficiency.
△ Less
Submitted 5 June, 2026;
originally announced June 2026.
-
JA-SIREN: Deterministic Initialization for Sinusoidal Networks via Spectral Matching
Authors:
Mohammed Alsakabi,
Kejia Hu,
John M. Dolan,
Ozan K. Tonguz
Abstract:
Existing implicit neural representation (INR) approaches suffer from stochastic initialization that does not guarantee consistent or high-quality performance across runs, with variations reaching more than 2.5 dB (78%) in image regression. This variation is problematic for scientific computing and simulation, where result reproducibility is crucial. To address this problem, we present Jacobi-Anger…
▽ More
Existing implicit neural representation (INR) approaches suffer from stochastic initialization that does not guarantee consistent or high-quality performance across runs, with variations reaching more than 2.5 dB (78%) in image regression. This variation is problematic for scientific computing and simulation, where result reproducibility is crucial. To address this problem, we present Jacobi-Anger Sinusoidal Representation Network (JA-SIREN), a deterministic initialization scheme for sinusoidal networks grounded in classical spectral analysis. By computing the Discrete Sine Transform (DST) of the target signal and leveraging the Jacobi-Anger expansion, we derive closed-form weights for a two-layer sinusoidal MLP that analytically match the network's initial spectral response to the target signal, requiring no random seed or additional hyperparameter tuning. On the Kodak dataset, JA-SIREN achieves a mean PSNR of 67.18 dB, a 21.30 dB improvement over the best baseline. This is achieved with zero run-to-run variance, confirming that spectrally-informed initialization is a more effective and reproducible alternative to stochastic initialization for sinusoidal INRs.
△ Less
Submitted 4 June, 2026;
originally announced June 2026.
-
Predictions for the X-ray polarisation modulation in Cygnus X-1 from reflection off the stellar companion and its wind
Authors:
Bert Vander Meulen,
Kun Hu,
Victoria Grinberg,
Henric Krawczynski
Abstract:
Context. Cyg X-1 is one of the brightest X-ray binaries and has been observed multiple times with the Imaging X-ray Polarimetry Explorer (IXPE). Recent studies report tentative evidence for a polarisation modulation with the orbital period P, but a half-period (P/2) signal, expected from reflection off the companion star and its stellar wind, has not been reported. Aims. We aim to quantify the ref…
▽ More
Context. Cyg X-1 is one of the brightest X-ray binaries and has been observed multiple times with the Imaging X-ray Polarimetry Explorer (IXPE). Recent studies report tentative evidence for a polarisation modulation with the orbital period P, but a half-period (P/2) signal, expected from reflection off the companion star and its stellar wind, has not been reported. Aims. We aim to quantify the reflection-induced variations of the polarisation degree PD and polarisation angle PA in Cyg X-1 as a function of orbital phase and energy, and interpret these in terms of binary geometry and wind structure. Methods. We set up a radiative transfer model combining a general relativistic description of the polarised source emission (kerrC) with a focussed stellar wind model for the binary medium. Using the 3D X-ray radiative transfer code SKIRT, we simulate broadband Stokes I, Q, and U fluxes, surface brightness maps, and linear polarisation maps over one binary orbit. Results. We find a prominent double-peaked (P/2) polarisation modulation, with a peak-to-peak PD amplitude of 0.25, 0.81, and 1.24 percentage points in the 2-4, 4-6, and 6-8 keV bands, respectively, with a strong energy dependence. The PA modulation is more modest, with |ΔPA| < 4.6°. Crucially, X-ray reprocessing reduces the overall PD relative to the source polarisation. Conclusions. The modulation is driven by reflection off the companion star and the focussed wind, which induces a polarisation signal that alternately reinforces and counteracts the source polarisation throughout the orbit. The diffuse scattering halo surrounding the source systematically reduces the PD, an effect that should be accounted for in all wind-fed XRBs. The PD amplitude increases with energy as absorption disproportionately attenuates the distant-reflection signal; as the extinction drops, the reflection signal becomes increasingly important.
△ Less
Submitted 2 June, 2026;
originally announced June 2026.
-
SRENet: Spectral Re-Entry Network for Point Cloud Action Recognition
Authors:
Qiuxia Wu,
Jiarui Lan,
Wenxiong Kang,
Zhiyong Wang,
Kun Hu
Abstract:
Recognizing human actions from point cloud sequences is critical for 3D perception driven applications such as autonomous driving and human-computer interaction. However, the irregular structure and temporal inconsistency of point clouds pose unique challenges for spatio-temporal representation learning, especially in capturing both global motion context and fine-grained temporal dynamics. We prop…
▽ More
Recognizing human actions from point cloud sequences is critical for 3D perception driven applications such as autonomous driving and human-computer interaction. However, the irregular structure and temporal inconsistency of point clouds pose unique challenges for spatio-temporal representation learning, especially in capturing both global motion context and fine-grained temporal dynamics. We propose SRENet, a spectral-aware framework designed to explicitly learn both global context and fine-grained temporal dynamics of motion from a frequency perspective for action recognition. SRENet introduces a Spectral Decomposition Block (SDeBlock) that performs wavelet-based analysis along temporal and spatial axes, disentangling features into low- and high-frequency components with frequency-specific attention. To recover residual dynamics and re-align temporal frequency structures distorted during semantic fusion, a Spectral Re-entry Block (SReBlock) performs secondary temporal decomposition. Furthermore, a spectral-aware learning strategy is devised to enhance discriminability in both frequency subspaces via contrastive loss and a curriculum schedule that gradually shifts focus from low- to high-frequency spaces in line with coarse to detailed motion patterns. Extensive experiments on MSR-Action3D, NTU-RGBD and NTU-RGBD120 demonstrate that SRENet achieves state-of-the-art performance, validating the effectiveness of frequency modeling in point cloud-based action understanding.
△ Less
Submitted 2 June, 2026;
originally announced June 2026.
-
Regulating EV Charging Markets for Fairness: Incentives for Pricing and Capacity Decisions
Authors:
Ruiting Wang,
Kita Hu,
Yitong Yu,
Scott Moura
Abstract:
The transition to electric mobility calls for charging infrastructure that is both efficient and socially equitable. This paper examines fairness in electric vehicle (EV) charging station pricing and capacity through a game-theoretic perspective. We model a non-cooperative market in which competing charging service providers set prices and capacities while customers choose stations based on genera…
▽ More
The transition to electric mobility calls for charging infrastructure that is both efficient and socially equitable. This paper examines fairness in electric vehicle (EV) charging station pricing and capacity through a game-theoretic perspective. We model a non-cooperative market in which competing charging service providers set prices and capacities while customers choose stations based on generalized cost, leading to a market equilibrium. We then benchmark this decentralized outcome against an idealized planner solution that jointly optimizes efficiency and equity. To align market outcomes with socially desirable goals, we design targeted incentives that guide operators toward more fair charger placement. Case studies demonstrate that unregulated competition tends to exacerbate disparities in charger access across demographic groups, whereas carefully calibrated incentives can reduce inequities without significant efficiency loss. The framework provides insights for policymakers on reconciling free-market dynamics with the broader societal goals of fairness in electrified mobility systems.
△ Less
Submitted 31 May, 2026;
originally announced June 2026.
-
MiCU: End-to-End Smart Home Command Understanding with Large Language Model
Authors:
Haowei Han,
Kexin Hu,
Weiwei Cai,
Debiao Zhang,
Bin Qin,
Yuxiang Wang,
Jiawei Jiang,
Xiao Yan,
Bo Du
Abstract:
Command understanding systems in smart home ecosystems can automate device control and substantially improve user experience. However, while they perform well on precise utterances (e.g., "turn on the bedroom light"), they struggle with ambiguous or misaligned commands (e.g., "make the bedroom cozy"). Large language models (LLMs) generalize well across various domains and can outperform traditiona…
▽ More
Command understanding systems in smart home ecosystems can automate device control and substantially improve user experience. However, while they perform well on precise utterances (e.g., "turn on the bedroom light"), they struggle with ambiguous or misaligned commands (e.g., "make the bedroom cozy"). Large language models (LLMs) generalize well across various domains and can outperform traditional rule-based systems on such tasks, but their effectiveness is often constrained by scarce domain-specific data, insufficient task-specific adaptation, and high computational costs. In this paper, we propose an automated training data synthesis workflow using user logs and LLMs; then we build MiCU, a domain-specific LLM that excels at command understanding. Specifically, we employ curriculum learning to inject domain knowledge into the base LLM, then we enhance its reasoning ability via cold-start training combined with reinforcement learning (RL) guided by domain-specific thinking rules. Additionally, we introduce a token compression technique that condenses device description into a single special token, substantially reducing inference overhead and enabling \model-fast, an efficient variant optimized for long inputs. Extensive experiments show that MiCU significantly outperforms baselines, with an average accuracy gain of 20.01% across all device categories. We have deployed MiCU in the Xiaomi Home app, receiving approximately 1.7 million page views per day. Production evaluations show that MiCU reduces user correction rate by 1.57% and increases human audited accuracy by 32.05%. Our data and code are available at https://github.com/xiaomi-research/iot_spec_llm
△ Less
Submitted 31 May, 2026;
originally announced June 2026.
-
TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents
Authors:
Weiyi Chen,
Shuaixiong Wang,
Ziyun Gao,
Kaichun Hu,
Wangze Ni,
Shimin Di,
Chen Jason Zhang,
Lei Chen
Abstract:
The development of Large Language Models (LLMs) has significantly improved travel planning applications, yet evaluating such models is limited by existing benchmarks' limitations: 1) overemphasis on constraint compliance, neglecting multi-dimensional qualities like spatio-temporal cost; 2) datasets lacking real-world authenticity and coverage in key areas (e.g., lodging, transport); and 3) isolate…
▽ More
The development of Large Language Models (LLMs) has significantly improved travel planning applications, yet evaluating such models is limited by existing benchmarks' limitations: 1) overemphasis on constraint compliance, neglecting multi-dimensional qualities like spatio-temporal cost; 2) datasets lacking real-world authenticity and coverage in key areas (e.g., lodging, transport); and 3) isolated daily plan assessments that miss critical details (e.g., the impact of daily accommodation and visit pacing) needed for entire plan's evaluation. To address this gap, we introduce TravelEval, a realistic and comprehensive benchmark. TravelEval features 1) a novel six-dimensional evaluation framework to holistically assess plans across accuracy, compliance, temporality, spatiality, economy, and utility dimensions; 2) a highly realistic data sandbox with precise accommodation pricing and authentic intercity transportation data; and 3) a simulation-based global evaluation method that emulates complete travel plans with API-integrated geographic information and fine-grained queuing time. Evaluating 12 mainstream approaches with TravelEval reveals several valuable insights, such that LLMs struggle with globally-optimized multi-dimensional planning (especially in spatio-temporal reasoning and budget compliance), and agentic reasoning strategies offer no consistent improvement. Concisely, TravelEval facilitates travel plan evaluation via grounded spatio-temporal emulation and comprehensive metrics, providing a robust foundation for advancing LLM-powered travel planning research and applications.
△ Less
Submitted 31 May, 2026;
originally announced June 2026.
-
Sample-Efficient Diffusion-based Reinforcement Learning with Critic Guidance
Authors:
Shutong Ding,
Zejia Zhong,
Zhongyi Wang,
Ke Hu,
Bikang Pan,
Jingya Wang,
Ye Shi
Abstract:
Recent advances in reinforcement learning (RL) have achieved great successes by leveraging the multimodality and exploration capability of diffusion policies. Among these approaches, one representative branch focuses on the sampling-based policy optimization. This design enables better exploration capability of the diffusion model, particularly at the beginning of training, but suffer from low exp…
▽ More
Recent advances in reinforcement learning (RL) have achieved great successes by leveraging the multimodality and exploration capability of diffusion policies. Among these approaches, one representative branch focuses on the sampling-based policy optimization. This design enables better exploration capability of the diffusion model, particularly at the beginning of training, but suffer from low exploitation in Q-value information, resulting in a slow policy convergence. Another branch pays attention to gradient-based policy optimization, which sufficiently exploits the gradient of the Q function yet tends to collapse into a unimodal policy with low diversity. To address this issue, we propose CGPO, \textbf{C}ritic-\textbf{G}uided diffusion \textbf{P}olicy \textbf{O}ptimization, which effectively balances exploration and exploitation with the training-free guidance technique integrated into the denoising process of diffusion policy. Concretely, CGPO steers action generation toward high-value regions defined by the critic network and uses the guided actions as regression objectives. In this manner, CGPO reduces the time required to obtain high-quality actions and improves final performance with better balance between the exploration-exploitation tradeoff. We validate the effectiveness of CGPO on 5 MuJoCo locomotion tasks, and CGPO achieves state-of-the-art performance compared with existing diffusion-based RL methods. Notably, CGPO is the first success to incorporate diffusion policy into real-world RL, with its superior performance on Franka robot arm grasping tasks. Our official page is released at https://dingsht.tech/cgpo-webpage.
△ Less
Submitted 28 May, 2026;
originally announced May 2026.
-
X-ray Polarization Signatures from Comptonization by Magnetic Reconnection Plasmoids
Authors:
John Groger,
Kun Hu,
Henric Krawczynski
Abstract:
Emission from X-ray binaries in the hard spectral state is dominated by high-energy radiation attributed to the Compton scattering of seed photons. The prevalent model of the Comptonization by hot electrons or pairs faces the problem of rapid radiative cooling of the emitting particles. A proposed alternative mechanism is the Comptonization by scattering off fast plasmoids formed during magnetic r…
▽ More
Emission from X-ray binaries in the hard spectral state is dominated by high-energy radiation attributed to the Compton scattering of seed photons. The prevalent model of the Comptonization by hot electrons or pairs faces the problem of rapid radiative cooling of the emitting particles. A proposed alternative mechanism is the Comptonization by scattering off fast plasmoids formed during magnetic reconnection. In this work, we simulate a simplified model of the plasmoid chain with Monte Carlo radiation transport and report on spectropolarimetric properties. We find that the Comptonization off trans-relativistic bulk plasmoids is not only able to reproduce the 100 keV spectral cutoff, but furthermore produces X-rays that are above 1 keV strongly polarized perpendicular to the reconnection layer. The polarization is stronger than that from the Comptonization by an isotropic hot plasma owing to the confinement of the motion of the scattering plasmoids in the plane of the reconnection layer. The dependence of polarization on azimuthal viewing angle is discussed, along with possible locations for the plasmoid chain in an equatorial current sheet or the sheath of the black hole's relativistic jet.
△ Less
Submitted 17 August, 2026; v1 submitted 25 May, 2026;
originally announced May 2026.
-
PhenoYieldNet: Learning Crop-Aware Phenological Responses for Multi-Crop Yield Prediction
Authors:
Yu Luo,
Xiaogang Zhu,
Shan Zeng,
Wei Xiang,
Thomas Francis Bishop,
Zhiyong Wang,
Kun Hu
Abstract:
Accurate crop yield prediction is crucial for sustainable agriculture and global food security. While existing methods are predominantly developed for single-crop prediction, they often struggle to generalize across diverse crop types, without addressing the unique crop phenological responses that are dynamically modulated by complex weather patterns. In this paper, we propose PhenoYieldNet, a mul…
▽ More
Accurate crop yield prediction is crucial for sustainable agriculture and global food security. While existing methods are predominantly developed for single-crop prediction, they often struggle to generalize across diverse crop types, without addressing the unique crop phenological responses that are dynamically modulated by complex weather patterns. In this paper, we propose PhenoYieldNet, a multi-crop yield prediction framework that learns crop-specific phenology by explicitly modeling their responses with temporal drivers. Specifically, we develop a crop-aware temporal decoder consisting of a Crop Phenology Bank (CPB) and a Crop Phenology Attention (CPA) module. The CPB integrates a set of learnable embeddings, which leverage a query to guide the CPA module to learn the most relevant phenology patterns for the specific crop. And the CPA module explicitly captures multi-scale trend and variation components to construct temporal contexts, enabling the model to dynamically adjust the attention across different phenological stages. To learn robust and generalizable features for multi-crop prediction, the encoder is initialized with a pre-trained foundation model, and further adapted via a self-supervised Temporal Contrastive Adaptation strategy to align with agricultural temporal dynamics. Extensive experiments conducted on multi-crop datasets indicate that our proposed method significantly outperforms state-of-the-art methods, exhibiting strong generalization capabilities across different regions and crops.
△ Less
Submitted 22 May, 2026;
originally announced May 2026.
-
Physics-Informed Generative Solver: Bridging Data-Driven Priors and Conservation Laws for Stable Spatiotemporal Field Reconstruction
Authors:
Ziyuan Zhu,
Keyu Hu,
Zhifei Chen,
Yuhao Shi,
Ming Bao,
Jing Zhao,
Gang Wang,
Haitan Xu,
Jiadong Li,
Qijun Zhao,
Xiaodong Li,
Minghui Lu,
Yanfeng Chen
Abstract:
Reconstructing continuous physical fields from sparse measurements is a central inverse problem, but data-driven generative models can produce states that violate governing dynamics. We introduce a physics-informed generative solver that separates stable prior learning from inference-time enforcement of conservation laws. Martingale-Regularized Score Matching regularizes score pretraining with a S…
▽ More
Reconstructing continuous physical fields from sparse measurements is a central inverse problem, but data-driven generative models can produce states that violate governing dynamics. We introduce a physics-informed generative solver that separates stable prior learning from inference-time enforcement of conservation laws. Martingale-Regularized Score Matching regularizes score pretraining with a Score Fokker-Planck constraint, yielding a dynamically stable prior. Physics-Informed Implicit Score Sampling then guides denoising trajectories by gradients of physical residuals, projecting samples toward admissible manifolds without retraining. In acoustics, the method co-generates pressure and particle velocity from sparse sensors, enabling dense virtual arrays that suppress spatial aliasing. The same framework generalizes to real-world ERA5 meteorological fields under extreme sparsity. Together, this work establishes a rigorous and generalizable paradigm for solving high-dimensional inverse problems, bridging the gap between generative artificial intelligence and first-principles science.
△ Less
Submitted 21 May, 2026;
originally announced May 2026.
-
Magnetar Fireballs and Short Bursts: Curved Spacetime Lensing, QED Effects, Spectra, Polarization, and Impulse Responses
Authors:
Zorawar Wadiasingh,
Hoa Dinh Thi,
Constantinos Kalapotharakos,
Kun Hu,
Matthew G. Baring,
Alice K. Harding,
George Younes,
Sebastien Guillot,
Andrea Sanna,
Michela Negro,
Jeremy D. Schnittman,
Oliver J. Roberts,
Eric Burns,
Chin-Ping Hu,
Ersin Göğüş
Abstract:
Magnetar short bursts (SBs) are hard X-ray transients of durations $0.01-1$ s peaking at $\sim 10-100$ keV, and are prime targets for new high-energy missions and polarimeters. The recent association of SBs with bright radio bursts in SGR 1935+2154 has broadened interest in SB physics. We present new advanced fireball models combining general relativistic light bending, polarized transport in magn…
▽ More
Magnetar short bursts (SBs) are hard X-ray transients of durations $0.01-1$ s peaking at $\sim 10-100$ keV, and are prime targets for new high-energy missions and polarimeters. The recent association of SBs with bright radio bursts in SGR 1935+2154 has broadened interest in SB physics. We present new advanced fireball models combining general relativistic light bending, polarized transport in magnetized photospheres, magnetic photon splitting attenuation, and magnetospheric vacuum birefringence. These models also have relevance to trapped fireballs in magnetar giant flare pulsating tails. We adopt confined flux tube geometries consistent with adiabatic fireballs, and anisotropic/polarized emergent intensities to produce spectra and polarizations, and energy-time Stokes impulse responses. We predict that most fireballs are highly linearly polarized, especially when vacuum birefringence is important. There is rich potential for diagnostics: coexisting direct and lensed delayed images, gaps by occultation of the neutron star surface, and Shapiro+Rømer delay with temporal caustics. These effects can imprint spin phase dependence of the spectral and polarization character of bursts. Predicted signatures depend strongly on viewing geometry, fireball configuration, and photon splitting assumptions, yielding large variance in model high-energy spectral shapes and cutoffs, and energy-dependent polarization. The models can reproduce established double-blackbody SB spectral phenomenology, and we find that the unusual April 2020 radio-associated SB from SGR 1935+2154 is broadly consistent with a footpoint close to the magnetic pole, and possibly near pole-on viewing geometry. Our models motivate reverberation-style analyses for SBs and suggest that high-quality data might constrain source geometry, burst crustal footpoints, and, potentially, neutron star masses and radii.
△ Less
Submitted 6 August, 2026; v1 submitted 21 May, 2026;
originally announced May 2026.
-
Predicting Performance of Symbolic and Prompt Programs with Examples
Authors:
Chengqi Zheng,
Keya Hu,
Shuzhi Liu,
Tao Wu,
Kevin Ellis,
Yewen Pu
Abstract:
LLM prompting is widely used for naturally stated tasks, yet it is unreliable it may succeed on a few test cases but fail at deployment time. We study performance prediction: given a program, either symbolic (e.g. Python) or a prompt executed on an LLM, and a few in-domain examples, predict its performance on unseen tasks from the same domain. We use a simple coin-flip model, treating each pass/fa…
▽ More
LLM prompting is widely used for naturally stated tasks, yet it is unreliable it may succeed on a few test cases but fail at deployment time. We study performance prediction: given a program, either symbolic (e.g. Python) or a prompt executed on an LLM, and a few in-domain examples, predict its performance on unseen tasks from the same domain. We use a simple coin-flip model, treating each pass/fail program execution as a Bernoulli random variable, whose success probability is the programs unknown performance. In this model, performance depends entirely on: 1) the observed execution outcomes on test cases, and 2) a prior over performances. We compile empirical performance priors from a corpus of diverse programs and tasks, and find that performance for symbolic programs (e.g., Python) are all or nothing, while prompt programs have a diffuse prior with many nearly-correct programs. This difference explains why a few passing tests can certify symbolic programs but not prompt programs. Building on this insight, we develop RAP (Retrieved Approximate Prior), which retrieves similar tasks and prompt programs from an existing corpus to construct a proxy prior, which is then used to predict performance. We show RAP achieves solid performances.
△ Less
Submitted 15 May, 2026;
originally announced May 2026.
-
You Can't Fool Us: Understanding the Resilience of LLM-driven Agent Communities to Misinformation
Authors:
Chichen Lin,
Yijie Jin,
Kangbo Hu,
Weijian Fan,
Han Xiao,
Yongbin Wang,
Zhihui Ying,
Zhanzhan Zhao
Abstract:
Misinformation resilience is a dynamic community process: communities differ not only in whether they initially trust false claims, but also in how they recover through interaction, questioning, correction, and support withdrawal. We study this process with an LLM-based agent simulation that constructs synthetic communities along two theoretically motivated dimensions: Actively Open-minded Thinkin…
▽ More
Misinformation resilience is a dynamic community process: communities differ not only in whether they initially trust false claims, but also in how they recover through interaction, questioning, correction, and support withdrawal. We study this process with an LLM-based agent simulation that constructs synthetic communities along two theoretically motivated dimensions: Actively Open-minded Thinking (AOT), which captures evidence-seeking and willingness to revise beliefs, and Political Ideology (PI), which captures identity-based interpretation of contested claims. These two traits allow us to examine how evidence-oriented reasoning and ideological alignment jointly shape community responses to credible misinformation shocks. Across systematically varied AOT-PI communities, we find that higher AOT improves both resistance to misinformation uptake and recovery after trust peaks. PI shapes the recovery pathway: ideologically moderate communities recover more reliably, while polarized communities retain more residual support. Stance-level analysis shows that resilience depends on whether agents move from questioning a claim to denying or correcting it and withdrawing prior support. Intervention experiments further show that persuasion and fact checking better support post-peak correction, whereas accuracy prompts mainly induce early caution and source warnings have weaker effects. Together, this work provides a mechanism-level account of community misinformation resilience, showing how psychological composition and intervention design shape whether communities move from misinformation exposure toward correction or persistent support.
△ Less
Submitted 17 May, 2026;
originally announced May 2026.
-
Stable Attention Response for Reliable Precipitation Nowcasting
Authors:
Penghui Wen,
Zexin Hu,
Sen Zhang,
Patrick Filippi,
Xiaogang Zhu,
Allen Benter,
Thomas Bishop,
Zhiyong Wang,
Kun Hu
Abstract:
Precipitation nowcasting remains challenging due to the highly localized, rapidly evolving, and heterogeneous nature of atmospheric dynamics. Although recent methods increasingly adopt attention-based architectures in both unimodal and multimodal settings, they mainly emphasize stronger representation learning and prediction capacity, while paying less attention to the stability of attention respo…
▽ More
Precipitation nowcasting remains challenging due to the highly localized, rapidly evolving, and heterogeneous nature of atmospheric dynamics. Although recent methods increasingly adopt attention-based architectures in both unimodal and multimodal settings, they mainly emphasize stronger representation learning and prediction capacity, while paying less attention to the stability of attention responses across samples. In this work, we show that cross-sample instability of attention-response energy is an important and previously underexplored source of forecasting unreliability. Empirically, inaccurate forecasts are associated with larger attention-response energy variance across heads and layers. Theoretically, we show that cross-sample variability can propagate through self-attention, and enlarge a lower bound on prediction error. Based on this insight, we propose HARECast, a Head-wise Attention Response Energy-regulated framework for precipitation nowcasting. HARECast explicitly models head-wise attention-response energy and stabilizes it through a group-wise regularization objective that reduces cross-sample fluctuations. The proposed formulation is generic and applicable to both unimodal and multimodal nowcasting architectures. We instantiate HARECast in a standard forecasting pipeline with reconstruction branches and a diffusion-based predictor, and evaluate it on commonly used benchmarks--SEVIR and MeteoNet. Experimental results demonstrate that HARECast achieves state-of-the-art performance.
△ Less
Submitted 5 August, 2026; v1 submitted 13 May, 2026;
originally announced May 2026.
-
ELF: Embedded Language Flows
Authors:
Keya Hu,
Linlu Qiu,
Yiyang Lu,
Hanhong Zhao,
Tianhong Li,
Yoon Kim,
Jacob Andreas,
Kaiming He
Abstract:
Diffusion and flow-based models have become the de facto approaches for generating continuous data, e.g., in domains such as images and videos. Their success has attracted growing interest in applying them to language modeling. Unlike their image-domain counterparts, today's leading diffusion language models (DLMs) primarily operate over discrete tokens. In this paper, we show that continuous DLMs…
▽ More
Diffusion and flow-based models have become the de facto approaches for generating continuous data, e.g., in domains such as images and videos. Their success has attracted growing interest in applying them to language modeling. Unlike their image-domain counterparts, today's leading diffusion language models (DLMs) primarily operate over discrete tokens. In this paper, we show that continuous DLMs can be made effective with minimal adaptation to the discrete domain. We propose Embedded Language Flows (ELF), a class of diffusion models in continuous embedding space based on continuous-time Flow Matching. Unlike existing DLMs, ELF predominantly stays within the continuous embedding space until the final time step, where it maps to discrete tokens using a shared-weight network. This formulation makes it straightforward to adapt established techniques from image-domain diffusion models, e.g., classifier-free guidance (CFG). Experiments show that ELF substantially outperforms leading discrete and continuous DLMs, achieving better generation quality with fewer sampling steps. These results suggest that ELF offers a promising path toward effective continuous DLMs.
△ Less
Submitted 25 June, 2026; v1 submitted 11 May, 2026;
originally announced May 2026.
-
KAN Text to Vision? The Exploration of Kolmogorov-Arnold Networks for Multi-Scale Sequence-Based Pose Animation from Sign Language Notation
Authors:
Guanyi Du,
Lintao Wang,
Kun Hu,
Ziyang Wang
Abstract:
Sign language production from symbolic notation offers a scalable route to accessible sign animation. We present KANMultiSign, a multi-scale sequence generator that translates HamNoSys notation into two-dimensional human pose sequences. Our framework makes two complementary contributions. First, we introduce a coarse-to-fine generation strategy with multi-scale supervision: the model is first guid…
▽ More
Sign language production from symbolic notation offers a scalable route to accessible sign animation. We present KANMultiSign, a multi-scale sequence generator that translates HamNoSys notation into two-dimensional human pose sequences. Our framework makes two complementary contributions. First, we introduce a coarse-to-fine generation strategy with multi-scale supervision: the model is first guided by an intermediate body--hand--face scaffold to encourage global structural coherence, and then refines fine-grained hand articulation to improve finger-level detail. Second, we investigate integrating Kolmogorov--Arnold Network modules into a Transformer backbone, using learnable univariate function primitives to model the highly non-linear mapping from discrete phonological symbols to continuous body kinematics with a compact parameterization. Experiments on multiple public corpora spanning Polish, German, Greek, and French sign languages show consistent reductions in dynamic time warping based joint error compared with a strong notation-to-pose baseline, while using substantially fewer parameters. Controlled ablations further indicate that KAN-based variants substantially reduce parameter count while maintaining competitive performance when coupled with multi-scale supervision, rather than serving as the main driver of accuracy gains. These findings position multi-scale supervision as the key mechanism for improving notation-conditioned pose generation, with KAN offering a compact alternative for efficient modeling. Our code will be publicly available.
△ Less
Submitted 10 May, 2026;
originally announced May 2026.
-
ECTraj: Enhanced Consistency Training for Multi-Agent Trajectory Prediction
Authors:
Alen Mrdovic,
Qingze,
Liu,
Danrui Li,
Mathew Schwartz,
Kaidong Hu,
Sejong Yoon,
Mubbasir Kapadia,
Vladimir Pavlovic
Abstract:
Diffusion models for multi-agent trajectory prediction are limited by iterative denoising, which causes inference latency that hinders their use in time-critical settings like autonomous driving. Fast-sampling variants using DDIM and informed initial noise distributions partially alleviate this issue, but they either fail to achieve true single-step generation or are constrained by the chosen nois…
▽ More
Diffusion models for multi-agent trajectory prediction are limited by iterative denoising, which causes inference latency that hinders their use in time-critical settings like autonomous driving. Fast-sampling variants using DDIM and informed initial noise distributions partially alleviate this issue, but they either fail to achieve true single-step generation or are constrained by the chosen noise distribution. Consistency Models (CMs) offer high-quality one-step generation by mapping noise directly to data, but are difficult to train from scratch. We propose ECTraj, an enhanced CM pipeline with improved training and conditional generation for trajectory prediction. Our framework extends the student-teacher consistency training scheme: the student produces standard outputs, while the teacher explicitly fuses its predictions with parts of the ground truth to give stronger supervision. We also exploit CMs' direct denoising for top-K multi-shot generation during training. Combining conditional generation with this enhanced consistency objective yields faster inference and improved prediction accuracy, establishing competitive new benchmarks on the large-scale Argoverse 2 dataset.
△ Less
Submitted 5 July, 2026; v1 submitted 8 May, 2026;
originally announced May 2026.
-
Resonant Inverse Compton Scattering and Hard X-ray Emission in Magnetar Magnetospheres
Authors:
Kun Hu,
Nicholas Rackers,
Alexander Y. Chen
Abstract:
Magnetars are a subclass of neutron stars with ultra-strong surface magnetic fields. Some magnetars exhibit persistent hard X-ray emission, characterized by power-law tails with photon indices around 1--1.5, extending from ${\sim}$10 keV to several hundred keV. The leading explanation for this hard X-ray component is resonant Compton scattering, in which the thermal seed photons are upscattered by…
▽ More
Magnetars are a subclass of neutron stars with ultra-strong surface magnetic fields. Some magnetars exhibit persistent hard X-ray emission, characterized by power-law tails with photon indices around 1--1.5, extending from ${\sim}$10 keV to several hundred keV. The leading explanation for this hard X-ray component is resonant Compton scattering, in which the thermal seed photons are upscattered by relativistic electron-positron pairs flowing along magnetic field lines in the magnetosphere. In this work, we adopt the pair outflow framework of the magnetar magnetosphere and calculate the resonant Compton scattering opacity, as well as the spectrum and polarization of the upscattered emission. We find that resonant cooling can substantially modify the magnetospheric plasma density and impose strong optical depth constraints on the hard X-ray emission regions. Under the viewing geometry inferred from IXPE, an equatorial twist near the stellar surface provides a viable configuration for the NuSTAR hard X-ray spectrum of 4U 0142+61, while a polar-twist geometry is disfavored. Joint spectral, timing, and polarimetric modeling will be essential for distinguishing between the magnetospheric scattering geometries and understanding the physical properties of the pair plasma.
△ Less
Submitted 6 August, 2026; v1 submitted 8 May, 2026;
originally announced May 2026.
-
Strong coupling between quantized magnon modes in a YIG microstucture and microwaves in a superconducting resonator
Authors:
Seth W. Kurfman,
Philipp Geyer,
Anoop Kamalasanan,
Karl Heimrich,
Kwangyul Hu,
Tharnier O. Puel,
Frank Heyroth,
Michael Flatté,
Georg Schmidt
Abstract:
Strong-coupling experiments based on magnons enable the exploration into on-chip demonstrations involving numerous long-lived excitations. Yttrium iron garnet (YIG) has been considered for decades as a gold standard material for magnonics due to its low-loss magnonic properties. While YIG has successfully demonstrated strong-coupling in macroscopic device geometries, the strong coupling of magnons…
▽ More
Strong-coupling experiments based on magnons enable the exploration into on-chip demonstrations involving numerous long-lived excitations. Yttrium iron garnet (YIG) has been considered for decades as a gold standard material for magnonics due to its low-loss magnonic properties. While YIG has successfully demonstrated strong-coupling in macroscopic device geometries, the strong coupling of magnons in truly sub-10 micron YIG structures to date has not yet been realized. This obstacle is due to the difficulty producing large enough effective magnonic mode volume necessary primarily due to thickness limitations of YIG deposition and device fabrication techniques. Here, we demonstrate the use of a microplatelet of YIG, manufactured from a single crystal of YIG via focused ion beam (FIB) techniques, placed on a constricted inductive line of an optimized superconducting lumped element LC resonator to achieve strong coupling between numerous magnon modes and the LC resonator photons. These experimental findings are qualitatively backed by micromagnetic simulations and quantitatively supported by analytical calculations to identify the magnon modes corresponding to the experimentally observed anti-crossings in the microwave transmission signal. Further, we show that these anti-crossings remain even at incredibly low device input powers ($\leq 10$ fW). The fabrication techniques and device geometry enable the deterministic use of numerous confined magnon modes in micron-scale YIG structures for various magnetic field strengths and orientations at substantially reduced device powers. The results here establish a foundational path forward to achieving efficient magnon-based strong-coupling experiments in micron-scale YIG magnetic elements for effective on-chip studies.
△ Less
Submitted 11 May, 2026; v1 submitted 30 April, 2026;
originally announced April 2026.
-
From Hypotheses to Factors: Constrained LLM Agents in Cryptocurrency Markets
Authors:
Yikuan Huang,
Zheqi Fan,
Kaiqi Hu,
Yifan Ye
Abstract:
LLM agents are promising tools for empirical discovery, but their flexibility can also turn discovery into uncontrolled search. We study how to use agents under a reproducible protocol through cryptocurrency factor discovery. Our framework casts the task as sequential hypothesis search: an agent reads an append-only experiment trace, proposes falsifiable factor hypotheses, and maps them to executa…
▽ More
LLM agents are promising tools for empirical discovery, but their flexibility can also turn discovery into uncontrolled search. We study how to use agents under a reproducible protocol through cryptocurrency factor discovery. Our framework casts the task as sequential hypothesis search: an agent reads an append-only experiment trace, proposes falsifiable factor hypotheses, and maps them to executable recipes, while a deterministic engine enforces fixed data splits, selection gates, transaction costs, and portfolio tests. Candidate actions are restricted to a point-in-time factor DSL, making both successful and failed hypotheses auditable. A ridge-combined portfolio trained only on 2020--2022 data achieves a 44.55% annualized return and Sharpe ratio of 1.55 in the 2024--2026 pure out-of-sample period after a 5 basis point one-way trading cost.
△ Less
Submitted 29 April, 2026;
originally announced April 2026.
-
When AI reviews science: Can we trust the referee?
Authors:
Jialiang Wang,
Yuchen Liu,
Hang Xu,
Kaichun Hu,
Shimin Di,
Wangze Ni,
Linan Yue,
Min-Ling Zhang,
Kui Ren,
Lei Chen
Abstract:
The volume of scientific submissions continues to climb, outpacing the capacity of qualified human referees and stretching editorial timelines. At the same time, modern large language models (LLMs) offer impressive capabilities in summarization, fact checking, and literature triage, making the integration of AI into peer review increasingly attractive -- and, in practice, unavoidable. Yet early de…
▽ More
The volume of scientific submissions continues to climb, outpacing the capacity of qualified human referees and stretching editorial timelines. At the same time, modern large language models (LLMs) offer impressive capabilities in summarization, fact checking, and literature triage, making the integration of AI into peer review increasingly attractive -- and, in practice, unavoidable. Yet early deployments and informal adoption have exposed acute failure modes. Recent incidents have revealed that hidden prompt injections embedded in manuscripts can steer LLM-generated reviews toward unjustifiably positive judgments. Complementary studies have also demonstrated brittleness to adversarial phrasing, authority and length biases, and hallucinated claims. These episodes raise a central question for scholarly communication: when AI reviews science, can we trust the AI referee? This paper provides a security- and reliability-centered analysis of AI peer review. We map attacks across the review lifecycle -- training and data retrieval, desk review, deep review, rebuttal, and system-level. We instantiate this taxonomy with four treatment-control probes on a stratified set of ICLR 2025 submissions, using two advanced LLM-based referees to isolate the causal effects of prestige framing, assertion strength, rebuttal sycophancy, and contextual poisoning on review scores. Together, this taxonomy and experimental audit provide an evidence-based baseline for assessing and tracking the reliability of AI peer review and highlight concrete failure points to guide targeted, testable mitigations.
△ Less
Submitted 26 April, 2026;
originally announced April 2026.
-
JoyAI-RA 0.1: A Foundation Model for Robotic Autonomy
Authors:
Tianle Zhang,
Zhihao Yuan,
Dafeng Chi,
Peidong Liu,
Dongwei Li,
Kejun Hu,
Likui Zhang,
Junnan Nie,
Ziming Wei,
Zengjue Chen,
Yili Tang,
Jiayi Li,
Zhiyuan Xiang,
Mingyang Li,
Tianci Luo,
Hanwen Wan,
Ao Li,
Linbo Zhai,
Zhihao Zhan,
Xiaodong Bai,
Jiakun Cai,
Peng Cao,
Kangliang Chen,
Siang Chen,
Yixiang Dai
, et al. (37 additional authors not shown)
Abstract:
Robotic autonomy in open-world environments is fundamentally limited by insufficient data diversity and poor cross-embodiment generalization. Existing robotic datasets are often limited in scale and task coverage, while relatively large differences across robot embodiments impede effective behavior knowledge transfer. To address these challenges, we propose JoyAI-RA, a vision-language-action (VLA)…
▽ More
Robotic autonomy in open-world environments is fundamentally limited by insufficient data diversity and poor cross-embodiment generalization. Existing robotic datasets are often limited in scale and task coverage, while relatively large differences across robot embodiments impede effective behavior knowledge transfer. To address these challenges, we propose JoyAI-RA, a vision-language-action (VLA) embodied foundation model tailored for generalizable robotic manipulation. JoyAI-RA presents a multi-source multi-level pretraining framework that integrates web data, large-scale egocentric human manipulation videos, simulation-generated trajectories, and real-robot data. Through training on heterogeneous multi-source data with explicit action-space unification, JoyAI-RA effectively bridges embodiment gaps, particularly between human manipulation and robotic control, thereby enhancing cross-embodiment behavior learning. JoyAI-RA outperforms state-of-the-art methods in both simulation and real-world benchmarks, especially on diverse tasks with generalization demands.
△ Less
Submitted 23 April, 2026; v1 submitted 21 April, 2026;
originally announced April 2026.
-
Explainable AML Triage with LLMs: Evidence Retrieval and Counterfactual Checks
Authors:
Dorothy Torres,
Wei Cheng,
Ke Hu
Abstract:
Anti-money laundering (AML) transaction monitoring generates large volumes of alerts that must be rapidly triaged by investigators under strict audit and governance constraints. While large language models (LLMs) can summarize heterogeneous evidence and draft rationales, unconstrained generation is risky in regulated workflows due to hallucinations, weak provenance, and explanations that are not f…
▽ More
Anti-money laundering (AML) transaction monitoring generates large volumes of alerts that must be rapidly triaged by investigators under strict audit and governance constraints. While large language models (LLMs) can summarize heterogeneous evidence and draft rationales, unconstrained generation is risky in regulated workflows due to hallucinations, weak provenance, and explanations that are not faithful to the underlying decision. We propose an explainable AML triage framework that treats triage as an evidence-constrained decision process. Our method combines (i) retrieval-augmented evidence bundling from policy/typology guidance, customer context, alert triggers, and transaction subgraphs, (ii) a structured LLM output contract that requires explicit citations and separates supporting from contradicting or missing evidence, and (iii) counterfactual checks that validate whether minimal, plausible perturbations lead to coherent changes in both the triage recommendation and its rationale. We evaluate on public synthetic AML benchmarks and simulators and compare against rules, tabular and graph machine-learning baselines, and LLM-only/RAG-only variants. Results show that evidence grounding substantially improves auditability and reduces numerical and policy hallucination errors, while counterfactual validation further increases decision-linked explainability and robustness, yielding the best overall triage performance (PR-AUC 0.75; Escalate F1 0.62) and strong provenance and faithfulness metrics (citation validity 0.98; evidence support 0.88; counterfactual faithfulness 0.76). These findings indicate that governed, verifiable LLM systems can provide practical decision support for AML triage without sacrificing compliance requirements for traceability and defensibility.
△ Less
Submitted 6 June, 2026; v1 submitted 22 March, 2026;
originally announced April 2026.
-
Cross-Stock Predictability via LLM-Augmented Semantic Networks
Authors:
Yikuan Huang,
Zheqi Fan,
Kaiqi Hu,
Yifan Ye
Abstract:
Text-based financial networks are increasingly used to study cross-stock return predictability. A common approach constructs links from similarities in firms' disclosure embeddings, but such networks often contain spurious edges because textual proximity does not necessarily imply economic connection. We propose a two-stage framework that first builds a sparse candidate graph from 10-K embeddings…
▽ More
Text-based financial networks are increasingly used to study cross-stock return predictability. A common approach constructs links from similarities in firms' disclosure embeddings, but such networks often contain spurious edges because textual proximity does not necessarily imply economic connection. We propose a two-stage framework that first builds a sparse candidate graph from 10-K embeddings and then uses a large language model to classify and filter candidate edges according to their economic relations. The refined graph is used to aggregate pair-level mean-reversion signals into stock-level trading signals with relation-aware and distance-based weights. In a backtest on S&P 500 constituents from 2011 to 2019, LLM-based edge filtering improves the long-short Sharpe ratio from 0.742 to 0.820 and reduces maximum drawdown from $-$10.47% to $-$7.85%. These results suggest that LLM-based reasoning can improve the economic fidelity of text-derived financial networks and strengthen cross-stock predictability.
△ Less
Submitted 26 April, 2026; v1 submitted 21 April, 2026;
originally announced April 2026.