Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 579 results for author: Hu, K

.
  1. arXiv:2608.18417  [pdf, ps, other

    stat.ME

    Centroid-Referenced Mahalanobis Matching (CRM): A Scalable, Representation-Based Framework for Causal Inference in Large Observational Studies

    Authors: Keming Hu, Yingpei He

    Abstract: Matching for causal inference can be computationally expensive at scale and can silently change the target population when overlap is limited. We propose Centroid-Referenced Mahalanobis Matching (CRM), which replaces global pairwise search with stratified sampling in two reference coordinates: each unit's Mahalanobis distance from the treated centroid and its Fisher coordinate along the treated-co… ▽ More

    Submitted 18 August, 2026; originally announced August 2026.

    Comments: Accepted by Transactions on Machine Learning Research (TMLR), 2026. Code: https://github.com/KemingHu-D/crm-matching

    Journal ref: Transactions on Machine Learning Research, 2026

  2. arXiv:2608.18278  [pdf

    cond-mat.mtrl-sci

    Writing and erasing skyrmions by single ultrafast laser pulses in monolayer Janus 2D magnets

    Authors: Guangyao Miao, Yonglong Ga, Chang Liu, Pan Chen, Yichen Jin, Florian Kronast, Wenxin Cheng, Zhaoqing Ding, Kai Hu, Zongnan Zhang, Nikolai Severin, Chenxi Meng, Patil Shubhada, Sergio Valencia, Meng Meng, Qinlin Guo, Xiaoran Liu, Jiandi Zhang, Yangmu Li, Carlos-Andres Palma, Jürgen P. Rabe, Hongxin Yang, Weihua Wang, Jiandong Guo

    Abstract: Skyrmions in 2D magnets are promising candidates for nonvolatile, low-power, and high-density spintronic memories. However, their experimental realization at the 2D limit remains challenging, owing to the difficulty in engineering the required chiral magnetic interactions. Here, we report the creation and direct imaging of Néel-type skyrmions in Janus 2D chromium chalcogenides using synchrotron X-… ▽ More

    Submitted 18 August, 2026; originally announced August 2026.

  3. arXiv:2608.13831  [pdf, ps, other

    eess.AS cs.CL

    VoiceChat-TTS: A Low-Latency Continuous Speech Synthesis Model for Interactive Agents

    Authors: Edresson Casanova, Jaehyeon Kim, Mariana Graterol Fuenmayor, Shehzeen Hussain, Viacheslav Klimkov, Valentin Mendelev, Mikyas Desta, Paarth Neekhara, Piotr Zelasko, Chen Chen, Elena Rastorgueva, Ke Hu, Ankita Pasad, Xuesong Yang, Aya Alja'fari, Rajarshi Roy, Rohan Badlani, Jason Roche, Jason Li, Zhehuai Chen

    Abstract: Spoken dialogue is a natural form of human--computer interaction, yet most speech language models remain limited to turn-based operation and lack real-time adaptability, such as user barge-in. Recent duplex speech-to-speech and speech-to-text models reduce latency by replacing multi-stage pipelines, but often compromise speech quality because accurate ASR, interruption handling, and high-fidelity… ▽ More

    Submitted 13 August, 2026; originally announced August 2026.

  4. arXiv:2608.11587  [pdf, ps, other

    eess.AS cs.CL cs.LG

    Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning

    Authors: Xulin Fan, Jialu Li, Mohammad Nur Hossain Khan, Kexin Hu, Bashima Islam, Mark Hasegawa-Johnson, Nancy L. McElwain

    Abstract: Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challenging due to limited labeled data, low signal-to-noise ratio, and cross-family domain shifts. We present a family-conditioned, multi-tier audio tagger that combines a LoRA-finetuned Whisper encoder with a lightweight, targe… ▽ More

    Submitted 11 August, 2026; originally announced August 2026.

    Comments: Accepted to Interspeech 2026

  5. arXiv:2608.08659  [pdf, ps, other

    cs.CV

    JSGS: JPEG State-Guided Supervision for 3D Gaussian Splatting from Mixed-Quality Views

    Authors: Jinhua Cui, Anhong Wang, Kai Hu, Donghan Bu, Peihao Li, Tammam Tillo, Hao Jing, Shiao Xu

    Abstract: Standard 3D Gaussian Splatting (3DGS) assumes that every input image faithfully samples scene radiance. However, mixed-quality JPEG images violate this assumption because compression-induced blocking and ringing artifacts can corrupt updates to Gaussians shared across views. To address this problem, we propose JPEG State-Guided Supervision for 3D Gaussian Splatting from Mixed-Quality Views (JSGS).… ▽ More

    Submitted 9 August, 2026; originally announced August 2026.

  6. arXiv:2608.05188  [pdf, ps, other

    cs.CL cs.AI

    Position: It's Time to Optimize LLMs for Self-Consistency

    Authors: Itamar Pres, Belinda Z. Li, Laura Ruis, Zifan Carl Guo, Keya Hu, Mehul Damani, Isha Puri, Ekdeep Singh Lubana, Jacob Andreas

    Abstract: Despite ever-increasing sophistication in language model (LM) pre- and post-training pipelines, many important failures persist: models overcondition on user framing ("sycophancy"), exhibit incomplete logical generalization, and produce confident but incorrect responses. We argue that these failures arise from a modeling assumption permeating all aspects of the pipeline: that behavior can be speci… ▽ More

    Submitted 31 July, 2026; originally announced August 2026.

    Comments: Accepted at the 43rd International Conference on Machine Learning (ICML 2026), Position Paper Track

  7. arXiv:2607.27180  [pdf, ps, other

    cs.CV cs.RO

    HumanCLAW: Can Vision-Language Models Act Through a Body?

    Authors: Li Siyao, Jiawei Gu, Shuai Liu, Kairui Hu, Zekun Li, Linjie Li, Chengcheng Tang, Po-Chen Wu, Ivan Shugurov, Lingni Ma, Michael Zollhoefer, Sizhe An, Abhay Mittal, Amy Zhao, Ranjay Krishna, Manling Li, Ziwei Liu, Chuan Guo

    Abstract: Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed to execute it, e.g., losing balance and falling. In this work, we introduce HumanCLAW, an evaluation framework that decou… ▽ More

    Submitted 3 August, 2026; v1 submitted 29 July, 2026; originally announced July 2026.

    Comments: Project page: https://human-claw.github.io/

  8. arXiv:2607.22932  [pdf, ps, other

    math.NA

    Optimal block preconditioners for a mass-conserving mixed stress formulation of Stokes flow

    Authors: Kaibo Hu, Jongho Park, Jindong Wang

    Abstract: We present optimal block diagonal and triangular preconditioners for a mass-conserving mixed stress formulation of Stokes flow. The algebraic formulation leads to a double saddle point system with unknowns corresponding to discrete stress, velocity, vorticity, and pressure. MINRES equipped with a block diagonal preconditioner for an augmented Lagrangian formulation of this system is analyzed and s… ▽ More

    Submitted 24 July, 2026; originally announced July 2026.

    MSC Class: 65F08; 65N22; 65N30; 76D07

  9. arXiv:2607.22478  [pdf, ps, other

    math.NA

    ReLU$^k$ Neural de Rham Complexes

    Authors: Kaibo Hu, Jindong Wang, Jinchao Xu

    Abstract: We construct finite-dimensional de Rham subcomplexes generated by fixed-neuron shallow ReLU$^k$ neural networks, a class of spaces known to provide optimal approximation rates. For neurons of the form $s_i(x)=ω_i\cdot x+b_i$, we introduce spaces of neural differential forms: differential $p$-forms whose coefficients are the ReLU$^k$ ridge functions $σ_{k-p}(s_i)$. These spaces are compatible with… ▽ More

    Submitted 24 July, 2026; originally announced July 2026.

    MSC Class: 41A30; 41A46; 55U15; 58A10; 65N25; 65N30; 68T07

  10. arXiv:2607.16401  [pdf, ps, other

    cs.CV

    Apple-$π$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

    Authors: Runmao Yao, Kairui Hu, Yukang Cao, Ruisi Wang, Shulin Tian, Ziang Cao, Weichen Fan, Ziqi Huang, Yuhao Dong, Hao Li, Zhaoxi Chen, Zhongang Cai, Lei Yang, Ziwei Liu

    Abstract: Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without verifying whether the model arrives there through a faithful, law-grounded reasoning process. We introduce Apple-PI, the first benchmark that anchors video-model evaluation explic… ▽ More

    Submitted 17 July, 2026; originally announced July 2026.

  11. arXiv:2607.14642  [pdf, ps, other

    cs.AI cs.SE

    MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers

    Authors: Huanxi Liu, Kun Hu, Jiaqi Liao, Qiang Wang, Pengfei Qian, YuanZhao Zhai, Dawei Feng, Bo Ding, Huaimin Wang

    Abstract: As Model Context Protocol (MCP) servers emerge as the core infrastructure for connecting LLMs with external tools, existing benchmarks leverage real-world MCP servers to evaluate LLM agents' tool-using capabilities. However, these benchmarks overlook the continuous evolution of tool interfaces and functionalities within MCP servers, resulting in flawed assessments that fail to capture the agent's… ▽ More

    Submitted 16 July, 2026; originally announced July 2026.

  12. arXiv:2607.13965  [pdf, ps, other

    cs.SE cs.CR

    ProfMalPlus: Agent-Coordinated Detection of Malicious NPM Packages via Static-Dynamic Analysis Synergy

    Authors: Yiheng Huang, Zhijia Zhao, Bihuan Chen, Susheng Wu, Zhuotong Zhou, Yiheng Cao, Kun Hu, Xin Hu, Xin Peng

    Abstract: Open source software is vulnerable to supply-chain attacks through transitive dependencies, especially malicious code injected into NPM packages. Existing detectors often inadequately model obfuscated behavior, overlook JavaScript's object-centric features, poorly coordinate static and dynamic analysis, and lose semantic information during behavior abstraction. We propose ProfMalPlus, a malicious… ▽ More

    Submitted 15 July, 2026; originally announced July 2026.

  13. arXiv:2607.11215  [pdf, ps, other

    cs.CL cs.MM

    Q-BridgeNet: A Quantization Network for Cross-Lingual Sign Language Translation

    Authors: Liqian Feng, Lintao Wang, Xiaochen Liu, Anusha Withana, Ken-Tye Yong, Dehui Kong, Zhiyong Wang, Kun Hu

    Abstract: Most sign language translation (SLT) methods focus on isolated native sign-spoken pairs (e.g., American Sign Language - English). Extending language-specific SLT models to multilingual translation would improve accessibility by enabling communication across diverse sign and spoken language communities. However, existing multilingual SLT approaches still struggle to learn a unified model that minim… ▽ More

    Submitted 13 July, 2026; originally announced July 2026.

  14. A Multi-Frequency Input-Admittance Model of Locomotive Rectifier Considering PWM Sideband Harmonic Coupling in Electrical Railways

    Authors: Xiangyu Meng, Zhigang Liu, Guorong Li, Xunjun Chen, Siqi Wu, Keting Hu

    Abstract: Electrical railway harmonic instability issues are common in the high-frequency range. The effective frequency of the traditional converter's small-signal averaging model is below 1/2 switching frequency since the pulse width modulation (PWM) sideband harmonic components are ignored. In this article, the dynamic propagations of perturbation frequency and the generated PWM sideband components are c… ▽ More

    Submitted 10 July, 2026; originally announced July 2026.

    Comments: 11 pages. Accepted manuscript

    Journal ref: IEEE Transactions on Transportation Electrification, vol. 8, no. 3, pp. 3848-3858, 2022

  15. arXiv:2607.02983  [pdf, ps, other

    cs.AI

    Reinforcement Learning for Evidence-Seeking Diagnostic Reasoning with Large Language Models

    Authors: Shengyi Hua, Kangzhe Hu, Conghui He, Xiaofan Zhang, Shaoting Zhang

    Abstract: Recent reasoning-centric Large Language Models (LLMs) have made significant strides, yet they predominantly operate on a passive-inference pattern that assumes complete information. In contrast, real-world clinical intelligence is inherently an iterative investigative process requiring strategic evidence acquisition. To bridge this gap, we formalize medical diagnosis as an Iterative Evidence-Seeki… ▽ More

    Submitted 3 July, 2026; originally announced July 2026.

  16. arXiv:2607.00174  [pdf, ps, other

    cs.CV cs.LG

    Steal the Patch Size: Adversarially Manipulate Vision-Language Models

    Authors: Kai Hu, Akash Bharadwaj, Weichen Yu, Matt Fredrikson

    Abstract: We present a black-box model-stealing attack that recovers private vision-tokenizer configurations of deployed vision-language models (VLMs), including the visual patch size and input preprocessing pipeline. The key idea is a task-level side channel induced by ViT-style patchification: when a synthetic grid image is aligned with the hidden patch grid, boundary cues are erased at tokenization, caus… ▽ More

    Submitted 30 June, 2026; originally announced July 2026.

    Journal ref: ICML 2026

  17. arXiv:2606.30894  [pdf, ps, other

    astro-ph.HE

    Energy-Resolved Limits on Orbital X-ray Polarization Modulation in Cygnus X-1

    Authors: Sohee Chun, Bert Vander Meulen, Kun Hu, Henric Krawczynski

    Abstract: Reflection off the companion star and its focused stellar wind is predicted to modulate the X-ray polarization of black hole X-ray binaries at half the orbital period ($P_{\rm orb}/2$), with an energy-dependent amplitude. We test this prediction against all publicly available IXPE observations of Cygnus X-1, comprising 26 one-day bins from 12 observation IDs spanning 2022-2024. Since the normalize… ▽ More

    Submitted 29 June, 2026; originally announced June 2026.

  18. arXiv:2606.30102  [pdf, ps, other

    quant-ph

    Enhanced Magnon Synchronization in Coupled WGM Optomagnonic Resonators with Phase-Dependent Photon Hopping

    Authors: Le-Ji Xue, Ying-Jian Zhu, Jaspal Singh, Ahmad Zahia, Kong-Ming Hu, Jia-Xin Peng, S. K. Singh

    Abstract: We investigate quantum synchronization in a coupled cavity optomagnonic system which consists of two spatially separated optical whispering-gallery-mode (WGM) resonators and each resonator is also coupled to a yttrium iron garnet (YIG) sphere through the optomagnonic interaction. Phase-dependent single-photon hopping factor couples the two optical resonators and provides an indirect interaction be… ▽ More

    Submitted 29 June, 2026; originally announced June 2026.

  19. arXiv:2606.26570  [pdf, ps, other

    math.CO math.GR

    Classification of regular Cayley maps of skew-type three on semidihedral groups

    Authors: Kan Hu, Tao Qiu

    Abstract: It is well known that every regular Cayley map $M = \CM(G,X,p)$ on a finite group $G$ with respect to an inverse-closed generating set $X$ of $G$ and a specified cyclic permutation $p$ on $X$ corresponds to a skew morphism $\varphi$ on $G$ such that the restriction of $\varphi$ to $X$ is $p$. The skew-type of the map $M$ is defined as the index $[G:\Ker \varphi]$, which equals the number of distin… ▽ More

    Submitted 24 June, 2026; originally announced June 2026.

    Comments: 24pages

    MSC Class: 05E18; 20B25; 57M15

  20. arXiv:2606.24338  [pdf, ps, other

    cs.RO

    RoBoSR: Structured Scene Representations for Embodied Robotic Reasoning

    Authors: Kewei Hu, Wanchan Yu, Fangwen Chen, Jing Jiajian, Zimeng Li, Ying Wei, Tianhao Liu, Michael Zhang, Hanwen Kang

    Abstract: Despite rapid progress, embodied reasoning under real-world variability remains challenging. Existing approaches rely on demonstration-driven sequential biases, limiting flexibility in open-ended and long-horizon tasks that require structured reasoning over evolving states. We introduce RoBoSR, an intermediate structural representation that formulates manipulation as step-wise state transitions… ▽ More

    Submitted 23 June, 2026; originally announced June 2026.

  21. arXiv:2606.18286  [pdf, ps, other

    cs.LG

    CODEBLOCK: Learning to Supervise Code at the Right Granularity

    Authors: Zhijie Deng, Ling Li, Jinlong Pang, Kaiqin Hu, Qi Xuan, Zhaowei Zhu, Jiaheng Wei

    Abstract: Supervised fine-tuning of code LLMs typically applies uniform cross-entropy loss to all response tokens, implicitly assuming that every token provides equally useful learning signal. Recent token-level selection methods challenge this assumption in natural-language SFT by supervising only high-value tokens. However, directly transferring token-level masking to code can break syntactically and sema… ▽ More

    Submitted 10 June, 2026; originally announced June 2026.

  22. arXiv:2606.16572  [pdf, ps, other

    cs.RO

    Steering Generative Reinforcement Learning into Stable Robotic Controller

    Authors: Yixuan Wang, Shutong Ding, Ke Hu, Tianxiang Gui, Jingya Wang, Ye Shi

    Abstract: Diffusion and flow-based generative policies provide a powerful policy class for reinforcement learning by inducing rich stochastic exploration through iterative action generation. However, the stochasticity of diffusion policies is not suitable for stable and precise control in high-dimensional robotic systems, where small action variations can accumulate into inconsistent motion and reduced robu… ▽ More

    Submitted 15 June, 2026; originally announced June 2026.

  23. arXiv:2606.15079  [pdf, ps, other

    cs.CL cs.AI

    Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale

    Authors: Ang Li, Ben Liu, Bin Han, Bin Hu, Bin Jing, Binbin Hu, Bing Li, Cai Chen, Caizhi Tang, Changxin Tian, Chao Huang, Chao Zhang, Chen Liang, Chen Qian, Chengfu Tang, Chengyao Wen, Chilin Fu, Chunwei Wu, Cong Zhang, Cunyin Peng, Daixin Wang, Dalong Zhang, Deng Zhao, Dingnan Jin, Dingyuan Zhu , et al. (193 additional authors not shown)

    Abstract: Efficient and scalable agentic intelligence requires models that can deliver both low-latency responses and strong reasoning capabilities while remaining practical to train, serve, and deploy. In this report, we present Ling-2.6 and Ring-2.6, a family of models designed to address this challenge at scale. Ling-2.6 is optimized for instant response generation and high capability per output token, w… ▽ More

    Submitted 12 June, 2026; originally announced June 2026.

  24. arXiv:2606.14760  [pdf, ps, other

    cs.CV cs.AI

    GeoRoPE: Ground-Aware Rotary Adaptation for Remote Sensing Foundation Models

    Authors: Yu Luo, Kun Hu, Mengwei He, Xiaogang Zhu, Shan Zeng, Allen Benter, Wei Xiang, Patrick Filippi, Thomas Francis Bishop, Zhiyong Wang

    Abstract: Remote-sensing foundation models (RSFMs) benefit from pretraining on imagery from multiple sensors and ground sampling distances (GSDs), but such exposure alone does not resolve scale mismatch during downstream adaptation. A fixed token-grid offset can correspond to different ground distances across sensors, making grid-based positional priors physically inconsistent. Meanwhile, heterogeneous spat… ▽ More

    Submitted 8 June, 2026; originally announced June 2026.

  25. arXiv:2606.08544  [pdf, ps, other

    math.OC cs.NI

    Block coordinate descent for joint delay-energy optimization in multi-hop D2D networks

    Authors: Kai-Xiang Hu, Jacek Gondzio, Caixia Kou

    Abstract: In multi-hop device-to-device (D2D) networks, the optimization of network-level metrics is particularly difficult due to the tight coupling between network-layer routing and physical-layer resource allocation. Departing from traditional average-performance metrics, this paper addresses the joint optimization of routing paths, transmission power, and bandwidth allocation. We formulate a generalized… ▽ More

    Submitted 7 June, 2026; originally announced June 2026.

  26. arXiv:2606.06967  [pdf, ps, other

    cs.LG

    GenPO++: Generative Policy Optimization with Jacobian-free Likelihood Ratios

    Authors: Ke Hu, Shutong Ding, Panxin Tao, Jingya Wang, Ye Shi

    Abstract: Generative policies provide expressive and multimodal action distributions, making them attractive for reinforcement learning (RL) in complex continuous-control tasks. Among them, flow-based policies are especially appealing because they generate actions through deterministic transport maps. However, applying such generative policies to likelihood-based on-policy learning remains limited by the di… ▽ More

    Submitted 5 June, 2026; originally announced June 2026.

  27. arXiv:2606.06671  [pdf, ps, other

    cs.CV

    JA-SIREN: Deterministic Initialization for Sinusoidal Networks via Spectral Matching

    Authors: Mohammed Alsakabi, Kejia Hu, John M. Dolan, Ozan K. Tonguz

    Abstract: Existing implicit neural representation (INR) approaches suffer from stochastic initialization that does not guarantee consistent or high-quality performance across runs, with variations reaching more than 2.5 dB (78%) in image regression. This variation is problematic for scientific computing and simulation, where result reproducibility is crucial. To address this problem, we present Jacobi-Anger… ▽ More

    Submitted 4 June, 2026; originally announced June 2026.

  28. arXiv:2606.04159  [pdf, ps, other

    astro-ph.HE

    Predictions for the X-ray polarisation modulation in Cygnus X-1 from reflection off the stellar companion and its wind

    Authors: Bert Vander Meulen, Kun Hu, Victoria Grinberg, Henric Krawczynski

    Abstract: Context. Cyg X-1 is one of the brightest X-ray binaries and has been observed multiple times with the Imaging X-ray Polarimetry Explorer (IXPE). Recent studies report tentative evidence for a polarisation modulation with the orbital period P, but a half-period (P/2) signal, expected from reflection off the companion star and its stellar wind, has not been reported. Aims. We aim to quantify the ref… ▽ More

    Submitted 2 June, 2026; originally announced June 2026.

    Comments: 10 pages, 10 figures, submitted to A&A

  29. SRENet: Spectral Re-Entry Network for Point Cloud Action Recognition

    Authors: Qiuxia Wu, Jiarui Lan, Wenxiong Kang, Zhiyong Wang, Kun Hu

    Abstract: Recognizing human actions from point cloud sequences is critical for 3D perception driven applications such as autonomous driving and human-computer interaction. However, the irregular structure and temporal inconsistency of point clouds pose unique challenges for spatio-temporal representation learning, especially in capturing both global motion context and fine-grained temporal dynamics. We prop… ▽ More

    Submitted 2 June, 2026; originally announced June 2026.

    Comments: 13 pages, 11 figures. Accepted by IEEE Transactions on Circuits and Systems for Video Technology

  30. arXiv:2606.01266  [pdf, ps, other

    eess.SY

    Regulating EV Charging Markets for Fairness: Incentives for Pricing and Capacity Decisions

    Authors: Ruiting Wang, Kita Hu, Yitong Yu, Scott Moura

    Abstract: The transition to electric mobility calls for charging infrastructure that is both efficient and socially equitable. This paper examines fairness in electric vehicle (EV) charging station pricing and capacity through a game-theoretic perspective. We model a non-cooperative market in which competing charging service providers set prices and capacities while customers choose stations based on genera… ▽ More

    Submitted 31 May, 2026; originally announced June 2026.

  31. MiCU: End-to-End Smart Home Command Understanding with Large Language Model

    Authors: Haowei Han, Kexin Hu, Weiwei Cai, Debiao Zhang, Bin Qin, Yuxiang Wang, Jiawei Jiang, Xiao Yan, Bo Du

    Abstract: Command understanding systems in smart home ecosystems can automate device control and substantially improve user experience. However, while they perform well on precise utterances (e.g., "turn on the bedroom light"), they struggle with ambiguous or misaligned commands (e.g., "make the bedroom cozy"). Large language models (LLMs) generalize well across various domains and can outperform traditiona… ▽ More

    Submitted 31 May, 2026; originally announced June 2026.

  32. TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents

    Authors: Weiyi Chen, Shuaixiong Wang, Ziyun Gao, Kaichun Hu, Wangze Ni, Shimin Di, Chen Jason Zhang, Lei Chen

    Abstract: The development of Large Language Models (LLMs) has significantly improved travel planning applications, yet evaluating such models is limited by existing benchmarks' limitations: 1) overemphasis on constraint compliance, neglecting multi-dimensional qualities like spatio-temporal cost; 2) datasets lacking real-world authenticity and coverage in key areas (e.g., lodging, transport); and 3) isolate… ▽ More

    Submitted 31 May, 2026; originally announced June 2026.

    Comments: 31pages, 8 figures, accepted by KDD 2026

  33. arXiv:2605.30056  [pdf, ps, other

    cs.RO cs.LG

    Sample-Efficient Diffusion-based Reinforcement Learning with Critic Guidance

    Authors: Shutong Ding, Zejia Zhong, Zhongyi Wang, Ke Hu, Bikang Pan, Jingya Wang, Ye Shi

    Abstract: Recent advances in reinforcement learning (RL) have achieved great successes by leveraging the multimodality and exploration capability of diffusion policies. Among these approaches, one representative branch focuses on the sampling-based policy optimization. This design enables better exploration capability of the diffusion model, particularly at the beginning of training, but suffer from low exp… ▽ More

    Submitted 28 May, 2026; originally announced May 2026.

    Comments: accepted by ICML2026

  34. arXiv:2605.26065  [pdf, ps, other

    astro-ph.HE

    X-ray Polarization Signatures from Comptonization by Magnetic Reconnection Plasmoids

    Authors: John Groger, Kun Hu, Henric Krawczynski

    Abstract: Emission from X-ray binaries in the hard spectral state is dominated by high-energy radiation attributed to the Compton scattering of seed photons. The prevalent model of the Comptonization by hot electrons or pairs faces the problem of rapid radiative cooling of the emitting particles. A proposed alternative mechanism is the Comptonization by scattering off fast plasmoids formed during magnetic r… ▽ More

    Submitted 17 August, 2026; v1 submitted 25 May, 2026; originally announced May 2026.

    Comments: ApJL, accepted

  35. arXiv:2605.23478  [pdf, ps, other

    cs.CV cs.AI

    PhenoYieldNet: Learning Crop-Aware Phenological Responses for Multi-Crop Yield Prediction

    Authors: Yu Luo, Xiaogang Zhu, Shan Zeng, Wei Xiang, Thomas Francis Bishop, Zhiyong Wang, Kun Hu

    Abstract: Accurate crop yield prediction is crucial for sustainable agriculture and global food security. While existing methods are predominantly developed for single-crop prediction, they often struggle to generalize across diverse crop types, without addressing the unique crop phenological responses that are dynamically modulated by complex weather patterns. In this paper, we propose PhenoYieldNet, a mul… ▽ More

    Submitted 22 May, 2026; originally announced May 2026.

    Comments: Accepted by CVPR2026

  36. arXiv:2605.22338  [pdf, ps, other

    cs.LG

    Physics-Informed Generative Solver: Bridging Data-Driven Priors and Conservation Laws for Stable Spatiotemporal Field Reconstruction

    Authors: Ziyuan Zhu, Keyu Hu, Zhifei Chen, Yuhao Shi, Ming Bao, Jing Zhao, Gang Wang, Haitan Xu, Jiadong Li, Qijun Zhao, Xiaodong Li, Minghui Lu, Yanfeng Chen

    Abstract: Reconstructing continuous physical fields from sparse measurements is a central inverse problem, but data-driven generative models can produce states that violate governing dynamics. We introduce a physics-informed generative solver that separates stable prior learning from inference-time enforcement of conservation laws. Martingale-Regularized Score Matching regularizes score pretraining with a S… ▽ More

    Submitted 21 May, 2026; originally announced May 2026.

  37. arXiv:2605.22323  [pdf, ps, other

    astro-ph.HE astro-ph.SR

    Magnetar Fireballs and Short Bursts: Curved Spacetime Lensing, QED Effects, Spectra, Polarization, and Impulse Responses

    Authors: Zorawar Wadiasingh, Hoa Dinh Thi, Constantinos Kalapotharakos, Kun Hu, Matthew G. Baring, Alice K. Harding, George Younes, Sebastien Guillot, Andrea Sanna, Michela Negro, Jeremy D. Schnittman, Oliver J. Roberts, Eric Burns, Chin-Ping Hu, Ersin Göğüş

    Abstract: Magnetar short bursts (SBs) are hard X-ray transients of durations $0.01-1$ s peaking at $\sim 10-100$ keV, and are prime targets for new high-energy missions and polarimeters. The recent association of SBs with bright radio bursts in SGR 1935+2154 has broadened interest in SB physics. We present new advanced fireball models combining general relativistic light bending, polarized transport in magn… ▽ More

    Submitted 6 August, 2026; v1 submitted 21 May, 2026; originally announced May 2026.

    Comments: Accepted for publication in ApJS

  38. arXiv:2605.21515  [pdf, ps, other

    cs.LG cs.AI

    Predicting Performance of Symbolic and Prompt Programs with Examples

    Authors: Chengqi Zheng, Keya Hu, Shuzhi Liu, Tao Wu, Kevin Ellis, Yewen Pu

    Abstract: LLM prompting is widely used for naturally stated tasks, yet it is unreliable it may succeed on a few test cases but fail at deployment time. We study performance prediction: given a program, either symbolic (e.g. Python) or a prompt executed on an LLM, and a few in-domain examples, predict its performance on unseen tasks from the same domain. We use a simple coin-flip model, treating each pass/fa… ▽ More

    Submitted 15 May, 2026; originally announced May 2026.

  39. arXiv:2605.17353  [pdf, ps, other

    cs.CY

    You Can't Fool Us: Understanding the Resilience of LLM-driven Agent Communities to Misinformation

    Authors: Chichen Lin, Yijie Jin, Kangbo Hu, Weijian Fan, Han Xiao, Yongbin Wang, Zhihui Ying, Zhanzhan Zhao

    Abstract: Misinformation resilience is a dynamic community process: communities differ not only in whether they initially trust false claims, but also in how they recover through interaction, questioning, correction, and support withdrawal. We study this process with an LLM-based agent simulation that constructs synthetic communities along two theoretically motivated dimensions: Actively Open-minded Thinkin… ▽ More

    Submitted 17 May, 2026; originally announced May 2026.

    Comments: 26 pages, 7 figures, 1 table

  40. arXiv:2605.13181  [pdf, ps, other

    cs.LG cs.AI

    Stable Attention Response for Reliable Precipitation Nowcasting

    Authors: Penghui Wen, Zexin Hu, Sen Zhang, Patrick Filippi, Xiaogang Zhu, Allen Benter, Thomas Bishop, Zhiyong Wang, Kun Hu

    Abstract: Precipitation nowcasting remains challenging due to the highly localized, rapidly evolving, and heterogeneous nature of atmospheric dynamics. Although recent methods increasingly adopt attention-based architectures in both unimodal and multimodal settings, they mainly emphasize stronger representation learning and prediction capacity, while paying less attention to the stability of attention respo… ▽ More

    Submitted 5 August, 2026; v1 submitted 13 May, 2026; originally announced May 2026.

  41. arXiv:2605.10938  [pdf, ps, other

    cs.CL cs.AI cs.LG

    ELF: Embedded Language Flows

    Authors: Keya Hu, Linlu Qiu, Yiyang Lu, Hanhong Zhao, Tianhong Li, Yoon Kim, Jacob Andreas, Kaiming He

    Abstract: Diffusion and flow-based models have become the de facto approaches for generating continuous data, e.g., in domains such as images and videos. Their success has attracted growing interest in applying them to language modeling. Unlike their image-domain counterparts, today's leading diffusion language models (DLMs) primarily operate over discrete tokens. In this paper, we show that continuous DLMs… ▽ More

    Submitted 25 June, 2026; v1 submitted 11 May, 2026; originally announced May 2026.

    Comments: Tech report. arXiv v2: add distillation results in Appendix B. https://linlu-qiu.github.io/assets/html/elf_pd.html

  42. arXiv:2605.09572  [pdf, ps, other

    cs.CV cs.AI cs.MM

    KAN Text to Vision? The Exploration of Kolmogorov-Arnold Networks for Multi-Scale Sequence-Based Pose Animation from Sign Language Notation

    Authors: Guanyi Du, Lintao Wang, Kun Hu, Ziyang Wang

    Abstract: Sign language production from symbolic notation offers a scalable route to accessible sign animation. We present KANMultiSign, a multi-scale sequence generator that translates HamNoSys notation into two-dimensional human pose sequences. Our framework makes two complementary contributions. First, we introduce a coarse-to-fine generation strategy with multi-scale supervision: the model is first guid… ▽ More

    Submitted 10 May, 2026; originally announced May 2026.

    Comments: Accepted at Neurocomputing

  43. arXiv:2605.08572  [pdf, ps, other

    cs.CV

    ECTraj: Enhanced Consistency Training for Multi-Agent Trajectory Prediction

    Authors: Alen Mrdovic, Qingze, Liu, Danrui Li, Mathew Schwartz, Kaidong Hu, Sejong Yoon, Mubbasir Kapadia, Vladimir Pavlovic

    Abstract: Diffusion models for multi-agent trajectory prediction are limited by iterative denoising, which causes inference latency that hinders their use in time-critical settings like autonomous driving. Fast-sampling variants using DDIM and informed initial noise distributions partially alleviate this issue, but they either fail to achieve true single-step generation or are constrained by the chosen nois… ▽ More

    Submitted 5 July, 2026; v1 submitted 8 May, 2026; originally announced May 2026.

  44. Resonant Inverse Compton Scattering and Hard X-ray Emission in Magnetar Magnetospheres

    Authors: Kun Hu, Nicholas Rackers, Alexander Y. Chen

    Abstract: Magnetars are a subclass of neutron stars with ultra-strong surface magnetic fields. Some magnetars exhibit persistent hard X-ray emission, characterized by power-law tails with photon indices around 1--1.5, extending from ${\sim}$10 keV to several hundred keV. The leading explanation for this hard X-ray component is resonant Compton scattering, in which the thermal seed photons are upscattered by… ▽ More

    Submitted 6 August, 2026; v1 submitted 8 May, 2026; originally announced May 2026.

    Comments: 22 pages, 13 figures. Accepted for publication in ApJ. Revised following referee comments

    Journal ref: The Astrophysical Journal, Volume 1006, Number 1 (2026)

  45. arXiv:2604.28145  [pdf, ps, other

    cond-mat.mtrl-sci

    Strong coupling between quantized magnon modes in a YIG microstucture and microwaves in a superconducting resonator

    Authors: Seth W. Kurfman, Philipp Geyer, Anoop Kamalasanan, Karl Heimrich, Kwangyul Hu, Tharnier O. Puel, Frank Heyroth, Michael Flatté, Georg Schmidt

    Abstract: Strong-coupling experiments based on magnons enable the exploration into on-chip demonstrations involving numerous long-lived excitations. Yttrium iron garnet (YIG) has been considered for decades as a gold standard material for magnonics due to its low-loss magnonic properties. While YIG has successfully demonstrated strong-coupling in macroscopic device geometries, the strong coupling of magnons… ▽ More

    Submitted 11 May, 2026; v1 submitted 30 April, 2026; originally announced April 2026.

  46. arXiv:2604.26747  [pdf, ps, other

    q-fin.PM q-fin.GN q-fin.TR

    From Hypotheses to Factors: Constrained LLM Agents in Cryptocurrency Markets

    Authors: Yikuan Huang, Zheqi Fan, Kaiqi Hu, Yifan Ye

    Abstract: LLM agents are promising tools for empirical discovery, but their flexibility can also turn discovery into uncontrolled search. We study how to use agents under a reproducible protocol through cryptocurrency factor discovery. Our framework casts the task as sequential hypothesis search: an agent reads an append-only experiment trace, proposes falsifiable factor hypotheses, and maps them to executa… ▽ More

    Submitted 29 April, 2026; originally announced April 2026.

  47. When AI reviews science: Can we trust the referee?

    Authors: Jialiang Wang, Yuchen Liu, Hang Xu, Kaichun Hu, Shimin Di, Wangze Ni, Linan Yue, Min-Ling Zhang, Kui Ren, Lei Chen

    Abstract: The volume of scientific submissions continues to climb, outpacing the capacity of qualified human referees and stretching editorial timelines. At the same time, modern large language models (LLMs) offer impressive capabilities in summarization, fact checking, and literature triage, making the integration of AI into peer review increasingly attractive -- and, in practice, unavoidable. Yet early de… ▽ More

    Submitted 26 April, 2026; originally announced April 2026.

    Journal ref: The Innovation Informatics 2:100030 (2026)

  48. arXiv:2604.20100  [pdf, ps, other

    cs.RO

    JoyAI-RA 0.1: A Foundation Model for Robotic Autonomy

    Authors: Tianle Zhang, Zhihao Yuan, Dafeng Chi, Peidong Liu, Dongwei Li, Kejun Hu, Likui Zhang, Junnan Nie, Ziming Wei, Zengjue Chen, Yili Tang, Jiayi Li, Zhiyuan Xiang, Mingyang Li, Tianci Luo, Hanwen Wan, Ao Li, Linbo Zhai, Zhihao Zhan, Xiaodong Bai, Jiakun Cai, Peng Cao, Kangliang Chen, Siang Chen, Yixiang Dai , et al. (37 additional authors not shown)

    Abstract: Robotic autonomy in open-world environments is fundamentally limited by insufficient data diversity and poor cross-embodiment generalization. Existing robotic datasets are often limited in scale and task coverage, while relatively large differences across robot embodiments impede effective behavior knowledge transfer. To address these challenges, we propose JoyAI-RA, a vision-language-action (VLA)… ▽ More

    Submitted 23 April, 2026; v1 submitted 21 April, 2026; originally announced April 2026.

  49. arXiv:2604.19755  [pdf

    cs.AI cs.LG

    Explainable AML Triage with LLMs: Evidence Retrieval and Counterfactual Checks

    Authors: Dorothy Torres, Wei Cheng, Ke Hu

    Abstract: Anti-money laundering (AML) transaction monitoring generates large volumes of alerts that must be rapidly triaged by investigators under strict audit and governance constraints. While large language models (LLMs) can summarize heterogeneous evidence and draft rationales, unconstrained generation is risky in regulated workflows due to hallucinations, weak provenance, and explanations that are not f… ▽ More

    Submitted 6 June, 2026; v1 submitted 22 March, 2026; originally announced April 2026.

  50. arXiv:2604.19476  [pdf, ps, other

    q-fin.PM q-fin.GN q-fin.ST

    Cross-Stock Predictability via LLM-Augmented Semantic Networks

    Authors: Yikuan Huang, Zheqi Fan, Kaiqi Hu, Yifan Ye

    Abstract: Text-based financial networks are increasingly used to study cross-stock return predictability. A common approach constructs links from similarities in firms' disclosure embeddings, but such networks often contain spurious edges because textual proximity does not necessarily imply economic connection. We propose a two-stage framework that first builds a sparse candidate graph from 10-K embeddings… ▽ More

    Submitted 26 April, 2026; v1 submitted 21 April, 2026; originally announced April 2026.