-
Tilted $p$-wave magnet candidate CeNiAsO
Authors:
Zhuo Wang,
Zheng Liu,
Shuo Zou,
Hua-Xun Li,
Jin-Xin Hu,
Zhuolun Qiu,
Ze Wang,
Jiamin Gong,
Lucheng Wei,
Kangjian Luo,
Hai Zeng,
Meng Zhang,
Chao Dong,
Chuanyin Xi,
Junfeng Wang,
Jiakun Fang,
Xiaotao Han,
Guang-Han Cao,
Liang Li,
Yongkang Luo
Abstract:
The unexpectedly small ordered moments of CeNiAsO, a candidate for correlated $p$-wave magnet, have posed a serious challenge to the precise determination of its magnetic structure, hindering the understanding of its fundamental properties. By leveraging the high sensitivity to local internal fields, our $^{75}$As nuclear quadrupole / magnetic resonance experiments reveal a commensurate antiferrom…
▽ More
The unexpectedly small ordered moments of CeNiAsO, a candidate for correlated $p$-wave magnet, have posed a serious challenge to the precise determination of its magnetic structure, hindering the understanding of its fundamental properties. By leveraging the high sensitivity to local internal fields, our $^{75}$As nuclear quadrupole / magnetic resonance experiments reveal a commensurate antiferromagnetic order with a small out-of-plane moment $m_z\approx0.05$ $μ_{\mathrm{B}}$. This tilted magnetic configuration not only rotates the spin polarization axis away from the crystallographic $\mathbf{c}$-axis, but also enhances the non-relativistic spin splitting. We refer to this rare paradigm as a \textit{tilted $p$-wave magnet}.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
An ultra-bright, highly-scalable, squeezed light source for hybrid quantum photonics
Authors:
Kai-Hong Luo,
Denis Kopylov,
Florian Lütkewitte,
Jan-Lucas Eickmann,
Simone Atzeni,
Fabian Schlue,
Benjamin Brecht,
Torsten Meier,
Polina Sharapova,
Michael Stefszky,
Christine Silberhorn
Abstract:
Hybrid quantum photonics seeks to combine the complementary advantages of continuous- and discrete-variable quantum optics. This typically entails photon-counting measurements on entangled states generated by interfering many single-mode squeezed-vacuum (SMSV) states. However, because conventional photon-counting schemes are mode-insensitive, it is critical that the SMSV states occupy a single, we…
▽ More
Hybrid quantum photonics seeks to combine the complementary advantages of continuous- and discrete-variable quantum optics. This typically entails photon-counting measurements on entangled states generated by interfering many single-mode squeezed-vacuum (SMSV) states. However, because conventional photon-counting schemes are mode-insensitive, it is critical that the SMSV states occupy a single, well-defined mode. Achieving this requires careful engineering of the process, which determines both the spatial and spectro-temporal properties of the generated state. In addition, the ideal source must be massively scalable, capable of efficiently generating strong squeezing, and remain compatible with existing detection schemes and fiber networks. Although many platforms address one or more of these requirements, satisfying all of them simultaneously remains challenging.
Here, we present a source that meets all of these requirements: a single-pass, periodically poled, Type-II potassium titanyl phosphate (KTP) waveguide optimized for scalable hybrid quantum-photonic architectures. The SMSV state produced by the source has a measured effective mode number of 1.24. Furthermore, the source is extremely bright (producing up to 40 000 photons per pulse) and operates at a central wavelength of 1546nm, optimized for fiber-network compatibility and which, in combination with picosecond duration, also enables intrinsic photon-number resolution in superconducting nanowire single-photon detectors. Although this source constitutes an ideal source in a simplified picture, the ultimate limitations of any source will be governed by complex dynamics that arise when the system is driven at high-gain or due to unavoidable loss during state generation. We have therefore developed a complete theoretical framework that enables a comprehensive photon-counting-based characterization of the source.
△ Less
Submitted 19 August, 2026;
originally announced August 2026.
-
Connected Subspace Clustering: Hardness, a Scalable Heuristic, and an Application to Sea Level Geodesy
Authors:
Johanna Hillebrand,
Jan Höckendorff,
Jürgen Kusche,
Kelin Luo,
Heiko Röglin,
Melanie Schmidt,
Christian Sohler,
Bernd Uebbing
Abstract:
Constrained optimization extends classical optimization by integrating side information, making it widely applicable across scientific and engineering domains. Consider a setting where we measure variables at different physical locations. When grouping these measurements, we often want clusters that are both internally similar and physically coherent. Thus, we have a constrained clustering problem…
▽ More
Constrained optimization extends classical optimization by integrating side information, making it widely applicable across scientific and engineering domains. Consider a setting where we measure variables at different physical locations. When grouping these measurements, we often want clusters that are both internally similar and physically coherent. Thus, we have a constrained clustering problem where the constraint models coherence. Motivated by an application in geodesy, where contiguous regions of the sea surface must be identified for principal component analysis, we introduce the Connected Subspace Clustering problem: given high-dimensional points and a connectivity graph, partition them into $k$ connected clusters, minimizing their total squared distance to the clusters' best-fit $m'$-dimensional affine subspaces. We prove that, even for $m' = 0$ and a grid graph with holes, the problem is NP-hard to approximate within $Ω(n^{1/2-\varepsilon})$ for every $\varepsilon>0$, where $n$ is the number of measurements. We then introduce an efficient Lloyd-style heuristic that alternates subspace fitting with an iterative merging procedure to enforce connectivity. Our method returns exactly $k$ connected regions by construction, whereas unconstrained methods leave up to $1{,}966$ disconnected fragments at higher cost. In a study of 160 configurations on global sea level time series, our merging-based repair is the strongest of four strategies in $73.75\%$ of cases, and consistently outperforms competitors such as (connected) Ward's method across all tested cluster counts. The resulting regions isolate signals aligning with climate indices such as the El Nino-Southern Oscillation and Indian Ocean Dipole. Although developed for geodesy, the approach applies to other spatially embedded multivariate time series, such as climate fields, remote sensing, neuroimaging, and sensor networks.
△ Less
Submitted 14 August, 2026;
originally announced August 2026.
-
Controlling the dynamics of an electric-field-driven droplet on a lubricant-infused micropillar surface
Authors:
Geng Wang,
Junyu Yang,
Timan Lei,
Jin Chen,
Halim Kusumaatmaja,
Kai Li,
Kai H. Luo
Abstract:
As a non-contact control approach, electric field (EF) can be utilised to drive droplet dynamics on a lubricant-infused surface (LIS), with numerous potential applications ranging from drug manufacturing to 3D printing. However, the resulting droplet dynamics remain poorly understood, especially as there are several possible droplet lubrication states on LIS. Here, we develop a lattice Boltzmann s…
▽ More
As a non-contact control approach, electric field (EF) can be utilised to drive droplet dynamics on a lubricant-infused surface (LIS), with numerous potential applications ranging from drug manufacturing to 3D printing. However, the resulting droplet dynamics remain poorly understood, especially as there are several possible droplet lubrication states on LIS. Here, we develop a lattice Boltzmann scheme that fully captures the interplay between the interfacial flows and electrohydrodynamics and harness it to investigate EF driven droplets on micropillar LIS. Combining simulations and analytical calculations, we establish quantitative expressions for the drag force and the electric force acting on a moving droplet. We demonstrate that the models can accurately capture droplet dynamics during programmable manipulation, including periodic motion and long-distance transport. Such reliable theoretical models can potentially transform precision control of droplet dynamics by removing the reliance on trial and error tests.
△ Less
Submitted 13 August, 2026;
originally announced August 2026.
-
VICBench: A Multi-Language Benchmark for Code Vulnerability Detection
Authors:
Jin Lu,
Xuening Han,
Yang Zhong,
Lin Tan,
Kevin Luo,
Andrew Gacek,
Neha Rungta
Abstract:
Evaluating security vulnerability detection tools requires benchmark datasets with vulnerability-inducing commits (VICs) - the commits that first introduce vulnerabilities into codebases. VICs are essential for determining the full range of vulnerable software versions. Existing vulnerability datasets suffer from limited programming language coverage, restricted patch complexity, and narrow projec…
▽ More
Evaluating security vulnerability detection tools requires benchmark datasets with vulnerability-inducing commits (VICs) - the commits that first introduce vulnerabilities into codebases. VICs are essential for determining the full range of vulnerable software versions. Existing vulnerability datasets suffer from limited programming language coverage, restricted patch complexity, and narrow project scope. Through our dual annotation by human experts and an agentic workflow, we create a benchmark - VICBench - of 100 verified VICs for 100 CVEs across 88 projects in Python, Java, and C++, covering 48 CWE types. VICBench features complex real-world vulnerability fixes averaging 38.6 lines and corresponding VICs of 252.5 lines - significantly larger than prior work. Our evaluation shows that state-of-the-art algorithms V-SZZ and LLM4SZZ achieve only 33.3%-40.1% F1, confirming that using existing approaches still entails significant manual effort. VICBench enables robust evaluation of vulnerability detection approaches.
△ Less
Submitted 12 August, 2026;
originally announced August 2026.
-
From Product Search to Preference Articulation: The Economics of Agentic Commerce
Authors:
Lingxiu Dong,
Kaiwen Luo,
Fasheng Xu
Abstract:
Generative AI is shifting digital commerce from browsing toward agentic search, in which consumers delegate product discovery to AI agents. We compare manual search, which accurately evaluates a limited product set, with agentic search, which screens a broad catalog through noisy representations of preferences and products. Preference complexity is the number of satisfaction-relevant dimensions th…
▽ More
Generative AI is shifting digital commerce from browsing toward agentic search, in which consumers delegate product discovery to AI agents. We compare manual search, which accurately evaluates a limited product set, with agentic search, which screens a broad catalog through noisy representations of preferences and products. Preference complexity is the number of satisfaction-relevant dimensions that are difficult to articulate before search but readily evaluated upon inspection. Consumers have finite attention and choose search intensity: products inspected manually or preference-refinement depth with an agent. We obtain three findings. First, manual search collapses beyond a finite complexity threshold: inspection ceases, mismatch reaches the no-search benchmark, and platform revenue falls to zero. Agentic search avoids this collapse. Once refinement becomes worthwhile, it remains worthwhile as complexity rises; mismatch stays below the no-search benchmark and revenue remains positive, although articulation effort and mismatch may increase. Second, platforms rank the regimes by conversion revenue, whereas consumers also bear search expenditure. When manual inspection is sufficiently inexpensive, agentic search becomes revenue-superior before consumers voluntarily adopt it, creating an adoption lag in which consumers rationally continue manual search. Third, conditional on agentic participation, platforms may assign lower fidelity to consumers with larger attention budgets because they can offset noisier representations through additional refinement, yielding an inverted fidelity allocation. Agentic commerce thus shifts scarcity from product inspection to preference articulation, making consumers' willingness and ability to interact central to voluntary use and platform fidelity design.
△ Less
Submitted 8 August, 2026;
originally announced August 2026.
-
Scenix: Sparse-View 3D Scene Reconstruction via Executable Scene Programs
Authors:
Kai Li,
Lutao Jiang,
Zhenyang Li,
Jiayu Dong,
Jierui Zhang,
Yingda Yin,
Runze Zhang,
Kai Yan,
Xiaoyang Huang,
Keyang Luo,
Xin Wang,
Xiangyu Zhao,
Weikai Chen
Abstract:
Synthesizing a structured and editable 3D indoor scene from a few uncalibrated RGB views requires more than generating high-quality individual assets: a system must infer the room structure, associate objects across incomplete observations, and recover a globally consistent spatial configuration. Previous methods mainly focus on 3D scene generation with text input or require continuous visual inpu…
▽ More
Synthesizing a structured and editable 3D indoor scene from a few uncalibrated RGB views requires more than generating high-quality individual assets: a system must infer the room structure, associate objects across incomplete observations, and recover a globally consistent spatial configuration. Previous methods mainly focus on 3D scene generation with text input or require continuous visual inputs with additional priors, \ e.g., human-annotated masks or accurate 3D layouts, which makes these methods labor demanding and hard to apply in general cases. We present \textsc{Scenix}, a sparse-view 3D scene reconstruction framework via executable scene programs, a structured representation that can be directly instantiated into editable 3D scenes. Given sparse views, \textsc{Scenix} predicts executable scene programs through perception-grounded asset instantiation and closed-loop spatial refinement. % We present \method, a framework that predicts an executable scene representation from sparse views and realizes it through perception-grounded asset instantiation and closed-loop spatial refinement. To support this task, we construct \dataset, a dataset of approximately 110,000 synthetic and real indoor scenes with multiview imagery, room structures, object-centric descriptions, and metric spatial annotations. We further introduce observation-consistent supervision that aligns each target scene with the visual evidence available in its input views. Experiments on held-out \textsc{XScene} scenes, real indoor images, and out-of-distribution SpatialGen cases evaluate structured scene prediction, object grounding, and spatial refinement.
△ Less
Submitted 7 August, 2026;
originally announced August 2026.
-
Principles of Robot Autonomy
Authors:
Daniele Gammelli,
Joseph Lorenzetti,
Katie Luo,
Gioele Zardini,
Marco Pavone
Abstract:
Autonomous robots are moving rapidly from research labs into everyday life - on roads, in the air, in warehouses, and in space. Robot autonomy is no longer solely an academic pursuit, but a collection of mature, field-tested methods and tools that practitioners rely on in real-world deployments. This book offers a clear, unified introduction to the methods that make this possible. Built on decades…
▽ More
Autonomous robots are moving rapidly from research labs into everyday life - on roads, in the air, in warehouses, and in space. Robot autonomy is no longer solely an academic pursuit, but a collection of mature, field-tested methods and tools that practitioners rely on in real-world deployments. This book offers a clear, unified introduction to the methods that make this possible. Built on decades of teaching at Stanford, the text develops the core elements of modern autonomy stacks within a single conceptual framework, bridging classical robotics and modern physical AI. Every major topic is paired with hands-on Jupyter notebooks and implementation-driven exercises, so readers build practical intuition alongside theoretical understanding. The result is a principled, accessible, and deployment-aware foundation for anyone seeking to design, analyze, or contribute to the next generation of autonomous systems. This is a comprehensive resource for students, engineers, and researchers entering one of today's fastest-growing fields.
△ Less
Submitted 4 August, 2026;
originally announced August 2026.
-
Iterative minimization in reduced density matrix functional theory for periodic systems
Authors:
Kai Luo,
Jingang Han,
Peize Lin,
Daye Zheng,
Mohan Chen,
Xinguo Ren
Abstract:
Reduced density matrix functional theory (RDMFT) offers a route beyond Kohn-Sham density functional theory for strongly correlated systems, yet practical calculations for periodic solids are still out of reach. We formulate RDMFT for extended systems in a basis-independent way and present a planewave implementation using iterative minimization for periodic solids, evaluating nonlocal exchange-corr…
▽ More
Reduced density matrix functional theory (RDMFT) offers a route beyond Kohn-Sham density functional theory for strongly correlated systems, yet practical calculations for periodic solids are still out of reach. We formulate RDMFT for extended systems in a basis-independent way and present a planewave implementation using iterative minimization for periodic solids, evaluating nonlocal exchange-correlation functionals through the existing adaptive compressed exchange machinery. Natural occupations are optimized under N-representability constraints with a spectral projected gradient (SPG) method or an first-order explicit-by-implicit (EBI) map, while natural orbitals are updated by Riemannian optimization on the complex Stiefel manifolds. Benchmarks on typical systems of \ce{H2}, silicon, and sodium with the Hartree-Fock functional show that SPG reproduces converged hybrid references, whereas EBI can stall when occupations approach $0$ or $1$. With the power and Müller functionals, SPG yields lower energies and more stable convergence than EBI. Applications to fractionally charged \ce{LiH}, dissociating \ce{H2} and \ce{N2} molecules, and equation of state of silicon show that the algorithm presented in this implementation is reliable and robust.
△ Less
Submitted 30 July, 2026;
originally announced July 2026.
-
Certain functional identities on matrix rings
Authors:
Kaijia Luo,
Jiankui Li
Abstract:
Let $D$ be a noncommutative division ring and let $R = M_{m}(D)$ with $m > 1$. We characterize additive mappings $f,$ $g:R\rightarrow R$ satisfying the identity $f(X) = X^{n}g(X^{-1})$ for every invertible element $X$ in $R$, where $n$ is a nonnegative integer. We show that the solutions are precisely given by $f = g$ and $f(X) = Xf(I)$ for all $X$ in $R$. Moreover, if $n \neq2$, both $f$ and $g$…
▽ More
Let $D$ be a noncommutative division ring and let $R = M_{m}(D)$ with $m > 1$. We characterize additive mappings $f,$ $g:R\rightarrow R$ satisfying the identity $f(X) = X^{n}g(X^{-1})$ for every invertible element $X$ in $R$, where $n$ is a nonnegative integer. We show that the solutions are precisely given by $f = g$ and $f(X) = Xf(I)$ for all $X$ in $R$. Moreover, if $n \neq2$, both $f$ and $g$ are identically zero.
△ Less
Submitted 9 July, 2026;
originally announced July 2026.
-
AREX: Towards a Recursively Self-Improving Agent for Deep Research
Authors:
Shuqi Lu,
Chaofan Li,
Kun Luo,
Zhang Zhang,
Hui Wang,
Hongwang Xiao,
Lei Xiong,
Jiahao Wang,
Sen Wang,
Xiyan Jiang,
Wanli Li,
Yuyang Hu,
Hongjin Qian,
Bingyu Yan,
Jianlyu Chen,
Ziyi Xia,
Yingxia Shao,
Kang Liu,
Zhicheng Dou,
Di He,
Chaozhuo Li,
Qiwei Ye,
Zhongyuan Wang,
Zheng Liu
Abstract:
Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This discovery--verification asymmetry suggests that a research agent should do more than simply search longer: it should recursively improve its current answer by verifying intermed…
▽ More
Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This discovery--verification asymmetry suggests that a research agent should do more than simply search longer: it should recursively improve its current answer by verifying intermediate results and using the partially verified state to guide subsequent refinement. We introduce AREX, a family of Recursively Self-Improving (RSI) deep research agents. AREX alternates between an inner research loop that gathers evidence and constructs a provisional answer, and an outer self-improvement loop that audits the answer constraint-wise, identifies unresolved claims, and launches targeted follow-up research. To sustain RSI over long horizons, AREX learns an autonomous context-update tool that compresses growing interaction history into a compact improvement state preserving verified evidence and unresolved constraints, without relying on an external model. We train AREX on verified synthetic tasks and high-quality trajectories through agentic mid-training and long-horizon reinforcement learning. To mitigate sparse final rewards during long horizon learning, we emphasize key steps where decisive evidence is acquired or erroneous research directions are corrected. We instantiate a dense 4B model and a 122B-A10B Mixture-of-Experts model. Across BrowseComp, WideSearch, DeepSearchQA, Humanity's Last Exam (HLE), and other reasoning and tool-use benchmarks, AREX substantially outperforms comparable-scale baselines and remains competitive with models using substantially more activated parameters.
△ Less
Submitted 23 July, 2026; v1 submitted 23 July, 2026;
originally announced July 2026.
-
High-accuracy ultrasonic positioning of calibration sources in the Jiangmen Underground Neutrino Observatory
Authors:
Ziqian Xiang,
Rongcheng Chen,
Zhangmin Chen,
Qian Chen,
Diwash Ghimire,
Jiaqi Hui,
Junting Huang,
Junjie Jiang,
Daijin Li,
Haojing Lai,
Kai Luo,
Rui Li,
Yilin Liao,
Jianglai Liu,
Yue Meng,
Yazhen Shi,
Duo Teng,
Linwei Tao,
Qi Wang,
Changsheng Ye,
Guolei Zhu,
Ping Zhang,
Tao Zhang
Abstract:
Precise source positioning is essential for detector calibration in large liquid scintillator detectors such as JUNO, particularly in regions where purely mechanical control is insufficient. An ultrasonic positioning system has been developed to reconstruct the three-dimensional coordinates of a calibration source without interfering with photon collection or contaminating the liquid scintillator.…
▽ More
Precise source positioning is essential for detector calibration in large liquid scintillator detectors such as JUNO, particularly in regions where purely mechanical control is insufficient. An ultrasonic positioning system has been developed to reconstruct the three-dimensional coordinates of a calibration source without interfering with photon collection or contaminating the liquid scintillator. The method combines a sound-speed modeling based on dedicated laboratory measurements and in-detector temperature profiles, waveform-based arrival-time reconstruction, and an in-situ calibration of the effective receiver geometry using central-axis deployments. With six active receivers, central-axis positioning yields a mean error of 1.23 cm relative to the known deployment reference. For off-axis operation in the Cable Loop System calibration plane, a detector-realistic simulation that includes timing resolution, sound-speed variation, and receiver-coordinate smearing predicts a positioning uncertainty of 2.40 cm. These results demonstrate that ultrasonic positioning can provide centimetre-level source accuracy for large liquid scintillator detectors and can support off-axis calibration in JUNO-like experiments.
△ Less
Submitted 22 July, 2026;
originally announced July 2026.
-
StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis
Authors:
Kaicheng Luo,
Xuefei Gong,
Yutao Sun,
Jinling He,
Yujie Hou,
Xiaoyang Xing,
Huiyan Li,
Bing Han,
Yanmin Qian
Abstract:
The trade-off between robustness, latency, and prosody critically challenges text-to-speech (TTS) systems. Autoregressive models, despite fidelity, are slow and error-prone; non-autoregressive (NAR) alternatives, while fast, often sacrifice prosodic naturalness via rigid alignments. This paper introduces StellarTTS, a novel mobile-optimized NAR TTS framework based on a sparse temporal embedding st…
▽ More
The trade-off between robustness, latency, and prosody critically challenges text-to-speech (TTS) systems. Autoregressive models, despite fidelity, are slow and error-prone; non-autoregressive (NAR) alternatives, while fast, often sacrifice prosodic naturalness via rigid alignments. This paper introduces StellarTTS, a novel mobile-optimized NAR TTS framework based on a sparse temporal embedding strategy, enabling granular control of phoneme duration, pronunciation, and prosody. Furthermore, we propose a semantic-aware codec that facilitates efficient single-stage decoding. Conditioned on the sparse temporal embedding, our 83M-parameter lightweight masked generative transformer achieves a real-time factor (RTF) of 0.08. Experiments demonstrate that StellarTTS attains lower latency and stronger robustness compared to state-of-the-art TTS systems, while maintaining competitive performance in audio quality, prosodic naturalness, and speaker similarity.
△ Less
Submitted 22 July, 2026;
originally announced July 2026.
-
Suppression of Non-Hermitian Skin Effect by Pseudomagnetic Field in Honeycomb Lattice
Authors:
Kai Shao,
Kun Luo
Abstract:
Magnetic suppression of the non-Hermitian skin effect (NHSE) offers a viable route for its control. While the NHSE has been realized in various classical-wave platforms, only pseudomagnetic fields (PMFs), which preserve time-reversal symmetry, can be engineered in such systems; however, their interplay with the NHSE remains underexplored. Here, we investigate this interplay in a non-Hermitian hone…
▽ More
Magnetic suppression of the non-Hermitian skin effect (NHSE) offers a viable route for its control. While the NHSE has been realized in various classical-wave platforms, only pseudomagnetic fields (PMFs), which preserve time-reversal symmetry, can be engineered in such systems; however, their interplay with the NHSE remains underexplored. Here, we investigate this interplay in a non-Hermitian honeycomb lattice by considering two distinct mechanisms for generating PMFs: monotonically increasing strain and spatially modulated gain and loss. We show that in both scenarios, PMFs can efficiently suppress the NHSE by driving skin modes into the bulk, accompanied by a reduction of the skin topological area and a contraction of the complex-energy spectrum under periodic boundary conditions. This mechanism is insensitive to boundary details and holds for various edge terminations, including zigzag, bearded, armchair, and twig edges. Our results establish PMFs as a versatile and effective means to control the NHSE, and point toward feasible implementations in a broad range of artificial platforms, including photonic and acoustic metamaterials as well as topolectrical circuits.
△ Less
Submitted 20 July, 2026;
originally announced July 2026.
-
Radial velocity statistics of cosmic voids as a probe of interacting dark energy
Authors:
Kin Ho Luo,
Ming-chung Chu,
Kwan Chuen Chan,
Wangzheng Zhang
Abstract:
Due to their vast sizes and extremely low densities, the dynamics of cosmic voids are largely decoupled from complex, small-scale baryonic physics and are highly sensitive to the background expansion of the Universe. This makes them clean and sensitive probes of dark energy's properties. Using N-body simulations, we show that the void radial velocity and velocity dispersion profiles are sensitive…
▽ More
Due to their vast sizes and extremely low densities, the dynamics of cosmic voids are largely decoupled from complex, small-scale baryonic physics and are highly sensitive to the background expansion of the Universe. This makes them clean and sensitive probes of dark energy's properties. Using N-body simulations, we show that the void radial velocity and velocity dispersion profiles are sensitive to the Type 3 interacting dark energy model parameters, the momentum coupling $β$ ($<0$) and the scalar field parameter $λ$. Within the $1σ$ range of the best-fit values of $β$ and $λ$ constrained by Planck CMB, DESI BAO, and DES-Y5 supernova data, we find up to $\sim 30\%$ deviations in the void radial-velocity and velocity-dispersion profile spans relative to the uncoupled scenario, which can be well-approximated by a 4-parameter quadratic regression model. This demonstrates that the void radial velocity statistics provide an independent and observationally accessible probe of the dark-sector interaction in the Type 3 model.
△ Less
Submitted 14 July, 2026;
originally announced July 2026.
-
A Glimpse into Long-term Physical Coexistence with Intelligent Robots
Authors:
Weiqi Jin,
Peijun Tang,
Kuncheng Luo,
Baifu Huang,
Binyan Sun,
Haotian Yang,
Shangjin Xie,
Jianan Wang
Abstract:
Long-term physical coexistence with intelligent robots requires more than capable robot policies. A persistent robotic assistant must support diverse user-facing interfaces, maintain long-horizon memory of people and preferences, coordinate across robot embodiments, and translate human intent into safe physical execution. We introduce PHILIA, a multi-robot agent built around a robot gateway abstra…
▽ More
Long-term physical coexistence with intelligent robots requires more than capable robot policies. A persistent robotic assistant must support diverse user-facing interfaces, maintain long-horizon memory of people and preferences, coordinate across robot embodiments, and translate human intent into safe physical execution. We introduce PHILIA, a multi-robot agent built around a robot gateway abstraction. PHILIA retains the rich interaction and tool ecosystem of OpenClaw while exposing robot-local runtimes, onboard perception, navigation, speaker, and robot policies through a unified capability interface. This design decouples low-frequency, high-semantic agent reasoning from high-frequency, low-level robot execution, enabling plug-and-play integration of user interfaces, robot embodiments, and policy backends. As a result, the user experience becomes compositional: advances in user interfaces, robot embodiments, robot policies, navigation, or interaction algorithms can improve the overall experience without redesigning the system. We validate the architecture on Astribot S1 robots while designing the robot gateway contract to support future heterogeneous robot platforms through a shared capability interface for observation, task execution, navigation, speech playback, status monitoring, and task cancellation. We present representative use cases in which agent memory and scene understanding are grounded in robot actions. These span interactive household scenarios, ranging from simple organization to challenging long-horizon and dexterous service tasks, such as packing a backpack and lifting a garbage bag. We highlight the human-robot interaction flow, where contextual understanding of user intent and preferences, together with human-in-the-loop confirmation or adjustment during execution, is essential for effective assistance.
△ Less
Submitted 13 July, 2026;
originally announced July 2026.
-
Towards Predictive, Aligned, and Scalable Robot Learning
Authors:
Peijun Tang,
Shangjin Xie,
Baifu Huang,
Binyan Sun,
Haotian Yang,
Kuncheng Luo,
Weiqi Jin,
Shilin Fang,
Jianan Wang
Abstract:
Learning, at its core, extends beyond memorization to the ability to reason and solve novel problems by navigating a space of possibilities. We introduce Lumo-2, a latent world-action model that generates actions by reasoning over world dynamics in latent space. The learned latent world dynamics capture physically grounded visual transitions, naturally encoding future possibilities and providing a…
▽ More
Learning, at its core, extends beyond memorization to the ability to reason and solve novel problems by navigating a space of possibilities. We introduce Lumo-2, a latent world-action model that generates actions by reasoning over world dynamics in latent space. The learned latent world dynamics capture physically grounded visual transitions, naturally encoding future possibilities and providing a unified substrate for cross-modal alignment. This formulation enables predictive reasoning akin to world modelling while remaining lightweight and focused on physical dynamics relevant to control. Central to our approach is the hypothesis that action generation quality is governed by the geometry of the latent space. We observe that standard reconstruction-based action tokenization objectives induce representations biased toward low-level signal fidelity, leading to misalignment between reconstruction quality and downstream control performance. To address this limitation, we propose a multi-stage modality pre-alignment strategy in which action representations are progressively aligned with latent world dynamics, vision, and language. This process enforces cross-modal consistency, promotes abstraction, and induces a structured latent space for predictive reasoning. We provide a systematic empirical study of latent world modelling and modality alignment, analyzing their roles in scaling laws and out-of-distribution generalization. Results show that Lumo-2 consistently outperforms strong vision-language-action (VLA) and world-action model (WAM) baselines, with gains on challenging real-world tasks requiring temporal reasoning, physical understanding, or high control complexity, including long-horizon and dexterous manipulation. These findings suggest that structured multimodal alignment and predictive reasoning are fundamental principles for advancing embodied intelligence.
△ Less
Submitted 13 July, 2026;
originally announced July 2026.
-
Approximation Algorithms for the Traveling Thief Problem
Authors:
Jan Eube,
Kelin Luo,
Heiko Röglin,
Sarah Sturm
Abstract:
The Traveling Thief Problem (TTP) combines the Traveling Salesperson Problem with the Knapsack Problem. In this problem, a finite metric space is given, and at each location an item with some profit and weight is placed. An agent seeks to collect a subset of the items. To do so, the agent must decide which items to collect and to determine a cyclic tour visiting the corresponding locations. While…
▽ More
The Traveling Thief Problem (TTP) combines the Traveling Salesperson Problem with the Knapsack Problem. In this problem, a finite metric space is given, and at each location an item with some profit and weight is placed. An agent seeks to collect a subset of the items. To do so, the agent must decide which items to collect and to determine a cyclic tour visiting the corresponding locations. While collecting an item yields its profit as a reward, the agent's speed decreases as more weight is picked up. The problem involves two competing objectives: maximizing the total profit of the collected items and minimizing the travel time of the tour.
While many heuristics and exact algorithms (with a non-polynomial running time) have been developed, no approximation algorithms are known for any variant of the TTP. We aim at computing an $(α_1,α_2)$-approximate Pareto set that, for every solution, contains another solution collecting at least a $\frac{1}{α_1}$ fraction of its profit while requiring at most $α_2$ times its travel time. Our main result is an algorithm that calculates a $(9 + ε,9 + ε)$-approximate Pareto set in polynomial time.
We also consider the setting in which the set of items to be collected is given in advance, so that the agent only has to compute a tour through the corresponding locations that minimizes the total travel time. This is the so-called Weighted TSP. For this setting, we present a $(2e + ε)$-approximation algorithm.
△ Less
Submitted 6 July, 2026;
originally announced July 2026.
-
ProWAFT: A ROMA-LPD Instance for Workload-Aware and Dynamic Fault Tolerance in FPGA-Based CNN Accelerators
Authors:
Xinxin Chen,
Haoran Qiao,
Yiming Guo,
Kecheng Luo,
Siyuan Feng,
Jingwen Ma
Abstract:
SRAM-based FPGAs provide an attractive platform for energy- and latency-constrained CNN inference at the network edge, yet transient faults can lead to silent errors that compromise reliability. Always-on redundancy (e.g., full TMR) improves correctness but incurs substantial performance and energy overhead, while reactive recovery may introduce unacceptable latency on the critical path. We propos…
▽ More
SRAM-based FPGAs provide an attractive platform for energy- and latency-constrained CNN inference at the network edge, yet transient faults can lead to silent errors that compromise reliability. Always-on redundancy (e.g., full TMR) improves correctness but incurs substantial performance and energy overhead, while reactive recovery may introduce unacceptable latency on the critical path. We propose \textbf{ProWAFT}, a proactive workload-aware fault-tolerance framework for FPGA-based CNN accelerators that uses partial reconfiguration to selectively apply TMR across reconfigurable partitions. ProWAFT quantifies workload criticality, models fault propagation and reconfiguration overhead, and selects configurations that minimize a composite objective over latency, energy, and reliability risk. Implemented on a Xilinx Zynq UltraScale+ ZCU104 platform with six reconfigurable regions and evaluated on a 500-task trace derived from ResNet-18, MobileNetV2, and EfficientNet-Lite under time-varying SEU injection, ProWAFT achieves lower composite cost than static TMR and reactive reconfiguration while maintaining high task success rate and near-baseline throughput with low online decision overhead.
△ Less
Submitted 1 July, 2026;
originally announced July 2026.
-
From Personas to Plot: Character-Grounded Multi-Agent Story Generation for Long-Form Narratives
Authors:
Aayush Aluru,
Chloe Ho,
Muhammad Hammouri,
Kerry Luo,
Myra Malik,
Ryan Lagasse,
Arjun Bahuguna,
Vasu Sharma
Abstract:
Although large language models (LLMs) have demonstrated impressive creative fiction generation, they struggle to maintain narrative consistency and coherent plot lines in long-form stories. In this work, we introduce a unified framework for long-form narrative generation and verification. MAGNET, a multi-agent goal-driven narrative engine for storytelling, generates stories with persona-grounded c…
▽ More
Although large language models (LLMs) have demonstrated impressive creative fiction generation, they struggle to maintain narrative consistency and coherent plot lines in long-form stories. In this work, we introduce a unified framework for long-form narrative generation and verification. MAGNET, a multi-agent goal-driven narrative engine for storytelling, generates stories with persona-grounded character agents that propose actions based on a shared world state and evolving story goals, while ATLAS is a graph-based pipeline that compares scene-level world representations across a generated story to detect hallucinations. By evaluating MAGNET using an LLM editor, pairwise rubric scoring, and ATLAS, we show that our framework produces coherent narratives compared to single-model prompting and IBSEN. At 100 pages, MAGNET reduced annotations and hallucinations by 41 and 50%, respectively, compared to the single model baseline and by 34 and 45%, respectively, compared to IBSEN, with pairwise rubric evaluation showing similar results. These results suggest that long-form narratives can emerge from explicit world-state tracking and goal-driven multi-agent generation, providing a foundation for controllable and structurally coherent long-form narrative generation.
△ Less
Submitted 1 July, 2026;
originally announced July 2026.
-
CSTrader: A Testbed for Language-Grounded Trading in a Community-Driven Virtual Asset Market
Authors:
Yao Shi,
Kingfung Luo,
Nan Tang,
Yuyu Luo
Abstract:
Niche asset markets, such as Counter-Strike 2 (CS2) weapon skins, are small, volatile, and heavily driven by community discussions and platform rules. These properties make them hard for traditional quantitative models, but provide an ideal testbed for studying how large language models (LLMs) turn unstructured text into trading actions. We present CSTrader, a multi-agent framework for language-gr…
▽ More
Niche asset markets, such as Counter-Strike 2 (CS2) weapon skins, are small, volatile, and heavily driven by community discussions and platform rules. These properties make them hard for traditional quantitative models, but provide an ideal testbed for studying how large language models (LLMs) turn unstructured text into trading actions. We present CSTrader, a multi-agent framework for language-grounded trading in the CS2 skin market. The system first integrates heterogeneous signals from various sources, then uses specialized agents for technical analysis, liquidity, events, and (reversed) sentiment, and finally applies risk control, transaction friction, and portfolio management agents to produce buy, sell, or hold decisions under realistic trading frictions. We build a live-like evaluation environment with real CS2 data from a highly volatile period and evaluate several recent LLM backbones. Across models, CSTrader consistently outperforms both a falling market index (-15.62%) and simple single-prompt LLM baselines, achieving up to a 7.58% cumulative return with controlled risk. Ablation studies show that liquidity, reversed sentiment, and transaction friction agents are crucial for turning noisy language signals into stable profits, suggesting that niche, language-driven markets are a useful benchmark for future language-to-action research. Code is available at: https://github.com/IatomicreactorI/CSGOTrading?tab=readme-ov-file#quick-start
△ Less
Submitted 30 June, 2026;
originally announced June 2026.
-
PS-MOT: Cultivating Instance Awareness from Point Seeds for Multi-Object Tracking
Authors:
Kai Luo,
Fei Teng,
Mengfei Duan,
Wanjun Jia,
Xu Wang,
Hao Shi,
Kunyu Peng,
Zhiyong Li,
Kailun Yang
Abstract:
We introduce Point-supervised Multi-Object Tracking (PS-MOT) as a cost-effective alternative to traditional bounding box supervision, shifting the focus from spatial fitting to topological center-driven representation. However, PS-MOT faces challenges, e.g., spatial ambiguity and identity drift due to the lack of explicit geometric structure and scale constraints. To address these, we propose PS-T…
▽ More
We introduce Point-supervised Multi-Object Tracking (PS-MOT) as a cost-effective alternative to traditional bounding box supervision, shifting the focus from spatial fitting to topological center-driven representation. However, PS-MOT faces challenges, e.g., spatial ambiguity and identity drift due to the lack of explicit geometric structure and scale constraints. To address these, we propose PS-Track, a hierarchical pipeline transitioning from points to instances across data, model, and loss levels. At the data level, we introduce Temporal-Feedback Prompting (TFP) to evolve points into temporally consistent pseudo-labels using negative spatial cues and motion priors. At the model level, we design the Point-Excited Wavelet Attention (PEWA) module, which leverages semantic correlations to activate high-frequency components, ``hallucinating'' object boundaries. At the loss level, Uncertainty-Guided Gaussian Learning (UGL) models pseudo-labels as probabilistic distributions, dynamically calibrating supervision intensity. Experiments on DanceTrack, EmboTrack, SportsMOT, and JRDB demonstrate that PS-Track provides a feasible and effective point-supervised alternative across diverse tracking scenarios, establishing a new state-of-the-art for point-supervised tracking. The source code is available at https://github.com/xifen523/PS-MOT.
△ Less
Submitted 29 June, 2026;
originally announced June 2026.
-
CylindTrack: Depth-Aware Cylindrical Motion Modeling for Panoramic Multi-Object Tracking
Authors:
Buyin Deng,
Kai Luo,
Lingxin Huang,
Xinqi Liu,
Fei Cheng,
Hang Zheng,
Liming Yin,
Kailun Yang
Abstract:
Multi-Object Tracking (MOT) is a core capability for embodied perception, and panoramic cameras are attractive for embodied systems because their 360° field of view reduces blind spots and keeps surrounding targets observable for longer durations. However, panoramic MOT is not a straightforward extension of perspective MOT. In equirectangular panoramic videos, the horizontal image domain is period…
▽ More
Multi-Object Tracking (MOT) is a core capability for embodied perception, and panoramic cameras are attractive for embodied systems because their 360° field of view reduces blind spots and keeps surrounding targets observable for longer durations. However, panoramic MOT is not a straightforward extension of perspective MOT. In equirectangular panoramic videos, the horizontal image domain is periodic rather than Euclidean, which breaks planar motion assumptions and makes IoU-based association unreliable near the 0°/360° seam. Meanwhile, large-FoV scenes often contain more objects, stronger scale variation, and more frequent interactions, making online association particularly sensitive to unstable frame-wise depth cues. To address these issues, we propose CylindTrack, a depth-aware cylindrical tracking-by-detection framework for panoramic MOT. CylindTrack first introduces Depth-Temporal Trajectory Modeling (DTM), which promotes instance depth from an isolated frame-wise cue to a temporally filtered trajectory-level state. To improve the reliability of depth observations, we further develop Spherical Spatio-Temporal Consistency Learning (SSTC), which combines a Temporal Mixer and Spherical Geometry-aware Attention to enhance temporal coherence and panoramic geometric alignment in depth-aware representations. Finally, we design a Topology-Aware Cylindrical Motion Model (TCMM) that lifts horizontal motion into a continuous angular state space and performs seam-consistent motion prediction and association in the periodic panoramic domain. By jointly modeling trajectory-level depth consistency and panoramic topology, CylindTrack improves identity preservation and trajectory continuity in challenging panoramic scenes. The source code will be released at https://github.com/warriordby/CylindTrack.
△ Less
Submitted 29 June, 2026;
originally announced June 2026.
-
Rigidity of Closed Minimal Hypersurfaces in $\mathbb{S}^5$
Authors:
Jianquan Ge,
Tong Liu,
Keyan Luo,
Wenjiao Yan
Abstract:
The celebrated Chern conjecture asserts that any closed minimal hypersurface in $\mathbb{S}^{n+1}$ with constant scalar curvature is isoparametric. In this paper, we resolve this conjecture in the affirmative for $M^4 \subset \mathbb S^5$ under the assumption that the Gauss-Kronecker curvature $K$ is constant.
This result breaks the traditional reliance on consecutive trace conditions, demonstra…
▽ More
The celebrated Chern conjecture asserts that any closed minimal hypersurface in $\mathbb{S}^{n+1}$ with constant scalar curvature is isoparametric. In this paper, we resolve this conjecture in the affirmative for $M^4 \subset \mathbb S^5$ under the assumption that the Gauss-Kronecker curvature $K$ is constant.
This result breaks the traditional reliance on consecutive trace conditions, demonstrating that the nonconsecutive spectral invariant set $\{H, S, K\}$ is sufficient to yield complete geometric rigidity. To overcome the analytical singular locus, we construct two novel weighted $3$-forms adapted to $S$ and $K$. Crucially, the global curvature estimates required to close our analysis are obtained unconditionally by proving the Euler characteristic $χ(M)=0$. This local-to-global approach provides a new paradigm for higher-dimensional rigidity problems.
△ Less
Submitted 28 June, 2026;
originally announced June 2026.
-
Rethinking Training & Inference for Forecasting: Linking Winner-Take-All back to GMMs
Authors:
Qiyuan Wu,
Katie Z Luo,
Bharath Hariharan,
Wei-Lun Chao,
Mark Campbell
Abstract:
Trajectory forecasting for autonomous driving has advanced rapidly, yet representative models often produce uninformative posteriors over forecast modes, causing problems for mode pruning. We trace this to a modeling-training mismatch: forecasters are typically modeled as conditional Gaussian mixture models (GMMs) but trained with a winner-take-all (WTA) loss that assigns each sample to its neares…
▽ More
Trajectory forecasting for autonomous driving has advanced rapidly, yet representative models often produce uninformative posteriors over forecast modes, causing problems for mode pruning. We trace this to a modeling-training mismatch: forecasters are typically modeled as conditional Gaussian mixture models (GMMs) but trained with a winner-take-all (WTA) loss that assigns each sample to its nearest mode. We argue that this K-means-like hard assignment (one-hot), while preventing mode collapse, is the source of uninformative mode probabilities: it over-segments the trajectory space, ignores relatedness among nearby modes, and yields assignment instability under small perturbations. Guided by this lens, we introduce two post-hoc treatments: (1) test-time posterior-weighted merging that aggregates nearby candidate trajectories; and (2) a one-step expectation-maximization (EM) update that replaces hard labels with soft responsibilities, sharing probability mass across neighboring modes. Across several WTA-trained architectures, these lightweight steps produce more informative, faithfully ranked mode posteriors and strengthen final forecasts on popular displacement metrics -- without retraining. Our analysis unifies recent design choices through a GMM-vs-K-means perspective and offers principled, practical corrections that better align training objectives with inference.
△ Less
Submitted 24 June, 2026;
originally announced June 2026.
-
OmniContact: Chaining Meta-Skills via Contact Flow for Generalizable Humanoid Loco-Manipulation
Authors:
Runyi Yu,
Xiaoyi Lin,
Ji Ma,
Yinhuai Wang,
Koukou Luo,
Jiahao Ji,
Huayi Wang,
Wenjia Wang,
Runhan Zhang,
Ping Tan,
Ting Wu,
Ruoli Dai,
Qifeng Chen,
Lei Han
Abstract:
Learning long-horizon humanoid loco-manipulation poses a dual challenge: it requires not only the robust execution of meta-skills but also their seamless, closed-loop chaining equipped with autonomous recovery. Existing approaches remain limited: explicit humanoid-object interaction representations offer precision but are notoriously difficult for high-level planning, whereas implicit skill embedd…
▽ More
Learning long-horizon humanoid loco-manipulation poses a dual challenge: it requires not only the robust execution of meta-skills but also their seamless, closed-loop chaining equipped with autonomous recovery. Existing approaches remain limited: explicit humanoid-object interaction representations offer precision but are notoriously difficult for high-level planning, whereas implicit skill embeddings are compact but lack the interpretability required for reliable composition. We propose \ours, a hierarchical framework centered on \textbf{contact flow (CF)}, a compact representation consisting of key body trajectories and time-series binary contact signals. Leveraging this shared interface, our low-level policy \textbf{CF-Track} learns a unified library of loco-manipulation skills, while our high-level module \textbf{CF-Gen} heuristically synthesizes future contact-flow sequences. To support this setting, we additionally collect the OmniContact dataset, a MoCap-based HOI corpus for humanoid loco-manipulation (Appendix~\ref{sec:dataset}). Together, they enable robust execution, autonomous failure recovery, and flexible composition of meta-skills for long-horizon tasks. Experiments show that OmniContact achieves \(98.7\%\) success on \textit{Carry Box} and \(76.5\%\) on \textit{Push-Stack Boxes}, outperforming prior baselines by average margins of \(40.9\%\) in meta-skill and \(66.5\%\) in skill chaining. Besides, our framework naturally integrates with VLMs for semantic task decomposition, enabling complex, semantically grounded loco-manipulation behaviors, such as arranging scattered boxes into a heart shape.
△ Less
Submitted 24 June, 2026;
originally announced June 2026.
-
Anomalous charge density wave in a two-dimensional superatomic superconductor
Authors:
Boqin Song,
Shuaishuai Sun,
Zhongxu Wei,
Xinbo Wang,
Xiaoping Ma,
Kaifa Luo,
Lei Wang,
Jun Deng,
Xu Chen,
Tian Qian,
Shuya Xing,
Zhihai Cheng,
Jiangang Guo,
Tianping Ying,
Xiaolong Chen
Abstract:
The spatial modulation of electron density into a wave-like pattern, known as charge density wave (CDW), represents a fundamental quantum state that often coexists with superconductivity, quantum Hall states, axion insulating phases and etc. Conventional CDWs are mediated by longitudinal acoustic phonons, exhibit picometer-scale lattice distortions ($10^{-12}$--$10^{-11}$ m), and typically vanish…
▽ More
The spatial modulation of electron density into a wave-like pattern, known as charge density wave (CDW), represents a fundamental quantum state that often coexists with superconductivity, quantum Hall states, axion insulating phases and etc. Conventional CDWs are mediated by longitudinal acoustic phonons, exhibit picometer-scale lattice distortions ($10^{-12}$--$10^{-11}$ m), and typically vanish approaching the atomic limit. Here, we report a series of anomalous CDW behaviors in the 2D superatomic superconductor Au$_6$Te$_{12}$Se$_8$. Remarkably, its CDW is governed by transverse phonons, accompanied by an extraordinarily high real-space displacement of $\sim 4$ Ångström. Furthermore, we observe an exotic dimensional response persisting up to micrometer-scale thickness, a regime where other materials are already considered as bulk. Through liquid helium-temperature transmission electron microscopy, ultrafast pump-probe spectroscopy and transport measurements, we demonstrate a dramatic enhancement of the CDW transition temperature ($T_{\text{CDW}}$) from $<2$ K in the bulk to 110 K in approaching the ``superatomic limit''. Our findings not only reveal novel facets of both CDW and superatomic materials, but the competition between this anomalous CDW and superconductivity opens avenues for exploring unconventional electron-phonon interactions.
△ Less
Submitted 22 June, 2026;
originally announced June 2026.
-
AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation
Authors:
Huadai Liu,
Kaicheng Luo,
Wen Wang,
Qian Chen,
Bin Ma,
Xiangang Li,
Wei Xue
Abstract:
Unifying speech, sound, and music generation in one model is hindered by tradeoffs between fidelity, end-to-end training, in-context conditioning, and variable-length synthesis that no current paradigm fully resolves. To address this challenge, we present AudioCALM, a universal audio generation framework that extends autoregressive (AR) next-token prediction from discrete tokens to continuous audi…
▽ More
Unifying speech, sound, and music generation in one model is hindered by tradeoffs between fidelity, end-to-end training, in-context conditioning, and variable-length synthesis that no current paradigm fully resolves. To address this challenge, we present AudioCALM, a universal audio generation framework that extends autoregressive (AR) next-token prediction from discrete tokens to continuous audio latents: a thin flow-matching head replaces the softmax to predict rectified-flow velocities at each position, and a block-causal AR-Flow attention pattern produces arbitrary-length output. Joint training of multiple audio generation tasks faces an asymmetric text--audio mismatch: speech transcripts align to specific time spans and demand tight, time-aligned attention, whereas sound and music captions describe only overall semantics and rely on diffuse, holistic attention; mixing the two disproportionately degrades sound and music generation. We address this asymmetry at two levels: a data reformulation strategy that unifies all three tasks under a single description-style conditioning interface, and a novel architecture Asymmetric Mixture-of-Modality-Experts (A-MoME), which adds a dedicated residual expert for speech while sound and music share the backbone, incurring no inference overhead on non-speech inputs. Experimental results demonstrate that AudioCALM matches modality-specific state-of-the-art and outperforms prior unified baselines on speech, sound, and music generation benchmarks.
△ Less
Submitted 22 June, 2026;
originally announced June 2026.
-
STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation
Authors:
Huadai Liu,
Wen Wang,
Kaicheng Luo,
Qian Chen,
Xiangang Li,
Wei Xue
Abstract:
Continuous Variational Autoencoders (VAEs) serve as the fundamental continuous tokenizer for modern neural audio generation systems, enabling high-fidelity reconstruction while providing a compact, smooth latent space for downstream generative priors. However, continuous VAEs face a fundamental conflict among compression rate, reconstruction fidelity, and latent space topology, which we formalize…
▽ More
Continuous Variational Autoencoders (VAEs) serve as the fundamental continuous tokenizer for modern neural audio generation systems, enabling high-fidelity reconstruction while providing a compact, smooth latent space for downstream generative priors. However, continuous VAEs face a fundamental conflict among compression rate, reconstruction fidelity, and latent space topology, which we formalize as the Rate-Distortion-Regularity Trilemma. This trilemma stems from a topological mismatch: the isotropic Gaussian prior in standard VAEs imposes a flat latent geometry that fails to accommodate audio's hierarchical nature, where low-frequency components are structured and compressible while high-frequency components are stochastic and incompressible, leading to disordered information packing in which crucial semantic features are interleaved with high-entropy noise. To address this challenge, we propose Structured Topology-Aware Regularization (STAR), a general training strategy that reshapes latent space geometry by imposing a growth-based constraint field, routing structural and textural information into channel subspaces with matching capacities. STAR is applicable to any VAE architecture and effectively resolves the trilemma, as demonstrated in CNN-based VAEs. We further present STAR-VAE, which combines STAR with a hybrid CNN-Mamba architecture for local feature extraction and linear-complexity global context modeling, and STAR-Gen, an LLM-based Flow Matching framework that leverages STAR-VAE's structured latent space for high-fidelity generation without vector quantization artifacts. Experiments across diverse audio domains show that STAR-VAE achieves state-of-the-art reconstruction fidelity and enhanced semantic information preservation, while the structured latent space improves both traditional diffusion models and STAR-Gen for text-to-audio generation.
△ Less
Submitted 22 June, 2026;
originally announced June 2026.
-
DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners?
Authors:
Jadelynn Dao,
Milan Ganai,
Yasmina Abukhadra,
Ajay Sridhar,
Mozhgan Nasr Azadani,
Katie Luo,
Clark Barrett,
Jiajun Wu,
Chelsea Finn,
Marco Pavone
Abstract:
Vision-Language Models (VLMs) are increasingly deployed as high-level planners for embodied agents, with an emerging strategy of scaling test-time compute to improve capability. However, we observe that doing so increases latency, token usage, and FLOPs while yielding uneven, often diminishing gains in downstream success, limiting where embodied agents can be deployed. We argue that choosing when…
▽ More
Vision-Language Models (VLMs) are increasingly deployed as high-level planners for embodied agents, with an emerging strategy of scaling test-time compute to improve capability. However, we observe that doing so increases latency, token usage, and FLOPs while yielding uneven, often diminishing gains in downstream success, limiting where embodied agents can be deployed. We argue that choosing when and where to spend test-time compute is central to bringing frontier performance to the real world. We introduce DIRECT, a routing framework that uses multimodal scene context to allocate compute per prompt, improving the success--cost Pareto frontier over fixed model selection. Across three dominant scaling axes, namely chain-of-thought depth, model size, and memory history, our experiments on VLABench and RoboMME show that test-time compute is not a uniform lever: different axes yield qualitatively distinct capability gains. We validate these insights on a physical Franka arm in a DROID setup spanning zero-shot manipulation and long-horizon chaining, where our router matches or exceeds a stronger model's success rate at up to 65% lower average latency. Ultimately, our results show that naively scaling test-time compute is wasteful, and that DIRECT can provide frontier-level embodied planning in robotic systems at a fraction of the cost. Project page can be found at jadee-dao.github.io/direct/.
△ Less
Submitted 10 June, 2026;
originally announced June 2026.
-
Adaptive Patching Is Harder Than It Looks For Time-Series Forecasting
Authors:
Federico Zucchi,
Yi Xie,
Chao Zhang,
Keyuan Luo,
Thomas Lampert,
Ziyue Li
Abstract:
Adaptive patching is a recent and compelling proposal for time-series Transformers: allocate finer patches where the sequence looks locally informative. This paper asks under what conditions a content-adaptive patching operator should outperform a tuned uniform one. Local heterogeneity alone is not enough: under pointwise forecasting losses, a complex-looking region is not automatically one where…
▽ More
Adaptive patching is a recent and compelling proposal for time-series Transformers: allocate finer patches where the sequence looks locally informative. This paper asks under what conditions a content-adaptive patching operator should outperform a tuned uniform one. Local heterogeneity alone is not enough: under pointwise forecasting losses, a complex-looking region is not automatically one where finer patching reduces the loss. We model patching as a budgeted bitrate allocation and derive an explicit threshold that a dynamic patching rule must satisfy to beat a well-tuned uniform baseline, then bound the achievable improvement both locally (a quadratic surrogate) and globally (a strong-convexity bound under the model's assumptions). Two structural results follow: without a coupling constraint, scalar local complexity cannot produce a non-uniform optimum under a common loss landscape; and once the backbone is trained to its representation-aware optimum, the alignment gain collapses around a well-tuned uniform patch size. To test these predictions, we run a controlled isolation study on three representative architectures, replacing each adaptive mechanism with a uniform patch-size sweep while keeping the backbone, data, and training protocol fixed. On standard long-horizon forecasting benchmarks, the validation-selected uniform baseline is competitive with the dynamic counterpart, with per-setting effects concentrated near zero and no consistent directional advantage once results are aggregated by dataset. The larger gains we do observe are method- and dataset-specific. Adaptive patching should therefore be evaluated against a tuned uniform baseline; its value depends on whether a cheap and reliable routing signal can identify where finer patches actually reduce forecasting loss.
△ Less
Submitted 2 June, 2026;
originally announced June 2026.
-
EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs
Authors:
Liang Lin,
Chunxi Luo,
Kaiwen Luo,
Jie Zhang,
Jin Wang,
Yuanhe Zhang,
Cai Yuchen,
Qiankun Li,
Gongli Xi,
Zhenhong Zhou,
Kun Wang,
Junhao Dong
Abstract:
Audio Large Language Models (ALLMs) are highly vulnerable to real-world noise, which often induces severe semantic drift and hallucinations. Existing robustness methods primarily rely on waveform-level acoustic enhancement, answer-level supervision, or the internal suppression of noise representations. To address these issues, we propose echodistill, an alignment-based noisy-to-clean self-distilla…
▽ More
Audio Large Language Models (ALLMs) are highly vulnerable to real-world noise, which often induces severe semantic drift and hallucinations. Existing robustness methods primarily rely on waveform-level acoustic enhancement, answer-level supervision, or the internal suppression of noise representations. To address these issues, we propose echodistill, an alignment-based noisy-to-clean self-distillation framework. Echodistill leverages a frozen clean-audio teacher to provide semantic references for an inference-time noisy-audio student. Specifically, the student samples candidate responses under noisy conditions to expose its test-time behavior. These trajectories are then optimized via group-relative policy optimization (GRPO), where the token-level consistency with the teacher acts as a reward bonus. By aligning the noisy student's candidate responses with clean semantic evidence, and applying audio-aware reward shaping, our method encourages reasoning trajectories that are both correct and genuinely acoustically grounded. Echodistill significantly improves the semantic reliability and task performance of Audio LLMs under complex noise, without introducing any additional inference costs. Extensive experiments show that: (I) Compared with the strongest baseline, echodistill achieves average improvements of 4.18\%$\uparrow$ in GSR under strong noise. (II) Ablation results on Qwen-Omni further show that echodistill improves over the GRPO-only variant by 3.02\%$\uparrow$ in Acc, 3.89\%$\uparrow$ in Noisy, and 4.53\%$\uparrow$ in GSR on average. Our codes are available at https://anonymous.4open.science/r/echodistill-10DE.
△ Less
Submitted 11 May, 2026;
originally announced May 2026.
-
Bridging Structure and Language: Graph-Based Visual Reasoning for Autonomous Road Understanding
Authors:
Lena Wild,
Katie Z Luo,
Marco Pavone
Abstract:
Structured road understanding of lane geometry, topology, and traffic element relationships is foundational to safe autonomous driving. While vision-language models (VLMs) offer promising semantic flexibility, they lack the geometric and relational grounding required for precise road reasoning. Conversely, traditional modular systems, e.g., HD maps and topological road graphs, provide structural p…
▽ More
Structured road understanding of lane geometry, topology, and traffic element relationships is foundational to safe autonomous driving. While vision-language models (VLMs) offer promising semantic flexibility, they lack the geometric and relational grounding required for precise road reasoning. Conversely, traditional modular systems, e.g., HD maps and topological road graphs, provide structural precision but remain semantically rigid. To bridge this gap, we introduce the Combined Road Substrate (CRS), a graph-grounded framework that makes geometric road structure and open-vocabulary semantics jointly executable in a single representation. CRS enables the automatic generation of compositionally complex and linguistically varied question-answer pairs via recursive graph queries, augmented with a "grounding for free" mechanism that ensures logical traceability to specific map elements, and procedurally extracted chain-of-thought supervision traces. We demonstrate that state-of-the-art VLMs - including large, closed-source models - struggle significantly with structured road reasoning, yet training a small 2- or 4-billion-parameter model with as few as 20 to 80 CRS-enriched scenes yields stable gains in compositional reasoning tasks of varying depth. Analysis of model behavior via verifiable reasoning traces reveals a systematic shift in failure modes: whereas baseline models fail at relational scene understanding, CRS-trained models reduce failures to attribute recognition, suggesting that the primary bottleneck in road understanding is not model scale, but the absence of structured supervision.
△ Less
Submitted 20 May, 2026;
originally announced May 2026.
-
A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook
Authors:
Kaiwen Luo,
Zhenhong Zhou,
Leyan Wang,
Liang Lin,
Tianyu Shao,
Yuanhe Zhang,
Yang Xiao,
Yuxuan Li,
Miao Yu,
Kailin Lyu,
Jiaming Zhang,
Li Sun,
Songze Li,
Yueming Wu,
Ting Dang,
Xiaojun Jia,
Dongrui Liu,
Kai Li,
Rohan Kumar Das,
Siyuan Liang,
Xinfeng Li,
Qiankun Li,
Jing Chen,
Xingjun Ma,
Kun Wang
, et al. (10 additional authors not shown)
Abstract:
Advances in Large Language Models (LLMs) have paved the way for Multimodal Large Language Models (MLLMs). Among these, Large Audio Language Models (LALMs) are essential for realizing universal auditory intelligence. Despite their remarkable performance, the escalation of LALMs' capabilities has significantly outpaced the development of systemic frameworks to ensure their trustworthiness. This surv…
▽ More
Advances in Large Language Models (LLMs) have paved the way for Multimodal Large Language Models (MLLMs). Among these, Large Audio Language Models (LALMs) are essential for realizing universal auditory intelligence. Despite their remarkable performance, the escalation of LALMs' capabilities has significantly outpaced the development of systemic frameworks to ensure their trustworthiness. This survey provides a comprehensive investigation into the endogenous mechanisms of LALMs, detailing the architectural innovations and alignment algorithms that facilitate emergent reasoning. Specifically, we analyze how the transition to unified end-to-end frameworks and the integration of continuous acoustic signals expand the attack surface. To rigorously evaluate the risks within these paradigms, we establish a comprehensive taxonomy of trustworthiness, categorizing critical vulnerabilities such as cross-modal jailbreaking, latent acoustic backdoors, and biometric privacy leakage. We review the state-of-the-art LALMs through six analytical pillars: hallucination, robustness, safety, privacy, fairness, and authentication. The pronounced imbalance between a mature offensive landscape and underdeveloped defenses highlights persistent trustworthiness gaps and multidimensional risks in audio-centric intelligence. Finally, we propose a roadmap advocating for ``Defense-in-Depth'' architectures, causal auditory world modeling, and intrinsic representation engineering to support the development of more reliable and trustworthy audio intelligence. Our project has been uploaded to GitHub https://github.com/Kwwwww74/Awesome-Trustworthy-AudioLLMs.
△ Less
Submitted 3 August, 2026; v1 submitted 18 May, 2026;
originally announced May 2026.
-
Grouped Annulus-Modulated Transceiver Is Almost Full DoF-Achieving for RIS-Assisted Symbiotic Radios Over Spatial-Correlated Channels
Authors:
Ruo-Qi Sun,
Jianfeng Shi,
Yonggang Zhu,
Mingliang Xie,
Kang Luo,
Yifu Sun,
Ru-Han Chen,
Kang An
Abstract:
This paper considers a RIS-assisted symbiotic communication system, where additional information is conveyed by the passive reconfigurable intelligent surface (RIS). In existing schemes, individual phase modulation is usually adopted at the RIS elements, which severely limits exploiting all extra multiplexing gains brought by the RIS. To address the issue, we propose a novel matrix decomposition a…
▽ More
This paper considers a RIS-assisted symbiotic communication system, where additional information is conveyed by the passive reconfigurable intelligent surface (RIS). In existing schemes, individual phase modulation is usually adopted at the RIS elements, which severely limits exploiting all extra multiplexing gains brought by the RIS. To address the issue, we propose a novel matrix decomposition algorithm that transforms the equivalent channel into a structured form while effectively suppressing the decomposition residual. Based on this, a novel transceiver architecture employing grouped annulus modulation (GAM) with a hexagonal-lattice-based constellation is developed, which is capable of achieving the full degrees of freedom (DoFs) when the decomposition algorithm performs as expected. Numerical results demonstrate that the proposed transceiver achieves much higher communication rates, thereby leading to higher spectral efficiency, compared to the conventional phase-only modulation scheme, while maintaining comparable error performance.
△ Less
Submitted 13 May, 2026;
originally announced May 2026.
-
Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion
Authors:
ShiYing Huang,
Liang Lin,
Yuer Li,
Kaiwen Luo,
Zhenhong Zhou,
An Zhang,
Junhao Dong,
Kun Wang,
Zhigang Zeng
Abstract:
In the realm of multi-objective alignment for large language models, balancing disparate human preferences often manifests as a zero-sum conflict. Specifically, the intrinsic tension between competing goals dictates that aggressively optimizing for one metric (e.g., helpfulness) frequently incurs a substantial penalty on another (e.g., harmlessness). While prior work mainly focuses on data selecti…
▽ More
In the realm of multi-objective alignment for large language models, balancing disparate human preferences often manifests as a zero-sum conflict. Specifically, the intrinsic tension between competing goals dictates that aggressively optimizing for one metric (e.g., helpfulness) frequently incurs a substantial penalty on another (e.g., harmlessness). While prior work mainly focuses on data selection, parameter merging, or algorithmic balancing during training, these approaches merely force compromises between divergent preferences along a fixed Pareto frontier, failing to fundamentally resolve the inherent trade-off. In this work, we approach this problem from a novel perspective of multi-dimensional rewards. By scaling up the model's rollouts and analyzing the outputs across different reward dimensions, we arrive at a critical conclusion: the conflict among multiple objectives stems from the fact that the prompt itself inherently restricts the achievable multi-dimensional rewards. Based on this core observation, we propose MORA: Multi-Objective Reward Assimilation. Specifically, MORA isolates single-reward prompts through pre-sampling and expands their reward diversity by rewriting the original questions to incorporate multi-dimensional intents. Extensive experiments demonstrate that: (1) in sequential alignment, MORA achieves single-preference improvements ranging from 5% to 12.4%, with exceptional gains in harmlessness, after multiple-preference alignment across helpful, harmless, and truthful dimensions. (2) In simultaneous alignment, MORA achieves an average overall reward improvement of 4.6%. Our codes are available at https://github.com/Shiying-Huang/MORA-MPA.
△ Less
Submitted 13 May, 2026; v1 submitted 12 May, 2026;
originally announced May 2026.
-
Contour-Native Bridge Defect Detection and Compact Digital Archiving with Frequency-Supervised Fourier Contours
Authors:
Jin Liu,
Wang Wang,
Hongxu Pu,
Zhen Cao,
Yasong Wang,
Hu Wang,
Kunming Luo
Abstract:
AI-assisted bridge defect inspection often produces bounding boxes with crude geometry or raster masks that are costly to store, transmit, and reuse. This study investigates how detected defects can be represented as compact, recoverable contour-level vector records in image space. We propose Frequency-Supervised Fourier Series Detection (FS-FSD), which directly regresses Fourier contour descripto…
▽ More
AI-assisted bridge defect inspection often produces bounding boxes with crude geometry or raster masks that are costly to store, transmit, and reuse. This study investigates how detected defects can be represented as compact, recoverable contour-level vector records in image space. We propose Frequency-Supervised Fourier Series Detection (FS-FSD), which directly regresses Fourier contour descriptors and evaluates boxes, masks, and contours under a unified polygon-space protocol. On 3,767 UAV-collected bridge images with 42,346 defect instances, FS-FSD achieves higher polygon-space accuracy and better matched-TP geometric quality than representative detection, segmentation, and contour baselines. These results show that, compared with bounding boxes and raster masks, Fourier contour records preserve defect-boundary geometry in a more compact, recoverable, and shareable form for engineering review and downstream information workflows. Future work will study the modeling of multi-region, fragmented, and adjacent bridge-defect boundaries and extend the framework toward long-term bridge-defect tracking and lifecycle-oriented management.
△ Less
Submitted 9 May, 2026;
originally announced May 2026.
-
When to Trust Imagination: Adaptive Action Execution for World Action Models
Authors:
Rui Wang,
Yue Zhang,
Jiehong Lin,
Kuncheng Luo,
Jianan Wang,
Zhongrui Wang,
Xiaojuan Qi
Abstract:
World Action Models (WAMs) have recently emerged as a promising paradigm for robotic manipulation by jointly predicting future visual observations and future actions. However, current WAMs typically execute a fixed number of predicted actions after each model inference, leaving the robot blind to whether the imagined future remains consistent with the actual physical rollout. In this work, we form…
▽ More
World Action Models (WAMs) have recently emerged as a promising paradigm for robotic manipulation by jointly predicting future visual observations and future actions. However, current WAMs typically execute a fixed number of predicted actions after each model inference, leaving the robot blind to whether the imagined future remains consistent with the actual physical rollout. In this work, we formulate adaptive WAM execution as a future-reality verification problem: the robot should execute longer when the WAM-predicted future remains reliable, and replan earlier when reality deviates from imagination. To this end, we propose Future Forward Dynamics Causal Attention (FFDC), a lightweight verifier that jointly reasons over predicted future actions, predicted visual dynamics, real observations, and language instructions to estimate whether the remaining action rollout can still be trusted. FFDC enables adaptive action chunk sizes as an emergent consequence of prediction-observation consistency, preserving the efficiency of long-horizon execution while restoring responsiveness in contact-rich or difficult phases. We further introduce Mixture-of-Horizon Training to improve long-horizon trajectory coverage for adaptive execution. Experiments on the RoboTwin benchmark and in the real world demonstrate that our method achieves a strong robustness-efficiency trade-off: on RoboTwin, it reduces WAM forward passes by 69.10% and execution time by 34.02%, while improving success rate by 2.54% over the short-chunk baseline; in real-world experiments, it improves success rate by 35%.
△ Less
Submitted 9 May, 2026; v1 submitted 7 May, 2026;
originally announced May 2026.
-
From Context to Skills: Can Language Models Learn from Context Skillfully?
Authors:
Shuzheng Si,
Haozhe Zhao,
Yu Lei,
Qingyi Wang,
Dingwei Chen,
Zhitong Wang,
Zhenhailong Wang,
Kangyang Luo,
Zheng Wang,
Gang Chen,
Fanchao Qi,
Minjia Zhang,
Maosong Sun
Abstract:
Many real-world tasks require language models (LMs) to reason over complex contexts that exceed their parametric knowledge. This calls for context learning, where LMs directly learn relevant knowledge from the given context. An intuitive solution is inference-time skill augmentation: extracting the rules and procedures from context into natural-language skills. However, constructing such skills fo…
▽ More
Many real-world tasks require language models (LMs) to reason over complex contexts that exceed their parametric knowledge. This calls for context learning, where LMs directly learn relevant knowledge from the given context. An intuitive solution is inference-time skill augmentation: extracting the rules and procedures from context into natural-language skills. However, constructing such skills for context learning scenarios faces two challenges: the prohibitive cost of manual skill annotation for long, technically dense contexts, and the lack of external feedback for automated skill construction. In this paper, we propose Ctx2Skill, a self-evolving framework that autonomously discovers, refines, and selects context-specific skills without human supervision or external feedback. At its core, a multi-agent self-play loop has a Challenger that generates probing tasks and rubrics, a Reasoner that attempts to solve them guided by an evolving skill set, and a neutral Judge that provides binary feedback. Crucially, both the Challenger and the Reasoner evolve through accumulated skills: dedicated Proposer and Generator agents analyze failure cases and synthesize them into targeted skill updates for both sides, enabling automated skill discovery and refinement. To prevent adversarial collapse caused by increasingly extreme task generation and over-specialized skill accumulation, we further introduce a Cross-time Replay mechanism that identifies the skill set achieving the best balance across representative cases for the Reasoner side, ensuring robust and generalizable skill evolution. The resulting skills can be plugged into any language model to obtain better context learning capability. Evaluated on four context learning tasks from CL-bench, Ctx2Skill consistently improves solving rates across backbone models.
△ Less
Submitted 27 July, 2026; v1 submitted 30 April, 2026;
originally announced April 2026.
-
AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
Authors:
Lei Xiong,
Kun Luo,
Ziyi Xia,
Wenbo Zhang,
Jin-Ge Yao,
Zheng Liu,
Jingying Shao,
Jianlyu Chen,
Hongjin Qian,
Xi Yang,
Qian Yu,
Hao Li,
Chen Yue,
Xiaan Du,
Yuyang Wang,
Yesheng Liu,
Haiyu Xu,
Zhicheng Dou
Abstract:
Autonomous scientific research is significantly advanced thanks to the development of AI agents. One key step in this process is finding the right scientific literature, whether to explore existing knowledge for a research problem, or to acquire evidence for verifying assumptions and supporting claims. To assess AI agents' capability in driving this process, we present AutoResearchBench, a dedicat…
▽ More
Autonomous scientific research is significantly advanced thanks to the development of AI agents. One key step in this process is finding the right scientific literature, whether to explore existing knowledge for a research problem, or to acquire evidence for verifying assumptions and supporting claims. To assess AI agents' capability in driving this process, we present AutoResearchBench, a dedicated benchmark for autonomous scientific literature discovery. AutoResearchBench consists of two complementary task types: (1) Deep Research, which requires tracking down a specific target paper through a progressive, multi-step probing process, and (2) Wide Research, which requires comprehensively collecting a set of papers satisfying given conditions. Compared to previous benchmarks on agentic web browsing, AutoResearchBench is distinguished along three dimensions: it is research-oriented, calling for in-depth comprehension of scientific concepts; literature-focused, demanding fine-grained utilization of detailed information; and open-ended, involving an unknown number of qualified papers and thus requiring deliberate reasoning and search throughout. These properties make AutoResearchBench uniquely suited for evaluating autonomous research capabilities, and extraordinarily challenging. Even the most powerful LLMs, despite having largely conquered general agentic web-browsing benchmarks such as BrowseComp, achieve only 9.39% accuracy on Deep Research and 9.31% IoU on Wide Research, while many other strong baselines fall below 5%. We publicly release the dataset and evaluation pipeline to facilitate future research in this direction. We publicly release the dataset, evaluation pipeline, and code at https://github.com/CherYou/AutoResearchBench.
△ Less
Submitted 28 April, 2026;
originally announced April 2026.
-
Functional Dismantling of Network Relaxation through Slow-Branch Susceptibility
Authors:
Kaiming Luo,
Huiying Zhou
Abstract:
Robustness of relaxation on asymmetric networks is not determined by connectivity alone, because the slow collective mode can be complex and may change its spectral identity under adaptive damage. We introduce a slow-branch susceptibility framework for functional dismantling of network relaxation. Starting from the projected relaxation dynamics, we show that the relevant robustness observable is t…
▽ More
Robustness of relaxation on asymmetric networks is not determined by connectivity alone, because the slow collective mode can be complex and may change its spectral identity under adaptive damage. We introduce a slow-branch susceptibility framework for functional dismantling of network relaxation. Starting from the projected relaxation dynamics, we show that the relevant robustness observable is the real part of the selected nonzero Laplacian branch, which controls the long-time decay of the nonstationary sector. Node deletion is then treated as a dimension-changing compression of the operator, leading to a modal susceptibility (MS) score that estimates the first-order reduction of the branch-tracked relaxation rate from the biorthogonal support of the slow mode. In the reciprocal limit, the same construction reduces to the weighted Fiedler sector, placing directed and weighted-undirected networks within a common spectral-response formulation. Tests on synthetic and real-world networks show that MS identifies vulnerability patterns that differ from standard centrality-based attacks and edge-level spectral proxies. These results resolve a modal-selection ambiguity in non-Hermitian robustness analysis and provide a spectral basis for functional dismantling in asymmetric networks.
△ Less
Submitted 27 April, 2026;
originally announced April 2026.
-
KARL: Mitigating Hallucinations in LLMs via Knowledge-Boundary-Aware Reinforcement Learning
Authors:
Cheng Gao,
Cheng Huang,
Kangyang Luo,
Ziqing Qiao,
Shuzheng Si,
Huimin Chen,
Chaojun Xiao,
Maosong Sun
Abstract:
Enabling large language models (LLMs) to appropriately abstain from answering questions beyond their knowledge is crucial for mitigating hallucinations. While existing reinforcement learning methods foster autonomous abstention, they often compromise answer accuracy because their static reward mechanisms, agnostic to models' knowledge boundaries, drive models toward excessive caution. In this work…
▽ More
Enabling large language models (LLMs) to appropriately abstain from answering questions beyond their knowledge is crucial for mitigating hallucinations. While existing reinforcement learning methods foster autonomous abstention, they often compromise answer accuracy because their static reward mechanisms, agnostic to models' knowledge boundaries, drive models toward excessive caution. In this work, we propose KARL, a novel framework that continuously aligns an LLM's abstention behavior with its evolving knowledge boundary. KARL introduces two core innovations: a Knowledge-Boundary-Aware Reward that performs online knowledge boundary estimation using within-group response statistics, dynamically rewarding correct answers or guided abstention; and a Two-Stage RL Training Strategy that first explores the knowledge boundary and bypasses the "abstention trap", and subsequently converts incorrect answers beyond the knowledge boundary into abstentions without sacrificing accuracy. Extensive experiments on multiple benchmarks demonstrate that KARL achieves a superior accuracy-hallucination trade-off, effectively suppressing hallucinations while maintaining high accuracy across both in-distribution and out-of-distribution scenarios.
△ Less
Submitted 3 April, 2026;
originally announced April 2026.
-
Synchrotron polarization of anisotropic electron distribution in GRB prompt emission
Authors:
Kang-Fa Cheng,
Kai-Xian Luo,
Xiao-Hong Zhao,
Jirong Mao,
Hong-Bang Liu,
Yu-Hang Mo,
Jin-Rong Huang,
Rong-Li Weng,
Wen-Jie Xie,
Gao-Jin Yu
Abstract:
In gamma-ray bursts (GRBs), the electron pitch angle ($α$) is usually assumed to be isotropically distributed. However, recent numerical simulations indicate that only the high-energy electrons (with Lorentz factors $γ>γ_{iso}$) are distributed isotropically, whereas the low-energy electrons (with $γ<γ_{iso}$) follow an energy-dependent anisotropic distribution during magnetic reconnection. The me…
▽ More
In gamma-ray bursts (GRBs), the electron pitch angle ($α$) is usually assumed to be isotropically distributed. However, recent numerical simulations indicate that only the high-energy electrons (with Lorentz factors $γ>γ_{iso}$) are distributed isotropically, whereas the low-energy electrons (with $γ<γ_{iso}$) follow an energy-dependent anisotropic distribution during magnetic reconnection. The mean value of $\sin^2 α$ approximately follows the relation $\langle \sin^2 α\rangle \propto γ^{m}$ for $γ<γ_{iso}$. In principle, polarization measurements may help us constrain the pitch-angle distribution of electrons in GRBs, since different pitch-angle distributions produce distinct synchrotron polarization signatures. The polarization of GRBs produced by isotropically distributed electrons has been extensively studied. In this paper, we investigate synchrotron polarization produced by anisotropically distributed electrons within a globally toroidal magnetic field in GRB prompt emission. Our results show that the synchrotron PDs in the $γ$-ray and X-ray bands produced by anisotropically distributed electrons are systematically lower than those produced by isotropically distributed electrons, while the PD in the optical band could be either lower or higher than that of isotropically distributed electrons, depending primarily on the value of the energy slope $m$. In addition, we compared our numerical results with observational data, and the comparison suggests that an anisotropic distribution of electrons may offer a potential explanation for the PD and spectral data of some GRBs.
△ Less
Submitted 24 April, 2026;
originally announced April 2026.
-
Effective Traveling for Metric Instances of the Traveling Thief Problem
Authors:
Jan Eube,
Kelin Luo,
Aneta Neumann,
Frank Neumann,
Heiko Röglin
Abstract:
The Traveling Thief Problem (TTP) is a multi-component optimization problem that captures the interplay between routing and packing decisions by combining the classical Traveling Salesperson Problem (TSP) and the Knapsack Problem (KP). The TTP has gained significant attention in the evolutionary computation literature and a wide range of approaches have been developed over the last 10 years. Judgi…
▽ More
The Traveling Thief Problem (TTP) is a multi-component optimization problem that captures the interplay between routing and packing decisions by combining the classical Traveling Salesperson Problem (TSP) and the Knapsack Problem (KP). The TTP has gained significant attention in the evolutionary computation literature and a wide range of approaches have been developed over the last 10 years. Judging the performance of these algorithms in particular in terms of how close the get to optimal solutions is a very challenging task as effective exact methods are not available due to the highly challenging traveling component. In this paper, we study the tour-optimization component of TTP under a fixed packing plan. We formulate this task as a weighted variant of the TSP, where travel costs depend on the cumulative weight of collected items, and investigate how different distance metrics and cost functions affect computational complexity. We present an $(O(n^2))$-time dynamic programming algorithm for the path metric with general cost functions, prove that the problem is NP-hard even on a star metric, and develop constant-factor approximation algorithms for star metrics. Finally, we also develop an approximation algorithm for the problem under a general metric with a linear cost function.
We complement our theoretical results with experimental evaluations on standard TTP instances adjusted to a path metric. Our experimental results demonstrate the practical effectiveness of our approaches by comparing it to solutions produced by popular iterative search algorithms. The results show that our methods are able to significantly improve the quality of solutions for some benchmark instances by optimizing the traveling part while pointing out the optimality of the travel component for other solutions obtained by iterative search methods.
△ Less
Submitted 21 April, 2026;
originally announced April 2026.
-
Negative Advantages Is a Double-Edged Sword: Calibrating advantages in GRPO for Search Agents
Authors:
Jiayi Wu,
Ruobing Xie,
Zeqian Huang,
Lei Jiang,
Can Xu,
Kangyang Luo,
Bochen Lin,
Ming Gao,
Xiang Li
Abstract:
Search agents achieve strong question-answering performance through multi-turn interactions with search engines, with Group Relative Policy Optimization (GRPO) being a widely used training algorithm. However, GRPO-style algorithms still face several challenges in multi-hop search settings. First, correct intermediate steps are often penalized when the final answer is wrong. Second, training is hig…
▽ More
Search agents achieve strong question-answering performance through multi-turn interactions with search engines, with Group Relative Policy Optimization (GRPO) being a widely used training algorithm. However, GRPO-style algorithms still face several challenges in multi-hop search settings. First, correct intermediate steps are often penalized when the final answer is wrong. Second, training is highly unstable, often causing degradation of natural language ability or even catastrophic training collapse. Our analysis attributes these issues to coarse-grained advantage assignment and an imbalance between positive and negative advantages. To address these problems, we propose CalibAdv, an advantage calibration method specifically designed for search agents that enables more accurate and more stable modeling of penalties and rewards. Specifically, CalibAdv leverages the correctness of intermediate steps to downscale excessive negative advantages at a fine-grained level. It then further rebalances positive and negative advantages to improve training stability. Importantly, CalibAdv adopts a lightweight design that calibrates advantages from standard rollout signals, making it simple and easy to deploy. Extensive experiments across three models and seven benchmarks demonstrate that CalibAdv improves both model performance and training stability. Our code is available at https://github.com/wujwyi/CalibAdv.
△ Less
Submitted 27 May, 2026; v1 submitted 20 April, 2026;
originally announced April 2026.
-
Location of the liquid-vapor critical point in aluminum
Authors:
Xuyang Long,
Kai Luo
Abstract:
The precise location of the liquid-vapor critical point in aluminum has remained elusive for decades, with reported critical temperatures spanning nearly 4000 K. Here we resolve this long-standing uncertainty by combining deep potential molecular dynamics with large-scale simulations trained on high-fidelity electronic-structure data. We benchmark multiple exchange-correlation functionals against…
▽ More
The precise location of the liquid-vapor critical point in aluminum has remained elusive for decades, with reported critical temperatures spanning nearly 4000 K. Here we resolve this long-standing uncertainty by combining deep potential molecular dynamics with large-scale simulations trained on high-fidelity electronic-structure data. We benchmark multiple exchange-correlation functionals against experimental liquid densities and identify PBEsol as providing the most consistent description. Using complementary approaches -- spinodal analysis of the equation of state and direct coexistence simulations with Gaussian mixture phase identification -- we converge on a critical temperature of 6531-6576 $^\circ$K, a critical density of $0.637$ g/cm$^{3}$, and a critical pressure of $1.6$ kbar. The precision of these values, with temperature uncertainties of $\sim$50 K, represents a marked improvement over previous estimates. Our framework establishes a transferable strategy for predicting critical phenomena in metals, with implications for laser ablation, shock compression, and planetary modeling under extreme conditions.
△ Less
Submitted 2 May, 2026; v1 submitted 12 April, 2026;
originally announced April 2026.
-
Shear, Not Coherence, Organizes chaotic response under Higher-Order Coupling
Authors:
Kaiming Luo
Abstract:
What dynamical quantity is actually controlled by higher-order interactions in chaotic oscillator networks remains unclear. In amplitude-active systems, chaos is often interpreted through coherence, yet coherence is not the quantity that governs instability. In this work, we study a minimal globally coupled quartet of nonisochronous Stuart-Landau oscillators with pairwise and symmetric three-body…
▽ More
What dynamical quantity is actually controlled by higher-order interactions in chaotic oscillator networks remains unclear. In amplitude-active systems, chaos is often interpreted through coherence, yet coherence is not the quantity that governs instability. In this work, we study a minimal globally coupled quartet of nonisochronous Stuart-Landau oscillators with pairwise and symmetric three-body interactions. The pairwise baseline already supports a connected chaotic branch, and higher-order coupling reconstructs rather than creates this irregular dynamics. We show that chaos is organized not by phase coherence but by effective-frequency shear: higher-order coupling regulates amplitude heterogeneity, which nonisochronicity converts into shear, and shear controls how chaos is expressed under higher-order coupling. The Lyapunov response collapses onto a reduced shear-based description, revealing an indirect control pathway. These results establish that higher-order interactions control chaos only indirectly, by regulating an amplitude-shear mechanism rather than acting directly on synchrony.
△ Less
Submitted 9 April, 2026;
originally announced April 2026.
-
LongCat-Next: Lexicalizing Modalities as Discrete Tokens
Authors:
Meituan LongCat Team,
Bin Xiao,
Chao Wang,
Chengjiang Li,
Chi Zhang,
Chong Peng,
Hang Yu,
Hao Yang,
Haonan Yan,
Haoze Sun,
Haozhe Zhao,
Hong Liu,
Hui Su,
Jiaqi Zhang,
Jiawei Wang,
Jing Li,
Kefeng Zhang,
Manyuan Zhang,
Minhao Jing,
Peng Pei,
Quan Chen,
Taofeng Xue,
Tongxin Pan,
Xiaotong Li,
Xiaoyang Li
, et al. (64 additional authors not shown)
Abstract:
The prevailing Next-Token Prediction (NTP) paradigm has driven the success of large language models through discrete autoregressive modeling. However, contemporary multimodal systems remain language-centric, often treating non-linguistic modalities as external attachments, leading to fragmented architectures and suboptimal integration. To transcend this limitation, we introduce Discrete Native Aut…
▽ More
The prevailing Next-Token Prediction (NTP) paradigm has driven the success of large language models through discrete autoregressive modeling. However, contemporary multimodal systems remain language-centric, often treating non-linguistic modalities as external attachments, leading to fragmented architectures and suboptimal integration. To transcend this limitation, we introduce Discrete Native Autoregressive (DiNA), a unified framework that represents multimodal information within a shared discrete space, enabling a consistent and principled autoregressive modeling across modalities. A key innovation is the Discrete Native Any-resolution Visual Transformer (dNaViT), which performs tokenization and de-tokenization at arbitrary resolutions, transforming continuous visual signals into hierarchical discrete tokens. Building on this foundation, we develop LongCat-Next, a native multimodal model that processes text, vision, and audio under a single autoregressive objective with minimal modality-specific design. As an industrial-strength foundation model, it excels at seeing, painting, and talking within a single framework, achieving strong performance across a wide range of multimodal benchmarks. In particular, LongCat-Next addresses the long-standing performance ceiling of discrete vision modeling on understanding tasks and provides a unified approach to effectively reconcile the conflict between understanding and generation. As an attempt toward native multimodality, we open-source the LongCat-Next and its tokenizers, hoping to foster further research and development in the community. GitHub: https://github.com/meituan-longcat/LongCat-Next
△ Less
Submitted 29 March, 2026;
originally announced March 2026.
-
AutoWeather4D: Autonomous Driving Video Weather Conversion via G-Buffer Dual-Pass Editing
Authors:
Tianyu Liu,
Weitao Xiong,
Kunming Luo,
Manyuan Zhang,
Peng Li,
Yuan Liu,
Ping Tan
Abstract:
Generative video models have significantly advanced the photorealistic synthesis of adverse weather for autonomous driving; however, they consistently demand massive datasets to learn rare weather scenarios. While 3D-aware editing methods alleviate these data constraints by augmenting existing video footage, they are fundamentally bottlenecked by costly per-scene optimization and suffer from inher…
▽ More
Generative video models have significantly advanced the photorealistic synthesis of adverse weather for autonomous driving; however, they consistently demand massive datasets to learn rare weather scenarios. While 3D-aware editing methods alleviate these data constraints by augmenting existing video footage, they are fundamentally bottlenecked by costly per-scene optimization and suffer from inherent geometric and illumination entanglement. In this work, we introduce AutoWeather4D, a feed-forward 3D-aware weather editing framework designed to explicitly decouple geometry and illumination. At the core of our approach is a G-buffer Dual-pass Editing mechanism. The Geometry Pass leverages explicit structural foundations to enable surface-anchored physical interactions, while the Light Pass analytically resolves light transport, accumulating the contributions of local illuminants into the global illumination to enable dynamic 3D local relighting. Extensive experiments demonstrate that AutoWeather4D achieves comparable photorealism and structural consistency to generative baselines while enabling fine-grained parametric physical control, serving as a practical data engine for autonomous driving.
△ Less
Submitted 1 April, 2026; v1 submitted 27 March, 2026;
originally announced March 2026.
-
Hidden Higher-Order Vulnerabilities in Simplicial Complexes Revealed by Branch-Consistent Functional Robustness
Authors:
Kaiming Luo
Abstract:
Robustness of higher-order networks is often quantified by the instantaneous smallest positive eigenvalue of the Hodge $1$-Laplacian under simplex deletion. We show that this observable is generically ill-defined: along a deletion trajectory, eigenvalue branches can switch, so the quantity being monitored may correspond to different nonharmonic modes at different steps. The primary issue is theref…
▽ More
Robustness of higher-order networks is often quantified by the instantaneous smallest positive eigenvalue of the Hodge $1$-Laplacian under simplex deletion. We show that this observable is generically ill-defined: along a deletion trajectory, eigenvalue branches can switch, so the quantity being monitored may correspond to different nonharmonic modes at different steps. The primary issue is therefore definitional rather than algorithmic. We resolve it by fixing the first nonharmonic branch of the intact complex and following that same branch throughout the damage process, which defines a branch-consistent functional robustness. Triangle sensitivities then follow directly from first-order perturbation theory, making the resulting mode-sensitive deletion protocol a consequence of the observable itself rather than an independent heuristic. Across synthetic and empirical clique complexes, removing only a small fraction of triangles is sufficient to drive the tracked mode to collapse, while graph-level observables remain unchanged because the $1$-skeleton is exactly preserved. The same framework also reveals bridge-like localization of functionally critical simplices and provides a compact predictor of dynamical timescales.
△ Less
Submitted 25 March, 2026;
originally announced March 2026.