-
Local B-site chemistry controls oxygen-vacancy energetics in Ca-Ce-Ti-Mn perovskites for thermochemical hydrogen production
Authors:
Manish Kumar,
Natalia Ali,
Matthew D. Witman,
Shang Zhai,
James E. Miller,
Ivan Ermanoski,
Ellen B. Stechel,
Robert B. Wexler
Abstract:
Two-step thermochemical water splitting driven by concentrated solar heat is a scalable route to renewable hydrogen, but it requires oxides whose oxygen-vacancy formation energies balance facile reduction with favorable reoxidation. Perovskite solid solutions can tune this balance, but the relationship between bulk stoichiometry and local defect energetics remains poorly understood. Here we map ox…
▽ More
Two-step thermochemical water splitting driven by concentrated solar heat is a scalable route to renewable hydrogen, but it requires oxides whose oxygen-vacancy formation energies balance facile reduction with favorable reoxidation. Perovskite solid solutions can tune this balance, but the relationship between bulk stoichiometry and local defect energetics remains poorly understood. Here we map oxygen-vacancy formation energetics across Ca-Ce-Ti-Mn (CCTM) perovskites by combining first-principles calculations with a coverage-constrained special quasirandom structure approach that realizes all fifteen symmetry-distinct oxygen nearest-neighbor environments, an interpretable crystal-feature model whose fitted coefficients directly encode the underlying Born-Haber thermochemistry, and a fine-tuned defect graph neural network. Local B-site chemistry dominates the oxygen-vacancy formation energy $E_\mathrm{v}$: varying the nearest-neighbor Mn fraction shifts $E_\mathrm{v}$ by 1.0-1.5 eV depending on local Ce content, whereas A-site Ce variation contributes a smaller, Mn-dependent shift of 0.2-0.6 eV. Short-range B-site cation order, if it can be established and kinetically retained through processing, is therefore a candidate means of tuning redox performance without changing bulk composition. Composition-space maps identify a Ce/Mn-balanced region ($X_\mathrm{Ce}$ = 0.29-0.33, $X_\mathrm{Mn}$ = 0.58-0.67) combining a high fraction of vacancy sites within the targeted $E_\mathrm{v}$ window with phase stability and solubility, whose predicted redox cycle capacity matches or exceeds the ceria benchmark at 1350 $^\circ$C rather than the roughly 1600 $^\circ$C ceria requires. Measurements on three CCTM compositions show cycle capacity increasing monotonically with Ce content under protocols close to the model conditions. The design rules are expected to transfer to related perovskite families.
△ Less
Submitted 30 July, 2026;
originally announced July 2026.
-
KG2Code: Bridging Knowledge Graphs and Large Language Models via Executable Code for Question Answering
Authors:
Yike Wu,
Nan Hu,
Guilin Qi,
Guohui Xiao,
Chen Jiang,
Xinchun Zou,
Yuchen Lu,
Songlin Zhai,
Yongrui Chen,
Yuyang Zhang,
Xiaoguang Li,
Lifeng Shang,
Jiaoyan Chen,
Jeff Z. Pan
Abstract:
Recent research has explored the integration of knowledge graphs (KGs) with large language models (LLMs) to enhance their performance on downstream knowledge-intensive tasks, particularly knowledge graph question answering (KGQA). Existing approaches primarily combine LLMs with KGs through retrieval-augmented generation (RAG)-based, agent-based, and SPARQL-based methods. Although these methods hav…
▽ More
Recent research has explored the integration of knowledge graphs (KGs) with large language models (LLMs) to enhance their performance on downstream knowledge-intensive tasks, particularly knowledge graph question answering (KGQA). Existing approaches primarily combine LLMs with KGs through retrieval-augmented generation (RAG)-based, agent-based, and SPARQL-based methods. Although these methods have achieved notable success, they still suffer from several limitations, including structural information loss, unfaithful reasoning, and limited flexibility and generalization. To address these challenges, this paper proposes KG2Code, a novel approach that transforms knowledge graphs into a code-based representation, preserving structural semantics while naturally aligning with the code-aware pretraining of modern LLMs. Based on KG2Code, KG2Code-QA is further introduced as a KGQA framework that formulates KGQA as a code generation task. This formulation enables the generation of verifiable reasoning traces and executable code, thereby substantially mitigating the impact of hallucinations. In addition, an automated pipeline is developed to construct a large-scale, high-quality code corpus for effectively training open-source LLMs on KG2Code-QA. After training, LLMs are able to perform KGQA in zero-shot scenarios. Extensive experiments demonstrate that the proposed approach significantly outperforms existing KG-enhanced LLM methods for KGQA, while exhibiting strong generalization to unseen KGs. The code and data are available at Github.
△ Less
Submitted 26 June, 2026;
originally announced July 2026.
-
Binary quadratic forms and elliptic curves with analytic rank one
Authors:
Tong Wei,
Shuai Zhai
Abstract:
Given an elliptic curve with Weierstrass equation $y^2=f(x)$, and a positive definite binary quadratic form $Q(u, v)$. We show that there are infinitely many $d$ in the set represented by the quadratic forms in the genus of $Q$ such that the twisted elliptic curve $dy^2=f(x)$ has analytic rank one.
Given an elliptic curve with Weierstrass equation $y^2=f(x)$, and a positive definite binary quadratic form $Q(u, v)$. We show that there are infinitely many $d$ in the set represented by the quadratic forms in the genus of $Q$ such that the twisted elliptic curve $dy^2=f(x)$ has analytic rank one.
△ Less
Submitted 21 July, 2026;
originally announced July 2026.
-
Radio-detected Lya emitters at 1.88 < z < 3.52: AGN fraction and Lya emission
Authors:
Sai Zhai,
Huub Röttgering,
Anniek J. Gloudemans,
Erin Mentuch Cooper,
Maya H. Debski,
Gregory Zeimann,
Matt J. Jarvis,
Leah K. Morabito,
Donald P. Schneider,
Daniel J. Farrow,
Gary J. Hill,
Caryl Gronwall,
Yuming Fu
Abstract:
Lya emitters (LAEs) are galaxies with strong Lya emission, tracing early star formation and ionizing radiation. Their connection to active galactic nuclei (AGNs) is key to understanding the mechanisms behind (extended) Lya emission. In this work, we measure the fraction of LAEs identified as radio-emitting AGN (fAGN,radio) and the fraction of radio sources that exhibit Lya emission (fLya) to inves…
▽ More
Lya emitters (LAEs) are galaxies with strong Lya emission, tracing early star formation and ionizing radiation. Their connection to active galactic nuclei (AGNs) is key to understanding the mechanisms behind (extended) Lya emission. In this work, we measure the fraction of LAEs identified as radio-emitting AGN (fAGN,radio) and the fraction of radio sources that exhibit Lya emission (fLya) to investigate the connection between radio AGN activity and Lya emission at 1.88 < z < 3.52. We identify 928 sources detected in both the Hobby-Eberly Telescope Dark Energy Experiment (HETDEX) and the LOw Frequency ARray (LOFAR) surveys. These matches are drawn from 55,109 spectroscopically confirmed LAEs and 27,625 radio sources. After applying completeness corrections, we obtain fAGN,radio = 1.77 $\pm$ 0.04% and fLya = 18.15 $\pm$ 0.14%. The fraction fAGN,radio increases from 0.4 $\pm$ 0.1% to 9.7 $\pm$ 1.3% with increasing Lya luminosity, while fLya rises from 0.7 $\pm$ 0.1% to 55.8 $\pm$ 14.5% with radio luminosity. `LAEs with radio AGN' and `optical AGN with Lya emission' show similar radio luminosities above the AGN threshold, although optical AGN have higher Lya luminosities. We find no significant correlation between Lya luminosity and either radio luminosity or spectral index. Lya line width increases with Lya luminosity but shows no correlation with radio size. Our results show that most Lya emission at 1.88 < z < 3.52 is powered by star formation, with radio AGN activity confined to a small luminous subset (1.77 $\pm$ 0.04%). The absence of correlations between Lya and radio properties suggests that Lya emission is governed primarily by host-galaxy gas properties rather than direct AGN jet coupling.
△ Less
Submitted 17 July, 2026;
originally announced July 2026.
-
HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation
Authors:
Shaopeng Zhai,
Qi Zhang,
Tianyi Zhang,
Haoran Zhang,
Fuxian Huang,
Zhanhui Lin,
Zijun Xu,
Weinan Zhang
Abstract:
When adapting Vision Language Action (VLA) models to downstream tasks, multiple rounds of post-training are often required to progressively address policy weaknesses. In this report, we focus on maximizing human efficiency during this iterative process, measured by policy improvement and task throughput per unit of human labor and time.
We propose HELP, a Human-Efficient Large-scale robot Post-t…
▽ More
When adapting Vision Language Action (VLA) models to downstream tasks, multiple rounds of post-training are often required to progressively address policy weaknesses. In this report, we focus on maximizing human efficiency during this iterative process, measured by policy improvement and task throughput per unit of human labor and time.
We propose HELP, a Human-Efficient Large-scale robot Post-training pipeline in which two specialized operators supervise twelve robots concurrently. A trained Teleoperator provides high-value remote interventions and recovery demonstrations, while a Floor Operator monitors the robot fleet, triggers takeovers, and performs physical resets. This role specialization improves human efficiency by reducing task switching, lowering operator training costs, and expanding robot interaction coverage. Beyond increasing rollout volume, concurrent supervision also broadens the range of policy behaviors observed by the human team, making recurring failure modes easier to identify and enabling more targeted takeovers, resets, and recovery demonstrations. To efficiently utilize the large and mixed-quality rollout data, HELP incorporates \vlac, an automatic rollout segmentation critic specifically designed for this setting. It separates autonomous trajectories into progress-making, idle, failure-inducing, and recovery segments. Useful rollout segments are retained and combined with Human-in-the-Loop data for the next post-training round.
Across four real-world manipulation tasks, HELP achieves 80\%--95\% success rates and improves task throughput by 1.7$\times$--4.2$\times$ over the base model. Under matched HITL recovery budgets, VLAC-CUT further amplifies throughput gains by 1.20$\times$--3.43$\times$ and success-rate gains by 1.50$\times$--3.00$\times$ over HITL-only updates.
△ Less
Submitted 15 July, 2026; v1 submitted 7 July, 2026;
originally announced July 2026.
-
PrISM-IQA: Image Quality Assessment Made Practical for Smartphone Photography
Authors:
Shuyan Zhai,
Jiaqi He,
Weixia Zhang,
Liang Wang,
Zhenjie Lee,
Zufeng Zhang,
Kede Ma
Abstract:
Existing smartphone image quality assessment (IQA) methods commonly reduce perceptual quality to a single score. However, this scalar formulation is poorly aligned with practical image signal processor (ISP) tuning, where engineers must identify specific quality issues, estimate their severities, and determine whether they are acceptable or require intervention. In this work, we introduce a Practi…
▽ More
Existing smartphone image quality assessment (IQA) methods commonly reduce perceptual quality to a single score. However, this scalar formulation is poorly aligned with practical image signal processor (ISP) tuning, where engineers must identify specific quality issues, estimate their severities, and determine whether they are acceptable or require intervention. In this work, we introduce a Practical ISP-aware Structured Model for IQA (PrISM-IQA), which reformulates smartphone IQA as a multi-issue ordinal diagnosis problem. Rather than regressing a single quality score, PrISM-IQA predicts an \textit{ordered} severity level -- absent, minor, severe, or critical -- for each ISP-relevant issue, covering both global image-level artifacts and local content-dependent defects. To produce logically consistent predictions, PrISM-IQA combines cumulative ordinal encoding with structured inference that captures within-issue monotonicity as well as cross-issue subsumption and exclusion relations. We evaluate PrISM-IQA on a reconstructed SPAQ benchmark annotated with $53$ ISP-relevant quality issues and on a small-scale expert-annotated real-world dataset. Experimental results demonstrate the effectiveness of PrISM-IQA for practical issue-level diagnosis, reveal transferable perceptual quality representations through linear probing, and further show how its predictions can support actionable and meaningful ISP tuning.
△ Less
Submitted 30 June, 2026;
originally announced June 2026.
-
Mask to Concept: Auto-Promptable SAM3 via Efficient Test-Time Concept Embedding Search for Few-Shot Annotation
Authors:
Quan Zhou,
Shaoqing Zhai,
Qiang Hu,
Jia Chen,
Qiang Li,
Zhiwei Wang
Abstract:
Transforming foundation segmentation models from human-prompted tools into auto-promptable annotators is critical for scalable medical data annotation. Current methods commonly depend on external feature matchers or auxiliary networks to automate geometric prompting, but introducing architectural overhead and limiting performance scalability. Although SAM3 natively supports concept segmentation vi…
▽ More
Transforming foundation segmentation models from human-prompted tools into auto-promptable annotators is critical for scalable medical data annotation. Current methods commonly depend on external feature matchers or auxiliary networks to automate geometric prompting, but introducing architectural overhead and limiting performance scalability. Although SAM3 natively supports concept segmentation via reusable text prompts, its direct use in medical imaging is hindered by a lack of fine-grained clinical knowledge and the ambiguity of human-written descriptions. In this work, we propose Mask to Concept (M2C), an efficient framework that adapts SAM3 for medical few-shot annotation without external modules, parameter retraining, or manual text engineering. Using only a few labeled images, M2C enables SAM3 to automatically search for transferable visual concepts entirely within its frozen architecture: it initializes a learnable concept embedding, uses it to prompt segmentation, and updates the embedding by gradients of minimizing the concept segmentation error. We further introduce a Hybrid Uncertainty Estimation (HUE) module that calculates the prediction entropy and maps concept predictions back to the box prompts, measuring concept-geometry prompting inconsistency. Highly uncertain samples are flagged actively for human correction, and the corrected masks are then fed back to M2C to continuously search for more precise concept embeddings, forming a self-enhancing annotation loop with minimal expert effort. Experiments on medical segmentation benchmarks show that our method achieves SOTA few-shot segmentation performance and outstanding annotation efficiency, offering a practical and efficient pathway toward scalable medical image labeling. Codes are at https://github.com/Huster-Hq/M2C.
△ Less
Submitted 30 June, 2026; v1 submitted 25 June, 2026;
originally announced June 2026.
-
Probing Nuclear Effects with Transverse Kinematic Imbalance in Muon-neutrino Induced Charged-Current $π^0$ Production on Argon with the MicroBooNE Detector
Authors:
MicroBooNE collaboration,
P. Abratenko,
D. Andrade Aldana,
J. Asaadi,
A. Ashkenazi,
S. Balasubramanian,
B. Baller,
A. Barnard,
G. Barr,
D. Barrow,
J. Barrow,
V. Basque,
J. Bateman,
B. Behera,
O. Benevides Rodrigues,
S. Berkman,
A. Bhat,
V. Bhelande,
M. Bhattacharya,
A. Binau,
M. Bishai,
A. Blake,
B. Bogart,
T. Bolton,
M. B. Brunetti
, et al. (170 additional authors not shown)
Abstract:
Neutrino-nucleus cross-section measurements are needed to improve interaction modeling and to enable precision neutrino oscillation measurements in upcoming experiments such as the Deep Underground Neutrino Experiment (DUNE), Hyper-Kamiokande, and the Short-Baseline Neutrino program. Baryon-resonance neutrino interactions constitute a dominant contribution near the peak of the DUNE neutrino energy…
▽ More
Neutrino-nucleus cross-section measurements are needed to improve interaction modeling and to enable precision neutrino oscillation measurements in upcoming experiments such as the Deep Underground Neutrino Experiment (DUNE), Hyper-Kamiokande, and the Short-Baseline Neutrino program. Baryon-resonance neutrino interactions constitute a dominant contribution near the peak of the DUNE neutrino energy spectrum. We present the first measurement of muon neutrino charged-current resonance-like interactions on argon using transverse kinematic imbalance variables with the MicroBooNE detector. These observables are highly sensitive to the modeling of final-state interactions. This measurement probes kinematic imbalances using the reconstructed momenta of the muon, leading proton, and neutral pion. A comprehensive characterization of the $π^0$-proton final state is presented; however, none of the models considered are able to simultaneously reproduce all measured observables.
△ Less
Submitted 22 June, 2026;
originally announced June 2026.
-
MotionVLA: Vision-Language-Action Model for Humanoid Motion
Authors:
Nonghai Zhang,
Siyu Zhai,
Yanjun Li,
Zeyu Zhang,
Zhihan Yin,
Yandong Guo,
Boxin Shi,
Hao Tang
Abstract:
Generating realistic humanoid motion from scene images and text involves both low-frequency pose semantics and high-frequency physical dynamics. However, many existing methods tokenize motion with a single shared codebook, forcing heterogeneous motion signals into the same quantization space. Our frequency-domain analysis of human motion data reveals a clear mismatch between single-codebook quanti…
▽ More
Generating realistic humanoid motion from scene images and text involves both low-frequency pose semantics and high-frequency physical dynamics. However, many existing methods tokenize motion with a single shared codebook, forcing heterogeneous motion signals into the same quantization space. Our frequency-domain analysis of human motion data reveals a clear mismatch between single-codebook quantization and motion statistics: five DCT coefficients capture 93% of joint-position energy but only 37% of joint-velocity energy, which can bias quantization toward pose statistics and under-represent high-frequency velocity components. A second challenge lies in adapting a standard autoregressive model to effectively model high-frequency physical signals in motion sequences. Therefore, we propose DSFT, a dual-stream frequency tokenizer that separates motion into Base and physical streams and compresses them independently with DCT truncation and BPE. Furthermore, we present MotionVLA, a Qwen3.5-based model that arranges Base and physical tokens in a unified sequence, where Phys tokens are predicted after Base tokens. Experiments on HumanML3D and MBench show that, despite using a lightweight 2B backbone, MotionVLA reduces the Diversity gap to real data by over 50% on HumanML3D and improves Motion-Condition Consistency by 3.8% on MBench, supporting frequency-aware dual-stream decoupling as an effective formulation for autoregressive motion generation. Code: https://github.com/AIGeeksGroup/MotionVLA. Website: https://aigeeksgroup.github.io/MotionVLA.
△ Less
Submitted 13 June, 2026;
originally announced June 2026.
-
CoMo3R-SLAM: Collaborative Monocular Dense SLAM with Learned 3D Reconstruction Priors for Outdoor Multi-Agent Systems
Authors:
Zhihao Cao,
Qi Shao,
Shuhao Zhai,
Feng Tian,
Anh Nguyen,
Hesheng Wang,
Baoru Huang
Abstract:
Collaborative dense SLAM is essential for multi-robot teams to achieve scalable and consistent 3D perception across large-scale outdoor environments. Existing systems typically depend on depth sensors, incurring significant payload, power, and calibration costs. Monocular RGB cameras are a lightweight alternative, but collaborative monocular dense SLAM remains difficult due to scale ambiguity, unr…
▽ More
Collaborative dense SLAM is essential for multi-robot teams to achieve scalable and consistent 3D perception across large-scale outdoor environments. Existing systems typically depend on depth sensors, incurring significant payload, power, and calibration costs. Monocular RGB cameras are a lightweight alternative, but collaborative monocular dense SLAM remains difficult due to scale ambiguity, unreliable inter-agent data association, especially in outdoor scenes where low overlap and repetitive structures make traditional feature matching unreliable, motivating robust geometric information. We propose CoMo3R-SLAM, the first collaborative monocular dense RGB SLAM system that leverages robust learned feed-forward 3D reconstruction priors for outdoor multi-agent mapping. Each agent runs a prior-guided front-end for real-time tracking and local dense fusion, while a coordinator performs dense pointmap matching for cross-agent verification, closed-form Sim(3) gauge synchronization, and GPU-accelerated global bundle adjustment with segment-level depth optimization. Requiring neither depth sensors nor parametric intrinsics, our system produces robust cross-agent constraints and globally consistent metric maps from monocular RGB alone. On Tanks and Temples and Waymo sequences, CoMo3R-SLAM achieves the best ATE on three of four Tanks and Temples scenes and competitive Waymo accuracy, matching or exceeding state-of-the-art RGB-D methods while running online at 8 FPS.
△ Less
Submitted 28 May, 2026;
originally announced May 2026.
-
Securing LLM Agents Need Intent-to-Execution Integrity
Authors:
Wenjie Qu,
Ming Xu,
Peiran Wang,
Shengfang Zhai,
Jiaheng Zhang,
Dawn Song
Abstract:
This position paper argues that securing LLM agents requires first defining an end-to-end correctness property that specifies when an agent's execution faithfully reflects the user's intent. Modern LLM agents operate over an \emph{intent-to-execution pipeline}, where natural-language instructions are translated into concrete system operations such as tool calls, API requests, and code execution. W…
▽ More
This position paper argues that securing LLM agents requires first defining an end-to-end correctness property that specifies when an agent's execution faithfully reflects the user's intent. Modern LLM agents operate over an \emph{intent-to-execution pipeline}, where natural-language instructions are translated into concrete system operations such as tool calls, API requests, and code execution. While recent defenses have made progress in constraining how agents construct tool calls, most existing formulations implicitly assume that tools are trusted. The emergence of systems such as OpenClaw, with open ecosystems of third-party skills and direct access to user environments, breaks this assumption and exposes new failure modes, including malicious or over-privileged components in the execution pipeline.
Despite rapid progress in defense mechanisms, there is no adequate correctness property that defines what ``secure'' means for LLM agents, nor a principled way to evaluate the coverage of existing defenses. We observe that LLM agents are structurally analogous to compilers, where security violations correspond to mis-executions that do not preserve user intent. Drawing on this analogy, we identify two fundamental problem sources -- untrusted data ingestion and untrusted tool execution -- and derive four integrity properties that must hold simultaneously: \emph{Tool Integrity}, \emph{Instruction Integrity}, \emph{Judgment Integrity}, and \emph{Data Flow Integrity}. We call their conjunction \emph{intent-to-execution integrity}.
Analyzing existing agentic defenses against these properties reveals that current systems provide only partial and non-compositional coverage, leaving fundamental gaps in securing modern LLM agents.
△ Less
Submitted 16 May, 2026;
originally announced May 2026.
-
TTE-Flash: Accelerating Reasoning-based Multimodal Representations via Think-Then-Embed Tokens
Authors:
Jianpeng Cheng,
Xian Wu,
Jiangfan Zhang,
Wentao Bao,
Chaitanya Ahuja,
Shlok Kumar Mishra,
Hanchao Yu,
Yang Gao,
Fan Xia,
Qi Guo,
Shaodan Zhai,
Xiangjun Fan,
Jun Xiao
Abstract:
Recent research has demonstrated that Universal Multimodal Embedding (UME) benefits significantly from Chain-of-Thought (CoT) reasoning. In this paradigm, a generative model produces explicit reasoning traces for a multimodal query, with the final representation extracted from an <eos> embedding token attending to both the query and the reasoning. Despite its effectiveness, the computational overh…
▽ More
Recent research has demonstrated that Universal Multimodal Embedding (UME) benefits significantly from Chain-of-Thought (CoT) reasoning. In this paradigm, a generative model produces explicit reasoning traces for a multimodal query, with the final representation extracted from an <eos> embedding token attending to both the query and the reasoning. Despite its effectiveness, the computational overhead of generating explicit CoT traces is often prohibitive. In this work, we propose replacing explicit CoT with latent think tokens, which are interpreted as latent variables that can produce explicit CoT traces as observed variables. By optimizing think tokens using CoT generation loss and subsequent embedding tokens using contrastive loss, we produce high-performance, reasoning-aware representations at a constant inference cost. Our study investigates two key architectural designs: 1) how think and embeddings tokens should be extracted from the same LLM backbone. 2) how the tokens should be trained as two dependent tasks. We introduce TTE-Flash-2B, a reasoning-aware multimodal representation model that outperforms its explicit-CoT counterpart on the MMEB-v2 benchmark, while producing latent think tokens that are interpretable both textually and visually. Furthermore, zero-shot evaluation across 15 video datasets reveals scaling behavior as the number of think tokens increases, and motivating a pilot study of adaptive think budget allocation based on task requirements.
△ Less
Submitted 15 May, 2026;
originally announced May 2026.
-
MAGS-SLAM: Monocular Multi-Agent Gaussian Splatting SLAM for Geometrically and Photometrically Consistent Reconstruction
Authors:
Zhihao Cao,
Qi Shao,
Shuhao Zhai,
Jing Zhang,
Anh Nguyen,
Hesheng Wang,
Baoru Huang
Abstract:
Collaborative photorealistic 3D reconstruction from multiple agents enables rapid large-scale scene capture for virtual production and cooperative multi-robot exploration. While recent 3D Gaussian Splatting (3DGS) SLAM algorithms can generate high-fidelity real-time mapping, most of the existing multi-agent Gaussian SLAM methods still rely on RGB-D sensors to obtain metric depth and simplify cross…
▽ More
Collaborative photorealistic 3D reconstruction from multiple agents enables rapid large-scale scene capture for virtual production and cooperative multi-robot exploration. While recent 3D Gaussian Splatting (3DGS) SLAM algorithms can generate high-fidelity real-time mapping, most of the existing multi-agent Gaussian SLAM methods still rely on RGB-D sensors to obtain metric depth and simplify cross-agent alignment, limiting their deployment on low-cost or power-constrained robotic platforms, especially given the wider availability of RGB cameras. To address this challenge, we propose MAGS-SLAM, the first RGB-only multi-agent 3DGS SLAM framework for collaborative scene reconstruction. Each agent independently builds local monocular Gaussian submaps and transmits compact submap summaries rather than raw observations or dense maps. To facilitate robust collaboration in the presence of monocular scale ambiguity, our framework integrates compact submap communication, geometry- and appearance-aware loop verification, and occupancy-aware Gaussian fusion, enabling coherent global reconstruction without active depth sensors. We further introduce ReplicaMultiagent Plus, a benchmark containing larger robot teams for evaluating collaborative Gaussian SLAM. Extensive experiments on synthetic and real-world datasets show that MAGS-SLAM achieves tracking accuracy and rendering quality competitive with or superior to those of state-of-the-art RGB-D collaborative Gaussian SLAM methods using RGB images alone.
△ Less
Submitted 27 July, 2026; v1 submitted 11 May, 2026;
originally announced May 2026.
-
BadDLM: Backdooring Diffusion Language Models with Diverse Targets
Authors:
Shengfang Zhai,
Xiaoyang Ji,
Yuling Shi,
Haoran Gao,
Fanyu Meng,
Yan Zeng,
Yuejian Fang,
Yinpeng Dong,
Jiaheng Zhang
Abstract:
Diffusion language models (DLMs) have recently emerged as an alternative modeling paradigm to autoregressive (AR) language models, enabling parallel generation and bidirectional context modeling. Yet their security implications, particularly their vulnerability to backdoor attacks, remain underexplored. We propose BadDLM, a unified framework for studying backdoor attacks against DLMs with diverse…
▽ More
Diffusion language models (DLMs) have recently emerged as an alternative modeling paradigm to autoregressive (AR) language models, enabling parallel generation and bidirectional context modeling. Yet their security implications, particularly their vulnerability to backdoor attacks, remain underexplored. We propose BadDLM, a unified framework for studying backdoor attacks against DLMs with diverse targets. We introduce a trigger-aware training objective that emphasizes target-relevant positions in poisoned samples, and theoretically prove that this objective is equivalent to training under an induced forward masking distribution. Unlike backdoors in autoregressive models, which typically manipulate next-token prediction, this characterization indicates that BadDLM can implant backdoors by exploiting the forward masking process. We instantiate BadDLM across different target levels: concept injection (BadDLM_Concept), semantic attribute steering (BadDLM_Attribute), alignment bypass (BadDLM_Align), and code payload injection (BadDLM_Payload). Experiments on mainstream open-source DLMs show that BadDLM achieves strong attack effectiveness across diverse targets while largely preserving benign utility, and remains effective against defenses designed for AR backdoors. Our findings expose a new class of security risks in diffusion-based language generation and call for defenses tailored to DLM denoising dynamics.
△ Less
Submitted 10 May, 2026;
originally announced May 2026.
-
On the coefficients of the Taylor expansion of $L$-functions of elliptic curves
Authors:
Tong Wei,
Shuai Zhai
Abstract:
In this paper, we investigate the coefficients of the Taylor expansion of the complex $L$-series of any elliptic curve over $\mathbb{Q}$. We prove that, in the family of quadratic twists by all the discriminants $d$, these coefficients are nonvanishing under GRH when $d$ is sufficiently large. Unconditionally, we obtain a general lower bound for the number of nonvanishing coefficients in the famil…
▽ More
In this paper, we investigate the coefficients of the Taylor expansion of the complex $L$-series of any elliptic curve over $\mathbb{Q}$. We prove that, in the family of quadratic twists by all the discriminants $d$, these coefficients are nonvanishing under GRH when $d$ is sufficiently large. Unconditionally, we obtain a general lower bound for the number of nonvanishing coefficients in the family of quadratic twists, through a series of results from the moments of the central values of the derivatives of quadratic twists of modular $L$-function.
△ Less
Submitted 9 May, 2026;
originally announced May 2026.
-
Preserving Foundational Capabilities in Flow-Matching VLAs through Conservative SFT
Authors:
Tianyi Zhang,
Shaopeng Zhai,
Haoran Zhang,
Fuxian Huang,
Qi Zhang
Abstract:
Unconstrained fine-tuning of flow-matching Vision-Language-Action (VLA) models drives dense parameter overwrites, degrading pre-trained capabilities. We present Conservative Supervised Fine-Tuning (ConSFT), an optimization objective that adapts to target distributions while mitigating catastrophic forgetting, requiring zero prior data or architectural overhead. By dynamically scaling learning sign…
▽ More
Unconstrained fine-tuning of flow-matching Vision-Language-Action (VLA) models drives dense parameter overwrites, degrading pre-trained capabilities. We present Conservative Supervised Fine-Tuning (ConSFT), an optimization objective that adapts to target distributions while mitigating catastrophic forgetting, requiring zero prior data or architectural overhead. By dynamically scaling learning signals based on model confidence, ConSFT suppresses excessive gradients from low-confidence samples to prevent disproportionate parameter updates, thereby bounding the intrinsic parameter disruption risk. Inspired by reinforcement learning's trust-region clipping, this formulation establishes a progressive learning dynamic to secure target convergence and prior capability retention, maintaining sparse parameter updates without relying on the parallel reference networks required by explicit regularization. We evaluate ConSFT on the LIBERO and RoboTwin benchmarks across state-of-the-art flow-matching VLAs ($π_0$, $π_{0.5}$, and GR00T-N1.6-3B). The method outperforms vanilla SFT in capability retention by an average absolute margin of over 20\%, matching the efficacy of data-heavy Experience Replay in a prior-data-free regime. Real-world robotic deployments confirm that ConSFT precludes spatial overfitting during downstream adaptation, preserving pre-trained physical skills while acquiring sequential target tasks.
△ Less
Submitted 19 May, 2026; v1 submitted 9 May, 2026;
originally announced May 2026.
-
Normalizing Trajectory Models
Authors:
Jiatao Gu,
Tianrong Chen,
Ying Shen,
David Berthelot,
Shuangfei Zhai,
Josh Susskind
Abstract:
Diffusion-based models decompose sampling into many small Gaussian denoising steps -- an assumption that breaks down when generation is compressed to a few coarse transitions. Existing few-step methods address this through distillation, consistency training, or adversarial objectives, but sacrifice the likelihood framework in the process. We introduce Normalizing Trajectory Models (NTM), which mod…
▽ More
Diffusion-based models decompose sampling into many small Gaussian denoising steps -- an assumption that breaks down when generation is compressed to a few coarse transitions. Existing few-step methods address this through distillation, consistency training, or adversarial objectives, but sacrifice the likelihood framework in the process. We introduce Normalizing Trajectory Models (NTM), which models each reverse step as an expressive conditional normalizing flow with exact likelihood training. Architecturally, NTM combines shallow invertible blocks within each step with a deep parallel predictor across the trajectory, forming an end-to-end network trainable from scratch or initializable from pretrained flow-matching models. Its exact trajectory likelihood further enables self-distillation: a lightweight denoiser trained on the model's own score produces high-quality samples in four steps. On text-to-image benchmarks, NTM matches or outperforms strong image generation baselines in just four sampling steps while uniquely retaining exact likelihood over the generative trajectory.
△ Less
Submitted 12 May, 2026; v1 submitted 8 May, 2026;
originally announced May 2026.
-
STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation
Authors:
Ying Shen,
Tianrong Chen,
Yuan Gao,
Yizhe Zhang,
Yuyang Wang,
Miguel Ángel Bautista,
Shuangfei Zhai,
Joshua M. Susskind,
Jiatao Gu
Abstract:
Deep generative models have advanced rapidly across text and vision, motivating unified multimodal systems that can understand, reason over, and generate interleaved text-image sequences. Most existing approaches combine autoregressive language modeling with diffusion-based image generators, inheriting a structural mismatch between causal text generation and iterative visual denoising. We observe…
▽ More
Deep generative models have advanced rapidly across text and vision, motivating unified multimodal systems that can understand, reason over, and generate interleaved text-image sequences. Most existing approaches combine autoregressive language modeling with diffusion-based image generators, inheriting a structural mismatch between causal text generation and iterative visual denoising. We observe that autoregressive normalizing flows are autoregressive Transformers--sharing the same causal mask, KV-cache mechanism, and left-to-right structure as LLMs--making them the most natural paradigm for true unified multimodal generation. We present STARFlow2, built on the Pretzel architecture that vertically interleaves a pretrained VLM stream with a TarFlow stream via residual skip connections, both operating under the same causal mask. Combined with a deep-shallow flow design and a unified FAE latent space, STARFlow2 enables cache-friendly interleaved generation where both text and visual outputs directly enter the KV-cache without re-encoding. Experiments demonstrate strong performance across image generation and multimodal understanding benchmarks, validating autoregressive flows as a viable foundation for unified multimodal modeling.
△ Less
Submitted 8 May, 2026;
originally announced May 2026.
-
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
Authors:
Fanqing Meng,
Lingxiao Du,
Zijian Wu,
Guanzheng Chen,
Xiangyan Liu,
Jiaqi Liao,
Chonghe Jiang,
Zhenglin Wan,
Jiawei Gu,
Pengfei Zhou,
Rui Huang,
Ziqi Zhao,
Shengyuan Ding,
Ailing Yu,
Bo Peng,
Bowei Xia,
Hao Sun,
Haotian Liang,
Ji Xie,
Jiajun Chen,
Jiajun Song,
Liu Yang,
Ming Xu,
Qionglin Qiu,
Runhao Fu
, et al. (24 additional authors not shown)
Abstract:
Language-model agents are increasingly used as persistent coworkers that assist users across multiple working days. During such workflows, the surrounding environment may change independently of the agent: new emails arrive, calendar entries shift, knowledge-base records are updated, and evidence appears across images, scanned PDFs, audio, video, and spreadsheets. Existing benchmarks do not adequa…
▽ More
Language-model agents are increasingly used as persistent coworkers that assist users across multiple working days. During such workflows, the surrounding environment may change independently of the agent: new emails arrive, calendar entries shift, knowledge-base records are updated, and evidence appears across images, scanned PDFs, audio, video, and spreadsheets. Existing benchmarks do not adequately evaluate this setting because they typically run within a single static episode and remain largely text-centric. We introduce \bench{}, a benchmark for coworker agents built around multi-turn multi-day tasks, a stateful sandboxed service environment whose state evolves between turns, and rule-based verification. The current release contains 100 tasks across 13 professional scenarios, executed against five stateful sandboxed services (filesystem, email, calendar, knowledge base, spreadsheet) and scored by 1537 deterministic Python checkers over post-execution service state; no LLM-as-judge is invoked during scoring. We benchmark seven frontier agent systems. The strongest model reaches 75.8 weighted score, but the best strict Task Success is only 20.0\%, indicating that partial progress is common while complete end-to-end workflow completion remains rare. Turn-level analysis shows that performance drops after the first exogenous environment update, highlighting adaptation to changing state as a key open challenge. We release the benchmark, evaluation harness, and construction pipeline to support reproducible coworker-agent evaluation.
△ Less
Submitted 5 May, 2026; v1 submitted 26 April, 2026;
originally announced April 2026.
-
Normalizing Flows with Iterative Denoising
Authors:
Tianrong Chen,
Jiatao Gu,
David Berthelot,
Joshua Susskind,
Shuangfei Zhai
Abstract:
Normalizing Flows (NFs) are a classical family of likelihood-based methods that have received revived attention. Recent efforts such as TARFlow have shown that
NFs are capable of achieving promising performance on image modeling tasks, making them viable alternatives to other methods such as diffusion models.
In this work, we further advance the state of Normalizing Flow generative models by i…
▽ More
Normalizing Flows (NFs) are a classical family of likelihood-based methods that have received revived attention. Recent efforts such as TARFlow have shown that
NFs are capable of achieving promising performance on image modeling tasks, making them viable alternatives to other methods such as diffusion models.
In this work, we further advance the state of Normalizing Flow generative models by introducing iterative TARFlow (iTARFlow). Unlike diffusion models, iTARFlow maintains a fully end-to-end, likelihood-based objective during training. During sampling, it performs autoregressive generation followed by an iterative denoising procedure inspired by diffusion-style methods. Through extensive experiments, we show that iTARFlow achieves competitive performance across ImageNet resolutions of 64, 128, and 256 pixels, demonstrating its potential as a strong generative model and advancing the frontier of Normalizing Flows. In addition, we analyze the characteristic artifacts produced by iTARFlow, offering insights that may shed light on future improvements. Code is available at https://github.com/apple/ml-itarflow.
△ Less
Submitted 21 April, 2026;
originally announced April 2026.
-
Abstain-R1: Calibrated Abstention and Post-Refusal Clarification via Verifiable RL
Authors:
Skylar Zhai,
Jingcheng Liang,
Dongyeop Kang
Abstract:
Reinforcement fine-tuning improves the reasoning ability of large language models, but it can also encourage them to answer unanswerable queries by guessing or hallucinating missing information. Existing abstention methods either train models to produce generic refusals or encourage follow-up clarifications without verifying whether those clarifications identify the key missing information. We stu…
▽ More
Reinforcement fine-tuning improves the reasoning ability of large language models, but it can also encourage them to answer unanswerable queries by guessing or hallucinating missing information. Existing abstention methods either train models to produce generic refusals or encourage follow-up clarifications without verifying whether those clarifications identify the key missing information. We study queries that are clear in meaning but cannot be reliably resolved from the given information, and argue that a reliable model should not only abstain, but also explain what is missing. We propose a clarification-aware RLVR reward that, while rewarding correct answers on answerable queries, jointly optimizes explicit abstention and semantically aligned post-refusal clarification on unanswerable queries. Using this reward, we train Abstain-R1, a 3B model that improves abstention and clarification on unanswerable queries while preserving strong performance on answerable ones. Experiments on Abstain-Test, Abstain-QA, and SelfAware show that Abstain-R1 substantially improves over its base model and achieves unanswerable-query behavior competitive with larger systems including DeepSeek-R1, suggesting that calibrated abstention and clarification can be learned through verifiable rewards rather than emerging from scale alone.
△ Less
Submitted 18 April, 2026;
originally announced April 2026.
-
Horseshoe Predictive Inference
Authors:
Percy S. Zhai,
Veronika Ročková
Abstract:
Predictive inference in the sparse Gaussian sequence model has received considerably less attention than its non-sparse, finite-sample counterpart. Existing work has largely been confined to discrete mixture priors. In this paper, we study predictive inference under a widely used continuous mixture prior, the Horseshoe. We provide new theoretical results establishing exact asymptotic minimax optim…
▽ More
Predictive inference in the sparse Gaussian sequence model has received considerably less attention than its non-sparse, finite-sample counterpart. Existing work has largely been confined to discrete mixture priors. In this paper, we study predictive inference under a widely used continuous mixture prior, the Horseshoe. We provide new theoretical results establishing exact asymptotic minimax optimality of the predictive Bayes estimator when the sparsity level is known. Furthermore, through a Gaussian-mixture representation of the posterior predictive density (which we term Horseshoe spectroscopy), the phase-transition in the local shrinkage scale is inherited by the predictive mechanism, producing behavior similar to that of previous thresholding/switching estimators. When sparsity is unknown, we adopt a fully Bayesian approach using a hierarchical Horseshoe prior and show that it performs adaptive, as opposed to manual, switching. Under a theta-min condition, the resulting predictive risk admits an upper bound over a restricted parameter class that is sharper than the minimax rate over the full class. We demonstrate the practical value of predictive Horseshoe shrinkage on data such as images and time series that can be naturally modeled as sparse Gaussian sequences. We illustrate this approach on facial recognition across varying facial expressions and study region-wise atypical brain lateralization in autism spectrum disorder.
△ Less
Submitted 17 April, 2026;
originally announced April 2026.
-
Simultaneous Inference for Covariance and Precision Matrices of Long-Range Dependent Time Series
Authors:
Percy S. Zhai,
Mladen Kolar,
Wei Biao Wu
Abstract:
For time series with long-range temporal dependence, inference for covariance and precision matrices is non-trivial. We propose a Berry-Esseen type Gaussian approximation result that gives a finite-sample bound for the Kolmogorov distance between the infinity norms of the estimation error of sample covariance matrix and the corresponding Gaussian approximation. The method utilizes martingale and m…
▽ More
For time series with long-range temporal dependence, inference for covariance and precision matrices is non-trivial. We propose a Berry-Esseen type Gaussian approximation result that gives a finite-sample bound for the Kolmogorov distance between the infinity norms of the estimation error of sample covariance matrix and the corresponding Gaussian approximation. The method utilizes martingale and m-dependent approximation and relies on constructing triadic blocks. We also establish a bootstrapping result with block sampling method, which preserves validity despite strong temporal dependence. Our results on covariance allow ultra-high-dimensional settings where the dimension of time series can grow sub-exponentially with sample size. Similar results can be built for precision matrix under low-dimensional settings. No assumption is required on the structure of covariance and precision matrices.
△ Less
Submitted 17 April, 2026;
originally announced April 2026.
-
Seedance 2.0: Advancing Video Generation for World Complexity
Authors:
Team Seedance,
De Chen,
Liyang Chen,
Xin Chen,
Ying Chen,
Zhuo Chen,
Zhuowei Chen,
Feng Cheng,
Tianheng Cheng,
Yufeng Cheng,
Mojie Chi,
Xuyan Chi,
Jian Cong,
Qinpeng Cui,
Fei Ding,
Qide Dong,
Yujiao Du,
Haojie Duanmu,
Junliang Fan,
Jiarui Fang,
Jing Fang,
Zetao Fang,
Chengjian Feng,
Yu Gao,
Diandian Gu
, et al. (146 additional authors not shown)
Abstract:
Seedance 2.0 is a new native multi-modal audio-video generation model, officially released in China in early February 2026. Compared with its predecessors, Seedance 1.0 and 1.5 Pro, Seedance 2.0 adopts a unified, highly efficient, and large-scale architecture for multi-modal audio-video joint generation. This allows it to support four input modalities: text, image, audio, and video, by integrating…
▽ More
Seedance 2.0 is a new native multi-modal audio-video generation model, officially released in China in early February 2026. Compared with its predecessors, Seedance 1.0 and 1.5 Pro, Seedance 2.0 adopts a unified, highly efficient, and large-scale architecture for multi-modal audio-video joint generation. This allows it to support four input modalities: text, image, audio, and video, by integrating one of the most comprehensive suites of multi-modal content reference and editing capabilities available in the industry to date. It delivers substantial, well-rounded improvements across all key sub-dimensions of video and audio generation. In both expert evaluations and public user tests, the model has demonstrated performance on par with the leading levels in the field. Seedance 2.0 supports direct generation of audio-video content with durations ranging from 4 to 15 seconds, with native output resolutions of 480p and 720p. For multi-modal inputs as reference, its current open platform supports up to 3 video clips, 9 images, and 3 audio clips. In addition, we provide Seedance 2.0 Fast version, an accelerated variant of Seedance 2.0 designed to boost generation speed for low-latency scenarios. Seedance 2.0 has delivered significant improvements to its foundational generation capabilities and multi-modal generation performance, bringing an enhanced creative experience for end users.
△ Less
Submitted 15 April, 2026;
originally announced April 2026.
-
LoViF 2026 Challenge on Human-oriented Semantic Image Quality Assessment: Methods and Results
Authors:
Xin Li,
Daoli Xu,
Wei Luo,
Guoqiang Xiang,
Haoran Li,
Chengyu Zhuang,
Zhibo Chen,
Jian Guan,
Weiping Li,
Weixia Zhang,
Wei Sun,
Zhihua Wang,
Dandan Zhu,
Chengguang Zhu,
Ayush Gupta,
Rachit Agarwal,
Shouvik Das,
Biplab Ch Das,
Amartya Ghosh,
Kanglong Fan,
Wen Wen,
Shuyan Zhai,
Tianwu Zhi,
Aoxiang Zhang,
Jianzhao Liu
, et al. (5 additional authors not shown)
Abstract:
This paper reviews the LoViF 2026 Challenge on Human-oriented Semantic Image Quality Assessment. This challenge aims to raise a new direction, i.e., how to evaluate the loss of semantic information from the human perspective, intending to promote the development of some new directions, like semantic coding, processing, and semantic-oriented optimization, etc. Unlike existing datasets of quality as…
▽ More
This paper reviews the LoViF 2026 Challenge on Human-oriented Semantic Image Quality Assessment. This challenge aims to raise a new direction, i.e., how to evaluate the loss of semantic information from the human perspective, intending to promote the development of some new directions, like semantic coding, processing, and semantic-oriented optimization, etc. Unlike existing datasets of quality assessment, we form a dataset of human-oriented semantic quality assessment, termed the SeIQA dataset. This dataset is divided into three parts for this competition: (i) training data: 510 pairs of degraded images and their corresponding ground truth references; (ii) validation data: 80 pairs of degraded images and their corresponding ground-truth references; (iii) testing data: 160 pairs of degraded images and their corresponding ground-truth references. The primary objective of this challenge is to establish a new and powerful benchmark for human-oriented semantic image quality assessment. There are a total of 58 teams registered in this competition, and 6 teams submitted valid solutions and fact sheets for the final testing phase. These submissions achieved state-of-the-art (SOTA) performance on the SeIQA dataset.
△ Less
Submitted 3 August, 2026; v1 submitted 13 April, 2026;
originally announced April 2026.
-
Ψ-Map: Panoptic Surface Integrated Mapping Enables Real2Sim Transfer
Authors:
Xuan Yu,
Yuxuan Xie,
Changjian Jiang,
Shichao Zhai,
Rong Xiong,
Yu Zhang,
Yue Wang
Abstract:
Open-vocabulary panoptic reconstruction is essential for advanced robotics perception and simulation. However, existing methods based on 3D Gaussian Splatting (3DGS) often struggle to simultaneously achieve geometric accuracy, coherent panoptic understanding, and real-time inference frequency in large-scale scenes. In this paper, we propose a comprehensive framework that integrates geometric reinf…
▽ More
Open-vocabulary panoptic reconstruction is essential for advanced robotics perception and simulation. However, existing methods based on 3D Gaussian Splatting (3DGS) often struggle to simultaneously achieve geometric accuracy, coherent panoptic understanding, and real-time inference frequency in large-scale scenes. In this paper, we propose a comprehensive framework that integrates geometric reinforcement, end-to-end panoptic learning, and efficient rendering. First, to ensure physical realism in large-scale environments, we leverage LiDAR data to construct plane-constrained multimodal Gaussian Mixture Models (GMMs) and employ 2D Gaussian surfels as the map representation, enabling high-precision surface alignment and continuous geometric supervision. Building upon this, to overcome the error accumulation and cumbersome cross-frame association inherent in traditional multi-stage panoptic segmentation pipelines, we design a query-guided end-to-end learning architecture. By utilizing a local cross-attention mechanism within the view frustum, the system lifts 2D mask features directly into 3D space, achieving globally consistent panoptic understanding. Finally, addressing the computational bottlenecks caused by high-dimensional semantic features, we introduce Precise Tile Intersection and a Top-K Hard Selection strategy to optimize the rendering pipeline. Experimental results demonstrate that our system achieves superior geometric and panoptic reconstruction quality in large-scale scenes while maintaining an inference rate exceeding 40 FPS, meeting the real-time requirements of robotic control loops.
△ Less
Submitted 13 April, 2026;
originally announced April 2026.
-
Fast-SegSim: Real-Time Open-Vocabulary Segmentation for Robotics in Simulation
Authors:
Xuan Yu,
Yuxuan Xie,
Shichao Zhai,
Shuhao Ye,
Rong Xiong,
Yue Wang
Abstract:
Open-vocabulary panoptic reconstruction is crucial for advanced robotics and simulation. However, existing 3D reconstruction methods, such as NeRF or Gaussian Splatting variants, often struggle to achieve the real-time inference frequency required by robotic control loops. Existing methods incur prohibitive latency when processing the high-dimensional features required for robust open-vocabulary s…
▽ More
Open-vocabulary panoptic reconstruction is crucial for advanced robotics and simulation. However, existing 3D reconstruction methods, such as NeRF or Gaussian Splatting variants, often struggle to achieve the real-time inference frequency required by robotic control loops. Existing methods incur prohibitive latency when processing the high-dimensional features required for robust open-vocabulary segmentation. We propose Fast-SegSim, a novel, simple, and end-to-end framework built upon 2D Gaussian Splatting, designed to realize real-time, high-fidelity, and 3D-consistent open-vocabulary segmentation reconstruction. Our core contribution is a highly optimized rendering pipeline that specifically addresses the computational bottleneck of high-channel segmentation feature accumulation. We introduce two key optimizations: Precise Tile Intersection to reduce rasterization redundancy, and a novel Top-K Hard Selection strategy. This strategy leverages the geometric sparsity inherent in the 2D Gaussian representation to greatly simplify feature accumulation and alleviate bandwidth limitations, achieving render rates exceeding 40 FPS. Fast-SegSim provides critical value in robotic applications: it serves both as a high-frequency sensor input for simulation platforms like Gazebo, and its 3D-consistent outputs provide essential multi-view 'ground truth' labels for fine-tuning downstream perception tasks. We demonstrate this utility by using the generated labels to fine-tune the perception module in object goal navigation, successfully doubling the navigation success rate. Our superior rendering speed and practical utility underscore Fast-SegSim's potential to bridge the sim-to-real gap.
△ Less
Submitted 12 April, 2026;
originally announced April 2026.
-
The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents
Authors:
Xuwei Ding,
Skylar Zhai,
Linxin Song,
Jiate Li,
Taiwei Shi,
Nicholas Meade,
Siva Reddy,
Jian Kang,
Jieyu Zhao
Abstract:
Computer-use agents (CUAs) can now autonomously complete complex tasks in real digital environments, but when misled, they can also be used to automate harmful actions programmatically. Existing safety evaluations largely target explicit threats such as misuse and prompt injection, but overlook a subtle yet critical setting where user instructions are entirely benign and harm arises from the task…
▽ More
Computer-use agents (CUAs) can now autonomously complete complex tasks in real digital environments, but when misled, they can also be used to automate harmful actions programmatically. Existing safety evaluations largely target explicit threats such as misuse and prompt injection, but overlook a subtle yet critical setting where user instructions are entirely benign and harm arises from the task context or execution outcome. We introduce OS-BLIND, a benchmark that evaluates CUAs under unintended attack conditions, comprising 300 human-crafted tasks across 12 categories, 8 applications, and 2 threat clusters: environment-embedded threats and agent-initiated harms. Our evaluation on frontier models and agentic frameworks reveals that most CUAs exceed 90% attack success rate (ASR), and even the safety-aligned Claude 4.5 Sonnet reaches 73.0% ASR. More interestingly, this vulnerability becomes even more severe, with ASR rising from 73.0% to 92.7% when Claude 4.5 Sonnet is deployed in multi-agent systems. Our analysis further shows that existing safety defenses provide limited protection when user instructions are benign. Safety alignment primarily activates within the first few steps and rarely re-engages during subsequent execution. In multi-agent systems, decomposed subtasks obscure the harmful intent from the model, causing safety-aligned models to fail. We will release our OS-BLIND to encourage the broader research community to further investigate and address these safety challenges.
△ Less
Submitted 17 April, 2026; v1 submitted 12 April, 2026;
originally announced April 2026.
-
Purify Once, Edit Freely: Breaking Image Protections under Model Mismatch
Authors:
Qichen Zhao,
Shengfang Zhai,
Xinjian Bai,
Qingni Shen,
Qiqi Lin,
Yansong Gao,
Zhonghai Wu
Abstract:
Diffusion models enable high-fidelity image editing but can also be misused for unauthorized style imitation and harmful content generation. To mitigate these risks, proactive image protection methods embed small, often imperceptible adversarial perturbations into images before sharing to disrupt downstream editing or fine-tuning. However, in realistic post-release scenarios, content owners cannot…
▽ More
Diffusion models enable high-fidelity image editing but can also be misused for unauthorized style imitation and harmful content generation. To mitigate these risks, proactive image protection methods embed small, often imperceptible adversarial perturbations into images before sharing to disrupt downstream editing or fine-tuning. However, in realistic post-release scenarios, content owners cannot control downstream processing pipelines, and protections optimized for a surrogate model may fail when attackers use mismatched diffusion pipelines. Existing purification methods can weaken protections but often sacrifice image quality and rarely examine architectural mismatch. We introduce a unified post-release purification framework to evaluate protection survivability under model mismatch. We propose two practical purifiers: VAE-Trans, which corrects protected images via latent-space projection, and EditorClean, which performs instruction-guided reconstruction with a Diffusion Transformer to exploit architectural heterogeneity. Both operate without access to protected images or defense internals. Across 2,100 editing tasks and six representative protection methods, EditorClean consistently restores editability. Compared to protected inputs, it improves PSNR by 3-6 dB and reduces FID by 50-70 percent on downstream edits, while outperforming prior purification baselines by about 2 dB PSNR and 30 percent lower FID. Our results reveal a purify-once, edit-freely failure mode: once purification succeeds, the protective signal is largely removed, enabling unrestricted editing. This highlights the need to evaluate protections under model mismatch and design defenses robust to heterogeneous attackers.
△ Less
Submitted 13 March, 2026;
originally announced March 2026.
-
Exclusive Self Attention
Authors:
Shuangfei Zhai
Abstract:
We introduce exclusive self attention (XSA), a simple modification of self attention (SA) that improves Transformer's sequence modeling performance. The key idea is to constrain attention to capture only information orthogonal to the token's own value vector (thus excluding information of self position), encouraging better context modeling. Evaluated on the standard language modeling task, XSA con…
▽ More
We introduce exclusive self attention (XSA), a simple modification of self attention (SA) that improves Transformer's sequence modeling performance. The key idea is to constrain attention to capture only information orthogonal to the token's own value vector (thus excluding information of self position), encouraging better context modeling. Evaluated on the standard language modeling task, XSA consistently outperforms SA across model sizes up to 2.7B parameters and shows increasingly larger gains as sequence length grows.
△ Less
Submitted 9 March, 2026;
originally announced March 2026.
-
The Coupling Within: Flow Matching via Distilled Normalizing Flows
Authors:
David Berthelot,
Tianrong Chen,
Jiatao Gu,
Marco Cuturi,
Laurent Dinh,
Bhavik Chandna,
Michal Klein,
Josh Susskind,
Shuangfei Zhai
Abstract:
Flow models have rapidly become the go-to method for training and deploying large-scale generators, owing their success to inference-time flexibility via adjustable integration steps. A crucial ingredient in flow training is the choice of coupling measure for sampling noise/data pairs that define the flow matching (FM) regression loss. While FM training defaults usually to independent coupling, re…
▽ More
Flow models have rapidly become the go-to method for training and deploying large-scale generators, owing their success to inference-time flexibility via adjustable integration steps. A crucial ingredient in flow training is the choice of coupling measure for sampling noise/data pairs that define the flow matching (FM) regression loss. While FM training defaults usually to independent coupling, recent works show that adaptive couplings informed by noise/data distributions (e.g., via optimal transport, OT) improve both model training and inference. We radicalize this insight by shifting the paradigm: rather than computing adaptive couplings directly, we use distilled couplings from a different, pretrained model capable of placing noise and data spaces in bijection -- a property intrinsic to normalizing flows (NF) through their maximum likelihood and invertibility requirements. Leveraging recent advances in NF image generation via auto-regressive (AR) blocks, we propose Normalized Flow Matching (NFM), a new method that distills the quasi-deterministic coupling of pretrained NF models to train student flow models. These students achieve the best of both worlds: significantly outperforming flow models trained with independent or even OT couplings, while also improving on the teacher AR-NF model.
△ Less
Submitted 9 March, 2026;
originally announced March 2026.
-
Lyα Nebulae in HETDEX: The Largest Statistical Census Bridging Lyα Halos and Blobs across Cosmic Noon
Authors:
Erin Mentuch Cooper,
Karl Gebhardt,
Dustin Davis,
Robin Ciardullo,
Chris Byrohl,
Chenxu Liu,
Maya H. Debski,
Óscar A. Chávez Ortiz,
Maximilian Fabricius,
Daniel J. Farrow,
Steven L. Finkelstein,
Caryl Gronwall,
Gary J. Hill,
Maja Lujan Niemeyer,
Brianna McKay,
Shiro Mukae,
Masami Ouchi,
Huub Röttgering,
Donald P. Schneider,
Sarah Tuttle,
Lutz Wisotzki,
Gregory Zeimann,
Sai Zhai
Abstract:
The Hobby-Eberly Dark Energy Experiment (HETDEX) is an untargeted ~540 deg^2 spectroscopic survey of Lyα emission in the 1.9 < z < 3.5 Universe. In surface brightness, this survey reaches 1σ Lyα sensitivities of approximately 2-5 x 10^-18 erg s^-1 cm^-2 arcsec^-2, allowing large samples of extended Lyα nebulae (LAN) to be studied. We selected a sample of 70,691 Lyα-emitting galaxies (LAEs) with an…
▽ More
The Hobby-Eberly Dark Energy Experiment (HETDEX) is an untargeted ~540 deg^2 spectroscopic survey of Lyα emission in the 1.9 < z < 3.5 Universe. In surface brightness, this survey reaches 1σ Lyα sensitivities of approximately 2-5 x 10^-18 erg s^-1 cm^-2 arcsec^-2, allowing large samples of extended Lyα nebulae (LAN) to be studied. We selected a sample of 70,691 Lyα-emitting galaxies (LAEs) with an emission-line signal-to-noise ratio greater than 6 and modeled the Lyα emission as a point-source component with an optional exponential envelope. Half (~47.5%) of the LAE sample (33,612 objects) exhibits significant extended emission and is best fit by the two-component model. The fraction of resolved sources increases with Lyα flux and luminosity. Their isophotal areas range from 10-130 arcsec^2 (median 15 arcsec^2), with integrated Lyα fluxes from 6-2000 x 10^-17 erg s^-1 cm^-2 (median 20 x 10^-17 erg s^-1 cm^-2). Comparison between point-spread-function-weighted and isophotal flux measurements shows that the HETDEX pipeline underestimates the total Lyα flux by ~30% on average, reflecting the substantial halo contribution in extended sources. Approximately 420 LANs are found per deg^2 over 79.5 deg^2 of non-contiguous sky. About 12% of resolved sources show active galactic nuclei signatures and are bright in Lyα and continuum. The remaining 88% span a wide range of morphologies and often lack continuum counterparts. Exponential scale lengths show no strong correlation with Lyα flux or luminosity (median 11.6 +/- 1.9 kpc). Only 2.9% of the full S/N > 6 LAE population with ancillary data have radio counterparts, but 64% of those are found to be extended, with the radio fraction increasing with Lyα size.
△ Less
Submitted 6 March, 2026;
originally announced March 2026.
-
IMMACULATE: A Practical LLM Auditing Framework via Verifiable Computation
Authors:
Yanpei Guo,
Wenjie Qu,
Linyu Wu,
Shengfang Zhai,
Lionel Z. Wang,
Ming Xu,
Yue Liu,
Binhang Yuan,
Dawn Song,
Jiaheng Zhang
Abstract:
Commercial large language models are typically deployed as black-box API services, requiring users to trust providers to execute inference correctly and report token usage honestly. We present IMMACULATE, a practical auditing framework that detects economically motivated deviations-such as model substitution, quantization abuse, and token overbilling-without trusted hardware or access to model int…
▽ More
Commercial large language models are typically deployed as black-box API services, requiring users to trust providers to execute inference correctly and report token usage honestly. We present IMMACULATE, a practical auditing framework that detects economically motivated deviations-such as model substitution, quantization abuse, and token overbilling-without trusted hardware or access to model internals. IMMACULATE selectively audits a small fraction of requests using verifiable computation, achieving strong detection guarantees while amortizing cryptographic overhead. Experiments on dense and MoE models show that IMMACULATE reliably distinguishes benign and malicious executions with under 1% throughput overhead. Our code is published at https://github.com/guo-yanpei/Immaculate.
△ Less
Submitted 26 February, 2026;
originally announced February 2026.
-
MemPot: Defending Against Memory Extraction Attack with Optimized Honeypots
Authors:
Yuhao Wang,
Shengfang Zhai,
Guanghao Jin,
Yinpeng Dong,
Linyi Yang,
Jiaheng Zhang
Abstract:
Large Language Model (LLM)-based agents employ external and internal memory systems to handle complex, goal-oriented tasks, yet this exposes them to severe extraction attacks, and effective defenses remain lacking. In this paper, we propose MemPot, the first theoretically verified defense framework against memory extraction attacks by injecting optimized honeypots into the memory. Through a two-st…
▽ More
Large Language Model (LLM)-based agents employ external and internal memory systems to handle complex, goal-oriented tasks, yet this exposes them to severe extraction attacks, and effective defenses remain lacking. In this paper, we propose MemPot, the first theoretically verified defense framework against memory extraction attacks by injecting optimized honeypots into the memory. Through a two-stage optimization process, MemPot generates trap documents that maximize the retrieval probability for attackers while remaining inconspicuous to benign users. We model the detection process as Wald's Sequential Probability Ratio Test (SPRT) and theoretically prove that MemPot achieves a lower average number of sampling rounds compared to optimal static detectors. Empirically, MemPot significantly outperforms state-of-the-art baselines, achieving a 50% improvement in detection AUROC and an 80% increase in True Positive Rate under low False Positive Rate constraints. Furthermore, our experiments confirm that MemPot incurs zero additional online inference latency and preserves the agent's utility on standard tasks, verifying its superiority in safety, harmlessness, and efficiency.
△ Less
Submitted 7 February, 2026;
originally announced February 2026.
-
On the gauge invariance of the Kuperberg invariant of certain high genus framed 3-manifolds
Authors:
Liang Chang,
Yilong Wang,
Saifei Zhai
Abstract:
We show that the Kuperberg invariant of the Weeks manifold with any framing is a gauge invariant of finite-dimensional Hopf algebras, which provides the first example of gauge invariants of general finite-dimensional Hopf algebras via hyperbolic 3-manifolds. We also show that the Kuperberg invariant of the 3-torus is gauge invariant, which further supports the idea of systematically producing gaug…
▽ More
We show that the Kuperberg invariant of the Weeks manifold with any framing is a gauge invariant of finite-dimensional Hopf algebras, which provides the first example of gauge invariants of general finite-dimensional Hopf algebras via hyperbolic 3-manifolds. We also show that the Kuperberg invariant of the 3-torus is gauge invariant, which further supports the idea of systematically producing gauge invariants of Hopf algebras via topological methods proposed in \cite{CNW25}.
△ Less
Submitted 27 January, 2026;
originally announced January 2026.
-
Estimating Dense-Packed Zone Height in Liquid-Liquid Separation: A Physics-Informed Neural Network Approach
Authors:
Mehmet Velioglu,
Song Zhai,
Alexander Mitsos,
Adel Mhamdi,
Andreas Jupke,
Manuel Dahmen
Abstract:
Separating liquid-liquid dispersions in gravity settlers is critical in chemical, pharmaceutical, and recycling processes. The dense-packed zone height is an important performance and safety indicator but it is often expensive and impractical to measure due to optical limitations. We propose a framework to estimate phase heights by combining a PINN model with readily available volume flow measurem…
▽ More
Separating liquid-liquid dispersions in gravity settlers is critical in chemical, pharmaceutical, and recycling processes. The dense-packed zone height is an important performance and safety indicator but it is often expensive and impractical to measure due to optical limitations. We propose a framework to estimate phase heights by combining a PINN model with readily available volume flow measurements, without requiring phase height measurements during deployment. To this end, a physics-informed neural network (PINN) is first pretrained on synthetic data and physics equations derived from a low-fidelity (approximate) mechanistic model to reduce the need for extensive experimental data. While the mechanistic model is used to generate synthetic training data, only volume balance equations are used in the PINN, as incorporating droplet coalescence and sedimentation submodels would be computationally prohibitive. The pretrained PINN is then fine-tuned with scarce experimental phase height and flow-rate data to capture the actual dynamics of the separator. We then deploy the differentiable PINN as a predictive model in an Extended Kalman Filter inspired state estimation framework, enabling the phase heights to be tracked and updated using flow-rate measurements only. We first test the two-stage trained PINN by forward simulation from a known initial state against the mechanistic model and a non-pretrained PINN. We then evaluate phase height estimation performance with the filter, comparing the two-stage trained PINN with a two-stage trained purely data-driven neural network. All model types are trained and evaluated using ensembles to account for model parameter uncertainty. In all evaluations, the two-stage trained PINN yields the most accurate phase-height estimates.
△ Less
Submitted 27 April, 2026; v1 submitted 26 January, 2026;
originally announced January 2026.
-
UniPR-3D: Towards Universal Visual Place Recognition with Visual Geometry Grounded Transformer
Authors:
Tianchen Deng,
Xun Chen,
Ziming Li,
Hongming Shen,
Shuhao Zhai,
Danwei Wang,
Javier Civera,
Hesheng Wang
Abstract:
Visual Place Recognition (VPR) has been traditionally formulated as a single-image retrieval task. Using multiple views offers clear advantages, yet this setting remains relatively underexplored and existing methods often struggle to generalize across diverse environments. In this work we introduce UniPR-3D, the first VPR architecture that effectively integrates information from multiple views. Un…
▽ More
Visual Place Recognition (VPR) has been traditionally formulated as a single-image retrieval task. Using multiple views offers clear advantages, yet this setting remains relatively underexplored and existing methods often struggle to generalize across diverse environments. In this work we introduce UniPR-3D, the first VPR architecture that effectively integrates information from multiple views. UniPR-3D builds on a VGGT backbone capable of encoding multi-view 3D representations, which we adapt by designing feature aggregators and fine-tune for the place recognition task. To construct our descriptor, we jointly leverage the 3D tokens and intermediate 2D tokens produced by VGGT. Based on their distinct characteristics, we design dedicated aggregation modules for 2D and 3D features, allowing our descriptor to capture fine-grained texture cues while also reasoning across viewpoints. To further enhance generalization, we incorporate both single- and multi-frame aggregation schemes, along with a variable-length sequence retrieval strategy. Our experiments show that UniPR-3D sets a new state of the art, outperforming both single- and multi-view baselines and highlighting the effectiveness of geometry-grounded tokens for VPR. Our code and models will be made publicly available on Github https://github.com/dtc111111/UniPR-3D.
△ Less
Submitted 29 June, 2026; v1 submitted 24 December, 2025;
originally announced December 2025.
-
ETP-R1: Evolving Topological Planning with Reinforcement Fine-tuning for Vision-Language Navigation in Continuous Environments
Authors:
Shuhao Ye,
Sitong Mao,
Yuxiang Cui,
Xuan Yu,
Shichao Zhai,
Wen Chen,
Shunbo Zhou,
Rong Xiong,
Yue Wang
Abstract:
Vision-Language Navigation in Continuous Environments (VLN-CE) requires an embodied agent to navigate towards target in continuous environments, following natural language instructions. While current graph-based methods offer an efficient, structured approach by abstracting the environment into a topological map and simplifying the action space to waypoint selection, they lag behind methods based…
▽ More
Vision-Language Navigation in Continuous Environments (VLN-CE) requires an embodied agent to navigate towards target in continuous environments, following natural language instructions. While current graph-based methods offer an efficient, structured approach by abstracting the environment into a topological map and simplifying the action space to waypoint selection, they lag behind methods based on Large Vision-Language Models (LVLMs) in leveraging large-scale data and advanced training paradigms. In this paper, we try to bridge this gap by introducing ETP-R1, a framework that applies the paradigm of scaling up data and Reinforcement Fine-Tuning (RFT) to a graph-based VLN-CE model. To build a strong foundation, we first construct a high-quality, large-scale pretraining dataset using the Gemini API. This dataset consists of diverse, low-hallucination instructions for topological trajectories, providing rich supervision for our graph-based policy to map language to topological paths. This foundation is further strengthened by unifying data from both R2R and RxR tasks for joint pretraining. Building on this, we introduce a three-stage training paradigm, which culminates in the first application of closed-loop, online RFT to a graph-based VLN-CE model, powered by the Group Relative Policy Optimization (GRPO) algorithm. Extensive experiments demonstrate that our approach is highly effective, establishing new state-of-the-art performance across all major metrics on both the R2R-CE and RxR-CE benchmarks. Our code is available at https://github.com/Cepillar/ETP-R1.
△ Less
Submitted 23 December, 2025;
originally announced December 2025.
-
SAVeD: A First-Person Social Media Video Dataset for ADAS-equipped vehicle Near-Miss and Crash Event Analyses
Authors:
Shaoyan Zhai,
Mohamed Abdel-Aty,
Chenzhu Wang,
Rodrigo Vena Garcia
Abstract:
The advancement of safety-critical research in driving behavior in ADAS-equipped vehicles require real-world datasets that not only include diverse traffic scenarios but also capture high-risk edge cases such as near-miss events and system failures. However, existing datasets are largely limited to either simulated environments or human-driven vehicle data, lacking authentic ADAS (Advanced Driver…
▽ More
The advancement of safety-critical research in driving behavior in ADAS-equipped vehicles require real-world datasets that not only include diverse traffic scenarios but also capture high-risk edge cases such as near-miss events and system failures. However, existing datasets are largely limited to either simulated environments or human-driven vehicle data, lacking authentic ADAS (Advanced Driver Assistance System) vehicle behavior under risk conditions. To address this gap, this paper introduces SAVeD, a large-scale video dataset curated from publicly available social media content, explicitly focused on ADAS vehicle-related crashes, near-miss incidents, and disengagements. SAVeD features 2,119 first-person videos, capturing ADAS vehicle operations in diverse locations, lighting conditions, and weather scenarios. The dataset includes video frame-level annotations for collisions, evasive maneuvers, and disengagements, enabling analysis of both perception and decision-making failures. We demonstrate SAVeD's utility through multiple analyses and contributions: (1) We propose a novel framework integrating semantic segmentation and monocular depth estimation to compute real-time Time-to-Collision (TTC) for dynamic objects. (2) We utilize the Generalized Extreme Value (GEV) distribution to model and quantify the extreme risk in crash and near-miss events across different roadway types. (3) We establish benchmarks for state-of-the-art VLLMs (VideoLLaMA2 and InternVL2.5 HiCo R16), showing that SAVeD's detailed annotations significantly enhance model performance through domain adaptation in complex near-miss scenarios.
△ Less
Submitted 19 December, 2025;
originally announced December 2025.
-
Supermassive Black Holes with High Accretion Rates in Active Galactic Nuclei. XV. Reverberation Mapping of Mg II Emission Lines
Authors:
Hua-Rui Bai,
Pu Du,
Chen Hu,
Yong-Jie Chen,
Zhu-Heng Yao,
Yan-Rong Li,
Yi-Xin Fu,
Yi-Lin Wang,
Yu Zhao,
Hao Zhang,
Jun-Rong Liu,
Sen Yang,
Yue-Chang Peng,
Feng-Na Fang,
Yu-Yang Songsheng,
Ming Xiao,
Shuo Zhai,
Sha-Sha Li,
Kai-Xing Lu,
Zhi-Xiang Zhang,
Dong-Wei Bao,
Wei-Jian Guo,
Jia-Qi Feng,
Yi-Peng Zhao,
Jesús Aceituno
, et al. (3 additional authors not shown)
Abstract:
As the 15th paper in a series reporting on a large reverberation mapping (RM) campaign of super-Eddington accreting massive black holes (SEAMBHs) in active galactic nuclei (AGNs), we present the results of measurements of the Mg II lines in 18 SEAMBHs monitored spectroscopically from 2017 to 2024. Among these, the time lags of Mg II have been successfully determined for 8 of the 18 objects, thereb…
▽ More
As the 15th paper in a series reporting on a large reverberation mapping (RM) campaign of super-Eddington accreting massive black holes (SEAMBHs) in active galactic nuclei (AGNs), we present the results of measurements of the Mg II lines in 18 SEAMBHs monitored spectroscopically from 2017 to 2024. Among these, the time lags of Mg II have been successfully determined for 8 of the 18 objects, thereby expanding the current Mg II RM sample, particularly at higher accretion rates. By incorporating measurements of the line widths, we determine the masses of their central supermassive black holes. Based on these new measurements, we update the relation between the Mg II radius and the monochromatic luminosity at 3000 $\mathring{\mathrm{A}}$ ($R_{\rm MgII}-L_{3000}$ relation), yielding a slope of $0.24 \pm 0.03$, which is slightly shallower than, yet still consistent with, previously reported values. Similar to the H$β$ lines, the Mg II time lags in SEAMBHs are shorter than those of AGNs with normal accretion rates at comparable luminosities. The deviation of AGNs from the best-fit $R_{\rm MgII}-L_{3000}$ relation shows a strong correlation with the accretion rate, while no significant correlation is found between the deviation and the flux ratio of UV iron to Mg II.
△ Less
Submitted 8 December, 2025;
originally announced December 2025.
-
Cultural Prompting Improves the Empathy and Cultural Responsiveness of GPT-Generated Therapy Responses
Authors:
Serena Jinchen Xie,
Shumenghui Zhai,
Yanjing Liang,
Jingyi Li,
Xuehong Fan,
Trevor Cohen,
Weichao Yuwen
Abstract:
Large Language Model (LLM)-based conversational agents offer promising solutions for mental health support, but lack cultural responsiveness for diverse populations. This study evaluated the effectiveness of cultural prompting in improving cultural responsiveness and perceived empathy of LLM-generated therapeutic responses for Chinese American family caregivers. Using a randomized controlled exper…
▽ More
Large Language Model (LLM)-based conversational agents offer promising solutions for mental health support, but lack cultural responsiveness for diverse populations. This study evaluated the effectiveness of cultural prompting in improving cultural responsiveness and perceived empathy of LLM-generated therapeutic responses for Chinese American family caregivers. Using a randomized controlled experiment, we compared GPT-4o and Deepseek-V3 responses with and without cultural prompting. Thirty-six participants evaluated input-response pairs on cultural responsiveness (competence and relevance) and perceived empathy. Results showed that cultural prompting significantly enhanced GPT-4o's performance across all dimensions, with GPT-4o with cultural prompting being the most preferred, while improvements in DeepSeek-V3 responses were not significant. Mediation analysis revealed that cultural prompting improved empathy through improving cultural responsiveness. This study demonstrated that prompt-based techniques can effectively enhance the cultural responsiveness of LLM-generated therapeutic responses, highlighting the importance of cultural responsiveness in delivering empathetic AI-based therapeutic interventions to culturally and linguistically diverse populations.
△ Less
Submitted 18 October, 2025;
originally announced December 2025.
-
STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flows
Authors:
Jiatao Gu,
Ying Shen,
Tianrong Chen,
Laurent Dinh,
Yuyang Wang,
Miguel Angel Bautista,
David Berthelot,
Josh Susskind,
Shuangfei Zhai
Abstract:
Normalizing flows (NFs) are end-to-end likelihood-based generative models for continuous data, and have recently regained attention with encouraging progress on image generation. Yet in the video generation domain, where spatiotemporal complexity and computational cost are substantially higher, state-of-the-art systems almost exclusively rely on diffusion-based models. In this work, we revisit thi…
▽ More
Normalizing flows (NFs) are end-to-end likelihood-based generative models for continuous data, and have recently regained attention with encouraging progress on image generation. Yet in the video generation domain, where spatiotemporal complexity and computational cost are substantially higher, state-of-the-art systems almost exclusively rely on diffusion-based models. In this work, we revisit this design space by presenting STARFlow-V, a normalizing flow-based video generator with substantial benefits such as end-to-end learning, robust causal prediction, and native likelihood estimation. Building upon the recently proposed STARFlow, STARFlow-V operates in the spatiotemporal latent space with a global-local architecture which restricts causal dependencies to a global latent space while preserving rich local within-frame interactions. This eases error accumulation over time, a common pitfall of standard autoregressive diffusion model generation. Additionally, we propose flow-score matching, which equips the model with a light-weight causal denoiser to improve the video generation consistency in an autoregressive fashion. To improve the sampling efficiency, STARFlow-V employs a video-aware Jacobi iteration scheme that recasts inner updates as parallelizable iterations without breaking causality. Thanks to the invertible structure, the same model can natively support text-to-video, image-to-video as well as video-to-video generation tasks. Empirically, STARFlow-V achieves strong visual fidelity and temporal consistency with practical sampling throughput relative to diffusion-based baselines. These results present the first evidence, to our knowledge, that NFs are capable of high-quality autoregressive video generation, establishing them as a promising research direction for building world models. Code and generated samples are available at https://github.com/apple/ml-starflow.
△ Less
Submitted 25 November, 2025; v1 submitted 25 November, 2025;
originally announced November 2025.
-
Chemical evolution of bulges of active galactic nuclei in the early Universe: roles of accreting stars
Authors:
Shuo Zhai,
Jian-Min Wang,
Yan-Rong Li,
Wei-Jian Guo,
Gang Zhao
Abstract:
JWST/NIRCam observations reveal dense stellar cores in high-redshift galactic bulges, indicative of sustained star formation and potential stellar accretion. We introduce accretion-modified star (AMS) as a new component in the chemical evolution of high-redshift bulges hosting active galactic nuclei (AGNs). The gas-phase chemical evolution of bulge environments containing AMS is modeled within 1 G…
▽ More
JWST/NIRCam observations reveal dense stellar cores in high-redshift galactic bulges, indicative of sustained star formation and potential stellar accretion. We introduce accretion-modified star (AMS) as a new component in the chemical evolution of high-redshift bulges hosting active galactic nuclei (AGNs). The gas-phase chemical evolution of bulge environments containing AMS is modeled within 1 Gyr by combining population evolution and galactic chemical evolution formalisms, and observational signatures are tracked via photoionization modeling on Baldwin-Phillips-Terlevich (BPT) diagrams. Sustained high accretion onto AMSs leads to rapid gas-phase metal enrichment of the bulge, producing abundance peaks up to five times solar metallicity within 0.1 Gyr and significantly modifying elemental ratios in the gas phase. Atypical gas-phase abundance patterns during early, high-accretion phases and gradually diminish as the accretion rate declines. In BPT diagrams, high-AMS-accretion scenarios shift the modeled emission-line sequence toward the local AGN branch and extend into the high-metallicity regime. Super-solar narrow-line regions observed in AGNs at z>15 may reflect such AMS-driven gas-phase enrichment of host bulge under extreme gas densities. While direct detection of AMSs within AGN bulges remains challenging, the model provides testable predictions for future spectroscopic surveys and motivates further exploration of non-canonical stellar populations in AGN host bulges.
△ Less
Submitted 20 November, 2025;
originally announced November 2025.
-
Changing-look Active Galactic Nuclei from the Dark Energy Spectroscopic Instrument. IV. Broad Emission Line Evolution Sequence Among Hα, Mg II, and Hβ
Authors:
Wei-Jian Guo,
Victoria A. Fawcett,
Małgorzata Siudek,
Yan-Rong Li,
Cheng Cheng,
Swayamtrupta Panda,
Zhiwei Pan,
Shengxiu Sun,
Claire L. Greenwell,
David M. Alexander,
John Moustakas,
Shuo Zhai,
Jun-Jie Jin,
Huaqing Cheng,
Jingwei Hu,
Yong-Jie Chen,
Zhi-Xiang Zhang,
Jian-Min Wang
Abstract:
From a parent catalog of 561 changing-look active galactic nuclei (CL-AGNs) identified by Guo et al. (2025), we investigate the evolutionary sequence of broad emission lines using a redshift-selected subset (0.35 < z < 0.45) of 54 CL-AGNs whose Dark Energy Spectroscopic Instrument (DESI) spectra simultaneously cover the Hα, H\b{eta}, and Mg II emission lines. To provide a baseline for comparison,…
▽ More
From a parent catalog of 561 changing-look active galactic nuclei (CL-AGNs) identified by Guo et al. (2025), we investigate the evolutionary sequence of broad emission lines using a redshift-selected subset (0.35 < z < 0.45) of 54 CL-AGNs whose Dark Energy Spectroscopic Instrument (DESI) spectra simultaneously cover the Hα, H\b{eta}, and Mg II emission lines. To provide a baseline for comparison, we construct a control sample of 19,897 normal Type 1 AGNs within the same redshift range from the DESI Year 1 data. Through stacked spectral analysis and line-continuum luminosity correlations, we identify a clear evolutionary sequence in all AGN where broad H\b{eta} fades first, followed by Mg II, and then Hα, as the AGN luminosity declines - consistent with expectations from reverberation mapping. This trend reflects a radially stratified broad line region (BLR), where each line's responsivity depends on its ionization potential and radial distance from the central engine. In addition, we find that more massive supermassive black holes (SMBHs) require lower Eddington ratios to fully suppress broad emission lines, suggesting that the critical accretion threshold for the CL phenomenon is mass-dependent. Our results present the first statistical confirmation of a stratified broad line fading sequence in AGNs, reinforcing the central role of accretion state in shaping BLR structure and visibility.
△ Less
Submitted 19 November, 2025;
originally announced November 2025.
-
Mip-NeWRF: Enhanced Wireless Radiance Field with Hybrid Encoding for Channel Prediction
Authors:
Yulin Fu,
Jiancun Fan,
Shiyu Zhai,
Zhibo Duan,
Jie Luo
Abstract:
Recent work on wireless radiance fields represents a promising deep learning approach for channel prediction, however, in complex environments these methods still exhibit limited robustness, slow convergence, and modest accuracy due to insufficiently refined modeling. To address this issue, we propose Mip-NeWRF, a physics-informed neural framework for accurate indoor channel prediction based on sp…
▽ More
Recent work on wireless radiance fields represents a promising deep learning approach for channel prediction, however, in complex environments these methods still exhibit limited robustness, slow convergence, and modest accuracy due to insufficiently refined modeling. To address this issue, we propose Mip-NeWRF, a physics-informed neural framework for accurate indoor channel prediction based on sparse channel measurements. The framework operates in a ray-based pipeline with coarse-to-fine importance sampling: frustum samples are encoded, processed by a shared multilayer perceptron (MLP), and the outputs are synthesized into the channel frequency response (CFR). Prior to MLP input, Mip-NeWRF performs conical-frustum sampling and applies a scale-consistent hybrid positional encoding to each frustum. The scale-consistent normalization aligns positional encodings across scene scales, while the hybrid encoding supplies both scale-robust, low-frequency stability to accelerate convergence and fine spatial detail to improve accuracy. During training, a curriculum learning schedule is applied to stabilize and accelerate convergence of the shared MLP. During channel synthesis, the MLP outputs, including predicted virtual transmitter presence probabilities and amplitudes, are combined with modeled pathloss and surface interaction attenuation to enhance physical fidelity and further improve accuracy. Simulation results demonstrate the effectiveness of the proposed approach: in typical scenarios, the normalized mean square error (NMSE) is reduced by 14.3 dB versus state-of-the-art baselines.
△ Less
Submitted 31 July, 2026; v1 submitted 12 November, 2025;
originally announced November 2025.
-
Conditional Flow Matching for Bayesian Posterior Inference
Authors:
Percy S. Zhai,
So Won Jeong,
Veronika Ročková
Abstract:
We propose a generative multivariate posterior sampler via flow matching. It offers a simple training objective, and does not require access to likelihood evaluation. The method learns a dynamic, block-triangular velocity field in the joint space of data and parameters, which results in a deterministic transport map from a source distribution to the desired posterior. The inverse map, named vector…
▽ More
We propose a generative multivariate posterior sampler via flow matching. It offers a simple training objective, and does not require access to likelihood evaluation. The method learns a dynamic, block-triangular velocity field in the joint space of data and parameters, which results in a deterministic transport map from a source distribution to the desired posterior. The inverse map, named vector rank, is accessible by reversibly integrating the velocity over time. It is advantageous to leverage the dynamic design: proper constraints on the velocity yield a monotone map, which leads to a conditional Brenier map, enabling a fast and simultaneous generation of Bayesian credible sets whose contours correspond to level sets of Monge-Kantorovich data depth. Our approach is computationally lighter compared to GAN-based and diffusion-based counterparts, and is capable of capturing complex posterior structures. Finally, frequentist theoretical guarantee on the consistency of the recovered posterior distribution, and of the corresponding Bayesian credible sets, is provided.
△ Less
Submitted 31 March, 2026; v1 submitted 10 October, 2025;
originally announced October 2025.
-
DMark: Order-Agnostic Watermarking for Diffusion Large Language Models
Authors:
Linyu Wu,
Linhao Zhong,
Wenjie Qu,
Yuexin Li,
Yue Liu,
Shengfang Zhai,
Chunhua Shen,
Jiaheng Zhang
Abstract:
Diffusion large language models (dLLMs) offer faster generation than autoregressive models while maintaining comparable quality, but existing watermarking methods fail on them due to their non-sequential decoding. Unlike autoregressive models that generate tokens left-to-right, dLLMs can finalize tokens in arbitrary order, breaking the causal design underlying traditional watermarks. We present DM…
▽ More
Diffusion large language models (dLLMs) offer faster generation than autoregressive models while maintaining comparable quality, but existing watermarking methods fail on them due to their non-sequential decoding. Unlike autoregressive models that generate tokens left-to-right, dLLMs can finalize tokens in arbitrary order, breaking the causal design underlying traditional watermarks. We present DMark, the first watermarking framework designed specifically for dLLMs. DMark introduces three complementary strategies to restore watermark detectability: predictive watermarking uses model-predicted tokens when actual context is unavailable; bidirectional watermarking exploits both forward and backward dependencies unique to diffusion decoding; and predictive-bidirectional watermarking combines both approaches to maximize detection strength. Experiments across multiple dLLMs show that DMark achieves 92.0-99.5% detection rates at 1% false positive rate while maintaining text quality, compared to only 49.6-71.2% for naive adaptations of existing methods. DMark also demonstrates robustness against text manipulations, establishing that effective watermarking is feasible for non-autoregressive language models.
△ Less
Submitted 3 October, 2025;
originally announced October 2025.
-
Graph-Based Spatio-temporal Attention and Multi-Scale Fusion for Clinically Interpretable, High-Fidelity Fetal ECG Extraction
Authors:
Chang Wang,
Ming Zhu,
Shahram Latifi,
Buddhadeb Dawn,
Shengjie Zhai
Abstract:
Congenital Heart Disease (CHD) is the most common neonatal anomaly, highlighting the urgent need for early detection to improve outcomes. Yet, fetal ECG (fECG) signals in abdominal ECG (aECG) are often masked by maternal ECG and noise, challenging conventional methods under low signal-to-noise ratio (SNR) conditions. We propose FetalHealthNet (FHNet), a deep learning framework that integrates Grap…
▽ More
Congenital Heart Disease (CHD) is the most common neonatal anomaly, highlighting the urgent need for early detection to improve outcomes. Yet, fetal ECG (fECG) signals in abdominal ECG (aECG) are often masked by maternal ECG and noise, challenging conventional methods under low signal-to-noise ratio (SNR) conditions. We propose FetalHealthNet (FHNet), a deep learning framework that integrates Graph Neural Networks with a multi-scale enhanced transformer to dynamically model spatiotemporal inter-lead correlations and extract clean fECG signals. On benchmark aECG datasets, FHNet consistently outperforms long short-term memory (LSTM) models, standard transformers, and state-of-the-art models, achieving R2>0.99 and RMSE = 0.015 even under severe noise. Interpretability analyses highlight physiologically meaningful temporal and lead contributions, supporting model transparency and clinical trust. FHNet illustrates the potential of AI-driven modeling to advance fetal monitoring and enable early CHD screening, underscoring the transformative impact of next-generation biomedical signal processing.
△ Less
Submitted 5 September, 2025;
originally announced September 2025.
-
A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning
Authors:
Shaopeng Zhai,
Qi Zhang,
Tianyi Zhang,
Fuxian Huang,
Haoran Zhang,
Ming Zhou,
Shengzhe Zhang,
Litao Liu,
Sixu Lin,
Jiangmiao Pang
Abstract:
Robotic real-world reinforcement learning (RL) with vision-language-action (VLA) models is bottlenecked by sparse, handcrafted rewards and inefficient exploration. We introduce VLAC, a general process reward model built upon InternVL and trained on large scale heterogeneous datasets. Given pairwise observations and a language goal, it outputs dense progress delta and done signal, eliminating task-…
▽ More
Robotic real-world reinforcement learning (RL) with vision-language-action (VLA) models is bottlenecked by sparse, handcrafted rewards and inefficient exploration. We introduce VLAC, a general process reward model built upon InternVL and trained on large scale heterogeneous datasets. Given pairwise observations and a language goal, it outputs dense progress delta and done signal, eliminating task-specific reward engineering, and supports one-shot in-context transfer to unseen tasks and environments. VLAC is trained on vision-language datasets to strengthen perception, dialogic and reasoning capabilities, together with robot and human trajectories data that ground action generation and progress estimation, and additionally strengthened to reject irrelevant prompts as well as detect regression or stagnation by constructing large numbers of negative and semantically mismatched samples. With prompt control, a single VLAC model alternately generating reward and action tokens, unifying critic and policy. Deployed inside an asynchronous real-world RL loop, we layer a graded human-in-the-loop protocol (offline demonstration replay, return and explore, human guided explore) that accelerates exploration and stabilizes early learning. Across four distinct real-world manipulation tasks, VLAC lifts success rates from about 30\% to about 90\% within 200 real-world interaction episodes; incorporating human-in-the-loop interventions yields a further 50% improvement in sample efficiency and achieves up to 100% final success.
△ Less
Submitted 19 September, 2025;
originally announced September 2025.
-
Enhancing Privacy in Decentralized Min-Max Optimization: A Differentially Private Approach
Authors:
Yueyang Quan,
Chang Wang,
Shengjie Zhai,
Minghong Fang,
Zhuqing Liu
Abstract:
Decentralized min-max optimization allows multi-agent systems to collaboratively solve global min-max optimization problems by facilitating the exchange of model updates among neighboring agents, eliminating the need for a central server. However, sharing model updates in such systems carry a risk of exposing sensitive data to inference attacks, raising significant privacy concerns. To mitigate th…
▽ More
Decentralized min-max optimization allows multi-agent systems to collaboratively solve global min-max optimization problems by facilitating the exchange of model updates among neighboring agents, eliminating the need for a central server. However, sharing model updates in such systems carry a risk of exposing sensitive data to inference attacks, raising significant privacy concerns. To mitigate these privacy risks, differential privacy (DP) has become a widely adopted technique for safeguarding individual data. Despite its advantages, implementing DP in decentralized min-max optimization poses challenges, as the added noise can hinder convergence, particularly in non-convex scenarios with complex agent interactions in min-max optimization problems. In this work, we propose an algorithm called DPMixSGD (Differential Private Minmax Hybrid Stochastic Gradient Descent), a novel privacy-preserving algorithm specifically designed for non-convex decentralized min-max optimization. Our method builds on the state-of-the-art STORM-based algorithm, one of the fastest decentralized min-max solutions. We rigorously prove that the noise added to local gradients does not significantly compromise convergence performance, and we provide theoretical bounds to ensure privacy guarantees. To validate our theoretical findings, we conduct extensive experiments across various tasks and models, demonstrating the effectiveness of our approach.
△ Less
Submitted 10 August, 2025;
originally announced August 2025.