-
Three-dimensional imaging of oxygen dopant distribution in Sr$_2$CuO$_{3+δ}$ by electron ptychography
Authors:
Hongbin Yang,
Jinkwon Kim,
Desheng Ma,
Dasol Yoon,
Darrell G. Schlom,
David A. Muller
Abstract:
Oxygen dopants play a critical role in tuning the properties of cuprate superconductors, yet it is challenging to visualize them at the atomic scale. Here, we use multislice electron ptychography to directly image oxygen dopants in a Sr2CuO3+delta film. We observe oxygen dopants at interstitial sites between the Cu-O chains, with a strong preference for clustering in tensile-strained regions, whic…
▽ More
Oxygen dopants play a critical role in tuning the properties of cuprate superconductors, yet it is challenging to visualize them at the atomic scale. Here, we use multislice electron ptychography to directly image oxygen dopants in a Sr2CuO3+delta film. We observe oxygen dopants at interstitial sites between the Cu-O chains, with a strong preference for clustering in tensile-strained regions, which are often associated with dislocations and interfacial steps. These findings indicate that the oxygen dopant distribution in cuprates is not random but rather sensitive to strain field, suggesting strain as a doping tuning parameter.
△ Less
Submitted 20 August, 2026;
originally announced August 2026.
-
Hundred-hertz quantum circuit iteration rate in a reusable neutral-atom array
Authors:
Liang Chen,
Wen-Yi Zhu,
Dong-Qi Ma,
Tian-Yang Zhang,
Zi-Jie Chen,
Yi-Chen Zhang,
Hong-Jie Fan,
Guang-Jie Chen,
Qing-Xuan Jie,
Wei-Zhou Cai,
Tian-Cai Zhang,
Luyan Sun,
Yan-Lei Zhang,
Xi-Feng Ren,
Guang-Can Guo,
Zhu-Bo Wang,
Ya-Dong Hu,
Gang Li,
Chang-Ling Zou
Abstract:
Neutral-atom quantum processors have rapidly advanced in scale and coherence, yet their practical performance remains constrained by limited quantum circuit iteration rates (qCIRs) and information throughput. Here we experimentally demonstrate a high-throughput neutral-atom system based on non-destructive readout and atom reuse. By integrating a chip-based photonic interface with a 10-qubit array,…
▽ More
Neutral-atom quantum processors have rapidly advanced in scale and coherence, yet their practical performance remains constrained by limited quantum circuit iteration rates (qCIRs) and information throughput. Here we experimentally demonstrate a high-throughput neutral-atom system based on non-destructive readout and atom reuse. By integrating a chip-based photonic interface with a 10-qubit array, we implement non-destructive readout with a retention probability of 99.7%, and further achieve a raw qCIR of 101Hz and a post-selected qCIR of 74.8Hz. More importantly, we verify a general throughput optimization methodology and obtain a normalized Fisher information rate of 57.7Hz, improving the achievable throughput by more than one order of magnitude compared with conventional methods. Our results establish a practical route toward high-throughput neutral-atom quantum processors.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
A scalable chip-integrated single-photon source array based on 50 individually addressable neutral atoms
Authors:
Ya-Dong Hu,
Tian-Yang Zhang,
Dong-Qi Ma,
Yi-Chen Zhang,
Liang Chen,
Wen-Yi Zhu,
Hong-Jie Fan,
Yan-Lei Zhang,
Zhu-Bo Wang,
Gang Li,
Xi-Feng Ren,
Guang-Can Guo,
Chang-Ling Zou
Abstract:
Scalable arrays of identical single-photon sources are a central resource for photonic quantum information processing, quantum networks and quantum metrology. Neutral atoms provide intrinsically identical emitters that can be assembled and rearranged in optical tweezers, but a many-channel fiber interface to individually trapped atoms has remained a major technical challenge. Here we demonstrate a…
▽ More
Scalable arrays of identical single-photon sources are a central resource for photonic quantum information processing, quantum networks and quantum metrology. Neutral atoms provide intrinsically identical emitters that can be assembled and rearranged in optical tweezers, but a many-channel fiber interface to individually trapped atoms has remained a major technical challenge. Here we demonstrate a chip-interfaced single-photon source array based on 50 individually addressable $^{87}\mathrm{Rb}$ atoms. A glass waveguide fan-out converts the \SI{5}{\micro m} pitch of the optical-tweezer array to the \SI{127}{\micro m} pitch of a commercial fiber array, mapping each atom to its own waveguide, fiber and single-photon detector. We resolve all 50 channels with an average nearest-neighbor cross-talk of $0.4\%$ and a uniform insertion loss of \SI{2.9}{dB}, and verify single-photon emission with $g^{(2)}(0)=0.29$, presently limited by detector dark counts and residual cooling-light scattering. Combining per-channel atom discrimination, rearrangement and reservoir replenishment, we prepare source subarrays of up to 24 atoms with a $93\%$ fill fraction. For small target numbers, atom loss is repaired from the reservoir at the detection-limited rate of \SI{118}{Hz}. We further fabricate a 784-channel waveguide chip, showing that the photonic interface can be extended well beyond the present number. This architecture establishes a fiber-native neutral-atom platform for larger arrays of identical single-photon sources.
△ Less
Submitted 16 August, 2026;
originally announced August 2026.
-
LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures
Authors:
Yunfei Zhang,
Boyu Feng,
Changhua Pei,
Zexin Wang,
Zhihuang Peng,
Xinlong Liu,
Hengyue Jiang,
Difeng Ma,
Jiayi Zhang,
Yongzhou Yao,
Yanan Zhao,
Fei Sun,
Yintong Huo,
Zhaoyang Liu,
Jingjing Li,
Gaogang Xie,
Dan Pei
Abstract:
When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory. Developers must then inspect the full execution to identify the responsible role and localize the earliest decisive root-cause step. Existing failure-attribution benchmarks largely focus on shorter traces, leaving diagnosis across hundreds of…
▽ More
When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory. Developers must then inspect the full execution to identify the responsible role and localize the earliest decisive root-cause step. Existing failure-attribution benchmarks largely focus on shorter traces, leaving diagnosis across hundreds of recorded steps underexplored. We introduce LongRCA Bench, comprising 1,140 failed trajectories across five domains without injected errors. It provides independently scored human labels for the responsible role and earliest decisive root-cause step. The median trajectory contains 145 steps, and the strongest baseline reaches only 13.2% exact root-step accuracy. We further present Root-Cause Trajectory Attribution (RCTA), a training-free method that retrieves candidate error steps from segment summaries and traces them to available earlier handoff instructions. Using the same backbone, benchmark instances, and scoring protocol, RCTA reaches 51.1% responsible-role accuracy and 24.1% exact root-step accuracy. These results highlight the need to evaluate responsible-role attribution and exact root-step localization as separate targets in long-trajectory failure diagnosis.
△ Less
Submitted 20 August, 2026; v1 submitted 15 August, 2026;
originally announced August 2026.
-
From Diagnosis to Correction: Benchmarking and Improving Real-World Table Parsing
Authors:
Jutao Xiao,
Yuan Qu,
Dongsheng Ma,
Fan Wu,
Tianyao He,
Weihong Li,
Jie Yang,
Yu Qiao,
Bin Wang,
Conghui He
Abstract:
Recent document parsers achieve table TEDS scores above 93 on OmniDocBench v1.6, yet community feedback and our audit reveal persistent failures on complex real-world tables. To quantify this gap, we introduce TableParseMap, a diagnostic benchmark of 916 real-world tables organized into five challenging scenarios and nine failure types. The strongest evaluated parser achieves only 85.03 TEDS, show…
▽ More
Recent document parsers achieve table TEDS scores above 93 on OmniDocBench v1.6, yet community feedback and our audit reveal persistent failures on complex real-world tables. To quantify this gap, we introduce TableParseMap, a diagnostic benchmark of 916 real-world tables organized into five challenging scenarios and nine failure types. The strongest evaluated parser achieves only 85.03 TEDS, showing that aggregate benchmark scores conceal substantial weaknesses. Our analysis attributes these failures to three complementary limitations: large tables exceed the reliable processing scale of a single pass, weak or ambiguous visual cues hinder structure perception, and the reconstructed table may remain visually inconsistent with the image. We therefore propose DEC (Decompose--Enhance--Correct), a visual-consistency-guided agentic framework that improves frozen table parsers without retraining. DEC uses a general VLM as the controller: Decompose partitions large tables along structure-aware boundaries, Enhance exposes weak visual evidence and reparses transformed views, and Correct diagnoses and repairs residual errors. A Visual Consistency Gate (VC-Gate) selectively triggers intervention, while a Visual Consistency Ranker (VC-Ranker) verifies candidate updates and supports rollback without ground-truth HTML at inference time. We further derive a 1,977-table Consensus-Hard Set from 4,556 candidates through offline metrics and cross-model consensus. Across three frozen parsers, DEC improves TEDS by 1.57 points on average; on TableParseMap, gains reach 1.89 points overall, 2.62 on structural errors, and 5.66 on large tables.
△ Less
Submitted 10 August, 2026;
originally announced August 2026.
-
ToolVision: Learning When and How to Use Visual Tools with Capability-Aligned Supervision
Authors:
Delin Mao,
Chenghao Sun,
Jingwei Song,
Chishui Chen,
Linfeng Zhang
Abstract:
Thinking with images allows a multimodal model to compensate for limited perception by invoking visual tools through code. Yet the prevailing SFT-then-RL recipe creates a different supervision misalignment at each stage. SFT is expected to teach how to use tools, but trajectories from stronger teachers may succeed through perceptual capabilities that a smaller student cannot reliably reproduce or…
▽ More
Thinking with images allows a multimodal model to compensate for limited perception by invoking visual tools through code. Yet the prevailing SFT-then-RL recipe creates a different supervision misalignment at each stage. SFT is expected to teach how to use tools, but trajectories from stronger teachers may succeed through perceptual capabilities that a smaller student cannot reliably reproduce or exploit, causing the student to imitate tool-call patterns without learning how to make them useful. RL is expected to teach when to use tools, but outcome-only rewards make fallible tool execution a liability and suppress tool use, whereas a blanket bonus for every correct tool-using trajectory encourages valid but ineffective operations. To address these two misalignments, we introduce ToolVision. During SFT, a multi-agent pipeline explores candidate trajectories, and a committee including student-scale models scores stepwise evidence gain to rank and prune the search branches. Only successfully executed trajectories with correct final answers are retained for SFT. Before RL, ToolVision compares the learner's performance with and without tools, then rewards successful tool use only on questions where tools provide a clear benefit. Both signals are constructed automatically from public task data without additional human annotations of tool use or necessity. ToolVision-8B improves over its base on all seven main benchmarks, surpasses Thyme-7B, CodeVision-8B, and CodeDance-7B on all three high-resolution benchmarks, and outperforms Qwen3-VL-32B-Thinking on V* and HRBench 8K. We will publicly release the datasets and source code.
△ Less
Submitted 9 August, 2026;
originally announced August 2026.
-
Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression
Authors:
Tianyu Liang,
Xiangxi Zheng,
Yilin Wang,
Dongxing Mao
Abstract:
Vision-Text Compression (VTC) renders long texts into images and encodes them through the vision encoder (ViT), compressing thousands of text tokens into far fewer visual tokens. However, since the ViT is pretrained predominantly on natural images, it captures visual attributes (glyphs, font sizes, layout) rather than linguistic semantics, causing rendered-image representations to diverge from nat…
▽ More
Vision-Text Compression (VTC) renders long texts into images and encodes them through the vision encoder (ViT), compressing thousands of text tokens into far fewer visual tokens. However, since the ViT is pretrained predominantly on natural images, it captures visual attributes (glyphs, font sizes, layout) rather than linguistic semantics, causing rendered-image representations to diverge from native-text representations. We term this cross-path inconsistency and show, via rendering perturbation experiments, that it is a critical yet overlooked bottleneck of VTC. We propose SPIRAL (Self-improving Path Integration and Realignment), a self-supervised alignment framework that closes this gap using only the model's own text-path behavior as supervision, requiring no external teachers or additional annotations. SPIRAL operates at two complementary granularities: token-level on-policy distillation (OPD) for local faithfulness, and sequence-level preference optimization (DPO) for global coherence. On VTCBench, SPIRAL improves the overall score of Qwen3-VL-8B from 35.10 to 54.02, approaching the native text-input performance (55.60) and outperforming models up to 30x larger. The two granularities exhibit complementary strengths: OPD excels at retrieval and is sample-efficient, while DPO is stronger on reasoning and memory and scales better with data. SPIRAL's benefits also generalize to out-of-domain benchmarks, confirming that effective VTC hinges on aligning rendered-image representations back to native-text semantics.
△ Less
Submitted 3 August, 2026;
originally announced August 2026.
-
Look Ahead Before You Distill: Future Trajectory Validation of Teacher Guidance for Agentic On-Policy Distillation
Authors:
Chishui Chen,
Yaoyou Fan,
Te Sun,
Yi Yang,
Chenghao Sun,
Delin Mao,
Hongbo Qiao,
Zuowei Zhang,
Junxi Wang,
Chenxing Sun,
Yangen Hu,
Lu Pan,
Xuyang Liu,
Linfeng Zhang
Abstract:
On-policy distillation (OPD) provides teacher supervision on states visited by the student, reducing the distribution gap between training and inference. However, in multi-turn agentic tasks, student deviations may accumulate over time, gradually moving the trajectory away from states where teacher guidance remains effective. Our quantitative analysis further shows that high-disagreement states of…
▽ More
On-policy distillation (OPD) provides teacher supervision on states visited by the student, reducing the distribution gap between training and inference. However, in multi-turn agentic tasks, student deviations may accumulate over time, gradually moving the trajectory away from states where teacher guidance remains effective. Our quantitative analysis further shows that high-disagreement states offer promising opportunities for teacher guidance, but determining whether such guidance is beneficial requires examining its effect on subsequent student trajectories. We propose FutureBridge-OPD (FTB), which executes a short teacher bridge at a high disagreement state and uses the resulting student continuation to assess whether the bridge increases the density of positive distillation signals relative to the teacher. On ALFWorld, WebShop, and ScienceWorld, under the main Qwen3-32B teacher to Qwen3-1.7B student setting, FTB outperforms vanilla OPD and TCOD by an average of 16.6 and 7.6 points, respectively, and remains effective across student scales and teacher settings. Our code is publicly available at https://github.com/ChenChiShui/FutureBridge-OPD.
△ Less
Submitted 5 August, 2026; v1 submitted 3 August, 2026;
originally announced August 2026.
-
CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation
Authors:
Xianjing Han,
Yuhan Su,
Yang Deng,
Dong Ma,
Wee Peng Tay,
Bin Zhu
Abstract:
Text-to-video (T2V) generation models have advanced rapidly, yet their ability to represent diverse cultural contexts remains underexplored. Existing benchmarks mainly focus on perceptual quality, physical plausibility, and text-video alignment, but do not directly assess whether generated videos capture culturally specific objects, actions, rituals, visible text, or audio cues. We introduce Cultu…
▽ More
Text-to-video (T2V) generation models have advanced rapidly, yet their ability to represent diverse cultural contexts remains underexplored. Existing benchmarks mainly focus on perceptual quality, physical plausibility, and text-video alignment, but do not directly assess whether generated videos capture culturally specific objects, actions, rituals, visible text, or audio cues. We introduce CultureVidBench, a comprehensive benchmark for evaluating cultural understanding in T2V generation. CultureVidBench contains 1,000 curated prompts covering 12 countries, 6 continents, 8 cultural regions, and 14 cultural aspects organized into three categories: material culture, social practice & performance, and ritual & ceremony. Designed specifically for video generation, CultureVidBench emphasizes dynamic and multimodal cultural representation, including social interactions, ritual procedure, and culturally appropriate visible text and audio. We evaluate seven representative T2V models through human user studies and MLLM-based automatic assessment across cultural faithfulness, multimodal cultural rendering, semantic adherence, and perceptual quality. Results show that although current models achieve strong semantic adherence and visual quality, they often fail to faithfully capture fine-grained cultural details, particularly for underrepresented regions, rituals, and multimodal cultural cues.
△ Less
Submitted 3 August, 2026;
originally announced August 2026.
-
LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing
Authors:
Wen Zan,
Jiaqi Zhang,
Jianchao Tan,
Hong Liu,
Cunguang Wang,
Xiang Li,
Duyue Ma,
Guanyu Wu,
Yifan Lu,
Fengcun Li,
Yerui Sun,
Peng Pei,
Yuchen Xie,
Xunliang Cai
Abstract:
DeepSeek Sparse Attention (DSA) enables efficient long-context modeling through its Lightning Indexer. However, practical deployment remains constrained by the indexer's expensive $O(L^2)$ scoring overhead and the hardware-inefficient, discontinuous memory-access patterns induced by its outputs. To address these system-level bottlenecks, we introduce LongCat Sparse Attention (LSA), a hardware-algo…
▽ More
DeepSeek Sparse Attention (DSA) enables efficient long-context modeling through its Lightning Indexer. However, practical deployment remains constrained by the indexer's expensive $O(L^2)$ scoring overhead and the hardware-inefficient, discontinuous memory-access patterns induced by its outputs. To address these system-level bottlenecks, we introduce LongCat Sparse Attention (LSA), a hardware-algorithm co-designed framework comprising three complementary and orthogonal strategies: (1) Streaming-Aware Indexing, which selectively converts scattered KV entries into hardware-aligned contiguous layouts to enable coalesced HBM access; (2) Cross-Layer Indexing, which amortizes indexing overhead by reusing the results produced by a single layer across consecutive layers, supported by cross-layer distillation; and (3) Hierarchical Indexing, which adopts a coarse-to-fine scoring scheme to progressively narrow the candidate set for each query, thereby substantially reducing indexing computation. Extensive scaling experiments, ranging from 69B-A3B to 560B-A27B models, demonstrate that LSA consistently achieves performance on par with full attention across both general-purpose and long-context benchmarks. Moreover, LSA supports native training with context lengths of up to one million tokens and underpins the development of LongCat-2.0 (1.6T-A48B). To facilitate further research, we also introduce and open-source LongCat-Flash-Lite-Sparse (69B-A3B), which integrates LSA into LongCat-Flash-Lite and incorporates an updated long-context training corpus.
△ Less
Submitted 4 August, 2026; v1 submitted 2 August, 2026;
originally announced August 2026.
-
Representation Transfer of Foundation Models for Ultra-Widefield Retinal Imaging
Authors:
Mingya Alexa Gong,
Da Ma,
Lovre Antonio Budimir,
Ivana Matovinovic,
Sven Loncaric,
Myeong Jin Ju,
Yukun Zhou,
Siegfried K. Wagner,
Pearse A. Keane,
Marinko V. Sarunic
Abstract:
Despite the widespread adoption of foundation models as feature extractors for medical imaging, relatively little is understood about how different pretraining strategies influence the transferability of learned representations to weakly supervised ophthalmic imaging tasks. We investigate this question in ultra-widefield (UWF) retinal imaging by evaluating foundation model representations within a…
▽ More
Despite the widespread adoption of foundation models as feature extractors for medical imaging, relatively little is understood about how different pretraining strategies influence the transferability of learned representations to weakly supervised ophthalmic imaging tasks. We investigate this question in ultra-widefield (UWF) retinal imaging by evaluating foundation model representations within a patch-based multiple instance learning (MIL) framework for disease classification on UWF images. We compare Vision Transformer encoders pretrained with supervised, Masked Autoencoder (MAE), and self-distillation objectives, while keeping the downstream aggregation architecture unchanged. Within a controlled comparison of ViT-B encoders pretrained on ImageNet-1k, the choice of pretraining objective substantially influenced frozen representation transfer, with supervised and self-distillation-based models outperforming MAE. A contemporary DINOv3 model pretrained at a larger scale achieved the strongest overall performance, with a quadratic weighted kappa of 0.863 for five-class diabetic retinopathy grading, comparable with DINOv1. Attention analysis further revealed distinct patch-aggregation behaviours associated with the different pretrained representations, while partial fine-tuning substantially reduced the performance gap for MAE. These findings suggest that pretraining strategy influences both representation transferability and the subsequent aggregation of patch-level evidence within MIL, resulting in differences in downstream classification performance.
△ Less
Submitted 1 August, 2026;
originally announced August 2026.
-
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
Authors:
Yifan Ding,
Xincheng Wei,
Yoshua Y. Li,
Ziheng Li,
Yuquan Lu,
Siyu Zhang,
Dongsheng Ma,
Rongxiang Weng,
Xunliang Cai,
Yun Chen
Abstract:
Reinforcement learning with verifiable rewards (RLVR) broadcasts a single response-level reward to every token, while on-policy distillation (OPD) scores each token against a stronger teacher for a dense advantage but caps performance at teacher quality and discourages exploration beyond it. Their complementarity makes combining RLVR and OPD promising, but we find that fusing the two advantages wi…
▽ More
Reinforcement learning with verifiable rewards (RLVR) broadcasts a single response-level reward to every token, while on-policy distillation (OPD) scores each token against a stronger teacher for a dense advantage but caps performance at teacher quality and discourages exploration beyond it. Their complementarity makes combining RLVR and OPD promising, but we find that fusing the two advantages with a fixed coefficient triggers entropy collapse from two miscalibrations: a magnitude mismatch, where token-level OPD advantages can spike far beyond the bounded RLVR advantage and erase its signal, and a temporal mismatch, where sustained full-strength OPD keeps pulling the student toward the teacher and limits exploration needed to surpass it. We propose SAF, a Stable Advantage Fusion framework that resolves both issues via a lightweight, four-stage pipeline applied only to the OPD advantage: a sparsify-then-compress mechanism for magnitude control paired with a warm-up-then-anneal mechanism for temporal control, with each stage independently switchable and adding negligible overhead. Instantiating RLVR with GRPO, we evaluate SAF across seven mathematical reasoning and code generation benchmarks with Qwen3-1.7B/4B/8B: SAF avoids entropy collapse and consistently outperforms fixed-coefficient GRPO+OPD fusion, improving the aggregate score by 0.51-2.70% across all six model-domain settings while achieving more stable training.
△ Less
Submitted 31 July, 2026;
originally announced July 2026.
-
Charge-Density-Wave Phase Selection by Janus-Induced Intrinsic Strain in Monolayer NbSSiAs$_2$
Authors:
Chun-Jie Zhang,
Bing Zhang,
Dongliang Mao,
Yapeng Wu,
Xiao-Ping Li,
Lei Wang
Abstract:
Controlling phase selection among competing charge-density-wave (CDW) instabilities remains challenging in two-dimensional materials. Here, first-principles calculations show that Janus-induced intrinsic tensile strain redirects the off-M soft-mode tendency of NbS$_2$ to the M point in NbSSiAs$_2$, selecting a $2\times2$ CDW reconstruction. Electron-phonon coupling analysis identifies momentum-sel…
▽ More
Controlling phase selection among competing charge-density-wave (CDW) instabilities remains challenging in two-dimensional materials. Here, first-principles calculations show that Janus-induced intrinsic tensile strain redirects the off-M soft-mode tendency of NbS$_2$ to the M point in NbSSiAs$_2$, selecting a $2\times2$ CDW reconstruction. Electron-phonon coupling analysis identifies momentum-selective coupling between Nb-derived states and a longitudinal acoustic mode as the origin of the M-point instability. The reconstructed phase hosts two nearly degenerate Nb-trimerized configurations whose relative stability is tuned by biaxial strain. Both configurations retain phonon-mediated superconductivity on the 6-7 K scale, indicating the coexistence of CDW order and superconductivity. Compressive strain favors the 1+3-hollow configuration and induces a band-inverted, $Z_2$-nontrivial state while preserving superconductivity. Together, these results identify Janus-induced intrinsic strain as an internal structural route for CDW phase selection, whereas external strain provides access to a regime in which CDW order, topology, and superconductivity coexist.
△ Less
Submitted 31 July, 2026;
originally announced July 2026.
-
DLAM: Distributional Latent Actions with Temporal Constraints
Authors:
Zuojin Tang,
Feifan Luo,
Haoyun Liu,
Botai Yuan,
Dekang Qi,
Ronghan Chen,
Yandan Yang,
Tong Lin,
Xinyuan Chang,
Mu Xu,
Bin Liu,
De Ma,
Zhiheng Ma
Abstract:
Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant observations of physical change. Latent action models can extract such priors, but reconstruction-trained codes may predict future observations without the structure required for joint generation with robot actions. Existing structured methods add temporal constrain…
▽ More
Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant observations of physical change. Latent action models can extract such priors, but reconstruction-trained codes may predict future observations without the structure required for joint generation with robot actions. Existing structured methods add temporal constraints but retain deterministic transition points, so residual errors in locally inferred transitions may propagate and compound under recursive composition. We introduce DLAM, a distributional latent-action model that represents each transition as a diagonal Gaussian. Reconstruction conditioned on the reference frame grounds the mean in observed visual change, while normalized composition and reversal over equal-gap triplets constrain both the mean and dimension-wise variance. Variance composition uses a lightweight shared-correlation coefficient to account for dependence between adjacent transitions that share an intermediate frame, whereas reversal negates the mean and preserves the variance. For downstream policy learning, we freeze the encoder and train a flow-matching policy to jointly generate mean transition sequences and robot actions. On held-out transitions, DLAM learns more temporally consistent latent dynamics than existing latent-action baselines and achieves stronger direct and cumulative reconstruction on held-out videos. Under the same controlled $π_0$ transfer protocol, it also improves policy performance on MetaWorld MT50, LIBERO, and real-world manipulation tasks. Controlled ablations show that normalized mean constraints account for most of the reconstruction gain, while learned variance and correlation-aware composition provide complementary improvements in downstream control.
△ Less
Submitted 29 July, 2026;
originally announced July 2026.
-
Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation
Authors:
Xiangbo Gao,
Siyuan Yang,
Ping He,
Mingyang Wu,
Yuheng Wu,
Yushen Zuo,
Jiongze Yu,
Ryan Cui,
Hongyuan Hua,
Devin Ma,
Xiao Jin,
Yubo Yuan,
Qing Yin,
Jie Yang,
Zhengzhong Tu
Abstract:
We present Visko Orbis 1.0, a Live Model for real-time, interactive long-video generation. Users can change the prompt at any moment during generation, and the update becomes visible in real time. Visko Orbis 1.0 supports long-form text-to-video, image-to-video, and video continuation, with multilingual prompts and prompt switching while generation is in progress. A bounded multi-scale memory pres…
▽ More
We present Visko Orbis 1.0, a Live Model for real-time, interactive long-video generation. Users can change the prompt at any moment during generation, and the update becomes visible in real time. Visko Orbis 1.0 supports long-form text-to-video, image-to-video, and video continuation, with multilingual prompts and prompt switching while generation is in progress. A bounded multi-scale memory preserves subjects, scenes, and style across chunks, sustaining hour-scale rollouts without evident quality or color drift. Built on a distilled chunk-wise streaming generator and a streaming video upscaler, Visko Orbis 1.0 delivers real-time 4K video generation at 24 FPS using an optimized GPU serving engine. In long-form Arena comparisons, Visko Orbis 1.0 obtains the highest overall-preference and temporal-stability ratings among state-of-the-art real-time interactive video-generation systems.
△ Less
Submitted 17 August, 2026; v1 submitted 29 July, 2026;
originally announced July 2026.
-
Beyond GDPR: Examining Disclosure Gaps in Mobile AR Privacy Policies under U.S. State Privacy Laws
Authors:
Hong Chen,
Xueling Zhang,
Hong-Ning Dai,
Huashan Chen,
Qin Yu,
Tiange Xie,
Duohe Ma,
Feng Liu
Abstract:
Mobile Augmented Reality (MAR) apps can collect and process highly sensitive data such as spatial maps and biometrics, yet their privacy policies remain largely understudied. Prior audits of app privacy policies have typically focused on a single legal framework, such as the GDPR. Meanwhile, 20 U.S. states have comprehensive privacy laws in effect, creating a fragmented and rapidly evolving set of…
▽ More
Mobile Augmented Reality (MAR) apps can collect and process highly sensitive data such as spatial maps and biometrics, yet their privacy policies remain largely understudied. Prior audits of app privacy policies have typically focused on a single legal framework, such as the GDPR. Meanwhile, 20 U.S. states have comprehensive privacy laws in effect, creating a fragmented and rapidly evolving set of privacy policy obligations. To date, no study has systematically audited privacy policies against this emerging body of state-level legislation.
In this paper, we present the first large-scale audit of MAR privacy policies under U.S. state privacy laws. We construct a dataset covering the MAR ecosystem, including 8,013 Google Play MAR app metadata records worldwide, and a U.S.-based subset with 6,620 APKs and 6,426 privacy policy files. We further derive an auditable disclosure taxonomy with 5 baseline requirements, 10 triggered requirements, and 4 logic chains, and build a validated four-stage automated pipeline that produces traceable, evidence-grounded disclosure judgments.
Our audit reveals widespread disclosure gaps: 44.62\% of audited policies exhibit severe disclosure omissions, with each missing more than eight requirements, and four privacy-policy requirements have violation rates above 90\%. These findings suggest that MAR privacy disclosures are not keeping pace with the growing complexity of U.S. state privacy regulation. We release our dataset, taxonomy, and auditing pipeline to support future research on scalable privacy compliance auditing.
△ Less
Submitted 27 July, 2026;
originally announced July 2026.
-
An on-chip programmable mechano-quantum transducer
Authors:
Xinrui Zhang,
Wei Liu,
Duanyu Ma,
Lin-Ke Xie,
Nai-Jie Guo,
Zhongtao Gou,
Yifan Wang,
Jianxin Xu,
Xiaoguang Luo,
Zhao Mu,
Honglong Chang,
Weizheng Yuan,
Jian-Shun Tang,
Chuan-Feng Li,
Guangcan Guo,
Tao Ye
Abstract:
Solid-state spin defects encode local perturbations as measurable shifts in spin-transition frequencies, but mechanical actuation and quantum readout remain physically separated, resulting in a discrete measurement setup. Integrating these functions requires an on-site mechano-quantum interface that programs the lattice state of a defect host and quantitatively maps it onto the spin Hamiltonian. H…
▽ More
Solid-state spin defects encode local perturbations as measurable shifts in spin-transition frequencies, but mechanical actuation and quantum readout remain physically separated, resulting in a discrete measurement setup. Integrating these functions requires an on-site mechano-quantum interface that programs the lattice state of a defect host and quantitatively maps it onto the spin Hamiltonian. Here we first report an on-chip programmable mechano-quantum transducer (OCPMQT) that integrates voltage-defined micromechanical actuation with in situ spin-frequency readout in a two-dimensional van der Waals quantum-defect host. Mechanically programmed lattice states are encoded as shifts in the axial zero-field splitting parameter and resolved by optically detected magnetic resonance (ODMR) spectroscopy. Within a chip volume of 2.05*10^-2 cm^3, the transducer accesses ODMR-inferred strains as low as 0.0080% and delivers a volumetric force density of approximately 2.6*10^4 N*m^-3. A micromechanical-to-spin-Hamiltonian framework links on-chip electromechanics, interfacial strain transfer, and strain-spin coupling, enabling the electrical control micromechanical input to be measured directly as spin-frequency response.
△ Less
Submitted 26 July, 2026; v1 submitted 23 July, 2026;
originally announced July 2026.
-
Workflow-Localized Mechanism Learning: Attribution-Guided Repair and Knowledge Reuse for Structured Agent Skills
Authors:
Zibin Lin,
Shengli Zhang,
Taotao Wang,
Yihan Xia,
Deen Ma,
Guofu Liao
Abstract:
Agent Skills package reusable procedural knowledge as external artifacts for frozen language-model agents, yet existing optimizers do not jointly resolve where a failure occurs in a workflow, which mechanism caused it, and how relevant knowledge from third-party Skills should be reused locally. We introduce Workflow-Localized Mechanism Learning (WML). Its Node--Mechanism Attribution identifies the…
▽ More
Agent Skills package reusable procedural knowledge as external artifacts for frozen language-model agents, yet existing optimizers do not jointly resolve where a failure occurs in a workflow, which mechanism caused it, and how relevant knowledge from third-party Skills should be reused locally. We introduce Workflow-Localized Mechanism Learning (WML). Its Node--Mechanism Attribution identifies the failed workflow node, implicated mechanisms, and smallest valid edit target, routing single-mechanism defects to L3 resources and relational defects across mechanisms to L2 composition protocols. A six-module Workflow-Guided Skill Optimization (WGSO) loop then selects provenance- and scope-aware third-party knowledge, applies bounded patches, evaluates candidates, and stores verified outcomes in optimizer-side memory. On SpreadsheetBench, WML reaches 90.33 +/- 1.53 and 74.67 +/- 3.51 Hard Accuracy with DeepSeek and Qwen3.6-Flash, respectively; without additional optimization, the learned Skills transfer to WikiTableQuestions with 84.00 +/- 2.00 and 83.00 +/- 2.00 Denotation Accuracy. On Compiler-Supported50, WML attains both the highest hard-PASS rate and the lowest cost per successful task; compiled execution sharply reduces tokens and calls relative to a direct SkillAgent while retaining most of its successful tasks. Code and artifacts are available at https://github.com/xiaolin9595/workflow-localized-mechanism-learning.
△ Less
Submitted 23 July, 2026;
originally announced July 2026.
-
When 2D Cues Fail: Improving Image Manipulation Localization with Reliable 3D Geometry
Authors:
Guofeng Yu,
Zhiqing Guo,
Dan Ma,
Gaobo Yang
Abstract:
Existing image manipulation localization (IML) methods rely heavily on 2D forensic cues, such as low-level artifacts, noise traces, and semantic inconsistencies in the manipulated image. While effective in many cases, these cues become much less discriminative when manipulated regions are well blended with their surrounding context in appearance. In such cases, a manipulated region may remain loca…
▽ More
Existing image manipulation localization (IML) methods rely heavily on 2D forensic cues, such as low-level artifacts, noise traces, and semantic inconsistencies in the manipulated image. While effective in many cases, these cues become much less discriminative when manipulated regions are well blended with their surrounding context in appearance. In such cases, a manipulated region may remain locally appearance-consistent, but still violate the geometric structure of the surrounding scene. This limitation motivates us to go beyond purely 2D evidence and introduce geometric reasoning into IML. To this end, we leverage monocular reconstruction to obtain auxiliary geometric cues, including depth and surface normals. However, a key challenge lies in the fact that reconstructed geometry on manipulated images is inherently noisy and cannot be used naively. Rather than treating depth and normals as direct evidence, we estimate their reliability and exploit them selectively for localization. Based on this principle, we design a geometry-aware framework (GFrame) that fuses reliable geometric cues with RGB features and propagates them across scales to improve fine-grained localization. Extensive experiments show that the proposed method achieves excellent performance under limited budget constraints. These results indicate that reliable 3D geometry provides complementary forensic evidence beyond traditional 2D cues for IML. Related code will be released.
△ Less
Submitted 20 July, 2026;
originally announced July 2026.
-
Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors
Authors:
Yilin Wang,
Xiangxi Zheng,
Dongxing Mao,
Linjie Li,
Zhengyuan Yang,
Ping Yu,
Rui Yan,
Yuan Yao,
Alex Jinpeng Wang
Abstract:
Understanding long videos with multimodal large language models (MLLMs) requires selecting a compact set of frames from thousands of candidates, yet identifying the right frames seemingly requires understanding the video first. We resolve this circular dependency with a simple observation: cross-modal attention at validation-selected extraction layers in MLLMs already provides query-relevant frame…
▽ More
Understanding long videos with multimodal large language models (MLLMs) requires selecting a compact set of frames from thousands of candidates, yet identifying the right frames seemingly requires understanding the video first. We resolve this circular dependency with a simple observation: cross-modal attention at validation-selected extraction layers in MLLMs already provides query-relevant frame evidence without requiring autoregressive generation. We exploit this property to build DAFS (Dynamic Attention-based Budget-aware Frame Selection), a training-free frame selector. A lightweight MLLM selector, even with only 2B parameters, can extract frame-level evidence by converting selected-layer attention into relevance scores through query-conditioned aggregation. This enables cross-frame comparison without autoregressive decoding. To handle the selector's own context constraint, we formulate the joint allocation of candidate pool size and per-frame token budget as a discrete optimization problem solved by dynamic programming. Under a 32-frame budget, our selector improves over uniform sampling by up to 6.4 points on Video-MME and outperforms prior training-based selectors under matched frame budgets, while generalizing across selector and answerer backbones, and across tasks, without retraining.
△ Less
Submitted 17 July, 2026;
originally announced July 2026.
-
Don't Predict, Prioritize: Rethinking GPU Reliability Assessment
Authors:
Difeng Ma,
Changhua Pei,
Yuanwei Lu,
Quan Zhou,
Zexin Wang,
Yibo Zhu,
Daxin Jiang,
Dan Pei,
Jingjing Li,
Gaogang Xie
Abstract:
The reliability of Graphics Processing Units (GPUs) is a criticalbottleneck for modern large-scale AI infrastructure, where a sin-gle node failure can disrupt synchronous training jobs and causesignificant financial losses. While predictive maintenance is widelyused in other hardware domains, we demonstrate that accuratelypredicting the exact timing of GPU failures is inherently difficult.Through…
▽ More
The reliability of Graphics Processing Units (GPUs) is a criticalbottleneck for modern large-scale AI infrastructure, where a sin-gle node failure can disrupt synchronous training jobs and causesignificant financial losses. While predictive maintenance is widelyused in other hardware domains, we demonstrate that accuratelypredicting the exact timing of GPU failures is inherently difficult.Through an in-depth analysis of telemetry data from a productioncluster, we find that major GPU failures, including Double Bit Er-rors (DBEs) and GPU Lost events, exhibit strong stochasticity andlow signal-to-noise ratios in time-series telemetry, which makesconventional time-based prediction ineffective.
This insight motivates a paradigm shift: instead of attempting topredict the absolute timing of a failure, we propose a more robustapproach focused on ranking nodes by their relative failure risk. Wepropose HeaRank (Health Rank), a Learning-to-Rank (LTR) frame-work that leverages stable historical failure patterns to computea global risk ranking of GPU nodes. Evaluated on a production-scale cluster with thousands of GPUs, HeaRank achieves an AUCof 0.83, significantly outperforming both heuristic baselines andstate-of-the-art ranking algorithms. In online deployment, HeaRanksuccessfully captures 64% of future failures within the top 5% ofranked nodes, compared to only 21% by the incumbent productionsystem. These results suggest that relative risk ranking can serveas a robust alternative in environments where absolute failure pre-diction is inherently limited. Our work highlights the importanceof risk-aware scheduling and proactive resource management inmodern GPU clusters.
△ Less
Submitted 16 July, 2026;
originally announced July 2026.
-
Cotton-SF YOLO: Learning Structural and Frequency Cues for Early Cotton Square Detection in Complex Field Environments
Authors:
Chengjia Zhang,
Yu Li,
Feiri Ali,
Yan Zhang,
Xin Chen,
Longke He,
Daokun Ma,
Liting Gao
Abstract:
Cotton squares are important phenotypic indicators of the early reproductive growth of cotton, and automatic field detection of cotton squares provides an important basis for cotton growth monitoring and precision cultivation management. However, early cotton square detection in complex field environments remains insufficiently explored, as cotton squares are small, frequently occluded, easily blu…
▽ More
Cotton squares are important phenotypic indicators of the early reproductive growth of cotton, and automatic field detection of cotton squares provides an important basis for cotton growth monitoring and precision cultivation management. However, early cotton square detection in complex field environments remains insufficiently explored, as cotton squares are small, frequently occluded, easily blurred, subject to illumination variations, and exhibit low contrast against surrounding cotton leaves. To address these challenges, we propose a task-oriented framework based on YOLO26m, named Cotton-SF YOLO, for cotton square detection under natural field conditions. To improve the perception of small and irregular cotton square boundaries, we introduce Dynamic Snake Convolution into the detector, enabling adaptive extraction of deformable edge features. Furthermore, a frequency-domain feature modulation module is designed by incorporating spectral enhancement into the C2f structure, which recalibrate frequency-domain representations and strengthen discriminative edge and texture cues while reducing interference from complex cotton leaf backgrounds. Trained and evaluated on our newly constructed and annotated field dataset with manually annotated cotton squares, the proposed model achieves mAP$_{50}$, mAP$_{50:95}$, and recall values of 0.8196, 0.4942, and 0.7939, improving over the baseline YOLO26m by 1.25%, 3.45%, and 2.96%, respectively. Ablation experiments and visualization demonstrate that the best performance is achieved with the complementary effects of structural and frequency cues.
△ Less
Submitted 15 July, 2026;
originally announced July 2026.
-
Record Loss Sets a Rare-Trajectory Limit on Quantum Purification
Authors:
Jiaxin Liu,
Zuoxian Wang,
Feng Li,
Danyue Ma
Abstract:
Continuous quantum feedback uses time-resolved measurement records to steer monitored systems toward pure states. Yet how the information available to a controller determines the ultimate purification speed remains unresolved. We establish this relation for a qubit under fixed-spectrum Hermitian monitoring with detector loss, obtaining the exact long-time impurity-moment spectrum optimized over ca…
▽ More
Continuous quantum feedback uses time-resolved measurement records to steer monitored systems toward pure states. Yet how the information available to a controller determines the ultimate purification speed remains unresolved. We establish this relation for a qubit under fixed-spectrum Hermitian monitoring with detector loss, obtaining the exact long-time impurity-moment spectrum optimized over causal basis controls at each horizon. Rare records with nearly canceled evidence then make all moments from half order upward decay at the Bhattacharyya information rate between two quantum nondemolition record laws. Aligned quantum nondemolition monitoring preserves that binary distinguishability and attains the limit, while complete detection restores an order-dependent branch. The mechanism extends to higher dimensions, where an attainable rank-two ceiling lies above the full-rank qutrit upper bound over a finite moment interval, establishing retained record distinguishability as a purification resource.
△ Less
Submitted 29 July, 2026; v1 submitted 10 July, 2026;
originally announced July 2026.
-
Low-latency FPGA-based electronic control system for fast preparation of defect-free atom arrays
Authors:
Ya-Dong Hu,
Dong-Qi Ma,
Tian-Yang Zhang,
Liang Chen,
Yi-Chen Zhang,
Xiao-Kang Zhong,
Wen-Yi Zhu,
Hong-Jie Fan,
Qing-Xuan Jie,
Yan-Lei Zhang,
Gang Li,
Xi-Feng Ren,
Xu-Liang Zhang,
Guang-Can Guo,
Zhu-Bo Wang,
Chang-Ling Zou
Abstract:
The scalability of neutral atom quantum computing demands integrated electronic control systems with low latency, modular architecture, and real-time feedback capability. Here, we present an FPGA-based electronic control system that eliminates the PC from the feedback loop, integrating photon counting, real-time decision-making, and waveform generation within a unified PXIe architecture. The syste…
▽ More
The scalability of neutral atom quantum computing demands integrated electronic control systems with low latency, modular architecture, and real-time feedback capability. Here, we present an FPGA-based electronic control system that eliminates the PC from the feedback loop, integrating photon counting, real-time decision-making, and waveform generation within a unified PXIe architecture. The system achieves a total feedback latency of $282\,\mathrm{μs}$ and is validated in practical experiments by assembling defect-free atom arrays from 24 stochastically loaded optical tweezers. A single-round rearrangement achieves a filling fraction of $\sim96\%$, while feedback-controlled iterative rearrangement over five rounds boosts the success probability for generating a 10-atom defect-free array from $65.7\%$ to $95.4\%$. This system establishes the electronic infrastructure necessary for mid-circuit measurement and real-time quantum error correction on neutral-atom platforms.
△ Less
Submitted 9 July, 2026;
originally announced July 2026.
-
Efficiency-Induced Freezing in Quantum-State Purification
Authors:
Jiaxin Liu,
Zuoxian Wang,
Feng Li,
Danyue Ma
Abstract:
Any nonzero detection loss qualitatively changes feedback-controlled purification under diffusive monitoring. In every finite dimension, we prove a sharp, dimension-independent ceiling on the decay of trajectory-averaged impurity moments, uniformly over admissible predictable feedback protocols.Below unit efficiency, this ceiling becomes independent of moment order above a critical value and is at…
▽ More
Any nonzero detection loss qualitatively changes feedback-controlled purification under diffusive monitoring. In every finite dimension, we prove a sharp, dimension-independent ceiling on the decay of trajectory-averaged impurity moments, uniformly over admissible predictable feedback protocols.Below unit efficiency, this ceiling becomes independent of moment order above a critical value and is attained on extremal rank-two quantum-nondemolition (QND) faces. For generic observable spectra, a determinant-root law precludes every full-rank state from attaining this boundary rate over an explicit moment-order interval. For qubits at $0<η<1$, the frozen rate is the exact optimum, set by rare, persistently mixed trajectories. Parameter-free finite-action scaling functions resolve both the rounded QND moment-order transition and the near-unit QND--always-unbiased crossover.
△ Less
Submitted 22 July, 2026; v1 submitted 8 July, 2026;
originally announced July 2026.
-
Governing Generative AI Across Financial Institutions: A Framework for Generative AI Risk Control
Authors:
Dennis Mao,
Alessandra Lin,
Yixin Kang,
Yiqing Wang
Abstract:
Generative artificial intelligence is moving from general-purpose experimentation toward specialized applications across banking, capital markets, insurance, payments, and wealth management. Its main contribution is not limited to conversational interfaces. Modern generative systems can synthesize large document collections, extract information from unstructured data, generate software and analyti…
▽ More
Generative artificial intelligence is moving from general-purpose experimentation toward specialized applications across banking, capital markets, insurance, payments, and wealth management. Its main contribution is not limited to conversational interfaces. Modern generative systems can synthesize large document collections, extract information from unstructured data, generate software and analytical code, create scenario narratives, support research workflows, and coordinate multi-step tasks. These capabilities make generative AI especially relevant to finance, where decisions often depend on combining quantitative data with contracts, policies,filings, news, customer communications, and expert judgment. This paper presents an application-oriented view of generative AI in finance. It organizes potential uses around five capability patterns, including knowledge synthesis, content generation, analytical assistance, interaction, and workflow orchestration, and maps them to major financia functions. Representative applications include investment research, customer service, lending support, fraud investigation, financial reporting, operations automation, software development, and personalized financial guidance. The paper also discusses common technical architectures, such as retrieval-augmented generation, tool-using assistants, multimodal models, and agentic workflows, and identifies practical factors that shape business value. The resulting landscape provides a foundation for researchers and practitioners seeking to understand where generative AI may produce the greatest operational and analytical impact in financial services
△ Less
Submitted 15 July, 2026; v1 submitted 4 July, 2026;
originally announced July 2026.
-
iVISION-2DCD: A Long-Term Change Detection Dataset for Large-Scale Outdoor Construction Monitoring
Authors:
Dayou Mao,
Yuchen Lin,
Ashkan Ebadi,
John Zelek,
Alexander Wong,
Yuhao Chen
Abstract:
Automation in construction is essential for reducing costs and human errors in large-scale projects. We approach the construction progress monitoring from the aspect of detecting changes in construction sites. As construction buildings continue to evolve in geometry and appearance over time, change detection need to be performed from arbitrary camera viewpoints. This necessitates developing 2D Cha…
▽ More
Automation in construction is essential for reducing costs and human errors in large-scale projects. We approach the construction progress monitoring from the aspect of detecting changes in construction sites. As construction buildings continue to evolve in geometry and appearance over time, change detection need to be performed from arbitrary camera viewpoints. This necessitates developing 2D Change Detection (2DCD) algorithms that operate robustly across diverse camera perspectives at construction sites. While developing and evaluating such systems is data-intensive, no open-source benchmark dataset exists at the intersection of 2D change detection and construction automation research. Data collection using Unmanned Aerial Vehicles (UAVs) is gaining its popularity in outdoor large-scale surveying. However, in active construction sites conducting drone missions equipped with high-end sensors imposes safety concerns. Flight trajectory and collected camera viewpoints can be significantly limited. To address this critical gap, we introduce iVISION-2DCD, a large-scale synthetically generated dataset from dense LiDAR point clouds with photorealistic input images and accurate ground truth annotations. Our dataset formally defines the problem of viewpoint-robust 2DCD at construction sites and captures the inherent complexities of real-world deployment. In this paper, we present our systematic methodology for synthetic data generation, developing novel view synthesis techniques to overcome bi-temporal alignment and viewpoint diversity challenges, and implementing semi-automated semantic segmentation with change label generation while preserving challenging real-world cases. Benchmark evaluations using state-of-the-art 2DCD algorithms demonstrate that iVISION-2DCD poses novel research challenges for the computer vision and robotics communities.
△ Less
Submitted 3 July, 2026;
originally announced July 2026.
-
Higher-order noise statistics restore Heisenberg scaling under collective dephasing
Authors:
Jiaxin Liu,
Xing Heng,
Zuoxian Wang,
Danyue Ma
Abstract:
Noisy-metrology theory characterizes decoherence by its two-point correlation function, equivalently the single-atom coherence time or noise spectrum. We show this is insufficient for entangled probes: two collective baths with identical single-atom $T_2$ but different higher-order statistics yield opposite entanglement-enhanced scaling. Under Gaussian Markovian collective dephasing a Greenberger-…
▽ More
Noisy-metrology theory characterizes decoherence by its two-point correlation function, equivalently the single-atom coherence time or noise spectrum. We show this is insufficient for entangled probes: two collective baths with identical single-atom $T_2$ but different higher-order statistics yield opposite entanglement-enhanced scaling. Under Gaussian Markovian collective dephasing a Greenberger--Horne--Zeilinger (GHZ) probe reaches an atom-number-independent sensitivity floor. For a fully Markovian compound-Poisson bath, in which collective dephasing is generated by a finite-rate sequence of unitary phase kicks, a Dicke coherence of order $q$ (a difference of $J_z$ eigenvalues) decays at $Γ_q=Γ[1-\mathrm{Re}\,\varphi(q)]$, with $\varphi$ the kick characteristic function; for any absolutely continuous kick law this rate saturates at large $q$ instead of growing as $q^2$, and a GHZ probe recovers Heisenberg scaling $δω\propto1/N$ over the window in which collective finite-rate noise dominates residual independent decoherence. We prove that the Gaussian floor is the exact worst case: at fixed single-atom coherence time every finite-rate kick statistics strictly beats it, and for arbitrary Lévy phase noise the asymptotic entangled-probe sensitivity is set exclusively by the diffusive component. A converse bound shows that no input state, ancilla, or measurement improves on the GHZ scaling. The mechanism is purely exponential and CP-divisible, distinct from the Zeno, non-Markovian, nonlinear-generator, and error-correction routes. A dissipative analogue caps the Dicke superradiant burst. The full counting statistics of common noise thus emerge as a control axis for noisy quantum metrology, beyond the spectrum.
△ Less
Submitted 3 July, 2026;
originally announced July 2026.
-
DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation
Authors:
Siyu Yan,
Yizhen Gao,
Yilin Wang,
Dongxing Mao,
Alex Jinpeng Wang
Abstract:
Text-rich image generation is one of the most challenging settings in image generation, since models must simultaneously produce visually realistic images and render legible, semantically aligned, and layout-consistent text. Existing data pipelines usually follow a static crawl-filter-freeze paradigm. They collect candidate samples, filter them once, and freeze the accepted data for training. Howe…
▽ More
Text-rich image generation is one of the most challenging settings in image generation, since models must simultaneously produce visually realistic images and render legible, semantically aligned, and layout-consistent text. Existing data pipelines usually follow a static crawl-filter-freeze paradigm. They collect candidate samples, filter them once, and freeze the accepted data for training. However, rejected samples are usually discarded, although they often contain useful failure signals such as OCR errors and semantic mismatches. As a result, later construction rounds may repeat the same failure modes. To address these limitations, we propose DataEvolver, a self-evolving multi-agent framework for text-rich image data construction. DataEvolver treats data construction as feedback-driven construction policy evolution. A Retriever collects candidate samples, a Verifier assigns quality scores and rejection causes, a Critic summarizes round-level feedback into semantic feedback, and a Generator completes under-covered regions through targeted synthesis. The updated feedback memory then guides the next construction round. Experiments on text-rich image generation benchmarks show that DataEvolver produces more useful training data than fixed-dataset baselines under matched data budgets. At the 0.75M scale on PixArt-alpha, DataEvolver improves OCR-F1 over the strongest baseline by 85.3 percent on TextScenesHQ and 35.3 percent on LongTextBench. The improvements are consistent across both evaluated benchmarks and also transfer to Show-o2, indicating that the benefit of DataEvolver is not tied to a single downstream generator. These results suggest that rejected samples can provide actionable feedback for improving text-rich image data construction.
△ Less
Submitted 30 June, 2026;
originally announced June 2026.
-
High-Resolution Flood Mapping With Sentinel-1 and Sentinel-2 via Misalignment-Robust Cross-Sensor Learning and Generative Despeckling
Authors:
David Ma,
Jeremy Feinstein,
Shreya Pandit,
Arkaprabha Ganguli,
Eugene Yan
Abstract:
Reliable high-resolution flood extent mapping from satellite imagery remains constrained by limited data fidelity and sensor-specific artifacts. Multispectral optical imagery is degraded by clouds, shadows, and urban confounders, while synthetic aperture radar (SAR) imagery is affected by speckle noise and sensor co-registration uncertainty. This work presents an integrated flood mapping framework…
▽ More
Reliable high-resolution flood extent mapping from satellite imagery remains constrained by limited data fidelity and sensor-specific artifacts. Multispectral optical imagery is degraded by clouds, shadows, and urban confounders, while synthetic aperture radar (SAR) imagery is affected by speckle noise and sensor co-registration uncertainty. This work presents an integrated flood mapping framework that jointly addresses these limitations through curated datasets and novel learning strategies. We introduce a new Sentinel-2 (S2) and Sentinel-1 (S1) dataset covering the contiguous United States, featuring pixel-accurate 10 m water masks with emphasis on challenging weather conditions and urban environments that are underrepresented in existing benchmarks. High-quality S2 annotations are manually produced using rigorous geospatial labeling protocols and transferred to SAR imagery through weakly labeled temporally coincident acquisitions. To address SAR-specific artifacts, a shift-invariant loss function is employed to tolerate residual geolocation uncertainty between SAR imagery and optical-derived labels, and a Conditional Variational Autoencoder (CVAE) is trained on multitemporal SAR composites to suppress speckle while preserving flood-relevant spatial structure. Experiments using UNet and UNet++ architectures demonstrate strong multispectral performance (AUPRC up to 0.956) and statistically significant improvements in SAR flood mapping when using shift-invariant loss and CVAE-based despeckling compared to classical filters. These results underscore the importance of dataset fidelity, misalignment-robust training, and demonstrate the viability of generative despeckling for operational flood mapping.
△ Less
Submitted 29 June, 2026;
originally announced June 2026.
-
Concept-Constrained Prompt Learning for Few-Shot CLIP Adaptation
Authors:
Na Sang,
Ding Ma,
Rui Sang,
Yuxuan Liu
Abstract:
Few-shot prompt learning is an effective strategy for adapting CLIP to downstream tasks, but class-only prompt optimization can overfit base-class supervision and weaken transfer to unseen classes. We propose Concept-Constrained Prompt Learning (CCPL), a lightweight regularization framework that anchors learnable class prompts to frozen concept-level text prototypes without updating CLIP encoders.…
▽ More
Few-shot prompt learning is an effective strategy for adapting CLIP to downstream tasks, but class-only prompt optimization can overfit base-class supervision and weaken transfer to unseen classes. We propose Concept-Constrained Prompt Learning (CCPL), a lightweight regularization framework that anchors learnable class prompts to frozen concept-level text prototypes without updating CLIP encoders. CCPL learns a set of shared context tokens, instantiates class prompts by appending class names, and constructs frozen concept prototypes from a class-level concept bank. During training, a text-space cosine consistency objective aligns learnable class-prompt embeddings with frozen concept prototypes; concept dropout provides additional regularization against over-reliance on fixed concept lists. At inference, CCPL optionally fuses class-prompt logits with concept-prototype logits using a controllable ensemble weight alpha. Our default configuration uses text-space concept regularization lambda = 0.5, concept dropout p = 0.3 and weak concept-guided fusion (alpha = 0.1), with no KL-based prediction consistency term. Experiments under identical automatically-generated fallback splits show that CCPL improves the base-to-new harmonic mean on DTD (+0.6) and EuroSAT (+2.9) compared with CoOp, while remaining near-neutral on OxfordPets (-0.1). Ablations indicate that text-space concept regularization is consistently beneficial, while the best concept-guided inference strength is dataset- and protocol-sensitive. These results suggest concept constraints are most effective when concept prototypes align naturally with dataset semantics, and identify fine-grained categories as a current boundary condition. The code is released at: https://github.com/richael-sang/concept-constrained-prompt-learning.
△ Less
Submitted 21 June, 2026;
originally announced June 2026.
-
Three-dimensional Foliated Fractional Quantum Hall Phases
Authors:
Sahana Das,
Navketan Batra,
Andrea Kouta Dagnino,
Dan Mao,
Nicolas Regnault,
Glenn Wagner,
Titus Neupert
Abstract:
Foliated topological orders in three dimensions are layered systems in which anyons are free to move within a layer but cannot hop between them. A simple model with such a phase is a stack of decoupled two-dimensional electron gases in a strong magnetic field, each in the same fractional quantum Hall state. By focusing on the case of filling $ν=1/3$ of the lowest Landau level in each layer, we sho…
▽ More
Foliated topological orders in three dimensions are layered systems in which anyons are free to move within a layer but cannot hop between them. A simple model with such a phase is a stack of decoupled two-dimensional electron gases in a strong magnetic field, each in the same fractional quantum Hall state. By focusing on the case of filling $ν=1/3$ of the lowest Landau level in each layer, we show that (i) the limit of decoupled Laughlin states is stable upon introducing interlayer interactions and (ii) the system can enter a spontaneously layer-trimerized foliated non-Abelian Fibonacci phase. We support our claims by numerical exact diagonalization of up to 10 layers as well as perturbative analytical calculations. Specifically, we show that the foliated Fibonacci phase exists in the 9-layer system with pseudopotential interactions within and between neighboring layers. We identify the phase via quasihole counting and by calculating the overlap with a model wave function which we derive from the associated conformal field theory. Our numerical results suggest the possibility of realizing these phases in layered van der Waals crystals in strong magnetic fields, as well as in multilayer heterostructures.
△ Less
Submitted 17 June, 2026;
originally announced June 2026.
-
Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness
Authors:
Zijian Wang,
Hanqi Li,
Ziyue Yang,
Zijian Hu,
Shenghan Zuo,
Yunzhe Zhang,
Da Ma,
Danyu Luo,
Chenrun Wang,
Jing Peng,
Tiancheng Huang,
Sijia Guo,
Huayang Wang,
Zichen Zhu,
Senyu Han,
Yilu Cao,
Bo Chen,
Xin Chen,
Kai Yu,
Lu Chen
Abstract:
AI systems can increasingly automate scientific workflows, but the reasoning that links prior evidence, generated ideas, experiments and final claims often remains implicit inside model inference. Here we introduce Xcientist, a research harness that externalizes research synthesis and experimental validation into inspectable, contract-governed processes. Xcientist organizes literature evidence, id…
▽ More
AI systems can increasingly automate scientific workflows, but the reasoning that links prior evidence, generated ideas, experiments and final claims often remains implicit inside model inference. Here we introduce Xcientist, a research harness that externalizes research synthesis and experimental validation into inspectable, contract-governed processes. Xcientist organizes literature evidence, idea states, implementation plans, ablation records and repair traces as persistent research artifacts, so that generated mechanisms can be grounded, executed, tested and revised without losing their evidential basis. We identify claim drift as a failure mode of automated research, where runnable artifacts no longer support the mechanism originally claimed. Across training-free memory systems, graph-structured traffic forecasting and multi-scale physics-informed neural networks, Xcientist preserves traceable trajectories from problem formulation to mechanism design, validation and bounded revision. These results suggest that AI scientists should be evaluated not only by their final artifacts, but by whether their synthesis and validation processes remain attributable, inspectable and scientifically accountable.
△ Less
Submitted 20 July, 2026; v1 submitted 17 June, 2026;
originally announced June 2026.
-
Distributed General-Purpose Agent Networks: Architecture, Key Mechanisms, and Prototypes
Authors:
Shengli Zhang,
Deen Ma,
Zibin Lin,
Taotao Wang
Abstract:
Large language models have accelerated the transition from passive conversational assistants to autonomous agents that can understand goals, plan actions, invoke tools, and execute multi-step tasks. Yet the capability of a single agent remains constrained by its local data, tool permissions, runtime environment, and governance boundary. This paper studies distributed general-purpose agent networks…
▽ More
Large language models have accelerated the transition from passive conversational assistants to autonomous agents that can understand goals, plan actions, invoke tools, and execute multi-step tasks. Yet the capability of a single agent remains constrained by its local data, tool permissions, runtime environment, and governance boundary. This paper studies distributed general-purpose agent networks: open peer-to-peer networks in which heterogeneous agents deployed on personal devices, edge nodes, or autonomous computing environments can discover one another, establish trust, negotiate cooperation rules, and execute open-ended tasks. We argue that such networks cannot be obtained by simply combining existing peer-to-peer overlays with conventional multi-agent systems. Unlike traditional P2P networks, agent networks must propagate semantic declarations about intentions, capabilities, states, and cooperation constraints. We therefore propose a layered architecture centered on a protocol adaptation layer that connects upper-level task semantics with lower-level network operations. Based on this architecture, the paper identifies three core mechanism problems: semantic announcement propagation for collaborator discovery, verifiable identity and multi-topic reputation for cooperation governance, and semantic-gradient mechanism design for open task execution. For each problem, we present a technical route, including bodyless gossip with sequential logs, BAID-based identity binding with MG-EigenTrust reputation, and a Stackelberg-style mechanism-generation loop driven by semantic attribution feedback. We further report prototype overhead results for BAID-style tiered verification and mechanism-level simulations of MG-EigenTrust under cross-topic disguise-collusion attacks. The resulting framework provides a system-level foundation for open, trustworthy, and scalable agent collaboration.
△ Less
Submitted 15 June, 2026;
originally announced June 2026.
-
Bayesian joint modelling using semiparametric accelerated failure time approaches
Authors:
Ding Ma,
Patrick Maher,
Andrew Martin
Abstract:
Longitudinal clinical studies often collect repeated measurements of biomarkers or health-related quality of life together with a time-to-event outcome. These processes are intrinsically linked: longitudinal trajectories may predict event risk, while event occurrence, or its anticipation, can induce informative censoring of the longitudinal process. Joint models provide a principled framework for…
▽ More
Longitudinal clinical studies often collect repeated measurements of biomarkers or health-related quality of life together with a time-to-event outcome. These processes are intrinsically linked: longitudinal trajectories may predict event risk, while event occurrence, or its anticipation, can induce informative censoring of the longitudinal process. Joint models provide a principled framework for handling this dependence, but most existing formulations rely on proportional hazards assumptions that may be restrictive and offer limited interpretability on the time scale. We propose a class of semiparametric accelerated failure time joint models that directly model covariate effects on event timing while flexibly capturing longitudinal-event associations. The survival component is specified through an accelerated failure time model with the baseline component represented by a flexible basis expansion, allowing a broad class of smooth baseline specifications. We illustrate the framework using Bernstein polynomial baseline representations and introduce rescaling strategies to improve numerical stability and parameter identifiability under time-warping. Estimation is conducted within a Bayesian framework, enabling joint inference for longitudinal, survival, and association parameters. Simulation studies reflecting realistic longitudinal trajectories, censoring mechanisms, and dependence structures are used to evaluate finite-sample performance. The proposed models show improved recovery of longitudinal treatment effects compared with a standalone linear mixed model when event risk depends on the underlying longitudinal process. Overall, the framework extends existing joint modelling methodology by offering a flexible and interpretable alternative to proportional hazards-based approaches.
△ Less
Submitted 14 June, 2026;
originally announced June 2026.
-
LLM Agents Can See Code Repositories
Authors:
Dongjian Ma,
Silin Chen,
Yufei Yang,
Yuling Shi,
Yanfu Yan,
Xiaodong Gu
Abstract:
Coding agents powered by large language models have demonstrated strong performance on software engineering tasks. Yet most agents consume repositories almost entirely as text, which differs from how human developers use visual structure such as folder hierarchies and dependency relationships to orient themselves in large codebases. With multimodal large language models (MLLMs), it is an open ques…
▽ More
Coding agents powered by large language models have demonstrated strong performance on software engineering tasks. Yet most agents consume repositories almost entirely as text, which differs from how human developers use visual structure such as folder hierarchies and dependency relationships to orient themselves in large codebases. With multimodal large language models (MLLMs), it is an open question whether agents can effectively benefit from visual representations of repositories. This paper presents the first systematic empirical study of visual repository representations for LLM-based agents on repository-level issue resolution. We evaluate four recent multimodal models. Our results show that a strictly vision-only setup degrades accuracy and increases token cost, because agents lack sufficient symbolic detail and compensate with repeated visual queries. In contrast, integrating visual graphs of repository structure as a supplementary modality alongside standard text interfaces helps agents understand structure more efficiently: input token consumption decreases by up to 26% while issue-resolution accuracy is maintained or improved. Visualization is most useful during fault localization and when the agent autonomously controls exploration depth. These findings point to a practical hybrid text-and-vision design for next-generation coding agents.
△ Less
Submitted 3 August, 2026; v1 submitted 11 June, 2026;
originally announced June 2026.
-
Beyond Patches: Superpixel Token-based Transformers for Attribute-Specific Fashion Retrieval
Authors:
Shuili Zhang,
Hongzhang Mu,
Wenyuan Zhang,
Duohe Ma,
Tingwen Liu
Abstract:
Attribute-Specific Fashion Retrieval (ASFR) aims to improve fine-grained image retrieval by focusing on specific attributes. However, existing patch-based attention and Transformer methods often misalign with irregular attribute regions and are prone to background noise, limiting their ability to capture subtle, pixel-level microstructures. To tackle these challenges, we propose SuperFashion, the…
▽ More
Attribute-Specific Fashion Retrieval (ASFR) aims to improve fine-grained image retrieval by focusing on specific attributes. However, existing patch-based attention and Transformer methods often misalign with irregular attribute regions and are prone to background noise, limiting their ability to capture subtle, pixel-level microstructures. To tackle these challenges, we propose SuperFashion, the first ASFR framework that adopts superpixel tokens within a Transformer architecture. SuperFashion initially employs an attribute-guided attention mechanism to extract attribute-related features, which in turn guide the cropping of semantically meaningful image regions. Superpixel segmentation is then leveraged on these regions to generate compact, semantically coherent superpixel tokens. By incorporating modality-specific embeddings for both attribute and superpixel tokens, the superpixel token-based Transformer facilitates adaptive interaction and fusion, thereby enhancing attribute localization and discrimination. Extensive experiments on FashionAI, DARN, and DeepFashion demonstrate relative overall MAP improvements of 1.84%, 9.27%, and 9.35% over prior SOTA. SuperFashion offers a new solution for web-based image retrieval.
△ Less
Submitted 9 June, 2026;
originally announced June 2026.
-
FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention
Authors:
Yan Wang,
Qifan Zhang,
Jiachen Yu,
Tian Liang,
Dongyang Ma,
Xiang Hu,
Zibo Lin,
Chunyang Li,
Zhichao Wang,
Miao Peng,
Nuo Chen,
Jia Li,
Yujiu Yang,
Haitao Mi,
Dong Yu
Abstract:
Conventional LLMs keep the full KV cache loaded during decoding, causing a severe GPU memory bottleneck for ultra-long context serving. In this report, we propose \textbf{Lookahead Sparse Attention (LSA)}, a novel inference paradigm powered by a Neural Memory Indexer built upon the DeepSeek-V4 architecture. Rather than passively attending to all historical tokens, LSA proactively predicts future c…
▽ More
Conventional LLMs keep the full KV cache loaded during decoding, causing a severe GPU memory bottleneck for ultra-long context serving. In this report, we propose \textbf{Lookahead Sparse Attention (LSA)}, a novel inference paradigm powered by a Neural Memory Indexer built upon the DeepSeek-V4 architecture. Rather than passively attending to all historical tokens, LSA proactively predicts future context demands and preserves only the query-critical KV chunks in the GPU memory. Crucially, we instantiate this architecture via a \textbf{backbone-free decoupled training} strategy. By formulating the indexer as a standard dual-encoder architecture, we train it independently using standard retrieval training frameworks without ever loading the massive backbone model into GPU memory.
We demonstrate that this ``less is more'' paradigm significantly maximizes serving efficiency while acting as an effective attention denoiser in tasks that rely on long-term global memory. Across primary long-context evaluation suites (e.g., LongBench-v2, LongMemEval, and RULER), \texttt{FM-DS-V4} compresses the average physical KV cache footprint down to merely 13.5\% of the full-context baseline, while consistently preserving or slightly elevating downstream accuracy (+0.6\% absolute margin on average). At 1M context, per-decode-token compute drops to 0.30$\times$ of the baseline and GPU KV cache shrinks by 90\% (3.73$\to$0.37 GB), translating into \textbf{2.8$\times$ aggregate throughput and 2.7$\times$ concurrency gains} in PD-disaggregated serving on 8$\times$H20 GPUs.
△ Less
Submitted 20 July, 2026; v1 submitted 8 June, 2026;
originally announced June 2026.
-
Macro Economists in the Machine: A Multi-Agent LLM Framework for Commodity-Related ETF Portfolio Construction
Authors:
Yiqing Wang,
Dehao Dai,
Ding Ma,
Kerui Geng
Abstract:
We test whether large language models (LLMs) add value in commodity portfolio construction when the information set and implementation rules are held fixed across strategies. A Hawkish Agent (inflation-tightening prior), a Dovish Agent (growth-easing prior), a Debate Agent, and a deterministic z-score Rule Agent each receive identical FRED macro z-scores and route their tilt signals through the sa…
▽ More
We test whether large language models (LLMs) add value in commodity portfolio construction when the information set and implementation rules are held fixed across strategies. A Hawkish Agent (inflation-tightening prior), a Dovish Agent (growth-easing prior), a Debate Agent, and a deterministic z-score Rule Agent each receive identical FRED macro z-scores and route their tilt signals through the same portfolio engine. Across 124 weekly rebalancing dates spanning the 2023 U.S. rate peak and the 2024-2025 soft landing, all three LLM strategies outperform the Rule Agent in Sharpe terms; the Hawkish and Debate Agents record the largest gains (ΔSharpe = +0.044 and +0.040, both p < 0.10 under a block bootstrap) and preserve a net-of-cost advantage over the passive inverse-volatility benchmark at one-way trading costs up to 30 basis points, while the Rule Agent's thin margin over passive disappears at approximately 5 basis points.The Debate Agent does not outperform the best single agent (ΔSharpe = -0.004, p = 0.769); its contribution is bias correction -- averaging out the Dovish Agent's miscalibrated prior -- rather than deliberation-generated return. The performance advantage is concentrated in the soft-landing sub-period, the evaluation window spans a single rate cycle, and the reported $p$-values are unadjusted for multiple comparisons. Within these limits, the results suggest that an LLM acting as a constrained macro-interpretation function can add modest but economically meaningful value over a transparent rule layer, though the margin is small and its persistence beyond this sample is unknown.
△ Less
Submitted 6 June, 2026;
originally announced June 2026.
-
Learning Emotion-discriminative Representations for Zero-Shot Cross-lingual Speech Emotion Recognition
Authors:
Jinyi Mi,
Ding Ma,
Tomoki Toda
Abstract:
Zero-shot cross-lingual speech emotion recognition (SER) remains challenging due to distribution mismatches across languages and the lack of emotion annotations in target language. Under such conditions, models trained solely on source-language data frequently suffer from degraded generalization when evaluated on unseen target languages. To address this limitation, we propose an emotion-discrimina…
▽ More
Zero-shot cross-lingual speech emotion recognition (SER) remains challenging due to distribution mismatches across languages and the lack of emotion annotations in target language. Under such conditions, models trained solely on source-language data frequently suffer from degraded generalization when evaluated on unseen target languages. To address this limitation, we propose an emotion-discriminative representation learning method that integrates supervised contrastive learning and speaker adversarial learning. The contrastive learning promotes cross-lingual emotion alignment, while speaker adversarial learning suppresses speaker-related cues to encourage speaker-invariant representations. Experimental results under a zero-shot cross-lingual SER setting demonstrate that the proposed method significantly improves SER performance over conventional training strategies.
△ Less
Submitted 4 June, 2026;
originally announced June 2026.
-
Message Tuning Outshines Graph Prompt Tuning: A Prismatic Space Perspective
Authors:
Yancheng Chen,
Dun Ma,
Shuai Zhang,
Yang Liu,
Xixun Lin,
Xiangyu Zhao,
Wenguo Yang,
Wei Chen,
Chuan Zhou
Abstract:
Graph Foundation Models (GFMs), built upon the Pre-training and Adaptation paradigm, have emerged as a research hotspot in graph learning. For GNN-based GFMs, graph prompt tuning has become the prevailing adaptation method for downstream tasks. Although recent methods explain why graph prompt tuning works, how to rigorously measure its adaptation capacity remains an open problem. Addressing this p…
▽ More
Graph Foundation Models (GFMs), built upon the Pre-training and Adaptation paradigm, have emerged as a research hotspot in graph learning. For GNN-based GFMs, graph prompt tuning has become the prevailing adaptation method for downstream tasks. Although recent methods explain why graph prompt tuning works, how to rigorously measure its adaptation capacity remains an open problem. Addressing this problem is critical for understanding the capability limits of graph prompt tuning and for developing more powerful adaptation methods. In this paper, we propose Prismatic Space Theory (PS-Theory), a novel mathematical framework to quantify the capacity of adaptation methods, while focusing on establishing the upper bound for the adaptation capacity of graph prompt tuning. Building upon the proposed PS-Theory, we further introduce Message Tuning for GFMs (MTG), a lightweight approach that injects a small set of learnable message prototypes into each layer of the GNN backbone to adaptively guide message fusion without updating pre-trained weights. Through our PS-Theory, we prove that the adaptation capacity of MTG can exceed the theoretical upper bound of graph prompt tuning. Extensive experiments demonstrate that MTG consistently outperforms graph prompt baselines across diverse benchmark datasets, providing strong empirical support for our theoretical findings.
△ Less
Submitted 2 June, 2026;
originally announced June 2026.
-
Cosmos 3: Omnimodal World Models for Physical AI
Authors:
NVIDIA,
:,
Aditi,
Niket Agarwal,
Arslan Ali,
Jon Allen,
Martin Antolini,
Adeline Aubame,
Alisson Azzolini,
Junjie Bai,
Maciej Bala,
Yogesh Balaji,
Josh Bapst,
Aarti Basant,
Mukesh Beladiya,
Mohammad Qazim Bhat,
Zaid Pervaiz Bhat,
Dan Blick,
Vanni Brighella,
Han Cai,
Tiffany Cai,
Eric Cameracci,
Jiaxin Cao,
Yulong Cao,
Mark Carlson
, et al. (271 additional authors not shown)
Abstract:
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-transformers architecture. By supporting highly flexible input-output configurations, Cosmos 3 seamlessly unifies critical modalities for Physical AI -- effectively subsuming vision-language models, video generators, worl…
▽ More
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-transformers architecture. By supporting highly flexible input-output configurations, Cosmos 3 seamlessly unifies critical modalities for Physical AI -- effectively subsuming vision-language models, video generators, world simulators, and world-action models into a single framework. Our evaluation demonstrates that Cosmos 3 establishes a new state-of-the-art across a diverse suite of understanding and generation tasks, demonstrating omnimodal world models as scalable, general-purpose backbones for embodied agents. Our post-trained Cosmos 3 models were ranked as the best open-source Text-to-Image and Image-to-Video models by Artificial Analysis, and the best policy model by RoboArena at the time the technical report was written. To accelerate open research and deployment in Physical AI, we make our code, model checkpoints, curated synthetic datasets, and evaluation benchmark available under the Linux Foundation's OpenMDW-1.1 License at https://github.com/nvidia/cosmos and https://huggingface.co/collections/nvidia/cosmos3. The project website is available at https://research.nvidia.com/labs/cosmos-lab/cosmos3.
△ Less
Submitted 23 June, 2026; v1 submitted 1 June, 2026;
originally announced June 2026.
-
Pancyclicity of graphs perturbed by a random $F$-factor
Authors:
Dingjia Mao,
Feihong Yuan,
Wenling Zhou
Abstract:
We determine the sharp minimum-degree threshold for Hamiltonicity in graphs perturbed by a uniformly random $K_r$-factor, resolving a conjecture of Espuny Díaz and Girão [Random Structures Algorithms, 2023]. In fact, we prove the stronger pancyclic statement. Let $α^*(K_r)$ and $α_{\text{pan}}^*(K_r)$ denote the Hamiltonicity and pancyclicity thresholds, respectively. We show that…
▽ More
We determine the sharp minimum-degree threshold for Hamiltonicity in graphs perturbed by a uniformly random $K_r$-factor, resolving a conjecture of Espuny Díaz and Girão [Random Structures Algorithms, 2023]. In fact, we prove the stronger pancyclic statement. Let $α^*(K_r)$ and $α_{\text{pan}}^*(K_r)$ denote the Hamiltonicity and pancyclicity thresholds, respectively. We show that $α^*(K_r)=α_{\text{pan}}^*(K_r)=ρ_r$, where $ρ_r$ is the unique positive solution of $x^r+rx-1=0$. The proof is obtained from a general framework for perturbations by a uniformly random $F$-factor, where $F$ is an arbitrary fixed connected graph.
△ Less
Submitted 28 June, 2026; v1 submitted 1 June, 2026;
originally announced June 2026.
-
Residual Decoder Adapter: ID-Preserving Tokenizer Adaption for Autoregressive Text Rendering
Authors:
Dongxing Mao,
Jinpeng Wang,
Jiahao Tang,
Kevin Qinghong Lin,
Linjie Li,
Zhengyuan Yang,
Lijuan Wang,
Min Li,
Jingru Tan
Abstract:
Visual Autoregressive (AR) models generate images by predicting discrete tokens that are decoded by a visual tokenizer. Despite demonstrating strong overall image generation ability, they still underperform on text rendering with blur strokes and disrupt letter shapes. In this work, we trace this limitation to the visual tokenizer, which struggles to reconstruct fine-grained detail. Improving the…
▽ More
Visual Autoregressive (AR) models generate images by predicting discrete tokens that are decoded by a visual tokenizer. Despite demonstrating strong overall image generation ability, they still underperform on text rendering with blur strokes and disrupt letter shapes. In this work, we trace this limitation to the visual tokenizer, which struggles to reconstruct fine-grained detail. Improving the tokenizer is straightforward but expensive, as it necessitates retraining both the tokenizer and the AR model. Can we improve text rendering performance of AR models without retraining the existing tokenizer and AR model? To achieve this, we propose the Residual Decoder Adapter(RDA) that upgrades an existing tokenizer post-hoc without changing its token space. Specifically, it refines the decoder output of the visual tokenizer by introducing two novel components: (i) a paired codebook that shares the token distribution with the original one; (ii) a parallel branch to learn the tiny differences (residual) between the reconstructed image and the ground-truth images in the pixel space. This residual design allows us to enhance the tokenizer non-invasively while preserving compatibility with prior AR models. RDA substantially improves text rendering significantly by a large margin. For instance, we boost finetuned Janus-Pro OCR accuracy rises from 24.52% to 58.26% (TextVisionBlend), from 12.75% to 36.81% (StyledTextSynth) on competitive TextAtlas benchmark. The code is available at https://github.com/CSU-JPG/RDA
△ Less
Submitted 1 June, 2026;
originally announced June 2026.
-
Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning
Authors:
Ding Ma,
Jinyi Mi,
Fengji Li,
Lester Phillip Violeta,
Jiajun He,
Wenchin Huang,
Kazuhiro Kobayashi,
Tomoki Toda
Abstract:
Objective: laryngectomees depend on an electromechanical device to generate electrolaryngeal (EL) speech. Compared with normal speech, EL speech suffers from severe distortion, limited phonetic variation, unnatural prosody, and temporal shifts, degrading naturalness and intelligibility. Although sequence-to-sequence (seq2seq) voice conversion (VC) based EL-speech-to-normal-speech conversion (EL2SP…
▽ More
Objective: laryngectomees depend on an electromechanical device to generate electrolaryngeal (EL) speech. Compared with normal speech, EL speech suffers from severe distortion, limited phonetic variation, unnatural prosody, and temporal shifts, degrading naturalness and intelligibility. Although sequence-to-sequence (seq2seq) voice conversion (VC) based EL-speech-to-normal-speech conversion (EL2SP) is promising, substantial mismatches between EL and normal speech inevitably cause cumulative mapping errors that limit performance. To address this, we describe a novel representation learning framework integrating speech and text representations to improve mapping and reconstruction quality within a seq2seq VC model. Methods: our methodology comprises two main stages: 1) representation integration and learning, and 2) reconstruction training. A network capable of incorporating auxiliary text information is first constructed with pretrained modules to learn speech--text-based integrated representations. Then, an autoencoder-style reconstruction strategy finalizes EL2SP model to inherit these representations without increasing model complexity. We introduce three fusion strategies including middle-, input-, and hybrid-level fusion strategies that progressively enhance learning. Moreover, besides standard seq2seq VC objectives, an additional reconstruction loss on the integrated representation is introduced to refine representation transfer. Results: experiments under different EL2SP datasets consistently demonstrate that our methods, combined with data augmentations, outperform baselines relying solely on speech representations. Furthermore, progressive improvements with system design depth validate the effectiveness of our methods. Significance: the proposed methods provide an extensible and practical methodology for EL speech enhancement and assistive communication technologies.
△ Less
Submitted 1 June, 2026;
originally announced June 2026.
-
A Unified and Reproducible Experimentation Framework for Speech Understanding
Authors:
Jing Peng,
Junhao Du,
Chenghao Wang,
Hanqi Li,
Yi Yang,
Yixuan Wang,
Xiaoyu Gu,
Guanyu Chen,
Yucheng Wang,
Jiang Li,
Zhangjie Zhao,
Haoran Wang,
Wenming Tu,
Haoyu Li,
Duo Ma,
Lirong Qian,
Yu Xi,
Wen Wen,
Jiaqi Guo,
Hui Zhang,
Shuai Fan,
Wenbin Jiang,
Shuai Wang,
Kai Yu
Abstract:
Speech foundation models and Speech LLMs have advanced speech understanding, yet deployment-oriented model selection is hindered by non-comparable evaluations caused by mismatched post-processing, and by training results that are hard to reproduce across data scales and pipelines. We present SURE, a unified experimentation framework that standardizes prediction formats, normalization, and scoring.…
▽ More
Speech foundation models and Speech LLMs have advanced speech understanding, yet deployment-oriented model selection is hindered by non-comparable evaluations caused by mismatched post-processing, and by training results that are hard to reproduce across data scales and pipelines. We present SURE, a unified experimentation framework that standardizes prediction formats, normalization, and scoring. SURE evaluates strong systems across paradigms, from conventional pipelines to Speech LLMs, on representative tasks under realistic acoustic and linguistic stressors. Beyond evaluation, SURE introduces an agent-assisted training conversion flow that maps paper and code into versioned, runnable training pipelines under a unified protocol on matched open-data subsets. Overall, SURE improves comparability and reproducibility for deployment-oriented evaluation.
△ Less
Submitted 29 May, 2026;
originally announced May 2026.
-
Seeing Before Agreeing: Aligning Multi-Agent Consensus with Visual Evidence
Authors:
Yuhan Wang,
Shuochen Chang,
Yalin Feng,
Dongsheng Ma,
Yuanzi Li,
Zhengren Wang,
Yinglong Yang,
Yufei Chen,
Yikang Wang,
Shaoxu Sun,
Wentao Zhang
Abstract:
Vision-language models (VLMs) have achieved strong performance on visual question answering (VQA). To mitigate individual hallucinations and blind spots, aggregating diverse perspectives via multi-agent collaboration has emerged as a promising paradigm. While this approach has shown great success in textual QA, its potential in the multimodal domain remains under-explored. Existing multi-agent VQA…
▽ More
Vision-language models (VLMs) have achieved strong performance on visual question answering (VQA). To mitigate individual hallucinations and blind spots, aggregating diverse perspectives via multi-agent collaboration has emerged as a promising paradigm. While this approach has shown great success in textual QA, its potential in the multimodal domain remains under-explored. Existing multi-agent VQA methods predominantly adapt text-centric protocols, focusing on textual discussions while ignoring the alignment of visual information. In this work, we reveal a key insight: answer-level agreement is insufficient for reliable multi-agent VQA; \textit{aligned visual evidence} -- shared support from the image regions agents rely on -- is essential for trustworthy consensus. To leverage this insight, we propose EAGLE (\textbf{E}vidence-\textbf{A}ligned \textbf{G}rounded mu\textbf{L}ti-agent r\textbf{E}asoning), a training-free evidence-centered framework for coordinating multiple VLM agents. EAGLE explicitly exposes each agent's grounding regions as visual evidence, enables mutual verification over the evidence, and uses evidence consistency to guide final decision-making. Experiments on six VQA benchmarks show that EAGLE achieves best average performance across domains while remaining lightweight, interpretable, and practical for deployment.
△ Less
Submitted 28 May, 2026;
originally announced May 2026.
-
DeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey Generation
Authors:
Ziyue Yang,
Da Ma,
Hanqi Li,
Zijian Wang,
Tiancheng Huang,
Zijian Hu,
Chenrun Wang,
Yunzhe Zhang,
Xiaobao Wu,
Kai Yu,
Lu Chen
Abstract:
As scientific literature grows rapidly, automated survey generation has become a key capability for AI scientists and human researchers. However, existing systems suffer from limited analytical depth due to reliance on abstracts and isolated paper processing, and unreliable citations from imprecise retrieval and post-hoc grounding, producing superficial surveys and may mislead researchers. We pres…
▽ More
As scientific literature grows rapidly, automated survey generation has become a key capability for AI scientists and human researchers. However, existing systems suffer from limited analytical depth due to reliance on abstracts and isolated paper processing, and unreliable citations from imprecise retrieval and post-hoc grounding, producing superficial surveys and may mislead researchers. We present DeepSurvey, an agentic system that addresses both. To enhance depth, DeepSurvey extracts structured keynotes from full-text papers, models cross-paper relationships through clustering and comparative analysis, and integrates code-repository analysis to recover implementation-level details. To fortify reliability, it combines citation-graph expansion with hybrid filtering for topic-focussed retrieval, enforces evidence-constrained citation assignment, and deploys multi-granularity agentic refinement to validate citation-claim alignment. Experiments show that DeepSurvey achieves the highest content score (8.644/10) and citation quality (12.3% and 9.3% recall and precision gains over the strongest baseline), generalizes more robustly across domains (0.14 vs 0.22 to 0.69 CS-to-non-CS drop), and is preferred over human-written surveys by domain experts (83.3% overall quality, 100% content depth).
△ Less
Submitted 28 May, 2026;
originally announced May 2026.
-
Zero-Field Thermal Hall Effect in Insulator
Authors:
Zhe Cui,
Haoran Fan,
Wenjiang Zhou,
Xianghong Jin,
Yuchen Gu,
Da Ma,
Cong Xiao,
Hua Jiang,
Xincheng Xie,
Bai Song,
Yuan Li,
Xi Lin
Abstract:
Fourier's law dictates that heat flow is usually parallel to the applied temperature gradient. However, under a high magnetic field, heat flow carried by both electrons in conductors and phonons in insulators can be deflected, a phenomenon known as thermal Hall effect. Intriguingly, we observe at zero field a spontaneous thermal Hall effect in an antiferromagnetic insulator. Despite a vanishingly…
▽ More
Fourier's law dictates that heat flow is usually parallel to the applied temperature gradient. However, under a high magnetic field, heat flow carried by both electrons in conductors and phonons in insulators can be deflected, a phenomenon known as thermal Hall effect. Intriguingly, we observe at zero field a spontaneous thermal Hall effect in an antiferromagnetic insulator. Despite a vanishingly small uncompensated magnetization, the magnitude of this effect is surprisingly large, comparable to typical responses induced by several teslas of external field. This zero-field behavior indicates that charge-neutral heat carriers can be governed by an intrinsic effective field arising from the unique spin arrangement. Our discovery challenges the centuries-old preconception of heat conduction and open up new avenues for exploring non-trivial topological responses in quantum materials.
△ Less
Submitted 27 May, 2026;
originally announced May 2026.
-
The Daily Dose: Workflow-Integrated Large Language Model Automation for Clinical Summarization and Trial Identification in Radiation Oncology
Authors:
Jason Holmes,
Federico Mastroleo,
Mariana Borras-Osorio,
Srinivas Seetamsetty,
Satomi Shiraishi,
Mirek Fatyga,
Judy C. Boughey,
Cornelius A. Thiels,
William G. Breen,
Daniel J. Ma,
Daniel K. Ebner,
David M. Routman,
Brady S. Laughlin,
Carlos E. Vargas,
Samir H. Patel,
Sujay A. Vora,
Nadia N. Laack,
Andrew Y. K. Foong,
Wei Liu,
Mark R. Waddle
Abstract:
Objective: To describe the design and early clinical evaluation of The Daily Dose (TDD), an LLM-driven, automated clinical summarization and clinical-trial identification system integrated into routine radiation oncology practice. Design: Mixed-methods evaluation using a cross-sectional, anonymous clinician survey administered after 1 month of system deployment. Exposure: Daily automated delivery…
▽ More
Objective: To describe the design and early clinical evaluation of The Daily Dose (TDD), an LLM-driven, automated clinical summarization and clinical-trial identification system integrated into routine radiation oncology practice. Design: Mixed-methods evaluation using a cross-sectional, anonymous clinician survey administered after 1 month of system deployment. Exposure: Daily automated delivery of physician-specific email summaries generated using RadOnc-GPT, including patient schedules, concise EHR-derived clinical-status summaries, and automated identification of potentially relevant clinical trials for new or consult visits. Main Outcomes and Measures: Primary outcomes included self-reported usability, satisfaction, perceived usefulness, perceived impact on workflow, time savings, and intention for continued use. Internal consistency reliability was assessed using Cronbach's $α$. Results: Among 55 respondents, 52 (94.5\%) worked in radiation oncology, and 38 (69.1\%) were attending physicians. Most participants (83.6\%) reported using TDD daily or several times per week. Mean (SD) scores were 3.89 (1.04) for usability and satisfaction, 3.43 (1.24) for perceived usefulness, and 3.80 (1.17) for impact and future use (5-point Likert scale). Overall satisfaction was positively associated with perceived time savings ($p < .001$). Participants reported variable time savings, with 27\% estimating $\geq 10$ minutes saved per day. The questionnaire demonstrated excellent internal consistency (overall Cronbach's $α$ = 0.97).
△ Less
Submitted 25 May, 2026;
originally announced May 2026.