-
Field deployment of a laser wakefield accelerator for on-site application
Authors:
Bo Guo,
Xiaonan Ning,
Dexiang Liu,
Yue Ma,
Weiwang Zeng,
Mingyuan Wei,
Shengtai Wei,
Jianfei Hua,
Yang Wan,
Wei Lu
Abstract:
Successive innovations in particle accelerators have continually expanded the frontiers of scientific discovery. Laser wakefield accelerators promise to transform science, medicine, and industry, yet moving them from laboratory demonstrations to reliable real-world operation has remained a central, long-standing challenge. Here we report a field-deployable system that produced 100-MeV-class electr…
▽ More
Successive innovations in particle accelerators have continually expanded the frontiers of scientific discovery. Laser wakefield accelerators promise to transform science, medicine, and industry, yet moving them from laboratory demonstrations to reliable real-world operation has remained a central, long-standing challenge. Here we report a field-deployable system that produced 100-MeV-class electron beams with 1%-level energy stability during 72 hours of continuous operation and supported routine full-power use throughout a seven-month field trial in an industrial setting. Applied to in situ micro-nondestructive testing, the system generated tens-of-MeV bremsstrahlung X-rays that enabled three-dimensional microtomography of dense materials at sub-50-μm spatial resolution and revealed 100-μm-scale internal defects in large composite structures, extending the capabilities beyond those of existing high-energy X-ray sources. These results mark a transition of laser wakefield acceleration from laboratory proof of concept toward practical deployment in scientific and industrial applications.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution
Authors:
Maolin Ran,
Xiaoyang Lu,
Jiaqi Liu,
Jian Wang,
Weiwen Liu,
Jianghao Lin,
Yong Yu,
Weinan Zhang
Abstract:
Storyboards turn screenplays into visual shot plans for automated short drama production. Professional storyboarding relies on tacit directorial expertise and remains an industrial bottleneck. Large language models can automate this step, but methods for supplying directing knowledge face three challenges: (1) Knowledge acquisition: the craft remains implicit in exemplars or must be written manual…
▽ More
Storyboards turn screenplays into visual shot plans for automated short drama production. Professional storyboarding relies on tacit directorial expertise and remains an industrial bottleneck. Large language models can automate this step, but methods for supplying directing knowledge face three challenges: (1) Knowledge acquisition: the craft remains implicit in exemplars or must be written manually. (2) Knowledge refinement: authored knowledge is not evaluated against execution outcomes, and opaque generation prevents feedback attribution to the knowledge behind each decision. (3) Knowledge injection: injecting all knowledge exceeds usable context, while manual selection for every narrative group does not scale. We present SAGE (Skill with Attribution-Guided Evolution), a deployed framework that learns, attributes, evolves, and routes directing knowledge from expert demonstrations. SAGE derives rules that are independent of episode content by contrasting each training screenplay with its expert storyboard. During generation, the model records each narrative group's adopted rules. Combining these records with localized feedback enables targeted updates to individual rules. Evolved rules form scenario packages with a routing index, so each group retrieves only a bounded set appropriate to its situation without expert intervention. On 18 test episodes across three genres, SAGE scored 77.8 on a rubric validated by experts, versus 77.1 for professional directors. Deployed for 14 days on Virtual Film Studio, SAGE produced 1,344 narrative group outputs; 87.2 percent were accepted without substantive edits, and the production team recorded over 83 percent less authoring time per episode. We release PROSE, the first public dataset pairing screenplays with storyboards by professional directors across 68 episodes: https://github.com/creDreams/PROSE.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
NGS-Marker: Robust Native Watermarking for 3D Gaussian Splatting
Authors:
Hao Qin,
Yukai Sun,
Luyuan Chen,
Mengxu Lu,
Feng Zhang,
Ming Kong,
Zhenhong Du,
Qiang Zhu
Abstract:
With the rapid development and adoption of 3D Gaussian Splatting (3DGS), the need for effective copyright protection has become increasingly critical. Existing watermarking techniques for 3DGS mainly focus on protecting rendered images via pre-trained decoders, leaving the underlying 3D Gaussian primitives vulnerable to misuse. In particular, they are ineffective against Partial Infringement, wher…
▽ More
With the rapid development and adoption of 3D Gaussian Splatting (3DGS), the need for effective copyright protection has become increasingly critical. Existing watermarking techniques for 3DGS mainly focus on protecting rendered images via pre-trained decoders, leaving the underlying 3D Gaussian primitives vulnerable to misuse. In particular, they are ineffective against Partial Infringement, where an adversary extracts and reuses only a subset of Gaussians. In this paper, we propose NGS-Marker, a novel native watermarking framework for 3DGS. It integrates a jointly trained watermark injector and message decoder, and employs a gradientbased progressive injection strategy to ensure full-scene coverage. This enables robust ownership decoding from any local region. We further extend NGS-Marker with hybrid protection (combining native and indirect watermarks) and support for multimodal watermarking. Extensive experiments demonstrate that NGS-Marker effectively defends against partial infringement while offering practical flexibility for real-world deployment.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Counterfactual Anatomy-guided Spatial-Temporal Decoding for Annotation-Free Hallucination Mitigation in Medical VLMs
Authors:
Yifan Lu,
Adinath Dukre,
Abhijit Das,
Ziyun Zou,
Haolin Yang,
Yutong Xie,
Imran Razzak
Abstract:
Medical vision-language models (Med-VLMs) have demonstrated strong performance on medical visual question answering, yet they remain prone to hallucination, generating clinically unsupported statements that are insufficiently grounded in image evidence. Mitigation methods applied during decoding offer a practical solution, but they typically lack anatomical awareness or rely heavily on ground trut…
▽ More
Medical vision-language models (Med-VLMs) have demonstrated strong performance on medical visual question answering, yet they remain prone to hallucination, generating clinically unsupported statements that are insufficiently grounded in image evidence. Mitigation methods applied during decoding offer a practical solution, but they typically lack anatomical awareness or rely heavily on ground truth annotations, which limits their applicability. We propose Counterfactual Anatomy-guided Spatial-Temporal decoding (CAST), a framework that operates entirely during inference and requires no manual annotations for anatomically grounded hallucination mitigation. CAST automatically discovers anatomical regions relevant to the given query through broad medical segmentation. It then selects a compact, causally informative area using counterfactual intervention based on the drop in answer likelihood under occlusion. Guided by this chosen region, CAST performs a unified contrastive decoding process, combining classifier-free guidance to correct spatial attention with stepwise temporal contrast to regulate generation dynamics. Experiments on the SLAKE and MIMIC-CXR datasets across three Med-VLMs demonstrate that CAST consistently outperforms strong baselines and surpasses decoding strategies reliant on ground truth. Our results indicate that compact, automatically selected regions provide highly effective contrastive guidance without expert annotations, offering a practical and generalizable solution for improving spatial grounding and reducing hallucinations. Code is available at https://github.com/csyifan/CAST.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Prism-GRPO: Faster VLA Policy Optimization via Splitting Same-outcome Groups
Authors:
Zeyun Deng,
Yuzhe Lu,
Yawei Wang,
Linbo Liu,
Qing Ping,
Han Ding,
Guande Wu,
Panpan Xu,
Jun Huan
Abstract:
GRPO is increasingly used for reinforcement learning of vision-language-action (VLA) policies because, unlike PPO, it does not require training a critic. This simplification comes with a sampling cost: group-relative advantages require multiple rollouts from each scene. Under binary success rewards, groups whose rollouts all succeed or all fail have zero advantage and are discarded by dynamic samp…
▽ More
GRPO is increasingly used for reinforcement learning of vision-language-action (VLA) policies because, unlike PPO, it does not require training a critic. This simplification comes with a sampling cost: group-relative advantages require multiple rollouts from each scene. Under binary success rewards, groups whose rollouts all succeed or all fail have zero advantage and are discarded by dynamic sampling. These groups are especially common early in training, when most rollouts fail, wasting much of the expensive robotic rollout budget. We introduce Prism-GRPO, which augments binary outcome reward with a weighted trajectory-level execution-quality score. By splitting same-outcome groups into a quality spectrum, Prism-GRPO recovers training signal while ensuring that every success still outranks every failure. Quality scores can be derived from simulator contacts, executed actions, or visual observations, avoiding task-specific progress rewards. We prove that Prism-GRPO never increases the probability that a sampled group is discarded for having zero advantages, and derive a gradient-alignment condition under which its combined update remains a local ascent direction for task success. Across four RoboTwin tasks spanning different horizons and coordination patterns, Prism-GRPO improves success and quality at matched rollout budgets and reaches target success rates with up to 56% fewer rollouts. It also suppresses a reward-hacking shortcut, with the cleaner behavior transferring under direct deployment to a real robot. Through ablations, we show consistent gains across contact-, smoothness-, and VLM-derived quality signals.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Green function asymptotics for the $σ_2$-Yamabe problem
Authors:
Bin Deng,
Han Lu
Abstract:
We establish, at every pole, the asymptotic expansion of the normalized $Γ_2$-Green function with a finite pole set in every dimension $n\ge5$. In dimensions $5\le n\le7$ the first correction is a unique nonnegative constant, whereas dimension $8$ exhibits a Weyl-driven $\sqrt{\log}$ term. For $n>8$ we construct the finite local curvature parametrix through the first positive indicial resonance an…
▽ More
We establish, at every pole, the asymptotic expansion of the normalized $Γ_2$-Green function with a finite pole set in every dimension $n\ge5$. In dimensions $5\le n\le7$ the first correction is a unique nonnegative constant, whereas dimension $8$ exhibits a Weyl-driven $\sqrt{\log}$ term. For $n>8$ we construct the finite local curvature parametrix through the first positive indicial resonance and identify the first genuinely global coefficient by a relative Newton flux.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding
Authors:
Bonan Zhang,
Shiyu Dong,
Quan Hung Tran,
Katharina Gschwind,
Shuqi Yang,
Sijia Chen,
Adel Ahmadyan,
Seungwhan Moon,
Lu Zhang,
Ahmed Kirmani,
Babak Damavandi,
Anuj Kumar
Abstract:
Vision encoders are a critical component of vision-language models, and scaling their capacity effectively improves performance. However, dense scaling increases compute cost and inference latency. Mixture-of-Experts (MoE) architectures offer a compelling alternative, having enabled efficient scaling in LLMs, yet the MoE design space for CLIP-style vision encoders remains underexplored at State-of…
▽ More
Vision encoders are a critical component of vision-language models, and scaling their capacity effectively improves performance. However, dense scaling increases compute cost and inference latency. Mixture-of-Experts (MoE) architectures offer a compelling alternative, having enabled efficient scaling in LLMs, yet the MoE design space for CLIP-style vision encoders remains underexplored at State-of-the-Art (SOTA) levels. In this work, we systematically study MoE designs for vision encoder scaling and find that fine-grained MoE topologies yield substantial gains over both dense and standard MoE counterparts. We further propose an auxiliary-loss-free balancing variant for better expert utilization, and design a specialized MoE kernel to mitigate inference latency overhead. To enhance video capabilities while preserving image knowledge, we introduce frame-level distillation paired with a novel freezing mechanism. We pretrain a series of Mixture-of-Experts Vision Encoders (MoE-ViE) across a range of sizes, all consistently outperforming their dense counterparts. Our largest model matches the zero-shot performance of a SOTA encoder 1.7x its size at 76% of its latency. When aligned with an LLM, MoE-ViE surpasses all compared encoders on image and video benchmarks, including those with up to 5x more activated parameters. Code is available at https://github.com/facebookresearch/moe_vie.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Universal CKM for Environment-Aware Wireless Networks: Enabling Cross-Device and Cross-Task Channel Knowledge Transfer
Authors:
Haiquan Lu,
Yong Zeng,
Cheng-Xiang Wang,
Xiqi Gao,
Rui Zhang
Abstract:
Channel knowledge map (CKM) is a promising technology for environment-aware sixth-generation (6G) wireless networks. However, most existing CKMs are tightly coupled with wireless devices and downstream tasks, which limit their scalability and reusability in wireless networks. To address these limitations, this article proposes the concept of universal CKM (uCKM) as a foundational wireless environm…
▽ More
Channel knowledge map (CKM) is a promising technology for environment-aware sixth-generation (6G) wireless networks. However, most existing CKMs are tightly coupled with wireless devices and downstream tasks, which limit their scalability and reusability in wireless networks. To address these limitations, this article proposes the concept of universal CKM (uCKM) as a foundational wireless environment prior, which aims to enable cross-device and cross-task channel knowledge transfer for environment-aware wireless networks. We first revisit the representative CKMs and discuss their limitations. Then, the uCKM-enabled new paradigm for environment-aware wireless networks is introduced, and its benefits are highlighted from the perspectives of uCKM construction and utilization phases, for which we propose the visions of ``All for uCKM'' and ``uCKM for All'', i.e., the data acquired by all devices and tasks should contribute to the construction of uCKM, and vice versa. Subsequently, we discuss the main challenges of uCKM and propose potential solutions. Last, we provide simulation results to demonstrate the feasibility and performance gains brought by uCKM and outline future research directions.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets
Authors:
Zhida He,
Xiaoyu Wen,
Han Qi,
Ziyuan Zhou,
Peng Yu,
Jiajia Li,
Chaochao Lu,
Qiaosheng Zhang
Abstract:
Reliable jailbreak evaluation is essential for assessing LLM safety, but most existing studies rely solely on attack success rate (ASR) without accounting for its dependence on attack budgets, resulting in unfair comparisons across methods. Existing compute-aware evaluations reduce heterogeneous resources into FLOPs, which is difficult to estimate for black-box models and fails to capture resource…
▽ More
Reliable jailbreak evaluation is essential for assessing LLM safety, but most existing studies rely solely on attack success rate (ASR) without accounting for its dependence on attack budgets, resulting in unfair comparisons across methods. Existing compute-aware evaluations reduce heterogeneous resources into FLOPs, which is difficult to estimate for black-box models and fails to capture resource-specific constraints. To provide a comparable evaluation basis, we introduce Fair-ASR, an evaluation protocol for black-box jailbreak attacks under shared target-call budgets B, using target calls as a directly observable and method-agnostic comparison axis while tracking attacker calls separately for efficiency analysis. We re-evaluate 11 representative attacks under the Fair-ASR protocol and find that attack rankings change substantially across target-call budgets, simple stochastic perturbations and hand-crafted templates remain highly competitive under equal target access, and no evaluated LLM-driven method is efficient in both target and attacker calls. Motivated by this efficiency gap, we introduce ReCode, a compositional budget-efficient attack that combines desensitization rewriting with two effective low-cost primitives identified by Fair-ASR. Under a budget of 20 target calls, ReCode achieves 85% ASR on GPT-5 while requiring only 7.19 attacker calls per request on average, showing strong efficiency in both target and attacker calls.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
B-Spline Embedded Structure Learning for 3D Tooth Segmentation
Authors:
Xianghan Wei,
Jianwen Lou,
Zhiguo Lu,
Hairong Jin,
Haihua Zhu
Abstract:
Accurate 3D tooth segmentation forms the cornerstone of digital dentistry, yet it remains a formidable challenge due to the inherent intricacy of real-world dentitions, such as crowding, misaligned teeth and high morphological similarity between adjacent teeth. To resolve this, we present B-Spline Embedded Structure Learning, a novel framework that distills the inherent sequential arrangement of t…
▽ More
Accurate 3D tooth segmentation forms the cornerstone of digital dentistry, yet it remains a formidable challenge due to the inherent intricacy of real-world dentitions, such as crowding, misaligned teeth and high morphological similarity between adjacent teeth. To resolve this, we present B-Spline Embedded Structure Learning, a novel framework that distills the inherent sequential arrangement of teeth into a continuous structural constraint to regularize representation space. Our approach parameterizes the global dental topology by fitting a parametric B-spline trajectory to tooth centers, assigning each point a continuous structural embedding that forces the shared backbone to capture global arch organization. To fully exploit these embedded priors, we introduce a Structure-Aware Dynamic Classifier (SADC) to substitute rigid static templates with adaptive, case-calibrated decision boundaries. SADC regularizes dynamic prototype pooling via a localized Gaussian proximity gate and contextually co-evolves them through an attention block modeling spatial relations and bilateral symmetries across teeth. Extensive evaluations on the 3DTeethSeg22 benchmark demonstrate that our method establishes a new state-of-the-art accuracy with exceptional structural robustness and efficiency in computational overhead, markedly enhancing the model's capacity to handle complex dental configurations.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
When Agents Act on Web3: An Attack-Surface Survey of MCP, Skills, and Tool Calling
Authors:
Rabimba Karanjai,
Yang Lu,
Nour Diallo,
Wujie Xiong,
Lei Xu,
Weidong,
Shi
Abstract:
AI agents increasingly act rather than merely read: across the Model Context Protocol (MCP) ecosystem, the share of deployed tools that modify external state has risen from 27% to 65% of tool use. When agents exercise this authority on public blockchains through MCP, skills, and tool calling, the consequences of an attack are governed by the blockchain execution layer rather than by conventional s…
▽ More
AI agents increasingly act rather than merely read: across the Model Context Protocol (MCP) ecosystem, the share of deployed tools that modify external state has risen from 27% to 65% of tool use. When agents exercise this authority on public blockchains through MCP, skills, and tool calling, the consequences of an attack are governed by the blockchain execution layer rather than by conventional software assumptions. This survey argues that four properties of that layer (irreversibility, signing authority, continuous autonomy, and sequence-level composition) qualitatively change the threat model, turning the recoverable failures of generic agent security into a standing, irreversible loss. We organize the fragmented MCP-security literature into an attack-surface taxonomy, then contribute a Web3 risk-mapping matrix that ties each attack class to its amplified impact, the responsible amplifiers, a representative mitigation, and the residual gap. We synthesize defenses, including emerging blockchain-based mechanisms, and find them improving but insufficient: measured protections stop fewer than 30% of attacks, and model-level safety refuses fewer than 3%. We close by positioning the work against adjacent surveys and deriving a research agenda from the matrix's open cells.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
PACE: Policy-Attested Contract Execution for Safe AI Agents in Decentralized Finance
Authors:
Rabimba Karanjai,
Yang Lu,
Richard Williamson,
Hemanth Hm,
Prakhar Mehrotra,
Lei Xu,
Weidong,
Shi
Abstract:
Autonomous AI agents are emerging as interfaces for decentralized finance (DeFi) actions such as swaps, lending operations, and yield management. Because these agents rely on large language models (LLMs) to plan transactions, they inherit the LLM's susceptibility to prompt injection and lack of mechanisms to bind a verifier's approval to the exact transaction ultimately submitted on-chain. We pres…
▽ More
Autonomous AI agents are emerging as interfaces for decentralized finance (DeFi) actions such as swaps, lending operations, and yield management. Because these agents rely on large language models (LLMs) to plan transactions, they inherit the LLM's susceptibility to prompt injection and lack of mechanisms to bind a verifier's approval to the exact transaction ultimately submitted on-chain. We present PACE (Policy-Attested Contract Execution), a transaction-level authorization framework that interposes between an LLM-based agent and on-chain execution. PACE introduces typed transaction intents, a deterministic policy verifier, and signed Policy Decision Records (PDRs) that cryptographically bind the approved intent, policy, and simulation report to the exact execution bytes, with replay and expiration protection. A Solidity smart account enforces PDR signatures on-chain with a measured overhead of 29,826-31,822 gas. We evaluate PACE against six baselines on 40 tasks spanning four attack categories plus benign utility (2,800 trials, 10 seeds). In our deterministic sandbox, PACE achieves a 0.00 unsafe execution rate and 0.00 false-positive rate on benign tasks, compared to 0.80 for the unguarded baseline. Ablation studies identify permissive policy settings (+57.5 pp) and the touched-contract allowlist (+12.5 pp) as the dominant safety components. To test whether the same deterministic floor holds for real model outputs, the artifact additionally provides a three-model live-LLM evaluation over the full task suite with repeated runs. A mainnet-fork harness is included for archive-RPC deployments, but fork results are reported only when the corresponding artifacts are generated. These auxiliary studies are separate from, and never substitute for, the deterministic benchmark. We frame our claims as logic-level safety within a reproducible benchmark rather than deployment-ready DeFi security.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Fast Nondestructive Readout for High-Clock-Rate Atom Array Quantum Processor
Authors:
Xu-Zhao-Qiu Zeng,
Chang You,
Qing-Wei Wang,
Zi-Feng Li,
Yi Ji,
Dong An,
Chao Yu,
Jia-Rui Liu,
Zi-Mo He,
Jia-Rui Gu,
Yuhao Mei,
Hao-Wen Cheng,
Yu-Chen Zhang,
Rui Lin,
Zhan Wu,
Jun Rui,
Jun Zhang,
Ming-Cheng Chen,
Yu-Hao Deng,
Chao-Yang Lu,
Jian-Wei Pan
Abstract:
Neutral-atom arrays have rapidly advanced to support thousands of qubits and execute high-fidelity logical operations. However, these processors remain severely throttled by their slowest fundamental operation: nondestructive qubit measurement, which requires milliseconds and fundamentally limits the system's clock rate. This bottleneck arises from both an inherent photon-budget dilemma---sufficie…
▽ More
Neutral-atom arrays have rapidly advanced to support thousands of qubits and execute high-fidelity logical operations. However, these processors remain severely throttled by their slowest fundamental operation: nondestructive qubit measurement, which requires milliseconds and fundamentally limits the system's clock rate. This bottleneck arises from both an inherent photon-budget dilemma---sufficient fluorescence for reliable state discrimination must be collected without excessive heating or loss---and frame-based imaging, which imposes one common exposure and decision latency on intrinsically independent, site-local measurements. Here, we overcome these limitations with a fast, nondestructive readout architecture based on real-time, site-resolved adaptive protection. By integrating continuous photon counting with a dynamic feedforward framework, we decode qubit states with sub-microsecond latency and instantly shield atoms from redundant scattering. Demonstrated in parallel across a 100-qubit reconfigurable atom array, with adaptive protection on a 25-site subarray, this dynamic decision protocol reduces the average probe time to just $15\ μ\text{s}$. Model-free benchmarking yields a discrimination infidelity of $4.1 \times 10^{-5}$ and an atom loss of $2.1 \times 10^{-4}$, simultaneously setting new performance records for atom arrays. Exploiting this capability, we operate repeated quantum circuits at an unprecedented 1.7 kHz clock rate with atoms reused over 120 consecutive rounds---nearly sevenfold higher than the previous record---and enter the sub-millisecond cycle regime for the first time. By removing nondestructive readout as the dominant cycle-time bottleneck, this work unlocks high-clock-rate mid-circuit syndrome extraction, paving the way for high-throughput, fault-tolerant quantum computation.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Balancing Safety and Autonomy: Accessibility-Oriented Interventions in Generative AI for Cognitive Impairment
Authors:
Yibo Meng,
Jingruo Chen,
Lyumanshan Ye,
Bingyi Liu,
Zhicong Lu
Abstract:
Generative AI systems are increasingly used by older adults with cognitive impairment for everyday tasks such as information seeking, health management, and communication. While these systems provide flexible, language-based support, their open-ended outputs introduce risks of over-reliance, misinterpretation, and inappropriate decision-making. Prior work has focused on usability and adoption, wit…
▽ More
Generative AI systems are increasingly used by older adults with cognitive impairment for everyday tasks such as information seeking, health management, and communication. While these systems provide flexible, language-based support, their open-ended outputs introduce risks of over-reliance, misinterpretation, and inappropriate decision-making. Prior work has focused on usability and adoption, with limited attention to how system design shapes users' participation in decision-making and the distribution of agency in care contexts. We present a qualitative study of 45 individuals with cognitive impairment and their caregivers. We identify five accessibility-oriented mechanisms: AI Capability Constraint, Human Oversight Embedding, Cognitive Engagement Maintenance, Human-AI Relationship Regulation, and Risk Transparency and Control, through which systems structure interaction. These mechanisms both support and constrain users by redistributing decision-making across users and caregivers. We show that their effects vary by impairment level: while protective mechanisms support users with severe impairment, they can restrict autonomy for those with mild impairment. As impairment progresses, tensions become less visible as user participation diminishes. Our findings highlight the need for dynamic designs that balance safety and autonomy in AI-supported care.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Clinical Pathways Matter for Multimodal Deep Learning in Early Alzheimers Disease Detection
Authors:
Yao Lu,
Solveig Kristina Hammonds,
Alvaro Fernandez-Quilez
Abstract:
Identifying individuals at risk of Alzheimer's disease (AD), particularly in the preclinical and early stages, remains challenging. Although deep learning approaches based on structural MRI show promise as a non-invasive biomarker, existing multimodal models require task-specific training and depend on biomarkers that are not routinely available in clinical practice. Here, we propose a zero-shot m…
▽ More
Identifying individuals at risk of Alzheimer's disease (AD), particularly in the preclinical and early stages, remains challenging. Although deep learning approaches based on structural MRI show promise as a non-invasive biomarker, existing multimodal models require task-specific training and depend on biomarkers that are not routinely available in clinical practice. Here, we propose a zero-shot multimodal framework based on SigLIP that combines structural MRI embeddings with text embeddings of routinely collected clinical variables for early AD risk stratification in individuals at preclinical or mild cognitive impairment (MCI) stages. We evaluated the approach in 416 individuals from the ADNI cohort (age: 72.73 +- 6.7). SigLIP was used without fine-tuning to extract MRI and clinical text embeddings, which were combined into multimodal representations for individual-level AD risk prediction within 4 years. We further compared the model performance in a single-visit and two-visit settings to assess the value of longitudinal information and framework scalability. In the single-visit setting, combining MRI embeddings with MMSE, age, and sex achieved an AUC of 0.91 +- 0.02, outperforming both a CSF A\b{eta}42-based model (AUC 0.73 +- 0.08) and an MMSE-based model (AUC 0.85 +- 0.22). In the two-visit setting, performance was maintained or improved, supporting the scalability of the approach to longitudinal data. These findings suggest that zero-shot multimodal fusion of structural MRI and routinely collected clinical variables may provide a practical and scalable strategy for early AD risk stratification without task-specific retraining.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
A diffusion model for time-dependent compositional data
Authors:
Lu Chen,
Omar De la Cruz Cabrera,
Oana Mocioalca
Abstract:
We introduce a stochastic process for modeling the evolution in time of compositional measurements (i.e., a vector of non-negative values that add up to a total of 1). This model is a diffusion, as it is defined as the solution for a stochastic differential equation in the Ito sense, and it has a Dirichlet distribution as its steady distribution. We have named this process Dirichlet Diffusion (DD)…
▽ More
We introduce a stochastic process for modeling the evolution in time of compositional measurements (i.e., a vector of non-negative values that add up to a total of 1). This model is a diffusion, as it is defined as the solution for a stochastic differential equation in the Ito sense, and it has a Dirichlet distribution as its steady distribution. We have named this process Dirichlet Diffusion (DD).
As the process is confined to a manifold and the coefficients of the equation are not globally Lipschitz, the usual theorems do not apply directly and establishing the existence and properties of solutions requires a somewhat delicate analysis. We establish the existence of strong solutions under the assumption that all the parameters of the Dirichlet distribution are greater than 2; for the general case we were only able to establish the existence of weak solutions, but also that these solutions remain confined to the closed simplex without need for reflecting boundaries.
A useful feature of DD that it inherits from the Dirichlet distribution is the property of aggregation: If components are combined to create a coarser composition, the resulting process is also a DD. This makes it useful, for example, for jointly modeling the evolution of a microbiome grouping the microbe species at different taxonomic levels.
△ Less
Submitted 14 August, 2026;
originally announced August 2026.
-
Data-Driven Generation of Compact Quasi-Isodynamic Stellarators
Authors:
Yang Han,
Hanlin Chen. Shuai Cao,
Zhiyuan Lu,
Dehong Chen,
Guosheng Xu,
Baonian Wan
Abstract:
Stellarator design explores a vast space of three-dimensional plasma boundaries, only a small fraction of which yields usable equilibria. Data-driven models can narrow this search by learning from existing optimized configurations. Building on the ConStellaration database, we extend conditional boundary generation to four-field-period QI configurations, focusing on the sparsely sampled low-aspect-…
▽ More
Stellarator design explores a vast space of three-dimensional plasma boundaries, only a small fraction of which yields usable equilibria. Data-driven models can narrow this search by learning from existing optimized configurations. Building on the ConStellaration database, we extend conditional boundary generation to four-field-period QI configurations, focusing on the sparsely sampled low-aspect-ratio regime. The approach learns the common geometric structure of known QI equilibria and then adapts it using a small high-fidelity compact dataset, allowing target magnetic properties to guide generation beyond the original data distribution. This adaptation reduces the compact-domain test loss by approximately 87% and yields converged ultra-compact candidates consistent with the prescribed conditions. Several candidates show favorable confinement indicators, and one provides a useful seed for further QI optimization and finite-beta assessment. The method therefore serves as a data-informed front end to high-fidelity physics and optimization.
△ Less
Submitted 11 August, 2026;
originally announced August 2026.
-
Hierarchical Data Selection via Manifold Coverage and Sparse Feature Coverage in LLM Post-training
Authors:
Peng Sun,
Yi Yang,
Antong Zhang,
Chunxiao Li,
Yanbo Wang,
Dianbo Liu,
xin chen,
Kai Yu,
Lu Chen,
Tianfan Fu
Abstract:
As supervised fine-tuning data continues to scale, selecting high-value subsets from large candidate pools is crucial for reducing training cost and improving model performance. Existing methods often measure diversity directly in the original embedding space, where geometric metrics entangle dominant semantic directions, fine-grained supervision differences, and local noise. We address this limit…
▽ More
As supervised fine-tuning data continues to scale, selecting high-value subsets from large candidate pools is crucial for reducing training cost and improving model performance. Existing methods often measure diversity directly in the original embedding space, where geometric metrics entangle dominant semantic directions, fine-grained supervision differences, and local noise. We address this limitation by formulating data selection as a coarse-to-fine hierarchical coverage problem and propose MASS. MASS learns low-dimensional principal manifold coordinates with a dense autoencoder for coarse semantic grouping, and then performs quality-aware sparse feature coverage within each group using a TopK sparse autoencoder. Experiments on Vision Flan and LLaVA-CoT show that MASS consistently outperforms strong data selection baselines across multiple budgets, and in several settings matches or surpasses full data training with only a small subset of data.
△ Less
Submitted 5 August, 2026;
originally announced August 2026.
-
Data-DPO: Direct Preference Optimization for Target Model Data Selection in LLM Post-Training
Authors:
Peng Sun,
Yi Yang,
Antong Zhang,
Chunxiao Li,
Yanbo Wang,
Dianbo Liu,
xin chen,
Kai Yu,
Lu Chen,
Tianfan Fu
Abstract:
Data selection in supervised fine-tuning aims to select a small set of effective samples from large-scale candidate data, reducing training cost while preserving model performance. However, existing methods usually treat data value as a relatively static property, and pay limited attention to the compatibility between data and the capability distribution of the target model. To address this issue,…
▽ More
Data selection in supervised fine-tuning aims to select a small set of effective samples from large-scale candidate data, reducing training cost while preserving model performance. However, existing methods usually treat data value as a relatively static property, and pay limited attention to the compatibility between data and the capability distribution of the target model. To address this issue, we propose Data-DPO, a target model-oriented SFT data selection method. Data-DPO observes the local training feedback of the target model on different samples through one-step probing, transforms activation differences among samples into pairwise data preferences, and trains a lightweight reward model to learn target-model-aware data preferences. In the final selection stage, Data-DPO further combines target model preference, external quality scores, and marginal diversity to construct a more stable and effective training subset. Experimental results on Vision-Flan and LLaVA-CoT show that Data-DPO consistently outperforms existing data selection baselines under multiple data budgets and stably surpasses full data training performance.
△ Less
Submitted 5 August, 2026;
originally announced August 2026.
-
Adaptive Repulsive Pheromone Clustering for Foraging Robot Swarms
Authors:
Carlos Pena-Caballero,
Constantine Tarawneh,
Qi Lu
Abstract:
The Central Place Foraging Algorithm (CPFA) combines site fidelity, pheromone-guided navigation, and uninformed random search to enable decentralized resource collection in robot swarms. However, CPFA often revisits previously explored regions while leaving other areas insufficiently searched, reducing efficiency as resources become scarce. In this paper, we propose Adaptive Repulsive Pheromone Cl…
▽ More
The Central Place Foraging Algorithm (CPFA) combines site fidelity, pheromone-guided navigation, and uninformed random search to enable decentralized resource collection in robot swarms. However, CPFA often revisits previously explored regions while leaving other areas insufficiently searched, reducing efficiency as resources become scarce. In this paper, we propose Adaptive Repulsive Pheromone Clustering (ARPC), a bio-inspired method in which robots deposit repulsive pheromone waypoints to mark previously explored locations. These waypoints are clustered around the nest to estimate low-value search regions, allowing robots to be redirected toward likely unvisited areas. By integrating the exploitation of known resources with systematic avoidance of redundant exploration, ARPC improves search diversity and resource discovery efficiency. Extensive simulations in ARGoS across varying arena sizes, resource densities, and clustered, random, and power-law spatial distributions demonstrate that ARPC consistently outperforms CPFA and the Grid-Based CPFA (GPFA). In particular, ARPC yields significant gains during both early discovery (10\%) and late-stage (up to 60\%) collection, where conventional methods typically degrade. These results indicate that ARPC provides a scalable and robust strategy for large-scale heterogeneous swarm foraging environments.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Sharp $L^2$-Caffarelli--Kohn--Nirenberg and weighted Poincaré inequalities on half-spaces and orthants and their stability
Authors:
Nguyen Lam,
Yukta Lodha,
Guozhen Lu,
Ambar N. Sengupta
Abstract:
Though the sharp $L^{2}$-Caffarelli--Kohn--Nirenberg (CKN) inequalities have been extensively studied in the entire Euclidean spaces, the corresponding problem on domains whose boundary contains the origin remains largely unexplored. We investigate the sharp $L^{2}$-CKN inequalities on half-spaces and orthants $\mathbb R^{n}_{k,+}$ by computing explicitly the optimal constants, determining all pos…
▽ More
Though the sharp $L^{2}$-Caffarelli--Kohn--Nirenberg (CKN) inequalities have been extensively studied in the entire Euclidean spaces, the corresponding problem on domains whose boundary contains the origin remains largely unexplored. We investigate the sharp $L^{2}$-CKN inequalities on half-spaces and orthants $\mathbb R^{n}_{k,+}$ by computing explicitly the optimal constants, determining all possible extremal functions, and establishing exact identities for the deficits. Since the singular weights $|x|^{-2b}$ rule out the lifting argument that is available for the simpler Heisenberg Uncertainty Principle, we develop an approach based on the transformations $u(x)=|x|^{m}v(x)$ for an appropriately chosen $m$ combined with spherical harmonic decompositions and weighted identities. Moreover, we establish weighted Poincaré inequalities associated with measures of the form \[ e^{-δ|x|^τ}|x|^β\bigl(\prod_{i=n-k+1}^{n}x_i^{2}\bigr)\,dx, \] together with their sharp constants, extremizers and stability estimates, which substantially extend those of the classical Gaussian Poincaré inequality. On the full orthant, the linear modes cease to be admissible competitors, since all odd spherical harmonics are annihilated by the lifting; the first non-radial mode is then of degree two, and both the sharp constant and the manifold of optimizers change accordingly. Finally, we establish several stability estimates, and second-order stability estimates, of the CKN inequalities on the half-spaces and orthants throughout the full parameter range.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Steering the Flow: Inverting Face Recognition Models via Gradient-Guided Flow Matching
Authors:
Ye Lu,
Shen Wang,
Zhaoyang Zhang,
Yihan Yan,
Li Liu,
Runze Liu,
Fanghui Sun
Abstract:
Model Inversion Attacks (MIAs) aim to reconstruct representative training samples of target identities from face recognition models, exposing critical security vulnerabilities. Existing methods typically rely on indirect guidance or highly stochastic guidance, making it difficult to stably optimize generation trajectories toward target facial images. In this paper, we propose Steering Flow Model I…
▽ More
Model Inversion Attacks (MIAs) aim to reconstruct representative training samples of target identities from face recognition models, exposing critical security vulnerabilities. Existing methods typically rely on indirect guidance or highly stochastic guidance, making it difficult to stably optimize generation trajectories toward target facial images. In this paper, we propose Steering Flow Model Inversion (SFMI), a novel two-stage white-box model inversion method that reformulates inversion as a trajectory-steering task. Specifically, Step I, Learning a Generic Flow Matching Prior, pre-trains a generic unconditional Flow Matching model to encode the manifold of human faces as a robust prior. Step II, Attacking with Progressive Guidance Scheduler (PGS), injects time-dependent target-specific gradients during sampling. By backpropagating through the target model to obtain gradients from intermediate generated states, PGS progressively injects adaptive guidance signals into the vector field. This process effectively steers the current generative flow from random noise toward the high-density regions of the target class. Under an identity-disjoint cross-evaluation setting using the CelebA dataset, SFMI achieves an ACC of 0.9248, an FID of 22.61, and an LPIPS of 0.3874 on the ArcFace target. Extensive experiments on multiple target models demonstrate that SFMI achieves competitive state-of-the-art performance in attack success and visual fidelity under the evaluated white-box protocol.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
A note on the $\mho$-quiver
Authors:
Xue-Song Lu
Abstract:
A two-parameter hierarchy of classes of finitely generated modules over an artin algebra is described by using the $\mho$-quiver. As an application, it is shown that the stable category of reflexive modules is equivalent to two other categories via the mutually quasi-inverse equivalences induced by $Ω$ and $\mho$.
A two-parameter hierarchy of classes of finitely generated modules over an artin algebra is described by using the $\mho$-quiver. As an application, it is shown that the stable category of reflexive modules is equivalent to two other categories via the mutually quasi-inverse equivalences induced by $Ω$ and $\mho$.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
NP-LEAP: Nonparametric Latent Exchangeability Prior for Model-Lean Borrowing from Historical Data
Authors:
Ethan M. Alt,
Miheer Dewaskar,
Jacob M. Maronge,
Yuelin Lu,
Matthew A. Psioda
Abstract:
Bayesian dynamic borrowing (BDB) methods leverage historical data to reduce treatment effect uncertainty, yet existing approaches rely on parametric outcome models susceptible to misspecification. We propose the nonparametric latent exchangeability prior (NP-LEAP), an outcome-agnostic, assumption-lean framework to borrow information from historical data. The NP-LEAP performs individual-level excha…
▽ More
Bayesian dynamic borrowing (BDB) methods leverage historical data to reduce treatment effect uncertainty, yet existing approaches rely on parametric outcome models susceptible to misspecification. We propose the nonparametric latent exchangeability prior (NP-LEAP), an outcome-agnostic, assumption-lean framework to borrow information from historical data. The NP-LEAP performs individual-level exchangeability assessment, inducing Bayesian model averaging over all possible partitions of the historical data into exchangeable and nonexchangeable subsets. Although applicable to a variety of data types with choice of appropriate kernel, the NP-LEAP is particularly well-suited for studies with time-to-event outcomes, where parametric BDB is potentially triply misspecified - imposing a parametric baseline hazard, the proportional hazards structure, and blanket exchangeability. We establish posterior consistency under mild regularity conditions. Simulation studies demonstrate favorable operating characteristics relative to parametric borrowing methods and nonborrowing semiparametric frequentist methods. We illustrate the method by augmenting the control arm in a randomized trial of patients with non-small cell lung cancer.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Physics-Aligned Deep Learning Enables SERS Resolving and Sequencing of Dynamic Single-Molecule DNA Oligomers in Plasmonic Nanocavity
Authors:
Kuo Zhan,
Peilin Xin,
Yingqi Zhao,
Han Gu,
Enock Adjei Agyekum,
Zhou Chen,
Shuai Li,
Jian Ye,
Lu Cheng,
Jian-an Huang
Abstract:
Single-molecule surface-enhanced Raman spectroscopy (SM-SERS) captures dynamic molecular behavior with ultrahigh sensitivity, but its biopolymer analysis is hindered by strong spectral heterogeneity, transient hotspot sampling, and background interference. Here, we develop a physics-aligned deep learning framework integrating contrastive attention-based multiple-instance learning (CAMIL), a tri-ch…
▽ More
Single-molecule surface-enhanced Raman spectroscopy (SM-SERS) captures dynamic molecular behavior with ultrahigh sensitivity, but its biopolymer analysis is hindered by strong spectral heterogeneity, transient hotspot sampling, and background interference. Here, we develop a physics-aligned deep learning framework integrating contrastive attention-based multiple-instance learning (CAMIL), a tri-channel multi-kernel CNN classifier, and trajectory-level transition-guided sequence reconstruction to decode single-molecule DNA oligomer dynamics in a plasmonic nanocavity. In this work, CAMIL mines informative spectra enriched with chain-embedded nucleotide and dinucleotide signatures from contrastive positive and negative DNA trajectory bags, generating a single-molecule DNA-segment spectral-states library. This library trains a 12-class CNN to assign query time-resolved DNA SM-SERS frames to high-confidence DNA-segment states, enabling trajectory-level analysis of state composition, dwell length, entropy, switching frequency, and transition matrix. The transition matrix is further converted into state-transition-edge evidence for candidate sequence scoring, enabling asymmetric DNA sequences inference from confined stochastic sampling dynamics. This framework transforms SM-SERS heterogeneity into quantitative analytical information, advancing dynamic single-molecule decoding.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Bessel-Debiased Pseudo-Marginal MCMC for Generalised Bayesian Inference
Authors:
Yingkai Lu,
Jeong Eun Lee,
Geoff K. Nicholls
Abstract:
Generalized Bayesian inference uses weights of the form $\exp\{-β_n\ell_{n}(θ)\}$, even when the loss is available only through simulation, numerical integration, or subsampling. Exponentiating an unbiased loss estimate changes the target, and when $β_n\asymp n$ an ordinary Monte Carlo (MC) loss estimate with $M^{-1}$ variance needs a per-proposal budget of order $n^2$ to keep the leading log-weig…
▽ More
Generalized Bayesian inference uses weights of the form $\exp\{-β_n\ell_{n}(θ)\}$, even when the loss is available only through simulation, numerical integration, or subsampling. Exponentiating an unbiased loss estimate changes the target, and when $β_n\asymp n$ an ordinary Monte Carlo (MC) loss estimate with $M^{-1}$ variance needs a per-proposal budget of order $n^2$ to keep the leading log-weight variance bounded. We introduce the Sign-Corrected Bessel Debiasing (SCBD) algorithm, a signed pseudo-marginal method based on $K$ independent block estimates of the loss, and study its independently randomized Randomised Quasi-Monte-Carlo (RQMC) implementation. Under an iid Gaussian block model, a Bessel factor constructed from the block sample variance exactly removes the Gaussian exponential bias. For general finite MC or RQMC blocks, signed averages instead target a density proportional to $π_n(θ)w_M(θ)$, where $w_M$ is the finite-block target factor remaining after the signed correction; its posterior variation is assessed separately from sign efficiency. If an RQMC block estimator, based on $d$-dimensional randomized inputs, has variance $O\{B^{-α}(\log B)^{d-1}\}$, a sufficient budget for bounded leading log-weight variance has order $n^{2/α}$, up to logarithmic factors. We give a finite-block total-variation bound on the approximation and show the separate improvements from the debiasing correction and RQMC. The numerical examples show that, when the RQMC representation is favorable, the selected budget for stabilizing the estimated log weight can inherit this order, and that the Bessel correction can matter even when the RQMC variance rate is close to that of ordinary Monte Carlo.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Automatic Cephalometric Landmark Localization on CBCT-Derived Digitally Reconstructed Radiographs for Skeletal Malocclusion Classification
Authors:
Benjamin Hou,
Konstantinia Almpani,
Janice S. Lee,
Zhiyong Lu
Abstract:
Manual cephalometric landmark annotation is important for craniofacial assessment but is labor-intensive and difficult to scale. We introduce CephViT, a Vision Transformer-based model for automated 2D lateral cephalometric landmark localization, and evaluate its use in downstream skeletal malocclusion classification. CephViT was trained and benchmarked on a public lateral cephalogram dataset, achi…
▽ More
Manual cephalometric landmark annotation is important for craniofacial assessment but is labor-intensive and difficult to scale. We introduce CephViT, a Vision Transformer-based model for automated 2D lateral cephalometric landmark localization, and evaluate its use in downstream skeletal malocclusion classification. CephViT was trained and benchmarked on a public lateral cephalogram dataset, achieving a mean radial error of 1.28 +/- 1.42 mm and a successful detection rate of 92.0% at 3.0 mm. Because the private evaluation cohort consisted of 3D CBCT scans, lateral cephalogram-like digitally reconstructed radiographs (DRRs) were generated from each volume and used as 2D inputs to the landmark localization model. Landmark coordinates were normalized into a common coordinate frame, and skeletal malocclusion classification was performed using landmarks shared between the reference and DRR-based pipelines. Classification performance using DRR-localized landmarks was comparable to that obtained using manually annotated reference landmarks, with accuracies of 70.0% and 68.3%, respectively. These results support the feasibility of automated cephalometric analysis on CBCT-derived DRRs for skeletal malocclusion assessment.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
HiFi-BRep: High-Fidelity Latent Representation for Robust B-Rep Generation
Authors:
Junhao Hou,
Chenqi Luo,
Pufan Wang,
Jiaying Lu,
Yusheng Liu,
Feiwei Qin,
Meie Fang,
Kun Zhou
Abstract:
Boundary representation (B-Rep) generation is a fundamental task in computer-aided design, yet the direct synthesis of high-fidelity and structurally valid B-Reps remains a major challenge. Existing deep generative methods suffer from two forms of brittleness: representation brittleness, caused by padding noise and feature contamination in the latent space, and generation brittleness, stemming fro…
▽ More
Boundary representation (B-Rep) generation is a fundamental task in computer-aided design, yet the direct synthesis of high-fidelity and structurally valid B-Reps remains a major challenge. Existing deep generative methods suffer from two forms of brittleness: representation brittleness, caused by padding noise and feature contamination in the latent space, and generation brittleness, stemming from sequential error propagation and a train-inference mismatch due to non-differentiable validity enforcement. We propose HiFi-BRep, a novel framework that addresses these limitations through two synergistic contributions. First, a topology-aware encoder constructs a high-fidelity latent representation by eliminating padding via learnable queries and preventing feature contamination with topology-guided attention. Second, a single-stage decoder jointly predicts geometry and topology in parallel, embedding core manifold constraints as a differentiable learning objective. This design ensures mutual guidance between geometry and topology while avoiding cascaded errors. Extensive experiments show that HiFi-BRep significantly outperforms state-of-the-art methods in both structural validity and geometric fidelity, providing a robust solution for high-quality B-Rep synthesis. Code and models are publicly available at https://github.com/1nnoh/HiFi-BRep.
△ Less
Submitted 17 August, 2026; v1 submitted 17 August, 2026;
originally announced August 2026.
-
JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills
Authors:
Xiaoyu Wen,
Jiajia Li,
Zhida He,
Peng Yu,
Chenxu Wang,
Han Qi,
Ziyuan Zhou,
Cheng Jin,
Ying Wen,
Xingcheng Xu,
Shuyue Hu,
Tianhang Zheng,
Chaochao Lu,
Qiaosheng Zhang
Abstract:
Automated red-teaming has produced a growing collection of attack strategies, yet they typically remain scattered across prompts and workflows, making them difficult to systematically integrate, reuse, and improve at scale. We introduce \textsc{JailbreakSkill}, a skill-centric framework for scaling automated red-teaming through reusable and continuously evolving attack capabilities. \textsc{Jailbr…
▽ More
Automated red-teaming has produced a growing collection of attack strategies, yet they typically remain scattered across prompts and workflows, making them difficult to systematically integrate, reuse, and improve at scale. We introduce \textsc{JailbreakSkill}, a skill-centric framework for scaling automated red-teaming through reusable and continuously evolving attack capabilities. \textsc{JailbreakSkill} packages existing attack strategies into modular, agent-ready skills that can be directly reused and adaptively selected across tasks and target models. Beyond reuse, it closes the loop between attacking and learning: attack experience is used to diagnose, refine, combine, and discover new skills, which are added back to an ever-growing skill library. This evolution lifts macro-average ASR by 17.5 percentage points on AdvBench and 13.4 points on HarmBench, including a 48.6-point gain against GPT-5.4 on AdvBench, while yielding novel attack strategies such as reframing a direct request as an unfinished document-completion task. Several evolved skills also generalize to unseen prompts and target models without further adaptation. Our code is available at https://github.com/BattleWen/JailbreakSkill.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Field-controlled breaking and restoration of parity-time symmetry in Josephson interference
Authors:
Yi-Chen Tsai,
Yung-Yeh Chang,
Tao-Yi Hsu,
Thomas Kuo,
Chia-Nung Kuo,
Chin-Shan Lue,
Kuei-Lin Chiu,
Chen-Hsuan Hsu,
Chung-Ting Ke
Abstract:
Symmetry plays a fundamental role in determining the phases and physical properties of quantum matter. Controlling symmetry in mesoscopic superconducting devices provides a route to reconfigure their phase-coherent transport. Here we demonstrate symmetry-selective Josephson interferometry in lateral NbTi/PtTe2/NbTi junctions by controlling the relative orientations of the current and magnetic fiel…
▽ More
Symmetry plays a fundamental role in determining the phases and physical properties of quantum matter. Controlling symmetry in mesoscopic superconducting devices provides a route to reconfigure their phase-coherent transport. Here we demonstrate symmetry-selective Josephson interferometry in lateral NbTi/PtTe2/NbTi junctions by controlling the relative orientations of the current and magnetic field. From the supercurrent interference patterns, we construct a field-current symmetry map that identifies configurations exhibiting or violating the device-level parity (\mathcal{P}), time-reversal (\mathcal{T}) and their combined \mathcal{P}\mathcal{T} symmetry. In the absence of an in-plane field, the junction exhibits a symmetric Fraunhofer pattern. An in-plane field parallel to the current produces a pronounced side-lobe asymmetry, whereas reversing both the current and the complete magnetic-field configuration restores a generalized \mathcal{T} relation. Remarkably, orienting the in-plane field perpendicular to the current restores the \mathcal{P}\mathcal{T}-symmetric Fraunhofer response even at substantial field strengths. A microscopic model attributes this behavior to the interplay between disorder-induced potential variations and flux dipoles generated by in-plane-field Meissner focusing near the superconducting electrodes. Our results establish a reconfigurable Josephson interferometer in which the field-current geometry selects the symmetry operation being probed and switches the device between symmetry-broken and symmetry-restored interference states.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Step-Level On-Policy Distillation: Interpolating Between On-Policy Distillation and Supervised Fine-Tuning
Authors:
Changhui Sun,
Lanbo Liu,
Hang Lei,
Tong Ling,
Jiahang Xie,
Zhiyong Zheng,
Yujia Wang,
Hao Liu,
Feng Xiao,
Lu Liu,
Yanlong Du,
Zifeng Cheng,
Ziwei Jiang,
Qing Gu
Abstract:
On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories. This approach has achieved strong empirical gains and can often surpass conventional off-policy distillation with substantially less data. However, standard token-level OPD can provide only fragmented corrections along an erroneous student trajectory and cannot unfold a comple…
▽ More
On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories. This approach has achieved strong empirical gains and can often surpass conventional off-policy distillation with substantially less data. However, standard token-level OPD can provide only fragmented corrections along an erroneous student trajectory and cannot unfold a complete and correct repair path. Motivated by this limitation, we propose \emph{Step-Level On-Policy Distillation} (SOPD), which combines the long-horizon correction of supervised fine-tuning (SFT) with the on-policy advantage of OPD to provide step-level supervision over complete student-generated trajectories. We show that, at different limits of step length, SOPD reduces to SFT or approximates OPD. Compared with SFT, the teacher responses in SOPD are conditioned on student trajectories and therefore align more closely with student-visited states; compared with OPD, SOPD provides longer-horizon corrections rather than fragmented token-level guidance. Across both reasoning and agent tasks, SOPD substantially outperforms conventional SFT and OPD. For example, on ALFWorld, SOPD improves the average success rate by 13.4 points over Vanilla OPD. We hope this work offers a new perspective for future research on distillation methods.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Booster-based beam recycling for swap-out injection at the High Energy Photon Source
Authors:
Zhe Duan,
Jinhui Chen,
Yaoyao Du,
Yuanyuan Guo,
Jun He,
Xiyang Huang,
Daheng Jia,
Jingyi Li,
Fang Liu,
Peng Liu,
Zhi Liu,
Xiaohan Lu,
Yanhua Lu,
Cai Meng,
Yuemei Peng,
Saike Tian,
Guanwen Wang,
Jiuqing Wang,
Na Wang,
Yuanyuan Wei,
Gang Xu,
Haisheng Xu,
Yaliang Zhao,
Ying Zhao,
Yi Jiao
, et al. (1 additional authors not shown)
Abstract:
Fourth-generation synchrotron light sources employ ultralow-emittance storage rings with stringent injection requirements. On-axis swap-out injection alleviates the dependence on storage-ring dynamic aperture, but high-charge operation requires an efficient injector architecture capable of producing high-charge replacement bunches. This paper presents the accelerator physics design and performance…
▽ More
Fourth-generation synchrotron light sources employ ultralow-emittance storage rings with stringent injection requirements. On-axis swap-out injection alleviates the dependence on storage-ring dynamic aperture, but high-charge operation requires an efficient injector architecture capable of producing high-charge replacement bunches. This paper presents the accelerator physics design and performance analysis of a booster-based beam-recycling swap-out injection scheme implemented at the High Energy Photon Source (HEPS). In this approach, the full-energy booster serves as both an injector and a high-energy accumulator. An extracted storage-ring bunch is returned to the booster, merged with a low-charge bunch previously injected from the linac and accelerated to full energy. Following high-energy damping, the merged bunch is reinjected into the original storage-ring bucket. The scheme avoids the need for a dedicated accumulator ring while enabling high-charge bunch replacement. The recycling scheme was commissioned through staged machine studies. Full recycling-chain simulations, commissioning studies, and measured performance analysis are presented. The measured results characterize the recycling operation and quantify the transmission efficiency and performance limitations of the complete recycling loop. These results demonstrate the feasibility of the booster-based beam-recycling architecture and establish its operational basis for high-charge swap-out injection in future fourth-generation synchrotron light sources.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
First measurements of the branching fractions of $J/ψ$ and $ψ(3686) \to Σ^{0} \barΣ^{0}η$
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
C. S. Akondi,
R. Aliberti,
A. Amoroso,
Q. An,
Y. H. An,
Y. Bai,
O. Bakina,
H. R. Bao,
X. L. Bao,
M. Barbagiovanni,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko
, et al. (750 additional authors not shown)
Abstract:
Based on $(10087 \pm 44) \times 10^6$ $J/ψ$ and $(2712 \pm 14) \times 10^6$ $ψ(3686)$ events collected with the BESIII detector at the BEPCII collider, the hadronic decays $J/ψ\to Σ^{0} \barΣ^{0} η$ and $ψ(3686) \to Σ^{0} \barΣ^{0} η$ are observed for the first time. The corresponding branching fractions are measured to be…
▽ More
Based on $(10087 \pm 44) \times 10^6$ $J/ψ$ and $(2712 \pm 14) \times 10^6$ $ψ(3686)$ events collected with the BESIII detector at the BEPCII collider, the hadronic decays $J/ψ\to Σ^{0} \barΣ^{0} η$ and $ψ(3686) \to Σ^{0} \barΣ^{0} η$ are observed for the first time. The corresponding branching fractions are measured to be $\mathcal{B}(J/ψ\to Σ^{0} \barΣ^{0}η)= (7.5 \pm 0.3 \pm 0.8) \times 10^{-5}$ and $\mathcal{B}(ψ(3686) \to Σ^{0} \barΣ^{0}η)= (1.3\pm 0.1 \pm 0.1) \times 10^{-5}$, respectively, where the first uncertainties are statistical, and the second systematic. The ratio $\text{Q} \approx \frac{\mathcal{B}(ψ(3686) \to Σ^{0} \barΣ^{0} η)}{\mathcal{B}(J/ψ\to Σ^{0} \barΣ^{0} η)}$ is determined to be $(17.3 \pm 1.5 \pm 1.7)\%$, which is con sistent with the 12\%-rule within 3.0$σ$.~No significant intermediate states or threshold enhancements are observed in the $Σ^0$($\barΣ^{0}$)$η$ and $Σ^0$$\barΣ^{0}$ invariant mass spectra.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
An FFT-Accelerated Boundary Integral Equation Method for Wave Scattering by Smooth Surfaces in Three Dimensions
Authors:
Wenmao Hua,
Jun Lai,
Huiyi Li,
Wangtao Lu
Abstract:
For wave scattering by axisymmetric surfaces, the fast Fourier transform (FFT) method provides an effective tool to accelerate standard boundary integral equation (BIE) solvers. Surface integral equations can be decoupled into a series of curve integral equations on the generating curve, due to the convolution-like integral operators. The Fourier coefficients of the three-dimensional fundamental k…
▽ More
For wave scattering by axisymmetric surfaces, the fast Fourier transform (FFT) method provides an effective tool to accelerate standard boundary integral equation (BIE) solvers. Surface integral equations can be decoupled into a series of curve integral equations on the generating curve, due to the convolution-like integral operators. The Fourier coefficients of the three-dimensional fundamental kernels can be rapidly computed through three-term recurrence relations based on Miller's algorithm. Such well-established techniques break down for nonaxisymmetric surfaces.
This paper proposes a novel FFT-accelerated boundary integral method for wave scattering by smooth surfaces of arbitrary shapes. The Fourier coefficients of the singular kernels now satisfy higher-order recurrence relations. Although they can be solved with an optimal linear complexity by the standard Olver's algorithm, it turns out that a singularity swapping approach, that rewrites each kernel as the product of a smooth function and an axisymmetric-related singular factor, is realistically much faster. Consequently, Miller's algorithm together with the standard FFT convolution yields an ${\cal O}(M\log M)$ approach for evaluating the ${\cal O}(M)$ Fourier modes of the kernels, attaining exactly the same order of complexity for axisymmetric surfaces! With such FFT-based efficient procedures, we rewrite the surface integral equations in terms of ${\cal O}(M)$ weakly singular curve integrals, discretize them by panel-based generalized Gaussian quadratures, and obtain highly accurate linear systems to approximate the wavefields. Extensive numerical experiments are carried out to demonstrate the effectiveness of the new approach.
△ Less
Submitted 18 August, 2026; v1 submitted 17 August, 2026;
originally announced August 2026.
-
Beyond Asking: A Pipeline for Personalized Game Generation that Reads Players from Behavior
Authors:
Yifan Lu,
Xiaopeng Yuan,
Haohan Wang
Abstract:
Personalized game generation requires inferring a player's abilities and behavioral style from how they play. Large language models have made this inference more attainable than ever: an LLM can read a raw gameplay transcript and produce a fluent, plausible profile of the player. Plausible, however, is not verified, and verification is precisely what the field lacks: latent traits are unobservable…
▽ More
Personalized game generation requires inferring a player's abilities and behavioral style from how they play. Large language models have made this inference more attainable than ever: an LLM can read a raw gameplay transcript and produce a fluent, plausible profile of the player. Plausible, however, is not verified, and verification is precisely what the field lacks: latent traits are unobservable; questionnaires provide noisy proxies and become circular when self-reports are used to validate behavior-based inference; and behavior itself is ambiguous without context -- a player who never collects an item may not want it, or may never have had the chance. We address both problems. First, we construct a synthetic player population whose traits are ground truth by construction: each trait is an explicit bot parameter, accepted only after controlled manipulation produces consistent, trait-specific behavioral change. Unlike prior parameter-recovery work that inverts a known decision model, our benchmark evaluates policy-agnostic inference from behavioral transcripts alone. Second, we introduce an opportunity-aware decision-moment representation that disentangles preference from the chance to express it; ablating it selectively degrades opportunity-dependent traits. On this benchmark, few-shot LLM inference outperforms embedding- and rule-based baselines on most traits, though feature-based supervised regressors remain stronger overall. Finally, we close the loop: inferred profiles drive difficulty adaptation, evaluated against ground-truth references and mismatched-profile controls, and an exploratory human study examines whether these findings transfer to real players.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Unified Condition-Action Modeling for Accurate One-Step Action Generation
Authors:
Xinyu Zhou,
Zikun Cai,
Kuangji Zuo,
Gen Li,
Boyu Ma,
Yanshuo Lu,
Yutong Song,
Mingqi Yuan,
Jiayu Chen,
Jianfei Yang
Abstract:
Robot manipulation requires policies that are both accurate and efficient, as robot control must respond to changing observations under tight latency constraints. Recent diffusion and flow policies are promising, but they often treat conditions as auxiliary signals rather than jointly evolving them with action trajectories. We find that this limitation can be effectively mitigated by a \textbf{sim…
▽ More
Robot manipulation requires policies that are both accurate and efficient, as robot control must respond to changing observations under tight latency constraints. Recent diffusion and flow policies are promising, but they often treat conditions as auxiliary signals rather than jointly evolving them with action trajectories. We find that this limitation can be effectively mitigated by a \textbf{simple yet effective unified condition-action modeling design} that represents conditions and actions in a shared token space, allowing a compact model to achieve high performance while improving both inference speed and accuracy. Therefore, we propose UCA-Flow, a unified condition-action modeling framework for accurate one-step action generation. Our method unifies observation conditions, timestep conditions, interval conditions, and action tokens into a single sequence, and processes them with a Unified Condition-Action Transformer for joint condition-action representation learning. As a result, condition representations are dynamically reconstructed according to the current generation stage, highlighting information most relevant for action refinement. Furthermore, we introduce an improved dual-pass supervision scheme over $u$ and $v$ for stronger optimization of unified condition-action modeling. UCA-Flow improves the average success rate by 9.3 percentage points over the strongest baseline, while achieving $45.6\times$ and $33.4\times$ speedups over DP3 and Simple DP3, and remaining $4.3\times$ and $2.3\times$ faster than one-step FlowPolicy and MP1, respectively.
△ Less
Submitted 19 August, 2026; v1 submitted 17 August, 2026;
originally announced August 2026.
-
Equality Cases for the Face-Degree Majorization Theorem on Simplicial Complexes
Authors:
Yueli Han,
Lu Lu
Abstract:
The Grone--Merris--Bai theorem states that the Laplacian spectrum of a
simple graph is majorized by its conjugate degree sequence. Recently,
Zhang, Song, and Fan extended this result to simplicial complexes by
establishing a majorization relation between the spectrum of the
$(r-1)$-dimensional up-Laplacian and the conjugate $(r-1)$-degree sequence.
In this paper, we characterize all equa…
▽ More
The Grone--Merris--Bai theorem states that the Laplacian spectrum of a
simple graph is majorized by its conjugate degree sequence. Recently,
Zhang, Song, and Fan extended this result to simplicial complexes by
establishing a majorization relation between the spectrum of the
$(r-1)$-dimensional up-Laplacian and the conjugate $(r-1)$-degree sequence.
In this paper, we characterize all equality cases in the partial-sum
inequalities of this higher-dimensional majorization theorem. For every
$r$-dimensional simplicial complex $X$ with $r\ge2$, we prove that
\[
\sum_{i=1}^{q}λ_{r-1,i}(X)
=
\sum_{i=1}^{q}d_{r-1,i}^{\top}(X)
\]
if and only if
\[
q\ge \max\{\operatorname{rank}B_r(X),Δ_{r-1}(X)\}.
\]
Thus, unlike the graph case, equality can occur only after both sequences
have exhausted all their nonzero terms. As consequences, equality in the
first partial sum and equality between the entire sequences are both equivalent
to $X$ containing a unique $r$-simplex. The proof is based on the local
down-Laplacian decomposition and the equality case of the Ky Fan inequality.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Topologically Configurable Nonlinear Vortex Generation at van der Waals Heterostructures
Authors:
Hongwei Wang,
Yuda Wan,
Kai Wang,
Shuzheng Chen,
Hao Yan,
Xu Jiang,
Xiaodan Lyu,
Chang-Yin Ji,
Weibo Gao,
Peixiang Lu,
Guangwei Hu
Abstract:
van der Waals (vdW) materials offer a highly tunable and efficient platform at nanoscale for nonlinear and quantum optics. Twist-stacked vdW heterostructures enable elegant control of symmetry and interlayer coupling. Prior studies mainly focus on planar twisted interfaces, while neglecting the naturally formed and mandatory defects in such vdW heterostructures. Here, we demonstrate nonlinear sing…
▽ More
van der Waals (vdW) materials offer a highly tunable and efficient platform at nanoscale for nonlinear and quantum optics. Twist-stacked vdW heterostructures enable elegant control of symmetry and interlayer coupling. Prior studies mainly focus on planar twisted interfaces, while neglecting the naturally formed and mandatory defects in such vdW heterostructures. Here, we demonstrate nonlinear singular optics with topologically configurable nonlinear vortex generation at the corner singularity of vdW heterostructures. By tailoring azimuthally discrete second-harmonic phase gradients at each interface, we obtain programmable nonlinear vortex emitters with dominant target OAM components. Nonlinear OAM beams with topological charge $\ell = 1$ and $\ell = -2$ are experimentally realized, respectively. Our work unlocks the untapped potentials of nonlinear singular optics in twisted vdW materials as a reconfigurable and lithography-free platform for nonlinear structured light generation, important in quantum nonlinear optics and related fields.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Measurement of Branching Fraction and Transition Magnetic Moment of the Hyperon Dalitz Decay $Σ^0 \rightarrow Λe^+e^-$
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
R. Aliberti,
A. Amoroso,
Q. An,
Y. Bai,
O. Bakina,
Y. Ban,
H. -R. Bao,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. B. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko,
R. A. Briere,
A. Brueggemann,
H. Cai
, et al. (683 additional authors not shown)
Abstract:
Based on a data sample of 10 billion $J/ψ$ events collected with the BESIII detector operating at the BEPCII collider, the Dalitz decay $Σ^0 \rightarrow Λe^+e^-$ is studied experimentally for the first time. The $Σ^0$ hyperons are produced through the process $J/ψ\rightarrow Σ^0\barΣ^0$ and analyzed using a double-tag method. The absolute branching fraction is measured to be…
▽ More
Based on a data sample of 10 billion $J/ψ$ events collected with the BESIII detector operating at the BEPCII collider, the Dalitz decay $Σ^0 \rightarrow Λe^+e^-$ is studied experimentally for the first time. The $Σ^0$ hyperons are produced through the process $J/ψ\rightarrow Σ^0\barΣ^0$ and analyzed using a double-tag method. The absolute branching fraction is measured to be $\mathcal{B}(Σ^0 \rightarrow Λe^+e^-) = (6.34 \pm 0.25_{\rm stat.} \pm 0.23_{\rm syst.}) \times 10^{-3}$. This result shows a $2σ$ discrepancy from the theoretical calculation quoted in the PDG, where the uncertainties are statistical and systematic, respectively. In addition to the branching fraction, the transition magnetic moment $μ$ is determined to be $(1.74 \pm 0.03_{\rm stat.} \pm 0.09_{\rm syst.})\,μ_N$, where $μ_N=e/(2m_p)$ represents the nucleon magnetic moment, providing valuable insight into the intrinsic structure of the $Σ^0$ hyperon.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
Skill2Query: Exploiting Skill Structure to Generate Pseudo-Queries for Agent Skill Retrieval
Authors:
Lihui Ding,
Zihan Guo,
Bingwei Lu,
Chenyu Zhou,
Yuanjian Zhou,
Weinan Zhang,
Jianghao Lin,
Dongdong Ge
Abstract:
Pseudo-query generation can alleviate the supervision bottleneck for agent skill retrieval, but existing document-level approaches typically leave the rich internal relations among capabilities, parameters, and usage examples implicit. As a result, generated queries may be topically relevant to a skill while lacking capability grounding and parameter consistency, raising the question of whether ex…
▽ More
Pseudo-query generation can alleviate the supervision bottleneck for agent skill retrieval, but existing document-level approaches typically leave the rich internal relations among capabilities, parameters, and usage examples implicit. As a result, generated queries may be topically relevant to a skill while lacking capability grounding and parameter consistency, raising the question of whether explicitly exploiting a skill document's internal structure can produce more effective retrieval signals. We therefore propose Skill2Query, a framework that first parses a skill document into a Skill Knowledge Graph and then generates pseudo-queries through a three-stage process including style mimicking, query template generation, and parameter filling. The generated queries can be used for offline index augmentation, online query expansion, and retriever training. Four benchmarks (TheoremQA, LogicBench, ToolQA, and CHAMP) are used to evaluate Skill2Query with large-scale skill candidate pools across multiple downstream applications, including skill retrieval, retriever training, and end-to-end agent execution. Using nearly 30K skills across diverse domains, we generate 700K category-diverse pseudo-queries. Skill2Query consistently improves sparse, dense, and skill-routing retrieval, with an average Recall@1 gain of 6.70 percentage points across retrieval settings. Skill2Query-generated training data also achieves the best Recall@1 and nDCG@1 among the evaluated generation baselines. Further evaluations with multiple LLM backends demonstrate that improved skill retrieval translates into higher agent task success rates. Code and resources are available at https://github.com/MatZaharia/Skill2Query.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
SiMUSation: An Interactive Visitor Experience Simulation Framework to Support Museum Exhibition Design
Authors:
Huanchen Wang,
Qiuming Chen,
Zhonghao Ji,
Ruqi Sun,
Zhichao Lu,
Yuxin Ma
Abstract:
Understanding how diverse audiences engage with narratives and content is central to exhibition design, yet designers often rely on intuition. Existing experience evaluation methods are typically retrospective, costly, and offer limited access to visitors' internal states, hindering early-stage iterative refinement. Rather than relying only on post-implementation evaluation with real visitors, we…
▽ More
Understanding how diverse audiences engage with narratives and content is central to exhibition design, yet designers often rely on intuition. Existing experience evaluation methods are typically retrospective, costly, and offer limited access to visitors' internal states, hindering early-stage iterative refinement. Rather than relying only on post-implementation evaluation with real visitors, we explore LLM-driven persona simulation as a reference for early-stage design. Following this idea, we present SiMUSation, an interactive framework designed to support early-stage exhibition design. SiMUSation models diverse visitor personas and simulates their exhibition experiences through a dual-layer representation that couples observable behaviors, such as movement and gaze, with corresponding internal responses, such as confusion and narrative engagement. Designers can steer simulations, inspect feedback from simulated visits, and iteratively revise layouts, content, and narrative flow to further examine how changes reshape visitor experience. We implemented a prototype and evaluated it through a user study (N=12), showing that SiMUSation provides insights for reflection and refinement in early-stage exhibition design. Our findings further highlight the potential of persona-driven simulation to support audience-informed evaluation and iterative decision-making across design tasks.
△ Less
Submitted 16 August, 2026;
originally announced August 2026.
-
On the Local Linear Convergence of Operator Splitting Methods for Conic Programming
Authors:
Lijun Ding,
Haihao Lu,
Jinwen Yang
Abstract:
Operator-splitting methods such as the primal-dual hybrid gradient method (PDHG) and the alternating direction method of multipliers (ADMM) often exhibit linear convergence on conic programs, although general theory guarantees only sublinear rates. We identify two geometric conditions -- strict complementarity and quadratic facial violation -- that explain this local behavior: under these conditio…
▽ More
Operator-splitting methods such as the primal-dual hybrid gradient method (PDHG) and the alternating direction method of multipliers (ADMM) often exhibit linear convergence on conic programs, although general theory guarantees only sublinear rates. We identify two geometric conditions -- strict complementarity and quadratic facial violation -- that explain this local behavior: under these conditions, PDHG and ADMM converge linearly to an optimal solution when initialized sufficiently close to the converging strictly complementary solution. We establish this result through a unified and verifiable primal-dual error-bound framework. First, we show that strict complementarity, together with a quadratic facial-violation property of the associated complementary faces, implies uniform quadratic growth of both the primal and dual augmented Lagrangians near a strictly complementary solution. Second, we prove the local equivalence of three regularity conditions: uniform quadratic growth of the augmented Lagrangians, quadratic growth of a localized smoothed primal-dual gap, and metric subregularity of the saddle-point mapping. This equivalence clarifies the relationship among previously proposed conditions for local linear convergence. Third, using a unified formulation, we give a concise analysis showing that these equivalent conditions yield local linear convergence of PDHG and ADMM. We verify the quadratic facial-violation property for standard polyhedral and symmetric cones, as well as relevant faces of exponential and power cones, and show that it is preserved under Cartesian products. We also obtain an improved local rate using a restarted Halpern scheme. Finally, we extend the framework to convex composite optimization through a quadratic subdifferential-violation condition, which generalizes the quadratic facial-violation.
△ Less
Submitted 16 August, 2026;
originally announced August 2026.
-
TR-GS: High-Fidelity Sparse-View CT Volumetric Rendering via t-Distribution Gaussian Splatting and Ray-Confidence Modeling
Authors:
Zedong Xiao,
Yiren Wang,
Zhou Liu,
Xiaolin Liu,
Zhangji Lu
Abstract:
High-fidelity 3D medical visualization supports applications such as clinical assessment and surgical planning. Sparse-view computed tomography (CT) can reduce projection requirements and associated radiation exposure, but limited observations may introduce structural artifacts and reconstruction uncertainty. Although 3D Gaussian Splatting (3DGS) provides an efficient explicit representation for v…
▽ More
High-fidelity 3D medical visualization supports applications such as clinical assessment and surgical planning. Sparse-view computed tomography (CT) can reduce projection requirements and associated radiation exposure, but limited observations may introduce structural artifacts and reconstruction uncertainty. Although 3D Gaussian Splatting (3DGS) provides an efficient explicit representation for volumetric rendering, existing CT methods based on standard Gaussian primitives may be sensitive to unreliable observations under sparse-view acquisition. We present TR-GS, a Gaussian-splatting framework for sparse view CT volumetric rendering. TR-GS replaces standard Gaussian primitives with projectable Student's t-distribution primitives and introduces a ray-confidence model that regulates their degrees of freedom according to local ray observability. Confidence-guided 3D wavelet regularization is further used to balance high-frequency detail preservation and noise suppression. This work is licensed under a Creative Commons Attribution 4.0 International License. Experiments on synthetic and real-world datasets show that TR-GS improves over representative baselines in most evaluated settings and remains competitive in the remaining cases. The resulting volumetric representations may support downstream medical multimedia applications, including XR-based visualization and interactive clinical rendering.
△ Less
Submitted 16 August, 2026;
originally announced August 2026.
-
RagGAD: Rationale-Aware Conditional Gaussian Mixture Normalizing Flow for Unsupervised Graph Anomaly Detection
Authors:
Junxin Lu,
Jing Zhao,
Shiliang Sun
Abstract:
Graph anomaly detection aims to identify nodes that deviate from normal behavioral patterns within graphs. However, existing methods largely rely on the homophily assumption, which makes it difficult to distinguish spurious affinities and to capture the diverse behaviors of normal nodes,limiting their robustness in complex real-world scenarios. To address this problem, we propose RagGAD, an unsupe…
▽ More
Graph anomaly detection aims to identify nodes that deviate from normal behavioral patterns within graphs. However, existing methods largely rely on the homophily assumption, which makes it difficult to distinguish spurious affinities and to capture the diverse behaviors of normal nodes,limiting their robustness in complex real-world scenarios. To address this problem, we propose RagGAD, an unsupervised graph anomaly detection framework based on rationale-aware conditional Gaussian mixture normalizing flow. RagGAD introduces an adaptive rationale disentangler to disentangle stable rationales from spurious correlations within node interrelationships, and further decomposes stable rationales into robust and fragile components. The learned rationales capture underlying interaction patterns that characterize normal behaviors under varying conditions, while anomalies emerge as deviations associated with unstable or spurious correlations. To model the intricate distributions of normal and abnormal nodes, RagGAD integrates rationale-non-rationale Gaussian mixture modeling with a robust-fragile rationale mixture learning strategy. By mitigating spurious homophilic correlations and embracing the heterogeneity of normal patterns, RagGAD identifies anomalies as low-density regions within a structure-aware distribution space. Extensive experiments on multiple benchmark datasets demonstrate that RagGAD outperforms state-of-the-art methods.
△ Less
Submitted 16 August, 2026;
originally announced August 2026.
-
From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents
Authors:
Zhengzhao Ma,
Boxi Cao,
Yaojie Lu,
Hongyu Lin,
Xianpei Han,
Le Sun
Abstract:
Reliable uncertainty quantification (UQ) is essential for deploying large language model (LLM) agents in complex interactive environments. Existing UQ methods largely rely on local signals, such as token probabilities, predictive entropy, or per-step confidence, and therefore overlook the long-range dependencies through which errors accumulate across an execution trajectory. As a result, they may…
▽ More
Reliable uncertainty quantification (UQ) is essential for deploying large language model (LLM) agents in complex interactive environments. Existing UQ methods largely rely on local signals, such as token probabilities, predictive entropy, or per-step confidence, and therefore overlook the long-range dependencies through which errors accumulate across an execution trajectory. As a result, they may fail to identify agent failures whose causes originate several reasoning or interaction steps before the final answer. We propose RUPA (Relational Uncertainty Propagation for Agents), a trajectory-level UQ framework for LLM agents. RUPA represents an execution history as a directed trajectory graph in which reasoning states, tool interactions, and environment feedback are nodes connected by temporal and semantic dependency edges. It then propagates uncertainty over this graph to capture how execution risk accumulates and transfers across interaction steps. The propagated signal is combined with trajectory-level behavioral features and goal-alignment information to produce a confidence estimate for the full agent trajectory. We evaluate RUPA on representative agent benchmarks, including $τ$-2, Terminal-Bench-2, and GAIA, using 6 open-source LLMs spanning multiple model families. Experimental results show that RUPA consistently outperforms existing UQ methods by providing more accurate uncertainty estimates, enabling earlier failure detection, and improving uncertainty-guided agent execution across diverse agent tasks. These results demonstrate that explicitly modeling relational dependency is crucial to reliable UQ for long-horizon LLM agents, providing a practical foundation for trustworthy agent execution.
△ Less
Submitted 18 August, 2026; v1 submitted 16 August, 2026;
originally announced August 2026.
-
SEER: Long-Context Reasoning via Selective Visual-Text Compression
Authors:
Jiawei Xu,
Zhilin Zhai,
Jinrui Fang,
Ruohan Xu,
Mingfei Lu,
Yi Zhang,
Guanchu Wang,
Tianlong Chen,
Ying Ding
Abstract:
Long-context reasoning remains computationally expensive for large language models due to the quadratic complexity of attention over text tokens. Visual-text compression offers a promising alternative by rendering text into images and processing them with vision-language models, often reducing token usage. However, existing approaches apply uniform compression regardless of query relevance, potent…
▽ More
Long-context reasoning remains computationally expensive for large language models due to the quadratic complexity of attention over text tokens. Visual-text compression offers a promising alternative by rendering text into images and processing them with vision-language models, often reducing token usage. However, existing approaches apply uniform compression regardless of query relevance, potentially sacrificing precision where detailed extraction is required. We present SEER, a framework that learns to select query-relevant images through visual scanning and retrieve textual content only where needed, combining the efficiency of visual compression with the precision of text-based reasoning. Through supervised fine-tuning on tool-interaction trajectories, SEER learns adaptive tool invocation for selection and retrieval. Experiments on long-context benchmarks show that SEER improves extraction precision through selective text retrieval while retaining average prompt-token savings relative to full-text baselines. On LongBench, SEER achieves 51.11% average accuracy, outperforming the visual-text baseline Glyph-9B by 2.33 points and Qwen3-8B by 3.49 points. Code can be accessed at https://github.com/jiaweixu98/SEER
△ Less
Submitted 16 August, 2026;
originally announced August 2026.
-
COOL: A Cooling-Aware Point Transformer Framework for Thermal Prediction in Advanced 3D/3.5D IC Packaging
Authors:
Yao Lu,
Zhicheng Guo,
Qijun Zhang,
Shang Liu,
Wenji Fang,
Wenkai Li,
Zhiyao Xie
Abstract:
Advanced 3D and 3.5D IC packaging significantly improves integration density but elevates thermal management challenges due to cross-layer heat coupling and complex cooling structures. Traditional solvers deliver high fidelity but are too slow for iterative design flows, while existing learning-based methods either fail to capture inter-die thermal coupling or treat cooling structures as static co…
▽ More
Advanced 3D and 3.5D IC packaging significantly improves integration density but elevates thermal management challenges due to cross-layer heat coupling and complex cooling structures. Traditional solvers deliver high fidelity but are too slow for iterative design flows, while existing learning-based methods either fail to capture inter-die thermal coupling or treat cooling structures as static components, limiting their applicability in real packaging co-design scenarios. In this work, we introduce COOL, a cooling-aware point transformer framework that represents heterogeneous assemblies (dies, interposers, TIMs, heat spreaders) as annotated 3D point clouds embedding geometric, material and power attributes. COOL explicitly encodes geometric boundaries and cooling structures, and introduces a physics-informed boundary condition (PI-BC) loss to enforce thermal consistency at material interfaces and cooling boundaries. Extensive experiments demonstrate that COOL achieves a remarkable 2.4\% NMAE on our constructed benchmark of multi-package thermal designs, substantially outperforming existing learning-based approaches while providing over 15.7x speedup compared to commercial FEM solvers.
△ Less
Submitted 16 August, 2026;
originally announced August 2026.
-
Grouping Auction-Consensus Algorithm for Decentralized Task Allocation in Multi-Robot Systems
Authors:
Jose Rodriguez,
Sven Koenig,
Wenjie Dong,
Qi Lu
Abstract:
Decentralized multi-robot task allocation (MRTA) is essential for scalable and resilient autonomous systems. The Consensus-Based Bundle Algorithm (CBBA) is a widely adopted decentralized baseline. However, its individual task-level bidding is poorly aligned with the min-sum objective of minimizing total team travel distance, leading to suboptimal allocations in spatially distributed environments.…
▽ More
Decentralized multi-robot task allocation (MRTA) is essential for scalable and resilient autonomous systems. The Consensus-Based Bundle Algorithm (CBBA) is a widely adopted decentralized baseline. However, its individual task-level bidding is poorly aligned with the min-sum objective of minimizing total team travel distance, leading to suboptimal allocations in spatially distributed environments. This paper introduces the Grouping Auction-Consensus Algorithm (GACA). This decentralized MRTA framework adopts the two-phase auction-consensus architecture of CBBA while fundamentally redesigning its bidding mechanism to reason over groups of spatially proximate tasks. A nearest-neighbor preprocessing step partitions tasks into spatially coherent groups before allocation. Agents then iteratively propose structured group-level actions: claiming unassigned groups, acquiring partial groups, or contesting groups held by other agents. Competing actions are resolved through a consensus phase. Operating in the MT-SR-IA problem class, GACA is evaluated against CBBA using a Mixed-Integer Linear Program as the ground-truth optimality reference. Across four swarm sizes and 4,000 test worlds, GACA achieves a median percent optimality of approximately 97% compared to 81--84% for CBBA, while converging in equal or fewer iterations. A scalability evaluation over 3,280 additional problem instances spanning swarm sizes of 5 to 20 agents and task counts of 10 to 50 confirms that these gains generalize robustly across a wide range of problem configurations.
△ Less
Submitted 16 August, 2026;
originally announced August 2026.
-
Dear Algo: A Precision-First Agentic Intent Layer for Unified Search and Recommendation
Authors:
Rui Wang,
Jiazhou Wang,
Zheng Wei,
Chenglin Lu,
Fangcheng Sun,
Ivy Sun,
Jin Sun,
Hui Geng,
Lillian Zhang,
Chao Yang,
Lei Chen,
Shahin Sefati,
Reem Helou,
Joe Zhou,
Babak Shakibi,
Yiyi Pan,
Bi Xue,
Hong Yan,
Shujian Bu
Abstract:
Search and recommendation serve a shared discovery objective but encode intent differently. We study this boundary through Dear Algo on Threads, a deployed product where open-ended requests such as \emph{more NBA news} or \emph{less politics} steer subsequent feed recommendations rather than return a one-shot result list. Its agentic intent layer compiles explicit, inferred, negative, and compound…
▽ More
Search and recommendation serve a shared discovery objective but encode intent differently. We study this boundary through Dear Algo on Threads, a deployed product where open-ended requests such as \emph{more NBA news} or \emph{less politics} steer subsequent feed recommendations rather than return a one-shot result list. Its agentic intent layer compiles explicit, inferred, negative, and compound intent into a grounded executable plan, then invokes conventional retrieval and optional semantic or multimodal reranking. The layer shares an intent-to-retrieval contract without requiring one model or serving path across search-like and recommendation-like modes.
We evaluate Dear Algo under a precision-first objective. In a blinded audit of 300 public request-item pairs (296 evaluable), a strict categorical LLM-as-a-judge gate achieved 94.4\% exact-Relevant precision [88.8\%, 98.9\%]. Across 72 normalized request clusters, the full configuration produced 7.73 judge-qualified candidates per 20 slots versus 6.61 for an LLM-derived-query baseline, a gain of 1.11 [0.12, 2.12]. In a candidate-randomized serving-path study restricted to the reranker path's first 72 eligible hours, the user-weighted judge-Irrelevant share among judged admissions was 2.80\% versus 4.78\% off (-1.97 points [-3.02, -0.94]), while Exact-Relevant share was 2.24 points higher [0.08, 4.41].
Together, these studies show how explicit natural-language intent can be carried into feed recommendation under a precision-first evaluation framework
△ Less
Submitted 16 August, 2026;
originally announced August 2026.
-
GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture
Authors:
GigaBrain Team,
Angen Ye,
Axiang Sun,
Can Jin,
Chenxi Cheng,
Chong Shi,
Dengke Shang,
Dingqian Zhang,
Guan Huang,
Guangqiang Wang,
Guangqing Ding,
Guo Li,
Hangcong Li,
Hengyu Zhong,
Hongtao Lu,
Jianbo Qin,
Jiming Mao,
Jing Zhu,
Jindi Lv,
Jingzhi Cui,
Junjie Xie,
Junyi Bao,
Kai Liu,
Lei Yuan,
Limin Long
, et al. (34 additional authors not shown)
Abstract:
Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit from more effective architectural design, scale to substantially larger and more heterogeneous data regimes, and achieve broader generalizatio…
▽ More
Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit from more effective architectural design, scale to substantially larger and more heterogeneous data regimes, and achieve broader generalization across tasks and embodiments. To this end, we present GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments. Specifically, GigaBrain-0.7 unifies understanding, prediction, and action through a three-system architecture, scales pretraining to over 37,000 hours of heterogeneous embodied data, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation. Compared with the preceding GigaBrain-0 series and prior state-of-the-art models including $π_{0.5}$, GigaBrain-0.7 achieves substantial improvements in foundation zero-shot capabilities, language-conditioned instruction following, and post-training task success rates. In particular, on our in-house Maker H01 platform and mainstream robot embodiments, GigaBrain-0.7 demonstrates strong task adaptability and completion ability across both home and industrial scenarios. All training code and pretrained model weights will be released.
△ Less
Submitted 16 August, 2026;
originally announced August 2026.