-
SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation
Authors:
Jinsheng Quan,
Jianhua Li,
Siyi Xie,
Xuanke Shi,
Kewang Deng,
Zukai Chen,
Feifei Shao,
Lei Yang,
Quan Wang,
Yawei Luo
Abstract:
Spatial perception and reasoning from visual observations require recovering geometric structure, establishing correspondences, and understanding spatial relations. Existing approaches typically address these capabilities separately using task-specific architectures or external geometric modules, limiting knowledge transfer among complementary representations of the same physical scene. We introdu…
▽ More
Spatial perception and reasoning from visual observations require recovering geometric structure, establishing correspondences, and understanding spatial relations. Existing approaches typically address these capabilities separately using task-specific architectures or external geometric modules, limiting knowledge transfer among complementary representations of the same physical scene. We introduce SPARGen, a unified multimodal framework that casts 3D reconstruction, dense correspondence, and spatial reasoning as instruction-conditioned generation tasks. SPARGen serializes compact structured and linguistic outputs as token sequences while generating dense geometric fields in image-aligned forms, enabling spatial supervision to jointly shape shared representations within a native multimodal generative model. Experiments across benchmarks for 3D reconstruction, correspondence, and spatial reasoning show that SPARGen achieves competitive performance across heterogeneous spatial tasks within a single native multimodal generative framework.
△ Less
Submitted 14 August, 2026;
originally announced August 2026.
-
Symmetry Rules for Cavity Materials Engineering with Linearly Polarized Vacuum Fields
Authors:
Jingkai Quan,
Chongxiao Fan,
Benshu Fan,
I-Te Lu,
Dante M. Kennes,
Angel Rubio
Abstract:
Cavity materials engineering, aiming to manipulate material properties by coupling to vacuum fluctuations inside a cavity, is a rapidly advancing field. Despite significant progress, most studies to date have focused on specific materials and cavity configurations. Here, through a comprehensive group-theoretical analysis, we establish general symmetry rules for cavity materials engineering with li…
▽ More
Cavity materials engineering, aiming to manipulate material properties by coupling to vacuum fluctuations inside a cavity, is a rapidly advancing field. Despite significant progress, most studies to date have focused on specific materials and cavity configurations. Here, through a comprehensive group-theoretical analysis, we establish general symmetry rules for cavity materials engineering with linearly polarized cavity photon modes. By analyzing the symmetry of the effective photon-free quantum-electrodynamics Hamiltonian, we provide a complete classification of the symmetry-breaking patterns induced by cavity modes for all crystallographic point groups. The power of this framework is then demonstrated by quantum-electrodynamical density functional theory calculations. In particular, we explain the distinct cavity-induced lifting of band degeneracies in cubic BaTiO$_3$ for different cavity mode configurations, and the cavity-modified infrared and Raman spectra of monolayer MoS$_2$ due to symmetry breaking. Our results highlight the central role of symmetry in cavity materials engineering and provide general guidelines for future studies in this field.
△ Less
Submitted 30 July, 2026;
originally announced July 2026.
-
Macroscopic Polarization and Magnetization from Cavity Vacuum Fluctuations
Authors:
Jingkai Quan,
Chongxiao Fan,
Benshu Fan,
I-Te Lu,
Dante M. Kennes,
Angel Rubio
Abstract:
Cavity light-matter interaction has recently emerged as a new avenue for manipulating material properties without driving fields. Here, we demonstrate that cavity vacuum fluctuations can induce macroscopic polarization (magnetization), even in materials that lack spontaneous polarization (net magnetization) in free space. Starting from the effective photon-free quantum-electrodynamics Hamiltonian,…
▽ More
Cavity light-matter interaction has recently emerged as a new avenue for manipulating material properties without driving fields. Here, we demonstrate that cavity vacuum fluctuations can induce macroscopic polarization (magnetization), even in materials that lack spontaneous polarization (net magnetization) in free space. Starting from the effective photon-free quantum-electrodynamics Hamiltonian, we identify all crystallographic (magnetic) point groups that allow such cavity-induced responses. We derive the form of the corresponding response tensors based on symmetry analysis, whose elements can be obtained by quantum electrodynamical density functional theory (QEDFT) calculations. As representative examples, we show that the cavity-induced polarization in $α$-quartz can be continuously controlled by rotating the cavity. For antiferromagnetic Mn$_3$Sn, we demonstrate that cavity-induced symmetry breaking generates an out-of-plane magnetization, accompanied by an anomalous Hall conductivity component that is forbidden outside the cavity. Our work establishes symmetry as a guiding principle for cavity materials engineering and provides a route for controlling polarization and magnetization through quantum vacuum fluctuations, i.e., cavity materials engineering.
△ Less
Submitted 21 July, 2026;
originally announced July 2026.
-
Symmetry-Based Microscopic Theory of the Unconventional Pairing Mechanism in La$_5$Ni$_3$O$_{11}$
Authors:
Guan-Hao Feng,
Jun Quan
Abstract:
Recent experiments report high-temperature superconductivity in the hybrid nickelate $\mathrm{La}_5\mathrm{Ni}_3\mathrm{O}_{11}$, which is composed of alternating stacks of bilayer $\mathrm{La}_3\mathrm{Ni}_2\mathrm{O}_7$ and monolayer $\mathrm{La}_2\mathrm{NiO}_4$. However, the superconducting transition temperature $T_c \approx 64~\mathrm{K}$ for $\mathrm{La}_5\mathrm{Ni}_3\mathrm{O}_{11}$ is re…
▽ More
Recent experiments report high-temperature superconductivity in the hybrid nickelate $\mathrm{La}_5\mathrm{Ni}_3\mathrm{O}_{11}$, which is composed of alternating stacks of bilayer $\mathrm{La}_3\mathrm{Ni}_2\mathrm{O}_7$ and monolayer $\mathrm{La}_2\mathrm{NiO}_4$. However, the superconducting transition temperature $T_c \approx 64~\mathrm{K}$ for $\mathrm{La}_5\mathrm{Ni}_3\mathrm{O}_{11}$ is remarkably lower than the $80~\mathrm{K}$ observed for pressurized $\mathrm{La}_3\mathrm{Ni}_2\mathrm{O}_7$. Thus, an unified microscopic theory is required to address the difference in the pairing mechanisms between these systems. Here, we develop a phenomenological symmetry-based approach to systematically analyze the low-energy physics in $\mathrm{La}_5\mathrm{Ni}_3\mathrm{O}_{11}$, which is obtained by a charge self-consistent density functional theory plus dynamical mean-field theory method. We show that the superconductivity in $\mathrm{La}_5\mathrm{Ni}_3\mathrm{O}_{11}$ exhibits a two-gap nature, consisting of a leading interlayer pairing between the $d_{z^2}$ orbitals and a subleading intralayer pairing between the $d_{x^2-y^2}$ orbitals. The reduction of $T_c$ can be attributed to the diminished contribution of the interlayer pairing, as reflected by the hopping parameter ratio $|t_{\perp}^z/t_{\parallel}^{x}|$. Base on this unified picture, we discuss the possible pairing mechanism and the role of $γ$ pocket for the superconductivity in the bilayer NiO$_2$ planes of nickelate superconductors.
△ Less
Submitted 21 July, 2026; v1 submitted 20 July, 2026;
originally announced July 2026.
-
Vision as Unified Multimodal Generation
Authors:
Xiaoyang Han,
Jianhua Li,
Kewang Deng,
Zukai Chen,
Xuanke Shi,
Sihan Wang,
Boxuan Li,
Linyan Wang,
Siyi Xie,
Xin You,
Jinsheng Quan,
Zhongang Cai,
Haiwen Diao,
Ziwei Liu,
Lei Yang,
Dahua Lin,
Quan Wang
Abstract:
We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal model, without task-specific architectures. Under this formulation, SenseNova-Vision uses natural-language instructions and optional visual prompts to specify tasks, target regions or views, and decoding conventions, an…
▽ More
We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal model, without task-specific architectures. Under this formulation, SenseNova-Vision uses natural-language instructions and optional visual prompts to specify tasks, target regions or views, and decoding conventions, and generates responses as text for symbolic outputs, images for dense spatial predictions, or mixed text-and-image outputs for compositional tasks. To support large-scale training, we convert diverse computer vision annotations into instruction-response examples compatible with these generation spaces, resulting in the SenseNova-Vision Corpus, a computer-vision instruction-response corpus spanning text, image, and mixed targets. Starting from an off-the-shelf pretrained unified multimodal model, SenseNova-Vision is trained primarily on this corpus, with auxiliary multimodal data used as a capability-preserving mixture, and requires no task-specific prediction heads or architectural modifications. The resulting model covers a broad range of vision tasks, including detection, OCR, keypoint estimation, segmentation, depth estimation, surface normal prediction, point maps, and camera pose estimation, while supporting language-defined variants that combine category, color, region, and other visual cues. Experiments show that a single unified model can match leading task-specialized systems across structured visual understanding, dense geometric prediction, segmentation, and multi-view visual geometry. These results suggest unified multimodal generation as a scalable route for integrating computer vision capabilities into general-purpose foundation models. The model and corpus are publicly available.
△ Less
Submitted 7 July, 2026;
originally announced July 2026.
-
Schema-Agnostic Process Trace Construction: From Raw Tables to Execution Behavior
Authors:
Joel Lim Zhi Quan,
Tan Kar Way,
Lau Hoong Chuin
Abstract:
Traditional information systems (IS) engineering assumes stable schemas, explicit keys, and curated event logs. In modern OLTP environments, schemas drift, keys are sparse, and execution traces are dispersed across loosely connected tables, making manual process trace construction costly and error prone. We propose a schema-agnostic pipe-line that automatically reconstructs process execution trace…
▽ More
Traditional information systems (IS) engineering assumes stable schemas, explicit keys, and curated event logs. In modern OLTP environments, schemas drift, keys are sparse, and execution traces are dispersed across loosely connected tables, making manual process trace construction costly and error prone. We propose a schema-agnostic pipe-line that automatically reconstructs process execution traces directly from raw relational data. The pipeline (i) identifies columns that function like keys or timestamps, (ii) discovers table-to-table connections using statistical signals rather than predefined schemas, (iii) assembles and orders events for each case while accommodating multiple date fields, and (iv) learns likely ordering and flow relations across systems using a Temporal Convolutional Network which models long-range dependencies and patterns. Evaluations on TPC-H/E benchmarks, synthetic corpora, and a real industry dataset show that our pipeline reconstructs high-fidelity event traces and accurate trace orderings, correctly predicting the next event with 85% accuracy and recovering about 82% of ground-truth precedence relations. By eliminating dependence on predefined schemas, ER diagrams and domain templates, our work offers a generalizable and scalable pathway for automated reconstruction of execution behaviour in dynamic and continuously evolving IS environments.
△ Less
Submitted 9 June, 2026;
originally announced June 2026.
-
Same Evidence, Different Answers: Canonical-Context On-Policy Distillation for Multi-Turn Language Models
Authors:
Zizhuo Lin,
Quanling Liu,
Jinsheng Quan,
Chao Zhang,
Yifan Zhu,
Xing Shi,
Jingtao Xu,
Zhihui Li,
Yawei Luo
Abstract:
Large language models (LLMs) often solve a task when all instructions are given in a single prompt, but fail when the same information is revealed gradually across turns. When a clean FULL prompt and a RAW-SHARDED conversation contain the same complete user evidence, the model should still arrive at the same answer. We argue that a key reason for this gap is self-anchored drift: responses produced…
▽ More
Large language models (LLMs) often solve a task when all instructions are given in a single prompt, but fail when the same information is revealed gradually across turns. When a clean FULL prompt and a RAW-SHARDED conversation contain the same complete user evidence, the model should still arrive at the same answer. We argue that a key reason for this gap is self-anchored drift: responses produced under partial information introduce unsupported assumptions, and those assumptions later distort the final answer. To reduce this effect, we propose Canonical-Context On-Policy Distillation (CCOPD). During training, the same base model is used in two roles: a frozen teacher conditioned on the clean FULL prompt and a trainable student that receives the same evidence incrementally through a multi-turn conversation; CCOPD aligns the student's behavior on its own trajectories with the teacher's canonical full-context behavior. Trained only on math problem conversations, CCOPD yields a 32\% average relative improvement in RAW-SHARDED performance over the original base model across math and five zero-shot out-of-domain task families, while largely preserving full-context performance. Further analyses suggest that CCOPD strengthens grounding in user evidence and reduces sensitivity to contamination from earlier assistant turns.
△ Less
Submitted 28 May, 2026;
originally announced May 2026.
-
Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives
Authors:
Konstantinos Oikonomidis,
Jan Quan,
Kimon Antonakopoulos,
Antonio Silveti-Falls,
Volkan Cevher,
Panagiotis Patrinos
Abstract:
In this work, we develop proximal preconditioned gradient methods with a focus on spectral gradient methods providing a proximal extension to the Muon and Scion optimizers. We introduce a family of stochastic algorithms that can handle a wide variety of convex and nonconvex constraints and study its convergence under heavy-tailed noise, through a novel analysis tailored to the geometry of the prop…
▽ More
In this work, we develop proximal preconditioned gradient methods with a focus on spectral gradient methods providing a proximal extension to the Muon and Scion optimizers. We introduce a family of stochastic algorithms that can handle a wide variety of convex and nonconvex constraints and study its convergence under heavy-tailed noise, through a novel analysis tailored to the geometry of the proposed methods. We further propose a variance-reduced version, which achieves faster convergence under standard noise assumptions. Finally, we show that the polynomial iterations used in Muon are more accurately captured by a nonlinear preconditioner than by the ideal matrix sign, leading to a convergence analysis that more faithfully reflects practical implementations.
△ Less
Submitted 12 May, 2026;
originally announced May 2026.
-
MDM-Prime-v2: Binary Encoding and Index Shuffling Enable Scaling of Diffusion Language Models
Authors:
Chen-Hao Chao,
Wei-Fang Sun,
Junwei Quan,
Chun-Yi Lee,
Rahul G. Krishnan
Abstract:
Masked diffusion models (MDM) exhibit superior generalization when learned using a Partial masking scheme (Prime). This approach converts tokens into sub-tokens and models the diffusion process at the sub-token level. We identify two limitations of the MDM-Prime framework. First, we find that the functional form of the subtokenizer significantly increases the cross-entropy loss in the objective wh…
▽ More
Masked diffusion models (MDM) exhibit superior generalization when learned using a Partial masking scheme (Prime). This approach converts tokens into sub-tokens and models the diffusion process at the sub-token level. We identify two limitations of the MDM-Prime framework. First, we find that the functional form of the subtokenizer significantly increases the cross-entropy loss in the objective when paired with commonly used Byte-Pair-Encoding (BPE) tokenizers. Second, we lack tools to guide the hyperparameter choice of the token granularity in the subtokenizer. To address these limitations, we analyze the optimal design of the subtokenizer that minimizes MDM-Prime training objective and develop MDM-Prime-v2, a masked diffusion language model which incorporates Binary Encoding and Index Shuffling. Our analysis characterizes how token granularity and sub-token entropy influence the training objective and downstream performance, providing principled criteria for subtokenizer design. When extending the model size to 1.1B parameters, MDM-Prime-v2 demonstrates superior average zero-shot accuracy across eight commonsense reasoning benchmarks, outperforming similar-sized baselines including GPT-Neo, OPT, Pythia, Bloom, SMDM, and TinyLLaMA.
△ Less
Submitted 20 May, 2026; v1 submitted 16 March, 2026;
originally announced March 2026.
-
Learning to Generate Secure Code via Token-Level Rewards
Authors:
Jiazheng Quan,
Xiaodong Li,
Bin Wang,
Guo An,
Like Liu,
Degen Huang,
Lin Liu,
Chengbin Hou
Abstract:
Large language models (LLMs) have demonstrated strong capabilities in code generation, yet they remain prone to producing security vulnerabilities. Existing approaches commonly suffer from two key limitations: the scarcity of high-quality security data and coarse-grained reinforcement learning reward signals. To address these challenges, we propose Vul2Safe, a new secure code generation framework…
▽ More
Large language models (LLMs) have demonstrated strong capabilities in code generation, yet they remain prone to producing security vulnerabilities. Existing approaches commonly suffer from two key limitations: the scarcity of high-quality security data and coarse-grained reinforcement learning reward signals. To address these challenges, we propose Vul2Safe, a new secure code generation framework that leverages LLM self-reflection to construct high-confidence repair pairs from real-world vulnerabilities, and further generates diverse implicit prompts to build the PrimeVul+ dataset. Meanwhile, we introduce SRCode, a novel training framework that pioneers the use of token-level rewards in reinforcement learning for code security, which enables the model to continuously attend to and reinforce critical fine-grained security patterns during training. Compared with traditional instance-level reward schemes, our approach allows for more precise optimization of local security implementations. Extensive experiments show that PrimeVul+ and SRCode substantially reduce security vulnerabilities in generated code while improving overall code quality across multiple benchmarks.
△ Less
Submitted 26 February, 2026;
originally announced February 2026.
-
City Editing: Hierarchical Agentic Execution for Dependency-Aware Urban Geospatial Modification
Authors:
Rui Liu,
Steven Jige Quan,
Zhong-Ren Peng,
Zijun Yao,
Han Wang,
Zhengzhang Chen,
Kunpeng Liu,
Yanjie Fu,
Dongjie Wang
Abstract:
As cities evolve over time, challenges such as traffic congestion and functional imbalance increasingly necessitate urban renewal through efficient modification of existing plans, rather than complete re-planning. In practice, even minor urban changes require substantial manual effort to redraw geospatial layouts, slowing the iterative planning and decision-making procedure. Motivated by recent ad…
▽ More
As cities evolve over time, challenges such as traffic congestion and functional imbalance increasingly necessitate urban renewal through efficient modification of existing plans, rather than complete re-planning. In practice, even minor urban changes require substantial manual effort to redraw geospatial layouts, slowing the iterative planning and decision-making procedure. Motivated by recent advances in agentic systems and multimodal reasoning, we formulate urban renewal as a machine-executable task that iteratively modifies existing urban plans represented in structured geospatial formats. More specifically, we represent urban layouts using GeoJSON and decompose natural-language editing instructions into hierarchical geometric intents spanning polygon-, line-, and point-level operations. To coordinate interdependent edits across spatial elements and abstraction levels, we propose a hierarchical agentic framework that jointly performs multi-level planning and execution with explicit propagation of intermediate spatial constraints. We further introduce an iterative execution-validation mechanism that mitigates error accumulation and enforces global spatial consistency during multi-step editing. Extensive experiments across diverse urban editing scenarios demonstrate significant improvements in efficiency, robustness, correctness, and spatial validity over existing baselines.
△ Less
Submitted 26 February, 2026; v1 submitted 22 February, 2026;
originally announced February 2026.
-
Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability
Authors:
Shobhita Sundaram,
John Quan,
Ariel Kwiatkowski,
Kartik Ahuja,
Yann Ollivier,
Julia Kempe
Abstract:
RL methods for scaling large reasoning models stall on datasets with low initial success rates, and thus little training signal. We investigate a fundamental question: Can a pretrained LLM leverage latent knowledge to generate an automated curriculum for problems it cannot solve? We explore this with SOAR: An asymmetric self-play framework that uses meta-RL to surface these pedagogical signals. A…
▽ More
RL methods for scaling large reasoning models stall on datasets with low initial success rates, and thus little training signal. We investigate a fundamental question: Can a pretrained LLM leverage latent knowledge to generate an automated curriculum for problems it cannot solve? We explore this with SOAR: An asymmetric self-play framework that uses meta-RL to surface these pedagogical signals. A teacher model proposes synthetic problems for a student model, and is rewarded with its improvement on a subset of hard problems, thus grounding the curriculum in real student progress rather than intrinsic proxy rewards. Our study on the hardest subsets of math benchmarks (0/128 success) reveals three core findings. First, it is possible to realize bilevel meta-RL that unlocks learning under sparse, binary rewards by sharpening a latent capacity of pretrained models to generate useful problems. Second, grounded rewards outperform intrinsic learnability rewards used in prior LLM self-play, reliably avoiding typical instability and diversity collapse modes. Third, the structure and well-posedness of questions are more critical for learning progress than solution correctness. Our results suggest that the ability to generate useful stepping stones does not require the preexisting ability to solve the hard problems, paving a principled path to escape reasoning plateaus without additional curated data
△ Less
Submitted 30 June, 2026; v1 submitted 26 January, 2026;
originally announced January 2026.
-
First Submillimeter Lights from Dome A: Tracing the Carbon Cycle in the Feedback of Massive Stars
Authors:
Yan Gong,
Jiaqiang Zhong,
Yuan Ren,
Yilong Zhang,
Daizhong Liu,
Yiping Ao,
Qijun Yao,
Wen Zhang,
Wei Miao,
Zhenhui Lin,
Wenying Duan,
Dong Liu,
Kangmin Zhou,
Jie Liu,
Zheng Wang,
Junda Jin,
Kun Zhang,
Feng Wu,
Jinpeng Li,
Boliang Liu,
Xuan Zhang,
Zhengheng Luo,
Jiameng Wang,
Huiqian Hao,
Xingming Lu
, et al. (16 additional authors not shown)
Abstract:
The cycling of carbon between its ionized, atomic, and molecular phases shapes the chemical compositions and physical conditions of the interstellar medium (ISM). However, ground-based studies of the full carbon cycle have been limited by atmospheric absorption. Dome~A, the most promising site for submillimeter astronomy, has long resisted successful submillimeter astronomical observations. Using…
▽ More
The cycling of carbon between its ionized, atomic, and molecular phases shapes the chemical compositions and physical conditions of the interstellar medium (ISM). However, ground-based studies of the full carbon cycle have been limited by atmospheric absorption. Dome~A, the most promising site for submillimeter astronomy, has long resisted successful submillimeter astronomical observations. Using the 60~cm Antarctic Terahertz Explorer, we present the first successful CO ($4-3$) and [CI] ($^3P_1 - ^3P_0$) mapping observations of two archetypal triggered massive star-formation regions at Dome~A. These data, together with archival [CII], provide the first complete characterization of all three carbon phases in these environments. We find elevated C$^{0}$/CO abundance ratios in high-extinction regions, plausibly driven by deep penetration of intense radiation fields from massive stars into a clumpy ISM. These findings mark a major milestone for submillimeter astronomy at Dome~A and offer valuable insights into the impact of massive star feedback on the surrounding ISM.
△ Less
Submitted 11 January, 2026;
originally announced January 2026.
-
Relational Mediators: LLM Chatbots as Boundary Objects in Psychotherapy
Authors:
Jiatao Quan,
Ziyue Li,
Tian Qi Zhu,
Yuxuan Li,
Baoying Wang,
Wanda Pratt,
Nan Gao
Abstract:
As large language models (LLMs) are embedded into mental health technologies, they are often framed either as tools assisting therapists or autonomous therapeutic systems. Such perspectives overlook their potential to mediate relational complexities in therapy, particularly for systemically marginalized clients. Drawing on in-depth interviews with 12 therapists and 12 marginalized clients in China…
▽ More
As large language models (LLMs) are embedded into mental health technologies, they are often framed either as tools assisting therapists or autonomous therapeutic systems. Such perspectives overlook their potential to mediate relational complexities in therapy, particularly for systemically marginalized clients. Drawing on in-depth interviews with 12 therapists and 12 marginalized clients in China, including LGBTQ+ individuals or those from other marginalized backgrounds, we identify enduring relational challenges: difficulties building trust amid institutional barriers, the burden clients carry in educating therapists about marginalized identities, and challenges sustaining authentic self-disclosure across therapy and daily life. We argue that addressing these challenges requires AI systems capable of actively mediating underlying knowledge gaps, power asymmetries, and contextual disconnects. To this end, we propose the Dynamic Boundary Mediation Framework, which reconceptualizes LLM-enhanced systems as adaptive boundary objects that shift mediating roles across therapeutic stages. The framework delineates three forms of mediation: Epistemic (reducing knowledge asymmetries), Relational (rebalancing power dynamics), and Contextual (bridging therapy-life discontinuities). This framework offers a pathway toward designing relationally accountable AI systems that center the lived realities of marginalized users and more effectively support therapeutic relationships.
△ Less
Submitted 26 December, 2025;
originally announced December 2025.
-
Reflection-Driven Control for Trustworthy Code Agents
Authors:
Bin Wang,
Jiazheng Quan,
Xingrui Yu,
Hansen Hu,
Yuhao,
Ivor Tsang
Abstract:
Contemporary large language model (LLM) agents are remarkably capable, but they still lack reliable safety controls and can produce unconstrained, unpredictable, and even actively harmful outputs. To address this, we introduce Reflection-Driven Control, a standardized and pluggable control module that can be seamlessly integrated into general agent architectures. Reflection-Driven Control elevates…
▽ More
Contemporary large language model (LLM) agents are remarkably capable, but they still lack reliable safety controls and can produce unconstrained, unpredictable, and even actively harmful outputs. To address this, we introduce Reflection-Driven Control, a standardized and pluggable control module that can be seamlessly integrated into general agent architectures. Reflection-Driven Control elevates "self-reflection" from a post hoc patch into an explicit step in the agent's own reasoning process: during generation, the agent continuously runs an internal reflection loop that monitors and evaluates its own decision path. When potential risks are detected, the system retrieves relevant repair examples and secure coding guidelines from an evolving reflective memory, injecting these evidence-based constraints directly into subsequent reasoning steps. We instantiate Reflection-Driven Control in the setting of secure code generation and systematically evaluate it across eight classes of security-critical programming tasks. Empirical results show that Reflection-Driven Control substantially improves the security and policy compliance of generated code while largely preserving functional correctness, with minimal runtime and token overhead. Taken together, these findings indicate that Reflection-Driven Control is a practical path toward trustworthy AI coding agents: it enables designs that are simultaneously autonomous, safer by construction, and auditable.
△ Less
Submitted 21 December, 2025;
originally announced December 2025.
-
A Hierarchical, Model-Based System for High-Performance Humanoid Soccer
Authors:
Quanyou Wang,
Mingzhang Zhu,
Ruochen Hou,
Kay Gillespie,
Alvin Zhu,
Shiqi Wang,
Yicheng Wang,
Gaberiel I. Fernandez,
Yeting Liu,
Colin Togashi,
Hyunwoo Nam,
Aditya Navghare,
Alex Xu,
Taoyuanmin Zhu,
Min Sung Ahn,
Arturo Flores Alvarez,
Justin Quan,
Ethan Hong,
Dennis W. Hong
Abstract:
The development of athletic humanoid robots has gained significant attention as advances in actuation, sensing, and control enable increasingly dynamic, real-world capabilities. RoboCup, an international competition of fully autonomous humanoid robots, provides a uniquely challenging benchmark for such systems, culminating in the long-term goal of competing against human soccer players by 2050. Th…
▽ More
The development of athletic humanoid robots has gained significant attention as advances in actuation, sensing, and control enable increasingly dynamic, real-world capabilities. RoboCup, an international competition of fully autonomous humanoid robots, provides a uniquely challenging benchmark for such systems, culminating in the long-term goal of competing against human soccer players by 2050. This paper presents the hardware and software innovations underlying our team's victory in the RoboCup 2024 Adult-Sized Humanoid Soccer Competition. On the hardware side, we introduce an adult-sized humanoid platform built with lightweight structural components, high-torque quasi-direct-drive actuators, and a specialized foot design that enables powerful in-gait kicks while preserving locomotion robustness. On the software side, we develop an integrated perception and localization framework that combines stereo vision, object detection, and landmark-based fusion to provide reliable estimates of the ball, goals, teammates, and opponents. A mid-level navigation stack then generates collision-aware, dynamically feasible trajectories, while a centralized behavior manager coordinates high-level decision making, role selection, and kick execution based on the evolving game state. The seamless integration of these subsystems results in fast, precise, and tactically effective gameplay, enabling robust performance under the dynamic and adversarial conditions of real matches. This paper presents the design principles, system architecture, and experimental results that contributed to ARTEMIS's success as the 2024 Adult-Sized Humanoid Soccer champion.
△ Less
Submitted 10 December, 2025;
originally announced December 2025.
-
Nonlinearly preconditioned gradient flows
Authors:
Konstantinos Oikonomidis,
Alexander Bodard,
Jan Quan,
Panagiotis Patrinos
Abstract:
We study a continuous-time dynamical system which arises as the limit of a broad class of nonlinearly preconditioned gradient methods. Under mild assumptions, we establish existence of global solutions and derive Lyapunov-based convergence guarantees. For convex costs, we prove a sublinear decay in a geometry induced by some reference function, and under a generalized gradient-dominance condition…
▽ More
We study a continuous-time dynamical system which arises as the limit of a broad class of nonlinearly preconditioned gradient methods. Under mild assumptions, we establish existence of global solutions and derive Lyapunov-based convergence guarantees. For convex costs, we prove a sublinear decay in a geometry induced by some reference function, and under a generalized gradient-dominance condition we obtain exponential convergence. We further uncover a duality connection with mirror descent, and use it to establish that the flow of interest solves an infinite-horizon optimal-control problem of which the value function is the Bregman divergence generated by the cost. These results clarify the structure and optimization behavior of nonlinearly preconditioned gradient flows and connect them to known continuous-time models in non-Euclidean optimization.
△ Less
Submitted 17 April, 2026; v1 submitted 25 November, 2025;
originally announced November 2025.
-
Scaled relative graphs for pairs of operators beyond classical monotonicity
Authors:
Jan Quan,
Alexander Bodard,
Konstantinos Oikonomidis,
Panagiotis Patrinos
Abstract:
We introduce a generalization of the scaled relative graph (SRG) to pairs of operators, enabling the visualization of their relative incremental properties. This novel SRG framework provides the geometric counterpart for the study of nonlinear resolvents based on paired monotonicity conditions. We demonstrate that these conditions apply to linear operators composed with monotone mappings, a class…
▽ More
We introduce a generalization of the scaled relative graph (SRG) to pairs of operators, enabling the visualization of their relative incremental properties. This novel SRG framework provides the geometric counterpart for the study of nonlinear resolvents based on paired monotonicity conditions. We demonstrate that these conditions apply to linear operators composed with monotone mappings, a class that notably includes NPN transistors, allowing us to compute the response of multivalued, nonsmooth and highly nonmonotone electrical circuits.
△ Less
Submitted 25 November, 2025;
originally announced November 2025.
-
Rethinking PCA Through Duality
Authors:
Jan Quan,
Johan Suykens,
Panagiotis Patrinos
Abstract:
Motivated by the recently shown connection between self-attention and (kernel) principal component analysis (PCA), we revisit the fundamentals of PCA. Using the difference-of-convex (DC) framework, we present several novel formulations and provide new theoretical insights. In particular, we show the kernelizability and out-of-sample applicability for a PCA-like family of problems. Moreover, we unc…
▽ More
Motivated by the recently shown connection between self-attention and (kernel) principal component analysis (PCA), we revisit the fundamentals of PCA. Using the difference-of-convex (DC) framework, we present several novel formulations and provide new theoretical insights. In particular, we show the kernelizability and out-of-sample applicability for a PCA-like family of problems. Moreover, we uncover that simultaneous iteration, which is connected to the classical QR algorithm, is an instance of the difference-of-convex algorithm (DCA), offering an optimization perspective on this longstanding method. Further, we describe new algorithms for PCA and empirically compare them with state-of-the-art methods. Lastly, we introduce a kernelizable dual formulation for a robust variant of PCA that minimizes the $l_1$ deviation of the reconstruction errors.
△ Less
Submitted 20 October, 2025;
originally announced October 2025.
-
Nonlinearly Preconditioned Gradient Methods: Momentum and Stochastic Analysis
Authors:
Konstantinos Oikonomidis,
Jan Quan,
Panagiotis Patrinos
Abstract:
We study nonlinearly preconditioned gradient methods for smooth nonconvex optimization problems, focusing on sigmoid preconditioners that inherently perform a form of gradient clipping akin to the widely used gradient clipping technique. Building upon this idea, we introduce a novel heavy ball-type algorithm and provide convergence guarantees under a generalized smoothness condition that is less r…
▽ More
We study nonlinearly preconditioned gradient methods for smooth nonconvex optimization problems, focusing on sigmoid preconditioners that inherently perform a form of gradient clipping akin to the widely used gradient clipping technique. Building upon this idea, we introduce a novel heavy ball-type algorithm and provide convergence guarantees under a generalized smoothness condition that is less restrictive than traditional Lipschitz smoothness, thus covering a broader class of functions. Additionally, we develop a stochastic variant of the base method and study its convergence properties under different noise assumptions. We compare the proposed algorithms with baseline methods on diverse tasks from machine learning including neural network training.
△ Less
Submitted 13 October, 2025;
originally announced October 2025.
-
Test-Time Graph Search for Goal-Conditioned Reinforcement Learning
Authors:
Evgenii Opryshko,
Junwei Quan,
Claas Voelcker,
Yilun Du,
Igor Gilitschenski
Abstract:
Offline goal-conditioned reinforcement learning (GCRL) often struggles with long-horizon tasks, where errors in value estimation accumulate and produce unreliable policies. It is typically assumed that effective long-term planning is infeasible without specialized training. In contrast, our work demonstrates that existing GCRL policies can complete long-horizon tasks when combined with a lightweig…
▽ More
Offline goal-conditioned reinforcement learning (GCRL) often struggles with long-horizon tasks, where errors in value estimation accumulate and produce unreliable policies. It is typically assumed that effective long-term planning is infeasible without specialized training. In contrast, our work demonstrates that existing GCRL policies can complete long-horizon tasks when combined with a lightweight, training-free planning wrapper. We find that standard goal-conditioned value functions encode locally consistent geometric structure sufficient for planning. Our approach, Test-Time Graph Search (TTGS), constructs a graph over the offline dataset and employs an adaptive subgoal selection strategy. To address unreliable value estimates during shortest-path search, we propose a novel mechanism that softly penalizes long-distance transitions. Our method incurs negligible computational overhead and requires no additional supervision or parameter updates. On the OGBench benchmark, TTGS significantly boosts success rates across multiple base learners and tasks, with primary gains on challenging long-horizon locomotion tasks where some success rates are improved from near-zero to over 90\%, often matching or outperforming methods that require complex auxiliary training. Code and videos can be found at https://ktolnos.github.io/ttgs.
△ Less
Submitted 22 May, 2026; v1 submitted 8 October, 2025;
originally announced October 2025.
-
Rethinking Thinking Tokens: LLMs as Improvement Operators
Authors:
Lovish Madaan,
Aniket Didolkar,
Suchin Gururangan,
John Quan,
Ruan Silva,
Ruslan Salakhutdinov,
Manzil Zaheer,
Sanjeev Arora,
Anirudh Goyal
Abstract:
Reasoning training incentivizes LLMs to produce long chains of thought (long CoT), which among other things, allows them to explore solution strategies with self-checking. This results in higher accuracy, but inflates context length, token/compute cost, and answer latency. We ask: Can current models leverage their metacognition to provide other combinations on this Pareto frontier, e.g., better ac…
▽ More
Reasoning training incentivizes LLMs to produce long chains of thought (long CoT), which among other things, allows them to explore solution strategies with self-checking. This results in higher accuracy, but inflates context length, token/compute cost, and answer latency. We ask: Can current models leverage their metacognition to provide other combinations on this Pareto frontier, e.g., better accuracy with lower context length and/or latency? Abstractly, we view the model as an improvement operator on its own "thoughts" with a continuum of possible strategies. We identify an interesting inference family Parallel-Distill-Refine (PDR), which performs the following: (i) generate diverse drafts in parallel; (ii) distill them into a bounded, textual workspace; and (iii) refine conditioned on this workspace, producing an output that seeds the next round. Importantly, context length (hence compute cost) is controllable via degree of parallelism, and is no longer conflated with the total number of generated tokens. We report PDR instantiations of current models that give better accuracy than long CoT while incurring lower latency. Setting degree of parallelism to 1 yields an interesting subcase, Sequential Refinement (SR) (iteratively improve a single candidate answer) which provides performance superior to long CoT. Success of such model orchestrations raises the question whether further training could shift the Pareto frontier. To this end, we train an 8B thinking model with Reinforcement Learning (RL) to make it consistent with PDR as the inference method. On math tasks with verifiable answers, iterative pipelines surpass single-pass baselines at matched sequential budgets, with PDR delivering the largest gains (e.g., +11% on AIME 2024 and +9% on AIME 2025).
△ Less
Submitted 1 October, 2025;
originally announced October 2025.
-
Large Language Models as Virtual Survey Respondents: Evaluating Sociodemographic Response Generation
Authors:
Jianpeng Zhao,
Chenyu Yuan,
Weiming Luo,
Haoling Xie,
Guangwei Zhang,
Steven Jige Quan,
Zixuan Yuan,
Pengyang Wang,
Denghui Zhang
Abstract:
Questionnaire-based surveys are foundational to social science research and public policymaking, yet traditional survey methods remain costly, time-consuming, and often limited in scale. Although prior work has explored large language models (LLMs) as virtual survey respondents, existing studies often address narrow task settings, focus on single sociological domains, or lack a unified evaluation…
▽ More
Questionnaire-based surveys are foundational to social science research and public policymaking, yet traditional survey methods remain costly, time-consuming, and often limited in scale. Although prior work has explored large language models (LLMs) as virtual survey respondents, existing studies often address narrow task settings, focus on single sociological domains, or lack a unified evaluation framework that enables systematic comparison across diverse datasets and models. To address these gaps, we introduce two complementary task abstractions: Partial Attribute Simulation (PAS), where LLMs predict missing attributes from incomplete respondent profiles, and Full Attribute Simulation (FAS), where LLMs generate complete synthetic datasets under zero-context and context-enhanced conditions. Both are framed as diagnostic and exploratory tools rather than replacements for human data collection. We curate LLM-S^3 (Large Language Model-based Sociodemographic Survey Simulation), a benchmark spanning 11 real-world public datasets across four sociological domains, and evaluate GPT-3.5/4 Turbo and LLaMA 3.0/3.1-8B under zero-shot and few-shot settings. Our evaluation reveals consistent performance trends across model families, highlights failure modes in structured output generation, and demonstrates how context and prompt design affect simulation fidelity. Our code and dataset are available at: https://github.com/dart-lab-research/LLM-S-Cube-Benchmark
△ Less
Submitted 27 April, 2026; v1 submitted 8 September, 2025;
originally announced September 2025.
-
A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code
Authors:
Keke Lian,
Bin Wang,
Lei Zhang,
Libo Chen,
Junjie Wang,
Ziming Zhao,
Yujiu Yang,
Miaoqian Lin,
Haotong Duan,
Haoran Zhao,
Shuang Liao,
Mingda Guo,
Jiazheng Quan,
Yilu Zhong,
Chenhao He,
Zichuan Chen,
Jie Wu,
Haoling Li,
Zhaoxuan Li,
Jiongchi Yu,
Hui Li,
Dong Zhang
Abstract:
The increasing adoption of large language models (LLMs) in software engineering necessitates rigorous security evaluation of their generated code. However, existing benchmarks often lack relevance to real-world AI-assisted programming scenarios, making them inadequate for assessing the practical security risks associated with AI-generated code in production environments. To address this gap, we in…
▽ More
The increasing adoption of large language models (LLMs) in software engineering necessitates rigorous security evaluation of their generated code. However, existing benchmarks often lack relevance to real-world AI-assisted programming scenarios, making them inadequate for assessing the practical security risks associated with AI-generated code in production environments. To address this gap, we introduce A.S.E (AI Code Generation Security Evaluation), a repository-level evaluation benchmark designed to closely mirror real-world AI programming tasks, offering a comprehensive and reliable framework for assessing the security of AI-generated code. Our evaluation of leading LLMs on A.S.E reveals several key findings. In particular, current LLMs still struggle with secure coding. The complexity in repository-level scenarios presents challenges for LLMs that typically perform well on snippet-level tasks. Moreover, a larger reasoning budget does not necessarily lead to better code generation. These observations offer valuable insights into the current state of AI code generation and help developers identify the most suitable models for practical tasks. They also lay the groundwork for refining LLMs to generate secure and efficient code in real-world applications.
△ Less
Submitted 18 September, 2025; v1 submitted 25 August, 2025;
originally announced August 2025.
-
From Screen to Stage: Kid Cosmo, A Life-Like, Torque-Controlled Humanoid for Entertainment Robotics
Authors:
Havel Liu,
Mingzhang Zhu,
Arturo Moises Flores Alvarez,
Yuan Hung Lo,
Conrad Ku,
Federico Parres,
Justin Quan,
Colin Togashi,
Aditya Navghare,
Quanyou Wang,
Dennis W. Hong
Abstract:
Humanoid robots represent the cutting edge of robotics research, yet their potential in entertainment remains largely unexplored. Entertainment as a field prioritizes visuals and form, a principle that contrasts with the purely functional designs of most contemporary humanoid robots. Designing entertainment humanoid robots capable of fluid movement presents a number of unique challenges. In this p…
▽ More
Humanoid robots represent the cutting edge of robotics research, yet their potential in entertainment remains largely unexplored. Entertainment as a field prioritizes visuals and form, a principle that contrasts with the purely functional designs of most contemporary humanoid robots. Designing entertainment humanoid robots capable of fluid movement presents a number of unique challenges. In this paper, we present Kid Cosmo, a research platform designed for robust locomotion and life-like motion generation while imitating the look and mannerisms of its namesake character from Netflix's movie The Electric State. Kid Cosmo is a child-sized humanoid robot, standing 1.45 m tall and weighing 25 kg. It contains 28 degrees of freedom and primarily uses proprioceptive actuators, enabling torque-control walking and lifelike motion generation. Following worldwide showcases as part of the movie's press tour, we present the system architecture, challenges of a functional entertainment robot and unique solutions, and our initial findings on stability during simultaneous upper and lower body movement. We demonstrate the viability of performance-oriented humanoid robots that prioritize both character embodiment and technical functionality.
△ Less
Submitted 15 August, 2025;
originally announced August 2025.
-
Efficient Band Structure Unfolding with Atom-centered Orbitals: General Theory and Application
Authors:
Jingkai Quan,
Nikita Rybin,
Matthias Scheffler,
Christian Carbogno
Abstract:
Band structure unfolding is a key technique for analyzing and simplifying the electronic band structure of large, internally distorted supercells that break the primitive cell's translational symmetry. In this work, we present an efficient band unfolding method for atomic orbital (AO) basis sets that explicitly accounts for both the non-orthogonality of atomic orbitals and their atom-centered natu…
▽ More
Band structure unfolding is a key technique for analyzing and simplifying the electronic band structure of large, internally distorted supercells that break the primitive cell's translational symmetry. In this work, we present an efficient band unfolding method for atomic orbital (AO) basis sets that explicitly accounts for both the non-orthogonality of atomic orbitals and their atom-centered nature. Unlike existing approaches that typically rely on a plane-wave representation of the (semi-)valence states, we here derive analytical expressions that recasts the primitive cell translational operator and the associated Bloch-functions in the supercell AO basis. In turn, this enables the accurate and efficient unfolding of conduction, valence, and core states in all-electron codes, as demonstrated by our implementation in the all-electron ab initio simulation package FHI-aims, which employs numeric atom-centered orbitals. We explicitly demonstrate the capability of running large-scale unfolding calculations for systems with thousands of atoms and showcase the importance of this technique for computing temperature-dependent spectral functions in strongly anharmonic materials using CuI as example.
△ Less
Submitted 8 January, 2026; v1 submitted 26 June, 2025;
originally announced June 2025.
-
Unconventional Superconductivity in $\mathrm{La_{3}Ni_{2}O_{7}}$ from the Perspective of Symmetry
Authors:
Guan-Hao Feng,
Jun Quan,
Yusheng Hou
Abstract:
The recently discovered superconductor $\mathrm{La_{3}Ni_{2}O_{7}}$ has attracted significant attention due to its remarkably high transition temperature ($T_{c}$) under high pressure. Shortly after this discovery, thin-film $\mathrm{La_{3}Ni_{2}O_{7}}$ was demonstrated to exhibit ambient-pressure superconductivity; however, the corresponding $T_c$ is only about half that of the pressurized bulk m…
▽ More
The recently discovered superconductor $\mathrm{La_{3}Ni_{2}O_{7}}$ has attracted significant attention due to its remarkably high transition temperature ($T_{c}$) under high pressure. Shortly after this discovery, thin-film $\mathrm{La_{3}Ni_{2}O_{7}}$ was demonstrated to exhibit ambient-pressure superconductivity; however, the corresponding $T_c$ is only about half that of the pressurized bulk material. This striking difference raises questions about the underlying mechanisms governing superconductivity in these two structures. To address this issue, we develop a phenomenological symmetry-based method to investigate the superconducting gap structure in $\mathrm{La_{3}Ni_{2}O_{7}}$. Using density-functional theory methods (DFT+$U$), together with the experimentally determined $T_c$ and structural symmetry, we find that both pressurized bulk and thin-film $\mathrm{La_{3}Ni_{2}O_{7}}$ exhibit $s_{\pm}$-wave pairing symmetry and two-gap superconductivity, yet their dominant microscopic pairing configurations are distinct. In the pressurized bulk, superconductivity is dominated by the out-of-plane pairing of the Ni-$d_{z^2}$ orbitals, while in the thin film, the in-plane pairing of the Ni-$d_{x^2-y^2}$ orbitals prevails. Furthermore, the observed reduction in $T_c$ can be attributed to this transition of the dominant pairing type, driven by the decreased ratio of inter-layer to intra-layer hoppings in the thin film. Our result sheds lights on the microscopic pairing in $\mathrm{La_{3}Ni_{2}O_{7}}$ and reveals the significance of the symmetry. This method can potentially be generalized to a broader range of unconventional superconductors.
△ Less
Submitted 4 May, 2026; v1 submitted 2 June, 2025;
originally announced June 2025.
-
Ultrafast dynamics of local charge order in a THz-induced metastable quantum state
Authors:
Luis E. Parra López,
Jingkai Quan,
Alkisti Vaitsi,
Vivien Sleziona,
Fabian Schulz,
Angel Rubio,
Martin Wolf,
Melanie Müller
Abstract:
Controlling quantum materials with ultrafast light pulses enables access to transient and metastable states that are inaccessible under equilibrium conditions. Yet their local dynamics remain poorly understood due to the challenge of resolving ultrafast processes with angstrom-scale spatial resolution. Here, we use terahertz scanning tunnelling microscopy (THz-STM) to probe coherent collective dyn…
▽ More
Controlling quantum materials with ultrafast light pulses enables access to transient and metastable states that are inaccessible under equilibrium conditions. Yet their local dynamics remain poorly understood due to the challenge of resolving ultrafast processes with angstrom-scale spatial resolution. Here, we use terahertz scanning tunnelling microscopy (THz-STM) to probe coherent collective dynamics within a THz-induced metastable state in the layered charge density wave (CDW) material 1T-TaS2. Following ultrafast photoexcitation, we locally resolve coherent oscillations of the CDW amplitude mode at 2.5 THz together with two previously unreported modes at 1.3 THz and 0.7 THz. Comparison with phonon calculations identifies these as interlayer breathing and shear vibrations that are sensitive to the stacking configuration. These coherent dynamics are observed within a THz-induced metastable state that exhibits long-lived and spatially inhomogeneous modifications of the local density of states within the insulating gap, while higher THz fields drive a local redistribution and disordering of Star-of-David clusters near defects and domain boundaries. Our results suggest that the THz-induced metastable state involves a modification of the local interlayer stacking configuration, and demonstrate the role of interlayer degrees of freedom in the ultrafast dynamics of light-induced phases in layered quantum materials.
△ Less
Submitted 24 May, 2026; v1 submitted 26 May, 2025;
originally announced May 2025.
-
ParticleGS: Learning Neural Gaussian Particle Dynamics from Videos for Prior-free Physical Motion Extrapolation
Authors:
Jinsheng Quan,
Qiaowei Miao,
Yichao Xu,
Zizhuo Lin,
Ying Li,
Wei Yang,
Zhihui Li,
Yawei Luo
Abstract:
The ability to extrapolate dynamic 3D scenes beyond the observed timeframe is fundamental to advancing physical world understanding and predictive modeling. Existing dynamic 3D reconstruction methods have achieved high-fidelity rendering of temporal interpolation, but typically lack physical consistency in predicting the future. To overcome this issue, we propose ParticleGS, a physics-based framew…
▽ More
The ability to extrapolate dynamic 3D scenes beyond the observed timeframe is fundamental to advancing physical world understanding and predictive modeling. Existing dynamic 3D reconstruction methods have achieved high-fidelity rendering of temporal interpolation, but typically lack physical consistency in predicting the future. To overcome this issue, we propose ParticleGS, a physics-based framework that reformulates dynamic 3D scenes as physically grounded systems. ParticleGS comprises three key components: 1) an encoder that decomposes the scene into static properties and initial dynamic physical fields; 2) an evolver based on Neural Ordinary Differential Equations (Neural ODEs) that learns continuous-time dynamics for motion extrapolation; and 3) a decoder that reconstructs 3D Gaussians from evolved particle states for rendering. Through this design, ParticleGS integrates physical reasoning into dynamic 3D representations, enabling accurate and consistent prediction of the future. Experiments show that ParticleGS achieves state-of-the-art performance in extrapolation while maintaining rendering quality comparable to leading dynamic 3D reconstruction methods.
△ Less
Submitted 27 November, 2025; v1 submitted 26 May, 2025;
originally announced May 2025.
-
Roadmap on Advancements of the FHI-aims Software Package
Authors:
Joseph W. Abbott,
Carlos Mera Acosta,
Alaa Akkoush,
Alberto Ambrosetti,
Viktor Atalla,
Alexej Bagrets,
Jörg Behler,
Daniel Berger,
Hannah Bertschi,
Björn Bieniek,
Jonas Björk,
Volker Blum,
Saeed Bohloul,
Connor L. Box,
Nicholas Boyer,
Danilo Simoes Brambila,
Gabriel A. Bramley,
Kyle R. Bryenton,
María Camarasa-Gómez,
Christian Carbogno,
Fabio Caruso,
Sucismita Chutia,
Michele Ceriotti,
Gábor Csányi,
William Dawson
, et al. (181 additional authors not shown)
Abstract:
Electronic-structure theory is the foundation of the description of materials including multiscale modeling of their properties and functions. Obviously, without sufficient accuracy at the base, reliable predictions are unlikely at any level that follows. The software package FHI-aims has proven to be a game changer for accurate free-energy calculations because of its scalability, numerical precis…
▽ More
Electronic-structure theory is the foundation of the description of materials including multiscale modeling of their properties and functions. Obviously, without sufficient accuracy at the base, reliable predictions are unlikely at any level that follows. The software package FHI-aims has proven to be a game changer for accurate free-energy calculations because of its scalability, numerical precision, and its efficient handling of density functional theory (DFT) with hybrid functionals and van der Waals interactions. It treats molecules, clusters, and extended systems (solids and liquids) on an equal footing. Besides DFT, FHI-aims also includes quantum-chemistry methods, descriptions for excited states and vibrations, and calculations of various types of transport. Recent advancements address the integration of FHI-aims into an increasing number of workflows and various artificial intelligence (AI) methods. This Roadmap describes the state-of-the-art of FHI-aims and advancements that are currently ongoing or planned.
△ Less
Submitted 20 April, 2026; v1 submitted 30 April, 2025;
originally announced May 2025.
-
Omni-AD: Learning to Reconstruct Global and Local Features for Multi-class Anomaly Detection
Authors:
Jiajie Quan,
Ao Tong,
Yuxuan Cai,
Xinwei He,
Yulong Wang,
Yang Zhou
Abstract:
In multi-class unsupervised anomaly detection(MUAD), reconstruction-based methods learn to map input images to normal patterns to identify anomalous pixels. However, this strategy easily falls into the well-known "learning shortcut" issue when decoders fail to capture normal patterns and reconstruct both normal and abnormal samples naively. To address that, we propose to learn the input features i…
▽ More
In multi-class unsupervised anomaly detection(MUAD), reconstruction-based methods learn to map input images to normal patterns to identify anomalous pixels. However, this strategy easily falls into the well-known "learning shortcut" issue when decoders fail to capture normal patterns and reconstruct both normal and abnormal samples naively. To address that, we propose to learn the input features in global and local manners, forcing the network to memorize the normal patterns more comprehensively. Specifically, we design a two-branch decoder block, named Omni-block. One branch corresponds to global feature learning, where we serialize two self-attention blocks but replace the query and (key, value) with learnable tokens, respectively, thus capturing global features of normal patterns concisely and thoroughly. The local branch comprises depth-separable convolutions, whose locality enables effective and efficient learning of local features for normal patterns. By stacking Omni-blocks, we build a framework, Omni-AD, to learn normal patterns of different granularity and reconstruct them progressively. Comprehensive experiments on public anomaly detection benchmarks show that our method outperforms state-of-the-art approaches in MUAD. Code is available at https://github.com/easyoo/Omni-AD.git
△ Less
Submitted 28 March, 2025; v1 submitted 26 March, 2025;
originally announced March 2025.
-
Advances in 4D Generation: A Survey
Authors:
Qiaowei Miao,
Kehan Li,
Jinsheng Quan,
Zhiyuan Min,
Shaojie Ma,
Yichao Xu,
Yi Yang,
Ping Liu,
Yawei Luo
Abstract:
Generative artificial intelligence has recently progressed from static image and video synthesis to 3D content generation, culminating in the emergence of 4D generation-the task of synthesizing temporally coherent dynamic 3D assets guided by user input. As a burgeoning research frontier, 4D generation enables richer interactive and immersive experiences, with applications ranging from digital huma…
▽ More
Generative artificial intelligence has recently progressed from static image and video synthesis to 3D content generation, culminating in the emergence of 4D generation-the task of synthesizing temporally coherent dynamic 3D assets guided by user input. As a burgeoning research frontier, 4D generation enables richer interactive and immersive experiences, with applications ranging from digital humans to autonomous driving. Despite rapid progress, the field lacks a unified understanding of 4D representations, generative frameworks, basic paradigms, and the core technical challenges it faces. This survey provides a systematic and in-depth review of the 4D generation landscape. To comprehensively characterize 4D generation, we first categorize fundamental 4D representations and outline associated techniques for 4D generation. We then present an in-depth analysis of representative generative pipelines based on conditions and representation methods. Subsequently, we discuss how motion and geometry priors are integrated into 4D outputs to ensure spatio-temporal consistency under various control schemes. From an application perspective, this paper summarizes 4D generation tasks in areas such as dynamic object/scene generation, digital human synthesis, editable 4D content, and embodied AI. Furthermore, we summarize and multi-dimensionally compare four basic paradigms for 4D generation: End-to-End, Generated-Data-Based, Implicit-Distillation-Based, and Explicit-Supervision-Based. Concluding our analysis, we highlight five key challenges-consistency, controllability, diversity, efficiency, and fidelity-and contextualize these with current approaches.By distilling recent advances and outlining open problems, this work offers a comprehensive and forward-looking perspective to guide future research in 4D generation.
△ Less
Submitted 23 July, 2025; v1 submitted 18 March, 2025;
originally announced March 2025.
-
Nonlinearly Preconditioned Gradient Methods under Generalized Smoothness
Authors:
Konstantinos Oikonomidis,
Jan Quan,
Emanuel Laude,
Panagiotis Patrinos
Abstract:
We analyze nonlinearly preconditioned gradient methods for solving smooth minimization problems. We introduce a generalized smoothness property, based on the notion of abstract convexity, that is broader than Lipschitz smoothness and provide sufficient first- and second-order conditions. Notably, our framework encapsulates algorithms associated with the gradient clipping method and brings out nove…
▽ More
We analyze nonlinearly preconditioned gradient methods for solving smooth minimization problems. We introduce a generalized smoothness property, based on the notion of abstract convexity, that is broader than Lipschitz smoothness and provide sufficient first- and second-order conditions. Notably, our framework encapsulates algorithms associated with the gradient clipping method and brings out novel insights for the class of $(L_0,L_1)$-smooth functions that has received widespread interest recently, thus allowing us to extend beyond already established methods. We investigate the convergence of the proposed method in both the convex and nonconvex setting.
△ Less
Submitted 16 June, 2025; v1 submitted 12 February, 2025;
originally announced February 2025.
-
LEGOS-SLEEC: Tool for Formalizing and Analyzing Normative Requirements
Authors:
Kevin Kolyakov,
Lina Marsso,
Nick Feng,
Junwei Quan,
Marsha Chechik
Abstract:
Systems interacting with humans, such as assistive robots or chatbots, are increasingly integrated into our society. To prevent these systems from causing social, legal, ethical, empathetic, or cultural (SLEEC) harms, normative requirements specify the permissible range of their behaviors. These requirements encompass both functional and non-functional aspects and are defined with respect to time.…
▽ More
Systems interacting with humans, such as assistive robots or chatbots, are increasingly integrated into our society. To prevent these systems from causing social, legal, ethical, empathetic, or cultural (SLEEC) harms, normative requirements specify the permissible range of their behaviors. These requirements encompass both functional and non-functional aspects and are defined with respect to time. Typically, these requirements are specified by stakeholders from a broad range of fields, such as lawyers, ethicists, or philosophers, who may lack technical expertise. Because such stakeholders often have different goals, responsibilities, and objectives, ensuring that these requirements are well-formed is crucial. SLEEC DSL, a domain-specific language resembling natural language, has been developed to formalize these requirements as SLEEC rules. In this paper, we present LEGOS-SLEEC, a tool designed to support interdisciplinary stakeholders in specifying normative requirements as SLEEC rules, and in analyzing and debugging their well-formedness. LEGOS-SLEEC is built using four previously published components, which have been shown to be effective and usable across nine case studies. Reflecting on this experience, we have significantly improved the user interface of LEGOS-SLEEC and its diagnostic support, and demonstrate the effectiveness of these improvements using four interdisciplinary stakeholders. Showcase video URL is: https://youtu.be/LLaBLGxSi8A
△ Less
Submitted 21 January, 2025;
originally announced January 2025.
-
Valley-mediated singlet- and triplet-polaron interactions and quantum dynamics in a doped WSe$_2$ monolayer
Authors:
Yue Ni,
Di Huang,
Danfu Liang,
Albert Liu,
Xiaohui Liu,
Kevin Sampson,
Zhida Liu,
Jianmin Quan,
Kenji Watanabe,
Takashi Taniguchi,
Dmitry K. Efimkin,
Jesper Levinsen,
Meera M. Parish,
Xiaoqin Li
Abstract:
In doped transition metal dichalcogenides, optically created excitons (bound electron-hole pairs) can strongly interact with a Fermi sea of electrons to form Fermi polaron quasiparticles. When there are two distinct Fermi seas, as is the case in WSe$_2$, there are two flavors of lowest-energy (attractive) polarons -- singlet and triplet -- where the exciton is coupled to the Fermi sea in the same…
▽ More
In doped transition metal dichalcogenides, optically created excitons (bound electron-hole pairs) can strongly interact with a Fermi sea of electrons to form Fermi polaron quasiparticles. When there are two distinct Fermi seas, as is the case in WSe$_2$, there are two flavors of lowest-energy (attractive) polarons -- singlet and triplet -- where the exciton is coupled to the Fermi sea in the same or opposite valley, respectively. Using two-dimensional coherent electronic spectroscopy, we analyze how their quantum decoherence evolves with doping density and determine the condition under which stable Fermi polarons form. Because of the large oscillator strength associated with these resonances, intrinsic quantum dynamics of polarons as well as valley coherence between coupled singlet- and triplet polarons occur on sub-picosecond time scales. Surprisingly, we find that a dark-to-bright state conversion process leads to a particularly long-lived singlet polaron valley polarization, persisting up to 200-800 ps. Valley coherence between the singlet- and triplet polaron is correlated with their energy fluctuations. Our finding provides valuable guidance for the electrical and optical control of spin and valley indexes in atomically thin semiconductors.
△ Less
Submitted 4 January, 2025;
originally announced January 2025.
-
Light-matter interactions in layered materials and heterostructures: from moiré physics and magneto-optical effects to ultrafast dynamics and hybrid meta-photonics
Authors:
Luca Sortino,
Marcos H. D. Guimarães,
Alejandro Molina-Sánchez,
Jiamin Quan,
Denis Garoli,
Nicolò Maccaferri
Abstract:
Layered two-dimensional (2D) materials have revolutionized how we approach light-matter interactions, offering unprecedented optical and electronic properties with the potential for vertical heterostructures and manipulation of spin-valley degrees of freedom. The discovery of moiré physics in twisted heterostructures has further unlocked new possibilities for controlling the band structure of tail…
▽ More
Layered two-dimensional (2D) materials have revolutionized how we approach light-matter interactions, offering unprecedented optical and electronic properties with the potential for vertical heterostructures and manipulation of spin-valley degrees of freedom. The discovery of moiré physics in twisted heterostructures has further unlocked new possibilities for controlling the band structure of tailored semiconductor heterostructures. In parallel, the integration of 2D materials with hybrid photonic structures and ultrafast studies on their optical and spin-valley properties has revealed a wealth of novel physical phenomena. This perspective highlights the recent advances in our understanding of light-matter interactions in moiré and 2D systems, with a particular emphasis on ultrafast processes and the integration of these materials into photonic platforms. We explore the implications for optoelectronics and emerging photonic technologies, positioning 2D materials as a transformative tool for next-generation devices.
△ Less
Submitted 2 December, 2024;
originally announced December 2024.
-
Scaled Relative Graphs for Nonmonotone Operators with Applications in Circuit Theory
Authors:
Jan Quan,
Brecht Evens,
Rodolphe Sepulchre,
Panagiotis Patrinos
Abstract:
The scaled relative graph (SRG) is a powerful graphical tool for analyzing the properties of operators, by mapping their graph onto the complex plane. In this work, we study the SRG of two classes of nonmonotone operators, namely the general class of semimonotone operators and a class of angle-bounded operators. In particular, we provide an analytical description of the SRG of these classes and sh…
▽ More
The scaled relative graph (SRG) is a powerful graphical tool for analyzing the properties of operators, by mapping their graph onto the complex plane. In this work, we study the SRG of two classes of nonmonotone operators, namely the general class of semimonotone operators and a class of angle-bounded operators. In particular, we provide an analytical description of the SRG of these classes and show that membership of an operator to these classes can be verified through geometric containment of its SRG. To illustrate the importance of these results, we provide several examples in the context of electrical circuits. Most notably, we show that the Ebers-Moll transistor belongs to the class of angle-bounded operators and use this result to compute the response of a common-emitter amplifier using Chambolle-Pock, despite the underlying nonsmoothness and multi-valuedness, leveraging recent convergence results for this algorithm in the nonmonotone setting.
△ Less
Submitted 20 October, 2025; v1 submitted 26 November, 2024;
originally announced November 2024.
-
Temperature-dependent Electronic Spectral Functions from Band-Structure Unfolding
Authors:
Jingkai Quan,
Min-Ye Zhang,
Nikita Rybin,
Marios Zacharias,
Xinguo Ren,
Hong Jiang,
Matthias Scheffler,
Christian Carbogno
Abstract:
The electronic band structure, describing the periodic dependence of electronic quantum states on lattice momentum in reciprocal space, is a fundamental concept in solid-state physics. However, it's only well-defined for static nuclei. To account for thermodynamic effects, this concept must be generalized by introducing the temperature-dependent spectral function, which characterizes the finite-wi…
▽ More
The electronic band structure, describing the periodic dependence of electronic quantum states on lattice momentum in reciprocal space, is a fundamental concept in solid-state physics. However, it's only well-defined for static nuclei. To account for thermodynamic effects, this concept must be generalized by introducing the temperature-dependent spectral function, which characterizes the finite-width distributions of electronic quantum states at each reciprocal vector. Many-body perturbation theory can compute spectral functions and associated observables, but it approximates the dynamics of nuclei and its coupling to the electrons using the harmonic approximation and linear-order electron-phonon coupling elements, respectively. These approximations may fail at elevated temperatures or for mobile atoms. To avoid inaccuracies, the electronic spectral function can be obtained non-perturbatively, capturing higher-order couplings between electrons and vibrational degrees of freedom. This process involves recovering the representation of supercell bands in the first Brillouin zone of the primitive cell, a process known as unfolding. In this contribution, we describe the implementation of the band-structure unfolding technique in the electronic-structure theory package FHI-aims and the updates made since its original development.
△ Less
Submitted 7 November, 2024;
originally announced November 2024.
-
Magnon-mediated exciton-exciton interaction in a van der Waals antiferromagnet
Authors:
Biswajit Datta,
Pratap Chandra Adak,
Sichao Yu,
Agneya V. Dharmapalan,
Siedah J. Hall,
Anton Vakulenko,
Filipp Komissarenko,
Egor Kurganov,
Jiamin Quan,
Wei Wang,
Kseniia Mosina,
Zdeněk Sofer,
Dimitar Pashov,
Mark van Schilfgaarde,
Swagata Acharya,
Akashdeep Kamra,
Matthew Y. Sfeir,
Andrea Alù,
Alexander B. Khanikaev,
Vinod M. Menon
Abstract:
Excitons are fundamental excitations that govern the optical properties of semiconductors. Interacting excitons can lead to various emergent phases of matter and large nonlinear optical responses. In most semiconductors, excitons interact via exchange interaction or phase space filling. Correlated materials that host excitons coupled to other degrees of freedom offer hitherto unexplored pathways f…
▽ More
Excitons are fundamental excitations that govern the optical properties of semiconductors. Interacting excitons can lead to various emergent phases of matter and large nonlinear optical responses. In most semiconductors, excitons interact via exchange interaction or phase space filling. Correlated materials that host excitons coupled to other degrees of freedom offer hitherto unexplored pathways for controlling these interactions. Here, we demonstrate magnon-mediated excitonic interactions in CrSBr, an antiferromagnetic semiconductor. This interaction manifests as the dependence of exciton energy on exciton density via a magnonic adjustment of the spin canting angle. Our study demonstrates the emergence of quasiparticle-mediated interactions in correlated quantum materials, leading to large nonlinear optical responses and potential device concepts such as magnon-mediated quantum transducers.
△ Less
Submitted 27 September, 2024;
originally announced September 2024.
-
Verifiable cloud-based variational quantum algorithms
Authors:
Junhong Yang,
Banghai Wang,
Junyu Quan,
Qin Li
Abstract:
Variational quantum algorithms (VQAs) have shown potential for quantum advantage with noisy intermediate-scale quantum (NISQ) devices for quantum machine learning (QML). However, given the high cost and limited availability of quantum resources, delegating VQAs via cloud networks is a more practical solution for clients with limited quantum capabilities. Recently, Shingu et al.[Physical Review A,…
▽ More
Variational quantum algorithms (VQAs) have shown potential for quantum advantage with noisy intermediate-scale quantum (NISQ) devices for quantum machine learning (QML). However, given the high cost and limited availability of quantum resources, delegating VQAs via cloud networks is a more practical solution for clients with limited quantum capabilities. Recently, Shingu et al.[Physical Review A, 105, 022603 (2022)] proposed a variational secure cloud quantum computing protocol, utilizing ancilla-driven quantum computation (ADQC) for cloud-based VQAs with minimal quantum resource consumption. However, their protocol lacks verifiability, which exposes it to potential malicious behaviors by the server. Additionally, channel loss requires frequent re-delegation as the size of the delegated variational circuit grows, complicating verification due to increased circuit complexity. This paper introduces a new protocol to address these challenges and enhance both verifiability and tolerance to channel loss in cloud-based VQAs.
△ Less
Submitted 3 September, 2024; v1 submitted 24 August, 2024;
originally announced August 2024.
-
Carrier Mobility of Strongly Anharmonic Materials from First Principles
Authors:
Jingkai Quan,
Christian Carbogno,
Matthias Scheffler
Abstract:
First-principle approaches for phonon-limited electronic transport are typically based on many-body perturbation theory and transport equations. With that, they rely on the validity of the quasi-particle picture for electrons and phonons, which is known to fail in strongly anharmonic systems. In this work, we demonstrated the relevance of effects beyond the quasi-particle picture by combining ab i…
▽ More
First-principle approaches for phonon-limited electronic transport are typically based on many-body perturbation theory and transport equations. With that, they rely on the validity of the quasi-particle picture for electrons and phonons, which is known to fail in strongly anharmonic systems. In this work, we demonstrated the relevance of effects beyond the quasi-particle picture by combining ab initio molecular dynamics and the Kubo-Greenwood (KG) formalism to establish a non-perturbative, stochastic method to calculate carrier mobilities while accounting for all orders of anharmonic and electron-vibrational couplings. In particular, we propose and exploit several numerical strategies that overcome the notoriously slow convergence of the KG formalism for both electronic and nuclear degree of freedom in crystalline solids. The capability of this method is demonstrated by calculating the temperature-dependent electron mobility of the strongly anharmonic oxide perovskites SrTiO3 and BaTiO3 across a wide range of temperatures. We show that the temperature-dependence of the mobility is largely driven by anharmonic, higher-order coupling effects and rationalize these trends in terms of the non-perturbative electronic spectral functions.
△ Less
Submitted 23 August, 2024;
originally announced August 2024.
-
QIris: Quantum Implementation of Rainbow Table Attacks
Authors:
Lee Jun Quan,
Tan Jia Ye,
Goh Geok Ling,
Vivek Balachandran
Abstract:
This paper explores the use of Grover's Algorithm in the classical rainbow table, uncovering the potential of integrating quantum computing techniques with conventional cryptographic methods to develop a Quantum Rainbow Table Proof-of-Concept. This leverages on Quantum concepts and algorithms which includes the principle of qubit superposition, entanglement and teleportation, coupled with Grover's…
▽ More
This paper explores the use of Grover's Algorithm in the classical rainbow table, uncovering the potential of integrating quantum computing techniques with conventional cryptographic methods to develop a Quantum Rainbow Table Proof-of-Concept. This leverages on Quantum concepts and algorithms which includes the principle of qubit superposition, entanglement and teleportation, coupled with Grover's Algorithm to enable a more efficient search through the rainbow table. The paper also details on the hardware constraints and the work around to produce better results in the implementation stages. Through this work we develop a working prototype of quantum rainbow table and demonstrate how quantum computing could significantly improve the speed of cyber tools such as password crackers and thus impact the cyber security landscape.
△ Less
Submitted 31 March, 2025; v1 submitted 13 August, 2024;
originally announced August 2024.
-
Emergent Trion Resonance Driven by Lattice Reconstruction in a Moiré Superlattice
Authors:
Zhida Liu,
Haonan Wang,
Xiaohui Liu,
Yue Ni,
Hongtao Yan,
Frank Y. Gao,
Saba Arash,
Hyunsue Kim,
Dong Seob Kim,
Xiangcheng Liu,
Xiaoxiao Yu,
Yongxin Zeng,
Jiamin Quan,
Di Huang,
Kenji Watanabe,
Takashi Taniguchi,
Edoardo Baldini,
Keji Lai,
Allan H. MacDonald,
Chih-Kang Shih,
Jamie Warner,
Li Yang,
Xiaoqin Li
Abstract:
We investigate how many-electron excited states emerge in twisted MoSe2 homobilayers when the lattice reconstructions evolve. Notably, we identify a new trion resonance that arises in the transition regime of lattice reconstruction, where gradual changes in atomic alignment between the layers occur. Magnetic field-dependent measurements, supported by first-principles calculations, indicate that th…
▽ More
We investigate how many-electron excited states emerge in twisted MoSe2 homobilayers when the lattice reconstructions evolve. Notably, we identify a new trion resonance that arises in the transition regime of lattice reconstruction, where gradual changes in atomic alignment between the layers occur. Magnetic field-dependent measurements, supported by first-principles calculations, indicate that the exciton forms at the K valley while the doped hole resides in the Gamma valley. First-principles calculations further indicate that two nearly degenerate exciton resonances can arise, localized at different sites within the moiré supercell. We propose that the new trion resonance is a "charge-transfer" trion, in which the electron-hole pair is spatially separated from the doped hole. The emergence of these complex excited states stems from the distinct moiré potentials acting on holes and excitons, resulting in their different spatial distribution within the superlattice.
△ Less
Submitted 1 July, 2026; v1 submitted 24 July, 2024;
originally announced July 2024.
-
Diffuse X-ray Explorer: a high-resolution X-ray spectroscopic sky surveyor on the China Space Station
Authors:
Hai Jin,
Junjie Mao,
Liubiao Chen,
Naihui Chen,
Wei Cui,
Bo Gao,
Jinjin Li,
Xinfeng Li,
Jiejia Liu,
Jia Quan,
Chunyang Jiang,
Guole Wang,
Le Wang,
Qian Wang,
Sifan Wang,
Aimin Xiao,
Shuo Zhang
Abstract:
DIffuse X-ray Explorer (DIXE) is a proposed high-resolution X-ray spectroscopic sky surveyor on the China Space Station (CSS). DIXE will focus on studying hot baryons in the Milky Way. Galactic hot baryons like the X-ray emitting Milky Way halo and eROSITA bubbles are best observed in the sky survey mode with a large field of view. DIXE will take advantage of the orbital motion of the CSS to scan…
▽ More
DIffuse X-ray Explorer (DIXE) is a proposed high-resolution X-ray spectroscopic sky surveyor on the China Space Station (CSS). DIXE will focus on studying hot baryons in the Milky Way. Galactic hot baryons like the X-ray emitting Milky Way halo and eROSITA bubbles are best observed in the sky survey mode with a large field of view. DIXE will take advantage of the orbital motion of the CSS to scan a large fraction of the sky. High-resolution X-ray spectroscopy, enabled by superconducting microcalorimeters based on the transition-edge sensor (TES) technology, will probe the physical properties (e.g., temperature, density, elemental abundances, kinematics) of the Galactic hot baryons. This will complement the high-resolution imaging data obtained with the eROSITA mission. Here we present the preliminary design of DIXE. The payload consists mainly of a detector assembly and a cryogenic cooling system. The key components of the detector assembly are a microcalorimeter array and frequency-domain multiplexing readout electronics. To provide a working temperature for the detector assembly, the cooling system consists of an adiabatic demagnetization refrigerator and a mechanical cryocooler system.
△ Less
Submitted 14 June, 2024;
originally announced June 2024.
-
PLA4D: Pixel-Level Alignments for Text-to-4D Gaussian Splatting
Authors:
Qiaowei Miao,
JinSheng Quan,
Kehan Li,
Yawei Luo
Abstract:
Previous text-to-4D methods have leveraged multiple Score Distillation Sampling (SDS) techniques, combining motion priors from video-based diffusion models (DMs) with geometric priors from multiview DMs to implicitly guide 4D renderings. However, differences in these priors result in conflicting gradient directions during optimization, causing trade-offs between motion fidelity and geometry accura…
▽ More
Previous text-to-4D methods have leveraged multiple Score Distillation Sampling (SDS) techniques, combining motion priors from video-based diffusion models (DMs) with geometric priors from multiview DMs to implicitly guide 4D renderings. However, differences in these priors result in conflicting gradient directions during optimization, causing trade-offs between motion fidelity and geometry accuracy, and requiring substantial optimization time to reconcile the models. In this paper, we introduce \textbf{P}ixel-\textbf{L}evel \textbf{A}lignment for text-driven \textbf{4D} Gaussian splatting (PLA4D) to resolve this motion-geometry conflict. PLA4D provides an anchor reference, i.e., text-generated video, to align the rendering process conditioned by different DMs in pixel space. For static alignment, our approach introduces a focal alignment method and Gaussian-Mesh contrastive learning to iteratively adjust focal lengths and provide explicit geometric priors at each timestep. At the dynamic level, a motion alignment technique and T-MV refinement method are employed to enforce both pose alignment and motion continuity across unknown viewpoints, ensuring intrinsic geometric consistency across views. With such pixel-level multi-DM alignment, our PLA4D framework is able to generate 4D objects with superior geometric, motion, and semantic consistency. Fully implemented with open-source tools, PLA4D offers an efficient and accessible solution for high-quality 4D digital content creation with significantly reduced generation time.
△ Less
Submitted 18 November, 2024; v1 submitted 30 May, 2024;
originally announced May 2024.
-
Deep Learning Driven Buffer-Aided Cooperative Networks for B5G/6G: Challenges, Solutions, and Future Opportunities
Authors:
Peng Xu,
Gaojie Chen,
Jianping Quan,
Chong Huang,
Ioannis Krikidis,
Kai-Kit Wong,
Chan-Byoung Chae
Abstract:
Buffer-aided cooperative networks (BACNs) have garnered significant attention due to their potential applications in beyond fifth generation (B5G) or sixth generation (6G) critical scenarios. This article explores various typical application scenarios of buffer-aided relaying in B5G/6G networks to emphasize the importance of incorporating BACN. Additionally, we delve into the crucial technical cha…
▽ More
Buffer-aided cooperative networks (BACNs) have garnered significant attention due to their potential applications in beyond fifth generation (B5G) or sixth generation (6G) critical scenarios. This article explores various typical application scenarios of buffer-aided relaying in B5G/6G networks to emphasize the importance of incorporating BACN. Additionally, we delve into the crucial technical challenges in BACN, including stringent delay constraints, high reliability, imperfect channel state information (CSI), transmission security, and integrated network architecture. To address the challenges, we propose leveraging deep learning-based methods for the design and operation of B5G/6G networks with BACN, deviating from conventional buffer-aided relay selection approaches. In particular, we present two case studies to demonstrate the efficacy of centralized deep reinforcement learning (DRL) and decentralized DRL in buffer-aided non-terrestrial networks. Finally, we outline future research directions in B5G/6G that pertain to the utilization of BACN.
△ Less
Submitted 2 January, 2024;
originally announced January 2024.
-
Vision-Language Models as a Source of Rewards
Authors:
Kate Baumli,
Satinder Baveja,
Feryal Behbahani,
Harris Chan,
Gheorghe Comanici,
Sebastian Flennerhag,
Maxime Gazeau,
Kristian Holsheimer,
Dan Horgan,
Michael Laskin,
Clare Lyle,
Hussain Masoom,
Kay McKinney,
Volodymyr Mnih,
Alexander Neitz,
Dmitry Nikulin,
Fabio Pardo,
Jack Parker-Holder,
John Quan,
Tim Rocktäschel,
Himanshu Sahni,
Tom Schaul,
Yannick Schroecker,
Stephen Spencer,
Richie Steigerwald
, et al. (2 additional authors not shown)
Abstract:
Building generalist agents that can accomplish many goals in rich open-ended environments is one of the research frontiers for reinforcement learning. A key limiting factor for building generalist agents with RL has been the need for a large number of reward functions for achieving different goals. We investigate the feasibility of using off-the-shelf vision-language models, or VLMs, as sources of…
▽ More
Building generalist agents that can accomplish many goals in rich open-ended environments is one of the research frontiers for reinforcement learning. A key limiting factor for building generalist agents with RL has been the need for a large number of reward functions for achieving different goals. We investigate the feasibility of using off-the-shelf vision-language models, or VLMs, as sources of rewards for reinforcement learning agents. We show how rewards for visual achievement of a variety of language goals can be derived from the CLIP family of models, and used to train RL agents that can achieve a variety of language goals. We showcase this approach in two distinct visual domains and present a scaling trend showing how larger VLMs lead to more accurate rewards for visual goal achievement, which in turn produces more capable RL agents.
△ Less
Submitted 12 July, 2024; v1 submitted 14 December, 2023;
originally announced December 2023.
-
Delegated variational quantum algorithms based on quantum homomorphic encryption
Authors:
Qin Li,
Junyu Quan,
Jinjing Shi,
Shichao Zhang,
Xuelong Li
Abstract:
Variational quantum algorithms (VQAs) are considered as one of the most promising candidates for achieving quantum advantages on quantum devices in the noisy intermediate-scale quantum (NISQ) era. They have been developed for numerous applications such as image processing and solving linear systems of equations. The application of VQAs can be greatly enlarged if users with limited quantum capabili…
▽ More
Variational quantum algorithms (VQAs) are considered as one of the most promising candidates for achieving quantum advantages on quantum devices in the noisy intermediate-scale quantum (NISQ) era. They have been developed for numerous applications such as image processing and solving linear systems of equations. The application of VQAs can be greatly enlarged if users with limited quantum capabilities can run them on remote powerful quantum computers. But the private data of clients may be leaked to quantum servers in such a quantum cloud model. To solve the problem, a novel quantum homomorphic encryption (QHE) scheme which is client-friendly and suitable for VQAs is constructed for quantum servers to calculate encrypted data. Then delegated VQAs are proposed based on the given QHE scheme, where the server can train the ansatz circuit using the client's data even without knowing the real input and the output of the client. Furthermore, a delegated variational quantum classifier to identify handwritten digit images is given as a specific example of delegated VQAs and simulated on the cloud platform of Original Quantum to show its feasibility.
△ Less
Submitted 25 January, 2023;
originally announced January 2023.
-
Magneto-optics in a van der Waals magnet tuned by self-hybridized polaritons
Authors:
Florian Dirnberger,
Jiamin Quan,
Rezlind Bushati,
Geoffrey Diederich,
Matthias Florian,
Julian Klein,
Kseniia Mosina,
Zdenek Sofer,
Xiaodong Xu,
Akashdeep Kamra,
Francisco J. García-Vidal,
Andrea Alù,
Vinod M. Menon
Abstract:
Controlling quantum materials with light is of fundamental and technological importance. By utilizing the strong coupling of light and matter in optical cavities (1-3), recent studies were able to modify some of their most defining features (4-6). In this work, we study the magneto-optical properties of a van der Waals magnet that supports strong coupling of photons and excitons even in the absenc…
▽ More
Controlling quantum materials with light is of fundamental and technological importance. By utilizing the strong coupling of light and matter in optical cavities (1-3), recent studies were able to modify some of their most defining features (4-6). In this work, we study the magneto-optical properties of a van der Waals magnet that supports strong coupling of photons and excitons even in the absence of external cavity mirrors. In this material - the layered magnetic semiconductor CrSBr - emergent light-matter hybrids called polaritons are shown to significantly increase the spectral bandwidth of correlations between the magnetic, electronic, and optical properties, enabling largely tunable optical responses to applied magnetic fields and magnons. Our results highlight the importance of exciton-photon self-hybridization in van der Waals magnets and motivate novel directions for the manipulation of quantum material properties by strong light-matter coupling.
△ Less
Submitted 23 September, 2024; v1 submitted 18 January, 2023;
originally announced January 2023.
-
Interaction-driven transport of dark excitons in 2D semiconductors with phonon-mediated optical readout
Authors:
Saroj B. Chand,
John M. Woods,
Jiamin Quan,
Enrique Mejia,
Takashi Taniguchi,
Kenji Watanabe,
Andrea Alù,
Gabriele Grosso
Abstract:
The growing field of quantum information technology requires propagation of information over long distances with efficient readout mechanisms. Excitonic quantum fluids have emerged as a powerful platform for this task due to their straightforward electro-optical conversion. In two-dimensional transition metal dichalcogenides, the coupling between spin and valley provides exciting opportunities for…
▽ More
The growing field of quantum information technology requires propagation of information over long distances with efficient readout mechanisms. Excitonic quantum fluids have emerged as a powerful platform for this task due to their straightforward electro-optical conversion. In two-dimensional transition metal dichalcogenides, the coupling between spin and valley provides exciting opportunities for harnessing, manipulating and storing bits of information. However, the large inhomogeneity of single layers cannot be overcome by the properties of bright excitons, hindering spin-valley transport. Nonetheless, the rich band structure supports dark excitonic states with strong binding energy and longer lifetime, ideally suited for long-range transport. Here we show that dark excitons can diffuse over several micrometers and prove that this repulsion-driven propagation is robust across non-uniform samples. The long-range propagation of dark states with an optical readout mediated by chiral phonons provides a new concept of excitonic devices for applications in both classical and quantum information technology.
△ Less
Submitted 10 April, 2023; v1 submitted 1 December, 2022;
originally announced December 2022.