-
ProjFormer: Point Cloud Completion via Geometric-Projective Transformer and Cross-Modal Semantic Constraints
Authors:
Sheng Liu,
Meng Wang,
Ruihui Li,
Huilong Pi,
Zhuo Tang,
Kenli Li
Abstract:
Point cloud completion is inherently ill-posed due to severe sparsity and ambiguity in partial observations. Existing multi-view methods alleviate this by incorporating 2D semantics, but often rely on learned attention and fixed fusion, which lack geometric consistency and adaptability. We propose ProjFormer, a cross-modal framework that enforces geometry-consistent 2D-3D interaction through expli…
▽ More
Point cloud completion is inherently ill-posed due to severe sparsity and ambiguity in partial observations. Existing multi-view methods alleviate this by incorporating 2D semantics, but often rely on learned attention and fixed fusion, which lack geometric consistency and adaptability. We propose ProjFormer, a cross-modal framework that enforces geometry-consistent 2D-3D interaction through explicit projection and adaptive feature routing. A Projective Guided View Attention module aligns 3D points with multi-view features via deterministic projection, enabling efficient and geometrically consistent aggregation. Building on this, a geometry-aware routing network performs point-wise adaptive fusion of structural and observation-driven features for progressive refinement. Experiments show that, under a lightweight design, ProjFormer delivers competitive performance with improved structural completeness.
△ Less
Submitted 15 August, 2026;
originally announced August 2026.
-
RayLift: Lifting Complementary Ray-Wise Evidence with 3D Geometry Priors for Semantic Scene Completion
Authors:
Meng Wang,
Hongxia Yu,
Wenzhe He,
Xingdong Song,
Huilong Pi,
Jiapeng Zhang,
Ruihui Li
Abstract:
Camera-based 3D semantic scene completion (SSC) provides comprehensive scene understanding for autonomous driving and robotics. However, existing methods often treat stereo depth estimates as deterministic geometric constraints, causing depth uncertainty and local correspondence errors to propagate directly into voxel representations. To address this issue, we propose RayLift, a framework that use…
▽ More
Camera-based 3D semantic scene completion (SSC) provides comprehensive scene understanding for autonomous driving and robotics. However, existing methods often treat stereo depth estimates as deterministic geometric constraints, causing depth uncertainty and local correspondence errors to propagate directly into voxel representations. To address this issue, we propose RayLift, a framework that uses stereo geometry as a metric reference while incorporating complementary ray evidence to recover reliable 3D structures adaptively. RayLift first employs a Complementary Context Encoder that extracts geometry-aware priors from a frozen 3D vision foundation model, thereby enriching the scene context. It then introduces a Depth Ray Evidence Lifter module that jointly models geometric dissimilarity, depth confidence, and spatial uncertainty to adaptively sample and weight candidate surface locations along each camera ray. Finally, a Semantic-Aware Voxel Integrator injects the resulting ray evidence into voxel features by explicitly modeling their spatial support. Extensive experiments on SemanticKITTI and SSCBench-KITTI-360 demonstrate that RayLift achieves competitive performance and consistently outperforms existing methods.
△ Less
Submitted 9 August, 2026;
originally announced August 2026.
-
Organizing Principles for Moiré Quantum Matter
Authors:
Qiaoling Xu,
Yifan Gao,
Tao Zhang,
Ammon Fischer,
Yi Jiang,
Hanqi Pi,
Zike Fan,
Dongdong An,
Kun Zhou,
Yingjian Li,
Yongqing Li,
Yuhao Fu,
Lei Wang,
Lijun Zhang,
B. Andrei Bernevig,
Dante M. Kennes,
Enge Wang,
Angel Rubio,
Lede Xian
Abstract:
Moiré flat bands in van der Waals bilayers are usually discussed through a small set of mechanisms associated with the $Γ$ and $K$ valleys of hexagonal crystals, and more recently with $M$-valleys systems. Here we show that this view is incomplete. The momentum-space location and effective local orbital character of the monolayer's band edge, in conjunction with the moiré symmetry and the symmetry…
▽ More
Moiré flat bands in van der Waals bilayers are usually discussed through a small set of mechanisms associated with the $Γ$ and $K$ valleys of hexagonal crystals, and more recently with $M$-valleys systems. Here we show that this view is incomplete. The momentum-space location and effective local orbital character of the monolayer's band edge, in conjunction with the moiré symmetry and the symmetry representations of the resulting bands, provide a general set of organizing variables for the emergent low-energy moiré Hamiltonian. Applying fully relaxed first-principles calculations, band unfolding and symmetry-representation analysis to more than 600 commensurate twisted bilayers spanning all 2D lattice classes, we identify several routes to moiré quantum matter beyond the conventional single-orbital paradigm. The resulting flat bands realize trigonal, honeycomb, square, checkerboard and kagome-like Hubbard models with single-orbital, multi-orbital and multi-site Hilbert spaces; spin-orbit-coupled multi-orbital flat bands exhibit symmetry-indicated topology beyond the conventional $K$-valley setting; and nonsymmorphic moiré symmetries enforce semimetallic flat-band connectivity. Analogous quasi-one-dimensional flat-band structures are found in $M$-valley hexagonal systems and $X$-valley square or rectangular systems resulting from emergent momentum-space nonsymmorphic symmetries. Separately, coupled multi-valley manifolds with kagome-like connectivity are identified in several systems whose parent band edges lie at non-high-symmetry points. These results establish a valley-orbital-symmetry framework for connecting parent-material electronic structure to emergent moiré Hamiltonians relevant to correlated, topological and symmetry-enforced moiré phases.
△ Less
Submitted 27 July, 2026;
originally announced July 2026.
-
Quantum geometry and critical temperature enhancement in MgB$_2$ superconductivity
Authors:
Yi Jiang,
Haoyu Hu,
Dumitru Călugăru,
Kaja H. Hiorth,
Junze Deng,
Hanqi Pi,
Handong Chen,
Maia G. Vergniory,
Ion Errea,
Emilia Morosan,
Leslie M. Schoop,
Claudia Felser,
Miguel A. L. Marques,
Päivi Törmä,
Daniel Agterberg,
B. Andrei Bernevig
Abstract:
MgB$_2$, a phonon-mediated superconductor with record-high critical temperature $T_c\simeq 39$ K, is revisited to obtain a comprehensive theory of electrons, phonons, and their coupling with minimal ab initio input. We construct compact analytic models for the electronic structure, phonons, and electron-phonon coupling (EPC) of MgB$_2$. We show that strong in-plane B $sp^2$ bonding realizes an obs…
▽ More
MgB$_2$, a phonon-mediated superconductor with record-high critical temperature $T_c\simeq 39$ K, is revisited to obtain a comprehensive theory of electrons, phonons, and their coupling with minimal ab initio input. We construct compact analytic models for the electronic structure, phonons, and electron-phonon coupling (EPC) of MgB$_2$. We show that strong in-plane B $sp^2$ bonding realizes an obstructed band structure whose natural description is a bond-centered kagome lattice, yielding small quasi-2D $σ$-band Fermi-surface cylinders and pronounced quantum-geometric effects. The phonon spectrum is found to closely track that of a graphene-like boron layer, but the heavy intercalated Mg atoms dominate the three acoustic branches and rigidly lift the boron modes into the optical sector, while the in-plane B-B bond-stretching mode exhibits a pronounced softening along $Γ$-A. By symmetry, this $Γ$-point bond-stretching mode is the only $Γ$ phonon that can couple to the $σ$ Fermi surface, explaining its dominant contribution to the EPC. Upon electron doping toward the doubly degenerate band edge of the $σ$ sheets, we find that a reduced density of states competes with enhanced EPC matrix elements. At light electron doping, ab initio calculations show that the EPC enhancement dominates, leading to an increase in $T_c$ (within the clean doping limit without disorder effects). Using the Gaussian approximation for the EPC tensor, we further show that this enhancement is overwhelmingly quantum geometric in origin, arising from a geometric EPC contribution of the small $σ$ Fermi surface peaked at $Γ$. Overall, our results provide a transparent, symmetry-based account of superconductivity in MgB$_2$ and suggest that quantum-geometric effects can be essential for shaping doping trends in phonon-mediated superconductors.
△ Less
Submitted 21 July, 2026;
originally announced July 2026.
-
Autoregressive B-Rep Shape Generation with Parametric Surfaces
Authors:
Dafei Qin,
Rui Xu,
Zeyu Shen,
Kaichun Qiao,
Hongyang Lin,
Qixuan Zhang,
Huaijin Pi,
Lan Xu,
Jingyi Yu,
Wenping Wang,
Taku Komura
Abstract:
Generative CAD modeling has broad design and application potential. Despite significant advances in Boundary Representation (B-Rep) generation, the dominant representation in CAD, existing methods largely depend on uniformly sampled point- or grid-based geometry representations, sacrificing native surface types and parameters and thereby limiting geometric fidelity and downstream usability. We pre…
▽ More
Generative CAD modeling has broad design and application potential. Despite significant advances in Boundary Representation (B-Rep) generation, the dominant representation in CAD, existing methods largely depend on uniformly sampled point- or grid-based geometry representations, sacrificing native surface types and parameters and thereby limiting geometric fidelity and downstream usability. We present ParaCAD, an autoregressive framework for point-cloud-conditioned B-Rep generation that directly operates on native parametric surfaces. ParaCAD introduces a surface-centric tokenization that explicitly encodes each face by its exact surface type and continuous parameters, preserving the intrinsic semantics of CAD geometry. Our model first generates parametric surfaces with constrained UV domains, and then constructs a valid B-Rep by globally intersecting these surfaces to recover edges and vertices. ParaCAD places point-cloud-conditioned generation at the core of B-Rep synthesis, making it practical for user-guided reconstruction and seamless integration into existing 3D generation pipelines. Extensive experiments demonstrate that ParaCAD produces accurate B-Reps with faithful point-cloud alignment, outperforming point-based baselines in geometric precision, robustness, watertightness and downstream usability.
△ Less
Submitted 19 July, 2026;
originally announced July 2026.
-
Engineering topological flat bands in $Γ$-valley moiré systems with Ising-type SOC: twisted 1T-ZrS$_2$ and 1T-SnSe$_2$
Authors:
Hanqi Pi,
Yves H. Kwan,
Haoyu Hu,
Yi Jiang,
Dumitru Călugăru,
Jie Shan,
Kin Fai Mak,
Miguel M. Ugeda,
Dmitri K. Efetov,
Maia G. Vergniory,
B. Andrei Bernevig
Abstract:
Twisted moiré superlattices hosting topological flat bands provide a platform to explore the interplay between topology and correlations. Here we investigate topological band structures in $Γ$-valley moiré systems based on 1T-ZrS$_2$ and 1T-SnSe$_2$. Using large-scale ab initio calculations and continuum modelling, we demonstrate that both materials exhibit an approximate spin-$U(1)$ symmetry and…
▽ More
Twisted moiré superlattices hosting topological flat bands provide a platform to explore the interplay between topology and correlations. Here we investigate topological band structures in $Γ$-valley moiré systems based on 1T-ZrS$_2$ and 1T-SnSe$_2$. Using large-scale ab initio calculations and continuum modelling, we demonstrate that both materials exhibit an approximate spin-$U(1)$ symmetry and host isolated topological moiré valence bands, including quantum spin Hall and high spin Chern states. By constructing a hierarchy of $Γ$-valley moiré continuum models, we show that isolated moiré bands carry a trivial $C_3$ symmetry indicator when the low-energy physics is described by a single effective orbital and a single layer-hybridized branch, either bonding or antibonding. Topological bands therefore arise from inter-branch and/or inter-orbital coupling. Moreover, we determine interaction-driven phase diagrams using Hartree--Fock and exact diagonalization, finding various phases tunable by twist angle, interaction strength, and displacement field. We identify specific conditions under which fractional Chern insulators are favored. Together with previous work showing that the moiré conduction bands of 1T-ZrS$_2$ and 1T-SnSe$_2$ realize $M$-valley twisting and host quasi-one-dimensional physics, our results establish these systems as ideal platforms for strongly correlated moiré physics and provide a systematic framework for understanding topological band structures in $Γ$-valley moiré materials.
△ Less
Submitted 13 May, 2026;
originally announced May 2026.
-
Contact Matrix: Enhancing Dance Motion Synthesis with Precise Interaction Modeling
Authors:
Xuhai Chen,
Zhi Cen,
Huaijin Pi,
Sida Peng,
Xiaowei Zhou,
Yong Liu
Abstract:
Generating realistic reactive motions, in which one person reacts to the fixed motions of others, is challenging due to strict interaction constraints and a limited feasible solution space. This paper focuses on a typical scenario: duet dance, where high-quality data is scarce, motion patterns are complex, and the details of human interactions are both intricate and abundant. To tackle these chall…
▽ More
Generating realistic reactive motions, in which one person reacts to the fixed motions of others, is challenging due to strict interaction constraints and a limited feasible solution space. This paper focuses on a typical scenario: duet dance, where high-quality data is scarce, motion patterns are complex, and the details of human interactions are both intricate and abundant. To tackle these challenges, we propose a novel two-stage framework. In the first stage, we introduce a motion VQ-VAE with separate body-part encoders and a joint decoder, enabling specialized codebooks to enhance representation capacity while dynamically modeling dependencies across body parts during decoding, thereby preventing inconsistencies in the generated motions. In the second stage, we propose a contact-aware diffusion model for reactive motion generation that jointly generates motion and a contact matrix between individuals, enabling explicit interaction modeling and providing guidance toward more precise and constrained interaction dynamics during sampling. Experiments show that our method outperforms Duolando with lower $\text{FID}_k$ (8.89 vs. 25.30) and $\text{FID}_{cd}$ (8.01 vs. 9.97), as well as a higher BED (0.4606 vs. 0.2858), indicating improved interaction fidelity and rhythmic synchronization.
△ Less
Submitted 6 May, 2026;
originally announced May 2026.
-
Strips as Tokens: Artist Mesh Generation with Native UV Segmentation
Authors:
Rui Xu,
Dafei Qin,
Kaichun Qiao,
Qiujie Dong,
Huaijin Pi,
Qixuan Zhang,
Longwen Zhang,
Lan Xu,
Jingyi Yu,
Wenping Wang,
Taku Komura
Abstract:
Recent advancements in autoregressive transformers have demonstrated remarkable potential for generating artist-quality meshes. However, the token ordering strategies employed by existing methods typically fail to meet professional artist standards, where coordinate-based sorting yields inefficiently long sequences, and patch-based heuristics disrupt the continuous edge flow and structural regular…
▽ More
Recent advancements in autoregressive transformers have demonstrated remarkable potential for generating artist-quality meshes. However, the token ordering strategies employed by existing methods typically fail to meet professional artist standards, where coordinate-based sorting yields inefficiently long sequences, and patch-based heuristics disrupt the continuous edge flow and structural regularity essential for high-quality modeling. To address these limitations, we propose Strips as Tokens (SATO), a novel framework with a token ordering strategy inspired by triangle strips. By constructing the sequence as a connected chain of faces that explicitly encodes UV boundaries, our method naturally preserves the organized edge flow and semantic layout characteristic of artist-created meshes. A key advantage of this formulation is its unified representation, enabling the same token sequence to be decoded into either a triangle or quadrilateral mesh. This flexibility facilitates joint training on both data types: large-scale triangle data provides fundamental structural priors, while high-quality quad data enhances the geometric regularity of the outputs. Extensive experiments demonstrate that SATO consistently outperforms prior methods in terms of geometric quality, structural coherence, and UV segmentation. Project page: https://ruixu.me/html/SATO/index.html
△ Less
Submitted 1 May, 2026; v1 submitted 10 April, 2026;
originally announced April 2026.
-
Gloria: Consistent Character Video Generation via Content Anchors
Authors:
Yuhang Yang,
Fan Zhang,
Huaijin Pi,
Shuai Guo,
Guowei Xu,
Wei Zhai,
Yang Cao,
Zheng-Jun Zha
Abstract:
Digital characters are central to modern media, yet generating character videos with long-duration, consistent multi-view appearance and expressive identity remains challenging. Existing approaches either provide insufficient context to preserve identity or leverage non-character-centric information as the memory, leading to suboptimal consistency. Recognizing that character video generation inher…
▽ More
Digital characters are central to modern media, yet generating character videos with long-duration, consistent multi-view appearance and expressive identity remains challenging. Existing approaches either provide insufficient context to preserve identity or leverage non-character-centric information as the memory, leading to suboptimal consistency. Recognizing that character video generation inherently resembles an outside-looking-in scenario. In this work, we propose representing the character visual attributes through a compact set of anchor frames. This design provides stable references for consistency, while reference-based video generation inherently faces challenges of copy-pasting and multi-reference conflicts. To address these, we introduce two mechanisms: Superset Content Anchoring, providing intra- and extra-training clip cues to prevent duplication, and RoPE as Weak Condition, encoding positional offsets to distinguish multiple anchors. Furthermore, we construct a scalable pipeline to extract these anchors from massive videos. Experiments show our method generates high-quality character videos exceeding 10 minutes, and achieves expressive identity and appearance consistency across views, surpassing existing methods.
△ Less
Submitted 31 March, 2026;
originally announced March 2026.
-
New Crystal Structures Hide in Plain Sight: A Stress Test for AI-Guided Materials Discovery
Authors:
Xin Zhang,
Scott B. Lee,
Sudipta Chatterjee,
Hanqi Pi,
Yi Jiang,
Fatmagül Katmer,
Emily G. Ward,
Daniel E. Widdowson,
Charles C. Tam,
Sarah Schwarz,
Connor J. Pollak,
Jaime M. Moya,
Grigorii Skorupskii,
Vitaliy A. Kurlin,
Stephen D. Wilson,
B. Andrei Bernevig,
Leslie M. Schoop
Abstract:
New types of crystal structures are discovered only rarely, and the artificial intelligence (AI) models now reshaping materials discovery have so far produced new chemical compositions within known structural families rather than genuinely new structures. We report GdNiSn4 and LuNiSn4, intermetallics that adopt a previously unreported structure type, found not by computation but by exploratory syn…
▽ More
New types of crystal structures are discovered only rarely, and the artificial intelligence (AI) models now reshaping materials discovery have so far produced new chemical compositions within known structural families rather than genuinely new structures. We report GdNiSn4 and LuNiSn4, intermetallics that adopt a previously unreported structure type, found not by computation but by exploratory synthesis. Single-crystal diffraction shows that the structure is an intergrowth of two known structural units. We then use this system as a benchmark for two leading generative models, MatterGen and DiffCSP++. For DiffCSP++, the benchmark is performed in its crystallographically constrained setting, using the required space-group and Wyckoff-position inputs. Under our sampling budget, neither model recovers the experimentally reported monoclinic structure within the structural-matching tolerance. The generated structures are evaluated without further structural relaxation using the nonmagnetic analog LuNiSn4, where we rule out 4f magnetism as the cause. Because the new structure is built from familiar building blocks, it should be derivable. We argue that encoding chemical reasoning, such as the stacking of known motifs, is a concrete path toward AI that can discover structurally novel materials.
△ Less
Submitted 18 June, 2026; v1 submitted 5 March, 2026;
originally announced March 2026.
-
EmbodMocap: In-the-Wild 4D Human-Scene Reconstruction for Embodied Agents
Authors:
Wenjia Wang,
Liang Pan,
Huaijin Pi,
Yuke Lou,
Xuqian Ren,
Yifan Wu,
Zhouyingcheng Liao,
Lei Yang,
Rishabh Dabral,
Christian Theobalt,
Taku Komura
Abstract:
Human behaviors in the real world naturally encode rich, long-term contextual information that can be leveraged to train embodied agents for perception, understanding, and acting. However, existing capture systems typically rely on costly studio setups and wearable devices, limiting the large-scale collection of scene-conditioned human motion data in the wild. To address this, we propose EmbodMoca…
▽ More
Human behaviors in the real world naturally encode rich, long-term contextual information that can be leveraged to train embodied agents for perception, understanding, and acting. However, existing capture systems typically rely on costly studio setups and wearable devices, limiting the large-scale collection of scene-conditioned human motion data in the wild. To address this, we propose EmbodMocap, a portable and affordable data collection pipeline using two moving iPhones. Our key idea is to jointly calibrate dual RGB-D sequences to reconstruct both humans and scenes within a unified metric world coordinate frame. The proposed method allows metric-scale and scene-consistent capture in everyday environments without static cameras or markers, bridging human motion and scene geometry seamlessly. Compared with optical capture ground truth, we demonstrate that the dual-view setting exhibits a remarkable ability to mitigate depth ambiguity, achieving superior alignment and reconstruction performance over single iphone or monocular models. Based on the collected data, we empower three embodied AI tasks: monocular human-scene-reconstruction, where we fine-tune on feedforward models that output metric-scale, world-space aligned humans and scenes; physics-based character animation, where we prove our data could be used to scale human-object interaction skills and scene-aware motion tracking; and robot motion control, where we train a humanoid robot via sim-to-real RL to replicate human motions depicted in videos. Experimental results validate the effectiveness of our pipeline and its contributions towards advancing embodied AI research.
△ Less
Submitted 1 April, 2026; v1 submitted 26 February, 2026;
originally announced February 2026.
-
LiNeXt: Revisiting LiDAR Completion with Efficient Non-Diffusion Architectures
Authors:
Wenzhe He,
Xiaojun Chen,
Ruiqi Wang,
Ruihui Li,
Huilong Pi,
Jiapeng Zhang,
Zhuo Tang,
Kenli Li
Abstract:
3D LiDAR scene completion from point clouds is a fundamental component of perception systems in autonomous vehicles. Previous methods have predominantly employed diffusion models for high-fidelity reconstruction. However, their multi-step iterative sampling incurs significant computational overhead, limiting its real-time applicability. To address this, we propose LiNeXt-a lightweight, non-diffusi…
▽ More
3D LiDAR scene completion from point clouds is a fundamental component of perception systems in autonomous vehicles. Previous methods have predominantly employed diffusion models for high-fidelity reconstruction. However, their multi-step iterative sampling incurs significant computational overhead, limiting its real-time applicability. To address this, we propose LiNeXt-a lightweight, non-diffusion network optimized for rapid and accurate point cloud completion. Specifically, LiNeXt first applies the Noise-to-Coarse (N2C) Module to denoise the input noisy point cloud in a single pass, thereby obviating the multi-step iterative sampling of diffusion-based methods. The Refine Module then takes the coarse point cloud and its intermediate features from the N2C Module to perform more precise refinement, further enhancing structural completeness. Furthermore, we observe that LiDAR point clouds exhibit a distance-dependent spatial distribution, being densely sampled at proximal ranges and sparsely sampled at distal ranges. Accordingly, we propose the Distance-aware Selected Repeat strategy to generate a more uniformly distributed noisy point cloud. On the SemanticKITTI dataset, LiNeXt achieves a 199.8x speedup in inference, reduces Chamfer Distance by 50.7%, and uses only 6.1% of the parameters compared with LiDiff. These results demonstrate the superior efficiency and effectiveness of LiNeXt for real-time scene completion.
△ Less
Submitted 29 November, 2025; v1 submitted 13 November, 2025;
originally announced November 2025.
-
Chern-Selective multi-valley Flat Bands in Twisted Mono-Bilayer and Mono-Trilayer MoTe$_2$
Authors:
Ziyue Qi,
Hanqi Pi,
Yan Zhang,
Jiaxuan Liu,
Nicolas Regnault,
Hongming Weng,
B. Andrei Bernevig,
Jiabin Yu,
Quansheng Wu
Abstract:
The interplay between moiré flat bands originating from different valleys can give rise to a variety of exotic quantum phases. In this work, we investigate the electronic properties of twisted mono-bilayer (A-AB) and mono-trilayer (A-ABA) MoTe$_2$ using first-principles calculations and continuum models. Unlike previous studies on twisted bilayer systems, in which low-energy flat bands originate s…
▽ More
The interplay between moiré flat bands originating from different valleys can give rise to a variety of exotic quantum phases. In this work, we investigate the electronic properties of twisted mono-bilayer (A-AB) and mono-trilayer (A-ABA) MoTe$_2$ using first-principles calculations and continuum models. Unlike previous studies on twisted bilayer systems, in which low-energy flat bands originate solely from the $K/K'$ valleys, in A-AB and A-ABA twisted MoTe$_2$ (\tmt) the moiré bands at low energies arise from both the $Γ$ and $K/K'$ valleys, with spin Chern numbers $C_s=0$ (for $Γ$) and $C_{\uparrow/\downarrow}=\pm1$ (for $K/K'$), respectively. We show that the multi-valley moiré flat bands are governed by interlayer-hybridization effects, and that different stacking configurations and thicknesses tune the relative energy alignment between the $Γ$ and $K$ valley moiré flat bands. By constructing valley-resolved continuum models and performing Wannierization for the low-energy moiré bands, we further uncover that the Berry curvature and quantum metric distributions can be effectively tuned by the layer number and stacking configuration. Unlike other moiré systems, where only one kind of valley influenced the low energy physics, the simultaneous appearance of two distinct types of valleys, with different symmetries, establish A-AB and A-ABA \tmt\ as ideal platforms for studying layer-controlled multi-valley physics.
△ Less
Submitted 14 October, 2025;
originally announced October 2025.
-
Self-Supervised Point Cloud Completion based on Multi-View Augmentations of Single Partial Point Cloud
Authors:
Jingjing Lu,
Huilong Pi,
Yunchuan Qin,
Zhuo Tang,
Ruihui Li
Abstract:
Point cloud completion aims to reconstruct complete shapes from partial observations. Although current methods have achieved remarkable performance, they still have some limitations: Supervised methods heavily rely on ground truth, which limits their generalization to real-world datasets due to the synthetic-to-real domain gap. Unsupervised methods require complete point clouds to compose unpaired…
▽ More
Point cloud completion aims to reconstruct complete shapes from partial observations. Although current methods have achieved remarkable performance, they still have some limitations: Supervised methods heavily rely on ground truth, which limits their generalization to real-world datasets due to the synthetic-to-real domain gap. Unsupervised methods require complete point clouds to compose unpaired training data, and weakly-supervised methods need multi-view observations of the object. Existing self-supervised methods frequently produce unsatisfactory predictions due to the limited capabilities of their self-supervised signals. To overcome these challenges, we propose a novel self-supervised point cloud completion method. We design a set of novel self-supervised signals based on multi-view augmentations of the single partial point cloud. Additionally, to enhance the model's learning ability, we first incorporate Mamba into self-supervised point cloud completion task, encouraging the model to generate point clouds with better quality. Experiments on synthetic and real-world datasets demonstrate that our method achieves state-of-the-art results.
△ Less
Submitted 26 September, 2025;
originally announced September 2025.
-
Enhancing Gradient Variance and Differential Privacy in Quantum Federated Learning
Authors:
Duc-Thien Phan,
Minh-Duong Nguyen,
Quoc-Viet Pham,
Huilong Pi
Abstract:
Upon integrating Quantum Neural Network (QNN) as the local model, Quantum Federated Learning (QFL) has recently confronted notable challenges. Firstly, exploration is hindered over sharp minima, decreasing learning performance. Secondly, the steady gradient descent results in more stable and predictable model transmissions over wireless channels, making the model more susceptible to attacks from a…
▽ More
Upon integrating Quantum Neural Network (QNN) as the local model, Quantum Federated Learning (QFL) has recently confronted notable challenges. Firstly, exploration is hindered over sharp minima, decreasing learning performance. Secondly, the steady gradient descent results in more stable and predictable model transmissions over wireless channels, making the model more susceptible to attacks from adversarial entities. Additionally, the local QFL model is vulnerable to noise produced by the quantum device's intermediate noise states, since it requires the use of quantum gates and circuits for training. This local noise becomes intertwined with learning parameters during training, impairing model precision and convergence rate. To address these issues, we propose a new QFL technique that incorporates differential privacy and introduces a dedicated noise estimation strategy to quantify and mitigate the impact of intermediate quantum noise. Furthermore, we design an adaptive noise generation scheme to alleviate privacy threats associated with the vanishing gradient variance phenomenon of QNN and enhance robustness against device noise. Experimental results demonstrate that our algorithm effectively balances convergence, reduces communication costs, and mitigates the adverse effects of intermediate quantum noise while maintaining strong privacy protection. Using real-world datasets, we achieved test accuracy of up to 98.47\% for the MNIST dataset and 83.85\% for the CIFAR-10 dataset while maintaining fast execution times.
△ Less
Submitted 4 September, 2025;
originally announced September 2025.
-
Emergent Interacting Phases in the Strong Coupling Limit of Twisted M-Valley Moiré Systems: Application to SnSe${}_2$
Authors:
Ming-Rui Li,
Dumitru Calugaru,
Yi Jiang,
Hanqi Pi,
Ammon Fischer,
Henning Schlömer,
Lennart Klebl,
Maia G. Vergniory,
Dante M. Kennes,
Siddharth A. Parameswaran,
Hong Yao,
B. Andrei Bernevig,
Haoyu Hu
Abstract:
We construct an interacting Wannier model for both AA-stacked and AB-stacked twisted SnSe2, revealing a rich landscape of correlated quantum phases. For the AA-stacked case, the system is effectively described by a three-orbital triangular lattice model, where each orbital corresponds to a valley and exhibits an approximate one-dimensional hopping structure due to a new momentum-space non-symmorph…
▽ More
We construct an interacting Wannier model for both AA-stacked and AB-stacked twisted SnSe2, revealing a rich landscape of correlated quantum phases. For the AA-stacked case, the system is effectively described by a three-orbital triangular lattice model, where each orbital corresponds to a valley and exhibits an approximate one-dimensional hopping structure due to a new momentum-space non-symmorphic symmetry. By exploring the interacting phase diagram using a combination of theoretical methods, including Hartree-Fock mean-field theory and exact solutions of the spin model in certain limits, we identify several exotic quantum phases. These include a dimerized phase with finite residual entropy, valence bond solids, and quantum paramagnetism. In the AB-stacked case, the system realizes an interacting kagome lattice model, where the Wannier orbitals associated with the three valleys form three sublattices. In the strong coupling regime, we use cluster mean-field methods to demonstrate the emergence of a classical spin liquid phase due to the frustrated lattice structure. The high tunability of the moiré system, which allows control over both the filling and interaction strength (via twist angle), renders twisted SnSe2 a versatile platform for realizing a wide range of exotic correlated quantum phases.
△ Less
Submitted 13 August, 2025;
originally announced August 2025.
-
CoDA: Coordinated Diffusion Noise Optimization for Whole-Body Manipulation of Articulated Objects
Authors:
Huaijin Pi,
Zhi Cen,
Zhiyang Dou,
Taku Komura
Abstract:
Synthesizing whole-body manipulation of articulated objects, including body motion, hand motion, and object motion, is a critical yet challenging task with broad applications in virtual humans and robotics. The core challenges are twofold. First, achieving realistic whole-body motion requires tight coordination between the hands and the rest of the body, as their movements are interdependent durin…
▽ More
Synthesizing whole-body manipulation of articulated objects, including body motion, hand motion, and object motion, is a critical yet challenging task with broad applications in virtual humans and robotics. The core challenges are twofold. First, achieving realistic whole-body motion requires tight coordination between the hands and the rest of the body, as their movements are interdependent during manipulation. Second, articulated object manipulation typically involves high degrees of freedom and demands higher precision, often requiring the fingers to be placed at specific regions to actuate movable parts. To address these challenges, we propose a novel coordinated diffusion noise optimization framework. Specifically, we perform noise-space optimization over three specialized diffusion models for the body, left hand, and right hand, each trained on its own motion dataset to improve generalization. Coordination naturally emerges through gradient flow along the human kinematic chain, allowing the global body posture to adapt in response to hand motion objectives with high fidelity. To further enhance precision in hand-object interaction, we adopt a unified representation based on basis point sets (BPS), where end-effector positions are encoded as distances to the same BPS used for object geometry. This unified representation captures fine-grained spatial relationships between the hand and articulated object parts, and the resulting trajectories serve as targets to guide the optimization of diffusion noise, producing highly accurate interaction motion. We conduct extensive experiments demonstrating that our method outperforms existing approaches in motion quality and physical plausibility, and enables various capabilities such as object pose control, simultaneous walking and manipulation, and whole-body generation from hand-only data.
△ Less
Submitted 27 May, 2025;
originally announced May 2025.
-
Leveraging Pre-trained Large Language Models with Refined Prompting for Online Task and Motion Planning
Authors:
Huihui Guo,
Huilong Pi,
Yunchuan Qin,
Zhuo Tang,
Kenli Li
Abstract:
With the rapid advancement of artificial intelligence, there is an increasing demand for intelligent robots capable of assisting humans in daily tasks and performing complex operations. Such robots not only require task planning capabilities but must also execute tasks with stability and robustness. In this paper, we present a closed-loop task planning and acting system, LLM-PAS, which is assisted…
▽ More
With the rapid advancement of artificial intelligence, there is an increasing demand for intelligent robots capable of assisting humans in daily tasks and performing complex operations. Such robots not only require task planning capabilities but must also execute tasks with stability and robustness. In this paper, we present a closed-loop task planning and acting system, LLM-PAS, which is assisted by a pre-trained Large Language Model (LLM). While LLM-PAS plans long-horizon tasks in a manner similar to traditional task and motion planners, it also emphasizes the execution phase of the task. By transferring part of the constraint-checking process from the planning phase to the execution phase, LLM-PAS enables exploration of the constraint space and delivers more accurate feedback on environmental anomalies during execution. The reasoning capabilities of the LLM allow it to handle anomalies that cannot be addressed by the robust executor. To further enhance the system's ability to assist the planner during replanning, we propose the First Look Prompting (FLP) method, which induces LLM to generate effective PDDL goals. Through comparative prompting experiments and systematic experiments, we demonstrate the effectiveness and robustness of LLM-PAS in handling anomalous conditions during task execution.
△ Less
Submitted 30 April, 2025;
originally announced April 2025.
-
Theory of Superconductivity in LaRu$_3$Si$_2$ and Predictions of New Kagome Flat Band Superconductors
Authors:
Junze Deng,
Yi Jiang,
Tiago F. T. Cerqueira,
Haoyu Hu,
Eeli O. Lamponen,
Dumitru Călugăru,
Hanqi Pi,
Zhijun Wang,
Maia G. Vergniory,
Emilia Morosan,
Titus Neupert,
S. Blanco-Canosa,
Claudia Felser,
Kristjan Haule,
Miguel A. L. Marques,
Päivi Törmä,
B. Andrei Bernevig
Abstract:
We present a comprehensive investigation of the flat-band kagome superconductor LaRu$_3$Si$_2$, which has recently been reported to host charge density wave (CDW) order above room temperature ($T_{CDW} \simeq 400$ K). The stable crystal structure above the CDW transition is identified via soft phonon condensation and confirmed to be harmonically stable through ab initio calculations, consistent wi…
▽ More
We present a comprehensive investigation of the flat-band kagome superconductor LaRu$_3$Si$_2$, which has recently been reported to host charge density wave (CDW) order above room temperature ($T_{CDW} \simeq 400$ K). The stable crystal structure above the CDW transition is identified via soft phonon condensation and confirmed to be harmonically stable through ab initio calculations, consistent with recent X-ray diffraction refinements. The electron-phonon coupling (EPC) in LaRu$_3$Si$_2$ is found to be mode-selective, primarily driven by strong interactions between Ru-$B_{3u}$ phonons (local $x$-direction, pointing toward the hexagon center) and Ru-$A_g$ electrons (local $d_{x^2-y^2}$ orbital) within the kagome lattice. Using a spring-ball model, we identify this mode-selective EPC as a universal feature of kagome materials. Employing the newly developed Gaussian approximation of the hopping parameters, we derive an analytical expression for the EPC and demonstrate that superconductivity in LaRu$_3$Si$_2$ is mostly driven by the coupling between the kagome $B_{3u}$ phonons and the $A_g$ electrons. The impact of doping is also investigated, revealing that light hole doping (approximately one hole per unit cell) significantly enhances the superconducting critical temperature $T_c$ by 50%, whereas heavy doping induces structural instability and ferromagnetism. Furthermore, high-throughput screening identifies 3063 stable 1:3:2 kagome materials, of which 428 are predicted to exhibit superconductivity with $T_c > 1$ K, and the highest $T_c$ reaching 15 K. These findings establish LaRu$_3$Si$_2$ and related materials as promising platforms for exploring the interplay among kagome flat bands, EPC, and superconductivity. Additionally, they may offer valuable insights into potential limitations on the $T_c$ of flat-band superconductivity in real materials.
△ Less
Submitted 26 March, 2025;
originally announced March 2025.
-
MotionStreamer: Streaming Motion Generation via Diffusion-based Autoregressive Model in Causal Latent Space
Authors:
Lixing Xiao,
Shunlin Lu,
Huaijin Pi,
Ke Fan,
Liang Pan,
Yueer Zhou,
Ziyong Feng,
Xiaowei Zhou,
Sida Peng,
Jingbo Wang
Abstract:
This paper addresses the challenge of text-conditioned streaming motion generation, which requires us to predict the next-step human pose based on variable-length historical motions and incoming texts. Existing methods struggle to achieve streaming motion generation, e.g., diffusion models are constrained by pre-defined motion lengths, while GPT-based methods suffer from delayed response and error…
▽ More
This paper addresses the challenge of text-conditioned streaming motion generation, which requires us to predict the next-step human pose based on variable-length historical motions and incoming texts. Existing methods struggle to achieve streaming motion generation, e.g., diffusion models are constrained by pre-defined motion lengths, while GPT-based methods suffer from delayed response and error accumulation problem due to discretized non-causal tokenization. To solve these problems, we propose MotionStreamer, a novel framework that incorporates a continuous causal latent space into a probabilistic autoregressive model. The continuous latents mitigate information loss caused by discretization and effectively reduce error accumulation during long-term autoregressive generation. In addition, by establishing temporal causal dependencies between current and historical motion latents, our model fully utilizes the available information to achieve accurate online motion decoding. Experiments show that our method outperforms existing approaches while offering more applications, including multi-round generation, long-term generation, and dynamic motion composition. Project Page: https://zju3dv.github.io/MotionStreamer/
△ Less
Submitted 7 August, 2025; v1 submitted 19 March, 2025;
originally announced March 2025.
-
VLScene: Vision-Language Guidance Distillation for Camera-Based 3D Semantic Scene Completion
Authors:
Meng Wang,
Huilong Pi,
Ruihui Li,
Yunchuan Qin,
Zhuo Tang,
Kenli Li
Abstract:
Camera-based 3D semantic scene completion (SSC) provides dense geometric and semantic perception for autonomous driving. However, images provide limited information making the model susceptible to geometric ambiguity caused by occlusion and perspective distortion. Existing methods often lack explicit semantic modeling between objects, limiting their perception of 3D semantic context. To address th…
▽ More
Camera-based 3D semantic scene completion (SSC) provides dense geometric and semantic perception for autonomous driving. However, images provide limited information making the model susceptible to geometric ambiguity caused by occlusion and perspective distortion. Existing methods often lack explicit semantic modeling between objects, limiting their perception of 3D semantic context. To address these challenges, we propose a novel method VLScene: Vision-Language Guidance Distillation for Camera-based 3D Semantic Scene Completion. The key insight is to use the vision-language model to introduce high-level semantic priors to provide the object spatial context required for 3D scene understanding. Specifically, we design a vision-language guidance distillation process to enhance image features, which can effectively capture semantic knowledge from the surrounding environment and improve spatial context reasoning. In addition, we introduce a geometric-semantic sparse awareness mechanism to propagate geometric structures in the neighborhood and enhance semantic information through contextual sparse interactions. Experimental results demonstrate that VLScene achieves rank-1st performance on challenging benchmarks--SemanticKITTI and SSCBench-KITTI-360, yielding remarkably mIoU scores of 17.52 and 19.10, respectively.
△ Less
Submitted 8 March, 2025;
originally announced March 2025.
-
Mocap-2-to-3: Multi-view Lifting for Monocular Motion Recovery with 2D Pretraining
Authors:
Zhumei Wang,
Zechen Hu,
Ruoxi Guo,
Huaijin Pi,
Ziyong Feng,
Liang Zhang,
Mingtao Pei,
Siyuan Huang
Abstract:
Human motion recovery for real-world interaction demands both precise action details and metric-scale trajectories. Recovering absolute human pose from monocular input presents a viable solution, but faces two main challenges: (1) models' reliance on 3D training data from constrained environments limits their out-of-distribution generalization; and (2) the inherent difficulty of estimating metric-…
▽ More
Human motion recovery for real-world interaction demands both precise action details and metric-scale trajectories. Recovering absolute human pose from monocular input presents a viable solution, but faces two main challenges: (1) models' reliance on 3D training data from constrained environments limits their out-of-distribution generalization; and (2) the inherent difficulty of estimating metric-scale poses from monocular observations. This paper introduces Mocap-2-to-3, a novel framework that differs from prior HMR methods by recovering absolute poses from monocular input and leveraging abundant 2D data to enhance 3D motion recovery. To effectively utilize the action priors and diversity in large-scale 2D datasets, we reformulate 3D motion as a multi-view synthesis process and divide the training into two stages: a single-view diffusion model is first pre-trained on extensive 2D data, followed by multi-view fine-tuning on 3D data, thus achieving a combination of strong priors and geometric constraints. Furthermore, to recover absolute poses, we introduce a novel human motion representation that decouples the learning of local pose and global movements, while encoding ground geometric priors to accelerate convergence, thereby yielding more precise positioning in the physical world. Experiments on in-the-wild benchmarks show that our method outperforms state-of-the-art approaches in both camera-space motion realism and world-grounded human positioning, while exhibiting strong generalization capability.
△ Less
Submitted 13 March, 2026; v1 submitted 5 March, 2025;
originally announced March 2025.
-
CSubBT: A Self-Adjusting Execution Framework for Mobile Manipulation System
Authors:
Huihui Guo,
Huizhang Luo,
Huilong Pi,
Mingxing Duan,
Kenli Li,
Chubo Liu
Abstract:
With the advancements in modern intelligent technologies, mobile robots equipped with manipulators are increasingly operating in unstructured environments. These robots can plan sequences of actions for long-horizon tasks based on perceived information. However, in practice, the planned actions often fail due to discrepancies between the perceptual information used for planning and the actual cond…
▽ More
With the advancements in modern intelligent technologies, mobile robots equipped with manipulators are increasingly operating in unstructured environments. These robots can plan sequences of actions for long-horizon tasks based on perceived information. However, in practice, the planned actions often fail due to discrepancies between the perceptual information used for planning and the actual conditions. In this paper, we introduce the {\itshape Conditional Subtree} (CSubBT), a general self-adjusting execution framework for mobile manipulation tasks based on Behavior Trees (BTs). CSubBT decomposes symbolic action into sub-actions and uses BTs to control their execution, addressing any potential anomalies during the process. CSubBT treats common anomalies as constraint non-satisfaction problems and continuously guides the robot in performing tasks by sampling new action parameters in the constraint space when anomalies are detected. We demonstrate the robustness of our framework through extensive manipulation experiments on different platforms, both in simulation and real-world settings.
△ Less
Submitted 28 February, 2025;
originally announced February 2025.
-
Ready-to-React: Online Reaction Policy for Two-Character Interaction Generation
Authors:
Zhi Cen,
Huaijin Pi,
Sida Peng,
Qing Shuai,
Yujun Shen,
Hujun Bao,
Xiaowei Zhou,
Ruizhen Hu
Abstract:
This paper addresses the task of generating two-character online interactions. Previously, two main settings existed for two-character interaction generation: (1) generating one's motions based on the counterpart's complete motion sequence, and (2) jointly generating two-character motions based on specific conditions. We argue that these settings fail to model the process of real-life two-characte…
▽ More
This paper addresses the task of generating two-character online interactions. Previously, two main settings existed for two-character interaction generation: (1) generating one's motions based on the counterpart's complete motion sequence, and (2) jointly generating two-character motions based on specific conditions. We argue that these settings fail to model the process of real-life two-character interactions, where humans will react to their counterparts in real time and act as independent individuals. In contrast, we propose an online reaction policy, called Ready-to-React, to generate the next character pose based on past observed motions. Each character has its own reaction policy as its "brain", enabling them to interact like real humans in a streaming manner. Our policy is implemented by incorporating a diffusion head into an auto-regressive model, which can dynamically respond to the counterpart's motions while effectively mitigating the error accumulation throughout the generation process. We conduct comprehensive experiments using the challenging boxing task. Experimental results demonstrate that our method outperforms existing baselines and can generate extended motion sequences. Additionally, we show that our approach can be controlled by sparse signals, making it well-suited for VR and other online interactive environments.
△ Less
Submitted 27 February, 2025;
originally announced February 2025.
-
Tertiary EOR-like microfluidic experiments: influence of viscosity ratio on oil clusters mobilization
Authors:
Haohong Pi,
Abdelaziz Omari,
Giuseppe Sciumè
Abstract:
Understanding the pore-scale dynamics of immiscible two-phase flow in porous media is crucial 9 for optimizing EOR strategies. In this work, we investigate the mobilization dynamics of oil clusters by 10 means of microfluidic devices that allow pore scale direct characterization of flow in water-wet chips. We varied both flow rates during waterflooding and the viscosity ratio by injecting Glycerol…
▽ More
Understanding the pore-scale dynamics of immiscible two-phase flow in porous media is crucial 9 for optimizing EOR strategies. In this work, we investigate the mobilization dynamics of oil clusters by 10 means of microfluidic devices that allow pore scale direct characterization of flow in water-wet chips. We varied both flow rates during waterflooding and the viscosity ratio by injecting Glycerol/water mixtures of various compositions right after the waterflooding period. During waterflooding, the flow rate has only a limited impact on residual oil. With a subsequent injection of a Glycerol/water mixture, the oil recovery is significantly enhanced. To better understand the recovery mechanisms, oil clusters were categorized into droplets, blobs and ganglia. Increasing the viscosity of the injected mixture resulted in only a slight reduction in the number of ganglia but significantly decreased their total volume, thus reducing overall oil saturation. This is due to ganglia breakup into smaller ganglia, blobs and droplets that are subsequently mobilized and transported away, while remaining parts of original ganglia still remain trapped. As long as droplets and blobs are considered, their number is seen to only weakly change by the increase of mixture viscosity and even their number may temporarily increase as they result from ganglia rupture. So, the process can be separated in two main steps: ganglia breakage that feed the medium in blobs and droplets and a second step where such moving oil entities are transported. The characteristic time for oil transport is believed to be longer than that required for ganglia breakage.
△ Less
Submitted 4 January, 2025;
originally announced January 2025.
-
Feedback Regulated Opto-Mechanical Soft Robotic Actuators
Authors:
Jianfeng Yang,
Haotian Pi,
Zixuan Deng,
Hongshuang Guo,
Wan Shou,
Hang Zhang,
Hao Zeng
Abstract:
Natural organisms can convert environmental stimuli into sensory feedback to regulate their body and realize active adaptivity. However, realizing such a feedback-regulation mechanism in synthetic material systems remains a grand challenge. It is believed that achieving complex feedback mechanisms in responsive materials will pave the way toward autonomous, intelligent structure and actuation with…
▽ More
Natural organisms can convert environmental stimuli into sensory feedback to regulate their body and realize active adaptivity. However, realizing such a feedback-regulation mechanism in synthetic material systems remains a grand challenge. It is believed that achieving complex feedback mechanisms in responsive materials will pave the way toward autonomous, intelligent structure and actuation without complex electronics. Inspired by living systems, we report a general principle to design and construct such feedback loops in light-responsive materials. Specifically, we design a baffle-actuator mechanism to incorporate programmed feedback into the opto-mechanical responsiveness. By simply addressing the baffle position with respect to the incident light beam, positive and negative feedback are programmed. We demonstrate the transformation of a light-bending strip into a switcher, where the intensity of light determines the energy barrier under positive feedback, realizing multi-stable shape-morphing. By leveraging the negative feedback and associated homeostasis, we demonstrate two soft robots, i.e., a locomotor and a swimmer. Furthermore, we unveil the ubiquity of feedback in light-responsive materials, which provides new insight into self-regulated robotic matters.
△ Less
Submitted 20 December, 2024;
originally announced December 2024.
-
Motion-2-To-3: Leveraging 2D Motion Data for 3D Motion Generations
Authors:
Ruoxi Guo,
Huaijin Pi,
Zehong Shen,
Qing Shuai,
Zechen Hu,
Zhumei Wang,
Yajiao Dong,
Ruizhen Hu,
Taku Komura,
Sida Peng,
Xiaowei Zhou
Abstract:
Text-driven human motion synthesis has showcased its potential for revolutionizing motion design in the movie and game industry. Existing methods often rely on 3D motion capture data, which requires special setups, resulting in high costs for data acquisition, ultimately limiting the diversity and scope of human motion. In contrast, 2D human videos offer a vast and accessible source of motion data…
▽ More
Text-driven human motion synthesis has showcased its potential for revolutionizing motion design in the movie and game industry. Existing methods often rely on 3D motion capture data, which requires special setups, resulting in high costs for data acquisition, ultimately limiting the diversity and scope of human motion. In contrast, 2D human videos offer a vast and accessible source of motion data, covering a wider range of styles and activities. In this paper, we explore the use of 2D human motion extracted from videos as an alternative data source to improve text-driven 3D motion generation. Our approach introduces a novel framework that disentangles local joint motion from global movements, enabling efficient learning of local motion priors from 2D data. We first train a single-view 2D local motion generator on a large dataset of text-2D motion pairs. Then we fine-tune the generator with 3D data, transforming it into a multi-view generator that predicts view-consistent local joint motion and root dynamics. Evaluations on the well-acknowledged dataset and novel text prompts demonstrate that our method can efficiently utilize 2D data, supporting a wider range of realistic 3D human motion generation. Our code is publicly available at https://zju3dv.github.io/Motion-2-to-3/.
△ Less
Submitted 19 May, 2026; v1 submitted 17 December, 2024;
originally announced December 2024.
-
A New Moiré Platform Based on M-Point Twisting
Authors:
Dumitru Călugăru,
Yi Jiang,
Haoyu Hu,
Hanqi Pi,
Jiabin Yu,
Maia G. Vergniory,
Jie Shan,
Claudia Felser,
Leslie M. Schoop,
Dmitri K. Efetov,
Kin Fai Mak,
B. Andrei Bernevig
Abstract:
We introduce a new class of moiré systems and materials based on monolayers with triangular lattices and low-energy states at the M points of the Brillouin zone. These M-point moiré materials are fundamentally distinct from those derived from $Γ$- or K-point monolayers, featuring three time-reversal-preserving valleys related by three-fold rotational symmetry. We propose twisted bilayers of experi…
▽ More
We introduce a new class of moiré systems and materials based on monolayers with triangular lattices and low-energy states at the M points of the Brillouin zone. These M-point moiré materials are fundamentally distinct from those derived from $Γ$- or K-point monolayers, featuring three time-reversal-preserving valleys related by three-fold rotational symmetry. We propose twisted bilayers of experimentally exfoliable 1T-SnSe$_2$ and 1T-ZrS$_2$ as realizations of this new class. Using extensive ab initio simulations, we develop quantitative continuum models and analytically show that the corresponding M-point moiré Hamiltonians exhibit emergent momentum-space non-symmorphic symmetries and a kagome plane-wave lattice in momentum space. This represents the first experimentally viable realization of a projective representation of crystalline space groups in a non-magnetic system. With interactions, these materials represent six-flavor Hubbard simulators with Mott physics, as can be seen by their flat Wilson loops. Furthermore, the presence of a non-symmorphic momentum-space in-plane mirror symmetry makes some of the M-point moiré Hamiltonians quasi-one-dimensional in each valley, suggesting the possibility of realizing Luttinger liquid physics. We predict the twist angles at which a series of (conduction) flat bands appear, provide a faithful continuum Hamiltonian, analyze its topology and charge density and briefly discuss several aspects of the physics of this new platform.
△ Less
Submitted 27 November, 2024;
originally announced November 2024.
-
2D Theoretically Twistable Material Database
Authors:
Yi Jiang,
Urko Petralanda,
Grigorii Skorupskii,
Qiaoling Xu,
Hanqi Pi,
Dumitru Călugăru,
Haoyu Hu,
Jiaze Xie,
Rose Albu Mustaf,
Peter Höhn,
Vicky Haase,
Maia G. Vergniory,
Martin Claassen,
Luis Elcoro,
Nicolas Regnault,
Jie Shan,
Kin Fai Mak,
Dmitri K. Efetov,
Emilia Morosan,
Dante M. Kennes,
Angel Rubio,
Lede Xian,
Claudia Felser,
Leslie M. Schoop,
B. Andrei Bernevig
Abstract:
The study of twisted two-dimensional (2D) materials, where twisting layers create moiré superlattices, has opened new opportunities for investigating topological phases and strongly correlated physics. While systems such as twisted bilayer graphene (TBG) and twisted transition metal dichalcogenides (TMDs) have been extensively studied, the broader potential of a seemingly infinite set of other twi…
▽ More
The study of twisted two-dimensional (2D) materials, where twisting layers create moiré superlattices, has opened new opportunities for investigating topological phases and strongly correlated physics. While systems such as twisted bilayer graphene (TBG) and twisted transition metal dichalcogenides (TMDs) have been extensively studied, the broader potential of a seemingly infinite set of other twistable 2D materials remains largely unexplored. In this paper, we define "theoretically twistable materials" as single- or multi-layer structures that allow for the construction of simple continuum models of their moiré structures. This excludes, for example, materials with a "spaghetti" of bands or those with numerous crossing points at the Fermi level, for which theoretical moiré modeling is unfeasible. We present a high-throughput algorithm that systematically searches for theoretically twistable semimetals and insulators based on the Topological 2D Materials Database. By analyzing key electronic properties, we identify thousands of new candidate materials that could host rich topological and strongly correlated phenomena when twisted. We propose representative twistable materials for realizing different types of moiré systems, including materials with different Bravais lattices, valleys, and strength of spin-orbital coupling. We provide examples of crystal growth for several of these materials and showcase twisted bilayer band structures along with simplified twisted continuum models. Our results significantly broaden the scope of moiré heterostructures and provide a valuable resource for future experimental and theoretical studies on novel moiré systems.
△ Less
Submitted 14 November, 2024;
originally announced November 2024.
-
Universal Moiré-Model-Building Method without Fitting: Application to Twisted MoTe$_2$ and WSe$_2$
Authors:
Yan Zhang,
Hanqi Pi,
Jiaxuan Liu,
Wangqian Miao,
Ziyue Qi,
Nicolas Regnault,
Hongming Weng,
Xi Dai,
B. Andrei Bernevig,
Quansheng Wu,
Jiabin Yu
Abstract:
We develop a comprehensive method to construct analytical continuum models for moiré systems directly from first-principle calculations without any parameter fitting. The core idea of this method is to interpret the terms in the continuum model as a basis, allowing us to determine model parameters as coefficients of this basis through Gram-Schmidt orthogonalization. We apply our method to twisted…
▽ More
We develop a comprehensive method to construct analytical continuum models for moiré systems directly from first-principle calculations without any parameter fitting. The core idea of this method is to interpret the terms in the continuum model as a basis, allowing us to determine model parameters as coefficients of this basis through Gram-Schmidt orthogonalization. We apply our method to twisted MoTe$_2$ and WSe$_2$ with twist angles ranging from 2.13$^\circ$ to 3.89$^\circ$, producing continuum models that exhibit excellent agreement with both energy bands and wavefunctions obtained from first-principles calculations. We further propose a strategy to integrate out the higher-energy degrees of freedom to reduce the number of the parameters in the model without sacrificing the accuracy for low-energy bands. Our findings reveal that decreasing twist angles typically need an increasing number of harmonics in the moiré potentials to accurately replicate first-principles results. We provide parameter values for all derived continuum models, facilitating further robust many-body calculations. Our approach is general and applicable to any commensurate moiré materials accessible by first-principles calculations.
△ Less
Submitted 12 November, 2024;
originally announced November 2024.
-
World-Grounded Human Motion Recovery via Gravity-View Coordinates
Authors:
Zehong Shen,
Huaijin Pi,
Yan Xia,
Zhi Cen,
Sida Peng,
Zechen Hu,
Hujun Bao,
Ruizhen Hu,
Xiaowei Zhou
Abstract:
We present a novel method for recovering world-grounded human motion from monocular video. The main challenge lies in the ambiguity of defining the world coordinate system, which varies between sequences. Previous approaches attempt to alleviate this issue by predicting relative motion in an autoregressive manner, but are prone to accumulating errors. Instead, we propose estimating human poses in…
▽ More
We present a novel method for recovering world-grounded human motion from monocular video. The main challenge lies in the ambiguity of defining the world coordinate system, which varies between sequences. Previous approaches attempt to alleviate this issue by predicting relative motion in an autoregressive manner, but are prone to accumulating errors. Instead, we propose estimating human poses in a novel Gravity-View (GV) coordinate system, which is defined by the world gravity and the camera view direction. The proposed GV system is naturally gravity-aligned and uniquely defined for each video frame, largely reducing the ambiguity of learning image-pose mapping. The estimated poses can be transformed back to the world coordinate system using camera rotations, forming a global motion sequence. Additionally, the per-frame estimation avoids error accumulation in the autoregressive methods. Experiments on in-the-wild benchmarks demonstrate that our method recovers more realistic motion in both the camera space and world-grounded settings, outperforming state-of-the-art methods in both accuracy and speed. The code is available at https://zju3dv.github.io/gvhmr/.
△ Less
Submitted 10 September, 2024;
originally announced September 2024.
-
Discovery of a metallic room-temperature d-wave altermagnet KV2Se2O
Authors:
Bei Jiang,
Mingzhe Hu,
Jianli Bai,
Ziyin Song,
Chao Mu,
Gexing Qu,
Wan Li,
Wenliang Zhu,
Hanqi Pi,
Zhongxu Wei,
Yujie Sun,
Yaobo Huang,
Xiquan Zheng,
Yingying Peng,
Lunhua He,
Shiliang Li,
Jianlin Luo,
Zheng Li,
Genfu Chen,
Hang Li,
Hongming Weng,
Tian Qian
Abstract:
Beyond conventional ferromagnetism and antiferromagnetism, altermagnetism is a recently discovered unconventional magnetic phase characterized by time-reversal symmetry breaking and spin-split band structures in materials with zero net magnetization. This distinct magnetic phase not only enriches the understanding of fundamental physical concepts but also has profound impacts on condense-matter ph…
▽ More
Beyond conventional ferromagnetism and antiferromagnetism, altermagnetism is a recently discovered unconventional magnetic phase characterized by time-reversal symmetry breaking and spin-split band structures in materials with zero net magnetization. This distinct magnetic phase not only enriches the understanding of fundamental physical concepts but also has profound impacts on condense-matter physics research and practical device applications. Spin-polarized band structures have been recently observed in semiconductors MnTe and MnTe2 with vanishing net magnetization, confirming the existence of this unconventional magnetic order. Metallic altermagnets have unique advantages for exploring novel physical phenomena related to low-energy quasiparticle excitations and for applications in spintronics as electrical conductivity in metals allows the direct manipulation of spin current through electric field. Here, through comprehensive characterization and analysis of the magnetic and electronic structures of KV2Se2O, we have unambiguously demonstrated a metallic room-temperature altermaget with d-wave spin-momentum locking. The highly anisotropic spin-polarized Fermi surfaces and the spin-density-wave order emerging in the altermagnetic phase make it an extraordinary platform for designing high-performance spintronic devices and studying many-body effects coupled with the unconventional magnetism.
△ Less
Submitted 13 August, 2024; v1 submitted 1 August, 2024;
originally announced August 2024.
-
Generating Human Motion in 3D Scenes from Text Descriptions
Authors:
Zhi Cen,
Huaijin Pi,
Sida Peng,
Zehong Shen,
Minghui Yang,
Shuai Zhu,
Hujun Bao,
Xiaowei Zhou
Abstract:
Generating human motions from textual descriptions has gained growing research interest due to its wide range of applications. However, only a few works consider human-scene interactions together with text conditions, which is crucial for visual and physical realism. This paper focuses on the task of generating human motions in 3D indoor scenes given text descriptions of the human-scene interactio…
▽ More
Generating human motions from textual descriptions has gained growing research interest due to its wide range of applications. However, only a few works consider human-scene interactions together with text conditions, which is crucial for visual and physical realism. This paper focuses on the task of generating human motions in 3D indoor scenes given text descriptions of the human-scene interactions. This task presents challenges due to the multi-modality nature of text, scene, and motion, as well as the need for spatial reasoning. To address these challenges, we propose a new approach that decomposes the complex problem into two more manageable sub-problems: (1) language grounding of the target object and (2) object-centric motion generation. For language grounding of the target object, we leverage the power of large language models. For motion generation, we design an object-centric scene representation for the generative model to focus on the target object, thereby reducing the scene complexity and facilitating the modeling of the relationship between human motions and the object. Experiments demonstrate the better motion quality of our approach compared to baselines and validate our design choices.
△ Less
Submitted 13 May, 2024;
originally announced May 2024.
-
Near-infrared metalens empowered dual-mode high resolution and large FOV microscope
Authors:
Chuang Sun,
Hailong Pi,
Kian Shen Kiang,
Jize Yan,
Jun-Yu Ou
Abstract:
The spiral phase contrast microscope can clearly distinguish the morphological information of the low contrast objects (i.e., biological samples) because of the isotropic edge-enhancement effect, while the bright field microscope can image the overall morphology of amplitude objects. However, the imaging resolution, magnification, and field of view of conventional spiral phase contrast microscopes…
▽ More
The spiral phase contrast microscope can clearly distinguish the morphological information of the low contrast objects (i.e., biological samples) because of the isotropic edge-enhancement effect, while the bright field microscope can image the overall morphology of amplitude objects. However, the imaging resolution, magnification, and field of view of conventional spiral phase contrast microscopes based on 4f filtering configuration are limited by the system's complexity. Here, we reported compact dual-mode microscopes working at near-infrared using the engineered metalens which can be tuned between the spiral phase contrast imaging and bright field imaging by polarization control. The metalens combines the high-resolution objective lens and polarization-controlled phase filter into a single-layer nanofins array. We demonstrated two infinity-corrected microscope systems to achieve subwavelength resolution (0.7 times of wavelength), large magnification (58X), and large field of view (600um times 800um). Unstained onion epidermal is imaged by the microscope to show the dual-mode imaging ability for the biological sample. Finally, a singlet dual-mode microscope system is demonstrated to show the edge-detection application for industrial standards. Our results could open new opportunities in applications of biological imaging, industrial machine vision, and semiconductor inspection.
△ Less
Submitted 18 February, 2024;
originally announced February 2024.
-
First-principles methodology for studying magnetotransport in narrow-gap semiconductors: an application to Zirconium Pentatelluride ZrTe5
Authors:
Hanqi Pi,
Shengnan Zhang,
Yang Xu,
Zhong Fang,
Hongming Weng,
Quansheng Wu
Abstract:
The origin of anomalous resistivity peak and accompanied sign reversal of Hall resistivity of ZrTe$_5$ has been under debate for a long time. Although various theoretical models have been proposed to account for these intriguing transport properties, a systematic study from first principles view is still lacking. In this work, we present a first principles calculation combined with Boltzmann trans…
▽ More
The origin of anomalous resistivity peak and accompanied sign reversal of Hall resistivity of ZrTe$_5$ has been under debate for a long time. Although various theoretical models have been proposed to account for these intriguing transport properties, a systematic study from first principles view is still lacking. In this work, we present a first principles calculation combined with Boltzmann transport theory to investigate the transport properties in narrow-gap semiconductors at different temperatures and doping densities within the relaxation time approximation. Regarding the sensitive temperature-dependent chemical potential and relaxation time of semiconductors, we take proper approximation to simulate these two variables, and then comprehensively study the transport properties of ZrTe$_5$ both in the absence and presence of an applied magnetic field. Without introducing topological phases and correlation interactions, we qualitatively reproduced crucial features observed in experiments, including zero-field resistivity anomaly, nonlinear Hall resistivity with sign reversal, and non-saturating magnetoresistance at high temperatures. Our calculation allows a systematic interpretation of the observed properties in terms of multi-carrier and Fermi surface geometry. Our method can be extended to other narrow-gap semiconductors and further pave the way to explore interesting and novel transport properties of this field.
△ Less
Submitted 26 January, 2024;
originally announced January 2024.
-
Complex field-, temperature-, and angle-dependent Hall effects from intrinsic Fermi surface revealed by first-principles calculations
Authors:
ShengNan Zhang,
Zhihao Liu,
Hanqi Pi,
Zhong Fang,
Hongming Weng,
QuanSheng Wu
Abstract:
The Hall effect, ever intriguing since its discovery, has spurred the exploration of its phenomena, intensified by advances in topology and novel materials. Differentiating the ordinary Hall effect from extraordinary properties like the anomalous Hall effect (AHE) is challenging, especially in materials with topological origins. In our study, we leverage semiclassical Boltzmann transport theory an…
▽ More
The Hall effect, ever intriguing since its discovery, has spurred the exploration of its phenomena, intensified by advances in topology and novel materials. Differentiating the ordinary Hall effect from extraordinary properties like the anomalous Hall effect (AHE) is challenging, especially in materials with topological origins. In our study, we leverage semiclassical Boltzmann transport theory and first-principles calculations within the relaxation time approximation to analyze Hall effects comprehensively. We have found that the complex magnetic field dependence of ordinary Hall effect, including the sign reversals, appearing of plateau and nonlinearity, can be understood and reproduced by our approach both for multiband models and realistic topological materials of ZrSiS and PtTe2. The Hall resistivity versus temperature and magnetic fields can be well scaled, similar to Kohler's rule for longitudinal resistivity. This methodology can also accurately model the angular dependent Hall effects such as planar Hall effects of bismuth. These findings indicate that the dependencies of various Hall effects and magnetoresistance on magnetic fields are mainly determined by the details of Fermi surface and the relaxation time. The intrinsic Fermi surface determines the carriers' density, type, and velocity, while the later is mostly influenced by extrinsic factors, such as quality of sample with defects, impurities, and domains. This insight might simplify the understanding of several seemingly complex transport phenomena in nonmagnetic materials, with no need for hypotheses of other sophisticated mechanisms, such as magnetization-induced AHE, Lifshitz transition-induced changes in carrier type, exotic orders like charge density wave, or some delicate scattering of carriers with chiral or nonreciprocal dependence. Finally, we also discussed the Hall effects contribute from the Berry curvature.
△ Less
Submitted 20 September, 2025; v1 submitted 26 January, 2024;
originally announced January 2024.
-
Tunable on-chip optical traps for levitating particles based on single-layer metasurface
Authors:
Chuang Sun,
Hailong Pi,
Kian Shen Kiang,
Tiberius S. Georgescu,
Jun-Yu Ou,
Hendrik Ulbricht,
Jize Yan
Abstract:
Optically levitated multiple nanoparticles has emerged as a platform for studying complex fundamental physics such as non-equilibrium phenomena, quantum entanglement, and light-matter interaction, which could be applied for sensing weak forces and torques with high sensitivity and accuracy. An optical trapping landscape of increased complexity is needed to engineer the interaction between levitate…
▽ More
Optically levitated multiple nanoparticles has emerged as a platform for studying complex fundamental physics such as non-equilibrium phenomena, quantum entanglement, and light-matter interaction, which could be applied for sensing weak forces and torques with high sensitivity and accuracy. An optical trapping landscape of increased complexity is needed to engineer the interaction between levitated particles beyond the single harmonic trap. However, existing platforms based on spatial light modulators for studying interactions between levitated particles suffered from low efficiency, instability at focal points, the complexity of optical systems, and the scalability for sensing applications. Here, we experimentally demonstrated that a metasurface which forms two diffraction-limited focal points with a high numerical aperture (0.9) and high efficiency (31%) can generate tunable optical potential wells without any intensity fluctuations. A bistable potential and double potential wells were observed in the experiment by varying the focal points distance, and two nanoparticles were levitated in double potential wells for hours, which could be used for investigating the levitated particles nonlinear dynamics, thermal dynamics, and optical binding. This would pave the way for scaling the number of levitated optomechanical devices or realizing paralleled levitated sensors.
△ Less
Submitted 16 January, 2024;
originally announced January 2024.
-
Hierarchical Generation of Human-Object Interactions with Diffusion Probabilistic Models
Authors:
Huaijin Pi,
Sida Peng,
Minghui Yang,
Xiaowei Zhou,
Hujun Bao
Abstract:
This paper presents a novel approach to generating the 3D motion of a human interacting with a target object, with a focus on solving the challenge of synthesizing long-range and diverse motions, which could not be fulfilled by existing auto-regressive models or path planning-based methods. We propose a hierarchical generation framework to solve this challenge. Specifically, our framework first ge…
▽ More
This paper presents a novel approach to generating the 3D motion of a human interacting with a target object, with a focus on solving the challenge of synthesizing long-range and diverse motions, which could not be fulfilled by existing auto-regressive models or path planning-based methods. We propose a hierarchical generation framework to solve this challenge. Specifically, our framework first generates a set of milestones and then synthesizes the motion along them. Therefore, the long-range motion generation could be reduced to synthesizing several short motion sequences guided by milestones. The experiments on the NSM, COUCH, and SAMP datasets show that our approach outperforms previous methods by a large margin in both quality and diversity. The source code is available on our project page https://zju3dv.github.io/hghoi.
△ Less
Submitted 3 October, 2023;
originally announced October 2023.
-
Non-centrosymmetric, transverse structural modulation in SrAl4, and elucidation of its origin in the BaAl4 family of compounds
Authors:
Sitaram Ramakrishnan,
Surya Rohith Kotla,
Hanqi Pi,
Bishal Baran Maity,
Jia Chen,
Jin-Ke Bao,
Zhaopeng Guo,
Masaki Kado,
Harshit Agarwal,
Claudio Eisele,
Minoru Nohara,
Leila Noohinejad,
Hongming Weng,
Srinivasan Ramakrishnan,
Arumugam Thamizhavel,
Sander van Smaalen
Abstract:
At ambient conditions SrAl4 adopts the BaAl4 structure type with space group I4/mmm. It undergoes a charge-density-wave (CDW) transition at TCDW = 243 K, followed by a structural transition at TS = 87 K. Temperature-dependent single-crystal X-ray diffraction (SXRD) leads to the observation of incommensurate superlattice reflections at q = σc* with σ= 0.1116 at 200 K. The CDW has orthorhombic symme…
▽ More
At ambient conditions SrAl4 adopts the BaAl4 structure type with space group I4/mmm. It undergoes a charge-density-wave (CDW) transition at TCDW = 243 K, followed by a structural transition at TS = 87 K. Temperature-dependent single-crystal X-ray diffraction (SXRD) leads to the observation of incommensurate superlattice reflections at q = σc* with σ= 0.1116 at 200 K. The CDW has orthorhombic symmetry with the acentric superspace group F222(00sigma)00s, where F222 is a subgroup of Fmmm as well as of I4/mmm. Atomic displacements mainly represent a transverse wave, with displacements that are 90 deg out of phase between the two diagonal directions of the I-centered unit cell, resulting in a helical wave. Small longitudinal displacements are provided by the second harmonic modulation. The orthorhombic phase realized in SrAl4 is similar to that found in EuAl4. Electronic structure calculations and phonon calculations by density functional theory (DFT) have failed to reveal the mechanism of CDW formation. However, DFT reveals that Al atoms dominate the density of states near the Fermi level, thus, corroborating the SXRD measurements. SrAl4 remains incommensurately modulated at the structural transition, where the symmetry lowers from orthorhombic to b-unique monoclinic. We have identified a simple criterion, that correlates the presence of a phase transition with the interatomic distances. Only those compounds XAl4-xGax(X = Ba, Eu, Sr, Ca; 0 < x <4) undergo phase transitions, for which the ratio c/a falls within the narrow range 2.51 < c/a < 2.54.
△ Less
Submitted 16 March, 2024; v1 submitted 16 September, 2023;
originally announced September 2023.
-
Gate-tunable multiband transport in ZrTe5 thin devices
Authors:
Yonghe Liu,
Hanqi Pi,
Kenji Watanabe,
Takashi Taniguchi,
Genda Gu,
Qiang Li,
Hongming Weng,
Quansheng Wu,
Yongqing Li,
Yang Xu
Abstract:
Interest in ZrTe5 has been reinvigorated in recent years owing to its potential for hosting versatile topological electronic states and intriguing experimental discoveries. However, the mechanism of many of its unusual transport behaviors remains controversial, for example, the characteristic peak in the temperature-dependent resistivity and the anomalous Hall effect. Here, through employing a cle…
▽ More
Interest in ZrTe5 has been reinvigorated in recent years owing to its potential for hosting versatile topological electronic states and intriguing experimental discoveries. However, the mechanism of many of its unusual transport behaviors remains controversial, for example, the characteristic peak in the temperature-dependent resistivity and the anomalous Hall effect. Here, through employing a clean dry-transfer fabrication method under inert environment, we successfully obtain high-quality ZrTe5 thin devices that exhibit clear dual-gate tunability and ambipolar field effects. Such devices allow us to systematically study the resistance peak as well as the Hall effect at various doping densities and temperatures, revealing the contribution from electron-hole asymmetry and multiple-carrier transport. By comparing with theoretical calculations, we suggest a simplified semiclassical two-band model to explain the experimental observations. Our work helps to resolve the long-standing puzzles on ZrTe5 and could potentially pave the way for realizing novel topological states in the two-dimensional limit.
△ Less
Submitted 15 May, 2023;
originally announced May 2023.
-
A Comprehensive Comparison of Projections in Omnidirectional Super-Resolution
Authors:
Huicheng Pi,
Senmao Tian,
Ming Lu,
Jiaming Liu,
Yandong Guo,
Shunli Zhang
Abstract:
Super-Resolution (SR) has gained increasing research attention over the past few years. With the development of Deep Neural Networks (DNNs), many super-resolution methods based on DNNs have been proposed. Although most of these methods are aimed at ordinary frames, there are few works on super-resolution of omnidirectional frames. In these works, omnidirectional frames are projected from the 3D sp…
▽ More
Super-Resolution (SR) has gained increasing research attention over the past few years. With the development of Deep Neural Networks (DNNs), many super-resolution methods based on DNNs have been proposed. Although most of these methods are aimed at ordinary frames, there are few works on super-resolution of omnidirectional frames. In these works, omnidirectional frames are projected from the 3D sphere to a 2D plane by Equi-Rectangular Projection (ERP). Although ERP has been widely used for projection, it has severe projection distortion near poles. Current DNN-based SR methods use 2D convolution modules, which is more suitable for the regular grid. In this paper, we find that different projection methods have great impact on the performance of DNNs. To study this problem, a comprehensive comparison of projections in omnidirectional super-resolution is conducted. We compare the SR results of different projection methods. Experimental results show that Equi-Angular cube map projection (EAC), which has minimal distortion, achieves the best result in terms of WS-PSNR compared with other projections. Code and data will be released.
△ Less
Submitted 13 April, 2023;
originally announced April 2023.
-
Clique Densification in Networks
Authors:
Haochen Pi,
Keith Burghardt,
Allon G. Percus,
Kristina Lerman
Abstract:
Real-world networks are rarely static. Recently, there has been increasing interest in both network growth and network densification, in which the number of edges scales superlinearly with the number of nodes. Less studied but equally important, however, are scaling laws of higher-order cliques, which can drive clustering and network redundancy. In this paper, we study how cliques grow with networ…
▽ More
Real-world networks are rarely static. Recently, there has been increasing interest in both network growth and network densification, in which the number of edges scales superlinearly with the number of nodes. Less studied but equally important, however, are scaling laws of higher-order cliques, which can drive clustering and network redundancy. In this paper, we study how cliques grow with network size, by analyzing several empirical networks from emails to Wikipedia interactions. Our results show superlinear scaling laws whose exponents increase with clique size, in contrast to predictions from a previous model. We then show that these results are in qualitative agreement with a new model that we propose, the Local Preferential Attachment Model, where an incoming node links not only to a target node but also to its higher-degree neighbors. Our results provide new insights into how networks grow and where network redundancy occurs.
△ Less
Submitted 7 April, 2023;
originally announced April 2023.
-
A Joint Modeling of Vision-Language-Action for Target-oriented Grasping in Clutter
Authors:
Kechun Xu,
Shuqi Zhao,
Zhongxiang Zhou,
Zizhang Li,
Huaijin Pi,
Yue Wang,
Rong Xiong
Abstract:
We focus on the task of language-conditioned grasping in clutter, in which a robot is supposed to grasp the target object based on a language instruction. Previous works separately conduct visual grounding to localize the target object, and generate a grasp for that object. However, these works require object labels or visual attributes for grounding, which calls for handcrafted rules in planner a…
▽ More
We focus on the task of language-conditioned grasping in clutter, in which a robot is supposed to grasp the target object based on a language instruction. Previous works separately conduct visual grounding to localize the target object, and generate a grasp for that object. However, these works require object labels or visual attributes for grounding, which calls for handcrafted rules in planner and restricts the range of language instructions. In this paper, we propose to jointly model vision, language and action with object-centric representation. Our method is applicable under more flexible language instructions, and not limited by visual grounding error. Besides, by utilizing the powerful priors from the pre-trained multi-modal model and grasp model, sample efficiency is effectively improved and the sim2real problem is relived without additional data for transfer. A series of experiments carried out in simulation and real world indicate that our method can achieve better task success rate by less times of motion under more flexible language instructions. Moreover, our method is capable of generalizing better to scenarios with unseen objects and language instructions. Our code is available at https://github.com/xukechun/Vision-Language-Grasping
△ Less
Submitted 31 October, 2024; v1 submitted 24 February, 2023;
originally announced February 2023.
-
Magnetic bulk photovoltaic effect as a probe of magnetic structures of $EuSn_2As_2$
Authors:
Hanqi Pi,
Shuai Zhang,
Hongming Weng
Abstract:
The bulk photovoltaic effect (BPVE) is a second-order optical process in noncentrosymmetric materials that converts the light into DC currents. BPVE is classified into shift current and injection current according to the generation mechanisms, whose dependence on the polarization of light is sensitive to the spatial and time-reversal symmetry of materials. In this work, we present a comprehensive…
▽ More
The bulk photovoltaic effect (BPVE) is a second-order optical process in noncentrosymmetric materials that converts the light into DC currents. BPVE is classified into shift current and injection current according to the generation mechanisms, whose dependence on the polarization of light is sensitive to the spatial and time-reversal symmetry of materials. In this work, we present a comprehensive study on the BPVE response of $EuSn_2As_2$ with different magnetic structures through symmetry analysis and first-principles calculation. We demonstrate that the interlayer antiferromagnetic (AFM) $EuSn_2As_2$ of even-layer breaks the inversion symmetry and has the second-order optical responses. Moreover, the bilayer AFM $EuSn_2As_2$ not only displays distinct BPVE responses when magnetic moments align in different directions, but also shows symmetry-related responses in two phases which have mutually perpendicular in-plane magnetic moments. Due to the dependence of BPVE responses on the polarization of light and magnetic symmetry, these magnetic structures can be distinguished by the circular polarized light with well-designed experiments. Our work demonstrates the feasibility of the BPVE response as a tool to probe the magnetic structure.
△ Less
Submitted 19 February, 2023;
originally announced February 2023.
-
Optical spectroscopy and band structure calculations of structural phase transition in the Vanadium-based kagome metal ScV$_6$Sn$_6$
Authors:
Tianchen Hu,
Hanqi Pi,
Shuxiang Xu,
Li Yue,
Qiong Wu,
Qiaomei Liu,
Sijie Zhang,
Rongsheng Li,
Xinyu Zhou,
Jiayu Yuan,
Dong Wu,
Tao Dong,
Hongming Weng,
Nanlin Wang
Abstract:
In condensed matter physics, materials with kagome lattice display a range of exotic quantum states, including charge density wave (CDW), superconductivity and magnetism. Recently, the intermetallic kagome metal ScV6Sn6 was discovered to undergo a first-order structural phase transition with the formation of a root3xroot3x3 CDW at around 92 K. The bulk electronic band properties are crucial to und…
▽ More
In condensed matter physics, materials with kagome lattice display a range of exotic quantum states, including charge density wave (CDW), superconductivity and magnetism. Recently, the intermetallic kagome metal ScV6Sn6 was discovered to undergo a first-order structural phase transition with the formation of a root3xroot3x3 CDW at around 92 K. The bulk electronic band properties are crucial to understanding the origin of the structural phase transition. Here, we conducted an optical spectroscopy study in combination with band structure calculations across the structural transition. Our findings showed abrupt changes in the optical reflectivity/conductivity spectra as a result of the structural transition, without any observable gap formation behavior. The optical measurements and band calculations actually reveal a sudden change of the band structure after transition. It is important to note that this phase transition is of the first-order type, which distinguishes it from conventional density-wave type condensations. Our results provide an insight into the origin of the structural phase transition in this new and unique kagome lattice.
△ Less
Submitted 17 February, 2023; v1 submitted 7 November, 2022;
originally announced November 2022.
-
Linear-in-Frequency Optical Conductivity over a broad range in the three-dimensional Dirac semimetal candidate Ir$_2$In$_8$Se
Authors:
S. X. Xu,
H. Q. Pi,
R. S. Li,
T. C. Hu,
Q. Wu,
D. Wu,
H. M. Weng,
N. L. Wang
Abstract:
The optical conductivity of the new Dirac semimetal candidate Ir$_2$In$_8$Se is measured in a frequency range from 40 to 30000 cm$^{-1}$ at temperatures from 300 K down to 10 K. The measurement reveals that the compound is a low carrier density metal. We find that the real part of the conductivity $σ_1(ω)$ is linear in frequency over a broad range from 500 to 4000 cm$^{-1}$ at 300 K and varies sli…
▽ More
The optical conductivity of the new Dirac semimetal candidate Ir$_2$In$_8$Se is measured in a frequency range from 40 to 30000 cm$^{-1}$ at temperatures from 300 K down to 10 K. The measurement reveals that the compound is a low carrier density metal. We find that the real part of the conductivity $σ_1(ω)$ is linear in frequency over a broad range from 500 to 4000 cm$^{-1}$ at 300 K and varies slightly with cooling. This linearity strongly suggests the presence of three-dimensional linear electronic bands with band crossings near the Fermi level. Band structure calculations indicate the presence of type-II Dirac points. By comparing our data with the optical conductivity computed from the band structure, we conclude that the observed linear dependence mainly originates from the Dirac cones and the transition between the Dirac cones and the next lower bands. In addition, a weak energy gap feature is resolved below the charge density wave phase transition temperature in reflectivity spectra. An enhanced structure arising from the imperfect Fermi surface nesting is identified in the electronic susceptibility function, suggesting a Fermi surface nesting driven instability.
△ Less
Submitted 1 September, 2022;
originally announced September 2022.
-
Highly in-plane anisotropic optical properties of fullerene monolayers
Authors:
Danwen Yuan,
Hanqi Pi,
Yi Jiang,
Yuefang Hu,
Liqin Zhou,
Yujin Jia,
Gang Su,
Zhong Fang,
Hongming Weng,
Xinguo Ren,
Wei Zhang
Abstract:
Both the intrinsic anisotropic optical materials and fullerene-assembled 2D materials have attracted a lot of interests in fundamental science and potential applications. The synthesis of a monolayer (ML) fullerene makes the combination of these two features plausible. In this work, using first-principles calculations, we systematically study the electronic structure, optical properties of quasi-h…
▽ More
Both the intrinsic anisotropic optical materials and fullerene-assembled 2D materials have attracted a lot of interests in fundamental science and potential applications. The synthesis of a monolayer (ML) fullerene makes the combination of these two features plausible. In this work, using first-principles calculations, we systematically study the electronic structure, optical properties of quasi-hexagonal phase (qHP) ML and quasi-tetragonal phase (qTP) ML fullerenes. The calculations of qHP ML show that it is a semi-conductor with small anisotropic optical absorption, which agrees with the recent experimental measurements. However, the results for qTP ML reveal that it is a semimetal with highly in-plane anisotropic absorption. The dichroic ratio, namely the absorption ratio of $x$- and $y$-polarized light $α$$_x$$_x$/$α$$_y$$_y$, is around 12 at photon energy of 0.29 eV. This anisotropy is much more pronounced when the photon energy is between 0.7 and 1.4 eV, where $α$$_x$$_x$ becomes nearly zero while $α$$_y$$_y$ is more than two orders of magnitude larger. This indicates qTP ML as a candidate for long-pursuit lossless metal and a potential material for atomically thin polarizer. We hope this will stimulate further experimental efforts in the study of qTP ML and other fullerene-assembled 2D materials.
△ Less
Submitted 22 July, 2022;
originally announced July 2022.
-
E-NeRV: Expedite Neural Video Representation with Disentangled Spatial-Temporal Context
Authors:
Zizhang Li,
Mengmeng Wang,
Huaijin Pi,
Kechun Xu,
Jianbiao Mei,
Yong Liu
Abstract:
Recently, the image-wise implicit neural representation of videos, NeRV, has gained popularity for its promising results and swift speed compared to regular pixel-wise implicit representations. However, the redundant parameters within the network structure can cause a large model size when scaling up for desirable performance. The key reason of this phenomenon is the coupled formulation of NeRV, w…
▽ More
Recently, the image-wise implicit neural representation of videos, NeRV, has gained popularity for its promising results and swift speed compared to regular pixel-wise implicit representations. However, the redundant parameters within the network structure can cause a large model size when scaling up for desirable performance. The key reason of this phenomenon is the coupled formulation of NeRV, which outputs the spatial and temporal information of video frames directly from the frame index input. In this paper, we propose E-NeRV, which dramatically expedites NeRV by decomposing the image-wise implicit neural representation into separate spatial and temporal context. Under the guidance of this new formulation, our model greatly reduces the redundant model parameters, while retaining the representation ability. We experimentally find that our method can improve the performance to a large extent with fewer parameters, resulting in a more than $8\times$ faster speed on convergence. Code is available at https://github.com/kyleleey/E-NeRV.
△ Less
Submitted 17 July, 2022;
originally announced July 2022.
-
Switchable Topological Phase Transition and Novel Nonlinear Optical Properties in ReC2H Monolayer
Authors:
Chunmei Zhang,
Hanqi Pi,
Liqin Zhou,
Si Li,
Jian Zhou,
Aijun Du,
Hongming Weng
Abstract:
Extensive investigations on topological phase transition (TPT) in three-dimensional compounds have been done. whereas, rare in two-dimensional systems, let alone noncentrosymmetric materials. In this work, based on first-principles calculations, we explore an inversion symmetry broken structural ReC2H monolayer. We reveal that it undergoes two TPTs, namely from normal insulator to Z2 topological i…
▽ More
Extensive investigations on topological phase transition (TPT) in three-dimensional compounds have been done. whereas, rare in two-dimensional systems, let alone noncentrosymmetric materials. In this work, based on first-principles calculations, we explore an inversion symmetry broken structural ReC2H monolayer. We reveal that it undergoes two TPTs, namely from normal insulator to Z2 topological insulator and back to normal insulator, at the critical biaxial strain of 2.3% and 7.8%, respectively. The band inversion occurs at the generic momentum in the first TPT, while at high symmetric K point in the second one. Usually, band inversion is identified by the exchange in the components or irreducible representations of the wavefunctions. These quantities can be easily obtained in theoretical calculation but hard to be detected in experimental techniques like Angle-resolved photoemission spectroscopy. It is well known that nonlinear optical (NLO) response is very sensitive to the components and symmetries of the engaged bands, which also incorporates information of band topology. Therefore, we study the shift current, one of the widely explored NLO responses in noncentrosymmetric systems, during the two TPTs. We find that in both cases band inversion leads to the sign change of shift vectors around the momenta where the bandgap closes and reopens. Whereas the shift current, as the overall contribution of shift vectors weighted by the absorption rate in the whole Brillouin zone, may keep its direction. This work offers insight that a scrutinized examination is highly demanded in utilizing shift current to detect TPT.
△ Less
Submitted 3 February, 2022;
originally announced February 2022.
-
Degradation Mechanism of Perovskite under High Charge Carrier Density Condition
Authors:
Guohui Li,
Huihui Pi,
Yanfu Wei,
Bolin Zhou,
Ya Gao,
Rong Wen,
Yuying Hao,
Han Zhang,
Beng S. Ong,
Yanxia Cui
Abstract:
Extensive studies have focused on degradation of perovskite at low charge carrier density (<10^16 cm^-3), but few have surveyed the degradation mechanism at high charge carrier density (~10^18 cm^-3). Here, we investigate the degradation mechanisms of perovskite under high charge carrier conditions. Unlike the observations in previous works, we find that MAPbI3 degradation starts at surface defect…
▽ More
Extensive studies have focused on degradation of perovskite at low charge carrier density (<10^16 cm^-3), but few have surveyed the degradation mechanism at high charge carrier density (~10^18 cm^-3). Here, we investigate the degradation mechanisms of perovskite under high charge carrier conditions. Unlike the observations in previous works, we find that MAPbI3 degradation starts at surface defects and progressing from the surface defects towards neighboring regions under high charge carrier density condition. By using PbI2 passivation, the defect-initiated degradation is significantly suppressed and the nanoplatelet degrades in a layer-by-layer way, enabling the MAPbI3 laser sustain for 4500 s (2.7*10^7 pulses), which is almost 3 times longer than that of the nanoplatelet laser without passivation. Meanwhile, the PbI2 passivated MAPbI3 nanoplatelet laser with the nanoplatelet cavity displaying a maximum quality factor up to ~7800, the highest reported for all MAPbI3 nanoplatelet cavities. Furthermore, a high stability MAPbI3 nanoplatelet laser that can last for 8500 s (5.1*10^7 pulses) is demonstrated based on a dual passivation strategy, by retarding the defect-initiated degradation and surface-initiated degradation, simultaneously. This work provides in-depth insights for understanding the degradation of perovskite at high charge carrier density.
△ Less
Submitted 16 December, 2021;
originally announced December 2021.