-
Fashion Outfit Generation via Unified Sequential Composition Models
Authors:
Kaicheng Pang,
Xingxing Zou,
Ruohan Xu,
Waikeung Wong
Abstract:
The task of synthesizing stylistically coherent fashion outfits from massive item libraries, known as fashion outfit generation, remains a non-trivial challenge, primarily due to the non-monotonic and implicit nature of aesthetic compatibility, coupled with the exponentially large combinatorial search space. In this paper, we formalize this task as Constrained Ensemble Generation (CEG) and model i…
▽ More
The task of synthesizing stylistically coherent fashion outfits from massive item libraries, known as fashion outfit generation, remains a non-trivial challenge, primarily due to the non-monotonic and implicit nature of aesthetic compatibility, coupled with the exponentially large combinatorial search space. In this paper, we formalize this task as Constrained Ensemble Generation (CEG) and model it as a finite-horizon deterministic Markov Decision Process. To address CEG in fashion, we propose the Unified Sequential Composition Model (USCM), which jointly models set-level compatibility and latent composition intents. Guided by USCM's learned priors, a Latent Expansion Monte Carlo Tree Search (LE-MCTS) mechanism is proposed to handle item retrieval during composition, balancing local aesthetic synergy with global structural balance. Extensive experiments on the Polyvore Outfits dataset, along with zero-shot evaluations on the iFashion and PolyvoreU datasets, demonstrate that our framework achieves state-of-the-art performance across independent human preference evaluations, automated aesthetic proxies, and structural validity metrics for constrained fashion outfit generation.
△ Less
Submitted 13 August, 2026;
originally announced August 2026.
-
On the extension of Kähler currents on compact complex manifolds
Authors:
Jiafu Ning,
Kai Pang,
Haoyuan Sun,
Zhiwei Wang,
Xiangyu Zhou
Abstract:
Let $(X,ω)$ be a compact Kähler manifold and let $V\subset X$ be a closed complex submanifold. Coman-Guedj-Zeriahi proposed the problem: is every $ω|_V$-plurisubharmonic function on $V$ the restriction of an $ω$-plurisubharmonic function on $X$? In this paper, we solve this problem affirmatively, even for a compact Hermitian manifold.
Let $(X,ω)$ be a compact Kähler manifold and let $V\subset X$ be a closed complex submanifold. Coman-Guedj-Zeriahi proposed the problem: is every $ω|_V$-plurisubharmonic function on $V$ the restriction of an $ω$-plurisubharmonic function on $X$? In this paper, we solve this problem affirmatively, even for a compact Hermitian manifold.
△ Less
Submitted 2 August, 2026;
originally announced August 2026.
-
Capacity Stability of Complex Monge-Ampère Equations with Moving Prescribed Singularities
Authors:
Kai Pang,
Haoyuan Sun,
Zhiwei Wang,
Xiangyu Zhou
Abstract:
For complex Monge-Ampère equations with moving big cohomology classes and prescribed model singularities of positive Monge-Ampère mass, we prove that, under total variation convergence of the right-hand side non-pluripolar positive Radon measures, convergence of the prescribed model potentials in Monge-Ampère capacity is equivalent to convergence in capacity of the associated normalized solutions.…
▽ More
For complex Monge-Ampère equations with moving big cohomology classes and prescribed model singularities of positive Monge-Ampère mass, we prove that, under total variation convergence of the right-hand side non-pluripolar positive Radon measures, convergence of the prescribed model potentials in Monge-Ampère capacity is equivalent to convergence in capacity of the associated normalized solutions. We further prove that the ceiling operator coincides with the singularity envelope for potentials associated to a big $(1,1)$-class, regardless of their Monge-Ampère mass, thereby resolving a conjecture of Darvas-Di Nezza-Lu. Consequently, the singularity envelope is idempotent without the positivity assumption on the mass.
△ Less
Submitted 14 July, 2026;
originally announced July 2026.
-
On the Modelling of the Hydrodynamic Drag of Mangroves
Authors:
Khang Ee Pang,
Zhi Yung Tay
Abstract:
Mangroves are increasingly promoted as nature-based solutions for coastal protection, yet many existing models neglect the vertical variation of vegetation biomass, leading to oversimplified representations of root-flow interactions. In this study, we introduce a generalised parametrisation of the mangrove vegetation profile that is applicable across multiple mangrove species and derive a wave att…
▽ More
Mangroves are increasingly promoted as nature-based solutions for coastal protection, yet many existing models neglect the vertical variation of vegetation biomass, leading to oversimplified representations of root-flow interactions. In this study, we introduce a generalised parametrisation of the mangrove vegetation profile that is applicable across multiple mangrove species and derive a wave attenuation model that explicitly accounts for the mangrove root characteristics. Based on this parametrisation, we propose a simplified mangrove representation that reproduces a prescribed drag force profile and is suitable for both computational fluid dynamics simulations and experimental fabrication. The hydrodynamic performance of the proposed model is evaluated using OpenFOAM simulations. Our results show that the wave attenuation effectiveness of mangroves is frequency-selective and species dependent. This nonlinear behaviour contrasts with classical vegetation models and reveals a previously unrecognized mechanism by which mangrove root characteristics govern coastal protection.
△ Less
Submitted 10 June, 2026;
originally announced June 2026.
-
Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation
Authors:
Zhifei Xie,
Kaiyu Pang,
Haobin Zhang,
Deheng Ye,
Xiaobin Hu,
Shuicheng Yan,
Chunyan Miao
Abstract:
Despite rapid advances in automatic speech recognition (ASR) and large audio-language models, robust recognition in real-world environments remains limited by an "acoustic robustness bottleneck": models often lose acoustic grounding and produce omissions or hallucinations under severe, compositional distortions. We propose Mega-ASR, a unified ASR-in-the-wild framework that combines scalable compou…
▽ More
Despite rapid advances in automatic speech recognition (ASR) and large audio-language models, robust recognition in real-world environments remains limited by an "acoustic robustness bottleneck": models often lose acoustic grounding and produce omissions or hallucinations under severe, compositional distortions. We propose Mega-ASR, a unified ASR-in-the-wild framework that combines scalable compound-data construction with progressive acoustic-to-semantic optimization. We introduce Voices-in-the-Wild-2M, covering 7 classic acoustic phenomena and 54 physically plausible compound scenarios, and train Mega-ASR with Acoustic-to-Semantic Progressive Supervised Fine-Tuning and Dual-Granularity WER-Gated Policy Optimization. Extensive experiments demonstrate that Mega-ASR achieves significant advantages over prior state-of-the-art systems on adverse-condition ASR benchmarks (45.69% vs. 54.01% on VOiCES R4-B-F, and 21.49% vs. 29.34% on NOIZEUS Sta-0). On complex compositional acoustic scenarios, Mega-ASR further delivers over 30% relative WER reduction against strong open- and closed-source baselines, establishing a scalable paradigm for robust ASR in-the-wild.
△ Less
Submitted 19 May, 2026;
originally announced May 2026.
-
NoiseRater: Meta-Learned Noise Valuation for Diffusion Model Training
Authors:
Fang Wu,
Haokai Zhao,
Da Xing,
Hanqun Cao,
Tinson Xu,
Yanchao Li,
Xiangru Tang,
Zehong Wang,
Aaron Tu,
Kuan Pang,
Hanchen Wang,
Hongbin Lin,
Zeqi Zhou,
Yinxi Li,
Peng Xia,
Li Erran Li,
Molei Tao,
Jure Leskovec,
Aditya Joshi,
Yejin Choi
Abstract:
Diffusion models have achieved remarkable success across a wide range of generative tasks, yet their training paradigm largely treats injected noise as uniformly informative. In this work, we challenge this assumption and introduce NoiseRater, a meta-learning framework for instance-level noise valuation in diffusion model training. We propose a parametric noise rater that assigns importance scores…
▽ More
Diffusion models have achieved remarkable success across a wide range of generative tasks, yet their training paradigm largely treats injected noise as uniformly informative. In this work, we challenge this assumption and introduce NoiseRater, a meta-learning framework for instance-level noise valuation in diffusion model training. We propose a parametric noise rater that assigns importance scores to individual noise realizations conditioned on data and timestep, enabling adaptive reweighting of the training objective. The rater is trained via bilevel optimization to improve downstream validation performance after inner-loop diffusion updates. To enable efficient deployment, we further design a decoupled two-stage pipeline that transitions from soft weighting during meta-training to hard noise selection during standard training. Extensive experiments on FFHQ and ImageNet demonstrate that not all noise samples contribute equally, and that prioritizing informative noise improves both training efficiency and generation quality. Our results establish noise valuation as a complementary and previously underexplored axis for improving diffusion model training. Our code is available at: https://anonymous.4open.science/r/NoiseRater-DEB116.
△ Less
Submitted 2 May, 2026;
originally announced May 2026.
-
Proteo-R1: Reasoning Foundation Models for De Novo Protein Design
Authors:
Fang Wu,
Weihao Xuan,
Heli Qi,
Hanqun Cao,
Heng-Jui Chang,
Zeqi Zhou,
Haokai Zhao,
Ma Jian,
Carl Ma,
Yu-Chi Cheng,
Kuan Pang,
Xiangru Tang,
Zehong Wang,
Guanlue Li,
Hanchen Wang,
Kejun Ying,
Pan Lu,
Chiho Im,
Seungju Han,
Peng Xia,
Tinson Xu,
Yinxi Li,
Deyao Zhu,
Pheng-Ann Heng,
Naoto Yokoya
, et al. (4 additional authors not shown)
Abstract:
Deep learning in de novo protein design has achieved atomic-level fidelity. However, existing models remain largely non-deliberative: they directly synthesize molecular geometries without explicitly reasoning about which residues or interactions are functionally essential. As a result, design decisions are entangled with continuous sampling dynamics, limiting interpretability, controllability, and…
▽ More
Deep learning in de novo protein design has achieved atomic-level fidelity. However, existing models remain largely non-deliberative: they directly synthesize molecular geometries without explicitly reasoning about which residues or interactions are functionally essential. As a result, design decisions are entangled with continuous sampling dynamics, limiting interpretability, controllability, and systematic reuse of biochemical knowledge. We introduce Proteo-R1, a reasoning-guided protein design framework that explicitly decouples molecular understanding from geometric generation. Proteo-R1 adopts a dual-expert architecture in which a multimodal large language model (MLLM) serves as an understanding expert, analyzing protein sequences, structures, and textual context to identify key functional residues that govern binding and specificity. These residue-level decisions are then passed as hard constraints to a separate diffusion-based generation expert, which performs conditional co-design while respecting the fixed interaction anchors. This factorization mirrors how human experts approach molecular engineering: first, reasoning about critical interactions, then optimizing geometry subject to those constraints. By operationalizing reasoning as explicit residue-level commitments rather than latent textual guidance, Proteo-R1 achieves stable, interpretable, and modular integration of LLM reasoning with state-of-the-art geometric generative models. Code, data, and demos are available at https://smiles724.github.io/r1/.
△ Less
Submitted 10 August, 2026; v1 submitted 1 May, 2026;
originally announced May 2026.
-
Provably Secure Steganography Based on List Decoding
Authors:
Kaiyi Pang,
Minhao Bai
Abstract:
Steganography embeds secret messages in seemingly innocuous carriers for covert communication under surveillance. Current Provably Secure Steganography (PSS) schemes based on language models can guarantee computational indistinguishability between the covertext and stegotext. However, achieving high embedding capacity remains a challenge for existing PSS. The inefficient entropy utilization render…
▽ More
Steganography embeds secret messages in seemingly innocuous carriers for covert communication under surveillance. Current Provably Secure Steganography (PSS) schemes based on language models can guarantee computational indistinguishability between the covertext and stegotext. However, achieving high embedding capacity remains a challenge for existing PSS. The inefficient entropy utilization renders them not well-suited for Large Language Models (LLMs), whose inherent low-entropy tendencies severely constrain feasible embedding capacity. To address this, we propose a provably secure steganography scheme with a theoretically proved high capacity. Our scheme is based on the concept of list decoding: it maintains a set of candidates that contain the correct secret message, instead of directly finding the correct message with more effort. This strategy fully utilizes the information content of the generated text, yielding higher capacity. To ensure the correctness of our scheme, we further introduce a suffix-matching mechanism to distinguish the correct secret message from the candidates. We provide theoretical proofs for both the security and correctness of our scheme, alongside a derivation of its theoretical capacity lower bound. Our approach is plug-and-play, requiring only a direct replacement of the model's standard random sampling module. Experiments on three LLMs and seven PSS baselines demonstrate that our method achieves computational efficiency comparable to prior PSS schemes while delivering a substantial improvement in embedding capacity.
△ Less
Submitted 28 April, 2026; v1 submitted 23 April, 2026;
originally announced April 2026.
-
Symmetry-protected four double-Weyl fermions and their topological phase transitions in nonmagnetic crystals
Authors:
Yun-Yun Bai,
Ke-Xin Pang,
Yan Gao
Abstract:
Realizing Weyl semimetals (WSMs) with the minimal number of Weyl points (WPs) fundamentally simplifies extracting intrinsic topological responses. While a minimum of four conventional ($|C|=1$) WPs in nonmagnetic crystals is well-established, the exact symmetry requirements and material realization for the unique configuration of four unconventional double-Weyl points (DWPs, $|C|=2$) remain unreso…
▽ More
Realizing Weyl semimetals (WSMs) with the minimal number of Weyl points (WPs) fundamentally simplifies extracting intrinsic topological responses. While a minimum of four conventional ($|C|=1$) WPs in nonmagnetic crystals is well-established, the exact symmetry requirements and material realization for the unique configuration of four unconventional double-Weyl points (DWPs, $|C|=2$) remain unresolved. Here, we establish rigorous crystalline symmetry constraints restricting the existence of exactly four symmetry-protected DWPs to merely 28 space groups in both nonmagnetic spinless and spinful systems. Guided by this classification, we identify an $sp$$^2$--$sp$$^3$ hybridized chiral carbon allotrope, THRLN-C$_{32}$, as an ideal candidate hosting precisely this four-DWP configuration near the Fermi level. These $C_4$-protected DWPs project extended or closed-loop Fermi arcs onto the surface Brillouin zone, providing unambiguous spectroscopic signatures. Furthermore, external strain drives profound topological phase transitions encapsulated in a unified evolution landscape: the pristine four-DWP state dissociates into two exotic three-terminal Weyl complexes, degenerates into eight conventional $|C|=1$ WPs, or collapses into a trivial insulator. This work provides a definitive theoretical framework for minimal double-WSMs in nonmagnetic spinful systems and introduces an optimal material platform for investigating strain-tunable topological quantum phenomena.
△ Less
Submitted 8 April, 2026;
originally announced April 2026.
-
Single-pair charge-2 Weyl-Dirac composite semimetals
Authors:
Hui-Jing Zheng,
Ke-Xin Pang,
Yun-Yun Bai,
Yanfeng Ge,
Yan Gao
Abstract:
The Nielsen--Ninomiya theorem requires that the total topological chiral charges in a crystal vanish, a constraint typically satisfied by identical nodes like Weyl--Weyl pairs. Whether a minimal heterogeneous configuration -- comprising a single Weyl point (WP) and a single Dirac point (DP) -- can exist in an electronic system has remained unresolved. Here, by systematically classifying all 1651 m…
▽ More
The Nielsen--Ninomiya theorem requires that the total topological chiral charges in a crystal vanish, a constraint typically satisfied by identical nodes like Weyl--Weyl pairs. Whether a minimal heterogeneous configuration -- comprising a single Weyl point (WP) and a single Dirac point (DP) -- can exist in an electronic system has remained unresolved. Here, by systematically classifying all 1651 magnetic space groups (MSGs), we reveal that only 14 MSGs without spin-orbit coupling (SOC) and 10 MSGs with SOC are compatible with this exotic state. Furthermore, for nonmagnetic crystals, this configuration is uniquely realized in the spinless limit of chiral space groups 92 and 96. Guided by this principle, we predict an ideal realization in chiral three-dimensional boron allotropes (SDHBN-B$_{28}$ enantiomers). First-principles calculations unveil a $|C|=2$ WP at the $Γ$ point and a $|C|=2$ DP at the $A$ point, which constitute the only fermions near the Fermi level within a large $2$ eV energy window. Strikingly, the structural chirality rigidly dictates the sign of the topological charges, yielding two ultralong Fermi arcs spanning the surface Brillouin zone. Our work provides a complete crystallographic classification and a definitive material platform for exploring minimal heterogeneous chiral fermions.
△ Less
Submitted 18 March, 2026;
originally announced March 2026.
-
AtomVLA: Scalable Post-Training for Robotic Manipulation via Predictive Latent World Models
Authors:
Xiaoquan Sun,
Zetian Xu,
Chen Cao,
Zonghe Liu,
Yihan Sun,
Jingrui Pang,
Ruijian Zhang,
Zhen Yang,
Kang Pang,
Dingxin He,
Mingqi Yuan,
Jiayu Chen
Abstract:
Vision-Language-Action (VLA) models demonstrate remarkable potential for generalizable robotic manipulation. The execution of complex multi-step behaviors in VLA models can be improved by robust instruction grounding, a critical component for effective control. However, current paradigms predominantly rely on coarse, high-level task instructions during supervised fine-tuning. This instruction grou…
▽ More
Vision-Language-Action (VLA) models demonstrate remarkable potential for generalizable robotic manipulation. The execution of complex multi-step behaviors in VLA models can be improved by robust instruction grounding, a critical component for effective control. However, current paradigms predominantly rely on coarse, high-level task instructions during supervised fine-tuning. This instruction grounding gap leaves models without explicit intermediate guidance, leading to severe compounding errors in long-horizon tasks. Therefore, bridging this instruction gap and providing scalable post-training for VLA models is urgent. To tackle this problem, we propose \method, the first subtask-aware VLA framework integrated with a scalable offline post-training pipeline. Our framework leverages a large language model to decompose high-level demonstrations into fine-grained atomic subtasks. This approach utilizes a pretrained predictive world model to score candidate action chunks against subtask goals in the latent space, mitigating error accumulation while significantly improving long-horizon robustness. Furthermore, this approach enables highly efficient Group Relative Policy Optimization without the prohibitive expenses associated with online rollouts on physical robots. Extensive simulations validate that our AtomVLA maintains strong robustness under perturbations. When evaluated against fundamental baseline models, it achieves an average success rate of 97.0\% on the LIBERO benchmark and 48.0\% on the LIBERO-PRO benchmark. Finally, experiments conducted in the real world using the Galaxea R1 Lite platform confirm its broad applicability across diverse tasks, especially long-horizon tasks. All datasets, checkpoints, and code will be released to the public domain following the acceptance of this work for future research.
△ Less
Submitted 9 March, 2026;
originally announced March 2026.
-
Symmetry-Protected Minimum of Four Conventional Weyl Points in Nonmagnetic Crystals
Authors:
Ze-Xin Xue,
Ke-Xin Pang,
Yun-Yun Bai,
Yanfeng Ge,
Yong Liu,
Yan Gao
Abstract:
Realizing nonmagnetic Weyl semimetals (WSMs) with the minimal number of conventional Weyl points (WPs) and a clean Fermi surface remains a central challenge. Here, combining symmetry analysis with first-principles calculations, we establish the definitive conditions under which a nonmagnetic crystal can host exactly four conventional ($C = \pm 1$) WPs, identifying 76 space groups in the spinless l…
▽ More
Realizing nonmagnetic Weyl semimetals (WSMs) with the minimal number of conventional Weyl points (WPs) and a clean Fermi surface remains a central challenge. Here, combining symmetry analysis with first-principles calculations, we establish the definitive conditions under which a nonmagnetic crystal can host exactly four conventional ($C = \pm 1$) WPs, identifying 76 space groups in the spinless limit and 83 in the spinful case that allow this minimal configuration. Guided by this framework, we predict two previously unknown boron allotropes, P6-B$_{48}$ and TBIN-B$_{48}$, as ideal WSMs. Both exhibits precisely four isolated WPs near the Fermi level, with exceptionally clean electronic structures. Notably, the WPs in P6-B$_{48}$ are pinned to high-symmetry points, while those in TBIN-B$_{48}$ lie along high-symmetry lines, leading to distinct and experimentally accessible surface states, including single and double Fermi arcs. Our work provides a complete symmetry-based foundation and pristine material platforms for minimal Weyl physics.
△ Less
Submitted 25 February, 2026;
originally announced February 2026.
-
SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic Guidance
Authors:
Minghan Yang,
Lan Yang,
Ke Li,
Honggang Zhang,
Kaiyue Pang,
Yizhe Song
Abstract:
Reconstructing dynamic visual experiences from brain activity provides a compelling avenue for exploring the neural mechanisms of human visual perception. While recent progress in fMRI-based image reconstruction has been notable, extending this success to video reconstruction remains a significant challenge. Current fMRI-to-video reconstruction approaches consistently encounter two major shortcomi…
▽ More
Reconstructing dynamic visual experiences from brain activity provides a compelling avenue for exploring the neural mechanisms of human visual perception. While recent progress in fMRI-based image reconstruction has been notable, extending this success to video reconstruction remains a significant challenge. Current fMRI-to-video reconstruction approaches consistently encounter two major shortcomings: (i) inconsistent visual representations of salient objects across frames, leading to appearance mismatches; (ii) poor temporal coherence, resulting in motion misalignment or abrupt frame transitions. To address these limitations, we introduce SemVideo, a novel fMRI-to-video reconstruction framework guided by hierarchical semantic information. At the core of SemVideo is SemMiner, a hierarchical guidance module that constructs three levels of semantic cues from the original video stimulus: static anchor descriptions, motion-oriented narratives, and holistic summaries. Leveraging this semantic guidance, SemVideo comprises three key components: a Semantic Alignment Decoder that aligns fMRI signals with CLIP-style embeddings derived from SemMiner, a Motion Adaptation Decoder that reconstructs dynamic motion patterns using a novel tripartite attention fusion architecture, and a Conditional Video Render that leverages hierarchical semantic guidance for video reconstruction. Experiments conducted on the CC2017 and HCP datasets demonstrate that SemVideo achieves superior performance in both semantic alignment and temporal consistency, setting a new state-of-the-art in fMRI-to-video reconstruction.
△ Less
Submitted 17 August, 2026; v1 submitted 25 February, 2026;
originally announced February 2026.
-
SynMind: Reducing Semantic Hallucination in fMRI-Based Image Reconstruction
Authors:
Lan Yang,
Minghan Yang,
Ke Li,
Honggang Zhang,
Kaiyue Pang,
Yi-Zhe Song
Abstract:
Recent advances in fMRI-based image reconstruction have achieved remarkable photo-realistic fidelity. Yet, a persistent limitation remains: while reconstructed images often appear naturalistic and holistically similar to the target stimuli, they frequently suffer from severe semantic misalignment -- salient objects are often replaced or hallucinated despite high visual quality. In this work, we ad…
▽ More
Recent advances in fMRI-based image reconstruction have achieved remarkable photo-realistic fidelity. Yet, a persistent limitation remains: while reconstructed images often appear naturalistic and holistically similar to the target stimuli, they frequently suffer from severe semantic misalignment -- salient objects are often replaced or hallucinated despite high visual quality. In this work, we address this limitation by rethinking the role of explicit semantic interpretation in fMRI decoding. We argue that existing methods rely too heavily on entangled visual embeddings which prioritize low-level appearance cues -- such as texture and global gist -- over explicit semantic identity. To overcome this, we parse fMRI signals into rich, sentence-level semantic descriptions that mirror the hierarchical and compositional nature of human visual understanding. We achieve this by leveraging grounded VLMs to generate synthetic, human-like, multi-granularity textual representations that capture object identities and spatial organization. Built upon this foundation, we propose SynMind, a framework that integrates these explicit semantic encodings with visual priors to condition a pretrained diffusion model. Extensive experiments demonstrate that SynMind outperforms state-of-the-art methods across most quantitative metrics. Notably, by offloading semantic reasoning to our text-alignment module, SynMind surpasses competing methods based on SDXL while using the much smaller Stable Diffusion 1.4 and a single consumer GPU. Large-scale human evaluations further confirm that SynMind produces reconstructions more consistent with human visual perception. Neurovisualization analyses reveal that SynMind engages broader and more semantically relevant brain regions, mitigating the over-reliance on high-level visual areas.
△ Less
Submitted 25 January, 2026;
originally announced January 2026.
-
Degenerate Complex Hessian type equations on compact Hermitian manifolds and Applications
Authors:
Kai Pang,
Haoyuan Sun,
Zhiwei Wang,
Xiangyu Zhou
Abstract:
The aim of this paper is to further develop the theory of the degenerate complex Hessian equations on compact Hermitian manifolds. Building upon the generalization of the Bedford-Taylor pluripotential theory to complex Hessian equations by Kołodziej-Nguyen, we solve these equations in the $(ω, m)$-positive cone, $(ω, m)$-big classes and in nef classes, where $ω$ is a reference Hermitian metric. Th…
▽ More
The aim of this paper is to further develop the theory of the degenerate complex Hessian equations on compact Hermitian manifolds. Building upon the generalization of the Bedford-Taylor pluripotential theory to complex Hessian equations by Kołodziej-Nguyen, we solve these equations in the $(ω, m)$-positive cone, $(ω, m)$-big classes and in nef classes, where $ω$ is a reference Hermitian metric. These results are also new in the Kähler case. Moreover, we adapt our techniques to solve complex Monge-Ampère equations in nef classes with mild singularities. The solutions we obtain, in the compact Kähler case, coincide with those for the complex Monge-Ampère equations in the sense of the non-pluripolar product introduced by Boucksom-Eyssidieux-Guedj-Zeriahi. One of the key ingredients in the proof is the adaption, to the Hermitian setting, of a new a priori $L^\infty$-estimate established by Guo-Phong-Tong and Guo-Phong-Tong-Wang.
△ Less
Submitted 7 December, 2025;
originally announced December 2025.
-
RoboTidy : A 3D Gaussian Splatting Household Tidying Benchmark for Embodied Navigation and Action
Authors:
Xiaoquan Sun,
Ruijian Zhang,
Kang Pang,
Bingchen Miao,
Yuxiang Tan,
Zhen Yang,
Ming Li,
Jiayu Chen
Abstract:
Household tidying is an important application area, yet current benchmarks neither model user preferences nor support mobility, and they generalize poorly, making it hard to comprehensively assess integrated language-to-action capabilities. To address this, we propose RoboTidy, a unified benchmark for language-guided household tidying that supports Vision-Language-Action (VLA) and Vision-Language-…
▽ More
Household tidying is an important application area, yet current benchmarks neither model user preferences nor support mobility, and they generalize poorly, making it hard to comprehensively assess integrated language-to-action capabilities. To address this, we propose RoboTidy, a unified benchmark for language-guided household tidying that supports Vision-Language-Action (VLA) and Vision-Language-Navigation (VLN) training and evaluation. RoboTidy provides 500 photorealistic 3D Gaussian Splatting (3DGS) household scenes (covering 500 objects and containers) with collisions, formulates tidying as an "Action (Object, Container)" list, and supplies 6.4k high-quality manipulation demonstration trajectories and 1.5k naviagtion trajectories to support both few-shot and large-scale training. We also deploy RoboTidy in the real world for object tidying, establishing an end-to-end benchmark for household tidying. RoboTidy offers a scalable platform and bridges a key gap in embodied AI by enabling holistic and realistic evaluation of language-guided robots.
△ Less
Submitted 18 November, 2025; v1 submitted 18 November, 2025;
originally announced November 2025.
-
4-Doodle: Text to 3D Sketches that Move!
Authors:
Hao Chen,
Jiaqi Wang,
Yonggang Qi,
Ke Li,
Kaiyue Pang,
Yi-Zhe Song
Abstract:
We present a novel task: text-to-3D sketch animation, which aims to bring freeform sketches to life in dynamic 3D space. Unlike prior works focused on photorealistic content generation, we target sparse, stylized, and view-consistent 3D vector sketches, a lightweight and interpretable medium well-suited for visual communication and prototyping. However, this task is very challenging: (i) no paired…
▽ More
We present a novel task: text-to-3D sketch animation, which aims to bring freeform sketches to life in dynamic 3D space. Unlike prior works focused on photorealistic content generation, we target sparse, stylized, and view-consistent 3D vector sketches, a lightweight and interpretable medium well-suited for visual communication and prototyping. However, this task is very challenging: (i) no paired dataset exists for text and 3D (or 4D) sketches; (ii) sketches require structural abstraction that is difficult to model with conventional 3D representations like NeRFs or point clouds; and (iii) animating such sketches demands temporal coherence and multi-view consistency, which current pipelines do not address. Therefore, we propose 4-Doodle, the first training-free framework for generating dynamic 3D sketches from text. It leverages pretrained image and video diffusion models through a dual-space distillation scheme: one space captures multi-view-consistent geometry using differentiable Bézier curves, while the other encodes motion dynamics via temporally-aware priors. Unlike prior work (e.g., DreamFusion), which optimizes from a single view per step, our multi-view optimization ensures structural alignment and avoids view ambiguity, critical for sparse sketches. Furthermore, we introduce a structure-aware motion module that separates shape-preserving trajectories from deformation-aware changes, enabling expressive motion such as flipping, rotation, and articulated movement. Extensive experiments show that our method produces temporally realistic and structurally stable 3D sketch animations, outperforming existing baselines in both fidelity and controllability. We hope this work serves as a step toward more intuitive and accessible 4D content creation.
△ Less
Submitted 29 October, 2025;
originally announced October 2025.
-
Generation then Reconstruction: Accelerating Masked Autoregressive Models via Two-Stage Sampling
Authors:
Feihong Yan,
Peiru Wang,
Yao Zhu,
Kaiyu Pang,
Qingyan Wei,
Huiqi Li,
Linfeng Zhang
Abstract:
Masked Autoregressive (MAR) models promise better efficiency in visual generation than autoregressive (AR) models for the ability of parallel generation, yet their acceleration potential remains constrained by the modeling complexity of spatially correlated visual tokens in a single step. To address this limitation, we introduce Generation then Reconstruction (GtR), a training-free hierarchical sa…
▽ More
Masked Autoregressive (MAR) models promise better efficiency in visual generation than autoregressive (AR) models for the ability of parallel generation, yet their acceleration potential remains constrained by the modeling complexity of spatially correlated visual tokens in a single step. To address this limitation, we introduce Generation then Reconstruction (GtR), a training-free hierarchical sampling strategy that decomposes generation into two stages: structure generation establishing global semantic scaffolding, followed by detail reconstruction efficiently completing remaining tokens. Assuming that it is more difficult to create an image from scratch than to complement images based on a basic image framework, GtR is designed to achieve acceleration by computing the reconstruction stage quickly while maintaining the generation quality by computing the generation stage slowly. Moreover, observing that tokens on the details of an image often carry more semantic information than tokens in the salient regions, we further propose Frequency-Weighted Token Selection (FTS) to offer more computation budget to tokens on image details, which are localized based on the energy of high frequency information. Extensive experiments on ImageNet class-conditional and text-to-image generation demonstrate 3.72x speedup on MAR-H while maintaining comparable quality (e.g., FID: 1.59, IS: 304.4 vs. original 1.59, 299.1), substantially outperforming existing acceleration methods across various model scales and generation tasks. Our codes will be released in https://github.com/feihongyan1/GtR.
△ Less
Submitted 26 January, 2026; v1 submitted 20 October, 2025;
originally announced October 2025.
-
OCELOT 2023: Cell Detection from Cell-Tissue Interaction Challenge
Authors:
JaeWoong Shin,
Jeongun Ryu,
Aaron Valero Puche,
Jinhee Lee,
Biagio Brattoli,
Wonkyung Jung,
Soo Ick Cho,
Kyunghyun Paeng,
Chan-Young Ock,
Donggeun Yoo,
Zhaoyang Li,
Wangkai Li,
Huayu Mai,
Joshua Millward,
Zhen He,
Aiden Nibali,
Lydia Anette Schoenpflug,
Viktor Hendrik Koelzer,
Xu Shuoyu,
Ji Zheng,
Hu Bin,
Yu-Wen Lo,
Ching-Hui Yang,
Sérgio Pereira
Abstract:
Pathologists routinely alternate between different magnifications when examining Whole-Slide Images, allowing them to evaluate both broad tissue morphology and intricate cellular details to form comprehensive diagnoses. However, existing deep learning-based cell detection models struggle to replicate these behaviors and learn the interdependent semantics between structures at different magnificati…
▽ More
Pathologists routinely alternate between different magnifications when examining Whole-Slide Images, allowing them to evaluate both broad tissue morphology and intricate cellular details to form comprehensive diagnoses. However, existing deep learning-based cell detection models struggle to replicate these behaviors and learn the interdependent semantics between structures at different magnifications. A key barrier in the field is the lack of datasets with multi-scale overlapping cell and tissue annotations. The OCELOT 2023 challenge was initiated to gather insights from the community to validate the hypothesis that understanding cell and tissue (cell-tissue) interactions is crucial for achieving human-level performance, and to accelerate the research in this field. The challenge dataset includes overlapping cell detection and tissue segmentation annotations from six organs, comprising 673 pairs sourced from 306 The Cancer Genome Atlas (TCGA) Whole-Slide Images with hematoxylin and eosin staining, divided into training, validation, and test subsets. Participants presented models that significantly enhanced the understanding of cell-tissue relationships. Top entries achieved up to a 7.99 increase in F1-score on the test set compared to the baseline cell-only model that did not incorporate cell-tissue relationships. This is a substantial improvement in performance over traditional cell-only detection methods, demonstrating the need for incorporating multi-scale semantics into the models. This paper provides a comparative analysis of the methods used by participants, highlighting innovative strategies implemented in the OCELOT 2023 challenge.
△ Less
Submitted 11 September, 2025;
originally announced September 2025.
-
Mini-Omni-Reasoner: Token-Level Thinking-in-Speaking in Large Speech Models
Authors:
Zhifei Xie,
Ziyang Ma,
Zihang Liu,
Kaiyu Pang,
Hongyu Li,
Jialin Zhang,
Yue Liao,
Deheng Ye,
Chunyan Miao,
Shuicheng Yan
Abstract:
Reasoning is essential for effective communication and decision-making. While recent advances in LLMs and MLLMs have shown that incorporating explicit reasoning significantly improves understanding and generalization, reasoning in LSMs remains in a nascent stage. Early efforts attempt to transfer the "Thinking-before-Speaking" paradigm from textual models to speech. However, this sequential formul…
▽ More
Reasoning is essential for effective communication and decision-making. While recent advances in LLMs and MLLMs have shown that incorporating explicit reasoning significantly improves understanding and generalization, reasoning in LSMs remains in a nascent stage. Early efforts attempt to transfer the "Thinking-before-Speaking" paradigm from textual models to speech. However, this sequential formulation introduces notable latency, as spoken responses are delayed until reasoning is fully completed, impairing real-time interaction and communication efficiency. To address this, we propose Mini-Omni-Reasoner, a framework that enables reasoning within speech via a novel "Thinking-in-Speaking" formulation. Rather than completing reasoning before producing any verbal output, Mini-Omni-Reasoner interleaves silent reasoning tokens with spoken response tokens at the token level. This design allows continuous speech generation while embedding structured internal reasoning, leveraging the model's high-frequency token processing capability. Although interleaved, local semantic alignment is enforced to ensure that each response token is informed by its preceding reasoning. To support this framework, we introduce Spoken-Math-Problems-3M, a large-scale dataset tailored for interleaved reasoning and response. The dataset ensures that verbal tokens consistently follow relevant reasoning content, enabling accurate and efficient learning of speech-coupled reasoning. Built on a hierarchical Thinker-Talker architecture, Mini-Omni-Reasoner delivers fluent yet logically grounded spoken responses, maintaining both naturalness and precision. On the Spoken-MQA benchmark, it achieves a +19.1% gain in arithmetic reasoning and +6.4% in contextual understanding, with shorter outputs and zero decoding latency.
△ Less
Submitted 20 September, 2025; v1 submitted 18 August, 2025;
originally announced August 2025.
-
SCORPION: Addressing Scanner-Induced Variability in Histopathology
Authors:
Jeongun Ryu,
Heon Song,
Seungeun Lee,
Soo Ick Cho,
Jiwon Shin,
Kyunghyun Paeng,
Sérgio Pereira
Abstract:
Ensuring reliable model performance across diverse domains is a critical challenge in computational pathology. A particular source of variability in Whole-Slide Images is introduced by differences in digital scanners, thus calling for better scanner generalization. This is critical for the real-world adoption of computational pathology, where the scanning devices may differ per institution or hosp…
▽ More
Ensuring reliable model performance across diverse domains is a critical challenge in computational pathology. A particular source of variability in Whole-Slide Images is introduced by differences in digital scanners, thus calling for better scanner generalization. This is critical for the real-world adoption of computational pathology, where the scanning devices may differ per institution or hospital, and the model should not be dependent on scanner-induced details, which can ultimately affect the patient's diagnosis and treatment planning. However, past efforts have primarily focused on standard domain generalization settings, evaluating on unseen scanners during training, without directly evaluating consistency across scanners for the same tissue. To overcome this limitation, we introduce SCORPION, a new dataset explicitly designed to evaluate model reliability under scanner variability. SCORPION includes 480 tissue samples, each scanned with 5 scanners, yielding 2,400 spatially aligned patches. This scanner-paired design allows for the isolation of scanner-induced variability, enabling a rigorous evaluation of model consistency while controlling for differences in tissue composition. Furthermore, we propose SimCons, a flexible framework that combines augmentation-based domain generalization techniques with a consistency loss to explicitly address scanner generalization. We empirically show that SimCons improves model consistency on varying scanners without compromising task-specific performance. By releasing the SCORPION dataset and proposing SimCons, we provide the research community with a crucial resource for evaluating and improving model consistency across diverse scanners, setting a new standard for reliability testing.
△ Less
Submitted 17 September, 2025; v1 submitted 28 July, 2025;
originally announced July 2025.
-
Annotation-Free Human Sketch Quality Assessment
Authors:
Lan Yang,
Kaiyue Pang,
Honggang Zhang,
Yi-Zhe Song
Abstract:
As lovely as bunnies are, your sketched version would probably not do them justice (Fig.~\ref{fig:intro}). This paper recognises this very problem and studies sketch quality assessment for the first time -- letting you find these badly drawn ones. Our key discovery lies in exploiting the magnitude ($L_2$ norm) of a sketch feature as a quantitative quality metric. We propose Geometry-Aware Classifi…
▽ More
As lovely as bunnies are, your sketched version would probably not do them justice (Fig.~\ref{fig:intro}). This paper recognises this very problem and studies sketch quality assessment for the first time -- letting you find these badly drawn ones. Our key discovery lies in exploiting the magnitude ($L_2$ norm) of a sketch feature as a quantitative quality metric. We propose Geometry-Aware Classification Layer (GACL), a generic method that makes feature-magnitude-as-quality-metric possible and importantly does it without the need for specific quality annotations from humans. GACL sees feature magnitude and recognisability learning as a dual task, which can be simultaneously optimised under a neat cross-entropy classification loss with theoretic guarantee. This gives GACL a nice geometric interpretation (the better the quality, the easier the recognition), and makes it agnostic to both network architecture changes and the underlying sketch representation. Through a large scale human study of 160,000 \doublecheck{trials}, we confirm the agreement between our GACL-induced metric and human quality perception. We further demonstrate how such a quality assessment capability can for the first time enable three practical sketch applications. Interestingly, we show GACL not only works on abstract visual representations such as sketch but also extends well to natural images on the problem of image quality assessment (IQA). Last but not least, we spell out the general properties of GACL as general-purpose data re-weighting strategy and demonstrate its applications in vertical problems such as noisy label cleansing. Code will be made publicly available at github.com/yanglan0225/SketchX-Quantifying-Sketch-Quality.
△ Less
Submitted 28 July, 2025;
originally announced July 2025.
-
Mic-hackathon 2024: Hackathon on Machine Learning for Electron and Scanning Probe Microscopy
Authors:
Utkarsh Pratiush,
Austin Houston,
Kamyar Barakati,
Aditya Raghavan,
Dasol Yoon,
Harikrishnan KP,
Zhaslan Baraissov,
Desheng Ma,
Samuel S. Welborn,
Mikolaj Jakowski,
Shawn-Patrick Barhorst,
Alexander J. Pattison,
Panayotis Manganaris,
Sita Sirisha Madugula,
Sai Venkata Gayathri Ayyagari,
Vishal Kennedy,
Ralph Bulanadi,
Michelle Wang,
Kieran J. Pang,
Ian Addison-Smith,
Willy Menacho,
Horacio V. Guzman,
Alexander Kiefer,
Nicholas Furth,
Nikola L. Kolev
, et al. (48 additional authors not shown)
Abstract:
Microscopy is a primary source of information on materials structure and functionality at nanometer and atomic scales. The data generated is often well-structured, enriched with metadata and sample histories, though not always consistent in detail or format. The adoption of Data Management Plans (DMPs) by major funding agencies promotes preservation and access. However, deriving insights remains d…
▽ More
Microscopy is a primary source of information on materials structure and functionality at nanometer and atomic scales. The data generated is often well-structured, enriched with metadata and sample histories, though not always consistent in detail or format. The adoption of Data Management Plans (DMPs) by major funding agencies promotes preservation and access. However, deriving insights remains difficult due to the lack of standardized code ecosystems, benchmarks, and integration strategies. As a result, data usage is inefficient and analysis time is extensive. In addition to post-acquisition analysis, new APIs from major microscope manufacturers enable real-time, ML-based analytics for automated decision-making and ML-agent-controlled microscope operation. Yet, a gap remains between the ML and microscopy communities, limiting the impact of these methods on physics, materials discovery, and optimization. Hackathons help bridge this divide by fostering collaboration between ML researchers and microscopy experts. They encourage the development of novel solutions that apply ML to microscopy, while preparing a future workforce for instrumentation, materials science, and applied ML. This hackathon produced benchmark datasets and digital twins of microscopes to support community growth and standardized workflows. All related code is available at GitHub: https://github.com/KalininGroup/Mic-hackathon-2024-codes-publication/tree/1.0.0.1
△ Less
Submitted 27 June, 2025; v1 submitted 9 June, 2025;
originally announced June 2025.
-
FashionM3: Multimodal, Multitask, and Multiround Fashion Assistant based on Unified Vision-Language Model
Authors:
Kaicheng Pang,
Xingxing Zou,
Waikeung Wong
Abstract:
Fashion styling and personalized recommendations are pivotal in modern retail, contributing substantial economic value in the fashion industry. With the advent of vision-language models (VLM), new opportunities have emerged to enhance retailing through natural language and visual interactions. This work proposes FashionM3, a multimodal, multitask, and multiround fashion assistant, built upon a VLM…
▽ More
Fashion styling and personalized recommendations are pivotal in modern retail, contributing substantial economic value in the fashion industry. With the advent of vision-language models (VLM), new opportunities have emerged to enhance retailing through natural language and visual interactions. This work proposes FashionM3, a multimodal, multitask, and multiround fashion assistant, built upon a VLM fine-tuned for fashion-specific tasks. It helps users discover satisfying outfits by offering multiple capabilities including personalized recommendation, alternative suggestion, product image generation, and virtual try-on simulation. Fine-tuned on the novel FashionRec dataset, comprising 331,124 multimodal dialogue samples across basic, personalized, and alternative recommendation tasks, FashionM3 delivers contextually personalized suggestions with iterative refinement through multiround interactions. Quantitative and qualitative evaluations, alongside user studies, demonstrate FashionM3's superior performance in recommendation effectiveness and practical value as a fashion assistant.
△ Less
Submitted 23 April, 2025;
originally announced April 2025.
-
Provable Secure Steganography Based on Adaptive Dynamic Sampling
Authors:
Kaiyi Pang,
Minhao Bai
Abstract:
The security of private communication is increasingly at risk due to widespread surveillance. Steganography, a technique for embedding secret messages within innocuous carriers, enables covert communication over monitored channels. Provably Secure Steganography (PSS), which ensures computational indistinguishability between the normal model output and steganography output, is the state-of-the-art…
▽ More
The security of private communication is increasingly at risk due to widespread surveillance. Steganography, a technique for embedding secret messages within innocuous carriers, enables covert communication over monitored channels. Provably Secure Steganography (PSS), which ensures computational indistinguishability between the normal model output and steganography output, is the state-of-the-art in this field. However, current PSS methods often require obtaining the explicit distributions of the model. In this paper, we propose a provably secure steganography scheme that only requires a model API that accepts a seed as input. Our core mechanism involves sampling a candidate set of tokens and constructing a map from possible message bit strings to these tokens. The output token is selected by applying this mapping to the real secret message, which provably preserves the original model's distribution. To ensure correct decoding, we address collision cases, where multiple candidate messages map to the same token, by maintaining and strategically expanding a dynamic collision set within a bounded size range. Extensive evaluations of three real-world datasets and three large language models demonstrate that our sampling-based method is comparable with existing PSS methods in efficiency and capacity.
△ Less
Submitted 12 February, 2026; v1 submitted 16 April, 2025;
originally announced April 2025.
-
A cryogenic test-mass suspension with flexures operating in compression for third-generation gravitational-wave detectors
Authors:
Fabián E. Peña Arellano,
Nelson L. Leon,
Leonardo González López,
Riccardo DeSalvo,
Harry Themann,
Esra Zerina Appavuravther,
Guerino Avallone,
Francesca Badaracco,
Mark A. Barton,
Alessandro Bertolini,
Christian Chavez,
Andy Damas,
Richard Damas,
Britney Gallego,
Eric Hennes,
Gerardo Iannone,
Seth Linker,
Marina Mondin,
Claudia Moreno,
Kevin Pang,
Stefano Selleri,
Mynor Soto,
Flavio Travasso,
Joris Van-Heijningen,
Fernando Velez
, et al. (1 additional authors not shown)
Abstract:
This paper presents an analysis of the conceptual design of a novel silicon suspension for the cryogenic test-mass mirrors of the low-frequency detector of the Einstein Telescope gravitational-wave observatory. In traditional suspensions, tensional stress is a severe limitation for achieving low thermal noise, safer mechanical margins and high thermal conductance simultaneously. In order to keep t…
▽ More
This paper presents an analysis of the conceptual design of a novel silicon suspension for the cryogenic test-mass mirrors of the low-frequency detector of the Einstein Telescope gravitational-wave observatory. In traditional suspensions, tensional stress is a severe limitation for achieving low thermal noise, safer mechanical margins and high thermal conductance simultaneously. In order to keep the tensional stress sufficiently low, we propose the use of rigid beams with large cross sections, combined with short flexures under compressional load. This configuration takes advantage of the many times higher strength of silicon in compression to respect to its strength in tension. The flexures are mechanically robust and at the same time soft in the working direction, thus producing low suspension thermal noise and, by being short, provide high thermal conductance for cryogenic cooling. The rigid beams, located between the test mass and an intermediate mass, allow the elimination of the recoil mass used conventionally for applying control forces for interferometer lock, and the use of optical anti-springs to reduce the pendulum resonant frequency to further improve the vibration isolation of the test mass. The configuration has the capability to reach a lower mirror operational temperature, which is expected to produce a substantial reduction of the thermal noise in the mirrors of the interferometer.
△ Less
Submitted 24 March, 2025;
originally announced March 2025.
-
Applications of Large Models in Medicine
Authors:
YunHe Su,
Zhengyang Lu,
Junhui Liu,
Ke Pang,
Haoran Dai,
Sa Liu,
Yuxin Jia,
Lujia Ge,
Jing-min Yang
Abstract:
This paper explores the advancements and applications of large-scale models in the medical field, with a particular focus on Medical Large Models (MedLMs). These models, encompassing Large Language Models (LLMs), Vision Models, 3D Large Models, and Multimodal Models, are revolutionizing healthcare by enhancing disease prediction, diagnostic assistance, personalized treatment planning, and drug dis…
▽ More
This paper explores the advancements and applications of large-scale models in the medical field, with a particular focus on Medical Large Models (MedLMs). These models, encompassing Large Language Models (LLMs), Vision Models, 3D Large Models, and Multimodal Models, are revolutionizing healthcare by enhancing disease prediction, diagnostic assistance, personalized treatment planning, and drug discovery. The integration of graph neural networks in medical knowledge graphs and drug discovery highlights the potential of Large Graph Models (LGMs) in understanding complex biomedical relationships. The study also emphasizes the transformative role of Vision-Language Models (VLMs) and 3D Large Models in medical image analysis, anatomical modeling, and prosthetic design. Despite the challenges, these technologies are setting new benchmarks in medical innovation, improving diagnostic accuracy, and paving the way for personalized healthcare solutions. This paper aims to provide a comprehensive overview of the current state and future directions of large models in medicine, underscoring their significance in advancing global health.
△ Less
Submitted 7 October, 2025; v1 submitted 24 February, 2025;
originally announced February 2025.
-
WMamba: Wavelet-based Mamba for Face Forgery Detection
Authors:
Siran Peng,
Tianshuo Zhang,
Li Gao,
Xiangyu Zhu,
Haoyuan Zhang,
Kai Pang,
Zhen Lei
Abstract:
The rapid evolution of deepfake generation technologies necessitates the development of robust face forgery detection algorithms. Recent studies have demonstrated that wavelet analysis can enhance the generalization abilities of forgery detectors. Wavelets effectively capture key facial contours, often slender, fine-grained, and globally distributed, that may conceal subtle forgery artifacts imper…
▽ More
The rapid evolution of deepfake generation technologies necessitates the development of robust face forgery detection algorithms. Recent studies have demonstrated that wavelet analysis can enhance the generalization abilities of forgery detectors. Wavelets effectively capture key facial contours, often slender, fine-grained, and globally distributed, that may conceal subtle forgery artifacts imperceptible in the spatial domain. However, current wavelet-based approaches fail to fully exploit the distinctive properties of wavelet data, resulting in sub-optimal feature extraction and limited performance gains. To address this challenge, we introduce WMamba, a novel wavelet-based feature extractor built upon the Mamba architecture. WMamba maximizes the utility of wavelet information through two key innovations. First, we propose Dynamic Contour Convolution (DCConv), which employs specially crafted deformable kernels to adaptively model slender facial contours. Second, by leveraging the Mamba architecture, our method captures long-range spatial relationships with linear complexity. This efficiency allows for the extraction of fine-grained, globally distributed forgery artifacts from small image patches. Extensive experiments show that WMamba achieves state-of-the-art (SOTA) performance, highlighting its effectiveness in face forgery detection.
△ Less
Submitted 21 October, 2025; v1 submitted 16 January, 2025;
originally announced January 2025.
-
Shifting-Merging: Secure, High-Capacity and Efficient Steganography via Large Language Models
Authors:
Minhao Bai,
Jinshuai Yang,
Kaiyi Pang,
Yongfeng Huang,
Yue Gao
Abstract:
In the face of escalating surveillance and censorship within the cyberspace, the sanctity of personal privacy has come under siege, necessitating the development of steganography, which offers a way to securely hide messages within innocent-looking texts. Previous methods alternate the texts to hide private massages, which is not secure. Large Language Models (LLMs) provide high-quality and explic…
▽ More
In the face of escalating surveillance and censorship within the cyberspace, the sanctity of personal privacy has come under siege, necessitating the development of steganography, which offers a way to securely hide messages within innocent-looking texts. Previous methods alternate the texts to hide private massages, which is not secure. Large Language Models (LLMs) provide high-quality and explicit distribution, which is an available mathematical tool for secure steganography methods. However, existing attempts fail to achieve high capacity, time efficiency and correctness simultaneously, and their strongly coupling designs leave little room for refining them to achieve better performance. To provide a secure, high-capacity and efficient steganography method, we introduce ShiMer. Specifically, ShiMer pseudorandomly shifts the probability interval of the LLM's distribution to obtain a private distribution, and samples a token according to the private bits. ShiMer produced steganographic texts are indistinguishable in quality from the normal texts directly generated by the language model. To further enhance the capacity of ShiMer, we design a reordering algorithm to minimize the occurrence of interval splitting during decoding phase. Experimental results indicate that our method achieves the highest capacity and efficiency among existing secure steganography techniques.
△ Less
Submitted 1 January, 2025;
originally announced January 2025.
-
A Plug-and-Play Method for Improving Imperceptibility and Capacity in Practical Generative Text Steganography
Authors:
Kaiyi Pang
Abstract:
Linguistic steganography embeds secret information into seemingly innocuous text to safeguard privacy under surveillance. Generative linguistic steganography leverages the probability distributions of language models (LMs) and applies steganographic algorithms during generation, and has attracted increasing attention with the rise of large language models (LLMs). To strengthen security, prior work…
▽ More
Linguistic steganography embeds secret information into seemingly innocuous text to safeguard privacy under surveillance. Generative linguistic steganography leverages the probability distributions of language models (LMs) and applies steganographic algorithms during generation, and has attracted increasing attention with the rise of large language models (LLMs). To strengthen security, prior work has focused on distribution-preserving steganographic algorithms that minimize the gap between stego sampling and random sampling from the model. However, their reliance on model distributions, which often deviate from real-world cover texts, leads to limited imperceptibility when facing steganalysis detectors in practical settings. Moreover, LLM distributions tend to be more deterministic, reducing entropy and thus lowering embedding capacity. In this paper, we propose a plug-and-play method that reconstructs the distributions of language models used for generative linguistic steganography. FreStega dynamically adjusts token probabilities from the language model at each step of autoregressive stego text generation, leveraging both sequential and spatial dimensions. Extensive experiments on four LLMs, three benchmark datasets, and four distribution-preserving steganographic baselines demonstrate that, by reforming the distribution, FreStega improves the imperceptibility of stego text in realistic scenarios and increases steganographic capacity by 15.41\%, without degrading the quality of the generated stegotext.
△ Less
Submitted 26 June, 2026; v1 submitted 27 December, 2024;
originally announced December 2024.
-
Spoof Trace Discovery for Deep Learning Based Explainable Face Anti-Spoofing
Authors:
Haoyuan Zhang,
Xiangyu Zhu,
Li Gao,
Jiawei Pan,
Kai Pang,
Guoying Zhao,
Zhen Lei
Abstract:
With the rapid growth usage of face recognition in people's daily life, face anti-spoofing becomes increasingly important to avoid malicious attacks. Recent face anti-spoofing models can reach a high classification accuracy on multiple datasets but these models can only tell people "this face is fake" while lacking the explanation to answer "why it is fake". Such a system undermines trustworthines…
▽ More
With the rapid growth usage of face recognition in people's daily life, face anti-spoofing becomes increasingly important to avoid malicious attacks. Recent face anti-spoofing models can reach a high classification accuracy on multiple datasets but these models can only tell people "this face is fake" while lacking the explanation to answer "why it is fake". Such a system undermines trustworthiness and causes user confusion, as it denies their requests without providing any explanations. In this paper, we incorporate XAI into face anti-spoofing and propose a new problem termed X-FAS (eXplainable Face Anti-Spoofing) empowering face anti-spoofing models to provide an explanation. We propose SPTD (SPoof Trace Discovery), an X-FAS method which can discover spoof concepts and provide reliable explanations on the basis of discovered concepts. To evaluate the quality of X-FAS methods, we propose an X-FAS benchmark with annotated spoof traces by experts. We analyze SPTD explanations on face anti-spoofing dataset and compare SPTD quantitatively and qualitatively with previous XAI methods on proposed X-FAS benchmark. Experimental results demonstrate SPTD's ability to generate reliable explanations.
△ Less
Submitted 5 September, 2025; v1 submitted 23 December, 2024;
originally announced December 2024.
-
VersaGen: Unleashing Versatile Visual Control for Text-to-Image Synthesis
Authors:
Zhipeng Chen,
Lan Yang,
Yonggang Qi,
Honggang Zhang,
Kaiyue Pang,
Ke Li,
Yi-Zhe Song
Abstract:
Despite the rapid advancements in text-to-image (T2I) synthesis, enabling precise visual control remains a significant challenge. Existing works attempted to incorporate multi-facet controls (text and sketch), aiming to enhance the creative control over generated images. However, our pilot study reveals that the expressive power of humans far surpasses the capabilities of current methods. Users de…
▽ More
Despite the rapid advancements in text-to-image (T2I) synthesis, enabling precise visual control remains a significant challenge. Existing works attempted to incorporate multi-facet controls (text and sketch), aiming to enhance the creative control over generated images. However, our pilot study reveals that the expressive power of humans far surpasses the capabilities of current methods. Users desire a more versatile approach that can accommodate their diverse creative intents, ranging from controlling individual subjects to manipulating the entire scene composition. We present VersaGen, a generative AI agent that enables versatile visual control in T2I synthesis. VersaGen admits four types of visual controls: i) single visual subject; ii) multiple visual subjects; iii) scene background; iv) any combination of the three above or merely no control at all. We train an adaptor upon a frozen T2I model to accommodate the visual information into the text-dominated diffusion process. We introduce three optimization strategies during the inference phase of VersaGen to improve generation results and enhance user experience. Comprehensive experiments on COCO and Sketchy validate the effectiveness and flexibility of VersaGen, as evidenced by both qualitative and quantitative results.
△ Less
Submitted 27 December, 2024; v1 submitted 16 December, 2024;
originally announced December 2024.
-
Weak convergence of complex Monge-Ampère operators on compact Hermitian manifolds
Authors:
Kai Pang,
Haoyuan Sun,
Zhiwei Wang
Abstract:
Let $(X,ω)$ be a compact Hermitian manifold and let $\{β\}\in H^{1,1}(X,\mathbb R)$ be a real $(1,1)$-class with a smooth representative $β$, such that $\int_Xβ^n>0$. Assume that there is a bounded $β$-plurisubharmonic function $ρ$ on $X$. First, we provide a criterion for the weak convergence of non-pluripolar complex Monge-Ampère measures associated to a sequence of $β$-plurisubharmonic function…
▽ More
Let $(X,ω)$ be a compact Hermitian manifold and let $\{β\}\in H^{1,1}(X,\mathbb R)$ be a real $(1,1)$-class with a smooth representative $β$, such that $\int_Xβ^n>0$. Assume that there is a bounded $β$-plurisubharmonic function $ρ$ on $X$. First, we provide a criterion for the weak convergence of non-pluripolar complex Monge-Ampère measures associated to a sequence of $β$-plurisubharmonic functions. Second, this criterion is utilized to solve a degenerate complex Monge-Ampère equation with an $L^1$-density. Finally, an $L^\infty$-estimate of the solution to the complex Monge-Ampère equation for a finite positive Radon measure is given.
△ Less
Submitted 16 December, 2024;
originally announced December 2024.
-
Semantic Steganography: A Framework for Robust and High-Capacity Information Hiding using Large Language Models
Authors:
Minhao Bai,
Jinshuai Yang,
Kaiyi Pang,
Yongfeng Huang,
Yue Gao
Abstract:
In the era of Large Language Models (LLMs), generative linguistic steganography has become a prevalent technique for hiding information within model-generated texts. However, traditional steganography methods struggle to effectively align steganographic texts with original model-generated texts due to the lower entropy of the predicted probability distribution of LLMs. This results in a decrease i…
▽ More
In the era of Large Language Models (LLMs), generative linguistic steganography has become a prevalent technique for hiding information within model-generated texts. However, traditional steganography methods struggle to effectively align steganographic texts with original model-generated texts due to the lower entropy of the predicted probability distribution of LLMs. This results in a decrease in embedding capacity and poses challenges for decoding stegos in real-world communication channels. To address these challenges, we propose a semantic steganography framework based on LLMs, which construct a semantic space and map secret messages onto this space using ontology-entity trees. This framework offers robustness and reliability for transmission in complex channels, as well as resistance to text rendering and word blocking. Additionally, the stegos generated by our framework are indistinguishable from the covers and achieve a higher embedding capacity compared to state-of-the-art steganography methods, while producing higher quality stegos.
△ Less
Submitted 14 December, 2024;
originally announced December 2024.
-
Applications and Novel Regularization of the Thin-Film Equation
Authors:
Khang Ee Pang
Abstract:
The classical no-slip boundary condition of the Navier-Stokes equations fails to describe the spreading motion of a droplet on a substrate due to the missing small-scale physics near the contact line. In this thesis, we introduce a novel regularization of the thin-film equation to model droplet spreading. The solution of the regularized thin-film equation -- the Geometric Thin-Film Equation is stu…
▽ More
The classical no-slip boundary condition of the Navier-Stokes equations fails to describe the spreading motion of a droplet on a substrate due to the missing small-scale physics near the contact line. In this thesis, we introduce a novel regularization of the thin-film equation to model droplet spreading. The solution of the regularized thin-film equation -- the Geometric Thin-Film Equation is studied and characterized. Two robust numerical solvers are discussed, notably, a fast and mesh-free numerical scheme for simulating thin-film flows in two and three spatial dimensions. Moreover, we prove the regularity and convergence of the numerical solutions. The existence and uniqueness of the solution of the Geometric Thin-Film Equation with respect to a wide range of measure-valued initial conditions are also discussed.
△ Less
Submitted 25 September, 2024;
originally announced September 2024.
-
Generalizing AI-driven Assessment of Immunohistochemistry across Immunostains and Cancer Types: A Universal Immunohistochemistry Analyzer
Authors:
Biagio Brattoli,
Mohammad Mostafavi,
Taebum Lee,
Wonkyung Jung,
Jeongun Ryu,
Seonwook Park,
Jongchan Park,
Sergio Pereira,
Seunghwan Shin,
Sangjoon Choi,
Hyojin Kim,
Donggeun Yoo,
Siraj M. Ali,
Kyunghyun Paeng,
Chan-Young Ock,
Soo Ick Cho,
Seokhwi Kim
Abstract:
Despite advancements in methodologies, immunohistochemistry (IHC) remains the most utilized ancillary test for histopathologic and companion diagnostics in targeted therapies. However, objective IHC assessment poses challenges. Artificial intelligence (AI) has emerged as a potential solution, yet its development requires extensive training for each cancer and IHC type, limiting versatility. We dev…
▽ More
Despite advancements in methodologies, immunohistochemistry (IHC) remains the most utilized ancillary test for histopathologic and companion diagnostics in targeted therapies. However, objective IHC assessment poses challenges. Artificial intelligence (AI) has emerged as a potential solution, yet its development requires extensive training for each cancer and IHC type, limiting versatility. We developed a Universal IHC (UIHC) analyzer, an AI model for interpreting IHC images regardless of tumor or IHC types, using training datasets from various cancers stained for PD-L1 and/or HER2. This multi-cohort trained model outperforms conventional single-cohort models in interpreting unseen IHCs (Kappa score 0.578 vs. up to 0.509) and consistently shows superior performance across different positive staining cutoff values. Qualitative analysis reveals that UIHC effectively clusters patches based on expression levels. The UIHC model also quantitatively assesses c-MET expression with MET mutations, representing a significant advancement in AI application in the era of personalized medicine and accumulating novel biomarkers.
△ Less
Submitted 30 July, 2024;
originally announced July 2024.
-
Provably Robust and Secure Steganography in Asymmetric Resource Scenario
Authors:
Minhao Bai,
Jinshuai Yang,
Kaiyi Pang,
Xin Xu,
Zhen Yang,
Yongfeng Huang
Abstract:
To circumvent the unbridled and ever-encroaching surveillance and censorship in cyberspace, steganography has garnered attention for its ability to hide private information in innocent-looking carriers. Current provably secure steganography approaches require a pair of encoder and decoder to hide and extract private messages, both of which must run the same model with the same input to obtain iden…
▽ More
To circumvent the unbridled and ever-encroaching surveillance and censorship in cyberspace, steganography has garnered attention for its ability to hide private information in innocent-looking carriers. Current provably secure steganography approaches require a pair of encoder and decoder to hide and extract private messages, both of which must run the same model with the same input to obtain identical distributions. These requirements pose significant challenges to the practical implementation of steganography, including limited access to powerful hardware and the intolerance of any changes to the shared input. To relax the limitation of hardware and solve the challenge of vulnerable shared input, a novel and practically significant scenario with asymmetric resource should be considered, where only the encoder is high-resource and accessible to powerful models while the decoder can only read the steganographic carriers without any other model's input. This paper proposes a novel provably robust and secure steganography framework for the asymmetric resource setting. Specifically, the encoder uses various permutations of distribution to hide secret bits, while the decoder relies on a sampling function to extract the hidden bits by guessing the permutation used. Further, the sampling function only takes the steganographic carrier as input, which makes the decoder independent of model's input and model itself. A comprehensive assessment of applying our framework to generative models substantiates its effectiveness. Our implementation demonstrates robustness when transmitting over binary symmetric channels with errors.
△ Less
Submitted 24 November, 2024; v1 submitted 18 July, 2024;
originally announced July 2024.
-
Cross-Slice Attention and Evidential Critical Loss for Uncertainty-Aware Prostate Cancer Detection
Authors:
Alex Ling Yu Hung,
Haoxin Zheng,
Kai Zhao,
Kaifeng Pang,
Demetri Terzopoulos,
Kyunghyun Sung
Abstract:
Current deep learning-based models typically analyze medical images in either 2D or 3D albeit disregarding volumetric information or suffering sub-optimal performance due to the anisotropic resolution of MR data. Furthermore, providing an accurate uncertainty estimation is beneficial to clinicians, as it indicates how confident a model is about its prediction. We propose a novel 2.5D cross-slice a…
▽ More
Current deep learning-based models typically analyze medical images in either 2D or 3D albeit disregarding volumetric information or suffering sub-optimal performance due to the anisotropic resolution of MR data. Furthermore, providing an accurate uncertainty estimation is beneficial to clinicians, as it indicates how confident a model is about its prediction. We propose a novel 2.5D cross-slice attention model that utilizes both global and local information, along with an evidential critical loss, to perform evidential deep learning for the detection in MR images of prostate cancer, one of the most common cancers and a leading cause of cancer-related death in men. We perform extensive experiments with our model on two different datasets and achieve state-of-the-art performance in prostate cancer detection along with improved epistemic uncertainty estimation. The implementation of the model is available at https://github.com/aL3x-O-o-Hung/GLCSA_ECLoss.
△ Less
Submitted 1 July, 2024;
originally announced July 2024.
-
Towards Next-Generation Steganalysis: LLMs Unleash the Power of Detecting Steganography
Authors:
Minhao Bai. Jinshuai Yang,
Kaiyi Pang,
Huili Wang,
Yongfeng Huang
Abstract:
Linguistic steganography provides convenient implementation to hide messages, particularly with the emergence of AI generation technology. The potential abuse of this technology raises security concerns within societies, calling for powerful linguistic steganalysis to detect carrier containing steganographic messages. Existing methods are limited to finding distribution differences between stegano…
▽ More
Linguistic steganography provides convenient implementation to hide messages, particularly with the emergence of AI generation technology. The potential abuse of this technology raises security concerns within societies, calling for powerful linguistic steganalysis to detect carrier containing steganographic messages. Existing methods are limited to finding distribution differences between steganographic texts and normal texts from the aspect of symbolic statistics. However, the distribution differences of both kinds of texts are hard to build precisely, which heavily hurts the detection ability of the existing methods in realistic scenarios. To seek a feasible way to construct practical steganalysis in real world, this paper propose to employ human-like text processing abilities of large language models (LLMs) to realize the difference from the aspect of human perception, addition to traditional statistic aspect. Specifically, we systematically investigate the performance of LLMs in this task by modeling it as a generative paradigm, instead of traditional classification paradigm. Extensive experiment results reveal that generative LLMs exhibit significant advantages in linguistic steganalysis and demonstrate performance trends distinct from traditional approaches. Results also reveal that LLMs outperform existing baselines by a wide margin, and the domain-agnostic ability of LLMs makes it possible to train a generic steganalysis model (Both codes and trained models are openly available in https://github.com/ba0z1/Linguistic-Steganalysis-with-LLMs).
△ Less
Submitted 15 May, 2024;
originally announced May 2024.
-
ModelShield: Adaptive and Robust Watermark against Model Extraction Attack
Authors:
Kaiyi Pang,
Tao Qi,
Chuhan Wu,
Minhao Bai,
Minghu Jiang,
Yongfeng Huang
Abstract:
Large language models (LLMs) demonstrate general intelligence across a variety of machine learning tasks, thereby enhancing the commercial value of their intellectual property (IP). To protect this IP, model owners typically allow user access only in a black-box manner, however, adversaries can still utilize model extraction attacks to steal the model intelligence encoded in model generation. Wate…
▽ More
Large language models (LLMs) demonstrate general intelligence across a variety of machine learning tasks, thereby enhancing the commercial value of their intellectual property (IP). To protect this IP, model owners typically allow user access only in a black-box manner, however, adversaries can still utilize model extraction attacks to steal the model intelligence encoded in model generation. Watermarking technology offers a promising solution for defending against such attacks by embedding unique identifiers into the model-generated content. However, existing watermarking methods often compromise the quality of generated content due to heuristic alterations and lack robust mechanisms to counteract adversarial strategies, thus limiting their practicality in real-world scenarios. In this paper, we introduce an adaptive and robust watermarking method (named ModelShield) to protect the IP of LLMs. Our method incorporates a self-watermarking mechanism that allows LLMs to autonomously insert watermarks into their generated content to avoid the degradation of model content. We also propose a robust watermark detection mechanism capable of effectively identifying watermark signals under the interference of varying adversarial strategies. Besides, ModelShield is a plug-and-play method that does not require additional model training, enhancing its applicability in LLM deployments. Extensive evaluations on two real-world datasets and three LLMs demonstrate that our method surpasses existing methods in terms of defense effectiveness and robustness while significantly reducing the degradation of watermarking on the model-generated content.
△ Less
Submitted 12 January, 2025; v1 submitted 3 May, 2024;
originally announced May 2024.
-
Learnable Linguistic Watermarks for Tracing Model Extraction Attacks on Large Language Models
Authors:
Minhao Bai,
Kaiyi Pang,
Yongfeng Huang
Abstract:
In the rapidly evolving domain of artificial intelligence, safeguarding the intellectual property of Large Language Models (LLMs) is increasingly crucial. Current watermarking techniques against model extraction attacks, which rely on signal insertion in model logits or post-processing of generated text, remain largely heuristic. We propose a novel method for embedding learnable linguistic waterma…
▽ More
In the rapidly evolving domain of artificial intelligence, safeguarding the intellectual property of Large Language Models (LLMs) is increasingly crucial. Current watermarking techniques against model extraction attacks, which rely on signal insertion in model logits or post-processing of generated text, remain largely heuristic. We propose a novel method for embedding learnable linguistic watermarks in LLMs, aimed at tracing and preventing model extraction attacks. Our approach subtly modifies the LLM's output distribution by introducing controlled noise into token frequency distributions, embedding an statistically identifiable controllable watermark.We leverage statistical hypothesis testing and information theory, particularly focusing on Kullback-Leibler Divergence, to differentiate between original and modified distributions effectively. Our watermarking method strikes a delicate well balance between robustness and output quality, maintaining low false positive/negative rates and preserving the LLM's original performance.
△ Less
Submitted 28 April, 2024;
originally announced May 2024.
-
Hierarchical Topological States in Thermal Diffusive Networks
Authors:
Bao Chen,
Kaiyun Pang,
Ru Zheng,
Feng Liu
Abstract:
The integration of topological concepts into electronic energy band theory has been a transformative development in condensed matter physics. Since then, this paradigm has broadened its reach, extending to a variety of physical systems, including open ones. In this study, we employ analogues of the generalized $n$-dimensional Su-Schrieffer-Heeger model, a cornerstone in understanding topological i…
▽ More
The integration of topological concepts into electronic energy band theory has been a transformative development in condensed matter physics. Since then, this paradigm has broadened its reach, extending to a variety of physical systems, including open ones. In this study, we employ analogues of the generalized $n$-dimensional Su-Schrieffer-Heeger model, a cornerstone in understanding topological insulators and higher-order topological states, to unveil a dimensional hierarchy of topological states within thermal diffusive networks. Unlike their electronic counterparts, the topological states in these networks are characterized by confined temperature profiles of dimension $(n-d)$ with constant diffusive rates, where $n$ represents the system's dimension and $d$ is the order of the topological state. Our findings demonstrate the existence of topological corner states in thermal diffusive systems up to $n=3$, along with surface and hinge states. We also identify and discuss an intermediate-order topological phase in the case $n=3$, characterized by the presence of hinge states but the absence of corner states. Furthermore, our work delves into the influence of chiral symmetry in these thermal networks, particularly focusing on topological thermal states with a near-zero diffusion rate. This research lays the foundation for advanced thermal management strategies that utilize topological states in multiple dimensions.
△ Less
Submitted 21 December, 2023;
originally announced December 2023.
-
Wired Perspectives: Multi-View Wire Art Embraces Generative AI
Authors:
Zhiyu Qu,
Lan Yang,
Honggang Zhang,
Tao Xiang,
Kaiyue Pang,
Yi-Zhe Song
Abstract:
Creating multi-view wire art (MVWA), a static 3D sculpture with diverse interpretations from different viewpoints, is a complex task even for skilled artists. In response, we present DreamWire, an AI system enabling everyone to craft MVWA easily. Users express their vision through text prompts or scribbles, freeing them from intricate 3D wire organisation. Our approach synergises 3D Bézier curves,…
▽ More
Creating multi-view wire art (MVWA), a static 3D sculpture with diverse interpretations from different viewpoints, is a complex task even for skilled artists. In response, we present DreamWire, an AI system enabling everyone to craft MVWA easily. Users express their vision through text prompts or scribbles, freeing them from intricate 3D wire organisation. Our approach synergises 3D Bézier curves, Prim's algorithm, and knowledge distillation from diffusion models or their variants (e.g., ControlNet). This blend enables the system to represent 3D wire art, ensuring spatial continuity and overcoming data scarcity. Extensive evaluation and analysis are conducted to shed insight on the inner workings of the proposed system, including the trade-off between connectivity and visual aesthetics.
△ Less
Submitted 13 June, 2024; v1 submitted 26 November, 2023;
originally announced November 2023.
-
CSAM: A 2.5D Cross-Slice Attention Module for Anisotropic Volumetric Medical Image Segmentation
Authors:
Alex Ling Yu Hung,
Haoxin Zheng,
Kai Zhao,
Xiaoxi Du,
Kaifeng Pang,
Qi Miao,
Steven S. Raman,
Demetri Terzopoulos,
Kyunghyun Sung
Abstract:
A large portion of volumetric medical data, especially magnetic resonance imaging (MRI) data, is anisotropic, as the through-plane resolution is typically much lower than the in-plane resolution. Both 3D and purely 2D deep learning-based segmentation methods are deficient in dealing with such volumetric data since the performance of 3D methods suffers when confronting anisotropic data, and 2D meth…
▽ More
A large portion of volumetric medical data, especially magnetic resonance imaging (MRI) data, is anisotropic, as the through-plane resolution is typically much lower than the in-plane resolution. Both 3D and purely 2D deep learning-based segmentation methods are deficient in dealing with such volumetric data since the performance of 3D methods suffers when confronting anisotropic data, and 2D methods disregard crucial volumetric information. Insufficient work has been done on 2.5D methods, in which 2D convolution is mainly used in concert with volumetric information. These models focus on learning the relationship across slices, but typically have many parameters to train. We offer a Cross-Slice Attention Module (CSAM) with minimal trainable parameters, which captures information across all the slices in the volume by applying semantic, positional, and slice attention on deep feature maps at different scales. Our extensive experiments using different network architectures and tasks demonstrate the usefulness and generalizability of CSAM. Associated code is available at https://github.com/aL3x-O-o-Hung/CSAM.
△ Less
Submitted 26 November, 2023; v1 submitted 7 November, 2023;
originally announced November 2023.
-
BodyFormer: Semantics-guided 3D Body Gesture Synthesis with Transformer
Authors:
Kunkun Pang,
Dafei Qin,
Yingruo Fan,
Julian Habekost,
Takaaki Shiratori,
Junichi Yamagishi,
Taku Komura
Abstract:
Automatic gesture synthesis from speech is a topic that has attracted researchers for applications in remote communication, video games and Metaverse. Learning the mapping between speech and 3D full-body gestures is difficult due to the stochastic nature of the problem and the lack of a rich cross-modal dataset that is needed for training. In this paper, we propose a novel transformer-based framew…
▽ More
Automatic gesture synthesis from speech is a topic that has attracted researchers for applications in remote communication, video games and Metaverse. Learning the mapping between speech and 3D full-body gestures is difficult due to the stochastic nature of the problem and the lack of a rich cross-modal dataset that is needed for training. In this paper, we propose a novel transformer-based framework for automatic 3D body gesture synthesis from speech. To learn the stochastic nature of the body gesture during speech, we propose a variational transformer to effectively model a probabilistic distribution over gestures, which can produce diverse gestures during inference. Furthermore, we introduce a mode positional embedding layer to capture the different motion speeds in different speaking modes. To cope with the scarcity of data, we design an intra-modal pre-training scheme that can learn the complex mapping between the speech and the 3D gesture from a limited amount of data. Our system is trained with either the Trinity speech-gesture dataset or the Talking With Hands 16.2M dataset. The results show that our system can produce more realistic, appropriate, and diverse body gestures compared to existing state-of-the-art approaches.
△ Less
Submitted 6 September, 2023;
originally announced October 2023.
-
PartDiff: Image Super-resolution with Partial Diffusion Models
Authors:
Kai Zhao,
Alex Ling Yu Hung,
Kaifeng Pang,
Haoxin Zheng,
Kyunghyun Sung
Abstract:
Denoising diffusion probabilistic models (DDPMs) have achieved impressive performance on various image generation tasks, including image super-resolution. By learning to reverse the process of gradually diffusing the data distribution into Gaussian noise, DDPMs generate new data by iteratively denoising from random noise. Despite their impressive performance, diffusion-based generative models suff…
▽ More
Denoising diffusion probabilistic models (DDPMs) have achieved impressive performance on various image generation tasks, including image super-resolution. By learning to reverse the process of gradually diffusing the data distribution into Gaussian noise, DDPMs generate new data by iteratively denoising from random noise. Despite their impressive performance, diffusion-based generative models suffer from high computational costs due to the large number of denoising steps.In this paper, we first observed that the intermediate latent states gradually converge and become indistinguishable when diffusing a pair of low- and high-resolution images. This observation inspired us to propose the Partial Diffusion Model (PartDiff), which diffuses the image to an intermediate latent state instead of pure random noise, where the intermediate latent state is approximated by the latent of diffusing the low-resolution image. During generation, Partial Diffusion Models start denoising from the intermediate distribution and perform only a part of the denoising steps. Additionally, to mitigate the error caused by the approximation, we introduce "latent alignment", which aligns the latent between low- and high-resolution images during training. Experiments on both magnetic resonance imaging (MRI) and natural images show that, compared to plain diffusion-based super-resolution methods, Partial Diffusion Models significantly reduce the number of denoising steps without sacrificing the quality of generation.
△ Less
Submitted 21 July, 2023;
originally announced July 2023.
-
Symmetry-Breaking in Point-Heated Droplets
Authors:
Khang Ee Pang,
Charles Cuvillier,
Yutaku Kita,
Lennon Ó Náraigh
Abstract:
We investigate theoretically the stability of thermo-capillary convection within a droplet when heated by a point source from below. To model the droplet, we use a mathematical model based on lubrication theory. We formulate a base-state droplet profile, and we examine its respect to small-amplitude perturbations in the azimuthal direction. Such linear stability analysis reveals that the base stat…
▽ More
We investigate theoretically the stability of thermo-capillary convection within a droplet when heated by a point source from below. To model the droplet, we use a mathematical model based on lubrication theory. We formulate a base-state droplet profile, and we examine its respect to small-amplitude perturbations in the azimuthal direction. Such linear stability analysis reveals that the base state is stable across a wide parameter space. We carry out transient simulations in three spatial dimensions: the simulations reveal that when the heating is slightly off-centered with respect to the droplet center, vortices develop within the droplet. The vortices persist when the contact line is pinned. These findings are consistent with experimental studies of point-heated sessile droplets.
△ Less
Submitted 18 July, 2023;
originally announced July 2023.
-
Spatially Resolved Gene Expression Prediction from H&E Histology Images via Bi-modal Contrastive Learning
Authors:
Ronald Xie,
Kuan Pang,
Sai W. Chung,
Catia T. Perciani,
Sonya A. MacParland,
Bo Wang,
Gary D. Bader
Abstract:
Histology imaging is an important tool in medical diagnosis and research, enabling the examination of tissue structure and composition at the microscopic level. Understanding the underlying molecular mechanisms of tissue architecture is critical in uncovering disease mechanisms and developing effective treatments. Gene expression profiling provides insight into the molecular processes underlying t…
▽ More
Histology imaging is an important tool in medical diagnosis and research, enabling the examination of tissue structure and composition at the microscopic level. Understanding the underlying molecular mechanisms of tissue architecture is critical in uncovering disease mechanisms and developing effective treatments. Gene expression profiling provides insight into the molecular processes underlying tissue architecture, but the process can be time-consuming and expensive. We present BLEEP (Bi-modaL Embedding for Expression Prediction), a bi-modal embedding framework capable of generating spatially resolved gene expression profiles of whole-slide Hematoxylin and eosin (H&E) stained histology images. BLEEP uses contrastive learning to construct a low-dimensional joint embedding space from a reference dataset using paired image and expression profiles at micrometer resolution. With this approach, the gene expression of any query image patch can be imputed using the expression profiles from the reference dataset. We demonstrate BLEEP's effectiveness in gene expression prediction by benchmarking its performance on a human liver tissue dataset captured using the 10x Visium platform, where it achieves significant improvements over existing methods. Our results demonstrate the potential of BLEEP to provide insights into the molecular mechanisms underlying tissue architecture, with important implications in diagnosis and research of various diseases. The proposed approach can significantly reduce the time and cost associated with gene expression profiling, opening up new avenues for high-throughput analysis of histology images for both research and clinical applications.
△ Less
Submitted 27 October, 2023; v1 submitted 2 June, 2023;
originally announced June 2023.
-
SketchXAI: A First Look at Explainability for Human Sketches
Authors:
Zhiyu Qu,
Yulia Gryaditskaya,
Ke Li,
Kaiyue Pang,
Tao Xiang,
Yi-Zhe Song
Abstract:
This paper, for the very first time, introduces human sketches to the landscape of XAI (Explainable Artificial Intelligence). We argue that sketch as a ``human-centred'' data form, represents a natural interface to study explainability. We focus on cultivating sketch-specific explainability designs. This starts by identifying strokes as a unique building block that offers a degree of flexibility i…
▽ More
This paper, for the very first time, introduces human sketches to the landscape of XAI (Explainable Artificial Intelligence). We argue that sketch as a ``human-centred'' data form, represents a natural interface to study explainability. We focus on cultivating sketch-specific explainability designs. This starts by identifying strokes as a unique building block that offers a degree of flexibility in object construction and manipulation impossible in photos. Following this, we design a simple explainability-friendly sketch encoder that accommodates the intrinsic properties of strokes: shape, location, and order. We then move on to define the first ever XAI task for sketch, that of stroke location inversion SLI. Just as we have heat maps for photos, and correlation matrices for text, SLI offers an explainability angle to sketch in terms of asking a network how well it can recover stroke locations of an unseen sketch. We offer qualitative results for readers to interpret as snapshots of the SLI process in the paper, and as GIFs on the project page. A minor but interesting note is that thanks to its sketch-specific design, our sketch encoder also yields the best sketch recognition accuracy to date while having the smallest number of parameters. The code is available at \url{https://sketchxai.github.io}.
△ Less
Submitted 23 April, 2023;
originally announced April 2023.
-
OCELOT: Overlapped Cell on Tissue Dataset for Histopathology
Authors:
Jeongun Ryu,
Aaron Valero Puche,
JaeWoong Shin,
Seonwook Park,
Biagio Brattoli,
Jinhee Lee,
Wonkyung Jung,
Soo Ick Cho,
Kyunghyun Paeng,
Chan-Young Ock,
Donggeun Yoo,
Sérgio Pereira
Abstract:
Cell detection is a fundamental task in computational pathology that can be used for extracting high-level medical information from whole-slide images. For accurate cell detection, pathologists often zoom out to understand the tissue-level structures and zoom in to classify cells based on their morphology and the surrounding context. However, there is a lack of efforts to reflect such behaviors by…
▽ More
Cell detection is a fundamental task in computational pathology that can be used for extracting high-level medical information from whole-slide images. For accurate cell detection, pathologists often zoom out to understand the tissue-level structures and zoom in to classify cells based on their morphology and the surrounding context. However, there is a lack of efforts to reflect such behaviors by pathologists in the cell detection models, mainly due to the lack of datasets containing both cell and tissue annotations with overlapping regions. To overcome this limitation, we propose and publicly release OCELOT, a dataset purposely dedicated to the study of cell-tissue relationships for cell detection in histopathology. OCELOT provides overlapping cell and tissue annotations on images acquired from multiple organs. Within this setting, we also propose multi-task learning approaches that benefit from learning both cell and tissue tasks simultaneously. When compared against a model trained only for the cell detection task, our proposed approaches improve cell detection performance on 3 datasets: proposed OCELOT, public TIGER, and internal CARP datasets. On the OCELOT test set in particular, we show up to 6.79 improvement in F1-score. We believe the contributions of this paper, including the release of the OCELOT dataset at https://lunit-io.github.io/research/publications/ocelot are a crucial starting point toward the important research direction of incorporating cell-tissue relationships in computation pathology.
△ Less
Submitted 23 March, 2023; v1 submitted 23 March, 2023;
originally announced March 2023.