Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 92 results for author: Paeng, K

.
  1. arXiv:2608.13888  [pdf, ps, other

    cs.LG

    Fashion Outfit Generation via Unified Sequential Composition Models

    Authors: Kaicheng Pang, Xingxing Zou, Ruohan Xu, Waikeung Wong

    Abstract: The task of synthesizing stylistically coherent fashion outfits from massive item libraries, known as fashion outfit generation, remains a non-trivial challenge, primarily due to the non-monotonic and implicit nature of aesthetic compatibility, coupled with the exponentially large combinatorial search space. In this paper, we formalize this task as Constrained Ensemble Generation (CEG) and model i… ▽ More

    Submitted 13 August, 2026; originally announced August 2026.

  2. arXiv:2608.01330  [pdf, ps, other

    math.CV

    On the extension of Kähler currents on compact complex manifolds

    Authors: Jiafu Ning, Kai Pang, Haoyuan Sun, Zhiwei Wang, Xiangyu Zhou

    Abstract: Let $(X,ω)$ be a compact Kähler manifold and let $V\subset X$ be a closed complex submanifold. Coman-Guedj-Zeriahi proposed the problem: is every $ω|_V$-plurisubharmonic function on $V$ the restriction of an $ω$-plurisubharmonic function on $X$? In this paper, we solve this problem affirmatively, even for a compact Hermitian manifold.

    Submitted 2 August, 2026; originally announced August 2026.

    Comments: 9 pages, comments welcome

  3. arXiv:2607.12797  [pdf, ps, other

    math.CV math.AP

    Capacity Stability of Complex Monge-Ampère Equations with Moving Prescribed Singularities

    Authors: Kai Pang, Haoyuan Sun, Zhiwei Wang, Xiangyu Zhou

    Abstract: For complex Monge-Ampère equations with moving big cohomology classes and prescribed model singularities of positive Monge-Ampère mass, we prove that, under total variation convergence of the right-hand side non-pluripolar positive Radon measures, convergence of the prescribed model potentials in Monge-Ampère capacity is equivalent to convergence in capacity of the associated normalized solutions.… ▽ More

    Submitted 14 July, 2026; originally announced July 2026.

    Comments: 59pages. Comments welcome!

  4. arXiv:2606.11653  [pdf, ps, other

    physics.flu-dyn

    On the Modelling of the Hydrodynamic Drag of Mangroves

    Authors: Khang Ee Pang, Zhi Yung Tay

    Abstract: Mangroves are increasingly promoted as nature-based solutions for coastal protection, yet many existing models neglect the vertical variation of vegetation biomass, leading to oversimplified representations of root-flow interactions. In this study, we introduce a generalised parametrisation of the mangrove vegetation profile that is applicable across multiple mangrove species and derive a wave att… ▽ More

    Submitted 10 June, 2026; originally announced June 2026.

    Comments: 10 pages, 12 figures. Presented at the International Conference of Hydrodynamics (ICHD) 2026

    MSC Class: 76B15

  5. arXiv:2605.19833  [pdf, ps, other

    cs.SD cs.AI cs.CL cs.MM eess.AS

    Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation

    Authors: Zhifei Xie, Kaiyu Pang, Haobin Zhang, Deheng Ye, Xiaobin Hu, Shuicheng Yan, Chunyan Miao

    Abstract: Despite rapid advances in automatic speech recognition (ASR) and large audio-language models, robust recognition in real-world environments remains limited by an "acoustic robustness bottleneck": models often lose acoustic grounding and produce omissions or hallucinations under severe, compositional distortions. We propose Mega-ASR, a unified ASR-in-the-wild framework that combines scalable compou… ▽ More

    Submitted 19 May, 2026; originally announced May 2026.

    Comments: Project page: https://xzf-thu.github.io/Mega-ASR/. Code, models, and dataset will be released. A robust ASR framework targeting in-the-wild and compositional acoustic scenarios where conventional ASR systems fail

  6. arXiv:2605.08144  [pdf, ps, other

    cs.LG cs.AI cs.CV

    NoiseRater: Meta-Learned Noise Valuation for Diffusion Model Training

    Authors: Fang Wu, Haokai Zhao, Da Xing, Hanqun Cao, Tinson Xu, Yanchao Li, Xiangru Tang, Zehong Wang, Aaron Tu, Kuan Pang, Hanchen Wang, Hongbin Lin, Zeqi Zhou, Yinxi Li, Peng Xia, Li Erran Li, Molei Tao, Jure Leskovec, Aditya Joshi, Yejin Choi

    Abstract: Diffusion models have achieved remarkable success across a wide range of generative tasks, yet their training paradigm largely treats injected noise as uniformly informative. In this work, we challenge this assumption and introduce NoiseRater, a meta-learning framework for instance-level noise valuation in diffusion model training. We propose a parametric noise rater that assigns importance scores… ▽ More

    Submitted 2 May, 2026; originally announced May 2026.

  7. arXiv:2605.02937  [pdf, ps, other

    cs.LG cs.AI cs.CE

    Proteo-R1: Reasoning Foundation Models for De Novo Protein Design

    Authors: Fang Wu, Weihao Xuan, Heli Qi, Hanqun Cao, Heng-Jui Chang, Zeqi Zhou, Haokai Zhao, Ma Jian, Carl Ma, Yu-Chi Cheng, Kuan Pang, Xiangru Tang, Zehong Wang, Guanlue Li, Hanchen Wang, Kejun Ying, Pan Lu, Chiho Im, Seungju Han, Peng Xia, Tinson Xu, Yinxi Li, Deyao Zhu, Pheng-Ann Heng, Naoto Yokoya , et al. (4 additional authors not shown)

    Abstract: Deep learning in de novo protein design has achieved atomic-level fidelity. However, existing models remain largely non-deliberative: they directly synthesize molecular geometries without explicitly reasoning about which residues or interactions are functionally essential. As a result, design decisions are entangled with continuous sampling dynamics, limiting interpretability, controllability, and… ▽ More

    Submitted 10 August, 2026; v1 submitted 1 May, 2026; originally announced May 2026.

    Journal ref: ICML 2026

  8. arXiv:2604.21394  [pdf, ps, other

    cs.CR

    Provably Secure Steganography Based on List Decoding

    Authors: Kaiyi Pang, Minhao Bai

    Abstract: Steganography embeds secret messages in seemingly innocuous carriers for covert communication under surveillance. Current Provably Secure Steganography (PSS) schemes based on language models can guarantee computational indistinguishability between the covertext and stegotext. However, achieving high embedding capacity remains a challenge for existing PSS. The inefficient entropy utilization render… ▽ More

    Submitted 28 April, 2026; v1 submitted 23 April, 2026; originally announced April 2026.

  9. arXiv:2604.07301  [pdf

    cond-mat.mtrl-sci

    Symmetry-protected four double-Weyl fermions and their topological phase transitions in nonmagnetic crystals

    Authors: Yun-Yun Bai, Ke-Xin Pang, Yan Gao

    Abstract: Realizing Weyl semimetals (WSMs) with the minimal number of Weyl points (WPs) fundamentally simplifies extracting intrinsic topological responses. While a minimum of four conventional ($|C|=1$) WPs in nonmagnetic crystals is well-established, the exact symmetry requirements and material realization for the unique configuration of four unconventional double-Weyl points (DWPs, $|C|=2$) remain unreso… ▽ More

    Submitted 8 April, 2026; originally announced April 2026.

    Comments: 30 pages, 6 figures

  10. arXiv:2603.17349  [pdf

    cond-mat.mtrl-sci

    Single-pair charge-2 Weyl-Dirac composite semimetals

    Authors: Hui-Jing Zheng, Ke-Xin Pang, Yun-Yun Bai, Yanfeng Ge, Yan Gao

    Abstract: The Nielsen--Ninomiya theorem requires that the total topological chiral charges in a crystal vanish, a constraint typically satisfied by identical nodes like Weyl--Weyl pairs. Whether a minimal heterogeneous configuration -- comprising a single Weyl point (WP) and a single Dirac point (DP) -- can exist in an electronic system has remained unresolved. Here, by systematically classifying all 1651 m… ▽ More

    Submitted 18 March, 2026; originally announced March 2026.

    Comments: 25 pages, 4 figures

  11. arXiv:2603.08519  [pdf, ps, other

    cs.RO

    AtomVLA: Scalable Post-Training for Robotic Manipulation via Predictive Latent World Models

    Authors: Xiaoquan Sun, Zetian Xu, Chen Cao, Zonghe Liu, Yihan Sun, Jingrui Pang, Ruijian Zhang, Zhen Yang, Kang Pang, Dingxin He, Mingqi Yuan, Jiayu Chen

    Abstract: Vision-Language-Action (VLA) models demonstrate remarkable potential for generalizable robotic manipulation. The execution of complex multi-step behaviors in VLA models can be improved by robust instruction grounding, a critical component for effective control. However, current paradigms predominantly rely on coarse, high-level task instructions during supervised fine-tuning. This instruction grou… ▽ More

    Submitted 9 March, 2026; originally announced March 2026.

  12. arXiv:2602.22566  [pdf

    cond-mat.mtrl-sci

    Symmetry-Protected Minimum of Four Conventional Weyl Points in Nonmagnetic Crystals

    Authors: Ze-Xin Xue, Ke-Xin Pang, Yun-Yun Bai, Yanfeng Ge, Yong Liu, Yan Gao

    Abstract: Realizing nonmagnetic Weyl semimetals (WSMs) with the minimal number of conventional Weyl points (WPs) and a clean Fermi surface remains a central challenge. Here, combining symmetry analysis with first-principles calculations, we establish the definitive conditions under which a nonmagnetic crystal can host exactly four conventional ($C = \pm 1$) WPs, identifying 76 space groups in the spinless l… ▽ More

    Submitted 25 February, 2026; originally announced February 2026.

    Comments: 6 figures

  13. arXiv:2602.21819  [pdf, ps, other

    cs.CV cs.AI

    SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic Guidance

    Authors: Minghan Yang, Lan Yang, Ke Li, Honggang Zhang, Kaiyue Pang, Yizhe Song

    Abstract: Reconstructing dynamic visual experiences from brain activity provides a compelling avenue for exploring the neural mechanisms of human visual perception. While recent progress in fMRI-based image reconstruction has been notable, extending this success to video reconstruction remains a significant challenge. Current fMRI-to-video reconstruction approaches consistently encounter two major shortcomi… ▽ More

    Submitted 17 August, 2026; v1 submitted 25 February, 2026; originally announced February 2026.

  14. arXiv:2601.17857  [pdf, ps, other

    cs.CV

    SynMind: Reducing Semantic Hallucination in fMRI-Based Image Reconstruction

    Authors: Lan Yang, Minghan Yang, Ke Li, Honggang Zhang, Kaiyue Pang, Yi-Zhe Song

    Abstract: Recent advances in fMRI-based image reconstruction have achieved remarkable photo-realistic fidelity. Yet, a persistent limitation remains: while reconstructed images often appear naturalistic and holistically similar to the target stimuli, they frequently suffer from severe semantic misalignment -- salient objects are often replaced or hallucinated despite high visual quality. In this work, we ad… ▽ More

    Submitted 25 January, 2026; originally announced January 2026.

  15. arXiv:2512.07084  [pdf, ps, other

    math.CV

    Degenerate Complex Hessian type equations on compact Hermitian manifolds and Applications

    Authors: Kai Pang, Haoyuan Sun, Zhiwei Wang, Xiangyu Zhou

    Abstract: The aim of this paper is to further develop the theory of the degenerate complex Hessian equations on compact Hermitian manifolds. Building upon the generalization of the Bedford-Taylor pluripotential theory to complex Hessian equations by Kołodziej-Nguyen, we solve these equations in the $(ω, m)$-positive cone, $(ω, m)$-big classes and in nef classes, where $ω$ is a reference Hermitian metric. Th… ▽ More

    Submitted 7 December, 2025; originally announced December 2025.

    Comments: 88pages. Comments Welcome!

  16. arXiv:2511.14161  [pdf, ps, other

    cs.RO cs.CV

    RoboTidy : A 3D Gaussian Splatting Household Tidying Benchmark for Embodied Navigation and Action

    Authors: Xiaoquan Sun, Ruijian Zhang, Kang Pang, Bingchen Miao, Yuxiang Tan, Zhen Yang, Ming Li, Jiayu Chen

    Abstract: Household tidying is an important application area, yet current benchmarks neither model user preferences nor support mobility, and they generalize poorly, making it hard to comprehensively assess integrated language-to-action capabilities. To address this, we propose RoboTidy, a unified benchmark for language-guided household tidying that supports Vision-Language-Action (VLA) and Vision-Language-… ▽ More

    Submitted 18 November, 2025; v1 submitted 18 November, 2025; originally announced November 2025.

  17. arXiv:2510.25319  [pdf, ps, other

    cs.GR cs.AI

    4-Doodle: Text to 3D Sketches that Move!

    Authors: Hao Chen, Jiaqi Wang, Yonggang Qi, Ke Li, Kaiyue Pang, Yi-Zhe Song

    Abstract: We present a novel task: text-to-3D sketch animation, which aims to bring freeform sketches to life in dynamic 3D space. Unlike prior works focused on photorealistic content generation, we target sparse, stylized, and view-consistent 3D vector sketches, a lightweight and interpretable medium well-suited for visual communication and prototyping. However, this task is very challenging: (i) no paired… ▽ More

    Submitted 29 October, 2025; originally announced October 2025.

  18. arXiv:2510.17171  [pdf, ps, other

    cs.CV

    Generation then Reconstruction: Accelerating Masked Autoregressive Models via Two-Stage Sampling

    Authors: Feihong Yan, Peiru Wang, Yao Zhu, Kaiyu Pang, Qingyan Wei, Huiqi Li, Linfeng Zhang

    Abstract: Masked Autoregressive (MAR) models promise better efficiency in visual generation than autoregressive (AR) models for the ability of parallel generation, yet their acceleration potential remains constrained by the modeling complexity of spatially correlated visual tokens in a single step. To address this limitation, we introduce Generation then Reconstruction (GtR), a training-free hierarchical sa… ▽ More

    Submitted 26 January, 2026; v1 submitted 20 October, 2025; originally announced October 2025.

    Comments: 12 pages, 6 figures

  19. OCELOT 2023: Cell Detection from Cell-Tissue Interaction Challenge

    Authors: JaeWoong Shin, Jeongun Ryu, Aaron Valero Puche, Jinhee Lee, Biagio Brattoli, Wonkyung Jung, Soo Ick Cho, Kyunghyun Paeng, Chan-Young Ock, Donggeun Yoo, Zhaoyang Li, Wangkai Li, Huayu Mai, Joshua Millward, Zhen He, Aiden Nibali, Lydia Anette Schoenpflug, Viktor Hendrik Koelzer, Xu Shuoyu, Ji Zheng, Hu Bin, Yu-Wen Lo, Ching-Hui Yang, Sérgio Pereira

    Abstract: Pathologists routinely alternate between different magnifications when examining Whole-Slide Images, allowing them to evaluate both broad tissue morphology and intricate cellular details to form comprehensive diagnoses. However, existing deep learning-based cell detection models struggle to replicate these behaviors and learn the interdependent semantics between structures at different magnificati… ▽ More

    Submitted 11 September, 2025; originally announced September 2025.

    Comments: This is the accepted manuscript of an article published in Medical Image Analysis (Elsevier). The final version is available at: https://doi.org/10.1016/j.media.2025.103751

    Journal ref: Medical Image Analysis 106 (2025) 103751

  20. arXiv:2508.15827  [pdf, ps, other

    cs.CL cs.AI cs.LG eess.AS

    Mini-Omni-Reasoner: Token-Level Thinking-in-Speaking in Large Speech Models

    Authors: Zhifei Xie, Ziyang Ma, Zihang Liu, Kaiyu Pang, Hongyu Li, Jialin Zhang, Yue Liao, Deheng Ye, Chunyan Miao, Shuicheng Yan

    Abstract: Reasoning is essential for effective communication and decision-making. While recent advances in LLMs and MLLMs have shown that incorporating explicit reasoning significantly improves understanding and generalization, reasoning in LSMs remains in a nascent stage. Early efforts attempt to transfer the "Thinking-before-Speaking" paradigm from textual models to speech. However, this sequential formul… ▽ More

    Submitted 20 September, 2025; v1 submitted 18 August, 2025; originally announced August 2025.

    Comments: Technical report; Work in progress. Project page: https://github.com/xzf-thu/Mini-Omni-Reasoner

  21. arXiv:2507.20907  [pdf, ps, other

    cs.CV cs.AI

    SCORPION: Addressing Scanner-Induced Variability in Histopathology

    Authors: Jeongun Ryu, Heon Song, Seungeun Lee, Soo Ick Cho, Jiwon Shin, Kyunghyun Paeng, Sérgio Pereira

    Abstract: Ensuring reliable model performance across diverse domains is a critical challenge in computational pathology. A particular source of variability in Whole-Slide Images is introduced by differences in digital scanners, thus calling for better scanner generalization. This is critical for the real-world adoption of computational pathology, where the scanning devices may differ per institution or hosp… ▽ More

    Submitted 17 September, 2025; v1 submitted 28 July, 2025; originally announced July 2025.

    Comments: Accepted in UNSURE 2025 workshop in MICCAI

  22. Annotation-Free Human Sketch Quality Assessment

    Authors: Lan Yang, Kaiyue Pang, Honggang Zhang, Yi-Zhe Song

    Abstract: As lovely as bunnies are, your sketched version would probably not do them justice (Fig.~\ref{fig:intro}). This paper recognises this very problem and studies sketch quality assessment for the first time -- letting you find these badly drawn ones. Our key discovery lies in exploiting the magnitude ($L_2$ norm) of a sketch feature as a quantitative quality metric. We propose Geometry-Aware Classifi… ▽ More

    Submitted 28 July, 2025; originally announced July 2025.

    Comments: Accepted by IJCV

  23. arXiv:2506.08423  [pdf

    cond-mat.mtrl-sci cs.LG physics.ins-det

    Mic-hackathon 2024: Hackathon on Machine Learning for Electron and Scanning Probe Microscopy

    Authors: Utkarsh Pratiush, Austin Houston, Kamyar Barakati, Aditya Raghavan, Dasol Yoon, Harikrishnan KP, Zhaslan Baraissov, Desheng Ma, Samuel S. Welborn, Mikolaj Jakowski, Shawn-Patrick Barhorst, Alexander J. Pattison, Panayotis Manganaris, Sita Sirisha Madugula, Sai Venkata Gayathri Ayyagari, Vishal Kennedy, Ralph Bulanadi, Michelle Wang, Kieran J. Pang, Ian Addison-Smith, Willy Menacho, Horacio V. Guzman, Alexander Kiefer, Nicholas Furth, Nikola L. Kolev , et al. (48 additional authors not shown)

    Abstract: Microscopy is a primary source of information on materials structure and functionality at nanometer and atomic scales. The data generated is often well-structured, enriched with metadata and sample histories, though not always consistent in detail or format. The adoption of Data Management Plans (DMPs) by major funding agencies promotes preservation and access. However, deriving insights remains d… ▽ More

    Submitted 27 June, 2025; v1 submitted 9 June, 2025; originally announced June 2025.

    Journal ref: Mach. Learn.: Sci. Technol. 6 (2025) 040701

  24. arXiv:2504.17826  [pdf, other

    cs.CV cs.AI

    FashionM3: Multimodal, Multitask, and Multiround Fashion Assistant based on Unified Vision-Language Model

    Authors: Kaicheng Pang, Xingxing Zou, Waikeung Wong

    Abstract: Fashion styling and personalized recommendations are pivotal in modern retail, contributing substantial economic value in the fashion industry. With the advent of vision-language models (VLM), new opportunities have emerged to enhance retailing through natural language and visual interactions. This work proposes FashionM3, a multimodal, multitask, and multiround fashion assistant, built upon a VLM… ▽ More

    Submitted 23 April, 2025; originally announced April 2025.

  25. arXiv:2504.12579  [pdf, ps, other

    cs.CR cs.CL

    Provable Secure Steganography Based on Adaptive Dynamic Sampling

    Authors: Kaiyi Pang, Minhao Bai

    Abstract: The security of private communication is increasingly at risk due to widespread surveillance. Steganography, a technique for embedding secret messages within innocuous carriers, enables covert communication over monitored channels. Provably Secure Steganography (PSS), which ensures computational indistinguishability between the normal model output and steganography output, is the state-of-the-art… ▽ More

    Submitted 12 February, 2026; v1 submitted 16 April, 2025; originally announced April 2025.

  26. arXiv:2503.19169  [pdf, other

    gr-qc astro-ph.IM physics.ins-det

    A cryogenic test-mass suspension with flexures operating in compression for third-generation gravitational-wave detectors

    Authors: Fabián E. Peña Arellano, Nelson L. Leon, Leonardo González López, Riccardo DeSalvo, Harry Themann, Esra Zerina Appavuravther, Guerino Avallone, Francesca Badaracco, Mark A. Barton, Alessandro Bertolini, Christian Chavez, Andy Damas, Richard Damas, Britney Gallego, Eric Hennes, Gerardo Iannone, Seth Linker, Marina Mondin, Claudia Moreno, Kevin Pang, Stefano Selleri, Mynor Soto, Flavio Travasso, Joris Van-Heijningen, Fernando Velez , et al. (1 additional authors not shown)

    Abstract: This paper presents an analysis of the conceptual design of a novel silicon suspension for the cryogenic test-mass mirrors of the low-frequency detector of the Einstein Telescope gravitational-wave observatory. In traditional suspensions, tensional stress is a severe limitation for achieving low thermal noise, safer mechanical margins and high thermal conductance simultaneously. In order to keep t… ▽ More

    Submitted 24 March, 2025; originally announced March 2025.

  27. Applications of Large Models in Medicine

    Authors: YunHe Su, Zhengyang Lu, Junhui Liu, Ke Pang, Haoran Dai, Sa Liu, Yuxin Jia, Lujia Ge, Jing-min Yang

    Abstract: This paper explores the advancements and applications of large-scale models in the medical field, with a particular focus on Medical Large Models (MedLMs). These models, encompassing Large Language Models (LLMs), Vision Models, 3D Large Models, and Multimodal Models, are revolutionizing healthcare by enhancing disease prediction, diagnostic assistance, personalized treatment planning, and drug dis… ▽ More

    Submitted 7 October, 2025; v1 submitted 24 February, 2025; originally announced February 2025.

  28. arXiv:2501.09617  [pdf, ps, other

    cs.CV

    WMamba: Wavelet-based Mamba for Face Forgery Detection

    Authors: Siran Peng, Tianshuo Zhang, Li Gao, Xiangyu Zhu, Haoyuan Zhang, Kai Pang, Zhen Lei

    Abstract: The rapid evolution of deepfake generation technologies necessitates the development of robust face forgery detection algorithms. Recent studies have demonstrated that wavelet analysis can enhance the generalization abilities of forgery detectors. Wavelets effectively capture key facial contours, often slender, fine-grained, and globally distributed, that may conceal subtle forgery artifacts imper… ▽ More

    Submitted 21 October, 2025; v1 submitted 16 January, 2025; originally announced January 2025.

    Comments: Accepted by ACM MM 2025

  29. arXiv:2501.00786  [pdf, other

    cs.CR

    Shifting-Merging: Secure, High-Capacity and Efficient Steganography via Large Language Models

    Authors: Minhao Bai, Jinshuai Yang, Kaiyi Pang, Yongfeng Huang, Yue Gao

    Abstract: In the face of escalating surveillance and censorship within the cyberspace, the sanctity of personal privacy has come under siege, necessitating the development of steganography, which offers a way to securely hide messages within innocent-looking texts. Previous methods alternate the texts to hide private massages, which is not secure. Large Language Models (LLMs) provide high-quality and explic… ▽ More

    Submitted 1 January, 2025; originally announced January 2025.

  30. arXiv:2412.19652  [pdf, ps, other

    cs.CR

    A Plug-and-Play Method for Improving Imperceptibility and Capacity in Practical Generative Text Steganography

    Authors: Kaiyi Pang

    Abstract: Linguistic steganography embeds secret information into seemingly innocuous text to safeguard privacy under surveillance. Generative linguistic steganography leverages the probability distributions of language models (LMs) and applies steganographic algorithms during generation, and has attracted increasing attention with the rise of large language models (LLMs). To strengthen security, prior work… ▽ More

    Submitted 26 June, 2026; v1 submitted 27 December, 2024; originally announced December 2024.

  31. arXiv:2412.17541  [pdf, ps, other

    cs.CV cs.AI

    Spoof Trace Discovery for Deep Learning Based Explainable Face Anti-Spoofing

    Authors: Haoyuan Zhang, Xiangyu Zhu, Li Gao, Jiawei Pan, Kai Pang, Guoying Zhao, Zhen Lei

    Abstract: With the rapid growth usage of face recognition in people's daily life, face anti-spoofing becomes increasingly important to avoid malicious attacks. Recent face anti-spoofing models can reach a high classification accuracy on multiple datasets but these models can only tell people "this face is fake" while lacking the explanation to answer "why it is fake". Such a system undermines trustworthines… ▽ More

    Submitted 5 September, 2025; v1 submitted 23 December, 2024; originally announced December 2024.

    Comments: Accepted by IJCB 2025. Keywords: explainable artificial intelligence, face anti-spoofing, explainable face anti-spoofing, interpretable

  32. arXiv:2412.11594  [pdf, other

    cs.CV

    VersaGen: Unleashing Versatile Visual Control for Text-to-Image Synthesis

    Authors: Zhipeng Chen, Lan Yang, Yonggang Qi, Honggang Zhang, Kaiyue Pang, Ke Li, Yi-Zhe Song

    Abstract: Despite the rapid advancements in text-to-image (T2I) synthesis, enabling precise visual control remains a significant challenge. Existing works attempted to incorporate multi-facet controls (text and sketch), aiming to enhance the creative control over generated images. However, our pilot study reveals that the expressive power of humans far surpasses the capabilities of current methods. Users de… ▽ More

    Submitted 27 December, 2024; v1 submitted 16 December, 2024; originally announced December 2024.

    Comments: The paper has been accepted by AAAI 2025. Paper code: https://github.com/FelixChan9527/VersaGen_official

    ACM Class: I.4.9; I.4.10

  33. arXiv:2412.11547  [pdf, ps, other

    math.CV

    Weak convergence of complex Monge-Ampère operators on compact Hermitian manifolds

    Authors: Kai Pang, Haoyuan Sun, Zhiwei Wang

    Abstract: Let $(X,ω)$ be a compact Hermitian manifold and let $\{β\}\in H^{1,1}(X,\mathbb R)$ be a real $(1,1)$-class with a smooth representative $β$, such that $\int_Xβ^n>0$. Assume that there is a bounded $β$-plurisubharmonic function $ρ$ on $X$. First, we provide a criterion for the weak convergence of non-pluripolar complex Monge-Ampère measures associated to a sequence of $β$-plurisubharmonic function… ▽ More

    Submitted 16 December, 2024; originally announced December 2024.

    Comments: 36pages, Comments welcome!

  34. arXiv:2412.11043  [pdf, other

    cs.CR

    Semantic Steganography: A Framework for Robust and High-Capacity Information Hiding using Large Language Models

    Authors: Minhao Bai, Jinshuai Yang, Kaiyi Pang, Yongfeng Huang, Yue Gao

    Abstract: In the era of Large Language Models (LLMs), generative linguistic steganography has become a prevalent technique for hiding information within model-generated texts. However, traditional steganography methods struggle to effectively align steganographic texts with original model-generated texts due to the lower entropy of the predicted probability distribution of LLMs. This results in a decrease i… ▽ More

    Submitted 14 December, 2024; originally announced December 2024.

  35. arXiv:2409.17187  [pdf, other

    physics.flu-dyn math.NA

    Applications and Novel Regularization of the Thin-Film Equation

    Authors: Khang Ee Pang

    Abstract: The classical no-slip boundary condition of the Navier-Stokes equations fails to describe the spreading motion of a droplet on a substrate due to the missing small-scale physics near the contact line. In this thesis, we introduce a novel regularization of the thin-film equation to model droplet spreading. The solution of the regularized thin-film equation -- the Geometric Thin-Film Equation is stu… ▽ More

    Submitted 25 September, 2024; originally announced September 2024.

    Comments: 157 pages

    MSC Class: 76A20

  36. arXiv:2407.20643  [pdf

    cs.CV

    Generalizing AI-driven Assessment of Immunohistochemistry across Immunostains and Cancer Types: A Universal Immunohistochemistry Analyzer

    Authors: Biagio Brattoli, Mohammad Mostafavi, Taebum Lee, Wonkyung Jung, Jeongun Ryu, Seonwook Park, Jongchan Park, Sergio Pereira, Seunghwan Shin, Sangjoon Choi, Hyojin Kim, Donggeun Yoo, Siraj M. Ali, Kyunghyun Paeng, Chan-Young Ock, Soo Ick Cho, Seokhwi Kim

    Abstract: Despite advancements in methodologies, immunohistochemistry (IHC) remains the most utilized ancillary test for histopathologic and companion diagnostics in targeted therapies. However, objective IHC assessment poses challenges. Artificial intelligence (AI) has emerged as a potential solution, yet its development requires extensive training for each cancer and IHC type, limiting versatility. We dev… ▽ More

    Submitted 30 July, 2024; originally announced July 2024.

  37. arXiv:2407.13499  [pdf, other

    cs.CR

    Provably Robust and Secure Steganography in Asymmetric Resource Scenario

    Authors: Minhao Bai, Jinshuai Yang, Kaiyi Pang, Xin Xu, Zhen Yang, Yongfeng Huang

    Abstract: To circumvent the unbridled and ever-encroaching surveillance and censorship in cyberspace, steganography has garnered attention for its ability to hide private information in innocent-looking carriers. Current provably secure steganography approaches require a pair of encoder and decoder to hide and extract private messages, both of which must run the same model with the same input to obtain iden… ▽ More

    Submitted 24 November, 2024; v1 submitted 18 July, 2024; originally announced July 2024.

  38. arXiv:2407.01146  [pdf, other

    eess.IV cs.CV

    Cross-Slice Attention and Evidential Critical Loss for Uncertainty-Aware Prostate Cancer Detection

    Authors: Alex Ling Yu Hung, Haoxin Zheng, Kai Zhao, Kaifeng Pang, Demetri Terzopoulos, Kyunghyun Sung

    Abstract: Current deep learning-based models typically analyze medical images in either 2D or 3D albeit disregarding volumetric information or suffering sub-optimal performance due to the anisotropic resolution of MR data. Furthermore, providing an accurate uncertainty estimation is beneficial to clinicians, as it indicates how confident a model is about its prediction. We propose a novel 2.5D cross-slice a… ▽ More

    Submitted 1 July, 2024; originally announced July 2024.

  39. arXiv:2405.09090  [pdf, other

    cs.CR

    Towards Next-Generation Steganalysis: LLMs Unleash the Power of Detecting Steganography

    Authors: Minhao Bai. Jinshuai Yang, Kaiyi Pang, Huili Wang, Yongfeng Huang

    Abstract: Linguistic steganography provides convenient implementation to hide messages, particularly with the emergence of AI generation technology. The potential abuse of this technology raises security concerns within societies, calling for powerful linguistic steganalysis to detect carrier containing steganographic messages. Existing methods are limited to finding distribution differences between stegano… ▽ More

    Submitted 15 May, 2024; originally announced May 2024.

  40. arXiv:2405.02365  [pdf, other

    cs.CR

    ModelShield: Adaptive and Robust Watermark against Model Extraction Attack

    Authors: Kaiyi Pang, Tao Qi, Chuhan Wu, Minhao Bai, Minghu Jiang, Yongfeng Huang

    Abstract: Large language models (LLMs) demonstrate general intelligence across a variety of machine learning tasks, thereby enhancing the commercial value of their intellectual property (IP). To protect this IP, model owners typically allow user access only in a black-box manner, however, adversaries can still utilize model extraction attacks to steal the model intelligence encoded in model generation. Wate… ▽ More

    Submitted 12 January, 2025; v1 submitted 3 May, 2024; originally announced May 2024.

  41. arXiv:2405.01509  [pdf, other

    cs.CR cs.AI cs.CL

    Learnable Linguistic Watermarks for Tracing Model Extraction Attacks on Large Language Models

    Authors: Minhao Bai, Kaiyi Pang, Yongfeng Huang

    Abstract: In the rapidly evolving domain of artificial intelligence, safeguarding the intellectual property of Large Language Models (LLMs) is increasingly crucial. Current watermarking techniques against model extraction attacks, which rely on signal insertion in model logits or post-processing of generated text, remain largely heuristic. We propose a novel method for embedding learnable linguistic waterma… ▽ More

    Submitted 28 April, 2024; originally announced May 2024.

    Comments: not decided

  42. Hierarchical Topological States in Thermal Diffusive Networks

    Authors: Bao Chen, Kaiyun Pang, Ru Zheng, Feng Liu

    Abstract: The integration of topological concepts into electronic energy band theory has been a transformative development in condensed matter physics. Since then, this paradigm has broadened its reach, extending to a variety of physical systems, including open ones. In this study, we employ analogues of the generalized $n$-dimensional Su-Schrieffer-Heeger model, a cornerstone in understanding topological i… ▽ More

    Submitted 21 December, 2023; originally announced December 2023.

    Journal ref: Phys. Rev. B 109, 054312 (2024)

  43. arXiv:2311.15421  [pdf, other

    cs.CV cs.AI

    Wired Perspectives: Multi-View Wire Art Embraces Generative AI

    Authors: Zhiyu Qu, Lan Yang, Honggang Zhang, Tao Xiang, Kaiyue Pang, Yi-Zhe Song

    Abstract: Creating multi-view wire art (MVWA), a static 3D sculpture with diverse interpretations from different viewpoints, is a complex task even for skilled artists. In response, we present DreamWire, an AI system enabling everyone to craft MVWA easily. Users express their vision through text prompts or scribbles, freeing them from intricate 3D wire organisation. Our approach synergises 3D Bézier curves,… ▽ More

    Submitted 13 June, 2024; v1 submitted 26 November, 2023; originally announced November 2023.

    Comments: CVPR 2024

  44. arXiv:2311.04942  [pdf, other

    eess.IV cs.CV

    CSAM: A 2.5D Cross-Slice Attention Module for Anisotropic Volumetric Medical Image Segmentation

    Authors: Alex Ling Yu Hung, Haoxin Zheng, Kai Zhao, Xiaoxi Du, Kaifeng Pang, Qi Miao, Steven S. Raman, Demetri Terzopoulos, Kyunghyun Sung

    Abstract: A large portion of volumetric medical data, especially magnetic resonance imaging (MRI) data, is anisotropic, as the through-plane resolution is typically much lower than the in-plane resolution. Both 3D and purely 2D deep learning-based segmentation methods are deficient in dealing with such volumetric data since the performance of 3D methods suffers when confronting anisotropic data, and 2D meth… ▽ More

    Submitted 26 November, 2023; v1 submitted 7 November, 2023; originally announced November 2023.

  45. arXiv:2310.06851  [pdf, other

    cs.CV cs.AI cs.GR

    BodyFormer: Semantics-guided 3D Body Gesture Synthesis with Transformer

    Authors: Kunkun Pang, Dafei Qin, Yingruo Fan, Julian Habekost, Takaaki Shiratori, Junichi Yamagishi, Taku Komura

    Abstract: Automatic gesture synthesis from speech is a topic that has attracted researchers for applications in remote communication, video games and Metaverse. Learning the mapping between speech and 3D full-body gestures is difficult due to the stochastic nature of the problem and the lack of a rich cross-modal dataset that is needed for training. In this paper, we propose a novel transformer-based framew… ▽ More

    Submitted 6 September, 2023; originally announced October 2023.

    Comments: 12 pages, 13 figures

  46. arXiv:2307.11926  [pdf, other

    eess.IV cs.CV

    PartDiff: Image Super-resolution with Partial Diffusion Models

    Authors: Kai Zhao, Alex Ling Yu Hung, Kaifeng Pang, Haoxin Zheng, Kyunghyun Sung

    Abstract: Denoising diffusion probabilistic models (DDPMs) have achieved impressive performance on various image generation tasks, including image super-resolution. By learning to reverse the process of gradually diffusing the data distribution into Gaussian noise, DDPMs generate new data by iteratively denoising from random noise. Despite their impressive performance, diffusion-based generative models suff… ▽ More

    Submitted 21 July, 2023; originally announced July 2023.

  47. arXiv:2307.09310  [pdf, other

    physics.flu-dyn

    Symmetry-Breaking in Point-Heated Droplets

    Authors: Khang Ee Pang, Charles Cuvillier, Yutaku Kita, Lennon Ó Náraigh

    Abstract: We investigate theoretically the stability of thermo-capillary convection within a droplet when heated by a point source from below. To model the droplet, we use a mathematical model based on lubrication theory. We formulate a base-state droplet profile, and we examine its respect to small-amplitude perturbations in the azimuthal direction. Such linear stability analysis reveals that the base stat… ▽ More

    Submitted 18 July, 2023; originally announced July 2023.

    Comments: 15 figures

  48. arXiv:2306.01859  [pdf, other

    cs.CV cs.AI

    Spatially Resolved Gene Expression Prediction from H&E Histology Images via Bi-modal Contrastive Learning

    Authors: Ronald Xie, Kuan Pang, Sai W. Chung, Catia T. Perciani, Sonya A. MacParland, Bo Wang, Gary D. Bader

    Abstract: Histology imaging is an important tool in medical diagnosis and research, enabling the examination of tissue structure and composition at the microscopic level. Understanding the underlying molecular mechanisms of tissue architecture is critical in uncovering disease mechanisms and developing effective treatments. Gene expression profiling provides insight into the molecular processes underlying t… ▽ More

    Submitted 27 October, 2023; v1 submitted 2 June, 2023; originally announced June 2023.

  49. arXiv:2304.11744  [pdf, other

    cs.CV

    SketchXAI: A First Look at Explainability for Human Sketches

    Authors: Zhiyu Qu, Yulia Gryaditskaya, Ke Li, Kaiyue Pang, Tao Xiang, Yi-Zhe Song

    Abstract: This paper, for the very first time, introduces human sketches to the landscape of XAI (Explainable Artificial Intelligence). We argue that sketch as a ``human-centred'' data form, represents a natural interface to study explainability. We focus on cultivating sketch-specific explainability designs. This starts by identifying strokes as a unique building block that offers a degree of flexibility i… ▽ More

    Submitted 23 April, 2023; originally announced April 2023.

    Comments: CVPR 2023

  50. arXiv:2303.13110  [pdf, other

    eess.IV cs.CV

    OCELOT: Overlapped Cell on Tissue Dataset for Histopathology

    Authors: Jeongun Ryu, Aaron Valero Puche, JaeWoong Shin, Seonwook Park, Biagio Brattoli, Jinhee Lee, Wonkyung Jung, Soo Ick Cho, Kyunghyun Paeng, Chan-Young Ock, Donggeun Yoo, Sérgio Pereira

    Abstract: Cell detection is a fundamental task in computational pathology that can be used for extracting high-level medical information from whole-slide images. For accurate cell detection, pathologists often zoom out to understand the tissue-level structures and zoom in to classify cells based on their morphology and the surrounding context. However, there is a lack of efforts to reflect such behaviors by… ▽ More

    Submitted 23 March, 2023; v1 submitted 23 March, 2023; originally announced March 2023.

    Comments: Accepted for publication at CVPR'23