Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–6 of 6 results for author: Tuo, Y

Searching in archive cs. Search in all archives.
.
  1. arXiv:2604.19858  [pdf, ps, other

    cs.CV

    Wan-Image: Pushing the Boundaries of Generative Visual Intelligence

    Authors: Chaojie Mao, Chen-Wei Xie, Chongyang Zhong, Haoyou Deng, Jiaxing Zhao, Jie Xiao, Jinbo Xing, Jingfeng Zhang, Jingren Zhou, Jingyi Zhang, Jun Dan, Kai Zhu, Kang Zhao, Keyu Yan, Minghui Chen, Pandeng Li, Shuangle Chen, Tong Shen, Yu Liu, Yue Jiang, Yulin Pan, Yuxiang Tuo, Zeyinzi Jiang, Zhen Han, Ang Wang , et al. (33 additional authors not shown)

    Abstract: We present Wan-Image, a unified visual generation system explicitly engineered to paradigm-shift image generation models from casual synthesizers into professional-grade productivity tools. While contemporary diffusion models excel at aesthetic generation, they frequently encounter critical bottlenecks in rigorous design workflows that demand absolute controllability, complex typography rendering,… ▽ More

    Submitted 23 April, 2026; v1 submitted 21 April, 2026; originally announced April 2026.

  2. arXiv:2603.25706  [pdf, ps, other

    cs.CV

    Wan-Weaver: Interleaved Multi-modal Generation via Decoupled Training

    Authors: Jinbo Xing, Zeyinzi Jiang, Yuxiang Tuo, Chaojie Mao, Xiaotang Gai, Xi Chen, Jingfeng Zhang, Yulin Pan, Zhen Han, Jie Xiao, Keyu Yan, Chenwei Xie, Chongyang Zhong, Kai Zhu, Tong Shen, Lianghua Huang, Yu Liu, Yujiu Yang

    Abstract: Recent unified models have made unprecedented progress in both understanding and generation. However, while most of them accept multi-modal inputs, they typically produce only single-modality outputs. This challenge of producing interleaved content is mainly due to training data scarcity and the difficulty of modeling long-range cross-modal context. To address this issue, we decompose interleaved… ▽ More

    Submitted 29 March, 2026; v1 submitted 26 March, 2026; originally announced March 2026.

    Comments: CVPR 2026 Camera-ready, Webpage: https://doubiiu.github.io/projects/WanWeaver

  3. arXiv:2501.09503  [pdf, other

    cs.CV

    AnyStory: Towards Unified Single and Multiple Subject Personalization in Text-to-Image Generation

    Authors: Junjie He, Yuxiang Tuo, Binghui Chen, Chongyang Zhong, Yifeng Geng, Liefeng Bo

    Abstract: Recently, large-scale generative models have demonstrated outstanding text-to-image generation capabilities. However, generating high-fidelity personalized images with specific subjects still presents challenges, especially in cases involving multiple subjects. In this paper, we propose AnyStory, a unified approach for personalized subject generation. AnyStory not only achieves high-fidelity perso… ▽ More

    Submitted 1 May, 2025; v1 submitted 16 January, 2025; originally announced January 2025.

    Comments: Tech report; Project page: https://aigcdesigngroup.github.io/AnyStory/

  4. arXiv:2411.15245  [pdf, other

    cs.CV

    AnyText2: Visual Text Generation and Editing With Customizable Attributes

    Authors: Yuxiang Tuo, Yifeng Geng, Liefeng Bo

    Abstract: As the text-to-image (T2I) domain progresses, generating text that seamlessly integrates with visual content has garnered significant attention. However, even with accurate text generation, the inability to control font and color can greatly limit certain applications, and this issue remains insufficiently addressed. This paper introduces AnyText2, a novel method that enables precise control over… ▽ More

    Submitted 21 November, 2024; originally announced November 2024.

  5. arXiv:2311.03054  [pdf, other

    cs.CV

    AnyText: Multilingual Visual Text Generation And Editing

    Authors: Yuxiang Tuo, Wangmeng Xiang, Jun-Yan He, Yifeng Geng, Xuansong Xie

    Abstract: Diffusion model based Text-to-Image has achieved impressive achievements recently. Although current technology for synthesizing images is highly advanced and capable of generating images with high fidelity, it is still possible to give the show away when focusing on the text area in the generated image. To address this issue, we introduce AnyText, a diffusion-based multilingual visual text generat… ▽ More

    Submitted 21 February, 2024; v1 submitted 6 November, 2023; originally announced November 2023.

  6. arXiv:2310.05227  [pdf, other

    cs.LG cs.AI physics.flu-dyn

    Physics-aware Machine Learning Revolutionizes Scientific Paradigm for Machine Learning and Process-based Hydrology

    Authors: Qingsong Xu, Yilei Shi, Jonathan Bamber, Ye Tuo, Ralf Ludwig, Xiao Xiang Zhu

    Abstract: Accurate hydrological understanding and water cycle prediction are crucial for addressing scientific and societal challenges associated with the management of water resources, particularly under the dynamic influence of anthropogenic climate change. Existing reviews predominantly concentrate on the development of machine learning (ML) in this field, yet there is a clear distinction between hydrolo… ▽ More

    Submitted 12 July, 2024; v1 submitted 8 October, 2023; originally announced October 2023.

    Comments: 44 pages, 6 figures