Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 111 results for author: Mai, J

.
  1. arXiv:2608.09491  [pdf, ps, other

    math.AP

    Global boundedness and stabilization for a three-component reaction-diffusion model with dual-dependent motility

    Authors: Hai-Yang Jin, Jingyi Mai

    Abstract: In this paper, we consider the initial-boundary value problem of a three-component reaction-diffusion system with dual-dependent motility, which depends on both the chemical concentration and the nutrient level. We systematically establish the global existence, boundedness, and asymptotic behavior of the classical solutions to the system with no-flux boundary conditions through classified discussi… ▽ More

    Submitted 10 August, 2026; originally announced August 2026.

  2. arXiv:2608.04780  [pdf, ps, other

    quant-ph

    Demonstrating advantages of dynamic quantum circuits on a hybrid superconducting qubit-cavity processor

    Authors: Hongbo Wu, Ling Hu, Jiasheng Mai, Munan Zhang, Libo Zhang, Yanyan Cai, Xiaowei Deng, Pan Zheng, Zhongchu Ni, Song Liu, Kun Fang, Dapeng Yu, Yuan Xu

    Abstract: Dynamic quantum circuits (DQCs) provide a hardware-efficient route to quantum computing by reducing physical-qubit overhead and compressing circuit topology through mid-circuit measurements, qubit reset and reuse, and classical feed-forward control. Here, we demonstrate the advantages of DQCs on a single hybrid superconducting qubit-cavity processor by implementing a hierarchy of algorithms with i… ▽ More

    Submitted 5 August, 2026; originally announced August 2026.

    Comments: 17 pages, 10 figures

  3. arXiv:2607.24560  [pdf, ps, other

    cs.CV cs.AI

    EgoPlay: Event-Triggered Video Editing for Egocentric Streams

    Authors: Jinjie Mai, Gordon Guocheng Qian, Willi Menapace, Arpit Sahni, Chaoyang Wang, Ashkan Mirzaei, Runjia Li, Sergey Tulyakov, Bernard Ghanem, Peter Wonka, Rameen Abdal

    Abstract: We introduce EgoPlay, an event-triggered video-to-video editor for egocentric streams, obtained by fine-tuning a pretrained V2V diffusion transformer on event-conditioned data built primarily from Ego4D. Given a monocular video and an event-triggered prompt of the form "when X happens, do Y," EgoPlay infers whether and when event X occurs, preserves pre-event frames, and applies edit Y only to the… ▽ More

    Submitted 27 July, 2026; originally announced July 2026.

    Comments: Accepted to SIGGRAPH Asia 2026 as a Conference Paper. Project page: https://egoplay2026.github.io/egoplay

  4. arXiv:2607.18147  [pdf, ps, other

    eess.SY cs.AI

    LLMs and Agentic AI Systems for Smart Grids: A Tutorial on Architectures and Applications

    Authors: Daniela Rojas, Abdulwahab Albassam, Aidan G. Leung, Jett Ngo, Ryan Luo, Peter R. Quawas, Junpyung Kim, Kangkai Liang, Mansi Nanavati, Jonathan Mai, Meng-Chi Tsai, Yun-Tong Tsai, Yize Chen, Yuanyuan Shi

    Abstract: Large language models (LLMs) and agentic AI systems have evolved from natural language tasks to using external tools to plan, retrieve, and act in technical domains. In smart grids, recent work applies agentic schemes to forecasting, optimization, and control, wrapping trusted solvers behind language interfaces and orchestrating multi-step workflows. The literature lacks a unified approach to desi… ▽ More

    Submitted 20 July, 2026; originally announced July 2026.

    Comments: 28 pages, 11 figures, 6 tables; plus supplementary material. Review/tutorial article

  5. arXiv:2607.17509  [pdf, ps, other

    physics.ins-det hep-ex

    Final assessment of radioactive impurities in the JUNO detector

    Authors: Thomas Adam, Fengpeng An, Costas Andreopoulos, Giuseppe Andronico, Nikolay Anfimov, Vito Antonelli, Tatiana Antoshkina, João Pedro Athayde Marcondes de André, Didier Auguste, Nikita Balashov, Andrea Barresi, Davide Basilico, Eric Baussan, Marco Beretta, Antonio Bergnoli, Nikita Bessonov, Daniel Bick, Lukas Bieger, Svetlana Biktemerova, Thilo Birkenfeld, Simon Blyth, Manuel Böhles, Anastasia Bolshakova, Mathieu Bongrand, Matteo Borghesi , et al. (549 additional authors not shown)

    Abstract: The Jiangmen Underground Neutrino Observatory (JUNO) collaboration has completed the construction of the 20,000-ton liquid scintillator detector and the associated muon veto detector system. To meet the physics objectives, the materials used in the detector must exhibit low radioactive contamination. The single-event rate in the fiducial volume (R $<$ 17.2 m) of the scintillator is required to be… ▽ More

    Submitted 19 July, 2026; originally announced July 2026.

  6. arXiv:2607.13427  [pdf, ps, other

    hep-ex

    A Low-energy Threshold and Multi-messenger Trigger System for the JUNO Experiment

    Authors: Thomas Adam, Fengpeng An, Costas Andreopoulos, Giuseppe Andronico, Nikolay Anfimov, Vito Antonelli, Tatiana Antoshkina, João Pedro Athayde Marcondes de André, Didier Auguste, Nikita Balashov, Andrea Barresi, Davide Basilico, Eric Baussan, Marco Beretta, Antonio Bergnoli, Nikita Bessonov, Daniel Bick, Lukas Bieger, Svetlana Biktemerova, Thilo Birkenfeld, Simon Blyth, Manuel Boehles, Anastasia Bolshakova, Mathieu Bongrand, Matteo Borghesi , et al. (543 additional authors not shown)

    Abstract: The Jiangmen Underground Neutrino Observatory (JUNO) is a 20-kiloton liquid scintillator neutrino detector, located 650 meters (1800 m.w.e.) underground in Jiangmen, Guangdong, China. JUNO is primarily designed for reactor neutrino measurements and has been taking data since 2025. With the largest mass of its kind and an excellent energy resolution, JUNO is a leading observatory for high-precision… ▽ More

    Submitted 15 July, 2026; originally announced July 2026.

    Comments: 29 pages, 19 figures, 3 tables

  7. arXiv:2607.06461  [pdf, ps, other

    eess.AS cs.CL cs.SD

    WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS

    Authors: Sihang Nie, Jinxin Ji, Xiaofen Xing, Deyi Tuo, Chengbin Jin, Jialong Mai, Xiangmin Xu

    Abstract: While recent Large Language Model (LLM)-based Text-to-Speech (TTS) systems have achieved remarkable naturalness, they predominantly rely on implicit end-to-end generation paradigms, resulting in coarse-grained control. In scenarios demanding precise stylistic interventions and strict temporal alignment, such as audiobook narration and video dubbing, the inability to explicitly manipulate word-leve… ▽ More

    Submitted 7 July, 2026; originally announced July 2026.

    Comments: 10 pages, 4 figures, 6 tables; Preprint

  8. arXiv:2607.04765  [pdf, ps, other

    cs.NE

    A Large-Scale Sparse Multiobjective Optimization Algorithm Based on Optimal Performance Scores

    Authors: Jia-Lin Mai, Min-Rong Chen, Guo-Qiang Zeng, Xiang Liu, Jian Weng

    Abstract: Large-scale sparse multiobjective optimization problems (LSSMOPs) involve a large number of decision variables and Pareto optimal solutions with only a few nonzero variables. However, as the number of decision variables grows, it becomes increasingly challenging to accurately identify the nonzero variables, and optimization performance is adversely affected. To address these issues, this paper pro… ▽ More

    Submitted 6 July, 2026; originally announced July 2026.

    Comments: 21 pages, 16 figures

  9. arXiv:2606.29531  [pdf, ps, other

    cs.CV cs.AI

    MotionAtlas: Detailed Region Captioning for Motion-Centric Videos

    Authors: Weisong Liu, Haochen Wang, Kuan Gao, Yuhao Wang, Yikang Zhou, Zhongwei Ren, Jacky Mai, Anna Wang, Yanwei Li, Jason Li, Zhaoxiang Zhang

    Abstract: We propose MotionAtlas, a system for detailed captioning of motion-centric videos, comprising (1) a dedicated human-annotated benchmark, (2) a scalable, high-quality pipeline to construct training samples, and (3) a family of powerful Video-MLLMs. Unlike conventional global motion captioning datasets, we focus on region-aware motion captioning: given a video and a spatiotemporal mask, the model ge… ▽ More

    Submitted 28 June, 2026; originally announced June 2026.

    Comments: Accepted to ECCV 2026. Project page: https://kagura-0001.github.io/projects/MotionAtlas

  10. arXiv:2606.19534  [pdf, ps, other

    cs.CV cs.AI cs.CL

    PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models

    Authors: Yueyi Sun, Yuhao Wang, Jason Li, Ye Tian, Tao Zhang, Jacky Mai, Yihan Wang, Haochen Wang, Jinbin Bai, Ling Yang, Yunhai Tong

    Abstract: Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks. However, most existing MLLMs rely on autoregressive generation, which limits their efficiency for perception tasks that require captioning multiple regions. In this work, we propose PerceptionDLM, a multimodal diffusion language model optimized for efficient parallel region perception. Built u… ▽ More

    Submitted 17 June, 2026; originally announced June 2026.

    Comments: Code available at https://github.com/MSALab-PKU/PerceptionDLM

  11. arXiv:2606.15888  [pdf, ps, other

    cs.SD cs.AI eess.AS

    NVMOS: Non-Verbal Vocalization Quality Assessment in Speech

    Authors: Jialong Mai, Jinxin Ji, Xiaofen Xing, Wencui Liu, Xiangmin Xu

    Abstract: Non-verbal vocalizations (NVs), such as laughter, sighs, and coughs, are important acoustic cues for emotion and intent. Existing speech quality assessment methods typically focus on overall naturalness, while non-verbal TTS evaluations mainly examine whether a target NV appears with the correct type and position. However, the perceptual quality of NV events themselves remains underexplored. To ad… ▽ More

    Submitted 14 June, 2026; originally announced June 2026.

    Comments: 6 pages. Code and model: https://github.com/yongaifadian1/NVMOS

  12. arXiv:2606.11704  [pdf, ps, other

    cs.RO

    Improving Human Diving Endurance with a Field-Deployable, Untethered Exoskeleton

    Authors: Zhihao Zhou, Zhenmeng Ju, Rui Yang, Chenxi Zhang, Zhihao Zhou, Ming Xu, Enhao Zheng, Dongjie Jiang, Lecheng Ruan, Jingeng Mai, Qining Wang

    Abstract: Human endurance in underwater locomotion is fundamentally restricted by high energetic demands to overcome drag and the finite supply of self-contained breathing gas. While exoskeleton technology can reduce the metabolic cost of humans in terrestrial locomotion, its potential to enhance human endurance during underwater diving remains entirely unexplored. Here, we present DiveMate, a field-deploya… ▽ More

    Submitted 10 June, 2026; originally announced June 2026.

  13. arXiv:2606.04101  [pdf, ps, other

    cs.DC cs.LG

    UltraEP: Unleash MoE Training and Inference on Rack-Scale Nodes with Near-Optimal Load Balancing

    Authors: Xinming Wei, Chao Jin, Tuo Dai, Yinmin Zhong, Shan Yu, Chengxu Yang, Bingyang Wu, Zili Zhang, Jing Mai, Qianchao Zhu, Zhouyang Li, Yuliang Liu, Guojie Luo

    Abstract: Large-scale expert parallelism (EP) is becoming pivotal for training and serving frontier MoE models, but it also amplifies device-level expert load imbalance into compute stragglers, token all-to-all bottlenecks, and activation-memory spikes. Existing balancers redistribute experts periodically based on historical load, which becomes unreliable for production deployments with non-stationary load… ▽ More

    Submitted 18 June, 2026; v1 submitted 2 June, 2026; originally announced June 2026.

  14. arXiv:2605.19639  [pdf, ps, other

    cs.CV

    Benchmarking and Evolving Reason-Reflect-Rectify for Reflective Visual Generation

    Authors: Junjie Wang, Xinghua Lou, Jason Li, Ye Tian, Keyu Chen, Yulin Li, Bin Kang, Jacky Mai, Yanwei Li, Zhuotao Tian, Liqiang Nie

    Abstract: Text-to-Image (T2I) models and Unified Multimodal Models (UMMs) have achieved remarkable progress in visual generation. However, their reliance on a single-pass generation paradigm limits their ability to handle complex prompts requiring iterative refinement. To enable multi-round Reflective Visual Generation (RVG), we formalize the Reason-Reflect-Rectify (R^3) loop as a core framework and introdu… ▽ More

    Submitted 19 May, 2026; originally announced May 2026.

  15. arXiv:2604.25835  [pdf, ps, other

    physics.ins-det hep-ex

    Embedded underwater front-end electronics for the 3-inch photomultipliers in the JUNO experiment

    Authors: Cédric Cerna, Miao He, Xiaoshan Jiang, Juan Pedro Ochoa-Ricoux, Frédéric Perrot, Angel Abusleme, Thomas Adam, Fengpeng An, Costas Andreopoulos, Giuseppe Andronico, João Pedro Athayde Marcondes de André, Nikolay Anfimov, Vito Antonelli, Tatiana Antoshkina, Didier Auguste, Nikita Balashov, Andrea Barresi, Davide Basilico, Eric Baussan, Marco Beretta, Antonio Bergnoli, Nikita Bessonov, Daniel Bick, Lukas Bieger, Svetlana Biktemerova , et al. (576 additional authors not shown)

    Abstract: The Jiangmen Underground Neutrino Observatory (JUNO) is a 20-kton liquid scintillator-based, low-radioactivity, multi-purpose neutrino detector located 693 meters (1800 m.w.e.) underground in the Guangdong province, China. To detect scintillation light produced in the target, the detector is equipped with 17,612 20-inch photomultipliers (PMTs), forming the Large PMT system (LPMT). In addition, 25,… ▽ More

    Submitted 1 June, 2026; v1 submitted 28 April, 2026; originally announced April 2026.

    Comments: Submitted to Nucl. Instrum. Methods Phys. Res. A

  16. arXiv:2604.21164  [pdf, ps, other

    cs.SD

    MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control

    Authors: Jialong Mai, Xiaofen Xing, Xiangmin Xu

    Abstract: Fine-grained local timing control is still absent from modern text-to-speech systems: existing approaches typically provide only utterance-level duration or global speaking-rate control, while precise token-level timing manipulation remains unavailable. To the best of our knowledge, MAGIC-TTS is the first TTS model with explicit local timing control over token-level content duration and pause. MAG… ▽ More

    Submitted 27 April, 2026; v1 submitted 22 April, 2026; originally announced April 2026.

    Comments: Release MAGIC-TTS code, pretrained models, and demo: https://github.com/yongaifadian1/MAGIC-TTS, https://huggingface.co/maimai11/MAGIC-TTS, https://yongaifadian1.github.io/MAGIC-TTS/

  17. arXiv:2602.03086  [pdf, ps, other

    cs.LG cs.CV

    Neural Predictor-Corrector: Solving Homotopy Problems with Reinforcement Learning

    Authors: Jiayao Mai, Bangyan Liao, Zhenjun Zhao, Yingping Zeng, Haoang Li, Javier Civera, Tailin Wu, Yi Zhou, Peidong Liu

    Abstract: The Homotopy paradigm, a general principle for solving challenging problems, appears across diverse domains such as robust optimization, global optimization, polynomial root-finding, and sampling. Practical solvers for these problems typically follow a predictor-corrector (PC) structure, but rely on hand-crafted heuristics for step sizes and iteration termination, which are often suboptimal and ta… ▽ More

    Submitted 2 February, 2026; originally announced February 2026.

  18. arXiv:2601.21616  [pdf, ps, other

    quant-ph

    A biased-erasure cavity qubit with hardware-efficient quantum error detection

    Authors: Jiasheng Mai, Qiyu Liu, Xiaowei Deng, Yanyan Cai, Zhongchu Ni, Libo Zhang, Ling Hu, Pan Zheng, Song Liu, Yuan Xu, Dapeng Yu

    Abstract: Erasure qubits are beneficial for quantum error correction due to their relaxed threshold requirements. While dual-rail erasure qubits have been demonstrated with a strong error hierarchy in circuit quantum electrodynamics, biased-erasure qubits -- where erasures originate predominantly from one logical basis state -- offer further advantages. Here, we realize a hardware-efficient biased-erasure q… ▽ More

    Submitted 29 January, 2026; originally announced January 2026.

    Comments: Main text: 10 pages, 4 figures; Supplementary material: 14 pages, 9 figures, 2 tables

  19. arXiv:2512.16920  [pdf, ps, other

    cs.CV cs.AI

    EasyV2V: A High-quality Instruction-based Video Editing Framework

    Authors: Jinjie Mai, Chaoyang Wang, Guocheng Gordon Qian, Willi Menapace, Sergey Tulyakov, Bernard Ghanem, Peter Wonka, Ashkan Mirzaei

    Abstract: While image editing has advanced rapidly, video editing remains less explored, facing challenges in consistency, control, and generalization. We study the design space of data, architecture, and control, and introduce \emph{EasyV2V}, a simple and effective framework for instruction-based video editing. On the data side, we compose existing experts with fast inverses to build diverse video pairs, l… ▽ More

    Submitted 18 December, 2025; originally announced December 2025.

    Comments: Project page: https://snap-research.github.io/easyv2v/

  20. arXiv:2511.14593  [pdf, ps, other

    hep-ex

    First measurement of reactor neutrino oscillations at JUNO

    Authors: Angel Abusleme, Thomas Adam, Kai Adamowicz, David Adey, Shakeel Ahmad, Rizwan Ahmed, Timo Ahola, Sebastiano Aiello, Fengpeng An, Guangpeng An, Costas Andreopoulos, Giuseppe Andronico, João Pedro Athayde Marcondes de André, Nikolay Anfimov, Vito Antonelli, Tatiana Antoshkina, Burin Asavapibhop, Didier Auguste, Margherita Buizza Avanzini, Andrej Babic, Jingzhi Bai, Weidong Bai, Nikita Balashov, Roberto Barbera, Andrea Barresi , et al. (1114 additional authors not shown)

    Abstract: Neutrino oscillations, a quantum effect manifesting at macroscopic scales, are governed by lepton flavor mixing angles and neutrino mass-squared differences that are fundamental parameters of particle physics, representing phenomena beyond the Standard Model. Precision measurements of these parameters are essential for testing the completeness of the three-flavor framework, determining the mass or… ▽ More

    Submitted 18 November, 2025; originally announced November 2025.

    Comments: 30 pages, 11 figures

  21. arXiv:2511.14590  [pdf, ps, other

    hep-ex physics.ins-det

    Initial performance results of the JUNO detector

    Authors: Angel Abusleme, Thomas Adam, Kai Adamowicz, David Adey, Shakeel Ahmad, Rizwan Ahmed, Timo Ahola, Sebastiano Aiello, Fengpeng An, Guangpeng An, Costas Andreopoulos, Giuseppe Andronico, João Pedro Athayde Marcondes de André, Nikolay Anfimov, Vito Antonelli, Tatiana Antoshkina, Burin Asavapibhop, Didier Auguste, Margherita Buizza Avanzini, Andrej Babic, Jingzhi Bai, Weidong Bai, Nikita Balashov, Roberto Barbera, Andrea Barresi , et al. (1114 additional authors not shown)

    Abstract: The Jiangmen Underground Neutrino Observatory (JUNO) started physics data taking on 26 August 2025. JUNO consists of a 20-kton liquid scintillator central detector, surrounded by a 35 kton water pool serving as a Cherenkov veto, and almost 1000 m$^2$ of plastic scintillator veto on top. The detector is located in a shallow underground laboratory with an overburden of 1800 m.w.e. This paper present… ▽ More

    Submitted 18 November, 2025; originally announced November 2025.

    Comments: 38 pages, 23 figures

  22. arXiv:2511.07227  [pdf, ps, other

    hep-ex physics.geo-ph

    Prospects for geoneutrino detection with JUNO

    Authors: Thomas Adam, Shakeel Ahmad, Rizwan Ahmed, Fengpeng An, João Pedro Athayde Marcondes de André, Costas Andreopoulos, Giuseppe Andronico, Nikolay Anfimov, Vito Antonelli, Tatiana Antoshkina, Didier Auguste, Marcel Büchner, Weidong Bai, Nikita Balashov, Andrea Barresi, Davide Basilico, Eric Baussan, Marco Beretta, Antonio Bergnoli, Nikita Bessonov, Daniel Bick, Lukas Bieger, Svetlana Biktemerova, Thilo Birkenfeld, Simon Blyth , et al. (605 additional authors not shown)

    Abstract: Geoneutrinos, which are antineutrinos emitted during the decay of long-lived radioactive elements inside Earth, serve as a unique tool for studying the composition and heat budget of our planet. The Jiangmen Underground Neutrino Observatory (JUNO) experiment in China, which has recently completed construction, is expected to collect a sample comparable in size to the entire existing world geoneutr… ▽ More

    Submitted 10 November, 2025; originally announced November 2025.

    Comments: 32 pages, with 13 figures and 5 tables

  23. arXiv:2511.07133  [pdf

    physics.optics

    Scattering Induced Mode Chirality in Ring Resonators

    Authors: Haochen Yan, Xu Guo, Arghadeep Pal, Xiaoyuan Huang, Alekhya Ghosh, Lewis Hill, Shuangyou Zhang, Nivedita Vishnukumar, Toby Bi, Masoud Kheyri, Jianming Mai, Hao Zhang, Yaojing Zhang, Jolly Xavier, Haihua Fan, Kok Wai Cheah, Peter Littlewood, Pascal DelHaye

    Abstract: Non-Hermitian physics can be used to break time reversal symmetry and is important for interactions in a wide range of systems, from active matter and neural networks to metamaterials and non-equilibrium thermodynamics. In integrated photonic devices, non-Hermitian physics can be used for direction-dependent light propagation, reconfigurable light paths, selective energy localization and optical i… ▽ More

    Submitted 10 November, 2025; originally announced November 2025.

  24. arXiv:2510.06616  [pdf, ps, other

    physics.ins-det hep-ex

    Design, waterproofing, and mass production of the 3-inch PMT frontend system of JUNO

    Authors: Jilei Xu, Miao He, Cédric Cerna, Yongbo Huang, Thomas Adam, Shakeel Ahmad, Rizwan Ahmed, Fengpeng An, Costas Andreopoulos, Giuseppe Andronico, João Pedro Athayde Marcondes de André, Nikolay Anfimov, Vito Antonelli, Tatiana Antoshkina, Didier Auguste, Weidong Bai, Nikita Balashov, Andrea Barresi, Davide Basilico, Eric Baussan, Marco Beretta, Antonio Bergnoli, Nikita Bessonov, Daniel Bick, Lukas Bieger , et al. (609 additional authors not shown)

    Abstract: Over 25,600 3-inch photomultiplier tubes (PMTs) have been instrumented for the central detector of the Jiangmen Underground Neutrino Observatory. Each PMT is equipped with a high-voltage divider and a frontend cable with waterproof sealing. Groups of sixteen PMTs are connected to the underwater frontend readout electronics via specialized multi-channel waterproof connectors. This paper outlines th… ▽ More

    Submitted 22 January, 2026; v1 submitted 7 October, 2025; originally announced October 2025.

  25. arXiv:2509.26042  [pdf, ps, other

    quant-ph

    Autonomous quantum error correction beyond break-even and its metrological application

    Authors: Zhongchu Ni, Ling Hu, Yanyan Cai, Libo Zhang, Jiasheng Mai, Xiaowei Deng, Pan Zheng, Song Liu, Shi-Biao Zheng, Yuan Xu, Dapeng Yu

    Abstract: The ability to extend the lifetime of a logical qubit beyond that of the best physical qubit available within the same system, i.e., the break-even point, is a prerequisite for building practical quantum computers. So far, this point has been exceeded through active quantum error correction (QEC) protocols, where a logical error is corrected by measuring its syndrome and then performing an adaptiv… ▽ More

    Submitted 30 September, 2025; originally announced September 2025.

    Comments: Main text: 10 pages, 4 figures; Supplementary material: 18 pages, 13 figures, 2 tables

  26. arXiv:2509.20864  [pdf, ps, other

    cs.CV

    SD-RetinaNet: Topologically Constrained Semi-Supervised Retinal Lesion and Layer Segmentation in OCT

    Authors: Botond Fazekas, Guilherme Aresta, Philipp Seeböck, Julia Mai, Ursula Schmidt-Erfurth, Hrvoje Bogunović

    Abstract: Optical coherence tomography (OCT) is widely used for diagnosing and monitoring retinal diseases, such as age-related macular degeneration (AMD). The segmentation of biomarkers such as layers and lesions is essential for patient diagnosis and follow-up. Recently, semi-supervised learning has shown promise in improving retinal segmentation performance. However, existing methods often produce anatom… ▽ More

    Submitted 25 September, 2025; originally announced September 2025.

  27. arXiv:2509.18196  [pdf, ps, other

    cs.SD cs.AI eess.AS

    MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech

    Authors: Jialong Mai, Jinxin Ji, Xiaofen Xing, Chen Yang, Weidong Chen, Jingyuan Xing, Xiangmin Xu

    Abstract: Mainstream Automatic Speech Recognition (ASR) systems excel at transcribing lexical content, but largely fail to recognize nonverbal vocalizations (NVs) embedded in speech, such as sighs, laughs, and coughs. This capability is important for a comprehensive understanding of human communication, as NVs convey crucial emotional and intentional cues. Progress in NV-aware ASR has been hindered by the l… ▽ More

    Submitted 24 September, 2025; v1 submitted 19 September, 2025; originally announced September 2025.

    Comments: Official dataset available at: https://github.com/yongaifadian1/MNV-17. Submitted to ICASSP 2026

  28. arXiv:2508.12564  [pdf, ps, other

    cs.RO cs.CV

    Temporal and Rotational Calibration for Event-Centric Multi-Sensor Systems

    Authors: Jiayao Mai, Xiuyuan Lu, Kuan Dai, Shaojie Shen, Yi Zhou

    Abstract: Event cameras generate asynchronous signals in response to pixel-level brightness changes, offering a sensing paradigm with theoretically microsecond-scale latency that can significantly enhance the performance of multi-sensor systems. Extrinsic calibration is a critical prerequisite for effective sensor fusion; however, the configuration that involves event cameras remains an understudied topic.… ▽ More

    Submitted 17 August, 2025; originally announced August 2025.

    Comments: 8 pages, 5 figures

    ACM Class: I.2.9

  29. arXiv:2508.04141  [pdf, ps, other

    eess.AS cs.SD

    Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech

    Authors: Jingyuan Xing, Zhipeng Li, Jialong Mai, Xiaofen Xing, Xiangmin Xu

    Abstract: Advances in speech representation and large language models have enhanced zero-shot text-to-speech (TTS) performance. However, existing zero-shot TTS models face challenges in capturing the complex correlations between acoustic and semantic features, resulting in a lack of expressiveness and similarity. The primary reason lies in the complex relationship between semantic and acoustic features, whi… ▽ More

    Submitted 28 August, 2025; v1 submitted 6 August, 2025; originally announced August 2025.

    Comments: Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP)

  30. arXiv:2507.23538  [pdf, ps, other

    quant-ph hep-ex hep-ph

    Quantum-Enhanced Dark Matter Search Using Cat States

    Authors: Pan Zheng, Yanyan Cai, Bin Xu, Shengcheng Wen, Libo Zhang, Zhongchu Ni, Jiasheng Mai, Yanjie Zeng, Lin Lin, Ling Hu, Xiaowei Deng, Song Liu, Jing Shu, Yuan Xu, Dapeng Yu

    Abstract: Quantum metrology has recently emerged as a powerful approach for dark matter (DM) searches, particularly using nonclassical bosonic states in microwave cavities that are sensitive to weak signals. Nonclassical cat states - macroscopic superpositions of coherent states featuring sub-Planck interference structures - offer promising advantages for high-precision measurements. However, their practica… ▽ More

    Submitted 8 May, 2026; v1 submitted 31 July, 2025; originally announced July 2025.

    Comments: Main text: 8 pages, 4 figures; Supplement: 15 pages, 13 figures, 5 tables

    Journal ref: Phys. Rev. Lett. 136, 171002 (2026)

  31. arXiv:2507.21138  [pdf, ps, other

    cs.CL cs.AI cs.LG cs.SD eess.AS

    TTS-1 Technical Report

    Authors: Oleg Atamanenko, Anna Chalova, Joseph Coombes, Nikki Cope, Phillip Dang, Zhifeng Deng, Jimmy Du, Michael Ermolenko, Feifan Fan, Yufei Feng, Cheryl Fichter, Pavel Filimonov, Louis Fischer, Kylan Gibbs, Valeria Gusarova, Pavel Karpik, Andreas Assad Kottner, Ian Lee, Oliver Louie, Jasmine Mai, Mikhail Mamontov, Suri Mao, Nurullah Morshed, Igor Poletaev, Florin Radu , et al. (7 additional authors not shown)

    Abstract: We introduce Inworld TTS-1, a set of two Transformer-based autoregressive text-to-speech (TTS) models. Our largest model, TTS-1-Max, has 8.8B parameters and is designed for utmost quality and expressiveness in demanding applications. TTS-1 is our most efficient model, with 1.6B parameters, built for real-time speech synthesis and on-device use cases. By scaling train-time compute and applying a se… ▽ More

    Submitted 22 July, 2025; originally announced July 2025.

    Comments: 20 pages, 10 figures. For associated modeling and training code, see https://github.com/inworld-ai/tts

  32. arXiv:2507.19182  [pdf, ps, other

    cs.AI cs.LO

    Faster Lifting for Ordered Domains with Predecessor Relations

    Authors: Kuncheng Zou, Jiahao Mai, Yonggang Zhang, Yuyi Wang, Ondřej Kuželka, Yuanhong Wang, Yi Chang

    Abstract: We investigate lifted inference on ordered domains with predecessor relations, where the elements of the domain respect a total (cyclic) order, and every element has a distinct (clockwise) predecessor. Previous work has explored this problem through weighted first-order model counting (WFOMC), which computes the weighted sum of models for a given first-order logic sentence over a finite domain. In… ▽ More

    Submitted 25 July, 2025; originally announced July 2025.

  33. arXiv:2507.11296  [pdf, ps, other

    cs.RO

    Diffusion-Based Imaginative Coordination for Bimanual Manipulation

    Authors: Huilin Xu, Jian Ding, Jiakun Xu, Ruixiang Wang, Jun Chen, Jinjie Mai, Yanwei Fu, Bernard Ghanem, Feng Xu, Mohamed Elhoseiny

    Abstract: Bimanual manipulation is crucial in robotics, enabling complex tasks in industrial automation and household services. However, it poses significant challenges due to the high-dimensional action space and intricate coordination requirements. While video prediction has been recently studied for representation learning and control, leveraging its ability to capture rich dynamic and behavioral informa… ▽ More

    Submitted 15 July, 2025; originally announced July 2025.

    Comments: 15 pages, including 10 figures and 16 tables. Accepted at ICCV 2025

  34. arXiv:2507.09076  [pdf, ps, other

    cs.CL cs.AI

    Dynamic Parameter Memory: Temporary LoRA-Enhanced LLM for Long-Sequence Emotion Recognition in Conversation

    Authors: Jialong Mai, Xiaofen Xing, Yawei Li, Weidong Chen, Zhipeng Li, Jingyuan Xing, Xiangmin Xu

    Abstract: Recent research has focused on applying speech large language model (SLLM) to improve speech emotion recognition (SER). However, the inherently high frame rate in speech modality severely limits the signal processing and understanding capabilities of SLLM. For example, a SLLM with a 4K context window can only process 80 seconds of audio at 50Hz feature sampling rate before reaching its capacity li… ▽ More

    Submitted 24 September, 2025; v1 submitted 11 July, 2025; originally announced July 2025.

    Comments: submitted to ICLR 2026

    MSC Class: 68T50 ACM Class: I.2.7; H.5.2

  35. arXiv:2507.05847   

    cs.NE

    A Universal Framework for Large-Scale Multi-Objective Optimization Based on Particle Drift and Diffusion

    Authors: Jia-Cheng Li, Min-Rong Chen, Guo-Qiang Zeng, Jian Weng, Man Wang, Jia-Lin Mai

    Abstract: Large-scale multi-objective optimization poses challenges to existing evolutionary algorithms in maintaining the performances of convergence and diversity because of high dimensional decision variables. Inspired by the motion of particles in physics, we propose a universal framework for large-scale multi-objective optimization based on particle drift and diffusion to solve these challenges in this… ▽ More

    Submitted 18 September, 2025; v1 submitted 8 July, 2025; originally announced July 2025.

    Comments: There are several details related to operators are imprecise.To uphold the principle of accuracy, we have decided to retract the article for now

  36. arXiv:2506.02658  [pdf, ps, other

    cs.SE

    Computational Thinking Reasoning in Large Language Models

    Authors: Kechi Zhang, Ge Li, Jia Li, Huangzhao Zhang, Jingjing Xu, Hao Zhu, Lecheng Wang, Jia Li, Yihong Dong, Jing Mai, Bin Gu, Zhi Jin

    Abstract: While large language models (LLMs) have demonstrated remarkable reasoning capabilities, they often struggle with complex tasks that require specific thinking paradigms, such as divide-and-conquer and procedural deduction, \etc Previous researches integrate external, reliable tools to alleviate logical inconsistencies and hallucinations in LLMs' problem-solving processes. However, we argue that the… ▽ More

    Submitted 3 June, 2025; v1 submitted 3 June, 2025; originally announced June 2025.

  37. arXiv:2505.07916  [pdf, ps, other

    eess.AS cs.SD

    MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

    Authors: Bowen Zhang, Congchao Guo, Geng Yang, Hang Yu, Haozhe Zhang, Heidi Lei, Jialong Mai, Junjie Yan, Kaiyue Yang, Mingqi Yang, Peikai Huang, Ruiyang Jin, Sitan Jiang, Weihua Cheng, Yawei Li, Yichen Xiao, Yiying Zhou, Yongmao Zhang, Yuan Lu, Yucen He

    Abstract: We introduce MiniMax-Speech, an autoregressive Transformer-based Text-to-Speech (TTS) model that generates high-quality speech. A key innovation is our learnable speaker encoder, which extracts timbre features from a reference audio without requiring its transcription. This enables MiniMax-Speech to produce highly expressive speech with timbre consistent with the reference in a zero-shot manner, w… ▽ More

    Submitted 12 May, 2025; originally announced May 2025.

  38. arXiv:2505.01545  [pdf, other

    q-bio.NC

    Burstiness and interpersonal foraging between human infants and caregivers in the vocal domain

    Authors: VPS Ritwika, Sara Schneider, Lukas D. Lopez, Jeffrey Mai, Ajay Gopinathan, Christopher T. Kello, Anne S. Warlaumont

    Abstract: Vocal responses from caregivers are believed to promote more frequent and more advanced infant vocalizations. However, studies that examine this relationship typically do not account for the fact that infant and adult vocalizations are distributed in hierarchical clusters over the course of the day. These bursts and lulls create a challenge for accurately detecting the effects of adult input at im… ▽ More

    Submitted 20 May, 2025; v1 submitted 2 May, 2025; originally announced May 2025.

    Comments: 48 pages total containing main text (Figures 1-4, 17 pages including references) and supplemental pdf (Appendices 1-7, including Figures 1-31 and Tables 1-13)

  39. arXiv:2503.21082  [pdf, other

    cs.CV

    Can Video Diffusion Model Reconstruct 4D Geometry?

    Authors: Jinjie Mai, Wenxuan Zhu, Haozhe Liu, Bing Li, Cheng Zheng, Jürgen Schmidhuber, Bernard Ghanem

    Abstract: Reconstructing dynamic 3D scenes (i.e., 4D geometry) from monocular video is an important yet challenging problem. Conventional multiview geometry-based approaches often struggle with dynamic motion, whereas recent learning-based methods either require specialized 4D representation or sophisticated optimization. In this paper, we present Sora3R, a novel framework that taps into the rich spatiotemp… ▽ More

    Submitted 26 March, 2025; originally announced March 2025.

  40. arXiv:2503.17827  [pdf, other

    cs.CV

    4D-Bench: Benchmarking Multi-modal Large Language Models for 4D Object Understanding

    Authors: Wenxuan Zhu, Bing Li, Cheng Zheng, Jinjie Mai, Jun Chen, Letian Jiang, Abdullah Hamdi, Sara Rojas Martinez, Chia-Wen Lin, Mohamed Elhoseiny, Bernard Ghanem

    Abstract: Multimodal Large Language Models (MLLMs) have demonstrated impressive 2D image/video understanding capabilities. However, there are no publicly standardized benchmarks to assess the abilities of MLLMs in understanding the 4D objects (3D objects with temporal evolution over time). In this paper, we introduce 4D-Bench, the first benchmark to evaluate the capabilities of MLLMs in 4D object understand… ▽ More

    Submitted 22 March, 2025; originally announced March 2025.

  41. Quantum squeezing amplification with a weak Kerr nonlinear oscillator

    Authors: Yanyan Cai, Xiaowei Deng, Libo Zhang, Zhongchu Ni, Jiasheng Mai, Peihao Huang, Pan Zheng, Ling Hu, Song Liu, Yuan Xu, Dapeng Yu

    Abstract: Quantum squeezed states, with reduced quantum noise, have been widely utilized in quantum sensing and quantum error correction applications. However, generating and manipulating these nonclassical states with a large squeezing degree typically requires strong nonlinearity, which inevitably induces additional decoherence that diminishes the overall performance. Here, we demonstrate the generation a… ▽ More

    Submitted 11 March, 2025; originally announced March 2025.

    Comments: Main text: 8 pages, 4 figures; Supplementary material: 9 pages, 11 figures, 2 tables

    Journal ref: Nat. Commun. 17, 970 (2026)

  42. arXiv:2503.00968  [pdf, other

    physics.ins-det hep-ex

    Simulation of the Background from $^{13}$C$(α, n)^{16}$O Reaction in the JUNO Scintillator

    Authors: JUNO Collaboration, Thomas Adam, Kai Adamowicz, Shakeel Ahmad, Rizwan Ahmed, Sebastiano Aiello, Fengpeng An, Costas Andreopoulos, Giuseppe Andronico, Nikolay Anfimov, Vito Antonelli, Tatiana Antoshkina, João Pedro Athayde Marcondes de André, Didier Auguste, Weidong Bai, Nikita Balashov, Andrea Barresi, Davide Basilico, Eric Baussan, Marco Beretta, Antonio Bergnoli, Nikita Bessonov, Daniel Bick, Lukas Bieger, Svetlana Biktemerova , et al. (608 additional authors not shown)

    Abstract: Large-scale organic liquid scintillator detectors are highly efficient in the detection of MeV-scale electron antineutrinos. These signal events can be detected through inverse beta decay on protons, which produce a positron accompanied by a neutron. A noteworthy background for antineutrinos coming from nuclear power reactors and from the depths of the Earth (geoneutrinos) is generated by ($α, n$)… ▽ More

    Submitted 2 May, 2025; v1 submitted 2 March, 2025; originally announced March 2025.

    Comments: 25 pages, 14 figures, 4 tables

  43. arXiv:2502.16789  [pdf, ps, other

    cs.CE cs.AI

    AlphaAgent: LLM-Driven Alpha Mining with Regularized Exploration to Counteract Alpha Decay

    Authors: Ziyi Tang, Zechuan Chen, Jiarui Yang, Jiayao Mai, Yongsen Zheng, Keze Wang, Jinrui Chen, Liang Lin

    Abstract: Alpha mining, a critical component in quantitative investment, focuses on discovering predictive signals for future asset returns in increasingly complex financial markets. However, the pervasive issue of alpha decay, where factors lose their predictive power over time, poses a significant challenge for alpha mining. Traditional methods like genetic programming face rapid alpha decay from overfitt… ▽ More

    Submitted 8 June, 2025; v1 submitted 23 February, 2025; originally announced February 2025.

    Comments: 9 pages; Code is available at: https://github.com/RndmVariableQ/AlphaAgent

  44. arXiv:2412.00535  [pdf, other

    cs.AI cs.SE

    FullStack Bench: Evaluating LLMs as Full Stack Coders

    Authors: Bytedance-Seed-Foundation-Code-Team, :, Yao Cheng, Jianfeng Chen, Jie Chen, Li Chen, Liyu Chen, Wentao Chen, Zhengyu Chen, Shijie Geng, Aoyan Li, Bo Li, Bowen Li, Linyi Li, Boyi Liu, Jiaheng Liu, Kaibo Liu, Qi Liu, Shukai Liu, Siyao Liu, Tianyi Liu, Tingkai Liu, Yongfei Liu, Rui Long, Jing Mai , et al. (31 additional authors not shown)

    Abstract: As the capabilities of code large language models (LLMs) continue to expand, their applications across diverse code intelligence domains are rapidly increasing. However, most existing datasets only evaluate limited application domains. To address this gap, we have developed a comprehensive code evaluation dataset FullStack Bench focusing on full-stack programming, which encompasses a wide range of… ▽ More

    Submitted 12 May, 2025; v1 submitted 30 November, 2024; originally announced December 2024.

    Comments: 26 pages

  45. arXiv:2408.10739  [pdf, other

    cs.CV

    TrackNeRF: Bundle Adjusting NeRF from Sparse and Noisy Views via Feature Tracks

    Authors: Jinjie Mai, Wenxuan Zhu, Sara Rojas, Jesus Zarzar, Abdullah Hamdi, Guocheng Qian, Bing Li, Silvio Giancola, Bernard Ghanem

    Abstract: Neural radiance fields (NeRFs) generally require many images with accurate poses for accurate novel view synthesis, which does not reflect realistic setups where views can be sparse and poses can be noisy. Previous solutions for learning NeRFs with sparse views and noisy poses only consider local geometry consistency with pairs of views. Closely following \textit{bundle adjustment} in Structure-fr… ▽ More

    Submitted 20 August, 2024; originally announced August 2024.

    Comments: ECCV 2024 (supplemental pages included)

  46. arXiv:2407.08410  [pdf, other

    cs.AI

    Specialized curricula for training vision-language models in retinal image analysis

    Authors: Robbie Holland, Thomas R. P. Taylor, Christopher Holmes, Sophie Riedl, Julia Mai, Maria Patsiamanidi, Dimitra Mitsopoulou, Paul Hager, Philip Müller, Hendrik P. N. Scholl, Hrvoje Bogunović, Ursula Schmidt-Erfurth, Daniel Rueckert, Sobha Sivaprasad, Andrew J. Lotery, Martin J. Menten

    Abstract: Clinicians spend a significant amount of time reviewing medical images and transcribing their findings regarding patient diagnosis, referral and treatment in text form. Vision-language models (VLMs), which automatically interpret images and summarize their findings as text, have enormous potential to alleviate clinical workloads and increase patient access to high-quality medical care. While found… ▽ More

    Submitted 24 February, 2025; v1 submitted 11 July, 2024; originally announced July 2024.

    Comments: Under review at npj Digital Medicine

  47. arXiv:2407.08023  [pdf, other

    cs.CV

    Hybrid Structure-from-Motion and Camera Relocalization for Enhanced Egocentric Localization

    Authors: Jinjie Mai, Abdullah Hamdi, Silvio Giancola, Chen Zhao, Bernard Ghanem

    Abstract: We built our pipeline EgoLoc-v1, mainly inspired by EgoLoc. We propose a model ensemble strategy to improve the camera pose estimation part of the VQ3D task, which has been proven to be essential in previous work. The core idea is not only to do SfM for egocentric videos but also to do 2D-3D matching between existing 3D scans and 2D video frames. In this way, we have a hybrid SfM and camera reloca… ▽ More

    Submitted 10 July, 2024; originally announced July 2024.

    Comments: 1st place winner of the 2024 Ego4D-Ego-Exo4D Challenge in VQ3D

  48. arXiv:2407.06890  [pdf, ps, other

    math.DS

    $(ω, α, n)$-sensitivity and limit sets of zero entropy homeomorphisms on the square

    Authors: Jiehua Mai, Enhui Shi, Kesong Yan, Fanping Zeng

    Abstract: For a homeomorphism $f$ of a compact metric space $X$ and a positive integer $n\geq 2$, we introduce the notion of $(ω, α, n)$-sensitivity of $f$, which describes such a kind of chaos: there is some $c>0$ such that for any $x\in X$ and any open neighborhood $U$ of $x$, there are points $\{x_i\}_{i=1}^n$ and $\{y_i\}_{i=1}^n$ in $U$ such that both the collection of $ω$-limit sets $ω(x_i, f)$ and th… ▽ More

    Submitted 9 July, 2024; originally announced July 2024.

  49. arXiv:2406.17243  [pdf, ps, other

    math.DS

    A new construction of counterexamples to the bounded orbit conjecture

    Authors: Jiehua Mai, Enhui Shi, Kesong Yan, Fanping Zeng

    Abstract: The bounded orbit conjecture says that every homeomorphism on the plane with each of its orbits being bounded must have a fixed point. Brouwer's translation theorem asserts that the conjecture is true for orientation preserving homeomorphisms, but Boyles' counterexample shows that it is false for the orientation reversing case. In this paper, we give a more comprehensible construction of counterex… ▽ More

    Submitted 9 April, 2025; v1 submitted 24 June, 2024; originally announced June 2024.

  50. arXiv:2406.08659  [pdf, other

    cs.CV

    Vivid-ZOO: Multi-View Video Generation with Diffusion Model

    Authors: Bing Li, Cheng Zheng, Wenxuan Zhu, Jinjie Mai, Biao Zhang, Peter Wonka, Bernard Ghanem

    Abstract: While diffusion models have shown impressive performance in 2D image/video generation, diffusion-based Text-to-Multi-view-Video (T2MVid) generation remains underexplored. The new challenges posed by T2MVid generation lie in the lack of massive captioned multi-view videos and the complexity of modeling such multi-dimensional distribution. To this end, we propose a novel diffusion-based pipeline tha… ▽ More

    Submitted 12 June, 2024; originally announced June 2024.

    Comments: Our project page is at https://hi-zhengcheng.github.io/vividzoo/