Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–19 of 19 results for author: Watanabe, A

Searching in archive cs. Search in all archives.
.
  1. arXiv:2605.05165  [pdf, ps, other

    cs.IR

    Interests Burn-down Diffusion Process for Personalized Collaborative Filtering

    Authors: Yifang Qin, Zhaobin Li, Arisa Watanabe, Wei Ju, Zhiping Xiao, Ming Zhang

    Abstract: Generative methods have gained widespread attention in Collaborative Filtering (CF) tasks for their ability to produce high-quality personalized samples aligned with users' interests. Among them, diffusion generative models have raised increasing attention in recommendation field. Despite that the pioneering efforts have applied the conventional diffusion process to model diffusive user interests,… ▽ More

    Submitted 6 May, 2026; originally announced May 2026.

  2. arXiv:2602.22742  [pdf, ps, other

    cs.CV

    ProjFlow: Projection Sampling with Flow Matching for Zero-Shot Exact Spatial Motion Control

    Authors: Akihisa Watanabe, Qing Yu, Edgar Simo-Serra, Kent Fujiwara

    Abstract: Generating human motion with precise spatial control is a challenging problem. Existing approaches often require task-specific training or slow optimization, and enforcing hard constraints frequently disrupts motion naturalness. Building on the observation that many animation tasks can be formulated as a linear inverse problem, we introduce ProjFlow, a training-free sampler that achieves zero-shot… ▽ More

    Submitted 26 February, 2026; originally announced February 2026.

  3. arXiv:2602.22594  [pdf, ps, other

    cs.CV

    Causal Motion Diffusion Models for Autoregressive Motion Generation

    Authors: Qing Yu, Akihisa Watanabe, Kent Fujiwara

    Abstract: Recent advances in motion diffusion models have substantially improved the realism of human motion synthesis. However, existing approaches either rely on full-sequence diffusion models with bidirectional generation, which limits temporal causality and real-time applicability, or autoregressive models that suffer from instability and cumulative errors. In this work, we present Causal Motion Diffusi… ▽ More

    Submitted 25 February, 2026; originally announced February 2026.

    Comments: Accepted to CVPR 2026, Project website: https://yu1ut.com/CMDM-HP/

  4. arXiv:2509.20927  [pdf, ps, other

    cs.CV

    SimDiff: Simulator-constrained Diffusion Model for Physically Plausible Motion Generation

    Authors: Akihisa Watanabe, Jiawei Ren, Li Siyao, Yichen Peng, Erwin Wu, Edgar Simo-Serra

    Abstract: Generating physically plausible human motion is crucial for applications such as character animation and virtual reality. Existing approaches often incorporate a simulator-based motion projection layer to the diffusion process to enforce physical plausibility. However, such methods are computationally expensive due to the sequential nature of the simulator, which prevents parallelization. We show… ▽ More

    Submitted 25 September, 2025; originally announced September 2025.

  5. arXiv:2507.00808  [pdf, ps, other

    cs.SD cs.CL eess.AS

    Multi-interaction TTS toward professional recording reproduction

    Authors: Hiroki Kanagawa, Kenichi Fujita, Aya Watanabe, Yusuke Ijima

    Abstract: Voice directors often iteratively refine voice actors' performances by providing feedback to achieve the desired outcome. While this iterative feedback-based refinement process is important in actual recordings, it has been overlooked in text-to-speech synthesis (TTS). As a result, fine-grained style refinement after the initial synthesis is not possible, even though the synthesized speech often d… ▽ More

    Submitted 2 July, 2025; v1 submitted 1 July, 2025; originally announced July 2025.

    Comments: 7 pages,6 figures, Accepted to Speech Synthesis Workshop 2025 (SSW13)

  6. arXiv:2505.00755  [pdf, other

    cs.CV cs.AI

    P2P-Insole: Human Pose Estimation Using Foot Pressure Distribution and Motion Sensors

    Authors: Atsuya Watanabe, Ratna Aisuwarya, Lei Jing

    Abstract: This work presents P2P-Insole, a low-cost approach for estimating and visualizing 3D human skeletal data using insole-type sensors integrated with IMUs. Each insole, fabricated with e-textile garment techniques, costs under USD 1, making it significantly cheaper than commercial alternatives and ideal for large-scale production. Our approach uses foot pressure distribution, acceleration, and rotati… ▽ More

    Submitted 1 May, 2025; originally announced May 2025.

  7. arXiv:2410.20233  [pdf, other

    cs.IT math.CO quant-ph

    Characterization of $n$-Dimensional Toric and Burst-Error-Correcting Quantum Codes from Lattice Codes

    Authors: Cibele Cristina Trinca, Reginaldo Palazzo Jr., J. Carmelo Interlando, Ricardo Augusto Watanabe, Clarice Dias de Albuquerque, Edson Donizete de Carvalho, Antonio Aparecido de Andrade

    Abstract: Quantum error correction is essential for the development of any scalable quantum computer. In this work we introduce a generalization of a quantum interleaving method for combating clusters of errors in toric quantum error-correcting codes. We present new $n$-dimensional toric quantum codes, where $n\geq 5$, which are featured by lattice codes and apply the proposed quantum interleaving method to… ▽ More

    Submitted 26 October, 2024; originally announced October 2024.

    Comments: This manuscript is the proposed generalization of the work New Three and Four-Dimensional Toric and Burst-Error-Correcting Quantum Codes

  8. arXiv:2407.04270  [pdf, other

    eess.AS cs.SD

    Who Finds This Voice Attractive? A Large-Scale Experiment Using In-the-Wild Data

    Authors: Hitoshi Suda, Aya Watanabe, Shinnosuke Takamichi

    Abstract: This paper introduces CocoNut-Humoresque, an open-source large-scale speech likability corpus that includes speech segments and their per-listener likability scores. Evaluating voice likability is essential to designing preferable voices for speech systems, such as dialogue or announcement systems. In this study, we let 885 listeners rate 1800 speech segments of a wide range of speakers regarding… ▽ More

    Submitted 5 July, 2024; originally announced July 2024.

    Comments: Accepted at Interspeech 2024

  9. arXiv:2403.13353  [pdf, other

    cs.SD eess.AS

    Building speech corpus with diverse voice characteristics for its prompt-based representation

    Authors: Aya Watanabe, Shinnosuke Takamichi, Yuki Saito, Wataru Nakata, Detai Xin, Hiroshi Saruwatari

    Abstract: In text-to-speech synthesis, the ability to control voice characteristics is vital for various applications. By leveraging thriving text prompt-based generation techniques, it should be possible to enhance the nuanced control of voice characteristics. While previous research has explored the prompt-based manipulation of voice characteristics, most studies have used pre-recorded speech, which limit… ▽ More

    Submitted 20 March, 2024; originally announced March 2024.

    Comments: Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing. arXiv admin note: text overlap with arXiv:2309.13509

  10. Open-Domain Dialogue Quality Evaluation: Deriving Nugget-level Scores from Turn-level Scores

    Authors: Rikiya Takehi, Akihisa Watanabe, Tetsuya Sakai

    Abstract: Existing dialogue quality evaluation systems can return a score for a given system turn from a particular viewpoint, e.g., engagingness. However, to improve dialogue systems by locating exactly where in a system turn potential problems lie, a more fine-grained evaluation may be necessary. We therefore propose an evaluation approach where a turn is decomposed into nuggets (i.e., expressions associa… ▽ More

    Submitted 30 September, 2023; originally announced October 2023.

    Journal ref: In Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region (SIGIR-AP `23), November 26-28, 2023, Beijing, China. ACM, New York, NY, USA, 6 pages

  11. arXiv:2309.13509  [pdf, other

    cs.SD eess.AS

    Coco-Nut: Corpus of Japanese Utterance and Voice Characteristics Description for Prompt-based Control

    Authors: Aya Watanabe, Shinnosuke Takamichi, Yuki Saito, Wataru Nakata, Detai Xin, Hiroshi Saruwatari

    Abstract: In text-to-speech, controlling voice characteristics is important in achieving various-purpose speech synthesis. Considering the success of text-conditioned generation, such as text-to-image, free-form text instruction should be useful for intuitive and complicated control of voice characteristics. A sufficiently large corpus of high-quality and diverse voice samples with corresponding free-form d… ▽ More

    Submitted 23 September, 2023; originally announced September 2023.

    Comments: Submitted to ASRU2023

  12. arXiv:2307.06241  [pdf, other

    cs.IT quant-ph

    New Three and Four-Dimensional Toric and Burst-Error-Correcting Quantum Codes

    Authors: Cibele Cristina Trinca, Reginaldo Palazzo Jr., Ricardo Augusto Watanabe, Clarice Dias de Albuquerque, José Carmelo Interlando, Antônio Aparecido de Andrade

    Abstract: Ongoing research and experiments have enabled quantum memory to realize the storage of qubits. On the other hand, interleaving techniques are used to deal with burst of errors. Effective interleaving techniques for combating burst of errors by using classical error-correcting codes have been proposed in several articles found in the literature, however, to the best of our knowledge, little is know… ▽ More

    Submitted 12 July, 2023; originally announced July 2023.

    Comments: 16 pages, 2 figures. arXiv admin note: substantial text overlap with arXiv:2205.13582

  13. arXiv:2210.09916  [pdf, other

    cs.SD eess.AS

    Mid-attribute speaker generation using optimal-transport-based interpolation of Gaussian mixture models

    Authors: Aya Watanabe, Shinnosuke Takamichi, Yuki Saito, Detai Xin, Hiroshi Saruwatari

    Abstract: In this paper, we propose a method for intermediating multiple speakers' attributes and diversifying their voice characteristics in ``speaker generation,'' an emerging task that aims to synthesize a nonexistent speaker's naturally sounding voice. The conventional TacoSpawn-based speaker generation method represents the distributions of speaker embeddings by Gaussian mixture models (GMMs) condition… ▽ More

    Submitted 18 October, 2022; originally announced October 2022.

    Comments: Submitted to ICASSP 2023. Demo: https://sarulab-speech.github.io/demo_mid-attribute-speaker-generation

  14. arXiv:2207.00157  [pdf

    eess.IV cs.CV cs.LG q-bio.QM

    Improving Disease Classification Performance and Explainability of Deep Learning Models in Radiology with Heatmap Generators

    Authors: Akino Watanabe, Sara Ketabi, Khashayar, Namdar, Farzad Khalvati

    Abstract: As deep learning is widely used in the radiology field, the explainability of such models is increasingly becoming essential to gain clinicians' trust when using the models for diagnosis. In this research, three experiment sets were conducted with a U-Net architecture to improve the classification performance while enhancing the heatmaps corresponding to the model's focus through incorporating hea… ▽ More

    Submitted 28 June, 2022; originally announced July 2022.

  15. arXiv:2205.13582  [pdf, other

    cs.IT math.CO

    On the Construction of New Toric Quantum Codes and Quantum Burst-Error Correcting Codes

    Authors: Cibele Cristina Trinca, J. Carmelo Interlando, Reginaldo Palazzo Jr., Antonio Aparecido de Andrade, Ricardo Augusto Watanabe

    Abstract: A toric quantum error-correcting code construction procedure is presented in this work. A new class of an infinite family of toric quantum codes is provided by constructing a classical cyclic code on the square lattice $\mathbb{Z}_{q}\times \mathbb{Z}_{q}$ for all odd integers $q\geq 5$ and, consequently, new toric quantum codes are constructed on such square lattices regardless of whether $q$ can… ▽ More

    Submitted 26 May, 2022; originally announced May 2022.

    Comments: Submitted to "Journal of Algebra, Combinatorics, Discrete Structures and Applications"

  16. A Novel Approach to Analyze Fashion Digital Archive from Humanities

    Authors: Satoshi Takahashi, Keiko Yamaguchi, Asuka Watanabe

    Abstract: Fashion styles adopted every day are an important aspect of culture, and style trend analysis helps provide a deeper understanding of our societies and cultures. To analyze everyday fashion trends from the humanities perspective, we need a digital archive that includes images of what people wore in their daily lives over an extended period. In fashion research, building digital fashion image archi… ▽ More

    Submitted 10 September, 2021; v1 submitted 17 July, 2021; originally announced July 2021.

    Comments: In Proceedings of 'The 23rd International Conference on Asia-Pacific Digital Libraries' 17 pages, 8 figures. arXiv admin note: text overlap with arXiv:2009.13395

    Journal ref: In International Conference on Asian Digital Libraries (pp. 179-194). Springer, Cham (2021)

  17. arXiv:2106.02836  [pdf, other

    cs.LG

    Constrained Generalized Additive 2 Model with Consideration of High-Order Interactions

    Authors: Akihisa Watanabe, Michiya Kuramata, Kaito Majima, Haruka Kiyohara, Kensho Kondo, Kazuhide Nakata

    Abstract: In recent years, machine learning and AI have been introduced in many industrial fields. In fields such as finance, medicine, and autonomous driving, where the inference results of a model may have serious consequences, high interpretability as well as prediction accuracy is required. In this study, we propose CGA2M+, which is based on the Generalized Additive 2 Model (GA2M) and differs from it in… ▽ More

    Submitted 22 November, 2021; v1 submitted 5 June, 2021; originally announced June 2021.

  18. arXiv:2009.13395  [pdf, ps, other

    cs.CV cs.DB cs.DL

    CAT STREET: Chronicle Archive of Tokyo Street-fashion

    Authors: Satoshi Takahashi, Keiko Yamaguchi, Asuka Watanabe

    Abstract: The analysis of daily-life fashion trends can provide us a profound understanding of our societies and cultures. However, no appropriate digital archive exists that includes images illustrating what people wore in their daily lives over an extended period. In this study, we propose a new fashion image archive, Chronicle Archive of Tokyo Street-fashion (CAT STREET), to shed light on daily-life fash… ▽ More

    Submitted 29 April, 2021; v1 submitted 28 September, 2020; originally announced September 2020.

    Comments: 19 pages, 17 figures

  19. arXiv:2003.10784  [pdf, other

    cs.NI cs.LG stat.AP stat.ML

    Recovery command generation towards automatic recovery in ICT systems by Seq2Seq learning

    Authors: Hiroki Ikeuchi, Akio Watanabe, Tsutomu Hirao, Makoto Morishita, Masaaki Nishino, Yoichi Matsuo, Keishiro Watanabe

    Abstract: With the increase in scale and complexity of ICT systems, their operation increasingly requires automatic recovery from failures. Although it has become possible to automatically detect anomalies and analyze root causes of failures with current methods, making decisions on what commands should be executed to recover from failures still depends on manual operation, which is quite time-consuming. To… ▽ More

    Submitted 24 March, 2020; originally announced March 2020.

    Comments: accepted for IEEE/IFIP Network Operations and Management Symposium 2020 (NOMS2020)