-
freelance
- Tokyo
-
09:35
(UTC +09:00) - https://huggingface.co/hironow
- https://cursor.com/@hironow
- @hironow
- in/hironow
ai-vocal
zero-shot voice conversion & singing voice conversion, with real-time support
VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
Separation voice and delete files with majority silence
リアルタイムボイスチェンジャー Realtime Voice Changer
Port of OpenAI's Whisper model in C/C++
SoftVC VITS Singing Voice Conversion
so-vits-svc fork with realtime support, improved interface and more features.
Implementation of Natural Speech 2, Zero-shot Speech and Singing Synthesizer, in Pytorch
Core Engine of Singing Voice Conversion & Singing Voice Clone
Real-time end-to-end singing voice conversion system based on DDSP (Differentiable Digital Signal Processing)
*CREPE+HYBRID TRAINING* A very experimental fork of the Retrieval-based-Voice-Conversion-WebUI repo that incorporates a variety of other f0 methods, along with a hybrid f0 nanmedian method.
Community interface for generative AI
A multi-voice TTS system trained with an emphasis on quality
Eleven Labs text to speech package for NodeJS. You can use the official package at: https://www.npmjs.com/package/elevenlabs
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
Audiocraft is a library for audio processing and generation with deep learning. It features the state-of-the-art EnCodec audio compressor / tokenizer, along with MusicGen, a simple and controllable…
vits2 backbone with multilingual-bert
💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
Faster Whisper transcription with CTranslate2
An environment where you can try out faster-whisper immediately.
A simple FastAPI Server to run XTTSv2