Collection of resources on the applications of Large Language Models (LLMs) in Audio AI.
-
Updated
Aug 8, 2026
8000
Collection of resources on the applications of Large Language Models (LLMs) in Audio AI.
Rapida is an open-source, end-to-end voice AI orchestration platform for building real-time conversational voice agents with audio streaming, STT, TTS, VAD, multi-channel integration, agent state management, and observability.
A New End-to-end Framework for Evaluating Voice Agents
🇺🇦 Open Source Ukrainian Text-to-Speech datasets
A Docker-based OpenAI-compatible Text-to-Speech API server powered by Kyutai's TTS models with GPU acceleration support.
Just a simple multimodal avatar interaction platform
A voice-based AI chat interface built with Next.js and ElevenLabs. Start and stop real-time conversations with an animated UI that reflects agent status. Fully responsive and deployable via Vercel with environment-based agent configuration.
MLX Porting Toolkit — an agent-guided, evidence-gated pipeline (scaffold → convert → parity → benchmark) plus a portable skill for porting PyTorch/Hugging Face models to Apple MLX.
A unified benchmarking framework for evaluating Voice AI agents across conversational quality, audio realism, latency metrics, and safety guardrails with scalable multi-language stress testing.
Open-source real-time Voice AI infrastructure in Go. Stream audio via WebRTC or WebSocket, connect STT → LLM → TTS pipelines, and build scalable voice agents and conversational AI applications.
🇺🇦 Ukrainian RAD-TTS++ models (decoder + models with 3 voices) and HiFiGAN model
A voice-based AI chat interface built with Next.js and ElevenLabs. Start and stop real-time conversations with an animated UI that reflects agent status. Fully responsive and deployable via Vercel with environment-based agent configuration.
Legacy Speech AI examples with migration links to the current Brainiall TTS and transcription services.
A source-linked directory of free and trial LLM APIs, multimodal models, embeddings, speech, translation, safety, and other inference endpoints. Companion catalog for freellmapi.io.
A curated list of the best Text-to-Speech, speech synthesis, and voice-cloning research — models, papers, benchmarks, and toolkits, focused on 2025–2026.
中文 ASR 评测工具箱 · micro-CER 对比 FunASR/Whisper/llama.cpp · 一条命令出报告 · 自带迷你测试集 · Mandarin ASR benchmark toolkit
Code-switching ASR adaptation for strong multilingual speech recognition models. Synthetic CSW data generation, Whisper adaptation, Bayesian LoRA (BLoRA), and robust multilingual ASR evaluation.
Interruptible voice-agent runtime for structured interview prototypes, with VAD-based interruption handling and modular speech backends.
Interactive documentation helper for Sarvam AI APIs — grounded answers from official docs for multilingual speech & language products
Add a description, image, and links to the speech-ai topic page so that developers can more easily learn about it.
To associate your repository with the speech-ai topic, visit your repo's landing page and select "manage topics."