AirPlay and AirPlay 2 audio player
-
Updated
Aug 9, 2026 - C
AirPlay and AirPlay 2 audio player
EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine
PyTorch implementation of convolutional neural networks-based text-to-speech synthesis models
Real-time voice assistant — WebRTC streaming, faster-whisper ASR, local LLM, Vui Nano (300M) TTS. OpenAI Realtime API compatible. Voice cloning, barge-in, ~9× realtime on a 4090. Apache 2.0.
Chinese Mandarin tts text-to-speech 中文 (普通话) 语音 合成 , by fastspeech 2 , implemented in pytorch, using waveglow as vocoder, with biaobei and aishell3 datasets
VoxNovel: generate audiobooks giving each character a different voice actor.
A Non-Autoregressive Transformer based Text-to-Speech, supporting a family of SOTA transformers with supervised and unsupervised duration modelings. This project grows with the research community, aiming to achieve the ultimate TTS
Two-talker Speech Separation with LSTM/BLSTM by Permutation Invariant Training method.
A Non-Autoregressive End-to-End Text-to-Speech (text-to-wav), supporting a family of SOTA unsupervised duration modelings. This project grows with the research community, aiming to achieve the ultimate E2E-TTS
Draft to Take beta: local-first AI audio production studio powered by IndexTTS2, Docker, Qwen, OmniVoice, SFX, ambience, and music sidecars.
This is the official implementation of our multi-channel multi-speaker multi-spatial neural audio codec architecture.
Adaptive and Focusing Neural Layers for Multi-Speaker Separation Problem
PyTorch Implementation of Google's Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions. This implementation supports both single-, multi-speaker TTS and several techniques to enforce the robustness and efficiency of the model.
🎵 Complete offline audio transcription system with speaker diarization using OpenAI Whisper and PyAnnote. Features automatic audio cleaning, precise timestamps, multiple output formats (JSON/TXT/Markdown), and support for 20+ audio formats. No external APIs required - works entirely offline.
Multi-Speaker FastSpeech2 applicable to Korean. Description about train and synthesize in detail.
DAVE: audio-visual speech enhancement & target speaker extraction for real-world meetings — audio-only TIGER-M backbone (2.56M) + four-cue speaker attribution. ISCSLP 2026 AVSE Challenge: Track1 4th, Track2 3rd.
An Algorithm for Speaker Recognition in a Multi-Speaker Environment
Urdu Speech Recognition using Kaldi ASR, by training Triphone Acoustic GMMs using the PRUS dataset.
Free, offline text-to-speech for conversations — give it a transcript, pick a voice per speaker, get one audio file. Runs on CPU.
基于 VoxCPM2 的 HTTP TTS 服务,面向 Legado 与有声书场景,支持零样本音色克隆、多角色路由、流式语音合成、ASR 提示词自动生成及跨平台部署。
Add a description, image, and links to the multi-speaker topic page so that developers can more easily learn about it.
To associate your repository with the multi-speaker topic, visit your repo's landing page and select "manage topics."