Voice-Pro — синтез, клонирование и перевод речи
★ 12.7K
Voice-Pro is an open-source web application that combines speech recognition, translation, text-to-speech, and zero-shot voice cloning in a single Gradio interface. Reach for it when you want to transcribe a YouTube video and re-dub it in another language without stitching scripts together, when you need to separate vocals from a music track before editing, when you're producing a podcast or audiobook and want a synthesised voice without studio recording, or when you need accurate subtitles generated automatically. Built on Python; the stack includes faster-whisper 1.2.1 and openai-whisper for ASR, Edge-TTS and kokoro for synthesis, F5-TTS / E2-TTS / CosyVoice for zero-shot voice cloning, Demucs and MDX-Net for stem separation, and yt-dlp for source downloads. Deep-Translator covers 100+ languages for translation, with optional Azure TTS and Azure Translator via your own keys. Runs on Windows via start.bat and requires an NVIDIA GPU; Linux and Mac are not officially verified. The pipeline moves from video or audio input toward finished dubbed output — not real-time streaming recognition. Cloning someone else's voice is permitted only with their consent and in compliance with applicable law. Not suitable for live transcription workflows or machines without a discrete NVIDIA GPU.
- #Speech