VoiceStudio — локальный клон голоса и синтез речи
VoiceStudio is an open-source tool for voice cloning, text-to-speech synthesis, dictation, video dubbing, and audiobook production that runs entirely on local hardware — no cloud, no account, no subscription required. Reach for it when you need to dub a video in your own cloned voice without sending files to third-party servers, when you want to produce an audiobook or podcast offline, when recordings are too sensitive for cloud processing, or when recurring TTS subscription costs have become a problem. It ships with 16 TTS engines, 11 ASR engines, and a 646-language catalogue, packaged as a desktop app for macOS 13.3+ (Apple Silicon), Windows 10/11 x64, and Linux x86_64 with glibc 2.39+. Acceleration covers CUDA, Apple Silicon MPS/MLX, ROCm on Linux, and CPU-only. Interfaces include a local REST/SSE/WebSocket API, an OpenAI-compatible audio endpoint, and an MCP Server. All voices, projects, and outputs stay on the machine by default. Direction is speech synthesis, voice cloning, and dubbing — not primarily speech-to-text transcription. Voice cloning and voice deepfakes must be used lawfully, only on your own material and with the consent of the people involved. Works well for fully offline pipelines; note it is in active beta, so use the latest release for stable operation.