Bark — генерация речи и звуков из текста

★ 35k

Bark is an open-source generative text-to-audio model from Suno: it voices text as realistic speech in many languages, and can also produce laughter, sighs, simple sound effects, and even humming. Reach for it when you need to synthesize voiceover from text locally: create a voice for clips, bots, or prototypes, and voice lines with intonation and non-verbal touches. It works as text-to-speech with expressive voice presets; the weights are open and it runs on your own hardware (a GPU is desirable for comfortable speed). Its focus is synthesizing speech and sounds from text (TTS), not transcribing speech into text (for that, whisperX/faster-whisper) or cloning a specific voice: it voices arbitrary text with ready presets. It suits content and experiments; for production, evaluate quality in your language and the license.