ElevenLabs Scribe — транскрипция и диаризация аудио

★ 7.7 · content

elevenlabs-stt is a Claude Code skill that delivers high-accuracy audio transcription using ElevenLabs Scribe v1 and v2 models — with 98%+ accuracy across 90+ languages, including automatic language detection. It runs via the inference.sh CLI (`belt`) and supports speaker diarization, audio event tagging (laughter, applause, music), word-level timestamps, and forced alignment that maps known text to audio at word and character precision. Forced alignment is particularly useful for subtitle timing, lip-sync animation, and karaoke lyrics. The skill fits workflows like meeting and interview transcription, podcast transcript generation, timed video captioning, and making audio content searchable and accessible.