FireRedVAD — детектор голосовой активности

★ 491

FireRedVAD is an industrial-grade voice-activity detector (VAD) and audio-event detector (AED): it identifies where speech (as well as singing and music) occurs in audio, in streaming and non-streaming modes, across 100+ languages. Reach for it as a utility component in audio pipelines: to slice a long recording into speech segments before transcription, cut out silence and noise, save on recognition, and detect singing/music. By reported metrics it outperforms Silero-VAD, TEN-VAD, FunASR-VAD, and WebRTC-VAD (97.57% F1 on FLEURS-VAD-102); weights and inference code are provided. Its focus is precisely finding speech boundaries and sound types in audio, not transcribing speech into text (ASR like whisperX does that) or generating sound: it is a preprocessor placed before ASR or audio analytics. It is useful to developers of voice pipelines for improving the accuracy and speed of subsequent steps.