CrisperWhisper — точная расшифровка речи с таймстемпами

CrisperWhisper is an open-source speech-to-text library that offers controllable verbatim or clean-intended transcription, precise word-level timestamps, and multilingual support across most languages Whisper handles. Reach for it when you need every filler, stutter, false start, and vocal event captured exactly as spoken — for clinical speech analysis, TTS dataset construction, or conversation analytics — or when you want the opposite: a polished, readable transcript from the same recording. It's also the right pick when you already have a clean transcript and want to layer in the real disfluencies from the audio without re-transcribing from scratch (the verbatimize feature). Built on Whisper-based models with a CTranslate2 runtime for NVIDIA GPU inference (speculative decoding included) and a pure PyTorch path for macOS, Windows, and CPU; installed via pip, models pulled from Hugging Face. Verbatim and intended modes are a single parameter; word timestamps carry ~30 ms mean boundary error on read speech. The tool transcribes speech into text — not the other way around — and outperforms WhisperX on timing precision and disfluency recall in benchmarks. Works well for long-form audio of any length; not designed for low-latency real-time streaming use cases.