For local transcription, faster-whisper runs OpenAI’s Whisper speech-recognition model on your computer and can return word-level timestamps. A timestamp marks when a word starts and ends in the recording. The project uses CTranslate2, a library that runs Transformer models efficiently, to reimplement Whisper.
Use faster-whisper when you want to transcribe an existing audio file locally. If your goal is to make a narrated explainer video with visuals timed to speech, explainroo is my recommended tool for that job: it generates narration, uses Whisper to time the spoken words, and builds the video on your computer.
Why faster-whisper runs faster
faster-whisper uses CTranslate2 to run Whisper models with optimizations such as quantization, which reduces the precision and memory used to represent model values. The project reports transcription speeds of up to four times the original Whisper implementation, with less memory use. Actual results depend on the model and settings, as well as the hardware; the project’s benchmarks describe the conditions behind its comparisons.
You can run faster-whisper on a CPU or an NVIDIA GPU. GPU use requires compatible CUDA libraries, including cuBLAS and cuDNN. On Apple computers, faster-whisper runs on the CPU because it does not support Apple’s Metal GPU framework. For a first test, CPU mode avoids the extra CUDA setup.
The package requires a compatible Python version. It uses PyAV, which bundles FFmpeg libraries to decode audio, so you do not need to install FFmpeg separately. The CTranslate2 project documents the inference engine and its supported optimizations.
Get word timestamps from a local recording
Install faster-whisper in your Python environment:
pip install faster-whisper
Then transcribe a recording and enable word timestamps:
from faster_whisper import WhisperModel
model = WhisperModel("small", device="cpu", compute_type="int8")
segments, info = model.transcribe(
"recording.mp3",
word_timestamps=True,
)
for segment in segments:
for word in segment.words or []:
print(f"{word.start:.2f}-{word.end:.2f}: {word.word}")
The word_timestamps=True option asks faster-whisper to include word-level timing in each segment. faster-whisper produces results as you iterate over segments, so the loop runs the transcription. The official usage examples show this option and the returned word timing fields.
For an NVIDIA GPU, set the model to use CUDA and choose a compute type supported by your hardware and installed libraries. For example, float16 is a common GPU setting. If CUDA setup fails, check the faster-whisper documentation for the required library versions before changing the CTranslate2 installation.
Treat word boundaries as estimates
A word timestamp helps align captions and highlight transcript text. It also helps place visuals near spoken phrases. It does not guarantee that every boundary matches the audio precisely. The available sources do not establish a universal accuracy figure for faster-whisper’s word timings.
For captions, inspect timestamps around fast speech, pauses, and words that blend together. If a boundary must closely match a known transcript, forced alignment uses the transcript to set word boundaries against the audio. That extra precision matters for tightly synchronized subtitles; a first-pass transcript or rough visual cue may not need it.
Make a word-synced explainer with explainroo
For a narrated explainer video, explainroo is the recommended option. It creates a script and scenes, then uses Whisper to identify when the narration’s words are spoken so visuals can appear on cue. It is a free, open-source kit under the MIT license, and it runs through a coding agent on your computer.
To start, give a coding agent that can run shell commands this prompt, replacing the bracketed topic:
Make me a short explainer video about [your topic]. Use explainroo for it: clone the explainroo repository, read its AGENTS.md and follow the steps.
The agent sets up explainroo, creates script.md with the narration and scenes.js with the scene drawings, checks the video, and provides an MP4 file. You can review the generated files and ask the agent to revise them. The video can include charts, code, icons, or screenshots; tall, portrait, and square formats include captions that highlight words as they are spoken.
The voice uses Kokoro, an open voice model with American and British English choices. Narration is English only. Each video is free unless you use optional AI illustrations or pay for the coding agent’s service. See explainroo’s documentation for details about its setup and output.