You can generate subtitles from a video with Python and FFmpeg by using FFmpeg’s Whisper filter to transcribe the audio, save the result as an editable SRT file, and optionally render a reviewed caption file into a new video. This guide builds that local workflow step by step, including prerequisites, error handling, and the difference between sidecar, burned-in, and selectable subtitles.
What you need before generating subtitles
- Python to orchestrate the steps.
- FFmpeg available as
ffmpegon your PATH, or an explicit path to its executable. - A whisper.cpp model file. FFmpeg’s Whisper filter requires a model file; the filter performs automatic speech recognition using the OpenAI Whisper model. See the FFmpeg Whisper filter documentation.
- A video file with speech and a writable output directory.
FFmpeg can read media, apply filters, transcode, and write output files. Whisper support and filter syntax depend on the FFmpeg build, so check that your installed executable recognizes the filter before relying on it in an application. The filter reference documents its options, including language, destination format, queue, maximum segment length, and optional voice-activity-detection controls.
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
The AI Income Generator: Subtitle: From GPT to Midjourney: Learn the Prompts, Tools, and Workflows | $0.99 | Buy on Amazon |
| 2 |
|
Intermediate Python | $41.63 | Buy on Amazon |
How to generate an SRT file automatically with Python and FFmpeg
1. Check the input and output paths
Validate that the video and model exist, and that the destination directory can be written to. Keeping the SRT as a separate intermediate file makes it easy to inspect and correct recognition or timing before you render captions into a video.
2. Call FFmpeg with a Python argument list
Python’s documentation recommends subprocess.run() for subprocess use cases it can handle. Pass the command as a list rather than building a shell command string; shell=False is the default and avoids unnecessary shell interpretation.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
from pathlib import Path
import subprocess
def generate_srt(video: Path, model: Path, srt: Path, language: str = "en") -> None:
command = [
"ffmpeg", "-y", "-i", str(video), "-vn",
"-af",
f"whisper=model={model}:language={language}:"
f"destination={srt}:format=srt",
"-f", "null", "-",
]
subprocess.run(
command,
check=True,
capture_output=True,
text=True,
timeout=3600,
)
The filter sends the audio through Whisper, writes SRT to the specified destination, and directs FFmpeg’s regular media output to the null muxer rather than creating a video. The language shown is a configuration value; set it for the speech being transcribed and confirm the options supported by your FFmpeg build. The model path is mandatory.
Here, check=True raises subprocess.CalledProcessError if FFmpeg exits unsuccessfully. capture_output=True retains stdout and stderr for diagnostics, and timeout prevents an unattended process from running indefinitely. Python documents run(), its options, and security considerations for shell=True in the subprocess documentation.
3. Handle failures and path edge cases
Catch FileNotFoundError to report that FFmpeg is missing or unavailable at the configured executable path. Catch subprocess.CalledProcessError to report an FFmpeg failure, using captured stderr to help diagnose it. Also handle subprocess.TimeoutExpired if the process exceeds the chosen timeout.
FFmpeg filter strings have their own parsing rules. Model paths and filenames containing spaces or special characters can therefore need careful escaping even when Python receives an argument list. Test those cases against the exact FFmpeg build you deploy; do not assume that a Python list eliminates FFmpeg’s filter-level parsing concerns. Make the model path configurable rather than embedding it in application code.
Recommended Free Tools
4. Keep output safe and reproducible
- Write the SRT to a temporary file and rename it to its final name only after FFmpeg succeeds, so a failed run does not leave a misleading partial output.
- Keep the original video untouched. Render to a new output file when burning captions in.
- Preserve stderr for troubleshooting, but avoid exposing sensitive local paths in logs shared outside your environment.
- Record the FFmpeg version and model identifier with each job so you can reproduce the configuration later.
Choose the subtitle format and output mode
SubRip (SRT) is a practical first format because it is plain text and easy to review. FFmpeg also documents WebVTT and SSA/ASS subtitle formats; choose based on where the captions will be used and how much styling matters. The FFmpeg formats documentation lists supported formats.
| Choice | Best fit | What the viewer receives |
|---|---|---|
| SRT | Easy editing and broad use as a caption sidecar | A separate text subtitle file |
| WebVTT | A web player destination | A separate subtitle file suitable for web workflows |
| ASS/SSA | Styling and positioning are central | A subtitle format with styling and positioning capabilities |
| Sidecar file | Captions need to remain editable or independently loadable | A separate file, such as captions.srt, alongside the video |
| Burned-in video | Captions must appear as part of the picture | Text rendered into the video image |
| Muxed subtitle track | Viewers should be able to select captions in a compatible player | A subtitle stream packaged with the media |
Format support and exact output behavior depend on the task and FFmpeg build. Consult the format table for format details and FFmpeg’s command documentation for stream mapping and output behavior.
Rank #2
How to burn subtitles into an MP4
Review and correct the SRT first, then use FFmpeg’s subtitles video filter to render it into a new file:
ffmpeg -i input.mp4 -vf "subtitles=captions.srt" -c:a copy output-burned.mp4
This example copies the audio stream and applies the subtitles filter to the video. The text is part of the resulting picture, so the viewer cannot turn it off. The subtitles filter requires an FFmpeg build configured with libass; verify that support before running this step. See the FFmpeg subtitles filter documentation.
If captions should remain selectable, mux a subtitle stream instead of applying a video filter. Use explicit stream mapping to choose the video, audio, and subtitle streams appropriate for your input and target container; FFmpeg’s command documentation explains mapping. Container and player support affect whether a muxed subtitle track can be selected.
Local versus hosted transcription: what to weigh
The example above keeps transcription in a local FFmpeg workflow and does not require an API key. You manage the model file and the compute environment yourself. A hosted transcription service may reduce model-management work, but introduces account, network, privacy, pricing, and regional-availability considerations. AWS Transcribe documents SRT and WebVTT subtitle output in its subtitle documentation; check the service’s current terms and availability before adopting it.
CPU or GPU use, processing time, and transcription quality vary with the model, language, hardware, audio, and settings. There is no universal accuracy, speed, or cost figure that applies to all Python-and-FFmpeg subtitle jobs. Choose based on your privacy needs, operational capacity, intended caption format, and the actual results on your media.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




