Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
The reliable way to turn a YouTube video into Markdown depends on what access you have. For a video you control, use YouTube Data API OAuth: call captions.list, download the selected track with captions.download, parse SRT/VTT (or another supported format), and render Markdown. That download method requires permission to edit the video, so it is not a universal downloader for arbitrary public videos. If no usable captions exist, send an audio file you are permitted to process to a speech-to-text service, or use a hosted transcript API that documents caption extraction and asynchronous ASR.
Choose the right transcript path
| Situation | Best approach | Important limitation |
|---|---|---|
| You own or manage the video | YouTube Data API OAuth, captions.list, then captions.download |
The authorized user must be allowed to edit the video. |
| A public video has captions but you lack edit permission | A hosted provider that documents YouTube URL extraction | Check its current terms, retention, rate limits and permission requirements. |
| No captions, but you have a permitted audio file | Upload the file to an ASR endpoint such as OpenAI’s transcription API | The documented endpoint accepts uploaded files, not a YouTube URL. |
These are different authorization and data-handling models. Keep the video ID, caption-track ID, language, original format and retrieval time with every generated document so a reader can audit where the Markdown came from.
Official YouTube workflow for videos you can edit
1. Extract and validate the video ID
Accept common URL forms, but store only the 11-character ID. Reject malformed input before making an API request.
Free tools Windows power users keep installed
One-click scans. No signup required.
from urllib.parse import urlparse, parse_qs
import re
def youtube_id(value: str) -> str:
value = value.strip()
if re.fullmatch(r"[A-Za-z0-9_-]{11}", value):
return value
parsed = urlparse(value)
if parsed.hostname in {"youtu.be"}:
candidate = parsed.path.lstrip("/").split("/")[0]
elif parsed.hostname and parsed.hostname.endswith("youtube.com"):
candidate = parse_qs(parsed.query).get("v", [""])[0]
else:
candidate = ""
if not re.fullmatch(r"[A-Za-z0-9_-]{11}", candidate):
raise ValueError("Invalid YouTube video ID or URL")
return candidate
2. Authorize with OAuth 2.0
Create credentials in Google Cloud, enable YouTube Data API v3, and run an OAuth consent flow for an account that can edit the target video. An API key alone is not sufficient for caption downloads. Store refresh tokens securely and request only the scope your application needs. Do not ask users to paste a token into logs or source control.
#1 Best Overall
3. Discover tracks with captions.list
The list response contains track IDs and metadata such as language and status; it does not contain the caption text. Reject tracks whose status indicates processing or failure. Prefer the requested language, then a documented fallback language, and record the choice.
from googleapiclient.discovery import build
youtube = build("youtube", "v3", credentials=credentials)
video_id = youtube_id("https://www.youtube.com/watch?v=VIDEO_ID")
response = youtube.captions().list(
part="id,snippet",
videoId=video_id
).execute()
tracks = []
for item in response.get("items", []):
snippet = item["snippet"]
if snippet.get("status") == "serving":
tracks.append({
"id": item["id"],
"language": snippet.get("language"),
"name": snippet.get("name"),
"kind": snippet.get("trackKind")
})
if not tracks:
raise RuntimeError("No serving caption track is available")
4. Download the selected track
Call captions.download with the track ID. Google documents SRT, VTT, TTML, SBV and SCC output through tfmt; tlang can request a translated track where available. The documented quota cost is 200 units per download, so avoid downloading repeatedly when a cached source is still valid.
download = youtube.captions().download(
id=tracks[0]["id"],
tfmt="vtt"
).execute()
with open("captions.vtt", "wb") as file:
file.write(download)
If the API returns a forbidden error, the OAuth user probably cannot edit that video. Invalid-value errors usually indicate a bad track ID or unsupported parameter; not-found can mean the video or track was removed. Treat these as actionable errors rather than silently producing an empty transcript.
Convert SRT or VTT into readable Markdown
Subtitle files are timed cue data, not prose. Remove sequence numbers, timecode lines and formatting tags, then merge adjacent fragments without erasing meaningful paragraph breaks. Keep timestamps when your use case requires auditability.
Rank #2
import html
import re
from datetime import datetime, timezone
def subtitle_text(raw: str) -> list[str]:
raw = raw.replace("rn", "n").replace("r", "n")
blocks = re.split(r"ns*n", raw.strip())
paragraphs = []
for block in blocks:
lines = [line.strip() for line in block.split("n") if line.strip()]
if not lines:
continue
if lines[0].upper() == "WEBVTT" or re.fullmatch(r"d+", lines[0]):
lines = lines[1:]
lines = [line for line in lines if not re.match(
r"^d{2}:d{2}(?::d{2})?[.,]d{3}s+-->s+", line)]
text = " ".join(lines)
text = re.sub(r"<[^>]+>", "", text)
text = html.unescape(re.sub(r"s+", " ", text)).strip()
if text:
paragraphs.append(text)
return paragraphs
def markdown_document(title, source_url, language, track_id, raw):
paragraphs = subtitle_text(raw.decode("utf-8", errors="replace"))
retrieved = datetime.now(timezone.utc).isoformat()
safe_title = title.replace("\", "\\").replace("[", "\[").replace("]", "\]")
lines = [
f"# {safe_title}", "", f"- Source: {source_url}",
f"- Language: {language}", f"- Caption track: `{track_id}`",
f"- Retrieved: {retrieved}", "", "## Transcript", ""
]
lines.extend(p + "n" for p in paragraphs)
return "n".join(lines)
markdown = markdown_document(
"Video title", "https://www.youtube.com/watch?v=VIDEO_ID",
tracks[0]["language"], tracks[0]["id"], download
)
open("transcript.md", "w", encoding="utf-8").write(markdown)
The parser deliberately strips cue timing. For legal, editorial or research workflows, retain the original VTT/SRT beside the Markdown and add timestamp links or a cue table instead of discarding timing data.
Complete request examples and API boundaries
cURL for a hosted transcript endpoint
A hosted service can accept a YouTube URL and return caption-based text or queue an ASR job when captions are missing. Follow that provider’s current schema; the following shape reflects the documented POST /api/v2/transcribe route of YouTubeTranscript.dev, not the Google API.
curl -X POST https://youtubetranscript.dev/api/v2/transcribe
-H "Content-Type: application/json"
-H "Authorization: Bearer YOUR_API_KEY"
-d '{"url":"https://www.youtube.com/watch?v=VIDEO_ID","language":"en","format":"text"}'
Expect either an immediate caption response or an asynchronous job when ASR is required. Verify the provider’s current authentication, response fields, retention and pricing before production use.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Python upload for permitted audio
OpenAI’s transcription endpoint documents file uploads and can return plain text or verbose, timestamped output where supported. It does not accept a YouTube URL directly.
Rank #3
from openai import OpenAI
client = OpenAI()
with open("permitted-audio.mp3", "rb") as audio:
result = client.audio.transcriptions.create(
model="whisper-1",
file=audio,
response_format="verbose_json"
)
text = getattr(result, "text", "")
open("transcript.md", "w", encoding="utf-8").write(
"## Transcriptnn" + text.strip() + "n"
)
The OpenAI documentation says callers must send a supported audio file. Acquiring or downloading audio from YouTube is a separate operation that requires permission and suitable tooling. The Help Center documents a 25 MiB maximum request size for legacy whisper-1 uploads; that model-specific limit should be checked again before deployment. Split larger permitted recordings into overlapping chunks and reconcile duplicated words at boundaries.
Production design: reliability, cost and data handling
Retry only safe failures
- Retry transient HTTP failures with exponential backoff and an upper bound.
- Do not retry forbidden or invalid-value responses without changing authorization or parameters.
- Cache the original caption file keyed by video ID, track ID, format and retrieval time. This reduces repeat quota usage.
- For asynchronous ASR, persist the job ID and poll according to the provider’s guidance rather than hammering the endpoint.
Keep provenance and privacy explicit
Record source URL, video ID, track ID, language, caption kind (human or automatic when exposed), format, retrieval timestamp, parser version and any translation request. Decide whether transcripts or uploaded audio may be retained by a vendor, and provide deletion controls if your application handles personal or confidential material.
Control output quality
- Normalize Unicode and whitespace, but do not remove speaker labels or non-speech cues that matter to meaning.
- Escape Markdown punctuation when caption text contains headings, links, list markers or code.
- For long videos, stream parsing or process chunks to avoid excessive memory use.
- Compare the selected language with the user’s requested language; do not silently substitute a translation.
Troubleshooting common failures
“The API key works, but captions.download is forbidden”
Use OAuth credentials for an account with edit permission on the video. A public video is not automatically downloadable through Google’s caption method.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitches“captions.list returns no usable tracks”
Check track status, language filters and the video ID. If captions are genuinely absent, switch to a permitted-audio ASR path or a hosted provider that documents asynchronous fallback.
Rank #4
“The Markdown is empty or full of timestamps”
Inspect the raw SRT/VTT first. Ensure your parser handles the file’s line endings, removes cue timing only after identifying it, and preserves text lines that contain punctuation or HTML entities.
“A speech-to-text request rejects my YouTube URL”
That is expected for an upload-based endpoint. Obtain an audio file lawfully, then upload it in a supported format and within the model’s current size limit.
“The transcript contains duplicated phrases”
Overlapping ASR chunks commonly repeat boundary words. Remove the longest matching suffix/prefix overlap, then run a human or automated review before publishing.
Or skip the browser setup
If your workflow also needs a clean image of the source page, ScreenshotNeo provides a single-call website screenshot API. It accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server exposes take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for options such as PNG, JPEG or WebP, full-page capture, device and retina settings, custom CSS or JavaScript, selectors, waits, headers, cookies, geolocation, PDF output, caching, signed links, asynchronous jobs and bulk capture. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
Best Value
Implementation checklist
- Validate and store the 11-character video ID.
- Determine whether your OAuth user can edit the video.
- List serving tracks and select language deliberately.
- Download once in VTT or SRT, retaining the original file.
- Parse cues, remove timing markup and escape Markdown characters.
- Add title, source URL, language, track ID and retrieval time.
- Use a hosted extractor or permitted-audio ASR only when its authorization and data policies fit your use case.
- Test forbidden, missing-track, timeout, oversized-file and malformed-subtitle cases.
Frequently Asked Questions
Does captions.list return the transcript text?
No. It returns caption-track metadata and IDs. The text comes from a separate captions.download request.
Can I download captions from any public YouTube video with Google’s API?
No. The documented download method requires OAuth and permission to edit the video.
Which subtitle formats can captions.download produce?
Google documents SRT, VTT, TTML, SBV and SCC through the tfmt parameter.
What should I save besides the Markdown file?
Keep the original subtitle file and provenance fields: video ID, track ID, language, format and retrieval timestamp.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

