YouTube Transcript Scraper — Captions + Speech AI, SRT & VTT avatar

YouTube Transcript Scraper — Captions + Speech AI, SRT & VTT

Pricing

from $1.20 / 1,000 transcripts

Go to Apify Store
YouTube Transcript Scraper — Captions + Speech AI, SRT & VTT

YouTube Transcript Scraper — Captions + Speech AI, SRT & VTT

YouTube videos, Shorts and live VODs to text: captions first, built-in speech-to-text when a video has none, so caption-less videos still return real text. JSON, plain text, SRT or VTT. From $1.20 per 1,000 transcripts, no API key, no cookies. Never charged for a video we can't transcribe.

Pricing

from $1.20 / 1,000 transcripts

Rating

0.0

(0)

Developer

Steadyfetch Team

Steadyfetch Team

Maintained by Community

Actor stats

0

Bookmarked

5

Total users

5

Monthly active users

16 hours ago

Last modified

Categories

Share

Never charged for a video we can't transcribe — or for one you already have. Captions first, built-in speech-to-text when a video has none, so caption-less videos still return real text. Videos, Shorts and live VODs, as JSON, plain text, SRT or VTT.

Just want to see it work? Click Start with nothing set and the run transcribes one short real video (youtube.com/watch?v=jNQXAC9IVRw, "Me at the zoo"), charged like any run at your plan's per-transcript price. Set only settings — a different Output format, a caption language, speech-to-text off, a limit — and leave Video URLs or IDs empty, and that same sample runs under your settings, charged like any run. Put your own links or IDs in Video URLs or IDs for your own run.

YouTube Transcript Scraper input form in the Apify console: Video URLs or IDs prefilled with two videos, above the collapsed transcript-options and cost-control sections

Issues answered in a couple of hours. Unofficial: steadyfetch is not affiliated with, endorsed by, or sponsored by YouTube or Google. "YouTube" is a trademark of Google LLC, used here only to say what this actor reads.

Every row carries charged and statusReason, so you can reconcile the invoice from the dataset itself without opening the console. Only rows with charged: true were billed. If your Maximum cost per run is reached while a video's transcript is already in hand, that transcript still ships in full with charged: false and a statusReason naming the cap — your cap is never exceeded, and work already done is never thrown away.


What a row looks like

Dataset table of a real run: one row — “Me at the zoo” (jNQXAC9IVRw), manual captions, English, 19 seconds, with the transcript text and the video URL

Real output, unedited apart from trimming the segment list:

{
"status": "ok",
"charged": true,
"statusReason": null,
"videoId": "jNQXAC9IVRw",
"url": "https://www.youtube.com/watch?v=jNQXAC9IVRw",
"inputUrl": "https://www.youtube.com/watch?v=jNQXAC9IVRw",
"title": "Me at the zoo",
"channelName": "jawed",
"channelId": "UC4QobU6STFB0P71PMvOGN5A",
"durationSeconds": 19,
"viewCount": 406463143,
"isLiveContent": false,
"language": "en",
"source": "captions",
"captionKind": "manual",
"text": "All right, so here we are, in front of the elephants the cool thing about these guys is that they have really... really really long trunks and that's cool (baaaaaaaaaaahhh!!) and that's pretty much all there is to say",
"segments": [
{ "start": 1.2, "end": 3.36, "text": "All right, so here we are, in front of the elephants" },
{ "start": 5.318, "end": 7.974, "text": "the cool thing about these guys is that they have really..." }
],
"srt": null,
"vtt": null,
"chargeEvents": { "transcript": 1, "speechMinutes": 0 }
}
fieldnotes
text · segments[{start,end,text}]the transcript, and the same text with timestamps in seconds
srt · vttfilled only when you ask for that format; otherwise null
sourcecaptions or speech_ai — how this transcript was produced
captionKindmanual (uploaded by the channel) or auto (YouTube's own auto-captions); null on the speech route
languageISO code, same vocabulary on both routes
title · channelName · channelId · durationSeconds · viewCount · isLiveContentas YouTube reports them
url · inputUrl · sourceIndexthe canonical watch URL, the exact link you passed, and its position in your input
charged · statusReason · chargeEventsthe reconciliation trio
repeat · firstSeenAt · firstSeenRunIdtrue when this account already had this transcript: it was handed back from the run named here, and nothing was charged for it
retryableon a row that did not deliver: true means YouTube refused us this time and the same input is worth running again, false means the answer will not change. It always agrees with the note in statusReason, so code can branch on the column instead of parsing the sentence

Sample dataset — six real transcripts from one run, no sign-in: open the JSON

Agent / API paste-block

Actor: steadyfetch/youtube-transcript-scraper
Required: videoUrls (array of YouTube video links or 11-character video IDs)
Optional: language (string, e.g. "en", "es", "pt-BR" — preferred caption track)
format (json | text | srt | vtt, default json)
enableSpeechFallback (boolean, default true — off = captions only, no speech minutes)
maxSpeechMinutes (integer, default 60 — hard cap for the whole run)
maxItems (integer, default 100 — hard cap on videos)
Charges: transcript once per delivered transcript, captions or speech-to-text
speech_minute per started minute, only when speech-to-text actually ran
Build spec: https://apify.com/steadyfetch/youtube-transcript-scraper/api
Token: https://console.apify.com/settings/integrations
MCP: pin this actor in any MCP client with https://mcp.apify.com?tools=steadyfetch/youtube-transcript-scraper
curl -X POST "https://api.apify.com/v2/acts/steadyfetch~youtube-transcript-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H 'Content-Type: application/json' \
-d '{"videoUrls":["https://www.youtube.com/watch?v=jNQXAC9IVRw","https://youtu.be/dQw4w9WgXcQ"],"format":"srt"}'

Calling from an agent or MCP client: every optional field accepts an explicit null and reads it as "use the default", so a template that renders one body per run and leaves the unset fields null still runs — including a body where every optional field is null. videoUrls is the one required field and needs a real value. If you send an output format we do not recognise, the run transcribes nothing, charges nothing, and returns one row naming the formats we do support.

Sending the same list under a sibling actor's field name works too: urls and channels are read as this actor's video list, and one uncharged note row tells you the field is called videoUrls here. A channel link inside either field still gets the row that points you at steadyfetch/youtube-channel-transcripts, which is the actor that walks a whole channel.

MCP: one link pins this actor in Claude, Cursor, or any MCP client — https://mcp.apify.com?tools=steadyfetch/youtube-transcript-scraper — or ask Apify's MCP server for "youtube transcript".


Watch links, youtu.be links, /shorts/, /live/, /embed/, m.youtube.com, music.youtube.com, youtube-nocookie.com, and bare 11-character video IDs. Casing in the link does not matter — chat apps and link shorteners rewrite it constantly — but the video ID itself is case-sensitive, because YouTube treats it that way.

A channel link gets an uncharged row pointing you at steadyfetch/youtube-channel-transcripts, which does whole channels in one run. A playlist link gets an uncharged row too — paste the individual video links here instead.

What you are charged for

Pricing: from $1.20/1,000 transcripts. Two events, and nothing else:

  • Transcript — once per video that returns real text, from captions or from speech-to-text.
  • Speech-to-text minute — per started minute, and only when a video had no usable captions so speech-to-text had to run. A captioned video never triggers it, and a caption-less video too long for the speech route (see audio_too_long_for_speech below) is reported uncharged instead of being charged for a partial transcript.

What can fail, and what it costs you: nothing. The two most common non-deliveries are no_speech (the audio is music or silence, so there is no transcript to sell) and blocked_retry (YouTube challenged or throttled the fetch, or was still processing the video). Both come back as a row with charged: false naming the reason — as do no_captions, private_or_members_only, removed_or_unavailable, age_restricted, region_blocked and live_no_transcript_yet. You pay for delivered transcripts and nothing else.

Switch "Use speech-to-text when a video has no captions" off and you will never be charged a speech-to-text minute; caption-less videos come back as uncharged rows instead. Max speech-to-text minutes and Max videos are hard stops, not suggestions: the run finishes successfully and each skipped row names the limit that stopped it.

You are never charged twice for the same video

Re-run the same list, poll a video every week, paste a link you already had months ago — a transcript this account has already been charged for is handed back from the run that produced it, with nothing charged for it a second time. Nothing is fetched and no speech-to-text minute is spent: the check happens before any of that.

Those rows carry repeat: true, charged: false, firstSeenAt (when you first got it) and firstSeenRunId (which run), and statusReason says the same thing in words. They do not count against Max videos either — your cap buys new videos.

The memory lives in your own account, in a key-value store called yt-transcripts-account on your Storage tab. Delete it to start over and be charged again. Transcripts drop out of it after 90 days on their own. If it cannot be read on a given run, the run still delivers — it charges as it always did and the status line says the check was unavailable, so you know that run could have billed a repeat.

When a video is not charged

statuswhat happenedtemporary?
no_speechthe audio is music or silence — there is no transcript to sellno
no_audio_streamthe video exposes no audio track at allno
no_captionsno captions, and you switched speech-to-text off for this runno
asr_unavailableour speech-to-text service refused this actor's access mid-run — that is on us, not you. Captioned videos still delivered; caption-less ones came back unchargedyes — try again later
audio_too_long_for_speechno captions, and the video is too long for the speech-to-text route — YouTube does not release enough of its audio for a complete transcript, and we never sell a partial one as whole. Captioned videos are unaffected at any lengthno
private_or_members_onlyprivate or members-onlyno
removed_or_unavailableYouTube says the video is gone; its own words are in the rowno
age_restrictedYouTube requires a signed-in, age-verified accountno
region_blockedthe uploader has not published it in the country we fetched fromno
live_no_transcript_yeta live stream — no transcript exists until it endsre-run after it ends
blocked_retryYouTube challenged, throttled or was still processingyes — re-run
skipped_too_large · skipped_budget · skipped_deadline · skipped_speech_capa limit stopped it before the video was fetched, and the row names which one: the file-size guard, your maximum cost per run, the run timeout, or the speech-minute cap (a transcript already in hand when the cost cap is reached is delivered instead, uncharged)yes
input_errorthe link was not a YouTube video linkfix and re-run
sample_noteyou set settings but no videos, so the sample ran under them — the row names what you set, and is not charged

A temporary problem is never reported as a permanent one. Bare video IDs are the one thing we cannot sanity-check: a typo in an 11-character ID is indistinguishable from a real ID that has been deleted, so it comes back as removed_or_unavailable, uncharged.

This actor may fail when the platform changes things — failed items are never charged.

FAQ

How do I get a YouTube transcript without an API key? Paste the video links and run it. There is no YouTube API key, no cookies and no sign-in.

Can I download YouTube subtitles as SRT? Set format to srt (or vtt) and each row carries a ready-to-save subtitle string.

What if a video has no captions? Speech-to-text runs on the audio and you get real text, marked source: "speech_ai". If the audio has no speech at all, the row comes back no_speech and uncharged. The speech route works on short videos only: past a few minutes YouTube stops releasing the audio to anything but its own player, so a long caption-less video comes back audio_too_long_for_speech and uncharged rather than half-transcribed. Videos that have captions — the large majority, including nearly every spoken upload — are unaffected at any length.

Can I pick the caption language? Set language. If that language is not published for the video, the default track is used and the row's language field tells you what you actually got.

Can I transcribe a whole channel? Use steadyfetch/youtube-channel-transcripts — paste a channel URL, @handle or channel ID and it returns every video's transcript in one run. Playlists are not supported by either actor yet; for a playlist, paste its video links here.

Can I use this through an MCP server? Yes. It is a standard Apify actor, so any MCP client that can call Apify actors can call it.

Why does a run cost more than the transcripts? Apify bills platform usage (compute and proxy) for what a run actually consumes, separately from these events. Maximum cost per run is the ceiling that covers both.


Steadyfetch YouTube suite

Same transcript engine, different way in. All-inclusive pay per event, no start fee, charged only on delivery.

What you pasteActor
Video URLs or IDsthis actor
A channel URL, @handle or channel IDYouTube Channel Transcript Scraper — Every Video, Shorts & Live

The rest of the steadyfetch shelf — same contract everywhere: all-inclusive pay per event, no start fee, charged only on delivery.

FamilyActors
Ad creative intelligenceFacebook · Google Ads video · TikTok · LinkedIn · Google Ads text & OCR
Trends & keywordsGoogle Trends · Trends Now · Breakout keywords · Autocomplete keywords · Keyword volume & CPC · Social trends
YouTube transcriptsYouTube videos · YouTube channels
InstagramReel transcripts · Profile posts
JobsIndeed · Career sites by domain · Glassdoor · Multi-board · Google Jobs
AmazonProducts · Search · Bestsellers · Sellers
Any media fileSpeech to Text · any link or file

Feedback & support

Found an issue? Open it on the Issues tab — we usually reply within a couple of hours, always within a day. And if this actor earned its keep, a rating helps other buyers find it, and saving it keeps it one click away.