YouTube Transcript Scraper - Any Video or Whole Channel
Pricing
from $0.50 / 1,000 videos
YouTube Transcript Scraper - Any Video or Whole Channel
Transcripts for one YouTube video or a whole channel, with views, dates, comments and Shorts in the same run - one Actor, not three. No API key, no quota. Exact counts, not 3.1M. Cached, failed and filtered rows are free, so re-runs cost almost nothing.
Pricing
from $0.50 / 1,000 videos
Rating
0.0
(0)
Developer
Lorenzo
Maintained by CommunityActor stats
0
Bookmarked
4
Total users
2
Monthly active users
13 days ago
Last modified
Categories
Share
Download YouTube Transcripts — Every Video from a Channel
Get the transcript of a YouTube video, or of every video a channel has ever published, as a spreadsheet. Full text, the language it was written in, and timed segments if you want them. No API key, no quota, nothing to install.
Paste a channel, a playlist, a search term or a single video URL. Transcripts, view counts, publish dates, likes and comments all come back from the same run, so one job gives you the transcript and everything around it.
{ "inputs": ["https://www.youtube.com/@veritasium"],"includeTranscript": true, "transcriptLanguages": ["en"] }
Videos with subtitles switched off come back as NO_CAPTIONS, not blank —
and you are not charged for them.
Callable from any MCP client. Claude, Cursor, VS Code and anything else that
speaks Model Context Protocol can find and run this through Apify's MCP server
at mcp.apify.com — search-actors to discover it, call-actor to run it. No
setup on your side and no separate integration.
Built for pipelines that run more than once
Most scrapers charge you the same on Monday as they did on Sunday. This one remembers what it already fetched.
Measured on 27 August 2026 — one channel, 500 videos, same input both times:
| first run | run it again | |
|---|---|---|
| rows delivered | 500 | 500 |
| videos charged | 489 | 11 |
| requests to YouTube | 525 | 30 |
| time | 174s | 46s |
Videos already fetched inside the cache window come back from stored state and are not charged. Eleven were charged the second time because those eleven had failed the first — you are not billed for a row that came back empty, and not billed again for one you already have.
And it is one Actor, not three. Video metadata, transcripts and comments
arrive from the same run, in the same dataset, on one bill. No joining three
outputs on video_id and no three integrations to keep alive.
What you get back
| video_type | title | views | published_at | duration_s |
|---|---|---|---|---|
video | Google Pixel 11/Pro/Fold Impressions: It Is What It Is | 3138780 | 2026-08-12T07:00:35-07:00 | 673 |
short | This ZOOM is Insane! | 3567393 | 2026-07-20T13:39:08-07:00 | 66 |
stream | 5 New Phone Updates + Giveaway Update! | 940608 | 2017-11-29T09:15:12-08:00 | 2501 |
Real rows from a real run. Exact counts, not 3.1M. Full timestamps, not
7 days ago. Both survive sorting.
How do I download the transcript of every video on a channel?
{ "inputs": ["https://www.youtube.com/@veritasium"],"includeTranscript": true, "transcriptLanguages": ["en", "es"] }
Each row gains the transcript as text, its language, whether it was auto-generated, and how many timed cues it has.
Timed cues are a separate switch: includeTranscriptSegments. Turn it on
and each row also carries transcript_segments — one entry per cue with
start_s, duration_s and text — for citing a moment in the video. It costs
nothing extra, because the cues arrive in the same request as the text. It is
off by default because every export format flattens them: a 1,000-cue video
becomes transcript_segments/0/text … /999/text, and a spreadsheet of a few
videos turns into thousands of columns. transcript_cue_count ships either way.
transcriptLanguages is a preference order and human-written subtitles beat
auto-generated ones. Subtitles off returns NO_CAPTIONS, free.
includeLikes adds a like count from one small extra request, not charged
separately.
What a row actually looks like
One video, with transcripts, timed segments and comments all switched on. Real
output from a real run — long fields trimmed with …, nothing else changed.
{"_row_type": "video","video_id": "o4SSoURPODY","video_type": "video","title": "Google Pixel 11/Pro/Fold Impressions: It Is What It Is","channel": "Marques Brownlee","channel_id": "UCBJycsmduvYEL83R_U4JriQ","views": 3148445,"likes": 84617,"published_at": "2026-08-12T07:00:35-07:00","category": "Science & Technology","duration_s": 673,"description": "Every year, a new Pixel, and new hopes and dreams... …","keywords": ["Pixel 11", "Pixel 11 Pro", "Pixel 11 Pro Fold", "…"],"thumbnail": "https://i.ytimg.com/vi_webp/o4SSoURPODY/maxresdefault.webp","subscriber_count": "21.1M subscribers","is_live": false,"is_private": false,"is_crawlable": true,"transcript": "[music] >> You know, I've been using Pixel phones for a long time. …","transcript_segments": [{ "start_s": 2.619, "duration_s": 2.741, "text": "[music]" }, "…"],"transcript_cue_count": 331,"transcript_language": "en","transcript_is_generated": true,"transcript_status": "OK","_source_url": "https://www.youtube.com/watch?v=o4SSoURPODY","_cached": false,"_error_kind": null}
A comment from the same run — its own row, not a nested field:
{"_row_type": "comment","video_id": "o4SSoURPODY","comment_id": "Ugz7VRjqhsV-60jGjwR4AaABAg","comment_text": "Adding a physical RGB LED to the back of the phone just to …","comment_author": "@Mosesplusofficial","comment_author_channel_id": "UCr3n8BfGuDIzhO8ALIxMe4w","comment_author_is_verified": false,"comment_published_text": "8 days ago","comment_like_count_text": "13K","comment_reply_count_text": "200","comment_is_reply": false}
Three things that make this different
🟢 Failures are typed, never blank
Subtitles off returns
NO_CAPTIONS. A broken fetch returnsFETCH_FAILED. Different values, so you always know which happened.
🟢 Transcripts and comments in the same run
One actor, one pass. Transcript text, timed segments, comments and replies — alongside the metadata, not in a separate job.
🟢 Failed, filtered and cached rows cost nothing
Charged only when a row has data. A deleted video is free. So is one your filter removed, or one served from cache.
How do I scrape all videos from a YouTube channel?
{ "inputs": ["https://www.youtube.com/@mkbhd"], "maxResultsPerSource": 500 }
@handle, a channel URL, or a /channel/UC… ID all work.
You get all three tabs. Long videos, Shorts and live streams sit in separate
tabs that do not overlap — one channel measured 30, 48 and 5. video_type says
which is which, and the tabs are interleaved so asking for 30 returns a mix.
channelTabs narrows it.
Channel rows also carry subscriber_count and channel_description, free —
same request that lists the videos.
What happens when something fails?
You get a typed reason in the row, never an empty cell.
_error_kind | means |
|---|---|
REJECTED_INPUT | not a recognised channel, playlist, video or search |
RESOLVE_FAILED | the listing could not be read |
NO_VIDEOS | it was read and contained none |
FETCH_FAILED | the video request failed |
NO_DATA | it loaded but carried no video details |
Blocked requests retry automatically on a fresh IP. In a 36-video run that turned 3 blocks into 0 failures.
How do I export a YouTube playlist to CSV?
{ "inputs": ["https://www.youtube.com/playlist?list=PLbpi6ZahtOH6Blw3RGYpWkSByi_T7Rygb"],"maxResultsPerSource": 100 }
Then hit Export — CSV, JSON, Excel, or an API endpoint. Playlists return 100 videos per request, so multiples of 100 go furthest.
How do I get YouTube comments with replies?
{ "inputs": ["https://www.youtube.com/@mkbhd"], "includeComments": true,"maxCommentsPerVideo": 100, "maxRepliesPerComment": 5 }
Comments arrive as their own rows, not extra columns — _row_type is
video or comment. Folding 100 comments into a video row would export as 600
unusable columns.
Replies are rows too (comment_is_reply: true), capped because each comment
with replies costs an extra request.
How do I search YouTube and export the results?
{ "inputs": ["drone review", "sourdough starter"] }
One quirk: a bare 11-character word is read as a video ID, because that is
exactly how long a YouTube ID is — helicopters is eleven characters. Write
search:helicopters to force a search.
Can I filter, or work in another language?
| input | effect |
|---|---|
publishedAfter / publishedBefore | YYYY-MM-DD, inclusive |
durationMinSeconds | set 61 to exclude Shorts |
durationMaxSeconds | upper bound in seconds |
language | en, de, ja, es — text YouTube generates |
region | US, DE, GB — how results are ranked |
Filters run after each video is read, so they are exact — and filtered-out videos are never charged.
subscriber_count follows the language as YouTube wrote it: 21.1M subscribers,
21,1 Mio. Abonnenten, チャンネル登録者数 2110万人. Titles, descriptions and
category are unaffected.
What people use this for
📈 Track a competitor's channel every week
- Run with the competitor's channel and
maxResultsPerSource: 50,maxAgeHours: 0so view counts are fresh. - Schedule it weekly in the Apify Console.
- Join runs on
video_idand diffviewsto get per-video growth, andpublished_atto get posting cadence.
Set maxAgeHours: 0 here — the cache is a time window, and stale view counts
would flatten exactly the trend you are measuring.
🔔 Watch a channel and get only what is new
- Run with the channel and
publishedWithinDays: 7. - Schedule it weekly in the Apify Console. That is the whole setup.
publishedWithinDays is a rolling window, counted from the day each run
starts — so a saved task keeps returning what is new instead of the same seven
days forever. publishedAfter is an absolute date and is the wrong tool for a
schedule: set once, it returns the same videos every week until someone edits it.
If you set both, the explicit date wins and the run log says which applied.
This is the cheapest thing this Actor does. Everything it has already fetched comes back from stored state and is not charged, and anything outside the window is filtered and free. Measured on a 500-video channel: the first run charged 489 videos, the second charged 11.
Use publishedWithinDays a little wider than your schedule — 8 or 9 days for a
weekly run — so a late-publishing channel or a missed run does not leave a hole.
How this differs from a YouTube RSS feed. Every channel has a free RSS feed
at youtube.com/feeds/videos.xml?channel_id=…, and if all you need is a title
and a link, use it — it is free and instant. It gives you the most recent 15
videos and nothing else: no transcript, no view count, no duration, no
description, no comments, and no way to reach further back. This Actor returns
new videos with the full record, transcripts included, for any number of
channels at once.
🤖 Build an LLM corpus from a channel's transcripts
- Run with the channel,
includeTranscript: true,maxResultsPerSource: 500. - Export as JSON and use
transcriptfor plain text. AddincludeTranscriptSegments: truewhen you need timestamps to cite back to a moment in the video — it is free, but it makes the rows much wider. - Filter on
transcript_status == "OK"—NO_CAPTIONSmeans that video never had subtitles, and it cost you nothing.
transcript_is_generated tells you whether a human wrote it. Auto-generated
captions carry more errors, which matters if you are grounding an answer on
them.
💬 Find what an audience actually asks about
- Run with the channel,
includeComments: true,maxCommentsPerVideo: 100, andmaxRepliesPerComment: 5to catch the threads under the popular ones. - Filter rows to
_row_type == "comment". - Group by
video_id, sort bycomment_like_count_text, read the top of each.
Useful for support-content planning, sponsorship research, and finding the question a whole audience keeps asking that nobody has answered yet.
What fields do I get?
| field | example |
|---|---|
video_id · video_type | dQw4w9WgXcQ · video, short, stream |
title · channel · channel_id | Rick Astley - Never Gonna Give You Up · Rick Astley · UCuAXFkgsw1L7xaCfnd5JJOw |
views · likes | 1806183998 · 136841 |
published_at · duration_s | 2009-10-24T23:57:33-07:00 · 213 |
category · keywords · thumbnail · description | Music · array · highest-res URL · full text |
subscriber_count · channel_description | from channel inputs |
is_live · is_private · is_crawlable | booleans |
| every row | _row_type _input _source_url _site _cached _cached_at _error_kind _error |
| with transcripts | transcript transcript_cue_count transcript_language transcript_is_generated transcript_status |
with includeTranscriptSegments | transcript_segments — one entry per cue, start_s duration_s text |
| comment rows | comment_id comment_text comment_author comment_author_channel_id comment_author_is_verified comment_author_is_creator comment_published_text comment_like_count_text comment_reply_count_text comment_is_reply |
Three comment fields end in _text because YouTube publishes them rounded and
relative — 13K, 7 days ago — and they pass through as given. 13K is not
expanded to 13000: it could be 13,499, and inventing digits hands you a wrong
number that looks precise.
Not included: no sort order, no video or audio downloads, no total comment count per video.
Does it re-fetch videos I already pulled?
Any video fetched in the last 24 hours comes back from stored state — and is
not charged. Marked _cached: true with _cached_at.
This is a time window, not change detection. It checks how long ago a video
was read, not whether it changed — so set maxAgeHours to 0 when sorting on
views or measuring growth. The window is per actor: yesterday's run counts for
today's.
Cached rows still follow this run's settings. Turn transcripts on and a cached video gets its transcript fetched and charged; turn them off and the row carries no transcript fields at all, whatever an earlier run stored. The cache never adds a field you did not ask for, and never withholds one you did.
What does it cost?
| you get | you pay |
|---|---|
| a video row | $0.50 / 1,000 |
| a transcript | $10.00 / 1,000 |
| a comment or reply | $0.20 / 1,000 |
| a failed, filtered or cached row | nothing |
Transcripts and comments are billed separately — each is a separate request to YouTube, not an extra field.
Send everything in one run, not one run each: fixed per-run cost is ~86% of a
single-video run and ~6% at a hundred. Requests run 8 at a time, one per second
each; concurrency adjusts that between 1 and 16.
Will it quietly break?
Every run emits a schema fingerprint — which fields resolved, which came back empty, the row count — compared against a stored baseline. A field that starts returning empty is caught as drift, not discovered months later.
Scrapers rarely break loudly. They thin out.
Frequently asked questions
How do I get a transcript from a YouTube video?
Paste the video URL into inputs and set includeTranscript to true. You get
the full transcript as text in the transcript field, along with its language
and whether YouTube auto-generated it. No API key and no browser extension.
{ "inputs": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],"includeTranscript": true }
How do I get the script of a YouTube video?
The same way — what YouTube calls captions or subtitles is what most people mean
by the script. transcript holds it as one block of text, ready to paste into a
document or feed to a model.
How do I download YouTube transcripts for a whole channel at once?
Give it the channel instead of the video. inputs accepts an @handle, a
channel URL or a /channel/UC… ID, and maxResultsPerSource decides how many
videos deep to go. Every video comes back as a row with its own transcript.
Can I get the transcript with timestamps?
Yes. Turn on includeTranscriptSegments and each row also carries
transcript_segments — one entry per cue with start_s, duration_s and
text — so you can cite the exact moment something was said. It costs nothing
extra. It is off by default because spreadsheets flatten it into thousands of
columns; transcript_cue_count ships either way.
What happens if a video has no subtitles?
transcript_status comes back as NO_CAPTIONS and you are not charged for
that video's transcript. It is a different value from FETCH_FAILED, so you
can always tell "this video has no subtitles" apart from "we could not reach
it" — and re-run only the second kind.
Do I need the YouTube Data API or an API key?
No. There is no key to obtain, no project to create, and no daily quota to run out of. That is the main reason people use this instead of the official API.
Is there a free way to do this?
Not here — this Actor is paid per transcript. What it does give you is that failed, filtered and cached rows cost nothing, so you only pay for transcripts you actually receive, and never twice for the same video inside the cache window.
Can I get transcripts in another language?
transcriptLanguages is a preference order, for example ["en", "es"].
Human-written subtitles are preferred over auto-generated ones when both exist.
transcript_language and transcript_is_generated tell you which you got.
Can I use this from Claude, Cursor or another AI assistant?
Yes. It is exposed through Apify's MCP server at mcp.apify.com, so any client
that speaks Model Context Protocol can discover it with search-actors and run
it with call-actor under the name aquixlabs/youtube-channel-scraper. Nothing
to install and nothing to configure — every public Apify Actor is reachable this
way, and this one is not in either excluded category (it is
LIMITED_PERMISSIONS, and pay-per-event rather than rental).
A useful shape for an agent: ask it for the transcript of a video or a channel, and it can call this directly rather than scraping the page itself.
Can I get the comments as well as the transcript?
Yes, in the same run. Set includeComments and each comment arrives as its own
row — not a column — tagged _row_type: comment. The dataset's Comments
view is where to read them.
Terms
You are responsible for your use of the data you collect, including under YouTube's Terms of Service and any applicable data protection law.
More from aquixlabs
Other actors from this account will be listed here.