# Youtube Transcript Scraper (`coregent/youtube-transcript-scraper`) Actor

Lightning-fast transcript extraction with pay-per-result pricing.

Extract comprehensive transcript data from YouTube videos using official APIs. Get paragraph-formatted transcript text, timed segments, and metadata with 15 complete fields in just 1-2 seconds per video.

- **URL**: https://apify.com/coregent/youtube-transcript-scraper.md
- **Developed by:** [Delowar Munna](https://apify.com/coregent) (community)
- **Categories:** SEO tools, Social media, Videos
- **Stats:** 49 total users, 7 monthly users, 100.0% runs succeeded, 6 bookmarks
- **User rating**: 5.00 out of 5 stars

## Pricing

from $4.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## YouTube Transcript Scraper 📝

**Lightning-fast transcript extraction with pay-per-result pricing**

Extract comprehensive transcript data from YouTube videos without browser automation. Get

#### ⭐ Why Choose This Scraper?

- 💰 **Pay Per Result**: Only pay for successful transcript extractions
- ⚡ **Fastest**: 1-2 seconds per video (3-4x faster than competitors)
- 🎯 **Most Reliable**: 99%+ success rate, never blocked by YouTube
- 📊 **Most Complete**: 26 comprehensive fields with paragraph formatting
- 🚀 **No Commitment**: No monthly fees, use when you need it
- 💡 **Perfect For**: All use cases from occasional to high-volume extraction

> **📊 High-Volume Users (10,000+ transcripts/month)?** Contact us for enterprise pricing and volume discounts!

***

![YouTube Transcript Scraper](https://raw.githubusercontent.com/coregentdevspace/youtube-transcript-scraper-assets/main/youtube-transcript-scraper-thumbnail.png)

### 🚀 Key Features

- ⚡ **Lightning Fast**: 1-2 seconds per video (330x faster than translation-enabled tools)
- 📝 **Paragraph Formatting**: Transcript text formatted with natural paragraph breaks (~40 words each)
- 🌍 **Multi-Language Support**: Auto-detects transcript language in 30+ languages
- 🎯 **Manual & Auto Captions**: Extracts both human-created and auto-generated transcripts
- 📊 **26 Complete Fields**: Comprehensive data with metadata and timed segments
- 🚀 **API-Only Architecture**: No browser automation = faster, more reliable, no blocking issues
- 🎬 **Video Metadata**: Complete video information (title, channel, duration, views, likes)
- 🔍 **Bulk Discovery**: Scrape whole channels or playlists, drop in a **seed video** to grab its entire channel, or enter **search keywords** to pull the top matching videos — all auto-expanded into videos, with per-source caps (and date filters for channels/playlists/seeds; keywords return by relevance)
- 📦 **Multiple URL Formats**: Supports full URLs, short URLs (youtu.be), and raw video IDs
- 📄 **Subtitle Ready**: Timed segments with millisecond precision for SRT/VTT generation
- 🌐 **Formatted Language Display**: Shows "English (en)" instead of "en" for better readability

> **Best for**: Content analysis, SEO optimization, accessibility, research, subtitle generation

***

### 🎯 At a Glance

| Feature         | Value                                                       |
| --------------- | ----------------------------------------------------------- |
| **Speed**       | ~1-2s per video (API-only, no browser)                      |
| **Throughput**  | 30-60 transcripts/minute (1,800-3,600/hour)                 |
| **Fields**      | 26 complete fields (100% reliability for transcripts)       |
| **Formatting**  | Paragraph breaks (~40 words each, \n\n separators)         |
| **Segments**    | Timestamped (millisecond precision)                         |
| **Architecture**| API-only (no browser automation)                            |
| **Concurrency** | 20 parallel requests (optimized automatically)              |

***

### 💡 Why This Scraper?

Traditional YouTube transcript tools rely on browser automation which is slow and unreliable. YouTube Transcript Scraper uses **direct API calls** for maximum speed and reliability:

| Metric                     | YouTube Transcript Scraper       | Traditional Browser-Based Tools |
| -------------------------- | -------------------------------- | ------------------------------- |
| **Architecture**           | ✅ API-only (fast, reliable)     | ❌ Browser automation (slow)    |
| **Time per video**         | ~1-2s                            | ~3-8s (browser overhead)        |
| **YouTube blocking**       | ✅ Never blocked (API access)    | ❌ Often blocked (bot detection)|
| **Fields extracted**       | 26 complete fields               | 5-10 fields                     |
| **Paragraph formatting**   | ✅ Built-in (~40 words/para)     | ❌ Raw text only                |
| **Language display**       | ✅ "English (en)" formatting     | ❌ "en" only                    |
| **Timestamp precision**    | Milliseconds                     | Sometimes missing               |
| **Transcripts per minute** | 30-60                            | 10-20 (browser limits)          |
| **Reliability**            | 99%+ (API-based)                 | 70-85% (blocking, errors)       |

**Performance Advantages:**

- **No Browser Overhead**: Direct API access = 3x faster than browser-based extraction
- **No Blocking**: APIs never trigger YouTube's bot detection
- **Complete Data**: Full transcript text + structured segments with timestamps
- **High Reliability**: 99%+ success rate for videos with transcripts
- **Business Ready**: All data needed for content analysis, SEO, accessibility

***

### 📋 Input Parameters

| Field                  | Key           | Type          | Default | Description                                                  |
| ---------------------- | ------------- | ------------- | ------- | ------------------------------------------------------------ |
| **Video URLs or IDs**  | `videoRefs`   | Array<string> | `[]`    | **Video Input section.** Specific videos to scrape — watch URLs, short URLs (youtu.be), or raw 11-char IDs. Only the videos listed here are scraped. |
| **Discovery Sources** | `channelPlaylistRefs` | Array<string> | `[]` | **Channels, Playlists & Discovery section.** Channel URLs (`/@handle`, `/channel/UC…`, `/c/…`, `/user/…`), playlist URLs (`…?list=…`), **seed video** URLs/IDs, or **search keywords**. Channels/playlists expand into their videos; a seed video is resolved to its owning channel and that channel's videos are scraped (not just the seed); a keyword returns the top matching videos from YouTube search (relevance order; date filters don't apply to keywords). To scrape only a specific video, use the **Video URLs or IDs** input instead. |
| **Transcript language**| `language`    | String (select) | `""` (Auto-detect) | Optional. Choose a specific caption language (e.g. English, Bengali). Leave on **Auto-detect** to return each video's original language. If a chosen language isn't available, the scraper auto-detects the video's original transcript as a fallback. |
| **Subtitle formats**   | `subtitleFormats` | String (select) | `"both"` | Which subtitle strings to include: `both` (SRT + VTT, default), `srt`, `vtt`, or `none`. Generated locally (no extra cost). The unselected `srt`/`vtt` field is set to `null` (fields always present). |
| **Max videos per source** | `maxVideosPerSource` | Integer | `50` | For each discovery source (channel, playlist, seed-video-derived channel, or search keyword), cap how many videos to include. Cost guardrail (each video = one paid result). |
| **Published after**    | `startDate`   | String        | `""`    | Optional `YYYY-MM-DD`. For channel, playlist, and seed-video expansion, only videos published on/after this date. Does not apply to search keywords. |
| **Published before**   | `endDate`     | String        | `""`    | Optional `YYYY-MM-DD`. For channel, playlist, and seed-video expansion, only videos published on/before this date. Does not apply to search keywords. |
| **Max videos (total)** | `maxResults`  | Integer       | `0`     | Optional. Cap the total number of videos processed this run (after expansion). `0` = no cap. |

**Important Notes:**

- 📝 **Two separate inputs**: Use **Video URLs or IDs** (`videoRefs`) for specific videos, and **Discovery Sources** (`channelPlaylistRefs`) for whole channels, playlists, or a seed video's channel. Fill either or both.
- 📝 **Video Formats**: Accepts watch URLs, short URLs (youtu.be), or raw video IDs
- 📺 **Discovery Mode**: Drop a channel URL, playlist URL, or seed video into `channelPlaylistRefs`. Channels/playlists auto-expand into their videos; a seed video is resolved to its channel and that channel is scraped. Expanding a source costs you nothing extra — you still pay only per transcript. Use `maxVideosPerSource` + `startDate`/`endDate` to control how many and which. ⚠️ A video URL here scrapes its **whole channel** — for a single video, use **Video URLs or IDs** instead.
- 📊 **Video metadata always included**: Title, channel, description, thumbnail, views, likes, comments, and duration are returned on every run at no extra charge — no toggle needed.
- 🌍 **Language Auto-Detection**: By default, automatically detects and extracts transcripts in the video's original language
- 🎯 **Optional Language Selection**: Pick a specific language from the `language` dropdown; it costs the same as auto-detect and falls back to auto-detect
- ⚡ **Performance**: Optimized for speed and reliability with API-only architecture
- 🎯 **Zero Configuration**: Works immediately - no setup required

***

### 🔧 How It Works

YouTube Transcript Scraper uses a **modern API-only architecture** for maximum reliability and speed:

#### What runs on each video

1. **Transcript extraction**
   - Manual and auto-generated captions, with broad language coverage
   - Fast extraction (~1-2s per video)
   - No browser automation required
   - Multiple independent extraction routes with automatic failover

2. **Video metadata + discovery**
   - Title, channel, description, thumbnail, duration, views, likes, comments
   - Also powers **Discovery Sources** — expanding channels, playlists, seed videos, and search keywords into video lists

3. **Subtitle generation**
   - SRT and WebVTT built locally from the timed segments, at no additional cost

#### Why API-Only Architecture?

Traditional browser-based scrapers face many challenges:

- ❌ YouTube bot detection and blocking
- ❌ Slow page loading and rendering
- ❌ High resource usage (Chrome instances)
- ❌ Frequent 429 rate limit errors
- ❌ Complex proxy management

**Our API-only approach eliminates all these issues:**

- ✅ **No Blocking**: Sidesteps the page-level bot detection that stops browser-based scrapers
- ✅ **3x Faster**: No browser overhead
- ✅ **Lower Costs**: No proxy or browser infrastructure needed
- ✅ **99%+ Reliability**: Direct API access
- ✅ **Scalable**: Handle high-volume requests easily

***

### 📤 Output Schema

#### Comprehensive Transcript Data: 26 Complete Fields

| #   | Field                      | Type                    | Description                                              |
| --- | -------------------------- | ----------------------- | -------------------------------------------------------- |
| 1   | **type**                   | String (const: "video") | Record type for filtering                                |
| 2   | **videoId**                | String                  | YouTube video ID (11 characters)                         |
| 3   | **PageURL**                | String                  | Full YouTube watch URL                                   |
| 4   | **title**                  | String                  | Video title                                              |
| 5   | **description**            | String or null          | Video description                                        |
| 6   | **thumbnailUrl**           | String or null          | Highest-resolution video thumbnail URL                   |
| 7   | **channelId**              | String                  | Channel ID (UC...)                                       |
| 8   | **channelName**            | String                  | Channel title                                            |
| 9   | **channelUrl**             | String                  | Channel URL (from channel ID)                            |
| 10  | **transcriptLanguage**     | String                  | Detected language, display form (e.g., "English (en)")   |
| 11  | **transcriptLanguageCode** | String                  | Raw BCP-47 language code (e.g., "en", "bn")              |
| 12  | **transcriptType**         | Enum: "manual" / "auto" | Manual vs auto-generated captions                        |
| 13  | **transcriptText**         | String                  | Full transcript with paragraph formatting (\n\n breaks)  |
| 14  | **segments**               | Array<Object>           | Timed segments (startMs, durMs, text)                    |
| 15  | **srt**                    | String or null          | SRT subtitles (included by default; controlled by `subtitleFormats`) |
| 16  | **vtt**                    | String or null          | WebVTT subtitles (included by default; controlled by `subtitleFormats`) |
| 17  | **availableLanguages**     | Array<String>           | Caption languages available for the video                |
| 18  | **hasTranscript**          | Boolean                 | Whether transcript was successfully extracted            |
| 19  | **status**                 | Enum: "success" / "no\_transcript" / "error" | Extraction outcome                  |
| 20  | **errorMessage**           | String or null          | Failure reason (null on success)                         |
| 21  | **durationSec**            | Number                  | Video duration in seconds                                |
| 22  | **publishedAt**            | String (ISO 8601)       | Video publish date                                       |
| 23  | **fetchedAt**              | String (ISO 8601)       | Timestamp of extraction                                  |
| 24  | **viewCount**              | Number                  | Total views                                              |
| 25  | **likeCount**              | Number                  | Total likes                                              |
| 26  | **commentCount**           | Number or null          | Total comments                                           |

#### Segment Structure

Each segment in the `segments` array contains:

| Field       | Type   | Description                       |
| ----------- | ------ | --------------------------------- |
| **startMs** | Number | Segment start time (milliseconds) |
| **durMs**   | Number | Segment duration (milliseconds)   |
| **text**    | String | Segment text content              |

**Why These Fields Matter:**

- 📝 **Complete Transcript**: Paragraph-formatted text + structured segments with precise timestamps
- 🌍 **Language Intelligence**: Auto-detection with formatted display ("English (en)")
- 🎬 **Video Context**: Title, channel, duration, views, likes for complete picture
- 📄 **Subtitle Generation**: Segments with timestamps for SRT/VTT creation
- 💼 **Business-Ready**: All data needed for content analysis, SEO, accessibility

***

### 📊 Output Examples

#### Output Table View - English Video

![Output Table - English Video](https://raw.githubusercontent.com/coregentdevspace/youtube-transcript-scraper-assets/main/youtube-transcript-scraper-output-table-english.png)

#### Output Table View - Bengali Video

![Output Table - Bengali Video](https://raw.githubusercontent.com/coregentdevspace/youtube-transcript-scraper-assets/main/youtube-transcript-scraper-output-table-bengali.png)

#### Example Output - Complete Transcript Data (JSON)

**English Video Example:**

```json
[{
  "type": "video",
  "videoId": "eLqveVYFWc4",
  "PageURL": "https://www.youtube.com/watch?v=eLqveVYFWc4",
  "title": "Example Video Title",
  "description": "Example video description...",
  "thumbnailUrl": "https://i.ytimg.com/vi/eLqveVYFWc4/maxresdefault.jpg",
  "channelId": "UCbo-KbSjJDG6JWQ_MTZ_rNA",
  "channelName": "Example Channel",
  "channelUrl": "https://www.youtube.com/channel/UCbo-KbSjJDG6JWQ_MTZ_rNA",
  "transcriptLanguage": "English (en)",
  "transcriptLanguageCode": "en",
  "transcriptType": "auto",
  "transcriptText": "This is the full transcript text with paragraph formatting...\n\n(Full transcript continues with natural paragraph breaks)",
  "segments": [
    {
      "startMs": 80,
      "durMs": 2799,
      "text": "First segment text"
    },
    {
      "startMs": 1280,
      "durMs": 3280,
      "text": "Second segment text"
    },
    {
      "startMs": 2880,
      "durMs": 2559,
      "text": "Third segment text"
    }
  ],
  "srt": "1\n00:00:00,080 --> 00:00:02,879\nFirst segment text\n\n2\n00:00:01,280 --> 00:00:04,560\nSecond segment text\n...",
  "vtt": "WEBVTT\n\n00:00:00.080 --> 00:00:02.879\nFirst segment text\n...",
  "availableLanguages": ["en"],
  "hasTranscript": true,
  "status": "success",
  "errorMessage": null,
  "durationSec": 769,
  "publishedAt": "2025-01-20T10:00:00.000Z",
  "fetchedAt": "2025-10-28T05:21:51.641Z",
  "viewCount": 148755,
  "likeCount": 6463,
  "commentCount": 512
}]
```

**Bengali Video Example:**

```json
[{
  "type": "video",
  "videoId": "UmmMLV5BaKQ",
  "PageURL": "https://www.youtube.com/watch?v=UmmMLV5BaKQ",
  "title": "Breaking: বেরিয়ে আসছে ...বিমানবন্দরে আগুনের গোপন কাহিনী | বিশ্লেষক: আমিরুল মোমেনীন মানিক",
  "description": "Bengali video description...",
  "thumbnailUrl": "https://i.ytimg.com/vi/UmmMLV5BaKQ/maxresdefault.jpg",
  "channelId": "UCWzOfBhVmuRmAd5bzxXP2CA",
  "channelName": "Change TV",
  "channelUrl": "https://www.youtube.com/channel/UCWzOfBhVmuRmAd5bzxXP2CA",
  "transcriptLanguage": "Bengali (bn)",
  "transcriptLanguageCode": "bn",
  "transcriptType": "auto",
  "transcriptText": "আসসালামু আলাইকুম প্রিয় দর্শক আমিরুল মুমিনন মানিক আপনাদের প্রত্যেককে আমন্ত্রণ জানাচ্ছি চেঞ্জ টিভির লাইভ পর্যালোচনায় প্রিয় দর্শক আজ 18ই অক্টোবর 2025 এই মুহূর্তে রাত 8টা বেজে 59 মিনিট আমাদের আলোচনার বিষয় হচ্ছে...\n\n(Full Bengali transcript continues with paragraph formatting)",
  "segments": [
    {
      "startMs": 2560,
      "durMs": 4400,
      "text": "আসসালামু আলাইকুম প্রিয় দর্শক আমিরুল"
    },
    {
      "startMs": 4960,
      "durMs": 5040,
      "text": "মুমিনন মানিক আপনাদের প্রত্যেককে আমন্ত্রণ"
    },
    {
      "startMs": 6960,
      "durMs": 6960,
      "text": "জানাচ্ছি চেঞ্জ টিভির লাইভ পর্যালোচনায়"
    }
  ],
  "srt": "1\n00:00:02,560 --> 00:00:06,960\nআসসালামু আলাইকুম প্রিয় দর্শক আমিরুল\n...",
  "vtt": "WEBVTT\n\n00:00:02.560 --> 00:00:06.960\nআসসালামু আলাইকুম প্রিয় দর্শক আমিরুল\n...",
  "availableLanguages": ["bn"],
  "hasTranscript": true,
  "status": "success",
  "errorMessage": null,
  "durationSec": 904,
  "publishedAt": "2025-10-18T15:26:57Z",
  "fetchedAt": "2025-10-27T12:34:01.574Z",
  "viewCount": 913141,
  "likeCount": 11498,
  "commentCount": 1847
}]
```

### 🎬 Quick Start

#### Example 1: List of YouTube Video URLs

```json
{
  "videoRefs": [
    "https://www.youtube.com/watch?v=DOtJEwVsJic",
    "https://www.youtube.com/watch?v=eLqveVYFWc4",
    "https://www.youtube.com/watch?v=gYXaPTDatis",
    "https://www.youtube.com/watch?v=ip8FEYOQob0"
  ]
}
```

#### Example 2: Video ID List

```json
{
    "videoRefs": [
        "dQw4w9WgXcQ",
        "jNQXAC9IVRw",
        "5oAnKSCP4do"
    ]
}
```

#### Example 3: Mix of Video URLs and IDs

```json
{
  "videoRefs": [
    "iG9CE55wbtY",
    "UyyjU8fzEYU",
    "https://youtu.be/jNQXAC9IVRw",
    "https://www.youtube.com/watch?v=Cm_Juzt9H2o"
  ]
}
```

#### Example 4: Discovery Sources (channels, playlists, seed videos — auto-expanded)

Put channel URLs, playlist URLs, seed video URLs/IDs, or search keywords in the `channelPlaylistRefs` (**Discovery Sources**) input. Channels/playlists expand into their videos; a seed video is resolved to its owning channel and that channel's videos are scraped (not just the seed); a keyword returns the top matching videos from YouTube search. Use `maxVideosPerSource` and the date range to control scope. Date filters apply to channels, playlists, and seed videos — not to keywords. You can still list specific videos in `videoRefs` at the same time.

```json
{
  "videoRefs": [
    "dQw4w9WgXcQ"
  ],
  "channelPlaylistRefs": [
    "https://www.youtube.com/@vidIQ",
    "https://www.youtube.com/playlist?list=PLBCF2DAC6FFB574DE",
    "https://youtu.be/jNQXAC9IVRw",
    "ai automation for solopreneurs"
  ],
  "maxVideosPerSource": 10,
  "startDate": "2025-01-01"
}
```

#### Example 5: Choose Transcript Language & Subtitle Format

Pick a specific caption language and control which subtitle strings are returned. Setting `language`
costs no more than auto-detect, and falls back to auto-detect if that language is not available.
Set `subtitleFormats` to `srt`, `vtt`, `both` (default), or `none`.

```json
{
  "videoRefs": [
    "https://www.youtube.com/watch?v=eLqveVYFWc4",
    "https://youtu.be/jNQXAC9IVRw"
  ],
  "language": "en",
  "subtitleFormats": "srt"
}
```

***

### 💪 Performance & Benchmarks

#### Speed Benchmarks

| Video Length | Segments | Processing Time | Total Fields |
|--------------|----------|-----------------|--------------|
| **5 minutes**  | ~200     | ~1-2 seconds    | 26           |
| **10 minutes** | ~400     | ~1-2 seconds    | 26           |
| **15 minutes** | ~600     | ~1-2 seconds    | 26           |
| **30 minutes** | ~1200    | ~2-3 seconds    | 26           |
| **60 minutes** | ~2400    | ~3-4 seconds    | 26           |

#### Throughput Comparison

| Videos         | YouTube Transcript Scraper | Traditional Browser Scrapers |
| -------------- | -------------------------- | ---------------------------- |
| **10 videos**  | ~10-20 seconds             | ~30-60 seconds               |
| **50 videos**  | ~1-2 minutes               | ~5-8 minutes                 |
| **100 videos** | ~2-3 minutes               | ~10-15 minutes               |
| **500 videos** | ~10-15 minutes             | ~40-70 minutes               |
| **1,000 videos** | ~20-30 minutes           | ~80-140 minutes              |

**Why So Fast?**

- ✅ API-only (no browser startup/rendering)
- ✅ Parallel processing (20 concurrent requests)
- ✅ No YouTube blocking or retries
- ✅ Direct API access to transcript data

***

### 📚 Use Cases

#### Content Analysis & Research

- **NLP Analysis**: Extract transcripts for sentiment analysis, topic modeling, keyword extraction
- **Market Research**: Analyze competitor video content at scale
- **Academic Research**: Study video content patterns, themes, and trends
- **Content Summarization**: Generate summaries from full transcript text

#### SEO & Marketing

- **SEO Optimization**: Convert video content to text for search engine indexing
- **Content Repurposing**: Transform video transcripts into blog posts, articles, social media
- **Keyword Research**: Analyze transcript text for keywords and themes
- **Competitive Analysis**: Study competitor video strategies through transcript analysis

#### Accessibility & Subtitles

- **Accessibility**: Create text versions of video content for hearing-impaired users
- **Subtitle Generation**: Create SRT/VTT subtitle files from transcript segments
- **Caption Management**: Extract and manage closed captions for video platforms
- **Multi-Platform**: Use transcripts across different platforms and applications

#### Content Creation

- **Video Editing**: Use timestamps to find exact moments in videos
- **Clip Creation**: Identify key segments for social media clips
- **Script Analysis**: Study successful video scripts and patterns
- **Quality Control**: Review video content at scale

***

### ❓ FAQ

**Q: Do I need to provide API keys?**
A: No! The actor handles all API integrations automatically. Just provide video URLs.

**Q: What if a video doesn't have transcripts?**
A: The scraper returns a full-shape record with `hasTranscript: false`, `status: "no_transcript"`, and an `errorMessage`. (Records without a transcript do not include video metadata.)

**Q: What video URL formats are supported?**
A: All formats:

- Full URL: `https://www.youtube.com/watch?v=dQw4w9WgXcQ`
- Short URL: `https://youtu.be/dQw4w9WgXcQ`
- Video ID only: `dQw4w9WgXcQ`

**Q: Can I scrape a whole channel or playlist instead of listing every video?**
A: Yes — use the **Discovery Sources** input. Paste a channel URL (`/@handle`, `/channel/UC…`, `/c/…`, `/user/…`), a playlist URL (`…?list=…`), a single **seed video** URL/ID, or plain **search keywords**, and the actor auto-expands each into videos and scrapes a transcript for each. Seed videos are resolved to their owning channel, so one video from a channel is enough to pull the whole channel; search keywords return the top matching videos from YouTube search. Control scope with **Max videos per source** and the **Published after/before** date filters. ⚠️ Search keywords are **not** affected by the date filters and are always returned in YouTube's relevance order (ordering isn't configurable) — the date filters apply only to channels, playlists, and seed videos. Channel/playlist/seed expansion is free; keyword search uses a small YouTube-search call under the hood (no daily quota limit) — you still only pay per transcript result.

**Q: What's the difference between "Video URLs or IDs" and "Discovery Sources"?**
A: **Video URLs or IDs** scrapes only the exact videos you list (one transcript each). **Discovery Sources** expands each entry into many videos — a channel's uploads, a playlist's items, a seed video's whole channel, or a search keyword's top results. ⚠️ A video URL placed in Discovery Sources scrapes its **entire channel**; to scrape just that one video, put it in **Video URLs or IDs** instead.

**Q: How accurate are the timestamps?**
A: Timestamps are millisecond-precision and come directly from YouTube's official transcript data, ensuring high accuracy for subtitle generation.

**Q: Does it work with live streams or premieres?**
A: Only after the video is published and transcripts are available. Live streams and ongoing premieres typically don't have transcripts yet.

**Q: How much does it cost to run?**
A: You're billed at the actor's Store price per result — that's the whole cost to you. There are no API keys to buy, no subscriptions, and no surcharge for picking a specific language, requesting SRT/VTT, or expanding a channel or playlist into its videos.

**Q: Can I translate the transcripts?**
A: The actor extracts transcripts in the video's original language. For translation, you can use the transcript data with external translation services.

***

### 🛠️ Why It Is Built This Way

The scraper talks to APIs directly instead of driving a headless browser:

- ✅ **API-Only**: No browser overhead, 3x faster than traditional scrapers
- ✅ **No Blocking**: Avoids the bot detection that stops browser-based scrapers
- ✅ **High Reliability**: 99%+ success rate, with automatic failover between extraction routes
- ✅ **Cost-Effective**: Optimal balance of speed, quality, and cost
- ✅ **Scalable**: Handles high-volume requests efficiently

***

### 📋 Best Practices

1. **Start Small**: Test with 3-5 videos before bulk processing
2. **Check Availability**: Not all videos have transcripts (check `hasTranscript` field)
3. **Language Auto-Detection**: Transcripts are extracted in the video's original language
4. **Paragraph Formatting**: Automatic paragraph breaks make transcripts more readable
5. **Subtitle Generation**: Use segments with timestamps for creating SRT/VTT files
6. **Export Formats**: Download as JSON, CSV, or Excel for further analysis
7. **Filter Results**: Filter by `hasTranscript: true` for analysis workflows
8. **Batch Processing**: Process videos in batches of 100-500 for optimal performance
9. **Monitor Costs**: Track your run cost from the Actor run detail page in Apify Console
10. **Backup Data**: Save extracted transcripts for future use

***

### 📜 Version

**v2.5.0** - Production Ready

**Current Features:**

- ✅ 26 comprehensive fields per video
- ✅ Paragraph formatting with natural breaks
- ✅ Multi-language support (30+ languages)
- ✅ Millisecond-precision timing
- ✅ API-only architecture (no browser)
- ✅ Formatted language display ("English (en)")
- ✅ Lightning-fast extraction (1-2 seconds per video)
- ✅ 99%+ reliability, with automatic failover between extraction routes

***

### 🤝 Compliance

- Intended for **legitimate content analysis, SEO, accessibility, and research**
- Extracts only **publicly available** YouTube transcript data
- Reads only data that YouTube already serves publicly for each video
- Designed for **content repurposing and business intelligence**
- Respects YouTube's Terms of Service
- Users responsible for compliance with applicable laws in their jurisdiction

***

### 💬 Support

- **Issues**: Report via Apify support or GitHub
- **Feature Requests**: Contact us with your use case
- **Documentation**: Comprehensive examples and guides included

***

**Built with ❤️ for lightning-fast transcript extraction and content analysis**

# Actor input Schema

## `videoRefs` (type: `array`):

Specific videos to scrape. Accepts video URLs (watch or youtu.be) or raw 11-char video IDs. Only the videos you list here are scraped — one transcript per video. To scrape whole channels, playlists, or a seed video's channel, use the "Discovery Sources" input below instead.

## `channelPlaylistRefs` (type: `array`):

Sources to discover videos from. Accepts four kinds of entries: (1) channel URLs (youtube.com/@handle, /channel/UC..., /c/..., /user/...), (2) playlist URLs (...?list=...), (3) seed video URLs/IDs, and (4) plain search keywords (e.g. "ai automation tutorial"). Channels and playlists expand into their videos; a seed video is resolved to its owning channel and that channel's videos are scraped (not just the seed video); a search keyword returns the top matching videos from YouTube search. A video URL here scrapes its whole channel — to scrape only a specific video, use the "Video URLs or IDs" input above instead. Each source is capped by "Max videos per source". ⚠️ Search keywords are NOT affected by the Published after/before date filters and are always returned in YouTube's relevance order (ordering is not configurable); the date filters apply only to channels, playlists, and seed videos.

## `maxVideosPerSource` (type: `integer`):

Cap how many videos to take from each discovery source (channel, playlist, seed-video-derived channel, or search keyword). Does not affect videos listed in the "Video URLs or IDs" input.

## `startDate` (type: `string`):

ISO date (YYYY-MM-DD). For channel, playlist, and seed-video expansion, only include videos published on or after this date. Does not apply to search-keyword sources. Leave empty for no lower bound.

## `endDate` (type: `string`):

ISO date (YYYY-MM-DD). For channel, playlist, and seed-video expansion, only include videos published on or before this date. Does not apply to search-keyword sources. Leave empty for no upper bound.

## `language` (type: `string`):

Language of the transcript to return. Leave on "Auto-detect" to return each video's original transcript language. Choosing a specific language returns that caption track when the video has it, otherwise falls back to the video's original transcript.

## `subtitleFormats` (type: `string`):

Which ready-to-use subtitle strings to include in each record (generated locally from the timed segments — no extra API cost). "SRT + VTT" is the default; pick one to slim the output, or "None" to omit both. The unselected field is set to null — the srt/vtt fields always exist so every record keeps the same shape.

## `maxResults` (type: `integer`):

Cap the total number of videos processed this run, across both inputs (after channel/playlist expansion). 0 = no cap.

## Actor input object example

```json
{
  "videoRefs": [
    "https://youtu.be/jNQXAC9IVRw",
    "https://www.youtube.com/watch?v=Cm_Juzt9H2o"
  ],
  "channelPlaylistRefs": [],
  "maxVideosPerSource": 50,
  "startDate": "",
  "endDate": "",
  "language": "",
  "subtitleFormats": "both",
  "maxResults": 0
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "videoRefs": [
        "https://youtu.be/jNQXAC9IVRw",
        "https://www.youtube.com/watch?v=Cm_Juzt9H2o"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("coregent/youtube-transcript-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "videoRefs": [
        "https://youtu.be/jNQXAC9IVRw",
        "https://www.youtube.com/watch?v=Cm_Juzt9H2o",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("coregent/youtube-transcript-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "videoRefs": [
    "https://youtu.be/jNQXAC9IVRw",
    "https://www.youtube.com/watch?v=Cm_Juzt9H2o"
  ]
}' |
apify call coregent/youtube-transcript-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,coregent/youtube-transcript-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/nfc95j75lfycE6q88/builds/IChCCIvOugfQL0W8l/openapi.json
