# YouTube Transcript Scraper - Bulk + Multi-language (`dltik/youtube-transcript-scraper`) Actor

Extract YouTube transcripts in bulk: any public video, manual + auto-generated captions, multi-language fallback. Outputs full text + segments with timestamps. HTTP-only, no API key. Pay $0.005/transcript.

- **URL**: https://apify.com/dltik/youtube-transcript-scraper.md
- **Developed by:** [Walid](https://apify.com/dltik) (community)
- **Categories:** AI, Developer tools, Marketing
- **Stats:** 10 total users, 2 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$25.00 / 1,000 transcript fetcheds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## YouTube Transcript Scraper — Bulk Extract Captions, Multi-language, with Timestamps

> Extract **YouTube transcripts in bulk** — any public video, any channel. Manual + auto-generated captions, multi-language fallback, full text + segments with timestamps. HTTP-only via [yt-dlp](https://github.com/yt-dlp/yt-dlp), better fingerprinting than `youtube-transcript-api`. No API key, no quota. **$0.005 per transcript** ($5 per 1,000).

⭐ **Bookmark this YouTube Transcript Scraper** — Apify ranks actors by bookmarks, so it directly helps the visibility of this scraper on the Apify Store.

### What is the YouTube Transcript Scraper?

The **YouTube Transcript Scraper** is an Apify actor that downloads YouTube video transcripts (subtitles / closed captions) in bulk, fast, and cheap. Drop in a list of YouTube URLs or video IDs and get back the full transcript text plus timestamped segments. Multi-language with fallback — request French, get auto-translated English if French isn't available.

Built on `yt-dlp` instead of the deprecated `youtube-transcript-api` Python lib, so it has better session/cookie management, lower 429 rate, and tolerates DC IPs in most cases. Add the Apify Residential proxy if you hit blocks on aggressive videos.

### Use cases

- **AI training data** — pull transcripts of 1,000+ podcast videos to fine-tune an LLM.
- **Content marketing research** — extract transcripts of competitor YouTube videos and analyze tone, keywords, hooks.
- **Multilingual subtitles** — get auto-generated captions in your target language for video re-localization.
- **Searchable video archives** — turn a 500-video channel into a full-text searchable corpus.
- **Podcast / interview pipelines** — feed YouTube transcripts into Claude / GPT for summarization, highlight extraction, social-media clip generation.
- **SEO research** — find which keywords appear in top-ranking video transcripts.

### Input

```json
{
  "videos": [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
    "https://youtu.be/jNQXAC9IVRw",
    "EZ8RFXNQK2g"
  ],
  "languages": ["en", "fr", "es"],
  "proxyConfig": { "useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"] }
}
```

The `videos` field accepts full URLs (`youtube.com/watch?v=…`, `youtu.be/…`, `youtube.com/shorts/…`, `youtube.com/embed/…`) or raw 11-character video IDs.

### Output

```json
{
  "video_id": "dQw4w9WgXcQ",
  "url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
  "language": "en",
  "is_auto_generated": false,
  "transcript_text": "We're no strangers to love. You know the rules and so do I...",
  "segments": [
    { "start": 18.4, "duration": 3.6, "text": "We're no strangers to love" },
    { "start": 22.0, "duration": 3.2, "text": "You know the rules and so do I" }
  ],
  "char_count": 1842,
  "word_count": 311
}
```

### Pricing

**PAY\_PER\_EVENT — $0.005 per transcript fetched** (= $5 per 1,000). Failed fetches are not charged. Compute and (optional) residential proxy bandwidth are billed by Apify on top — typically <$0.001/transcript with HTTP-only mode.

### FAQ — YouTube Transcript API alternatives

**Why use this scraper vs `youtube-transcript-api` (Python lib)?** That library breaks frequently when YouTube rotates internal endpoints, and it has zero anti-fingerprinting. This scraper uses `yt-dlp`, which is the most-maintained YouTube client (it powers most YouTube downloaders). It bypasses the issues `youtube-transcript-api` users hit weekly.

**Does it work without a proxy?** Often, yes — yt-dlp's session management bypasses most DC blocks. For aggressive videos (recently uploaded, high-view, or live-streamed), enable Apify Residential proxy to maintain >95% success rate.

**Does it support YouTube Shorts and live streams?** Yes — Shorts URLs (`youtube.com/shorts/…`) and live VOD pages are auto-detected.

**Bulk extraction — how fast?** Around 2-3 transcripts per second per actor instance. For 10K transcripts, run with 4-8 concurrent actor runs.

**Multi-language fallback — how does it pick?** It tries each language in your `languages` list in order; if none has a manual transcript, it falls back to auto-generated, then to auto-translated. The chosen language is in the `language` output field.

***

⭐ **Found this useful? Bookmark this YouTube Transcript Scraper** — it's the strongest signal for Apify Store ranking.

#### Related actors

- [Substack Scraper](https://apify.com/dltik/substack-scraper) — for written content + sentiment
- [HackerNews MCP Server](https://apify.com/dltik/mcp-server-hackernews) — tech-content analysis
- [Pappers MCP Server](https://apify.com/dltik/mcp-server-pappers) — French B2B intelligence

License: MIT · Author: [dltik](https://apify.com/dltik)

- [Vimeo Scraper](https://apify.com/dltik/vimeo-scraper) - Video metadata extractor
- [Dailymotion Scraper](https://apify.com/dltik/dailymotion-scraper) - Video metadata extractor

# Actor input Schema

## `urls` (type: `array`):

List of YouTube video URLs or 11-char video IDs. Supports youtube.com/watch, youtu.be, /shorts/, /live/, /embed/.

## `preferredLanguages` (type: `array`):

ISO 639-1 codes in priority order. Tries manual transcripts first, then auto-generated. Falls back to any available.

## `proxyConfig` (type: `object`):

YouTube blocks datacenter IPs. Apify Residential proxy is enabled by default.

## Actor input object example

```json
{
  "urls": [
    "https://www.youtube.com/watch?v=dQw4W9WgXcQ",
    "https://youtu.be/jNQXAC9IVRw"
  ],
  "preferredLanguages": [
    "en",
    "fr",
    "es",
    "de"
  ],
  "proxyConfig": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.youtube.com/watch?v=dQw4W9WgXcQ",
        "https://youtu.be/jNQXAC9IVRw"
    ],
    "proxyConfig": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("dltik/youtube-transcript-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": [
        "https://www.youtube.com/watch?v=dQw4W9WgXcQ",
        "https://youtu.be/jNQXAC9IVRw",
    ],
    "proxyConfig": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("dltik/youtube-transcript-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.youtube.com/watch?v=dQw4W9WgXcQ",
    "https://youtu.be/jNQXAC9IVRw"
  ],
  "proxyConfig": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call dltik/youtube-transcript-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,dltik/youtube-transcript-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/WTRuLEc8vlvWU7Rfk/builds/XYlDXDLXcyN7VPkUI/openapi.json
