# Threads Scraper - Extract Posts, Profiles & Engagement Data (`mellowed_headpiece/threads-scraper`) Actor

Extract public posts, profiles, and engagement data from Threads by Meta. Scrape profile feeds, individual threads, images, videos, likes, and reply counts, no login required. Perfect for brand monitoring, influencer analytics, AI training data, and competitor research. Export to JSON, CSV, or Excel

- **URL**: https://apify.com/mellowed\_headpiece/threads-scraper.md
- **Developed by:** [Trevor Smith](https://apify.com/mellowed_headpiece) (community)
- **Categories:** Social media, Automation, Lead generation
- **Stats:** 3 total users, 2 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## 🧵 Threads Scraper — Extract Posts, Profiles & Engagement Data

**The first and only Threads web scraper on the Apify Store.** Extract public posts, profile feeds, media, and engagement metrics from **Threads by Meta** — the fastest-growing social platform with 275M+ monthly active users.

***

### 🚀 Why Threads?

Threads is exploding. Brands, researchers, marketers, and AI companies are desperate for Threads data — but **nobody has built a scraper yet**. You're looking at the first mover advantage.

| Use Case | Who Needs It |
|----------|-------------|
| 📊 **Brand Monitoring** | Marketing teams tracking mentions and sentiment |
| 🧠 **AI Training Data** | LLM companies need fresh conversational data |
| 🔍 **Competitor Research** | Social media managers analyzing rival strategies |
| 📈 **Influencer Analytics** | Agencies tracking engagement rates |
| 🏷️ **Trend Detection** | Journalists and trend forecasters |
| 💼 **Lead Generation** | Sales teams finding prospects in conversations |

***

### ✨ Features

| Feature | Description |
|---------|-------------|
| 🔍 **Keyword Search** | Search Threads by keyword or hashtag — find posts about any topic |
| 📥 **Profile Scraping** | Extract all posts from any public Threads profile |
| 🔗 **Post URLs** | Scrape individual threads by direct URL |
| 💬 **Reply Thread Extraction** | Extract full conversation trees under any post |
| ♾️ **Infinite Scroll** | Automatically loads more posts as you scroll |
| 🖼️ **Media Extraction** | Get high-res image URLs (srcset) and video URLs from posts |
| 👤 **Profile Metadata** | Bio, follower count, display name — automatically captured |
| 📅 **Timestamps** | Every post includes exact creation date |
| ❤️ **Engagement Metrics** | Like, reply, and repost counts on every post |
| 📋 **Structured Output** | Clean JSON/CSV/Excel via Apify dataset |
| 📊 **Analytics Dashboard** | Free interactive dashboard with charts (see below) |
| ⏰ **Scheduled Runs** | Monitor profiles or keywords daily/weekly via Apify scheduler |
| 🔌 **Integrations** | Zapier, Make, n8n, LangChain, LlamaIndex |
| 🛡️ **Proxy Support** | Residential proxies to avoid rate limiting |

***

### 🔧 How to Use

#### 1️⃣ Quick Start (One Profile)

```
Input:
  usernames: ["zuck"]
  maxPostsPerProfile: 20
```

↓ **Output (JSON)**

```json
{
  "id": "C6Tt8xRRJXA",
  "authorUsername": "zuck",
  "text": "Threads now has 275M+ monthly active users! Thanks for building this community with us 🧵",
  "url": "https://www.threads.net/@zuck/post/C6Tt8xRRJXA",
  "createdAt": "2025-06-15T14:30:00.000Z",
  "likeCount": 124500,
  "replyCount": 8450,
  "repostCount": 3200,
  "imageUrls": ["https://...threads.net/media.jpg"],
  "type": "post",
  "profile": "zuck",
  "scrapedAt": "2025-06-15T15:00:00.000Z"
}
```

#### 2️⃣ Multiple Profiles + Replies

```
Input:
  usernames: ["zuck", "nike", "nasa", "elonmusk"]
  maxPostsPerProfile: 50
  includeReplies: true
  includeMediaInfo: true
```

#### 3️⃣ Search by Keyword

```
Input:
  searchKeywords: ["AI", "crypto", "photography"]
  maxPostsPerProfile: 30
  includeReplies: true
```

Returns all recent public posts matching each keyword, with full reply threads.

#### 4️⃣ Individual Posts

```
Input:
  postUrls: [
    "https://www.threads.net/@zuck/post/C6Tt8xRRJXA",
    "https://www.threads.net/@nike/post/C6Tt7xRRJXB"
  ]
```

***

### 📦 Output Fields

| Field | Type | Description |
|-------|------|-------------|
| `id` | string | Unique post identifier |
| `authorUsername` | string | Threads username (without @) |
| `text` | string | Post caption/content |
| `url` | string | Direct link to the post |
| `createdAt` | string | ISO 8601 timestamp |
| `likeCount` | integer | Number of likes |
| `replyCount` | integer | Number of replies |
| `repostCount` | integer | Number of reposts |
| `imageUrls` | array | List of image URLs |
| `videoUrls` | array | List of video URLs |
| `type` | string | "post" or "reply" |
| `profile` | string | Source profile |
| `replies` | array | Reply objects (if includeReplies=true) |
| `scrapedAt` | string | When data was extracted |

***

### 💰 Pricing

| Tier | Price | What You Get |
|------|-------|-------------|
| 🆓 **Free** | $0 | 10 profile runs per month (up to 20 posts each) |
| 🚀 **Starter** | **~$2/1,000 posts** | Regular extraction for small teams |
| 📊 **Pro** | Custom volume | Enterprise-scale monitoring |

*Platform compute costs are separate — the Scraper is optimized for efficiency.*

***

### 📊 Free Analytics Dashboard

Turn your scraped Threads data into **interactive charts and insights** — no signup, no backend, just paste your Dataset ID.

👉 **[Launch Dashboard →](https://titanmishmannah2000-cmd.github.io/threads-dashboard/)**

```
┌─────────────────────────────────────────────────────────┐
│  🧵 Threads Analytics                          📊 ONLINE │
├─────────────────────────────────────────────────────────┤
│  ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ │
│  │ 30   │ │ 12.5K│ │ 3.2K │ │ 1.8K │ │ 583  │ │ 15   │ │
│  │Posts │ │Likes │ │Replies│ │Repost│ │AvgEng│ │Media │ │
│  └──────┘ └──────┘ └──────┘ └──────┘ └──────┘ └──────┘ │
│                                                         │
│  📈 Engagement Over Time           🍩 Engagement Split  │
│  ╱╲    ╱╲                                                │
│ ╱  ╲  ╱  ╲                          ● Likes    75%     │
│╱    ╲╱    ╲                         ● Replies  18%     │
│                                     ● Reposts   7%     │
│  🔥 Top Posts                      📅 Posting Days     │
│  ┌─────────────────────┐           ██ ██ ██ █■■ ██    │
│  │"Just launched..."   │❤️12K     Mon Tue Wed Thu Fri  │
│  │"Big news today..."  │❤️8K                           │
│  │"New research..."    │❤️5K       📋 All Posts        │
│  └─────────────────────┘    Date │ Text │ ❤️ │ 💬 │ 🔄│
│                             06/23│"New AI..."│12K│4K│1K│
│  🖼️ Media Gallery              06/16│"500M..."│18K│5K│1K│
│  [img] [img] [img] [img]                               │
└─────────────────────────────────────────────────────────┘
```

**How to use:**

1. Run the Threads Scraper → copy the **Dataset ID** from the run's Storage tab
2. Open the **[Dashboard](https://titanmishmannah2000-cmd.github.io/threads-dashboard/)** in your browser
3. Paste the Dataset ID → click **Load Dashboard**
4. See 5 interactive charts + posts table + media gallery

*For private datasets, also paste your Apify API token (Settings → API). Or make the dataset public in Storage tab.*

***

### 🔗 Integrations

| Platform | Integration |
|----------|-------------|
| 🦜🔗 **LangChain / LlamaIndex** | Feed Threads data directly into RAG pipelines |
| ⚡ **Zapier / Make** | Connect to 5,000+ apps |
| 🔧 **n8n** | Self-hosted automation workflows |
| 🐍 **Python / JavaScript SDK** | Build custom applications |
| 🤖 **AI Agents** | MCP server support for AI agent data pipelines |

***

### ⚠️ Data Compliance

This Actor extracts **publicly available** Threads profile data — the same information visible to anyone visiting threads.net. Respect Threads' Terms of Service and Meta's Automated Data Collection Terms when using extracted data.

***

### 🧑‍💻 Need Help?

- [Apify Documentation](https://docs.apify.com)
- [Discord Community](https://discord.gg/apify)
- [Report an Issue](https://apify.com/your-username/threads-scraper/issues)

***

**Built with 🧵 for the Apify community.** First Threads scraper on the market — grab the opportunity while it's fresh!

# Actor input Schema

## `usernames` (type: `array`):

Threads usernames (without @) to scrape. Each profile's posts will be extracted.

## `postUrls` (type: `array`):

Full URLs to specific Threads posts (e.g. https://www.threads.net/@zuck/post/C6Tt8xRRJXA). Leave empty if scraping entire profiles.

## `maxPostsPerProfile` (type: `integer`):

Maximum number of posts to extract from each profile.

## `searchKeywords` (type: `array`):

Search Threads by keyword — finds recent public posts matching each term. Combine with usernames or use standalone.

## `includeReplies` (type: `boolean`):

For each scraped post, also extract all replies/comments. More data but slower runs.

## `includeMediaInfo` (type: `boolean`):

Extract image and video URLs from posts.

## `scrollTimeoutSecs` (type: `integer`):

How long to keep scrolling for new posts.

## `extractProfileMeta` (type: `boolean`):

Also extract profile info: bio, follower count, display name.

## `proxy` (type: `object`):

Use Apify proxy (recommended with RESIDENTIAL group) or your own.

## Actor input object example

```json
{
  "usernames": [
    "zuck"
  ],
  "postUrls": [],
  "maxPostsPerProfile": 20,
  "searchKeywords": [],
  "includeReplies": false,
  "includeMediaInfo": true,
  "scrollTimeoutSecs": 30,
  "extractProfileMeta": true,
  "proxy": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `id` (type: `string`):

No description

## `authorUsername` (type: `string`):

No description

## `text` (type: `string`):

No description

## `url` (type: `string`):

No description

## `createdAt` (type: `string`):

No description

## `likeCount` (type: `string`):

No description

## `replyCount` (type: `string`):

No description

## `repostCount` (type: `string`):

No description

## `imageUrls` (type: `string`):

No description

## `videoUrls` (type: `string`):

No description

## `type` (type: `string`):

No description

## `profile` (type: `string`):

No description

## `scrapedAt` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "usernames": [
        "zuck"
    ],
    "postUrls": [],
    "searchKeywords": [],
    "proxy": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("mellowed_headpiece/threads-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "usernames": ["zuck"],
    "postUrls": [],
    "searchKeywords": [],
    "proxy": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("mellowed_headpiece/threads-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "usernames": [
    "zuck"
  ],
  "postUrls": [],
  "searchKeywords": [],
  "proxy": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call mellowed_headpiece/threads-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,mellowed_headpiece/threads-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/mDQsFUdNh6VvMW5V8/builds/a33LlxcRZXQ0Gp5cn/openapi.json
