# Substack Scraper (`sheshinmcfly/substack-scraper`) Actor

Scrape posts from any Substack publication (subdomain or custom domain). Get title, subtitle, description, word count, reactions, restacks, comment counts, tags, authors, and publication metadata.

- **URL**: https://apify.com/sheshinmcfly/substack-scraper.md
- **Developed by:** [Sheshinmcfly](https://apify.com/sheshinmcfly) (community)
- **Categories:** News
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 1 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Substack Scraper

Scrape posts from any Substack publication — subdomain or custom domain. Get full post metadata: title, subtitle, description, word count, publish date, tags, authors, and engagement signals (reactions, restacks, comment count). No login, no account, no API key.

Give it one or more publications and get back their entire post archive with rich engagement data, ready for content analysis, competitive research, or newsletter monitoring.

### Data you get

**Per post**

| Field | Description |
|---|---|
| postId | Substack post identifier |
| title / subtitle / description | Post title, subtitle, and summary |
| slug / url | URL slug and canonical link |
| postDate | Publish date (ISO 8601) |
| type | `newsletter`, `podcast`, etc. |
| audience | `everyone` or `only_paid` |
| isPaid | Whether the post is paywalled |
| wordcount | Word count of the post |
| authors | List of author names |
| tags | Post tags |
| sectionName | Newsletter section, if any |
| reactionCount / reactions | Total reactions and breakdown by emoji |
| restacks | Number of restacks |
| commentCount | Number of comments |
| coverImage | Cover image URL |
| podcastUrl / podcastDuration | Podcast audio link and duration (if applicable) |
| bodyHtml | Full post HTML content (optional, opt-in) |

**Per comment** (optional, opt-in — includes nested replies)

| Field | Description |
|---|---|
| commentId / postId | Comment and parent post identifiers |
| body | Comment text |
| authorName / authorHandle / authorPhoto | Commenter info |
| date / editedAt | Comment timestamps |
| reactionCount / reactions | Reaction counts |
| restacks | Number of restacks |
| childrenCount | Number of replies |
| score | Comment ranking score |

**Per publication (metadata record)**

| Field | Description |
|---|---|
| title | Publication name |
| description | Publication description |
| coverImage | Publication cover/social image |

### How to use

1. Add one or more **publications** — use the exact domain the publication uses (e.g. `platformer.substack.com` for a Substack subdomain, or `www.lennysnewsletter.com` for a custom domain). Check the publication's own URL in your browser if unsure.
2. Choose the **sort order** (newest, top, or pinned).
3. Set **max posts per publication** and run.

### Use cases

- **Content research:** analyze a newsletter's posting frequency, topics, and engagement over time.
- **Competitive monitoring:** track competitor newsletters' output and reader engagement.
- **Creator economy analysis:** compare reactions, restacks, and comments across publications.
- **Content archiving:** build a structured archive of a publication's posts.
- **Lead generation:** identify active newsletter authors and publications in a niche.

### Input

| Field | Type | Description |
|---|---|---|
| publications | array | Substack domains to scrape (required) |
| sort | select | `new`, `top`, or `pinned` |
| maxPostsPerPublication | integer | Max posts per publication (default 50) |
| includePublicationMetadata | boolean | Output a publication-summary record |
| includeFullContent | boolean | Fetch full HTML content per post (extra request per post) |
| includeComments | boolean | Fetch all comments and replies per post (extra request per post) |

### Output example

```json
{
  "recordType": "post",
  "title": "Why founders should ignore most advice",
  "postDate": "2026-07-01T12:00:00.000Z",
  "wordcount": 2450,
  "reactionCount": 312,
  "restacks": 18,
  "commentCount": 45,
  "isPaid": false,
  "authors": ["Lenny Rachitsky"],
  "url": "https://www.lennysnewsletter.com/p/example-post"
}
```

### Performance & cost

Runs on lightweight requests (no browser), so it is fast and cheap: dozens to hundreds of posts per publication in seconds.

### Pricing

This actor uses **pay-per-result** pricing — you are charged per record returned. See the pricing shown on the actor page.

### Keywords

substack scraper, newsletter scraper, substack api, newsletter analytics, content scraper, creator economy data, substack posts, newsletter engagement data

### Related actors

- [Reddit Thread Scraper](https://apify.com/sheshinmcfly/reddit-thread-scraper)
- [Eventbrite Events Scraper](https://apify.com/sheshinmcfly/eventbrite-scraper)
- [App Store & Google Play Reviews Scraper](https://apify.com/sheshinmcfly/app-reviews-scraper)

### Disclaimer

This actor collects only publicly available post data from Substack. Use the data in compliance with the applicable terms and with data-protection laws in your jurisdiction.

### Changelog

| Version | Date | Changes |
|---|---|---|
| 1.0 | 2026-07-13 | Initial release — post archive with full engagement and content metadata |

# Actor input Schema

## `publications` (type: `array`):

Substack publications to scrape. Use the full domain the publication actually uses — either a subdomain (e.g. 'platformer.substack.com') or a custom domain (e.g. 'www.lennysnewsletter.com'). Check the publication's own URL in your browser if unsure.

## `sort` (type: `string`):

How to sort the post archive.

## `maxPostsPerPublication` (type: `integer`):

Maximum number of posts to fetch per publication.

## `includePublicationMetadata` (type: `boolean`):

Output a summary record per publication (title, description, cover image).

## `includeFullContent` (type: `boolean`):

Fetch the full HTML body of each post (adds one extra request per post — slower and increases usage).

## `includeComments` (type: `boolean`):

Fetch all comments (including nested replies) for each post as separate records (adds one extra request per post — slower and increases usage).

## Actor input object example

```json
{
  "publications": [
    "platformer.substack.com"
  ],
  "sort": "new",
  "maxPostsPerPublication": 50,
  "includePublicationMetadata": true,
  "includeFullContent": false,
  "includeComments": false
}
```

# Actor output Schema

## `dataset` (type: `string`):

All extracted post and publication records.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "publications": [
        "platformer.substack.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("sheshinmcfly/substack-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "publications": ["platformer.substack.com"] }

# Run the Actor and wait for it to finish
run = client.actor("sheshinmcfly/substack-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "publications": [
    "platformer.substack.com"
  ]
}' |
apify call sheshinmcfly/substack-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,sheshinmcfly/substack-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/kbe2JhbJGd7s1peVr/builds/P5WPgHuiAPNT1ZI3I/openapi.json
