Subreddit Scraper - Deep Crawl Posts, Media & Nested Comments avatar

Subreddit Scraper - Deep Crawl Posts, Media & Nested Comments

Pricing

from $1.60 / 1,000 results

Go to Apify Store
Subreddit Scraper - Deep Crawl Posts, Media & Nested Comments

Subreddit Scraper - Deep Crawl Posts, Media & Nested Comments

Extract complete subreddit feeds (Hot, New, Top, Rising, Controversial), historical post archives, media galleries, flairs, and nested discussion comment trees from any Reddit community. Download clean, structured JSON/CSV data with AI sentiment and content taxonomy enrichment.

Pricing

from $1.60 / 1,000 results

Rating

0.0

(0)

Developer

Mikolabs

Mikolabs

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

11 days ago

Last modified

Categories

Share

Subreddit Scraper โ€” Deep Crawl Posts, Media & Nested Comments

Extract complete subreddit feeds (Hot, New, Top, Rising, Controversial), historical post archives, media galleries, flairs, and nested discussion comment trees from any Reddit community. Download clean, structured JSON/CSV data with AI sentiment and content taxonomy enrichment.

Overview

Subreddit Scraper is an enterprise-grade data extraction tool built specifically for monitoring and deep-crawling entire Reddit communities. Provide one or more subreddit names (e.g. technology, startups, AskReddit), select your sorting order, and extract thousands of posts and discussion threads with automatic pagination.

When comment extraction is enabled, the actor also collects the full discussion thread for every post, including replies, author karma, upvote counts, and AI sentiment scoring.


Why Use Subreddit Scraper

  • Community & Niche Research: Analyze audience interests, common pain points, and viral topics across specific subreddits.
  • Brand & Product Monitoring: Track customer feedback, brand mentions, and product discussions inside industry communities.
  • Trend Forecasting: Monitor Rising and Hot feeds to spot emerging trends before they hit mainstream social media.
  • Content & SEO Ideation: Mine high-upvote posts (Top of the year/all-time) to generate engaging article and video ideas.
  • Longitudinal Datasets: Crawl historical subreddit archives for academic discourse research and market intelligence.

Pricing & Plans (No Hidden Fees)

Transparent and predictable pricing with no extra proxy costs, no setup fees, and no hidden maintenance charges.

Tiered Pricing Structure

Tier / Discount LevelPrice per 1,000 ItemsEffective SavingsMinimum Scrape
No Discount (Standard / Pay-As-You-Go)$4.00 / 1,000 itemsStandard Rate1 item
๐Ÿฅ‰ Bronze Discount$2.00 / 1,000 items50% OFF20 items
๐Ÿฅˆ Silver Discount$1.80 / 1,000 items55% OFF20 items
๐Ÿฅ‡ Gold Discount$1.60 / 1,000 items60% OFF20 items

Plan Comparison

FeatureFree TierSubscriber / Paid Tier
Free Daily Allowance20 items / run (4 runs / day free)Unlimited
Pricing$4.00 / 1,000 results (or free allowance)Down to $1.60 / 1,000 results
Additional Fees$0.00 (No extra fees)$0.00 (No extra fees)
Proxy / Bandwidth CostsIncluded ($0.00)Included ($0.00)
Deep Crawl Paginationโœ…โœ…
Comments Extraction per Postโœ…โœ…
AI Sentiment & Taxonomyโœ…โœ…
Granular Filtersโœ…โœ…
Run Summary Dashboardโœ…โœ…

Free users can extract up to 20 items per run (4 runs/day) completely free. Upgrade for volume discounts down to $1.60 / 1,000 items with zero hidden fees.


How to Use โ€” Step by Step

  1. Enter Subreddit Names: Input one or more subreddits in Subreddits to Scrape (e.g. ["startups", "SaaS"]).
  2. Choose Sort & Timeframe: Select feed sorting (hot, new, top, rising, controversial) and time window (all, year, month, week, day).
  3. Configure Comments Extraction (Optional): Toggle Scrape Comments for Each Post and specify how many comments per thread to collect.
  4. Set Filters & AI Analytics: Apply date ranges, score minimums, flair filters, and toggle AI Sentiment Analysis or AI Content Taxonomy.
  5. Click Start: Download results in JSON, CSV, Excel, XML, or HTML table format.

Input Parameters

ParameterTypeDefaultDescription
subredditsstring[]["technology", "startups"]List of subreddit names to scrape (without r/).
sortstringhotFeed sort: hot, new, top, rising, controversial.
timeframestringallTime window for top/controversial: all, year, month, week, day, hour.
fullSubredditModebooleanfalseDeep crawl mode โ€” paginates historical pages for maximum coverage.
maxPostsPerSubredditinteger100Maximum posts to collect per subreddit.
maxTotalItemsinteger500Safety ceiling for total items (posts + comments).
includeCommentsbooleanfalseExtract full comment threads for each collected post.
maxCommentsPerPostinteger25Maximum comments per post thread.
commentsSortstringconfidenceComment ranking order: confidence (Best), top, new, controversial, old, qa.
flattenCommentsbooleanfalseWhen true, pushes each comment as a separate dataset row.
sentiment_analysisbooleanfalseAdds sentiment score, confidence, and label to each post and comment.
content_analysisbooleanfalseClassifies posts against an enterprise topic taxonomy.
postTypestringallFilter by media format: all, text, image, video, gallery, link.
minScoreintegerโ€“Only keep posts with at least this upvote score.
minCommentsintegerโ€“Only keep posts with at least this many comments.
postsCreatedAfterstringโ€“Only keep posts created on or after date (YYYY-MM-DD).
postsCreatedBeforestringโ€“Only keep posts created on or before date (YYYY-MM-DD).
flairContainsstringโ€“Only keep posts matching this flair keyword.
titleContainsstringโ€“Only keep posts whose title contains this keyword.
excludeStickiedbooleanfalseExclude pinned moderator announcements.
excludeKeywordsstring[]โ€“Exclude posts containing any of these keywords.
includeNsfwbooleantrueInclude NSFW/18+ content in output.

Example Output: Post Record with Nested Comments

{
"kind": "post",
"id": "1hvoazn",
"title": "My best cheesecake so far",
"body": "Found my new favorite recipe (no water bath).",
"author": "ClearlyBulky",
"score": 3489,
"upvote_ratio": 1.0,
"num_comments": 43,
"subreddit": "Baking",
"created_utc": "2025-01-07T10:09:56.000Z",
"url": "https://www.reddit.com/r/Baking/comments/1hvoazn/my_best_cheesecake_so_far/",
"permalink": "/r/Baking/comments/1hvoazn/my_best_cheesecake_so_far/",
"flair": "Recipe",
"media_type": "gallery",
"sentiment_score": 2,
"sentiment_label": "positive",
"content_category_label": "Desserts & Baking",
"content_category_path": ["Food & Drink", "Desserts & Baking"],
"comments_count_scraped": 1,
"comments": [
{
"kind": "comment",
"id": "m5un6bj",
"author": "BakingFanatic",
"score": 76,
"depth": 0,
"body": "This looks absolutely incredible! Can you share the recipe?",
"sentiment_label": "positive"
}
]
}

API Access

from apify_client import ApifyClient
client = ApifyClient("YOUR_API_TOKEN")
run = client.actor("YOUR_ACTOR_ID").call(run_input={
"subreddits": ["startups", "SaaS"],
"sort": "top",
"timeframe": "month",
"includeComments": True,
"maxCommentsPerPost": 20,
"sentiment_analysis": True,
"content_analysis": True,
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(f"[{item.get('content_category_label')}] {item['title']} ({item['score']} upvotes)")

Support

For help or feature requests, use the Issues tab on the actor page in Apify Console.