# Trustpilot Scraper (`magicfingers/trustpilot-scraper`) Actor

Scrape Trustpilot company profiles, reviews, ratings, and category pages. Filter by star rating, date range, language, and verification status. Supports pagination for companies with thousands of reviews.

- **URL**: https://apify.com/magicfingers/trustpilot-scraper.md
- **Developed by:** [abdulrahman alrashid](https://apify.com/magicfingers) (community)
- **Categories:** Other
- **Stats:** 19 total users, 3 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: 1.00 out of 5 stars

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Trustpilot Scraper

Scrape company profiles, reviews, ratings, and category pages from [Trustpilot](https://www.trustpilot.com). Extract structured data for sentiment analysis, market research, competitor monitoring, and reputation tracking.

### Features

- **Search companies** by keyword or name
- **Scrape company profiles**: name, TrustScore, total reviews, category, location, website, response rate, claimed status, description
- **Scrape all reviews**: reviewer name, location, rating (1-5 stars), title, review text, date of experience, date posted, company reply, verified status, useful votes, language
- **Filter reviews** by star rating, date range, language, and verification status
- **Scrape category pages** (e.g., "Electronics & Technology")
- **Full pagination support** for companies with thousands of reviews
- **Three-layer extraction**: `__NEXT_DATA__` (primary) > JSON-LD structured data > HTML fallback
- **Anti-bot handling**: session rotation, rate limit detection, realistic browser headers

### Input Configuration

| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `action` | string | `scrapeReviews` | One of: `scrapeReviews`, `scrapeCompanyProfile`, `searchCompanies`, `scrapeCategory` |
| `companyUrls` | string\[] | `[]` | Company URLs or slugs (e.g., `https://www.trustpilot.com/review/example.com` or `example.com`) |
| `searchQuery` | string | | Keyword to search for companies |
| `categoryUrl` | string | | Full category page URL |
| `maxReviews` | integer | `100` | Max reviews per company (0 = unlimited) |
| `maxSearchResults` | integer | `50` | Max companies from search/category |
| `filterByStars` | integer\[] | `[]` | Only these star ratings (e.g., `[1, 2]`) |
| `filterDateFrom` | string | | Only reviews after this date (`YYYY-MM-DD`) |
| `filterDateTo` | string | | Only reviews before this date (`YYYY-MM-DD`) |
| `filterLanguage` | string | | Language code (e.g., `en`, `de`, `fr`) |
| `filterVerifiedOnly` | boolean | `false` | Only verified reviews |
| `includeCompanyProfile` | boolean | `true` | Output company profile with reviews |
| `proxyConfiguration` | object | Apify residential | Proxy settings |
| `maxConcurrency` | integer | `5` | Concurrent requests (1-20) |

### Usage Examples

#### Scrape reviews for a company

```json
{
    "action": "scrapeReviews",
    "companyUrls": ["https://www.trustpilot.com/review/amazon.com"],
    "maxReviews": 500,
    "filterByStars": [1, 2],
    "filterVerifiedOnly": true
}
```

#### Search for companies

```json
{
    "action": "searchCompanies",
    "searchQuery": "web hosting",
    "maxSearchResults": 100
}
```

#### Scrape company profile only

```json
{
    "action": "scrapeCompanyProfile",
    "companyUrls": ["amazon.com", "ebay.com", "walmart.com"]
}
```

#### Scrape a category page

```json
{
    "action": "scrapeCategory",
    "categoryUrl": "https://www.trustpilot.com/categories/electronics_technology",
    "maxSearchResults": 200
}
```

### Output Schema

Each result includes a `type` field: `company_profile`, `review`, `search_result`, or `category_result`.

#### Company Profile

```json
{
    "type": "company_profile",
    "companyName": "Amazon",
    "trustScore": 1.7,
    "totalReviews": 125432,
    "category": "Electronics & Technology",
    "location": "Seattle, WA, US",
    "website": "https://www.amazon.com/",
    "responseRate": 12,
    "claimed": true,
    "description": "...",
    "scrapedAt": "2025-01-15T10:30:00.000Z"
}
```

#### Review

```json
{
    "type": "review",
    "companySlug": "amazon.com",
    "companyName": "Amazon",
    "reviewId": "65abc123def456",
    "reviewerName": "John D.",
    "reviewerLocation": "US",
    "reviewerReviewCount": 5,
    "rating": 4,
    "title": "Great product, slow shipping",
    "reviewText": "The product quality was excellent but delivery took 2 weeks...",
    "dateOfExperience": "2025-01-10T00:00:00.000Z",
    "datePosted": "2025-01-12T14:22:00.000Z",
    "verified": true,
    "companyReply": "Thank you for your feedback...",
    "companyReplyDate": "2025-01-13T09:00:00.000Z",
    "usefulVotes": 3,
    "language": "en",
    "reviewUrl": "https://www.trustpilot.com/reviews/65abc123def456",
    "scrapedAt": "2025-01-15T10:30:00.000Z"
}
```

#### Search/Category Result

```json
{
    "type": "search_result",
    "companyName": "Bluehost",
    "slug": "bluehost.com",
    "trustScore": 3.8,
    "totalReviews": 4521,
    "category": "Web Hosting",
    "profileUrl": "https://www.trustpilot.com/review/bluehost.com",
    "scrapedAt": "2025-01-15T10:30:00.000Z"
}
```

### Pricing

**Pay-Per-Event**: $0.35 per 1,000 results ($0.00035 per result)

| Results | Cost |
|---------|------|
| 100 | $0.035 |
| 1,000 | $0.35 |
| 10,000 | $3.50 |
| 100,000 | $35.00 |

### Technical Details

- Uses **CheerioCrawler** from Crawlee for fast, lightweight HTML parsing
- Extracts data from **`__NEXT_DATA__`** (Next.js server-side props) as the primary data source
- Falls back to **JSON-LD structured data**, then **HTML DOM parsing** if needed
- Session pool with cookie persistence for anti-bot resilience
- Randomized User-Agent headers across Chrome versions and platforms
- Automatic retry with session rotation on 403/429 responses
- Detects and handles Cloudflare challenge pages

### Legal Notice

This Actor is intended for legitimate data collection purposes such as market research, sentiment analysis, and reputation monitoring. Always respect Trustpilot's Terms of Service. The responsibility for how the data is used lies with the end user.

# Actor input Schema

## `action` (type: `string`):

What to scrape from Trustpilot.

## `companyUrls` (type: `array`):

List of Trustpilot company URLs (e.g., https://www.trustpilot.com/review/example.com) or slugs (e.g., example.com). Used for scrapeReviews and scrapeCompanyProfile actions.

## `searchQuery` (type: `string`):

Keyword or company name to search for. Used with searchCompanies action.

## `categoryUrl` (type: `string`):

Trustpilot category page URL (e.g., https://www.trustpilot.com/categories/electronics\_technology). Used with scrapeCategory action.

## `maxReviews` (type: `integer`):

Maximum number of reviews to scrape per company. Set to 0 for all reviews.

## `maxSearchResults` (type: `integer`):

Maximum number of companies to return from search or category pages.

## `filterByStars` (type: `array`):

Only scrape reviews with these star ratings. Leave empty for all ratings.

## `filterDateFrom` (type: `string`):

Only scrape reviews posted on or after this date (YYYY-MM-DD).

## `filterDateTo` (type: `string`):

Only scrape reviews posted on or before this date (YYYY-MM-DD).

## `filterLanguage` (type: `string`):

Only scrape reviews in this language (e.g., 'en', 'de', 'fr'). Leave empty for all languages.

## `filterVerifiedOnly` (type: `boolean`):

If enabled, only scrape reviews marked as verified.

## `includeCompanyProfile` (type: `boolean`):

When scraping reviews, also output a company profile record.

## `proxyConfiguration` (type: `object`):

Select proxies to use for scraping.

## `maxConcurrency` (type: `integer`):

Maximum number of concurrent page requests.

## Actor input object example

```json
{
  "action": "scrapeReviews",
  "companyUrls": [
    "https://www.trustpilot.com/review/amazon.com"
  ],
  "maxReviews": 5,
  "maxSearchResults": 50,
  "filterByStars": [],
  "filterVerifiedOnly": false,
  "includeCompanyProfile": true,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  },
  "maxConcurrency": 5
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "action": "scrapeReviews",
    "companyUrls": [
        "https://www.trustpilot.com/review/amazon.com"
    ],
    "maxReviews": 5,
    "maxSearchResults": 50,
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    },
    "maxConcurrency": 5
};

// Run the Actor and wait for it to finish
const run = await client.actor("magicfingers/trustpilot-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "action": "scrapeReviews",
    "companyUrls": ["https://www.trustpilot.com/review/amazon.com"],
    "maxReviews": 5,
    "maxSearchResults": 50,
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
    "maxConcurrency": 5,
}

# Run the Actor and wait for it to finish
run = client.actor("magicfingers/trustpilot-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "action": "scrapeReviews",
  "companyUrls": [
    "https://www.trustpilot.com/review/amazon.com"
  ],
  "maxReviews": 5,
  "maxSearchResults": 50,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  },
  "maxConcurrency": 5
}' |
apify call magicfingers/trustpilot-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,magicfingers/trustpilot-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/5csThd9SJk10d38zF/builds/BQsuQ285vJtM6aH8t/openapi.json
