# Clutch.co Scraper (`morkerr/clutch-co-scraper`) Actor

Scrape business listings from Clutch.co directory pages. Extract company names, emails, ratings, reviews, project sizes, hourly rates, employee counts, locations, services, phone numbers, and websites.

- **URL**: https://apify.com/morkerr/clutch-co-scraper.md
- **Developed by:** [morkerr](https://apify.com/morkerr) (community)
- **Categories:** Automation, Lead generation
- **Stats:** 3 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 clutch.co listing results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Clutch.co Company & Agency Data Scraper

Extract verified B2B company listings from Clutch.co directory pages at scale. Collect company names, ratings, review counts, project budgets, hourly rates, employee ranges, structured location data, service offerings, phone numbers, logo images, website URLs, and listing metadata — all from the listing cards without the cost and overhead of visiting individual profile pages.

### Features

- **80+ directory categories** across Advertising & Marketing, Development, Design & Production, IT Services, and Business Services
- **21 data fields per listing** extracted directly from directory cards
- **Location filtering** by city, region, or country via on-page dropdown interaction
- **Unlimited pagination** — scrape all pages or set a maximum per category
- **Minimum rating filter** — only keep listings above a target rating threshold
- **Featured/excluded toggle** — include or skip promoted/sponsored results
- **Multi-category batch scraping** — select any combination of categories in one run
- **Custom URL support** — paste any Clutch.co directory URL not in the preset list
- **Optional email discovery** — visit each company's website to extract contact emails (disabled by default)
- **Proxy integration** — Apify proxy recommended to avoid rate limiting and Cloudflare challenges
- **Automatic Cloudflare resolution** — retries with proxy rotation on challenge pages
- **CSV/JSON/XML export** via Apify dataset

### How to Use

#### Quick Start

1. Select one or more **categories** from the dropdown (or enter custom URLs in the text area)
2. Optionally enter a **location** to filter results (e.g. `London, England`)
3. Set **Max Pages per Category** (default 3; set to `0` for unlimited)
4. Click **Start**

#### Input Reference

| Field | Type | Description |
|---|---|---|
| **Pick Categories** | Multi-select dropdown | Choose from 80+ Clutch.co directory categories organized by service group |
| **Custom URLs** | Textarea (one per line) | Additional directory URLs not in the preset list, e.g. `https://clutch.co/web-developers` |
| **Location Filter** | Free text | Filter by city or region. Single: `London, England`. Multiple: `[London] [New York]` |
| **Max Pages per Category** | Integer (0–5000) | How many pagination pages to scrape. `0` = no limit (scrapes until no more pages exist) |
| **Minimum Rating** | Decimal (0–5) | Only keep listings with at least this star rating. `0` = include all |
| **Exclude Featured** | Boolean | Skip promoted/sponsored listings at the top of each results page |
| **Scrape Emails** | Boolean | When enabled, visits each company's website to extract email addresses. Adds significant runtime |
| **Page Timeout** | Integer (ms) | Maximum wait time per page load before retrying |
| **Proxy Configuration** | Apify proxy object | Proxy pool for IP rotation and Cloudflare bypass |

#### Location Filter Details

The location filter interacts with Clutch.co's autocomplete dropdown. Enter a city or region, wait for the suggestion to appear, and the scraper selects the matching item. The page URL updates to reflect the location filter (e.g. `/uk/england/london`).

**Examples:**

- `London, England` — single city filter
- `New York, United States` — single city with country
- `[London, England] [New York, United States]` — scrape the same category twice, once per location

> **Note:** Location results vary by category. The filter only shows locations that have matching listings. If a dropdown item does not appear within the retry window, the category is skipped.

### Category Coverage

The actor supports 80+ directory categories across five service groups:

| Group | Example Categories |
|---|---|
| **Advertising & Marketing** | SEO, PPC, Social Media Marketing, Content Marketing, Branding, PR, Video Production, Email Marketing, Digital Strategy, Conversion Optimization, Media Buying |
| **Development** | Web Developers, Mobile App Development (iOS, Android), AI/ML, Blockchain, eCommerce, Ruby on Rails, WordPress, Drupal, Magento, .NET, PHP, IoT, AR/VR, Software Testing |
| **Design & Production** | Web Design, UX/UI, Graphic Design, Logo Design, Product Design, Packaging Design, Print Design, Interior Design |
| **IT Services** | Cybersecurity, Cloud Consulting, BI & Big Data Analytics, Staff Augmentation, Managed Service Providers (MSP) |
| **Business Services** | Call Centers, BPO, HR & Recruiting, Accounting & Payroll, Translation, Transcription, Real Estate, Legal, Logistics & Supply Chain |

### Output Schema

Each dataset record represents one company listing from a Clutch.co directory page.

| Field | Type | Description |
|---|---|---|
| `category` | string | Human-readable category name (e.g. `Full Service Digital (Advertising & Marketing)`) |
| `categoryUrl` | string | Clutch.co directory category URL this listing was scraped from |
| `currentPageUrl` | string | Exact page URL where the listing was found (includes location filter and pagination) |
| `name` | string | Company name |
| `profileUrl` | string | Full Clutch.co profile URL |
| `website` | string | Company website URL (redirects resolved, UTM parameters stripped) |
| `phone` | string | Publicly listed phone number |
| `rating` | string | Star rating displayed on the listing (e.g. `4.5`) |
| `reviewCount` | string | Number of reviews |
| `minProject` | string | Minimum project budget (e.g. `$1,000+`) |
| `hourlyRate` | string | Hourly rate range (e.g. `$50 - $99 / hr`) |
| `employees` | string | Employee count range (e.g. `50 - 249`) |
| `location` | string | Location text from the listing card |
| `addressLocality` | string | City from structured address metadata |
| `addressRegion` | string | State or region from structured address metadata |
| `addressCountry` | string | Country code from structured address metadata |
| `streetAddress` | string | Street address from structured address metadata |
| `postalCode` | string | Postal code from structured address metadata |
| `services` | string | Comma-separated list of services offered (from chart tooltip or services list) |
| `description` | string | Short description or project highlight text |
| `verified` | string | Whether the listing has a Clutch Verified checkmark (`true`/`false`) |
| `logo` | string | URL of the company logo image |
| `listingType` | string | `featured` for promoted/sponsored listings, `regular` for organic results |
| `emails` | string | Email addresses extracted from the company's website (only when email scraping is enabled) |

### Technical Details

#### Pagination

The scraper uses a two-tier content loading strategy for paginated results:

1. **Browser fetch (fast path)** — uses the page's own `fetch()` API to request the next page's HTML. If the response contains listing cards, it sets the content directly via `page.setContent()`.
2. **Full navigation (fallback)** — if the fetch response does not contain valid listing HTML, falls back to `page.goto()` with `waitForSelector` to ensure cards are rendered.

This minimizes page transitions and speeds up multi-page scraping.

#### Location Filter

The location feature simulates human interaction with Clutch.co's autocomplete input. The scraper:

1. Locates the location input field using multiple selectors
2. Types the query character by character with realistic delays
3. Polls for the dropdown to appear and match against the query text
4. Clicks the matching item using a trusted click event
5. Waits for the page URL to reflect the location change
6. Validates that listing cards load in the filtered view

If the dropdown does not appear after multiple retries with full page reloads, the category is skipped.

#### Email Scraping

When enabled, the actor visits each company's website (from the `website` field) and extracts email addresses using:

- `mailto:` links on the page
- Regex pattern matching for email addresses in the page text and HTML
- Deduplication and filtering of common false positives (`noreply@`, `example.com`, etc.)

Email scraping is processed concurrently (5 sites at a time) and is disabled by default due to the significant runtime it adds.

#### Proxy & Cloudflare Handling

- Proxy configuration is required for reliable operation at scale. Apify proxy with rotating IPs is recommended.
- The actor detects Cloudflare challenge pages by title and waits up to 15 seconds for resolution
- If a challenge is not resolved within the wait window, the proxy is rotated (up to 3 attempts)
- On repeated Cloudflare failures, the actor falls through to the next proxy group

#### Rate Limits & Timeouts

- Each page load has a configurable timeout (default 60s)
- Email scraping uses a 25s per-site timeout
- Location filter waits up to 20 seconds for URL change after selection
- Pagination retries once with a full reload if cards do not appear

### Use Cases

- **Lead generation** — build targeted prospect lists filtered by category, location, and minimum rating
- **Market research** — analyze the competitive landscape across service verticals and geographies
- **Sales intelligence** — enrich CRM data with company size, budget ranges, services, and contact information
- **Agency discovery** — find agencies, consultancies, and service providers by expertise and location
- **Data enrichment** — supplement existing datasets with Clutch.co verification, ratings, and review counts

# Actor input Schema

## `categories` (type: `array`):

Select one or more categories to scrape.

## `location` (type: `string`):

Filter results by location. Use bracket notation for multiple locations.

Single location:
London, England
New York, United States

Multiple locations (same category):
\[London, England] \[New York, United States] \[Chicago]

Leave empty for all locations.

## `categoryUrls` (type: `string`):

Additional Clutch.co directory URLs (one per line). Useful if the category you need isn't listed in the dropdown above.

Without location (scrapes all listings):
https://clutch.co/web-developers
https://clutch.co/developers/artificial-intelligence

With location (scope to a specific area):
https://clutch.co/us/ny/new-york/web-developers
https://clutch.co/uk/england/london/app-developers

Tip: The Location Filter field above also applies to these custom URLs.

## `maxPagesPerCategory` (type: `integer`):

Number of pagination pages to scrape per category (each page shows ~50 listings). Set to 0 for unlimited — scrapes all available pages until no more exist.

## `minRating` (type: `number`):

Only include companies with at least this star rating (e.g. 4.0, 3.5). Set to 0 to include all.

## `excludeFeatured` (type: `boolean`):

Skip promoted/sponsored listings at the top of each page.

## `scrapeEmails` (type: `boolean`):

When enabled, visits each company's website to extract email addresses. This significantly increases run time (several seconds per website).

## `pageTimeout` (type: `integer`):

Maximum time to wait for each directory page to fully load.

## `proxyConfiguration` (type: `object`):

Proxy settings for the scraper. Apify proxy recommended to avoid IP blocking.

## Actor input object example

```json
{
  "categories": [
    "https://clutch.co/agencies/digital"
  ],
  "location": "",
  "categoryUrls": "",
  "maxPagesPerCategory": 3,
  "minRating": 0,
  "excludeFeatured": false,
  "scrapeEmails": false,
  "pageTimeout": 60000,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "BUYPROXIES94952"
    ]
  }
}
```

# Actor output Schema

## `category` (type: `string`):

Human-readable name of the Clutch.co directory category (e.g. 'Full Service Digital (Advertising & Marketing)').

## `categoryUrl` (type: `string`):

The Clutch.co directory category URL this listing was scraped from.

## `currentPageUrl` (type: `string`):

The exact URL of the page where this listing was found, including any location or pagination parameters.

## `name` (type: `string`):

Name of the company listed on Clutch.co.

## `profileUrl` (type: `string`):

Direct URL to the company's Clutch.co profile page.

## `website` (type: `string`):

Company's website URL with UTM params stripped.

## `phone` (type: `string`):

Publicly listed phone number.

## `rating` (type: `string`):

Star rating displayed on the listing card (e.g. 4.5).

## `reviewCount` (type: `string`):

Number of reviews on the listing card.

## `minProject` (type: `string`):

Minimum project size displayed on the card (e.g. $1,000+).

## `hourlyRate` (type: `string`):

Hourly rate range (e.g. $50 - $99 / hr).

## `employees` (type: `string`):

Employee count range (e.g. 50 - 249).

## `location` (type: `string`):

Location text from the listing card.

## `addressLocality` (type: `string`):

City from the listing card's structured address data.

## `addressRegion` (type: `string`):

State or region from the listing card's structured address data.

## `addressCountry` (type: `string`):

Country code from the listing card's structured address data.

## `streetAddress` (type: `string`):

Street address from the listing card's structured data.

## `postalCode` (type: `string`):

Postal code from the listing card's structured address data.

## `services` (type: `string`):

Comma-separated list of services offered, extracted from the listing card's chart or list.

## `description` (type: `string`):

Short description or highlight text from the listing card.

## `verified` (type: `string`):

Whether the listing has a verified checkmark.

## `logo` (type: `string`):

URL of the company's logo image from the listing card.

## `listingType` (type: `string`):

Type of listing: 'featured' (sponsored/promoted) or 'regular'.

## `emails` (type: `string`):

Email addresses found on the company's website (only populated when email scraping is enabled).

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "location": "",
    "categoryUrls": ""
};

// Run the Actor and wait for it to finish
const run = await client.actor("morkerr/clutch-co-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "location": "",
    "categoryUrls": "",
}

# Run the Actor and wait for it to finish
run = client.actor("morkerr/clutch-co-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "location": "",
  "categoryUrls": ""
}' |
apify call morkerr/clutch-co-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,morkerr/clutch-co-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/OJlaZlCxzrOhV3NeA/builds/enWHxhbMHg6pSBlkM/openapi.json
