# Wikipedia Email Scraper - Keyword & Location Targeting (`scrapido/wikipedia-email-scraper`) Actor

📖 Wikipedia Email Scraper pulls contributor and referenced organisation emails by keyword and location. 🔓 Custom domain filters, hidden-address decoding and dedup. 📤 Export results to CSV, JSON or Excel for research teams.

- **URL**: https://apify.com/scrapido/wikipedia-email-scraper.md
- **Developed by:** [Scrapido](https://apify.com/scrapido) (community)
- **Categories:** Lead generation, Automation, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.50 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

### Wikipedia Email Scraper 🔍

**Wikipedia Email Scraper** helps marketers, recruiters, and sales pros quickly find relevant email addresses without manually hunting through pages one by one. It’s an email extraction tool designed to scrape Wikipedia for targeted Wikipedia email extractor results—perfect for extracting emails from Wikipedia and speeding up Wikipedia lead generation at scale and pace.

***

### 🌟 Key Features of Wikipedia Email Scraper

| Feature | Benefit |
|---|---|
| ✅ **Targeted Keyword Search** | Reach the exact Wikipedia audience you need using your chosen terms |
| ✅ **Location Filtering** | Narrow results to a specific area to improve relevance for email extraction |
| ✅ **Custom Domain Filter** | Focus on specific email address detection rules like `@gmail.com` or `@yahoo.com` |
| ✅ **Bulk Export (JSON / CSV)** | Drop results straight into your CRM or email tool for dataset building |
| ✅ **Proxy-Ready** | Built for reliable web scraping with support for uninterrupted runs |
| ✅ **Real-Time Saving** | Each result is saved incrementally to prevent data loss on longer scraping automations |
| ✅ **Resilient Execution** | Includes retries and fallbacks to keep scraping stable when pages are slow or blocked |
| ✅ **Deduplicated Emails** | Avoid repeated contacts so your contact information parsing stays clean |

***

### 📥 Input — Wikipedia Email Scraper Parameters

```json
{
  "keywords": ["manager", "founder"],
  "location": "",
  "customDomains": ["@gmail.com", "@yahoo.com"],
  "maxEmails": 20
}
```

| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
| `keywords` | Array | ✅ Yes | — | Keywords or queries used to search for Wikipedia pages that match your topic for email extraction |
| `location` | String | No | `""` | Optional location filter to focus results on a specific city, region, or country |
| `customDomains` | Array | No | `[]` | Optional list of email domains (like `@gmail.com`, `@yahoo.com`) to filter email address detection |
| `maxEmails` | Integer | No | `20` | Upper cap on total emails to collect for this run (helps control scraping time and cost) |

***

### 📤 Output — What Wikipedia Email Scraper Returns

The actor saves each result as a JSON record in your Apify dataset.

```json
{
  "keyword": "manager",
  "title": "Maria Jensen (Project Management)",
  "description": "Maria Jensen is a project manager with experience across construction and operations. Contact emails are listed in publicly available sources.",
  "url": "https://en.wikipedia.org/wiki/Maria_Jensen",
  "email": "maria.jensen@gmail.com"
}
```

| Field | Type | Description |
|---|---|---|
| `keyword` | String | The search term that surfaced this Wikipedia result during scraping Wikipedia |
| `title` | String | Wikipedia profile name or business title associated with the extracted contact |
| `description` | String | Bio or profile summary text used for contact information parsing and email address detection |
| `url` | String | Direct link to the Wikipedia page that the email was extracted from |
| `email` | String | The extracted email address that matched your customDomains filters |

***

### 💻 How to Use Wikipedia Email Scraper — Step-by-Step

1. **Open the Actor** — Find **Wikipedia Email Scraper** on [Apify Store](https://apify.com/store)
2. **Enter Keywords** — Add role titles, responsibilities, or business terms (e.g., “manager”, “founder”) for extracting emails from Wikipedia
3. **Set Location** *(optional)* — If you want geo-targeted Wikipedia lead generation, add a city or region
4. **Filter by Domain** *(optional)* — Limit results to specific domains like `@gmail.com` or `@yahoo.com`
5. **Set Max Emails** — Cap results to control dataset building cost and runtime
6. **Run the Actor** — Start the run and monitor progress in live logs
7. **Export Results** — Download your Wikipedia email list from the dataset tab (JSON / CSV)

***

### 💡 Best Use Cases for Wikipedia Email Scraper

- 🎯 **B2B Lead Generation** — Build targeted email lists from Wikipedia for outbound sales campaigns
- 📣 **Email Marketing** — Use scraped Wikipedia contacts to power outreach and newsletter sequences
- 🤝 **Recruitment** — Find and contact relevant professionals using email address detection from publicly available sources
- 🔬 **Market Research** — Identify industry voices and niche communities for data mining workflows
- 📊 **CRM Enrichment** — Add emails to existing records after scraping Wikipedia data into a clean dataset

***

### Disclaimer

This actor only accesses **publicly available data** on Wikipedia. It does not scrape private profiles, authenticated content, or password-protected pages. You are responsible for ensuring your use complies with Wikipedia’s Terms of Service, GDPR, CCPA, and all applicable anti-spam laws. This tool is intended for legitimate purposes only, including lead generation, research, and marketing in compliance with local regulations. For data-removal requests, contact 📧 <scrapidocontact@gmail.com>.

***

### 🆘 Support & Feedback

Have a question or found an issue with the **Wikipedia Email Scraper**? We're here to help.

- 🐞 **Bug Reports:** Open a ticket in the repository's Issues section
- ✨ **Custom Solutions & Feature Requests:** Reach out to our team
- 📧 **Email:** <scrapidocontact@gmail.com>

Your feedback shapes the roadmap — we read every message.

### Multiple Email Types

**Email Types** replaces the old single Audience Type choice: select as many
kinds of mailbox as you want and the run chases all of them together.

| Type | What it matches |
| --- | --- |
| Personal / free webmail | Gmail, Outlook, Yahoo, iCloud, AOL, Proton, ... |
| Business / corporate | Company domains - free webmail and institutions excluded |
| Education (.edu / .ac) | `.edu`, `.ac.uk`, `.edu.au`, `.ac.in` and other academic suffixes |
| Government (.gov / .mil) | `.gov`, `.mil`, `.gov.uk`, `.gc.ca`, ... |
| Non-profit (.org) | `.org`, `.ngo`, `.org.uk`, ... |

Each selected type contributes its own Google dork patterns *and* its own domain
test, so a result is only kept if it genuinely belongs to the type that found
it. Every row carries an `emailType` field recording which one that was.

Suffixes are matched as real domain suffixes, so `cs.mit.edu` counts as
Education while `notedu.com` does not.

Setting **Custom Email Domains** still overrides everything: an explicit domain
list is a manual override and replaces the type-driven patterns. The legacy
`audienceType` value is still accepted, so saved inputs keep working.

# Actor input Schema

## `keywords` (type: `array`):

A list of keywords or queries to search for.

## `audienceType` (type: `string`):

Business Emails runs contextual discovery patterns tuned for company contact pages ("email us at", "contact@", careers, bookings, ...) and filters out consumer webmail domains. Consumer Emails instead searches gmail.com, yahoo.com, outlook.com, hotmail.com and icloud.com directly.

## `maxEmails` (type: `integer`):

Maximum number of emails to collect. The scraper will stop once this limit is reached. Setting a higher limit allows for more potential results but doesn't guarantee reaching that number. This helps save costs by controlling scraping time.

## `customDomains` (type: `array`):

Optional manual override — provide specific email domains to search for (e.g. @hubspot.com) instead of using Audience Type. Leave empty to use Audience Type.

## `emailTypes` (type: `array`):

Which kinds of mailbox to hunt for. Pick as many as you like - each type contributes its own set of Google search patterns and its own domain filter, and every result records the type it was found as. Personal = free webmail (Gmail, Outlook, Yahoo, iCloud). Business = company domains, excluding free webmail and institutions. Education = .edu / .ac.uk and friends. Government = .gov / .mil. Non-profit = .org.

## Actor input object example

```json
{
  "keywords": [
    "manager",
    "founder"
  ],
  "audienceType": "Consumer Emails",
  "maxEmails": 20,
  "customDomains": [],
  "emailTypes": [
    "Personal",
    "Business"
  ]
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "keywords": [
        "manager",
        "founder"
    ],
    "customDomains": [],
    "emailTypes": [
        "Personal",
        "Business"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapido/wikipedia-email-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "keywords": [
        "manager",
        "founder",
    ],
    "customDomains": [],
    "emailTypes": [
        "Personal",
        "Business",
    ],
}

# Run the Actor and wait for it to finish
run = client.actor("scrapido/wikipedia-email-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "keywords": [
    "manager",
    "founder"
  ],
  "customDomains": [],
  "emailTypes": [
    "Personal",
    "Business"
  ]
}' |
apify call scrapido/wikipedia-email-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapido/wikipedia-email-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/hXBDB4AbOSSpLFdY2/builds/6DOKt7v61eqgmPmAm/openapi.json
