# Taobao Email Scraper - Keyword & Location Targeting (`scrapido/taobao-email-scraper`) Actor

🏮 Taobao Email Scraper pulls seller and supplier emails by keyword and location. 🔓 Domain filters, obfuscated-address decoding and duplicate handling. 📤 Export Taobao leads to CSV, JSON or Excel for China sourcing.

- **URL**: https://apify.com/scrapido/taobao-email-scraper.md
- **Developed by:** [Scrapido](https://apify.com/scrapido) (community)
- **Categories:** Lead generation, E-commerce, Automation
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.50 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

### Taobao Email Scraper 📬

**Taobao Email Scraper** is an Apify automated web scraping tool that extracts public email addresses from Taobao based on your keywords and filters. It solves the biggest pain point for marketers, recruiters, and data teams: manual contact research that can’t scale. Get faster lead generation from taobao with clean, export-ready results.

***

### What is Taobao Email Scraper? 🔍

**Taobao Email Scraper** is an Apify actor built for automated contact email harvesting from Taobao. It searches using the keywords you provide, then extracts publicly available email addresses that match your email-domain filters (like `@gmail.com`). If you’re looking for a taobao email extractor or a taobao contact scraper, this tool helps you turn time-consuming prospecting into repeatable web data extraction.

It’s a practical taobao email scraper tool for marketers, recruiters, sales teams, and analysts who need taobao seller email leads at scale—without manually hunting across pages. In short: it’s a lead generation tool that helps you compile thousands of taobao emails efficiently for CRM enrichment, outreach, and data pipelines.

***

### What Data Does a Taobao Email Scraper Collect? 📊

This actor focuses on contact discovery, keeping each lead record structured so you can use it downstream for email outreach and CRM enrichment. It captures contact and identity-style details plus the context used to find the contact.

| Data Category | Fields Extracted | Description |
|---|---|---|
| Contact | `email` | Public email address associated with the Taobao result |
| Identity | `title` | Profile title / business title (as shown in the scraped result) |
| Context | `description` | Profile description or snippet text that contains context around the contact |
| Discovery | `keyword` | Search keyword you used to surface the result |
| Navigation | `url` | Direct link to the result page |
| Location | `location` | Location associated with the result (if available in the scraped content) |

***

### What Do Results from Taobao Email Scraper Look Like? 👀

Each result is a structured JSON record saved to your Apify dataset. Here's a real example:

```json
{
  "keyword": "manager",
  "title": "Oceanstar Trading Co., Ltd.",
  "description": "Oceanstar Trading Co., Ltd. • Wholesale & procurement services • Contact via email for cooperation",
  "url": "https://item.taobao.com/item.htm?id=824193455211",
  "email": "sales@oceanstartrading.com",
  "location": "Guangzhou, Guangdong"
}
```

Your dataset is stored in Apify in JSON form (and you can export as CSV from the Apify Console).

***

#### Core Features: Taobao Email Scraper ⚡

| Feature | Benefit |
|---|---|
| ✅ Keyword-driven targeting | Uses your keywords to surface relevant Taobao contacts for ecommerce email harvesting |
| ✅ Location filter (`location`) | Helps you narrow results by region to improve lead generation from taobao |
| ✅ Custom email-domain filtering (`customDomains`) | Extract only the domains you want (e.g., `@gmail.com`) for cleaner contact email scraping |
| ✅ Configurable result cap (`maxEmails`) | Controls run size for cost and time management with predictable harvesting |
| ✅ Proxy support | Built-in proxy support helps keep scraping reliable on larger batches |
| ✅ Real-time dataset saving | Results are pushed incrementally to your Apify dataset |
| ✅ Structured output for analysis | Ready for CRM import, outreach workflows, or data extraction from ecommerce sites |
| ✅ Designed for resilience | Includes robustness for long runs so you can continue collecting leads without starting over |

***

### Getting Started with Taobao Email Scraper 🚀

1. **Open Apify Console** — Go to [apify.com/store](https://apify.com/store) and search **Taobao Email Scraper**
2. **Click Try for Free** — Sign in or create a free Apify account
3. **Open the Input Tab** — Configure your scraping parameters
4. **Add Keywords** — Enter role- or intent-based keywords (e.g., `manager`, `founder`)
5. **(Optional) Set Filters** — Add a `location` and/or specific `customDomains`
6. **Cap Your Results** — Use `maxEmails` to control how many taobao emails you collect
7. **Click Start** — Run the actor and monitor progress in real time
8. **Access Your Data** — Open the dataset (Scraped Leads) and export when ready

You can be productive quickly—set your keywords, optionally restrict domains, and start collecting contacts.

***

### Ways to Use Taobao Email Scraper 💡

- 🎯 **Lead generation from taobao**: Build targeted lists for outreach and partnership requests
- 📣 **Ecommerce email harvesting**: Collect public emails for supplier discovery and vendor onboarding
- 🤝 **Contact information mining**: Enrich an existing contact base with scraped emails and profile context
- 🔬 **Market research support**: Compare niches by keyword and measure email availability across results
- ⚙️ **Data pipelines**: Use exported JSON/CSV records for downstream automation and NLP email entity extraction workflows
- 📊 **Web scraping emails for segmentation**: Filter by domain and refine targeting for higher response rates

***

#### Input Parameters — Taobao Email Scraper

**Example input JSON (with real default values):**

```json
{
  "keywords": ["manager", "founder"],
  "location": "",
  "customDomains": ["@gmail.com", "@yahoo.com"],
  "maxEmails": 20
}
```

| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
| `keywords` | Array | ✅ Yes | — | Keywords used to search for Taobao results that may contain contact emails |
| `location` | String | No | `""` | Optional location to filter results (leave empty for no location filtering) |
| `customDomains` | Array | No | `["@gmail.com","@yahoo.com"]` | Only extract emails matching these domains (like restricting to `@gmail.com`) |
| `maxEmails` | Integer | No | `20` | Maximum number of emails to collect before the run stops |

***

#### Output Parameters — Taobao Email Scraper

**Dataset output example:**

```json
[
  {
    "keyword": "founder",
    "title": "Linhai Home & Living",
    "description": "Home & Living brand • Business inquiries welcome • Contact by email",
    "url": "https://item.taobao.com/item.htm?id=837204558901",
    "email": "hello@linhaihome.com"
  }
]
```

| Field | Label | Format | Description |
|---|---|---|---|
| `keyword` | Keyword | text | The search keyword that surfaced this result |
| `title` | Title | text | Profile name or business title as shown in the scraped result |
| `description` | Description | text | Profile description or snippet context where contact information may appear |
| `url` | Url | link | Direct URL to the scraped result page |
| `email` | Email | text | Extracted public email address |

***

### Why Choose This Taobao Email Scraper? 🏆

**Taobao Email Scraper** is built for practical ecommerce email harvesting: keyword-based discovery, domain filtering, and a clean dataset you can export for analysis. It’s designed for scaling contact email scraping beyond manual research, with support for reliable scraping at higher volumes and incremental dataset saving.

Compared to generic scraping approaches, you get structured output that’s immediately useful for CRM enrichment, lead generation from taobao, and email address parsing. If you want tailored workflows or custom enhancements, you can reach out via <scrapidocontact@gmail.com>.

***

### How Many Results Can You Scrape? 📈

Use `maxEmails` to set a cap from 1 up to 10,000. The actual number of returned records depends on how many Taobao results match your keywords and email-domain filters and contain publicly available email addresses. For larger lead generation projects, increase the cap and run time to capture more potential leads.

***

### Legal Guidelines for Scraping Taobao ⚖️

**Taobao Email Scraper** works with **publicly available data** on Taobao. It does not use logins or access private or password-protected content. You’re responsible for complying with applicable laws and Taobao’s Terms of Service, as well as privacy and anti-spam regulations (including GDPR/CCPA where relevant). Use extracted data only for legitimate business purposes.

For data removal requests, contact <scrapidocontact@gmail.com>.

***

### FAQ — Taobao Email Scraper ❓

#### How does the Taobao Email Scraper identify data?

The actor uses your provided keywords and optional filters to discover relevant Taobao results, then extracts public email addresses from the scraped content. It focuses on email extraction from publicly available sources—no authentication is required.

#### What Taobao profile types can I scrape?

You can scrape public Taobao results that expose contact emails in publicly visible content. If no public email is present for a given result, it won’t contribute to your extracted email list.

#### How did the Taobao Email Scraper perform in our tests?

Performance depends on keyword fit and how often public email addresses appear in matching results. To improve yield, use more targeted keywords and include relevant email domains in `customDomains`.

#### Why scrape Taobao for contacts?

Taobao hosts many businesses and sellers that publish contact emails for inquiries. Automating contact information mining can dramatically reduce the time required to find taobao seller email leads compared to manual searching.

#### How much does the Taobao Email Scraper cost?

Costs are controlled by how many results you request and how long you run. Use `maxEmails` to cap your run. Larger caps can increase scraping time, so start with a realistic target and refine after reviewing your first dataset export.

#### How does the Taobao Email Scraper help my business?

It helps you build segmented lead lists for outreach and CRM enrichment. The dataset output (keyword, title, description, url, email) is structured for analysis, deduplication workflows, and downstream automation.

#### What challenges should I expect when using the Taobao Email Scraper?

Not every Taobao result contains a public email address, so results vary by niche. If you get fewer emails than expected, try adjusting your keywords, adding related terms, or expanding your `customDomains` list.

#### How do I choose a high-performing Taobao Email Scraper?

Choose configurations that balance specificity and coverage: use focused keywords, optional `location` where relevant, domain filtering via `customDomains`, and a sensible `maxEmails` cap so you can iteratively improve your lead generation output.

***

### Conclusion 🏁

**Taobao Email Scraper** is a fast, structured way to extract taobao emails for lead generation and outreach workflows. Whether you’re enriching a CRM, building segmented ecommerce email harvesting lists, or powering data pipelines, start with clear keywords, filter email domains, and export your dataset from Apify.

### Multiple Email Types

**Email Types** replaces the old single Audience Type choice: select as many
kinds of mailbox as you want and the run chases all of them together.

| Type | What it matches |
| --- | --- |
| Personal / free webmail | Gmail, Outlook, Yahoo, iCloud, AOL, Proton, ... |
| Business / corporate | Company domains - free webmail and institutions excluded |
| Education (.edu / .ac) | `.edu`, `.ac.uk`, `.edu.au`, `.ac.in` and other academic suffixes |
| Government (.gov / .mil) | `.gov`, `.mil`, `.gov.uk`, `.gc.ca`, ... |
| Non-profit (.org) | `.org`, `.ngo`, `.org.uk`, ... |

Each selected type contributes its own Google dork patterns *and* its own domain
test, so a result is only kept if it genuinely belongs to the type that found
it. Every row carries an `emailType` field recording which one that was.

Suffixes are matched as real domain suffixes, so `cs.mit.edu` counts as
Education while `notedu.com` does not.

Setting **Custom Email Domains** still overrides everything: an explicit domain
list is a manual override and replaces the type-driven patterns. The legacy
`audienceType` value is still accepted, so saved inputs keep working.

# Actor input Schema

## `keywords` (type: `array`):

A list of keywords or queries to search for.

## `audienceType` (type: `string`):

Business Emails runs contextual discovery patterns tuned for company contact pages ("email us at", "contact@", careers, bookings, ...) and filters out consumer webmail domains. Consumer Emails instead searches gmail.com, yahoo.com, outlook.com, hotmail.com and icloud.com directly.

## `maxEmails` (type: `integer`):

Maximum number of emails to collect. The scraper will stop once this limit is reached. Setting a higher limit allows for more potential results but doesn't guarantee reaching that number. This helps save costs by controlling scraping time.

## `customDomains` (type: `array`):

Optional manual override — provide specific email domains to search for (e.g. @hubspot.com) instead of using Audience Type. Leave empty to use Audience Type.

## `emailTypes` (type: `array`):

Which kinds of mailbox to hunt for. Pick as many as you like - each type contributes its own set of Google search patterns and its own domain filter, and every result records the type it was found as. Personal = free webmail (Gmail, Outlook, Yahoo, iCloud). Business = company domains, excluding free webmail and institutions. Education = .edu / .ac.uk and friends. Government = .gov / .mil. Non-profit = .org.

## Actor input object example

```json
{
  "keywords": [
    "manager",
    "founder"
  ],
  "audienceType": "Consumer Emails",
  "maxEmails": 20,
  "customDomains": [],
  "emailTypes": [
    "Personal",
    "Business"
  ]
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "keywords": [
        "manager",
        "founder"
    ],
    "customDomains": [],
    "emailTypes": [
        "Personal",
        "Business"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapido/taobao-email-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "keywords": [
        "manager",
        "founder",
    ],
    "customDomains": [],
    "emailTypes": [
        "Personal",
        "Business",
    ],
}

# Run the Actor and wait for it to finish
run = client.actor("scrapido/taobao-email-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "keywords": [
    "manager",
    "founder"
  ],
  "customDomains": [],
  "emailTypes": [
    "Personal",
    "Business"
  ]
}' |
apify call scrapido/taobao-email-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapido/taobao-email-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/HNjzfVyhgwHIRveAU/builds/qCxHJ36YiLqoIeWMR/openapi.json
