# \[DEPRECATED v0.7] JD.com Scraper � Price Endpoint Blocked (`zhorex/jd-scraper`) Actor

DEPRECATED � JD.com pricing endpoints blocked at proxy-infrastructure level on Apify residential pool. Actor returns enrichment fields only (brand, title, category, images, stock) and does NOT charge � realtimePrice universally null. Do not subscribe. v1.0 with premium proxy on roadmap.

- **URL**: https://apify.com/zhorex/jd-scraper.md
- **Developed by:** [Sami](https://apify.com/zhorex) (community)
- **Categories:** AI, E-commerce, Automation
- **Stats:** 2 total users, 0 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $8.00 / 1,000 product detail extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## \[DEPRECATED v0.7] JD.com Scraper — Price Endpoint Blocked

> ## ⚠️ DEPRECATED — Do not subscribe
>
> JD.com's pricing endpoints (`p.3.cn`, `c0.3.cn`, `item-soa.jd.com`, `cd.jd.com/promotion`, `union-pim.jd.com`) are all blocked at the **proxy-infrastructure level** on Apify residential CN. Two distinct block patterns were verified across five endpoints and four proxy geographies (CN / HK / SG / no-country):
>
> - `*.jd.com` price aggregators return HTML error pages instead of JSON (WAF intercept)
> - `*.3.cn` hosts return `curl (56) CONNECT tunnel failed, response 590` (proxy refuses to tunnel)
>
> **As a result**: `realtimePrice` is universally `null` on this Actor today, and the PPE gate (added in v0.6.5) refuses to charge for any record that lacks a price. The Actor returns the enrichment fields (brand, title, category, images, stock, JD-self-run flag, service tags) for diagnostic visibility but **never bills**.
>
> Subscribing to this Actor today will give you **$0 in charges** but also **$0 in usable price data**. If you need JD product pricing data, contact the author about funding a Bright Data / Oxylabs / Soax premium residential pool — a v1.0 release with that integration is on the roadmap.
>
> Existing Saved Tasks continue to function (they receive the diagnostic records, not billing).

***

Extract enrichment fields from **JD.com (京东 / Jingdong)** product pages. Currently shipped fields (when not blocked): productTitle, brandName (3-layer fallback), categoryPath, images, isJdSelfRun, stockStatus, serviceTags, sellerId. **Real-time price extraction is currently non-functional** — see deprecation warning above.

> Part of the **Chinese Digital Intelligence Suite** by zhorex — pairs with the [Chinese Brand Monitor](https://apify.com/zhorex/chinese-brand-monitor) (cross-platform brand mention aggregator across Weibo + RedNote + Bilibili + Douban + Xueqiu, $0.045/mention) and the individual Weibo, RedNote, Xueqiu, Douban scrapers for full-stack China data coverage.

***

### What you get per product

For each JD product URL or SKU ID you submit, one record with:

- **Identifiers** — `productId`, `productUrl`, `brandName`, `categoryPath` (full breadcrumb)
- **Specs** — full JD specs panel as a `{key: value}` dict (商品名称, 商品编号, 上架时间, dimensions, color options…)
- **Images** — `primaryImageUrl` plus full `descriptionImages` array for catalog ingestion
- **Pricing** — `realtimePrice` queried fresh at scrape time (not stale page-load price)
- **Seller signal** — `sellerId`, `isJdSelfRun` flag (true = JD's own warehouse + warranty + return logistics; false = third-party merchant)
- **Stock** — `stockStatus` enum (`in_stock` / `low_stock` / `out_of_stock`), best-effort `stockCount`
- **Service tags** — JD's protection options decoded to human-readable English: `7-day return`, `JD Tianhua (warranty)`, `JD Plus exclusive`, `cash on delivery`, etc.
- **Origin** — `shippingCity` (defaults to Beijing — JD's central warehouse region)
- **Timestamp** — `scrapedAt` UTC ISO 8601

***

### Why this Actor, not a generic e-commerce scraper

- **`isJdSelfRun` flag** — JD's hybrid model means each SKU is either fulfilled by JD itself (own logistics, warranty, return path) or by a third-party merchant on the marketplace. Generic scrapers don't distinguish; this one surfaces the flag on every record.
- **Service tag decoding** — JD's protection options arrive as opaque numeric codes; this Actor maps them to human-readable English so downstream pipelines don't have to maintain the code table.
- **Real-time price** — `realtimePrice` is fetched fresh at scrape time, not parsed from cached HTML, so it captures flash-discount cycles that move within hours.

***

### Example input

```json
{
    "mode": "product_detail",
    "productUrls": [
        "https://item.jd.com/100037053980.html",
        "100066898260"
    ]
}
```

The Actor accepts a mix of full URLs and bare SKU IDs. Duplicates are removed automatically.

***

### Example output (truncated)

```json
{
    "mode": "product_detail",
    "productId": "100037053980",
    "productTitle": "得力（deli） 6018 剪刀剪子剪纸裁剪 办公文具用品",
    "brandName": "得力",
    "categoryPath": ["文教文化用品", "切屑文具", "办公/学生剪刀"],
    "specs": {"商品名称": "得力 6018", "商品编号": "100037053980"},
    "realtimePrice": "12.90",
    "priceCurrency": "CNY",
    "priceSource": "item-soa",
    "sellerId": "0",
    "isJdSelfRun": true,
    "stockStatus": "in_stock",
    "serviceTags": ["7-day return", "JD Tianhua (warranty)"],
    "primaryImageUrl": "https://img12.360buyimg.com/n0/jfs/...",
    "descriptionImages": ["...", "..."],
    "shippingCity": "北京",
    "productUrl": "https://item.jd.com/100037053980.html",
    "scrapedAt": "2026-05-16T10:00:00+00:00"
}
```

***

### Pricing (Pay-Per-Event)

| Event | Price |
|---|---|
| `product-detail-scraped` | **$0.008 / product** |

#### Realistic workflow costs

| Workflow | Volume | Cost |
|---|---|---|
| Daily price tracking on 200 SKUs (per month) | 6,000 details/mo | **~$48 / month** |
| Competitor SKU enrichment (one-time batch) | 1,000 products | **$8** |
| Hedge-fund-grade daily refresh, 5,000 SKUs | 150,000 details/mo | **~$1,200 / month** |
| Catalog migration / one-time pull | 50,000 SKUs | **$400** |

***

### Proxy

The Actor's input schema defaults to **Apify residential proxy with `apifyProxyCountry: "CN"`** — leave it on for production workflows. Apify residential proxy is billed separately from the per-event price (typically a few cents per MB transferred — see your Apify Billing → Proxy usage).

Most non-CN residential IPs also work because `item.jd.com` applies lighter rate-limiting than other JD subdomains. The default config is the safest bet but you can experiment.

***

### Status note (v0.6.3)

JD's classic price endpoint `p.3.cn/prices/mgets` is currently rate-limited at the edge for most shared residential proxy pools (Apify included). v0.6.3 ships with a **multi-endpoint price fallback** — the Actor tries `item-soa.jd.com/getWareBusiness` first (a modern aggregator that also returns brand info), then `c0.3.cn/stock`, then the legacy `p.3.cn`. The first endpoint that returns parseable data wins; price + brand are filled together when possible.

If all three price hosts are blocked AND the page parser can't recover a brand name from the HTML / title, the record is pushed with an `error: "missing_price_and_brand"` field and **no PPE event is charged for it**. You always see the diagnostic in the dataset, you never pay for an empty record.

Brand-name extraction has three fallback layers: the canonical `#parameter-brand` link, the `parameter2` 品牌 field, and a title-pattern heuristic (JD product titles wrap the brand inside the first `（...）` pair).

***

### What this Actor does NOT do (and why)

Earlier versions (v0.1–v0.5) shipped four modes: `product_detail`, `seller_store`, `product_search`, `product_reviews`. The latter three were removed in v0.6 after extensive testing because JD's WAF reliably blocks them on Apify's shared residential proxy pools, returning `系统繁忙` or silently redirecting to the JD homepage. We tested:

- Browser-impersonation HTTP client — blocked
- Playwright with full Chromium + JS execution + primed cookies — blocked
- Four proxy geographies (CN / HK / SG / no-country) — all blocked
- Mobile API endpoints — return 403 (require JD mobile-app signing scheme)

The block is at the IP-reputation layer, not anything fixable client-side. Rather than ship modes that return zero items and accidentally charge buyers, v0.6 ships only the mode that works reliably.

If you specifically need search, reviews, or seller-store data from JD, **contact the author** about integrating a premium residential proxy pool (Bright Data / Oxylabs / Soax) with cleaner reputation against JD specifically. That work is parked as a roadmap item; one paying buyer would unlock it.

Saved tasks that still call the removed modes get a clean rejection message — no PPE charge.

***

### Use cases

1. **Competitor pricing intelligence** — `realtimePrice` enables sub-hour competitor SKU price tracking. Pair with a daily-cron schedule via Apify's Saved Tasks.
2. **Catalog enrichment** — pull every JD product from your brand's SKU list to refresh title, specs, current price, stock status for downstream BI / dashboards.
3. **JD-self-run vs marketplace audit** — the `isJdSelfRun` flag identifies which of your SKUs are sold directly by JD (your authorized channel) vs by third-party merchants (potential gray-market / unauthorized resale).
4. **AI training data** — JD product titles, specs, and category paths are clean labeled data for product-classification, product-matching, and Chinese commerce NLP tasks.

***

### Part of the Chinese Digital Intelligence Suite

- 🆕 [Chinese Brand Monitor](https://apify.com/zhorex/chinese-brand-monitor) — Cross-platform brand mention aggregator (Weibo + RedNote + Bilibili + Douban + Xueqiu in one normalized feed, sentiment + dedup, $0.045/mention). **Pairs perfectly with this Actor**: track your competitors' JD product pricing here, then monitor consumer sentiment about those brands across all 5 social platforms in one call.
- [Weibo Scraper](https://apify.com/zhorex/weibo-scraper) — public sentiment, hot search, KOL posts
- [RedNote Scraper](https://apify.com/zhorex/rednote-scraper) — lifestyle / consumer brand reviews
- [Xueqiu Scraper](https://apify.com/zhorex/xueqiu-scraper) — Chinese stock discussion & cashtag sentiment
- [Douban Scraper](https://apify.com/zhorex/douban-scraper) — film / book / music reviews & ratings

***

### Compliance posture

- Only **public JD data** — same content any anonymous browser visitor sees on item pages.
- **No login bypass** — does not attempt authenticated-only content.
- **No personal data harvesting** — only the seller-identification metadata JD itself displays publicly.

Buyers running this commercially are responsible for downstream compliance with their own jurisdiction's data laws.

***

### Support

Found a bug? Need a field that's not extracted? Open an issue on the Actor page — typical turnaround 48 hours.

If this Actor saves you time, **a 30-second review is the single biggest thing that helps** — it brings the tool to other buyers and pays for continued maintenance.

# Actor input Schema

## `mode` (type: `string`):

Operation mode. This Actor extracts JD.com product detail pages. Earlier versions also offered product\_search, product\_reviews, and seller\_store — those were removed in v0.6 because JD's WAF reliably blocks them on shared residential proxy pools. See README for the full story.

## `productUrls` (type: `array`):

JD product URLs (e.g. 'https://item.jd.com/100037053980.html') or bare SKU IDs (e.g. '100037053980'). Mixed input is accepted; duplicates are removed.

## `proxyConfiguration` (type: `object`):

Proxy configuration. Default Apify residential proxy with apifyProxyCountry='CN' works reliably for product\_detail. Most non-CN residential IPs also work because item.jd.com applies lighter rate-limiting than other JD subdomains.

## Actor input object example

```json
{
  "mode": "product_detail",
  "productUrls": [
    "https://item.jd.com/100037053980.html"
  ],
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "CN"
  }
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "mode": "product_detail",
    "productUrls": [
        "https://item.jd.com/100037053980.html"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("zhorex/jd-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "mode": "product_detail",
    "productUrls": ["https://item.jd.com/100037053980.html"],
}

# Run the Actor and wait for it to finish
run = client.actor("zhorex/jd-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "mode": "product_detail",
  "productUrls": [
    "https://item.jd.com/100037053980.html"
  ]
}' |
apify call zhorex/jd-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,zhorex/jd-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/E25kbmnjS9cnk7T84/builds/uZXh0BiIx3GCqI3Vn/openapi.json
