# Amazon Australia Product & Reviews Scraper (`abotapi/amazon-au-scraper`) Actor

Scrape Amazon Australia (amazon.com.au) products and customer reviews. Search by keyword or paste product, search and category URLs. Returns title, brand, price, list price, rating, review count, star distribution, availability, features, images, categories, specs, best-seller rank and seller.

- **URL**: https://apify.com/abotapi/amazon-au-scraper.md
- **Developed by:** [Abot API](https://apify.com/abotapi) (community)
- **Categories:** E-commerce, Automation, Developer tools
- **Stats:** 9 total users, 0 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.50 / 1,000 product results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Amazon Australia Product & Reviews Scraper

Scrape **[amazon.com.au](https://www.amazon.com.au)** products, **deals** and customer reviews at scale. Search by keyword with filters, filter to **Deals & Discounts**, or paste product / search / deal links (or bare ASINs). Every product comes back as one flat record with current price, **was/original price and savings**, ratings, star distribution, images, specifications and - optionally - the customer reviews Amazon publishes on the product page.

### What it does

- **Two ways to start**
  - **Search** - one or more keywords, with sort, minimum-rating and price-range filters, walking forward through result pages.
  - **URL** - paste product pages (`/dp/<ASIN>`), search pages (`/s?k=...`), deal / category pages, or bare 10-character ASINs.
- **Deals & Discounts** - flip on **Deals & discounts only** to return just the products currently on a deal, price drop or coupon. Combine it with a keyword to find deals in a category, or leave keywords empty to browse the whole deals grid. Each discounted product carries its **was/original price**, **savings** amount, **discount percentage** and **promo label** (for example a coupon or a limited-time deal). Amazon Australia surfaces its specials through a single "Deals & Discounts" collection - there is no separate half-price / clearance sub-taxonomy - so this one switch covers the site's promotional view.
- **Full product detail** - open each product page for brand, list price, star distribution, availability, feature bullets, description, the full image gallery, category breadcrumb, specifications, best-seller rank and seller.
- **Customer reviews** - collect the reviews Amazon publishes on each product page: star rating, title, text, author, date, verified-purchase flag and helpful votes, plus aggregate rating statistics.
- **Resume & recurring updates** - turn on Incremental mode to get only NEW, UPDATED, and REAPPEARED products on every scheduled run, or resume one specific interrupted crawl with `resumeFromRunId`. See "Resume & recurring updates" below.
- **Send results into your apps** - optional Notion / Linear / Airtable / Apify export via MCP connectors, without changing the dataset.

### Reviews: what you get

Amazon Australia shows the **aggregate rating** (average score, total ratings, and the 5-star-to-1-star distribution) and a set of **top customer reviews** to anonymous visitors - this actor captures all of them.

> **Note:** amazon.com.au requires a signed-in account to open its **full paginated review list**. This actor does not sign in, so review collection returns the top reviews Amazon exposes publicly on the product page (typically the most relevant ones), each with full title, text, rating, author, date, verified-purchase flag and helpful-vote count. The aggregate rating, total rating count and star distribution are always captured in full.

### Input

| Field | Type | Description |
|-------|------|-------------|
| `mode` | select | `search` (keyword + filters) or `url` (paste links / ASINs). |
| `queries` | string\[] | Search keywords (search mode). Each keyword is searched separately. Leave empty with `specialsOnly` on to browse all deals. |
| `sortBy` | select | `relevance`, `price_asc`, `price_desc`, `newest`, `avg_rating`, `best_selling`. |
| `specialsOnly` | boolean | Only return products on Amazon Australia's **Deals & Discounts** (deals, price drops, coupons). Default `false`. |
| `minRating` | select | Keep products rated at least `3`, `4` or `5` stars (or any). |
| `minPrice` / `maxPrice` | integer | Keep products within an AUD price range. |
| `urls` | string\[] | Product / search / category URLs, or bare ASINs (url mode). |
| `fetchDetails` | boolean | Open each product page for full detail. Default `true`. |
| `fetchReviews` | boolean | Also collect customer reviews. Default `false`. |
| `maxReviewsPerProduct` | integer | Cap reviews per product (`0` = all available). Default `10`. |
| `maxItems` | integer | Max products for the whole run (`0` = unlimited). Default `20`. |
| `maxPages` | integer | Optional bound on result pages walked per keyword/URL. Empty walks every result page (does not cap products; `maxItems` does). |
| `resumeFromRunId` | string | Optional. ID of a previous run of this actor (or a dataset ID). Products already in that dataset are skipped, so this run returns only NEW products (a delta). Combine both runs' datasets for the full set. Max products then counts only the new products. For recurring daily monitoring of the same search, use `incrementalMode` instead - see "Resume & recurring updates" below. |
| `incrementalMode` | boolean | Daily/recurring monitoring of this same search/URL list. First run returns everything as `NEW`; later runs return only `NEW`/`UPDATED`/`REAPPEARED` by default. Default `false`. See "Resume & recurring updates" below. |
| `stateKey` | string | Optional name for a monitoring campaign, so its incremental state stays stable or is deliberately shared. Auto-derived from your mode/keywords/URLs/sort/filters/detail settings when left empty. |
| `emitUnchanged` | boolean | Incremental mode only. Also return products unchanged since the last run, marked `UNCHANGED`. Adds and bills extra rows you already have. Default `false`. |
| `emitExpired` | boolean | Incremental mode only. Also return products from a previous run no longer found, marked `EXPIRED`, once a run has fully scanned the search/URLs (not capped, not a resume). Adds and bills extra synthetic rows. Default `false`. |
| `proxy` | object | Proxy configuration (Australian residential recommended). |
| `mcpConnectors` | array | Optional MCP connectors to export results into. |

#### Example input

```json
{
  "mode": "search",
  "queries": ["echo dot"],
  "sortBy": "avg_rating",
  "minRating": "4",
  "minPrice": 30,
  "maxPrice": 150,
  "fetchDetails": true,
  "fetchReviews": true,
  "maxReviewsPerProduct": 10,
  "maxItems": 20,
  "proxy": { "useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"], "apifyProxyCountry": "AU" }
}
```

#### Example input - deals only

```json
{
  "mode": "search",
  "queries": ["coffee machine"],
  "specialsOnly": true,
  "fetchDetails": true,
  "maxItems": 20,
  "proxy": { "useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"], "apifyProxyCountry": "AU" }
}
```

### Output

One record per product. Example (values below are illustrative):

```json
{
  "asin": "B0EXAMPLE01",
  "title": "Example Smart Speaker (3rd Gen) with Voice Assistant - Charcoal",
  "url": "https://www.amazon.com.au/dp/B0EXAMPLE01",
  "brand": "ExampleBrand",
  "price": 79.0,
  "listPrice": 99.0,
  "currency": "AUD",
  "savings": 20.0,
  "savingsPercent": 20,
  "promoLabel": "Limited time deal",
  "isOnSpecial": true,
  "rating": 4.5,
  "reviewsCount": 1284,
  "ratingDistribution": { "5": 72, "4": 15, "3": 6, "2": 3, "1": 4 },
  "availability": "In stock",
  "inStock": true,
  "features": [
    "Rich, room-filling sound with clear vocals",
    "Ask the voice assistant to play music, set timers and control smart devices"
  ],
  "description": "A compact smart speaker for any room.",
  "categories": ["Electronics", "Smart Home", "Smart Speakers"],
  "specs": { "Model Number": "EX-3000", "Release Year": "2024", "Connectivity": "Wi-Fi, Bluetooth" },
  "bestSellersRank": 12,
  "seller": "Example Retail AU",
  "shipsFrom": "Amazon AU",
  "images": [
    "https://m.media-amazon.com/images/I/EXAMPLE01._AC_SL1000_.jpg",
    "https://m.media-amazon.com/images/I/EXAMPLE02._AC_SL1000_.jpg"
  ],
  "thumbnailImage": "https://m.media-amazon.com/images/I/EXAMPLE01._AC_SL1000_.jpg",
  "reviewsCollected": 8,
  "reviewStats": {
    "averageRating": 4.5,
    "totalRatings": 1284,
    "ratingDistribution": { "5": 72, "4": 15, "3": 6, "2": 3, "1": 4 }
  },
  "reviews": [
    {
      "reviewId": "REXAMPLE00001",
      "rating": 5.0,
      "title": "Great value",
      "body": "Easy to set up and the sound is excellent for the size.",
      "author": "Sample Reviewer",
      "date": "2025-02-14",
      "reviewedIn": "Australia",
      "verifiedPurchase": true,
      "helpfulCount": 6,
      "variation": "Colour Name: Charcoal",
      "reviewUrl": "https://www.amazon.com.au/gp/customer-reviews/REXAMPLE00001"
    }
  ],
  "searchMode": "search"
}
```

#### Field reference

| Field | Description |
|-------|-------------|
| `asin` | Amazon product identifier. |
| `title`, `brand`, `url` | Product name, brand, canonical product URL. |
| `price`, `listPrice`, `currency`, `savings`, `savingsPercent` | Current price, was/original (strike-through) price, AUD, and computed savings and discount percentage. |
| `promoLabel`, `isOnSpecial` | Promo / coupon / deal label shown on the listing (when present), and a boolean flag set when the product is currently discounted. |
| `rating`, `reviewsCount`, `ratingDistribution` | Average stars, total rating count, and the 5★→1★ percentage split. |
| `availability`, `inStock` | Stock text and a boolean. |
| `features`, `description` | Feature bullets and long description. |
| `categories`, `specs`, `bestSellersRank` | Category breadcrumb, specification table, best-seller rank. |
| `seller`, `shipsFrom` | Seller / dispatch information from the buy box. |
| `images`, `thumbnailImage` | Full image gallery and the primary image. |
| `reviews`, `reviewsCollected`, `reviewStats` | Collected reviews, their count, and aggregate statistics (present when reviews are enabled). |

Whenever a value is not published for a product, the field is omitted or `null` rather than guessed.

**Incremental mode only.** When `incrementalMode` is on, every returned record also carries:

| Field | Description |
|---|---|
| `changeType` | `NEW` | `UPDATED` | `UNCHANGED` | `REAPPEARED` | `EXPIRED` |
| `changedFields` | Top-level fields that changed since last seen; non-empty only for `UPDATED` |
| `firstSeenAt` | When this product was first observed by this monitoring campaign |
| `lastSeenAt` | When this product was last observed |

### Proxy

Amazon Australia serves results to Australian residential connections, so the actor uses an **Australian residential** exit by default and rotates a fresh exit per request. Apify Proxy is recommended for reliable results.

### Resume & recurring updates

There are two different things here — pick the one that matches what you're doing:

| Need | Use |
| --- | --- |
| A crawl stopped and should continue | `resumeFromRunId` / automatic checkpoint recovery |
| Run the same search/URLs every day and receive only changes | `incrementalMode` |
| Keep separate daily campaigns for similar searches | distinct `stateKey` values |
| Run a normal full snapshot | leave both off |

**Resume** (`resumeFromRunId`) continues one specific interrupted or previous large crawl: paste a run ID or dataset ID and this run skips products already collected there, returning only the remaining new products. An automatic same-run checkpoint also protects against platform migrations/Resurrects without any input needed.

**Incremental mode** (`incrementalMode`) is for a schedule (for example, daily): the actor remembers the previous run of the *same* search/URLs by itself, so you never paste a run ID. The first run returns everything as `NEW`. Later runs return only `NEW`, `UPDATED`, and `REAPPEARED` products by default — duplicates and unchanged products are suppressed (and not charged). Turn on `emitUnchanged` or `emitExpired` only when you also want those rows returned (and billed for). State is isolated per mode/keyword/URL/sort/filter/detail setup automatically; set `stateKey` to name or deliberately share a monitoring campaign.

Scheduled-run example — same search, run daily:

Day 1 (first run ever for this search):

```json
{ "mode": "search", "queries": ["echo dot"], "incrementalMode": true }
```

→ every product comes back with `"changeType": "NEW"`.

Day 2 (the schedule fires again, identical input):

```json
{ "mode": "search", "queries": ["echo dot"], "incrementalMode": true }
```

→ products whose price/stock/rank/etc. changed come back as `"changeType": "UPDATED"` with `changedFields` listing what changed, brand-new products come back as `"changeType": "NEW"`, products that vanished and came back come back as `"changeType": "REAPPEARED"` — and products that are still there, unchanged, are **not** returned at all (suppressed, not charged) unless `emitUnchanged` is on.

**What counts as a change.** `price`, `listPrice`, `savings`/`savingsPercent`, `promoLabel` and `isOnSpecial` are real data and always trigger `UPDATED` — that is the point of monitoring a retailer. A few fields move on Amazon's own schedule regardless of whether the listing itself changed, so they are excluded from change detection (they are still returned on every row, just not compared): average `rating`, `reviewsCount`, `ratingDistribution` and the nested `reviewStats` (move with every vote cast, independent of the listing), `bestSellersRank` (Amazon recomputes category rank roughly hourly from live sales), `sponsored` (which ad auction won this placement on this exact search request), and `deliveryDate` (an estimate computed from today's date, so it advances daily on its own). The stock **count** inside `availability` ("Only 3 left in stock") is normalized away for the same reason, but the surrounding stock **state** ("In stock" ↔ "Only N left in stock" ↔ "Currently unavailable") still triggers `UPDATED` — only the fluctuating number is ignored. `url` is compared in its canonical `/dp/<ASIN>` form (the trailing search-rank segment, e.g. `ref=sr_1_13`, is stripped before comparison) since that segment is the product's rank on that search, not the product itself — the raw URL is still returned unchanged on every row.

### Notes

- Prices, ratings and stock are captured as shown on the storefront at scrape time.
- Filters (`minRating`, `minPrice`, `maxPrice`) are applied to each product's own values, so exact numeric ranges work independently of Amazon's coarse on-site filters.
- The run periodically checkpoints its progress, so if the platform restarts the run (a server migration, or you use Resurrect) it picks up where it left off instead of starting over, with no duplicate products and no double charges. No input is needed for this; it's automatic.

# Actor input Schema

## `mode` (type: `string`):

Choose 'search' to use keywords and filters, or 'url' to scrape specific Amazon Australia product, search or category URLs (or bare ASINs) you paste below.

## `queries` (type: `array`):

One or more keywords to search, for example 'echo dot' or 'coffee machine'. Each keyword is searched separately.

## `sortBy` (type: `string`):

Result ordering, as offered by Amazon.

## `specialsOnly` (type: `boolean`):

Only return products currently on Amazon Australia's Deals & Discounts (deals, price drops and coupons). Combine with a keyword to find deals in a category, or leave keywords empty to browse the whole deals grid. Each result carries its was/original price, savings, discount percentage and promo label when discounted.

## `urls` (type: `array`):

Product, search or category URLs (or bare 10-character ASINs) to scrape, for example https://www.amazon.com.au/s?k=echo+dot or a product page ending in /dp/<ASIN>. Multiple entries supported.

## `minRating` (type: `string`):

Optional. Only keep products whose average customer rating is at least this many stars.

## `minPrice` (type: `integer`):

Optional. Only keep products priced at or above this amount.

## `maxPrice` (type: `integer`):

Optional. Only keep products priced at or below this amount.

## `fetchDetails` (type: `boolean`):

Open each product page to collect its full detail (brand, list price, star distribution, availability, feature bullets, description, all images, category path, specifications, best-seller rank, seller). Adds one page fetch per product.

## `fetchReviews` (type: `boolean`):

Also collect the customer reviews Amazon publishes on each product page (star rating, title, text, author, date, verified-purchase flag, helpful votes) plus aggregate rating statistics. Uses the same product-page fetch as details. Note: Amazon Australia gates its full paginated review list behind sign-in for anonymous visitors, so this returns the top reviews Amazon exposes publicly.

## `maxReviewsPerProduct` (type: `integer`):

Cap on reviews collected per product when 'Fetch customer reviews' is on. Use 0 for all publicly available reviews on the product page.

## `maxItems` (type: `integer`):

Maximum number of products to return across the whole run. Use 0 for unlimited.

## `maxPages` (type: `integer`):

Optional bound on result pages walked per keyword/URL. Leave empty to walk every result page. This does NOT cap the number of products; the run stops at Max products.

## `resumeFromRunId` (type: `string`):

Optional. ID of a previous run of this actor (or a dataset ID). Products already in that dataset are skipped, so this run returns only NEW products (a delta). Combine both runs' datasets for the full set. Max products then counts only the new products. For recurring daily monitoring of the same search, use Incremental mode below instead.

## `incrementalMode` (type: `boolean`):

Daily/recurring monitoring of this same search or URL list. The actor remembers the previous run itself, so you never paste a run ID. The first run returns everything as NEW; later runs return only NEW, UPDATED and REAPPEARED products by default (unchanged products are suppressed and not billed). Combine with 'Resume from a previous run' only when bootstrapping a brand-new monitoring campaign from an existing dataset. See resumeFromRunId above for continuing one specific interrupted run instead.

## `stateKey` (type: `string`):

Optional name for a monitoring campaign, so its incremental state stays stable or is deliberately shared across otherwise-different search setups. Auto-derived from your mode/keywords/URLs/sort/filters/detail settings when left empty.

## `emitUnchanged` (type: `boolean`):

Incremental mode only. Also return products unchanged since the last run, marked changeType 'UNCHANGED'. Adds and bills extra rows you already have.

## `emitExpired` (type: `boolean`):

Incremental mode only. Also return products from a previous run no longer found, marked changeType 'EXPIRED', once a run has fully scanned the search/URLs (not capped by Max products/Max pages, not a resume). Adds and bills extra synthetic rows.

## `proxy` (type: `object`):

Apify Proxy is required for reliable results. The actor connects through an Australian residential exit by default.

## `mcpConnectors` (type: `array`):

Optionally send results into the apps you already use, via Model Context Protocol (MCP) connectors. Authorize one under Apify, Settings, API & Integrations, then select it here. Notion gets a rich page-per-item export; other connectors get a best-effort write/digest. Leave empty to skip; never changes the dataset output. Supported: Notion (https://mcp.notion.com/mcp), Linear (https://mcp.linear.app/sse), Airtable (https://mcp.airtable.com/mcp), Apify (https://mcp.apify.com).

## `notionParentPageUrl` (type: `string`):

URL or id of the Notion page under which item pages are created. Required to enable the Notion export; ignored by other connectors.

## `maxNotifyListings` (type: `integer`):

Cap on items written to each connector per run. Does not affect the dataset.

## Actor input object example

```json
{
  "mode": "search",
  "queries": [
    "echo dot"
  ],
  "sortBy": "relevance",
  "specialsOnly": false,
  "urls": [
    "https://www.amazon.com.au/s?k=echo+dot"
  ],
  "minRating": "0",
  "fetchDetails": true,
  "fetchReviews": false,
  "maxReviewsPerProduct": 10,
  "maxItems": 20,
  "incrementalMode": false,
  "emitUnchanged": false,
  "emitExpired": false,
  "proxy": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "AU"
  },
  "maxNotifyListings": 50
}
```

# Actor output Schema

## `overview` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "mode": "search",
    "queries": [
        "echo dot"
    ],
    "urls": [
        "https://www.amazon.com.au/s?k=echo+dot"
    ],
    "proxy": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ],
        "apifyProxyCountry": "AU"
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("abotapi/amazon-au-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "mode": "search",
    "queries": ["echo dot"],
    "urls": ["https://www.amazon.com.au/s?k=echo+dot"],
    "proxy": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
        "apifyProxyCountry": "AU",
    },
}

# Run the Actor and wait for it to finish
run = client.actor("abotapi/amazon-au-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "mode": "search",
  "queries": [
    "echo dot"
  ],
  "urls": [
    "https://www.amazon.com.au/s?k=echo+dot"
  ],
  "proxy": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "AU"
  }
}' |
apify call abotapi/amazon-au-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,abotapi/amazon-au-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/wwVdY3x1WdS1uVPml/builds/YvyA7VV4OmwaCnw4f/openapi.json
