# Ozon Scraper - Products, Prices & Ratings (`scrapesage/ozon-scraper`) Actor

Scrape Ozon.ru category pages into clean rows: SKU, title, numeric price and original price, rating, review count, images and product URL. Built-in unblocker with automatic retries. Ideal for price monitoring, discount tracking and competitor research on Russia's largest marketplace.

- **URL**: https://apify.com/scrapesage/ozon-scraper.md
- **Developed by:** [Scrape Sage](https://apify.com/scrapesage) (community)
- **Categories:** E-commerce, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.20 / 1,000 product scrapeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Ozon Scraper - Products, Prices & Ratings

> **Disclaimer:** This Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by Ozon Holdings PLC or any of its subsidiaries. All trademarks mentioned are the property of their respective owners. "Ozon" is referenced only to describe the publicly available website this Actor collects data from.

Scrape **Ozon.ru** category and product pages into clean rows: title, current price, original price, rating, review count, images, SKU and the product URL. Built for **price monitoring, catalogue tracking and competitor research** on Russia's largest marketplace.

### Why this scraper?

| Typical approach | This actor |
|---|---|
| Datacenter proxies get refused outright | Uses a **paid unblocker** - the transport Ozon actually answers |
| One flaky attempt per page | **Retries every page up to 3×** - Ozon's unlock succeeds intermittently, so the actor absorbs that instead of handing you an empty run |
| Prices as display strings | **Numeric `price` and `originalPrice`** plus the original text, so you can sort and compare without parsing `64 705 ₽` yourself |
| Rating buried in nested widget JSON | **`rating` and `reviewCount` as plain numbers** |
| Paste URLs one at a time | Paste them, upload a file, or **link a remote .txt / .csv / Google Sheet** |

### Use cases

- **Price monitoring** - re-run a category on a schedule and track `price` against `originalPrice`.
- **Competitor & assortment research** - pull a whole category and compare ratings, review counts and pricing.
- **Discount discovery** - filter to rows where `originalPrice` is set to find what is actually reduced.
- **Catalogue enrichment** - resolve SKUs to titles, images and current prices.

### How to use

1. [Sign up for Apify](https://console.apify.com/sign-up) - the free plan is enough to try this actor.
2. Paste one or more **Ozon category URLs** (for example `https://www.ozon.ru/category/noutbuki-15692/`), set how many pages and products you want, and click **Start**.
3. Watch products stream into the dataset.
4. **Export** as JSON, CSV, Excel, XML or RSS - or pull results over the [Apify API](https://docs.apify.com/api/v2).

Leave the URL fields empty and the actor runs a short built-in sample on one real category so you can see the output shape first.

### Input

- **`startUrls`** - Ozon **category or product** URLs. All four Console shapes work: a plain string, `{"url": "..."}`, `{"requestsFromUrl": "..."}`, and file upload.
- **`urlsFromFile`** - paste a block of URLs (one per line) **or** one link to a `.txt` / `.csv` / Google Sheet / Drive file holding the list.
- **`maxResults`** - cap on total products written across every URL. Each product written is one billable result.
- **`maxPagesPerUrl`** - how many paginated pages to follow per category URL.

### Output

| Field | What it holds |
|---|---|
| `sku`, `title`, `url` | Ozon SKU, full product title, and the product page URL |
| `price`, `priceText`, `currency` | Current price as a **number**, the original display string, and `RUB` |
| `originalPrice` | The struck-through price, when the item is discounted |
| `rating`, `reviewCount`, `reviewsText` | Star rating and review count as numbers, plus Ozon's own wording |
| `images`, `imageCount` | Every product image URL on the tile |
| `brandLogo`, `isAdult` | Brand-store logo when the seller has one; adult-content flag |
| `sourceUrl`, `scrapedAt` | The URL this row came from, and when |

Some fields are populated **only when they apply**, which is Ozon's data and not missing extraction: `originalPrice` only on discounted items (measured 3 of 8 tiles), `rating`/`reviewCount`/`reviewsText` only once a product has reviews (5 of 8), and `brandLogo` only for brand-store sellers (Ozon ships the key as `null` otherwise).

### Reliability - read this before you buy

- **Scope: category and product URLs, not search.** Ozon does not serve its `/search/` results to this transport - the page comes back as a stub. Rather than sell a feature that silently returns nothing, `/search/` URLs are **skipped with an explicit message** telling you to use a category URL instead.
- **The unlock is intermittent, and the actor absorbs it.** Measured 2026-09-04, roughly half of first attempts returned an empty body while the same URL succeeded moments later. Every page is therefore retried up to 3× with backoff. If a page still cannot be fetched it is skipped, named in the run message, and **not billed**.
- **You are only charged for products actually written to your dataset.** A run that fetches nothing bills nothing.
- **Long jobs stop cleanly at the run's time budget** with partial results and a message saying so, rather than timing out.
- **Ozon pages are slow to unlock** - expect tens of seconds per page. Set the run timeout generously for large jobs.

### How much does it cost?

Pay-per-event, no monthly rental:

- **Product scraped** - charged once per product written to your dataset.

The current price is on the **Pricing** tab. Set a maximum total charge on the run for a hard spend ceiling, and use `maxResults` to cap a run.

### Automate & schedule

Run this Actor on autopilot and pull results into your own stack:

- **[Apify API](https://docs.apify.com/api/v2)** - start runs, fetch datasets and manage schedules over REST.
- **[apify-client for JavaScript](https://docs.apify.com/api/client/js/)** and **[apify-client for Python](https://docs.apify.com/api/client/python/)** - official SDKs.
- **[Schedules](https://docs.apify.com/platform/schedules)** - run it hourly, daily or weekly and keep your dataset current.
- **[Webhooks](https://docs.apify.com/platform/integrations/webhooks)** - trigger downstream actions (CRM import, Slack alert, email sequence) the moment a run finishes.

```js
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: 'MY_APIFY_TOKEN' });

const run = await client.actor('scrapesage/ozon-scraper').call({
    // your input - see the Input section above
});

const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(`Got ${items.length} records`);
```

### Integrate with any app

Connect the dataset to thousands of apps - no code required:

- **[Make](https://docs.apify.com/platform/integrations/make)** - multi-step automation scenarios.
- **[Zapier](https://docs.apify.com/platform/integrations/zapier)** - push new records straight into your CRM or spreadsheet.
- **[Slack](https://docs.apify.com/platform/integrations/slack)** - get notified when a scheduled run finds something new.
- **[Google Drive / Sheets](https://docs.apify.com/platform/integrations/drive)** - auto-export every run to a spreadsheet.
- **[Airbyte](https://docs.apify.com/platform/integrations/airbyte)** - pipe results into your data warehouse.
- **[GitHub](https://docs.apify.com/platform/integrations/github)** - trigger runs from commits or releases.

### Use with AI assistants (MCP)

The output is clean, LLM-ready JSON. You can call this Actor from Claude, ChatGPT or any agent framework through the **[Apify MCP server](https://docs.apify.com/platform/integrations/mcp)** - describe the data you need and let the assistant run this Actor for you.

### Agent-ready: autonomous payments (x402 & Skyfire)

This actor is **agent-ready** — AI agents can discover it, run it, and **pay for it autonomously**, with no Apify account and no human in the loop. It uses [pay-per-event](https://docs.apify.com/platform/actors/publishing/monetize/pay-per-event) pricing and [limited permissions](https://docs.apify.com/platform/actors/development/permissions), so it qualifies for Apify's agentic-payment standards:

- **[x402](https://docs.apify.com/platform/integrations/x402)** — an open, HTTP-native payment protocol. Agents pay per run in USDC on the Base network directly through the [Apify MCP server](https://docs.apify.com/platform/integrations/mcp) — no account, no API key.
- **[Skyfire](https://docs.apify.com/platform/integrations/skyfire)** — agent-to-service payments for fully autonomous AI-agent workflows.

Building an AI agent, MCP tool, or autonomous data pipeline? This scraper is ready to plug in and pay as it goes.

### Tips

- Start with `maxPagesPerUrl: 1` and a small `maxResults` to confirm the category is right, then scale up.
- Category URLs return more reliably than deep pagination - split a big job across several category URLs rather than one very deep crawl.
- Filter the dataset to rows where `originalPrice` is not null to get only discounted products.

### FAQ

**Can I scrape Ozon search results?** No - Ozon does not serve `/search/` to this transport. Use the equivalent category URL, which does work.

**Why did a page get skipped?** Ozon's unlock is intermittent. The actor retries 3× and then skips, naming it in the run message. Re-running usually picks it up.

**Are prices numbers?** Yes - `price` and `originalPrice` are numeric; `priceText` keeps Ozon's original string.

**Does it need a proxy of my own?** No. The unblocker transport is built in.

### Data & lawful use

This Actor reads public listings and the seller information shown on them. Sellers on Ozon include private individuals, so seller names, usernames and any contact details are personal data. If you are in the EU or UK you are the data controller for what you do with them: have a lawful basis, use the data for market research, price monitoring or lawful sourcing, honour deletion requests, and do not send private sellers unsolicited marketing.

Under [Apify's Standard Actor Contract](https://docs.apify.com/legal/standard-actor-contract), which governs your use of this Actor, you are the controller of any personal data in your input and output and scrapesage acts only as your processor: that data is processed solely to run your job, written only to your own Apify storage, never used for any other purpose and never shared onward. If you need help with a data-subject request that involves this Actor's output, open an issue on the Issues tab.

### Disclaimer

**This Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by Ozon Holdings PLC or any of its subsidiaries. All trademarks mentioned are the property of their respective owners.**

"Ozon" and any related marks are the property of their respective owners and are used here only in a descriptive, nominative sense - to identify the publicly accessible website from which this Actor collects data. This Actor is not an official Ozon product, is not authorised or certified by Ozon Holdings PLC, and does not distribute Ozon software. It collects only publicly available information; you are responsible for ensuring your use of that data complies with applicable laws, regulations and the terms of the source website.

### Need help?

Open an issue on the Actor's **Issues** tab, or visit the [Apify help center](https://help.apify.com/). Feature requests are welcome - this Actor is actively maintained.

# Actor input Schema

## `startUrls` (type: `array`):

Ozon.ru CATEGORY or PRODUCT URLs. Paste them, upload a file, or link a remote .txt/.csv/Google Sheet. Note: Ozon /search/ URLs are not supported - Ozon does not serve its search results to this transport, so category URLs are used instead.

## `urlsFromFile` (type: `string`):

Alternative to the field above: paste a block of Ozon URLs (one per line) OR one link to a .txt / .csv / Google Sheet / Drive file holding the list.

## `maxResults` (type: `integer`):

Cap on total products written across every URL. Each product written is one billable result.

## `maxPagesPerUrl` (type: `integer`):

How many paginated pages to follow per category URL.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://www.ozon.ru/category/noutbuki-15692/"
    }
  ],
  "maxResults": 100,
  "maxPagesPerUrl": 3
}
```

# Actor output Schema

## `results` (type: `string`):

Every product scraped in this run.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://www.ozon.ru/category/noutbuki-15692/"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapesage/ozon-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [{ "url": "https://www.ozon.ru/category/noutbuki-15692/" }] }

# Run the Actor and wait for it to finish
run = client.actor("scrapesage/ozon-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://www.ozon.ru/category/noutbuki-15692/"
    }
  ]
}' |
apify call scrapesage/ozon-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapesage/ozon-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/f3ldlTwCHtptdSTwt/builds/28AbIPLQ1vZmcL7XR/openapi.json
