# Structured Data Extractor - JSON-LD, Microdata, RDFa (`eliai/webpage-structured-data-extractor`) Actor

Extract all structured data from any page - JSON-LD, Microdata, and RDFa - plus the schema.org types found. Single or bulk. $0.0004 per page, cheaper than every paid structured-data actor measured; failed fetches are free.

- **URL**: https://apify.com/eliai/webpage-structured-data-extractor.md
- **Developed by:** [Broke to Built](https://apify.com/eliai) (community)
- **Categories:** SEO tools, Developer tools, Automation
- **Stats:** 1 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.32 / 1,000 scanned pages

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Structured Data Extractor - JSON-LD, Microdata & RDFa from any URL

Extract all **structured data** from a web page: **JSON-LD** (`<script type="application/ld+json">`),
**Microdata** (`itemscope`/`itemprop`), and **RDFa** (`typeof`/`property`), plus the list of
**schema.org types** found. Built for SEO audits, rich-result debugging, knowledge-graph building,
and AI agents that need a page's semantic data. Single URL or bulk.

**$0.0004 per page - cheaper than every paid structured-data actor measured on the Store.** Failed
fetches are recorded free.

### What you get

- **JSON-LD** objects, with `@graph` unwrapped into individual entities.
- **Microdata** items with nested scopes resolved into properties.
- **RDFa** typed resources and their properties.
- **`allTypes`** - every schema.org type detected (Product, Organization, BreadcrumbList, ...).
- **`entityCount`** across all three formats.
- **Fail-soft**: an unreachable URL returns `{ok:false, error}` and is **never charged**.

### Input

```json
{ "url": "https://www.apify.com/", "urls": ["https://example.com/product"], "maxUrls": 25 }
```

### Output (real run, 2026-08-07)

```json
{
  "ok": true,
  "url": "https://apify.com/",
  "status": 200,
  "jsonLd": [
    { "@context": "https://schema.org/", "@type": "Organization", "name": "Apify",
      "url": "https://apify.com/", "sameAs": ["https://github.com/apify", "https://x.com/apify"] }
  ],
  "microdata": [],
  "rdfa": [],
  "allTypes": ["ContactPoint", "Organization"],
  "entityCount": 1
}
```

### Pricing - $0.0004 per page

Prices below checked via the Apify Store API on 2026-08-07:

| Actor | Pricing | One page |
|---|---|---|
| **This actor** | **$0.0004 per page** | **$0.0004** |
| sync-network/schema-org-json-ld-extractor | $0.00005 start + $0.0005 per item | $0.00055 |
| andok/jsonld-extractor | $0.01 start + $0.001 per item | $0.011 |
| pink\_comic/schema-markup-extractor | $0.0001 start + $0.002 per item | $0.0021 |

### Limits (honest ones)

- Parses the HTML the server returns; structured data injected later by JavaScript is not seen (no headless browser).
- Reads the raw values as authored - it does not validate them against Google's rich-result requirements.
- `maxUrls` capped at 50 per run.

### FAQ

- **Which formats does it cover?** All three: JSON-LD, Microdata, and RDFa, in one record per page.
- **How do I extract JSON-LD from a URL?** Pass `{"url": "https://example.com/product"}`; every `<script type="application/ld+json">` block comes back parsed in the `jsonLd` array, with `@graph` unwrapped into individual entities.
- **Does it run JavaScript?** No - it reads the server HTML, which is where the vast majority of structured data lives.
- **What if a page has none?** You get `entityCount: 0` with empty arrays (a real answer), still one page scanned.
- **Can it do many pages?** Yes - pass `urls` (up to 50).
- **What schema.org types will it tell me a page uses?** `allTypes` lists every type found across all three formats, de-duplicated and sorted - `Product`, `Organization`, `BreadcrumbList`, `FAQPage`, `Article`, whatever the page declares.
- **Can an AI agent use this to read a page's semantic data?** Yes - it is callable through the Apify MCP server and the record shape is fixed, so a model can rely on the field names.

### When not to use this

- **You want to know if Google will show a rich result.** This reports the markup as authored. It does
  not validate against Google's rich-result requirements or flag missing required properties - use
  the Rich Results Test for that verdict.
- **Your markup is injected by JavaScript.** No headless browser here. A React app that adds JSON-LD
  on mount will read as empty; server-rendered markup is what this sees.
- **You want to crawl a site's structured data.** It reads exactly the URLs you pass, up to 50 per
  run. Generate the URL list yourself (a sitemap extractor pairs well with this).
- **You want Open Graph / Twitter Card tags.** Those are social meta tags, not structured data - see
  our Open Graph & Twitter Card Extractor.
- **You need the markup rewritten or generated.** This is read-only: it extracts, it does not author
  or repair schema.

### Use from code or AI agents

```bash
curl -X POST "https://api.apify.com/v2/acts/EliAI~webpage-structured-data-extractor/runs?token=YOUR_APIFY_TOKEN" \
  -H 'content-type: application/json' \
  -d '{"url":"https://www.apify.com/"}'
```

Callable as an agent tool through the Apify MCP server (`mcp.apify.com`).

# Actor input Schema

## `url` (type: `string`):

A single webpage URL to extract structured data from (JSON-LD, Microdata, RDFa).

## `urls` (type: `array`):

Multiple webpage URLs to process. Combined with the single URL above.

## `maxUrls` (type: `integer`):

Maximum number of URLs to process (1-50).

## Actor input object example

```json
{
  "url": "https://www.apple.com/shop/buy-iphone/iphone-15-pro",
  "urls": [],
  "maxUrls": 25
}
```

# Actor output Schema

## `results` (type: `string`):

Every item this run produced, as JSON.

## `resultsCsv` (type: `string`):

The same items as a spreadsheet-ready CSV.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "url": "https://www.apple.com/shop/buy-iphone/iphone-15-pro"
};

// Run the Actor and wait for it to finish
const run = await client.actor("eliai/webpage-structured-data-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "url": "https://www.apple.com/shop/buy-iphone/iphone-15-pro" }

# Run the Actor and wait for it to finish
run = client.actor("eliai/webpage-structured-data-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "url": "https://www.apple.com/shop/buy-iphone/iphone-15-pro"
}' |
apify call eliai/webpage-structured-data-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,eliai/webpage-structured-data-extractor"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/WnwV69cYuhfq9UDAl/builds/TWKhk06MZBHOa3h3f/openapi.json
