# Webpage Metadata Extractor - OpenGraph, Twitter Cards, JSON-LD (`eliai/webpage-metadata-extractor`) Actor

Extract webpage metadata from any URL. Input a URL, get one JSON record with the page title, meta tags, OpenGraph, Twitter Card, and JSON-LD structured data. $0.02 per page analyzed — for link previews, SEO checks, and AI agent pipelines.

- **URL**: https://apify.com/eliai/webpage-metadata-extractor.md
- **Developed by:** [Broke to Built](https://apify.com/eliai) (community)
- **Categories:** Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $16.00 / 1,000 metadata extractions

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Webpage Metadata Extractor

Get the **full machine-readable metadata** of any web page in one call — title, meta tags, OpenGraph, Twitter cards, **JSON-LD structured data**, favicon, canonical, feeds, hreflang, and page stats. No API key, no proxy setup, pay per page.

**▶ Live on the Apify Store** — run it instantly, or call it as an agent tool via Apify MCP.

### Why

Building link previews, knowledge graphs, SEO tools, or feeding clean page metadata to an LLM/agent? This returns everything a page declares about itself as structured JSON — including the `schema.org` JSON-LD that powers rich results.

### What it extracts

- **Core**: title, description, keywords, author, canonical, language, charset, viewport, robots, generator, theme-color, favicon
- **OpenGraph** (`og:*`) and **Twitter Card** (`twitter:*`) — every property, as objects
- **JSON-LD structured data** — fully parsed, plus the list of `schema.org` `@type`s found
- **Feeds** (RSS/Atom) and **hreflang** alternates
- **Stats**: H1 list, H2 count, image count, internal vs external link counts, word count

### Input

```json
{ "url": "https://example.com/" }
```

or bulk:

```json
{ "urls": ["https://a.com/", "https://b.com/"], "maxUrls": 50 }
```

### Output (per page)

```json
{
  "url": "https://example.com/",
  "title": "Example Domain",
  "description": "...",
  "canonical": "https://example.com/",
  "openGraph": { "og:title": "...", "og:image": "..." },
  "twitterCard": { "twitter:card": "summary_large_image" },
  "schemaTypes": ["Organization", "WebSite"],
  "jsonLd": [ { "@context": "https://schema.org/", "@type": "Organization" } ],
  "counts": { "h1": 1, "images": 12, "internalLinks": 30, "externalLinks": 5, "words": 820 }
}
```

### Notes

Reads only the public HTML of the URL you provide. Single fetch per page (plus follows redirects) — fast and reliable on any site.

### Pricing

**$0.02 per page analyzed** — billed as the `page-analyzed` event. Bulk runs are capped by `maxUrls`, which is also your budget cap.

### FAQ

**How do I get the Open Graph tags and title of a URL?** Pass it in — one record comes back with the title, every meta tag, all `og:*` and `twitter:*` properties as objects, canonical, favicon, and language: everything you need to render a link preview.

**Does it extract JSON-LD structured data?** Yes, fully parsed — plus a `schemaTypes` list of the schema.org `@type`s found, so you can see at a glance whether a page declares `Product`, `Article`, `Organization`, and so on.

**What else comes back beyond meta tags?** RSS/Atom feed links, hreflang alternates, and page stats: H1 list, H2 count, image count, internal vs external link counts, and word count.

**Can I use this to build link previews in my app?** Yes — that's the core use. Call it via the run-sync API endpoint, read `title`, `description`, and `openGraph["og:image"]`, and render. The fallback chain is your job, but every layer is in the record.

**Does it execute JavaScript?** No — it reads the served HTML in a single fetch (following redirects), which is exactly what social platforms and search engines read for metadata, and what keeps it fast at $0.02.

### For AI agents

This Actor is built to be called by software, not just by people.

- **Mount it directly as an MCP tool** — no Store search, no ranking, just this one tool:
  `https://mcp.apify.com/?actors=eliai/webpage-metadata-extractor`
- **Or call it over HTTP** and get the results in the same request:
  `POST https://api.apify.com/v2/acts/eliai~webpage-metadata-extractor/run-sync-get-dataset-items`
- **Pay with x402, without an Apify account.** This Actor is whitelisted for agentic payments, so an agent holding USDC on Base can buy a prepaid token and spend it here. The minimum purchase is $1, the token balance is an absolute spending cap, and it expires 14 days after purchase.
- **Costs are predictable before you call.** Pricing is pay-per-event (see Pricing above), so an agent can budget a run in advance instead of discovering the bill afterwards.
- **Send only the field you mean.** If you pass the bulk field, it is used on its own; the single-value field is a fallback, never merged into your request. You are charged for the items you sent and nothing else.

# Actor input Schema

## `url` (type: `string`):

A single page URL to extract metadata from.

## `urls` (type: `array`):

Multiple page URLs to process in one run.

## `maxUrls` (type: `integer`):

Safety cap on how many URLs to process.

## Actor input object example

```json
{
  "url": "https://apify.com/",
  "urls": [],
  "maxUrls": 50
}
```

# Actor output Schema

## `results` (type: `string`):

Every item this run produced, as JSON.

## `resultsCsv` (type: `string`):

The same items as a spreadsheet-ready CSV.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "url": "https://apify.com/",
    "urls": []
};

// Run the Actor and wait for it to finish
const run = await client.actor("eliai/webpage-metadata-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "url": "https://apify.com/",
    "urls": [],
}

# Run the Actor and wait for it to finish
run = client.actor("eliai/webpage-metadata-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "url": "https://apify.com/",
  "urls": []
}' |
apify call eliai/webpage-metadata-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,eliai/webpage-metadata-extractor"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/g4qRapWr1lNzmboua/builds/INHOael2OFtbid85n/openapi.json
