# URL Metadata & OpenGraph Extractor (`scrapeworks/url-metadata`) Actor

Extract clean metadata from any list of URLs: title, description, preview image, favicon, site name, and more. Reconciles OpenGraph, Twitter Card, and standard meta tags into one tidy record per URL.

- **URL**: https://apify.com/scrapeworks/url-metadata.md
- **Developed by:** [Nicolas van Arkens](https://apify.com/scrapeworks) (community)
- **Categories:** Automation, Other, Open source
- **Stats:** 6 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## URL Metadata & OpenGraph Extractor 🔗

Give it any list of URLs and get back clean, unified metadata for each page — **title, description, preview image, favicon, site name, canonical URL, author, and published date**. It reconciles **OpenGraph**, **Twitter Card**, and standard `<meta>` tags into one tidy record per URL, so you don't have to care which tags a given site happens to use.

Perfect for building link previews, social-media tools, bookmarking apps, content curation, and SEO audits.

### Why use it

- 🔀 **Reconciled metadata** — OpenGraph → Twitter Card → standard meta → `<title>`, in priority order
- 🖼️ **Preview image & favicon** — resolved to absolute URLs (with a sensible favicon fallback)
- 🔗 **Canonical URL** — and the final URL after redirects
- 🏷️ **Rich fields** — site name, type, author, published date, theme color, Twitter card type, keywords
- 📦 **Batch** — process many URLs in a single run
- 🪶 **Fast & light** — metadata-only extraction, no heavy rendering

### Use cases

- **Link previews** — build rich cards like Slack, Discord, and iMessage show
- **Social media tools** — pull share metadata for scheduling and preview
- **Bookmarking & read-later apps** — auto-title and illustrate saved links
- **Content curation & newsletters** — enrich link lists with titles and images
- **SEO audits** — check OG/Twitter tags across a set of pages at once

### Input

| Field | Description |
|-------|-------------|
| **URLs** | List of page URLs (scheme added automatically if omitted). |

### Output

```json
{
  "url": "https://example.com/page",
  "finalUrl": "https://example.com/page",
  "success": true,
  "statusCode": 200,
  "title": "OG Title Wins",
  "description": "The OpenGraph description.",
  "image": "https://cdn.example.com/img.jpg",
  "siteName": "Example Site",
  "type": "article",
  "canonicalUrl": "https://example.com/canonical-page",
  "favicon": "https://example.com/favicon.png",
  "author": "Jane Author",
  "publishedDate": "2026-05-01T10:00:00Z",
  "twitterCard": "summary_large_image"
}
```

Export to JSON, CSV, or Excel, or pull via the Apify API.

### Notes

- Extracts metadata from publicly accessible pages only. Some sites block automated requests or require JavaScript rendering, in which case fewer fields may be available.
- Independent tool; respects each site's access policy. Please use responsibly.

# Actor input Schema

## `urls` (type: `array`):

List of page URLs to extract metadata from. A scheme is added automatically if you omit it (e.g. 'example.com' becomes 'https://example.com/').

## Actor input object example

```json
{
  "urls": [
    "https://www.apify.com/",
    "https://github.com/",
    "https://news.ycombinator.com/"
  ]
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.apify.com/",
        "https://github.com/",
        "https://news.ycombinator.com/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapeworks/url-metadata").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://www.apify.com/",
        "https://github.com/",
        "https://news.ycombinator.com/",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("scrapeworks/url-metadata").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.apify.com/",
    "https://github.com/",
    "https://news.ycombinator.com/"
  ]
}' |
apify call scrapeworks/url-metadata --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapeworks/url-metadata"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/FmUFrTFopUJfk0Gtm/builds/bE81siDwTrSTM8Je7/openapi.json
