# CWTS Leiden Ranking Scraper (`parseforge/leiden-ranking-scraper`) Actor

Scrapes CWTS Leiden Ranking university data by publication period and optional country filter. Returns each university as a flat row with impact, collaboration, and open access indicators.

- **URL**: https://apify.com/parseforge/leiden-ranking-scraper.md
- **Developed by:** [ParseForge](https://apify.com/parseforge) (community)
- **Categories:** Developer tools, Other
- **Stats:** 2 total users, 1 monthly users, 86.7% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.43 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

[![ParseForge](https://raw.githubusercontent.com/ParseForge/apify-assets/main/banner.jpg)](https://apify.com/parseforge?fpr=vmoqkp)

### CWTS Leiden Ranking Scraper

**Scrape CWTS Leiden Ranking university data for any publication period and country, up to a million rows per run.** Each university comes with its scientific impact indicators, collaboration metrics, and open access stats. No login or API key. Export to CSV, JSON, Excel, or XML.

The CWTS Leiden Ranking is a leading source for university research performance, but its official website offers no bulk export and manual copying is slow. This scraper reads the public ranking pages directly, filtered by publication period and country, and returns each university in one fixed schema. It covers scientific impact, collaboration, and open access indicators for hundreds of universities worldwide.

| Who uses it | What they scrape CWTS Leiden Ranking for |
|---|---|
| University administrators | Benchmark their institution against peers on research impact and collaboration |
| Higher education consultants | Build comparative reports for clients on university performance |
| Academic researchers | Analyze trends in scientific output and open access across countries |
| Policy analysts | Track national research strengths and international collaboration patterns |

### What it does

This Actor collects CWTS Leiden Ranking university records by publication period and optional country filter, and returns each university as a flat row with its ranking metrics.

- 📊 **Publication period filter:** choose from 2019-2022, 2018-2021, or 2017-2020 to match your analysis window.
- 🌍 **Country filter:** narrow results to a single country by name, such as Germany or Japan.
- 🔢 **Maximum universities:** set a cap from 1 to 1,000,000 rows per run to control dataset size.
- 📥 **Structured output:** each university is returned as a flat row with all ranking indicators, ready for export.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

### What you can do with CWTS Leiden Ranking data

**🏫 Benchmark university performance.**

A university planning office runs the scraper for the latest period and their country to compare their institution's impact and collaboration scores against national peers.

**📈 Track research trends over time.**

A higher education analyst collects data for multiple publication periods to see how open access rates and international collaboration have changed across universities.

**🌐 Map global research strengths.**

A policy researcher scrapes all universities without a country filter to identify which nations lead in scientific impact and which are rising.

**📊 Feed dashboards and reports.**

A data engineer schedules the scraper to refresh a dashboard with the latest CWTS Leiden Ranking data for internal stakeholders.

### Why choose this scraper

|  | What you get |
|---|---|
| **No API key** | Access the public ranking data without registration or rate limits |
| **Bulk export** | Download hundreds or thousands of university records in one run |
| **Fixed schema** | Every row has the same fields, so you can merge runs and compare periods |
| **Up-to-date** | Scrapes the latest published ranking data for the selected period |

### How it compares

No other Store actor targets CWTS Leiden Ranking the same way, so the honest comparison is with the alternatives teams actually weigh.

| | CWTS Leiden Ranking Scraper | Build it in-house | By hand |
|---|---|---|---|
| Setup | Run it now, zero config | Days of engineering | None, but hours per pull |
| When CWTS Leiden Ranking changes | Maintained for you | You fix it | You re-learn the page |
| Proxies, retries, anti-bot | Built in | Your problem | Browser only |
| Output | Fixed JSON schema, CSV/Excel export | Whatever you build | Copy-paste |
| Cost | Pay per result | Engineering time | Analyst hours |

### Configure the run

Drive the Actor with a publication period and optional country filter, and set a maximum number of universities to collect per run. The Input tab lists every parameter.

A first run with the defaults:

```json
{
  "maxItems": 10
}
```

A larger pull:

```json
{
  "maxItems": 200
}
```

### Pricing

Pay-per-result: **$0.006 per result** collected. You pay only for the results written to your dataset.

| Results collected | Approximate cost |
|---|---|
| 100 results | $0.60 |
| 1,000 results | $6.00 |
| 10,000 results | $60.00 |

New Apify accounts start with $5 in free credit.

### Free users

Free-plan runs return up to 10 results as a preview. [Upgrade your Apify plan](https://console.apify.com/sign-up?fpr=vmoqkp) to collect up to 1,000,000 results per run.

### Run it

1. [Create a free Apify account with $5 in credit](https://console.apify.com/sign-up?fpr=vmoqkp).
2. Open the [CWTS Leiden Ranking Scraper](https://apify.com/parseforge/leiden-ranking-scraper?fpr=vmoqkp).
3. Set your inputs and any filters, then click **Start**.
4. Export the results as CSV, Excel, JSON, or XML from the **Dataset** tab.

Run it programmatically through the [Apify API](https://docs.apify.com/api/v2) (`run-sync-get-dataset-items`) or the [ApifyClient](https://docs.apify.com/api/client/js) for JavaScript and Python.

### Use with AI agents (MCP)

Give an AI agent live access to CWTS Leiden Ranking through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

```bash
claude mcp add --transport http apify "https://mcp.apify.com/?tools=parseforge/leiden-ranking-scraper"
```

Then prompt it in plain language to run the scraper and read back the results.

### Troubleshooting

**Why am I getting no results?**

Make sure the country filter is spelled correctly and matches the name used on the CWTS Leiden Ranking site. Also verify that the selected publication period is valid.

**The run stops before collecting all universities.**

Increase the 'Maximum universities' input to a higher number, up to 1,000,000. If you still hit a limit, check your Apify plan's memory and timeout settings.

**Some fields are empty in the output.**

Not all universities have data for every indicator. Empty fields mean the ranking site did not provide a value for that metric.

**The scraper fails with a timeout error.**

Try reducing the maximum number of universities or run during off-peak hours. You can also increase the timeout in the actor's run settings.

### FAQ

| Question | Answer |
|---|---|
| What is the CWTS Leiden Ranking? | It is an annual ranking of universities worldwide based on bibliometric indicators from Web of Science data, focusing on scientific impact, collaboration, and open access. |
| Do I need an API key or login? | No. The scraper reads the public ranking pages directly, so no registration or authentication is required. |
| Can I filter by country? | Yes, you can enter a country name in the optional country filter to return only universities from that country. |
| Which publication periods are available? | The scraper supports 2019-2022, 2018-2021, and 2017-2020. You can select one per run. |
| How many universities can I scrape in one run? | You can set a maximum from 1 to 1,000,000 universities. The default is 10, but you can increase it to collect the full ranking. |
| What data fields are returned for each university? | Each row includes the university name, country, and all ranking indicators such as scientific impact, collaboration, and open access metrics. The exact fields are shown in the sample output. |
| In what formats can I export the data? | You can export the dataset as CSV, JSON, Excel, or XML from the Apify platform. |
| Is the data updated automatically? | The scraper fetches the current data from the CWTS Leiden Ranking website each time you run it, so you always get the latest published figures. |
| Can I schedule regular scrapes? | Yes, you can set up a schedule in Apify to run the scraper daily, weekly, or monthly to keep your dataset fresh. |
| What if I get no results? | Check that your country filter matches the exact name used on the ranking site, and ensure the selected publication period is available. If the problem persists, contact support. |

### Related actors

Browse the full [ParseForge collection](https://apify.com/parseforge?fpr=vmoqkp) for more scrapers.

🆘 **Need help?** Email parseforge@protonmail.com with your run ID, your input, and what you expected.

⚠️ **Disclaimer.** This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by Centre for Science and Technology Studies, Leiden University. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.

# Actor input Schema

## `maxItems` (type: `integer`):

How many universities to collect per run.

## `period` (type: `string`):

Multi-year publication window (e.g. 2019-2022).

## `country` (type: `string`):

Optional country name filter.

## Actor input object example

```json
{
  "maxItems": 10,
  "period": "2019-2022"
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("parseforge/leiden-ranking-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "maxItems": 10 }

# Run the Actor and wait for it to finish
run = client.actor("parseforge/leiden-ranking-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "maxItems": 10
}' |
apify call parseforge/leiden-ranking-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,parseforge/leiden-ranking-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/XKdEkFvXAeE9cDrCe/builds/BPWripp2obLqWNu10/openapi.json
