# OpenAIRE Publications Scraper (`parseforge/openaire-scraper`) Actor

Scrapes open access research publications from OpenAIRE by search query and optional year range. Returns each publication as a flat row with title, authors, DOI, date, and access status.

- **URL**: https://apify.com/parseforge/openaire-scraper.md
- **Developed by:** [ParseForge](https://apify.com/parseforge) (community)
- **Categories:** Education, Other, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $19.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

[![ParseForge](https://raw.githubusercontent.com/ParseForge/apify-assets/main/banner.jpg)](https://apify.com/parseforge?fpr=vmoqkp)

### OpenAIRE Publications Scraper

**Scrape open access research publications from OpenAIRE by keyword, year range, or both, up to a million per run.** Every record comes with its title, authors, DOI, publication date, and access status. No API key or registration. Export to CSV, JSON, Excel, or XML.

OpenAIRE aggregates millions of open access publications from repositories, journals, and aggregators across Europe and beyond. Its official API requires registration and has rate limits. This Actor reads the public search results directly, filtered by search term and publication year, and returns each match in one fixed schema.

| Who uses it | What they scrape OpenAIRE for |
|---|---|
| Academic researchers | Building a literature review dataset for a specific topic |
| Librarians | Monitoring new open access publications in their institution's field |
| Data analysts | Tracking publication trends over time for reporting |
| Grant managers | Finding open access outputs from funded projects |

### What it does

This Actor collects research publications from OpenAIRE by search query and optional year range, and returns each one as a flat row.

- 🔍 **Keyword search:** any term, phrase, or boolean query OpenAIRE supports, e.g. 'machine learning' or 'climate change'.
- 📅 **Year range filter:** set fromYear and toYear to narrow results to a specific period.
- 📦 **Bulk collection:** collect up to 1,000,000 records per run with a single input.
- 📄 **Flat output:** each publication is one row with title, authors, DOI, date, and access status.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

### What you can do with OpenAIRE data

**📚 Build a literature review dataset.**

A PhD student enters a research topic and collects all matching open access publications from the last five years to seed their reference manager.

**📈 Track publication trends.**

A data analyst runs the Actor monthly with a fixed query and year range to count new publications and spot emerging topics.

**🔎 Monitor open access compliance.**

A librarian checks which publications from their institution's researchers are openly available in OpenAIRE.

**🌍 Map research output by region.**

A policy researcher collects publications mentioning a country name and analyzes the geographic distribution of authors.

### Why choose this scraper

| | What you get |
|---|---|
| **No API key** | No registration or OAuth, a search term |
| **Open access focus** | Returns publications that are freely available to read |
| **Year filtering** | Limit results to a specific publication window |
| **Scalable** | Collect up to a million records per run |

### How it compares

No other Store actor targets OpenAIRE the same way, so the honest comparison is with the alternatives teams actually weigh.

| | OpenAIRE Publications Scraper | Build it in-house | By hand |
|---|---|---|---|
| Setup | Run it now, zero config | Days of engineering | None, but hours per pull |
| When OpenAIRE changes | Maintained for you | You fix it | You re-learn the page |
| Proxies, retries, anti-bot | Built in | Your problem | Browser only |
| Output | Fixed JSON schema, CSV/Excel export | Whatever you build | Copy-paste |
| Cost | Pay per result | Engineering time | Analyst hours |

### Configure the run

Drive the Actor with a search query and optional year range, and filters run as each publication is read so only matches reach your dataset. The Input tab lists every parameter.

A first run with the defaults:

```json
{
 "searchQuery": "machine learning",
 "maxItems": 10
}
```

A larger pull:

```json
{
 "searchQuery": "machine learning",
 "maxItems": 200
}
```

### Pricing

Pay-per-result: **$0.021 per result** collected. You pay only for the results written to your dataset.

| Results collected | Approximate cost |
|---|---|
| 100 results | $2.10 |
| 1,000 results | $21.00 |
| 10,000 results | $210.00 |

New Apify accounts start with $5 in free credit.

### Free users

Free-plan runs return up to 10 results as a preview. [Upgrade your Apify plan](https://console.apify.com/sign-up?fpr=vmoqkp) to collect up to 1,000,000 results per run.

### Run it

1. [Create a free Apify account with $5 in credit](https://console.apify.com/sign-up?fpr=vmoqkp).
2. Open the [OpenAIRE Publications Scraper](https://apify.com/parseforge/openaire-scraper?fpr=vmoqkp).
3. Set your inputs and any filters, then click **Start**.
4. Export the results as CSV, Excel, JSON, or XML from the **Dataset** tab.

Run it programmatically through the [Apify API](https://docs.apify.com/api/v2) (`run-sync-get-dataset-items`) or the [ApifyClient](https://docs.apify.com/api/client/js) for JavaScript and Python.

### Use with AI agents (MCP)

Give an AI agent live access to OpenAIRE through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

```bash
claude mcp add --transport http apify "https://mcp.apify.com/?tools=parseforge/openaire-scraper"
```

Then prompt it in plain language to run the scraper and read back the results.

### Troubleshooting

**Why am I getting no results?**

Check your search query for typos or overly specific terms. Try a broader keyword or remove the year filter to see if any records exist.

**Why are results missing some fields?**

OpenAIRE metadata varies by source. Some publications may not have a DOI or author list. The Actor returns whatever is available.

**Why did the run stop before reaching maxItems?**

OpenAIRE may not have enough matching records. The Actor collects all available results up to your limit.

**Why is the run slow?**

Large result sets require pagination through OpenAIRE's search. Reduce maxItems or narrow your query to speed things up.

**Can I search for an exact phrase?**

Yes. Put the phrase in double quotes, like "climate change adaptation".

### FAQ

| Question | Answer |
|---|---|
| What is OpenAIRE? | OpenAIRE is a European open access infrastructure that aggregates metadata for millions of research publications from repositories, journals, and other sources. |
| Do I need an API key to use this Actor? | No. The Actor reads the public OpenAIRE search interface directly, so no registration or key is required. |
| What data does each result include? | Each row includes the publication title, authors, DOI, publication date, access status, and other metadata available from OpenAIRE. |
| Can I filter by publication year? | Yes. Set the fromYear and toYear inputs to restrict results to a specific range. |
| How many results can I collect? | You can set maxItems up to 1,000,000 per run. |
| What search queries does OpenAIRE support? | You can use simple keywords, phrases in quotes, and boolean operators like AND, OR, and NOT. |
| Does this Actor return only open access publications? | OpenAIRE focuses on open access content, but some records may be metadata-only or have restricted access. The access status field tells you. |
| Can I export the results? | Yes. You can export to CSV, JSON, Excel, or XML from the Apify dataset. |
| Is this Actor free to use? | The Actor itself is free, but you need an Apify account and may use free tier credits. Large runs may require a paid plan. |
| What is the difference between this and the OpenAIRE API? | The official API requires registration and has rate limits. This Actor handles pagination and rate limiting for you and returns a clean dataset. |

### Related actors

- [google-scholar-scraper](https://apify.com/parseforge/google-scholar-scraper?fpr=vmoqkp): Use this if you need citation counts and a broader scholarly index, including paywalled articles.
- [crossref-scraper](https://apify.com/parseforge/crossref-scraper?fpr=vmoqkp): Use this if you need DOI metadata from Crossref, including funding information and references.

Browse the full [ParseForge collection](https://apify.com/parseforge?fpr=vmoqkp) for more scrapers.

🆘 **Need help?** Email parseforge@protonmail.com with your run ID, your input, and what you expected.

⚠️ **Disclaimer.** This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by OpenAIRE A.M.K.E. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.

# Actor input Schema

## `searchQuery` (type: `string`):

Keywords to search for (e.g. 'climate change', 'machine learning', 'COVID-19')

## `maxItems` (type: `integer`):

How many research products to collect per run.

## `fromYear` (type: `integer`):

Filter publications from this year (e.g. 2020)

## `toYear` (type: `integer`):

Filter publications up to this year (e.g. 2024)

## Actor input object example

```json
{
  "searchQuery": "machine learning",
  "maxItems": 10
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQuery": "machine learning",
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("parseforge/openaire-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchQuery": "machine learning",
    "maxItems": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("parseforge/openaire-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQuery": "machine learning",
  "maxItems": 10
}' |
apify call parseforge/openaire-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,parseforge/openaire-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/BTd7yaCI2jDMtSFME/builds/oN3VLJKeQXe1YniEn/openapi.json
