# Crossref Scraper — 150M+ Papers, DOIs & Citations (`dami_studio/crossref-scraper`) Actor

Search Crossref for journal articles, preprints, books and datasets. Crossref registers DOIs, so its index holds 150M+ works. Each row has the DOI, title, authors, journal and publisher. You also get date, citation count, subjects, ISSN and abstract. No API key. $1.00 per 1,000 works.

- **URL**: https://apify.com/dami\_studio/crossref-scraper.md
- **Developed by:** [Dami's Studio](https://apify.com/dami_studio) (community)
- **Categories:** Developer tools, AI, Other
- **Stats:** 1 total users, 0 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: 5.00 out of 5 stars

## Pricing

$1.00 / 1,000 work returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Crossref Scholarly Works Scraper

Search the [Crossref](https://www.crossref.org/) catalog of 150M+ scholarly works (journal articles, preprints, books, datasets, and more) via its public REST API — no API key, no login, no anti-bot.

The actor is a **polite Crossref client**: it identifies itself with a contact `User-Agent` and a `mailto` query parameter so Crossref routes it to the faster "polite pool", and it uses **deep cursor pagination** (`cursor=*` → `next-cursor`) which is the only reliable way to page past 1,000 rows.

### Input

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `query` | string (required) | `deep learning` | Keywords searched across titles, authors, abstracts and metadata. |
| `filterType` | string | *all* | Restrict to a Crossref work type, e.g. `journal-article`. |
| `fromDate` | string `YYYY-MM-DD` | *none* | Only works published on/after this date. |
| `sort` | enum | `relevance` | `relevance`, `is-referenced-by-count` (most cited), or `published` (newest). |
| `maxItems` | integer | `100` | Max works to return (cursor pagination handles >100). |
| `proxyConfiguration` | object | *none* | Optional and off by default; Crossref is a public, no-key API with no anti-bot, so a proxy adds no benefit. Only enable it if you hit IP-level rate limits. |

### Output

Each successful row:

```json
{
  "ok": true,
  "doi": "10.1038/nature14539",
  "title": "Deep learning",
  "authors": ["Yann LeCun", "Yoshua Bengio", "Geoffrey Hinton"],
  "journal": "Nature",
  "publisher": "Springer Science and Business Media LLC",
  "type": "journal-article",
  "publishedDate": "2015-05-28",
  "citations": 70000,
  "subjects": ["Multidisciplinary"],
  "issn": ["0028-0836", "1476-4687"],
  "abstract": null,
  "url": "https://doi.org/10.1038/nature14539"
}
```

- `authors` are formatted `"Given Family"` (organizational authors fall back to their name).
- `publishedDate` is assembled from Crossref's `date-parts` (may be year-only or year-month for older records).
- `citations` is Crossref's `is-referenced-by-count`.
- `abstract` is the JATS-XML abstract stripped to plain text, or `null` when Crossref has none.
- **Nullable fields:** `title`, `journal`, `publisher`, `type`, `publishedDate`, `abstract`, and `url` may be `null`, and `authors`, `subjects`, and `issn` may be empty arrays, depending on what the publisher deposited with Crossref. `doi` is always present (rows without a DOI are dropped). `citations` defaults to `0` when absent.

Results are **deduplicated by DOI**.

### Pricing

**$1.00 per 1,000 works** ($0.001 each). Flat rate — no volume tiers, no plan gates — and you are charged only for works actually returned.

Charging is **per successful work** (`work` event). Diagnostic / empty / blocked rows (`ok: false` with an `errorCode`) are **never charged** — this includes `BAD_INPUT` (empty query or malformed `fromDate`), `NO_RESULTS`, and any network/block error. A run that finds nothing costs nothing.

### Troubleshooting

- **`BAD_INPUT` row, no results:** you left `query` empty or `fromDate` isn't `YYYY-MM-DD`. Fix the input and re-run — you were not charged.
- **`NO_RESULTS` row:** your query/filter combination matched nothing in Crossref. Try broader keywords or drop the type/date filters.
- **`RATE_LIMITED` / `BLOCKED` row:** rare for Crossref. The actor already retries with backoff; if it persists, enable a proxy to use a different IP.

### Notes

- Powered entirely by the public Crossref REST API (`https://api.crossref.org/works`). Please be considerate of the shared, free service.
- Citation counts and abstracts depend on what publishers deposit with Crossref; coverage varies by record.

# Actor input Schema

## `query` (type: `string`):

Keywords to search Crossref for across titles, authors, abstracts, and metadata (e.g. "deep learning", "CRISPR gene editing", "climate change adaptation").

## `fromDate` (type: `string`):

Only return works published on or after this date, in YYYY-MM-DD format (e.g. 2020-01-01). Leave empty for no date floor.

## `filterType` (type: `string`):

Only return works of this Crossref type. Leave empty for all types. "journal-article" is the most common for research papers.

## `sort` (type: `string`):

How to order results. "Relevance" matches the query best; "Most cited" surfaces influential papers; "Newest first" sorts by publication date descending.

## `maxItems` (type: `integer`):

Maximum number of scholarly works to return. Uses deep cursor pagination to fetch beyond 100 reliably.

## `notionConnector` (type: `string`):

Optional. Write each result as a page into your Notion when the run finishes. Authorize a Notion connector once in Settings → API & Integrations → MCP connectors, then pick it here. Leave empty to skip (default) — results are always saved to the dataset regardless.

## `notionParentId` (type: `string`):

Optional. The Notion data source ID of the database to write into (only used if a Notion connector is set). Leave empty to create the pages privately in your workspace instead.

## `proxyConfiguration` (type: `object`):

Optional proxy for reaching the Crossref API. Not needed: Crossref is a public, no-key API with no anti-bot, so a proxy adds no benefit and costs credits. Only enable it if you hit IP-level rate limits. Defaults to no proxy.

## Actor input object example

```json
{
  "query": "deep learning",
  "fromDate": "",
  "filterType": "",
  "sort": "relevance",
  "maxItems": 100,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

Scraped rows are stored in the default dataset (one row per result). Blocked/empty/error runs return a single uncharged diagnostic row instead.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "query": "deep learning"
};

// Run the Actor and wait for it to finish
const run = await client.actor("dami_studio/crossref-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "query": "deep learning" }

# Run the Actor and wait for it to finish
run = client.actor("dami_studio/crossref-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "query": "deep learning"
}' |
apify call dami_studio/crossref-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,dami_studio/crossref-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/vrPcFIdyyvywJyGDt/builds/c7gnSttBikdKR0y2I/openapi.json
