# Google Scholar Scraper (`parseforge/google-scholar-scraper`) Actor

Scrapes Google Scholar search results for a given query and returns each paper as a flat row with title, authors, publication venue, year, citations, and URL.

- **URL**: https://apify.com/parseforge/google-scholar-scraper.md
- **Developed by:** [ParseForge](https://apify.com/parseforge) (community)
- **Categories:** Education, AI
- **Stats:** 1 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.06 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

[![ParseForge](https://raw.githubusercontent.com/ParseForge/apify-assets/main/banner.jpg)](https://apify.com/parseforge?fpr=vmoqkp)

### Google Scholar Scraper

**Scrape Google Scholar search results for any query, filtered by year and language.** Each result includes title, authors, publication venue, year, citations, and URL. No API key or login required. Export to CSV, JSON, Excel, or XML.

Google Scholar has no official API, and scraping it yourself means handling CAPTCHAs, rate limits, and brittle HTML. This Actor reads the public search results directly, applies your year and language filters, and returns each paper as a clean, flat row. It is built for researchers, students, and anyone who needs scholarly data at scale.

| Who uses it | What they scrape Google Scholar for |
|---|---|
| Academic researchers | Building a literature review dataset for a specific topic |
| PhD students | Tracking new papers in their field every week |
| Data scientists | Collecting citation data for bibliometric analysis |
| Librarians | Compiling publication lists for faculty or departments |
| Market analysts | Monitoring research output of companies or institutions |

### What it does

This Actor collects Google Scholar search results for a given query and returns each paper as a flat row with title, authors, publication venue, year, citations, and URL.

- 🔍 **Search query:** any phrase you would type into Google Scholar, from 'machine learning' to 'site:nature.com'.
- 📅 **Year range filter:** set yearFrom and yearTo to limit results to a publication window.
- 🌐 **Language filter:** choose from English, French, German, Spanish, Chinese, Japanese, Portuguese, or Russian.
- 🔢 **Result cap:** set maxItems to control how many papers to scrape, up to the limit of Google Scholar pagination.
- 📄 **Flat output:** each paper becomes one row with title, authors, venue, year, citations, and URL.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

### What you can do with Google Scholar data

**📚 Build a literature review dataset.**

A researcher enters a query like 'deep learning in medicine' with yearFrom 2020 and yearTo 2025, then exports the results to CSV for screening in Excel.

**📈 Track new publications weekly.**

A PhD student schedules the Actor to run every Monday for their thesis topic, collecting the latest papers with citations and URLs.

**🔬 Analyze citation patterns.**

A data scientist scrapes thousands of papers for a bibliometric study, using the citation counts and publication years to map research trends.

**🏛️ Compile faculty publication lists.**

A librarian runs the Actor for each professor's name and affiliation, gathering their papers into a single spreadsheet for an annual report.

### Why choose this scraper

|  | What you get |
|---|---|
| **No official API** | Google Scholar does not offer a public API, so this Actor is the easiest way to get structured data. |
| **Clean, flat schema** | Every paper is returned as a single row with consistent fields, ready for analysis. |
| **Year and language filters** | Limit results to a specific publication window or language without post-processing. |
| **No login or API key** | The Actor handles the scraping, so you do not need to authenticate or manage cookies. |

### How it compares

This Actor focuses on Google Scholar search results, while the competitors below cover broader Google search types or images.

| Feature | ParseForge | Google Images Scraper | Google Search Scraper |
|---|---|---|---|
| Scrapes Google Scholar results | Yes | Not listed | Yes |
| Year range filter | Yes | Not listed | Not listed |
| Language filter | Yes | Not listed | Not listed |
| Returns citation counts | Yes | Not listed | Not listed |
| Returns paper authors and venue | Yes | Not listed | Not listed |
| Exports to CSV, JSON, Excel, XML | Yes | Yes | Yes |

### Configure the run

Drive the Actor with a search query, then narrow results by publication year and language. The maxItems field caps how many papers are returned. The Input tab lists every parameter.

A first run with the defaults:

```json
{
  "searchQuery": "example",
  "language": "en",
  "yearFrom": 2020,
  "yearTo": 2025,
  "maxItems": 10
}
```

A larger pull:

```json
{
  "searchQuery": "example",
  "language": "en",
  "yearFrom": 2020,
  "yearTo": 2025,
  "maxItems": 200
}
```

### Pricing

Pay-per-result: **$0.00449 per result** collected. You pay only for the results written to your dataset.

| Results collected | Approximate cost |
|---|---|
| 100 results | $0.45 |
| 1,000 results | $4.49 |
| 10,000 results | $44.90 |

New Apify accounts start with $5 in free credit.

### Free users

Free-plan runs return up to 10 results as a preview. [Upgrade your Apify plan](https://console.apify.com/sign-up?fpr=vmoqkp) to collect up to 1,000,000 results per run.

### Run it

1. [Create a free Apify account with $5 in credit](https://console.apify.com/sign-up?fpr=vmoqkp).
2. Open the [Google Scholar Scraper](https://apify.com/parseforge/google-scholar-scraper?fpr=vmoqkp).
3. Set your inputs and any filters, then click **Start**.
4. Export the results as CSV, Excel, JSON, or XML from the **Dataset** tab.

Run it programmatically through the [Apify API](https://docs.apify.com/api/v2) (`run-sync-get-dataset-items`) or the [ApifyClient](https://docs.apify.com/api/client/js) for JavaScript and Python.

### Use with AI agents (MCP)

Give an AI agent live access to Google Scholar through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

```bash
claude mcp add --transport http apify "https://mcp.apify.com/?tools=parseforge/google-scholar-scraper"
```

Then prompt it in plain language to run the scraper and read back the results.

### Troubleshooting

**Why am I getting no results?**

Your query may be too specific, or the year range may exclude all papers. Try a broader query, widen the year range, or remove the language filter.

**Why did the Actor stop before reaching maxItems?**

Google Scholar may have fewer results than your maxItems, or it may have blocked further pagination. Try a different query or reduce the number of results.

**Why are some fields empty?**

Not all papers have complete metadata on Google Scholar. For example, some may lack a publication venue or citation count. This is normal.

**Why am I getting a CAPTCHA or block?**

Google may temporarily block automated requests. Wait a few minutes and try again, or reduce the frequency of your runs.

**Can I search for a specific author?**

Yes, use the author's name in the search query, for example 'author:John Smith'. You can also combine with topic keywords.

### FAQ

| Question | Answer |
|---|---|
| Does Google Scholar have an official API? | No, Google Scholar does not provide a public API. This Actor scrapes the public search results directly, so you do not need an API key. |
| Can I filter results by publication year? | Yes, use the yearFrom and yearTo input fields to set a range. For example, yearFrom 2020 and yearTo 2025 returns only papers published in those years. |
| What languages are supported? | You can filter results by language using the language field. Options include English, French, German, Spanish, Chinese, Japanese, Portuguese, and Russian. |
| How many results can I scrape? | Set maxItems to the number of papers you want. Google Scholar typically shows up to 1000 results per query, but the Actor will stop at your cap. |
| What data do I get for each paper? | Each row includes the paper title, authors, publication venue, year, citation count, and URL. The exact fields are shown in the sample output. |
| Do I need to log in to Google? | No, the Actor does not require a Google account or login. It accesses the public search interface. |
| Can I export the data? | Yes, you can export the results to CSV, JSON, Excel, or XML directly from the Apify platform. |
| Is this legal? | Scraping public data from Google Scholar is generally allowed for personal research, but you should review Google's terms of service and respect robots.txt for your use case. |
| Can I schedule this Actor to run automatically? | Yes, you can set up a schedule in Apify to run the Actor daily, weekly, or at any interval you need. |
| What if I get no results? | Check your search query for typos, broaden your year range, or remove the language filter. Google Scholar may also block requests if you run too many in a short time. |

### Related actors

- [google-search-scraper](https://apify.com/parseforge/google-search-scraper?fpr=vmoqkp): Use this if you need general Google search results, including Scholar, Images, News, and more, in one Actor.

Browse the full [ParseForge collection](https://apify.com/parseforge?fpr=vmoqkp) for more scrapers.

🆘 **Need help?** Email parseforge@protonmail.com with your run ID, your input, and what you expected.

⚠️ **Disclaimer.** This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by Google LLC. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.

# Actor input Schema

## `searchQuery` (type: `string`):

Enter search terms for Google Scholar (e.g., machine learning).

## `language` (type: `string`):

Language of the results (e.g., en for English).

## `yearFrom` (type: `integer`):

Start year for publication date filter (e.g., 2020).

## `yearTo` (type: `integer`):

End year for publication date filter (e.g., 2025).

## `maxItems` (type: `integer`):

Maximum number of search results to scrape.

## Actor input object example

```json
{
  "searchQuery": "example",
  "language": "en",
  "yearFrom": 2020,
  "yearTo": 2025,
  "maxItems": 10
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQuery": "example",
    "language": "en",
    "yearFrom": 2020,
    "yearTo": 2025,
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("parseforge/google-scholar-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchQuery": "example",
    "language": "en",
    "yearFrom": 2020,
    "yearTo": 2025,
    "maxItems": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("parseforge/google-scholar-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQuery": "example",
  "language": "en",
  "yearFrom": 2020,
  "yearTo": 2025,
  "maxItems": 10
}' |
apify call parseforge/google-scholar-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,parseforge/google-scholar-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/9XhppdjvtXIulV2I2/builds/I7d9kOlL5Lagb4Q7c/openapi.json
