# Hugging Face Papers Scraper (`parseforge/huggingface-papers-scraper`) Actor

Scrapes Hugging Face papers by search query or trending feed. Returns each paper as a flat row with title, authors, abstract, upvotes, and URL. Export to CSV, JSON, Excel, or XML.

- **URL**: https://apify.com/parseforge/huggingface-papers-scraper.md
- **Developed by:** [ParseForge](https://apify.com/parseforge) (community)
- **Categories:** Education, Developer tools, Other
- **Stats:** 8 total users, 1 monthly users, 76.5% runs succeeded, 1 bookmarks
- **User rating**: No ratings yet

## Pricing

from $9.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

[![ParseForge](https://raw.githubusercontent.com/ParseForge/apify-assets/main/banner.jpg)](https://apify.com/parseforge?fpr=vmoqkp)

### Hugging Face Papers Scraper

**Scrape Hugging Face papers by search query or trending feed, up to a million per run.** Each paper comes with its title, authors, abstract, upvotes, and direct URL. No API key or login. Export to CSV, JSON, Excel, or XML.

Hugging Face's official API needs a token and rate-limits you. This reads the public papers feed directly, filtered by keyword or trending, and returns each match in one fixed schema. It is the fastest way to track artificial intelligence news and research papers without writing code.

| Who uses it | What they scrape Hugging Face for |
|---|---|
| AI researchers | Monitor new papers in their subfield without manually checking the site |
| Data scientists | Build datasets of paper metadata for trend analysis or literature reviews |
| Product managers | Track competitor research and emerging techniques in AI |
| Content creators | Find trending papers to write about or share with their audience |

### What it does

This Actor collects Hugging Face papers by search query or trending feed, and returns each one as a flat row.

- 🔍 **Search mode:** enter any keyword like 'transformer', 'diffusion model', or 'LLM' to get matching papers.
- 📈 **Trending mode:** fetch the current trending papers on Hugging Face with one click.
- 📊 **Flat output:** every paper is a single row with title, authors, abstract, upvotes, and URL, ready for spreadsheets.
- ⚙️ **Flexible limits:** set maxItems from 1 to 1,000,000 to control how many papers you collect per run.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

### What you can do with Hugging Face data

**📚 Build a research database.**

A PhD student runs the Actor weekly with the query 'diffusion model' and exports the results to CSV to maintain a personal library of relevant papers.

**📈 Track AI trends.**

A market analyst uses trending mode daily to see which papers are gaining traction, then shares a summary with their team to inform product strategy.

**🔎 Monitor competitor research.**

A startup founder sets up a scheduled run for keywords like 'LLM' and 'RAG' to stay aware of new techniques from competitors and academic labs.

**📰 Curate AI news.**

A newsletter writer scrapes trending papers each morning and picks the most interesting ones to feature in their daily AI digest.

### Why choose this scraper

|  | What you get |
|---|---|
| **No API key** | Scrape public paper feeds without registering an app or dealing with OAuth. |
| **Up-to-date** | Get the latest papers as they appear on Hugging Face, including trending items. |
| **Structured data** | Each paper is a flat object with title, authors, abstract, upvotes, and URL. |
| **Scalable** | Collect up to a million papers per run, suitable for large-scale analysis. |

### How it compares

This Actor is the only one in this group that scrapes Hugging Face papers. The others are OCR tools for receipts and PDFs, so they are not direct competitors for this use case.

| Feature | ParseForge | Receipt OCR API | Pdf OCR API | Artificial Intelligence News |
|---|---|---|---|---|
| Scrapes Hugging Face papers | Yes | Not listed | Not listed | Not listed |
| Search by keyword | Yes | Not listed | Not listed | Not listed |
| Trending papers feed | Yes | Not listed | Not listed | Not listed |
| Returns paper metadata (title, authors, abstract) | Yes | Not listed | Not listed | Not listed |
| No API key required | Yes | Not listed | Not listed | Not listed |

### Configure the run

Drive the Actor with a search query or switch to trending mode, and set the maximum number of papers to collect. Results are returned in a consistent schema for easy export. The Input tab lists every parameter.

A first run with the defaults:

```json
{
  "maxItems": 10,
  "searchQuery": "transformer"
}
```

A larger pull:

```json
{
  "maxItems": 200,
  "searchQuery": "transformer"
}
```

### Pricing

Pay-per-result: **$0.0095 per result** collected. You pay only for the results written to your dataset.

| Results collected | Approximate cost |
|---|---|
| 100 results | $0.95 |
| 1,000 results | $9.50 |
| 10,000 results | $95.00 |

New Apify accounts start with $5 in free credit.

### Free users

Free-plan runs return up to 10 results as a preview. [Upgrade your Apify plan](https://console.apify.com/sign-up?fpr=vmoqkp) to collect up to 1,000,000 results per run.

### Run it

1. [Create a free Apify account with $5 in credit](https://console.apify.com/sign-up?fpr=vmoqkp).
2. Open the [Hugging Face Papers Scraper](https://apify.com/parseforge/huggingface-papers-scraper?fpr=vmoqkp).
3. Set your inputs and any filters, then click **Start**.
4. Export the results as CSV, Excel, JSON, or XML from the **Dataset** tab.

Run it programmatically through the [Apify API](https://docs.apify.com/api/v2) (`run-sync-get-dataset-items`) or the [ApifyClient](https://docs.apify.com/api/client/js) for JavaScript and Python.

### Use with AI agents (MCP)

Give an AI agent live access to Hugging Face through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

```bash
claude mcp add --transport http apify "https://mcp.apify.com/?tools=parseforge/huggingface-papers-scraper"
```

Then prompt it in plain language to run the scraper and read back the results.

### Troubleshooting

**Why am I getting no results?**

Make sure your searchQuery is spelled correctly and not too specific. Try a broader term like 'AI' or 'model'. Also check that mode is set to 'search' if you are using a query.

**The Actor returns fewer papers than maxItems.**

This can happen if there are not enough matching papers on Hugging Face for your query or trending feed. Try a different query or increase the time range if available.

**I get an error about rate limiting.**

The Actor includes automatic retries and delays to avoid rate limits. If you still see errors, reduce the frequency of your runs or contact support.

**Some fields are empty in the output.**

Not all papers have every field populated. For example, some may lack an abstract or upvotes. This is normal and reflects the source data.

### FAQ

| Question | Answer |
|---|---|
| Do I need a Hugging Face API key? | No. This Actor reads the public papers feed directly, so no authentication or token is required. |
| What data does each paper include? | Each paper is returned as a flat object with fields like title, authors, abstract, upvotes, and URL. The exact fields are shown in the sample output. |
| Can I scrape trending papers? | Yes. Set the mode to 'trending' and the Actor will fetch the current trending papers on Hugging Face. |
| How many papers can I collect in one run? | You can set maxItems from 1 to 1,000,000. The Actor will stop after reaching that number. |
| Can I search by multiple keywords? | The searchQuery field accepts a single string. For multiple keywords, run the Actor multiple times or use a more complex query if supported by the site. |
| What export formats are supported? | You can export the results to CSV, JSON, Excel, or XML from the Apify dataset. |
| Is this Actor legal to use? | Yes, it only accesses publicly available data. However, you should respect Hugging Face's terms of service and robots.txt. |
| Can I schedule this Actor to run automatically? | Yes, you can set up a schedule in Apify to run it daily, weekly, or at any interval. |
| What if I get no results? | Check your search query for typos or try a broader term. Also ensure the mode is set correctly. |
| Does this Actor handle pagination? | Yes, it automatically paginates through results until it reaches maxItems or no more papers are available. |

### Related actors

Browse the full [ParseForge collection](https://apify.com/parseforge?fpr=vmoqkp) for more scrapers.

🆘 **Need help?** Email parseforge@protonmail.com with your run ID, your input, and what you expected.

⚠️ **Disclaimer.** This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by Hugging Face, Inc. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.

# Actor input Schema

## `maxItems` (type: `integer`):

How many papers to collect per run.

## `searchQuery` (type: `string`):

Search papers by keyword. Example: 'transformer', 'diffusion model', 'LLM'.

## `mode` (type: `string`):

Search or trending.

## Actor input object example

```json
{
  "maxItems": 10,
  "searchQuery": "transformer",
  "mode": "search"
}
```

# Actor output Schema

## `overview` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "maxItems": 10,
    "searchQuery": "transformer"
};

// Run the Actor and wait for it to finish
const run = await client.actor("parseforge/huggingface-papers-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "maxItems": 10,
    "searchQuery": "transformer",
}

# Run the Actor and wait for it to finish
run = client.actor("parseforge/huggingface-papers-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "maxItems": 10,
  "searchQuery": "transformer"
}' |
apify call parseforge/huggingface-papers-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,parseforge/huggingface-papers-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Bo0WbEAfhiN5kcWkH/builds/yuKjFlwnI9vGpsUmm/openapi.json
