# Research Papers & Scientific Preprints — arXiv (`jungle_synthesizer/arxiv-scraper`) Actor

Academic research papers and scientific preprints from arXiv.org — 2.5M+ open-access papers across physics, maths, computer science, biology and economics. Query by keyword, author, category or date range; returns title, authors, abstract, categories, publication dates and the PDF link.

- **URL**: https://apify.com/jungle\_synthesizer/arxiv-scraper.md
- **Developed by:** [BowTiedRaccoon](https://apify.com/jungle_synthesizer) (community)
- **Categories:** AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## arXiv Research Paper & Scientific Preprint Scraper

Academic research papers and scientific preprints from [arXiv.org](https://arxiv.org) — the leading open-access repository, 2.5 million+ papers across physics, mathematics, computer science, biology, economics, and quantitative finance. Query by keyword, author, category, or date range; returns title, authors, full abstract, primary and cross-listed categories, publication dates, and the PDF link.

This actor queries the **official ArXiv Atom API** (`export.arxiv.org/api/query`) — the method ArXiv officially supports for programmatic data access. No scraping, no JavaScript rendering, no account required.

### What you get

Each result includes:

- **arxiv\_id** — the canonical short ID (e.g. `2301.12345`)
- **abs\_url** — link to the abstract page
- **pdf\_url** — direct PDF download link
- **title** — full paper title
- **abstract** — complete abstract / summary
- **authors** — comma-separated author names
- **primary\_category** — primary subject category (e.g. `cs.AI`)
- **categories** — all subject categories, comma-separated
- **published** — original submission date (ISO 8601)
- **updated** — date of the latest version
- **comment** — author notes (page count, conference, etc.) if available

### Search query syntax

The `searchQuery` field supports ArXiv's full query language:

| Pattern | Example | Meaning |
|---------|---------|---------|
| Plain keyword | `machine learning` | Full-text search |
| Title | `ti:attention` | Papers with "attention" in the title |
| Author | `au:Hinton` | Papers by Hinton |
| Abstract | `abs:transformer` | Papers with "transformer" in abstract |
| Category | `cat:cs.AI` | Papers in the cs.AI category |
| Boolean | `cat:cs.LG AND ti:diffusion` | Category AND title filter |
| Date range | `submittedDate:[202301010000 TO 202312312359]` | Papers from 2023 |

See the [ArXiv query language reference](https://info.arxiv.org/help/api/user-manual.html#query_details) for the full syntax.

### Common arXiv categories

| Category | Field |
|----------|-------|
| `cs.AI` | Artificial Intelligence |
| `cs.LG` | Machine Learning |
| `cs.CL` | Computation and Language (NLP) |
| `cs.CV` | Computer Vision |
| `physics.hep-th` | High Energy Physics Theory |
| `math.CO` | Combinatorics |
| `q-bio.NC` | Neurons and Cognition |
| `econ.GN` | General Economics |

### Input parameters

| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `searchQuery` | string | **required** | ArXiv query expression |
| `maxItems` | integer | 50 | Maximum number of papers to return |
| `sortBy` | string | `submittedDate` | Sort field: `relevance`, `lastUpdatedDate`, `submittedDate` |
| `sortOrder` | string | `descending` | `ascending` or `descending` |

### Usage examples

**Fetch the 100 most recent cs.AI papers:**

```json
{
  "searchQuery": "cat:cs.AI",
  "maxItems": 100,
  "sortBy": "submittedDate",
  "sortOrder": "descending"
}
```

**Find papers by a specific author:**

```json
{
  "searchQuery": "au:LeCun",
  "maxItems": 50,
  "sortBy": "relevance"
}
```

**Search for diffusion model papers from 2024:**

```json
{
  "searchQuery": "ti:diffusion AND submittedDate:[202401010000 TO 202412312359]",
  "maxItems": 200
}
```

### Technical notes

- Uses the [ArXiv Atom API](https://info.arxiv.org/help/api/index.html) — ArXiv's official programmatic interface
- Pagination is handled automatically; set `maxItems` to any number
- Rate-limited to ~1 request/second per ArXiv usage guidelines
- No authentication required
- Results span all of arXiv's subject areas (2.5M+ papers total)

# Actor input Schema

## `sp_intended_usage` (type: `string`):

What will this data feed? E.g. lead lists, KYB checks, price tracking.

## `sp_improvement_suggestions` (type: `string`):

Provide any feedback or suggestions for improvements.

## `sp_contact` (type: `string`):

We'll personally help with your use case. No spam.

## `searchQuery` (type: `string`):

Search query in ArXiv query syntax. Examples: "machine learning", "au:Hinton", "cat:cs.AI", "ti:attention". Supports boolean operators (AND, OR, ANDNOT) and field-specific search (ti: title, au: author, abs: abstract, cat: category).

## `maxItems` (type: `integer`):

Maximum number of papers to return. Defaults to 50.

## `sortBy` (type: `string`):

Sort results by relevance, lastUpdatedDate, or submittedDate.

## `sortOrder` (type: `string`):

Sort order — ascending or descending.

## Actor input object example

```json
{
  "sp_intended_usage": "Describe your intended use...",
  "sp_improvement_suggestions": "Share your suggestions here...",
  "sp_contact": "Share your email here...",
  "searchQuery": "cat:cs.AI",
  "maxItems": 10,
  "sortBy": "submittedDate",
  "sortOrder": "descending"
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "sp_intended_usage": "Describe your intended use...",
    "sp_improvement_suggestions": "Share your suggestions here...",
    "sp_contact": "Share your email here...",
    "searchQuery": "cat:cs.AI",
    "maxItems": 10,
    "sortBy": "submittedDate",
    "sortOrder": "descending"
};

// Run the Actor and wait for it to finish
const run = await client.actor("jungle_synthesizer/arxiv-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "sp_intended_usage": "Describe your intended use...",
    "sp_improvement_suggestions": "Share your suggestions here...",
    "sp_contact": "Share your email here...",
    "searchQuery": "cat:cs.AI",
    "maxItems": 10,
    "sortBy": "submittedDate",
    "sortOrder": "descending",
}

# Run the Actor and wait for it to finish
run = client.actor("jungle_synthesizer/arxiv-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "sp_intended_usage": "Describe your intended use...",
  "sp_improvement_suggestions": "Share your suggestions here...",
  "sp_contact": "Share your email here...",
  "searchQuery": "cat:cs.AI",
  "maxItems": 10,
  "sortBy": "submittedDate",
  "sortOrder": "descending"
}' |
apify call jungle_synthesizer/arxiv-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,jungle_synthesizer/arxiv-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/3VepAHqYfsc0x4vqQ/builds/3vfBbTa0R9WQbtoxs/openapi.json
