arXiv Scraper avatar

arXiv Scraper

Pricing

from $2.50 / 1,000 results

Go to Apify Store
arXiv Scraper

arXiv Scraper

[πŸ’° $2.5 / 1K] Search arXiv and extract paper metadata β€” titles, authors, abstracts, subject categories, DOIs, journal references, submission dates, and PDF links. Search by keyword, title, author, or category, or fetch specific papers by arXiv ID.

Pricing

from $2.50 / 1,000 results

Rating

0.0

(0)

Developer

SolidCode

SolidCode

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Share

Search arXiv at scale and pull clean, structured paper metadata β€” titles, full author lists, abstracts, subject categories, DOIs, journal references, submission and revision dates, and direct PDF and abstract-page links. Search by keyword, by title, author, or abstract individually, by subject category, or fetch exact papers by arXiv ID. Built for researchers, data scientists, and librarians who need a ready-to-use arXiv dataset without manual copy-paste or wrestling with raw repository feeds one page at a time.

Why This Scraper?

  • 42 subject categories across 8 disciplines β€” pick from a labeled list spanning computer science, statistics, mathematics, physics, quantitative biology, quantitative finance, economics, and electrical engineering. Select cs.LG, stat.ML, math.PR, quant-ph and more with a checkbox β€” no codes to memorize.
  • Field-specific search, not just keywords β€” match words in the title, the author name, or the abstract as separate inputs, then combine them. Find "transformer" in the title by Vaswani in the cs.CL category in one run.
  • Direct arXiv-ID lookup, including legacy IDs β€” paste a list of IDs to fetch exact papers. Handles both modern (2310.06825) and legacy slash-style (cond-mat/0011267) identifiers, so decades-old preprints come back just as cleanly as last week's.
  • De-duplicated rows in the order you asked for β€” a 1,000-paper sweep across 10 pages returned zero duplicate arXiv IDs and zero date-ordering breaks at the page seams, so newest-first stays newest-first all the way to the last row.
  • DOI and journal reference for published-version cross-linking β€” arXiv carries these once a preprint reaches a journal, so they are field-dependent: a condensed-matter and applied-physics sweep came back 65% DOI and 44% journal reference, a quantum-computing sweep 53% and 53%. Both fields land in the row whenever arXiv has them, letting you join preprints to their peer-reviewed counterparts.
  • Direct PDF and abstract-page links on every paper β€” a pdfUrl for the full text and an absUrl for the human-readable landing page, so downstream tools can fetch or link without rebuilding URLs.
  • Sort by relevance, submission date, or last-updated date β€” newest-first or oldest-first, so you can surface the freshest preprints or build a chronological corpus.
  • Up to 50,000 papers per run β€” set the result cap to zero to sweep an entire topic, with a built-in safety ceiling so a broad query never runs away.

Use Cases

Academic Literature & Systematic Review

  • Assemble a complete reading list for a topic, sorted by relevance or recency
  • Narrow a survey to a single subject category to cut cross-field noise
  • Pull every preprint by a specific author for a focused author study
  • Track the latest submissions in a field by sorting on submission date

Research Trend & Citation Analysis

  • Measure publication volume in an emerging sub-field over time
  • Build co-author networks from the complete author list on every paper
  • Detect bursts of activity by sweeping recent submissions in a category
  • Build a chronological corpus to chart how terminology shifts year over year

Competitive R&D Intelligence

  • Monitor what a competing lab or research group is publishing on a topic
  • Benchmark a research group's output by author name across categories and years
  • Spot new directions before they reach peer-reviewed journals
  • Watch a category daily for the newest preprints in your space

ML & AI Dataset Building

  • Harvest abstracts at scale to train or fine-tune domain models
  • Build a labeled corpus by subject category for classification tasks
  • Collect title-abstract pairs for summarization and retrieval datasets
  • Gather a topic-specific text set for embeddings and semantic search

Bibliographic Database Enrichment

  • Cross-reference preprints to published versions via DOI and journal reference
  • Fill in missing abstracts, categories, and dates in an existing catalog
  • Resolve legacy slash-style IDs to current metadata
  • Enrich a reference manager export with version numbers and last-revision dates

Grant & Patent Prior-Art Search

  • Surface the earliest preprints describing a technique for prior-art review
  • Document the state of the art in a field for a grant proposal
  • Trace an idea back to its first submission date on arXiv
  • Compile a dated evidence trail across multiple subject categories

Getting Started

The simplest possible run β€” one topic, 50 papers:

{
"searchQuery": "large language models",
"maxResults": 50
}

Field-Specific Search by Category

Find recent computer-vision papers whose title mentions diffusion, newest first:

{
"title": "diffusion",
"categories": ["cs.CV", "cs.LG"],
"sortBy": "submittedDate",
"sortOrder": "descending",
"maxResults": 200
}

Fetch Specific Papers by ID

Pull exact papers β€” modern and legacy IDs together β€” ignoring all search fields:

{
"arxivIds": ["2310.06825", "1706.03762", "cond-mat/0011267"]
}

Author and Abstract Search Combined

Every author preprint mentioning reinforcement learning in the abstract:

{
"author": "Yann LeCun",
"abstract": "reinforcement learning",
"categories": ["cs.AI", "cs.LG", "stat.ML"],
"sortBy": "lastUpdatedDate",
"maxResults": 500
}

Input Reference

Combine any of these fields, or paste arXiv IDs to fetch exact papers.

ParameterTypeDefaultDescription
searchQuerystring"large language models"Free-text search across the whole paper (title, abstract, authors). Advanced users can use field prefixes like ti:, au:, abs:, cat: and boolean operators.
titlestringnullOnly include papers whose title contains these words.
authorstringnullOnly include papers by this author (e.g. "Yann LeCun" or "Hinton").
abstractstringnullOnly include papers whose abstract contains these words.
categoriesarray[]Restrict results to selected arXiv subject areas. Choose from 42 labeled categories across 8 disciplines; leave empty to search all subjects.
arxivIdsarray[]Fetch specific papers by arXiv ID (e.g. 2310.06825 or legacy cond-mat/0011267). When set, the search fields above are ignored.

Results

ParameterTypeDefaultDescription
maxResultsinteger50Maximum papers to return. Set to 0 to fetch all matches, with a safety cap of 50,000 so very broad searches don't run indefinitely. Ignored when fetching by ID.
sortByselectRelevanceOrder results by Relevance, Submission date, or Last updated date.
sortOrderselectNewest first (descending)Newest first (descending) or Oldest first (ascending). Most useful when sorting by date.

Output

Each paper is one flat row in the dataset. Here is a real result, exactly as it comes out of a run:

{
"arxivId": "2503.06391",
"version": 2,
"title": "Superconducting Coherence Peak in Near-Field Radiative Heat Transfer",
"authors": [
{ "name": "Wenbo Sun" },
{ "name": "Zhuomin M. Zhang" },
{ "name": "Zubin Jacob" }
],
"abstract": "Enhancement and peaks in near-field radiative heat transfer (NFRHT) typically arise due to surface phonon-polaritons, plasmon-polaritons, and electromagnetic (EM) modes in structured materials. However, the role of material quantum coherence in enhancing near-field radiative heat transfer remains unexplored...",
"primaryCategory": "cond-mat.mes-hall",
"categories": ["cond-mat.mes-hall", "cond-mat.supr-con", "physics.app-ph", "physics.optics"],
"publishedDate": "2025-03-09T01:54:08Z",
"updatedDate": "2025-07-02T04:12:47Z",
"doi": "10.1103/19ft-v715",
"journalRef": "Phys. Rev. B 112, 125423 (2025)",
"comments": "12 pages, 6 figures",
"pdfUrl": "https://arxiv.org/pdf/2503.06391v2",
"absUrl": "https://arxiv.org/abs/2503.06391v2"
}

This paper is a published one, so it carries a DOI and a journal reference. Preprints that have not reached a journal yet return null for both β€” the table below gives the real fill rate for every optional field.

Core Fields

FieldTypeDescription
titlestringPaper title, whitespace-normalized
authorsobject[]One record per author, in the order arXiv lists them. Always carries name; an affiliation sub-field is added on the rare papers where arXiv supplies one β€” about 1 author entry in 100
abstractstringFull abstract text
primaryCategorystringPrimary arXiv subject code (e.g. cs.CL)
categoriesstring[]All subject codes on the paper
commentsstring|nullAuthor comments (e.g. "12 pages, 6 figures") β€” filled on roughly 6 rows in 10

Identifiers & Cross-References

FieldTypeDescription
arxivIdstringarXiv identifier without version (e.g. 1706.03762)
versionintegerVersion number (v7 β†’ 7)
doistring|nullDOI once the paper has a published version registered β€” around 15% of preprints overall, and 65% on a condensed-matter and applied-physics sweep. null on preprints that have not been published
journalRefstring|nullJournal or venue citation for the published version β€” around 12% overall, rising above 50% in published-heavy fields like quantum physics. null while a paper is preprint-only
FieldTypeDescription
publishedDatestringFirst-submitted timestamp (ISO 8601)
updatedDatestringLast-updated timestamp (ISO 8601)
pdfUrlstringDirect link to the full-text PDF
absUrlstringLink to the arXiv abstract landing page

Tips for Best Results

  • Use field prefixes for precision. In searchQuery you can write ti:transformer to match only titles or cat:cs.CL to scope a subject β€” power users can build advanced boolean queries like ti:transformer AND abs:translation in a single field.
  • Narrow by category to cut noise. A broad term like "networks" spans biology, physics, and computer science. Selecting one or two subject categories sharpens results dramatically and lowers your result count.
  • Sort by submission date for the freshest preprints. Set sortBy to Submission date with Newest first to surface the very latest work in a field β€” ideal for daily monitoring and trend tracking.
  • Fetch by ID when you know exactly what you want. Pasting arXiv IDs is the fastest, most precise path β€” it skips search entirely and returns those exact papers, legacy slash-style IDs included.
  • Start small, then scale. Run with maxResults of 25–50 to confirm the data matches your needs, then raise the cap or set it to 0 to sweep a whole topic.
  • Expect DOI and journal reference to follow the field, not the query. Established physics areas publish through journals: a condensed-matter and applied-physics sweep came back 65% DOI and 44% journal reference, a quantum-physics sweep 53% and 53%. Fast-moving computer-science categories are mostly preprint-only and return single digits on both fields. Search where the published record lives if cross-linking is the goal.
  • Combine title, author, and abstract for laser-focused queries. The three field inputs are AND-joined, so a name in author plus a phrase in abstract returns only papers that satisfy both.

Pricing

From $2.50 per 1,000 results β€” a flat per-result rate that undercuts comparable arXiv extractors, with no hidden surcharges. Bronze, Silver, and Gold subscribers pay progressively less; the table below shows total cost at each discount tier.

ResultsNo discountBronzeSilverGold
100$0.30$0.28$0.265$0.25
1,000$3.00$2.80$2.65$2.50
10,000$30.00$28.00$26.50$25.00
100,000$300.00$280.00$265.00$250.00

A "result" is any paper row in the output dataset. No compute or time-based charges β€” you pay per result, plus a small fixed per-run start fee.

Integrations

Export data in JSON, CSV, Excel, XML, or RSS. Connect to 1,500+ apps via:

  • Zapier / Make / n8n β€” Workflow automation
  • Google Sheets β€” Direct spreadsheet export
  • Slack / Email β€” Notifications on new results
  • Webhooks β€” Trigger custom workflows on run completion
  • Apify API β€” Full programmatic access

arXiv content is openly accessible, and this actor is designed for legitimate academic research, literature review, bibliometrics, and dataset building. Each paper on arXiv is distributed under its own license chosen by the authors β€” respect those individual licenses when reusing abstracts or full text. Users are responsible for complying with applicable laws and arXiv's terms of use, including making reasonable-rate requests. Do not use extracted data for spam, harassment, or any illegal purpose.