# Medium Articles Scraper (`codingfrontend/medium-articles-scraper`) Actor

Scrape Medium articles by search query or topic/tag. Extracts title, author, publication, claps, responses, read time, tags, and more.

- **URL**: https://apify.com/codingfrontend/medium-articles-scraper.md
- **Developed by:** [Coding Frontned](https://apify.com/codingfrontend) (community)
- **Categories:** Developer tools, Automation, News
- **Stats:** 4 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.99 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Medium Articles RSS Scraper

Collect current article metadata and public excerpts from Medium's public tag RSS feeds. The Actor uses bounded direct HTTP requests: it does not launch a browser, configure a proxy, spoof a user agent, or generate fingerprints.

### Input

Provide `query`, `queries`, or both. Each value is normalized to one exact tag slug; the Actor does not silently expand it into loosely related tags.

```json
{
  "query": "machine learning",
  "mode": "tag",
  "maxItems": 3
}
```

- `query`: one topic or tag, up to 80 characters.
- `queries`: up to 10 additional non-empty topics or tags. Duplicates are removed case-insensitively.
- `mode`: `tag` or `search`. This is provenance metadata; both values use the normalized tag feed because Medium has no public RSS search endpoint.
- `maxItems`: maximum unique rows across all feeds, from 1 to 100. A feed can expose fewer items than requested.

Unknown fields and invalid bounds fail before any request is made.

### Dataset

Each row is one valid public RSS item. Core fields include:

```json
{
  "recordType": "mediumArticle",
  "recordId": "7e5902a33bbf84ca24a1e23b",
  "articleTitle": "An example article",
  "articleUrl": "https://medium.com/@example/an-example-article-12345678",
  "authorName": "Example author",
  "summary": "A public RSS excerpt.",
  "tags": ["Machine Learning"],
  "publishedAt": "2026-08-30T09:00:00.000Z",
  "feedUrl": "https://medium.com/feed/tag/machine-learning",
  "feedTag": "machine-learning",
  "sourceQuery": "machine learning",
  "scrapeMode": "tag",
  "sourceHttpStatus": 200,
  "extractionMethod": "medium-rss-feed",
  "proxyConfigured": false,
  "scrapedAt": "2026-08-30T10:00:00.000Z"
}
```

`recordId` is a deterministic 24-character SHA-256 prefix derived from the cleaned article URL. Optional fields are omitted rather than written as `null`. Scriptable or embedded elements are removed from excerpt HTML, tracking parameters are stripped from known URLs, and local, private, or credentialed URLs are rejected. Medium RSS generally exposes an excerpt, not the full paywalled article body.

### Run summary and failure behavior

`OUTPUT` contains status, counts, timing, direct-request provenance, and bounded per-feed diagnostics. Diagnostics are never mixed into the article dataset.

The Actor succeeds when at least one valid row is saved. If every feed fails or yields no valid item, it saves zero dataset rows, writes a diagnostic `OUTPUT`, and fails explicitly. A partially successful run completes with `SUCCEEDED_WITH_WARNINGS` in `OUTPUT`.

### Limits and responsible use

Medium may change or disable feeds, return fewer items, or omit optional author/publication fields. There is no external API fee, though Apify compute charges apply. Use the data only where you have a lawful basis, and respect Medium's terms, RSS policies, copyright, privacy rules, and applicable law.

### Development

```text
npm ci
npm test
npm run lint
npm run validate:schemas
npm audit --omit=dev
```

# Actor input Schema

## `query` (type: `string`):

A topic phrase normalized to one exact Medium tag feed.

## `queries` (type: `array`):

Up to 10 additional topics. Duplicates are removed case-insensitively.

## `mode` (type: `string`):

Stored as provenance. Both values use the exact normalized tag feed; Medium has no public RSS search endpoint.

## `maxItems` (type: `integer`):

Maximum unique article records saved across all requested feeds. A feed can expose fewer items.

## Actor input object example

```json
{
  "query": "artificial intelligence",
  "mode": "tag",
  "maxItems": 50
}
```

# Actor output Schema

## `dataset` (type: `string`):

Public Medium RSS article records.

## `output` (type: `string`):

Run status, counts, request mode, timing, and feed diagnostics.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "query": "artificial intelligence"
};

// Run the Actor and wait for it to finish
const run = await client.actor("codingfrontend/medium-articles-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "query": "artificial intelligence" }

# Run the Actor and wait for it to finish
run = client.actor("codingfrontend/medium-articles-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "query": "artificial intelligence"
}' |
apify call codingfrontend/medium-articles-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,codingfrontend/medium-articles-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/kv8aWRprl1QeMTDDx/builds/POODTeb6mrOgzky7I/openapi.json
