# GitHub Scraper - Repos, Users, Stars, Topics (`chrisp1211/github-scraper-max`) Actor

Scrape GitHub repositories by search query. Returns repo name, description, stars, forks, language, topics and last update. Optional token raises rate limits. Pay per repo; empty or failed runs cost nothing.

- **URL**: https://apify.com/chrisp1211/github-scraper-max.md
- **Developed by:** [Christian Pichichero](https://apify.com/chrisp1211) (community)
- **Categories:** Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$0.80 / 1,000 repo or user records

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## GitHub Scraper — Repos, Users, Stars, Topics

**GitHub Scraper — Repos, Users, Stars, Topics** — a fast, reliable github scraper that needs **no API key**. Scrape GitHub repositories and users — stars, forks, language, topics, license, profile data. You pay only for the results you get: failed or empty runs are always free.

This github scraper runs on the [Apify platform](https://apify.com), so you can call it from the API, run it on a schedule, or export results to JSON, CSV, Excel, or Google Sheets.

### What this scraper does

- Extracts structured **github** data with no browser or API key required
- Returns clean JSON, one record per result — ready for sheets, databases, or apps
- Pay-per-result pricing: you are never charged for a run that returns nothing
- Runs on demand or on a schedule, and integrates with 5,000+ apps via the Apify API and webhooks

### What data you get

Each result record includes fields such as:

- **Full Name** (`fullName`) — e.g. `"huggingface/transformers"`
- **Name** (`name`) — e.g. `"transformers"`
- **Owner** (`owner`) — e.g. `"huggingface"`
- **Owner Type** (`ownerType`) — e.g. `"Organization"`
- **Description** (`description`) — e.g. `"\ud83e\udd17 Transformers: the model-definition framewor...`
- **Url** (`url`) — e.g. `"https://github.com/huggingface/transformers"`
- **Homepage** (`homepage`) — e.g. `"https://huggingface.co/transformers"`
- **Stars** (`stars`) — e.g. `162448`
- **Forks** (`forks`) — e.g. `33858`
- **Watchers** (`watchers`) — e.g. `162448`
- **Open Issues** (`openIssues`) — e.g. `2477`
- **Language** (`language`) — e.g. `"Python"`
- **Topics** (`topics`) — e.g. `["audio", "deep-learning", "deepseek", "gemma", "glm", "h...`
- **License** (`license`) — e.g. `"Apache-2.0"`
- **Is Fork** (`isFork`) — e.g. `false`

### Input

| Field | Type | Description |
|---|---|---|
| `searchQueries` | array | GitHub search queries (e.g. 'web scraper language:python stars:>100'). |
| `repos` | array | Specific repos as owner/name or GitHub URLs (e.g. apify/crawlee). |
| `users` | array | GitHub usernames or org names (or profile URLs) to scrape. |
| `includeUserRepos` | boolean | Also scrape each user's public repositories. |
| `sort` | string | Sort order for repo search. |
| `maxResults` | integer | Maximum repos returned per search query (GitHub caps search at 1,000). |
| `minStars` | integer | Skip repos below this star count. Leave empty for no filter. |
| `githubToken` | string | A GitHub personal access token raises the rate limit from 60/hr to 5,000/hr. Create one at github.com/setti... |

### Example output

```json
{
  "type": "repo",
  "fullName": "huggingface/transformers",
  "name": "transformers",
  "owner": "huggingface",
  "ownerType": "Organization",
  "description": "\ud83e\udd17 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training. ",
  "url": "https://github.com/huggingface/transformers",
  "homepage": "https://huggingface.co/transformers",
  "stars": 162448,
  "forks": 33858,
  "watchers": 162448,
  "openIssues": 2477,
  "language": "Python",
  "topics": [
    "audio",
    "deep-learning",
    "deepseek",
    "gemma",
    "glm",
    "hacktoberfest",
    "llm",
    "machine-learning",
    "model-hub",
    "natural-language-processing",
    "nlp",
    "pretrained-models",
    "python",
    "pytorch",
    "pytorch-transformers",
    "qwen",
    "speech-recognition",
    "transformer",
    "vlm"
  ],
  "license": "Apache-2.0",
  "isFork": false,
  "isArchived": false,
  "defaultBranch": "main",
  "sizeKb": 515184,
  "createdAt": "2018-10-29T13:56:00Z",
  "updatedAt": "2026-07-10T15:52:03Z",
  "pushedAt": "2026-07-10T16:35:24Z",
  "scrapedAt": "2026-07-10T16:37:33Z"
}
```

### Use cases

- Monitor packages, repos, and dependencies
- Automate security and license audits
- Build developer dashboards and alerts
- Enrich internal tools and integrations

### Run it as a monitor (alerts on new github)

Schedule this github scraper to run automatically — hourly, daily, or on any cron schedule — with the built-in [Apify Scheduler](https://docs.apify.com/platform/schedules), and have results pushed to you by **webhook, email, Slack, or your own app** the moment they're ready. It's the easiest way to **monitor github for changes over time** and get alerted the instant something new appears — no manual runs. Because pricing is per result, a scheduled monitor only ever charges you for the data each run actually returns.

### Pricing

This actor uses **pay-per-result** pricing at **$0.0008 per record**. There is no monthly fee and no start fee — and **empty or failed runs cost $0**, so you only ever pay for data you actually receive.

### Frequently asked questions

**Do I need an API key or account for the source?** No. This github scraper works out of the box with no API key required.

**What happens if a run returns no results?** You are not charged. Billing is per result, so empty or failed runs are free.

**Can I run the github scraper on a schedule?** Yes. Use the Apify Scheduler to run it hourly, daily, or on any cron schedule, and get results by webhook or API.

**What export formats are supported?** Results can be exported as JSON, CSV, Excel, HTML, or pushed to Google Sheets, a database, or your own app via the Apify API.

**Is the data structured?** Yes. Every github result is a clean, flat JSON record you can use immediately.

# Actor input Schema

## `searchQueries` (type: `array`):

GitHub search queries (e.g. 'web scraper language:python stars:>100').

## `repos` (type: `array`):

Specific repos as owner/name or GitHub URLs (e.g. apify/crawlee).

## `users` (type: `array`):

GitHub usernames or org names (or profile URLs) to scrape.

## `includeUserRepos` (type: `boolean`):

Also scrape each user's public repositories.

## `sort` (type: `string`):

Sort order for repo search.

## `maxResults` (type: `integer`):

Maximum repos returned per search query (GitHub caps search at 1,000).

## `minStars` (type: `integer`):

Skip repos below this star count. Leave empty for no filter.

## `githubToken` (type: `string`):

A GitHub personal access token raises the rate limit from 60/hr to 5,000/hr. Create one at github.com/settings/tokens (no scopes needed for public data). Optional.

## `maxCostUsd` (type: `integer`):

HARD budget cap. The run stops cleanly with partial data before exceeding this.

## `proxyConfiguration` (type: `object`):

The GitHub API is open; no proxy needed. Default is fine.

## Actor input object example

```json
{
  "searchQueries": [
    "web scraper language:python"
  ],
  "repos": [],
  "users": [],
  "includeUserRepos": false,
  "sort": "stars",
  "maxResults": 50,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQueries": [
        "web scraper language:python"
    ],
    "repos": [],
    "users": [],
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("chrisp1211/github-scraper-max").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchQueries": ["web scraper language:python"],
    "repos": [],
    "users": [],
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("chrisp1211/github-scraper-max").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQueries": [
    "web scraper language:python"
  ],
  "repos": [],
  "users": [],
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call chrisp1211/github-scraper-max --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,chrisp1211/github-scraper-max"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/8CCfsvbvZUuop2ytW/builds/o9lyH2HTwXvjbd8iq/openapi.json
