# Y Combinator Companies Scraper (`parseforge/y-combinator-scraper`) Actor

Scrapes Y Combinator company profiles from the public directory. Returns each company as a flat row with optional founder details and open job listings, filterable by batch, industry, region, and status.

- **URL**: https://apify.com/parseforge/y-combinator-scraper.md
- **Developed by:** [ParseForge](https://apify.com/parseforge) (community)
- **Categories:** Social media, Lead generation, Other
- **Stats:** 73 total users, 6 monthly users, 100.0% runs succeeded, 1 bookmarks
- **User rating**: No ratings yet

## Pricing

from $11.99 / 1,000 result items

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

[![ParseForge](https://raw.githubusercontent.com/ParseForge/apify-assets/main/banner.jpg)](https://apify.com/parseforge?fpr=vmoqkp)

### Y Combinator Companies Scraper

**Scrape Y Combinator companies by batch, industry, region, or keyword, up to a million per run.** Every company comes with its description, founders, job listings, and status. No API key required. Export to CSV, JSON, Excel, or XML.

Y Combinator's startup directory is a goldmine of market intelligence, but manually browsing hundreds of company profiles is slow and unscalable. This Actor reads the public YC company pages directly, letting you filter by batch, industry, region, hiring status, and more. It returns each matched company in one consistent, flat schema ready for analysis.

| Who uses it | What they scrape Y Combinator for |
|---|---|
| Venture capital analysts | Screen new YC batches for investment targets by industry and region. |
| Market researchers | Map startup activity and emerging trends across Y Combinator's portfolio. |
| Recruiters | Find YC-backed companies that are actively hiring for specific roles. |
| Sales teams | Build lead lists of funded startups by technology vertical and growth stage. |

### What it does

This Actor collects Y Combinator company profiles from the public directory and returns each one as a flat row with optional founder details and job listings.

- 🎓 **Batch filtering:** target specific cohorts like W25, S24, or X25 using short codes or full names.
- 🏭 **Industry and subindustry filters:** narrow results to B2B, Fintech, Healthcare, and their specific sub-sectors.
- 🌍 **Region targeting:** focus on companies headquartered in the United States, Europe, Latin America, India, and more.
- 👤 **Founder profiles:** include each founder's name, title, bio, LinkedIn, and Twitter handle.
- 💼 **Job listings:** pull open roles with salary range, equity, visa sponsorship, and required skills.
- 📄 **Full job descriptions:** optionally fetch the complete text of every job posting.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

### What you can do with Y Combinator data

**📈 Monitor startup formation by sector.**

A VC analyst runs the scraper weekly on the latest batch, filtered by Fintech and United States, to spot new investment opportunities before they hit the news.

**🎯 Build a lead list of funded startups.**

A sales team scrapes all active B2B companies from the last three batches that are hiring, then exports the list to their CRM for outreach.

**💼 Source candidates from high-growth companies.**

A recruiter pulls every YC company with open engineering roles in Europe, including full job descriptions, to match candidates with funded startups.

**🔬 Analyze Y Combinator's portfolio composition.**

A researcher scrapes the entire directory with industry and region filters to quantify how YC's investment thesis has shifted across batches.

### Why choose this scraper

| | What you get |
|---|---|
| **Batch and status filters** | Screen only the most recent cohorts or filter for active, public, and acquired companies. |
| **Founder and job data** | Enrich company profiles with founder backgrounds and detailed open positions. |
| **Flexible output** | Download results as CSV, JSON, Excel, or XML for your existing workflows. |

### How it compares

No other Store actor targets Y Combinator the same way, so the honest comparison is with the alternatives teams actually weigh.

| | Y Combinator Companies Scraper | Build it in-house | By hand |
|---|---|---|---|
| Setup | Run it now, zero config | Days of engineering | None, but hours per pull |
| When Y Combinator changes | Maintained for you | You fix it | You re-learn the page |
| Proxies, retries, anti-bot | Built in | Your problem | Browser only |
| Output | Fixed JSON schema, CSV/Excel export | Whatever you build | Copy-paste |
| Cost | Pay per result | Engineering time | Analyst hours |

### Configure the run

Drive the Actor from a keyword search or a list of direct company URLs, and filters run as each company is read so only matches reach your dataset. The Input tab lists every parameter.

A first run with the defaults:

```json
{
 "maxItems": 10
}
```

A larger pull:

```json
{
 "maxItems": 200
}
```

### Pricing

Pay-per-result: **$0.01599 per result** collected. You pay only for the results written to your dataset.

| Results collected | Approximate cost |
|---|---|
| 100 results | $1.60 |
| 1,000 results | $15.99 |
| 10,000 results | $159.90 |

New Apify accounts start with $5 in free credit.

### Free users

Free-plan runs return up to 10 results as a preview. [Upgrade your Apify plan](https://console.apify.com/sign-up?fpr=vmoqkp) to collect up to 1,000,000 results per run.

### Run it

1. [Create a free Apify account with $5 in credit](https://console.apify.com/sign-up?fpr=vmoqkp).
2. Open the [Y Combinator Companies Scraper](https://apify.com/parseforge/y-combinator-scraper?fpr=vmoqkp).
3. Set your inputs and any filters, then click **Start**.
4. Export the results as CSV, Excel, JSON, or XML from the **Dataset** tab.

Run it programmatically through the [Apify API](https://docs.apify.com/api/v2) (`run-sync-get-dataset-items`) or the [ApifyClient](https://docs.apify.com/api/client/js) for JavaScript and Python.

### Use with AI agents (MCP)

Give an AI agent live access to Y Combinator through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

```bash
claude mcp add --transport http apify "https://mcp.apify.com/?tools=parseforge/y-combinator-scraper"
```

Then prompt it in plain language to run the scraper and read back the results.

### Troubleshooting

**Why am I getting no results?**

Check that your filters are not too restrictive. Try removing some filters or broadening your batch, industry, or region selections. Also verify that your keyword search term is spelled correctly.

**Why are job listings missing from my output?**

Ensure the Include Open Jobs option is enabled. If it is on and jobs are still missing, the company may not have any open positions listed on Y Combinator at this time.

**The scraper is running slowly. What can I do?**

Fetching full job descriptions adds one request per job listing. If speed is a concern, disable Include Job Descriptions and run the scraper with only the core company data and job summaries.

**I pasted a company URL but the scraper ignored my filters.**

When you provide direct URLs in the startUrls field, all filter settings are bypassed. The Actor scrapes exactly those URLs and nothing else. Remove the URLs if you want to use filters.

**Some founder social media links are missing.**

The scraper includes whatever LinkedIn and Twitter URLs are publicly listed on the YC profile. If a founder has not added them, those fields will be empty in your results.

### FAQ

| Question | Answer |
|---|---|
| Can I scrape a specific company by URL? | Yes. Paste one or more direct Y Combinator company URLs into the startUrls field. When you provide URLs, all other filters are ignored and only those companies are scraped. |
| How do I filter by a specific YC batch? | Use the Batches field with short codes like W25, S25, or F25, or full names like Winter 2025. You can enter multiple batches to cover several cohorts in one run. |
| Does this scraper require a Y Combinator account or API key? | No. It reads the public Y Combinator company directory, so no login, account, or API key is needed. |
| Can I get founder LinkedIn and Twitter profiles? | Yes. Enable the Include Founders option and each company row will contain founder names, titles, bios, and links to their LinkedIn and Twitter profiles when available. |
| How do I scrape job listings with salary information? | Turn on Include Open Jobs. The output will include each job's title, salary range, equity range, visa sponsorship status, and required skills. Enable Include Job Descriptions to also get the full posting text. |
| What is the maximum number of companies I can scrape? | You can set the maximum up to 1,000,000 companies per run. The actual number returned depends on how many match your filters. |
| Can I filter for only top YC companies like Airbnb or Stripe? | Yes. Check the Top Companies Only box to restrict results to the most notable Y Combinator alumni. |
| How do I find only companies that are currently hiring? | Enable the Hiring Only checkbox. This will return only companies that have at least one open job listing on their YC profile. |
| What export formats are supported? | You can export your results to CSV, JSON, Excel, or XML directly from the dataset tab after the run completes. |
| Can I search by company name or keyword? | Yes. Use the Company name or keyword field to search across company names and descriptions. This works together with all other filters. |

### Related actors

Browse the full [ParseForge collection](https://apify.com/parseforge?fpr=vmoqkp) for more scrapers.

🆘 **Need help?** Email parseforge@protonmail.com with your run ID, your input, and what you expected.

⚠️ **Disclaimer.** This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by Y Combinator Management, LLC. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.

# Actor input Schema

## `startUrls` (type: `array`):

Scrape specific YC companies by URL. Example: https://www.ycombinator.com/companies/airbnb. If provided, all filter fields below are ignored.

## `maxItems` (type: `integer`):

How many companies to collect per run.

## `query` (type: `string`):

Search by company name, description, or keyword. Example: AI assistant

## `batches` (type: `array`):

Filter by YC batch. Use short codes (W25, S25, X25, F25) or full names (Winter 2025, Spring 2025, Fall 2025). Leave empty for all batches.

## `industries` (type: `array`):

Filter by top-level industry. Options: B2B, Consumer, Healthcare, Fintech, Industrials, Real Estate and Construction, Education, Government.

## `subindustries` (type: `array`):

Filter by specific subindustry. Examples: B2B -> Engineering, Product and Design | Fintech -> Payments | Healthcare -> Drug Discovery and Delivery.

## `regions` (type: `array`):

Filter by region. Examples: United States of America, Europe, Latin America, South Asia, Southeast Asia, Africa, India, United Kingdom.

## `companyStatus` (type: `string`):

Filter by company status. Leave empty for all statuses.

## `isHiring` (type: `boolean`):

Only return companies with open job listings.

## `nonprofit` (type: `boolean`):

Only return nonprofit organizations.

## `topCompaniesOnly` (type: `boolean`):

Only return companies flagged as top YC companies (most notable alumni like Airbnb, Stripe, Coinbase).

## `scrapeFounders` (type: `boolean`):

Include founder profiles with name, title, bio, LinkedIn, and Twitter.

## `scrapeJobs` (type: `boolean`):

Include open job listings with salary range, equity range, visa sponsorship, and required skills.

## `scrapeJobDescriptions` (type: `boolean`):

Include full job description text for each open position. Requires scrapeJobs to be enabled. Adds one extra request per job listing.

## Actor input object example

```json
{
  "maxItems": 10,
  "scrapeFounders": true,
  "scrapeJobs": true,
  "scrapeJobDescriptions": false
}
```

# Actor output Schema

## `results` (type: `string`):

Full dataset of YC company profiles with founders and job listings

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("parseforge/y-combinator-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "maxItems": 10 }

# Run the Actor and wait for it to finish
run = client.actor("parseforge/y-combinator-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "maxItems": 10
}' |
apify call parseforge/y-combinator-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,parseforge/y-combinator-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/VriFm9wnPo2TSETDQ/builds/Xmo0JNVNYzvTmvvQg/openapi.json
