# LinkedIn & Indeed Job Scraper - Hiring Data API (`groupoject/indeed-jobs-scraper`) Actor

Scrape and monitor LinkedIn, Indeed, Google Jobs and optional global job boards. Rich company and salary data, hiring-signal scores, deduplication, and persistent new-job alerts. No login or API key.

- **URL**: https://apify.com/groupoject/indeed-jobs-scraper.md
- **Developed by:** [Group Oject](https://apify.com/groupoject) (community)
- **Categories:** Jobs, Lead generation
- **Stats:** 7 total users, 1 monthly users, 100.0% runs succeeded, 1 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## LinkedIn & Indeed Job Scraper - Hiring Data API

**Scrape and monitor LinkedIn, Indeed, Google Jobs, Glassdoor, ZipRecruiter, Bayt, and Naukri. Export rich, deduplicated jobs with company intelligence, normalized salary data, direct application URLs, and hiring-signal scores. No login or API key.**

Search once and get one structured dataset. Use it to find software engineer jobs, remote data jobs, AI/ML roles, product manager openings, nurse jobs, marketing jobs, sales openings, finance roles, and other hiring-market data.

This Actor is built to be the practical, lower-friction alternative to heavyweight job-board scrapers: **Indeed-first reliability**, optional extra boards, low default run size, and pay-per-result economics.

***

### What it does

- Searches one or many keywords across your chosen job boards and locations.
- Returns **one rich row per job** with salary components, company intelligence, descriptions, skills, emails, direct apply URLs, job age, and source metadata.
- **Deduplicates across boards** so the same job appears once.
- Filters by remote, job type, posting recency, and location.
- Supports radius, pagination, Easy Apply, LinkedIn company targeting, and optional full LinkedIn descriptions.
- Converts hourly, daily, weekly, and monthly compensation to comparable annual salary ranges.
- Scores each record from 0-100 as a **hiring signal** using freshness and data completeness.
- `onlyNewJobs` remembers prior results, making scheduled runs a clean incremental job feed.
- Streams results into the Apify dataset so API users can start consuming rows quickly.
- Works with Apify schedules, webhooks, API clients, Make, Zapier, n8n, Google Sheets, CRMs, and AI workflows.

**Indeed is the reliable default, and LinkedIn is verified as a second source.** Google Jobs uses Apify's dedicated Google SERP route but remains beta because that upstream service can time out. Glassdoor, ZipRecruiter, Bayt, and Naukri are optional beta sources that may return no data when challenged, even with a residential proxy. The Actor reports those board-level outcomes honestly while preserving successful results from other boards.

***

### Who it's for

- **Recruiters & sourcers** — pull fresh openings across boards into one sheet.
- **Job seekers & career coaches** — run repeatable searches for remote, local, or niche roles.
- **Job-board builders** — aggregate listings programmatically for niche job sites.
- **Sales & lead generation teams** — use new job postings as buying signals.
- **Labor-market researchers** — track roles, salaries, and hiring trends.
- **ATS / HR-tech** — feed live postings into your pipeline.
- **AI agents and MCP workflows** — give agents structured job listings on demand.

***

### Popular searches and use cases

Use this Actor for focused job-search and market-research workflows such as:

- **Software engineer jobs in New York** — collect engineering roles from Indeed, Google Jobs, and optional extra boards with salary and company details.
- **Remote data engineer jobs** — find remote data engineering openings across major job boards.
- **Remote AI and machine learning jobs** — search AI engineer, ML engineer, and machine learning engineer roles in one run.
- **Product manager jobs in London** — gather UK product roles with normalized job records.
- **Digital marketing jobs in the USA** — track SEO specialist, content marketing, and digital marketing manager openings.
- **Nurse jobs in California** — collect registered nurse, RN, and nurse practitioner jobs by location.
- **Recruiter sourcing lists** — build structured lead lists of current hiring companies by role, location, and board.

You can save any configuration as an Apify task for repeated runs, scheduled job monitoring, or API integrations.

***

### Use with AI agents and MCP

Apify Actors can be called from the Apify API and Apify MCP server, so this scraper can act as a live jobs tool for Claude, ChatGPT, Cursor, internal agents, or workflow automations.

Useful agent prompts:

- "Find remote senior Python jobs posted this week and return companies with salary ranges."
- "Which AI startups are hiring sales and partnerships roles in the US?"
- "Build a CSV of product manager jobs in London from Indeed and Google Jobs."
- "Monitor nurse jobs in California every morning and send new matches to a spreadsheet."

***

### Pricing and cost control

This Actor is pay-per-result: users pay only for job rows written to the dataset, plus Apify platform usage. Keep `maxResultsPerSite` and `maxItems` low while testing, then scale after the query is proven.

Good defaults:

- Quick test: `maxResultsPerSite: 20`, `maxItems: 50`
- Normal task page: `maxResultsPerSite: 50`, `maxItems: 250`
- Larger research run: `maxResultsPerSite: 100`, `maxItems: 1000`

The Actor defaults to 1 GB memory and is designed to keep normal Indeed-first runs cheap.

***

### Input

| Field | Type | Default | Description |
|---|---|---|---|
| `searchTerm` | string | — | Job title / keyword (e.g. "software engineer") |
| `searchTerms` | array | — | Run several searches in one go |
| `location` | string | — | City, state, or country (e.g. "New York, NY") |
| `jobBoards` | array | `[indeed]` | Indeed, Google, LinkedIn, Glassdoor, ZipRecruiter, Bayt, or Naukri |
| `maxResultsPerSite` | integer | `50` | Results per board per term |
| `maxItems` | integer | `0` | Hard cap on total jobs (0 = unlimited) |
| `hoursOld` | integer | — | Only jobs posted within N hours |
| `jobType` | enum | any | full-time / part-time / contract / internship / temporary |
| `remoteOnly` | boolean | `false` | Remote positions only |
| `countryIndeed` | string | `USA` | Country for Indeed / Glassdoor |
| `distance` | integer | `50` | Search radius in miles where supported |
| `offset` | integer | `0` | Pagination offset |
| `easyApply` | boolean | `false` | Easy Apply jobs only where supported |
| `linkedinFetchDescription` | boolean | `false` | Fetch full LinkedIn descriptions and direct URLs |
| `linkedinCompanyIds` | array | — | Restrict LinkedIn to company IDs |
| `normalizeAnnualSalary` | boolean | `false` | Convert salaries to annual estimates |
| `onlyNewJobs` | boolean | `false` | Emit only jobs not seen by this monitor before |
| `stateKey` | string | `default` | Separate persistent histories for multiple monitors |
| `proxyConfiguration` | object | Apify Proxy | Proxy settings. Leave default for the most reliable runs |

At least one of `searchTerm` / `searchTerms` is required.

#### Example input

```json
{
  "searchTerm": "software engineer",
  "location": "New York, NY",
  "jobBoards": ["indeed"],
  "maxResultsPerSite": 50,
  "remoteOnly": false
}
```

#### Remote AI jobs example

```json
{
  "searchTerms": ["ai engineer", "machine learning engineer", "ml engineer"],
  "location": "Remote",
  "remoteOnly": true,
  "jobBoards": ["indeed"],
  "maxResultsPerSite": 40,
  "maxItems": 300,
  "countryIndeed": "USA"
}
```

#### Recruiter sourcing example

```json
{
  "searchTerms": ["software engineer", "data engineer", "product manager"],
  "location": "United States",
  "jobBoards": ["indeed"],
  "maxResultsPerSite": 50,
  "maxItems": 500,
  "countryIndeed": "USA"
}
```

#### LinkedIn + Indeed sourcing example

```json
{
  "searchTerms": ["sales manager", "account executive", "business development"],
  "location": "United States",
  "jobBoards": ["indeed", "google", "linkedin"],
  "maxResultsPerSite": 30,
  "maxItems": 300,
  "countryIndeed": "USA"
}
```

#### Fresh remote software jobs example

```json
{
  "searchTerms": ["software engineer", "backend engineer", "full stack developer"],
  "location": "Remote",
  "remoteOnly": true,
  "jobBoards": ["indeed"],
  "hoursOld": 168,
  "maxResultsPerSite": 50,
  "maxItems": 300
}
```

#### Healthcare jobs example

```json
{
  "searchTerms": ["registered nurse", "nurse practitioner", "medical assistant"],
  "location": "California",
  "jobBoards": ["indeed"],
  "maxResultsPerSite": 50,
  "maxItems": 500,
  "countryIndeed": "USA"
}
```

#### Sales jobs lead-generation example

```json
{
  "searchTerms": ["account executive", "sales manager", "sales development representative"],
  "location": "United States",
  "jobBoards": ["indeed", "google", "zip_recruiter"],
  "maxResultsPerSite": 50,
  "maxItems": 300,
  "countryIndeed": "USA"
}
```

#### Finance jobs example

```json
{
  "searchTerms": ["financial analyst", "controller", "accounting manager"],
  "location": "New York, NY",
  "jobBoards": ["indeed", "google"],
  "maxResultsPerSite": 50,
  "maxItems": 250,
  "countryIndeed": "USA"
}
```

#### Cybersecurity jobs example

```json
{
  "searchTerms": ["cybersecurity analyst", "security engineer", "SOC analyst"],
  "location": "Remote",
  "remoteOnly": true,
  "jobBoards": ["indeed", "google", "zip_recruiter"],
  "maxResultsPerSite": 50,
  "maxItems": 300,
  "countryIndeed": "USA"
}
```

#### Construction jobs example

```json
{
  "searchTerms": ["construction project manager", "estimator", "superintendent"],
  "location": "Texas",
  "jobBoards": ["indeed", "google"],
  "maxResultsPerSite": 50,
  "maxItems": 250,
  "countryIndeed": "USA"
}
```

More in [`examples/`](examples/).

***

### Output

One dataset row per job:

```json
{
  "title": "Senior Software Engineer",
  "company": "Acme",
  "location": "New York, NY, US",
  "jobBoard": "indeed",
  "jobType": "fulltime",
  "isRemote": false,
  "salary": "$120,000 - $150,000 / year",
  "datePosted": "2026-06-12",
  "jobUrl": "https://www.indeed.com/viewjob?jk=...",
  "companyUrl": null,
  "description": "…",
  "searchTerm": "software engineer"
}
```

A `SUMMARY` key-value record holds totals, boards, and any warnings.

***

### Tips & limitations

- **Indeed** is the reliable default board. **Google Jobs** automatically uses the dedicated Apify Google SERP proxy and currently returns up to 10 results per search term. Google SERP requests are billed as platform proxy usage. **ZipRecruiter, LinkedIn, and Glassdoor** remain best-effort boards; use a residential proxy when they are blocked.
- For large or scheduled runs, keep Apify Proxy enabled to avoid rate limits.
- Salaries are only as complete as the board provides; many postings omit them.
- Built on the open-source [`python-jobspy`](https://github.com/cullenwatson/JobSpy) library.
- Scrapes publicly listed jobs only. Respect each board's terms of use.

***

### Changelog

See [CHANGELOG.md](CHANGELOG.md).

# Actor input Schema

## `searchTerm` (type: `string`):

Job title or keyword to search for, e.g. "software engineer".

## `searchTerms` (type: `array`):

Run several searches in one go (in addition to or instead of the single term above).

## `location` (type: `string`):

City, state, or country, e.g. "New York, NY" or "Remote".

## `jobBoards` (type: `array`):

Indeed and LinkedIn are currently the most reliable. Google Jobs uses a dedicated SERP route. Glassdoor, ZipRecruiter, Bayt, and Naukri are optional beta sources and may require residential proxy or return no data when challenged.

## `maxResultsPerSite` (type: `integer`):

How many results to request from each board per search term. Google Jobs currently returns up to 10 results per query.

## `maxItems` (type: `integer`):

Hard cap on total jobs pushed to the dataset (0 = unlimited).

## `hoursOld` (type: `integer`):

Only jobs posted within this many hours. Leave empty for no limit.

## `jobType` (type: `string`):

Filter by employment type.

## `remoteOnly` (type: `boolean`):

Only return remote positions.

## `countryIndeed` (type: `string`):

Country for Indeed/Glassdoor, e.g. USA, UK, Canada, Germany.

## `distance` (type: `integer`):

Search radius around the location where supported.

## `offset` (type: `integer`):

Skip the first N results for pagination where supported.

## `easyApply` (type: `boolean`):

Return quick-apply jobs where the selected board supports this filter.

## `linkedinFetchDescription` (type: `boolean`):

Fetch complete LinkedIn descriptions and direct application URLs. Slower but richer.

## `linkedinCompanyIds` (type: `array`):

Optional LinkedIn company IDs to restrict the search.

## `normalizeAnnualSalary` (type: `boolean`):

Convert hourly, daily, weekly, and monthly salary ranges to annual estimates.

## `onlyNewJobs` (type: `boolean`):

For scheduled runs, emit only jobs not seen in earlier runs using the same state key.

## `stateKey` (type: `string`):

Separates saved history for different recurring monitors.

## `proxyConfiguration` (type: `object`):

Apify Proxy is used by default when no proxy is provided. Residential proxy works best for LinkedIn, Glassdoor, and larger runs.

## `debugMode` (type: `boolean`):

Verbose logging.

## Actor input object example

```json
{
  "searchTerm": "software engineer",
  "location": "New York, NY",
  "jobBoards": [
    "indeed"
  ],
  "maxResultsPerSite": 50,
  "maxItems": 0,
  "jobType": "",
  "remoteOnly": false,
  "countryIndeed": "USA",
  "distance": 50,
  "offset": 0,
  "easyApply": false,
  "linkedinFetchDescription": false,
  "normalizeAnnualSalary": false,
  "onlyNewJobs": false,
  "stateKey": "default",
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "debugMode": false
}
```

# Actor output Schema

## `jobs` (type: `string`):

One rich record per job with salary, company intelligence, descriptions, skills, direct apply URL, age, and hiring-signal score.

## `summary` (type: `string`):

Search terms, boards, total + unique jobs, warnings.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchTerm": "software engineer",
    "location": "New York, NY",
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("groupoject/indeed-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchTerm": "software engineer",
    "location": "New York, NY",
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("groupoject/indeed-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchTerm": "software engineer",
  "location": "New York, NY",
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call groupoject/indeed-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,groupoject/indeed-jobs-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/LzsIpToRTUZSnqjbS/builds/3g5ACFYCeMx8JzwUN/openapi.json
