# LinkedIn Profile & Company Data Extractor (`second_coming/linkedin-data-extractor`) Actor

Extract company employee lists, decision-maker profiles, job postings, and company page data from LinkedIn. B2B lead generation tool for sales teams and recruiters.

- **URL**: https://apify.com/second\_coming/linkedin-data-extractor.md
- **Developed by:** [Richard P](https://apify.com/second_coming) (community)
- **Categories:** Lead generation, Business
- **Stats:** 2 total users, 0 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$0.02 / scan run

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## LinkedIn Profile & Company Data Extractor

An [Apify Actor](https://apify.com/actors) for extracting publicly available data from LinkedIn company pages and public profiles. Uses JSON-LD structured data, meta tags, and HTML parsing — no login required.

### Features

#### Company Data Extraction

- **Company Name** — extracted from JSON-LD, meta tags, and HTML
- **Industry** — parsed from the company about page
- **Company Size** — employee count or size range
- **Headquarters** — location from about section
- **Website** — from JSON-LD `sameAs` or visible links
- **Description** — company summary
- **Followers** — follower count from interaction statistics
- **Specialties** — comma/semicolon-separated list

#### Profile Data Extraction

- **Name & Headline** — from JSON-LD Person schema, meta tags, HTML
- **Location** — geographic location
- **About** — profile summary text
- **Experience** — work history (company, title, dates, description)
- **Education** — schools, degrees, dates
- **Connections Count** — if publicly visible

#### Job Listing Extraction

- For each company, fetches job postings from LinkedIn's public jobs search
- Extracts: title, location, posted date, description, employment type
- Configurable limit (default 20, max 50)

### Input

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `companyUrls` | `string[]` | `["https://www.linkedin.com/company/microsoft/"]` | LinkedIn company page URLs |
| `profileUrls` | `string[]` | `[]` | LinkedIn public profile URLs |
| `includeJobs` | `boolean` | `true` | Whether to fetch job listings |
| `maxJobs` | `integer` | `20` | Max job listings per company (max 50) |

#### Example Input

```json
{
  "companyUrls": [
    "https://www.linkedin.com/company/microsoft/",
    "https://www.linkedin.com/company/google/"
  ],
  "profileUrls": [
    "https://www.linkedin.com/in/sundarpichai/"
  ],
  "includeJobs": true,
  "maxJobs": 10
}
```

### Output

Each record is pushed to the Apify dataset. Records have a `type` field distinguishing companies from profiles.

#### Company Record Fields

| Field | Type | Description |
|-------|------|-------------|
| `type` | `string` | `"company"` |
| `companyName` | `string` | Company name |
| `linkedinUrl` | `string` | LinkedIn URL |
| `slug` | `string` | URL slug |
| `industry` | `string` | Industry |
| `companySize` | `string` | Employee count / range |
| `headquarters` | `string` | HQ location |
| `website` | `string` | Company website |
| `description` | `string` | Company description |
| `followers` | `integer` | LinkedIn followers |
| `specialties` | `string[]` | Specialty tags |
| `jobs` | `object[]` | Job listings (if requested) |
| `scrapedAt` | `string` | ISO 8601 timestamp |

#### Profile Record Fields

| Field | Type | Description |
|-------|------|-------------|
| `type` | `string` | `"profile"` |
| `name` | `string` | Full name |
| `headline` | `string` | Profile headline |
| `location` | `string` | Location |
| `about` | `string` | About section |
| `experience` | `object[]` | Work experience |
| `education` | `object[]` | Education |
| `connectionsCount` | `integer` | Connection count |
| `linkedinUrl` | `string` | LinkedIn URL |
| `scrapedAt` | `string` | ISO 8601 timestamp |

#### Job Entry Fields (inside `jobs` array)

| Field | Type | Description |
|-------|------|-------------|
| `title` | `string` | Job title |
| `location` | `string` | Job location |
| `postedDate` | `string` | Date posted |
| `description` | `string` | Job description |
| `employmentType` | `string` | Full-time, part-time, etc. |
| `seniority` | `string` | Seniority level |

### Technical Approach

1. **JSON-LD** — primary data source. Parse `script[type="application/ld+json"]` tags for `Organization`, `Corporation`, `Person`, and `JobPosting` schemas.
2. **Meta tags** — secondary source. Extract OG and standard meta tags.
3. **Visible HTML** — tertiary source. Regex patterns and CSS selectors for fields not in structured data.
4. **Rate limiting** — 3.5s delay between requests with exponential backoff on 429 responses.
5. **Google Cache fallback** — if the direct LinkedIn fetch fails, tries `webcache.googleusercontent.com`.
6. **Graceful abort** — listens for the Apify `ABORTING` event to stop mid-crawl.

### Limitations

> **Important**: LinkedIn is extremely protective of its data. This actor only extracts what is **publicly visible** without authentication. Many profiles and data fields require a LinkedIn login to access.

- **Profile details** (experience, education, about) are often only partially visible without login. Profiles with "public" visibility settings yield the most data.
- **Job listings** are fetched from the public jobs search page. Some companies may require login to view jobs.
- **HTML selectors** for experience/education sections are fragile — LinkedIn changes class names frequently. JSON-LD parsing is far more reliable.
- **Rate limiting** — aggressive scraping triggers LinkedIn's rate limits. The actor implements delays and backoff.
- **No login/cookie support** — this actor never authenticates. For authenticated scraping, consider using the Apify LinkedIn Actors with built-in session management.

### Local Development

```bash
## Install dependencies
pip install -r requirements.txt

## Run locally with Apify CLI
apify run --purge
```

### Deployment

```bash
apify push
```

### Input Schema

See `.actor/input_schema.json` for the full input specification.

# Actor input Schema

## `companyUrls` (type: `array`):

LinkedIn company page URLs to scrape (e.g., https://www.linkedin.com/company/microsoft/)

## `profileUrls` (type: `array`):

LinkedIn public profile URLs to scrape (e.g., https://www.linkedin.com/in/sundarpichai/)

## `includeJobs` (type: `boolean`):

Whether to fetch job listings for each company from LinkedIn's public jobs search

## `maxJobs` (type: `integer`):

Maximum number of job listings to extract per company (max 50)

## Actor input object example

```json
{
  "companyUrls": [
    "https://www.linkedin.com/company/microsoft/"
  ],
  "profileUrls": [],
  "includeJobs": true,
  "maxJobs": 20
}
```

# Actor output Schema

## `results` (type: `string`):

Dataset URL containing scraped LinkedIn data

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "companyUrls": [
        "https://www.linkedin.com/company/microsoft/"
    ],
    "profileUrls": []
};

// Run the Actor and wait for it to finish
const run = await client.actor("second_coming/linkedin-data-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "companyUrls": ["https://www.linkedin.com/company/microsoft/"],
    "profileUrls": [],
}

# Run the Actor and wait for it to finish
run = client.actor("second_coming/linkedin-data-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "companyUrls": [
    "https://www.linkedin.com/company/microsoft/"
  ],
  "profileUrls": []
}' |
apify call second_coming/linkedin-data-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,second_coming/linkedin-data-extractor"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/uskbJRZU30cAu6ySg/builds/FPMFgaMIL0rfJV5OQ/openapi.json
