# Personio Job Board Scraper — DACH SME Hiring Signals (`foxlabs/personio-job-board-scraper`) Actor

Scrape open roles from any Personio careers site by company slug or URL. Returns title, department, office, employment type, seniority, schedule, years of experience, description sections and apply URL.

- **URL**: https://apify.com/foxlabs/personio-job-board-scraper.md
- **Developed by:** [Berkan Kaplan](https://apify.com/foxlabs) (community)
- **Categories:** Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$2.00 / 1,000 job postings

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Personio Job Board Scraper — DACH SME Hiring Signals

Personio is the HR system of the German-speaking Mittelstand, and every customer publishes an XML job feed. That makes it the clearest public window into SME hiring across Germany, Austria and Switzerland — a segment the US-centric ATS feeds miss entirely.

**No API key · Official source · Pay only for delivered rows · Same schema across the series**

### What data do you get?

| Field | Description |
|---|---|
| `companyName` | Company slug, title-cased |
| `companyBoard` | Personio slug — the join key across runs |
| `jobId` | Personio position ID |
| `jobTitle` | Role title |
| `department` | Department |
| `location` | Office |
| `isRemote` | True when the office or title marks the role remote |
| `employmentType` | Permanent, intern, working student, temporary |
| `seniority` | Seniority level as published |
| `schedule` | Full-time or part-time |
| `yearsOfExperience` | Experience band where published |
| `description` | Description sections joined, HTML stripped |
| `occupationCategory` | Personio's occupation category |
| `applyUrl` | Direct application link |
| `sourceUrl` | Public posting URL |

Every row also carries `query` (what you asked for), `scrapedAt` (ISO timestamp) and, when a
lookup fails, `error` explaining why.

### Example output

```json
{
  "source": "Personio",
  "companyName": "Personio",
  "companyBoard": "personio",
  "jobId": "1834171",
  "jobTitle": "Senior Software Engineer (m/f/d)",
  "department": "Engineering",
  "location": "Munich",
  "employmentType": "permanent",
  "seniority": "experienced",
  "schedule": "full-time",
  "sourceUrl": "https://personio.jobs.personio.de/job/1834171"
}
```

### Input

```json
{
  "queries": ["personio","deskbird","https://deskbird.jobs.personio.com/"],
  "maxResultsPerQuery": 200,
  "maxConcurrency": 2,
  "includeRaw": false
}
```

| Input | What it does |
|---|---|
| `queries` | Company slugs (`personio`, `deskbird`) or careers URLs (`https://deskbird.jobs.personio.com`). Both `.de` and `.com` careers domains work. |
| `maxResultsPerQuery` | Caps how many rows one query may produce. |
| `maxConcurrency` | How many queries run at once. Lower it if the source starts throttling. |
| `includeRaw` | Attaches the source's untouched record under `raw`, for fields this actor does not map. |
| `requestDelayMs` | Politeness delay between requests. |
| `proxyConfiguration` | Personio rate-limits aggressively per IP. Turn the proxy on for runs over a handful of companies. |

### What people use it for

- **DACH SME prospecting** — Personio's base is German-speaking mid-market — the hardest segment to find hiring data for anywhere else.
- **Hiring-intent scoring** — a Mittelstand company opening its first data roles is starting a data project.
- **Recruiter sourcing** — collect open roles across a target list of DACH employers.

### Notes and limits

- Personio throttles hard. The defaults here are deliberately conservative — two companies at a time with a half-second delay — and the proxy input is there for larger runs.
- Feeds are usually German-language; field values like `permanent` and `full-time` come back in Personio's own English vocabulary regardless.
- Personio accounts sit on either `.jobs.personio.de` or `.jobs.personio.com`. Pass a bare slug and the actor tries both, so you do not have to know which.

### Where the data comes from

Every Personio careers site publishes its openings as an XML feed at `/xml`, with no key. Source: [Personio XML job feed](https://www.personio.com/)

### FAQ

#### Is this Personio job board scraper free?

The data source is free and needs no API key — you pay only for the rows the run delivers ($0.002 each). Failed or empty lookups are never charged.

#### Do I need an API key or a login?

No. Every Personio careers site publishes its openings as an XML feed at `/xml`, with no key.

#### What can I search by?

By company slug (`personio`) or a full careers URL (`https://sennder.jobs.personio.de`). Both `.de` and `.com` careers domains work.

#### How current is the data?

Every run queries the source live, so results are as fresh as the source itself. The feed is generated live by Personio.

#### How fast is it, and how many queries can I run?

Queries run concurrently (2 at a time by default, tunable in the input). A prefilled run finishes in seconds; large lists scale roughly linearly and stay well inside a normal run timeout.

#### Can I export the results to CSV, Excel or JSON?

Yes. Apify datasets export to CSV, Excel, JSON, XML and HTML, and can be pulled through the API or pushed to your own storage.

#### What happens when a query returns nothing?

You still get a row, carrying your original `query` and an `error` field explaining why. Nothing is silently dropped, and you are not charged for it.

#### Is scraping this data legal?

Yes. The XML feed is a syndication format Personio publishes on the company's behalf. It carries job data, not personal data.

### Related actors

- [Teamtailor Job Scraper](https://apify.com/foxlabs/teamtailor-job-board-scraper)
- [SmartRecruiters Job Scraper](https://apify.com/foxlabs/smartrecruiters-job-board-scraper)

***

Built by [Fox Labs](https://apify.com/foxlabs) — B2B company intelligence from public sources, as clean JSON.

# Actor input Schema

## `queries` (type: `array`):

Company slugs (`personio`, `deskbird`) or careers URLs (`https://deskbird.jobs.personio.com`). Both `.de` and `.com` careers domains work.

## `maxResultsPerQuery` (type: `integer`):

How many rows a single query may produce.

## `maxConcurrency` (type: `integer`):

How many queries to run at the same time. Lower it if the source throttles you.

## `includeRaw` (type: `boolean`):

Attach the source's untouched response under `raw`. Useful when you need a field this actor does not map.

## `requestDelayMs` (type: `integer`):

Politeness delay against a public source. Raise it for large runs.

## `proxyConfiguration` (type: `object`):

Personio rate-limits aggressively per IP. Turn the proxy on for runs over a handful of companies.

## Actor input object example

```json
{
  "queries": [
    "personio",
    "deskbird",
    "https://deskbird.jobs.personio.com/"
  ],
  "maxResultsPerQuery": 200,
  "maxConcurrency": 2,
  "includeRaw": false,
  "requestDelayMs": 500,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "personio",
        "deskbird",
        "https://deskbird.jobs.personio.com/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("foxlabs/personio-job-board-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "queries": [
        "personio",
        "deskbird",
        "https://deskbird.jobs.personio.com/",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("foxlabs/personio-job-board-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "personio",
    "deskbird",
    "https://deskbird.jobs.personio.com/"
  ]
}' |
apify call foxlabs/personio-job-board-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,foxlabs/personio-job-board-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/yt2dBF6fxOL304HF8/builds/uZoTAU9FQW2hC92xu/openapi.json
