# Gupy.io Scraper — Brazil Job Board (`unfenced-group/gupy-scraper`) Actor

Extract job listings from any Gupy.io company career page. No proxy, no API key. 64 fields per job including hiring pipeline stages, benefits, company social links and full structured address.

- **URL**: https://apify.com/unfenced-group/gupy-scraper.md
- **Developed by:** [Unfenced Group](https://apify.com/unfenced-group) (community)
- **Categories:** Jobs, Automation, Developer tools
- **Stats:** 1 total users, 0 monthly users, 100.0% runs succeeded, 1 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.19 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Gupy Scraper — Brazil Job Board

![Gupy.io Scraper — Brazil Job Board](https://api.apify.com/v2/key-value-stores/ClElVyZWvQgPQIuDL/records/gupy-scraper)

Extract job listings from any company career page on Gupy.io — Brazil's #1 HR and recruitment platform. **No proxy, no browser, no API key required.**

### Why this scraper?

#### 🚀 Complete data — 61 fields per job

Title, department, city, workplace type, full hiring pipeline stages, company social links, benefits, salary range, PCD/disability flag, video URLs and more.

#### ⚡ Fast listing mode — no detail fetches

Toggle `fetchDetails` off to retrieve all jobs in a single request per company — ideal for monitoring job counts, filtering by location or department, and building pipelines that don't need full descriptions.

#### 🌎 Multi-company in one run

Scrape Itaú, Ambev, Sicredi and any other Gupy-powered company simultaneously by listing their subdomains.

#### 🔒 Zero bot detection

Public career pages, no authentication, no proxy required. Fully reliable for scheduled runs.

***

### Input parameters

| Parameter | Type | Default | Description |
|---|---|---|---|
| `subdomains` | `string[]` | — | Company subdomains to scrape, e.g. `["vempra", "ambev", "sicredi"]`. Full URLs are also accepted. |
| `startUrls` | `RequestSource[]` | — | Alternative: full Gupy career URLs, e.g. `https://vempra.gupy.io`. |
| `fetchDetails` | `boolean` | `true` | When `true`, fetches each job's full detail page for description, prerequisites, responsibilities, benefits, salary, hiring pipeline, and all dates. When `false`, returns listing-level data only (faster and cheaper for large companies). |
| `maxJobsPerCompany` | `integer` | `0` | Cap results per company. `0` = unlimited. |
| `requestDelay` | `integer` | `200` | Mean delay in milliseconds between requests (Gaussian randomised). |
| `proxyConfiguration` | `Proxy` | — | Optional. Not required for Gupy — leave empty. |

***

### Output schema

Every field below is present on every record. Fields the source does not publish are returned as `null` rather than omitted.

| Field | Type | Description |
|---|---|---|
| `id` | string | Unique ID from the source. |
| `title` | string | Title as published. |
| `url` | string | Direct link to the listing. |
| `applyUrl` | string | Apply url. |
| `companyName` | string | Hiring company name. |
| `companyId` | string | Company id. |
| `subdomain` | string | Subdomain. |
| `type` | string | Type. |
| `department` | string | Department. |
| `workplaceType` | string | Workplace type. |
| `city` | string | City. |
| `state` | string | State. |
| `stateShortName` | string | State short name. |
| `country` | string | Country. |
| `quickApply` | boolean | Quick apply. |
| `scrapedAt` | string | Timestamp when this record was scraped. |
| `description` | string | Full description in plain text. |
| `descriptionMarkdown` | string | Description markdown. |
| `prerequisites` | string | Prerequisites. |
| `responsibilities` | string | Responsibilities. |
| `requirements` | string | Requirements. |
| `skills` | array | Skills. |
| `salary` | string | Salary as displayed (null if not published). |
| `salaryRange` | string | Salary range. |
| `publishedAt` | string | Date published. |
| `expiresAt` | string | Expires at. |
| `handicapped` | boolean | Handicapped. |
| `isInternalMobility` | boolean | Is internal mobility. |
| `jobType` | string | Job type. |
| `videoUrl` | string | Video url. |

```json
{
  "id": 11122229,
  "code": "G_CS_36",
  "jobUrl": "https://vempra.gupy.io/jobs/11122229",
  "applicationUrl": "https://vempra.gupy.io/jobs/11122229/apply",

  "companySubdomain": "vempra",
  "companyId": 2,
  "companyName": "Gupy",
  "companyCareerPageUrl": "https://vempra.gupy.io/jobs",
  "companyLogoUrl": "https://attachments.gupy.io/...",
  "companyBannerUrl": null,
  "companyWebsite": null,
  "companyAbout": "Somos uma empresa de tecnologia...",
  "companyAboutHtml": "<p>Somos uma empresa de tecnologia...</p>",
  "companyLinkedin": null,
  "companyFacebook": null,
  "companyInstagram": null,
  "companyGlassdoor": null,
  "companyTimezone": "America/Sao_Paulo",
  "companyLanguage": "pt",

  "title": "Consultor de Customer Success",
  "type": "vacancy_type_effective",
  "workplaceType": "hybrid",
  "isRemote": false,
  "publicationType": "external",
  "isInternalMobility": false,
  "quickApply": false,
  "department": null,
  "role": null,
  "jobLanguage": null,
  "skills": null,
  "reason": "staff_replacement",
  "isDisabilityJob": true,

  "fullAddress": "Av. Paulista, 1079, São Paulo, SP, 01311200",
  "city": "São Paulo",
  "state": "São Paulo",
  "stateCode": "SP",
  "country": "Brasil",
  "countryCode": "BR",

  "publishedAt": "2026-04-10T17:35:23.160Z",
  "expiresAt": "2026-05-04",
  "registerEndDate": "2026-06-09",

  "description": "Somos uma empresa de tecnologia...",

  "descriptionMarkdown": "## About\n\n...",
  "descriptionHtml": "<p>Somos uma empresa...</p>",
  "prerequisites": "Experiência prévia em Customer Success...",
  "prerequisitesHtml": "<ul><li>...</li></ul>",
  "responsibilities": "A área de Operações reúne...",
  "responsibilitiesHtml": "<p>A área de Operações...</p>",
  "benefits": "Vale-refeição no cartão Caju; Plano médico...",
  "benefitsHtml": "<ul><li>...</li></ul>",

  "jobSteps": [
    { "id": 65670466, "name": "Cadastro", "order": 0, "type": "register", "category": "registration" },
    { "id": 65670463, "name": "Entrevista", "order": 1, "type": "offline", "category": "interview" }
  ],
  "jobStepsCount": 7,

  "salaryRange": null,
  "journey": null,
  "contractType": null,

  "jobImageUrl": "https://attachments.gupy.io/...",
  "jobSocialImageUrl": "https://attachments.gupy.io/...",
  "videoUrl": null,
  "videoTitle": null,

  "isPostedOnGoogle": true,
  "status": "published",

  "scrapedAt": "2026-05-06T10:00:00.000Z"
}
```

***

### Examples

**Single company — full details**

```json
{
  "subdomains": ["vempra"],
  "fetchDetails": true
}
```

**Large company — listing only (fast mode)**

```json
{
  "subdomains": ["sicredi"],
  "fetchDetails": false
}
```

**Multi-company run**

```json
{
  "subdomains": ["vempra", "ambev", "suzano", "usiminas"],
  "fetchDetails": true,
  "maxJobsPerCompany": 100
}
```

**Finding a company's subdomain**

Visit the company's Gupy career page. The subdomain is the part before `.gupy.io`:

- `https://vempra.gupy.io` → `vempra`
- `https://ambev.gupy.io` → `ambev`
- `https://sicredi.gupy.io` → `sicredi`

You can also paste the full URL directly in `startUrls` and the subdomain is extracted automatically.

***

**Scheduled daily run:**

```json
{
  "startUrls": [
    {
      "url": "https://vempra.gupy.io/"
    }
  ],
  "maxResults": 500
}
```

Schedule this input in the Apify Scheduler (for example daily at 07:00) to keep an always-fresh dataset.

### 💰 Pricing

**$1.49 per 1,000 results** — you only pay for successfully retrieved listings.
Failed retries are never charged.

| Results | Cost |
|---|---|
| 100 | ~$0.15 |
| 1,000 | ~$1.49 |
| 10,000 | ~$14.90 |
| 100,000 | ~$149.00 |

> Flat-rate alternatives typically charge $29–$49/month regardless of usage.
> At 10,000 results/month, this scraper costs significantly less with no commitment.

Use the **Max jobs per company** cap to control your spend exactly.

***

### Performance

| Company size | Jobs | Mode | Approx. time |
|---|---|---|---|
| Small (< 20 jobs) | 10 | `fetchDetails: true` | ~5 s |
| Medium (50 jobs) | 50 | `fetchDetails: true` | ~20 s |
| Large (300+ jobs) | 300 | `fetchDetails: true` | ~2 min |
| Large (300+ jobs) | 300 | `fetchDetails: false` | ~3 s |

Memory: 256 MB.

***

### Known limitations

- **Apply URL:** Always points to the Gupy application form. Gupy does not expose external ATS redirect URLs.
- **Salary:** Not published by all employers — will be `null` when unavailable. Gupy does not require salary disclosure.
- **Company scope:** Each run is scoped to the subdomains you provide. There is no built-in discovery of all companies on Gupy.
- **Internal vacancies:** Jobs with `publicationType: "internal"` are visible when the career page is public. Internal-only pages (requiring login) cannot be scraped.

***

### Technical details

- **Source:** gupy.io — Brazil's leading ATS and job board platform
- **Memory:** 256 MB
- **Repost storage:** KeyValueStore `gupy-scraper-job-dedup`, 90-day TTL
- **Retry:** Automatic retry on network errors, exponential backoff, 3 attempts per request

***

### Additional services

Need a custom actor, additional filters, scheduled runs, or integration support?.nl]\(mailto:info@unfencedgroup.nl) — we build on request.

***

### Related scrapers

Other scrapers in our **Jobs — Americas** collection:

- [Bumeran Group Scraper — 8 LATAM Job Boards](https://apify.com/unfenced-group/bumeran-scraper)
- [Computrabajo.com Job Scraper](https://apify.com/unfenced-group/computrabajo-com-job-scraper)
- [Eluta.ca Scraper](https://apify.com/unfenced-group/eluta-ca-scraper)

***

### Run it on a schedule

This actor is built for repeat use. Set it to run daily, weekly, or hourly, and the data keeps flowing without you touching it.

- **Schedule runs** — open the actor, go to Schedules, and pick a cadence. Each run only charges you for the results it returns.
- **Connect it to your stack** — push results straight to Google Sheets, Slack, a webhook, or your database using Apify Integrations. No glue code needed.
- **Pull results via API** — every run writes a clean dataset you can fetch with one API call, ready for whatever you build on top of it.

Set it once and it runs on its own.

***

### Need a custom scraper?

**[Unfenced Group](https://www.unfencedgroup.nl)** builds Apify actors for any website — for free.

If the site you need isn't in our portfolio yet, just ask. We scope, build, and publish it at no cost to you. You only pay for results — we absorb the compute and proxy costs ourselves. Same pay-per-result pricing, same quality, same standards as every actor in this portfolio.

**Get in touch:** [www.unfencedgroup.nl](https://www.unfencedgroup.nl)

# Actor input Schema

## `subdomains` (type: `array`):

List of Gupy company subdomains to scrape. Example: \["vempra", "itau", "bayer"]. You can also paste the full URL (https://vempra.gupy.io) and the subdomain will be extracted automatically.

## `startUrls` (type: `array`):

Alternative to 'subdomains'. Paste full Gupy career page URLs here (e.g. https://vempra.gupy.io). Subdomains will be extracted automatically.

## `fetchDetails` (type: `boolean`):

When enabled (default), each job page is fetched for the complete data set: description, prerequisites, responsibilities, benefits, salary, hiring pipeline, and all dates. When disabled, only listing-level data is returned (title, department, location, workplace type) — much faster and cheaper for large companies.

## `maxJobsPerCompany` (type: `integer`):

Maximum number of jobs to scrape per company. Set to 0 for unlimited.

## `requestDelay` (type: `integer`):

Mean delay in milliseconds between requests when fetching job detail pages. Actual delay uses Gaussian randomisation. Default: 200ms.

## `proxyConfiguration` (type: `object`):

Optional. Not required for Gupy — the career pages are publicly accessible. Leave empty to run without proxy.

## `maxResults` (type: `integer`):

Maximum number of results to return.

## Actor input object example

```json
{
  "subdomains": [
    "vempra",
    "itau",
    "bayer"
  ],
  "startUrls": [
    {
      "url": "https://vempra.gupy.io/"
    }
  ],
  "fetchDetails": true,
  "maxJobsPerCompany": 0,
  "requestDelay": 200,
  "maxResults": 100
}
```

# Actor output Schema

## `results` (type: `string`):

Job listings scraped from Gupy-powered company career pages. Each item contains up to 57 fields including title, location, workplace type, hiring pipeline, company info, description, prerequisites, responsibilities, and benefits.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "subdomains": [
        "vempra"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("unfenced-group/gupy-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "subdomains": ["vempra"] }

# Run the Actor and wait for it to finish
run = client.actor("unfenced-group/gupy-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "subdomains": [
    "vempra"
  ]
}' |
apify call unfenced-group/gupy-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,unfenced-group/gupy-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/qoIbI4xRvprcxxMRT/builds/mXYI31mpmJOM12G4z/openapi.json
