# Local Business Leads Scraper: Verified Business Emails & Finder (`flash_scraper/local-business-leads`) Actor

Local business leads scraper, local business email finder and local business email scraper: any category, any city. Verified business emails (MX-checked), phones, socials on every row - local leads with emails, no API key, no proxy, $3 per 1,000. Business email finder for agencies.

- **URL**: https://apify.com/flash\_scraper/local-business-leads.md
- **Developed by:** [Flash Scrape](https://apify.com/flash_scraper) (community)
- **Categories:** Lead generation, Business, Marketing
- **Stats:** 38 total users, 26 monthly users, 99.4% runs succeeded, 6 bookmarks
- **User rating**: 5.00 out of 5 stars

## Pricing

from $2.40 / 1,000 business leads

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Local Business Leads Scraper: Verified Business Emails & Finder

**Local Business Leads Scraper** is a pay-per-result lead-generation actor that finds local businesses in any category and any city on earth and delivers MX-verified emails, phones, social profiles, website platform and a 0-100 lead score — built on OpenStreetMap, so it needs no API key, no proxy and no login.

**Try it:** [Find local business leads with emails by category](https://apify.com/flash_scraper/local-business-leads/examples/find-local-business-leads)

**Local business leads scraper and business email finder** — verified business emails, phones and socials for any category in any city; a local business email scraper built for agencies.

### What you get per lead

One row per business, 73 stable columns, with a Google Maps link and OpenStreetMap provenance on every row. Fill rates are measured on the reference run — `dentist` / `Austin, Texas`, `onlyWithWebsite: true`, n=55, 2026-08-08 (`google_maps_url` and `osm_url` also measured 100% on the same day's filter-off n=100 run); the full story is under *Measured, not promised* below.

| Field | Filled | What it is |
|---|---|---|
| MX-verified email + whose mailbox | 55% | `email` graded `deliverable` / `risky` / `undeliverable`; `email_type` says own domain, free inbox or the business's marketing agency |
| Phone | 96% | `phone` is the OSM number; the site's own `tel:` numbers ship in `phones` — a row with either counts as having a phone |
| Social profiles | Facebook 73% · Instagram 55% · YouTube 25% · LinkedIn 22% | profile URLs the business's own site links |
| Website platform + tech | 71% | WordPress, Wix, Shopify, Squarespace… plus Meta-pixel / Google-Ads-tag flags |
| Lead score + grade | 100% | 0-100 completeness/reachability score and an `A`-`F` grade — not a customer rating |
| Google Maps link | 100% | `google_maps_url` opens a Google Maps search for the business in one click; nothing is scraped from Google |
| OpenStreetMap provenance | 100% | `osm_url`, coordinates and the ODbL attribution string |

**$3 per 1,000 delivered leads** ($0.003 per lead) on the free plan; paid plans pay less (Pricing tab). Every filter — `onlyWithEmail`, `onlyVerifiedEmail`, `onlyWithWebsite`, `requirePhone`, `requireAnyContact` and the rest — runs before you are charged, MX verification is included in that price, and a run that finds nothing charges nothing. **A scheduled pricing record raises this to $0.005 per delivered lead on 2026-09-14** (notified 2026-08-30); the Pricing tab on this page is authoritative.

![Run report: leads with phone, verified email and website from a real run](https://api.apify.com/v2/key-value-stores/XtUUgXag7DIismvPm/records/local-business-leads-report.png)

This Actor takes a business category and a city and returns one row per local business, 73 stable columns wide and every row scored 0-100 — with an MX-verified email, whose mailbox that email is, the phone, the social profiles and the website platform wherever the business's own site exposes them (measured fill rates below). It costs **$3 per 1,000 delivered leads** ($0.003 per lead on the free plan; the Pricing tab always carries the current rate, and a change to $0.005 per lead is scheduled for 2026-09-14), and email verification is part of that price rather than an add-on.

Dentists, gyms, lawyers, plumbers, roofers, salons, real estate agencies, restaurants and 50+ more curated categories, plus any term you type. No API key, no proxy, no separate scraper per niche.

**Key facts:**

- **$3 per 1,000 delivered leads** ($0.003 per lead; rising to $0.005 per lead on 2026-09-14 under a scheduled pricing record) — MX email verification, mailbox-ownership classification and lead scoring are included in the single per-lead rate; filtered rows are dropped before billing, and a run that finds nothing charges nothing.
- **No API key, no proxy, no login** — discovery runs on OpenStreetMap and enrichment crawls each business's own public website, from datacenter IPs.
- **95 curated categories (220+ terms), any city on earth** — plus a radius search around a coordinate, or bring your own website list and skip discovery entirely.
- **73 stable columns on every row** — export as CSV, JSON or Excel, or connect the dataset to your CRM via API/webhook; column headers never shift mid-export.
- **Built-in monitoring and alerts** — `onlyNewBusinesses` turns a schedule into a new-business alert, and `webhookUrl` posts a Slack / Discord / JSON digest whenever a run delivers rows.

**What it does not do**, stated up front so nothing here is a surprise after you have paid:

- **It does not scrape Google Maps.** Listings come from OpenStreetMap, so there are no Google star ratings and no Google review counts. Pair it with our [Google Maps Places Scraper](https://apify.com/flash_scraper/google-maps-places) when you need those.
- **The `rating` and `review_count` columns are sparse.** They exist only when a business publishes a rating in its own website markup, measured at about **7% of rows** (4 of 55 on the 2026-08-08 reference run), and they are never a Google rating.
- **Van-based trades are thinly mapped.** OpenStreetMap held 171 dentists in the Austin bounding box but 9 plumbers and 5 electricians. The coverage section below names which categories are dense.
- **A guessed email is never sold as a verified one.** A pattern-guessed address stays in `email_guess` and is never promoted into `email`.

### Use from an AI agent

- **MCP:** point Claude, ChatGPT, Cursor or any MCP client at `https://mcp.apify.com?tools=flash_scraper/local-business-leads`; the tool is named after the Store slug and takes this actor's input unchanged. Keywords for the server's `search-actors` tool: *business leads*, *local business*, *lead generation*, *business emails*, *verified business emails*. Tool-name spellings, payment without an Apify token and measured timings: [Use it from an AI agent (MCP)](#use-it-from-an-ai-agent-mcp).
- **Smallest useful call** (Python `apify-client`; the same JSON works in the Console, the REST API and n8n/Make/Zapier):

```python
from apify_client import ApifyClient
client = ApifyClient("<APIFY_TOKEN>")
run = client.actor("flash_scraper/local-business-leads").call(run_input={"category": "dentist", "location": "Austin, Texas", "maxItems": 10, "crawlEmails": False})
rows = client.dataset(run["defaultDatasetId"]).list_items().items
```

- **Output contract:** the same 73 columns on every row, headers never shift mid-export; `email_guess` is never promoted into `email`. The full field list with measured fill rates is under [What data you get](#what-data-you-get); every run also writes a machine-readable `RUN_SUMMARY` record to its key-value store.

### What you get

- **Any category, any city** — 220+ category terms (95 curated categories plus their aliases) mapped to exact OpenStreetMap tags, any other term attempted directly and then rescued by a business-name search, and any city on earth geocoded (`Austin, Texas`, `casablanca morocco`, `Dubai UAE`).
- **Emails that are verified, not guessed** — every address is MX-checked over DNS-over-HTTPS and graded `deliverable` / `risky` / `undeliverable`. Verification is part of the price, not an upsell.
- **Whose mailbox it is** — `own_domain`, a free inbox, or the business's **marketing agency**. Emailing an agency mailbox never reaches the business, so this column decides whether a lead is worth a send.
- **Redesign pitch signals** — `website_platform` (WordPress, Wix, Squarespace, GoDaddy, Shopify, Webflow, Weebly, Duda and more), `mobile_viewport`, `copyright_year_stale`, `has_meta_pixel`, `has_google_ads_tag`. A builder-tier site with no tracking is the highest-intent pitch there is.
- **A score you can audit** — 0-100, an A-F grade, a hot/warm/cold tier, and a `score_breakdown` object showing the arithmetic that produced it.
- **Five ready-made Output views** — Overview, Email-ready, Web-agency targets, Map & source, All columns. Pick one above the results table instead of scrolling a wall of 73 raw columns.
- **Filters that cut your bill, not just the table** — every filter drops the row *before* it is pushed and *before* it is charged, and the run log names each filter and how many rows it removed.
- **Schedulable — only what is new** — `onlyNewBusinesses` turns a weekly schedule into a new-business alert: the first run is the baseline, every later run delivers only businesses that were not delivered before, and a run with nothing new delivers 0 rows and bills nothing. [How it works](#how-do-i-run-this-on-a-schedule-and-get-only-the-new-businesses).
- **One-click provenance** — `google_maps_url` (a Google Maps search for the business's name and address, so it opens the listing search rather than a bare map pin) and `osm_url` on every discovered row, plus the ODbL attribution the licence requires you to keep.

### Is there a Google Maps scraper alternative that does not scrape Google Maps?

Yes: this Actor discovers businesses on **OpenStreetMap** and then crawls each business's **own public website**, so no listing, rating or review is ever taken from Google Maps and no Google page is ever scraped. That is the whole design, not a setting you have to switch on.

On the 2026-08-08 reference run (dentist / Austin, Texas, onlyWithWebsite: true, n=55) that design delivered a phone on 96% of rows, an MX-verified email on 55% and a detected website platform on 71%, with no proxy input and no proxy cost, at $3 per 1,000 delivered leads (free-plan rate read 2026-08-29 and re-read 2026-09-05; a change to $0.005 per result is scheduled for 2026-09-14 — the Pricing tab is authoritative).

What that buys you:

- **No proxy bill and no proxy setup.** There is no proxy input on this Actor and no proxy cost in a run: OpenStreetMap's public endpoints and most business websites answer datacenter IPs, so runs go direct. A minority refuse them — measured at 8 of the 55 website-bearing rows in the reference run — and those rows say `site_blocked` rather than pretending the business is gone.
- **No Google Maps terms-of-service exposure.** What you receive is open-licensed map data plus pages the businesses publish themselves.
- **Fields a Maps listing does not carry.** MX-verified emails, mailbox ownership (`own_domain` / free inbox / the business's marketing agency), `website_platform`, `has_meta_pixel` and `mobile_viewport` all come from the business's own site, and the 0-100 lead score is computed from those signals together with the phone and website the map supplies.
- **A Google Maps link on every discovered row anyway.** `google_maps_url` is a constructed search link (`?api=1&query=<name>, <address>`), so you can open the real listing in one click. Nothing is read from Google to build it.

What you give up is equally concrete. There are no Google star ratings or review counts, and OpenStreetMap maps some trades thinly, which the coverage section below quantifies. If you need Google's own review data, run our [Google Maps Places Scraper](https://apify.com/flash_scraper/google-maps-places) alongside this one.

### Is this a small business leads scraper with MX verified emails?

Yes. It is a small business leads scraper for any category and any city: every row is a real local business found on OpenStreetMap, and every email is MX verified against the domain's mail servers before delivery, so what you get is verified email leads rather than guessed addresses. Local leads with emails, phones and social profiles arrive in one table, and rows removed by a filter are never billed.

### Measured, not promised

Every number below is from a real run of this Actor on `dentist` / `Austin, Texas`. Inputs and dates are given so you can reproduce them.

| Run | Result |
|---|---|
| Console form untouched, pressed **Save & start** | Today's untouched form is `dentist` / `Austin, Texas` at a cap of **25** (the form's starting value since 2026-08-29; it was 100 before) with **Website required**, **Skip closed** and **Widen to nearby areas** pre-ticked and **no keyword exclusions** — every delivered row lists a website. No count is quoted here on purpose: the 2026-08-08 measurement of this row was taken at cap 25 with a two-chain exclusion list (`Aspen Dental`, `Walmart`) the form pre-filled at the time and no longer does, so its figure cannot be reproduced from the form as it opens now. The next row is the same search with the same website filter at cap 100 (the API default) |
| Console form as it opened on 2026-08-08 (it then pre-filled `excludeKeywords: ["Aspen Dental", "Walmart"]`, which today's form does not), **Max businesses** raised to 100, pressed **Save & start** | **55 businesses, 100% contactable** — 96% with a phone, 55% with an MX-verified email. Under the 100 asked for, and the run says why: OpenStreetMap holds 171 `dentist` records in Austin and 55 of them are crawlable with a website inside the search area. That is the whole city, not a truncation |
| Bare `{}` from the API (no filters at all) — 2026-08-08 | 100 businesses, **50% contactable** — the honest filter-off case, see the fill-rate table below |
| **Cold-email list** preset — 2026-08-08 | 31 rows (every emailable dentist Austin has — the exact count varies by city and over time), **100% with an MX-verified email**, 94% with a phone |
| **Call list** preset, cap 100 — 2026-08-08 | **56 rows, 100% with a phone** — the same number on two independent runs, and the run reports it as the whole of Austin rather than a truncation |
| **Web-design prospects** preset, cap 100 — 2026-08-08 | **7 rows**, every one scored **≤45** *and* reachable on at least one channel. The score ceiling alone matches 68 businesses in Austin, but only 7 of those publish a phone, email or social profile — the preset drops the other 61 rather than bill you for map pins you cannot contact |
| **Full enrichment** preset — 2026-08-08 | 100 rows plus **23 `email_guess` addresses**, kept out of `email` and never billed as verified |
| Column stability across every run above | **one identical column tuple on every row** (61 columns on those runs; 73 on the current build — the 12 additions of 2026-08-29 are appended at the END, so existing CSV column positions are unchanged) — CSV headers never shift mid-export |

Every 2026-08-08 row that was started from the Console form (the preset rows included) ran with the two-chain `excludeKeywords` list the form pre-filled at the time and no longer does; the bare `{}` row did not. The counts stay as historical figures.

The gap between the Console-form rows and the bare `{}` row is the whole story of this Actor's honesty: the Console form pre-fills `onlyWithWebsite`, which is why it delivers 100% contactable businesses, while a bare API call keeps every mapped location including the ones with nothing to contact. Both numbers are published; neither is hidden. The same goes for the 55-of-100 line: a run that cannot reach your cap says so in its status message, with the counts that explain it.

**Max businesses starts at 25 in the Console form** — raise it once the first run looks right. The form also pre-ticks **`requireAnyContact`** (since 2026-08-29), so a first run never bills you for a business with no phone, no email and no social profile — untick it if you want every listing. At $3 per 1,000 that first run costs at most $0.075 in leads (at most $0.125 once the scheduled 2026-09-14 change to $5 per 1,000 lands) and finishes in a couple of minutes. The API default is unchanged at **100**: API calls, tasks and schedules that send no `maxItems` keep the ceiling they always had.

### Quick start

The shortest complete input is a business category and a location. Everything else has a working default:

```json
{
  "category": "dentist",
  "location": "Austin, Texas"
}
```

That is a complete run, and it is also exactly what the form already contains. Open the Actor, press **Save & start** without touching anything, and you get businesses. The defaults behind it: **Max businesses** starts at **25** in the form (raise it once the first run looks right; an API call that sends no `maxItems` gets 100), website crawling on, email verification on, 3 pages per website.

### The input form at a glance

Eight sections, in the order the form shows them. A first run only ever touches the first three; the rest are already tuned and safe to ignore.

| Section | What it is for |
|---|---|
| **Quick start** | One dropdown, `preset`, that fills in the rest of the form for one outreach job — see the presets below. **Custom** changes nothing |
| **What to search** | The business category and the city (dropdown or free text, or the `categories` / `locations` lists for several at once), an optional `countryCode` and `keyword` (matched against the name, and since 2026-08-29 the `cuisine` / `healthcare:speciality` tags too), **`maxItems`** — how many businesses this run may deliver, i.e. your cost ceiling; the form starts at 25, raise it once the first run looks right (API default 100) — and `expandNearby`, pre-ticked, which widens the search around the city when the city itself runs out (each widened row labelled in `query_location`) |
| **Filters — businesses are dropped BEFORE you are charged** | Website / email / phone / social floors, `excludeKeywords` (arrives empty — nothing is excluded until you type a keyword), `skipClosed` (pre-ticked), and the score and rating floors |
| **🔔 Monitoring** | `onlyNewBusinesses` — put the run on a schedule and get only the businesses that were not delivered before |
| **Enrichment** | What gets read from each business website: emails, social profiles, phones, optional pattern-guessed addresses, plus the `maxPagesPerSite` and `concurrency` crawl knobs |
| **Other ways to search** | A radius around a coordinate (`searchRadiusKm` + `centerLat` / `centerLon`) instead of a city, or bring your own list (`websiteList` / `startUrls`) and skip OpenStreetMap discovery entirely |
| **Output** | `outputFields` — which columns you get, and in what order — and `sortBy` |
| **🔔 Alerts** | `webhookUrl` — a Slack / Discord / JSON digest whenever a run delivers rows |

### Presets

The **Use-case preset** dropdown at the top configures the rest of the form for one specific outreach job. It is optional — leave it on **Custom** and nothing at all changes.

| Preset | What it sets | Measured on the default search (Austin dentists) |
|---|---|---|
| **Cold-email list** | `onlyWithEmail`, `onlyVerifiedEmail` | 31 leads (the whole of Austin; the count varies by city and over time), **100% with an MX-verified email**, 94% with a phone |
| **Web-design prospects** | `maxScore: 45`, `requireAnyContact` | 7 leads, all in the thin/neglected half — no published email, no marketing tech, often no mobile viewport — and all reachable. `maxScore: 45` on its own matches 68 businesses; the contact floor is what removes the 61 you could not pitch to |
| **Call list** | `requirePhone` | 56 leads, **100% with a phone number** |
| **Full enrichment** | `emailPatternGuess`, `maxPagesPerSite: 6` (every crawl toggle is already on by default) | 100 leads plus 23 guessed `email_guess` addresses, kept out of `email` |

**Anything you set yourself wins**, **no preset ever widens your bill**, and the run log names exactly what each preset applied and what it stood down on. The three guarantees are spelled out — and tested — under *Presets: the three guarantees* below.

### What the Output tab looks like

Six curated views ship with the Actor. Pick one above the results table; the CSV / JSON / Excel export is unaffected and always carries every column you asked for.

| View | Columns | Use it for |
|---|---|---|
| **Overview** *(default tab)* | `name`, `category`, `city`, `state`, `phone`, `email`, `email_status`, `website`, `rating`, `lead_grade`, `lead_score`, `google_maps_url` | The twelve columns that answer "is this a lead?" — plus one click to the live Google Maps listing |
| **Email-ready** | `name`, `email`, `email_status`, `email_type`, `phone`, `website`, `contact_page_url`, `lead_score` | Loading a cold-email sequence — deliverability grade and mailbox owner side by side |
| **Business details** | `name`, `category`, `speciality`, `cuisine`, `brand`, `is_chain`, `opening_hours`, `price_range`, `rating`, `review_count`, `wheelchair`, `operator`, `city` | What the business actually is: speciality / cuisine / brand tags, chain or independent, hours, price range, and the sparse rating / review count |
| **Web-agency targets** | `name`, `website`, `website_platform`, `mobile_viewport`, `copyright_year_stale`, `has_meta_pixel`, `has_google_ads_tag`, `phone`, `email`, `website_platform_status` | Building a redesign pitch list from the neglect signals |
| **Map & source** | `name`, `address`, `google_maps_url`, `osm_url`, `latitude`, `longitude`, `osm_last_edited`, `osm_check_date`, `attribution` | Verifying a row in one click, territory mapping, row freshness, keeping the ODbL attribution with the data |
| **All columns** | all 73, in CSV order | Everything, when you want the full table |

Three deliberate choices in those views. `website`, `contact_page_url`, `google_maps_url` (labelled **Map link**) and `osm_url` render as **clickable links**; `rating`, `review_count`, `lead_score`, `latitude` and `longitude` render as **numbers** (sortable). The tri-state flags — `mobile_viewport`, `copyright_year_stale`, `has_meta_pixel`, `has_google_ads_tag` — render as **text**, not as a checkbox, because they are `null` when the site was never crawled and a checkbox cannot tell "no mobile viewport" apart from "we never looked". `has_email` / `has_phone` / `has_website` / `is_chain` are never null, so those do get the real boolean widget.

### How do I find local businesses whose website is neglected enough to pitch a redesign?

Set `maxScore: 45` with `requireAnyContact` (the **Web-design prospects** preset), then read the neglect columns on the rows that come back. The score cap keeps both kinds of prospect — a business with a neglected site and a business with no site at all — so add `onlyWithWebsite: true` if you only want the ones that already have a site to replace. Four columns carry the pitch:

Measured 2026-08-08 on Austin dentists at cap 100: the Web-design prospects preset delivered 7 rows, every one scored 45 or below and reachable on at least one channel. The score ceiling alone matched 68 businesses; the contact floor dropped the other 61 — unbilled — rather than sell you map pins you cannot contact.

- `website_platform` — WordPress, Wix, Squarespace, GoDaddy, Shopify, Webflow, Weebly, Duda and more. A builder-tier platform usually means a self-built site.
- `mobile_viewport` — `false` means the homepage declares no mobile viewport tag, so the site predates responsive design.
- `copyright_year_stale` — a visibly out-of-date footer copyright year.
- `has_meta_pixel` — `false` means nobody is measuring anything on the site.

A builder-tier site with no mobile viewport, a stale copyright year and no tracking pixel is the strongest redesign signal this Actor can give you. The **Web-agency targets** output view shows exactly those columns and nothing else.

Read `email_type` before you send. It says whether the mailbox belongs to the business, to a free inbox, or to the marketing agency that already holds the account, so you can drop the leads where you would only be pitching a competitor.

### How do I find local businesses that have no website, for a web-design pitch list?

Type the trade and the city and tick **Businesses with no website** (`onlyWithoutWebsite: true`), and every delivered row is a business that lists no site at all. That is the first-website pitch list for web designers and local SEO agencies, verified at 15 of 15 rows without websites on a 15-row run.

Size the list honestly: on the filter-off benchmark of 2026-08-08 (n=100), 50 of the 52 businesses with no website had no phone, no email and no social profile either, so these rows are name, address and coordinates — which is also why nobody else can cold-email them.

Expect name, address, coordinates and occasionally a phone. With no site to crawl, no email, platform or tech enrichment is possible on these rows, which is also what keeps them uncrowded: nobody else can cold-email them either. Every row still carries `google_maps_url`, one click to a Google Maps search for the business's name and address.

The filter is mutually exclusive with `onlyWithWebsite`. Setting both stops the run with an explanatory status message before anything is charged, and nothing is billed. Like every other filter it runs before billing, and a thin city can be widened with `expandNearby`.

### What does it do?

It turns a business category and a city into one row per local business — 73 stable columns with an MX-verified email, whose mailbox it is, phone, socials, website platform and a 0-100 lead score — at $3 per 1,000 delivered leads. Measured 2026-08-08 on dentist / Austin, Texas with the website filter on: 55 businesses, 96% with a phone, 55% with an MX-verified email (free-plan rate read 2026-08-29 and re-read 2026-09-05; a change to $0.005 per result is scheduled for 2026-09-14 — the Pricing tab is authoritative).

This actor takes a **plain-English business category** (e.g. "dentist," "hair salon," "roofing contractor") and a **location**, maps the category to the right OpenStreetMap tags, and pulls every matching business in the area. It then **crawls each business's own website** — following the site's own contact/about/team links (including Shopify `/pages/contact` and non-English slugs) within the `maxPagesPerSite` budget — and extracts, from pages it has already downloaded:

- a **contact email**, MX-verified over DNS-over-HTTPS and graded `deliverable` / `risky` / `undeliverable`
- **whose mailbox it is** — the business's own domain, a free inbox, or its *marketing agency* (emailing that one never reaches the business)
- **social profiles** — Facebook, Instagram, LinkedIn, X/Twitter, YouTube
- **which website platform it runs on** — WordPress, Wix, Shopify, Squarespace, Webflow, GoDaddy, Weebly, Duda, Jimdo, Joomla, Drupal, HubSpot CMS, and a "Custom / Next.js" bucket for hand-built sites
- **marketing & booking tech** — Meta Pixel, Google Analytics/GTM, Calendly, NexHealth, Klaviyo, WooCommerce and ~25 more
- the homepage's **title and meta description** — drop-in mail-merge personalization fields (and a data tripwire: a title naming a different business exposes a wrong OSM `website` tag)
- **web-agency pitch signals** — a missing mobile viewport tag, a stale footer copyright year
- **one-click source links** — `google_maps_url`, a Google Maps search for the business's name and address, and `osm_url` to the OpenStreetMap element
- a **lead score 0-100**, an **A-F grade**, and a **hot/warm/cold tier**

### Why use it / who's it for

- **Web design & marketing agencies** — filter for a builder-tier platform (Wix, GoDaddy, Weebly) *and* `has_meta_pixel: false` to find businesses with a dated site and no tracking: the highest-intent redesign pitch there is. The `Custom / Next.js` value is the inverse signal — don't pitch a DIY-site rebuild to someone who already paid a developer. Build 0.1.21 adds two more neglect signals — `mobile_viewport` (missing = the site predates responsive design) and `copyright_year_stale` — plus an `onlyWithoutWebsite` filter that returns only businesses with **no site at all**: the first-website pitch list.
- **Freelancers on Fiverr/Upwork** — generate a "100 dentists in Austin with verified emails" list on demand for any client vertical without building a new scraper per niche.
- **SaaS sales teams** — any B2B tool sold to local businesses (booking software, payment processors, review management) can target by category and city, and `has_booking_widget` tells you who already has a competitor installed.
- **B2B lead-gen resellers** — one actor covers *any* category, replacing dozens of niche scrapers.
- **Franchise & market researchers** — count and map competitor density for a category in a target city (set `onlyWithWebsite: false` to keep every mapped location, including ones with no contact details).

### How to use it

The form is ready to run as it opens: hit **Save & start** without touching anything and you get Austin dentists that all list a website, at a cap of 25 (raise it once the first run looks right) with **Website required**, **Skip closed** and **Widen to nearby areas** pre-ticked and no keyword exclusions — at $3 per 1,000 that is at most $0.075 of leads (at most $0.125 after the scheduled 2026-09-14 change to $5 per 1,000). (The count that used to be quoted here was measured on 2026-08-08 at cap 25 with a chain-exclusion list the form pre-filled at the time and no longer does, so it is not repeated.)

Filtered runs used to come up short of the cap. The cause was on our side, not the map's: only the two email filters deepened the candidate pool, so any run using **Website required**, **Phone required** or **Social profile required** filtered a pool sized for an unfiltered run and quietly came up short. Every filter that can drop a row now deepens the pool, and the filters that can be decided from the map data alone (website present, name keywords, permanently closed) are applied *before* any site is crawled, so no crawl budget is spent on a row that was never going to ship. Same-day comparisons at cap 25 in Austin (Console form, 2026-08-08, which then pre-filled a two-chain exclusion list the form no longer carries) showed **Website required**, **Phone required** and **Email required** each filling the cap after the fix where they had come up short before; those counts are not repeated here because the form no longer reproduces them.

**Max businesses is still a ceiling, not a promise** — a thin area really can run out. When that happens the run now *says so* in its status message with the counts behind it, instead of returning fewer rows without comment. A live example, `dentist` in Laramie, Wyoming at cap 25: *"you asked for up to 25 and 1 could be delivered. That is the whole of Laramie, Albany County, Wyoming, United States, not a silent truncation: OpenStreetMap returned 4 'dentist' record(s) there, 1 became crawlable candidate(s), and 1 survived your filter(s) (onlyWithWebsite)."* The run only claims an area is exhausted when no search came back at its row ceiling **and** every candidate was actually checked; otherwise it says which limit it hit and what to change.

1. Pick a **Business category** from the dropdown, or leave it on **Custom** and type one in plain English (e.g. `medspa`, `HVAC contractor`, `funeral home`).
2. Pick a **City** from the dropdown, or leave it on **Custom** and type any city on earth — city and region/country (e.g. `Austin, Texas`). Capitalisation and the comma are optional: `RABAT MAROC` and `casablanca morocco` resolve too.
3. Leave **Website required** on unless you are doing density research — see the honest note below. (Web-design agencies can flip **Businesses with no website** instead to get the no-site prospect list; the two filters are mutually exclusive, and setting both stops the run with an explanatory status message (nothing is billed) with a clear error before anything is charged.)
4. Run the actor. It geocodes the location, pulls matching places from OpenStreetMap, then crawls each business website for contact details, tech and reviews.
5. Export as CSV, JSON, or Excel — or connect the dataset to your CRM/outreach tool via API/webhook.

#### Presets: the three guarantees

The preset table is at the top of this page. Behind it sit three guarantees, each of them tested:

- **Anything you set yourself wins.** A preset only fills in a field you left at its default. Set `maxScore: 100` alongside the web-design preset and you keep 100 — the run log says so explicitly (`left as you set them: maxScore (you chose 100)`).
- **A preset never widens your bill.** Every preset either adds a *filter* (fewer rows delivered, so fewer rows charged) or turns on enrichment that adds *columns* to rows you were already getting. None of them clears a filter you set or raises `maxItems`.
- **It tells you what it did.** One log line names the preset and lists exactly which settings it applied and which it left alone, and the run summary names the preset too.

Presets are Console *and* API: send `"preset": "cold_email"` from the API and you get the same behaviour. Existing API callers, tasks and schedules that send no `preset` are completely unaffected.

#### Which field wins: the precedence chain

Category and location each accept three inputs. Precedence is the same for both, highest first:

| Rank | Category | Location | Why |
|---|---|---|---|
| 1 | `categories` (list) | `locations` (list) | The plural list is an explicit multi-search request — it beats everything when non-empty. |
| 2 | `categorySelect` (dropdown) | `citySelect` (dropdown) | 20 categories with hand-verified OpenStreetMap tags; 25 metros verified against the live geocoder. |
| 3 | `category` (free text) | `location` (free text) | Anything else — over a hundred more category aliases are mapped, and any city on earth geocodes. |

Both dropdowns default to `""` ("Custom"), which is why every input that worked before this option existed still resolves to exactly the same search.

#### Three search modes

| Mode | Set | What happens |
|---|---|---|
| **City** (default) | `location` (or `locations`) + `category` (or `categories`) | The location is geocoded and every matching business inside its bounding box is returned. Several categories x several cities run as a matrix in one run. |
| **Radius** | `searchRadiusKm` + `centerLat` + `centerLon` | Searches a circle around a coordinate instead of a city box — sales territories, franchise catchment areas, "everything within 5 km of this address". Replaces the bounding box entirely; no geocoding happens, so `country` stays empty. |
| **Bring your own list** | `websiteList` (or `startUrls`) | OpenStreetMap discovery is **skipped entirely**. The actor runs only the crawl + email verification + scoring pipeline over the websites you supply. Enrich a CRM export, a conference exhibitor list, or a list you bought elsewhere. |

#### What happens when the city runs out of businesses?

Set **`expandNearby: true`** and the same search is widened in growing rings around the city centre until your `maxItems` is met or the region is genuinely exhausted. It is needed because a single city often holds fewer businesses than you asked for: measured live, **"dentist" in Austin, Texas tops out near 119 rows**, and Round Rock, Texas holds 21.

- The rings are sized to the city and reach up to about 150 km.
- **Every widened row is labelled**: its `query_location` reads `within ~16 km of Round Rock, Williamson County, Texas, United States` instead of the city name, so you can always tell expansion rows apart — or filter them out afterwards.
- The same dedup and every pre-charge filter apply to ring rows, and the status message reports exactly how many delivered rows came from outside the city.
- Measured (2026-08-15, discovery-only): `dentist` / `Round Rock, Texas` / `maxItems: 100` delivered **21 rows without the flag, 100 with it** — 79 labelled ring rows, 0 duplicates.

It is **off by default for API callers** (existing inputs keep a byte-identical pull and bill) and **pre-ticked in the Console form**. It never fires when you set your own radius (`searchRadiusKm`) or bring your own list, when the city pull was truncated (deepening, not widening, is the fix there — the status message tells you), or during a mirror outage.

#### Searching several categories and cities at once

`categories` and `locations` are the plural versions of `category` and `location`, and they cross into a matrix — `["dentist","orthodontist"] x ["Austin, Texas","Dallas, Texas"]` is four searches in one run. The single fields keep working exactly as before; the plural ones take precedence when non-empty.

Three things make the matrix safe rather than a footgun:

- **`maxItems` is the whole run's budget**, not a per-search one. The budget is split between the searches and results are **interleaved**, so the first city cannot eat the entire quota.
- **The same business found by two categories is delivered — and billed — once.** Deduplication is on the OpenStreetMap object identity (`osm_type` + `osm_id`), so a clinic tagged both `dentist` and `orthodontist` appears one time.
- **25 category x location combinations is the ceiling** for one run. Above it the run fails immediately with a message naming the numbers, before any network request and before any charge — OpenStreetMap's public mirrors are a free shared resource.

Every row carries `query_category` and `query_location`, so you always know which search produced it.

#### Bring your own list (skip discovery)

Already have the businesses and only need the emails, socials, platform and score? Put the domains in `websiteList` (bare domains and full URLs both work; `startUrls` is accepted as an alias):

```json
{ "websiteList": ["aloha-dental.com", "https://www.averyranchdental.com/", "typotes.com"] }
```

No map data is fetched at all. Those rows differ from discovered rows in exactly three honest ways:

- `source` is `user_supplied`, not `OpenStreetMap`;
- `attribution` is **null** — the rows are not OSM-derived, so stamping the ODbL notice on them would be a false licence claim. The *column* is still present, so a mixed export keeps one stable header row;
- `name` comes from the site's own `<title>`, falling back to the bare domain when the site does not answer. Nothing is invented; `latitude`, `longitude`, `osm_id` and `osm_url` stay empty.

Everything else — the contact crawl, MX verification, mailbox-owner classification, tech fingerprinting, scoring and every filter — behaves identically.

### How complete is the data? (measured, not estimated)

On the reference run of `dentist` in `Austin, Texas` with `onlyWithWebsite: true` (n=55), **96% of rows carried a phone, 55% an MX-verified email, 71% a detected website platform and 100% a website**. With every filter off (n=100) the same city measured 48% phone, 26% email and 34% platform, because roughly half of mapped businesses list no website to crawl.

Both reference runs were taken on 2026-08-08 (the website-filtered run and the bare API run in the Measured, not promised table above), and a roofing contractor / Denver run delivered emails on 6 of 8 rows (75%) — roughly three-quarters on website-verified trade categories.

Both reference runs are in the table below, so you can see the gaps before you pay rather than after:

| | `onlyWithWebsite: true` (n=55) | schema defaults, filter **off** (n=100) |
|---|---|---|
| `phone` | **96%** | 48% |
| `website` | **100%** | 48% |
| `email` | **55%** | 26% |
| `website_platform` | **71%** | 34% |
| `opening_hours` | **80%** | 44% |
| `facebook` | 73% | 33% |
| rows graded F | 0% | **51%** |

On that same filter-off n=100 run, the crawl-derived fields measured: `contact_page_url` 18%, `website_title` 40%, `website_description` 35%, `mobile_viewport` 40%, `copyright_year_stale` flagged on 3 rows; `google_maps_url`, `osm_url` and `attribution` sat at 100%.

With `onlyWithWebsite: true` the email fill runs far higher than the filter-off numbers — the reference run above measured 55% for dentists, and a `roofing contractor` / Denver run delivered emails on **6 of 8 rows (75%)**: roughly three-quarters on website-verified trade categories.

**Read that second column before you run.** OpenStreetMap has no website for a large share of businesses, and — measured on the filter-off run — **50 of the 52 rows with no website had no phone, no email and no social profile either** (the other 2 carried only an OSM phone): name, coordinates and usually an address, nothing contactable. With the filter off you pay for those rows. `onlyWithWebsite` and `onlyWithEmail` drop non-matching rows **before** you are charged, so they cut your bill rather than just tidying the output. Every run's status message now reports the contactable ratio it actually delivered.

**Every** filter now goes further: the actor **over-fetches**, deepening the candidate pool and — for filters that need the site crawled — crawling extra candidates in batches (up to 24× `maxItems`, hard-capped at 30,000 candidates — the ceiling that makes a thin area terminate) until it has `maxItems` surviving rows or the pool is spent. This used to apply to `onlyWithEmail` / `onlyVerifiedEmail` only, which is why `onlyWithWebsite`, `requirePhone` and `requireSocial` quietly returned short. Compared 2026-08-08 on `dentist` / `Austin, Texas` at cap 25, one filter at a time, through a Console form that then pre-filled a two-chain exclusion list it no longer carries: `onlyWithWebsite`, `requirePhone` and `onlyWithEmail` each filled the cap after the change where each had come up short before (the exact counts are not repeated because today's form cannot reproduce them). Where the pool genuinely runs out first, the status message reports the counts instead of leaving you to guess — e.g. `dentist` in Laramie, Wyoming returns 1 row and says the map holds 4 dentist records there, 1 of them with a website.

The default is `false` (not `true`) so that existing API callers, scheduled tasks and density-research use cases keep getting every mapped location. The Apify console pre-fills it to `true`.

There is also the mirror filter, **`onlyWithoutWebsite`**: keep only businesses that list **no** website — the prospect list for web-design agencies pitching a first site (verified: a 15-row run delivered 15/15 rows without websites). It is mutually exclusive with `onlyWithWebsite`; setting both stops the run with an explanatory status message (nothing is billed) with a clear error before anything is charged. Expect these rows to be name + address + coordinates (and occasionally a phone) — with no site to crawl, no email/platform enrichment is possible.

#### Which business categories does OpenStreetMap cover well, and which are sparse?

OpenStreetMap covers businesses with premises a mapper walks past, and is thin on trades run from a van or a home office: the Austin bounding box held **171 dentist records but 9 plumbers, 5 electricians, 0 chiropractors and about 12 roofers**. This is the honest limit of an OSM-based source, and it matters more than any field:

Two dated counts from the same city: 38 of the 60 pharmacies OpenStreetMap returned for Austin, Texas were brand-tagged chain outlets (measured 2026-08-29), and on a 163-element sample of named Austin dentists taken the same day, 12% had not been edited since before 2020 and 13% carried a mapper's on-the-ground check\_date.

- **Dense**: businesses with premises a mapper walks past — restaurants, cafés, dentists, pharmacies, hairdressers, gyms, hotels, shops, banks, clinics. (171 dentist records in the Austin bounding box.)
- **Sparse**: trades run from a van or a home office — plumbers (9 in the same box), electricians (5), chiropractors (0), roofers (~12). A metro of a million people can return single digits. That is what is mapped, not a bug.

95 categories (220+ terms with aliases) are curated and mapped to exact OSM tags — build 0.1.21 added **pest control, photographer, moving company, self storage, funeral home and dry cleaner**. Any other term is attempted as an OSM tag directly, and common phrasings are handled (`auto repair shop` → `shop=car_repair`, `landscaping company` → `craft=gardener`, `insurance agency` → `office=insurance`). Terms with no OSM tag now fall back to a **business-name keyword search** of OSM, and a mapped category that returns zero tagged places in the area is rescued by the same name-keyword query. Rows found only by name are labelled `category_match: "name_keyword"` (tag-matched rows say `category`) so you can filter them out if you only trust tag-confirmed rows. Measured: `pest control` in Denver returned 1 labelled row on 0.1.21 where the previous build returned 0 — sparse trades stay sparse, but no longer invisible. A term matching nothing at all still returns **0 rows with a suggestion and no charge** — you are never billed for a run that found nothing.

#### Can I use this as a business email finder for local businesses?

Yes. For any category and city it crawls each business's own website and returns the MX-verified business email wherever the site exposes one, plus the mailbox it belongs to (info@, owner name, etc.). The fill rate is measured, not promised — see "How complete is the data?" above.

Measured 2026-08-08 on dentist / Austin, Texas with onlyWithWebsite: true: 55% of 55 delivered rows carried an MX-verified email, and the Cold-email list preset (onlyWithEmail plus onlyVerifiedEmail) delivered 31 leads — the whole of Austin — 100% with an MX-verified email and 94% with a phone, at $3 per 1,000 with verification included (free-plan rate read 2026-08-29 and re-read 2026-09-05; a change to $0.005 per result is scheduled for 2026-09-14 — the Pricing tab is authoritative).

### Local business contacts, B2B email finder and local SEO leads

The output is a **local business contacts** table: one row per business with an MX-verified email, phone, socials and the site's platform. Used as a **B2B email finder** it answers the question "who do I email at businesses of this type in this city", and because every row carries the website, the platform and whether the copyright year is stale, it doubles as a **local SEO leads** list for agencies pitching redesigns. As a **company email scraper** it reads each business's own site rather than a purchased database, so the addresses are current and checkable. Good for **SMB leads** in any category OpenStreetMap covers.

### Filters — every one of them runs before you are charged

Filters here are not a tidying step applied to an invoice you have already run up: a filtered row is dropped **before** it is pushed to the dataset and before the charge, so filtering cuts your bill.

**Filtering does not shrink your delivery.** Every filter in this table deepens the candidate pool to compensate, so `maxItems` means "this many rows I can use", not "this many businesses considered". Filters decidable from the map data alone — `onlyWithWebsite`, `onlyWithoutWebsite`, `excludeKeywords`, `excludeChains`, `skipClosed` — are applied before any website is crawled, so no crawl budget is spent on a row that was never going to ship. The rest are applied batch by batch as sites are crawled, stopping the moment enough rows survive. The pool is bounded at 6x `maxItems` candidates so a thin area terminates instead of crawling forever; when that bound or the map itself is what stopped the run, the status message says so and gives you the counts.

| Filter | Keeps only |
|---|---|
| `onlyWithWebsite` | businesses that list a website (**recommended**, see the fill-rate table) |
| `onlyWithoutWebsite` | businesses with **no** site — the first-website pitch list. Mutually exclusive with the above |
| `onlyWithEmail` | rows where an email was found |
| `onlyVerifiedEmail` | rows whose email passed MX verification (`deliverable` / `risky`) |
| `requirePhone` | **New.** rows with a phone number in `phone` (the OSM number) **or** in `phones` (the `tel:` numbers the site publishes) — either one passes |
| `requireSocial` | **New.** rows with at least one Facebook / Instagram / LinkedIn / X / YouTube profile |
| `requireAnyContact` | **New.** rows reachable on **at least one** channel — phone *or* email *or* a social profile. The loosest contactability floor there is, and the one to reach for when you do not care *which* channel. It matters more than it sounds: on a bare `{}` run of the default search (100 rows, measured 2026-08-08) exactly **50 rows carried no phone, no email and no social profile at all**, and without this switch you are billed for them. Off by default for API calls and schedules (no existing run changes); **pre-ticked on the Console form since 2026-08-29** |
| `excludeKeywords` | drops businesses whose **name** contains any of your keywords (case-insensitive). The form arrives with the list empty. Only the name is matched, so `clinic` cannot knock out a business on Clinic Street. Use it to strip chains, franchises or your existing customers |
| `excludeChains` | **New (2026-08-29).** drops every business the map marks as a chain outlet — a `brand` or `brand:wikidata` tag, i.e. `is_chain: true` — leaving the independents. Decided from the map data alone, so a dropped outlet is never crawled or billed, and the status line says how many were removed. Off by default; rows you supply yourself have `is_chain: null` and are never dropped by it. In chain-dense categories the 6x pool bound can bite — 38 of the 60 pharmacies OpenStreetMap returned for Austin, Texas were brand-tagged (measured 2026-08-29), so a `maxItems: 10` run delivered 6 and said so; raise `maxItems` to deepen the pool |
| `skipClosed` | **New.** drops permanently-closed premises (see below) |
| `minScore` / `maxScore` | **`maxScore` is new.** A score *ceiling* is the web-design agency filter: a low score means a thin online presence, which is exactly the redesign pitch list |
| `minRating` / `minReviewCount` | **New — read the warning below before using these** |

#### `skipClosed`: what it actually removes

OpenStreetMap mappers retire a business without deleting it: the primary tag moves from `amenity=restaurant` to `disused:amenity=restaurant`, so the object keeps its name and address but no longer describes an operating business. `skipClosed` drops elements carrying a lifecycle-prefixed primary tag (`disused:`, `abandoned:`, `was:`, `removed:`, `demolished:`, `razed:`), a `disused=yes` / `abandoned=yes` flag, `opening_hours=closed`, or `shop=vacant`. A `disused:` tag alongside a *live* tag of the same kind (a former bank that is now a café) is **not** treated as closed.

Measured live on 2026-08-08 in the Austin bounding box: **199 such elements, 43 of them still carrying a business name** — e.g. `Chago's` with `disused:amenity=restaurant`, `Corner Store` with `abandoned:shop=fuel`. These cannot reach a normal tagged search (`["amenity"="restaurant"]` cannot match `disused:amenity`), but they **do** reach the business-name fallback used for unmapped categories, which is where dead businesses were being delivered as fresh leads. Verified end to end: with `skipClosed: false` that closed restaurant is delivered and billed; with it `true` it is dropped before billing.

It is **off by default** so existing runs, tasks and API callers are unchanged. The Apify console pre-fills it to `true`.

> **Honest warning about `minRating` / `minReviewCount`:** OpenStreetMap carries **no review data at all**. A rating only exists when the business publishes schema.org `aggregateRating` on its own website — measured at roughly **3-8% of rows**. So when you set a rating or review floor, **rows with no rating are DROPPED, not kept**: an unknown rating is not a passing rating, and nothing is ever invented to save a row. Setting either filter will cut your result count to a small fraction — an Austin restaurant run asking for 15 rows with `minRating: 4.0` crawled **86 candidates** looking for them and delivered **0 rows, charging nothing** (re-measured 2026-08-08 after the over-fetch was generalised, so this is the deep-search result, not a shallow one). If you need ratings on every business this is the wrong source; pair it with our [Google Maps Places Scraper](https://apify.com/flash_scraper/google-maps-places), or use the `google_maps_url` on every row.

### How do I run this on a schedule and get only the new businesses?

Create an Apify schedule for the run and set **`onlyNewBusinesses: true`**: the first run is the baseline, every later run delivers only businesses that were not delivered before, and a run with nothing new delivers 0 rows and bills nothing. Point a weekly schedule at "dentists in Austin" and you get the practices that appeared since the last run, and only those.

Monitoring was hardened on 2026-08-25: the memory lives in a named key-value store in your own account, entries are pruned after 90 days with the stamp refreshed on every run, each record is capped at 50,000 keys, and a schedule created before 2026-08-29 keeps its memory because excludeChains joins the key only when you set it.

**How it behaves**

| Run | What happens |
|---|---|
| First run of a watch | **Baseline.** Everything the search finds is delivered, and the status message says so in those words. |
| Later runs | Only businesses that were not delivered before. Already-delivered ones are dropped **before the crawl**, so they cost nothing and are never billed. |
| Nothing is new | **0 rows, SUCCEEDED, nothing billed.** The status message reads `Nothing new: all N business(es) this search found were already delivered by earlier runs of this watch ... You were not charged.` |

**What counts as the same business.** The memory key is the business's **OpenStreetMap object identity** (`node/2135639605`, `way/198109528`) — the same identity the run already uses to dedupe in-run, so a business that renames itself or changes domain does not come back as "new". Bring-your-own-list rows have no OSM object, so they are remembered by their registered domain — the same key `websiteList` is already deduped on.

**A row is remembered only after it has actually been delivered.** Marking happens *after* `push_data` succeeds, never at the point the candidate is chosen. That ordering is the whole safety property: every filter, the `maxItems` trim and above all the `maxTotalChargeUsd` budget trim can still remove a row between the two points, and a row marked early would be recorded as delivered, skipped by every future run, and never reach you.

**What identifies a watch.** Your search plus the filters that shape which businesses can survive it:

- `categories` / `category`, `locations` / `location`, `countryCode`, `keyword`
- `searchRadiusKm`, `centerLat`, `centerLon`, `expandNearby`
- `websiteList` / `startUrls`
- every option in the **Filters** section (`onlyWithWebsite`, `onlyWithoutWebsite`, `onlyWithEmail`, `verifyEmails`, `onlyVerifiedEmail`, `requirePhone`, `requireSocial`, `requireAnyContact`, `skipClosed`, `excludeKeywords`, `minScore`, `maxScore`, `minRating`, `minReviewCount`) — and `excludeChains`, which joins the key **only when you set it**, so a schedule created before 2026-08-29 keeps the memory it already has

Change any of those and you are running a **different watch**, with its own independent memory — the two never share state, and the new one starts from its own baseline. Category and location lists are compared case-insensitively and order-independently, so `["gym","dentist"]` and `["Dentist","GYM"]` are one watch, not two.

**Deliberately *not* part of a watch:** `maxItems`, `sortBy`, `outputFields`, `concurrency`, `maxPagesPerSite`. Those change how many rows you get, in what order and with which columns — not *which businesses exist* in the search — so tuning them never resets the memory and never re-bills you for businesses you already have.

**Where the memory lives.** A named key-value store called `local-business-leads-monitor` **in your own Apify account** (runs execute there), one record per watch, keyed `sig-<hash>`. Entries older than **90 days** are pruned on every write and each record is capped at 50,000 keys, so a long-running schedule cannot grow it without bound. The timestamp stored is *last* seen, not first, and it is refreshed on every run for businesses still returned by the search — so the 90-day prune drops businesses that have genuinely disappeared from OpenStreetMap rather than businesses that have simply been mapped for a long time (those would otherwise age out and be re-delivered, and re-billed, as "new"). If the store cannot be read the run treats itself as a first run and says so; monitoring never fails a run.

**Two honest caveats.**

1. **Runs take longer with it on.** Already-delivered businesses have to be skipped over, so the run searches a deeper candidate pool to still fill your `maxItems` with genuinely new rows.
2. **New businesses appear at the speed of OpenStreetMap.** This is a mapping database, not a live business registry — a new dentist shows up here when a mapper adds it. A weekly or monthly schedule matches that cadence; an hourly one will mostly report "Nothing new" and bill you nothing for the privilege.

```json
{
  "category": "dentist",
  "location": "Austin, Texas",
  "maxItems": 200,
  "onlyWithWebsite": true,
  "onlyNewBusinesses": true
}
```

### How do I get a Slack or Discord alert when new leads land?

Put your Slack or Discord incoming-webhook URL in **`webhookUrl`**, and every run that delivers at least one row POSTs a digest to it; a run that delivers nothing sends nothing.

Slack / Discord / JSON digests shipped on 2026-08-25; the JSON payload carries the delivered count, the run and dataset links and the first 20 rows, and a webhook that fails is named in the run's status message without failing the run.

A Slack incoming webhook and a Discord webhook each get a text message. Any other URL, such as an n8n, Make or Zapier catch hook or your own endpoint, gets JSON: `{actor, delivered, run_url, dataset_url, rows[:20], text}`.

Pair it with `onlyNewBusinesses` on a schedule and the Actor is a new-business alert service on its own, with no extra automation needed just to see the rows. Delivery is best-effort: a webhook that fails is reported in the run's status message and never fails the run.

```json
{
  "category": "dentist",
  "location": "Austin, Texas",
  "onlyNewBusinesses": true,
  "webhookUrl": "https://hooks.slack.com/services/T000/B000/XXXX"
}
```

### Every run comes with a report

Every run that delivers at least one row also stores a one-page HTML **report**: the headline numbers (businesses, % with email, % verified, % with website, average lead score), the A–F lead-grade split, the top categories or cities, the email-status split and the first 100 rows, plus the same notes the status message carries. It is saved as the `REPORT` record of the run's key-value store — open it from the run's **Output** tab (record `REPORT`) or through the `Report:` link in the run's status message. It is a single self-contained HTML file (no scripts, no external assets), so it is safe to forward, attach to an email or screenshot for a client. The dataset stays the source of truth: the report summarises what was delivered and never replaces the rows, and a run that delivers nothing writes no report.

### Choosing your columns and sort order

- **`outputFields`** — pick the columns you want (e.g. `["name","email","phone","website","lead_grade"]`) and the dataset carries only those. Every row still shares one identical column tuple, so CSV headers never shift mid-export, and the columns keep the documented ROW order regardless of the order you listed them in. `name` and `attribution` are **always** included whatever you choose — `attribution` because the OpenStreetMap licence (ODbL) has to travel with the data into your CSV. Column names are matched case-insensitively (and `-`/space count as `_`, so `Lead Grade` works). A single unknown name is logged and ignored; if **none** of the names you list exists, the run stops before billing rather than charging you full price for a name-only export.
- **`sortBy`** — `score_desc` (default, what every previous build did), `name_asc`, `review_count_desc` or `rating_desc`. Sorting never removes a row; all filtering already happened. Rows missing the sort value (no rating, no review count) are placed **last**, because a missing value is unknown rather than zero.

### Output fields

73 columns on every row (61 before 2026-08-29; the twelve new ones — `has_google_ads_tag`, `speciality`, `cuisine`, `wheelchair`, `operator`, `osm_description`, `brand`, `brand_wikidata`, `is_chain`, `city_source`, `osm_last_edited`, `osm_check_date` — are appended at the END of the column order, so an existing CSV import keeps its positions). Re-measured on a 100-row `dentist` / `Austin, Texas` run of the previous build with schema defaults: **all 100 rows carried the same stable column tuple**, so CSV headers don't shift mid-export — and if you narrow the export with `outputFields`, every row still shares one identical (smaller) tuple. Fill rates are from the `onlyWithWebsite: true` reference run above; anything conditional says so. Fields marked **New (0.1.21)** show fill rates from the 100-row filter-off benchmark instead — read them accordingly, since roughly half of those rows had no website to crawl.

In the Console, the **Output** tab opens on the Overview view; the **All columns** tab shows every field listed below, in this order. Views only change what the Console table renders — a CSV, JSON or Excel export always carries every column the run produced (or exactly the ones you named in `outputFields`).

#### Identity & location — from OpenStreetMap

| Field | Fill | Description |
|---|---|---|
| `name` | 100% | Business name |
| `category` | 100% | OSM category tag (e.g. dentist, hairdresser, lawyer) |
| `category_match` | 100% | **New (0.1.21).** How the row matched your category: `category` = matched the mapped OpenStreetMap tag; `name_keyword` = found by the business-name fallback search (used for unmapped terms, and as a rescue when a mapped tag returns zero places in the area). Filter on `category` if you only want tag-confirmed rows |
| `address` | 96% | Street address assembled from OSM address tags |
| `city` / `state` / `postal_code` | 93% / 93% / 93% | Address components from the OSM `addr:*` tags. Since 2026-08-29 a row that sits **inside** the searched city's outline but carries no `addr:city` tag gets `city` (and a null `state`) filled from the geocoder's breakdown of the searched location — zero extra requests, and `city_source` says which happened. For US locations the filled `state` is the postal abbreviation (`TX`, not `Texas`), the same form mappers write in `addr:state`, so one run never splits a state into two spellings. Rows admitted by `expandNearby` rings and radius searches (no outline) keep their nulls |
| `city_source` | 100% when `city` is set | **New (2026-08-29).** `osm_tag` = the mapper wrote `addr:city`; `search_area` = filled from the searched location because the row is inside its administrative outline; null when `city` is null. Appended at the end of the column order |
| `country` | 100% | From the geocode. Empty in radius mode and on user-supplied rows (no geocode happens there) |
| `query_category` | 100% | **New.** Which of your input categories produced this row — the column that makes a multi-category run readable. Null on user-supplied rows |
| `query_location` | 100% | **New.** Which of your input locations produced this row (the geocoder's resolved name, or the radius description). Null on user-supplied rows |
| `latitude` / `longitude` | 100% | Coordinates |
| `osm_type` / `osm_id` | 100% | OpenStreetMap source identifiers |
| `google_maps_url` | 100% | A Google Maps **listing search for the business** — `?api=1&query=<name>, <address>`, URL-encoded — so one click opens Google's search for that business rather than a bare coordinate pin (a discovered row with coordinates but no address falls back to the pin). The link carries none of this actor's data and nothing is scraped from Google: whatever rating you see there is Google's, not the `rating` column (see the warning below). Need Google's fields at scale? Pair with our [Google Maps Places Scraper](https://apify.com/flash_scraper/google-maps-places) |
| `osm_url` | 100% | **New (0.1.21).** Link to the row's source OpenStreetMap element — instant provenance, and the place to fix bad map data |
| `speciality` | varies by category | **New (2026-08-29).** The OSM `healthcare:speciality` tag as a list (`["orthodontics"]`, `["general", "paediatric"]`; first 10 values) — dentists, doctors and clinics; empty for every other trade. The `keyword` input matches it too |
| `cuisine` | varies by category | **New (2026-08-29).** The OSM `cuisine` tag as a list (`["pizza", "italian"]`; first 10 values) — restaurants, cafes, fast food; empty elsewhere. The `keyword` input matches it too |
| `wheelchair` | varies by category | **New (2026-08-29).** The OSM `wheelchair` tag as written (`yes`, `no`, `limited`) |
| `operator` | varies by category | **New (2026-08-29).** The OSM `operator` tag — the company running the premises, where the mapper recorded one |
| `osm_description` | varies by category | **New (2026-08-29).** The mapper's free-text `description` tag, clamped to 150 characters |
| `brand` / `brand_wikidata` | varies by category | **New (2026-08-29).** The OSM `brand` and `brand:wikidata` tags (`Aspen Dental` / `Q4807808`) — dense on pharmacies, banks and fast food, sparse on independents |
| `is_chain` | 100% on OSM rows | **New (2026-08-29).** `true` when either brand tag is present, `false` otherwise — the one-column chain flag; set `excludeChains: true` to drop those rows before billing (`excludeKeywords` still works by name). Null on user-supplied rows |

#### Contact

| Field | Fill | Description |
|---|---|---|
| `phone` | 96% | Primary phone — the OSM tag, and only the OSM tag; never filled from the site (see the licence note below) |
| `phones` | 73% | **New.** All `tel:` numbers found on the site. May include a call-tracking number — `phone` stays the trusted value. `requirePhone`, `requireAnyContact`, `has_phone` and the 20 phone points count a row that has either `phone` or `phones` |
| `website` | 100% | Business website URL. A social page in OSM's `website` tag is routed to that social column instead, so this is always a real site |
| `domain` | 100% | **New.** Bare registered domain — the field CRMs dedup on |
| `email` | 55% | Primary contact email, chosen by mailbox ownership then deliverability |
| `emails` | 55% | **All** emails for the row, primary first (previously excluded a primary that came from OSM) |
| `contact_page_url` | 42% | **New.** The exact page the primary email was found on, so you can spot-check it. Null when the email came from OpenStreetMap rather than from a crawled page |
| `email_guess` | 0% unless opted in | **New, opt-in, and deliberately not an email.** With `emailPatternGuess: true`, a business that has a website but publishes no address anywhere we crawled gets a **pattern-guessed** `info@<domain>` here. It is never promoted into `email`, never counts as `has_email`, never earns a lead-score point and never satisfies `onlyWithEmail` / `onlyVerifiedEmail`. Nobody checked that this mailbox exists — treat it as a lead, not an address |
| `email_guess_confidence` | same | **New.** A statement about the *domain*, never the mailbox: `mx_ok` (the domain does run mail servers), `no_mx` (it does not — the guess is almost certainly dead), `unchecked` (no lookup completed — email verification is off, the DNS lookup itself failed, or the run's time budget stopped it) |

#### Email quality — verification is included, not an add-on

| Field | Fill | Description |
|---|---|---|
| `email_status` | 100% | `deliverable` / `risky` / `undeliverable` when verification is on and a candidate exists; `missing` when no email was found; `found` / `missing` when `verifyEmails` is off. **Turning verification off changes the vocabulary** |
| `email_provider` | 55% | Mailbox host where identifiable (Google Workspace, Microsoft 365...). Only when `verifyEmails` is on and an email was found |
| `email_domain_match` | 55% | Whether the email's domain matches the website's |
| `email_type` | 55% | `own_domain` / `free_mail` / `third_party` / `unknown`. **`third_party` means the address belongs to the business's marketing agency or web designer** — it is deliverable but does not reach the business, and it is scored at half weight. Measured on the reference run: 12% of all harvested addresses were third-party, but only 10% of *primary* emails, because candidates are ranked by mailbox ownership before one is promoted |
| `email_types` | 55% | The same classification for every entry in `emails` |

#### Socials

| Field | Fill | Description |
|---|---|---|
| `facebook` | 73% | Facebook profile URL |
| `instagram` | 55% | Instagram profile URL |
| `youtube` | 25% | **New.** YouTube channel URL |
| `linkedin` | 22% | **New.** LinkedIn company/profile URL |
| `twitter` | 18% | **New.** X/Twitter profile URL |

Share buttons, tracking pixels and embedded posts are filtered out, so these are profile URLs rather than `facebook.com/tr` or `instagram.com/p/...`.

#### Website & marketing tech

| Field | Fill | Description |
|---|---|---|
| `website_platform` | 71% | CMS / site builder: WordPress, Wix, Shopify, Squarespace, Webflow, GoDaddy, Weebly, Duda, Jimdo, Site123, Framer, HubSpot CMS, BigCommerce, Joomla, Drupal, `Custom / Next.js`, `Custom / React`. WordPress plugins (Elementor, WP Rocket, Divi) report as **WordPress**, not as their own platform |
| `website_platform_status` | 100% | **New.** Why `website_platform` is what it is — and what actually happened to the site: `detected`, `unknown` (a page loaded, nothing recognisable on it), `site_blocked` (the site answered but refused us — a 403, a bot-check page or a login wall), `not_found` (every URL that answered came back 404 or 410 — the address in the listing is dead, the host serving it is not), `site_error` (the site answered with a server error), `no_page` (it answered, but never with a readable web page), `site_unreachable` (nothing answered at all — DNS failure, refused connection, broken TLS or a timeout), `no_website`, `not_crawled` (crawling was off), `crawl_error` (our own crawler failed on that site — worth reporting). Only `site_unreachable` means the business has no reachable website, and only `crawl_error` says nothing about the site; every other value tells you the site exists and why we could not read it, so a blocked or 404 site is still a live prospect. A null platform is explained rather than unexplained |
| `platform_version` | 27% | **New.** Only when the site's own `meta generator` names the platform *and* a version — never a plugin's version passed off as the platform's |
| `site_generator` | 42% | **New.** The raw `meta name="generator"` string |
| `tech` | 84% | **New.** Marketing/booking/commerce tech found on the page (Meta Pixel, Google Analytics, GTM, Calendly, NexHealth, Klaviyo, WooCommerce, Yelp Reviews, live chat...) |
| `has_meta_pixel` / `has_google_analytics` / `has_booking_widget` / `has_google_ads_tag` | see note | **New.** Booleans derived from `tech`. **Tri-state**: `true`/`false` when the site was read, `null` whenever no page was read — unreachable, blocked, a bot-check interstitial, or not crawled at all. `false` never means "we could not check", so a no-pixel filter cannot fill up with sites nobody ever saw. `has_google_ads_tag` (**new 2026-08-29**, appended at the end of the column order) is `true` when the Google Ads conversion tag (`googleadservices.com` / `gtag` `AW-…`) is on the page — a wiring signal, not proof of live spend (see the FAQ) |
| `website_title` | 40% | **New (0.1.21).** The crawled homepage's `<title>` — a drop-in mail-merge personalization field. Also a data tripwire: a title that clearly names a different business means OSM's `website` tag is wrong, and this column lets you catch that before you hit send |
| `website_description` | 35% | **New (0.1.21).** The homepage's meta description — the other personalization field, and the same tripwire |
| `mobile_viewport` | 40% | **New (0.1.21).** Whether the homepage declares a mobile viewport tag. A site without one predates responsive design — a concrete web-agency pitch signal |
| `copyright_year_stale` | 3% | **New (0.1.21).** Flags a visibly out-of-date copyright year in the footer — a small but unambiguous neglected-site signal (flagged on 3 of the 100 benchmark rows) |

> **Changed in this build — `website_platform_status` got more precise.** It used to fold every crawl failure into one value, `site_unreachable`, which read as "this business has no working website". Most of those sites were alive and simply refused a datacenter request. The value set now separates `site_blocked`, `not_found`, `site_error`, `no_page` and `crawl_error` from `site_unreachable`, which is reserved for hosts that produced no HTTP response at all. **If you have a saved filter, Make/Zapier step or script that matches `website_platform_status == "site_unreachable"`, it will now match fewer rows** — by design: the rows it stops matching are live websites. Match on the list above (or on `email_status`) to get the old, broader set. Pricing does not change, and no run delivers more rows than its `maxItems`. This build does read pages the previous one threw away (HTML served as `text/plain`, as XHTML, or with no `Content-Type` at all), so a site that used to yield no email can now yield one — with `onlyWithEmail` or `minScore` on, that can change *which* businesses fill your order, always within the same cap.

#### Business detail & reviews

| Field | Fill | Description |
|---|---|---|
| `opening_hours` | 80% | From OSM only; never filled from the site (see the licence note below) |
| `rating` | **7%** | **New, and sparse — see the warning below.** Star rating (1-5) the business publishes in its own schema.org markup |
| `review_count` | **7%** | **New, sparse.** Review count from the same markup |
| `price_range` | 24% | **New.** schema.org `priceRange` (e.g. `$$`) |

> **Honest warning about `rating` / `review_count`:** these are **not** Google Maps ratings. They are only present when a business publishes `aggregateRating` in its own JSON-LD, and most do not. Measured on-platform: **7%** of website-bearing rows for dentists in Austin (4/55) and **13%** for roofing contractors in Denver (1/8); the filter-off benchmark (n=100) measured **4%**. An earlier 12-row roofer sample hit 33%, so the rate swings wildly with category, city and sample size — assume **under 10%** and treat anything higher as luck. **Do not build a workflow that needs a rating on every row.** If you need star ratings and review counts for every business, this is the wrong source — pair it with our [Google Maps Places Scraper](https://apify.com/flash_scraper/google-maps-places); every row's `google_maps_url` also opens a Google Maps search for the business in one click. OpenStreetMap carries no review data at all, and this actor never invents a substitute: `lead_score` is a data-completeness score, **not** a customer rating.

#### Lead scoring

| Field | Fill | Description |
|---|---|---|
| `lead_score` | 100% | 0-100 completeness/reachability score |
| `lead_grade` | 100% | `A` ≥80, `B` ≥65, `C` ≥50, `D` ≥35, else `F` |
| `lead_tier` | 100% | `hot` ≥75, `warm` ≥50, else `cold` |
| `score_breakdown` | 100% | Per-component points, plus `signals_from` — the list of extra signals that actually scored — so the score is auditable |
| `has_email` / `has_phone` / `has_website` | 100% | Booleans for quick filtering |

**Scoring rubric** (sums to a true 100, so grade A is reachable — the filter-off benchmark's top row scored the full 100): email `deliverable` 40 / `risky` 25 / present-but-unverified 15 — **halved when `email_type` is `third_party`**; phone 20 (`phone` or `phones`); website 15; socials 5 each capped at 10; extra signals 5 each capped at 15, drawn from **website platform detected, opening hours, marketing tech found, and a published star rating** — `score_breakdown.signals_from` names the ones that counted.

#### Provenance

| Field | Fill | Description |
|---|---|---|
| `source` | 100% | `OpenStreetMap` for discovered rows, `user_supplied` for rows that came from your own `websiteList` |
| `attribution` | 100% on OSM rows | The ODbL attribution string, so the licence travels with an exported CSV. **Null on `user_supplied` rows** — they are not OpenStreetMap data, so attaching the notice would be a false licence claim. The column is always present either way |
| `osm_last_edited` | 100% of OSM rows (measured 2026-08-29 on three runs: 25/25 dentists in Austin, 20/20 restaurants in Chicago) | **New (2026-08-29).** ISO date of the element's last edit by any mapper, read from OpenStreetMap's own edit metadata (`out meta`) on the same request as before. Sort or filter on it to skip rows nobody has touched in years; the last column but one |
| `osm_check_date` | varies by category (measured 2026-08-29: 4/25 dentists in Austin, 3/20 restaurants in Chicago) | **New (2026-08-29).** The mapper's `check_date` tag — the day someone confirmed the business on the ground (used mostly for opening hours); present on a minority of rows. The last column |
| `enriched_from_website` | 33% | **New.** Lists only the fields where an *OSM-overlapping* value was taken from the site instead — `rating`, `review_count`, `price_range` (OSM carries none of these) plus `email` where OSM was empty and the site's schema.org data filled the gap; `phone` and `opening_hours` are never gap-filled, so they never appear here. It is **not** a full provenance map: `emails`, the five socials, `website_platform`, `platform_version`, `site_generator` and `tech` are *always* crawled from the site and are deliberately not repeated here |

#### Example output

Real rows from a run with `onlyWithWebsite: true`. **This is a best-case slice, not a typical one** — see the measured fill rates above; on the website-filtered reference run roughly half of rows carry an email.

| Name | Phone | Email | Status | Type | Platform | Tech | Score |
|---|---|---|---|---|---|---|---|
| Aloha Dental | +1-512-707-7300 | riverside@aloha-dental.com | deliverable | own\_domain | WordPress | Meta Pixel, GTM, Yelp | 95 (A) |
| Avery Ranch Dental | +1-512-246-7645 | smile@averyranchdental.com | deliverable | own\_domain | WordPress | GTM, reCAPTCHA | 95 (A) |
| Aviva Dental Care | +1 512 852 8528 | dr.apurva@avivadentalcare.com | deliverable | own\_domain | WordPress | Google Analytics | 90 (A) |

### How to read the output

One row per business, 73 columns, always in the same order. The Console's **Output** tab opens on the Overview view; the other views are the same rows with a different set of columns in front:

- **Overview** — the first-look table: Name, Category, City, State, Phone, Email, Email status, Website, Rating, Grade, Lead score, Map link. Read `lead_grade` (A–F) first, then `email_status`.
- **Email-ready** — only what a cold-email sequence needs: the address, its deliverability grade (`email_status`), whose mailbox it is (`email_type`), and the page it was found on.
- **Business details** — what the mapper wrote about the business: `speciality`, `cuisine`, `brand`, `is_chain`, opening hours, price range, and the rating / review count when the business's own site publishes one (roughly 3–8% of rows).
- **Web-agency targets** — the redesign-pitch signals: site builder, mobile viewport, stale copyright year, Meta Pixel and Google Ads tag, with `website_platform_status` saying whether the site was actually read.
- **Map & source** — where the row is (address, coordinates, Google Maps and OpenStreetMap links), how fresh it is (`osm_last_edited`, `osm_check_date`) and the ODbL attribution that has to travel with the data.
- **All columns** — every column, in CSV order.

Column labels follow one vocabulary across every Flash Scrape actor: **Name**, **Category**, **City**, **State**, **Phone**, **Email**, **Email status**, **Website**, **Lead score**, **Grade**, **Map link**. Booleans (`has_email`, `has_phone`, `has_website`, `is_chain`) are real true/false values; numbers (`rating`, `review_count`, `lead_score`, `latitude`, `longitude`) are real numbers, never strings; the two OSM dates are `YYYY-MM-DD`.

There are no derived or duplicated columns: `address` is already the full one-line postal address next to its `city` / `state` / `postal_code` / `country` parts, and `google_maps_url` is a ready-made link, so nothing needs assembling on your side.

**Exporting just one view.** Views change what the Console *shows*, not what its **Export** button downloads — a Console export always carries every stored column (Apify's dataset-schema docs: views only affect the Console display). To download only a view's columns use the `view` parameter of the dataset-items API, e.g. `https://api.apify.com/v2/datasets/<datasetId>/items?view=email_ready&format=csv` (view keys: `overview`, `email_ready`, `business_details`, `web_agency`, `map_source`, `all_columns`). Leave `view` off, or use `all_columns`, for the full table; `outputFields` in the input narrows what is stored in the first place.

**The run report and status message.** Every run that delivers rows ends with the same one-line status — `Done — N leads delivered for <what>. <coverage / filter notes> Report: <link>.` — and stores the `REPORT` HTML page whose table shows exactly the Overview columns (up to 10 of them: the sparse `rating` and the Map link are the first to be left out when the table is full), with tiles for businesses, % with email, % verified (when email verification ran), % with website and the average lead score.

### Input examples

**A preset plus the two dropdowns** — the shortest useful input there is:

```json
{
  "preset": "cold_email",
  "categorySelect": "hair salon",
  "citySelect": "Casablanca, Morocco",
  "maxItems": 100
}
```

The preset sets `onlyWithEmail` and `onlyVerifiedEmail`; everything else stays at its default. Add any field you want and yours wins — `{"preset": "web_design", "maxScore": 100}` keeps `maxScore: 100` and the log says which setting the preset stood down on.

**One category, one city** — the classic run, unchanged:

```json
{
  "category": "hair salon",
  "location": "Miami, Florida",
  "maxItems": 200,
  "crawlEmails": true,
  "onlyWithWebsite": true,
  "onlyWithEmail": false,
  "verifyEmails": true,
  "maxPagesPerSite": 3,
  "concurrency": 8
}
```

**Several categories across several cities**, contactable rows only, narrow columns:

```json
{
  "categories": ["dentist", "orthodontist"],
  "locations": ["Austin, Texas", "Dallas, Texas"],
  "maxItems": 200,
  "onlyWithWebsite": true,
  "requirePhone": true,
  "excludeKeywords": ["Aspen Dental"],
  "skipClosed": true,
  "outputFields": ["name", "email", "phone", "website", "lead_grade", "query_location"],
  "sortBy": "name_asc"
}
```

**A 5 km sales territory around one address:**

```json
{
  "category": "restaurant",
  "searchRadiusKm": 5,
  "centerLat": 30.2672,
  "centerLon": -97.7431,
  "maxItems": 150,
  "onlyWithEmail": true
}
```

**Web-design prospect list** — businesses with a site, but a *weak* one:

```json
{
  "category": "hair salon",
  "location": "Lyon, France",
  "countryCode": "fr",
  "onlyWithWebsite": true,
  "maxScore": 45,
  "skipClosed": true
}
```

**Enrich your own list** (no map data fetched at all):

```json
{
  "websiteList": ["aloha-dental.com", "averyranchdental.com", "typotes.com"],
  "verifyEmails": true,
  "emailPatternGuess": true
}
```

### How much does it cost?

This Actor is pay-per-result and costs **$3 per 1,000 delivered leads**, which is $0.003 per lead on the free plan and less on paid plans (the Pricing tab always carries the current rate). You are charged for delivered businesses only, not for API calls or compute. There is no subscription, and new Apify users get platform free credits to test with. **A scheduled pricing record raises this to $0.005 per delivered lead on 2026-09-14** (notified 2026-08-30); the Pricing tab on this page is authoritative.

Read from the live pricing on 2026-08-29 and re-read 2026-09-05: $0.003 per delivered lead on the free plan and $0.0027 (Bronze) down to $0.0021 (Diamond) on paid plans, so $5 buys 1,666 delivered leads; a change to $0.005 per result is scheduled for 2026-09-14 (notified 2026-08-30) and the Pricing tab is authoritative. For comparison, lukaskrivka/google-maps-with-contact-details charges $0.005 per place, $0.0025 per contact enrichment and $0.10 per verified email on the free plan (Store pricing read 2026-09-05), falling to $0.004 per verified email on Bronze.

**$5 buys 1,666 delivered leads** at that rate (1,000 after the scheduled 2026-09-14 change to $0.005 per lead). A free-plan visitor can run the form as it opens (25 dentists in Austin, at most $0.075 of leads plus a few cents of compute — well inside Apify's monthly free credit), download the CSV and judge every column before paying anything.

A run's total is simply rows delivered x the per-lead rate. Be aware what the rows contain: on the reference run with `onlyWithWebsite: true`, ~55% carried an email, so 500 rows ≈ 275 emailable leads (roughly 1.8x the per-lead rate per emailable lead). With the website filter off, 26% carried an email on the filter-off benchmark (roughly three-quarters on website-verified trade categories). Filters run before billing, so `onlyWithEmail: true` is the cheapest way to buy emails specifically — and the adaptive over-fetch delivers as close to `maxItems` email rows as the city allows (measured 2026-08-08: asked 25, delivered 25 in Austin; the same input delivered 19 before the pool-depth fix, and the whole city tops out at 31, which the run tells you when you ask for more).

### How it compares

Read from Apify's public Store API on **2026-09-05**. Nearly every alternative is a Google Maps scraper: [lukaskrivka/google-maps-with-contact-details](https://apify.com/lukaskrivka/google-maps-with-contact-details) (87,957 users, 4.63 from 221 reviews, $5 per 1,000 places, with contact enrichment at $2.50/1,000 and **email verification at $100/1,000 on Apify's free plan, falling to $4/1,000 on Bronze**), [s-r/google-maps-contact-details](https://apify.com/s-r/google-maps-contact-details) (166, no reviews, $4/1,000 plus $2/1,000 enrichment), [leadharbor](https://apify.com/leadharbor/google-maps-verified-email-scraper) (40, no reviews, $3/1,000 MX-checked), [code-node-tools](https://apify.com/code-node-tools/google-maps-lead-scraper) (33, no reviews, $1.10/1,000) and [jurassic\_jove](https://apify.com/jurassic_jove/google-maps-lead-generator) (99, no reviews, $20/1,000). Paid Apify plans get tiered discounts on several of these, this Actor included, so read every figure off the Pricing tab for the plan you are on.

This Actor (32 users on the public Store API cache, 5.0 from 3) is $3 per 1,000 with MX verification included — a change to $5 per 1,000 is scheduled for 2026-09-14 — but **discovery runs on OpenStreetMap, not Google Maps**, so there are no Google star ratings and van-based trades are thinly mapped.

For comparison, [lukaskrivka/google-maps-with-contact-details](https://apify.com/lukaskrivka/google-maps-with-contact-details) charges **$0.005 per place plus $0.0025 per contact enrichment plus $0.10 per verified email** on Apify's free plan (its live Store pricing record, read 2026-09-05) — the verification event alone is $100 per 1,000 there, falling to $4 per 1,000 on Bronze. Here, MX verification, mailbox-ownership classification and lead scoring are all included in the single per-lead rate.

### Where does the data come from, and what licence is it under?

Business listings come from **OpenStreetMap** under the **Open Database License (ODbL) v1.0**, and every contact detail comes from the business's own public website, which ODbL does not cover. **© OpenStreetMap contributors** — https://www.openstreetmap.org/copyright. Every row carries `source` and `attribution` fields; **keep them if you redistribute or publish the data**, as ODbL requires attribution. The columns added on 2026-08-29 from OSM tags — `speciality`, `cuisine`, `wheelchair`, `operator`, `osm_description`, `brand`, `brand_wikidata`, `is_chain`, `osm_last_edited`, `osm_check_date` — are OSM data under the same licence (`city_source` and `has_google_ads_tag` are ours).

The website-crawled fields — `email`, `emails`, socials, `website_platform`, `tech`, `rating`, `review_count`, `price_range`, `phones` — are **not** OSM-derived. They come from each business's own public website and are not covered by ODbL. Those fields are site-derived on **every** row. The `enriched_from_website` column is narrower than that: it flags only the fields that OSM *could* have supplied but didn't, so treat the list above — not that column — as the ODbL boundary.

### Frequently asked questions

#### How do I find local business leads with verified email addresses?

Run this Actor with a category, a city and `onlyWithEmail: true`, and every delivered row carries an email address; with verification left on (the default) each one is MX-checked over DNS-over-HTTPS and graded `deliverable`, `risky` or `undeliverable` in `email_status`. Measured on `dentist` / `Austin, Texas` with `onlyWithWebsite: true`, 55% of the 55 delivered rows carried an MX-verified email and 96% carried a phone; a `roofing contractor` / Denver run delivered emails on 6 of 8 rows.

Two things make those addresses usable rather than just present. `email_type` says whether the mailbox is the business's `own_domain`, a free inbox, or its **marketing agency** (mail to an agency mailbox never reaches the business). And `onlyWithEmail` and `onlyVerifiedEmail` run **before** billing, so a row without an email is dropped rather than charged, which makes them the cheapest way to buy emails specifically.

Those figures date from 2026-08-08, when the Cold-email list preset (onlyWithEmail plus onlyVerifiedEmail) delivered 31 leads for Austin dentists — the whole city that day — 100% with an MX-verified email and 94% with a phone.

#### Is it legal to scrape local business data?

This Actor reads only publicly available data: OpenStreetMap records, which are an open-data project licensed under ODbL, and each business's own public website, which is the same information anyone could read by visiting the site. There is no login, no private data, no anti-bot circumvention and no Google Maps terms-of-service exposure. Keep the `attribution` column with the data if you republish it, because ODbL requires attribution.

#### Does it work outside the US? How should I type the location?

Yes, it works anywhere OpenStreetMap covers, and you can type the location however you like. The raw string is tried first, and only if that finds nothing is it automatically re-spelled (Title Case, and a comma inserted before the trailing country word) before the run is given up on. `RABAT MAROC`, `Rabat Maroc`, `casablanca morocco` and `Rabat, Morocco` all resolve to the same place. If the run still cannot geocode, the status message lists every spelling it tried instead of a generic failure.

Verified: Rabat with countryCode mt resolves to Rabat, Western Region, Malta, while Morocco wins without the pin. The country column was fixed on 2026-08-08 — every Moroccan row used to ship a three-script run-on where Morocco belonged — and Austin, Texas and Lyon, France were unaffected.

If your city name exists in more than one country — `Rabat` is a city in **both Morocco and Malta**, `Cambridge` in both the UK and the US — set `countryCode` to an ISO 3166-1 alpha-2 code (`ma`, `mt`, `us`, `fr`, `gb`) to pin the geocoder. Verified: `location: "Rabat"` with `countryCode: "mt"` resolves to *Rabat, Western Region, Malta*; without it, Morocco wins. The pin is **enforced**: the fallback geocoder (Photon) has no country filter, so it is skipped whenever `countryCode` is set. A city that does not exist inside that country ends the run with a clear message and no charge, instead of quietly returning the same-named city elsewhere (`location: "Austin"` + `countryCode: "ma"` used to deliver Austin, **Texas**).

One non-US bug is fixed as of this build: the geocoder is now asked for English place names. It previously answered in the local language(s), and since the row's `country` column is derived from the geocoder's answer, **every Moroccan row used to ship `country: "Maroc ⵍⵎⵖⵔⵉⴱ المغرب"`** — a three-script run-on where `Morocco` belonged. Measured and fixed on 2026-08-08; `Austin, Texas` and `Lyon, France` are unaffected. Note that `city` and `address` still come straight from the OpenStreetMap tags a local mapper wrote, so a Moroccan row can legitimately read `Témara تمارة` — that is the map's own data and it is not rewritten.

#### How fresh is the data?

Every email, social profile, website platform and tech signal is crawled live from the business's website during the run, so that half of the row is current at run time; the OpenStreetMap half — the name, the address and the phone where the map supplies it — is only as fresh as the last mapper edit, which each row states in `osm_last_edited`.

That makes freshness something you can filter on rather than trust:

- `osm_last_edited` — the day any mapper last changed the element, read from the map's own edit metadata. On a 163-element sample of named `dentist` elements in the Austin bounding box (measured 2026-08-29), **12% had not been edited since before 2020**.
- `osm_check_date` — the mapper's `check_date` tag, set when someone confirmed the business on the ground. It was present on **13%** of that same sample, and the share varies by category and city.

Sort or filter on `osm_last_edited` to skip rows nobody has touched in years, and keep `skipClosed` on to drop premises mappers have retired.

#### How do I know whether a business email will bounce before I send to it?

Read the `email_status` column: with verification on (the default) every address is checked over DNS before delivery and graded `deliverable`, `risky` or `undeliverable`, and that verification is included in the per-lead price rather than sold as an add-on ($0.003 today, $0.005 from 2026-09-14 under a scheduled pricing record). Set `onlyVerifiedEmail: true` to keep only the rows that passed — `deliverable` and `risky` — and drop the ones proved dead. Turn `verifyEmails` off and the column falls back to `found` / `missing`; an address the run ran out of time to check also ships as `found`, unverified, and the status message says so.

On the 2026-08-08 reference run 55% of the 55 delivered rows carried an MX-verified email; 12% of all harvested addresses belonged to a third party (a marketing agency or web designer) but only 10% of primary emails did, because candidates are ranked by mailbox ownership before one is promoted.

Four checks run, all over DNS and none over SMTP: syntax, a mail-server (MX, or implicit-MX A record) lookup on the domain, a role-address check (`info@`, `sales@` and the like grade `risky`), and a disposable-domain list (grades `undeliverable`). The MX lookup goes over DNS-over-HTTPS to the public Google and Cloudflare resolvers, so no mail server is ever contacted and no proxy is needed. No mailbox is ever probed — that is what keeps verification proxy-free, fast and included in the price, and it is also what it cannot see: an address on a **catch-all** domain, or a **mailbox that was deleted** while the domain kept its mail servers, can still grade `deliverable`. Read `deliverable` as "the domain accepts mail and the address is not a known-bad shape", and warm a new list with a soft first send. When the MX lookup itself fails the row grades `risky` rather than `undeliverable`, so a busy resolver never deletes a good lead.

#### Can it tell me whether a business is running ads?

Partly: `has_google_ads_tag` and `has_meta_pixel` tell you whether the ads or conversion **tag** is present on the crawled page, which is a wiring signal rather than proof of live ad spend. Each is `true`, `false`, or `null` when no page was read. Sites keep the tag long after a campaign ends, and a business can run ads with no on-site tag at all. For actual live creatives, feed each row's `domain` into our [Competitor Google Ads Scraper](https://apify.com/flash_scraper/google-ads-transparency-scraper) (its `queries` input takes domains), which lists what a domain is currently running from Google's Ads Transparency Center.

has\_google\_ads\_tag was appended on 2026-08-29; both flags are tri-state (null when no page was read), and the underlying tech column — Meta Pixel, Google Analytics, GTM, Calendly, Klaviyo and about 25 more — was filled on 84% of rows on the 2026-08-08 reference run (n=55).

#### Why is `email` empty on some rows?

Three distinct reasons, and `email_status` plus `website_platform_status` tell you which: the business has no website in OSM (`no_website`), no page could be read from the site it does have (`site_blocked`, `not_found`, `site_error`, `no_page`, or `site_unreachable` — and only that last one means there is nothing there to reach), or the site simply does not publish an address anywhere on the pages crawled (`missing`). Cloudflare-obfuscated addresses **are** decoded, and since 0.1.21 the crawler follows the site's own contact/about/team links — including Shopify `/pages/contact` and non-English slugs — within the `maxPagesPerSite` budget. Raising `maxPagesPerSite` finds a few more.

Sized on 2026-08-08: with every filter off, 26% of 100 rows carried an email because roughly half of mapped businesses list no website; with onlyWithWebsite: true it was 55% of 55, and 8 of those 55 sites refused datacenter requests (site\_blocked) — a live business, not a dead one.

#### Does it get star ratings and review counts?

Only when a business publishes them in its own website markup, and these are never Google Maps ratings: measured at 7-33% of website-bearing rows depending on the category, and 4% on the latest filter-off benchmark. Do not build a workflow that needs a rating on every row; see the warning in the output section. OpenStreetMap has no review data, so if ratings are essential for every row, pair this with our [Google Maps Places Scraper](https://apify.com/flash_scraper/google-maps-places) — every row's `google_maps_url` opens a Google Maps search for the business in one click.

Measured 2026-08-08: 4 of 55 Austin dentists (7%) published a rating, and an Austin restaurant run asking for 15 rows with minRating: 4.0 crawled 86 candidates and delivered 0 rows, charging nothing — a rating floor drops unrated rows rather than keeping them.

#### How do I build a list of businesses running WordPress, Wix, Squarespace or Shopify in one city?

Run the category and city with `onlyWithWebsite: true` and read the `website_platform` column, which names the CMS or site builder behind each site (WordPress, Wix, Shopify, Squarespace, Webflow, GoDaddy, Weebly, Duda, Jimdo, Site123, Framer, HubSpot CMS, BigCommerce, Joomla, Drupal, and `Custom / Next.js` or `Custom / React` for hand-built sites). It was filled on **71% of rows** on the website-filtered reference run, and `website_platform_status` explains every row where it is empty.

There is no platform input to filter on, so the pattern is: pull the city, then filter the CSV on `website_platform`. Knowing a business runs on Wix, GoDaddy or Weebly rather than WordPress or a custom build lets you pre-qualify before picking up the phone, and combining it with `has_meta_pixel: false` and `has_booking_widget: false` narrows the list to businesses visibly under-invested in their web presence.

That 71% is the 2026-08-08 reference run (dentist / Austin, Texas, onlyWithWebsite: true, n=55); with every filter off the same city measured 34% of 100 rows, because roughly half of mapped businesses list no site to fingerprint.

#### Only independents? How do I leave the chains out?

Set **`excludeChains: true`**. OpenStreetMap marks a chain outlet with a `brand` tag (and usually `brand:wikidata`, both seeded from the name-suggestion-index — `brand=Aspen Dental`,

# Actor input Schema

## `preset` (type: `string`):

Pick one job and the rest of the form configures itself. A preset only fills in settings you left at their default value — <b>anything you set yourself always wins</b> — and the run log names the preset and lists exactly which settings it applied and which it skipped because you had already chosen them.<br><br><b>Cold-email list</b> — only businesses with an MX-verified, deliverable email.<br><b>Web-design prospects</b> — the thin/neglected half of the market (lead score at or below 45): no website at all, or a website that publishes no email and no marketing tech. It also requires at least one way to contact the business, because a score cap on its own selects the rows with no phone, no email and no socials — the ones you cannot pitch.<br><b>Call list</b> — only businesses with a phone number.<br><b>Full enrichment</b> — every enrichment signal on, 6 pages crawled per website, plus the opt-in guessed-email column.<br><br>A preset only ever <i>adds</i> filters or enrichment; it never removes a filter you set and never causes a business to be billed that your own settings exclude. Leave it on <code>Custom</code> to change nothing at all.

## `categorySelect` (type: `string`):

Pick a vertical whose OpenStreetMap tag has been verified end to end — <code>Dentist</code> maps to <code>amenity=dentist</code> + <code>healthcare=dentist</code>. Leave it on <b>Custom</b> to type anything you like in the box below: more than 200 terms are mapped, including <code>medspa</code>, <code>HVAC contractor</code>, <code>realtor</code> and <code>exterminator</code>.<br><br>Precedence: <b>Multiple categories</b> > this dropdown > the custom box.

## `category` (type: `string`):

Type the trade in plain English — <code>medspa</code>, <code>HVAC contractor</code>, <code>funeral home</code>, <code>photographer</code>. Used only while the dropdown above is on <b>Custom</b>. Common terms map to exact OpenStreetMap tags; unrecognised terms fall back to a broad tag guess PLUS a business-name match, and a mapped category with zero tagged places in the area is retried by business name too (businesses found only by name are labelled <code>category\_match='name\_keyword'</code> so you can filter them). OpenStreetMap coverage varies hugely by category: storefront businesses (dentists, salons, restaurants, shops) are dense, while van-based trades (plumbers, electricians, roofers) can return single digits even in a big city. A term matching nothing returns 0 businesses with a suggestion and no charge.

## `categories` (type: `array`):

Add SEVERAL categories to cover in one run — <code>dentist</code> + <code>orthodontist</code> + <code>dental clinic</code>. Every category is crossed with every location below (a matrix), the businesses are merged, and the same business found by two categories is delivered and billed <b>once</b> (deduplicated on its OpenStreetMap object id). Leave empty to use the dropdown or the custom box above — when this list is non-empty it beats both. Max 25 category x location combinations per run.

## `citySelect` (type: `string`):

Pick a metro that has been geocode-verified against OpenStreetMap, or leave it on <b>Custom</b> and type any city on earth in the box below — <code>Lyon, France</code>, <code>Marrakesh, Morocco</code>.<br><br>Precedence: <b>Multiple cities</b> > this dropdown > the custom box.

## `location` (type: `string`):

Type the city with its region or country — <code>Austin, Texas</code>, <code>Lyon, France</code>, <code>Casablanca, Morocco</code>. Used only while the dropdown above is on <b>Custom</b>. Capitalisation and the comma are optional: <code>RABAT MAROC</code> and <code>casablanca morocco</code> resolve too, because the raw string is tried first and then automatically re-spelled (<code>Rabat Maroc</code>, <code>RABAT, MAROC</code>, <code>Rabat, Maroc</code>) before the run is given up on.

## `locations` (type: `array`):

Add SEVERAL cities to cover in one run — <code>Austin, Texas</code> + <code>Dallas, Texas</code>. Each city is crossed with each category above, and the <b>Max businesses</b> budget is shared fairly between the searches (businesses are interleaved, so city 1 cannot eat the whole quota). Leave empty to use the dropdown or the custom box above — when this list is non-empty it beats both. Every business carries <code>query\_category</code> and <code>query\_location</code> so you can tell which search produced it. Max 25 category x location combinations per run. API callers: send a JSON array; a plain string is split on newlines and semicolons only, never on commas, so <code>Austin, Texas</code> stays one city.

## `countryCode` (type: `string`):

Pin the geocoder to one country with an ISO 3166-1 alpha-2 code — <code>ma</code> (Morocco), <code>us</code>, <code>fr</code>, <code>gb</code>. Use it when your city name exists in several countries: <i>Rabat</i> is a city in both Morocco and Malta, <i>Cambridge</i> in both the UK and the USA. Comma-separated codes are allowed (<code>gb,ie</code>). Leave empty for a worldwide search. The pin is enforced: if the city cannot be resolved inside that country the run stops with a clear message and no charge, rather than falling back to a worldwide lookup and delivering a same-named city somewhere else.

## `keyword` (type: `string`):

Keep only businesses whose NAME contains this word — <code>smile</code> narrows dentists to <i>Smile Studio</i> and <i>Bright Smiles</i>. Since 2026-08-29 it is also matched against the OpenStreetMap <code>cuisine</code> and <code>healthcare:speciality</code> tags, so <code>sushi</code> returns the sushi restaurants whose name never says sushi and <code>orthodontics</code> the dentists tagged with that speciality. Case-insensitive and matched anywhere in the value; the address, website and email are never matched. Leave empty to keep every business in the category.

## `maxItems` (type: `integer`):

Cap the whole run — you are charged per delivered business, so this is your cost ceiling. Anything from <b>1 to 10,000</b> is accepted. With several categories or cities the budget is <b>shared</b> between the searches, not multiplied by them. Businesses removed by a filter never count against it and are never billed. The Console form starts at 25 (about $0.075 of leads at $3 per 1,000); raise it once the first run looks right. API calls, tasks and schedules that send no <code>maxItems</code> keep the default of 100. Large orders: the run tells you and reduces the order if it cannot crawl and enrich that many in the time available, so you are not billed for rows the run had no time to enrich. Rows with no email are delivered and billed by default - only <code>onlyWithEmail</code> guarantees an email on every billed row. Depth is limited by what OpenStreetMap holds - a single category in a single city is often a few hundred businesses, so five figures needs several categories or cities.

## `expandNearby` (type: `boolean`):

A single city often holds fewer businesses than Max businesses asks for — measured live: 'dentist' in Austin tops out near 119. With this on, the run automatically widens the same search in growing rings around the city (up to ~150 km) until your cap is met or the region is genuinely exhausted. Every widened row is labelled in <code>query\_location</code> ('within ~40 km of Austin, Texas'), so you can always tell them apart or filter them out. Off = strict city limits, exactly the previous behaviour. Ignored when you set your own radius (searchRadiusKm) or supply your own website list.

## `onlyWithWebsite` (type: `boolean`):

Keep only businesses that list a website — this cuts your bill, because filtered businesses are dropped BEFORE you are charged. Strongly recommended for lead generation: OpenStreetMap businesses without a website almost never carry a phone, email or social profile either — on the reference run all 29 such businesses had none of the three, so they were name + coordinates + address only, and all graded F. The default is off so density research and existing API callers still get every mapped location; the Console form switches it on for you.

## `onlyWithoutWebsite` (type: `boolean`):

Keep only businesses that do NOT list a website — the prospect list for web-design and digital agencies pitching a first website. Mutually exclusive with <b>Website required</b>: setting both stops the run straight away with a message naming the clash, delivers nothing and charges nothing. Filtered businesses are dropped BEFORE billing. Note these businesses are name + address + coordinates (and occasionally a phone) only: no email, platform or tech enrichment is possible without a website to crawl.

## `onlyWithEmail` (type: `boolean`):

Keep only businesses where an email was found — the cheapest way to buy emails specifically, since businesses without one are dropped before billing and you pay only for contactable leads. With this on, the Actor over-fetches and keeps crawling extra candidates until it has <code>maxItems</code> businesses with an email or the area is exhausted, so asking for 100 emails delivers as close to 100 as the city allows rather than ~40. The candidate pool is 6x <b>Max businesses</b> for a single filter and deepens as you stack filters, up to 24x, and never past 30,000 candidates in one run — that bounds the crawling, not your bill: only delivered rows are charged. A guessed address (see <b>Guessed info@ address</b>) never satisfies this filter.

## `verifyEmails` (type: `boolean`):

Check every email's domain for real mail servers (MX) and flag role (<code>info@</code>, <code>sales@</code>) and disposable addresses — included at no extra charge, no key needed, and it protects your sender reputation. With this ON, <code>email\_status</code> is <code>deliverable</code>, <code>risky</code> or <code>undeliverable</code> when an email was found and <code>missing</code> when none was; with it OFF the values are <code>found</code> and <code>missing</code> instead, so turning verification off changes the vocabulary. Also populates <code>email\_provider</code> (Google Workspace, Microsoft 365...). It is a syntax + MX + role + disposable check over DNS, never an SMTP mailbox probe, so an address on a catch-all domain or a mailbox that was deleted can still grade <code>deliverable</code>.

## `onlyVerifiedEmail` (type: `boolean`):

Drop leads whose email fails deliverability verification, keeping only <code>email\_status</code> <code>deliverable</code> or <code>risky</code> (a domain with real mail servers); businesses with no email at all are dropped too. Requires <b>Email verification</b> to be on — set it without verification and the run turns verification on for you rather than delivering rows the filter is meant to remove. Like <b>Email required</b>, this over-fetches (a pool of 6x <b>Max businesses</b>, deepening to 24x as you stack filters, capped at 30,000 candidates) until <code>maxItems</code> verified leads exist or the area is exhausted.

## `requirePhone` (type: `boolean`):

Drop every business with no phone number, before billing. A business passes when either column carries a number: <code>phone</code> (the OpenStreetMap number) or <code>phones</code> (the <code>tel:</code> numbers its own website publishes) — if <b>Phone numbers from websites</b> is off, only map-sourced numbers count. Combine with <b>Email required</b> for a fully contactable list.

## `requireSocial` (type: `boolean`):

Drop every business with no Facebook, Instagram, LinkedIn, X/Twitter or YouTube profile, before billing — useful for social-media agencies and DM-first outreach. Needs website crawling (or an OpenStreetMap social tag) to find anything: with crawling off, almost every business is dropped.

## `requireAnyContact` (type: `boolean`):

Drop every business that has <b>no phone, no email and no social profile</b>, before billing. This is the loosest possible contactability floor — a business only needs ONE reachable channel to survive it, unlike <b>Phone number required</b> or <b>Email required</b>, which each demand a specific one.<br><br>Worth knowing: OpenStreetMap maps plenty of businesses as little more than a name and a map pin. On an unfiltered run of the default search roughly half the delivered rows carried no contact details of any kind, and you were billed for them. This switch removes exactly those rows. Left off by default so that nothing about an existing run, task or API call changes.

## `excludeKeywords` (type: `array`):

List the words that disqualify a business by NAME — add <code>Aspen Dental</code> to strip a chain out of a dentist list. Case-insensitive and matched anywhere in the name; only the name is matched — never the address, website or email — so <code>clinic</code> cannot knock out a business on Clinic Street. Typical uses: filter out chains and franchises, or your own existing customers. Businesses are dropped before crawling and before billing. The form arrives with this list empty, so nothing is excluded until you add a keyword; API callers who send nothing get no exclusions at all. Whenever it removes anything, the run logs how many businesses your exclusion list took out in total (one figure for the whole list, not one per keyword).

## `excludeChains` (type: `boolean`):

Drop every business OpenStreetMap marks as a chain outlet — a <code>brand</code> or <code>brand:wikidata</code> tag, which is what sets the <code>is\_chain</code> column to <code>true</code> — before crawling and before billing, leaving the independents. Decided from the map data alone, so no crawl budget is spent on a dropped outlet, and the run says how many it removed. Businesses you supply yourself (<b>websiteList</b> / <b>startUrls</b>) have no brand tag (<code>is\_chain</code> is null) and are never dropped by this. Off by default; on a schedule with <b>Only new businesses</b> it joins the watch's memory key only when you switch it on, so existing schedules keep their memory.

## `skipClosed` (type: `boolean`):

Drop OpenStreetMap elements the mappers have retired: a lifecycle-prefixed tag (<code>disused:amenity</code>, <code>abandoned:shop</code>, <code>was:\*</code>), a <code>disused=yes</code> / <code>abandoned=yes</code> flag, <code>opening\_hours=closed</code>, or <code>shop=vacant</code>. Measured live on 2026-08-08: 199 such elements in the Austin bounding box, 43 of them still carrying a business name. They cannot reach a normal tagged search, but they DO reach the business-name fallback used for unmapped categories — which is where dead businesses were being delivered as fresh leads. Off by default so existing runs are unchanged; the Console form switches it on for you.

## `minScore` (type: `integer`):

Keep only leads scoring at or above this value, 0-100, where <code>0</code> means no filter. Every lead is scored from its verified email (deliverable 40 / risky 25 / unverified 15 / verified-undeliverable 0, halved if the address belongs to a third-party marketing agency), phone 20, website 15, socials 5 each capped 10, and extra signals 5 each capped 15 (website platform detected, opening hours, marketing tech, star rating — <code>score\_breakdown.signals\_from</code> names the ones that counted). Grades: A>=80, B>=65, C>=50, D>=35, else F. This is a <b>data-completeness</b> score, NOT a customer rating.

## `maxScore` (type: `integer`):

Keep only leads scoring at or BELOW this value, 0-100 — the web-design agency filter. A low score means a thin online presence (no email published, no marketing tech, often a builder-tier website), which is exactly the redesign pitch list. Try <code>45</code> to get the neglected half. Accepts 0-100; leave empty for no upper limit. Setting a maximum below the minimum stops the run straight away with a message and charges nothing.

## `minRating` (type: `number`):

Keep only businesses with a published star rating at or above this value, 1-5. <b>READ THIS BEFORE USING IT:</b> OpenStreetMap carries NO review data. A rating only exists when the business publishes schema.org <code>aggregateRating</code> on its own website, measured at roughly 3-8% of businesses. Businesses with NO rating are DROPPED, not kept — an unknown rating is not a passing rating, and no rating is ever invented to save a business. Setting this will therefore slash the number of businesses you get to a small fraction. If you need ratings on every business, use a Google Maps source instead; every business here carries a <code>google\_maps\_url</code> to the live listing.

## `minReviewCount` (type: `integer`):

Keep only businesses whose own website publishes a review count at or above this value (1-100,000). Same warning as the minimum star rating: review counts come only from schema.org markup on the business's own website (roughly 3-8% of businesses), never from OpenStreetMap, and businesses with no review count are DROPPED rather than kept. Expect very few businesses to survive it.

## `onlyNewBusinesses` (type: `boolean`):

Off by default. The FIRST run of a watch is the baseline — it delivers everything it finds and says so. Every later run with the <b>same categories, locations and filters</b> delivers ONLY businesses that were not delivered before; already-delivered ones are dropped <b>before the crawl</b>, so they cost nothing and are never billed. A run where nothing is new delivers 0 rows, bills nothing, and says “Nothing new … You were not charged.”<br><br>A business is remembered by its OpenStreetMap object identity (<code>node/123456789</code>), so a rename or a new domain does not make it look new again; websiteList rows are remembered by their domain. A row is marked seen only <b>after</b> it has actually been delivered, so nothing you were not sent can be skipped next time.<br><br>The watch is identified by your search + the filters that shape it (categories, locations, countryCode, keyword, radius/centre, expandNearby, websiteList, and every <b>Filters</b> option). Change any of them and you start a separate watch with its own memory — <b>maxItems</b>, <b>sortBy</b> and <b>outputFields</b> are deliberately NOT part of it, so tuning them never re-bills you for businesses you already have. Memory lives in a named key-value store (<code>local-business-leads-monitor</code>) in your own account, is kept for 90 days, and holds the 50,000 most recent businesses per watch — past that the oldest are forgotten and could be delivered (and billed) again.<br><br>Note: with this on, the run searches a deeper pool of candidates (the already-delivered ones have to be skipped over), which takes longer than the same run without it.

## `crawlEmails` (type: `boolean`):

Visit each business's own website to extract a contact email, 5 social profiles, phone numbers, the website platform (WordPress, Wix, Squarespace, Shopify, GoDaddy...), marketing/booking tech, and any star rating it publishes. Turn it off and every website-derived column stays empty and the email filters become unusable. Included in the per-business price.<br><br>Every row records what happened to its website in <code>website\_platform\_status</code>: <code>detected</code> (platform identified), <code>unknown</code> (a page loaded, nothing recognisable on it), <code>site\_blocked</code> (the site answered but refused us — a 403, a bot-check page or a login wall), <code>not\_found</code> (every URL that answered came back 404 or 410 — the address in the listing is dead, the host serving it is not), <code>site\_error</code> (the site answered with a server error), <code>no\_page</code> (it answered, but never with a readable web page), <code>site\_unreachable</code> (nothing answered at all — DNS failure, refused connection, broken TLS or a timeout), <code>no\_website</code> (there was no website to crawl), <code>not\_crawled</code> (crawling was off) and <code>crawl\_error</code> (our own crawler failed on that site — worth reporting). Only <code>site\_unreachable</code> means the business has no reachable website, and only <code>crawl\_error</code> says nothing at all about the site (it is a fault on our side); every other value tells you the website exists and why we could not read it, so a blocked or 404 site is still a live prospect.

## `crawlSocialProfiles` (type: `boolean`):

Fill the <code>facebook</code>, <code>instagram</code>, <code>linkedin</code>, <code>twitter</code> and <code>youtube</code> columns from the crawled pages, with share buttons, tracking pixels and post embeds filtered out. On by default — that is the current behaviour. Turn it off to leave the five social columns empty and keep exports narrower. Social links that OpenStreetMap itself carries are unaffected.

## `extractPhonesFromSite` (type: `boolean`):

Collect every <code>tel:</code> number on the crawled pages into the <code>phones</code> column. <code>phone</code> always means "what OpenStreetMap said" and is never filled from the site; <code>phones</code> is where the site's own numbers live, and <b>Phone number required</b>, <code>has\_phone</code> and the lead score count either column. On by default. Turn it off if you only trust map-sourced numbers: <code>phones</code> then stays empty.

## `emailPatternGuess` (type: `boolean`):

Write a GUESSED <code>info@\<domain></code> into the separate <code>email\_guess</code> column when a business has a website but publishes no address anywhere that was crawled — OFF by default, and it stays out of your <code>email</code> column on purpose. <code>email\_guess\_confidence</code> is <code>mx\_ok</code> (the domain does run mail servers), <code>no\_mx</code> (it does not — the guess is almost certainly dead) or <code>unchecked</code> (email verification is off). A guess is never promoted to <code>email</code>, never counts as <code>has\_email</code>, never earns a lead-score point and never satisfies the email filters. Treat it as a lead, not a verified mailbox — nobody checked that this mailbox exists.

## `maxPagesPerSite` (type: `integer`):

Set how deep to look for an email on each website, 1-8 pages (default 3): fixed contact paths, plus up to 4 contact/about/team links discovered in the website's own menu (covers Shopify <code>/pages/contact</code> and non-English slugs), plus a bare-domain fallback. Higher finds more emails and takes longer; it does not change what you are charged per business.

## `concurrency` (type: `integer`):

Set how many business websites to crawl in parallel, 1-50 (default 8). Raise it to finish a large run faster - the run's own time estimate scales with it, so a 10,000-business order needs it well above the default to fit one run timeout - and lower it if websites start refusing connections. Requests to any one website stay capped at 2 in flight whatever you set here, and OpenStreetMap lookups are not affected by it.

## `searchRadiusKm` (type: `number`):

Search a circle around a point instead of a city's bounding box — <code>5</code> covers every business within 5 km of one address. Accepts 1-100 km; for anything wider, search by city with <b>Widen to nearby areas</b>, which rings out to ~150 km. Set this together with the latitude and longitude below. When set, it REPLACES the location/locations bounding box entirely (no geocoding happens, so the <code>country</code> column stays empty) and <b>Widen to nearby areas</b> is ignored, because you defined the area yourself. Leave empty to search by city.

## `centerLat` (type: `number`):

Enter the latitude of the radius centre, -90 to 90, decimals allowed — <code>30.2672</code> is downtown Austin. Required when a search radius is set, ignored otherwise.

## `centerLon` (type: `number`):

Enter the longitude of the radius centre, -180 to 180, decimals allowed — <code>-97.7431</code> is downtown Austin. Required when a search radius is set, ignored otherwise.

## `websiteList` (type: `array`):

Paste websites or bare domains you already have — <code>example.com</code>, <code>https://www.example.com/contact</code> — and business discovery is SKIPPED entirely: only the crawl + email verification + scoring pipeline runs over your list. Use it to enrich a CRM export, a conference exhibitor list or a competitor's client list. These businesses carry <code>source='user\_supplied'</code>, an empty attribution (they are not OpenStreetMap data, so no ODbL notice is attached), no coordinates and no OSM ids; their name comes from the website's own page title, falling back to the bare domain. Filters and scoring work exactly as on discovered businesses. Entries pointing at the same domain (<code>example.com</code>, <code>https://www.example.com</code>, <code>example.com/contact</code>) are merged, so one business is crawled, delivered and billed once. <b>Max businesses</b> caps how much of the list is used: with no row-dropping filter set, only the first <code>maxItems</code> websites are crawled — raise it to cover a longer list — and whenever the cap trims your list the run logs how many entries it actually used. If you supply a list and NONE of the entries is a usable website or domain, the run stops with an error instead of quietly falling back to an OpenStreetMap search you did not ask for — nothing is charged.

## `startUrls` (type: `array`):

Use this instead of <b>Your own website list</b> when your integration already speaks Apify's usual <code>startUrls</code> convention — it behaves identically. Both lists are merged and deduplicated by registered domain (not by URL), so <code>example.com</code> and <code>https://www.example.com/contact</code> are one business — crawled, delivered and billed once.

## `outputFields` (type: `array`):

List the columns you want in the dataset — <code>name</code>, <code>email</code>, <code>phone</code>, <code>website</code>, <code>lead\_grade</code>. Leave empty to get all 73. Every delivered business carries exactly the same columns in the documented order, so CSV headers never shift mid-export. <code>name</code> and <code>attribution</code> are always included whatever you choose — attribution because the OpenStreetMap licence (ODbL) has to travel with the data into your CSV. Names are matched case-insensitively and <code>-</code>/space count as <code>\_</code>, so <code>Lead Grade</code> and <code>email-status</code> resolve. An unknown name is reported in the log and ignored; but if NONE of the names you list exists the run stops before billing, because a name-only export at full price is not worth paying for.<br><br>Valid names: name, category, category\_match, address, city, state, postal\_code, country, latitude, longitude, query\_category, query\_location, phone, phones, website, domain, email, emails, email\_status, email\_provider, email\_domain\_match, email\_type, email\_types, contact\_page\_url, email\_guess, email\_guess\_confidence, facebook, instagram, linkedin, twitter, youtube, website\_platform, website\_platform\_status, platform\_version, site\_generator, tech, has\_meta\_pixel, has\_google\_analytics, has\_booking\_widget, website\_title, website\_description, mobile\_viewport, copyright\_year\_stale, opening\_hours, rating, review\_count, price\_range, lead\_score, lead\_grade, lead\_tier, score\_breakdown, has\_email, has\_phone, has\_website, enriched\_from\_website, source, attribution, google\_maps\_url, osm\_url, osm\_type, osm\_id, has\_google\_ads\_tag, speciality, cuisine, wheelchair, operator, osm\_description, brand, brand\_wikidata, is\_chain, city\_source, osm\_last\_edited, osm\_check\_date.

## `sortBy` (type: `string`):

Choose the order the businesses arrive in — <b>Lead score, best first</b> is the default and what previous builds always did. Sorting never removes a business; all filtering happened earlier. Businesses missing the sort value (no rating, no review count) are placed last, because a missing value is unknown rather than zero.

## `webhookUrl` (type: `string`):

Optional. When at least one row is delivered, the run POSTs a digest to this URL: a Slack incoming webhook gets a text message, a Discord webhook gets a message, any other URL (n8n / Make / Zapier catch hook, your own endpoint) gets JSON with the counts, console links and the first 20 rows. Quiet runs send nothing. Pair it with onlyNewBusinesses on a schedule and this actor becomes an alert service on its own. Delivery is best-effort: a webhook failure is reported in the status message and never fails the run.

## Actor input object example

```json
{
  "preset": "",
  "categorySelect": "",
  "category": "medspa",
  "categories": [
    "dentist",
    "orthodontist"
  ],
  "citySelect": "",
  "location": "Lyon, France",
  "locations": [
    "Austin, Texas",
    "Dallas, Texas"
  ],
  "countryCode": "ma",
  "keyword": "smile",
  "maxItems": 25,
  "expandNearby": true,
  "onlyWithWebsite": true,
  "onlyWithoutWebsite": false,
  "onlyWithEmail": false,
  "verifyEmails": true,
  "onlyVerifiedEmail": false,
  "requirePhone": false,
  "requireSocial": false,
  "requireAnyContact": true,
  "excludeKeywords": [
    "Aspen Dental",
    "Walmart"
  ],
  "excludeChains": false,
  "skipClosed": true,
  "minScore": 0,
  "maxScore": 45,
  "onlyNewBusinesses": false,
  "crawlEmails": true,
  "crawlSocialProfiles": true,
  "extractPhonesFromSite": true,
  "emailPatternGuess": false,
  "maxPagesPerSite": 3,
  "concurrency": 8,
  "centerLat": 30.2672,
  "centerLon": -97.7431,
  "outputFields": [
    "name",
    "email",
    "phone",
    "website",
    "lead_grade"
  ],
  "sortBy": "score_desc"
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "category": "dentist",
    "location": "Austin, Texas",
    "maxItems": 25,
    "expandNearby": true,
    "onlyWithWebsite": true,
    "requireAnyContact": true,
    "skipClosed": true,
    "onlyNewBusinesses": false
};

// Run the Actor and wait for it to finish
const run = await client.actor("flash_scraper/local-business-leads").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "category": "dentist",
    "location": "Austin, Texas",
    "maxItems": 25,
    "expandNearby": True,
    "onlyWithWebsite": True,
    "requireAnyContact": True,
    "skipClosed": True,
    "onlyNewBusinesses": False,
}

# Run the Actor and wait for it to finish
run = client.actor("flash_scraper/local-business-leads").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "category": "dentist",
  "location": "Austin, Texas",
  "maxItems": 25,
  "expandNearby": true,
  "onlyWithWebsite": true,
  "requireAnyContact": true,
  "skipClosed": true,
  "onlyNewBusinesses": false
}' |
apify call flash_scraper/local-business-leads --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,flash_scraper/local-business-leads"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/BRiuRAMHY1njVTRqQ/builds/OGGw6UbHnEIQLp2He/openapi.json
