# Google AI Overview & Brand Visibility Tracker (GEO) (`jongoose/geo-tracker-scraper`) Actor

Track brand visibility in Google's AI Overview (SGE) and organic search across a list of queries. Per query: detect the AI Overview, extract the source domains it cites, flag whether your domain is cited, and return the top-10 organic ranking + People Also Ask. For GEO/SEO.

- **URL**: https://apify.com/jongoose/geo-tracker-scraper.md
- **Developed by:** [James Scott](https://apify.com/jongoose) (community)
- **Categories:** SEO tools, Marketing
- **Stats:** 2 total users, 2 monthly users, 93.8% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Google AI Overview & Brand Visibility Tracker (GEO)

Track how visible your brand is in **Google's AI Overview** (the generative "AI Overview" answer box, formerly SGE) **and** in classic organic results - for any list of search queries. Built for **GEO (Generative Engine Optimization)** and SEO monitoring: know when an AI Overview appears, which sources it cites, and whether *your* domain is one of them.

### What you get (one record per query)

- **`ai_overview_present`** - is there an AI Overview on the page?
- **`ai_overview_text`** - the AI Overview's answer text (cleaned of citation badges)
- **`ai_overview_sources`** - the list of **domains the AI Overview cites** (the core GEO signal), plus `ai_overview_source_urls` (the full cited URLs)
- **`tracked_domain_present_in_overview`** - is your domain cited inside the AI Overview? (when `trackDomain` is set)
- **`organic_results`** - top 10 organic results: `position`, `title`, `url`, `domain`
- **`tracked_domain_organic_rank`** - your domain's organic position (or `null`)
- **`people_also_ask`** - the "People also ask" questions
- **`serp_state`** - `results` on success (or `js_challenge` / `consent` / `captcha` if a proxy misconfiguration got a gate instead of results)

### Input

- **`queries`** (required) - the searches to track, e.g. `["best crm for small business", "how to file an llc"]`. One SERP is fetched per query.
- **`trackDomain`** - your website domain, e.g. `acme.com`. Enables the "cited in AI Overview?" and organic-rank fields (subdomains match too).
- **`trackBrand`** - your brand name (stored on each record for reporting).
- **`gl`** / **`hl`** - country and language codes for localisation (default `us` / `en`).
- **`maxItems`** - cap on how many queries to process.
- **`proxyConfiguration`** - keep the default **`GOOGLE_SERP`** proxy group (see below).

#### Example input

```json
{
  "queries": ["best crm for small business", "generative engine optimization"],
  "trackBrand": "Zoho",
  "trackDomain": "zoho.com",
  "gl": "us",
  "hl": "en",
  "proxyConfiguration": { "useApifyProxy": true, "apifyProxyGroups": ["GOOGLE_SERP"] }
}
```

#### Example output (one query)

```json
{
  "query": "best crm for small business",
  "timestamp": "2026-07-10T14:40:00Z",
  "ai_overview_present": true,
  "ai_overview_text": "The best CRM for a small business depends on budget and needs...",
  "ai_overview_sources": ["hubspot.com", "zoho.com", "forbes.com"],
  "tracked_domain_present_in_overview": true,
  "organic_results": [
    { "position": 1, "title": "HubSpot CRM...", "url": "https://www.hubspot.com/products/crm", "domain": "hubspot.com" },
    { "position": 3, "title": "Zoho CRM...", "url": "https://www.zoho.com/crm/", "domain": "zoho.com" }
  ],
  "tracked_domain_organic_rank": 3,
  "people_also_ask": ["Which CRM is best for a small business?"],
  "serp_state": "results"
}
```

### Why this Actor

- **AI Overview-first.** Most SERP scrapers only give you organic results. This one is built around **detecting the AI Overview and extracting the domains it cites** - the metric GEO/SEO teams now care about most.
- **Durable parser.** Google rotates its CSS class names every few weeks, so detection keys on **stable signals** (the visible "AI Overview" label, `data-attrid`, and the `<a><h3>` organic pattern) with class-name hints only as a fallback.
- **Honest reporting.** Every record carries `serp_state`, so a proxy/consent gate is never silently reported as "no AI Overview".

### Proxy & how Google fetching works

Google serves a JavaScript/bot gate to raw IPs, so this Actor uses Apify's **`GOOGLE_SERP` proxy group** (available on the free plan, tuned for Google SERP scraping) to get clean results HTML. Keep the default proxy configuration. Requests go to `http://www.google.com/search` (the GOOGLE\_SERP proxy is HTTP-only and requires the `www.` host); localisation is via `gl`/`hl`.

**AI Overview caveat (important):** an AI Overview can be (1) rendered in the initial HTML, (2) *deferred* and injected by JavaScript after the page paints, or (3) absent. This Actor fetches HTML over HTTP, so it reliably catches states (1) and (3). A minority of overviews load only via JavaScript (state 2) and will read as `ai_overview_present: false` - this is a known limitation of every non-headless SERP scraper. Organic results, source detection, and rankings are unaffected.

This data is for **market/SEO research**. It is not a consumer report and must not be used for FCRA-covered purposes.

# Actor input Schema

## `queries` (type: `array`):

The search queries to track (one Google SERP is fetched per query). Use the queries your customers actually search - e.g. "best crm for small business", "how to file an llc". One record is returned per query.

## `trackBrand` (type: `string`):

Your brand / company name. Stored on each record for reporting; used for readability (domain matching drives visibility detection).

## `trackDomain` (type: `string`):

Your website domain, e.g. "acme.com". If given, each record flags whether this domain is cited inside the AI Overview and reports its organic rank (subdomains match too).

## `gl` (type: `string`):

Two-letter Google country code for result localisation (e.g. us, gb, ca, au, de).

## `hl` (type: `string`):

Two-letter Google interface/results language code (e.g. en, es, fr, de).

## `maxItems` (type: `integer`):

Process at most this many queries from the list (one record each).

## `proxyConfiguration` (type: `object`):

Google blocks raw IPs with a JavaScript/consent gate. Keep the GOOGLE\_SERP proxy group (Apify's Google-SERP proxy, available on the free plan) - it returns clean results HTML.

## Actor input object example

```json
{
  "queries": [
    "best crm for small business",
    "generative engine optimization",
    "how does an ai overview work"
  ],
  "trackBrand": "Acme CRM",
  "trackDomain": "acme.com",
  "gl": "us",
  "hl": "en",
  "maxItems": 50,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "GOOGLE_SERP"
    ]
  }
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "best crm for small business",
        "generative engine optimization",
        "how does an ai overview work"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("jongoose/geo-tracker-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "queries": [
        "best crm for small business",
        "generative engine optimization",
        "how does an ai overview work",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("jongoose/geo-tracker-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "best crm for small business",
    "generative engine optimization",
    "how does an ai overview work"
  ]
}' |
apify call jongoose/geo-tracker-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,jongoose/geo-tracker-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/wJlVwKxwnNKiW5CUR/builds/yPqd1xYqDLdPKkxgq/openapi.json
