# Shopify Catalog & Merchandising Intelligence (`zinin/shopify-store-intelligence`) Actor

Qualify Shopify merchant leads with evidence-backed public catalog coverage, observed price positioning, vendors, product categories, and merchandising recency. Includes confidence, data gaps, and a review-ready next action. No login or Shopify API key.

- **URL**: https://apify.com/zinin/shopify-store-intelligence.md
- **Developed by:** [Tim Zinin](https://apify.com/zinin) (community)
- **Categories:** E-commerce, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.10 / 1,000 store analyzeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Shopify Catalog & Merchandising Intelligence

Turn a list of public merchant websites into **evidence-backed Shopify qualification data** for
marketing, sales research, ecommerce consulting, and competitive monitoring. For every submitted
website, the Actor checks whether supported Shopify evidence is visible and, when the public feeds
are available, observes catalog coverage, variant list prices, vendors, product categories,
collections, and the newest product timestamp.

The result is designed for a decision, not just a scrape. Every successful observation includes
source evidence, count semantics, freshness, confidence reasons, explicit data gaps, and a
recommended review action. The Actor does **not** turn catalog size into a fictional revenue number:
public product counts and list prices do not prove sales volume, conversion, or revenue.

![Shopify Catalog and Merchandising Intelligence turns merchant domains into evidence-backed catalog, price-positioning, and merchandising signals for lead review](https://raw.githubusercontent.com/TimmyZinin/apify-actor-assets/7bac0e9d52219782d32a587479b3e15b6490ca5d/commercial115/shopify-store-intelligence/readme-hero.webp)

### The product in one sentence

Submit 1–100 public HTTP(S) websites and receive one review-ready Shopify qualification result per
site: platform evidence, exact-or-lower-bound catalog semantics, observed price positioning,
merchandising recency, provenance, confidence, gaps, and the next sensible human action.

### Why marketers need more than “this site uses Shopify”

A technology flag is useful, but it is rarely enough to decide whether an account belongs in a
campaign. Two merchants may both use Shopify while having very different public storefronts:

- one exposes twelve products in a narrow category and has not published a product recently;
- another exposes hundreds of products across several product types and vendors;
- one has observed variant list prices around USD 15;
- another has an observed median variant price above USD 100;
- one exposes a complete public feed;
- another reaches the 750-product scan cap, so the observed count is only a lower bound.

Those differences can help a marketer choose an appropriate segment, message, service package, or
manual research priority. They cannot prove revenue, budget, purchase intent, internal staffing, or
whether a merchant wants to buy anything. This Actor keeps that boundary visible in the data.

### What you receive

For each reachable website, the Actor can return:

- **Shopify evidence:** whether supported Shopify signatures were observed on the submitted public
  storefront response.
- **Catalog coverage:** products observed from the public `/products.json` feed, with `exact` or
  `lower_bound` semantics instead of an ambiguous total.
- **Observed price positioning:** minimum, maximum, average, and median public variant list prices,
  the number of variant prices observed, and storefront currency when it is visible.
- **Vendor and category mix:** the most common public `vendor` and `product_type` values, plus how
  many observed products supported each leading value.
- **Merchandising recency:** the newest public product creation/publication timestamp observed and
  its age at run time.
- **Collection coverage:** the number of collections observed and whether that number is exact,
  lower-bound, or unavailable.
- **Evidence and provenance:** source type, source URL, observation timestamp, supported fields,
  records observed, and coverage semantics.
- **Decision support:** coverage score, confidence score, confidence reasons, data gaps,
  `recommendedAction`, action priority, and an explanation.
- **Stable identifiers:** deterministic website entity ID plus observation and event IDs for joins,
  deduplication, warehousing, and repeated checks.
- **Operational truth:** `OUTPUT` in the default key-value store records requested, processed,
  confirmed-delivered, withheld, duplicate, and advisory counts, plus partial/budget/replay state.

Existing integrations remain supported. The original `productCount`, `collectionsCount`,
`priceRange`, `topProductTypes`, `topVendors`, `newestProductAt`, `catalogComplete`,
`catalogTruncatedReason`, `salesSignals`, and `revenueEstimate` fields remain present. Their meaning
is now more explicit:

- `productCount` remains the number of products scanned and is mirrored by `productsScanned`;
- `collectionsCount` remains the number observed and is mirrored by `collectionsObserved`;
- `revenueEstimate` is retained only for backward compatibility and returns
  `band: "unknown"`, `confidence: "insufficient_evidence"` on a confirmed Shopify result;
- no public Dataset view presents revenue as a qualification metric.

### Who this is for

#### Small marketing agencies

Qualify a prospect list before spending time on manual storefront research. Segment confirmed
Shopify stores by observed catalog band, price positioning, leading product type, or recent public
merchandising activity. Use the evidence fields to prepare a specific account note rather than a
generic “we help ecommerce brands” opener.

Example defensible note: “The public storefront check observed a large catalog lower bound, a
leading footwear category, and a recently published product. Review the current site before using
those facts in outreach.”

Example unsupported claim: “The company is growing quickly, earns more than $10M, needs an agency,
and has budget now.” None of those conclusions follows from a public catalog.

#### Freelancers and boutique ecommerce consultants

Use observed catalog complexity and merchandising cadence to prioritize accounts that resemble the
stores you serve. A catalog with many categories or vendors may be relevant to feed management,
merchandising operations, creative production, taxonomy, localization, or storefront QA services.
It is still a research signal, not proof that a problem exists.

#### Lead generation and RevOps teams

Enrich an existing, lawfully obtained website list before routing records to a campaign. Stable
`entityId` values make website-level joins easier. `productCountMode`, evidence, and data gaps make
it possible to avoid treating lower bounds as exact totals. `safeToAutomate: false` keeps the result
out of an unreviewed auto-send path.

This Actor does not find personal emails, infer identities, or determine a lawful basis for contact.
Use your own compliance process and outreach policy.

#### Ecommerce software and service providers

Research the public shape of a merchant account before a discovery call. A catalog-tool vendor may
care about observed product volume and category diversity; a pricing consultant may care about
observed variant price distribution; a creative studio may care about recent product publishing.
Choose only fields relevant to your product and verify the live storefront before making a claim.

#### Portfolio and competitive researchers

Store results on a schedule to build your own history of public observations. The Actor reports the
current observation; it does not claim that a change occurred unless you compare observations. Use
`observationId`, `observedAt`, and `freshness` to keep lineage, then calculate diffs in your own
warehouse or use the related Shopify Price Change Monitor.

#### Developers and data teams

Use the Apify API, schedules, webhooks, datasets, and exports to add bounded Shopify storefront
evidence to a larger workflow. The result is JSON-first and intentionally includes operational
state in `OUTPUT` so a consumer can distinguish complete runs from budget stops or ambiguous
delivery failures.

### Practical use cases

#### 1. Segment a Shopify prospect list

Input a list of domains from your CRM or market research. Filter results where `isShopify` is true,
then group by `catalogSizeBand.label`, `pricePositioning.median`, `topProductTypes`, or
`merchandisingActivity.status`. Keep `productCountMode` beside every size filter so a lower bound is
never presented as an exact catalog total.

#### 2. Build a manual personalization queue

Route records with a relevant observed category, vendor mix, or recent product timestamp into a
research queue. Show the reviewer `evidence`, `confidenceReasons`, `dataGaps`, and the public URL.
The reviewer confirms that the signal remains visible and decides whether it belongs in the final
message.

#### 3. Prioritize ecommerce audits

Consultants can identify stores whose public catalog shape matches a service speciality. For
example, many observed product types can justify a taxonomy review; many observed vendors can
justify a feed-governance conversation; recent product publication can justify checking launch
workflows. These are hypotheses for an audit, not diagnosed problems.

#### 4. Enrich an account table

Join `entityId`, `isShopify`, `productsScanned`, `productCountMode`, `pricePositioning`, leading
vendors/types, and `checkedAt` to a company record. Preserve the full result separately so the
decision can be traced back to the observation and its limitations.

#### 5. Monitor a known merchant set

Schedule the same domain list weekly or monthly. Store snapshots and compare exact fields yourself.
Do not compare `productCount` without checking `productCountMode`: a complete count and a lower
bound are not interchangeable.

#### 6. Separate Shopify-specific and non-Shopify campaigns

A reachable website without a supported Shopify signature returns `isShopify: false` and an
explicit gap: absence of the supported signature is not proof of the site's complete ecommerce
architecture. Use that result to exclude it from a Shopify-specific batch or send it to manual
technology verification.

### How the pipeline works

![Shopify catalog intelligence workflow from public storefront and Shopify JSON feeds through bounded processing to evidence, confidence, gaps, and review actions](https://raw.githubusercontent.com/TimmyZinin/apify-actor-assets/7bac0e9d52219782d32a587479b3e15b6490ca5d/commercial115/shopify-store-intelligence/readme-workflow.webp)

1. **Validate input.** The runtime requires an object with 1–100 website strings. It rejects blank
   values, non-string items, unsupported schemes, embedded credentials, overlong URLs, and invalid
   concurrency instead of silently dropping or coercing them.
2. **Normalize and deduplicate.** Missing schemes become HTTPS, fragments are removed, and
   case-equivalent normalized URLs are deduplicated before processing and billing.
3. **Fetch the submitted storefront safely.** The Actor accepts HTTP(S), resolves DNS, blocks
   private/loopback/link-local/metadata targets, pins verified addresses against DNS rebinding, and
   repeats the check on every redirect hop.
4. **Check supported Shopify signatures.** The initial public HTML, final URL, and selected public
   response header names are checked for a narrow Shopify signature set.
5. **Read public Shopify feeds.** Confirmed stores are queried through their public
   `/products.json` and `/collections.json` endpoints. The product scan uses Shopify's 250-item page
   size and stops after three pages or the real end of the feed.
6. **Calculate observable statistics.** The Actor calculates price distribution from observed
   public variants, leading vendors/types from observed products, and merchandising recency from
   the newest observed product timestamp.
7. **Attach semantics and evidence.** Counts are labelled `exact`, `lower_bound`, `not_available`,
   or `not_applicable`; evidence records identify the source and supported fields; coverage and
   confidence explain how much can be used.
8. **Recommend review.** The result suggests a conservative next action and always sets
   `safeToAutomate: false`.
9. **Deliver with billing safeguards.** Paid results are linked to Dataset delivery. Remaining
   budget is checked inside a concurrency-safe lock. An ambiguous delivery/charge receipt fails the
   run with `replaySafe: false` rather than pretending success.
10. **Persist completion truth.** `OUTPUT` records counts and whether the run is complete, partial,
    budget-stopped, or unsafe to replay.

### Public data sources

| Source | What is observed | What it does not prove |
|---|---|---|
| Submitted storefront HTML/final URL/public header names | Supported Shopify signatures and storefront currency when exposed | Store ownership, merchant identity, plan, GMV, revenue, conversion, or internal configuration |
| `/products.json` | Public product records, public variant list prices, vendor/type values, product timestamps | Orders, units sold, inventory truth, margins, discounts actually paid, customer demand, or complete Admin API data |
| `/collections.json?limit=250` | Up to 250 public collection records | A guaranteed total when the cap is reached or the feed is unavailable |

The Actor does not use a Shopify Admin API token, merchant login, browser session, proxy by default,
or private customer data. A storefront may disable a feed, return a custom response, rate-limit the
request, or expose only part of its internal catalog. Those conditions are returned as gaps rather
than filled with guesses.

### Input

#### Input schema

| Field | Type | Required | Limits | Meaning |
|---|---|---:|---|---|
| `websites` | array of strings | yes | 1–100; each ≤2,048 characters | Public HTTP(S) websites such as `allbirds.com` or `https://example.com/` |
| `maxConcurrency` | integer | no | 1–20; default 5 | Number of websites processed concurrently |

#### Example input

```json
{
  "websites": [
    "allbirds.com",
    "gymshark.com",
    "stripe.com"
  ],
  "maxConcurrency": 5
}
```

Input is strict by design. `ftp://example.com`, `https://user:password@example.com`, non-string
array entries, empty arrays, and a `maxConcurrency` such as `"5"` fail validation. This protects
downstream billing and avoids the common failure mode where invalid entries disappear silently.

### Output

Successful and failed per-site rows are written to the default Dataset. A machine-readable run
summary is written to the default key-value store as `OUTPUT`.

#### Representative successful Shopify result shape

The values below demonstrate the current schema and count semantics. Live values change with the
public storefront and should always be read from your run.

```json
{
  "url": "https://www.allbirds.com/",
  "requestedUrl": "https://allbirds.com/",
  "found": true,
  "error": null,
  "httpStatus": 200,
  "isShopify": true,
  "productsAccessible": true,
  "productCount": 750,
  "productsScanned": 750,
  "productCountLowerBound": 750,
  "productCountMode": "lower_bound",
  "collectionsCount": 250,
  "collectionsObserved": 250,
  "collectionsCountMode": "lower_bound",
  "priceRange": {
    "min": 3,
    "max": 160,
    "currency": "USD"
  },
  "pricePositioning": {
    "min": 3,
    "max": 160,
    "average": 78.44,
    "median": 75,
    "currency": "USD",
    "variantPricesObserved": 1287,
    "basis": "public_product_variant_list_prices"
  },
  "topProductTypes": ["Shoes", "Apparel", "Socks"],
  "topVendors": ["Allbirds"],
  "newestProductAt": "2026-02-04T17:38:18-08:00",
  "merchandisingActivity": {
    "status": "no_recent_product_publication_observed",
    "newestProductAgeDays": 187,
    "basis": "newest_created_or_published_timestamp_in_observed_public_products"
  },
  "catalogComplete": false,
  "catalogTruncatedReason": "scan limit reached (3 pages / 750 products cap)",
  "catalogSizeBand": {
    "label": "500_plus",
    "basis": "observed_lower_bound",
    "productsScanned": 750
  },
  "revenueEstimate": {
    "band": "unknown",
    "confidence": "insufficient_evidence",
    "method": "not_estimated_from_public_catalog",
    "note": "Public catalog size and listed prices do not prove sales volume or revenue."
  },
  "evidence": [
    {
      "evidenceType": "shopify_site_signature",
      "sourceUrl": "https://www.allbirds.com/",
      "observedAt": "2026-08-11T10:00:00.000Z",
      "supports": ["isShopify"]
    },
    {
      "evidenceType": "public_shopify_products_feed",
      "sourceUrl": "https://www.allbirds.com/products.json",
      "observedAt": "2026-08-11T10:00:00.000Z",
      "supports": ["productsScanned", "pricePositioning", "topVendors", "topProductTypes", "newestProductAt"],
      "recordsObserved": 750,
      "coverage": "lower_bound"
    }
  ],
  "coverageScore": 65,
  "coverageBand": "medium",
  "confidenceScore": 70,
  "confidenceBand": "medium",
  "dataGaps": [
    "PRODUCT_CATALOG_SCAN_TRUNCATED_OR_INCOMPLETE",
    "COLLECTION_COUNT_IS_AN_OBSERVED_LOWER_BOUND_OR_UNAVAILABLE",
    "SALES_VOLUME_NOT_PUBLIC",
    "REVENUE_NOT_OBSERVED_OR_ESTIMATED"
  ],
  "recommendedAction": "USE_AS_LOWER_BOUND_AND_VERIFY_CATALOG_SCOPE_BEFORE_HIGH_VALUE_OUTREACH",
  "actionPriority": "medium",
  "safeToAutomate": false,
  "summary": "https://www.allbirds.com/%20%E2%80%94%20confirmed%20Shopify%20storefront;%20750+%20public%20products%20observed;%20observed%20prices%20USD%203%E2%80%93160;%20leading%20observed%20category:%20Shoes.%20Revenue%20is%20not%20inferred%20from%20public%20catalog%20data.",
  "checkedAt": "2026-08-11T10:00:00.000Z"
}
```

This example deliberately shows a truncated catalog. `productCount: 750` is preserved for older
integrations, while `productCountMode: "lower_bound"` and `catalogComplete: false` prevent it from
being presented as the merchant's exact product total. Numeric values in a live run may differ.

#### Non-Shopify result

A reachable page without a supported Shopify signature is still a completed, billable check:

```json
{
  "url": "https://stripe.com/",
  "requestedUrl": "https://stripe.com/",
  "found": true,
  "isShopify": false,
  "productsAccessible": false,
  "productCount": 0,
  "productsScanned": 0,
  "productCountMode": "not_applicable",
  "evidenceScope": "submitted_storefront_response",
  "confidenceScore": 80,
  "dataGaps": [
    "ABSENCE_OF_SUPPORTED_SIGNATURE_IS_NOT_PROOF_OF_ECOMMERCE_PLATFORM_IDENTITY",
    "SUBPAGES_NOT_SCANNED"
  ],
  "recommendedAction": "EXCLUDE_FROM_SHOPIFY_SPECIFIC_CAMPAIGN_OR_VERIFY_MANUALLY",
  "safeToAutomate": false
}
```

#### Failure result

If the site cannot be observed, the Actor emits a free failure row where pricing permits it. The
row does not pretend that the site is non-Shopify:

```json
{
  "requestedUrl": "https://unavailable.example/",
  "found": false,
  "error": "dns lookup failed: ENOTFOUND",
  "isShopify": false,
  "productCountMode": "not_available",
  "evidence": [],
  "coverageScore": 0,
  "confidenceScore": 0,
  "dataGaps": [
    "NO_SUCCESSFUL_STOREFRONT_OBSERVATION",
    "SHOPIFY_STATUS_UNKNOWN",
    "CATALOG_UNKNOWN"
  ],
  "recommendedAction": "RETRY_AFTER_SOURCE_RECOVERY",
  "retryable": true,
  "safeToAutomate": false
}
```

#### `OUTPUT` run summary

```json
{
  "actor": "shopify-store-intelligence",
  "status": "COMPLETE",
  "inputWebsites": 3,
  "requestedUniqueSites": 3,
  "duplicatesRemoved": 0,
  "processedSites": 3,
  "confirmedDeliveredResults": 3,
  "paidResultsConfirmed": 3,
  "freeFailureRows": 0,
  "withheldOrUnconfirmedResults": 0,
  "unprocessedSites": 0,
  "advisoryRows": 0,
  "partial": false,
  "budgetStopped": false,
  "replaySafe": true
}
```

When a linked Dataset delivery returns an ambiguous billing receipt, the Actor marks the run failed
and sets `replaySafe: false`. Inspect the Dataset and charged events before retrying, because a blind
rerun could duplicate a delivered result or charge.

### Field guide

| Field | Meaning |
|---|---|
| `url` | Final observed storefront URL after guarded redirects |
| `requestedUrl` | Normalized submitted URL before redirects |
| `found` | Whether a public storefront response was successfully observed |
| `isShopify` | Whether the supported Shopify signature set matched |
| `productCount` | Backward-compatible number of products scanned; not automatically an exact total |
| `productsScanned` | Number of public product records observed during this run |
| `productCountLowerBound` | Minimum catalog size supported by this observation |
| `productCountMode` | `exact`, `lower_bound`, `not_available`, or `not_applicable` |
| `collectionsObserved` | Public collections observed, up to the endpoint cap |
| `collectionsCountMode` | Coverage semantics for the collection count |
| `priceRange` | Backward-compatible observed min/max and currency |
| `pricePositioning` | Observed min/max/average/median, currency, variant-price count, and basis |
| `topProductTypes` | Leading public product-type names in the scanned records |
| `topProductTypeSignals` | Leading names with supporting observed-product counts |
| `topVendors` | Leading public vendor names in the scanned records |
| `topVendorSignals` | Leading names with supporting observed-product counts |
| `newestProductAt` | Newest observed `created_at` or `published_at` timestamp |
| `merchandisingActivity` | Recency interpretation and its explicit timestamp basis |
| `catalogComplete` | Whether pagination reached a real short/empty terminal page |
| `catalogTruncatedReason` | Why the catalog is incomplete; `null` when complete |
| `catalogSizeBand` | Descriptive observed-size band with exact/lower-bound basis |
| `revenueEstimate` | Compatibility field that explicitly declines to estimate revenue |
| `evidence` | Source-level provenance records without raw private headers or bodies |
| `entityId` | Stable website entity key |
| `observationId` | Identifier for this timestamped observation |
| `eventId` | Identifier for this store-check event |
| `freshness` | Live-observation timestamp semantics and cache status |
| `coverageScore` | 0–100 measure of how much relevant public catalog scope was observed |
| `confidenceScore` | 0–100 confidence in the bounded claims actually made |
| `confidenceReasons` | Plain-English reasons behind confidence |
| `dataGaps` | Material information the run did not observe |
| `recommendedAction` | Conservative next review step |
| `actionPriority` | Low, medium, or high review priority based on observable signals |
| `safeToAutomate` | Always `false`; human review is required before consequential action |
| `failureType` / `retryable` | Failure classification and retry guidance |
| `summary` | Human-readable result with the same bounded semantics |

### Count semantics that prevent bad segmentation

#### Exact

`productCountMode: "exact"` means the public product feed reached an empty or short terminal page
within the scan boundary. It is exact for the public feed observed at that time, not a guarantee
about unpublished, market-restricted, authenticated, draft, or Admin-only records.

#### Lower bound

`productCountMode: "lower_bound"` means at least that many public records were observed, but the end
of the catalog was not proven. This happens when the 750-product cap is reached or a later page
fails. Use `750+`, not `750`, in a human-facing segment label.

#### Not available

The site was Shopify-confirmed or failed, but the public feed did not yield usable product data. Do
not substitute zero; unknown and empty are different states.

#### Not applicable

The submitted page was reachable but did not match the supported Shopify evidence. A Shopify
catalog count does not apply to this observation.

### Confidence is claim-specific

A row can have high confidence that a site is Shopify and low catalog coverage at the same time.
For example, a storefront signature may be clear while `/products.json` is disabled. The Actor
therefore keeps `confidenceScore` separate from `coverageScore` and records the reasons and gaps.

Scores are deterministic descriptions of this Actor's observation scope. They are not audited
accuracy rates, probabilities that a merchant will buy, or guarantees about the storefront.

### Pricing and cost control

The current pay-per-event configuration is **$0.005 per Actor start** plus **$0.006 per completed
site check**. The live Apify pricing panel is authoritative if pricing changes after this README is
published.

At the current rate:

- 1 completed site check: approximately `$0.011` including one run start;
- 10 completed site checks in one run: approximately `$0.065`;
- 100 completed site checks in one run: approximately `$0.605`.

Unreachable failure rows are written without the useful-result event where platform pricing allows
a free explanation. A reachable non-Shopify result is still a completed check and is billable.
Shopify-confirmed rows with an unavailable public products feed are also completed checks because
the platform evidence and diagnostic result were delivered.

The Actor checks the buyer's remaining maximum charge before each paid delivery inside a shared
lock, preventing concurrent workers from independently passing the same budget check. If the budget
ends before all requested sites are delivered, `OUTPUT` reports a partial budget stop and a free
advisory row explains the remainder. No advisory is added when the last requested result itself
reaches the event limit, because the run is already complete.

### Integrations

#### Apify API with cURL

```bash
curl "https://api.apify.com/v2/acts/zinin~shopify-store-intelligence/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
  -X POST \
  -H "Content-Type: application/json" \
  -d '{"websites":["allbirds.com","gymshark.com"],"maxConcurrency":5}'
```

Keep tokens in your secret manager or Apify integration settings. Do not commit them to source
control or paste them into public logs.

#### JavaScript client

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('zinin/shopify-store-intelligence').call({
  websites: ['allbirds.com', 'gymshark.com'],
  maxConcurrency: 5,
});

const { items } = await client.dataset(run.defaultDatasetId).listItems();
for (const item of items) {
  console.log(item.requestedUrl, item.isShopify, item.productCountMode, item.recommendedAction);
}
```

#### Python client

```python
import os
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("zinin/shopify-store-intelligence").call(run_input={
    "websites": ["allbirds.com", "gymshark.com"],
    "maxConcurrency": 5,
})

for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item["requestedUrl"], item["isShopify"], item["recommendedAction"])
```

#### CSV or Excel

Run the Actor, open the default Dataset, and export CSV or Excel. For a compact review table, keep
`requestedUrl`, `isShopify`, `productsScanned`, `productCountMode`, `pricePositioning`,
`topProductTypes`, `topVendors`, `merchandisingActivity`, `confidenceScore`, `dataGaps`, and
`recommendedAction`.

#### Webhooks and schedules

Use an Apify schedule for recurring observations and a run-succeeded webhook to trigger your data
pipeline. Read `OUTPUT` before accepting the batch. A run status alone is not enough for a robust
consumer: confirm `status: "COMPLETE"`, `partial: false`, and `replaySafe: true`.

#### CRM workflow

Map `entityId` to your account record, store `checkedAt`, and add only the fields your campaign
needs. Preserve evidence and gaps in a linked research record. Place results in a review queue; do
not auto-send based only on `actionPriority` or `recommendedAction`.

### Responsible use and privacy

The Actor reads public website responses submitted by the user. It does not log in, bypass access
controls, collect checkout/customer records, or return personal contact data. You are responsible
for the website list, retention, outreach policy, contractual restrictions, and applicable law.

Recommended practices:

- submit domains you are permitted to research;
- use reasonable schedules rather than aggressive repeated polling;
- preserve the source timestamp and data gaps;
- verify facts on the current storefront before external communication;
- do not infer sensitive traits, revenue, financial health, or purchase intent;
- do not use a public vendor name as proof of a corporate relationship beyond the observed feed;
- do not present `lower_bound` counts as exact;
- keep `safeToAutomate: false` in consequential workflows.

### Limitations

- Shopify themes and storefront behavior vary. The supported signature set is intentionally narrow
  and can miss Shopify sites that hide or transform the relevant evidence.
- Only the submitted page response is checked for the storefront signature. Subpages and rendered
  client-side state are not exhaustively crawled.
- Product collection is capped at 750 public records per store. Larger catalogs return a lower
  bound.
- Collection observation is capped at 250 records.
- Public feeds may be disabled, rate-limited, customized, geographically varied, or different from
  Shopify Admin data.
- Variant prices are public list-price observations. They do not prove final checkout price,
  discount usage, currency conversion, taxes, shipping, margin, units sold, or realized revenue.
- `newestProductAt` is based on timestamps in observed public records. It is not a complete history
  of merchandising work and does not prove business growth.
- Top vendors and product types reflect the scanned sample. On a truncated catalog, unseen records
  may change the ranking.
- A non-Shopify result means supported evidence was not observed on this response; it is not a
  universal technology audit.
- Stable IDs identify the normalized website and observation, not a legal entity or beneficial
  owner.
- Scores describe evidence quality and coverage, not buying propensity or model accuracy.
- The Actor is not a full catalog exporter, contact finder, financial estimator, legal opinion, or
  autonomous outreach system.

### Related Actors

| Actor | When to use it |
|---|---|
| [Shopify Price Change Monitor](https://apify.com/zinin/shopify-price-change-monitor) | Compare repeated Shopify catalog observations and monitor candidate changes over time |
| [Website Contact Extractor](https://apify.com/zinin/website-contact-extractor) | Separately observe public business contact channels after you have a legitimate enrichment use case |
| [Tech Stack Detector](https://apify.com/zinin/tech-stack-detector) | Observe a broader set of public website technology signatures |
| [Website Tech Stack Change Detector](https://apify.com/zinin/tech-stack-change-detector) | Compare a buyer-supplied technology baseline with a fresh bounded page observation |
| [Zid Store Products Scraper](https://apify.com/zinin/zid-store-products) | Collect public product data from Zid storefronts rather than Shopify |

### FAQ

#### Does this require a Shopify API key or merchant login?

No. It observes the submitted public storefront and public Shopify JSON feeds where available. It
does not access Shopify Admin data.

#### Does it estimate revenue?

No. The legacy `revenueEstimate` field remains only for compatibility and explicitly reports
`unknown / insufficient_evidence`. Catalog size and list prices do not prove units sold or revenue.

#### Is `productCount` the exact catalog size?

Only when `productCountMode` is `exact` and `catalogComplete` is true—and even then it is exact for
the public feed observed at that moment, not private Admin records. When the mode is `lower_bound`,
display the number with a plus sign.

#### Why is a non-Shopify result billable?

The Actor completed the requested public check and delivered a supported answer with evidence scope,
confidence, gaps, and a next action. Billing is for the check, not for a positive Shopify match.

#### Why can a Shopify-confirmed row have low coverage?

The storefront can expose a strong Shopify signature while disabling or limiting the public product
feed. Confidence in platform detection and coverage of catalog details are different questions.

#### Does it collect email addresses or owner names?

No. It focuses on store, catalog, price, vendor, category, and merchandising observations. Use a
separate, compliant contact-enrichment workflow if needed.

#### Can it scan more than 750 products?

Not in the current bounded product. The cap keeps requests predictable and affordable. A result that
reaches the cap clearly reports a lower bound and truncation reason.

#### Can I use the result to send automated outreach?

The output intentionally says `safeToAutomate: false`. Use it to prioritize and support human
research, then apply your own legal, quality, and relevance checks.

#### What should I do with duplicates?

Normalized case-equivalent URLs are deduplicated before work and billing. `OUTPUT.duplicatesRemoved`
reports how many submitted values were removed as duplicates.

#### What happens if my run budget is too low?

The Actor stops before an unaffordable paid delivery. `OUTPUT` reports partial status and counts;
when the free Dataset explanation channel is available, an advisory row explains how many results
were confirmed and recommends resubmitting only undelivered websites.

#### What does `replaySafe: false` mean?

The Actor could not prove the delivery/charge outcome. Inspect the Dataset and run charging events
before retrying. A blind rerun may duplicate data or charges.

#### Can an AI agent call it?

Yes. Use the Apify API or Apify MCP integration, but instruct the agent to inspect confidence,
coverage, data gaps, and `safeToAutomate` rather than treating the summary as an autonomous decision.

#### How do I report an incorrect result?

Use the issue/contact option on the Actor page with the submitted public URL, run ID, timestamp, and
the field you believe is incorrect. Do not include credentials or private customer data.

***

Built by [zinin](https://apify.com/zinin). Questions? Telegram
[@timzinin](https://t.me/timzinin).

# Actor input Schema

## `websites` (type: `array`):

Public HTTP(S) websites to check (for example `allbirds.com`). Invalid values, credentials, and non-HTTP schemes fail validation instead of being silently dropped.

## `maxConcurrency` (type: `integer`):

How many websites to check in parallel.

## Actor input object example

```json
{
  "websites": [
    "allbirds.com",
    "gymshark.com",
    "stripe.com"
  ],
  "maxConcurrency": 5
}
```

# Actor output Schema

## `results` (type: `string`):

API URL for the default dataset items produced by this run.

## `runSummary` (type: `string`):

API URL for OUTPUT: requested, processed, confirmed-delivered, withheld and advisory counts, plus budget and replay-safety status.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "allbirds.com",
        "gymshark.com",
        "stripe.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("zinin/shopify-store-intelligence").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "websites": [
        "allbirds.com",
        "gymshark.com",
        "stripe.com",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("zinin/shopify-store-intelligence").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "allbirds.com",
    "gymshark.com",
    "stripe.com"
  ]
}' |
apify call zinin/shopify-store-intelligence --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,zinin/shopify-store-intelligence"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/a6d9qF20wSYul2CSQ/builds/5xT9cemn2NlvlEyFv/openapi.json
