Similarweb Scraper - Traffic, Audience & Competitors avatar

Similarweb Scraper - Traffic, Audience & Competitors

Pricing

from $1.50 / 1,000 base data

Go to Apify Store
Similarweb Scraper - Traffic, Audience & Competitors

Similarweb Scraper - Traffic, Audience & Competitors

Get traffic & competitor intelligence for any website - ranks, monthly visits, engagement, traffic sources, top keywords, AI-referral traffic. Complete profile adds audience age & gender, referring domains, ad publishers, technologies, ranked competitors & rank history. No login. CSV, Excel, JSON.

Pricing

from $1.50 / 1,000 base data

Rating

0.0

(0)

Developer

Kelopr_bk

Kelopr_bk

Maintained by Community

Actor stats

0

Bookmarked

104

Total users

49

Monthly active users

3 days ago

Last modified

Share

Similarweb Scraper

Similarweb Scraper β€” traffic, competitors, AI visibility & monitoring

Explore website traffic, audiences and competitors from one list of domains. Collect growth, ranks, acquisition channels, keywords, AI referrals, WHOIS, company signals and homepage technologies. Compare websites, track changes between runs or build a competitive market map.

No Similarweb login or browser session is required. Existing API integrations using base_data, similar_sites, aitdk, or all remain compatible.

▢️ Your first run

  1. Open Input and select Website overview. This includes traffic sources, engagement, ranks, top countries, search keywords, and available AI-referral data.
  2. Under Domains, add the websites you want to analyze, one per entry. A domain such as apify.com or a full HTTP(S) website URL is accepted. Duplicate websites are removed automatically.
  3. Leave the other settings at their defaults and click Start. For a first run, use 1–3 domains.
  4. Open Output β†’ Results β†’ Overview for eight headline metrics. Switch to Traffic, Channels, Countries or Keywords for details.
  5. Use Export to download the data. Choose All fields when you want the full record. JSON preserves nested objects; list cells in the table expand to show their entries.

Copy this into Input β†’ JSON for a first run:

{
"domains": ["python.org", "docker.com"],
"mode": "base_data"
}

Choose Complete intelligence profile to add competitor discovery, domain analysis, demographics, referrals and rank history. Each website produces one main result. The countries, keywords and other lists inside that record are included in its collection event.

πŸš€ Pick your goal

WorkflowBest forWhat you receiveCollection event
base_dataFast website overviewTraffic history, growth, ranks, engagement, channels, geography, keywords, AI visibility, screenshotbase-data
similar_sitesCompetitor discoverySimilar websites, similarity grades, estimated visits, categories, tags, related appssimilar-sites
aitdkDomain due diligenceWHOIS/RDAP, domain age, expiry, homepage keyword density, company metadata, social links, detected technologieswhois-keywords
allComplete profileTraffic, competitor discovery and domain insights, plus in-depth audience, referrals, ads, technologies and rank historyall-in-one
compareSide-by-side benchmarkActual traffic shares, leaders, ranks, engagement position, and gaps versus the first domainbase-data
monitorRecurring intelligenceSaved baseline and material changes in traffic, ranks, engagement, acquisition shares, keywords and available AI referralsbase-data or all-in-one
market_mapCompetitive landscapeSeed sites, auto-discovered competitors, enriched traffic rows, market shares, and a separate relationship datasetbase-data

Current rates are listed in the Actor's Pricing tab. Successful and partial records use the workflow's collection event. Failed lookups go to Errors without a result-event charge; the platform's Actor start event still applies. Monitoring checks are charged even when onlyChanges filters their rows from the output.

🧠 In-depth audience & referral intelligence β€” included in the Complete profile

The Complete intelligence profile (all) adds these details within the same all-in-one result event:

GroupFields
AudienceAge distribution, largest age group, male/female split, audience topics, the other sites that audience visits
Competitive setSimilarweb's own ranked competitors with category rank and affinity, plus the websites immediately above and below in the global ranking
Referral graphIncoming referring domains and outgoing destinations with visit shares, category splits, and the real totals behind both
Paid & socialAdvertising publishers and network counts, and the split of social traffic per network
TechnographicsTechnology providers by category, and how many technologies are detected in each
FirmographicsLegal name, year founded, employee range, headquarters, revenue range, parent domain β€” when Similarweb publishes them
HistoryGlobal, category, and country rank month by month, movement versus the previous month, and traffic by country
PreviewsDesktop and mobile preview images

Select all with includeInDepth: true to collect this layer. Set the flag to false for a lighter Complete profile. Other workflows, including Monitor and Market map, return their own workflow data.

indepthStatus reports whether the profile was collected. If it is unavailable, the website's available data is still returned under the same collection event. indepthWithheldRows counts entries for which Similarweb did not disclose a domain; the returned lists contain the identifiable websites.

{
"domains": ["python.org", "docker.com"],
"mode": "all",
"includeInDepth": true
}

✨ What makes the output comfortable

  • Flat columns where they matter: latestMonthlyVisits, trafficChangePercent, bounceRatePercent, organicSearchPercent, and dozens more export cleanly to CSV and Excel.
  • Original objects remain: existing consumers can still use engagement, estimatedMonthlyVisits, trafficSources, countryRank, and other legacy fields.
  • Four output destinations: Results, Market links, Errors and Run summary. Topic tabs inside Results show the data for your chosen workflow.
  • Results appear as they finish: independent website profiles are saved immediately. Compare and Market Map calculate their cohort-wide shares after collecting the bounded comparison group.
  • Clean failures: no error rows mixed into paid result counts or downstream exports.
  • Clear finish state: every run writes an OUTPUT JSON summary with exact result, error, filter, duration, and charged-event counts.
  • Honest enrichment: unavailable values stay empty. The Actor does not invent demographics, traffic, technologies, or company data.

πŸ“Š Website overview fields

The base_data workflow returns the source-compatible objects plus useful derived fields.

Identity and rank

domain, websiteUrl, similarwebUrl, siteName, title, description, category, snapshotDate, screenshot, globalRank, countryRank, countryCode, countryRankValue, categoryRank, categoryRankValue, marketPosition

Traffic and engagement

engagement, estimatedMonthlyVisits, trafficHistory, latestMonthlyVisits, previousMonthlyVisits, trafficChangeMoM, trafficChangePercent, threeMonthGrowthPercent, trafficTrend, bounceRatePercent, pagesPerVisit, timeOnSiteSeconds

MetricMeaning and units
latestMonthlyVisitsEstimated visits in the latest available month, not unique visitors
engagement.month, engagement.yearPeriod of the engagement metrics
bounceRatePercentBounce rate on a 0–100 scale; engagement.bounceRate preserves the original 0–1 fraction
pagesPerVisitAverage pages per visit, not total pageviews
timeOnSiteSecondsAverage visit duration in seconds
trafficChangePercentCalculated month-over-month change; trafficChangeMoM retains the fractional form

Use the dates in trafficHistory and snapshotDate when comparing periods. An unavailable metric stays empty; a real zero stays zero.

Acquisition channels

trafficSources, directPercent, organicSearchPercent, paidSearchPercent, organicSocialPercent, paidSocialPercent, referralsPercent, mailPercent, displayAdsPercent, affiliatePercent, aiTrafficPercent, topTrafficChannel, topTrafficChannelPercent, organicPaidRatio, trafficConcentrationScore, acquisitionDiversityScore

The Channels view separates organic and paid search, organic and paid social, direct visits, referrals, email, display advertising, affiliates, and AI referrals. Flat *Percent fields use 0–100; values inside trafficSources use 0–1. Referral share measures the channel's contribution, not a list of referring URLs.

Geography, search, AI, and competitors

topCountryShares, topCountryCode, topCountrySharePercent, topKeywords, topKeyword, topKeywordVolume, aiTraffic, aiChatbotDistribution, aiTotalVisits, estimatedAiVisits, topAiPlatform, topAiPrompts, competitors, competitorCount, topCompetitor

Countries shows each available top country with countryCode, countryName, countryId, share, and sharePercent inside topCountryShares. Names are taken from the source's country dictionary. These are the leading countries, so their shares need not add up to 100%.

countryRank.countryCode / countryCode identify the country used for the country rank. topCountryCode identifies the first country in the audience breakdown. They can legitimately differ.

Keywords shows the leading term and an expandable topKeywords list with name, volume, cpc, and estimatedValue, as supplied by the source. Search-channel shares are split into organic and paid; the keyword list itself is not classified as OrganicKeywords or PaidKeywords. A CPC value does not prove that the website pays for that keyword. keywordDensity in Domain SEO & WHOIS is a separate analysis of words on the homepage, not search traffic.

For competitor discovery, use similar_sites or all: their similarSites lists include descriptions, estimated visits, categories and similarity signals. Complete profile also supplies Similarweb's ranked competitive set in similarwebCompetitors and uses it to populate competitors when the basic list is empty. Similarity is a discovery signal, not a guarantee that every listed site is a direct business competitor.

In Demographics, audienceMaleShare and audienceFemaleShare use a 0–1 scale: 0.4628 means 46.28%. Fields ending in Percent, such as topCountrySharePercent, use a 0–100 scale.

Illustrative country entry:

{
"countryCode": "US",
"countryId": 840,
"countryName": "United States",
"share": 0.2543,
"sharePercent": 25.43
}

Quality signals

dataCompletenessScore, riskFlags, isSmall, status

The legacy isDataFromGoogleAds key is retained for API compatibility only. It mirrors the source's IsDataFromGa flag and must not be used as evidence of Google Ads spending; use paidSearchPercent for paid-search traffic share.

riskFlags are deterministic signals such as traffic_decline, single_channel_dependence, high_bounce_rate, domain_expiring_soon, and limited_public_data. They are not an opaque AI score.

🧭 Competitor discovery fields

The similar_sites workflow returns:

{
"domains": ["docker.com", "notion.so"],
"mode": "similar_sites"
}

domain, title, description, category, categoryRank, totalVisits, thumbnail, screenshot, favicon, tags, similarSites, relatedApps, status

Each similarSites item can contain:

{
"site": "peer-site.test",
"description": "Example competitor website",
"category": "Business_and_Consumer_Services",
"similarityRank": 1,
"topCountryRank": 2400,
"totalVisits": 1850000,
"grade": 0.91,
"thumbnail": "https://example.test/example-preview.png"
}

Result fragments are illustrative. The runnable input examples use real websites.

πŸ” Domain, company & technology fields

The aitdk workflow combines domain registration data with analysis of the public homepage. Open Results β†’ Website to inspect it.

{
"domains": ["python.org", "docker.com"],
"mode": "aitdk"
}
  • whois: registrar, status, nameservers, DNSSEC, registration, expiration, and last-change dates
  • domainAgeDays, daysUntilDomainExpiration
  • keywordDensity: page title, visible token count, top tokens, counts, and frequencies
  • websiteIntelligence: page title, meta description, language, canonical URL, page image, company metadata, public social profiles, generator, and technology signals
  • companyName, companyLogo
  • technologies, technologyCount

Technology detection is based on visible homepage signatures and is labelled as detected or declared; it is not presented as a guaranteed full technology stack. A company logo is returned only when structured organization data explicitly identifies it as a logo β€” a generic Open Graph hero image is not mislabeled.

WHOIS fields are filled from the available registry data. When a registry record is unavailable, other collected details remain in the result and notes explains the missing part.

βš”οΈ Compare websites

Put the primary website first. Open Results β†’ Compare to see each site's share, position and gap against that primary website.

{
"domains": ["nytimes.com", "theguardian.com", "cnn.com"],
"mode": "compare",
"maxConcurrency": 10
}

Comparison fields include:

primaryDomain, compareRank, trafficSharePercent, trafficGapVsPrimaryPercent, globalRankGapVsPrimary, growthRank, engagementRank, competitivePositionScore, winnerMetrics

Illustrative result fragment:

{
"domain": "competitor-a.test",
"mode": "compare",
"primaryDomain": "primary-site.test",
"compareRank": 2,
"trafficSharePercent": 31.42,
"trafficGapVsPrimaryPercent": -38.7,
"growthRank": 1,
"engagementRank": 3,
"winnerMetrics": ["growth"]
}

πŸ”” Monitor changes

Run the same domains with the same Monitor name (monitorName) to compare against the previous successful check. Open Results β†’ Changes for the outcome. The older API spelling monitorKey is also accepted; use monitorName in new inputs.

{
"domains": ["python.org", "docker.com"],
"mode": "monitor",
"monitorName": "weekly-competitors",
"monitorDepth": "base_data",
"changeThreshold": "medium",
"onlyChanges": false
}

The first run returns changeStatus: "baseline". Later runs return changed or unchanged and can include:

checkedAt, previousCheckedAt, changeSeverity, changedMetrics, recommendedAction

Sensitivity presets:

  • low: catches smaller movements
  • medium: recommended default
  • high: reports only major movements

monitorDepth: "base_data" returns traffic data. monitorDepth: "all" also returns competitor discovery, WHOIS and homepage insights as context; in-depth demographics and referrals belong to the standalone Complete profile workflow. changedMetrics compares traffic, ranks, engagement, channel shares, search keywords and the available competitors and AI-platform lists. The discovery list similarSites, WHOIS and homepage insights remain contextual fields in the returned record.

Set onlyChanges: true to omit unchanged rows. For example, checking two unchanged websites produces zero result rows and two paid monitoring checks. Run summary reports the checks in filteredUnchanged and details.unchanged.

πŸ•ΈοΈ Build a market map

Market Map starts with one or more seed domains, discovers related sites, deduplicates them, enriches the strongest candidates, and calculates share inside the returned landscape.

{
"domains": ["docker.com"],
"mode": "market_map",
"marketMapMaxCompetitors": 2,
"maxConcurrency": 10
}

The main dataset contains site rows with:

marketRole, sourceSeedDomains, similarityGrade, similarityRank, marketTrafficSharePercent, marketRank, emergingCompetitor

Open Results β†’ Market for the website rows. The separate Market links destination contains one edge per returned seed-to-competitor connection:

{
"seedDomain": "seed-site.test",
"competitorDomain": "peer-site.test",
"similarityRank": 1,
"grade": 0.91,
"competitorVisits": 1850000,
"competitorGlobalRank": 43210,
"trafficChangePercent": 8.4,
"marketTrafficSharePercent": 17.6,
"category": "Business_and_Consumer_Services"
}

πŸ“₯ Input reference

FieldTypeDefaultNotes
domainsarray of stringsrequiredDomains or full URLs; normalized, deduplicated, and validated
modestringbase_dataOne of the seven workflows above
includeInDepthbooleantrueCollects the audience, referral, ads, technology, and rank-history layer in the all workflow, at no extra charge. Set to false for a faster, lighter Complete profile
maxDomainsinteger10000Safety limit for unexpectedly large pasted lists
monitorNamestringmy-monitorReuse the same name on later checks; monitorKey remains an accepted API alias
monitorDepthstringbase_dataall adds competitor and domain details to monitoring results; in-depth audience data is collected by standalone all
changeThresholdstringmediumlow, medium, or high
onlyChangesbooleanfalseHide unchanged monitor rows after evaluation
marketMapMaxCompetitorsinteger10Up to 100 unique discovered sites
maxConcurrencyinteger101–50 websites processed in parallel
proxyConfigurationobjectActor defaultOptional connection configuration; use the default in Apify Console

πŸ“¦ Outputs and table views

The output menu has four destinations. Choose Results to browse website data, then switch topics using the tabs above the table. Overview starts with eight headline metrics; detailed fields remain available in the topic views and dataset export.

DestinationContains
ResultsWebsite records and their topic views
Market linksSeed-to-competitor relationships from Market map
ErrorsFailed lookups, stored separately and never charged as results
Run summaryCounts, timing, filters and workflow details

Within Results, use these topic tabs. Start with Competitors for competitor discovery and Website for Domain SEO; traffic metrics belong to workflows that collect traffic data.

OutputContains
OverviewWebsite, monthly visits, growth, global rank, leading country, top channel, bounce rate and status
TrafficTraffic history, growth and engagement
ChannelsSource percentages, concentration, and diversity
CountriesLeading country plus country names, codes and audience shares
KeywordsLeading term plus keyword text, volume, CPC and source-estimated value
AI trafficAI share, chatbot distribution and available prompts
CompetitorsSimilar websites, direct competitors, tags, apps, and Similarweb's ranked competitive set
DemographicsAge and gender, audience topics and other visited sites
ReferralsIncoming referrers, outgoing destinations and totals
Ads & techAdvertising publishers, social networks and technology providers
RanksGlobal, country and category rankings and their history
WebsitePreview, category, company, WHOIS, technologies and collection details
CompareLeaders, shares and gaps between compared websites
ChangesBaselines and material changes from Monitor
MarketEnriched seed and competitor websites

Demographics, Referrals, Ads & tech and rank history use the Complete profile data. Compare, Changes and Market show their corresponding workflow results. Switching views does not collect data or create additional charges. Export the full dataset to keep all original fields; display presets do not remove stored values.

πŸ“€ Website overview example

Illustrative result fragment showing the field structure:

{
"domain": "demo-site.test",
"mode": "base_data",
"snapshotDate": "2026-07-01T00:00:00+00:00",
"globalRank": 18420,
"estimatedMonthlyVisits": {
"2026-05-01": 220000,
"2026-06-01": 238000,
"2026-07-01": 245000
},
"latestMonthlyVisits": 245000,
"previousMonthlyVisits": 238000,
"trafficChangePercent": 2.94,
"trafficTrend": "stable",
"bounceRatePercent": 41.2,
"pagesPerVisit": 3.84,
"timeOnSiteSeconds": 176.5,
"directPercent": 31.0,
"organicSearchPercent": 46.0,
"aiTrafficPercent": 2.1,
"topTrafficChannel": "Organic search",
"topCountryCode": "US",
"topKeyword": "demo analytics",
"topAiPlatform": "chatgpt.com",
"dataCompletenessScore": 92,
"riskFlags": [],
"status": "ok"
}

Traffic values are estimates tied to the source snapshot date, not real-time analytics or exact first-party measurements.

πŸ”Œ API examples

Set the APIFY_TOKEN environment variable to your Apify API token. These examples use Bash syntax.

Start a run:

curl -X POST "https://api.apify.com/v2/acts/trakk~similarweb-scraper/runs" \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"domains":["python.org","docker.com"],"mode":"base_data","maxConcurrency":2}'

Run synchronously and receive dataset items:

curl -X POST "https://api.apify.com/v2/acts/trakk~similarweb-scraper/run-sync-get-dataset-items" \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"domains":["nytimes.com","theguardian.com","cnn.com"],"mode":"compare"}'

For larger jobs, start an asynchronous run and consume the output URLs returned with the run. Dataset items can be exported as JSON, JSONL, CSV, Excel, XML, HTML, or RSS.

βš™οΈ Following a run

  • Completed website profiles appear in the dataset during collection.
  • Compare and Market map calculate shares across the collected group before saving their final rows.
  • Run summary shows saved results, partial records, errors, filtered checks and elapsed time.
  • ok identifies a successful collection; partial keeps the available data and explains missing parts in notes. Failed lookups appear in Errors.

Start with the default 512 MB run configuration. Duration depends on workflow depth, the number of websites and source response times.

The run summary exposes validDomains, scheduledDomains, skippedByChargeLimit, and details.chargeLimitReached. If a budget limits a comparison or market map, its shares describe the collected group, not the websites that were not processed.

❓ FAQ

Do old inputs still work?

Yes. base_data, similar_sites, aitdk, and all keep their existing values and legacy nested fields. The new flat fields are additive.

Why is a field empty?

Public coverage varies by domain, size, source, and snapshot. Optional fields remain null, empty arrays, or empty objects instead of being guessed.

What happens if I enter something incorrectly?

Invalid website entries are skipped with a warning. If none remain, or an input setting is invalid, the Actor exits normally with zero results and an explanation in Run summary (status: invalid_input, message, actionNeeded), without requesting website data. Apify's input form can also reject incorrectly typed settings before a run starts. Diagnostic messages are not inserted into the paid results dataset.

Does a country or keyword view create extra billable results?

No. These views display arrays already stored inside a website result. They do not create new dataset items or trigger extra collection events.

Are failed domains mixed into my export?

No. Valid and partial results go to the default dataset. Failures go to Errors with a short reason such as not_found, no_data, or source_unavailable; they are not charged as collection results.

Does onlyChanges make unchanged monitoring checks free?

No. It filters unchanged rows from the result dataset after the Actor checks the source and compares the saved baseline. The run summary tells you exactly how many checks were filtered.

What does in-depth data cost me?

Nothing beyond the Complete profile's own result event. Websites with no in-depth profile come back with indepthStatus: "unavailable" and a note, and are charged exactly like any other row.

Why did in-depth data stop partway through my run?

Open Run summary and inspect details.inDepth.stoppedBecause. If the reason is the run's spending limit, increase Maximum charge per run for the next collection. If a profile is unavailable from the source, review the returned notes and retry that website later. Other collected website data remains available.

Why are some referrers missing from the in-depth lists?

Similarweb withholds part of those lists from anonymous visitors and ships the withheld entries with an empty domain. Exporting them would put blank rows in your spreadsheet, so they are dropped and counted in indepthWithheldRows. The incomingReferrerCount and outgoingReferrerCount totals still reflect the full figure the source reports.

Is technology detection definitive?

No. It reports signatures present on the public homepage. Server-side and hidden technologies may not be visible.

Which examples can I run?

The input examples use real websites and can be copied into Input β†’ JSON. The result fragments use illustrative values and reserved .test domains to explain the field structure. A live run returns the latest data available from its sources.