# ImportYeti Scraper - US Customs Importer & Supplier Data (`alwaysprimedev/importyeti-scraper`) Actor

Pull structured US sea-shipment, importer, and supplier records from ImportYeti.com. Returns company addresses, country-of-origin breakdowns, top HS codes, monthly shipment history, trademarks, and tags.

- **URL**: https://apify.com/alwaysprimedev/importyeti-scraper.md
- **Developed by:** [Always Prime](https://apify.com/alwaysprimedev) (community)
- **Categories:** Automation, Developer tools, Lead generation
- **Stats:** 18 total users, 3 monthly users, 93.9% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.50 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## 🚢 ImportYeti Scraper — US Customs Importer & Supplier Data

[![Apify Actor](https://img.shields.io/badge/apify-actor-blue)](https://apify.com)
[![Python](https://img.shields.io/badge/python-3.11-blue)](https://www.python.org/)
[![Output](https://img.shields.io/badge/output-JSON%20%7C%20CSV%20%7C%20Excel-green)]()

> 🤖 **Structured supply-chain intelligence in seconds.** Pull every public datapoint from ImportYeti — importer addresses, country-of-origin breakdowns, monthly shipment history, top HS codes, trademarks and tags — straight into a JSON / CSV / Excel dataset.

### ✨ Why this scraper

- ⚡️ **Fast** — concurrent extraction, 5 records / second under default settings
- 📦 **Complete** — every visible field from each company/supplier page in one flat record
- 🏷️ **Both sides of the trade** — US importer "companies" *and* their foreign "supplier" counterparts
- 🌍 **Country-of-origin rollups** — pre-aggregated per country & continent
- 📈 **Monthly history** — shipments, weight (kg) and TEU broken down per month
- 🔁 **Incremental mode** — `since:` parameter skips records older than your last run
- 🛡️ **Cloudflare-resilient** — browser-fingerprint impersonation + your own residential/mobile proxy (via env var)
- 🧾 **3 export formats** — JSON, CSV, Excel

### 🚀 Quick start

1. Click **Try for free**
2. Type a search term (a company name, brand, product, or keyword) — or paste a list of `/company/<slug>` / `/supplier/<slug>` URLs
3. Hit **Start** and grab a coffee
4. Download the result as JSON, CSV, or Excel from the **Storage** tab

### 🛠️ Input

| Field | Type | Description |
|---|---|---|
| `query` | string | Free-text search term (brand, importer, HS-code keyword) |
| `entityType` | enum | `company` (US importer), `supplier` (foreign manufacturer), or both |
| `startUrls` | string\[] | Optional explicit list of `/company/<slug>` or `/supplier/<slug>` URLs |
| `maxItems` | integer | Hard cap on records; `0` = unlimited (default `50`) |
| `since` | datetime | Skip records whose most recent shipment is older than this ISO timestamp |
| `concurrency` | integer | Parallel HTTP requests (default `5`, max `25`) |
| `fetchDetails` | boolean | `true` (default) enriches each result from its detail page. `false` = **list-only mode**: search-surface fields only, **no detail-page traffic** (records flagged `partial: true`). |

> 🔌 **Proxy:** set via the `PROXY_URL` (or `HTTPS_PROXY` / `HTTP_PROXY`) **environment variable** on the actor — every request is routed through it. A **residential / mobile** proxy is recommended (ImportYeti is behind Cloudflare and blocks datacenter IPs). Format: `http://user:pass@host:port`.

### 💸 Saving proxy traffic

Residential/mobile proxies bill by the gigabyte, and detail pages are the cost driver. The actor minimises bytes for you, and you can trade depth for cost:

| Lever | How | Effect |
|---|---|---|
| **RSC fetch** (automatic) | Detail pages are pulled as the Next.js RSC payload, brotli-compressed | **~34 KB/record** on the wire vs ~60 KB for full HTML — ~44% less, no data loss |
| **List-only mode** | `fetchDetails: false` | Skips detail pages entirely — **~99% less** traffic; keeps name, address, country, totals, trademarks |
| **Incremental** | `since: <last run ISO>` | Old records are filtered from the search seed **before** any detail fetch — repeat runs only pay for new data |
| **Cap volume** | `maxItems`, narrower `query`/`entityType` | Fewer detail fetches = fewer bytes |

> Tip for recurring monitoring: run **`fetchDetails: false`** to list what's changed cheaply, then re-run **`fetchDetails: true`** with `startUrls` for just the records you care about.

### 📤 Sample output

```json
{
  "url": "https://www.importyeti.com/company/apple",
  "id": "company/apple",
  "scraped_at": "2026-05-10T20:00:10Z",
  "type": "company",
  "title": "Apple",
  "address": "568 Aldi Blvd, Mount Juliet, Tn 37122, Us",
  "country_code": "US",
  "phone_number": "XXX-XXX-X000",
  "website": "apple.com",
  "other_names_count": 2,
  "other_addresses_count": 89,
  "most_recent_shipment": "01/21/2026",
  "total_sea_shipments": 2449,
  "total_shipping_cost": 101746.81,
  "shipping_cost_coverage": 2.69,
  "avg_teu_per_month": 0.39,
  "avg_teu_per_shipment": 0.77,
  "database_updated": "05/06/2026",
  "multi_address": true,
  "dataset": "us",
  "uflpa": false,
  "location": {
    "state": "Tennessee",
    "county": "Wilson County",
    "city": "Mount Juliet",
    "district": "Mount Juliet"
  },
  "tags": [
    { "tag": "puter",    "shipments": 617 },
    { "tag": "computer", "shipments": 617 },
    { "tag": "lithium",  "shipments": 388 }
  ],
  "trademarks": [{ "name": "Apple", "trademarks_count": 1620 }],
  "imports_per_country": [
    { "country": "China",       "continent": "Asia",   "shipments": 2292 },
    { "country": "Hong Kong",   "continent": "Asia",   "shipments": 62 },
    { "country": "Vietnam",     "continent": "Asia",   "shipments": 37 },
    { "country": "Germany",     "continent": "Europe", "shipments": 11 }
  ],
  "top_hs_codes": ["8504.40","8517.62","8471.30","8544.42","8473.30"],
  "shipments_time_series": [
    { "period": "01/01/2024", "shipments": 32, "weight": 89421, "teu": 84 },
    { "period": "01/02/2024", "shipments": 41, "weight": 117850, "teu": 102 }
  ]
}
```

### 💼 Use cases

| Who | What for |
|---|---|
| **Procurement teams** | Find alternative suppliers for components shipped from specific countries |
| **B2B sales** | Build prospect lists by HS code, trade volume, or country of origin |
| **Market research** | Quantify competitor sourcing strategies — who buys what, from where, how often |
| **Investment analysts** | Track import volume as a leading indicator for retail / consumer-goods companies |
| **Logistics & freight** | Identify high-volume lanes and consolidation opportunities |
| **Trade compliance** | Screen suppliers against UFLPA and other forced-labour flags |

### 💡 Tips & tricks

- **Free-text queries** match titles, addresses and aliases — `"apple"` returns 27 hits across companies and suppliers, not just Apple Inc.
- For **company-only** datasets, set `entityType: "company"`. Suppliers are foreign manufacturers and have a different shape (no trademarks, often no website).
- The **`since` filter** uses the *most recent shipment* date on each record — perfect for nightly incremental runs.
- **Want comparable competitors?** Run one query per company name, then merge — each result carries a `tags` array you can use for affinity clustering.
- Large runs (10k+ records) benefit from `concurrency: 10` and unlimited `maxItems: 0`.

### ❓ FAQ

**Q: What data is in each record?**
A: Every field shown in the sample above — importer name, primary and alternate addresses, phone, website, total sea-shipment count, country-of-origin breakdown, top HS codes, monthly time series, trademarks and keyword tags.

**Q: How fresh is the data?**
A: ImportYeti's database update timestamp travels with each record in the `database_updated` field, so you always know how recent the underlying customs data is.

**Q: Can I get individual shipment-level records?**
A: This actor returns the public aggregated view (count + breakdowns + monthly history). Individual bill-of-lading records aren't part of the public page surface.

**Q: Will this work for non-US companies?**
A: Yes — the dataset itself is US sea imports, so US-side companies are *importers* and foreign-side companies are *suppliers*. Both show up.

**Q: How are records de-duplicated?**
A: By the canonical `id` field (`company/<slug>` or `supplier/<slug>`) — the same company appearing on multiple search pages or via multiple `startUrls` is pushed once.

**Q: What happens if a search slug doesn't have a real page?**
A: ImportYeti's search occasionally surfaces orphan slugs. The actor detects and silently skips them — they don't count against your run.

### 📊 Output

Three checkmarks live in the Storage tab — the dataset is exportable as **JSON**, **CSV**, and **Excel** out of the box. Use the *Open in Apify Console* link for an interactive table view with filtering and sort.

***

🛟 **Found a bug or want a field that's missing?** Open an issue from the actor page — we ship fixes fast.

# Actor input Schema

## `query` (type: `string`):

Free-text search term used to find companies or suppliers on ImportYeti (e.g. a brand name, importer name, product, or HS-code keyword). Leave empty if you supply Start URLs directly.

## `entityType` (type: `string`):

Restrict search results to either US importers (companies) or foreign manufacturers (suppliers). Leave empty to return both.

## `startUrls` (type: `array`):

Optional list of importyeti.com URLs to scrape directly. Accepts /company/<slug> or /supplier/<slug>. Use this for incremental runs over a known list of companies you already care about.

## `maxItems` (type: `integer`):

Hard cap on the number of company/supplier records pushed to the dataset. Set 0 for unlimited (the run will stop when search results run out).

## `since` (type: `string`):

Incremental mode. Provide an ISO-8601 timestamp (e.g. 2026-01-01T00:00:00Z) and the actor will skip detail fetches for companies whose most recent shipment is older than this date. Leave empty to scrape everything matching the query.

## `concurrency` (type: `integer`):

Maximum number of parallel HTTP requests to ImportYeti. Higher = faster but more likely to trigger rate limits. Defaults to 5; cap is 25.

## `fetchDetails` (type: `boolean`):

ON (default): enrich every result by fetching its detail page — HS codes, country-of-origin breakdown, monthly history, tags, phone, website. OFF: list-only mode — skip detail pages entirely to slash proxy traffic (~99% less). List-only records carry just the search-surface fields (name, address, country, total shipments, most recent shipment, trademarks) and are flagged `partial: true`. Applies to `query` searches.

## Actor input object example

```json
{
  "query": "apple",
  "entityType": "",
  "startUrls": [],
  "maxItems": 50,
  "concurrency": 5,
  "fetchDetails": true
}
```

# Actor output Schema

## `records` (type: `string`):

All importer & supplier records as a JSON array.

## `recordsCsv` (type: `string`):

Same data as a CSV spreadsheet — nested fields flattened.

## `recordsXlsx` (type: `string`):

Same data as an .xlsx workbook.

## `consoleView` (type: `string`):

Browse, filter, and re-export the dataset interactively in the Apify Console.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "query": "apple",
    "startUrls": [],
    "since": ""
};

// Run the Actor and wait for it to finish
const run = await client.actor("alwaysprimedev/importyeti-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "query": "apple",
    "startUrls": [],
    "since": "",
}

# Run the Actor and wait for it to finish
run = client.actor("alwaysprimedev/importyeti-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "query": "apple",
  "startUrls": [],
  "since": ""
}' |
apify call alwaysprimedev/importyeti-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,alwaysprimedev/importyeti-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/cnhtMhLU2tz17wCGa/builds/VUoDS9SnUCfPaRwbK/openapi.json
