# AI Image Captioner (`seemuapps/image-captioner`) Actor

Generate accurate text descriptions for any image using AI — bulk caption product photos, screenshots, or any image URL for SEO, accessibility, and content tagging.

- **URL**: https://apify.com/seemuapps/image-captioner.md
- **Developed by:** [Andrew](https://apify.com/seemuapps) (community)
- **Categories:** E-commerce, SEO tools, Social media
- **Stats:** 15 total users, 5 monthly users, 97.8% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $12.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## AI Image Captioner

Generate accurate, detailed text descriptions for any image using AI - bulk caption product photos, screenshots, or any image URL for SEO alt text, accessibility compliance, and content tagging.

### What you get

- Natural-language captions generated by Molmo 2, trained on 712,000+ human-described images
- Three detail levels: brief one-liner, balanced description, or full detailed paragraph
- Optional focus directive to target specific aspects (text, background, faces, objects, etc.)
- One output record per image with the caption and source URL
- Supports bulk processing - pass up to 50 image URLs and get all captions in a single run
- Export to JSON or CSV directly from the Apify console

### Use cases

- **E-commerce SEO** - generate alt text for thousands of product images automatically
- **Accessibility compliance** - add descriptive alt text to images on websites and apps
- **Content moderation** - understand what's in user-uploaded images before publishing
- **Dataset labeling** - annotate image datasets for machine learning pipelines
- **Digital asset management** - auto-tag and describe photos in large media libraries
- **Social media monitoring** - caption scraped images to make them searchable by content

### Examples

| Image | Detail Level | Caption |
|-------|-------------|---------|
| ![Nike sneaker on red background](https://i.imgur.com/JnmcDHP.jpeg) | **High** | A vibrant red Nike sneaker takes center stage in this striking advertisement, set against a bold red background that creates a visually cohesive and eye-catching composition. The shoe is positioned at an angle, giving the impression of motion and energy. The sneaker features a white Nike swoosh, darker red laces, and "Nike Free" branding on the white sole. The lighting is bright and even, highlighting the shoe's textures and details. |
| ![Nike sneaker on red background](https://i.imgur.com/JnmcDHP.jpeg) | **Medium** | A vibrant red Nike sneaker is displayed against a matching red background. The shoe features a white Nike swoosh and "Nike Free" branding on the sole. The laces are a darker shade of red, complementing the overall design. |
| ![Nike sneaker on red background](https://i.imgur.com/JnmcDHP.jpeg) | **Low** | A red Nike sneaker with white accents. |
| ![Food flatlay with three dishes](https://i.imgur.com/GdMlOMj.jpeg) | **High** | A top-down view of a rustic wooden table with three round bowls arranged in a triangular formation. The central bowl features slices of medium-rare steak garnished with fresh green leaves and a red chili pepper. The left bowl holds crispy fried fish topped with a creamy sauce and herbs. The right bowl contains a meat dish garnished with thinly sliced red onions and nuts. Scattered around the bowls are whole chili peppers, cashews, and a small bowl of brown dipping sauce. |
| ![Food flatlay with three dishes](https://i.imgur.com/GdMlOMj.jpeg) | **Medium** | Three bowls of food are arranged on a gray wooden table, creating a rustic dining scene. The central bowl contains sliced steak, while the left bowl holds fried fish topped with sauce and herbs. The right bowl features a meat dish garnished with onions and nuts. |
| ![Food flatlay with three dishes](https://i.imgur.com/GdMlOMj.jpeg) | **Low** | Three bowls of food on a wooden table with garnishes. |

### How to use

1. Paste one or more image URLs into the **Images** field (or upload files directly)
2. Choose a **Detail Level** - High gives the most descriptive output (recommended for SEO and accessibility)
3. Optionally add a **Focus** hint to direct the model's attention (e.g. "describe only the text visible")
4. Click **Run** - captions appear in the **Dataset** tab when complete
5. Export results as JSON or CSV, or connect to downstream actors via the Apify API

### Output format

Each dataset record:

```json
{
  "inputImageUrl": "https://example.com/product.jpg",
  "caption": "A white ceramic coffee mug sitting on a wooden table next to an open laptop. The mug has a minimalist logo on the front and steam rising from the top, suggesting the coffee is hot.",
  "detailLevel": "high",
  "status": "success",
  "error": null
}
```

### Input options

| Field | Type | Description |
|-------|------|-------------|
| Images | URL list | One or more `http/https` image URLs or base64 data URIs |
| Upload Images | File upload | Upload images directly from your computer |
| Detail Level | Select | `Low` (one-liner), `Medium` (balanced), `High` (detailed paragraph) - default: High |
| Focus | Text | Optional directive to focus the caption on a specific aspect of the image |

### Limits

- Maximum 50 images per run
- Each image must be a publicly accessible URL or a base64 data URI
- Processing time is typically 5-15 seconds per image

### Related AI image actors

Part of a complete AI image toolkit - explore the rest of the suite:

- [AI Image Background Remover](https://apify.com/seemuapps/image-background-remover) - Remove backgrounds to clean transparent PNGs
- [AI Image Upscaler](https://apify.com/seemuapps/image-upscaler) - Batch-upscale images to 4K or 8K
- [AI Image Watermark Remover](https://apify.com/seemuapps/image-watermark-remover) - Remove text and logo watermarks from images
- [Image OCR Scraper](https://apify.com/seemuapps/image-ocr-scraper) - Extract text from images in 109 languages
- [Photo Location Finder](https://apify.com/seemuapps/image-to-location) - Find where a photo was taken - no EXIF needed

# Actor input Schema

## `images` (type: `array`):

Images to caption. Each entry must be a publicly accessible http/https URL or a base64 data URI (data:image/jpeg;base64,...).

## `imageFiles` (type: `array`):

Upload image files directly. Combined with any URLs provided above.

## `detailLevel` (type: `string`):

How much detail to include in the caption. High produces the most descriptive output.

## `focus` (type: `string`):

Optionally direct the model's attention to a specific aspect of the image - e.g. 'describe only the text visible' or 'focus on the background environment'.

## Actor input object example

```json
{
  "images": [
    "https://i.imgur.com/TSem3Jf.jpeg",
    "https://i.imgur.com/or3U2Xx.jpeg"
  ],
  "detailLevel": "high"
}
```

# Actor output Schema

## `results` (type: `string`):

One record per image: inputImageUrl, caption, detailLevel, status (success|error), error.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "images": [
        "https://i.imgur.com/TSem3Jf.jpeg",
        "https://i.imgur.com/or3U2Xx.jpeg"
    ],
    "focus": ""
};

// Run the Actor and wait for it to finish
const run = await client.actor("seemuapps/image-captioner").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "images": [
        "https://i.imgur.com/TSem3Jf.jpeg",
        "https://i.imgur.com/or3U2Xx.jpeg",
    ],
    "focus": "",
}

# Run the Actor and wait for it to finish
run = client.actor("seemuapps/image-captioner").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "images": [
    "https://i.imgur.com/TSem3Jf.jpeg",
    "https://i.imgur.com/or3U2Xx.jpeg"
  ],
  "focus": ""
}' |
apify call seemuapps/image-captioner --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,seemuapps/image-captioner"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/cvFmIR8jDbhnJOvKz/builds/CFIyH92qGB6Jhbycn/openapi.json
