# Legislation.gov.uk Scraper (`crawlerbros/legislation-gov-uk-scraper`) Actor

Scrape UK statute law from legislation.gov.uk. Fetch specific Acts and Statutory Instruments, browse by type and year, get the newest legislation, or search by title. Returns clean metadata plus official XML/PDF links.

- **URL**: https://apify.com/crawlerbros/legislation-gov-uk-scraper.md
- **Developed by:** [Crawler Bros](https://apify.com/crawlerbros) (community)
- **Categories:** Automation, Developer tools, Integrations
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Legislation.gov.uk Scraper

Scrape **UK statute law** straight from the official [legislation.gov.uk](https://www.legislation.gov.uk) service — Acts of Parliament, Statutory Instruments, devolved legislation for Scotland, Wales and Northern Ireland, and more. Fetch a specific piece of legislation, browse everything of a given type and year, pull the newest published items, or search by title. HTTP-only, no login, no proxy.

### What this actor does

- **Four modes:** `browseByTypeYear`, `byLegislation`, `newest`, `search`
- **31 legislation types** — from UK Public General Acts (`ukpga`) and Statutory Instruments (`uksi`) to Acts of the Scottish Parliament (`asp`), Senedd Cymru (`asc`), the Northern Ireland Assembly (`nia`), and retained EU legislation (`eur`, `eudn`, `eudr`)
- **Clean metadata** parsed from the official Atom feeds and item XML
- **Official links** to each item's web page, XML, PDF and table of contents
- **Empty fields are omitted** — every field in a record is populated

### Modes

| Mode | What it does | Needs |
|---|---|---|
| `browseByTypeYear` | List all legislation of a type (optionally one year) | `legislationType` (+ optional `year`) |
| `byLegislation` | Fetch specific item(s) with full metadata | `legislationType`, `year`, `number` (or `numbers`) |
| `newest` | Most recently published items of a type | `legislationType` |
| `search` | Title search within a type (or across all types) | `titleQuery` (+ optional `legislationType`) |

### Output fields

Feed-based modes (`browseByTypeYear`, `newest`, `search`) return a lightweight record per item; `byLegislation` returns the full item metadata. Fields that aren't available for a given item are omitted.

- `legislationType` — type code (e.g. `ukpga`)
- `typeTitle` — human-readable type (e.g. `UK Public General Act`)
- `year`, `number` — the item's year and number
- `title` — the legislation title
- `summary` — long title / description (feed modes)
- `documentMainType` — canonical classification (e.g. `UnitedKingdomPublicGeneralAct`)
- `documentCategory` — `primary` or `secondary` (`byLegislation`)
- `documentStatus` — `revised`, `final`, `enacted` (`byLegislation`)
- `extent` — territorial extent, e.g. `E+W+S+N.I.` (`byLegislation`)
- `numberOfProvisions` — provision count (`byLegislation`)
- `statistics` — paragraph/image counts: `{ totalParagraphs, bodyParagraphs, scheduleParagraphs, attachmentParagraphs, totalImages }` (`byLegislation`)
- `enactmentDate`, `creationDate`, `madeDate` — key dates
- `validDate`, `modifiedDate`, `restrictStartDate` — revision dates (`byLegislation`)
- `alternativeNumber` — e.g. Welsh series number (`byLegislation`)
- `isbn` — ISBN of the printed publication
- `publisher` — publishing authority (`byLegislation`)
- `schemaVersion` — legislation XML schema version (`byLegislation`)
- `published`, `updated` — feed timestamps (feed modes)
- `url` — official web page for the item
- `idUri` — canonical identifier URI
- `xmlUrl`, `contentsUrl` — official XML and table-of-contents links
- `pdfUrl` — official print PDF, English edition when bilingual (`byLegislation`)
- `pdfSizeBytes` — size of the print PDF in bytes (`byLegislation`)
- `aknUrl` — Akoma Ntoso XML (`byLegislation`)
- `versionedXmlUrl` — point-in-time XML (feed modes)
- `recordType: "legislation"`, `sourceUrl`, `scrapedAt`

### Input

| Field | Type | Default | Description |
|---|---|---|---|
| `mode` | string | `browseByTypeYear` | `browseByTypeYear` / `byLegislation` / `newest` / `search` |
| `legislationType` | string | `ukpga` | One of 31 type codes |
| `year` | int | `2023` | Calendar year (required for `byLegislation`) |
| `number` | int | – | Item number (`byLegislation`) |
| `numbers` | array | – | Multiple item numbers (`byLegislation`) |
| `titleQuery` | string | `data protection` | Title search text (`search`) |
| `maxItems` | int | `50` | Hard cap (1–2000) |

#### Example: browse Acts of the Scottish Parliament for 2023

```json
{ "mode": "browseByTypeYear", "legislationType": "asp", "year": 2023, "maxItems": 50 }
```

#### Example: fetch the Finance Act 2023

```json
{ "mode": "byLegislation", "legislationType": "ukpga", "year": 2023, "number": 1 }
```

#### Example: newest UK Statutory Instruments

```json
{ "mode": "newest", "legislationType": "uksi", "maxItems": 100 }
```

#### Example: title search for "data protection"

```json
{ "mode": "search", "legislationType": "ukpga", "titleQuery": "data protection" }
```

### Use cases

- **Legal research** — build a corpus of Acts and Statutory Instruments by type and year
- **Compliance monitoring** — track the newest legislation affecting your sector
- **Legal tech** — feed clean statute metadata and official XML/PDF links into your product
- **Govtech & civic data** — mirror UK legislation for open-data projects
- **Academic research** — bulk-export legislative metadata for analysis

### FAQ

**What is legislation.gov.uk?**  The official home of UK legislation, managed by The National Archives. It publishes most UK legislation and is free to use and re-use under the Open Government Licence.

**What do the type codes mean?**  Each jurisdiction and instrument has a short code — `ukpga` (UK Public General Act), `uksi` (UK Statutory Instrument), `asp` (Act of the Scottish Parliament), `asc` (Act of Senedd Cymru), `nia` (Act of the Northern Ireland Assembly), and 26 more, all selectable in the input.

**How do I find an item's number?**  Legislation is identified by type + year + number, e.g. the Finance Act 2023 is `ukpga/2023/1`. Use `browseByTypeYear` or `search` to discover numbers, then `byLegislation` for full detail.

**What does `extent` mean?**  The territorial extent — where the legislation applies. `E+W+S+N.I.` means England, Wales, Scotland and Northern Ireland; `S` means Scotland only.

**What's the difference between `revised` and `enacted`?**  `documentStatus` reflects whether you're seeing the up-to-date revised text or the text as originally enacted/made.

**Is the data official?**  Yes — every record links back to the official legislation.gov.uk page, XML, PDF and table of contents.

**How fresh is the data?**  legislation.gov.uk publishes new legislation continuously; the `newest` mode surfaces the latest items and each item carries `published`/`updated` timestamps.

# Actor input Schema

## `mode` (type: `string`):

What to fetch.

## `legislationType` (type: `string`):

The legislation.gov.uk document-type code.

## `year` (type: `integer`):

Calendar year of the legislation (required for byLegislation; optional filter for browseByTypeYear).

## `number` (type: `integer`):

The item number within the type + year (e.g. `1` for the first Act of that year).

## `numbers` (type: `array`):

Multiple item numbers to fetch for the given type + year.

## `titleQuery` (type: `string`):

Free-text title search (e.g. `data protection`).

## `maxItems` (type: `integer`):

Hard cap on emitted records.

## Actor input object example

```json
{
  "mode": "browseByTypeYear",
  "legislationType": "ukpga",
  "year": 2023,
  "numbers": [],
  "titleQuery": "data protection",
  "maxItems": 20
}
```

# Actor output Schema

## `legislation` (type: `string`):

Dataset containing all scraped legislation.gov.uk records.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "mode": "browseByTypeYear",
    "legislationType": "ukpga",
    "year": 2023,
    "numbers": [],
    "titleQuery": "data protection",
    "maxItems": 20
};

// Run the Actor and wait for it to finish
const run = await client.actor("crawlerbros/legislation-gov-uk-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "mode": "browseByTypeYear",
    "legislationType": "ukpga",
    "year": 2023,
    "numbers": [],
    "titleQuery": "data protection",
    "maxItems": 20,
}

# Run the Actor and wait for it to finish
run = client.actor("crawlerbros/legislation-gov-uk-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "mode": "browseByTypeYear",
  "legislationType": "ukpga",
  "year": 2023,
  "numbers": [],
  "titleQuery": "data protection",
  "maxItems": 20
}' |
apify call crawlerbros/legislation-gov-uk-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,crawlerbros/legislation-gov-uk-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/GG0oKoi4Uk7XVgzC3/builds/WeAlgtwHklgseaCQL/openapi.json
