Extract website data with Python
Extract product titles and prices with a Python client, validate JSON, retain currency and handle rate limits, truncation and retries.
By Scraipe editorial. Checked 2026-10-03.
Extract records you can validate
To extract website data with Python, send a permitted page URL and the fields you need to Scraipe's POST /v1/scrape endpoint, then validate the returned records before saving them. This guide requests product titles and prices using Python's standard library. It preserves the source URL, currency text and request receipt so you can inspect what happened later.
If your application needs an article's text rather than product records, start with Choose Read or Extract. Here the reader task is narrower: turn one product-list page into a checked JSON result. The example deliberately stops when output is incomplete and sends no automatic retries.
Before you start
You need Python 3.10 or later, a working Scraipe API deployment and an API key with the scrape scope. The downloadable example and its fixtures were tested with Python 3.14.6 on 3 October 2026. Keep your key in your server environment or secret manager, outside source control and browser code.
Use a public page you are permitted to process. The command below uses Books to Scrape, a demonstration catalogue. It requests only that page: it does not discover every product in the catalogue or follow next-page links. Check the source page and your intended use before expanding to other sites.
Download the Python client and its local HTTP tests into the same folder. No pip packages are required.
Run the example without an API key first
bash
python3 test-extract-website-data.pyThis starts a temporary server on your own computer, sends fixture requests to it and shuts it down afterward. It does not contact Scraipe or Books to Scrape and spends no credits. The expected result is Ran 11 tests followed by OK.
The tests cover the outgoing headers and request body, retained currency text, empty results, invalid records, response truncation, a rate limit, a failure envelope, non-JSON gateway responses, redirects, an insecure remote API origin, and a simulated network failure. They also check that manually repeating an operation preserves its key and body. That last check proves client behaviour; it does not simulate the backend's deduplication guarantee.
Make one authenticated extraction
Configure SCRAIPE_API_URL with the HTTPS origin of your Scraipe API deployment, without /v1 or a trailing endpoint path. The client appends /v1/scrape. Configure SCRAIPE_API_KEY securely in the process environment. An API URL is separate from the website you want to extract.
Create a new operation ID for this logical request, retain it with your job record, then run:
bash
export REQUEST_ID="$(python3 -c 'import uuid; print(uuid.uuid4())')"
python3 extract-website-data.py \
'https://books.toscrape.com/' \
--request-id "$REQUEST_ID"The program prints JSON to standard output on success and returns a nonzero exit code on failure. Inspect a successful response before piping it into your storage system. If you retry this exact operation, reuse the saved ID; do not rerun the ID-generation command first. Generate a new ID when you intentionally request a fresh operation.
The request body is:
json
{
"url": "https://books.toscrape.com/",
"fields": ["title", "price"],
"maxTokens": 2000,
"egress": "direct"
}It also sends Authorization: Bearer …, JSON content headers and Idempotency-Key. The API reference documents field hints and the alternative JSON Schema input if your application needs another contract. This client is a small product-list example, not a general SDK: changing the requested fields also requires changing its validation rules.
Understand the output
A successful result contains sourceUrl, a UTC retrievedAt timestamp, the operation key, the server's requestId, rows and the API meta object under receipt. The timestamp records when this client accepted the response; it is not the website's publication time or proof of a fresh origin fetch.
The fixture's row looks like this:
json
{
"title": "Fixture book",
"price": "£12.50"
}This is invented test input, not an observed book listing. Real titles, prices, record counts and receipts depend on the source and request. The tests establish local client behaviour against the documented API contract; they are not a live-service extraction result, availability check or speed benchmark.
The validator requires ok: true, truncated: false, a list of records, and nonempty string values for every title and price. It preserves the pound sign rather than coercing a price to a floating-point number with an ambiguous currency. If you need arithmetic, perform a separate, explicitly tested currency and decimal conversion.
An empty list stays empty. A missing field stops the whole result instead of silently dropping a row or inventing a value. Passing these checks proves the data has the required shape; compare a sample with the website to confirm that the title and price refer to the intended product.
Treat incomplete output separately from website pagination
maxTokens bounds the response, not the number of products on a website or the total cost of an extraction. If the API reports truncated: true, this client rejects the partial dataset. Inspect the response and cursor in the playground before adding a continuation loop to your application.
A response cursor continues extraction output from the requested page. Following a website's next-page links is a separate operation, documented as follow in the API reference. Neither a complete response nor a nonempty list proves that you collected the entire catalogue.
When extending this example, bound the number of continuation requests and fetched pages, and track repeated cursors. A cursor changes the request body, so give that continuation its own operation key. Keep each page's source and receipt rather than merging records without their origin.
Handle errors without creating a retry loop
The program reports a safe error reason, request ID and retry delay when available. It does not dump the response body or the API key into an error message. Common decisions are:
| Response | Next step |
|---|---|
401 | Check the key and its permissions before another request. |
402 | Inspect the reason and account/network requirement; repeating the request does not fix a missing prerequisite. |
409 | Distinguish an operation still running from a key reused with a different body. |
429 | Respect retryAfter or Retry-After, then apply a bounded retry policy. |
| HTML, invalid JSON or a network timeout | Treat the outcome as unknown; preserve the operation key and inspect the request before retrying. |
| Truncated or invalid records | Fix the collection or validation requirement before importing data. |
The current backend scopes idempotency keys to the caller and retains them for ten minutes in process memory. Within that window, the same key and body can replay a stored success without another charge. Failed requests are not stored. Restarts can discard the window, so idempotency is not a durable job ledger or an unlimited guarantee against repeated work.
The client sets a 60-second timeout for blocking network operations. Python's timeout is not an end-to-end job deadline, and an upstream host can stop waiting earlier. It also limits its local response read to 1 MiB. A timeout or size rejection does not establish that the server cancelled its work.
Redirects from the API origin are rejected so the client cannot forward the bearer key to another host. Use the final trusted API origin directly. Remote API connections require HTTPS; HTTP is accepted only on loopback for local tests.
Check the receipt before increasing volume
The example selects egress: direct, so it does not opt into a paid residential fallback. A blocked site can therefore fail. Direct fetching does not make extractor compilation free: the API describes compiled-fresh, healed and llm as billable tiers. Inspect receipt.credits.compile, receipt.credits.egress and receipt.credits.total for the successful request rather than inferring price from the number of output rows.
Start with one page in the playground, verify a few records against the source and read the receipt. Then decide whether you need pagination, another schema or the Markdown reader. Keeping this first extraction small makes failures and unexpected data easier to diagnose.