Choose Read or Extract for a web-data task
Choose the right operation from the data your application needs, then validate the response before using it.
By Scraipe editorial. Checked 2026-10-02.
Start with the data your application needs
A website can supply two different kinds of useful input. A research assistant may need the words in an article, with its source URL attached. A catalog job may need a list of records with a title, price and availability field. These are different requirements even when the input URL is identical.
Use Read when the downstream task needs page content. Use Extract, exposed by the scrape endpoint, when the task needs named fields or records. Neither choice proves that the website is accessible or that the result is complete: inspect the response and validate it before using it.
Choose an operation
| Downstream requirement | Starting operation | What to validate |
|---|---|---|
| Read an article for a research summary | Read | Returned format, missing sections and source attribution |
| Store product titles and prices | Extract | Required fields, types, currency and record count |
| Build a knowledge-base ingestion pipeline | Read, then a bounded crawl if needed | Coverage, duplicates, freshness and source URLs |
| Compare changes to structured product records | Extract, then a watch workflow | Initial values and meaningful changes |
A workflow can need both representations. For example, an agent may read a product description for context while storing a structured price separately. Keep those outputs distinct instead of asking later steps to guess which part of a free-form answer is the price.
Try a deterministic selection example
The following helper demonstrates the decision using explicit requirements. It does not fetch a website or claim that a live extraction succeeded.
javascript
function chooseOperation({ needsFields, needsWholeText }) {
if (needsFields && needsWholeText) return ['read', 'scrape'];
if (needsFields) return ['scrape'];
if (needsWholeText) return ['read'];
throw new Error('Define the downstream data requirement first');
}Download the complete example with four assertions. Save it locally and run node choose-operation.mjs using Node 22 or later. The example makes no network calls and spends no credits. It checks field-only, text-only, combined and unspecified requirements.
Map the choice to the API
The API reference defines POST /v1/read for page content and POST /v1/scrape for structured extraction. Supply a supported public URL and the required authentication. For structured extraction, specify the fields or supported schema your consumer expects. Do not put an API key in frontend code or a public example.
Read can return Markdown, text or HTML depending on the request and supported behaviour. Markdown conversion may fall back to plain text, so inspect the returned format rather than assuming the requested format was always produced. A token or character limit can also mean that you have only part of a long document.
Extraction uses the page structure to locate values. Reusing an extractor can avoid another model call, but that does not mean every network request is free. Consult the current pricing and request receipt for compilation, repair and network charges. A selector finding a real element also does not prove it selected the correct business field.
Handle an incomplete result
Before passing output to another system, check that the operation succeeded and that its output satisfies your own requirements. For a record, validate required fields and retain the original currency. For an article, check that the expected section is present and attach the source URL and retrieval time.
If access is refused or the page requires unsupported interaction, surface that limitation. Do not turn an empty response into an invented value. For a rate limit, respect the retry guidance instead of repeatedly submitting the same request.
Continue with a small real workload
Try the JSON demo or the Markdown reader, then move to your own permitted URL in the playground. Compare the response with the source page before increasing the number of pages. This guide verifies the local decision example and documents the API contract; it does not present a live-service benchmark.