Skip to content

Menu

Choose Read or Extract for a web-data task

Choose the right operation from the data your application needs, then validate the response before using it.

By Scraipe editorial. Checked 2026-10-02.

Start with the data your application needs

A website can supply two different kinds of useful input. A research assistant may need the words in an article, with its source URL attached. A catalog job may need a list of records with a title, price and availability field. These are different requirements even when the input URL is identical.

Use Read when the downstream task needs page content. Use Extract, exposed by the scrape endpoint, when the task needs named fields or records. Neither choice proves that the website is accessible or that the result is complete: inspect the response and validate it before using it.

Choose an operation

Downstream requirementStarting operationWhat to validate
Read an article for a research summaryReadReturned format, missing sections and source attribution
Store product titles and pricesExtractRequired fields, types, currency and record count
Build a knowledge-base ingestion pipelineRead, then a bounded crawl if neededCoverage, duplicates, freshness and source URLs
Compare changes to structured product recordsExtract, then a watch workflowInitial values and meaningful changes

A workflow can need both representations. For example, an agent may read a product description for context while storing a structured price separately. Keep those outputs distinct instead of asking later steps to guess which part of a free-form answer is the price.

Try a deterministic selection example

The following helper demonstrates the decision using explicit requirements. It does not fetch a website or claim that a live extraction succeeded.

javascript

function chooseOperation({ needsFields, needsWholeText }) {
  if (needsFields && needsWholeText) return ['read', 'scrape'];
  if (needsFields) return ['scrape'];
  if (needsWholeText) return ['read'];
  throw new Error('Define the downstream data requirement first');
}

Download the complete example with four assertions. Save it locally and run node choose-operation.mjs using Node 22 or later. The example makes no network calls and spends no credits. It checks field-only, text-only, combined and unspecified requirements.

Map the choice to the API

The API reference defines POST /v1/read for page content and POST /v1/scrape for structured extraction. Supply a supported public URL and the required authentication. For structured extraction, specify the fields or supported schema your consumer expects. Do not put an API key in frontend code or a public example.

Read can return Markdown, text or HTML depending on the request and supported behaviour. Markdown conversion may fall back to plain text, so inspect the returned format rather than assuming the requested format was always produced. A token or character limit can also mean that you have only part of a long document.

Extraction uses the page structure to locate values. Reusing an extractor can avoid another model call, but that does not mean every network request is free. Consult the current pricing and request receipt for compilation, repair and network charges. A selector finding a real element also does not prove it selected the correct business field.

Handle an incomplete result

Before passing output to another system, check that the operation succeeded and that its output satisfies your own requirements. For a record, validate required fields and retain the original currency. For an article, check that the expected section is present and attach the source URL and retrieval time.

If access is refused or the page requires unsupported interaction, surface that limitation. Do not turn an empty response into an invented value. For a rate limit, respect the retry guidance instead of repeatedly submitting the same request.

Continue with a small real workload

Try the JSON demo or the Markdown reader, then move to your own permitted URL in the playground. Compare the response with the source page before increasing the number of pages. This guide verifies the local decision example and documents the API contract; it does not present a live-service benchmark.

Sources and verification

Try the next step

Try the playground