Skip to content

Menu

Extract website data with Node.js

Use native Node.js fetch to extract quote records, validate JSON and handle cancellation, incomplete results and retries with a tested client.

By Scraipe editorial. Checked 2026-10-08.

Make a server-side extraction request

A Node.js ingestion job often needs records it can validate before handing them to another service. For a page of quotations, that means a list of quote and author pairs, the source URL, and enough request information to investigate failures. Scraipe performs the website extraction; your Node process calls the API and decides whether the returned records satisfy your application.

This guide builds one bounded request with Node's native fetch. It also covers a detail that matters in a worker or AI agent: cancelling the client does not prove that remote work stopped. The downloadable client combines a timeout with a caller's cancellation signal, rejects incomplete output and leaves retry decisions to your job runner.

For page text instead of named fields, see Choose Read or Extract.

Prerequisites and the runnable files

Use Node 22 or later. This example was tested with Node 22.23.2 on 4 October 2026. It uses ES modules and built-in APIs, so there is no npm dependency to install. The .mjs extension lets you run it without changing your project's package configuration.

Download the extraction client and its test file into the same directory. Start with the tests:

bash

node --test extract-website-data.test.mjs

The expected result is 16 passing tests. They start temporary HTTP servers on loopback, use invented records and close the servers afterward. No Scraipe API key, external website request or credit spend is involved. They check the API request, data validation, receipts, redirects, limits, cancellation and the actual command-line entry point.

For a real request, you need a working Scraipe API deployment and a server-side key with the scrape scope. Set SCRAIPE_API_URL to that deployment's HTTPS origin, without /v1 or an endpoint path. Set SCRAIPE_API_KEY through your runtime's secret configuration. Keep the key out of a browser bundle: this is server-side JavaScript even though the HTTP interface resembles browser fetch.

Request quote and author records

The example uses Quotes to Scrape, a public demonstration site. Its source URL is separate from your API origin. Start with this single page, then compare a few returned records with what the page actually displays.

Generate an operation ID once and keep it with the logical job:

bash

export REQUEST_ID="$(node --input-type=module -e 'import { randomUUID } from "node:crypto"; console.log(randomUUID())')"
node extract-website-data.mjs \
  'https://quotes.toscrape.com/' \
  "$REQUEST_ID"

On success, the command prints one JSON object. On failure, it writes a diagnostic to standard error and exits nonzero, without emitting partial result JSON. The tests exercise both paths. Reuse the saved ID only when retrying the same operation with the same request; generating another ID starts a separate logical request.

The client sends this body to POST /v1/scrape:

json

{
  "url": "https://quotes.toscrape.com/",
  "fields": ["quote", "author"],
  "maxTokens": 2000,
  "egress": "direct"
}

It sets the bearer authorization header, JSON content headers and Idempotency-Key. Change both the requested fields and the validator if you adapt the client to another record type. The API contract also supports schema input when your application needs a different extraction shape.

Use it inside an existing worker

The same file exports extractQuotes. After configuring the environment and REQUEST_ID above, save the following next to the client as run-extraction.mjs and run it with Node:

javascript

import { extractQuotes } from './extract-website-data.mjs';

const controller = new AbortController();
const result = await extractQuotes({
  apiBase: process.env.SCRAIPE_API_URL,
  apiKey: process.env.SCRAIPE_API_KEY,
  url: 'https://quotes.toscrape.com/',
  operationKey: process.env.REQUEST_ID,
  signal: controller.signal,
  timeoutMs: 60_000,
});
console.log(JSON.stringify(result, null, 2));

In a real worker, pass its cancellation signal or call controller.abort() when that job is cancelled. The client uses AbortSignal.any to combine caller cancellation with AbortSignal.timeout. The command-line version connects Ctrl+C to its controller.

The signal remains attached while the response body is read. This matters when headers arrive promptly but the body stalls: a test deliberately sends an unfinished JSON body and verifies that the timeout ends the client request. Another test cancels a pending body from the caller. An already-cancelled signal prevents dispatch entirely.

These are client-side bounds. An abort after dispatch can leave the server outcome unknown, and a busy JavaScript event loop can delay when cancellation is handled. Store job identity and inspect the request before deciding whether to retry; do not treat cancellation as proof that no work or charge occurred.

Validate before importing data

The fixture contains this deliberately invented record:

json

{
  "quote": "An invented fixture quotation.",
  "author": "Fixture author"
}

It is not a quotation attributed to a real person or evidence of a live extraction. The fixture tests demonstrate client behaviour against the development API contract, not production availability, accuracy on arbitrary websites or extraction speed.

The client accepts only a successful API envelope, truncated: false, and an array whose rows contain nonempty strings for both fields. A missing author, numeric value or partial response stops the result before import. An empty array remains empty; the client never fills it with guessed content.

The returned object includes rows, sourceUrl, receivedAt, operationKey, the server requestId, receipt from the API's meta, and any supplied provenance. Preserve these alongside your dataset. receivedAt is the client's acceptance time, not the page's publication date or proof that the origin was fetched afresh.

Shape checks do not establish historical authorship or factual truth. Provenance can help locate the page element that supplied a value; compare the extracted pair with the intended source. If an AI agent consumes the text, keep it as source data rather than allowing it to change the agent's instructions.

Bound response size and collection scope

The client reads the response stream with a default 1 MiB byte limit. It stops accepting data when that limit is exceeded. This bounds retained response chunks, not the entire process's memory use: decoding JSON, a single incoming chunk and parsed objects also use memory. Choose limits appropriate to your workload.

maxTokens is a separate API output limit. If it truncates the output, inspect the response cursor in the playground before implementing continuation. A cursor changes the request body and needs a separate operation key. It continues extraction output; it is not the website's next-page URL.

The sample does not pass follow, so it does not traverse the demonstration site's Next link. A complete response for one page is not proof of complete site coverage. Add explicit page, request and cost bounds before extending it into a collection job.

Handle HTTP failures and unknown outcomes

A fulfilled fetch promise does not by itself make a request successful. This client checks the HTTP status and the API's ok field, then validates the data. It reports safe reason/request-ID tokens and retry guidance when available, without dumping response bodies or credentials.

ConditionApplication decision
401Check the key and required scope.
402Read the account or network prerequisite; repeated requests do not resolve it.
409Distinguish an in-progress operation from an ID reused with a different body.
429Respect the returned retry delay and apply your bounded job retry policy.
Invalid JSON or interrupted transportPreserve the operation ID and investigate the unknown outcome.
Missing fields or truncationResolve the collection requirement before importing records.

There is no retry loop in the client. The current backend keeps caller-scoped successful responses in a bounded, in-memory idempotency store with a ten-minute expiry. Restarts or capacity eviction can remove entries earlier. Repeating an identical request while its stored success is available can replay that success without another charge; failures are not stored. This is not a durable exactly-once job guarantee.

The HTTP call uses redirect: 'error'. It therefore refuses an API redirect instead of sending the bearer credential toward a new destination. Configure the trusted final API origin directly. Remote API origins require HTTPS; loopback HTTP is permitted for the fixture tests.

Inspect the cost before scaling

egress: direct opts out of residential fallback, so some sites can remain inaccessible. It does not make new extractor compilation free. The API identifies compiled-fresh, healed and llm as billable tiers; inspect receipt.credits.compile, receipt.credits.egress and receipt.credits.total instead of inferring cost from output length or HTTP status alone.

Use the JSON tool to inspect the workflow, then make one permitted request with your own API key. Validate the source records, receipt and failure behaviour before adding concurrency or automatic retries to your Node.js service.

Sources and verification

Try the next step

Try the playground

Continue building