Skip to content

Menu

Docs API reference

API reference

The endpoints the v1 server serves, as built on 2026-09-24. The examples are real responses from a running server, trimmed with … where they were long.

  • Base URL: http://localhost:5001, or your deployment. /api/v1/* is an alias of /v1/*.
  • The contract is backend/openapi.json (OpenAPI 3.1). Operations it describes that v1 does not serve are marked "x-status": "not-in-v1", and are listed at the end of this page. api-v2-changes.md is the changelog for the full contract.

Contents: Conventions · Networks · Streaming · Health, pricing, whoami · Scrape · Quote · Batch · Compile · Read · Map · Crawl and jobs · Watches · Registry · Network page · Demo · Opt-out · MCP · Account · Not in v1


Conventions

Envelope

Every JSON response has ok and a requestId, which also comes back in the X-Scraipe-Request-Id header. Success puts its data at the top level. Failure puts it in error:

json

{ "ok": false,
  "error": { "reason": "llm_unconfigured",
             "message": "No LLM is configured on this server (OPENAI_API_KEY is not set), so new extractors cannot be compiled.",
             "retryable": false,
             "hint": "Cached and compiled requests are unaffected. The operator needs to set OPENAI_API_KEY.",
             "charged": false },
  "meta": { "tier": "error", "credits": { "compile": 0, "egress": 0, "total": 0 }, "egress": { … } },
  "requestId": "req_muf0svwrea9cfe201edc8586" }
  • error.reason is stable and machine-readable. Branch on it, never on message.
  • error.charged is always false: a failed request is never charged.
  • A failure may also carry retryable, retryAfter (seconds), hint and subtype. When a network was involved it can carry class, vendor, code, status, attempts, rungsTried, estimatedCredits and limitBytes.
  • Metered failures (scrape, batch items, read, compile) also carry meta, with zero credits and the network attempts made.
  • /api/* responses still carry the old success field, and on failure a top-level reason and message, next to ok and error. Read ok and error.

Authentication and scopes

credentialheaderwhere
API keyAuthorization: Bearer sk_scr_… or X-API-Key: sk_scr_…/v1, and /mcp (Bearer only)
session token (7 days)Authorization: Bearer <jwt>/api/*, and every /v1 route (logged as source: "playground")
nonehealth, pricing, the demo, opt-out, register, login, password reset, get-prices

SCRAIPE_OPEN_MODE=true (local only) needs no credential at all.

scopegrants
scrapePOST /v1/scrape, /v1/batch, /v1/crawl; DELETE /v1/jobs/:id
readPOST /v1/read, /v1/map, /v1/quote; every authenticated GET under /v1 except whoami
compilePOST /v1/compile
watchescreate, update, delete and run watches
admin/v1/admin/*; operator file keys only, never account keys

New keys get all four scopes unless you ask for fewer. Keys minted before read and watches existed still hold them if they hold scrape.

Statuses

statusmeaningwhat to do
400the request is malformed: bad_request, invalid_url, url_not_allowed, bad_egress, bad_cursor, too_many_samples, too_many_urls, malformed_jsonfix the request
401credential missing or refused; error.reason says whichfix the key
402money: out_of_credits (a compile or a residential page), needs_residential (the page needs the paid network and this request cannot use it)add credits, or allow residential. Cached and compiled requests keep working
403the key lacks the scope (insufficient_scope), or demo_not_alloweduse a key with the scope
404no such resource, or it belongs to another account (never a 403, which would confirm it exists)check the id
409idempotency_key_reuse, idempotency_in_progress, already_running, email_taken, key_limit_reached, no_billing_accountsee the reason
413body_too_large (over 2 MB)send less
422understood, but the site refused, or the page does not fitread error.reason; do not retry blindly
429a temporary budget: heal_limit, layout_limit, concurrency_limit, demo_limit_reached, auth rate_limitedretry after Retry-After
503ours, not yours: starting, auth_unavailable, accounts_unavailable, llm_unavailable, llm_unconfigured, spend_cap_reached, egress_unavailable, billing_unavailable, reader_busy, payments_unavailableretry after Retry-After when it is set; your key is fine
  • 401 reasons: token_missing, key_invalid, key_revoked, token_invalid, token_expired, password_changed (a session issued before the password changed), account_deleted.
  • The site's refusals (422):

    • bot_protection: a challenge or a block page; vendor is named;
    • access_denied: a firewall or a ban, on every network tried;
    • geo_blocked (countryTried);
    • rate_limited (subtype: origin, retryAfter) and rate_limited_backoff;
    • robots_disallowed, site_opted_out, auth_required, unavailable_for_legal_reasons;
    • http_error (subtype: not_found for 404/410), origin_error, timeout, network_error.
  • The page (422):

    • egress_requires_https, too_heavy (limitBytes), too_large, too_complex;
    • needs_browser, render_failed, unsupported_content_type;
    • extraction_failed, compile_failed.

On Heroku, a response that has not started within 30 seconds is replaced by Heroku's own HTML error page. A first-time compile of a large page can hit it. The work still finishes and the extractor is stored, so a retry with the same Idempotency-Key is served from the registry. Stream the request, or pre-warm with POST /v1/compile.

Response headers

headerexamplemeaning
X-Scraipe-Request-Idreq_muf0…every response
X-Scraipe-Tiercompiledwhich tier answered
X-Scraipe-Latency-Ms214server time
X-Scraipe-Tokens0LLM tokens spent
X-Scraipe-Origin-Requests1times the site was contacted, every attempt included
X-Scraipe-Cost-Micros0our LLM cost, when SCRAIPE_PRICE_PER_MTOK is set
X-Scraipe-Credits-Charged0legacy: whole compile credits only
X-Scraipe-Credits-Total0.050everything charged: compile plus network, 3 decimals
X-Scraipe-Egressresidential/http<network>/<mode> that got the page, or none (cache)
X-Scraipe-Egress-Bytes48211wire bytes over all attempts
X-Scraipe-Egress-Attempts2network attempts made
X-Scraipe-Egress-Credits0.050network credits charged
X-Scraipe-Idempotent-Replaytruea stored replay of an earlier request

Scrape, read, map and compile send the metered set. Batch sends X-Scraipe-Credits-Total only. CORS exposes all of them, plus Retry-After, Content-Disposition and Mcp-Session-Id, to the console's origins.

Money in meta

json

"meta": {
  "tier": "compiled", "latencyMs": 214, "tokens": 0, "originRequests": 1,
  "credits": { "compile": 0, "egress": 0.05, "total": 0.05 },
  "creditsCharged": 0.05,
  "egress": { "rung": "residential", "mode": "http", "provider": "decodo", "country": "us",
              "wireBytes": 48211, "costMicros": 180, "credits": 0.05, "priceClass": "standard",
              "private": false,
              "attempts": [
                { "rung": "direct", "mode": "http", "class": "waf_block", "vendor": "cloudflare",
                  "code": "1020", "status": 403, "bytes": 3120, "ms": 180, "probe": false, … },
                { "rung": "residential", "mode": "http", "provider": "decodo", "country": "us",
                  "class": "ok", "status": 200, "bytes": 45091, "ms": 903, "probe": false, … } ] },
  "requestId": "req_…" }
  • credits.compile is 0 or 1 per page.
  • credits.egress is the residential price: 0.05, 0.15 or 0.20.
  • priceClass is standard or heavy, and null when nothing was charged.
  • private: true means the request carried your own credentials.

Idempotency

Send Idempotency-Key: <unique per operation> on POST /v1/scrape.

  • A retry with the same key and body replays the stored success (X-Scraipe-Idempotent-Replay: true) and is not re-billed.
  • A duplicate that arrives while the first is still running waits for it, up to 60 s, then gets 409 idempotency_in_progress with Retry-After: 5.
  • The same key with a different body (a different URL, fields, network or credentials) is 409 idempotency_key_reuse.
  • Scope and lifetime. Keys are scoped to the caller and live 10 minutes. Failures are not stored, so a retry after a failure is a genuine second attempt.

Lists

Lists that paginate take limit (default 50, max 200) and an opaque cursor, and return nextCursor (null at the end). These are GET /v1/watches, /v1/registry, /v1/egress/domains, /api/account/requests and /api/account/ledger. A cursor the endpoint did not issue is 400 bad_cursor. Timestamps are ISO-8601 UTC.


Networks: the egress option

Pages are fetched from our servers for free. When a site blocks them, a request may use the residential network: it is paid, and charged only on success. anti-blocking.md has the whole design.

Scrape, batch, read, crawl, watches and quote accept:

fieldvaluesdefault
egress"auto": the ladder decides; "direct": our servers only, never charged for a network; "residential": start on residential"auto" (crawl: "direct")
or the object { "max": "auto" | "direct" | "residential", "country": "de" }
countrytwo-letter exit country for residentialthe site's ccTLD, else us

Anything else is 400 bad_egress. Residential needs an account that has bought credits at least once. Otherwise a blocked page answers 402 needs_residential (subtype: not_entitled), with estimatedCredits.

residential price (charged only on success)credits
HTML up to 256 KB on the wire0.05
HTML up to 1 MB on the wire0.15
HTML over 1 MBrefused too_heavy, 0
browser render (1 MB budget)0.20

POST /v1/compile and the demo use our servers only in v1.


Streaming scrape (SSE)

Send POST /v1/scrape with Accept: text/event-stream to watch the work happen. The events are progress, attempt, then exactly one result or error. The stream then closes. A comment line (: ping) goes out every 10 s, and the ladder gets 60 s instead of 20.

Example

id: 1
event: progress
data: {"stage":"fetch","status":"start","message":"Fetching","elapsedMs":0}

id: 2
event: attempt
data: {"rung":"direct","mode":"http","class":"ok","status":200,"bytes":2469,"ms":248,"index":0,
       "elapsedMs":249,"next":null,"message":"Our servers · 200 · 2 KB", …}

id: 4
event: progress
data: {"stage":"compile","status":"start","message":"Learning the layout","elapsedMs":251}

id: 6
event: error
data: {"ok":false,"error":{"reason":"llm_unconfigured", …},"meta":{…},"requestId":"req_…","status":503}
  • Stages: queued, waiting (an identical request is already running), cache, fetch, render, compile, verify, heal, extract, page. Each has a status of start, done or skipped.
  • An attempt is one network attempt. next names the network and mode tried after it.
  • result carries the same body the plain call returns. error adds the HTTP status the plain call would have used.
  • Failures before the stream starts (validation, auth, an idempotency conflict) are plain JSON with their real status.
  • Streams are not resumable, so there is no Last-Event-ID.

Health, pricing, whoami

GET /v1/health

No credential. It always answers 200; ready says whether the rest of /v1 is serving.

json

{ "ok": true, "service": "scraipe", "version": 1, "ready": true, "phase": "ready", "requestId": "req_…" }

While the server restores state from Mongo, phase is connecting_database or restoring_state. During that time every stateful /v1 route answers 503 { reason: "starting", retryable: true } with Retry-After: 5. degraded: true means a store could not be restored.

GET /v1/pricing

Public. Packs come from Stripe and are cached for 5 minutes. When Stripe is unreachable, packs is the last snapshot and livePrices is false.

json

{ "ok": true, "currency": "usd", "creditUsd": 0.01, "freeCredits": 25, "freeCreditsOn": "signup",
  "compileCredits": 1,
  "packs": [ { "priceId": "price_…", "name": "…", "amount": 1000, "currency": "usd",
               "credits": 1000, "perCredit": 0.01, "badge": null, "recurring": null } ],
  "egressPrices": {
    "unit": "millicredits",
    "residential": { "http": { "standard": 50, "heavy": 150 }, "browser": { "standard": 200, "heavy": 200 } },
    "unblocker": { "available": false },
    "free": ["direct", "cache", "conditional"],
    "priceClasses": { "http": { "standardMaxWireBytes": 262144, "maxWireBytes": 1048576 },
                      "browser": { "standardMaxWireBytes": 1048576, "maxWireBytes": 1048576 } },
    "chargedOn": "success" },
  "policy": { "creditsExpire": false, "expiryDays": null,
              "refunds": "Blocked and failed requests are never charged.",
              "failedRequestsCharged": false, "residentialRequiresPurchase": true },
  "asOf": "2026-09-24T04:16:57.389Z", "livePrices": true, "requestId": "req_…" }

GET /v1/whoami

Any valid credential. It returns who you are and what you may spend.

json

{ "ok": true,
  "principal": { "type": "account_key", "keyId": "6ab4…", "name": "default", "prefix": "sk_scr_-S7Gy",
                 "scopes": ["scrape", "read", "compile", "watches"], "userId": "6ab4…", "email": "you@example.com" },
  "account": { "creditsRemaining": 25, "everPaid": false, "paidEgress": false },
  "caps": { "monthlyCap": null, "monthlyUsed": 0, "capResetsAt": null, "maxCreditsPerRequest": null,
            "networkCap": null, "dailyCapCredits": null, "dailySpent": 0 } }

everPaid (and paidEgress) say whether the residential network is open to this account. v1 has no per-key caps: the caps fields are always empty.


POST /v1/scrape

Extract structured data from a page. Scope scrape.

bash

curl -X POST localhost:5001/v1/scrape \
  -H "Authorization: Bearer sk_scr_..." -H 'Content-Type: application/json' \
  -d '{ "url": "https://quotes.toscrape.com/", "fields": ["quote", "author"], "maxTokens": 150 }'

json

{ "ok": true,
  "data": [ { "quote": "“The world as we have created it is a process of our thinking...”",
              "author": "Albert Einstein" } ],
  "returned": 4, "total": 10, "truncated": true, "cursor": "4",
  "provenance": [ { "quote":  { "selector": ".text",   "found": 1, "index": 0, "raw": "“The world as we...”" },
                    "author": { "selector": ".author", "found": 1, "index": 0, "raw": "Albert Einstein" } } ],
  "meta": { "tier": "cache", "latencyMs": 7, "tokens": 0, "grounded": true, "specVersion": 1,
            "credits": { "compile": 0, "egress": 0, "total": 0 }, "creditsCharged": 0,
            "egress": { "rung": "none", … }, "domain": "quotes.toscrape.com", "schemaHash": "sha256:…" },
  "requestId": "req_…" }
fieldtypenotes
urlstringrequired
fieldsarray | object["title","price"], or {"price": "in USD"} for hints. Suffix [] for repeating values: "tags[]"
schemaJSON Schemaalternative to fields, for strict typing
maxTokensnumbercap the response; you get a cursor instead of a silent cut
cursorstringcontinue a truncated response; pass it back verbatim
followobject{ "next": "auto", "maxPages": 20 }; see Pagination
headers / cookiesobjectyour own authorized session; see below
egress / countrysee Networks
noCachebooleanskip the cache tier
forceCompilebooleanrebuild the extractor even if one exists (1 credit)
forceBrowserbooleanforce headless rendering (otherwise detected)

Flags are on only for true, or the string "true".

What meta tells you:

fieldmeaning
tierwhich tier answered. Billable: compiled-fresh, healed, llm
variantwhich layout variant of this site's extractor ran (up to 8 per site and schema)
empty: truea known layout that legitimately has no records right now: data: [], free
coalesced: trueanother request was compiling this exact extractor; you waited and were served free
credits, egresswhat you paid, and how the page was fetched

Provenance. Every value carries the selector that produced it and the raw text it matched. A value with no selector behind it cannot exist.

Pagination. follow walks "next" links with the same extractor, so it costs one fetch per page and no extra tokens. next takes a CSS selector, or "auto" to detect the usual patterns. meta.stoppedBecause is maxPages or no next link, so you can tell "the list ended" from "you capped me".

Your own session. Scraipe does not defeat logins. Supply your authorized session instead: { "cookies": { "session": "abc123" }, "headers": { "Authorization": "Bearer your-token" } }.

  • Credentialed responses never enter the shared cache, and never shape the shared routing policy.
  • Credentials are dropped when a redirect leaves the requested host.
  • You cannot set User-Agent, Host or sec-ch-ua*.

POST /v1/quote

What a scrape would cost, before you run it. Scope read. Nothing is fetched.

json

{ "url": "https://books.toscrape.com/", "fields": ["title", "price"] }

json

{ "ok": true, "warm": false, "specVersion": null, "schemaHash": "sha256:80d8…", "variant": null, "cached": false,
  "estimatedCredits": { "compile": 1, "egress": 0, "max": 1 },
  "network": { "rung": "direct", "mode": "http", "priceClass": null, "successRate7d": null,
               "learned": false, "country": null },
  "requiresPurchase": false }
  • estimatedCredits.max is the worst case: learning the layout plus the dearest residential page.
  • network is what Scraipe learned about the host.
  • refused appears when the host is refusing us right now.
  • requiresPurchase: true means the page needs residential and the account has never bought credits.
  • successRate7d is shown only for hosts this account has requested.

POST /v1/batch

Many URLs, one schema, one call. Scope scrape.

json

{ "urls": ["https://site.com/1", "https://site.com/2"], "fields": ["title", "price"], "concurrency": 5 }

json

{ "ok": true, "count": 2, "succeeded": 2, "tiers": { "compiled": 2 }, "totalTokens": 0,
  "credits": { "compile": 0, "egress": 0, "total": 0 },
  "results": [ { "url": "https://site.com/1", "requestId": "req_…", "ok": true, "data": [ … ], "meta": { … } } ] }
  • At most 100 URLs. concurrency is clamped to 1–10, default 5.
  • Each URL is its own request-log entry, with its own requestId.
  • An entry that is not an absolute http(s) URL gets its own invalid_url result.
  • Per-origin rate limits still apply.
  • Each URL reserves its own credit, so a batch stops charging when the balance reaches zero. The remaining URLs get out_of_credits.
  • Accepts egress/country and maxTokens.

POST /v1/compile

Pre-warm an extractor, so the slow first call never blocks a real request. Scope compile.

json

{ "url": "https://site.com/product/1", "fields": ["title", "price"],
  "samples": ["https://site.com/product/2", "https://site.com/product/3"] }

json

{ "ok": true, "domain": "site.com", "specVersion": 1, "variant": "…", "groundedness": 1,
  "brittleness": { "ok": true, "checked": 2 }, "root": "div.product", "fields": ["title", "price"],
  "tokens": 2618, "latencyMs": 7420, "creditsCharged": 1, "meta": { "tier": "compiled-fresh", … } }
  • Pass samples (at most 5, each an absolute http(s) URL). Selectors are checked against them, which rejects extractors that only work on the page they were born on.
  • alreadyWarm: true means an extractor for this layout exists. Nothing is charged, and force: true rebuilds it.
  • Failures use the plain error envelope, plus top-level tokens, latencyMs and creditsCharged: 0.
  • Compile fetches from our servers only in v1. For a site that blocks them, run POST /v1/scrape instead: it compiles on first use.

POST /v1/read

A page as clean Markdown. Free from our servers: no model is involved. Scope read.

bash

curl -X POST localhost:5001/v1/read -H "Authorization: Bearer sk_scr_..." \
  -H 'Content-Type: application/json' -d '{ "url": "https://books.toscrape.com/", "maxChars": 500 }'

json

{ "ok": true, "url": "https://books.toscrape.com/",
  "title": "All products | Books to Scrape - Sandbox", "description": null, "author": null,
  "publishedAt": null, "siteName": "books.toscrape.com", "image": null, "language": "en-us",
  "content": "**Warning!** This is a demo website for web scraping purposes…", "format": "markdown",
  "links": [ { "href": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html", "text": "" }, … ],
  "stats": { "chars": 500, "words": 46, "estimatedTokens": 125, "readingTimeMinutes": 1,
             "extractedVia": "scored", "truncated": true },
  "meta": { "tier": "read", "fetchTier": "http", "latencyMs": 1400, "tokens": 0, "originRequests": 1,
            "credits": { "compile": 0, "egress": 0, "total": 0 }, "creditsCharged": 0,
            "egress": { "rung": "direct", "mode": "http", "wireBytes": 5841, "attempts": [ … ], … } } }
fieldnotes
formatmarkdown (default), text, html
mainOnlystrip navigation, ads and footer. Default true
maxCharstruncate (stats.truncated tells you)
includeLinksdefault true. With false, link text is kept as plain text
forceBrowser, headers, cookies, egress, countryas for scrape
  • A page fetched through residential costs its network price (0.05 / 0.15 / 0.20), and only if the read succeeds.
  • Markdown conversion has a 3 s budget. A page over it comes back as plain text: check format in the response. A read that takes more than 15 s is refused too_complex.
  • Each key may have 4 reads in flight at once (429 concurrency_limit). A full reader queue answers 503 reader_busy.

POST /v1/map

The URLs a site has. It reads the sitemap when there is one, which is one request instead of a crawl. Free. Scope read.

json

{ "url": "https://books.toscrape.com/", "limit": 3, "search": "/catalogue/" }

json

{ "ok": true, "source": "links", "total": 6, "truncated": true,
  "urls": [ { "url": "https://books.toscrape.com/index.html", "text": "Books to Scrape", "lastModified": null }, … ],
  "meta": { "tier": "map", "tokens": 0, "credits": { "compile": 0, "egress": 0, "total": 0 }, … } }
  • source is sitemap:<url>, links (the fallback) or none.
  • Results are freshest first when the sitemap has lastmod.
  • limit defaults to 500, maximum 5,000.
  • includeSubdomains is off by default.
  • Map uses our servers only.

POST /v1/crawl

Traverse a site. It returns a job at once. Scope scrape.

json

{ "url": "https://example.com/blog", "limit": 25, "maxDepth": 2, "mode": "read",
  "include": ["*/blog/*"], "exclude": ["*/tag/*"], "webhook": "https://your-app.com/hooks/done" }

json

{ "ok": true, "jobId": "job_wduvN0d7z-eo", "status": "queued", "poll": "/v1/jobs/job_wduvN0d7z-eo",
  "message": "Crawl started. Poll the job, or supply \"webhook\" to be notified.",
  "webhookSecret": "whsec_…" }
fieldnotes
moderead: Markdown per page, free from our servers. extract: structured, needs fields or schema
limit / maxDepthmax pages (default 25) / link depth (default 2)
include / excludeURL globs
formatfor read mode
egress / countrydefault "direct": a crawl uses residential only when you ask
webhookPOST the finished job here. webhookSecret is returned once, to verify signatures
  • The crawl seeds from the sitemap when there is one.
  • One bad page does not end the crawl: errors are collected and reported.
  • Each page is its own request-log entry.
  • Each key may have 2 crawls in progress at once (429 concurrency_limit).

Jobs

bash

GET    /v1/jobs?limit=&type=     # your jobs, newest first (summaries)
GET    /v1/jobs/:id              # the full job, with its result
DELETE /v1/jobs/:id              # cancel if running, and delete

json

{ "ok": true,
  "job": { "id": "job_…", "type": "crawl", "status": "completed",
           "progress": { "done": 2, "total": 2, "message": "crawled 2/2" },
           "input": { "url": "https://books.toscrape.com/", "limit": 2, "maxDepth": 1, "egress": "direct" },
           "result": { "pagesCrawled": 2, "urlsDiscovered": 20, "errorCount": 0, "pages": [ … ] },
           "creditsCharged": 0, "credits": { "compile": 0, "egress": 0, "total": 0 } } }
  • Status: queued → running → completed | failed | cancelled | interrupted. interrupted means the server restarted mid-job, and it is safe to resubmit.
  • Retention. Jobs expire after 7 days.
  • Webhook outcome. A job with a webhook records its delivery outcome as delivery: { at, ok, status, error }.
  • Paging. GET /v1/jobs does not page in v1. It returns up to limit (default 50) and no nextCursor.

Watches

A watch re-runs a scrape on a schedule and reports only what changed.

bash

curl -X POST localhost:5001/v1/watches -H "Authorization: Bearer sk_scr_..." \
  -H 'Content-Type: application/json' \
  -d '{ "url": "https://example.com/product/1", "fields": ["title", "price", "stock"],
        "keyFields": ["title"], "ignoreFields": ["viewCount"], "intervalMs": 86400000,
        "notifyOn": "changes", "webhook": "https://your-app.com/hooks/scraipe" }'

It answers 201 with watch, webhookSecret (shown once), estimatedCredits (first run, per run, monthly) and note.

fieldnotes
keyFieldswhat identifies a record across runs, for example ["sku"]. Inferred if omitted
ignoreFieldsfields whose changes are noise
intervalMsminimum 5 minutes. Default daily
notifyOnchanges (default), always, never. Failures notify unless never
labelat most 200 characters; defaults to the hostname
webhookan http(s) URL; null or "" removes it (on PATCH)
headers / cookiesyour own session, kept with the watch (create only)
egress / countrydefault "auto"

bash

GET    /v1/watches?status=active|paused|failing&cursor=&limit=   # list, with counts
GET    /v1/watches/:id            # detail + recent run history
GET    /v1/watches/:id/runs       # run history, newest first (cursor = a run id)
GET    /v1/watches/:id/snapshot   # the records the next run is diffed against (404 before the first success)
PATCH  /v1/watches/:id            # label, intervalMs, webhook, notifyOn, active, keyFields, ignoreFields, egress, country
DELETE /v1/watches/:id
POST   /v1/watches/:id/run        # check now

POST /v1/watches/:id/run reports what the run actually did:

statusmeaning
200the run succeeded: { ok, watch, run, changes }
402it needed a credit, or residential, and the account cannot pay
409the watch is already running (the scheduler or another click); retryable
422it ran and failed; error is the run's own reason (for example bot_protection)
429 / 503refused for a temporary budget or our outage, with Retry-After
500run_failed: our bug

A run returns the diff:

json

{ "run": { "id": "run_…", "trigger": "manual", "ok": true, "tier": "compiled", "tokens": 0,
           "credits": { "compile": 0, "egress": 0, "total": 0 }, "recordCount": 19,
           "diff": { "added": 0, "changed": 1, "removed": 0, "unchanged": 18 } },
  "changes": { "changed": [ { "key": "A Light in the Attic",
                              "fields": { "price": { "from": "£99.99", "to": "£51.77" } } } ],
               "added": [], "removed": [], "hasChanges": true } }

When a site refuses us, the watch backs off. After bot_protection, rate_limited or blocked, the next run waits twice as long per consecutive refusal. The wait is capped at a day (or the interval, if longer), and is never shorter than the site's Retry-After. refusedStreak counts refusals, and a success resets it.

Webhooks and signatures

Watch deliveries use User-Agent: Scraipe-Watch/1.0, and crawl deliveries Scraipe-Jobs/1.0. Both are POST with JSON, follow no redirects, and carry:

Example

X-Scraipe-Event: watch.changed            (watch.ran, watch.failed; job.completed, job.failed, …)
X-Scraipe-Delivery: dlv_…
X-Scraipe-Signature: t=1790223331,v1=5f2c…   (when the watch or crawl has a secret)

json

{ "id": "dlv_…", "event": "watch.changed", "createdAt": "…",
  "watch": { "id": "wch_…", "label": "poetry prices", "url": "…" },
  "run": { … },
  "changes": { "added": [ … ], "changed": [ … ], "removed": [ … ] } }

A crawl delivery is { id, event, createdAt, job, result }. Check the signature over the raw body, and reject a timestamp more than 5 minutes old. This is the verifier the test suite runs against real deliveries (webhookSign.test.js):

js

// Node 18+. rawBody is a Buffer: in Express, use express.raw({ type: 'application/json' }) on this route.
const crypto = require('crypto');
function verifyScraipe(rawBody, header, secret, toleranceS = 300) {
  const pairs = String(header).split(',').map((kv) => kv.trim().split('='));
  const t = Number((pairs.find(([k]) => k === 't') || [])[1]);
  if (!t || Math.abs(Date.now() / 1000 - t) > toleranceS) return false;
  const expected = crypto.createHmac('sha256', secret).update(`${t}.`).update(rawBody).digest();
  return pairs.filter(([k]) => k === 'v1').some(([, v]) => {
    const got = Buffer.from(v || '', 'hex');
    return got.length === expected.length && crypto.timingSafeEqual(got, expected);
  });
}

Each delivery is attempted once. The outcome is recorded on the run (run.delivery) or the job (job.delivery).


Registry

bash

GET /v1/registry?q=&status=healthy|repaired|failing|llm_fallback&cursor=&limit=
GET /v1/registry/:domain/:schemaHash
  • The list shows the extractors your account has used, most recently used first, with total and nextCursor. Each entry has fields, versions, variants, status, lastUsedAt, freeRequestsServed and creditsSpent.
  • The detail adds the current spec, its version history, its layout variants, the host's network line, and your watches that use it.

Which domains you scrape is competitive information, so it is never visible to other tenants:

  • an account key or session sees its own account's records;
  • a self-host key (no account) and open mode see records that belong to no account;
  • only an operator key with the admin scope sees everything. An admin scope on an account key is ignored.

A resource you cannot see is a 404, never a 403.


Network (hosts)

What the console's Network page reads: how Scraipe reaches each host you have requested. Scope read.

bash

GET /v1/egress/domains?q=&attention=true&cursor=&limit=
GET /v1/egress/domains/:host        # 404 unless this account requested the host in the last 30 days

json

{ "ok": true,
  "domain": { "host": "books.toscrape.com", "rung": "direct", "mode": "http", "status": "works_direct",
              "learned": null, "successRate7d": 1, "requests7d": 8, "blocks7d": 0, "credits7d": 0,
              "watchesAffected": 0, "nextProbeAt": null, "override": null,
              "byRung": [ { "rung": "direct", "mode": "http", "attempts7d": 9, "successRate7d": 1 } ],
              "priceClass": { "http": null, "browser": null },
              "costPer1k": { "credits": 0, "usd": 0 },
              "lastBlock": null,
              "robots": { "state": "absent", "checkedAt": "2026-09-24T04:16:57.421Z", "allowsRoot": true },
              "optedOut": false, "history": [] } }
  • status is works_direct, needs_residential, refused, opted_out or unknown.
  • The list also returns a 7-day summary: fetches7d, freeNetworkShare, residentialPages7d, residentialCredits7d and blocks7d.
  • attention=true keeps hosts that are refused, opted out, need residential, or were blocked this week.
  • Read-only in v1. The routing is learned; there are no per-host overrides.

Demo (no sign-up)

For the landing page. No credential, never billed, our servers only. It is limited per IP per UTC day: 60 scrapes and 20 reads.

bash

GET  /v1/demo/examples        # { examples: [{ id, label, url, mode, fields }], limits: { scrapePerDay, readPerDay } }
POST /v1/demo/scrape          # { "example": "books", "fields"?: [subset of the example's] }
POST /v1/demo/read            # { "url": any public page, "format"?: "markdown" | "text", "maxChars"? (≤ 20,000) }
  • Scrape runs only the listed examples. Anything else is 403 demo_not_allowed. New pages would mean paid model calls.
  • Read takes any public URL. robots.txt and opt-outs apply.
  • meta.demoRemaining counts down. Past the limit, 429 demo_limit_reached with Retry-After set to midnight UTC.
  • A site that blocks our servers answers 422 access_denied with a hint that an account can use residential.

Opt-out

POST /v1/optout is public. Site owners use it through the form on /bot. It accepts JSON or a form post; a form post gets a small HTML page back.

json

{ "domain": "example.com", "email": "owner@example.com", "reason": "…", "scope": "domain" }

json

{ "ok": true,
  "optout": { "id": "opt_9bWKJWETQgM", "domain": "example.com", "scope": "domain", "paths": [],
              "status": "pending_review",
              "verification": { "methods": ["dns"], "emailSentTo": null,
                                "dns": { "type": "TXT", "name": "_scraipe-optout.example.com",
                                         "value": "scraipe-optout=opt_9bWKJWETQgM" },
                                "expiresAt": "2026-09-27T04:17:06.013Z" } } }
  • It answers 202. A request takes effect only once the operator approves it.
  • scope: "paths" with paths: ["/private/"] limits the opt-out to those paths.
  • It is limited to 5 requests per IP per day.

Operator key (admin scope) only:

bash

GET  /v1/admin/optout?status=pending_review|approved
POST /v1/admin/optout/:id/approve      # { "note"?: "TXT record checked" }

Once approved, the domain and its subdomains are refused on every network, redirects included (422 site_opted_out).


Hosted MCP

POST /mcp is MCP over Streamable HTTP. It is stateless, and takes only Authorization: Bearer sk_scr_…. Send Accept: application/json, text/event-stream.

bash

curl -s localhost:5001/mcp -H "Authorization: Bearer sk_scr_..." \
  -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'
  • Tools: scrape, compile, batch, list_extractors, read, map_site, crawl, watch, egress_quote. Tools that fetch take egress and country.
  • Metering. Each call runs as the key's account: limited by its scopes, charged to its credits, and logged with source: "mcp".
  • Response size. Responses are capped at 8,000 tokens by default and return a cursor.
  • Refusals. A tool's refusal is a result with isError: true and readable text: the reason, whether a retry helps, and the vendor and networks tried.
  • Before dispatch, failures are the REST envelope: 401 with WWW-Authenticate: Bearer realm="scraipe", 503 starting, and 406 without the Accept header.
  • GET /mcp answers 405, and DELETE /mcp answers 204.

Account API

The console's API. It uses a session token, except where marked public. It needs a database; without one it answers 503 accounts_unavailable.

Auth

bash

POST /api/auth/register   { email, password, name? }   → 201 { token, user, defaultKey }   (public)
POST /api/auth/login      { email, password }          → { token, user }                   (public)
GET  /api/auth/user                                    → { user }
PUT  /api/auth/update     { name }
PUT  /api/auth/password   { currentPassword, newPassword } → { token }
POST /api/auth/request-password-reset  { email }       → always the same generic answer     (public)
POST /api/auth/reset-password          { token, newPassword }                              (public)
  • Register returns the account's first API key in defaultKey.key, once. The account gets 25 free credits, and the first ledger line records them.
  • user has creditsRemaining, everPaid and isPaidUser.
  • A wrong current password is 400 invalid_password, not a 401: the session is still valid.
  • Login, registration and reset are rate limited per IP, and current-password checks per account: 429 rate_limited with Retry-After.
  • Changing or resetting the password invalidates every earlier session, on /v1 too.
  • Without Mailgun configured, reset emails are not sent.

Keys

bash

GET    /api/account/keys                    → { keys: [{ id, name, prefix, scopes, createdAt, lastUsedAt, revoked, … }], maxKeys: 20 }
POST   /api/account/keys   { name, scopes? } → 201 { key: { …, key: "sk_scr_…" } }   ← the secret, shown ONCE
DELETE /api/account/keys/:id                → revoke

We store only a hash of each key, so it cannot be recovered. An account may have 20 active keys (409 key_limit_reached).

Activity

bash

GET /api/account/usage?from=YYYY-MM-DD&to=YYYY-MM-DD
GET /api/account/requests?source=&outcome=&host=&parentRequestId=&cursor=&limit=
GET /api/account/requests/:id
GET /api/account/ledger?cursor=&limit=
GET /api/account/extractors
  • usage is a daily series from the request log (default: the last 30 days). It has totals, with free, charged, residential, failed, blocked, credits and freeRatio, and breakdowns byKey, byDomain, bySource and egress. The old lifetime counters stay under usage for one release.
  • requests is the request log: one entry per URL, kept 30 days. An entry has source (api, mcp, watch, crawl, playground, demo), endpoint, tier, credits, egress, outcome and reason. /:id adds the full receipt, with every network attempt. The id is the requestId of the call.
  • ledger lists grants, purchases and charges, newest first, each with balanceAfter.

json

{ "ok": true, "entries": [ { "id": "led_signup_…", "at": "…", "type": "grant", "credits": 25, "balanceAfter": 25,
                             "description": "Free credits on sign-up", "ref": { "kind": "signup", … } } ],
  "balance": 25, "nextCursor": null }

Your data

bash

GET    /api/account/export               # everything we hold about you, as one JSON download
DELETE /api/account   { password }        # erase the account, its watches, request log and ledger

The shared extractor registry is not deleted: it holds no personal data, and other tenants use it. Your account is removed from its attribution. The answer says how many unused credits were forfeited.

Payments

bash

GET  /api/payment/get-prices                          # public: the packs, from Stripe
POST /api/payment/create-checkout-session { priceId, returnTo? } → { url }   # Stripe Checkout
GET  /api/payment/session/:id                         → { paid, applied, credits, balanceAfter, status, … }
POST /api/payment/portal { returnTo? }                → { url }   # Stripe's customer portal; 409 before a first purchase
POST /api/payment/webhook                             # Stripe only; signature-verified
  • The webhook grants credits and sets everPaid, which opens residential.
  • The checkout success page polls session/:id until applied: true.
  • Another account's session is a 404.
  • Without Stripe configured: 503 payments_unavailable.

Rate limits and politeness

  • Per origin, not per customer. A batch of 100 URLs from one site is still polite.
  • robots.txt and Crawl-delay are obeyed. robots.txt is read on the network the page is fetched from. Hacker News declares 30 s, so repeat requests there are slow on purpose.
  • Rate limits. When a site answers 429, or 503 with Retry-After, we back that origin off on every network and answer 422 rate_limited with retryAfter. A rate limit is never retried through another network.

Not in v1

backend/openapi.json describes these operations, and marks each "x-status": "not-in-v1". The v1 server answers 404 not_found for all of them.

areaoperations
discoveryGET /openapi.json, GET /.well-known/http-message-signatures-directory (Web Bot Auth)
jobsPOST /v1/jobs/:id/cancel (use DELETE), GET /v1/jobs/:id/download
watchesPOST /v1/watches/:id/test-webhook, GET /v1/watches/:id/deliveries, POST /v1/watches/:id/rotate-secret
registryDELETE /v1/registry/:domain/:schemaHash, POST /v1/registry/:domain/:schemaHash/recompile (use POST /v1/compile with force: true)
networksGET /v1/egress/quote (use POST /v1/quote)
opt-outGET /v1/optout/confirm, POST /v1/admin/optout/:id/reject
keysPATCH /api/account/keys/:id (renames and per-key caps)
accountGET /api/account/blocked; /api/account/network-policy (4 operations); /api/account/proxies (3); GET/PUT /api/account/egress; GET/PUT /api/account/billing-settings; GET/PUT /api/account/onboarding

The following request fields are not read in v1:

  • maxCredits, maxJobCredits and maxCreditsPerRun;
  • the egress object's sessionId, maxCredits, render, proxy and unblocker;
  • the datacenter and unblocker network values (400 bad_egress);
  • per-key caps on key creation;
  • egress and country on POST /v1/compile.