On this page
The endpoints the v1 server serves, as built on 2026-09-24. The examples are real responses from a running server, trimmed with … where they were long.
- Base URL:
http://localhost:5001, or your deployment./api/v1/*is an alias of/v1/*. - The contract is
backend/openapi.json(OpenAPI 3.1). Operations it describes that v1 does not serve are marked"x-status": "not-in-v1", and are listed at the end of this page. api-v2-changes.md is the changelog for the full contract.
Contents: Conventions · Networks · Streaming · Health, pricing, whoami · Scrape · Quote · Batch · Compile · Read · Map · Crawl and jobs · Watches · Registry · Network page · Demo · Opt-out · MCP · Account · Not in v1
Conventions
Envelope
Every JSON response has ok and a requestId, which also comes back in the X-Scraipe-Request-Id header. Success puts its data at the top level. Failure puts it in error:
json
{ "ok": false,
"error": { "reason": "llm_unconfigured",
"message": "No LLM is configured on this server (OPENAI_API_KEY is not set), so new extractors cannot be compiled.",
"retryable": false,
"hint": "Cached and compiled requests are unaffected. The operator needs to set OPENAI_API_KEY.",
"charged": false },
"meta": { "tier": "error", "credits": { "compile": 0, "egress": 0, "total": 0 }, "egress": { … } },
"requestId": "req_muf0svwrea9cfe201edc8586" }error.reasonis stable and machine-readable. Branch on it, never onmessage.error.chargedis alwaysfalse: a failed request is never charged.- A failure may also carry
retryable,retryAfter(seconds),hintandsubtype. When a network was involved it can carryclass,vendor,code,status,attempts,rungsTried,estimatedCreditsandlimitBytes. - Metered failures (scrape, batch items, read, compile) also carry
meta, with zero credits and the network attempts made. /api/*responses still carry the oldsuccessfield, and on failure a top-levelreasonandmessage, next tookanderror. Readokanderror.
Authentication and scopes
| credential | header | where |
|---|---|---|
| API key | Authorization: Bearer sk_scr_… or X-API-Key: sk_scr_… | /v1, and /mcp (Bearer only) |
| session token (7 days) | Authorization: Bearer <jwt> | /api/*, and every /v1 route (logged as source: "playground") |
| none | health, pricing, the demo, opt-out, register, login, password reset, get-prices |
SCRAIPE_OPEN_MODE=true (local only) needs no credential at all.
| scope | grants |
|---|---|
scrape | POST /v1/scrape, /v1/batch, /v1/crawl; DELETE /v1/jobs/:id |
read | POST /v1/read, /v1/map, /v1/quote; every authenticated GET under /v1 except whoami |
compile | POST /v1/compile |
watches | create, update, delete and run watches |
admin | /v1/admin/*; operator file keys only, never account keys |
New keys get all four scopes unless you ask for fewer. Keys minted before read and watches existed still hold them if they hold scrape.
Statuses
| status | meaning | what to do |
|---|---|---|
400 | the request is malformed: bad_request, invalid_url, url_not_allowed, bad_egress, bad_cursor, too_many_samples, too_many_urls, malformed_json | fix the request |
401 | credential missing or refused; error.reason says which | fix the key |
402 | money: out_of_credits (a compile or a residential page), needs_residential (the page needs the paid network and this request cannot use it) | add credits, or allow residential. Cached and compiled requests keep working |
403 | the key lacks the scope (insufficient_scope), or demo_not_allowed | use a key with the scope |
404 | no such resource, or it belongs to another account (never a 403, which would confirm it exists) | check the id |
409 | idempotency_key_reuse, idempotency_in_progress, already_running, email_taken, key_limit_reached, no_billing_account | see the reason |
413 | body_too_large (over 2 MB) | send less |
422 | understood, but the site refused, or the page does not fit | read error.reason; do not retry blindly |
429 | a temporary budget: heal_limit, layout_limit, concurrency_limit, demo_limit_reached, auth rate_limited | retry after Retry-After |
503 | ours, not yours: starting, auth_unavailable, accounts_unavailable, llm_unavailable, llm_unconfigured, spend_cap_reached, egress_unavailable, billing_unavailable, reader_busy, payments_unavailable | retry after Retry-After when it is set; your key is fine |
- 401 reasons:
token_missing,key_invalid,key_revoked,token_invalid,token_expired,password_changed(a session issued before the password changed),account_deleted. The site's refusals (422):
bot_protection: a challenge or a block page;vendoris named;access_denied: a firewall or a ban, on every network tried;geo_blocked(countryTried);rate_limited(subtype: origin,retryAfter) andrate_limited_backoff;robots_disallowed,site_opted_out,auth_required,unavailable_for_legal_reasons;http_error(subtype: not_foundfor 404/410),origin_error,timeout,network_error.
The page (422):
egress_requires_https,too_heavy(limitBytes),too_large,too_complex;needs_browser,render_failed,unsupported_content_type;extraction_failed,compile_failed.
On Heroku, a response that has not started within 30 seconds is replaced by Heroku's own HTML error page. A first-time compile of a large page can hit it. The work still finishes and the extractor is stored, so a retry with the same
Idempotency-Keyis served from the registry. Stream the request, or pre-warm withPOST /v1/compile.
Response headers
| header | example | meaning |
|---|---|---|
X-Scraipe-Request-Id | req_muf0… | every response |
X-Scraipe-Tier | compiled | which tier answered |
X-Scraipe-Latency-Ms | 214 | server time |
X-Scraipe-Tokens | 0 | LLM tokens spent |
X-Scraipe-Origin-Requests | 1 | times the site was contacted, every attempt included |
X-Scraipe-Cost-Micros | 0 | our LLM cost, when SCRAIPE_PRICE_PER_MTOK is set |
X-Scraipe-Credits-Charged | 0 | legacy: whole compile credits only |
X-Scraipe-Credits-Total | 0.050 | everything charged: compile plus network, 3 decimals |
X-Scraipe-Egress | residential/http | <network>/<mode> that got the page, or none (cache) |
X-Scraipe-Egress-Bytes | 48211 | wire bytes over all attempts |
X-Scraipe-Egress-Attempts | 2 | network attempts made |
X-Scraipe-Egress-Credits | 0.050 | network credits charged |
X-Scraipe-Idempotent-Replay | true | a stored replay of an earlier request |
Scrape, read, map and compile send the metered set. Batch sends X-Scraipe-Credits-Total only. CORS exposes all of them, plus Retry-After, Content-Disposition and Mcp-Session-Id, to the console's origins.
Money in meta
json
"meta": {
"tier": "compiled", "latencyMs": 214, "tokens": 0, "originRequests": 1,
"credits": { "compile": 0, "egress": 0.05, "total": 0.05 },
"creditsCharged": 0.05,
"egress": { "rung": "residential", "mode": "http", "provider": "decodo", "country": "us",
"wireBytes": 48211, "costMicros": 180, "credits": 0.05, "priceClass": "standard",
"private": false,
"attempts": [
{ "rung": "direct", "mode": "http", "class": "waf_block", "vendor": "cloudflare",
"code": "1020", "status": 403, "bytes": 3120, "ms": 180, "probe": false, … },
{ "rung": "residential", "mode": "http", "provider": "decodo", "country": "us",
"class": "ok", "status": 200, "bytes": 45091, "ms": 903, "probe": false, … } ] },
"requestId": "req_…" }credits.compileis 0 or 1 per page.credits.egressis the residential price: 0.05, 0.15 or 0.20.priceClassisstandardorheavy, andnullwhen nothing was charged.private: truemeans the request carried your own credentials.
Idempotency
Send Idempotency-Key: <unique per operation> on POST /v1/scrape.
- A retry with the same key and body replays the stored success (
X-Scraipe-Idempotent-Replay: true) and is not re-billed. - A duplicate that arrives while the first is still running waits for it, up to 60 s, then gets
409 idempotency_in_progresswithRetry-After: 5. - The same key with a different body (a different URL, fields, network or credentials) is
409 idempotency_key_reuse. - Scope and lifetime. Keys are scoped to the caller and live 10 minutes. Failures are not stored, so a retry after a failure is a genuine second attempt.
Lists
Lists that paginate take limit (default 50, max 200) and an opaque cursor, and return nextCursor (null at the end). These are GET /v1/watches, /v1/registry, /v1/egress/domains, /api/account/requests and /api/account/ledger. A cursor the endpoint did not issue is 400 bad_cursor. Timestamps are ISO-8601 UTC.
Networks: the egress option
Pages are fetched from our servers for free. When a site blocks them, a request may use the residential network: it is paid, and charged only on success. anti-blocking.md has the whole design.
Scrape, batch, read, crawl, watches and quote accept:
| field | values | default |
|---|---|---|
egress | "auto": the ladder decides; "direct": our servers only, never charged for a network; "residential": start on residential | "auto" (crawl: "direct") |
or the object { "max": "auto" | "direct" | "residential", "country": "de" } | ||
country | two-letter exit country for residential | the site's ccTLD, else us |
Anything else is 400 bad_egress. Residential needs an account that has bought credits at least once. Otherwise a blocked page answers 402 needs_residential (subtype: not_entitled), with estimatedCredits.
| residential price (charged only on success) | credits |
|---|---|
| HTML up to 256 KB on the wire | 0.05 |
| HTML up to 1 MB on the wire | 0.15 |
| HTML over 1 MB | refused too_heavy, 0 |
| browser render (1 MB budget) | 0.20 |
POST /v1/compile and the demo use our servers only in v1.
Streaming scrape (SSE)
Send POST /v1/scrape with Accept: text/event-stream to watch the work happen. The events are progress, attempt, then exactly one result or error. The stream then closes. A comment line (: ping) goes out every 10 s, and the ladder gets 60 s instead of 20.
Example
id: 1
event: progress
data: {"stage":"fetch","status":"start","message":"Fetching","elapsedMs":0}
id: 2
event: attempt
data: {"rung":"direct","mode":"http","class":"ok","status":200,"bytes":2469,"ms":248,"index":0,
"elapsedMs":249,"next":null,"message":"Our servers · 200 · 2 KB", …}
id: 4
event: progress
data: {"stage":"compile","status":"start","message":"Learning the layout","elapsedMs":251}
id: 6
event: error
data: {"ok":false,"error":{"reason":"llm_unconfigured", …},"meta":{…},"requestId":"req_…","status":503}- Stages:
queued,waiting(an identical request is already running),cache,fetch,render,compile,verify,heal,extract,page. Each has astatusofstart,doneorskipped. - An
attemptis one network attempt.nextnames the network and mode tried after it. resultcarries the same body the plain call returns.erroradds the HTTPstatusthe plain call would have used.- Failures before the stream starts (validation, auth, an idempotency conflict) are plain JSON with their real status.
- Streams are not resumable, so there is no
Last-Event-ID.
Health, pricing, whoami
GET /v1/health
No credential. It always answers 200; ready says whether the rest of /v1 is serving.
json
{ "ok": true, "service": "scraipe", "version": 1, "ready": true, "phase": "ready", "requestId": "req_…" }While the server restores state from Mongo, phase is connecting_database or restoring_state. During that time every stateful /v1 route answers 503 { reason: "starting", retryable: true } with Retry-After: 5. degraded: true means a store could not be restored.
GET /v1/pricing
Public. Packs come from Stripe and are cached for 5 minutes. When Stripe is unreachable, packs is the last snapshot and livePrices is false.
json
{ "ok": true, "currency": "usd", "creditUsd": 0.01, "freeCredits": 25, "freeCreditsOn": "signup",
"compileCredits": 1,
"packs": [ { "priceId": "price_…", "name": "…", "amount": 1000, "currency": "usd",
"credits": 1000, "perCredit": 0.01, "badge": null, "recurring": null } ],
"egressPrices": {
"unit": "millicredits",
"residential": { "http": { "standard": 50, "heavy": 150 }, "browser": { "standard": 200, "heavy": 200 } },
"unblocker": { "available": false },
"free": ["direct", "cache", "conditional"],
"priceClasses": { "http": { "standardMaxWireBytes": 262144, "maxWireBytes": 1048576 },
"browser": { "standardMaxWireBytes": 1048576, "maxWireBytes": 1048576 } },
"chargedOn": "success" },
"policy": { "creditsExpire": false, "expiryDays": null,
"refunds": "Blocked and failed requests are never charged.",
"failedRequestsCharged": false, "residentialRequiresPurchase": true },
"asOf": "2026-09-24T04:16:57.389Z", "livePrices": true, "requestId": "req_…" }GET /v1/whoami
Any valid credential. It returns who you are and what you may spend.
json
{ "ok": true,
"principal": { "type": "account_key", "keyId": "6ab4…", "name": "default", "prefix": "sk_scr_-S7Gy",
"scopes": ["scrape", "read", "compile", "watches"], "userId": "6ab4…", "email": "you@example.com" },
"account": { "creditsRemaining": 25, "everPaid": false, "paidEgress": false },
"caps": { "monthlyCap": null, "monthlyUsed": 0, "capResetsAt": null, "maxCreditsPerRequest": null,
"networkCap": null, "dailyCapCredits": null, "dailySpent": 0 } }everPaid (and paidEgress) say whether the residential network is open to this account. v1 has no per-key caps: the caps fields are always empty.
POST /v1/scrape
Extract structured data from a page. Scope scrape.
bash
curl -X POST localhost:5001/v1/scrape \
-H "Authorization: Bearer sk_scr_..." -H 'Content-Type: application/json' \
-d '{ "url": "https://quotes.toscrape.com/", "fields": ["quote", "author"], "maxTokens": 150 }'json
{ "ok": true,
"data": [ { "quote": "“The world as we have created it is a process of our thinking...”",
"author": "Albert Einstein" } ],
"returned": 4, "total": 10, "truncated": true, "cursor": "4",
"provenance": [ { "quote": { "selector": ".text", "found": 1, "index": 0, "raw": "“The world as we...”" },
"author": { "selector": ".author", "found": 1, "index": 0, "raw": "Albert Einstein" } } ],
"meta": { "tier": "cache", "latencyMs": 7, "tokens": 0, "grounded": true, "specVersion": 1,
"credits": { "compile": 0, "egress": 0, "total": 0 }, "creditsCharged": 0,
"egress": { "rung": "none", … }, "domain": "quotes.toscrape.com", "schemaHash": "sha256:…" },
"requestId": "req_…" }| field | type | notes |
|---|---|---|
url | string | required |
fields | array | object | ["title","price"], or {"price": "in USD"} for hints. Suffix [] for repeating values: "tags[]" |
schema | JSON Schema | alternative to fields, for strict typing |
maxTokens | number | cap the response; you get a cursor instead of a silent cut |
cursor | string | continue a truncated response; pass it back verbatim |
follow | object | { "next": "auto", "maxPages": 20 }; see Pagination |
headers / cookies | object | your own authorized session; see below |
egress / country | see Networks | |
noCache | boolean | skip the cache tier |
forceCompile | boolean | rebuild the extractor even if one exists (1 credit) |
forceBrowser | boolean | force headless rendering (otherwise detected) |
Flags are on only for true, or the string "true".
What meta tells you:
| field | meaning |
|---|---|
tier | which tier answered. Billable: compiled-fresh, healed, llm |
variant | which layout variant of this site's extractor ran (up to 8 per site and schema) |
empty: true | a known layout that legitimately has no records right now: data: [], free |
coalesced: true | another request was compiling this exact extractor; you waited and were served free |
credits, egress | what you paid, and how the page was fetched |
Provenance. Every value carries the selector that produced it and the raw text it matched. A value with no selector behind it cannot exist.
Pagination. follow walks "next" links with the same extractor, so it costs one fetch per page and no extra tokens. next takes a CSS selector, or "auto" to detect the usual patterns. meta.stoppedBecause is maxPages or no next link, so you can tell "the list ended" from "you capped me".
Your own session. Scraipe does not defeat logins. Supply your authorized session instead: { "cookies": { "session": "abc123" }, "headers": { "Authorization": "Bearer your-token" } }.
- Credentialed responses never enter the shared cache, and never shape the shared routing policy.
- Credentials are dropped when a redirect leaves the requested host.
- You cannot set
User-Agent,Hostorsec-ch-ua*.
POST /v1/quote
What a scrape would cost, before you run it. Scope read. Nothing is fetched.
json
{ "url": "https://books.toscrape.com/", "fields": ["title", "price"] }json
{ "ok": true, "warm": false, "specVersion": null, "schemaHash": "sha256:80d8…", "variant": null, "cached": false,
"estimatedCredits": { "compile": 1, "egress": 0, "max": 1 },
"network": { "rung": "direct", "mode": "http", "priceClass": null, "successRate7d": null,
"learned": false, "country": null },
"requiresPurchase": false }estimatedCredits.maxis the worst case: learning the layout plus the dearest residential page.networkis what Scraipe learned about the host.refusedappears when the host is refusing us right now.requiresPurchase: truemeans the page needs residential and the account has never bought credits.successRate7dis shown only for hosts this account has requested.
POST /v1/batch
Many URLs, one schema, one call. Scope scrape.
json
{ "urls": ["https://site.com/1", "https://site.com/2"], "fields": ["title", "price"], "concurrency": 5 }json
{ "ok": true, "count": 2, "succeeded": 2, "tiers": { "compiled": 2 }, "totalTokens": 0,
"credits": { "compile": 0, "egress": 0, "total": 0 },
"results": [ { "url": "https://site.com/1", "requestId": "req_…", "ok": true, "data": [ … ], "meta": { … } } ] }- At most 100 URLs.
concurrencyis clamped to 1–10, default 5. - Each URL is its own request-log entry, with its own
requestId. - An entry that is not an absolute
http(s)URL gets its owninvalid_urlresult. - Per-origin rate limits still apply.
- Each URL reserves its own credit, so a batch stops charging when the balance reaches zero. The remaining URLs get
out_of_credits. - Accepts
egress/countryandmaxTokens.
POST /v1/compile
Pre-warm an extractor, so the slow first call never blocks a real request. Scope compile.
json
{ "url": "https://site.com/product/1", "fields": ["title", "price"],
"samples": ["https://site.com/product/2", "https://site.com/product/3"] }json
{ "ok": true, "domain": "site.com", "specVersion": 1, "variant": "…", "groundedness": 1,
"brittleness": { "ok": true, "checked": 2 }, "root": "div.product", "fields": ["title", "price"],
"tokens": 2618, "latencyMs": 7420, "creditsCharged": 1, "meta": { "tier": "compiled-fresh", … } }- Pass
samples(at most 5, each an absolutehttp(s)URL). Selectors are checked against them, which rejects extractors that only work on the page they were born on. alreadyWarm: truemeans an extractor for this layout exists. Nothing is charged, andforce: truerebuilds it.- Failures use the plain error envelope, plus top-level
tokens,latencyMsandcreditsCharged: 0. - Compile fetches from our servers only in v1. For a site that blocks them, run
POST /v1/scrapeinstead: it compiles on first use.
POST /v1/read
A page as clean Markdown. Free from our servers: no model is involved. Scope read.
bash
curl -X POST localhost:5001/v1/read -H "Authorization: Bearer sk_scr_..." \
-H 'Content-Type: application/json' -d '{ "url": "https://books.toscrape.com/", "maxChars": 500 }'json
{ "ok": true, "url": "https://books.toscrape.com/",
"title": "All products | Books to Scrape - Sandbox", "description": null, "author": null,
"publishedAt": null, "siteName": "books.toscrape.com", "image": null, "language": "en-us",
"content": "**Warning!** This is a demo website for web scraping purposes…", "format": "markdown",
"links": [ { "href": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html", "text": "" }, … ],
"stats": { "chars": 500, "words": 46, "estimatedTokens": 125, "readingTimeMinutes": 1,
"extractedVia": "scored", "truncated": true },
"meta": { "tier": "read", "fetchTier": "http", "latencyMs": 1400, "tokens": 0, "originRequests": 1,
"credits": { "compile": 0, "egress": 0, "total": 0 }, "creditsCharged": 0,
"egress": { "rung": "direct", "mode": "http", "wireBytes": 5841, "attempts": [ … ], … } } }| field | notes |
|---|---|
format | markdown (default), text, html |
mainOnly | strip navigation, ads and footer. Default true |
maxChars | truncate (stats.truncated tells you) |
includeLinks | default true. With false, link text is kept as plain text |
forceBrowser, headers, cookies, egress, country | as for scrape |
- A page fetched through residential costs its network price (0.05 / 0.15 / 0.20), and only if the read succeeds.
- Markdown conversion has a 3 s budget. A page over it comes back as plain text: check
formatin the response. A read that takes more than 15 s is refusedtoo_complex. - Each key may have 4 reads in flight at once (
429 concurrency_limit). A full reader queue answers503 reader_busy.
POST /v1/map
The URLs a site has. It reads the sitemap when there is one, which is one request instead of a crawl. Free. Scope read.
json
{ "url": "https://books.toscrape.com/", "limit": 3, "search": "/catalogue/" }json
{ "ok": true, "source": "links", "total": 6, "truncated": true,
"urls": [ { "url": "https://books.toscrape.com/index.html", "text": "Books to Scrape", "lastModified": null }, … ],
"meta": { "tier": "map", "tokens": 0, "credits": { "compile": 0, "egress": 0, "total": 0 }, … } }sourceissitemap:<url>,links(the fallback) ornone.- Results are freshest first when the sitemap has
lastmod. limitdefaults to 500, maximum 5,000.includeSubdomainsis off by default.- Map uses our servers only.
POST /v1/crawl
Traverse a site. It returns a job at once. Scope scrape.
json
{ "url": "https://example.com/blog", "limit": 25, "maxDepth": 2, "mode": "read",
"include": ["*/blog/*"], "exclude": ["*/tag/*"], "webhook": "https://your-app.com/hooks/done" }json
{ "ok": true, "jobId": "job_wduvN0d7z-eo", "status": "queued", "poll": "/v1/jobs/job_wduvN0d7z-eo",
"message": "Crawl started. Poll the job, or supply \"webhook\" to be notified.",
"webhookSecret": "whsec_…" }| field | notes |
|---|---|
mode | read: Markdown per page, free from our servers. extract: structured, needs fields or schema |
limit / maxDepth | max pages (default 25) / link depth (default 2) |
include / exclude | URL globs |
format | for read mode |
egress / country | default "direct": a crawl uses residential only when you ask |
webhook | POST the finished job here. webhookSecret is returned once, to verify signatures |
- The crawl seeds from the sitemap when there is one.
- One bad page does not end the crawl: errors are collected and reported.
- Each page is its own request-log entry.
- Each key may have 2 crawls in progress at once (
429 concurrency_limit).
Jobs
bash
GET /v1/jobs?limit=&type= # your jobs, newest first (summaries)
GET /v1/jobs/:id # the full job, with its result
DELETE /v1/jobs/:id # cancel if running, and deletejson
{ "ok": true,
"job": { "id": "job_…", "type": "crawl", "status": "completed",
"progress": { "done": 2, "total": 2, "message": "crawled 2/2" },
"input": { "url": "https://books.toscrape.com/", "limit": 2, "maxDepth": 1, "egress": "direct" },
"result": { "pagesCrawled": 2, "urlsDiscovered": 20, "errorCount": 0, "pages": [ … ] },
"creditsCharged": 0, "credits": { "compile": 0, "egress": 0, "total": 0 } } }- Status:
queued→running→completed|failed|cancelled|interrupted.interruptedmeans the server restarted mid-job, and it is safe to resubmit. - Retention. Jobs expire after 7 days.
- Webhook outcome. A job with a webhook records its delivery outcome as
delivery: { at, ok, status, error }. - Paging.
GET /v1/jobsdoes not page in v1. It returns up tolimit(default 50) and nonextCursor.
Watches
A watch re-runs a scrape on a schedule and reports only what changed.
bash
curl -X POST localhost:5001/v1/watches -H "Authorization: Bearer sk_scr_..." \
-H 'Content-Type: application/json' \
-d '{ "url": "https://example.com/product/1", "fields": ["title", "price", "stock"],
"keyFields": ["title"], "ignoreFields": ["viewCount"], "intervalMs": 86400000,
"notifyOn": "changes", "webhook": "https://your-app.com/hooks/scraipe" }'It answers 201 with watch, webhookSecret (shown once), estimatedCredits (first run, per run, monthly) and note.
| field | notes |
|---|---|
keyFields | what identifies a record across runs, for example ["sku"]. Inferred if omitted |
ignoreFields | fields whose changes are noise |
intervalMs | minimum 5 minutes. Default daily |
notifyOn | changes (default), always, never. Failures notify unless never |
label | at most 200 characters; defaults to the hostname |
webhook | an http(s) URL; null or "" removes it (on PATCH) |
headers / cookies | your own session, kept with the watch (create only) |
egress / country | default "auto" |
bash
GET /v1/watches?status=active|paused|failing&cursor=&limit= # list, with counts
GET /v1/watches/:id # detail + recent run history
GET /v1/watches/:id/runs # run history, newest first (cursor = a run id)
GET /v1/watches/:id/snapshot # the records the next run is diffed against (404 before the first success)
PATCH /v1/watches/:id # label, intervalMs, webhook, notifyOn, active, keyFields, ignoreFields, egress, country
DELETE /v1/watches/:id
POST /v1/watches/:id/run # check nowPOST /v1/watches/:id/run reports what the run actually did:
| status | meaning |
|---|---|
200 | the run succeeded: { ok, watch, run, changes } |
402 | it needed a credit, or residential, and the account cannot pay |
409 | the watch is already running (the scheduler or another click); retryable |
422 | it ran and failed; error is the run's own reason (for example bot_protection) |
429 / 503 | refused for a temporary budget or our outage, with Retry-After |
500 | run_failed: our bug |
A run returns the diff:
json
{ "run": { "id": "run_…", "trigger": "manual", "ok": true, "tier": "compiled", "tokens": 0,
"credits": { "compile": 0, "egress": 0, "total": 0 }, "recordCount": 19,
"diff": { "added": 0, "changed": 1, "removed": 0, "unchanged": 18 } },
"changes": { "changed": [ { "key": "A Light in the Attic",
"fields": { "price": { "from": "£99.99", "to": "£51.77" } } } ],
"added": [], "removed": [], "hasChanges": true } }When a site refuses us, the watch backs off. After bot_protection, rate_limited or blocked, the next run waits twice as long per consecutive refusal. The wait is capped at a day (or the interval, if longer), and is never shorter than the site's Retry-After. refusedStreak counts refusals, and a success resets it.
Webhooks and signatures
Watch deliveries use User-Agent: Scraipe-Watch/1.0, and crawl deliveries Scraipe-Jobs/1.0. Both are POST with JSON, follow no redirects, and carry:
Example
X-Scraipe-Event: watch.changed (watch.ran, watch.failed; job.completed, job.failed, …)
X-Scraipe-Delivery: dlv_…
X-Scraipe-Signature: t=1790223331,v1=5f2c… (when the watch or crawl has a secret)json
{ "id": "dlv_…", "event": "watch.changed", "createdAt": "…",
"watch": { "id": "wch_…", "label": "poetry prices", "url": "…" },
"run": { … },
"changes": { "added": [ … ], "changed": [ … ], "removed": [ … ] } }A crawl delivery is { id, event, createdAt, job, result }. Check the signature over the raw body, and reject a timestamp more than 5 minutes old. This is the verifier the test suite runs against real deliveries (webhookSign.test.js):
js
// Node 18+. rawBody is a Buffer: in Express, use express.raw({ type: 'application/json' }) on this route.
const crypto = require('crypto');
function verifyScraipe(rawBody, header, secret, toleranceS = 300) {
const pairs = String(header).split(',').map((kv) => kv.trim().split('='));
const t = Number((pairs.find(([k]) => k === 't') || [])[1]);
if (!t || Math.abs(Date.now() / 1000 - t) > toleranceS) return false;
const expected = crypto.createHmac('sha256', secret).update(`${t}.`).update(rawBody).digest();
return pairs.filter(([k]) => k === 'v1').some(([, v]) => {
const got = Buffer.from(v || '', 'hex');
return got.length === expected.length && crypto.timingSafeEqual(got, expected);
});
}Each delivery is attempted once. The outcome is recorded on the run (run.delivery) or the job (job.delivery).
Registry
bash
GET /v1/registry?q=&status=healthy|repaired|failing|llm_fallback&cursor=&limit=
GET /v1/registry/:domain/:schemaHash- The list shows the extractors your account has used, most recently used first, with
totalandnextCursor. Each entry hasfields,versions,variants,status,lastUsedAt,freeRequestsServedandcreditsSpent. - The detail adds the current spec, its version history, its layout variants, the host's network line, and your watches that use it.
Which domains you scrape is competitive information, so it is never visible to other tenants:
- an account key or session sees its own account's records;
- a self-host key (no account) and open mode see records that belong to no account;
- only an operator key with the
adminscope sees everything. Anadminscope on an account key is ignored.
A resource you cannot see is a 404, never a 403.
Network (hosts)
What the console's Network page reads: how Scraipe reaches each host you have requested. Scope read.
bash
GET /v1/egress/domains?q=&attention=true&cursor=&limit=
GET /v1/egress/domains/:host # 404 unless this account requested the host in the last 30 daysjson
{ "ok": true,
"domain": { "host": "books.toscrape.com", "rung": "direct", "mode": "http", "status": "works_direct",
"learned": null, "successRate7d": 1, "requests7d": 8, "blocks7d": 0, "credits7d": 0,
"watchesAffected": 0, "nextProbeAt": null, "override": null,
"byRung": [ { "rung": "direct", "mode": "http", "attempts7d": 9, "successRate7d": 1 } ],
"priceClass": { "http": null, "browser": null },
"costPer1k": { "credits": 0, "usd": 0 },
"lastBlock": null,
"robots": { "state": "absent", "checkedAt": "2026-09-24T04:16:57.421Z", "allowsRoot": true },
"optedOut": false, "history": [] } }statusisworks_direct,needs_residential,refused,opted_outorunknown.- The list also returns a 7-day
summary:fetches7d,freeNetworkShare,residentialPages7d,residentialCredits7dandblocks7d. attention=truekeeps hosts that are refused, opted out, need residential, or were blocked this week.- Read-only in v1. The routing is learned; there are no per-host overrides.
Demo (no sign-up)
For the landing page. No credential, never billed, our servers only. It is limited per IP per UTC day: 60 scrapes and 20 reads.
bash
GET /v1/demo/examples # { examples: [{ id, label, url, mode, fields }], limits: { scrapePerDay, readPerDay } }
POST /v1/demo/scrape # { "example": "books", "fields"?: [subset of the example's] }
POST /v1/demo/read # { "url": any public page, "format"?: "markdown" | "text", "maxChars"? (≤ 20,000) }- Scrape runs only the listed examples. Anything else is
403 demo_not_allowed. New pages would mean paid model calls. - Read takes any public URL. robots.txt and opt-outs apply.
meta.demoRemainingcounts down. Past the limit,429 demo_limit_reachedwithRetry-Afterset to midnight UTC.- A site that blocks our servers answers
422 access_deniedwith a hint that an account can use residential.
Opt-out
POST /v1/optout is public. Site owners use it through the form on /bot. It accepts JSON or a form post; a form post gets a small HTML page back.
json
{ "domain": "example.com", "email": "owner@example.com", "reason": "…", "scope": "domain" }json
{ "ok": true,
"optout": { "id": "opt_9bWKJWETQgM", "domain": "example.com", "scope": "domain", "paths": [],
"status": "pending_review",
"verification": { "methods": ["dns"], "emailSentTo": null,
"dns": { "type": "TXT", "name": "_scraipe-optout.example.com",
"value": "scraipe-optout=opt_9bWKJWETQgM" },
"expiresAt": "2026-09-27T04:17:06.013Z" } } }- It answers
202. A request takes effect only once the operator approves it. scope: "paths"withpaths: ["/private/"]limits the opt-out to those paths.- It is limited to 5 requests per IP per day.
Operator key (admin scope) only:
bash
GET /v1/admin/optout?status=pending_review|approved
POST /v1/admin/optout/:id/approve # { "note"?: "TXT record checked" }Once approved, the domain and its subdomains are refused on every network, redirects included (422 site_opted_out).
Hosted MCP
POST /mcp is MCP over Streamable HTTP. It is stateless, and takes only Authorization: Bearer sk_scr_…. Send Accept: application/json, text/event-stream.
bash
curl -s localhost:5001/mcp -H "Authorization: Bearer sk_scr_..." \
-H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'- Tools:
scrape,compile,batch,list_extractors,read,map_site,crawl,watch,egress_quote. Tools that fetch takeegressandcountry. - Metering. Each call runs as the key's account: limited by its scopes, charged to its credits, and logged with
source: "mcp". - Response size. Responses are capped at 8,000 tokens by default and return a cursor.
- Refusals. A tool's refusal is a result with
isError: trueand readable text: the reason, whether a retry helps, and the vendor and networks tried. - Before dispatch, failures are the REST envelope:
401withWWW-Authenticate: Bearer realm="scraipe",503 starting, and406without the Accept header. GET /mcpanswers405, andDELETE /mcpanswers204.
Account API
The console's API. It uses a session token, except where marked public. It needs a database; without one it answers 503 accounts_unavailable.
Auth
bash
POST /api/auth/register { email, password, name? } → 201 { token, user, defaultKey } (public)
POST /api/auth/login { email, password } → { token, user } (public)
GET /api/auth/user → { user }
PUT /api/auth/update { name }
PUT /api/auth/password { currentPassword, newPassword } → { token }
POST /api/auth/request-password-reset { email } → always the same generic answer (public)
POST /api/auth/reset-password { token, newPassword } (public)- Register returns the account's first API key in
defaultKey.key, once. The account gets 25 free credits, and the first ledger line records them. userhascreditsRemaining,everPaidandisPaidUser.- A wrong current password is
400 invalid_password, not a 401: the session is still valid. - Login, registration and reset are rate limited per IP, and current-password checks per account:
429 rate_limitedwithRetry-After. - Changing or resetting the password invalidates every earlier session, on
/v1too. - Without Mailgun configured, reset emails are not sent.
Keys
bash
GET /api/account/keys → { keys: [{ id, name, prefix, scopes, createdAt, lastUsedAt, revoked, … }], maxKeys: 20 }
POST /api/account/keys { name, scopes? } → 201 { key: { …, key: "sk_scr_…" } } ← the secret, shown ONCE
DELETE /api/account/keys/:id → revokeWe store only a hash of each key, so it cannot be recovered. An account may have 20 active keys (409 key_limit_reached).
Activity
bash
GET /api/account/usage?from=YYYY-MM-DD&to=YYYY-MM-DD
GET /api/account/requests?source=&outcome=&host=&parentRequestId=&cursor=&limit=
GET /api/account/requests/:id
GET /api/account/ledger?cursor=&limit=
GET /api/account/extractorsusageis a daily series from the request log (default: the last 30 days). It hastotals, withfree,charged,residential,failed,blocked,creditsandfreeRatio, and breakdownsbyKey,byDomain,bySourceandegress. The old lifetime counters stay underusagefor one release.requestsis the request log: one entry per URL, kept 30 days. An entry hassource(api,mcp,watch,crawl,playground,demo),endpoint,tier,credits,egress,outcomeandreason./:idadds the full receipt, with every network attempt. The id is therequestIdof the call.ledgerlists grants, purchases and charges, newest first, each withbalanceAfter.
json
{ "ok": true, "entries": [ { "id": "led_signup_…", "at": "…", "type": "grant", "credits": 25, "balanceAfter": 25,
"description": "Free credits on sign-up", "ref": { "kind": "signup", … } } ],
"balance": 25, "nextCursor": null }Your data
bash
GET /api/account/export # everything we hold about you, as one JSON download
DELETE /api/account { password } # erase the account, its watches, request log and ledgerThe shared extractor registry is not deleted: it holds no personal data, and other tenants use it. Your account is removed from its attribution. The answer says how many unused credits were forfeited.
Payments
bash
GET /api/payment/get-prices # public: the packs, from Stripe
POST /api/payment/create-checkout-session { priceId, returnTo? } → { url } # Stripe Checkout
GET /api/payment/session/:id → { paid, applied, credits, balanceAfter, status, … }
POST /api/payment/portal { returnTo? } → { url } # Stripe's customer portal; 409 before a first purchase
POST /api/payment/webhook # Stripe only; signature-verified- The webhook grants credits and sets
everPaid, which opens residential. - The checkout success page polls
session/:iduntilapplied: true. - Another account's session is a
404. - Without Stripe configured:
503 payments_unavailable.
Rate limits and politeness
- Per origin, not per customer. A batch of 100 URLs from one site is still polite.
- robots.txt and
Crawl-delayare obeyed. robots.txt is read on the network the page is fetched from. Hacker News declares 30 s, so repeat requests there are slow on purpose. - Rate limits. When a site answers
429, or503withRetry-After, we back that origin off on every network and answer422 rate_limitedwithretryAfter. A rate limit is never retried through another network.
Not in v1
backend/openapi.json describes these operations, and marks each "x-status": "not-in-v1". The v1 server answers 404 not_found for all of them.
| area | operations |
|---|---|
| discovery | GET /openapi.json, GET /.well-known/http-message-signatures-directory (Web Bot Auth) |
| jobs | POST /v1/jobs/:id/cancel (use DELETE), GET /v1/jobs/:id/download |
| watches | POST /v1/watches/:id/test-webhook, GET /v1/watches/:id/deliveries, POST /v1/watches/:id/rotate-secret |
| registry | DELETE /v1/registry/:domain/:schemaHash, POST /v1/registry/:domain/:schemaHash/recompile (use POST /v1/compile with force: true) |
| networks | GET /v1/egress/quote (use POST /v1/quote) |
| opt-out | GET /v1/optout/confirm, POST /v1/admin/optout/:id/reject |
| keys | PATCH /api/account/keys/:id (renames and per-key caps) |
| account | GET /api/account/blocked; /api/account/network-policy (4 operations); /api/account/proxies (3); GET/PUT /api/account/egress; GET/PUT /api/account/billing-settings; GET/PUT /api/account/onboarding |
The following request fields are not read in v1:
maxCredits,maxJobCreditsandmaxCreditsPerRun;- the
egressobject'ssessionId,maxCredits,render,proxyandunblocker; - the
datacenterandunblockernetwork values (400 bad_egress); - per-key caps on key creation;
egressandcountryonPOST /v1/compile.
Something unclear? support@scraipe.co
Get a free API key