On this page
The whole product rests on one idea. This page explains it, then shows how the code implements it. You should be able to read this in ten minutes and know where everything lives.
The one idea
Most AI scrapers send the page to a language model every time you ask:
Example
request 1 → fetch page → LLM reads it → data $$ 5-15s
request 2 → fetch page → LLM reads it → data $$ 5-15s
request 500 → fetch page → LLM reads it → data $$ 5-15sCost and latency are flat forever. The five-hundredth page costs exactly what the first one did.
Scraipe uses the model once, to write down how to extract the data — a set of CSS selectors. After that it just runs the selectors:
Example
request 1 → fetch → LLM writes an extractor → data $$ 5-15s
request 2 → fetch → run the extractor → data $0 ~200ms
request 500 → fetch → run the extractor → data $0 ~200msThat's it. Everything else in the codebase exists to make that safe, correct, and self-repairing.
The consequence: your cost per page falls as you scrape the same sites repeatedly. Which is why the product is aimed at recurring monitoring rather than one-off scraping.
The tiers
A request walks down a ladder and stops at the first tier that can answer it. Each tier is cheaper and faster than the one below.
Example
POST /v1/scrape
│
├── Tier 0 CACHE ~1ms 0 tokens 0 requests to the site
│ we answered this exact question recently
│
├── Tier 0.5 CONDITIONAL ~50ms 0 tokens 1 request, no body
│ we asked "changed since last time?" and the site said no (HTTP 304)
│
├── Tier 1 COMPILED ~200ms 0 tokens 1 request
│ we already have an extractor for this site+schema — just run it
│ │
│ └── check it still works (free: schema match + field presence)
│ ├── works → return
│ └── broken → fall through and heal
│
├── Tier 2 COMPILE ~5-15s $ 1 request ← the only step that costs money
│ ask the model for an extractor, verify it, store it, use it forever
│
└── Tier 3 BROWSER +1-3s
not a tier of its own — an escalation when the HTML is a JavaScript
shell and there is nothing to extract from until it rendersEverything above Tier 2 is free. That is not generosity: those tiers cost us essentially nothing, so charging for them would be pretending.
Fetching is a second, independent axis: which network the page comes from. Our own servers are free. When a site blocks cloud servers, the page can come through a residential network, which costs 0.05–0.20 credit, and only when it succeeds. That ladder has its own page: anti-blocking.md.
Every response tells you which tier served it, and how the page was fetched:
Example
X-Scraipe-Tier: compiled
X-Scraipe-Tokens: 0
X-Scraipe-Credits-Total: 0.000
X-Scraipe-Egress: direct/httpWhy an extractor can't make things up
When the model writes an extractor, it emits data, not code:
json
{
"root": "div.product",
"fields": {
"title": { "selector": "h1.name", "extract": "text", "transform": ["trim"] },
"price": { "selector": "[data-price]", "extract": "attr:data-price", "transform": ["toNumber"] }
}
}Two things follow from that, and they are the reasons the format is shaped this way:
1. It cannot hallucinate. Every value is produced by running a selector against a real DOM node. If h1.name isn't on the page, the answer is null — the extractor has no way to invent a plausible title. A model emitting JSON directly can invent one, and you cannot tell the difference from the outside. Ours ships a receipt:
json
"provenance": { "price": { "selector": "[data-price]", "raw": "£47.82", "found": 1 } }2. It cannot run code. Transforms come from a closed allowlist (trim, toNumber, absoluteUrl, …). A name not in that table is dropped, not executed. So "the model wrote the extractor" never means "the model runs arbitrary code on our servers."
See extractor-spec.md for the full format.
How it notices a site changed
An extractor built for yesterday's HTML breaks when the site is redesigned. Detecting that cheaply is the hard part — and the part that decides whether the economics survive.
The trick: hash the page's structure, not its content. Content changes constantly (prices, dates, item counts). Structure changes only when someone redeploys the site.
Example
Monday: <div class="product"><h1>Widget</h1><span class="price">£10</span></div>
Tuesday: <div class="product"><h1>Gadget</h1><span class="price">£12</span></div>
→ different content, SAME structure → extractor still works, don't touch it
Wednesday: <div class="item"><h2>Gadget</h2><span class="cost">£12</span></div>
→ structure moved → the extractor is probably broken → check, and heal if soTwo details make this work in practice, and both were found by the live benchmark rather than by reasoning:
- Scope the hash to the part of the page you extract from. Hash the whole document and a rotating ad banner invalidates everything, healing fires on every request, and the savings evaporate.
- Collapse repeated siblings. A list with 4 tags and a list with 2 tags have the same shape. Counting them made the hash move on ordinary content changes.
And the hash is a trigger to re-check, not an order to recompile. Running the extractor costs ~5ms and answers the only question that matters — does it still work? So on drift we try first and only pay for a recompile if it genuinely broke. (Some sites encode content in class names — class="star-rating Three" — so structural hashes legitimately wobble.)
One site, several layouts. A domain rarely has one page shape: listing pages, product pages, a redesign rolling out behind an A/B test. So a (domain, schema) record holds up to 8 layout variants, each with its own extractor and the set of structural hashes it is known to work on. A page runs the variant that recognises its skeleton; failing that, the others are tried (same URL template first) and the first that passes adopts the new hash. Only when none fits does anything heal — the variant that broke, or a new one — so layout B no longer "heals" layout A's extractor away and back on alternate requests.
Healing has a budget. A page whose layout keeps moving could otherwise buy a recompile on every request. Each variant is rebuilt at most SCRAIPE_HEALS_PER_DAY (default 3) times per 24 hours; past that the request is refused with heal_limit and nothing is charged. And a list page that is legitimately empty today (data: [] on a layout we know) is served free, with meta.empty, instead of being "healed".
The module map
Example
backend/
├── index.js Express app, CORS, readiness gate, boot, background workers
│
├── core/ the engine: no HTTP routes, importable on its own
│ │
│ │ ── getting a page ──────────────────────────────────────────
│ ├── urlGuard.js SSRF rules: which hosts and addresses may be fetched
│ ├── safeHttp.js the ONE outbound HTTP client: per-hop SSRF checks, total deadline,
│ │ byte and wire caps, charset decoding, signed webhooks. CI fails on any other client.
│ ├── fetcher.js fetchPage(): the page as HTML, through the egress ladder; BYO credentials
│ ├── egress/ which NETWORK fetches a page (see anti-blocking.md)
│ │ ├── router.js the ladder: direct → residential → refusal; budgets; billing per attempt
│ │ ├── classify.js every response → one class (waf_block, rate_limited, js_shell, …) + vendor
│ │ ├── policy.js per-host learned start, promotion, re-tests, refusal memory
│ │ ├── robots.js robots.txt per network, RFC 9309
│ │ ├── optout.js site owners who asked not to be fetched
│ │ ├── providers.js the residential provider (Decodo) from env; config.js has the prices
│ │ ├── sessions.js sticky exit sessions and the exit country
│ │ ├── meter.js provider spend and the daily cap
│ │ ├── resolve.js resolve and vet a target once; connectAgent.js pins the CONNECT to it
│ │ └── guardProxy.js the local proxy every Chrome connection goes through
│ ├── browser.js pooled headless Chrome (only for JS-rendered pages)
│ ├── rateLimiter.js per-origin token bucket, shared by every network
│ │
│ │ ── understanding a page ────────────────────────────────────
│ ├── spec.js the extractor format + the interpreter that runs it
│ ├── compiler.js asks the model for a spec, then VERIFIES it against the DOM
│ ├── skeleton.js structural fingerprint: "has this page been redesigned?"
│ ├── reader.js page → clean markdown (no model involved), in worker threads
│ │
│ │ ── orchestration ───────────────────────────────────────────
│ ├── pipeline.js THE TIER LADDER. Start here to understand the system.
│ ├── store.js cache + extractor registry (on disk); mirror.js copies state to Mongo
│ ├── discover.js what URLs exist on this site (sitemap-first)
│ ├── crawler.js site traversal
│ ├── jobs.js async jobs for work too slow to hold a connection open
│ ├── watches.js scheduled monitoring + change detection + signed webhooks
│ │
│ │ ── who is asking, and what it cost ─────────────────────────
│ ├── apiKeys.js API-key + session auth, scopes, credit and residential gating
│ ├── accounts.js bridges accounts (Mongo) to the engine; atomic credit reservations
│ ├── requestLog.js one entry per URL fetched, 30 days: the receipts
│ └── ledger.js grants, purchases and charges
│
├── routes/
│ ├── v1.js the agent-facing API: every /v1 endpoint, SSE streaming
│ ├── demo.js the no-signup demo (/v1/demo)
│ ├── mcpHttp.js hosted MCP (POST /mcp)
│ ├── auth.js register (returns the first key) / login / password
│ ├── account.js API keys, usage, request log, ledger, export, delete
│ └── payment.js Stripe checkout, portal + webhooks
│
├── mcp/tools.js the MCP tools, shared by hosted MCP and mcp/server.js (stdio)
└── bench/run.js live benchmark against real sitesReading order for a newcomer: core/pipeline.js (the ladder) → core/spec.js (what an extractor is) → core/compiler.js (how one gets made) → routes/v1.js (how it's exposed).
The life of a request
POST /v1/scrape { url, fields }:
routes/v1.js— normalise the input.["title","price"],{title: "hint"}and a full JSON Schema all become the same internal schema.apiKeys.js— resolve the API key or session to an account; note whether it has credits (this does not block the request — see below).pipeline.js— walk the tiers:store.js— is this cached? → returnfetcher.js→egress/router.js— fetch on the network the host needs (our servers first, residential when they are blocked and the account may use it), sendingIf-None-Matchif we have it → 304 → return cachedstore.js— do we have an extractor for this (domain, schema)? Pick the layout variant whose structural hash matches the page.- yes →
spec.jsruns it → validate → return - no, or it broke →
compiler.jswrites and verifies one → store → run it. Compiles are single-flight: concurrent requests for the same extractor wait for the one compile and are served free (meta.coalesced), and only the leader reserves a credit.
- yes →
accounts.js— record usage against the account, charging a credit only if the request reached the compile step, plus the residential price only if the page came back that way.requestLog.jsstores the receipt;ledger.jsthe charge.- Return data + provenance + tier and network metadata.
Where the credit check happens matters. It gates the compile step, not the request. A customer at zero credits can keep using extractors they already have — those cost us nothing, so blocking them would make "your bill falls as your volume repeats" false exactly when someone would notice.
What compilation actually does
This is the expensive step, so it's worth being extravagant — it runs once and is amortised over every later request. A competitor paying per-request can never spend this much.
Example
1. CONDENSE strip scripts/styles, truncate long text, collapse repeated siblings
(a 50-item grid teaches the model nothing a 2-item grid doesn't,
and costs 25x the tokens)
2. ASK the model returns selectors AND the value it expects each to produce
3. GROUND run every selector against the real DOM
model said "£47.82", selector returned "£47.82" → grounded
disagreement → reject, hand back what actually happened, retry
4. CONFORM does the output match the schema the caller asked for?
(a model can be perfectly self-consistent and still wrong —
this check exists because that happened)
5. BRITTLENESS run it against OTHER pages of the same layout
a selector that only works on the page it was born on is worthless
6. STORE versioned, with its structural fingerprint, reusable foreverOnly a spec that passes all of it is stored. If nothing can be verified, the request falls back to direct LLM extraction and says so in the response — you get an answer, but it isn't grounded and it costs full price.
Politeness is the same thing as efficiency
Scrapers get blocked because they fetch too much, too fast, anonymously. Every mechanism that makes Scraipe cheap also makes it polite — they are not two goals:
| mechanism | cost win | politeness win |
|---|---|---|
| cache hit | no fetch | zero requests to the site |
| conditional request | no body transfer | 304 is the sanctioned way to poll |
| shared cache | one fetch serves many customers | our volume grows sub-linearly with customers |
| compiled extractor | no tokens | — |
| per-origin rate limit | — | never trips a limit in the first place |
| watch back-off | no paid runs against a wall | a refusing site is revisited at 2x, 4x, 8x the interval (up to a day), never before its Retry-After |
| a refusal remembered for 6 h | no paid retries | a site that refused us by name, with a challenge, or with a ban is not asked again, by any tenant, until then |
| the rate limit is per origin, on every network | — | a site that asks us to slow down is not retried from another address |
We identify ourselves honestly (Scraipe/1.0 (+https://scraipe.co/bot)) on every network, and obey robots.txt, including Crawl-delay, read on the network the page is fetched from.
What we will not do. We do not solve CAPTCHAs or interactive challenges, and we do no stealth fingerprinting. We do not switch networks after a rate limit, or after a site blocks Scraipe by name. When a site blocks cloud servers or bans our network, Scraipe may fetch through a residential network, under the same honest name and obeying robots.txt and opt-outs. When a site still refuses, you get a structured error naming the vendor and the networks tried, and nothing is charged. anti-blocking.md explains the ladder, and the risks of the one judgement call in it (explicit bans).
Where to go next
| you want to… | read |
|---|---|
| call the API | api.md |
| understand how pages are fetched and what the networks cost | anti-blocking.md |
| understand or audit an extractor | extractor-spec.md |
| run it in production | operations.md |
| see the numbers | cd backend && npm run bench |
Something unclear? support@scraipe.co
Get a free API key