scraipe

For site owners

Scraipe's crawler

Scraipe fetches pages for developers' AI agents and scheduled monitors. Here's how to recognise us, how we behave, and how to limit or block us.

Something wrong? Email abuse@scraipe.co. A person replies within 24 hours.

At a glance

as of 2026-09-24
robots.txt token
Scraipe
User-Agent
Scraipe/1.0
Name on every network
the same
Opt-out
every network
CAPTCHAs solved
never
Rotating to escape a block
never

01 How to recognise us

How to recognise Scraipe's requests
User-AgentScraipe/1.0 (+https://scraipe.co/bot; compiled-extractor scraper)
The same string on plain HTTP fetches and in our browser renders, on every network.
robots.txt tokenScraipe
Who asks for a pageA developer's agent calling our API, or a scheduled watch or crawl a developer set up. Logged-in pages are fetched only with that developer's own credentials for their own account.
Our addressesRequests start from our own cloud servers. Residential addresses used when a site blocks cloud servers aren't ours to list, so the User-Agent is the reliable way to recognise us.

02 Networks we use

  1. Our servers

    Every request starts here.

  2. Residential

    When a site blocks cloud servers wholesale, or bans our servers outright, we may fetch through a residential network. It is used only for developers who have paid, and never to get around a rate limit or a block of Scraipe by name.

On every network we send the same name, obey the same robots.txt and the same rate limits.

03 How we behave

  • robots.txt per RFC 9309, re-checked on the network actually used. A robots.txt we can't read because your server errored counts as "don't fetch" for now.
  • Crawl-delay honoured.
  • About 1 request per second to your site, shared by all our customers, with at most 2 at a time.
  • Retry-After and 429s honoured for everyone, on every network.
  • Conditional GET and a shared cache, so one fetch serves every customer asking for that page.
  • GET and HEAD only.

What we never do

On any network, for any customer.

04 Limit or block us

robots.txt is the quickest way. We re-check it on whichever network fetches the page.

Block us everywhere

robots.txt
User-agent: Scraipe
Disallow: /

Stops every request from us, on every network.

Slow us down

robots.txt
User-agent: Scraipe
Crawl-delay: 10

Seconds between our requests to your site.

Block at your WAF or CDN

Match the User-Agent token Scraipe.

It covers every network we use, residential included.

05 Opt out completely

Stop Scraipe fetching your domain for every customer. No account needed.

We use it only about this request.

Scope

One per line, each starting with /.

A person reviews every request. Once approved, it applies to every network we use, before robots.txt is even read.

How we check it's your site

  1. Send the form

    The page that comes back gives you a reference.

  2. Publish a DNS TXT record

    At _scraipe-optout.your-domain, with the value scraipe-optout=<reference>, within 72 hours.

  3. A person approves it

    Only the domain's owner can publish that record, so nobody else can switch you off.

What happens next

DNS record
within 72 h
Review
by a person
Covers
every network we use
Your email
only about this request

Form not going through? Email abuse@scraipe.co instead.

06 Contact a person

abuse@scraipe.co

Crawling, rate limits and opt-outs. A person replies within 24 hours.

security@scraipe.co

Security issues and vulnerability reports.