For site owners
Scraipe's crawler
Scraipe fetches pages for developers' AI agents and scheduled monitors. Here's how to recognise us, how we behave, and how to limit or block us.
Something wrong? Email abuse@scraipe.co. A person replies within 24 hours.
At a glance
as of 2026-09-24- robots.txt token
- Scraipe
- User-Agent
- Scraipe/1.0
- Name on every network
- the same
- Opt-out
- every network
- CAPTCHAs solved
- never
- Rotating to escape a block
- never
01 How to recognise us
| User-Agent | Scraipe/1.0 (+https://scraipe.co/bot; compiled-extractor scraper)The same string on plain HTTP fetches and in our browser renders, on every network. |
|---|---|
| robots.txt token | Scraipe |
| Who asks for a page | A developer's agent calling our API, or a scheduled watch or crawl a developer set up. Logged-in pages are fetched only with that developer's own credentials for their own account. |
| Our addresses | Requests start from our own cloud servers. Residential addresses used when a site blocks cloud servers aren't ours to list, so the User-Agent is the reliable way to recognise us. |
02 Networks we use
Our servers
Every request starts here.
Residential
When a site blocks cloud servers wholesale, or bans our servers outright, we may fetch through a residential network. It is used only for developers who have paid, and never to get around a rate limit or a block of Scraipe by name.
On every network we send the same name, obey the same robots.txt and the same rate limits.
03 How we behave
- robots.txt per RFC 9309, re-checked on the network actually used. A robots.txt we can't read because your server errored counts as "don't fetch" for now.
Crawl-delayhonoured.- About 1 request per second to your site, shared by all our customers, with at most 2 at a time.
Retry-Afterand 429s honoured for everyone, on every network.- Conditional GET and a shared cache, so one fetch serves every customer asking for that page.
GETandHEADonly.
What we never do
On any network, for any customer.- We never solve CAPTCHAs or interactive challenges.
- We never rotate addresses to get around a block or a rate limit.
- We never hide who we are: no stealth browsers, no forged fingerprints.
- We never log in unless a customer supplies their own credentials for their own account.
04 Limit or block us
robots.txt is the quickest way. We re-check it on whichever network fetches the page.
Block us everywhere
User-agent: Scraipe
Disallow: /Stops every request from us, on every network.
Slow us down
User-agent: Scraipe
Crawl-delay: 10Seconds between our requests to your site.
Block at your WAF or CDN
Match the User-Agent token Scraipe.
It covers every network we use, residential included.
05 Opt out completely
Stop Scraipe fetching your domain for every customer. No account needed.
How we check it's your site
Send the form
The page that comes back gives you a reference.
Publish a DNS TXT record
At
_scraipe-optout.your-domain, with the valuescraipe-optout=<reference>, within 72 hours.A person approves it
Only the domain's owner can publish that record, so nobody else can switch you off.
What happens next
- DNS record
- within 72 h
- Review
- by a person
- Covers
- every network we use
- Your email
- only about this request
Form not going through? Email abuse@scraipe.co instead.
06 Contact a person
Crawling, rate limits and opt-outs. A person replies within 24 hours.
Security issues and vulnerability reports.