Anti-bot & reverse engineering

Hard sites, solved.

Most scraping teams stop at a 403. We don't. Here is what the common protections do and how we work with them — for public data, professionally.

The protections

What each one does, and how we work with it.

Technique level, on purpose. The specifics change monthly and belong in an engagement, not on a web page.

Cloudflare

01
What it does
Sits in front of a large share of the web. Serves JavaScript challenges, Turnstile and managed rules that score each request on TLS, headers and behaviour before the origin ever sees it.
Our approach
Browser-accurate TLS and HTTP/2 fingerprints, consistent session handling, and residential or ISP proxies that match the expected traffic. Headless browsers only for the pages that truly need them.

Akamai Bot Manager

02
What it does
Collects sensor data from the browser (mouse, timing, device signals) and scores sessions over time. Cookies are issued per session and lose trust quickly when the signals look synthetic.
Our approach
Real browser sessions where sensor data is required, kept warm and reused across requests, with pacing that matches a person reading the site. We do not ship or sell sensor generators.

DataDome

03
What it does
Combines device fingerprinting, IP reputation and behavioural checks, and returns a captcha or a block page when the score drops. Decisions are made per request at the edge.
Our approach
Stable device profiles per session, clean IP pools scoped to the target, and endpoint selection that avoids the most heavily guarded routes when the same data is served elsewhere.

PerimeterX / HUMAN

04
What it does
Issues cookie-based challenges and the press-and-hold captcha. Tracks pointer movement and timing to separate people from scripts.
Our approach
Session reuse, human-paced request budgets, and browser automation only for the challenge step. Collection itself runs over plain HTTP once a session is established.

Kasada

05
What it does
Client-side proof-of-work tokens generated by obfuscated JavaScript that changes often. Requests without a fresh, valid token are rejected outright.
Our approach
Token generation inside a controlled pool of real browser sessions, monitored for script changes so breakage is noticed within hours rather than days.

reCAPTCHA / hCaptcha

06
What it does
Image and invisible captchas that gate forms, search and sometimes every page load. Scores depend on account history, IP and interaction.
Our approach
Avoid first: there is usually a route that does not trigger the captcha. Where one is unavoidable and the site's terms allow automation, we use solving services and tell you the per-result cost up front.

TLS / JA3 / HTTP/2 fingerprinting

07
What it does
Servers inspect the TLS handshake and HTTP/2 settings. A default Python client looks nothing like Chrome and is blocked before any HTML is served.
Our approach
HTTP clients that reproduce real browser handshakes and header order, matched to the declared User-Agent and kept current as browsers ship new versions.

Mobile & private APIs

08
What it does
Many sites serve cleaner, faster JSON to their own apps than to the browser. Those endpoints are usually signed or token-gated and undocumented.
Our approach
Traffic capture from the app, reverse engineering of the signing and token flow, then a lightweight client that calls the endpoint directly. Lower load on the site, higher data quality for you.

Login & session flows

09
What it does
Some data is only shown to signed-in users: dashboards, member directories, saved searches. Sessions expire, 2FA interrupts, and parallel logins get flagged.
Our approach
Your account, your credentials, our automation: cookie-based session management, 2FA handoff where needed, and strict per-account rate limits. We never create fake accounts or use credentials that are not yours.

Rate limits & IP reputation

10
What it does
Per-IP and per-session budgets, with blocks for data-centre ranges and known proxy networks. Hitting the limit often poisons the IP for days.
Our approach
Request budgets below the site's visible limits, proxy pools sized to the job rather than the other way round, and backoff that treats every 429 as a signal, not an obstacle.

The line

What we won't do.

Getting past a protection is a technical question. Whether we should is not, and we answer it before quoting.

  • No scraping behind someone else's login or paywall
  • No personal data beyond what the site shows publicly
  • No request rates that degrade the target site
  • No CAPTCHA farms on sites that prohibit automation in their terms
  • We review robots.txt and terms per job and tell you if we decline

Questions

Hard sites, answered.

Can you scrape a site protected by Cloudflare?

Usually, yes. Most Cloudflare configurations respond to accurate browser fingerprints and well-behaved sessions; the strictest ones need a real browser for part of the flow. We test the specific site before quoting and tell you which level of protection it runs and what that means for cost and speed.

Do you sell bypass tools or scripts?

No. We deliver data, or we deliver scrapers inside an engagement where our engineer runs them in your repo. We do not publish or sell standalone bypass kits, sensor generators or captcha tokens.

How do you keep a scraper working when the protection changes?

Every scraper ships with monitoring: row-count variance, empty-field rates and block-rate alerts. When a vendor updates its script we usually know within a run. The fix is covered inside the engagement, or inside the thirty-day window on a data project.

Will this get the target site in trouble or slow it down?

No. Request budgets are sized to stay well below anything that would affect a site's performance, and we back off on every 429 or 5xx. Public data at a human-like pace is the line we hold.

What do you need from me to assess a hard site?

The URL, the fields you want, the volume and how often you need it. An engineer runs a short probe, identifies the protection stack and replies within 24 hours with doable or not doable, the approach at a high level, and a price.

Stuck at a 403? Send us the URL.

An engineer probes the site, names the protection stack and replies within 24 hours with an approach and a price.