Technique

AI-assisted scraping: where models help, and where they don't

We use LLMs to infer schemas, extract from messy pages and draft scrapers faster. We don't use them to guess data. Here is how that works in practice.

Large language models have changed how quickly we can get a scraper from a blank page to a working sample, and they are good at reading documents that have no structure at all. They are also expensive per page and happy to invent a value when a field is missing. Our approach is to use models where they beat hand-written code, verify their output against the page, and fall back to deterministic parsers for anything that runs at volume.

In detail

How it works

01

Schema inference from a sample

Point us at a page type and a model proposes field names, types and selectors from a handful of examples. An engineer reviews and tightens the schema before anything runs at scale, so the output is stable and typed, not whatever the model felt like that day.

02

Extraction from unstructured pages

Job descriptions, court notices, property brochures and PDFs rarely share a layout. For these we pass cleaned text through a model with a strict output schema, validate every record and route failures to a human review queue instead of shipping them.

03

Faster scraper generation

Boilerplate, selector drafts and test fixtures are generated, then checked and rewritten by the engineer who owns the job. The time saved goes into the hard part: the anti-bot layer and the private API calls that no model can see from the HTML.

04

Guardrails on cost and accuracy

Model calls are limited to pages that need them; everything regular is parsed deterministically. We measure extraction accuracy on a labelled sample for each job and report it, and model costs appear as a separate line in the quote.

Schema

Fields we typically deliver

Starting point only. You name the columns; we keep the names stable across every run.

  • source_url
  • extracted_json
  • schema_version
  • confidence
  • validated
  • reviewer
  • scraped_at

Questions

About ai-assisted scraping

Does using AI make the data less reliable?

Not when it is used for the right step. Model output is validated against a fixed schema and spot-checked against pages, and anything regular is still parsed by code. We report measured accuracy per job rather than asking you to trust the model.

Which models do you use?

Whatever fits the page and the budget: hosted models from the major providers for difficult documents, smaller open models run on our side for high volume. Your data is not used to train anything.

Can AI get past anti-bot protections?

No. Protections are handled with fingerprinting, session handling and reverse engineering, which is engineering work. Models help afterwards, with reading what the page returns.

Is it more expensive than a normal scraper?

It can be cheaper to build and more expensive to run, because model calls cost per page. The quote shows both numbers so you can choose the mix that fits your volume.

Tell us the site. We'll tell you what it takes.

Send a URL and the fields you need. An engineer looks at it the same day and replies with a plan and a price.