Industries

What we scrape in e-commerce

Product catalogues, prices, stock, variants, reviews and sellers from marketplaces and brand stores, in a schema that matches across sites.

E-commerce is where most of our scraping hours go. Catalogue pages, product detail pages, search results, seller profiles and reviews each need their own handling, and nearly every large store sits behind Akamai, Cloudflare or a custom bot check. We extract products with every variant and its own price and stock state, keep images and attributes, and match the same product across stores on GTIN or MPN. Datasets range from one category on one site to full-catalogue crawls of marketplaces.

Sources

Typical sources

Public pages and documents we have scraped in this vertical. Named sites are examples, not an exhaustive list.

  • Amazon, eBay, Walmart and other marketplaces
  • Brand and DTC stores on Shopify, Magento and Salesforce Commerce
  • Grocery and pharmacy chains with store-level pricing
  • Fashion retailers with size and colour variants
  • Electronics and home-improvement chains
  • Price comparison engines and deal sites
  • Marketplace seller storefronts and ratings
  • Review widgets such as Bazaarvoice and Yotpo

Schema

Fields

A common starting schema. You decide the final columns and names; we keep them stable across runs.

  • sku
  • gtin
  • title
  • brand
  • category
  • variant
  • price
  • sale_price
  • in_stock
  • rating
  • review_count
  • scraped_at

Output

What a row looks like

Example row — structure only
skugtintitlebrandcategoryvariantpricesale_pricein_stockratingreview_countscraped_at
EX-10042-BLK0000000000000Example wireless headphonesExamplebrandAudio > HeadphonesBlack$129.00$99.00true4.41,2842026-10-07

Values are illustrative placeholders to show shape and types, not records from any client or source.

Use cases

Jobs we are usually asked for

Competitor catalogue and price pulls

Scrape the categories you compete in from the stores you care about, daily or weekly, with every variant's price and stock state. Snapshots accumulate in your bucket so price history builds up without a separate tool.

Product content collection

Titles, bullet points, specifications, images and category paths for a brand's products across retailers, so you can check how they are presented and where listings are incomplete, outdated or simply wrong.

Review and rating extraction

All reviews for a product set with rating, date, verified-purchase flag and text, deduplicated across syndication networks. Useful for product teams and for brands checking how retailers present them to shoppers.

Questions

About e-commerce scraping

Can you scrape Amazon?

Yes, for public product, search and review pages, within the volume limits that keep the crawl polite. We do not log in or scrape buyer accounts, and we explain the cadence that is realistic for your page count.

How do you match products across stores?

On GTIN, EAN, UPC or MPN when stores publish them, otherwise on brand plus normalised title and attributes with a confidence score. You see both the match key and the original identifiers from each store.

Can you capture regional or logged-in prices?

Regional prices, yes, by scraping from the right country and postcode. Member-only or logged-in prices are only collected where you hold the account and the terms allow it.

How big a catalogue can you handle?

Millions of product pages are routine when spread over the right schedule and proxy pool. The scoping reply sets a cadence that gets you the data without hurting the site.

Need e-commerce data? Tell us the sites.

Send the sources and the fields you need. An engineer looks at them the same day and replies with a plan and a price.