Delivery

Where your data lands: S3, GCS, BigQuery, Sheets or a webhook

Scraped data is only useful once it is in your systems. Every job ends with delivery to a destination you already use, in a layout you can load without cleanup.

We do not ask you to log into anything of ours to collect your data. When a run finishes, the files are written where you tell us: an S3 or GCS bucket you own, a BigQuery or Postgres table, a Google Sheet for small jobs, or a webhook that pushes each record into your pipeline. Layout, naming and partitioning follow the conventions those tools expect, so the data is queryable the moment it arrives, and historical runs accumulate without anyone moving files around.

In detail

How it works

01

Object storage: S3 and GCS

Files are written to your bucket under a date-partitioned path such as source=site/run_date=2026-10-07/, in Parquet or JSON Lines with a manifest. This is the Hive-style layout that Athena, BigQuery and Spark read directly, so no renaming or reshuffling is needed.

02

Warehouses and databases

For BigQuery, Snowflake, Postgres or MySQL we load directly into a table you nominate, with a stable schema and a run_id column on every row. Loads are idempotent, so a rerun replaces the same partition instead of duplicating rows.

03

Sheets, webhooks and downloads

Small or exploratory jobs go into a Google Sheet that refreshes on every run. Event-style data is pushed record by record to your webhook with retries and signed payloads. Everything can also be fetched from a signed download link that expires.

04

Ownership and access

Credentials you give us are scoped to one bucket or table and revoked when the job ends. Data never sits on our side longer than the run and its checks need, and you can ask for a retention statement in the contract.

Schema

Fields we typically deliver

Starting point only. You name the columns; we keep the names stable across every run.

  • run_id
  • run_date
  • source
  • record_count
  • file_path
  • format
  • checksum
  • delivered_at

Questions

About delivery: where your data lands

Which destinations do you support?

Amazon S3, Google Cloud Storage, Azure Blob, BigQuery, Snowflake, Postgres, MySQL, Google Sheets, webhooks and plain download links. If you use something else with an API or a bulk loader, we can usually add it within the project.

How are files named and partitioned?

By source and run date, following the Hive-style convention that most query engines detect automatically. Each run includes a manifest with row counts and checksums so you can verify a load.

Do you keep a copy of our data?

Only for as long as the run and its quality checks need, then it is deleted from our side. Your bucket or table is the system of record, and we can put that commitment in writing.

Can delivery be set up on an existing scraper?

Yes. If you already have scrapers producing files, we can add the storage layout, loads and webhooks as a small project or as part of an engineer month.

Tell us the site. We'll tell you what it takes.

Send a URL and the fields you need. An engineer looks at it the same day and replies with a plan and a price.