Topic and keyword feeds
A continuous stream of articles matching your keywords, companies or people from a list of outlets, deduplicated across syndication, delivered to your storage or by webhook shortly after each article is published.
Industries
Articles, press releases, broadcast transcripts and public social posts from the outlets you name, with full text, author, date and section.
News scraping is less about any one site and more about keeping hundreds of them running at once. Outlets change templates often, paywalls vary by region and the same wire story appears under fifty headlines. We maintain parsers per outlet, pull articles soon after publication through RSS, sitemaps and direct crawls, extract clean body text and metadata and group duplicates. You receive a continuous feed of articles, or a one-off archive for a date range and topic.
Sources
Public pages and documents we have scraped in this vertical. Named sites are examples, not an exhaustive list.
Schema
A common starting schema. You decide the final columns and names; we keep them stable across runs.
article_idoutlettitleauthorpublished_atsectionurllanguageword_countbody_textscraped_atOutput
article_id | outlet | title | author | published_at | section | url | language | word_count | body_text | scraped_at |
|---|---|---|---|---|---|---|---|---|---|---|
| NW-20261007-00042 | Example Times | Example headline about a market change | J. Example | 2026-10-07T06:40:00Z | Business | https://example.com/story/42 | en | 812 | (full article text) | 2026-10-07 |
Values are illustrative placeholders to show shape and types, not records from any client or source.
Use cases
A continuous stream of articles matching your keywords, companies or people from a list of outlets, deduplicated across syndication, delivered to your storage or by webhook shortly after each article is published.
Every article from a set of outlets for a date range, with clean text and metadata, for model training, linguistic research or coverage studies. Delivered as Parquet or JSON Lines partitioned by publication date.
Scrape company and regulator newsrooms directly so announcements are captured the moment they go live, before they are rewritten by the press, with the original wording and publication time kept.
Questions
We collect what is publicly available: headlines, metadata and the text that appears without a subscription. Full paywalled text is only collected where you hold a licence that permits it.
Outlets with RSS or sitemaps are polled every few minutes; others are crawled on a schedule that matches how often they publish. The delay from publication to delivery is usually minutes, not hours.
Yes. Text extraction is language-agnostic and we tag each article with its detected language. Translation can be added as a step if you need one working language.
Public posts on open sites such as Reddit, YouTube comments and public forums are in scope. Closed networks and anything requiring a login are not.
Send the sources and the fields you need. An engineer looks at them the same day and replies with a plan and a price.
Hi! Send me the site you need data from and I'll get an engineer to look at it.
Chat on WhatsApp Or get a quote