Industries

What we scrape in jobs and recruitment

Job postings, company career pages, salary pages and employer reviews, deduplicated into one feed with posting and expiry dates.

Job data is the vertical we have scraped the longest. Postings are duplicated across boards, expire silently, hide salary in free text and sit behind some of the strongest anti-bot setups on the web. We collect from boards, aggregators and employer career sites, parse titles, locations, salary ranges and skills into fields, deduplicate the same job across sources and record when each posting appeared and disappeared. Candidate profiles and anything behind a login are not part of the work.

Sources

Typical sources

Public pages and documents we have scraped in this vertical. Named sites are examples, not an exhaustive list.

  • Indeed, LinkedIn Jobs and Glassdoor job pages
  • Niche boards for tech, healthcare, trades and hospitality
  • Employer career sites on Workday, Greenhouse, Lever and SuccessFactors
  • Government and public-sector job portals
  • Freelance marketplaces such as Upwork for public job posts
  • Salary pages and compensation surveys
  • Employer review pages
  • Company pages for headcount and location signals

Schema

Fields

A common starting schema. You decide the final columns and names; we keep them stable across runs.

  • job_id
  • title
  • company
  • location
  • remote
  • salary_min
  • salary_max
  • currency
  • employment_type
  • posted_at
  • source
  • scraped_at

Output

What a row looks like

Example row — structure only
job_idtitlecompanylocationremotesalary_minsalary_maxcurrencyemployment_typeposted_atsourcescraped_at
JB-00912377Senior backend engineerExample CorpBerlin, DEhybrid7000090000EURfull-time2026-10-01Employer site2026-10-07

Values are illustrative placeholders to show shape and types, not records from any client or source.

Use cases

Jobs we are usually asked for

Job posting feeds

Daily or hourly collection of new postings for a set of titles, locations or companies, deduplicated across boards and employer sites, with first-seen and last-seen dates per job so expiry is measurable.

Salary and skills datasets

Parse salary ranges, seniority and required skills from posting text into structured fields across thousands of postings, so compensation and demand can be studied by role, city and employer over time.

Employer and hiring signals

Track open roles per company over time from career sites and boards, giving a public view of where companies are hiring, which teams are growing and which postings stay open the longest.

Questions

About job market & recruitment scraping

Can you scrape LinkedIn or Glassdoor?

Public job and company pages, within the limits those sites place on access and without logging in. Member profiles, messaging and anything behind a login are out of scope.

How do you deduplicate postings across boards?

On employer, normalised title, location and posting text similarity, with the source identifiers kept on every record. You see one job with a list of where it was found.

Can you extract salary from free text?

Yes. Ranges, currencies and pay periods are parsed into numeric fields with a flag for estimated versus stated, and the original text is kept for review.

Do you offer this as a recurring feed?

Most job clients run daily. Each run lands in your storage with a diff of new, changed and removed postings, priced per month after the first sample.

Need job market & recruitment data? Tell us the sites.

Send the sources and the fields you need. An engineer looks at them the same day and replies with a plan and a price.