← Personal projects Data engineering Personal project · 2026

A job-market pipeline that throws most of its input away

Aggregators return hundreds of postings for a metro and most of them are noise. This pipeline pulls from four job boards and four ATS APIs, then filters, deduplicates and loads to Postgres on a daily schedule — in testing, 313 raw postings became 62 real tech roles.

  • Python
  • FastAPI
  • PostgreSQL
  • APScheduler
  • pytest
313 → 62
Raw postings → kept
8
Source systems
55 companies
Employer registry
Daily, 06:00
Cadence

Architecture

  1. 01Sources4 job boards, 4 ATS APIs
  2. 02Validate + cleansalary, seniority, remote
  3. 03Filtermetro + tech-only
  4. 04Dedupehashed job_id
  5. 05Postgresjobs, runs, staging
  6. 06API + dashboardFastAPI, Chart.js, Leaflet

The problem

I wanted a straight answer to a simple question: what data and engineering jobs are actually open in the Phoenix metro today? Job boards answer a different question — what postings match these keywords — and the gap between the two is enormous. Search “data engineer, Phoenix” on an aggregator and you get auto-body technicians at a company that happens to have a data team, nursing roles at a health system, and postings whose location resolved to a different state.

So the interesting part of this project was never the scraping. It was deciding what to throw away.

Two kinds of source

The pipeline reads from two families of source, and they fail in opposite ways.

Aggregators — Indeed, LinkedIn, Glassdoor and ZipRecruiter, through python-jobspy. Broad coverage, noisy results, and locations you can’t trust.

ATS APIs — Greenhouse, Lever, SmartRecruiters and Workday all expose public JSON for their job boards. These are exact: structured fields, no guessing. But you have to know which company uses which system, so there’s a curated registry of 55 Phoenix-area tech employers with their ATS and board slug. Adding a company is appending one CSV row; the runner picks it up on the next run.

The ATS adapters give precision for the employers I know about. The aggregators give recall for the ones I don’t. Everything after that is about making the two agree.

The filter is the product

Every posting, from any source, goes through the same three gates before it’s stored.

Is it actually here? Each result is checked against the ten cities of the metro. Broad aggregator searches routinely return rows whose location turned out to be somewhere else, and those are dropped rather than trusted.

Is it actually a tech role? A classifier with a blocklist and an allowlist. Blocklist first — mechanic, cashier, retail, nursing — because the long tail of non-tech roles at large employers is what dominates raw results. Then anything matching a software, data or infrastructure term is kept, along with per-company keywords from the registry.

Have I seen it already? The same job appears on a company’s own board and on three aggregators. Each posting gets a hashed job_id and the table is unique on it.

In testing, one employer’s 190 raw postings collapsed to 13 tech roles and another’s 122 became 48. Across the run, 313 raw postings became 62. That ratio is the point: four fifths of what the sources return is not what was asked for.

Shape of the data

Postgres, schema created on first run:

  • jobs — the master table, unique on the job hash, with derived columns for city, seniority, remote flag, and salary normalized to annual USD.
  • scrape_runs — one row per run: what was fetched, what was kept, what failed.
  • salary_staging and job_skills — staging for normalization and skill extraction.

Raw responses are kept on disk per search, and each cleaned batch is written out before it’s loaded. When a number on the dashboard looks wrong, there’s a file to open.

Running it

APScheduler triggers the run daily at 06:00 Phoenix time. There’s a --skip-db mode that writes the processed batch to disk without touching Postgres, which is how I test changes to the filter without polluting the table. The unit tests cover the metro filter and the ETL helpers and need neither a database nor a network.

FastAPI serves the result to a single-page dashboard — charts for volume and salary, and a map with postings pinned to the metro.

What I’d do differently

The role filter should be measured, not asserted. It’s a keyword classifier and I tuned it by reading output. A few hundred hand-labelled postings would turn “this looks right” into a precision and recall number, and tell me whether a small text classifier would beat the rules.

Secrets were in the wrong place once. The first version had a database password fallback in source. The current one is environment-only with no default. It’s a small thing and exactly the kind of small thing that matters.

Workday is a moving target. Every tenant has its own host and site path, and the adapter needs per-company configuration. That’s the adapter I’d expect to break first.

Contact

Want the longer version?

Happy to walk through any of this in detail — the parts that broke are usually the interesting bit.