← Personal projects Data engineering Data engineering project · 2025

Attorney data extraction at scale

Every law firm publishes attorney bios differently, and none of them publish them cleanly. I built a Selenium scraper suite covering forty-plus U.S. firms, normalizing all of them into a single twenty-four-field schema, feeding an AWS ETL job that lands clean records on a schedule.

  • Python
  • Selenium
  • AWS
  • pandas
  • openpyxl
40+
Firm scrapers shipped
24
Fields normalized
Manual → scheduled
Collection cadence

Architecture

  1. 01Firm siteSelenium, headless Chrome
  2. 02Profile parserper-firm selectors
  3. 03Normalizer24-field schema
  4. 04Validatorsstates, honors, years
  5. 05Excel + S3openpyxl, AWS ETL

The problem

A sales team needs structured records for practicing attorneys in the United States: where they work, what they practice, where they went to school, which state bars they’re admitted to. That data is public. It’s on every firm’s website. It is also, from a data engineering standpoint, a worst case.

There is no standard. One firm renders bios as server-side HTML with clean semantic markup; the next builds them client-side from a JSON blob; the next hides half the biography behind a “Read More” toggle that only fires on click. Education appears as prose in one place and as a list in another. Some firms put “J.D., cum laude” in a single string; some split degree, school, honors, and year across four sibling elements; some use an <em> tag for honors and nothing else.

And the target wasn’t “get the data” — it was get the data in exactly one shape, because everything downstream assumed that shape.

The constraint that shaped everything

The output schema was fixed and non-negotiable: twenty-four fields, in a specified order, covering identity, firm and role, contact, biography, practice areas, two separate education records, bar admissions, languages, and source URLs. Every scraper, regardless of what the source site looked like, had to emit exactly that.

That constraint is the whole story of this project. It meant the per-firm work could only ever be extraction — the moment a scraper started making its own decisions about formatting, the dataset fractured.

So the architecture split hard along that line:

  • Per-firm extractors knew about selectors, pagination, and the specific ways one site was weird. Nothing else.
  • A shared normalization layer owned every formatting decision, and every scraper called it.

Normalization rules worth naming

Most of the real engineering was here, and most of it came from watching the data fail.

Education had to be split before it was parsed. Law school and undergraduate records are separate fields, and firms interleave them freely. Parsing had to identify which entry was which before touching either — and the school-name fields had to contain only the school name. Not “B.S., George Mason University,” not “George Mason University (2014).” Just George Mason University. Degrees, majors, years, honors, and punctuation all stripped out into their own fields or dropped.

Honors were matched, never inferred. Firms write honors in free text, and free text produces garbage categories. Honors were extracted from a specific nested element and then matched case-insensitively against a fixed validated list — cum laude, magna cum laude, summa cum laude, Order of the Coif, Law Review, clerkship, and a small set of secondary degrees. Anything that didn’t match the list didn’t become a value; it became UNASSIGNED. A known-unknown is a usable record. An invented category is a corrupted one.

Bar admissions were restricted to U.S. states. Firm bio pages list courts alongside states — district courts, circuit courts, the Supreme Court. Only state-level admissions (plus the District of Columbia) belong in that field, so courts were filtered out, and admission years were stripped from the values rather than left to pollute the string.

Middle names lost their punctuation. Trivial-sounding, and it mattered: J. and J are different keys, and deduplication downstream ran on name fields.

Biographies had to be complete. Where a “Read More” control existed, the scraper expanded it before reading. A truncated bio isn’t a shorter record, it’s a wrong one.

Some profiles were out of scope entirely. Non-U.S. offices — London in particular — were excluded at the extractor level rather than filtered later, so they never entered the pipeline.

Making forty scrapers maintainable

Forty independently-written scrapers is forty independent liabilities. The suite converged on a single template that every new firm started from:

  1. A data dictionary initialized with all twenty-four keys, so every record was structurally complete before extraction began — missing fields came out empty, never absent.
  2. A standard education-parsing block, calling the shared normalizer.
  3. Per-field try/except isolation.

That third one is the reason the suite scaled. A firm that redesigns its bio template breaks one selector, not one scraper. The record still lands, with one empty field and a logged failure, instead of the run dying at attorney fourteen of six hundred.

What it fed

Output was written through openpyxl into the fixed template structure, then picked up by an AWS-based ETL job that loaded it into the sales database. The manual collection process it replaced was a person opening bio pages and copying fields into a spreadsheet.

What I’d do differently

Two things.

Contract tests per firm. The try/except isolation keeps a broken selector from killing a run — but it also means a silently-empty field looks a lot like an attorney who genuinely has no LinkedIn. A small assertion suite per firm (“this page should yield a non-empty law school”) would have turned silent degradation into a loud failure.

Orchestration. Scheduled scraper runs are exactly the workload Airflow exists for — retries, per-firm task isolation, backfills, and a dependency graph that runs normalization after extraction rather than inside it. That’s the first thing I’d add.

What I took from it

The transformation is never the hard part. Ninety percent of the work in this project was upstream of any transformation: deciding what a valid value is, deciding what to do with something that isn’t one, and building the thing so that the answer is written down once instead of forty times.

Contact

Want the longer version?

Happy to walk through any of this in detail — the parts that broke are usually the interesting bit.