Work & projects

Work

Two parts: the jobs I've held — Astrion, PayPal and Siemens Healthcare — and the projects I've built on my own. Open any of them for the details: the problem first, then what I built, then what actually happened.

Work experience

Work experience Current role

Astrion

ML / AI Engineer

When
2025 — Now
Where
Huntsville, AL (remote)
Worked on
Mission Operations & Engineering Intelligence Platform

Predictive models, RAG and Graph RAG over engineering documents, and GPT-4 workflows for mission and engineering teams.

  • Built predictive workflows around equipment events and operating conditions with XGBoost, LightGBM and Random Forest, engineering features from historical event and maintenance records and tuning with Optuna and cross-validation.
  • Turned engineering notes, incident descriptions and maintenance comments into usable signals with BERT, spaCy and Sentence Transformers.
  • Built document retrieval with RAG, FAISS, OpenSearch and pgvector, and connected equipment, events and documents in Neo4j for Graph RAG answers a vector-only search would miss.

PythonPySparkXGBoostLightGBMBERTFAISSOpenSearchpgvectorNeo4jGPT-4 +5

In detail

Work experience

PayPal

Machine Learning Engineer

When
2023 — 2024
Where
Hyderabad, India
Worked on
Transaction Risk, Fraud Detection & Customer Intelligence Platform

Fraud and transaction-risk models, behavioural feature engineering, and a semantic case-search layer for risk analysts.

  • Designed behavioural features (transaction frequency, amounts, merchant activity, timing, account history) and built fraud and risk classifiers with XGBoost, LightGBM, CatBoost and Logistic Regression.
  • Evaluated with AUC-ROC, precision/recall and confusion matrices with a focus on false positives against legitimate customers; tuned with Optuna, grid and random search.
  • Moved high-volume transaction preparation to Spark / PySpark and reusable SQL + Python routines, so training sets refreshed without manual rework.

PythonSQLPySparkXGBoostCatBoostSHAPFAISSOpenSearchFastAPIDocker +4

In detail

Work experience

Siemens Healthcare

Data Scientist / ML Engineer

When
2022 — 2023
Where
Hyderabad, India
Worked on
Medical Imaging & Clinical Document Intelligence Platform

Medical image classification and detection, OCR on clinical documents, and semantic retrieval over healthcare records.

  • Prepared medical image datasets with OpenCV and trained CNN classifiers in TensorFlow / Keras; explored object detection with YOLO and Faster R-CNN.
  • Extracted text from scanned clinical documents with OCR, Azure Document Intelligence and Azure Cognitive Services.
  • Built clinical text classification with spaCy, NLTK and BERT, and semantic document retrieval with Sentence Transformers and FAISS, prototyping RAG grounded in healthcare documentation.

PythonOpenCVTensorFlowKerasPyTorchYOLOFaster R-CNNBERTspaCyFAISS +4

In detail

Personal projects

Separate from my jobs — projects I designed and built myself, most with the code on GitHub.

01 Data engineering Data engineering project · 2025

Attorney data extraction at scale

Every law firm publishes attorney bios differently, and none of them publish them cleanly. I built a Selenium scraper suite covering forty-plus U.S. firms, normalizing all of them into a single twenty-four-field schema, feeding an AWS ETL job that lands clean records on a schedule.

40+
Firm scrapers shipped
24
Fields normalized

Firm site Profile parser Normalizer Validators Excel + S3

In detail

02 Applied AI Personal project · 2026

An LLM pipeline where the model goes last

A local-first job-search tracker: it captures a posting, extracts the facts, scores a résumé against it and logs every status change. The design rule is deterministic first, model last — a local LLM only fills the blanks that APIs and structured data left empty.

3
Parsers tried before the LLM
6 GB VRAM
Runs in

Capture ATS API JSON-LD Readability Local LLM Event log

In detail

03 Data engineering Personal project · 2026

A job-market pipeline that throws most of its input away

Aggregators return hundreds of postings for a metro and most of them are noise. This pipeline pulls from four job boards and four ATS APIs, then filters, deduplicates and loads to Postgres on a daily schedule — in testing, 313 raw postings became 62 real tech roles.

313 → 62
Raw postings → kept
8
Source systems

Sources Validate + clean Filter Dedupe Postgres API + dashboard

In detail Code ↗

04 Machine learning Personal project · 2026

Finding incidents the metrics can't see

Anomaly detection over a microservice platform's metrics and distributed traces. Metrics alone caught a third of the labelled failures on the first day I evaluated; adding trace latency caught all 38 across three days — and the write-up is honest about what that number does and doesn't prove.

38 / 38
Labelled failures detected
3 / 9
Metrics-only baseline

Raw archives Time series Anomalies Incident windows Hybrid match Dashboard

In detail Code ↗

05 Applied AI Proof of concept · 2026

One assistant, three roles, and a tool layer that says no

A finance assistant where employees, clients and the owner all talk to the same chatbot and each sees only what their role allows. A small model routes each message to one of three agents, and authorization lives in the tool code, never in the prompt.

3
Specialist agents
3
Roles

Chat UI Auth context Router Agent Tools Postgres

In detail

06 Data engineering Graduate work · 2026

A streaming sensor pipeline, from MQTT to dashboard

End-to-end device telemetry: MQTT ingestion, schema validation at the edge, MongoDB as the raw store, and an aggregation layer that keeps queries fast as the collection grows. Built around the failure cases first, because those are the only parts that matter at 3 a.m.

MQTT → Mongo
Ingestion path
Idempotent
Write semantics

Devices Broker Consumer MongoDB Aggregation

In detail

07 Machine learning Graduate research · 2025–26

From DCGANs to Vision Transformers

A run through the modern vision stack, implemented rather than imported: DCGANs for generation, Fast R-CNN detection on PASCAL VOC, UNet segmentation, and ViT classification. Written to understand the architectures rather than to hit a leaderboard.

4
Architectures built
PASCAL VOC
Detection benchmark

DCGAN Fast R-CNN UNet ViT

In detail

Contact

Questions about any of these?

I'm looking for data science, machine learning and AI engineering roles in Phoenix or remote — and I'm always happy to talk about models, retrieval, or why yours is overfitting.