I'm a data scientist and ML engineer who builds models that hold up on real data, and the systems that serve them.
Data Scientist / ML Engineer with almost five years across Astrion, PayPal and Siemens Healthcare. Python, SQL, PyTorch, XGBoost, NLP, computer vision and RAG.
Fraud models, medical imaging and Graph RAG, plus the Spark pipelines and FastAPI services around them. The code lives on my GitHub ↗
I take models from messy data to production: features that capture behaviour, evaluation tuned for the errors that matter, and explanations people can act on.
experience.
- 2025 — Now
ML / AI Engineer
Astrion · Huntsville, AL (remote) Mission Operations & Engineering Intelligence Platform
- Built predictive workflows around equipment events and operating conditions with XGBoost, LightGBM and Random Forest, engineering features from historical event and maintenance records and tuning with Optuna and cross-validation.
- Turned engineering notes, incident descriptions and maintenance comments into usable signals with BERT, spaCy and Sentence Transformers.
- Built document retrieval with RAG, FAISS, OpenSearch and pgvector, and connected equipment, events and documents in Neo4j for Graph RAG answers a vector-only search would miss.
- Prototyped GenAI workflows with GPT-4, LangChain and LangGraph, kept answers tied to approved technical content, and traced them with LangSmith.
- Shipped components as FastAPI services in Docker on AWS (S3, SageMaker, EC2, Glue, Redshift), with MLflow tracking, GitHub Actions CI/CD and SHAP/LIME explanations.
PythonPySparkXGBoostLightGBMBERTFAISSOpenSearchpgvectorNeo4jGPT-4LangGraphLangSmithFastAPIAWSMLflow
In detail → - 2023 — 2024
Machine Learning Engineer
PayPal · Hyderabad, India Transaction Risk, Fraud Detection & Customer Intelligence Platform
- Designed behavioural features (transaction frequency, amounts, merchant activity, timing, account history) and built fraud and risk classifiers with XGBoost, LightGBM, CatBoost and Logistic Regression.
- Evaluated with AUC-ROC, precision/recall and confusion matrices with a focus on false positives against legitimate customers; tuned with Optuna, grid and random search.
- Moved high-volume transaction preparation to Spark / PySpark and reusable SQL + Python routines, so training sets refreshed without manual rework.
- Added SHAP and LIME explanations for analysts, and built a semantic case search over support text with BERT, Sentence Transformers, FAISS and OpenSearch.
- Served models through FastAPI and Docker on AWS (S3, SageMaker, ECS, Glue, Redshift), with MLflow tracking and Redis + Celery for async jobs.
PythonSQLPySparkXGBoostCatBoostSHAPFAISSOpenSearchFastAPIDockerAWSMLflowRedisCelery
In detail → - 2022 — 2023
Data Scientist / ML Engineer
Siemens Healthcare · Hyderabad, India Medical Imaging & Clinical Document Intelligence Platform
- Prepared medical image datasets with OpenCV and trained CNN classifiers in TensorFlow / Keras; explored object detection with YOLO and Faster R-CNN.
- Extracted text from scanned clinical documents with OCR, Azure Document Intelligence and Azure Cognitive Services.
- Built clinical text classification with spaCy, NLTK and BERT, and semantic document retrieval with Sentence Transformers and FAISS, prototyping RAG grounded in healthcare documentation.
- Handled class imbalance and inconsistent images with augmentation and tuning, and used SHAP and error analysis to see where predictions held up.
- Scaled preparation with PySpark, tracked experiments in MLflow, and served models behind Flask and FastAPI for application teams.
PythonOpenCVTensorFlowKerasPyTorchYOLOFaster R-CNNBERTspaCyFAISSAzurePySparkMLflowFlask
In detail →
Education
M.S. Data Science University of Alabama at Birmingham 2024 — 2026
Bachelor’s, Data Science Vignan Institute of Technology and Science
about.
I'm Chetan, a data scientist and machine learning engineer in Phoenix, with almost five years of building ML, deep learning, NLP, computer vision and generative AI systems across financial services, healthcare and engineering teams.
The thread through all of it is the full modelling cycle: preparing data at scale with Spark, evaluating carefully and explaining what a model does, and getting it out of a notebook into a service people actually use.
Outside of work I keep a running list of things I don't understand yet and slowly shorten it — lately LLM internals, and radio: a LoRa mesh across the West Valley, a satellite pass planner, and an ADS-B receiver on my desk.
I'm looking for data science, ML and AI engineering roles. If that's what you're hiring for, I'd like to hear from you.
status : open to work