Summary
Overview
Work History
Education
Skills
Certification
Interests
Timeline
Generic

Sanjay Bomma

Summary

AWS Python and Databricks engineer with 6+ years of experience in data platforms, delivering complex data solutions within Finance, Healthcare, and Retail. Engineered robust data pipelines and ETL workflows using AWS services like Lambda, S3, and Databricks for financial data governance and PCI compliance. Integrated Python and SQL to build data models, ensuring Data Quality and security for real-time transactional systems in the retail domain. Managed GitHub-based CI/CD pipelines, automating deployment processes and improving Test Driven Development practices with Databricks notebooks. Used Terraform to provision and manage cloud infrastructure on AWS, ensuring consistent, version-controlled environments for data platforms. Collaborated with cross-functional teams to define data requirements and enforce Data Governance policies across multiple financial reporting systems. Monitored and maintained complex data ecosystems, applying Python scripts to resolve Data Quality issues and ensure data integrity within Healthcare projects. Designed and executed data security protocols using AWS services and Python to protect sensitive customer information, adhering to HIPAA and GDPR regulations. Developed and maintained data catalogs using Unity Catalog to improve data discoverability and enforce Data Governance for diverse business units. Contributed to data strategy and roadmap planning, focusing on modernizing legacy data platforms using cloud-native services like AWS S3 and DynamoDB.

Diligent [Desired Position] with robust background in data engineering and proven ability to design and implement complex data pipelines. Successfully contributed to optimizing data architecture and enhancing data processing efficiencies. Demonstrated expertise in big data technologies and proficiency in Python and SQL.

Experienced with building and maintaining data pipelines to ensure seamless data flow. Utilizes advanced knowledge of big data technologies to drive data-driven decision-making. Track record of enhancing data architecture for improved performance and reliability.

Overview

5
5
years of professional experience
1
1
Certification

Work History

Senior Data Engineer

Northern Trust
07.2024 - 07.2025
  • Architected an AWS data pipeline for ingesting trade data, using Python and Spark on Databricks to manage large-scale financial datasets while enforcing PCI compliance.
  • Engineered the CI/CD pipeline with GitHub Actions, automating the deployment of Python functions and Databricks notebooks, which was a little nerve-wracking at first but I figured it out.
  • Used Terraform to provision and manage AWS resources, including S3 and Lambda, ensuring consistent infrastructure setup for various financial analytics environments.
  • Developed custom Python scripts to perform Data Quality checks and validation on incoming transaction data, addressing data integrity issues before they impacted reporting.
  • Collaborated with the Data Governance team to implement Unity Catalog for data lineage tracking and access control, ensuring regulatory compliance.
  • Integrated AWS Step Functions with Lambda to orchestrate complex ETL workflows, solving issues with dependency management between data processing steps.
  • Built a data security layer using AWS IAM and S3 bucket policies, addressing the critical problem of securing sensitive client financial data.
  • My team and I spent a few frustrating hours debugging a failed CI/CD run, but we eventually found a small error in the Terraform state file.
  • Implemented a new data ingestion pattern using AWS Event Bridge to trigger Lambda functions, which streamlined the data loading process for market data feeds.
  • Participated in daily stand-ups and sprint planning sessions, and it was great seeing how everyone's small contributions added up to a big project.
  • Worked on a particularly tricky bug where a Python script was miscalculating a financial metric, requiring a deep dive into the code and a lot of trial-and-error to fix.
  • Environment: AWS, Python, SQL, Databricks, GitHub, Terraform, CI/CD, S3, Lambda, Event Bridge, Step Functions, DynamoDB, Postgres, Unity Catalog, SonarQube, PCI

Software Developer

Mayo Clinic
12.2021 - 07.2023
  • Worked with Python to develop data ingestion pipelines for patient records, handling sensitive data and ensuring strict HIPAA compliance.
  • Collaborated with data scientists to prepare and clean healthcare data using SQL, ensuring Data Quality for their predictive modeling.
  • Deployed and managed data workflows on AWS, utilizing S3 to store large volumes of de-identified patient data in a secure, organized manner.
  • Authored and executed Test Driven Development scripts for new features, ensuring the integrity and accuracy of the data processing logic for clinical trials.
  • Integrated Databricks notebooks into the data platform, streamlining the process of data analysis and reporting for medical research teams.
  • Used Python to implement Data Governance rules on the data platform, ensuring proper classification and handling of protected health information (PHI).
  • Helped automate data migration processes from legacy systems to AWS S3, a project that was initially a bit intimidating due to the sheer volume of data.
  • Assisted in configuring AWS DynamoDB to create a fast, serverless database for storing metadata and application logs for the healthcare application.
  • Partnered with the security team to apply Data Security best practices, including encryption at rest and in transit for all patient data, a crucial part of our work.
  • Troubleshooted data pipeline failures, finding and fixing a bug in a SQL query that was causing an intermittent data loss issue.
  • Contributed to code reviews, providing feedback on Python and SQL scripts to maintain high coding standards and prevent future data issues.
  • Built a simple Postgres database on AWS for a small team to track and manage their research datasets, something I was still learning to do effectively.
  • Developed a small tool with Python to check the Guidewire CDA data schema for compatibility, which was a little tedious but totally necessary for data integrity.
  • Environment: AWS, Python, SQL, Databricks, GitHub, S3, Lambda, DynamoDB, Postgres, Guidewire CDA, HIPAA, GDPR

Software Developer

Lowe's
06.2020 - 12.2021
  • Developed ETL jobs using Python scripts to process retail sales data, ensuring timely updates for inventory and demand forecasting systems.
  • Engineered data pipelines on Azure to aggregate customer purchase data from multiple retail channels, which was a lot of data to handle at first.
  • Used SQL to create complex stored procedures and views for business intelligence dashboards, supporting decision-making for marketing campaigns.
  • Utilized Azure Synapse Analytics to build a data warehouse for storing and analyzing transactional data from the e-commerce platform.
  • Automated daily data ingestion from various sources into Azure Blob Storage, reducing manual effort and improving data availability for analysts.
  • Collaborated with the business team to understand data needs, helping to shape the structure of our data models for more effective reporting.
  • Assisted in debugging a data flow issue where a Python script was incorrectly handling certain product IDs, a frustrating but educational experience.
  • Contributed to the development of a Data Quality monitoring system using Azure functions to alert on inconsistencies in sales data.
  • Worked with the security team to implement role-based access control (RBAC) on Azure to protect sensitive customer and sales data.
  • Learned on the job about data synchronization between different systems, which was a key part of our project to unify customer data.
  • Integrated GitHub into our development workflow for version control, a process that was new to me but I quickly got the hang of it.
  • Developed a small web application with a team to visualize supply chain data, and I was responsible for building the API backend with Python.
  • My first big task was writing a SQL script to generate a daily sales report, and I remember feeling a sense of accomplishment when it ran successfully.
  • Environment: Azure, Azure Synapse Analytics, Python, SQL, GitHub, Retail regulations

Education

Master of Arts (MA) - Information Technology and Management

Webster University

Skills

  • Cloud Technologies: AWS, S3, Lambda, Event Bridge, Step Functions
  • Programming Languages: Python, SQL, Postgres
  • Data Platforms: Databricks, DynamoDB, Unity Catalog
  • CI/CD & DevOps: Terraform, GitHub, CI/CD, SonarQube
  • Data Management: Data Governance, Data Security, Data Quality, Test Driven Development
  • Domain & Regulatory: Guidewire CDA, HIPAA, PCI

Certification

  • Competency of SQL by LearnSQL.com
  • Generative AI for everyone by deeplearning.ai

Interests

  • Passionate about balancing physical health with mental and emotional wellness
  • I enjoy sketching and drawing, which helps improve my creativity and attention to detail
  • Cooking
  • Interested in Human psychology and Philosophy

Timeline

Senior Data Engineer

Northern Trust
07.2024 - 07.2025

Software Developer

Mayo Clinic
12.2021 - 07.2023

Software Developer

Lowe's
06.2020 - 12.2021

Master of Arts (MA) - Information Technology and Management

Webster University
Sanjay Bomma