- Designed, built, and deployed scalable data pipelines to support regulatory and enterprise analytics workloads, translating federal business requirements into production-ready data engineering solutions.
- Implemented end-to-end ETL and ELT workflows using Ab Initio, and PySpark, ensuring high performance, data integrity, and reliable batch processing at scale.
- Developed and maintained logical and physical data models for CCAR datasets, performing metadata scans, PG classification, and CCMS registration in alignment with enterprise data governance standards.
- Engineered data governance controls by validating data dictionary entries, enforcing classification policies, and reducing redundant data across databases, data lakes, and staging environments.
- Modernized legacy reporting pipelines by converting PL/SQL-based processes to distributed Spark jobs and building Tableau dashboards to automate CCAR reporting and analytics delivery.
- Built cloud-native ingestion pipelines using REST/PUT APIs with publisher–subscriber patterns, Control-M scheduling, and automated dynamic data loads for persistent datasets using Spark and Java.
- Designed and executed AWS-based data lake ingestion and extraction POCs, moving data from S3 into Exadata using Terraform-driven infrastructure provisioning and IaC best practices.
- Developed event-driven serverless workflows using AWS Lambda (Python), integrating with S3, Glue Crawlers, and Athena to enable scalable, on-demand ETL processing.
- Delivered high-performance analytics and reporting solutions on AWS using Apache Spark, focusing on scalability, reliability, and cost efficiency for large datasets.
- Drove cloud optimization and engineering best practices by performing cost-benefit analysis (CBA), recommending efficiency improvements across AWS, VSI, and GKP, maintaining ≥75% code coverage through automated APPFIT scans.
- Created Snowflake tables aligned with enterprise data governance, security classifications, and data modeling standards.
- Implemented Snowflake table structures to support structured and semi-structured data ingestion from cloud data lakes.
- Collaborated with cross-functional teams to implement solutions aligned with business objectives.
Skill Set: Oracle, PL/SQL, Python, Java, PySpark, AWS Lambda, S3, EC2, EKS, ELB, EMR, Glue, Athena, Airflow, Terraform, Snowflake, Control-M, Erwin, API, Tableau, GitHub, Maven, Jules, VS Code, Jupyter Notebook, IntelliJ, Kafka, Grafana, Ab initio, Cassandra – CQL, GenAI / LLM