Summary
Overview
Work History
Education
Skills
Timeline
Generic

P B

Dallas,TX

Summary

Senior engineering professional with deep expertise in data architecture, pipeline development, and big data technologies. Proven track record in optimizing data workflows, enhancing system efficiency, and driving business intelligence initiatives. Strong collaborator, adaptable to evolving project demands, with focus on delivering impactful results through teamwork and innovation. Skilled in SQL, Python, Spark, and cloud platforms, with strategic approach to data management and problem-solving.

Overview

8
8
years of professional experience

Work History

Senior Data Engineer

Client 1
06.2022 - Current
  • Built scalable ETL pipelines using Glue to ingest, transform, and load large-scale clinical and operational datasets into S3 and Redshift.
  • Designed and implemented an automated record retention framework in AWS for a global pharma client, applying country-specific data aging parameters to meet GDPR and HIPAA compliance across clinical and operational datasets.
  • Configured and managed Glue Crawlers for automatic schema discovery and cataloging of structured and semi-structured pharma data.
  • Orchestrated ETL workflows using Airflow DAGs, scheduling tasks across Glue, Lambda, API Gateway, and Redshift to automate pharma data pipelines.
  • Tuned Redshift queries and implemented SQL optimizations, partitioning, and indexing strategies to accelerate reporting on clinical and operational metrics.
  • Created views, stored procedures, and temporary tables in Redshift to support pharma dashboards, analytics, and regulatory reporting.
  • Designed and implemented automated data validation, unit, and integration tests using Python and SQL, ensuring high-quality ETL outputs for sensitive healthcare data.
  • Built interactive dashboards and visualizations using QuickSight to monitor clinical trial progress, patient data, and operational KPIs.
  • Developed and maintained CI/CD pipelines for Glue jobs, Lambda functions, and Airflow workflows, enabling automated deployment and rollback in pharma environments.
  • Implemented infrastructure-as-code using Terraform, integrated with CI/CD pipelines for reproducible and compliant data environments.
  • Configured CloudWatch monitoring, alerts, and logging to track ETL performance, failures, and system health, with proactive notifications for mission-critical pharma data.
  • Collaborated in Agile/Scrum environment, following best practices in modular PySpark coding, CI/CD automation, and reusable pipeline design for regulatory-compliant pharma data workflows.

Data Engineer

Client 2
03.2021 - 08.2021
  • Built ETL pipelines in Glue using PySpark to transform large on-premises datasets and store them in S3 and Redshift.
  • Utilized Glue Crawlers for automatic schema discovery and integrated Lambda for automated ETL processes, reducing.
  • Designed and orchestrated Airflow DAGs to schedule and monitor data pipelines across Glue, Lambda, and Redshift.
  • Monitored data pipelines with CloudWatch, creating alarms and custom metrics for Glue job failures, Airflow task retries, and Lambda errors.
  • Implemented serverless processing using Lambda for transformations, S3 event triggers, and schema validation.
  • Optimized SQL queries in Redshift and Athena by applying partitioning, compression, and query tuning techniques, reducing execution time.
  • Configured S3 lifecycle policies and data partitioning strategies, reducing storage cost and improving downstream query performance.
  • Developed Unit tests in Python to validate ETL transformations and SQL logic, improving pipeline reliability and reducing production defects.
  • Contributed to CI/CD pipelines using CodePipeline, CodeBuild, and GitHub Actions for automated deployment of Glue scripts, Lambda code, and Airflow DAGs.
  • Collaborated with data analysts to design optimized Redshift schemas and delivered self-service datasets through Athena and Spectrum.
  • Participated in Agile ceremonies (daily standups, sprint planning, retrospectives), ensuring timely delivery of user stories and backlog items.

Big Data Engineer

Client 3
05.2017 - 03.2021
  • Designed and developed ETL pipelines using PySpark to ingest, transform, and load data from MySQL, CSV, and JSON into HDFS for downstream analytics.
  • Implemented batch and incremental processing using Spark DataFrames and RDDs, reducing processing time by ~30% compared to MapReduce.
  • Built automated ingestion pipelines with Apache NiFi and Sqoop to extract data from on-premises databases into the Hadoop ecosystem (HDFS/Hive).
  • Created and optimized Hive tables and partitions for large datasets, improving query performance and reducing scan times.
  • Developed real-time streaming pipelines using Apache Kafka and Spark Streaming to capture IoT and log data for operational dashboards.
  • Monitored and tuned Spark jobs for memory management and execution efficiency on Amazon EMR clusters, reducing failures and improving throughput.
  • Integrated cloud storage solutions like AWS S3 for scalable storage and applied cost-effective retention policies.
  • Collaborated with analysts and scientists to deliver datasets for predictive modeling and BI tools such as Tableau and Power BI.
  • Implemented data lineage and logging mechanisms in ETL workflows to track transformations, monitor failures, and support compliance.
  • Documented ETL processes, data models, and pipeline architecture for team knowledge sharing and faster onboarding.
  • Contributed to CI/CD automation of Big Data jobs using Jenkins and Apache Airflow, improving pipeline reliability.

Education

Master of Science - BioMedical Engineering & Data Science

Case Western Reserve University
Cleveland, Ohio

Bachelor of Technology (B. Tech) - Information Technology

JNTUH

Skills

  • Python programming
  • ETL development
  • Git version control
  • Big data processing
  • Kafka streaming
  • NoSQL databases
  • Data pipeline design
  • Data modeling
  • API development
  • Performance tuning
  • Hadoop ecosystem

Timeline

Senior Data Engineer

Client 1
06.2022 - Current

Data Engineer

Client 2
03.2021 - 08.2021

Big Data Engineer

Client 3
05.2017 - 03.2021

Bachelor of Technology (B. Tech) - Information Technology

JNTUH

Master of Science - BioMedical Engineering & Data Science

Case Western Reserve University
P B