Big Data engineer building and tuning Spark, Hive, Hadoop, and Airflow pipelines on AWS EMR to reduce ETL pipeline runtime by 21–35% while processing 10 tb datasets with faster runtime and lower storage cost. Delivers real-time ingestion and analytics with Kinesis, Spark Streaming, Lambda, and Step Functions, while strengthening reliability through monitoring, alerting, and automated retries. Improves data accessibility with Presto, Athena, and Redshift for business and analytics users.