Summary
Overview
Work History
Education
Skills
Accomplishments
ACHIEVEMENTS
Languages
Timeline
Hi, I’m

Manideep Gongalla

Seattle,NJ

Summary

Big Data engineer building and tuning Spark, Hive, Hadoop, and Airflow pipelines on AWS EMR to reduce ETL pipeline runtime by 21–35% while processing 10 tb datasets with faster runtime and lower storage cost. Delivers real-time ingestion and analytics with Kinesis, Spark Streaming, Lambda, and Step Functions, while strengthening reliability through monitoring, alerting, and automated retries. Improves data accessibility with Presto, Athena, and Redshift for business and analytics users.

Overview

3
Languages
5
years of professional experience

Work History

Amazon

Big Data engineer
11.2024 - Current

Job overview

  • Achieved efficient processing of multi-terabyte datasets by developing and optimizing ETL pipelines with Hadoop Hive and Spark on AWS EMR. Facilitated scalability and fault tolerance through strategic configuration and fine-tuning of EMR clusters. Enhanced performance and reduced storage costs by implementing effective data partitioning and compression techniques.
  • Assisted in building real-time analytics frameworks utilizing Spark Streaming, AWS Kinesis, and Lambda for ongoing data integration and immediate insights. Supported automation of job scheduling and monitoring to maintain consistent data availability and minimize manual tasks.
  • Designed ETL workflows with Airflow to ensure seamless integration and continuous data availability from multiple sources. Enabled continuous data ingestion from real-time data sources and delivered actionable insights with minimal latency. Used Kinesis Analytics to query and transform streaming data, and integrated results with downstream data pipelines for further processing.
  • Optimized SQL performance across distributed datasets using Presto-based query engines, enhancing data accessibility for business users and analytics teams through interactive querying on large-scale datasets in S3.
  • Integrated Amazon Bedrock with MCP framework to enable LLM-based data insights and metadata summarization.
  • Leveraged Amazon Athena for ad-hoc querying and analytics on data stored in S3, reducing query response time. Enabled business users to explore datasets without the need for complex ETL processes or additional infrastructure. Automated error handling and retries within workflows to ensure high availability and fault tolerance.
  • Developed complex workflows using AWS Step Functions to orchestrate multi-step data processing pipelines.
  • Tuned Glue / EMR Spark jobs (dynamic partitions, compaction to 256 MB, optimized shuffles) to reduce runtime by 35%.
  • Designed config-driven MCP architecture for dynamic job monitoring, anomaly detection, and feedback loops.
  • Implemented Apache Iceberg tables on S3 to facilitate ACID transactions and schema evolution, enhancing query performance and reliability for large-scale analytical workloads.
  • Orchestrated end-to-end jobs using Airflow + Step Functions, monitored via CloudWatch EMF metrics and alerting dashboards.
  • Deployed Apache Airflow on AWS to manage and orchestrate data workflows, integrating it with other AWS services. Integrated Airflow with AWS services like S3, Glue, and Redshift to automate the execution of ETL jobs, improving pipeline reliability and scalability. Utilized Airflow’s DAG capabilities to ensure smooth handling of task dependencies and scheduling.
  • Designed and optimized data warehousing solutions using Amazon Redshift, enabling efficient data storage, retrieval, and querying.
  • Optimized ETL workflows to enhance data retrieval efficiency and accuracy.
  • Mentored junior engineers in best practices for big data technologies and methodologies.
  • Developed monitoring tools to ensure data quality and system reliability across platforms.

Freddie Mac

Data engineer
03.2023 - 11.2024

Job overview

  • Implemented a composable data platform architecture, leveraging tools like Apache Airflow and AWS Redshift, to ensure modular scalability and observability within the data infrastructure.
  • Designed and implemented complex data pipelines using Apache Airflow to orchestrate ETL processes, improving workflow efficiency by 30%.
  • Implemented AWS Lambda for serverless computing to increase microservices scalability and responsiveness while engaging in ETL processes and application development using Python and SQL.
  • Designed and implemented complex data pipelines using Apache Airflow to orchestrate ETL processes, achieving a 30% increase in workflow efficiency.
  • Implemented AWS Lambda for serverless computing, enhancing microservices scalability and responsiveness. Actively engaged in ETL processes and application development using Python and SQL.
  • Developed AWS Glue jobs using Spark to cleanse, normalize, and enrich raw data from multiple sources for optimized downstream analytics.
  • Translating architectural and design blueprints into fully functional big data systems, leveraging technologies such as Hadoop, Spark, Airflow to handle large-scale data processing tasks efficiently.
  • Implemented Spark using Python and Spark SQL for faster testing and processing of data. Loaded processed data into Redshift for advanced data warehousing and analytics capabilities.
  • Developed polished visualizations to share results of data analyses.

Tek gigzs

Data engineer
04.2022 - 03.2023

Job overview

  • Designed and implemented Flink-based streaming pipelines for real-time data analytics, enabling faster and more accurate insights from massive datasets.
  • Utilized Spark SQL API to extract and load data, performing SQL queries that enhanced data processing and analysis.
  • Monitored, alerted, and troubleshot ETL pipelines using AWS Glue job metrics and CloudWatch logs.
  • Executed data modeling, database design, and data mining activities to integrate and consolidate disparate datasets for comprehensive analysis.
  • Executed data modeling, database design, and data mining activities, contributing to the integration and consolidation of disparate datasets for comprehensive analysis.
  • Utilized AWS Glue job metrics and CloudWatch logs for monitoring, alerting and troubleshooting ETL pipelines.
  • Integrated Snowflake with ETL pipelines using tools like Apache Airflow or Informatica, orchestrating data ingestion, transformation, and loading processes.
  • Experienced in creating custom Airflow operators and sensors to integrate with various data sources, including databases, APIs, and file systems, enhancing pipeline flexibility and reliability.
  • Demonstrated expertise in Python, SQL, and data pipeline orchestration tools like Airflow, leveraging cloud data platforms such as AWS to enhance data processing capabilities.
  • Executed stored procedures in Amazon Redshift, applying corresponding DDL statements.

Amazon India

Quality engineer
03.2021 - 08.2021

Job overview

  • Business Intelligence Engineer Designed and developed data models and ETL processes to support the integration of data from multiple sources into a centralized data warehouse using Snowflake.
  • Developed and maintained BI dashboards in Tableau, delivering actionable insights to stakeholders and facilitating data-driven decision making.
  • Managed and optimized backend database storage systems including MS SQL Server, AWS Redshift, Snowflake, and MongoDB, improving data accessibility and performance for analytics processes.
  • Designed and developed data models and ETL processes to integrate data from various sources into a centralized data warehouse using AWS Glue and Redshift.

Education

state university of New York at buffalo
Buffalo, New York

Master’s from data science
02-2023

University Overview

  • Relevant coursework: Machine Learning, Probability and statistics, Deep Learning, SQL, Statistical Data mining, Numerical mathematics, Linear algebra and Optimizations, Data analysis
  • GPA: 3.8

Chaitanya Bharathi Institute of Technology
Hyderabad, Telangana

Bachelor of Technology from Mechanical Engineering
05-2020

University Overview

  • Relevant coursework: Object Oriented Programming, Database Management systems, C++, Data Structures&Algorithm
  • GPA: 3.8

Skills

  • Python
  • SQL
  • AWS
  • Docker
  • Airflow
  • Data modeling
  • Dimensional modeling
  • Data warehousing
  • Schema design
  • Data pipeline orchestration
  • Star schema design
  • Slowly changing dimensions
  • Spark SQL
  • Snowflake
  • Amazon Redshift
  • Apache Iceberg
  • ETL optimization
  • CI/CD pipelines

Accomplishments

Accomplishments
  • Performance Evaluation of Copper Coated Aluminum .Electrodes in EDM Process”. IIJISET - International Journal of Innovative Science, Engineering &Technology, Vol. 7 Issue 7, July 2020.

ACHIEVEMENTS

ACHIEVEMENTS
  • Performance Evaluation of Copper Coated Aluminum .Electrodes in EDM Process”. IIJISET - International Journal of Innovative Science, Engineering &Technology, Vol. 7 Issue 7, July 2020.

Languages

English
Native or Bilingual
Hindi
Full Professional
Telugu
Native or Bilingual

Timeline

Big Data engineer
Amazon
11.2024 - Current
Data engineer
Freddie Mac
03.2023 - 11.2024
Data engineer
Tek gigzs
04.2022 - 03.2023
Quality engineer
Amazon India
03.2021 - 08.2021
Chaitanya Bharathi Institute of Technology
Bachelor of Technology from Mechanical Engineering
state university of New York at buffalo
Master’s from data science
Manideep Gongalla