Summary
Overview
Work History
Education
Contact
Timeline
Generic

Hamza Al - U.S Citizen

Staten Island,NY

Summary

Highly skilled Data Engineer with extensive experience in designing, implementing, and optimizing data solutions using cutting-edge technologies. Proficient in Databricks on Azure, Medallion architecture, Delta Lake, Azure services, SQL, Python, and data migration/integration. Successfully led the migration of Hadoop-based data infrastructure to Azure Cloud, ensuring seamless compatibility, data quality, and performance. Proven track record in data modeling, ETL development, data governance, metadata management, and data lineage tracking. Strong analytical and problem-solving abilities coupled with excellent communication and collaboration skills. Passionate about leveraging innovative data solutions to drive business insights and enhance decision-making processes.

Overview

7
7
years of professional experience

Work History

Data Engineer

Delta Air Lines
09.2020 - Current
  • Designed and implemented scalable data pipelines using Databricks on Azure, optimizing performance and reliability.
  • Leveraged Medallion architecture for efficient data processing, storage, and retrieval, ensuring data quality and consistency.
  • Managed data lakes using Delta Lake, enabling version control, schema enforcement, and ACID transactions for large-scale data operations.
  • Utilized Azure services such as Azure Data Factory and Azure Databricks for end-to-end data integration and orchestration.
  • Developed and optimized complex SQL queries and scripts for data transformation, aggregation, and analysis.
  • Implemented Python scripts and libraries for data manipulation, validation, and cleansing, enhancing data accuracy and usability.
  • Successfully completed migration projects, ensuring seamless transfer of data between on-premises systems and cloud.
  • Developed and optimized large-scale data processing workflows using Apache Spark and PySpark on Azure Databricks, leveraging Scala and Python scripts for efficient data transformations.
  • Implemented data pipelines in Azure Data Factory to integrate data from various sources into Snowflake, utilizing Scala and Python for ETL processes and Pandas for data manipulation, with DBT for data modeling.

DATA ENGINEER

Fannie Mae
09.2019 - 08.2020
  • Designed and implemented real-time and batch data processing pipelines using Databricks on Azure, incorporating best practices for performance and scalability.
  • Utilized Databricks notebooks for ETL (Extract, Transform, Load) processes, data exploration, and model development, improving development efficiency and collaboration.
  • Engineered data architectures based on the Medallion architecture principles, emphasizing modularity, reusability, and scalability.
  • Implemented data governance frameworks within Medallion architecture, ensuring data lineage, metadata management, and compliance with data policies.
  • Led the adoption of Delta Lake for managing large-scale data lakes, enabling ACID transactions, schema evolution, and time travel capabilities for data versioning and auditing.
  • Optimized Delta Lake performance through partitioning strategies, indexing, and caching mechanisms, enhancing query performance and data accessibility.
  • Designed and deployed data solutions on Azure Cloud, leveraging services like Azure Data Factory, Azure Databricks, Azure SQL Database, and Azure Storage for end-to-end data processing and analytics.
  • Implemented Azure security features such as Azure Key Vault, Azure Active Directory, and network security groups to ensure data protection and regulatory compliance.
  • Developed complex SQL queries, stored procedures, and views for data transformation, aggregation, and reporting, optimizing query performance and data retrieval.
  • Utilized Python libraries such as Pandas, NumPy, and PySpark for data manipulation and automation of data workflows.
  • Completed the migration effort from Hadoop-based data infrastructure to Azure Cloud, including assessment, planning, execution, and post-migration validation.
  • Implemented data transfer tools like Azure Data Box, Azure Data Factory, and Azure Storage Explorer for efficient and secure data migration to Azure.
  • Designed and implemented data integration solutions for heterogeneous data sources, including relational databases, NoSQL databases, flat files, and APIs, ensuring data consistency and integrity.
  • Developed data validation scripts and automated testing frameworks to validate data accuracy, completeness, and conformity to business rules and standards.
  • Conducted performance tuning and optimization of data pipelines, SQL queries, and Spark jobs, analyzing execution plans, resource utilization, and bottlenecks to improve overall system efficiency.
  • Designed and deployed data solutions on Azure, including Azure Data Lake and Azure Synapse Analytics, to support advanced analytics and reporting with Spark, Databricks, and Terraform for infrastructure as code.
  • Utilized Azure Functions with Python scripts and Pandas for serverless data processing and automation, enabling real-time data ingestion and transformation.
  • Built and managed data lakes on Azure Data Lake Storage, integrating with Databricks and Snowflake for seamless data processing and analytics using Scala, PySpark, and Python.

DATA ENGINEER

Toyota.
New York, NY
01.2018 - 08.2019
  • Planned and executed the migration of data infrastructure from Hadoop to Azure Cloud, including data assessment, mapping, and transformation processes to ensure compatibility and data integrity.
  • Implemented Azure Data Factory pipelines to automate data ingestion, transformation, and loading (ETL) processes, enhancing data processing efficiency and reliability.
  • Leveraged Azure Blob Storage and Azure Data Lake Storage Gen2 for scalable and cost-effective storage solutions, optimizing data access and retrieval.
  • Designed and implemented data models and schemas on Azure SQL Database, ensuring data consistency, integrity, and performance for analytical workloads.
  • Utilized Azure Databricks for data exploration, analysis, and model development, integrating with Azure Machine Learning for advanced analytics and predictive modeling.
  • Conducted performance benchmarking and optimization of data processing workflows on Azure Databricks, improving job execution times and resource utilization.
  • Implemented security measures such as Azure Key Vault integration, encryption at rest and in transit, and role-based access control (RBAC) to protect sensitive data assets.
  • Developed custom Python scripts and U-SQL jobs for data cleansing, enrichment, and validation, ensuring data quality and accuracy throughout the migration process.
  • Collaborated with stakeholders to define data migration strategies, prioritize data sets for migration, and establish migration timelines and milestones.
  • Implemented monitoring and alerting solutions using Azure Monitor and Azure Log Analytics to track data migration progress, detect anomalies, and troubleshoot issues in real time.
  • Conducted post-migration validation and testing to verify data completeness, consistency, and correctness, ensuring a seamless transition to the Azure Cloud environment.
  • Documented migration processes, configurations, and data lineage on Azure Data Catalog and Azure Data Governance tools for transparency, compliance, and auditability.
  • Planned and executed the migration of data infrastructure from Hadoop to Azure Cloud, including data assessment, mapping, and transformation processes to ensure compatibility and data integrity.
  • Implemented Azure Data Factory pipelines to automate data ingestion, transformation, and loading (ETL) processes, enhancing data processing efficiency and reliability.
  • Leveraged Azure Blob Storage and Azure Data Lake Storage Gen2 for scalable and cost-effective storage solutions, optimizing data access and retrieval.
  • Designed and implemented data models and schemas on Azure SQL Database, ensuring data consistency, integrity, and performance for analytical workloads.
  • Utilized Azure Databricks for data exploration, analysis, and model development, integrating with Azure Machine Learning for advanced analytics and predictive modeling.
  • Conducted performance benchmarking and optimization of data processing workflows on Azure Databricks, improving job execution times and resource utilization.
  • Implemented security measures such as Azure Key Vault integration, encryption at rest and in transit, and role-based access control (RBAC) to protect sensitive data assets.
  • Developed custom Python scripts and U-SQL jobs for data cleansing, enrichment, and validation, ensuring data quality and accuracy throughout the migration process.
  • Designed and implemented metadata architecture for data cataloging, data lineage tracking, and metadata management using Azure Data Catalog and Azure Purview.
  • Collaborated with stakeholders to define data migration strategies, prioritize data sets for migration, and establish migration timelines and milestones.
  • Configured and optimized Spark clusters on Azure Databricks using Terraform, ensuring efficient resource allocation and high performance for big data applications.
  • Developed custom UDFs (User Defined Functions) in Scala for complex data transformations in Spark, enhancing data processing capabilities within Azure environments.
  • Implemented real-time data streaming solutions using Azure Event Hubs and Spark Streaming on Databricks, integrating with Python scripts and PySpark for data processing and analysis.

Education

Bachelor of Science - Engineering

College of Staten Island
Staten Island, NY

Contact

Staten Island, NY 10304

Timeline

Data Engineer

Delta Air Lines
09.2020 - Current

DATA ENGINEER

Fannie Mae
09.2019 - 08.2020

DATA ENGINEER

Toyota.
01.2018 - 08.2019

Bachelor of Science - Engineering

College of Staten Island
Hamza Al - U.S Citizen