Project Name: US based Diagnostics Company
Data Migration Project: Architected and implemented end-to-end cloud data migration solutions leveraging PySpark, Spark Scala, Databricks, and Snowflake, enabling the ingestion, transformation, and loading of structured and semi-structured data from AWS S3 into optimized Snowflake tables, Databricks based on business requirements.
- Snowflake - Snowpark Project: To obtain Cloud based Data warehousing Tool and high scalability opted Snowflake. Using different Objects of Snowflake, Migrated the Spark code to Snowpark. Handle CDC through Stream, Tasks in Snowflake.
- Databricks : Developed the business logic layer in Spark Scala to orchestrate Apache Airflow workflows, triggering Databricks jobs, pipelines and DBT for automated data processing.
Responsibilities:
- Migration - Architect the code movement from Different sources to Cloud- AWS & Snowflake, with Proper remediation need to the Code.
- Work in AWS S3, EC2, AWS Glue, Manged Apache Airflow, AWS Cloud Watch.
- Maintain the code base in Proper buckets with proper IAM roles in AWS S3.
- Implemented the Business Logic in different layers like Prestage, Stage, Target, Archive Tables and club it inside Stored Procedures.
- Snowflake Migration Project – Architect & Developed from scratch Snowflake Stage, Tables, File format, Functions, Views, Stored Procedures. Set up Connection with Local spark code using Snowpark for Interactive Testing.
- Architect & Create Snowpark Code from scratch to pull Data from different types of Sources like Flat Files, API, Web Scraping to Snowflake Stages, Tables.
- Code in PySpark to extract data from Files and Tables, Transforming, and onboarding the Data to Load Tables (Staging) and then enrich the data with static table and stored in Consolidation tables.
- Schedule the Jobs in IST/UTC/PST time in Apache Airflow Python Scripts as per the client requirement.
- Develop and enhance the Spark SQL code for many to one mapping as per the requirement.
- Developed scalable ETL pipelines leveraging Delta Tables to support Slowly Changing Dimensions (SCD), ensuring data consistency and historical tracking.
- Engineered a proactive monitoring and alerting solution using PySpark, SMTP, and email APIs, integrated with Databricks and emails to provide real-time pipeline completion and failure notifications.
- Leveraged Claude AI and Databricks Genie to develop AI-driven business solutions, AI Agents, automate workflows, and enhance decision-making processes.
Project Performance Optimizations:
- Enhanced the Spark Code by using Spark Optimization Techniques.
- Setting up Spark speed Test Logic using click House Jar.
- Improved test scenario efficiency by identifying and eliminating redundant test cases, increasing test coverage while reducing execution time.
- Recognized for optimizing the build process, reducing execution time from 11 min to 6 mins and enhancing overall delivery efficiency.
Process Optimization & Efficiency Improvement:
- Shared the CICD process to Implement Schemachange for Snowflake Objects.
- Collaborated closely with the DevOps team to implement AI-driven code review processes during GitHub code commits, improving code quality and compliance.
- Organized and standardized logging frameworks across Managed Apache Airflow, custom application loggers, and Snowflake Query History, Databricks Jobs&Pipeline logs improving monitoring and troubleshooting efficiency.
- Lead team in Migration Project. Help and guide the Colleagues.
- Responsible for resolving issues and queries raised by team members within the team.
Technologies: PySpark, Spark Scala, AWS S3, EC2, Glue, Airflow, Snowflake, Snowpark, Databricks.