PROFESSIONAL SUMMARY
Overview
Work History
Education
Skills
Languages
Interests
Timeline

DINESH KUMAR MOHAPATRA

Capgemini
KOLKATA,WestBengal
DINESH KUMAR MOHAPATRA
1
Language
9
years of professional experience
  • 11+ years of professional IT experience in Analysis, Design, Development, Enhancement, Support, Testing, and Implementation of Big Data and Data Analytics applications.
  • 8+ years of experience in designing, developing, and enhancing data pipelines using Spark Scala, PySpark, Hadoop, and Cloud platforms within the Big Data ecosystem.
  • Led migration projects, including solution architecture, workflow design, implementation, and delivery.
  • Architected and developed Snowflake data pipelines and Snowpark solutions from scratch. Designed and implemented Snowflake Tasks and Streams to support Change Data Capture (CDC) processes.
  • Hands-on experience with cloud technologies including AWS S3, EC2, Glue, CloudWatch, Airflow, and Microsoft Azure.
  • Skilled in processing large volumes of Structured, Semi-Structured, and Streaming data.
  • Strong expertise in data ingestion strategies such as Full Load, Incremental Load, Delta Load, CDC, Partitioning, and Bucketing for Spark SQL optimization.
  • Domain expertise in Medical, Diagnostics, Banking & Finance and Retail Domains.
  • Experience working with Git, Github, Actions, AI code Review, Azure DevOps CI/CD Pipelines, and Agile methodologies.
  • Skilled in analyzing business requirements, collaborating with stakeholders, performing source-to-target data mapping, solution design, and code reviews.
  • Strong problem-solving, analytical, and innovation skills with a focus on continuous improvement and operational excellence.
  • Proven leadership capabilities with strong interpersonal, communication, and team management skills.

Work History

Data Engineer Specialist

2 Years
Capgemini | 08.2024 - Current

Project Name: US based Diagnostics Company

Data Migration Project: Architected and implemented end-to-end cloud data migration solutions leveraging PySpark, Spark Scala, Databricks, and Snowflake, enabling the ingestion, transformation, and loading of structured and semi-structured data from AWS S3 into optimized Snowflake tables, Databricks based on business requirements.

  • Snowflake - Snowpark Project: To obtain Cloud based Data warehousing Tool and high scalability opted Snowflake. Using different Objects of Snowflake, Migrated the Spark code to Snowpark. Handle CDC through Stream, Tasks in Snowflake.
  • Databricks : Developed the business logic layer in Spark Scala to orchestrate Apache Airflow workflows, triggering Databricks jobs, pipelines and DBT for automated data processing.

Responsibilities:

  • Migration - Architect the code movement from Different sources to Cloud- AWS & Snowflake, with Proper remediation need to the Code.
  • Work in AWS S3, EC2, AWS Glue, Manged Apache Airflow, AWS Cloud Watch.
  • Maintain the code base in Proper buckets with proper IAM roles in AWS S3.
  • Implemented the Business Logic in different layers like Prestage, Stage, Target, Archive Tables and club it inside Stored Procedures.
  • Snowflake Migration Project – Architect & Developed from scratch Snowflake Stage, Tables, File format, Functions, Views, Stored Procedures. Set up Connection with Local spark code using Snowpark for Interactive Testing.
  • Architect & Create Snowpark Code from scratch to pull Data from different types of Sources like Flat Files, API, Web Scraping to Snowflake Stages, Tables.
  • Code in PySpark to extract data from Files and Tables, Transforming, and onboarding the Data to Load Tables (Staging) and then enrich the data with static table and stored in Consolidation tables.
  • Schedule the Jobs in IST/UTC/PST time in Apache Airflow Python Scripts as per the client requirement.
  • Develop and enhance the Spark SQL code for many to one mapping as per the requirement.
  • Developed scalable ETL pipelines leveraging Delta Tables to support Slowly Changing Dimensions (SCD), ensuring data consistency and historical tracking.
  • Engineered a proactive monitoring and alerting solution using PySpark, SMTP, and email APIs, integrated with Databricks and emails to provide real-time pipeline completion and failure notifications.
  • Leveraged Claude AI and Databricks Genie to develop AI-driven business solutions, AI Agents, automate workflows, and enhance decision-making processes.

Project Performance Optimizations:

  • Enhanced the Spark Code by using Spark Optimization Techniques.
  • Setting up Spark speed Test Logic using click House Jar.
  • Improved test scenario efficiency by identifying and eliminating redundant test cases, increasing test coverage while reducing execution time.
  • Recognized for optimizing the build process, reducing execution time from 11 min to 6 mins and enhancing overall delivery efficiency.

Process Optimization & Efficiency Improvement:

  • Shared the CICD process to Implement Schemachange for Snowflake Objects.
  • Collaborated closely with the DevOps team to implement AI-driven code review processes during GitHub code commits, improving code quality and compliance.
  • Organized and standardized logging frameworks across Managed Apache Airflow, custom application loggers, and Snowflake Query History, Databricks Jobs&Pipeline logs improving monitoring and troubleshooting efficiency.
  • Lead team in Migration Project. Help and guide the Colleagues.
  • Responsible for resolving issues and queries raised by team members within the team.

Technologies: PySpark, Spark Scala, AWS S3, EC2, Glue, Airflow, Snowflake, Snowpark, Databricks.

Big Data & Snowpark Developer

1 Year 9 Months
IBM | 11.2022 - 08.2024
  • Project Name: US based Retail Company
  • Hana to Snowflake Migration Project: Created the business logic in snowflake. Worked in SNP Glue connector to pull records. Created PySpark- Snowpark script for Automating Testing. Handled all test scenarios which were part of Testing tracks. Architect Dynamic Tables, Implemented CDC, maintained the code base.
  • Responsibilities
  • Hana challenges – Identified the slowness, challenges to Bigdata sets.
  • Identified the design at Hana and replicated the logic with Snowflake.
  • Implemented the Dynamic Tables which provides dependable and cost-effective.
  • Worked in AWS EC2, Airbyte to pull records to Snowflake.
  • Worked in complex join condition. Analysed the Data variance in Join condition.
  • Pivot the complex data to flatten with proper business logic.
  • Created the Automation script in PySpark- Snowpark to automate the testing tracks.
  • Derived new Business to Company with proper plans and Business optimization layouts.
  • Implemented CDC through Stream, Tasks in Snowflake. UAT Testing end to end post development.
  • Understand the Business requirement, prepare data model, and design documents for each Application Business stream.
  • Technologies: Hadoop, HDFS, Sqoop, Oozie, Scala, AWS S3, EC2, Glue, Airflow, Snowflake, Snowpark

Big Data Developer

1 Year 6 Months
IBM | 05.2021 - 11.2022
  • Project Name: US based Banking Company
  • This project involves in, creating Data pipeline in Spark Scala with Hadoop Ecosystem for obtaining Data from different SOR to Hadoop as per Business requirements in Agile Methodology.
  • Responsibilities
  • Create and test PySpark Framework using Azure services so that the Data is pushed from different storage (BLOB storage, ADLS) to Delta Lake to Hadoop layers.
  • Migrate from Hana to Delta Lake.
  • Implement the Acid properties in the Delta Lake.
  • Code the different layers of PySpark class and implement the business logic Create the Python Wheel file and deploy in Azure create jobs.
  • Schedule the Azure Notebook by assigning dynamic parameters in runtime.
  • Mapping, Data Quality checks for each Application.
  • Test in UAT to validate the framework and business functionalities before sharing it with the Business and perform SIT validation using Pytest.
  • Analyzing business requirements to formulate strategies for implementation of Big Data initiatives.
  • Check the code rating with Software like Pylint before creating PR requests.
  • Maintain the code in Git repo for CI/CD Pipeline.
  • Taking Technical Interview for the Organization.
  • Technologies: PySpark, Python, Azure Databricks, Blob Storage, ADLS, Azure Notebooks Scheduling

Data Engineer

6 Months
Wipro | 10.2020 - 04.2021
  • Project Name: Large US based Bank
  • This project involves in, creating Data pipeline in Spark Scala with AWS Environment for obtaining the Data from different application sources to AWS S3 as per Business requirements in Agile Methodology.
  • Responsibilities
  • Work with Client Data Modernization Team directly as an independent resource from the Offshore.
  • Create and test Spark Framework so that the Data is pushed from different Application like One Lake, Snowflake, Cerebro to AWS S3 layers.
  • Transformed the existing pipeline from Cerebro - AWS S3 to Snowflake - AWS S3.
  • Created Shell scripts to call Spark Submit Jobs Automated the applications by creating Arrow Jobs to call appropriate Shell Scripts as per Business scheduled Created the Documentation for STTMs, Extract scripts, File Selection Criteria, Data Mapping, Data Quality checks for each Application.
  • Work in Spark Shell to validate the output files generated in UAT as validation before sharing to the Business
  • Work in GitHub, maintaining different branches as per the Jira Story, creating PR.
  • Work in Jira, to Create story, Groom the story, create sub task and update the Workflow of the story.
  • Monitor and guide the juniors from Client Team both Technical and Business Functionalities.
  • Technologies: Spark Scala, Snowflake, Shell Script, AWS S3, Putty, Arrow, Git, GitHub, Jira, Confluence

Spark streaming and Spark-SQL Developer

2 Years 1 Month
Cognizant | 09.2018 - 10.2020
  • Project Name: Large Retail Company in Sweden
  • This project involves in, creating Data pipeline from scratch in which procedure was to schedule an automated SSIS Job to get the Data coming from file share to Data Lake Storage (DLS). Once Data is in DLS, we created a Spark based framework written in Scala to transfer the Data into Client's specific platform for further Analytics.
  • Responsibilities
  • Worked in Creating and testing Spark Framework so that the Data is pushed from Source Platform to Landing Layers.
  • Effectively articulated and worked with senior stakeholders regarding the proper structure of the Schema
  • Created and tested SSIS Job so that the Data is pushed to DLS from Source Destination
  • Scheduled the same SSIS Job by SQL server Agent for Daily so that it is pushed as when dropped in File share.
  • Created a Spark framework to push the Data from DLS to Base (Client's specific platform).
  • Implement some Business Transformations along with cleaning the Data if required in Spark framework.
  • Created schema, Temporary view to run the Spark SQL Queries. Automated the Spark Job in Airflow and Schedule the Directed Acyclic Graph (DAG).
  • Implemented some of the common functional Business Calculation in retail Data. Worked on source control repository management system, i.e., GIT, Bit Bucket, Followed Agile model using JIRA.
  • Enhanced customer retention by 28% and registered an increase in customer satisfaction scores from 2.5 to 4.2
  • Technologies: Spark Scala, SSIS (Visual Studio), SQL server, Python, Airflow, Maven, GIT, Bit bucket, Jira

Spark streaming and Spark-SQL Developer

7 Months
Infosys | 01.2018 - 08.2018
  • Project Name: Large US Based Financial Services
  • Responsibilities
  • Performed Import and Export of data into HDFS and Hive using Sqoop and managed data within the environment.
  • Involved in handling streaming Data from Kafka and processing it in Spark.
  • Was responsible for Optimizing Spark sql queries that helped in saving Cost to the project.
  • Handled streaming data using both Spark streaming and Spark structured streaming APIs.
  • Managed Kafka Consumption in an optimized way for better performance.
  • Exported necessary spark Jars to run in the cluster.
  • Involved in working on the Data Analysis, Data Quality and data profiling for handling the business that helped the Business team.
  • Loaded and transformed large sets of semi structured data like XML JSON, Avro, Parquet.
  • Code & peer review of assigned task. Unit testing and Bug fixing.
  • Technologies: Hadoop, HDFS, Hive, Sqoop, Spark SQL, Spark streaming, Spark structured streaming, Kafka
  • A community-based financial services company with $1.9 trillion in assets offering banking, investment and mortgage products and services, as well as consumer and commercial finance, through 8,050 locations.

Education

B. TECH - EE

ITER, SOA University | Bhubaneshwar | 04-2015

CGPA: 8.3

Skills

Big Data Ecosystem: Apache Spark Scala
PySpark
Hadoop
Sqoop
Hive
Languages: Scala
Python
Unix Shell Scripting
Databases: Oracle
MySQL
Snowflake
Snowflake
Snowpark
Databricks DBT
CICD
Schemchange
GitHub
Scheduling tools: Airflow
Snowflake Task & Streams
Autosys
AWS Cloud service: AWS S3
EC2
Glue
Lambda
Azure Databricks
Azure Storages
Delta Tables
Generative AI - AI agent
Claude code

Languages

English
Full Professional

Interests

Generative AI
Music

Timeline

Data Engineer Specialist

Capgemini
08.2024 - CurrentRead More

Big Data & Snowpark Developer

IBM
11.2022 - 08.2024Read More

Big Data Developer

IBM
05.2021 - 11.2022Read More

Data Engineer

Wipro
10.2020 - 04.2021Read More

Spark streaming and Spark-SQL Developer

Cognizant
09.2018 - 10.2020Read More

Spark streaming and Spark-SQL Developer

Infosys
01.2018 - 08.2018Read More

ITER, SOA University

B. TECH from EE
Read More
DINESH KUMAR MOHAPATRA