Summary
Overview
Work History
Education
Skills
Affiliations
Timeline
Generic
PRAGNA KATASANI

PRAGNA KATASANI

Carbondale,IL

Summary

Data Engineer with Around 4 years of professional experience in Data Extraction, Data Modelling, Data Mining, and Data Visualization and 2+ years of Research oriented experience in Data Analysis on Real time datasets of state government.

  • Extensive experience with Informatica (ETL Tool) for Data Extraction, Transformation and Loading. Skilled in importing and exporting data using Sqoop between HDFS and RDBMS and adapting the process according to client's requirements.
  • Experienced in developing data marts and warehousing with advanced transformation for ETL (Extract, Transform & Load Process) using SQL, PostgreSQL, HiveQL, SAS, Python (Pandas, NumPy) and PySpark.
  • Extensively used Azure Databricks for data validations and analysis on Cosmos structured steams.
  • Knowledge of Big Data tools and Hadoop ecosystem components like Map Reduce, HDFS, Hive, Sqoop, Apache Spark, and Kafka.
  • Hands-on experience in implementing LDA, and Naive Bayes and skilled in Random Forests, Decision Trees, Linear and Logistic Regression, SVM, Clustering, neural networks, and Principal Component Analysis.

Overview

5
5
years of professional experience
2
2

Research Experience

Work History

Data Engineer

Envision Technologies
NEWARK, Delaware
08.2023 - Current
  • Involved in analysis, specification, design, implementation, and testing phases of Software Development Life Cycle (SDLC) and used Agile methodology for developing applications.
  • Analyzed and gathered the requirements in the project documentation for business requirements for ETL processes by conducting workshops and meetings with various business users.
  • Performed data profiling and analysis making use of Informatica Data Explorer (IDE) and Informatica Data Quality (IDQ).
  • Worked on migrating an existing on-premises application to the AWS platform using various services and experienced in maintaining the Hadoop cluster on AWS EMR.
  • Worked on Amazon AWS concepts like EMR and EC2 web services for fast and efficient processing of Big Data.
  • Build data pipelines in airflow in GCP for ETL related jobs using different airflow operators.
  • Experience in GCP, GCS, Cloud functions, Big Query.

Data Engineer

APSIS Technologies Pvt, LTD.
Bangalore, India
07.2019 - 06.2021
  • Produced and maintained Tableau data sources and data extracts, improving data accuracy by 15% and reducing data processing time by 20%
  • Developing the Sqoop scripts to make the interaction between Pig and MySQL Database
  • Utilizing Hive to analyze the partitioned and bucketed data and compute various metrics for reporting
  • Working with SQOOP import and export functionalities to handle large data set transfer between DB2 database and HDFS
  • Responsible for operations and support of big data Analytics platform, Splunk, and Tableau visualization
  • Building predictive models including ensemble models using machine learning algorithms such as Logistic Regression, Random Forests, and KNN to predict customer churn
  • Developed, deployed, and maintained ETL pipelines in Azure Data Factory, ensuring timely and accurate data availability across the organization
  • Utilized Azure Databricks to perform data analytics, employing languages such as SQL and Python to derive insights and facilitate data-driven decision-making
  • Leveraged Informatica Intelligent Cloud Services (IICS) to design and implement scalable and efficient cloud-based integration solutions, enabling seamless data synchronization, transformation, and connectivity between diverse cloud and on-premises systems
  • Developed and deployed ETL workflows in IICS, optimizing data integration processes and ensuring data accuracy and integrity across multiple platforms
  • Collaborated with cross-functional teams to gather requirements and provide IICS-based solutions, enhancing overall system efficiency and business productivity
  • Create and implement data warehousing solutions using Python and data warehousing technologies, optimizing data storage and retrieval for complex analytical queries
  • Designing various Jenkins jobs to continuously integrate the processes and executed CI/CD pipeline using Jenkins
  • Using Snowflake functions to perform semi-structured data parsing entirely with SQL statements
  • Performing Code release from one environment to another environment using release management in Azure DevOps.

SE - Intern

Cybermatic Systems Pvt Ltd
, India
05.2017 - 04.2019
  • Performed unit testing and integration testing to ensure quality of the product before releasing it to customers.
  • Conducted research on the latest trends in software engineering best practices.
  • Resolved customer issues by establishing workarounds and solutions to debug and create defect fixes.
  • Produced supporting reports and documentation to help development team members complete project work.
  • Created technical documentation such as user manuals, flowcharts, and diagrams.

Education

Master of Science - Computer Science

Southern Illinois University
Carbondale, IL
05.2023

Skills

  • Scala, Python, SQL
  • PyCharm, Jupyter Notebook
  • Hadoop, MapReduce, Hive, Pig, Redshift, DynamoDB, BigQuery, Kubernetes
  • SSIS, Informatica, Informatica Cloud (IICS)
  • AWS, Azure (Azure DevOps pipelines, Databricks)
  • NumPy, Pandas, Matplotlib, SciPy, Scikit-learn, Seaborn, TensorFlow, Kafka, PySpark
  • Tableau, Power BI, SSRS, Jenkins
  • MS SQL Server, PostgreSQL, MongoDB, MySQL

Affiliations

Master’s Thesis:

Kaggle Data Analysis using Big Data Technologies:

  • In this project, I utilized Azure Databricks as the working environment, leveraging Pyspark, Scala, and DBFS for efficient data analysis.
  • The data sets from Kaggle were processed using parallel processing clusters, focusing on technologies such as Hadoop, Hive, Spark, and Scoop in Oracle Virtual Box. This approach allowed for the seamless loading and processing of data using Hive Query Language.

IPL Data Analysis (2008-2020):

  • The second project involved analyzing IPL data spanning from 2008 to 2020. Python was employed to clean the dataset before importing it into Power BI for creating insightful dashboards.
  • The emphasis was on identifying and understanding factors that contribute to a team's success in the IPL. The process involved data cleaning, manipulation, and visualization.

Personality and Substance Abuse Study:

  • This project delved into the interrelation between personality traits and substance abuse, specififically focusing on volatile substances. Python was used for generating visualizations like heat maps and bar diagrams.
  • SQL queries were employed for data retrieval, and machine learning models such as linear regression, logistic regression, and KNN were applied to classify user behavior and substance abuse patterns over certain periods.
  • Also, Principal Component Analysis (PCA) was used to decompose clusters into one component, enhancing understanding of the data structure.

Timeline

Data Engineer

Envision Technologies
08.2023 - Current

Data Engineer

APSIS Technologies Pvt, LTD.
07.2019 - 06.2021

SE - Intern

Cybermatic Systems Pvt Ltd
05.2017 - 04.2019

Master of Science - Computer Science

Southern Illinois University
PRAGNA KATASANI