Summary
Overview
Work History
Education
Skills
Timeline
Hi, I’m

Sandeep Sangem

Naperville,IL
Sandeep Sangem

Summary

· 9+ years of experience in data analysis, data engineering, and statistical modeling, including data extraction, manipulation, visualization, and validation techniques, and reporting on various projects.

· Experienced with full software development life cycle, architecting scalable platforms, object-oriented programming, database design, and agile methodologies.

· Worked on Scala codebase related to Apache Spark performing the Actions, Transformations on RDDs, Data Frames & Datasets using Spark SQL and Spark Streaming Contexts.

· Experience in data architecture including data ingestion pipeline design, Hadoop information architecture, data modeling, data mining, machine learning, and advanced data processing.

· Expertise in creating Pods using Kubernetes and worked with Jenkins pipelines to drive all microservices builds out to the Docker registry and then deployed to the Kubernetes cluster.

· Extensive experience in Hadoop-led development of enterprise-level solutions utilizing Hadoop components such as Apache Spark, MapReduce, HDFS, Sqoop, PIG, Hive, HBase, Oozie, Flume, NiFi, Druid, Kafka, Zookeeper, YARN.

· Profound experience in performing Data Ingestion, Data Processing (Transformations, enrichment, and aggregations).

· Working knowledge of modeling, loading, and optimizing large amounts of data into druid to be queried by low latency web applications via API.

· Working with GCP cloud using GCP Cloud storage, DataProc, Data Flow, Big Query, Cloud Composer, Cloud Pub/Sub.

· Expert in working with cloud PUB/SUB to replicate data in real-time from the source system to GCP Big Query.

· Experienced professional adept at leveraging Oracle Cloud services to deliver scalable and efficient solutions. Skilled in Oracle Cloud Infrastructure (OCI), Oracle Database, and Oracle Cloud Applications.

· Strong Knowledge of the Architecture of Distributed systems and Parallel processing, In-depth understanding of MapReduce programming paradigm and Spark execution framework.

· Experienced with the Spark improving the performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark-SQL, Dataframe API, Spark Streaming, MLlib, and Pair RDD, and worked explicitly on PySpark and Scala.

· Handled ingestion of data from different data sources into HDFS using Sqoop, and Flume and perform transformations using Hive, and Map Reduce and then loaded data into HDFS.

· Managed Sqoop jobs with incremental load to populate HIVE external tables. Experience in importing streaming data into HDFS using Flume sources, and Flume sinks and transforming the data using Flume interceptors.

· Experience in Oozie and workflow scheduler to manage Hadoop jobs by Direct Acyclic Graph (DAG) of actions with control flow.

· Good experience in Amazon Web Service (AWS) concepts like EMR and EC2 Webservices which provide fast and efficient processing of Teradata Big Data Analytics.

· Experience in Big Data/Hadoop, Data Analysis, and Data Modeling professional with applied information Technology.

· Good knowledge in Database Creation and maintenance of physical data models with Oracle, Teradata, Netezza, DB2, MongoDB, HBase, and SQL Server databases.

· Proficient in AWS Cloud Platform which includes services like EC2, S3, VPC, ELB, DynamoDB, Cloud Front, Cloud Watch, Route 53, Security Groups, Redshift, CloudWatch, and CloudFormation.

· Migrated an existing on-premises application to AWS. Used AWS services like EC2 and S3 for small data sets processing and storage, Experienced in Maintaining the Hadoop cluster on AWS EMR.

· Integrated Kafka with Spark Streaming for real-time data processing.

· Strong experience in the Analysis, design, development, testing, and Implementation of Business Intelligence solutions using Data Warehouse/Data Mart Design, ETL, BI, Client/Server applications, and writing ETL scripts using Regular Expressions and custom tools (Informatica, Pentaho, and Sync Sort) to ETL data.

· Extensive experience with Azure services like HDInsight, Stream Analytics, Active Directory, Blob Storage, Cosmos DB, and Storage Explorer.

· Strong Experience in implementing Data warehouse solutions in Confidential Redshift; Worked on various projects to migrate data from on-premise databases to Confidential Redshift, RDS, and S3.

· Experience in Cloud Databases and Data warehouses (SQL Azure and Confidential Redshift/RDS).

· Experience working in Azure Cloud, Azure DevOps, Azure Data Factory, Azure Data Lake Storage, Azure Synapse Analytics, Azure Analytical services, Azure Cosmos NO SQL DB, Azure HD Insight Bigdata Technologies (Hadoop and Apache Spark), and Data bricks.

· Worked on setting up Data Lake/Data catalog on AWS Glue.

· Experience with different file formats like Avro, parquet, ORC, JSON, and XML.

· Expertise in Creating, Debugging, Scheduling, and Monitoring jobs using Airflow and Oozie.

· Experienced with using most common Operators in Airflow - Python Operator, Bash Operator, Google Cloud Storage Download Operator, and Google Cloud Storage Object Sensor.

· Hands-on experience in handling database issues and connections with SQL and NoSQL databases such as MongoDB, HBase, Cassandra, SQL Server, and PostgreSQL.

· Created Java apps to handle data in MongoDB and HBase. Used Phoenix to create SQL layer on HBase.

· Experience in designing and creating RDBMS Tables, Views, User Created Data Types, Indexes, Stored Procedures, Cursors, Triggers, and Transactions.

· Expert in designing ETL data flows using creating mappings/workflows to extract data from SQL Server and Data Migration and Transformation from Oracle/Access/Excel Sheets using SQL Server SSIS.

· Expert in designing Parallel jobs using various stages like Join, Merge, Lookup, remove duplicates, Filter, Dataset, Lookup file set, Complex flat file, Modify, Aggregator, and XML.

· Hands-on experience with Amazon EC2, Amazon S3, Amazon RDS, VPC, IAM, Amazon Elastic Load Balancing, Auto Scaling, CloudWatch, SNS, SES, SQS, Lambda, EMR, and other services of the AWS family.

· Created and configured a new batch job in the Denodo scheduler with email notification capabilities Implemented Cluster setting for multiple Denodo nodes and created load balance for improving performance activity.

· Instantiated, created, and maintained CI/CD (continuous integration & deployment) pipelines and apply automation to environments and applications.

· Worked on various automation tools like GIT, Terraform, and Ansible. Experienced in fact dimensional modeling (Star schema, Snowflake schema), transactional modeling, and SCD (Slowly changing dimension)

· Experienced with JSON-based RESTful web services, and XML/QML-based SOAP web services and worked on various applications using python integrated IDEs like Sublime Text and PyCharm

· Efficient Cloud Engineer with years of experience assembling cloud infrastructure. Utilizes strong managerial skills by negotiating with vendors and coordinating tasks with other IT team members. Implements best practices to create cloud functions, applications, and databases.

Overview

10
years of professional experience

Work History

Northern Trust Bank

Sr. Data Engineer
12.2021 - Current

Job overview

  • Understand requirements, and building codes, and guide other developers during development activities to develop high-standard stable codes within the limits of Confidential and client processes, standards, and guidelines
  • Involved in migrating the client data warehouse architecture from on-premises into Azure cloud and implementation of data movements from on-premises to cloud in Azure
  • Develop, and design data models, data structures, and ETL jobs for data acquisition and manipulation purposes
  • Architect & implement medium to large-scale BI solutions on Azure using Azure Data Platform services (Azure Data Lake, Data Factory, Data Lake Analytics, Stream Analytics, Azure SQL DW, HDInsight/Databricks, NoSQL DB, Cosmos DB)
  • Create pipelines in ADF using linked services to extract, transform and load data from multiple sources like Azure SQL, Blob storage and Azure SQL Data warehouse
  • Creating storage accounts that involved with the end-to-end environment for running jobs
  • Develop batch processing solutions by using Data Factory and Azure Databricks
  • Implement Azure Data bricks clusters, notebooks, jobs, and auto-scaling
  • Performed ETL operations in Azure Databricks by connecting to different relational database source systems using JDBC connectors
  • Developed Python scripts to do file validations in Databricks and automated the process using ADF
  • Developed an automated process in Azure cloud that can ingest data daily from web service and load it into Azure SQL DB
  • Worked with Business Owners on perfecting the process, increasing the overall efficiency of the systems Responsible for building scalable distributed data solutions using Azure Data Lake, and Azure Databricks
  • Azure HD Insight & Azure Cosmos DB
  • Developed Streaming pipelines using Azure Event Hubs and Stream Analytics to analyze data for dealer efficiency and open table counts for data coming in Fiot-enabled poker and other pit tables
  • Established a formal EDM, MDM (Master Data Management) program that creates effective engagement between Business operations, EDM, delivery team, and IT
  • Analyzed data where it lives by Mounting Azure Data Lake and Blob to Databricks
  • Used Logic App to take decisional actions based on the workflow
  • Implement Copy activity and Custom Azure Data Factory Pipeline Activities
  • Design and implement database solutions in Azure SQL Data Warehouse, Azure SQL
  • Developed custom alerts using Azure Data Factory, SQLDB, and Logic App
  • Developed Databricks ETL pipelines using notebooks, Spark Data frames, SPARK SQL, and Python scripting
  • Design for data auditing and data masking Design for data encryption for data at rest and in transit Design relational and non-relational data stores on Azure Preparing ETL test strategy, designs, and test plans to execute test cases for ETL and BI systems
  • Creating ETL test scenarios and test cases and plans to execute test cases
  • Experience in building an efficient pipeline for moving data between GCP and Azure using Azure Data Factory
  • Developed multi-cloud strategies in better using GCP (for its PASS) and Azure (for its SAAS)
  • Experience in moving data between GCP and Azure using Azure Data Factory
  • Create Data Flows of existing SSIS Packages in Azure Data Factory
  • Designing and maintaining ADF pipelines with activities Copy, Lookup, For Each, Get Metadata, Execute Pipeline, Stored Procedure, if condition, Web, Wait, Delete, etc
  • Used Python and Shell scripts to Automate Teradata ELT and Admin activities
  • Performed Application-level DBA activities creating tables, and indexes, and monitored and tuned Teradata BETQ scripts using Teradata Visual Explain utility
  • Integrated both framework and CloudFormation to automate Azure environment creation along with the ability to deploy on Azure, using build scripts (Azure CLI) and automate solutions using Terraform
  • Extract Transform and Load data from Sources Systems to Azure Data Storage services using a combination of Azure Data Factory, T-SQL, Spark SQL, and U-SQL Azure Data Lake Analytics
  • Data Ingestion to one or more Azure Services - (Azure Data Lake, Azure Storage, Azure SQL, Azure DW) and processing the data in In Azure Databricks
  • Performance tuning, monitoring, UNIX shell scripting, and physical and logical database design
  • Performed ETL operation using SSIS and loaded the data into Secure DB
  • Good hands-on experience in Data Vault concepts, and data models, well-versed understanding, and implementation of Data warehousing concepts/Data Vault
  • Designed, reviewed, and created primary objects such as views, and indexes based on logical design models, user requirements, and physical constraints
  • Worked with stored procedures for data set results for use in Reporting Services to reduce report complexity and optimize the run time
  • Exported reports into various formats (PDF, Excel) and resolved formatting issues
  • Collaborate with application architects on infrastructure as a service (IaaS) application to Platform as a Service (PaaS).

Care first (BCPS fepoc )

AWS Data Engineer
08.2019 - 12.2021

Job overview

  • Worked on building the data pipelines (ELT/ETL Scripts), extracting the data from different sources (MySQL, AWS S3 files), transforming, and loading the data to the Data Warehouse (AWS Redshift)
  • Used Agile Scrum methodology/ Scrum Alliance for development
  • Extensive experience in working with AWS cloud Platform (EC2, S3, EMR, Redshift, Lambda and Glue)
  • Worked on developing & adding few Analytical dashboards using Looker product
  • Worked on building the data pipelines using PySpark (AWS EMR), processing the data files present in S3 and loading it to Redshift
  • Worked on building the aggregate tables & de-normalized tables, populating the data using ETL to improve the looker analytical dashboard performance and to help data scientist and analysts to speed up the ML model training & analysis
  • Played a lead role in gathering requirements, analysis of the entire system and providing estimation on development, testing efforts
  • Developed custom Jenkins jobs/pipelines that contained Bash shell scripts utilizing the AWS CLI to automate infrastructure provisioning
  • Experience in Developing Spark applications using Spark - SQL in Databricks for data extraction, transformation, and aggregation from multiple file formats for analyzing & transforming the data to uncover insights into the customer usage patterns
  • Responsible for estimating the cluster size, monitoring, and troubleshooting of the Spark Data Bricks cluster
  • Worked on scheduling all jobs using Airflow scripts using python
  • Adding different tasks to DAG's and dependencies between the tasks
  • Developed Spark code using Python and Spark-SQL/Streaming for faster testing and processing of data and Data Extraction, aggregations, and consolidation of Adobe data within AWS Glue using PySpark
  • Performed SQL queries on AWS Athena on the database from AWS Glue
  • Implemented Spark using Scala and SparkSQL for faster testing and processing of data
  • Developed a user-eligibility library using Python to accommodate the partner filters and exclude these users from receiving the credit products
  • Built the data pipelines to aggregate the user click stream session data using spark streaming module which reads the click stream data from Kinesis streams and store the aggregate results in S3 and data and eventually loaded to AWS Redshift warehouse
  • Worked on supporting & building the infrastructure for the core module of the Credit Sesame i.e., Approval Odds, started with Batch ETL, moved to micro-batches, and then converted to a real time predictions
  • Used Jira for ticketing and tracking issues and Jenkins for continuous integration and continuous deployment
  • Enforced standards and best practices around data catalog, data governance efforts
  • Created numerous ODI interfaces and load into Snowflake DB
  • Developing and writing SQLs and stored procedures in Teradata
  • Loading data into a snowflake and writing Snow SQL scripts
  • Worked on Amazon Redshift for shifting all Data warehouses into one Data warehouse
  • Good understanding of Cassandra architecture, replication strategy, gossip, snitches, etc
  • Designed columnar families in Cassandra and Ingested data from RDBMS, performed data transformations and then exported the transformed data to Cassandra as per the business requirement
  • Created DataStage jobs using different stages like Transformer, Aggregator, Sort, Join, Merge, Lookup, Data Set, Funnel, Remove Duplicates, Copy, Modify, Filter, Change Data Capture, Change Apply, Sample, Surrogate Key, Column Generator, Row Generator, Etc
  • Good experience with Version Control tools Bitbucket, GitHub, and GIT
  • Experience with Jira, Oozie, and Airflow scheduling tools
  • Experienced in Strong scripting skills in Python, Scala, and UNIX shell
  • Involved in writing Python, and Java APIs for Amazon Lambda functions to manage the AWS services
  • Used the Spark Data Cassandra Connector to load data to and from Cassandra
  • Worked from Scratch in Configurations of Kafka such as Managers and Brokers
  • Developed the AWS Lambda serverless scripts to handle ad-hoc requests
  • Performed Cost optimization and reduced infrastructure costs
  • Knowledge and experience in using Python NumPy, Pandas, Sci-kit Learn, Onnx & Machine Learning
  • Worked on scheduling all jobs using Airflow scripts using Python and added different tasks to DAG, and LAMBDA
  • Worked on adding the Rest API layer to the ML models built using Python, and Flask& deploying the models in AWS Beanstalk Environment using Docker containers
  • Other activities include supporting and keeping the data pipelines active, working with Product Managers, Analysts, and Data Scientists & addressing the requests coming from them, unit testing, load testing, and SQL optimizations.

Citi Bank

Big Data Engineer
09.2017 - 08.2019

Job overview

  • Implemented Spark using Scala and utilizing Data frames and Spark SQL API for faster processing of data
  • Ingested data from RDBMS and performed data transformations, and then export the transformed data to Cassandra as per the business requirement
  • Design and development of ETL processes using the Informatica ETL tool for dimension and fact file creation
  • Performed wide, and narrow transformations, actions like filter, Lookup, Join, count, etc
  • On Spark Data Frames
  • Worked with Parquet files and Impala using PySpark, and Spark Streaming with RDDs and Data Frames
  • Involved in Uploading Master and Transactional data from flat files and preparation of Test cases, Sub System Testing
  • Aggregated logs data from various servers and made them available in downstream systems for analytics by using Apache Kafka
  • Improved Kafka performance and implemented security
  • Developed batch and streaming processing apps using Spark APIs for functional pipeline requirements
  • Worked with Spark to create structured data from the pool of unstructured data received
  • Implemented intermediate functionalities like events or records count from the flume sinks or Kafka topics by writing Spark programs in java and python
  • Documented the requirements including the available code which should be implemented using Spark, Hive, HDFS
  • Experienced in transferring Streaming data, data from different data sources into HDFS, No SQL databases
  • Created ETL Mapping with Talend Integration Suite to pull data from Source, apply transformations, and load data into target database
  • Transformed data from different files (Text, CSV, JSON) using Python scripts in Spark
  • Loaded data from various sources like RDBMS (MySQL, Teradata) using Sqoop jobs
  • Well versed with the Database and Data Warehouse concepts like OLTP, OLAP, Star Schema
  • AWS provides a secure global infrastructure, plus a range of features that use to secure the data in the cloud
  • Worked and learned a great deal from Amazon Webservices (AWS) Cloud services like EC2, S3, EBS, RDS and VPC
  • Developed multiple Kafka Producers and Consumers from scratch to as per the software requirement
  • Performed advanced procedures like text analytics and processing, using the in-memory computing capabilities of Spark using python.

Sakshath Technologies

ETL Developer
07.2016 - 08.2017

Job overview

  • Involved in understanding the requirements of the End Users/Business Analysts and Developed Strategies for ETL processes
  • Developed mappings/Reusable Objects/Transformation by using a mapping designer, and transformation developer in Informatica Power Center
  • Extensively used Informatica Client Tools Source Analyzer, Warehouse Designer, Transformation Developer, Mapping Designer, Mapplet Designer, Informatica Repository
  • Designed and developed ETL Mappings to extract data from flat files, and Oracle to load the data into the target database
  • Developed SQL queries to develop the Interfaces to extract the data in regular intervals to meet the business requirements
  • Used various transformations like Unconnected/Connected Lookup, Aggregator, Expression Joiner, Sequence Generator, Router etc
  • Used ETL to load data using PowerCenter/Power Connect from source systems like Flat Files and Excel Files into staging tables and load the data into the target database
  • Developed complex mappings using multiple sources and targets in different databases, and flat files
  • Designed and developed mappings using Source Qualifier, Aggregator, Joiner, Lookup, Sequence Generator, Stored Procedure, Expression, Filter, and Rank transformations and validated the Data
  • Documentation of Technical specifications, business requirements, and functional specifications for the development of Informatica Extraction, Transformation, and Loading (ETL) mappings to load data into various tables.

Mindtree

SQL Developer
08.2014 - 07.2016

Job overview

  • Proficient working experience with SQL, PL/SQL, and Database objects like Stored Procedures, Functions, and Triggers and using the latest features to optimize the performance of Inline views and Global Temporary tables
  • Performed the data analysis and mapping database normalization, performance tuning, query optimization data extraction, transfer, loading (ETL), and clean up
  • Created SSIS Packages using SSIS Designer for export heterogeneous data from OLE DB Source (Oracle), Excel Spreadsheet to SQL Server
  • Extensive use of Triggers to implement business logic and for auditing changes to critical tables in the database
  • Experience in developing external Tables, Views, Joins, Cluster indexes, and Cursors
  • Defining data warehouse (star and snowflake schema), fact table, cubes, dimensions, and measures using SQL Server Analysis Services
  • Used Execution Plan, SQL Profiler, and Database Engine Tuning Advisor to optimize queries and enhance the performance of databases
  • Worked on the data warehouse design and analyzed various approaches for maintaining different dimensions and facts in the process of building a data warehousing application
  • Using reporting services (SSRS) generated various reports
  • Optimized query performance by creating Indexes.

Education

VNR Vignana Joythi
Hyderabad

Bachelor of Science from Electrical, Electronics Engineering Technologies
05.2015

University Overview

Skills

  • Advanced analytics
  • Data Governance
  • Hadoop Ecosystem
  • Advanced data mining

Timeline

Sr. Data Engineer
Northern Trust Bank
12.2021 - Current
AWS Data Engineer
Care first (BCPS fepoc )
08.2019 - 12.2021
Big Data Engineer
Citi Bank
09.2017 - 08.2019
ETL Developer
Sakshath Technologies
07.2016 - 08.2017
SQL Developer
Mindtree
08.2014 - 07.2016
VNR Vignana Joythi
Bachelor of Science from Electrical, Electronics Engineering Technologies
Sandeep Sangem