· 9+ years of experience in data analysis, data engineering, and statistical modeling, including data extraction, manipulation, visualization, and validation techniques, and reporting on various projects.
· Experienced with full software development life cycle, architecting scalable platforms, object-oriented programming, database design, and agile methodologies.
· Worked on Scala codebase related to Apache Spark performing the Actions, Transformations on RDDs, Data Frames & Datasets using Spark SQL and Spark Streaming Contexts.
· Experience in data architecture including data ingestion pipeline design, Hadoop information architecture, data modeling, data mining, machine learning, and advanced data processing.
· Expertise in creating Pods using Kubernetes and worked with Jenkins pipelines to drive all microservices builds out to the Docker registry and then deployed to the Kubernetes cluster.
· Extensive experience in Hadoop-led development of enterprise-level solutions utilizing Hadoop components such as Apache Spark, MapReduce, HDFS, Sqoop, PIG, Hive, HBase, Oozie, Flume, NiFi, Druid, Kafka, Zookeeper, YARN.
· Profound experience in performing Data Ingestion, Data Processing (Transformations, enrichment, and aggregations).
· Working knowledge of modeling, loading, and optimizing large amounts of data into druid to be queried by low latency web applications via API.
· Working with GCP cloud using GCP Cloud storage, DataProc, Data Flow, Big Query, Cloud Composer, Cloud Pub/Sub.
· Expert in working with cloud PUB/SUB to replicate data in real-time from the source system to GCP Big Query.
· Experienced professional adept at leveraging Oracle Cloud services to deliver scalable and efficient solutions. Skilled in Oracle Cloud Infrastructure (OCI), Oracle Database, and Oracle Cloud Applications.
· Strong Knowledge of the Architecture of Distributed systems and Parallel processing, In-depth understanding of MapReduce programming paradigm and Spark execution framework.
· Experienced with the Spark improving the performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark-SQL, Dataframe API, Spark Streaming, MLlib, and Pair RDD, and worked explicitly on PySpark and Scala.
· Handled ingestion of data from different data sources into HDFS using Sqoop, and Flume and perform transformations using Hive, and Map Reduce and then loaded data into HDFS.
· Managed Sqoop jobs with incremental load to populate HIVE external tables. Experience in importing streaming data into HDFS using Flume sources, and Flume sinks and transforming the data using Flume interceptors.
· Experience in Oozie and workflow scheduler to manage Hadoop jobs by Direct Acyclic Graph (DAG) of actions with control flow.
· Good experience in Amazon Web Service (AWS) concepts like EMR and EC2 Webservices which provide fast and efficient processing of Teradata Big Data Analytics.
· Experience in Big Data/Hadoop, Data Analysis, and Data Modeling professional with applied information Technology.
· Good knowledge in Database Creation and maintenance of physical data models with Oracle, Teradata, Netezza, DB2, MongoDB, HBase, and SQL Server databases.
· Proficient in AWS Cloud Platform which includes services like EC2, S3, VPC, ELB, DynamoDB, Cloud Front, Cloud Watch, Route 53, Security Groups, Redshift, CloudWatch, and CloudFormation.
· Migrated an existing on-premises application to AWS. Used AWS services like EC2 and S3 for small data sets processing and storage, Experienced in Maintaining the Hadoop cluster on AWS EMR.
· Integrated Kafka with Spark Streaming for real-time data processing.
· Strong experience in the Analysis, design, development, testing, and Implementation of Business Intelligence solutions using Data Warehouse/Data Mart Design, ETL, BI, Client/Server applications, and writing ETL scripts using Regular Expressions and custom tools (Informatica, Pentaho, and Sync Sort) to ETL data.
· Extensive experience with Azure services like HDInsight, Stream Analytics, Active Directory, Blob Storage, Cosmos DB, and Storage Explorer.
· Strong Experience in implementing Data warehouse solutions in Confidential Redshift; Worked on various projects to migrate data from on-premise databases to Confidential Redshift, RDS, and S3.
· Experience in Cloud Databases and Data warehouses (SQL Azure and Confidential Redshift/RDS).
· Experience working in Azure Cloud, Azure DevOps, Azure Data Factory, Azure Data Lake Storage, Azure Synapse Analytics, Azure Analytical services, Azure Cosmos NO SQL DB, Azure HD Insight Bigdata Technologies (Hadoop and Apache Spark), and Data bricks.
· Worked on setting up Data Lake/Data catalog on AWS Glue.
· Experience with different file formats like Avro, parquet, ORC, JSON, and XML.
· Expertise in Creating, Debugging, Scheduling, and Monitoring jobs using Airflow and Oozie.
· Experienced with using most common Operators in Airflow - Python Operator, Bash Operator, Google Cloud Storage Download Operator, and Google Cloud Storage Object Sensor.
· Hands-on experience in handling database issues and connections with SQL and NoSQL databases such as MongoDB, HBase, Cassandra, SQL Server, and PostgreSQL.
· Created Java apps to handle data in MongoDB and HBase. Used Phoenix to create SQL layer on HBase.
· Experience in designing and creating RDBMS Tables, Views, User Created Data Types, Indexes, Stored Procedures, Cursors, Triggers, and Transactions.
· Expert in designing ETL data flows using creating mappings/workflows to extract data from SQL Server and Data Migration and Transformation from Oracle/Access/Excel Sheets using SQL Server SSIS.
· Expert in designing Parallel jobs using various stages like Join, Merge, Lookup, remove duplicates, Filter, Dataset, Lookup file set, Complex flat file, Modify, Aggregator, and XML.
· Hands-on experience with Amazon EC2, Amazon S3, Amazon RDS, VPC, IAM, Amazon Elastic Load Balancing, Auto Scaling, CloudWatch, SNS, SES, SQS, Lambda, EMR, and other services of the AWS family.
· Created and configured a new batch job in the Denodo scheduler with email notification capabilities Implemented Cluster setting for multiple Denodo nodes and created load balance for improving performance activity.
· Instantiated, created, and maintained CI/CD (continuous integration & deployment) pipelines and apply automation to environments and applications.
· Worked on various automation tools like GIT, Terraform, and Ansible. Experienced in fact dimensional modeling (Star schema, Snowflake schema), transactional modeling, and SCD (Slowly changing dimension)
· Experienced with JSON-based RESTful web services, and XML/QML-based SOAP web services and worked on various applications using python integrated IDEs like Sublime Text and PyCharm
· Efficient Cloud Engineer with years of experience assembling cloud infrastructure. Utilizes strong managerial skills by negotiating with vendors and coordinating tasks with other IT team members. Implements best practices to create cloud functions, applications, and databases.
