Summary
Overview
Work History
Education
Skills
Timeline
Generic

Rois Ahmed

New Rochelle,NY

Summary

Data Engineer with 6+ years of IT experience working in a variety of environments with a vast number of tools and technologies, that included but were not limited to the programming languages of Python, Java and SQL. Worked as a Data Engineer with big data technologies including Hadoop, Spark, and cloud technologies with cross platform integration experience.

Overview

6
6
years of professional experience

Work History

Data Engineer

BNY Mellon
New York, New York
11.2024 - Current
  • Evaluated Hadoop requirements, designed, and deployed high-availability big data clusters.
  • Administered and managed Hadoop infrastructure implementation, ensuring operational stability and performance.
  • Collaborated with application teams to install OSS updates and manage Hadoop patches, enhancing system reliability.
  • Developed Pig programs for loading and filtering streaming data into HDFS using Flume.
  • Constructed HBase data models on HDFS for real-time analytics utilizing Java API.
  • Implemented secondary sorting in MapReduce to enhance reducer output organization.
  • Maintained and troubleshot UNIX and Linux environments, resolving issues to support seamless operations.
  • Analyzed system security threats and implemented safeguards against vulnerabilities.
  • Developed data pipelines for processing large datasets efficiently.
  • Collaborated with cross-functional teams to understand data requirements.
  • Designed and maintained database systems for accurate data storage.
  • Ensured data quality by implementing validation checks and monitoring processes.

Big Data Engineer

Healthfirst
New York, New York
03.2022 - 10.2024
  • Analyzed the impact on server performance CPU usage, server memory usage for the applications of varied numbers of multiple, simultaneous users.
  • Used Spark-Streaming APIs to perform necessary transformations and actions on the fly for building the common learner data model which gets the data from Kafka in near real time and Persists into Cassandra.
  • Developed Spark code using Scala and Spark-SQL/Streaming for faster testing and processing of data.
  • Worked on Sequence files, Map side joins, Bucketing, Static and Dynamic Partitioning for Hive performance enhancement and storage improvement.
  • Ingested data from RDBMS to Hive for transformations, then exported transformed data to Cassandra for enhanced data access and analysis.
  • Created Job management using Fair scheduler and Developed job processing scripts using Oozie workflow.
  • Implemented Spark Scripts using Scala, Spark SQL to access hive tables into Spark for faster processing of data.
  • Configured Spark Streaming with Kafka Streams to efficiently retrieve and store information in HDFS.
  • Handled large datasets by utilizing partitions, broadcasts, effective joins, and transformations during the ingestion process.
  • Developed Sqoop scripts to facilitate data import and export between RDBMS and HDFS.
  • Executed many performance tests using the Cassandra-stress tool to measure and improve the read and write performance of the cluster.
  • Configured various workflows to run on top of Hadoop using Oozie and these workflows comprise heterogeneous jobs like Pig, Hive, Sqoop and MapReduce.
  • Written Spark applications using Scala to interact with the MySQL database using Spark SQL Context and accessed Hive tables using Hive Context.
  • Developed Spark batch job to automate creation/metadata update of external Hive table created on top of datasets residing in HDFS.
  • Performed the migration of Hive and MapReduce Jobs from on-premise MapR to AWS cloud using EMR.
  • Optimized Hive queries and used Hive on top of Spark engine.

Data Engineer

Lyft
New York, New York
05.2020 - 02.2022
  • Owned complete application design with emphasis on Java and Hadoop integration for improved data processing.
  • Managed structured, semi-structured, and unstructured data through thorough Hadoop log file reviews, ensuring data integrity.
  • Extracted data from mainframes into Kafka for ingestion into HBase, facilitating analytics operations.
  • Created MapReduce jobs to extract HBase content and configured Oozie workflows for analytical report generation.
  • Implemented Spark RDD transformations for business analysis, enhancing overall data processing efficiency.
  • Assisted architect in analyzing existing systems, preparing design blueprints and application flow documentation to support development.
  • Participated in client business meetings to gather security requirements beyond standard requirement gathering.
  • Designed and developed prototype screens for search engine report analysis application using AngularJS and Bootstrap.

Education

Associate of Arts - Computer Science

Westchester Community College
Valhalla, NY
05-2026

Skills

  • Data analysis and visualization
  • Database design and management
  • Big data technologies and processing
  • Cloud computing solutions
  • Project management methodologies
  • Programming languages proficiency
  • Cross-functional collaboration skills
  • SQL expertise and optimization
  • Data migration strategies
  • Data pipeline control and design
  • Real-time analytics implementation
  • Data integration techniques
  • Hadoop ecosystem knowledge
  • Scripting languages proficiency
  • ETL development processes
  • Performance tuning strategies
  • Data curation practices
  • Data warehousing solutions
  • Spark framework utilization
  • SQL transactional replication techniques
  • RDBMS expertise
  • Data cleaning methodologies
  • Agile project management practices

Timeline

Data Engineer

BNY Mellon
11.2024 - Current

Big Data Engineer

Healthfirst
03.2022 - 10.2024

Data Engineer

Lyft
05.2020 - 02.2022

Associate of Arts - Computer Science

Westchester Community College
Rois Ahmed