Summary
Overview
Work History
Education
Skills
Certification
Activities
Timeline
Generic

Godfrey Agbro

Owings Mills,MD

Summary

Accomplished Senior Data Engineer ( Team Lead ) with a proven track record, specializing in Python and data pipeline optimization. Expertly reduced delivery timelines by 75% through innovative solutions and collaboration. Adept at enhancing data quality and fostering team success, driving impactful analytics and strategic decision-making.

Overview

1
1
Certification
11
11
years of professional experience

Work History

Senior Data Engineer

Interos
12.2024 - Current
  • Deployed and managed pipelines on a modern cloud stack (AWS, Snowflake, Databricks) with robust observability tools (Datadog, Monte Carlo) for end-to-end reliability.
  • Refactored central repository to integrate agentic workflows and AI-assisted tooling for pipeline development and data quality enforcement, reducing delivery timelines by ~75% (from 2 weeks to 3–4 days).
  • Engineered custom data pipelines integrating external and internal datasets, enhancing business analytics and impact tracking.
  • Developed a real-time risk scoring engine utilizing news feeds, predictive modeling, and clustering algorithms to ensure accurate, deduplicated metrics.
  • Designed a scalable entity resolution framework to seamlessly match and merge new records while preserving data integrity and eliminating duplicates.
  • Created modular, reusable ingestion framework standardizing third-party data onboarding, accelerating integration timelines.

Senior Data Engineer

Whatifmedia Group
11.2023 - 12.2024
  • Partnered with DevOps to establish Airflow on AWS EKS across multiple environments, facilitating test-driven development and advancing previously backlogged data pipeline projects.
  • Developed a batch ETL application in Python, orchestrated with Airflow, to replicate data from app team's PostgreSQL RDS instances to Snowflake. This solution allowed the team to deactivate unreliable AWS DMS replication tasks, mitigating issues with replication lag.
  • Revamped a PHP/cron data pipeline for revenue reporting, converting it to a Python application with Airflow for scheduling and orchestration. This overhaul enhanced reporting accuracy by 67% and streamlined the architecture, eliminating redundant technologies.
  • Collaborated with DevOps to replace existing CDC system with Debezium and MSK Kafka, improving data replication reliability.
  • Conducted research and proposed initiatives for data standardization, data contracts, and a centralized data catalog system, improving health of organization’s data ecosystem.
  • Led team of two developers, ensuring project success and enhancing collaboration.

Senior Data Engineer

Omadahealth
08.2022 - 11.2023
  • Led design and implementation of asynchronous file ingestion process for customer data from SFTP server, adapting incoming files to business rules using MSK Kafka, Pandas, and AWS Lambda.
  • Managed DevOps processes for data engineering projects, utilizing infrastructure-as-code principles with Terraform.
  • Implemented data quality checks and validation routines, enhancing data accuracy, consistency, and compliance with industry standards.
  • Collaborated closely with development and data engineering to define and implement standardized naming conventions for Kafka topics, ensuring seamless data integration and streamlined operations.
  • Designed and executed proof-of-concept projects, demonstrating AWS AppFlow's capabilities in secure data extraction, transformation, and loading into target destinations.
  • Collaborated with AWS technical support to address challenges and questions, gaining valuable insights and demonstrating a proactive approach to problem-solving.

Data Engineer

Omadahealth
12.2021 - 08.2022
  • Design and implement a serverless system using lambda, s3, and redshift to process member insights from meal records and physical activity that triggers moments in upstream application used for developing interventions for health coaches to advise members to take action on their health.
  • Created and maintained Airflow DAGs to extract and transform data assets from upstream source app database and external source API, ensuring reliable data flow for analytics.
  • Established a data lake for data assets extracted from upstream applications that will be stored and transformed using AWS resources like S3, Spectrum and Redshift.
  • Collaborated with data architects to develop validation procedures for upstream app data extraction and assessed data quality frameworks, data catalogs, and lineage tools to enhance data governance.
  • Optimize tableau reports generation time from ~8hr to ~2hrs that is interfacing with in-house reporting app built in ruby rails.

Data Engineer

Asurion
11.2020 - 12.2021
  • Developed custom Python scripts to automate data ingestion, transformation, and validation, resulting in a 75% reduction in manual effort.
  • Wrote and optimized complex SQL queries and stored procedures with multiple joins and window functions for data manipulation and merging from large volumes of historical data in SQL Server, validating ETL/ELT processed data in target database.
  • Design Dimensional model with SCD Type 1, 2 and Snapshot dimension tables with their related fact tables to support reporting and analytics that enforces data integrity and optimized data retrieval.
  • Designed and implemented HR Data Warehouse consolidating HR data to enable self-service reporting and analytics, saving HR analytics team an average of 1 week in manual data gathering before analysis.
  • Created CRUD application in Power App to streamline data collection and manipulation in the data warehouse.
  • Work with executive leadership team to strategize data architecture and next step plans for data enhancements.

Big Data Engineer

Charles Schwab
Westlake, USA
04.2019 - 11.2020
  • Designed and developed data pipeline that ingests data from a wide variety of source systems including Oracle, SQL server, Teradata and NAS into a data lake in MapR Hadoop eco-system using Talend, Bash scripts, Python scripts and Control-M.
  • Assisted the risk team with their reporting needs by creating reporting views in Hive and Drill data warehouse that have resulted in the reduction of 92% of joins in reporting queries replacing approximately 12 joins with 1 join with newly implemented process.
  • Utilized data modeling techniques, including Dimensional, Star Schema, Snowflake modeling, and Slowly Changing Dimensions (SCD Type 1, Type 2, Hybrid), to enhance data structure.
  • Designed data models for data mart using Erwin and Visio, clearly representing data grain, dimension, and fact tables.
  • Understood Hadoop architecture and core components, facilitating effective use of Name Node, Data Node, Resource Manager, Node Manager, and other distributed components.
  • Expertise in the Analysis, Design and Development of Software Applications and providing Data Integration/Warehousing solutions using Kimball Methodologies.
  • Mentored teammates on processes and standards, enhancing team knowledge and adherence to best practices.

Technical Intern /Associate Data Engineer

Fidelity Investment
Westlake, USA
05.2015 - 04.2019
  • Develop packages, stored procedures and functions that are utilized by restful web services, feeds and chat bot application.
  • Created, maintained, and deployed data warehouse objects for use by various downstream applications.
  • Enhanced response time of stored procedure queries by optimizing indexes, partitioning tables, and refining query structure.
  • Created Lambda functions using Boto3 for Python to split files greater than 50MB, process the files and archive unprocessed files.
  • Produced Hive table in Athena in AWS to query data in AWS S3 bucket.
  • Designed process that utilized AWS resources for moving on-premises application (AWB) to AWS.
  • Developed proof of concept for ELK stack to analyze long running SQL queries from Oracle database Active Session History.
  • Hands on experience with informatica for ETL of structured and unstructured data.

Education

Bachelor of Science - Information Systems

University of Texas at Arlington
01-2017

Skills

  • Data Integration
  • Data Modeling
  • Data Analytics
  • Data Querying
  • SQL Developer
  • PostgreSQL
  • Snowflake
  • Redshift
  • Databricks
  • Oracle 12c
  • T-SQL
  • MySQL
  • SQL Server
  • Teradata
  • MapR-Hadoop
  • Talend
  • Informatica
  • Airflow
  • CI/CD Pipelines
  • Process Automation
  • Custom Applications
  • Python
  • Java
  • Bash
  • PowerShell
  • AWS - Lambda
  • S3
  • RDS
  • Terraform
  • Event Streaming
  • Drill
  • Power Bi
  • AI Integration
  • Data Querying
  • AI Integration

Certification

  • AWS (Developer)
  • Teradata (Developer)

Activities

  • Association of Information Technology Professionals (AITP), Member
  • YMCA, Volunteer Soccer Coach, 02/01/23, Current, Mentored and managed soccer team of 6-10-year-olds., Planned practice times and built formations.

Timeline

Senior Data Engineer

Interos
12.2024 - Current

Senior Data Engineer

Whatifmedia Group
11.2023 - 12.2024

Senior Data Engineer

Omadahealth
08.2022 - 11.2023

Data Engineer

Omadahealth
12.2021 - 08.2022

Data Engineer

Asurion
11.2020 - 12.2021

Big Data Engineer

Charles Schwab
04.2019 - 11.2020

Technical Intern /Associate Data Engineer

Fidelity Investment
05.2015 - 04.2019

Bachelor of Science - Information Systems

University of Texas at Arlington
Godfrey Agbro