Summary
Overview
Work History
Education
Skills
Languages
Timeline
Generic

Anthony Wafula

Austin,USA

Summary

Software and Data Engineer with 3+ years of experience building scalable, distributed data platforms and ETL pipelines in cloud environments. Proficient in Apache Spark, Python, Scala, SQL, AWS, Airflow, Hive, and Trino, with experience processing over 10 billion records to power marketing analytics, customer segmentation, and business intelligence. Strong background in data quality, monitoring, and cross-functional collaboration to deliver reliable, production-ready data solutions.

Overview

1
1
Language
4
4
years of professional experience

Work History

Data Engineer II

Expedia Group
Austin, USA
08.2022 - Current
  • Build and maintain distributed data processing systems and distributed systems using Apache Spark, handling 10B+ records to power customer segmentation and marketing analytics products.
  • Design and implement scalable Spark-based ETL pipelines for data ingestion, data transformation, and delivery in production using Scala, Python, and Java.
  • Manage Apache Airflow workflows, optimizing DAG scheduling, dependency management, retries, and failure recovery to improve pipeline reliability and efficiency.
  • Implemented data modeling and data governance to ensure trusted sources, data lineage, and RBAC-secured access for enhanced data integrity.
  • Collaborated with stakeholders and ML/data teams to refine requirements, delivering reliable data solutions and curating clean datasets for warehousing and analytics that enabled informed decision-making.
  • Developed Tableau and Datadog dashboards to visualize pipeline health, system performance, and key business metrics for proactive monitoring.
  • Oversaw full software development lifecycle, managing architecture, API design, development, testing, deployment, and production support.
  • Created technical documentation and design specifications for ETL pipelines, workflow architecture, API design, and production support procedures.
  • Created runbooks for on-call responders detailing PagerDuty escalation steps and common remediation commands to streamline production support.
  • Leveraged AI-assisted tools: GitHub Copilot agents, ChatGPT, Codex, and Claude to accelerate development while maintaining production code quality.
  • Worked in an Agile (Scrum) environment, collaborating with cross-functional teams and using Jira to manage sprint planning, backlog refinement, task tracking, and feature delivery.

Education

BACHELOR'S OF ARTS - COMPUTER & INFORMATION SCIENCE

Berea College
Berea, Kentucky
05-2022

Skills

  • Data processing: Apache Spark, Spark SQL, Hive, Iceberg, Trino (Presto)
  • ETL development
  • Data warehousing
  • Data modeling
  • Workflow orchestration: Apache Airflow
  • Programming languages: Python, Scala, Java
  • Cloud services: AWS (Amazon S3, EMR)
  • Monitoring tools: Datadog, PagerDuty, Tableau
  • Development tools: GitHub, Docker, Git, GitLab CI/CD and Jenkins
  • Distributed systems
  • Critical thinking
  • AI tools: GitHub Copilot, ChatGPT, Codex, Claude
  • Agile (Scrum)
  • Jira
  • Cross-functional teamwork
  • Production support

Languages

English
Native/ Bilingual

Timeline

Data Engineer II

Expedia Group
08.2022 - Current

BACHELOR'S OF ARTS - COMPUTER & INFORMATION SCIENCE

Berea College
Anthony Wafula