Summary
Overview
Work History
Education
Skills
Timeline
Generic

Sonika Suggala

Summary

  • Data Engineer with 7+ years of experience building scalable, reliable data platforms across banking, healthcare, and retail analytics
  • Delivered end-to-end ETL/ELT pipelines and curated data products used for reporting, reconciliation, and operational decision-making
  • Strong in Python and SQL with production practices including reusable code, logging, testing, and performance tuning
  • Built distributed processing using Spark and PySpark (Spark SQL) for high-volume batch and near real-time workloads
  • Implemented lakehouse patterns using Delta Lake and Apache Iceberg to support schema evolution and reliable incremental processing
  • Built event-driven ingestion using Kafka for freshness-sensitive pipelines and downstream consumption
  • Automated workflows using Apache Airflow, Azure Data Factory, and Dagster with retries, SLAs, backfills, and dependency management
  • Standardized transformations using dbt (models, tests, snapshots) and dimensional modeling to deliver analytics-ready marts
  • Improved data trust using Great Expectations, Deequ, and Soda, plus lineage and governance with Unity Catalog, OpenLineage, DataHub, and data contracts
  • Strengthened reliability and delivery using Monte Carlo observability, Datadog/CloudWatch monitoring, and CI/CD with Jenkins/GitHub Actions/Azure DevOps

Overview

8
8
years of professional experience

Work History

Data Engineer

WELLS FARGO
Dallas, Texas, USA
01.2024 - Current
  • Built banking data pipelines (transactions, accounts, customers, reference) into curated datasets supporting risk, compliance, and regulatory reporting
  • Processed ~1.5 TB/day (~150M+ records/day) to deliver reliable daily and intra-day refreshes for downstream consumers
  • Developed Spark/PySpark (Spark SQL) transformations for standardization, deduplication, enrichment, and incremental loads
  • Orchestrated workflows in Apache Airflow with retries, SLAs, parameterized backfills, and production alerting
  • Managed lakehouse storage on AWS S3 using Apache Iceberg for schema evolution, partitioning, and reliable incremental processing
  • Implemented Deequ validation gates (schema, nulls, duplicates, referential integrity, thresholds), reducing recurring data issues by ~35%
  • Defined data contracts for critical datasets to standardize schema/quality expectations and prevent breaking downstream consumers
  • Captured lineage using OpenLineage and published metadata into DataHub to improve traceability, ownership, and faster incident triage
  • Implemented data observability using Monte Carlo with CloudWatch dashboards/alerts for freshness, volume drift, and SLA breaches
  • Automated releases using Jenkins and Terraform, and published certified datasets into Microsoft Fabric/OneLake for self-serve analytics

Data engineer

OPTUM
Hyderabad, Telangana, India
01.2021 - 01.2023
  • Built healthcare data products integrating scheduling, encounters, claims, and member/provider reference data to support care-gap closure workflows
  • Processed ~0.6 TB/day (~70M+ records/day) and delivered refreshed outreach worklists for care management and operations
  • Streamed appointment updates using Kafka to keep follow-up indicators current as schedules changed (reschedules, cancellations, no-shows)
  • Delivered analytics-ready datasets into BigQuery to enable fast self-serve reporting and operational worklists
  • Modeled transformation layers using dbt (models, tests, snapshots) to standardize business logic and publish curated marts
  • Orchestrated modular pipelines using Dagster with asset-driven dependencies to simplify reruns, backfills, and environment promotion
  • Implemented data quality gates using Soda (freshness, schema, duplicates, referential integrity, thresholds) to reduce noisy outputs
  • Optimized pipeline performance and query efficiency using partitioning/clustering strategies, cutting runtime from ~3 hours to ~40 minutes (~78% reduction)
  • Automated CI/CD using GitHub Actions and managed deployments using Terraform for repeatable releases across environments
  • Strengthened production monitoring using Datadog alerts tied to SLAs and dataset freshness (MTTR ~30 minutes)

Associate Data Engineer

TARGET
Hyderabad, Telangana, India
01.2018 - 01.2021
  • Engineered Databricks (PySpark/Spark SQL) pipelines to unify POS sales/voids, ASN receipts, transfers, returns/RTV, and cycle counts into an inventory-movement ledger powering daily on-hand reporting
  • Landed raw feeds into ADLS Gen2 and standardized them into Delta Lake (Parquet) tables partitioned by business_date/store for scalable processing and replay
  • Built transformation layers in Databricks to generate movement facts and daily on-hand snapshots with transaction-level traceability for audit and finance tie-outs
  • Orchestrated daily close workflows using Azure Data Factory with dependency control, retries, SLAs, and parameterized reruns for late-arriving store/DC loads
  • Implemented reconciliation logic in Spark SQL to compare computed on-hand vs snapshot/physical counts and produced exception outputs with variance categories (receiving gaps, transfer timing, shrink, mis-scan)
  • Standardized keys and hierarchies (UPC→SKU mapping, store/DC identifiers, dept/class/subclass) to enforce consistent grain and prevent double counting
  • Implemented data validation using Great Expectations (schema, nulls, duplicates, negative on-hand thresholds, spike detection) with quarantine and controlled reprocessing
  • Published curated marts to Azure Synapse Analytics (SQL DW) for inventory accuracy scorecards, shrink/variance trends, and receiving performance KPIs
  • Enabled ad-hoc investigation using Presto/Trino over curated lake tables to speed root-cause analysis of reconciliation exceptions
  • Delivered operational dashboards in Power BI and Tableau for inventory accuracy, daily-close health, and variance trending
  • Stabilized performance by optimizing partition strategy, controlling shuffle, and tuning joins to support high SKU/store volumes
  • Managed governance and dataset certification using Unity Catalog and releases via Azure DevOps (repos + pipelines) with environment configs (dev/test/prod)

Education

Master of Science - Data Science

The University of Texas At Arlington
Arlington, Texas, TX

Skills

Programming: Python, SQL
Data Processing: Apache Spark, PySpark, Spark SQL, Databricks
Orchestration: Apache Airflow, Azure Data Factory, Dagster
Streaming: Kafka (Confluent)
Lakehouse & Storage: AWS S3, ADLS Gen2, Delta Lake, Apache Iceberg, Parquet, Hive Metastore
Warehouses & Query: Snowflake, BigQuery, Azure Synapse Analytics, Amazon Redshift, Presto, Trino, Microsoft Fabric, OneLake
Data Quality, Governance & Lineage: Great Expectations, Deequ, Soda, Unity Catalog, OpenLineage, DataHub, Data Contracts, RBAC
DevOps, Observability & BI: Git, Azure DevOps, Jenkins, GitHub Actions, Terraform, Monte Carlo, Datadog, CloudWatch, Splunk, Power BI, Tableau

Timeline

Data Engineer

WELLS FARGO
01.2024 - Current

Data engineer

OPTUM
01.2021 - 01.2023

Associate Data Engineer

TARGET
01.2018 - 01.2021

Master of Science - Data Science

The University of Texas At Arlington
Sonika Suggala