Summary
Overview
Work History
Education
Skills
Websites
Timeline
Generic

Mihail Kavrakov

Senior Data Engineer
Štip

Summary

Software-engineering minded Senior Data Engineer with 7 years of experience building and running production data pipelines across GCP, AWS, and Azure. I care about fundamentals: clean architecture, sensible patterns, and code that is easy to test and maintain. I have worked with distributed batch/stream processing and I am comfortable owning delivery end-to-end, from Terraform and CI/CD to production ops and monitoring. Longer term, I want to move toward retrieval-focused data products and RAG, with an emphasis on data quality and evaluation.

Overview

7
7
years of professional experience

Work History

Senior Data Engineer

RTL Deutschland
09.2022 - Current
  • Built and maintained large-scale batch ETL pipelines on GCP Dataproc (orchestrated with Airflow on GKE) using PySpark, producing curated datasets in Parquet/Avro for downstream consumption.
  • Cut runtime of a heavy PySpark workload by introducing Spark + Scala (UDAF), reducing execution time by ~5-6 times.
  • Maintained and extended the core tracking engine (Core API) and the streaming ingestion path Kafka -> Dataflow -> partitioned GCS) that feeds downstream processing.
  • Helped shape platform architecture and handled provisioning/access with Terraform; managed secrets in Vault.
  • Maintained solid test coverage (unit + integration) for core logic and data contracts (pytest), enforced in GitLab CI/CD.
  • Automated deliveries to Decentriq using the SDK: created/validated datasets and datalabs and provisioned DCRs (Data Clean Rooms) as scheduled workflows.

Senior Data Engineer

BlueprintData
04.2025 - Current
  • Owned end-to-end delivery of a configurable ETL platform: architecture, repo structure, CI/CD workflows, multi-environment releases, and production operations.
  • Built a config-driven runtime ETL application that ingests from Azure Blob and applies per-source parsing/loading rules (data types, validations, and load modes like append/upsert/replace/truncate).
  • Implemented extensible parsers for CSV/XLSX/XLS/JSON, with a clean structure for adding new formats and transformations as new sources onboard.
  • Automated Azure provisioning with Terraform across multiple resource groups, subscriptions, and environments, including RBAC and Key Vault secrets.
  • Centralized shared Terraform templates and reusable build assets; published standardized test images to GitHub Container Registry (GHCR) for CI and integration testing.

Data Engineer

Lykke AG
02.2024 - 10.2024
  • Worked closely with the Quants team to deploy and run production prop trading and market-making algorithms, keeping releases stable and repeatable.
  • Improved code structure and delivery workflows (automation around builds/deployments and run pipelines) to reduce manual steps and make changes safer to ship.
  • Designed and maintained ETL pipelines that extracted trading data from InfluxDB, transformed it, and loaded it back to support continuous monitoring.
  • Built monitoring views in Grafana and automated daily PDF trade-performance reports delivered to stakeholders.

Data Engineer

Torrance Analytics
01.2022 - 02.2024
  • Built an end-to-end ingestion app to pull ticketing data from external sources, transform it in Python, and upsert into PostgreSQL.
  • Deployed on AWS using Terraform to provision required resources and orchestrated extraction + transformation with Lambda and Step Functions.
  • Improved table design and upsert behavior in PostgreSQL and kept ingestion stable with automated unit tests.
  • Consumed and processed messages from SQS; ran workers on EC2 (pm2) and added alerting via AWS SES.

Data Mining Engineer

Valuer AI
09.2021 - 09.2022
  • Built and maintained a reusable web-mining framework with clear abstractions (base classes + shared components) so new sources and transformations could be added quickly.
  • Delivered scrapers using Selenium with anti-bot handling (proxies/VPNs); stored staged results in MongoDB for downstream processing and re-runs.
  • Containerized workloads with Docker, deployed to Kubernetes, and added unit tests with pytest.

Data Engineer / Analyst

Sample Solutions BV
03.2019 - 09.2021
  • Built and maintained ETL pipelines enriching an existing database: pulled external data and loaded it into staging MongoDB for validation and enrichment (Companies + Contacts) data.
  • Added unit tests and practical data quality monitoring to catch inconsistencies early and keep the database accurate over time.

Education

Faculty of Computer Engineering And Technology, University of Goce Delchev
Shtip, North Macedonia
06-2019

Skills

  • Languages: Python, SQL, Scala
  • Processing & Orchestration: Airflow, PySpark, Dataproc, Dataflow, Kafka
  • GCP: BigQuery, GCS, Pub/Sub, Composer, Dataflow, Dataproc
  • AWS: Lambda, SQS, Step Functions, S3, EC2, SES, CloudWatch
  • Azure: Storage Accounts/Blob Containers, Container Apps Jobs, Container Registry (ACR), Key Vault, Key Vault Secrets
  • Databases: SQL Server, PostgreSQL, InfluxDB, MongoDB
  • Infra/DevOps: Terraform, Docker, GitHub Actions, GitLab CI, Vault

Timeline

Senior Data Engineer

BlueprintData
04.2025 - Current

Data Engineer

Lykke AG
02.2024 - 10.2024

Senior Data Engineer

RTL Deutschland
09.2022 - Current

Data Engineer

Torrance Analytics
01.2022 - 02.2024

Data Mining Engineer

Valuer AI
09.2021 - 09.2022

Data Engineer / Analyst

Sample Solutions BV
03.2019 - 09.2021

Faculty of Computer Engineering And Technology, University of Goce Delchev
Mihail KavrakovSenior Data Engineer