- Lead Data Engineer within Advanced Analytics/Data Science, owning architecture, delivery and production support end-to-end for enterprise telecom workloads.
- Designed production data-pipeline architectures for various clients across the US, Europe and Asia regions, handling terabytes of data and billions of records with horizontal scalability and reliability.
- Partnered with product, engineering and operations stakeholders to translate ambiguous business problems into scalable data/AI solutions, leading discovery, architecture, rapid POCs, production implementation and operational hand-off.
- Re-architected a real-time text/image event pipeline for customer highlights and flashbacks, delivering 40% faster processing while improving scalability, flexibility and operating efficiency; replaced a third-party service and saved multi-million dollars annually.
- Designed and implemented LLM-backed services directly in the production highlights/flashbacks pipeline: personalized title generation from theme/image metadata via PySpark UDFs; multilingual translation into Spanish, Japanese, Portuguese, French and German; and vision-model image tagging.
- Built self-hosted multi-model platform using open-source LLMs on on-prem and AWS GPU infrastructure; selected hosted vs. self-hosted approaches based on workload quality, latency, cost, and data privacy.
- Building AI-driven self-healing Airflow operations on Kubernetes: locally hosted models triage DAG failures, distinguish infrastructure/resource failures from code/data errors, automatically rerun recoverable failures from the last successful step, and send root-cause context to on-call engineers.
- Rolled out AI agents for development and testing, including automated first-pass code review on every Git commit against codified standards for formatting, syntax, code structure and data-writing functions; expanded AI-assisted development and test generation across the team.
- Built an NLP pipeline that reduced analysis effort from days to minutes with higher accuracy, replaced a third-party vendor and saved approximately $1M annually, enabling customer-experience and marketing teams to act on top issues faster.
- Developed churn and customer-retention pipelines/models that reduced AT&T zero-usage customers by approximately 20% and Virgin Mobile Latin America churn by approximately 13%.
- Contributed to applied marketing intelligence for offer optimization and response-rate optimization, supporting campaign lifts of 175% for telemarketing and 200% for direct mail.
- Built real-time IoT pipelines and dashboards tracking device runtime, energy use, temperature and humidity, contributing to approximately 3% energy savings for Rackspace and 5% for Yumi Ice Creams.
- Led migration of production applications from Hadoop-MapR to Kubernetes/AWS/Airflow, improving scalability and reducing operating cost.
- Created enriched datasets, scoring logic, and fact/dimension tables for millions of customers, enabling application UIs and Tableau dashboards for campaign optimization.
Environment: Python, PySpark, SQL, Spark, Hadoop/MapR, Kubernetes, Docker, Airflow, AWS, LLMs/GenAI, Prompt Engineering, GPU Inference, Git/Jira