Summary
Overview
Work History
Education
Skills
Accomplishments
Certification
Timeline
AI / GENAI CAPABILITY SNAPSHOT
Generic

Shilpa Dantuluri

Ashburn,VA

Summary

Accomplished AI & Data Engineering Lead with 19 years of experience in developing large-scale data platforms and advanced analytics for enterprise solutions. Specializes in Generative AI and LLM applications, with hands-on expertise in integrating models into production pipelines and optimizing performance across various metrics. Strong foundation in Python, PySpark, SQL, and cloud technologies, driving impactful data-driven solutions and operational efficiencies.

Overview

1
1
Certification
19
19
years of professional experience

Work History

Lead Data Engineer – Advanced Analytics, Data Science & AI / GenAI

Synchronoss Technologies
Reston, VA
12.2015 - Current
  • Lead Data Engineer within Advanced Analytics/Data Science, owning architecture, delivery and production support end-to-end for enterprise telecom workloads.
  • Designed production data-pipeline architectures for various clients across the US, Europe and Asia regions, handling terabytes of data and billions of records with horizontal scalability and reliability.
  • Partnered with product, engineering and operations stakeholders to translate ambiguous business problems into scalable data/AI solutions, leading discovery, architecture, rapid POCs, production implementation and operational hand-off.
  • Re-architected a real-time text/image event pipeline for customer highlights and flashbacks, delivering 40% faster processing while improving scalability, flexibility and operating efficiency; replaced a third-party service and saved multi-million dollars annually.
  • Designed and implemented LLM-backed services directly in the production highlights/flashbacks pipeline: personalized title generation from theme/image metadata via PySpark UDFs; multilingual translation into Spanish, Japanese, Portuguese, French and German; and vision-model image tagging.
  • Built self-hosted multi-model platform using open-source LLMs on on-prem and AWS GPU infrastructure; selected hosted vs. self-hosted approaches based on workload quality, latency, cost, and data privacy.
  • Building AI-driven self-healing Airflow operations on Kubernetes: locally hosted models triage DAG failures, distinguish infrastructure/resource failures from code/data errors, automatically rerun recoverable failures from the last successful step, and send root-cause context to on-call engineers.
  • Rolled out AI agents for development and testing, including automated first-pass code review on every Git commit against codified standards for formatting, syntax, code structure and data-writing functions; expanded AI-assisted development and test generation across the team.
  • Built an NLP pipeline that reduced analysis effort from days to minutes with higher accuracy, replaced a third-party vendor and saved approximately $1M annually, enabling customer-experience and marketing teams to act on top issues faster.
  • Developed churn and customer-retention pipelines/models that reduced AT&T zero-usage customers by approximately 20% and Virgin Mobile Latin America churn by approximately 13%.
  • Contributed to applied marketing intelligence for offer optimization and response-rate optimization, supporting campaign lifts of 175% for telemarketing and 200% for direct mail.
  • Built real-time IoT pipelines and dashboards tracking device runtime, energy use, temperature and humidity, contributing to approximately 3% energy savings for Rackspace and 5% for Yumi Ice Creams.
  • Led migration of production applications from Hadoop-MapR to Kubernetes/AWS/Airflow, improving scalability and reducing operating cost.
  • Created enriched datasets, scoring logic, and fact/dimension tables for millions of customers, enabling application UIs and Tableau dashboards for campaign optimization.

Environment: Python, PySpark, SQL, Spark, Hadoop/MapR, Kubernetes, Docker, Airflow, AWS, LLMs/GenAI, Prompt Engineering, GPU Inference, Git/Jira

Data Analyst / Application Support Lead

Aetna – Active Health Management
Chantilly, VA
04.2012 - 11.2015
  • Analyzed large volumes of real-time patient, member and claims data across ODS, Datamart, CareEngine and Active Advice using complex SQL/PLSQL to isolate data issues, dependencies and business impact.
  • Performed deep-dive analysis on recurring production inquiries; identified upstream/downstream patterns and designed validation and reconciliation scenarios that closed source-to-target gaps and reduced repeat incidents.
  • Conducted deep-dive analysis on recurring production inquiries, identifying upstream/downstream patterns and designing validation and reconciliation scenarios that closed source-to-target gaps.
  • Optimized SQL and PL/SQL stored procedures/functions while monitoring high-volume batch/reporting jobs to ensure performance and reliability.
  • Developed and automated business/customer insight reporting using SQL Server, Oracle, Power BI, and SSIS; standardized recurring requests into repeatable reports and contributed to analytics warehouse/ODS design and data quality/performance testing.

Environment: Oracle, SQL Server, Informatica, SQL, PL/SQL, UNIX, Power BI, SSIS, Quality Center, ServiceNow, BMC Remedy, TFS

Enterprise Testing & Metrics Reporting Lead

Aetna
Hartford, CT
03.2011 - 03.2012
  • Automated management reporting with Excel/VBScript and SSRS, scheduling daily, weekly, monthly, and quarterly reports for senior management, streamlining reporting processes.
  • Environment: SSRS, Excel/VBScript, SQL/Reporting, enterprise metrics
  • Self-motivated, with a strong sense of personal responsibility.
  • Worked effectively in fast-paced environments.
  • Skilled at working independently and collaboratively in a team environment.

Claims Tracking System – Data Analyst (Claims Data & ETL)

Aetna
Hartford, CT
01.2010 - 03.2011
  • Analyzed large claims and overpayment datasets using SQL across Oracle, DB2, and SQL Server to identify overpaid claims, recovery candidates, and data-quality issues, enhancing data accuracy and recovery strategies.
  • Traced claims lineage across legacy sources and target warehouse; validated Informatica mappings/transformations and reconciled record counts and financial totals, ensuring data integrity and compliance.
  • Supported ETL testing/UAT/production validation, including claims test datasets and flat-file testing, ensuring accurate ETL monitoring and comprehensive translation of requirements to mappings.

Environment: Informatica 8.x, SQL Server 2008, DB2, Oracle log, SQL, PL/SQL, HP Quality Center, IBM Rational ClearCase/ClearQuest

Member Individualization – Developer

Aetna
Chennai, India
11.2007 - 12.2009
  • Developed Member Individualization web application and Clerical Tool search module to enhance user experience and streamline member data retrieval.

Environment: IBM Mainframes – MVS/OS390, DB2, COBOL, JCL, CICS, REXX, VSAM, TSO/ISPF

Education

B.Tech. / Bachelor's Degree - Civil Engineering

Jawaharlal Nehru Technological University
Kakinada, AP, India
05-2007

Skills

  • LLMs & GenAI: hosted commercial APIs, self-hosted open-source LLMs, prompt engineering, structured generation, multilingual generation, model evaluation/selection, multimodal/vision models
  • Production AI: PySpark UDF model integration, GPU inference on on-prem/AWS, model-serving endpoints, latency/cost/privacy trade-offs, image tagging, image stylization
  • Agentic AI / AI Engineering: automated code review on Git commits, AI-assisted development and testing, test generation, Airflow failure triage, self-healing pipeline workflows
  • FDE / Solution Delivery: technical discovery, ambiguous-problem framing, architecture, POC/prototyping, productionization, troubleshooting, stakeholder collaboration and hand-off
  • Data & Streaming: Apache Spark, Hadoop/MapR, batch & real-time pipelines, ETL/ELT, Informatica, SSIS
  • Cloud & Infrastructure: AWS, Kubernetes, Docker, on-prem/cloud GPU infrastructure
  • Programming: Python, PySpark, SQL, PL/SQL, UNIX Shell, R
  • Databases & Warehousing: Oracle, Netezza, SQL Server, DB2, dimensional modeling, star/snowflake schemas, SCD
  • Analytics & BI: Churn, retention, propensity, segmentation, campaign optimization, Tableau, Power BI, SSRS
  • Engineering Tools: Git, Bitbucket, Jira, Confluence, Agile/Scrum, Quality Center, release/resource management

Accomplishments

  • 19+ years of engineering and analytics experience, with current leadership centered on production AI/GenAI and large-scale data engineering.
  • Enterprise-scale delivery: Production pipelines serving Verizon, AT&T, SoftBank and BT across multiple regions, processing terabytes and billions of records.
  • Business Impact: Multi-million dollar annual savings from real-time pipeline redesign and third-party service replacement; huge savings from NLP solution; measurable churn, zero-usage and marketing improvements.
  • Architecture: Distributed data processing, real-time event pipelines, cloud/on-prem infrastructure, Kubernetes orchestration, Airflow, GPU model serving and analytical data platforms.
  • Delivery: Technical leadership, stakeholder engagement, production support, troubleshooting, quality engineering, release management and operational ownership.

Certification

  • FAHM (Fellow, Academy for Healthcare Management) – AHIP
  • IBM DB2 UDB V8.1 Family Fundamentals
  • CSQA
  • ISTQB

Timeline

Lead Data Engineer – Advanced Analytics, Data Science & AI / GenAI

Synchronoss Technologies
12.2015 - Current

Data Analyst / Application Support Lead

Aetna – Active Health Management
04.2012 - 11.2015

Enterprise Testing & Metrics Reporting Lead

Aetna
03.2011 - 03.2012

Claims Tracking System – Data Analyst (Claims Data & ETL)

Aetna
01.2010 - 03.2011

Member Individualization – Developer

Aetna
11.2007 - 12.2009

B.Tech. / Bachelor's Degree - Civil Engineering

Jawaharlal Nehru Technological University

AI / GENAI CAPABILITY SNAPSHOT

  • LLM Applications | Title generation, multilingual translation, prompt construction, structured generation, model endpoint integration
  • Multimodal AI | Vision-model image tagging, image-byte processing, generative image stylization
  • Model Deployment | Hosted commercial APIs; self-hosted open-source models; on-prem/AWS GPU inference
  • AI Engineering | PySpark UDF integration, production pipeline embedding, workload-specific model selection
  • Agentic Workflows | Automated Git-commit code review, AI-assisted development/testing, Airflow failure triage & recovery
  • Production AI Operations | Latency, cost, quality and privacy trade-offs; operational hand-off and support
  • FDE Delivery | Discovery → architecture → rapid prototype → productionization → troubleshooting → hand-off
Shilpa Dantuluri