Summary
Overview
Work History
Education
Skills
LEADERSHIP HIGHLIGHTS
Timeline
Generic

HARSHITA GUNDI

Boston,MA

Summary

Data and AI Engineering Leader with over 20 years of experience in creating intelligent, scalable, and governed enterprise data ecosystems. Expertise in spearheading AI-first transformations through advanced technologies such as LLMs, Agentic AI frameworks, Retrieval-Augmented Generation (RAG), vector databases, and modern Lake house architectures to address complex data platform, governance, and reliability challenges. Proficient in utilizing Claude, GenAI copilots and autonomous workflow orchestration to enhance data operations, automate governance processes, boost developer productivity, and expedite enterprise decision-making. Committed to driving innovation and delivering impactful solutions that empower organizations to harness the full potential of their data assets.

Overview

16
16
years of professional experience

Work History

Director, Data and AI Engineering -

Dropbox
Boston, Massachusetts
08.2024 - Current
  • Led engineering organizations to accelerate enterprise AI adoption and build AI-native data ecosystems supporting intelligent applications, semantic retrieval, MCP-enabled integrations and GenAI workloads.
  • Architected enterprise-scale AI-ready Lakehouse platforms by modernizing legacy Hive and MapReduce systems into Databricks Delta Lake ecosystems optimized for AI training data, feature engineering, vector search, and real-time analytics.
  • Designed and operationalized enterprise RAG platforms integrating vector databases, semantic search, and Claude-powered knowledge retrieval systems to improve data discovery, developer productivity, and internal AI assistant capabilities.
  • Introduced Agentic AI frameworks to automate metadata enrichment, lineage tracking, data certification, anomaly detection, and governance workflows across streaming and batch ecosystems.
  • Leveraged Claude and enterprise LLMs to build AI copilots for data platform operations, enabling natural-language debugging, pipeline diagnostics, governance insights, and self-service analytics for engineering teams.
  • Established AI governance and Responsible AI standards including lineage-aware audibility, access controls, data quality SLAs, compliance automation, and cost governance for enterprise AI systems.
  • Drove Data Governance as a Service (DGaaS) initiatives using AI-assisted classification, policy enforcement, and automated certification pipelines for petabyte-scale data environments, significantly reducing compliance and operational risks.
  • Improved platform reliability through AI-driven observability, intelligent alerting, automated root cause analysis, and reliability scorecards across Airflow, Databricks, Presto, Kafka, and distributed compute systems.

Engineering Leader, Dataplatform and AI -

Spotify
Boston, Massachusetts
01.2022 - 08.2024
  • Directed data platform engineering strategy for Spotify for Artists, ensuring delivery of scalable AI-enabled data products and analytics platforms.
  • Collaborated with Data Science and ML teams to operationalize machine learning and GenAI workflows using MLFlow, SageMaker, Spark, and Delta Lake architectures.
  • Introduced AI-assisted observability and anomaly detection capabilities across distributed data pipelines to proactively identify reliability and data quality issues.
  • Mentored Engineering Managers and senior technical leaders in adopting AI-first platform engineering practices, automation strategies, and scalable data architecture patterns.
  • Built highly scalable Data Lake and Delta Lake solutions for structured and unstructured data processing at multi-terabyte scale using Spark, Scala, AWS, and distributed compute technologies.
  • Drove modernization initiatives incorporating LLM-based documentation generation, metadata summarization, and developer productivity enhancements.

Software Engineering Manager-

VMware
Boston, Massachusetts
06.2019 - 01.2022
  • Led engineering teams building cloud analytics and cost intelligence platforms leveraging Spark, distributed systems, and AI-assisted operational insights.
  • Developed scalable SaaS data platforms providing intelligent cloud utilization analytics and cost optimization recommendations for enterprise customers.
  • Introduced automation frameworks and AI-assisted operational tooling to improve cloud reporting accuracy, scalability, and engineering productivity.
  • Directed teams delivering large-scale Spark and Scala data pipelines processing complex cloud telemetry datasets with improved performance and operational reliability.
  • Established engineering best practices around platform scalability, CI/CD automation, distributed systems reliability, and data governance.
  • Partnered with product and architecture leadership to modernize analytics platforms for future AI and predictive analytics use cases.

TeamLead Analytics Engineering -

McGraw-Hill Education
Boston, Massachusetts
08.2016 - 06.2019
  • Architected cloud-native data infrastructure platforms capable of processing millions of real-time customer interaction events using Spark, NoSQL systems, Kafka, and distributed data processing frameworks.
  • Built event-driven microservices and scalable APIs supporting intelligent analytics and large-scale educational engagement platforms.
  • Led engineering teams delivering high-throughput streaming applications and scalable data processing systems.
  • Implemented foundational metadata and analytics frameworks that later enabled advanced machine learning and recommendation use cases.
  • Drove modernization of distributed systems architecture using open standards, cloud-native design principles, and scalable data engineering practices.

Web Application Engineer -

MathWorks
Boston, Massachusetts
01.2015 - 08.2016
  • Designed and implemented real-time and batch ingestion frameworks using Kafka, Kinesis, Lambda, DynamoDB, and Redis to support advanced analytics and enterprise reporting systems.
  • Built scalable CI/CD pipelines and developer tooling to improve deployment automation, testing reliability, and platform delivery velocity.
  • Developed frameworks integrating large-scale analytics platforms with distributed data infrastructure for enterprise decision-making.

Software Team Lead -

IBM
Boston, Massachusetts
05.2010 - 01.2015
  • Led teams designing and developing large-scale ETL and Hadoop-based data processing platforms for structured and unstructured enterprise data.
  • Drove architectural decisions, code quality initiatives, and scalable distributed data processing implementations.
  • Built optimized database models, analytical queries, and ingestion pipelines supporting enterprise analytics and reporting systems.
  • Collaborated across engineering and infrastructure teams to improve deployment automation, testing, platform reliability, and operational scalability.

Education

Master of Science (MS) - Computer Information Systems

Boston University
Boston, Massachusetts

Skills

  • Apache Spark Scala Python Databricks Delta Lake Apache Airflow Kubernetes Docker Apache Flink Kafka Google Dataflow Apache Storm Amazon Kinesis PostgreSQL Microservices Vector Databases RAG Implementation Claude APIs OpenAI APIs LangChain Agentic AI Frameworks MLFlow SageMaker Triton Inference Server KServe AI Governance Semantic Search LLMOps Data Lineage & Cataloging

LEADERSHIP HIGHLIGHTS

  • Led enterprise AI and data platform modernization initiatives supporting petabyte-scale analytics and GenAI adoption.
  • Operationalized enterprise RAG systems and vector search platforms for intelligent knowledge discovery.
  • Introduced Agentic AI workflows for governance automation, metadata intelligence, and operational efficiency.
  • Built AI governance frameworks balancing innovation, compliance, reliability, and responsible AI adoption.
  • Scaled and mentored high-performing engineering organizations across AI infrastructure, platform engineering, and distributed data systems.
  • Championed AI-assisted engineering productivity through Claude-powered copilots and intelligent platform automation.

Timeline

Director, Data and AI Engineering -

Dropbox
08.2024 - Current

Engineering Leader, Dataplatform and AI -

Spotify
01.2022 - 08.2024

Software Engineering Manager-

VMware
06.2019 - 01.2022

TeamLead Analytics Engineering -

McGraw-Hill Education
08.2016 - 06.2019

Web Application Engineer -

MathWorks
01.2015 - 08.2016

Software Team Lead -

IBM
05.2010 - 01.2015

Master of Science (MS) - Computer Information Systems

Boston University
HARSHITA GUNDI