Summary
Overview
Work History
Education
Skills
Websites
Timeline
Generic

Sandhya Itham

Santa Clara,CA

Summary

Data Engineer with a strong background in scalable data architectures and AI-driven analytics. Developed efficient ETL/ELT frameworks and real-time data processing workflows, integrating machine learning to enhance data ecosystems. Collaborates effectively with stakeholders, translating complex datasets into actionable insights that drive strategic business decisions.

Overview

5
5
years of professional experience

Work History

Data Engineer

Fidelity Investments
12.2023 - Current
  • Developed efficient EDA, and ETL in Python using Pandas and NumPy, handling millions of records with optimized data cleaning, aggregation, and transformation to support reporting and predictive modeling.
  • Implemented supervised and unsupervised ML models (e.g., regression, classification, clustering) with Scikit-learn in Python, including feature engineering, model training, evaluation (ROC, F1-score), and deployment-ready pipelines.
  • Proficient in transforming complex datasets into actionable insights by designing interactive dashboards and visual reports using tools like Tableau and Power BI, improving data accessibility and decision-making efficiency for business stakeholders.
  • Designed, built, and maintained scalable data models and transformation workflows using Snowflake, DBT, and SQL, enabling efficient data pipelines and reliable analytics for business stakeholders.
  • Developed and maintained complex SSIS packages integrated with T-SQL scripts to perform data extraction, transformation, and loading (ETL) tasks, ensuring data accuracy and optimizing query performance in SQL Server.
  • Led end-to-end requirements gathering and documentation for complex data initiatives, specifically focusing on data catalogue and metadata definitions essential for ingestion.
  • Led the design and implementation of an enterprise-scale AI-ready Customer Data Platform integrating customer data from 40+ source systems into a centralized Data Lake architecture on Azure Data Lake Storage and Databricks.
  • Developed scalable ETL/ELT pipelines using Python, PySpark, Azure Data Factory, and Delta Lake to process over 15 TB of customer data supporting analytics and AI initiatives.
  • Implemented data modeling frameworks using Star Schema and Data Vault methodologies to support business intelligence, customer analytics, and machine learning workloads.
  • Built Retrieval-Augmented Generation (RAG) pipelines using Azure OpenAI, LangChain, vector embeddings, and Pinecone to enable AI-powered customer service search capabilities.
  • Developed real-time streaming ingestion frameworks using Kafka and Spark Structured Streaming to process customer interactions and behavioral events.
  • Integrated Snowflake and Azure Synapse Analytics environments to support enterprise reporting, advanced analytics, and AI model consumption.
  • Implemented automated data quality validation, metadata management, and governance controls ensuring trusted datasets for AI and business users.
  • Created semantic search solutions leveraging vector databases and embeddings to improve enterprise knowledge discovery and customer support operations.
  • Designed CI/CD deployment pipelines using GitHub Actions, Terraform, and Azure DevOps for automated infrastructure and data pipeline deployments.
  • Collaborated with Data Scientists and AI Engineers to prepare feature stores and AI-ready datasets supporting predictive analytics and recommendation models.
  • Collaborated extensively with Business Owners, Product Owners, and Data Architects to meticulously define data sources, data lineage, data domains, and data mapping for critical enterprise data assets.
  • Provided data-driven consulting support to stakeholders by translating business challenges into analytical solutions and delivering actionable insights.
  • Developed and maintained integration solutions to connect systems and optimize data infrastructure, ensuring seamless data flow and improved operational efficiency.
  • Leveraged Azure Data Factory, Azure DevOps, and Power Automate to streamline data workflows, automate reporting processes, and support efficient data integration and analysis across cloud-based platforms.

Data Engineer

Atos India Pvt Ltd
Bangalore, India
04.2021 - 05.2022
  • Real-time data was extracted using web scraping techniques in Python. CSV export and data manipulation. Insights are easily incorporated into Power BI for visualization.
  • Dynamic Dashboard Creation: created dynamic dashboards, demonstrating proficiency in data extraction, analysis, and visualization. Gathered all the end user’s requirements for the application's design and development.
  • Collaborated with senior data analysts and scientists to gather, clean, and prepare data for complex projects.
  • Developed new tables, views, indexes, and relations to improve the current database schema.
  • Developed and delivered business intelligence reports aligned with client requirements, enhancing data-driven decision-making.
  • Designed a cloud-native enterprise Data Lake architecture on AWS integrating structured, semi-structured, and unstructured datasets from more than 100 business applications.
  • Developed Spark and AWS Glue-based ETL pipelines to transform and load multi-terabyte datasets into curated analytics zones.
  • Implemented Large Language Model (LLM) integration frameworks using AWS Bedrock and OpenAI APIs to support enterprise Generative AI use cases.
  • Built vector search architectures using Pinecone and embeddings pipelines for AI-powered document retrieval and knowledge management solutions.
  • Created scalable data transformation workflows using dbt, Snowflake, and Apache Airflow to support data engineering and AI workloads.
  • Implemented data governance controls, lineage tracking, and cataloging solutions using enterprise metadata management tools.
  • Developed automated data validation frameworks using Python and SQL to ensure accuracy and consistency across analytical datasets.
  • Designed machine learning feature engineering pipelines supporting customer segmentation and predictive analytics initiatives.
  • Built cloud-native monitoring and observability solutions using CloudWatch, Datadog, and Splunk for proactive pipeline management.
  • Partnered with business stakeholders to define AI use cases, data requirements, and enterprise analytics strategies.
  • Created ETL mappings and workflows in Talend Studio to load data into relational tables and retrieve data from diverse sources, contributing to the establishment of a robust ETL framework.
  • Participated in creating the ETL process that extracts data from various file sources, including JSON, CSV, Excel, and XML. Collaborated with the business team to collect requirements and understand their needs.
  • Developed main and staging tables for effective data loading into the database.
  • Before loading the main tables, a Talend job was created to extract the data from a database and place it in staging tables. Used Talend FTP components and created Talend jobs to copy files between servers.
  • Maintained SSIS packages and assisted in their deployment and troubleshooting for both internal and external feeds.
  • Developed ETL packages using a variety of data sources, including XML files, SQL Server, flat files, and Excel source files.

Education

Masters - Computer Science

George Mason University
Fairfax, Virginia, US
05-2024

Bachelor’ - Computer Science

Gandhi Institute of Technology and Management ( GITAM ) Deemed to be University
Visakhapatnam, India
04-2022

Skills

  • Python
  • SQL
  • PySpark
  • Shell Scripting
  • ETL/ELT
  • Data Pipelines
  • Data Modeling
  • Data Warehousing
  • Data Lakes
  • Lakehouse Architecture
  • Data Migration
  • Data Integration
  • Data Transformation
  • Apache Spark
  • Databricks
  • Hadoop
  • Hive
  • Kafka
  • Delta Lake
  • Structured Streaming
  • AWS
  • Azure
  • GCP
  • S3
  • Redshift
  • Glue
  • Lambda
  • Athena
  • EMR
  • ECS
  • EKS
  • CloudWatch
  • IAM
  • Azure Data Factory
  • Synapse Analytics
  • ADLS
  • Azure Databricks
  • Azure SQL Database
  • Azure Functions
  • BigQuery
  • Dataflow
  • Cloud Storage
  • Pub/Sub
  • Composer
  • Snowflake
  • SQL Server
  • Oracle
  • PostgreSQL
  • MySQL
  • MongoDB
  • OpenAI
  • AWS Bedrock
  • LangChain
  • LLMs
  • Prompt Engineering
  • Vector Databases
  • Embeddings
  • Semantic Search
  • AI Service Integrations
  • Pinecone
  • ChromaDB
  • FAISS
  • Weaviate
  • MDM
  • Data Governance
  • Metadata Management
  • Data Catalog
  • Data Lineage
  • Data Quality
  • Data Security
  • Apache Airflow
  • AWS Step Functions
  • Control-M
  • Git
  • GitHub
  • GitHub Actions
  • Jenkins
  • Terraform
  • Docker
  • Kubernetes
  • CI/CD Pipelines
  • Infrastructure Code
  • Power BI
  • Tableau
  • Looker
  • QuickSight
  • Agile
  • Scrum
  • SAFe
  • SDLC
  • DevOps
  • DataOps
  • MLOps
  • GitLab
  • Bitbucket
  • Jira
  • Confluence
  • Splunk
  • ELK Stack
  • Datadog
  • Azure Monitor

Timeline

Data Engineer

Fidelity Investments
12.2023 - Current

Data Engineer

Atos India Pvt Ltd
04.2021 - 05.2022

Masters - Computer Science

George Mason University

Bachelor’ - Computer Science

Gandhi Institute of Technology and Management ( GITAM ) Deemed to be University
Sandhya Itham