Summary
Overview
Work History
Education
Skills
Personal Information
ATS KEYWORDS
Certification
Timeline
Generic

Gopii Krishna

Dallas-Fort Worth,TX

Summary

Senior Data Engineer / Databricks Engineer with 15+ years of experience spanning database development, ETL engineering, AWS cloud data engineering, Databricks Lakehouse platforms, enterprise data integration, and analytics-ready data solutions. Progressive experience from database and ETL development into AWS Glue and Databricks engineering, with strong hands-on expertise in Python, PySpark, SQL, Apache Spark, Delta Lake, AWS, Talend, Snowflake, CI/CD, data migration, data quality, performance optimization, and production support. Experienced in designing scalable ETL/ELT pipelines, cloud-native ingestion frameworks, batch and near-real-time processing, CDC, dimensional modeling, enterprise data modernization, data governance, and technical leadership. Strong exposure to Generative AI, LLM, RAG, Agentic AI, vector search, MLflow, and Databricks Mosaic AI.

Overview

1
1
Certification
15
15
years of professional experience

Work History

Databricks Engineer

TIAA
06.2025 - Current
  • Design, develop, and support enterprise-scale Databricks Lakehouse solutions using Databricks, Apache Spark, Delta Lake, Python/PySpark, SQL, and modern cloud data technologies.
  • Developed and maintained scalable data ingestion, transformation, validation, and downstream analytics pipelines for enterprise workloads, ensuring data accessibility and reliability.
  • Design modular PySpark notebooks and reusable data-processing frameworks aligned with enterprise engineering standards and maintainability requirements.
  • Implement Delta Lake data models, transactional pipelines, data quality controls, and scalable ingestion frameworks supporting reliable analytics and reporting.
  • Develop and manage Databricks Workflows, Jobs, dependency-driven DAGs, scheduling, monitoring, and automated recovery processes.
  • Optimize Spark workloads using partitioning, Z-Ordering, caching, broadcast joins, adaptive query execution, file compaction, and cluster-resource optimization.
  • Collaborated with stakeholders, data architects, analysts, application teams, and business teams to translate requirements into production-ready Lakehouse solutions, enhancing cross-functional alignment.
  • Troubleshoot production issues involving Spark processing, workflow orchestration, data quality, pipeline failures, and cloud infrastructure performance.
  • Supported CI/CD automation for Databricks notebooks, workflows, code, and infrastructure using Git and DevOps practices, streamlining deployment processes and improving release efficiency.
  • Establish operational metrics, coding standards, performance benchmarks, and production support practices to improve reliability and SLA-driven delivery.

Databricks Engineer

Charles Schwab Bank
06.2022 - 05.2025
  • Developed and supported enterprise data engineering solutions using Databricks, Apache Spark, PySpark, SQL, Delta Lake, and cloud data technologies.
  • Built scalable ETL/ELT pipelines for ingestion, transformation, validation, and delivery of enterprise data to analytics and downstream applications.
  • Designed Lakehouse data-processing workflows using Delta Lake, Databricks Workflows, Jobs, Notebooks, and modular PySpark components.
  • Implemented batch and near-real-time ingestion patterns using Auto Loader, COPY INTO, Structured Streaming, Kafka, and CDC approaches.
  • Applied dimensional modeling including Star Schema and Snowflake Schema to support analytics, reporting, and machine-learning workloads.
  • Performed Spark and ETL performance tuning using AQE, broadcast joins, skew handling, partitioning, caching, Z-Ordering, OPTIMIZE, VACUUM, and Photon optimization.
  • Implemented data governance, metadata management, security, and access-control practices using Unity Catalog and Lakehouse governance capabilities.
  • Built reusable Python/PySpark ETL frameworks with logging, exception handling, configuration-driven execution, and standardized utilities.
  • Developed enterprise GenAI, LLM, RAG, and Agentic AI capabilities using OpenAI, Claude, LangChain, LangGraph, MCP, Databricks Mosaic AI, embeddings, and vector search frameworks.
  • Implemented CI/CD and DevOps pipelines using GitHub, GitHub Actions, Jenkins, Terraform, Databricks CLI, and Databricks Asset Bundles while supporting code reviews and technical delivery.
  • Troubleshoot pipeline failures, data-quality issues, and performance problems across development and production environments and implemented sustainable fixes.
  • Collaborated with data, application, architecture, and business teams to translate requirements into production-ready data engineering solutions.

AWS Glue Engineer

Cambia Health
05.2020 - 05.2022
  • Developed and supported cloud-based ETL pipelines using AWS Glue, Amazon S3, Glue Data Catalog, Lambda, Step Functions, SQL, Python, and PySpark.
  • Built cloud-native ingestion and transformation workflows to move, cleanse, validate, and process enterprise data for downstream applications and analytics.
  • Designed SQL-centric ETL transformations using joins, subqueries, window functions, CTEs, aggregations, ranking, filtering, and deduplication logic.
  • Applied PySpark within AWS Glue for data movement, transformations, parallel processing, and performance optimization.
  • Implemented incremental data loads and CDC processing using partitioning strategies to improve refresh efficiency and reduce processing time.
  • Optimized ETL performance using predicate pushdown, partition filtering, query rewriting, file optimization, and reduced full-table scans.
  • Designed Raw, Staging, and Curated transformation layers using dimensional modeling and standardized ETL patterns.
  • Built reusable, parameterized ETL frameworks and AWS Glue Jobs for consistent processing across multiple pipelines.
  • Monitored and troubleshoot ETL jobs, investigated failures, performed root-cause analysis, and implemented corrective actions.
  • Supported data integration, migration, reconciliation, data-quality validation, and parallel-run activities during cloud modernization initiatives.
  • Collaborated with technical teams to support production data workflows, operational monitoring, SLA compliance, and reliable cloud data delivery.

Senior ETL Developer

BCID
04.2019 - 04.2020
  • Designed, developed, and supported enterprise ETL processes for data integration and data warehouse environments using SQL and ETL frameworks.
  • Developed complex transformation logic to extract, cleanse, transform, validate, and load structured and semi-structured enterprise data.
  • Built reusable ETL components and modular workflows to improve maintainability, consistency, and delivery efficiency.
  • Implemented incremental loads, CDC patterns, data-quality checks, reconciliation, and historical-data processing.
  • Applied SQL optimization techniques including predicate pushdown, query rewriting, partition filtering, efficient joins, and reduced full-table scans.
  • Designed and maintained dimensional data models including Star Schema and Snowflake Schema for reporting and analytics.
  • Worked with business and technical teams to understand requirements, map source-to-target relationships, and translate requirements into data solutions.
  • Troubleshoot ETL failures, data discrepancies, pipeline errors, and production issues; performed root-cause analysis and implemented corrective actions.
  • Supported production deployments, release coordination, incident/change management, and operational monitoring.
  • Provided technical guidance on ETL development, data processing, performance optimization, troubleshooting, and issue resolution.

ETL Developer

Medicare
01.2016 - 03.2019
  • Developed and maintained enterprise ETL processes supporting data integration, reporting, data warehousing, and downstream analytics requirements.
  • Created complex SQL queries and transformation logic to extract, cleanse, standardize, transform, and load data across source and target systems.
  • Designed reusable ETL workflows and parameterized processing patterns for repeatable, scalable data delivery.
  • Implemented incremental data processing and CDC patterns to improve load efficiency and support historical data management.
  • Built analytics-ready datasets and transformation layers using dimensional modeling, Star Schema, and Snowflake Schema concepts.
  • Performed data reconciliation, row-level comparisons, data-quality validation, and source-to-target testing.
  • Supported data warehouse development, production ETL operations, scheduling, monitoring, and SLA-driven data delivery.
  • Investigated data discrepancies, pipeline failures, transformation issues, and production incidents and coordinated fixes with technical teams.
  • Performed SQL and ETL performance optimization through indexing considerations, query tuning, partition filtering, efficient joins, and reduced unnecessary data movement.
  • Collaborated with business analysts, application teams, and data stakeholders to translate business requirements into reliable data solutions.

Database Developer

Telstra
06.2011 - 12.2015
  • Developed and maintained database solutions supporting enterprise applications, data processing, reporting, and business data requirements.
  • Designed SQL queries, stored procedures, database objects, views, and transformation logic for application and data-processing needs.
  • Performed data analysis, source-to-target investigation, troubleshooting, reconciliation, and database support across development and production environments.
  • Built data extraction and transformation routines supporting enterprise ETL, reporting, and downstream analytics requirements.
  • Optimized SQL queries and database processing through query rewriting, efficient joins, filtering, indexing strategies, and reduction of unnecessary scans.
  • Supported data integration across relational databases, flat files, APIs, and enterprise application sources.
  • Investigated production incidents, data discrepancies, performance issues, and database failures and implemented sustainable corrective actions.
  • Collaborated with application, ETL, business, and technical teams to resolve database issues and deliver data-related enhancements.
  • Supported release, deployment, change-management, and production support activities while maintaining operational documentation.
  • Provided technical guidance on SQL development, database processing, data quality, troubleshooting, and performance optimization.

Education

Master of Science - Master of Information Systems

M I T - Melbourne Institute of Technolgy
Auburn, PA
05-2011

Bachelor of Science - Bachelor of Computer Applications

Kakatiya University
India
05-2009

Skills

  • Programming: Python, PySpark, SQL, Scala, Shell Scripting
  • Databricks & Lakehouse: Databricks Workspace, Delta Lake, Unity Catalog, Databricks Workflows, Jobs, Repos, SQL Warehouses, Photon Engine, Auto Loader, COPY INTO, Medallion Architecture
  • Big Data & Streaming: Apache Spark, Spark SQL, Spark Structured Streaming, Kafka, Hadoop, AWS Kinesis, CDC
  • AWS: AWS Glue, Amazon S3, Lambda, Step Functions, Glue Data Catalog, CloudWatch, CodePipeline
  • Azure: Azure Data Factory, ADLS Gen2, Azure Blob Storage, Azure Key Vault, Azure Event Hubs, Azure Monitor
  • ETL / Orchestration: Talend DI, AWS Glue, Azure Data Factory, Apache Airflow, Databricks ETL, DAG orchestration
  • Databases / Warehousing: Snowflake, SQL Server, PostgreSQL, Oracle, MongoDB, Athena
  • Data Modeling: Star Schema, Snowflake Schema, Data Vault, Dimensional Modeling, SCD Type 1/2
  • DevOps / CI-CD: Git, GitHub, GitHub Actions, Azure DevOps, Jenkins, Terraform, Docker, Kubernetes, Databricks CLI, Databricks Asset Bundles
  • GenAI / AI: LLMs, RAG, AI Agents, Agentic AI, Prompt Engineering, OpenAI API, Claude, Databricks Mosaic AI, Hugging Face
  • AI Frameworks / Vector: LangChain, LangGraph, MCP, Databricks Vector Search, FAISS, ChromaDB, Pinecone, Azure AI Search
  • MLOps / LLMOps: MLflow, Model Registry, Model Deployment, Model Monitoring, CI/CD for ML pipelines
  • Performance / Operations: AQE, Partitioning, Caching, Broadcast Joins, Skew Handling, Z-Ordering, OPTIMIZE, VACUUM, Predicate Pushdown, Spark UI, Log Analytics

Personal Information

Title: SENIOR DATA ENGINEER | DATABRICKS ENGINEER | AWS DATA ENGINEER

ATS KEYWORDS

Senior Data Engineer, Lead Data Engineer, Databricks Engineer, AWS Data Engineer, Databricks, AWS Databricks, Apache Spark, PySpark, Python, SQL, Spark SQL, Delta Lake, Unity Catalog, Lakehouse, Medallion Architecture, Databricks Workflows, Databricks Jobs, Auto Loader, COPY INTO, Structured Streaming, Kafka, CDC, ETL, ELT, AWS Glue, Amazon S3, AWS Lambda, AWS Step Functions, Glue Data Catalog, Azure Data Factory, ADLS Gen2, Talend, Snowflake, Data Warehousing, Data Modeling, Star Schema, Snowflake Schema, SCD Type 1, SCD Type 2, Data Quality, Data Governance, Data Lineage, CI/CD, Git, GitHub, GitHub Actions, Jenkins, Terraform, Databricks Asset Bundles, Apache Airflow, REST API, SOAP API, GenAI, Generative AI, LLM, RAG, Agentic AI, OpenAI, Claude, LangChain, LangGraph, MCP, Databricks Mosaic AI, Vector Search, MLflow, MLOps, LLMOps, Spark Performance Tuning, AQE, Partitioning, Caching, Broadcast Joins, Skew Handling, Z-Ordering, OPTIMIZE, VACUUM, Predicate Pushdown, Production Support, Incident Management, Change Management, Root Cause Analysis, Technical Leadership, Code Review, Mentoring.

Certification

  • AI Certified
  • AWS Solution Architect

Timeline

Databricks Engineer

TIAA
06.2025 - Current

Databricks Engineer

Charles Schwab Bank
06.2022 - 05.2025

AWS Glue Engineer

Cambia Health
05.2020 - 05.2022

Senior ETL Developer

BCID
04.2019 - 04.2020

ETL Developer

Medicare
01.2016 - 03.2019

Database Developer

Telstra
06.2011 - 12.2015

Master of Science - Master of Information Systems

M I T - Melbourne Institute of Technolgy

Bachelor of Science - Bachelor of Computer Applications

Kakatiya University
Gopii Krishna