Summary
Overview
Work History
Education
Skills
Work Preference
Timeline
Generic
Nick Moore
Open To Work

Nick Moore

Littleton,CO

Summary

Principal Site Reliability Engineer (SRE) and Kubernetes Architect with 20+ years of progressive infrastructure, networking, and cloud-native experience. Expert in designing hyper-scalable AWS platforms and enterprise-scale data platforms using Apache Spark within Azure Databricks and Microsoft Fabric ecosystems. Proven track record of delivering real-time analytics, ML workloads, and petabyte-scale data processing while integrating FinOps for significant cost optimization.

Overview

9
9
years of professional experience

Work History

Principal Site Reliability Engineer (SRE)

Talkdesk
01.2024 - Current
  • Pioneered FinOps-integrated Karpenter scaling on AWS EKS, achieving $500k annual savings.
  • Built AIOps anomaly detection (Datadog + Prometheus) resulting in 60% fewer false positives.
  • Partnered with Rust/C++ teams on subsecond latency microservices on Kubernetes.

Network Automation Architect

Pinterest
01.2022 - 12.2023
  • Managed petabyte-scale data flows on hybrid EKS with Four Golden Signals SLO monitoring.
  • Collaborated with Rust/C++ engineers to achieve 30% better P99 latency via Cilium CNI and mesh tuning.

VP, Site Reliability Engineering

City National Bank
Los Angeles, USA
01.2020 - 01.2022
  • Data Platform Architecture: Architected end-to-end lakehouse solutions using Microsoft Fabric and Azure Databricks, leveraging OneLake and Unity Catalog for unified governance.
  • ETL & Quality: Implemented a Medallion Architecture (Bronze/Silver/Gold) using Spark notebooks for ETL transformations, ensuring high-tier data quality for bank analytics.
  • Cloud Migration: Led zero-downtime migration of 5K users from on-prem to Azure.

Senior Cloud Engineer

Allscripts
San Francisco Bay Area, USA
12.2018 - 01.2020
  • Real-Time Analytics: Designed and deployed Spark Structured Streaming pipelines on Azure Databricks for processing millions of events per second with Delta Lake for ACID compliance.
  • ML & Scalability: Built feature engineering pipelines using PySpark and Databricks Feature Store to support ML model training across terabytes of historical data; integrated MLflow for experiment tracking.
  • Cost Optimization: Implemented autoscaling clusters with spot instance integration, reducing compute costs by 40% while maintaining strict SLAs.

Senior Platform Engineer

Gemini
Los Angeles, USA
03.2017 - 12.2018
  • Built Azure microservices platform for high-throughput crypto trading and low-latency order-matching engines.

Education

B.S. - Computer Science

University of Washington

Skills

  • Apache Spark (PySpark/SQL)
  • Azure Databricks
  • Microsoft Fabric
  • Delta Lake
  • Unity Catalog
  • MLflow
  • Structured Streaming
  • EKS
  • Karpenter
  • Lambda
  • Fargate
  • Multi-AZ/Region DR
  • Over-Provisioning
  • ML-driven anomaly detection
  • Predictive scaling
  • Datadog
  • Prometheus
  • Grafana
  • Golang Operators
  • Terraform
  • Helm
  • ArgoCD
  • GitOps
  • Service Mesh (Cilium/Istio)

Work Preference

Job Search Status

Open to work

Salary Range

$45000/yr - $200000/yr

Timeline

Principal Site Reliability Engineer (SRE)

Talkdesk
01.2024 - Current

Network Automation Architect

Pinterest
01.2022 - 12.2023

VP, Site Reliability Engineering

City National Bank
01.2020 - 01.2022

Senior Cloud Engineer

Allscripts
12.2018 - 01.2020

Senior Platform Engineer

Gemini
03.2017 - 12.2018

B.S. - Computer Science

University of Washington
Nick Moore