Summary
Overview
Work History
Education
Skills
Websites
Certification
Timeline
Generic

KANWARPREET SINGH

Tracy,CA

Summary

Accomplished Senior Site Reliability Engineer / Platform Engineer with 13 years of experience in designing and managing highly available AWS cloud platforms. Expertise in Kubernetes, GitLab CI/CD, Terraform, and Infrastructure as Code. Led standardization of CI/CD platforms, executed zero-downtime EKS upgrades across multiple production clusters, and optimized cloud infrastructure through automation.

Overview

16
16
years of professional experience
1
1
Certification

Work History

Senior Site Reliability Engineer

Bill.Com
San Jose
08.2020 - Current
  • Executed zero-downtime Amazon EKS upgrades (Control Plane, Node Groups, CoreDNS, VPC CNI, and add-ons) across 20+ production EKS clusters, improving security, stability, and Kubernetes feature adoption.
  • Owned the GitLab CI/CD platform as GitLab Administrator, driving platform-wide CI/CD improvements by designing reusable pipeline templates, integrating Terraform-based infrastructure provisioning, standardizing CI/CD pipelines to eliminate operational toil and improve the Developer Experience (DevEx), optimizing build and deployment performance, establishing CI/CD best practices, and mentoring 40+ engineering teams to improve developer productivity, increase deployment reliability, and accelerate time-to-production.
  • Reduced deployment time by 80%+ by transforming fragmented deployment workflows into One-Click deployment pipelines for major and minor releases, minimizing manual effort and deployment failures.
  • Led production release management for major and minor releases, enabling RBAC-based self-service deployments through Ansible Tower and improving release velocity across engineering teams.
  • Developed reusable Terraform modules to provision and manage AWS infrastructure, ensuring consistency, reducing configuration drift, and accelerating environment provisioning.
  • Developed a Go-based automation platform for monitoring and managing SSL/TLS and GPG certificates, automating certificate lifecycle management, proactive expiration alerts, and end-to-end certificate rotations with enterprise CAs including DigiCert and GoDaddy for critical banking partners (e.g., Citi, JPMC, BMO), eliminating manual operational effort.
  • Saved $20K+ per month in AWS infrastructure costs by optimizing NAT Gateway traffic through caching strategies and a DNAT-based proxy architecture, significantly reducing network egress charges while maintaining application performance.
  • Reduced Amazon ECR storage costs by 40% by implementing automated lifecycle policies, image retention strategies, and IAM governance, improving cost efficiency while strengthening security and compliance.
  • Led the migration of Jenkins CI/CD infrastructure from on-premises to AWS, improving scalability, reliability, and platform availability for growing engineering teams.
  • Designed and optimized monitoring and alerting for availability, latency (p50/p95/p99), error rates, resource utilization, and infrastructure drift, enabling faster incident detection and reducing MTTR.
  • Partnered with application teams to take microservices from design to production, providing Kubernetes platform engineering, CI/CD automation, infrastructure provisioning, and operational support.

Linux Administrator/Kubernetes Admin

SoundHound
San Jose
06.2019 - 08.2020
  • Administered and optimized enterprise GitLab and Jenkins CI/CD platforms, managing upgrades, plugin lifecycle, RBAC, security hardening, high availability, and platform governance for multiple engineering teams.
  • Designed, standardized, and maintained GitLab CI/CD and Jenkins pipelines, integrating Ansible and Infrastructure as Code (Terraform) to enable automated, repeatable application deployments across Development, QA, UAT, and Production environments.
  • Implemented GitLab CI/CD best practices, reusable pipeline templates, and developer enablement initiatives to improve deployment consistency, reduce operational toil, and accelerate software delivery.
  • Provisioned and managed cloud infrastructure using Terraform, creating reusable Infrastructure as Code modules to improve consistency and reduce manual provisioning.
  • Led production deployments for major releases, minor releases, and hotfixes using GitLab CI/CD and Jenkins, coordinating cross-functionally with Development and QA teams to ensure reliable, zero-impact releases.
  • Implemented disaster recovery capabilities by configuring storage snapshots, backup strategies, and recovery procedures for critical infrastructure.
  • Configured and administered Linux infrastructure, including NFS, LVM, storage management, and system administration across production environments.
  • Designed and implemented proactive monitoring and alerting using Prometheus and Grafana, partnering with development teams to improve application observability across Development, QA, and UAT environments.
  • Supported 24×7 production on-call rotations, resolving infrastructure and application incidents within SLA targets for Production and Sandbox environments

Senior DevOps Engineer

Verisk Analytics
San Francisco, USA
08.2016 - 06.2019
  • Provided L2/L3 production support across Development, QA, UAT, and Production environments, ensuring high availability and timely incident resolution for internal and customer-facing applications.
  • Played a key role in application migration from on-premises infrastructure to AWS, including migrating VMware virtual machines to Amazon EC2 and modernizing deployment workflows.
  • Provisioned and managed AWS infrastructure and operational tools, configuring CloudWatch, Amazon SNS, and Slack integrations for proactive monitoring and alerting.
  • Automated infrastructure management and operational workflows using Puppet, Ansible, and Python, reducing manual effort and improving operational consistency.
  • Administered and optimized Jenkins CI/CD infrastructure, managing masters, agents, plugins, RBAC, and pipeline automation to improve software delivery.
  • Strengthened application security through OS hardening, network security measures, application security controls, and regular security patching in Linux production environments.
  • Designed and configured NGINX reverse proxy and load balancing solutions for internal and external applications, improving application availability and traffic distribution.
  • Extended and maintained internal Platform-as-a-Service (PaaS) platform, standardizing application hosting and accelerating service delivery.
  • Installed, upgraded, patched, and administered enterprise middleware and infrastructure components including WSO2, Cassandra (DataStax), and NGINX on Red Hat Enterprise Linux (RHEL 6/7).

Senior Linux Administrator/DevOps

Ebay
San Jose, USA
08.2015 - 08.2016
  • Managing Physical servers & Virtual Machines for various environments for HLT, Engineering and Science Teams.
  • Maintenance of physical GPU machines and various open source packages used for Computational research and Machine learning packages like Caffe, Cuda, Cudnn, tensorflow etc.
  • Deploying Machine Translation models into Production developed by Science team manually and by using internal PAAS service.
  • Managing LVM for extending storage for huge GPU processing & closely working with Storage team to resolve storage issues.
  • Automated internal deployments using Puppet, Hiera on Openstack based private cloud.
  • Configured Jenkins build pipelines using Jenkins master and slave architecture.
  • Deploying immutable infrastructure using Terraform.
  • Worked with Storage team to setup high speed SAN mounts for the development clusters.

Linux Administrator

Ericsson
09.2012 - 08.2014
  • Providing L2 support for 250+ servers.
  • Coordinated with L1 team to resolve issues, enhancing service response.
  • Monitored daily system administrator activities, including performance monitoring and disk space management.
  • Managed daily backup processes, ensuring data integrity and availability.
  • MySQL new server installation, administration and migration.
  • Coordinated with incident management and configuration management teams to streamline processes, maintaining maximum uptime and minimizing escalations.
  • Configured clusters to provide failover, load balancing and deployment of Charging System application servers.
  • Working with senior administrators and developers to troubleshoot the core issues.
  • Troubleshooting of IN applications to ensure maximum uptime.

System Administration

Ericsson
07.2010 - 09.2012
  • Monitored 240+ Linux servers for performance, ensuring maximum uptime.
  • Managed and monitored IN nodes with Nagios to maintain infrastructure stability.
  • Investigated issues and performed first-level troubleshooting to restore service functionality.
  • Monitoring and providing support for Ericsson Charging System 3.0.
  • Adhered to defined processes to ensure compliance with service level agreements.
  • Recorded processes according to company policies and instructions.
  • Executing the process (running manual Linux commands) for checking services and raising trouble tickets.
  • Raising Trouble tickets to 2nd LA via Remedy tool.

Education

Masters of Science - Computer Science

Silicon Valley University
San Jose, CA, USA
01-2015

Bachelor in Technology - Electronics and Communications Engineering

Punjab Technical University
India
01-2010

Skills

  • Infrastructure as code
  • CI/CD optimization
  • Terraform module development
  • Cloud infrastructure management
  • Monitoring and alerting
  • GitLab administration
  • Zero-downtime upgrades

Certification

  • Certified Kubernetes Admin (CKA)
  • Puppet Certified Professional (PPT-203)
  • Redhat Certified Systems Administrator
  • AWS Certified Solutions Architect Associate

Timeline

Senior Site Reliability Engineer

Bill.Com
08.2020 - Current

Linux Administrator/Kubernetes Admin

SoundHound
06.2019 - 08.2020

Senior DevOps Engineer

Verisk Analytics
08.2016 - 06.2019

Senior Linux Administrator/DevOps

Ebay
08.2015 - 08.2016

Linux Administrator

Ericsson
09.2012 - 08.2014

System Administration

Ericsson
07.2010 - 09.2012

Masters of Science - Computer Science

Silicon Valley University

Bachelor in Technology - Electronics and Communications Engineering

Punjab Technical University
KANWARPREET SINGH