Site Reliability Engineer specializing in proactive monitoring and incident reduction. Developed and implemented observability dashboards that cut incident rates by 50% while maintaining compliance with industry standards. Skilled in leveraging DevSecOps practices to enhance automation and team productivity.
Overview
1
1
Certification
16
16
years of professional experience
Work History
Site Reliability Engineer
Bright Vision Technologies
06.2026 - Current
Handled production support and on-call responsibilities with SRE practices to minimize downtime and latency for critical customer services.
Created monitoring dashboards for application observability, security, business observability, infrastructure observability, log analytics, and threat detection; collaborated with cross-functional teams to ensure PHI and PII data compliance.
Developed automations and auto-scaling mechanisms within DevSecOps framework.
Managed JFrog Artifactory as a centralized artifact repository for Maven, NuGet, npm, Docker images, and generic artifacts, enabling secure storage and version management.
Integrated JFrog Artifactory with Azure DevOps CI/CD pipelines to publish build artifacts, manage Docker images, and maintain build traceability using Build Info.
Leveraged JFrog Artifactory for artifact lifecycle management, Docker image versioning, Build Info generation, and seamless integration with CI/CD workflows.
Agentic AI proof of concepts - using multimodal prompt engineering, building a health care intelligent RAG assistant, Research, Analysis & Generating Reports with Prompt Engineering, Build ReAct Banking text2SQL AI Agent with LangChain & LangGraph, Learnt Evaluating & Monitoring RAG & Agentic AI Systems, Analysis on Dynatrace – Agentic AI systems
Experience in working with Platform Team – capacity planning, database team - Data Consistency & Availability, implemented best security practices in Azure Devops Impact
Site Reliability Engineer Lead
Oak Street Health, Part of CVS Health
10.2024 - 04.2026
Apache Kafka: Integrated Confluent Kafka into Kubernetes production environment, enhancing data streaming capabilities.
Proficient in handling on production support and on-call with SRE Practices forefront to reduce downtime and latency for the critical customer services.
Delivered healthcare domain services, contributing to improved patient data management and compliance.
Experience in creating monitoring dashboards for Application Observability and Security, Business Observability, Infrastructure Observability, Log Analytics, Threat and Vulnerabilities detection. Worked with many cross-functional teams on PHI PII data compliance standards.
Utilized JFrog Artifactory to store, version, and distribute build artifacts and container images while supporting Build Info, artifact replication, and security scanning integration.
Automated projects on Incident Management using Azure Function, SlackBot, Jira, Kubernetes, Python APIs, shell scripting, sqlite, etc.
Administered Kubernetes cluster using observability and monitoring - Azure monitors, synthetics, prometheus, grafana, Grafana Tracing, Grafana logging Loki, OpenTelemetry.
Migrating on-premises applications to Azure Kubernetes cluster with docker containerized microservices architecture.
Proposed AIOps - architecting end-to-end monitoring infrastructure with advanced dashboards, observability, and proactive alerting.
Agentic AI proof of concepts - using multimodal prompt engineering, building a health care intelligent RAG assistant, Research, Analysis & Generating Reports with Prompt Engineering, Build ReAct Banking text2SQL AI Agent with LangChain & LangGraph, Learnt Evaluating & Monitoring RAG & Agentic AI Systems, Analysis on Dynatrace – Agentic AI systems.
Collaborated with Platform Team on capacity planning and database consistency, implementing security best practices in Azure DevOps.
Successfully achieved bringing the entire team to implement SRE principles to avoid customer downtime.
Increased automation rate to 75% in team, Reduced false positive IT alerts and incidents from 25% to 0%, human errors by 82%.
Reduced monitoring and observability (Grafana – Logs, Traces, Spans, Metrics) billing cost from 99% to 50%.
Instrumented OpenTelemetry for around 20 services.
Successfully migrated health care APIs from one third party to the other by thoroughly analyzing the user traffic.
Implemented the whole slack Bot automation.
Created proactive Anomaly Detection Observability dashboards and reduced incidents count from 80% to 40%.
Trained leadership team about SRE practices, Successfully led the team to eliminate any blockers on the assignments.
Migrated all Observability and Monitoring dashboards using Infrastructure as a Code – Terraform.
Site Reliability Engineer
Microsoft
02.2024 - 10.2024
Established Azure DevOps deployment pipeline end-to-end, incorporating best security practices.
Integrated GitHub, Azure DevOps Pipelines, Docker, and JFrog Artifactory to automate CI/CD, manage artifact versioning, and support consistent deployments across Dev, QA, and Production.
Experience in developing Infrastructure as a Code using ARM templates, Azure Kubernetes Cluster worker node version upgrades.
Reduced vulnerabilities from 70% to 10% - Microsoft Compliance and Governance.
Resolved Vulnerabilities from the services, Image Vulnerabilities in registries.
Automated compliance incident mitigation for 20 issues, reducing manpower needs by 3 personnel and allowing team to concentrate on new customer feature development.
Increased 25% developer productivity by offloading team’s repetitive work with my automations.
Developed statistical data using Kusto Query Language (KQL) for telemetry of customer-live quantum computing services.
Site Reliability Engineer
IBM
01.2021 - 01.2024
Managed full Kubernetes lifecycle on IBM Cloud (networking, Ingress, Istio, storage, upgrades, OS patching, security) and secured microservices. Ensured compliance with HIPAA, SOC2, FedRAMP, CIS.
SRE Lead observability using New Relic, Grafana, Prometheus, Sysdig, Instana, OpenTelemetry, synthetics.
Implemented CI/CD pipelines (Jenkins, Travis, Tekton) with Artifactory, Container Registry, Cloud Object Storage; automated ZAP/SonarQube scans, SSL updates, and compliance, Designed FinOps tools, SlackBot for GitOps workflows.
Ensured compliance with HIPAA, SOC2, FedRAMP, CIS.
Facilitated live demonstrations for SRE Guild members to enhance collaboration.
Improved reliability (MTTR, SLO/SLA) using SRE practices with service availability 99.99%.
Guided 15 teams, 30 microservices to migrate all monitoring dashboards and alerts New Relic to Sysdig.
Reduced human errors through automations by 30%.
Saved Monitoring tool migration $50K a month revenue.
Created highly significant 15 LOGDNA log analysis Dashboards.
Created Proactive monitoring and alerts thus reducing 40% of the customer issues and building happy customers.
Site Reliability Engineer - Oracle Cloud Infrastructure
ORACLE
01.2016 - 01.2021
Cloud Migration: Expertise in billing and metering enablement whitelisting for oracle cloud customers.
CI/CD: Created TeamCity pipelines, deployed Oracle Cloud Infrastructure IaaC with Terraform, OCI deployment - metering and billing.