
Results-driven Application Production Support and Site Reliability Engineering (SRE) with 12+ years of experience supporting mission-critical enterprise applications and high-availability systems. Proven track record of improving application reliability, operational stability, and service performance through automation, observability, and proactive incident management. Experienced in transforming traditional production support operations into scalable SRE practices that reduce operational toil and improve service resilience using automation. Skilled in enterprise monitoring and observability using Grafana, Splunk, and Apica, with strong expertise in relational databases and advanced SQL queries. Proficient in developing automation solutions using Python, AAAS workflows, AI-enabled capabilities, and Linux shell scripting to streamline compliance, incident response, and operational processes. Recognized for leading major incident response efforts, driving cross-functional collaboration, and reducing Mean Time to Recovery (MTTR) while maintaining high levels of service availability and customer satisfaction.
Production Support & SRE Evolution: Managed day-to-day production support operations for critical enterprise infrastructure, driving the shift toward SRE model by using Python, Linux Bash and other scripts to eliminate manual operational toil.
Observability & Telemetry Engineering: Configured and maintained end-to-end application and infrastructure monitoring using Grafana, Splunk, Apica, and Dynatrace. Designed comprehensive telemetry dashboards and fine-tuned intelligent alerting systems to prevent outages and meet SLOs.
Job Scheduling & Orchestration: Engineered and optimized automated batch schedules and critical data processing pipelines utilizing advanced AutoSys scheduling and Linux cron jobs, lowering system failure rates.
Resiliency Automation & Compliance: Infrastructure resiliency and disaster recovery (DR) test executions. Programmed custom screen-capture and evidence-collection scripts, reducing recovery time frames and manual engineering efforts during Data center network event or regular resiliency test.
Major Incident Bridge Leadership:Acted as the primary technical lead on high-severity incident bridges. Orchestrated cross-functional teams to isolate root causes, drive rapid system remediation, and consistently reduce MTTR.
Application Modernization: critical application modernization efforts, including database migrations, Java and cipher upgrades, and logging enhancements, improving system stability, monitoring accuracy, and platform reliability.
Cross-Functional Collaboration: Collaborated with Product, Engineering, SRE, and business stakeholders to drive process improvements, strengthen operational governance, and enhance the overall user experience.