Work Preference
Summary
Overview
Work History
Education
Skills
Certification
Timeline
Generic
Manisha Sona Mengu
Open To Work

Manisha Sona Mengu

SRE Engineer
Delaware,OH

Work Preference

Job Search Status

Open to work

Work Type

Full Time

Location Preference

HybridRemote

Summary

Site reliability engineer with extensive experience in cloud infrastructure management, specializing in Azure, AWS, and hybrid systems. Demonstrated ability to lead initiatives that improve reliability, scalability, and security of services. Successful in driving automation efforts that reduce operational overhead and enhance incident management processes.

Overview

1
1
Certification
9
9
years of professional experience

Work History

Senior Site Reliability Engineer (SRE)

Victoria's Secret
Reynoldsburg, Ohio
08.2022 - Current
  • Established enterprise-wide site reliability engineering practices for large-scale retail and ecommerce platforms across Azure, AWS, and hybrid cloud environments.
  • Drove reliability strategy by defining service level indicators (SLIs), service level objectives (SLOs), availability targets, and operational excellence standards.
  • Architected and managed observability platforms using Dynatrace, Grafana, Prometheus, ELK stack, and Fluentd for comprehensive visibility in distributed systems.
  • Implemented scalable monitoring, logging, alerting, and dashboard frameworks to enhance incident detection and reduce mean time to resolution (MTTR).
  • Developed infrastructure as code and automation frameworks with Terraform, Ansible, Chef, Azure DevOps, and Jenkins to standardize deployments and reduce operational overhead.
  • Coordinated cross-functional teams for incident management and production support of mission-critical applications during high-severity incidents, improving response effectiveness.
  • Integrated reliability and operational readiness into software development lifecycle by collaborating with software engineering, architecture, security, and platform teams.
  • Supported highly available distributed systems using Kubernetes, Docker, Kafka, Confluent, Nginx, MongoDB, Couchbase, and Azure Cosmos DB.

Site Reliability Engineer

Ellucian
DeBary, Florida
02.2021 - 08.2022
  • Ensured continuous operation of highly available OpenStack Linux environments with minimal downtime.
  • Developed automation solutions that streamlined deployment processes and improved service provisioning.
  • Managed Jenkins, Artifactory, Git platforms, and CI/CD tooling practices to ensure seamless integration and delivery.
  • Implemented monitoring protocols with Splunk and Grafana to maintain application performance and reliability.
  • Executed platform upgrades alongside security patch management and vulnerability remediation tasks.
  • Engineered cloud-native applications along with robust infrastructure platforms for enterprise education products.
  • Deployed secure authentication mechanisms using OAuth, OpenID Connect, and SAML technologies.
  • Contributed to 24x7 production support efforts including incident response initiatives.

SRE/Application System Engineer

Comcast
Philadelphia, Pennsylvania
08.2019 - 02.2021

Led coding, testing, deployment, and documentation of solutions for cloud-native infrastructure.

Utilized CI/CD tools including Jenkins, Maven, and Docker for automation and configuration management.

Maintained OpenStack Linux virtual machines with zero downtime while executing unit and system testing activities to ensure reliability.

Contributed to development of large-scale Java/Spring Batch/Hadoop systems using Docker, ensuring software met quality standards.

Designed, developed, and implemented web-based Java applications meeting enterprise business requirements.

Analyzed business requirements, creating technical design documents that aligned with company architecture to guide development.

Configured JFrog Artifactory and Git repositories in Bitbucket for version control management.

Developed security integration adhering to SAML v2.0, OAuth, and OpenID specifications.

DevOps Engineer

Change Healthcare
King of Prussia, Pennsylvania
07.2018 - 08.2019

Deployed and managed clustered ECS/EC2 instances, enhancing application availability and scalability.

Automated VMware infrastructure management using container-based solutions like Docker and Kubernetes, streamlining operations.

Implemented CI/CD pipeline automation through Jenkins, configuring Docker integration.

Created Ansible playbooks to automate AWS services, integrating seamlessly with the Apigee platform.

Wrote Terraform scripts for infrastructure automation, gaining hands-on experience with Chef and Puppet.

Developed API proxies using API Gateway technologies such as Apigee and Kong.

Oversaw operations, coordinating code deployments across development, testing, staging, and production environments to ensure smooth transitions.

Built real-time data pipelines utilizing Kafka producers and Spark streaming applications.

Software Engineer

Rivier University
Nashua, New Hampshire
08.2017 - 05.2018

Developed, tested, and implemented software programs that enhanced user functionality and experience.

Designed and updated software databases, including applications, websites, and user interfaces.

Tested, maintained, and recommended improvements for software optimization and performance.

Resolved complex technical design issues by conducting root cause analysis, improving system reliability.

Managed and streamlined maintenance of cloud servers and hosted services, ensuring optimal performance and uptime.

Utilized configuration management tools like Chef and Ansible for system efficiency.

Employed scripting languages such as Python and Bash to modify existing programs.

Collaborated with team on development processes using Jenkins, Git, and Bamboo.

Education

Master of Science - Management Information Systems

Rivier University
Nashua, NH
08-2018

Bachelor of Science - Computer And Information Sciences

Jawaharlal Nehru Technological University
India
05-2016

Skills

  • Infrastructure as code
  • Python scripting
  • AI/ML observability tools
  • Observability platforms
  • SLA/SLO/SLI management
  • CI/CD processes
  • Cloud architecture
  • Incident management
  • Dynatrace observability tools
  • Grafana visualization
  • Prometheus monitoring
  • ELK stack integration
  • Splunk analytics
  • Datadog performance monitoring
  • AppDynamics application performance
  • CloudWatch metrics
  • Fluentd data collection
  • OpsGenie incident response
  • Kafka for telemetry/event streaming
  • GCL management
  • Embrace platform integration
  • Sentry error tracking

Certification

  • WELD - Women for Economic and Leadership Development certification program
  • AWS Certified Solutions Architect - Associate
  • Azure Cloud Certified
  • AI & ML Certification
  • Academy Accreditation - Generative AI Fundamentals
  • Learning ITIL
  • Site Reliability Engineering: Service-Level Agreements and Objectives

Timeline

Senior Site Reliability Engineer (SRE)

Victoria's Secret
08.2022 - Current

Site Reliability Engineer

Ellucian
02.2021 - 08.2022

SRE/Application System Engineer

Comcast
08.2019 - 02.2021

DevOps Engineer

Change Healthcare
07.2018 - 08.2019

Software Engineer

Rivier University
08.2017 - 05.2018

Master of Science - Management Information Systems

Rivier University

Bachelor of Science - Computer And Information Sciences

Jawaharlal Nehru Technological University
Manisha Sona MenguSRE Engineer