Work Preference
Summary
Work History
Education
Skills
Timeline
web
Rich Fahey
Open To Work
Verified
This profile is verified using an email address.

Rich Fahey

Site Reliability Engineer
Belmont,MA

Work Preference

Desired Job Title

Site Reliability EngineerCloud EngineerCloud Infrastructure EngineerCloud DevOps EngineerPlatform Engineer

Work Type

Full Time

Location Preference

On-SiteRemoteHybrid
Location: Belmont, MA, US
Open to relocation: No

Summary

Senior Site Reliability and Platform Engineer specializing in highly available AWS and Kubernetes environments, infrastructure automation, observability, and operational reliability. Experienced in designing and scaling cloud platforms, optimizing infrastructure costs, leading complex incident response, and building secure, multi-tenant systems in regulated environments. Strong expertise in AWS, Terraform, Kubernetes, and modern observability, with hands-on experience applying LLMs and autonomous agents to automate engineering and operational workflows.

Work History

Site Reliability Engineer II

Mimecast
Lexington, MA
10.2022 - 04.2026
  • Developed AWS environments for new services with integrated cross-account networking and GCP Apigee API Gateway to AWS connectivity.
  • Authored Infrastructure as Code in Terraform covering AWS VPC, Route 53, EKS, IAM, ELB, S3, EFS, CloudWatch, SNS, and more.
  • Authored Helm charts for Kubernetes/Prometheus deployments, configured CloudWatch and Prometheus alerting to JSM and Slack via AWS SNS.
  • Implemented Jenkins CI/CD pipelines with GitLab triggers, enhancing deployment efficiency and security.
  • Conducted cost analysis, designed and implemented a tiered archiving service for migrating on-premises legacy data to S3 and S3 Glacier Deep Archive, achieving significant cost savings.
  • Designed and deployed autonomous AI agents using LLMs to automate workflows and increase operational efficiency.
  • Engineered custom software solutions by leveraging LLMs for rapid prototyping, code generation, and script optimization.
  • Defined SLOs/SLIs and created the relevant dashboards and alerting.
  • Participated in on-call rota. Led incident responses and incident reviews.

Site Reliability Engineer II

Everbridge
Burlington, MA
07.2021 - 05.2022
  • Collaborated with development teams to implement infrastructure-as-code, improving deployment consistency for cloud services.
  • Designed architecture template for internal routing of S3 traffic within VPCs to bolster security.
  • Monitored vulnerabilities, ensuring compliance with FedRAMP standards and mitigating potential security risks.
  • Developed deployment and rollback methodologies; executed deployments and conducted post-deployment smoke tests.
  • Performed troubleshooting of cloud service issues, identifying root causes and recommending monitoring strategies.
  • Audited systems to identify and eliminate obsolete infrastructure, enhancing overall system performance.
  • Participated in on-call rotation for comprehensive support of cloud services operations.

Network Engineer

Limelight Networks
Burlington, MA
03.2014 - 03.2021
  • Performed troubleshooting for streaming media, network storage ingress/egress, storage data replication and cloud application deployments.
  • Led bug resolution efforts across development teams, implementing workarounds and tracking fixes to enhance system reliability.
  • Developed and maintained Python-based network diagnostic tools, enabling operations team to streamline troubleshooting processes.
  • Measured and reported key performance indicators (KPIs) to identify improvement opportunities.
  • Defined and developed consistent release processes for all products to enhance efficiency.
  • Acted as the final escalation point for customer support incidents across 170 global CDN locations.

Education

Bachelor of Science - Computer Science

Merrimack College
North Andover, MA

Skills

  • Cloud and Infrastructure: AWS, Terraform, Helm, Puppet, Linux
  • Automation and Delivery: Kubernetes, Docker, Jenkins, GitLab, Maven, Artifactory, ECR, Python, Bash, Groovy
  • API & Gateway Management: GCP Apigee
  • Observability and Reliability: Prometheus, Thanos, Graphite, Grafana, Splunk, Defining SLOs/SLIs
  • Security and Compliance: SonarQube, Black Duck, FedRAMP
  • Operations & Collaboration: Incident Management, Jira, Confluence, JSM, ServiceNow
  • Emerging Technology: LLMs (automation agents, engineering)

Timeline

Site Reliability Engineer II

Mimecast
10.2022 - 04.2026

Site Reliability Engineer II

Everbridge
07.2021 - 05.2022

Network Engineer

Limelight Networks
03.2014 - 03.2021

Bachelor of Science - Computer Science

Merrimack College
Rich FaheySite Reliability Engineer