Provide infrastructure architecture, engineering, operations, and technical leadership across HPC, cloud, container, application, monitoring, and enterprise technology environments. Collaborate with infrastructure, networking, security, DevOps/SRE, application, and vendor teams to implement reliable, scalable, and secure platforms supporting multiple concurrent users and business-critical workloads.
HPC and Infrastructure Engineering
- Managed and supported multiple HPC clusters across the United States, maintaining availability, performance, operational consistency, and effective coordination across geographically distributed infrastructure.
- Supported complex HPC infrastructure spanning compute, operating systems, networking, storage, applications, monitoring, and facility dependencies.
- Gathered cross-functional requirements for next-generation HPC infrastructure, translating user, application, performance, reliability, and operational needs into actionable implementation plans.
- Implemented next-generation HPC infrastructure solutions, partnering with engineering and infrastructure teams during design, deployment, testing, and operational transition.
- Troubleshot performance and reliability issues at both the application and system levels, analyzing infrastructure behavior, system dependencies, monitoring data, logs, and workload characteristics to identify root causes.
- Provided technical input on infrastructure integration topics involving compute, networking, storage, security, monitoring, and application platforms.
Monitoring, Observability, and Analytics
- Monitored infrastructure facilities and environmental conditions supporting critical technology operations; identified and helped address critical cooling issues to reduce infrastructure risk and improve operational stability.
- Developed monitoring standards, dashboards, alerting capabilities, and operational views to improve proactive issue detection and incident response.
- Implemented monitoring and analytics solutions using the ELK Stack, creating dashboards and visualizations for infrastructure, application, and operational analysis.
- Correlated metrics, logs, and system behavior to support root-cause analysis and performance improvement initiatives.
Cloud, Security, and Systems Integration
- Acted as the subject matter expert for onboarding team applications and AWS accounts through Zscaler, coordinating application, cloud, networking, security, and access requirements.
- Worked with cross-functional teams to integrate cloud-based applications and accounts into enterprise security and connectivity controls.
- Supported hybrid infrastructure environments spanning on-premises systems, cloud services, enterprise security platforms, and application dependencies.
- Contributed to the design and implementation of scalable infrastructure solutions that improved operational consistency and simplified support across environments.
Container Platforms
- Designed and implemented Docker Enterprise infrastructure and onboarded dozens of applications to the platform.
- Served as a subject matter expert for the enterprise team during Docker Enterprise adoption, helping define onboarding practices, resolve technical issues, and establish operational processes.
- Partnered with application and engineering teams to support container adoption and improve consistency between development, test, and production environments.
Technical Leadership and Team Development
- Onboarded three new full-time employees, helping grow the team to four members and establishing effective knowledge-transfer and operational support practices.
- Mentored team members on infrastructure operations, application onboarding, monitoring, troubleshooting, and support procedures.
- Developed and communicated technical approaches to engineering, infrastructure, security, and enterprise stakeholders.
Enterprise Application Architecture
- Helped establish the initial architecture and implementation approach for the Atlassian platform, which was subsequently adopted by enterprise IT.
- Collaborated with enterprise stakeholders to improve platform adoption, scalability, supportability, and integration with organizational processes.
Vendors and External Partners
- Worked directly with vendors and external partners to support software and hardware solutions, troubleshoot technical issues, coordinate upgrades, and resolve operational problems.
- Served as a technical liaison between internal teams and vendors, helping clarify requirements, evaluate proposed solutions, track issues, and validate remediation.
- Contributed technical input to infrastructure decisions involving performance, reliability, supportability, and long-term operational requirements.