Seasoned Data Center Technician with proven ability to troubleshoot hardware and software issues, implement maintenance strategies, and maintain optimal system performance, aiming to contribute to the efficient and reliable operation of your data center.
Experienced with managing data center operations and ensuring system stability. Utilizes advanced knowledge in network configuration and hardware maintenance to optimize performance. Track record of successfully troubleshooting complex technical issues and maintaining operational continuity.
Overview
12
12
years of professional experience
Work History
Sr. Data Center Engineer
Penguin Solutions / Meta (Research Super cluster)
07.2023 - Current
Supported platform health for Meta production data center infrastructure by resolving tickets (~80/month) with durable solutions to prevent recurrence and maximize uptime.
Performed evidence-based Linux server triage via SSH/IPMI and hands-on troubleshooting, using logs and health signals to isolate and remediate OS, hardware, and network faults.
Executed end-to-end hardware break/fix of FRU components such as GPUs, SSDs, NICs, DIMMs: averaging ~30 swaps/month, with 85% SLA compliance and only 2% of tickets reopened.
Owned root cause analysis for a platform-introduction blocker, patched the underlying script logic, and regression-tested across new and legacy hardware to prevent recurrence.
Took proactive ownership of H200 expansion fallout; became the team's go-to resource for new-platform issues by documenting gotchas and unblocking peers to close tickets faster.
Drove ~17 NVIDIA RMAs/month using evidence-first diagnostics, concise technical narratives, and explicit approval tasks; reducing back-and-forth and achieving ~1-day approvals.
Used SLA remaining as a primary operational KPI to plan daily work, shifting effort to highest-risk tickets, unblocking dependencies early, and preventing avoidable SLA breaches.
Performed racking, stacking, cable management, and OS provisioning for new server deployments.
Secured physical data assets, including servers and networking hardware, against theft, tampering, and environmental threats.
Optimized data center airflow by implementing hot/cold aisle containment.
Performed physical security and inventory control, including secure removal and disposal of failed hard drives.
Sr. Data Center Engineer
Maintech / Bank of America
09.2014 - 06.2023
Diagnosing and resolving hardware and software issues on servers, storage, and networking equipment.
Handling Installations, Moves, Additions, and Changes related to data center equipment.
Providing remote or on-site technical assistance for network and other infrastructure-related tasks.
Identifying and resolving technical problems, often involving fault isolation and part replacement.
Maintaining accurate records of activities, including using ticketing systems and other internal tracking tools.
Keeping track of client-provided parts and inventory according to established procedures.
Adhering to client policies, procedures, and best practices.
Working with internal engineering teams, vendors, and other stakeholders.
Unbox, inspect for damage, and inventory all hardware components before installation.
Calculate power requirements and connect PDU, considering redundancy requirements.
Identify efficient routes, avoiding obstacles and keeping fiber separated from electromagnetic interference (EMI) sources like power cables.
Conducted monthly audits and physical risk assessments, addressing any critical vulnerabilities.
Performed preventative and predictive maintenance on electrical infrastructure to ensure maximum redundancy and eliminate unplanned outages.
Resolved high-priority hardware incidents (break/fix) for Dell PowerEdge and HP ProLiant servers.
Education
Bachelor's Degree - Information Systems and Cyber Security
ITT Technical Institute
Richmond, VA
10-2013
Military - undefined
U.S.M.C.
05-2010
Skills
Computer: SSH, ipmi, bash, tmux, GitHub, Office Suite
Experienced with Windows and CentOS Stream operating systems
Experience with HP, Dell, IBM, Tyan, and Nvidia systems
Operating system administration
Customer support
Hardware troubleshooting
Automation tools
Incident management
Asset management
Teamwork and collaboration
Critical thinking
Task prioritization
Timeline
Sr. Data Center Engineer
Penguin Solutions / Meta (Research Super cluster)
07.2023 - Current
Sr. Data Center Engineer
Maintech / Bank of America
09.2014 - 06.2023
Military - undefined
U.S.M.C.
Bachelor's Degree - Information Systems and Cyber Security