Data Center Technician
- Support hyperscale data center operations and AI/ML infrastructure with a focus on operational reliability and disciplined execution.
Data Center & Infrastructure Technician
Data Center Technician at Google, working on hyperscale AI/ML infrastructure.
I diagnose and restore compute and network hardware, resolve Layer 0/1 faults, and train the technicians who keep these environments running — then document the work so the fix holds the next time.
Current and recent roles in data center operations and IT infrastructure.
Amazon Web Services (AWS)
Amazon — Michigan, United States
Operations from the AWS role, with the reasoning behind each one.
Building operational capability across L2/L3/L4 teams.
A hyperscale AI/ML environment needs technicians who can diagnose and recover hardware consistently, not just follow a checklist.
Delivered in-class and hands-on coaching, wrote and refined SOPs, ran labs, and led scenario-based troubleshooting for L2, L3, and L4 technicians.
115+ technicians trained and mentored, with repeatable material left behind for the next cohort.
Rehearsing recovery before the real incident.
A rack going down is time-critical. Recovery quality depends on whether technicians have rehearsed the procedure beforehand.
Ran drills covering PSU failures, ToR reloads, ToR replacements, and emergent ToR replacement procedures.
50+ drills delivered, building familiarity with the highest-urgency failure modes.
Systematic audit across a large network fabric.
Port-level faults across a large fabric degrade capacity quietly, so they need systematic auditing rather than reactive ticket-by-ticket fixes.
Supported MCM audit and remediation across 113 bricks, 64 spines, 8 NDFs, and 50 fusion racks.
300 port issues resolved across the audited footprint.
Power infrastructure for a campus launch.
A campus launch requires power infrastructure staged and deployed correctly before capacity can come online.
Deployed infrastructure equipment supporting the SBN100 campus launch in South Bend, Indiana.
300+ BBUs and PSUs deployed in support of the launch.
Physical-layer readiness for GPU interconnect.
High-speed inter-rack GPU connectivity depends entirely on physical-layer work being correct and traceable.
Labeled SC/LC fiber and patched OSFP-XD cables supporting high-speed inter-rack GPU connectivity.
300+ cables labeled, with the physical layer documented for future troubleshooting.
Structured onboarding for incoming technicians.
Scaling a technician workforce requires structured onboarding rather than ad-hoc shadowing.
Helped develop and support MLZ OJT/ILT training delivery, trainer guidance, and onboarding structure.
A more consistent onboarding path for incoming technicians.
The technical ground covered day to day.
All certifications are publicly verifiable through Credly.
Penetration testing, vulnerability assessment, and reporting.
Security concepts, risk management, controls, and secure operations.
Networking concepts, infrastructure, troubleshooting, and operations.
IT support, hardware, operating systems, and device troubleshooting.
AWS cloud fundamentals, security, services, and architecture.
Security administration, access controls, and operational security.
Western Governors University
Jul 2022 — Jul 2023
Full record, most recent first.
Amazon Web Services (AWS)
Amazon
Amazon — full-time
Amazon — part-time
Gig platforms
Captain Jay's
NYX Inc.
Interests and current reading.
I keep up with developments in technology and read widely on where the field is heading.
I also follow markets closely — studying investing strategies and testing new approaches of my own.
Currently reading — 2026 focus: health
Happy to talk about infrastructure, reliability, or security work.
I'm interested in work where I can improve uptime, support critical systems, strengthen operational security, train technical teams, and build infrastructure processes that hold up over time.
Or reach me directly at muhammadmhoque@gmail.com