MHMuhammad Hoque

Data Center & Infrastructure Technician

Muhammad Hoque

Data Center Technician at Google, working on hyperscale AI/ML infrastructure.

I diagnose and restore compute and network hardware, resolve Layer 0/1 faults, and train the technicians who keep these environments running — then document the work so the fix holds the next time.

B.S. Cybersecurity — WGU/ 6 certifications/ Previously AWS

2,730+tickets resolved across AWS and Amazon environments
115+technicians trained on data center operations
5,200+devices managed across IT operations
$65K+saved through repair and lifecycle recovery
01

Experience

Current and recent roles in data center operations and IT infrastructure.

Data Center Technician

Google

  • Support hyperscale data center operations and AI/ML infrastructure with a focus on operational reliability and disciplined execution.
Data Center OperationsAI/ML InfrastructureHyperscaleOperational Reliability

Data Center Technician

Amazon Web Services (AWS)

  • Maintained uptime, performance, and security of racks, servers, storage, and network gear.
  • Diagnosed and restored AI/ML TRNT2 compute hosts and J5 network devices.
  • Resolved Layer 0/1 faults including cabling, optics, PDUs, power, cooling, and capacity recovery issues.
  • Resolved, triaged, and deep-dived 1,500+ tickets.
  • Trained 115+ L2/L3/L4 technicians through hands-on training, SOPs, labs, and operational guidance.
  • Ran 50+ rack-down drills covering ToR Reload, ToR Replacement, Emergent ToR Replacement, and PSU Failures.
  • Documented RCAs, published runbooks, updated SOPs/MOPs, and improved repeatable recovery workflows.
  • Supported security practices including asset chain-of-custody, drive sanitization, patch compliance, access reviews, and clean documentation.
  • Managed 100+ laptops across the SBN100 campus in Indiana, maintaining 100% compliance and zero audit discrepancies; trained 10 team members and repaired 15+ devices, saving $20K+ in replacement costs.
TRNT2J5 NetworkingToR ReloadToR ReplacementPSU FailuresRCASOP / MOPSecurity Compliance

IT Equipment Coordinator / Technical Ops Support

Amazon — Michigan, United States

  • Managed 400+ laptops, 1,200 thin clients, 1,000 printers, and 3,000 scanners for warehouse operations.
  • Supported 3,000+ users with 99.8% customer satisfaction and resolved 1,230+ support tickets.
  • Repaired 100+ laptops and 100+ Zebra printers, saving more than $45,000 in replacement costs.
  • Coordinated $75,000+ in equipment procurement, vendor partnerships, returns, and disposal processes.
Inventory ManagementTechnical SupportProcurementVendor CoordinationDevice Repair
02

Selected work

Operations from the AWS role, with the reasoning behind each one.

115+

Technicians trained

Building operational capability across L2/L3/L4 teams.

Context

A hyperscale AI/ML environment needs technicians who can diagnose and recover hardware consistently, not just follow a checklist.

Approach

Delivered in-class and hands-on coaching, wrote and refined SOPs, ran labs, and led scenario-based troubleshooting for L2, L3, and L4 technicians.

Outcome

115+ technicians trained and mentored, with repeatable material left behind for the next cohort.

50+

Rack-down drills

Rehearsing recovery before the real incident.

Context

A rack going down is time-critical. Recovery quality depends on whether technicians have rehearsed the procedure beforehand.

Approach

Ran drills covering PSU failures, ToR reloads, ToR replacements, and emergent ToR replacement procedures.

Outcome

50+ drills delivered, building familiarity with the highest-urgency failure modes.

300

Port issues resolved

Systematic audit across a large network fabric.

Context

Port-level faults across a large fabric degrade capacity quietly, so they need systematic auditing rather than reactive ticket-by-ticket fixes.

Approach

Supported MCM audit and remediation across 113 bricks, 64 spines, 8 NDFs, and 50 fusion racks.

Outcome

300 port issues resolved across the audited footprint.

300+

BBUs and PSUs deployed

Power infrastructure for a campus launch.

Context

A campus launch requires power infrastructure staged and deployed correctly before capacity can come online.

Approach

Deployed infrastructure equipment supporting the SBN100 campus launch in South Bend, Indiana.

Outcome

300+ BBUs and PSUs deployed in support of the launch.

300+

Fiber cables labeled

Physical-layer readiness for GPU interconnect.

Context

High-speed inter-rack GPU connectivity depends entirely on physical-layer work being correct and traceable.

Approach

Labeled SC/LC fiber and patched OSFP-XD cables supporting high-speed inter-rack GPU connectivity.

Outcome

300+ cables labeled, with the physical layer documented for future troubleshooting.

MLZ

Training program support

Structured onboarding for incoming technicians.

Context

Scaling a technician workforce requires structured onboarding rather than ad-hoc shadowing.

Approach

Helped develop and support MLZ OJT/ILT training delivery, trainer guidance, and onboarding structure.

Outcome

A more consistent onboarding path for incoming technicians.

03

Capabilities

The technical ground covered day to day.

Data center infrastructure

Rack & StackServersStorageOpticsPDUsCablingLayer 0/1

AI/ML operations

TRNT2 HostsJ5 Network DevicesFleet ImagingFirmwareCapacity Recovery

Rackdown & recovery

ToR ReloadToR ReplacementEmergent ToRPSU FailuresRecovery Sequencing

Cybersecurity

Access ReviewsDrive SanitizationChain-of-CustodyPatch ComplianceAudit Readiness

IT operations

InventoryTechnical SupportProcurementVendor ManagementAsset Lifecycle

Documentation

RCASOPMOPRunbooksAudit RecordsTraining Labs
05

History

Full record, most recent first.

Data Center Technician

Google

Current

Data Center Technician

Amazon Web Services (AWS)

IT Equipment Coordinator

Amazon

FC Associate

Amazon — full-time

FC Associate

Amazon — part-time

Delivery Driver

Gig platforms

Cook

Captain Jay's

Hi-Lo Driver

NYX Inc.

06

Outside work

Interests and current reading.

I keep up with developments in technology and read widely on where the field is heading.

I also follow markets closely — studying investing strategies and testing new approaches of my own.

Currently reading — 2026 focus: health

The Body: A Guide for OccupantsBill Bryson
Good EnergyDr. Casey Means
Eat, Drink, and Be HealthyDr. Walter Willett
Food RulesMichael Pollan
07

Contact

Happy to talk about infrastructure, reliability, or security work.

Let's build something reliable.

I'm interested in work where I can improve uptime, support critical systems, strengthen operational security, train technical teams, and build infrastructure processes that hold up over time.

Or reach me directly at muhammadmhoque@gmail.com