POSITION SUMMARY
We are seeking an experienced and highly motivated Contract Worker to support the installation, maintenance, and operational support of next-generation liquid-cooled AI rack infrastructure in a high-density datacenter environment. The ideal candidate will have direct hands-on experience with Direct Liquid Cooling (DLC) systems, high-density compute platforms — including Client Instinct GPU compute trays and EPYC server nodes — and the physical plant systems (CDUs, manifolds, busbar, networking fabric) required to sustain racks operating at up to 246 kW per rack.
This role supports the client Helios-class rack-scale AI infrastructure: open-standard double-wide (ORW) racks housing compute trays (4× Client Instinct MI455X GPUs + 1× EPYC Venice CPU each) and switch trays interconnected via a UALink / Ethernet fabric. The technician performs hands-on tray-level service, cooling system operations, firmware cycling, and physical cabling to sustain continuous availability of production AI training and inference workloads.
KEY RESPONSIBILITIES
Liquid Cooling Operations
Operate, monitor, and maintain Coolant Distribution Units (CDUs) including flow rate management (target: -385 L/min per rack system), temperature setpoints, and alarm response.
Install, replace, and pressure-test blind-mate quick-disconnect (QD) connectors on compute and switch trays without interrupting adjacent rack cooling circuits.
Perform leak detection, containment, and remediation in accordance with site EHS protocols.
Maintain secondary loop plumbing (supply/return manifolds, flexible hose assemblies, isolation valves); coordinate with facilities for primary loop support.
Support rack-level thermal assessments; contribute to coolant quality monitoring (pH, conductivity, biocide levels).
High-Density Rack & Server Hardware
Install and service client Helios-class compute trays (-77 kg each; 576 differential connections) using approved lift equipment and torque specifications.
Install and service switch trays (1,728 connections; -310 kg insertion force) using lever-assist handles per OEM guidance.
Replace field-replaceable units (FRUs): Client Instinct MI450/MI455X GPU modules, HBM4 memory, EPYC CPU+DIMM assemblies, NVMe drives, power modules, and fan trays.
Execute rack-level hot-swap operations on busbar and power shelf components following LOTO and high-voltage DC (HVDC) safety procedures.
Cable-manage high-speed optical interconnects (QSFP-DD / OSFP) and DAC/AOC assemblies for UALink, Ethernet, and out-of-band management fabrics.
Networking & Fabric Infrastructure
Patch and manage fiber/copper connections within the Broadcom Tomahawk 6-based switch fabric supporting single-hop, multi-plane GPU-to-GPU interconnect.
Validate optical power budgets and verify high-speed links following tray moves, adds, or changes.
Coordinate with network engineering for fabric topology changes and UALink / UEC configuration validation.
Monitoring, Diagnostics & Firmware
Use BMC/Redfish/IPMI tools to monitor GPU and CPU health, review system event logs, and identify hardware faults proactively.
Execute firmware update workflows (BIOS/UEFI, BMC, NIC, GPU firmware) following approved change management processes.
Run Client ROCm-based diagnostic utilities (rocm-smi, GPU benchmarks) to validate compute readiness post-maintenance.
Interface with DCIM and BMS platforms (e.g., Schneider EcoStruxure, Motivair CDU telemetry) to track PUE, coolant flow, and rack-level power draw (up to 246 kW/rack).
Documentation & Compliance
Maintain accurate asset inventory, maintenance logs, and incident records in the CMDB/ticketing system.
Follow EHS, physical security, and datacenter access procedures at all times.
Contribute to runbooks, SOP updates, and lessons-learned documentation after significant incidents or deployments.
Participate in on-call rotation for after-hours critical hardware incidents.
REQUIRED SKILLS & QUALIFICATIONS — SKILLS MATRIX
Skill Area - Details - Status
Liquid Cooling Systems - Direct Liquid Cooling (DLC), CDU operation, rear-door HX, manifold & distribution plumbing - Required
Blind-Mate Quick-Disconnects - Installation, replacement, and leak-testing of QD connectors under operational rack conditions - Required
High-Density Rack Platforms - OCP Open Rack / ORW form factor (client Helios-class); double-wide rack mechanics up to 246 kW - Required
GPU Compute Infrastructure - Client Instinct MI450/MI455X or NVIDIA equivalent; OAM/SXM mezzanine hardware; HBM4 awareness - Required
EPYC Server Hardware - EPYC (SP5/SP7) platforms; DIMM installation; BIOS/UEFI familiarity - Required
Electrical Safety (HV DC) - Hot-plug busbar systems, 48V/400V DC bus awareness, LOTO procedures - Required
Fiber & High-Speed Networking - QSFP-DD/OSFP optics, DAC/AOC cable mgmt, fabric interconnect (UALink, InfiniBand, RoCE) - Required
Physical Infrastructure - Rack installation, torque specs, cable management, PDU work, grounding - Required
BMC / IPMI / Redfish - Out-of-band management, firmware updates, power cycling, event log review - Required
ROCm Software Stack - Basic familiarity: rocm-smi, GPU health checks, diagnostic tooling - Preferred
OCP / Open Standards - OCP Open Rack, UALink, UEC familiarity; open telemetry standards - Preferred
DCIM / BMS Tools - Schneider EcoStruxure, Motivair, Vertiv, or equivalent DCIM / facility monitoring - Preferred
Certifications - BICSI RCDD/DCIS, CompTIA Server+, ASHRAE Class A4/W5 awareness, vendor certs (HPE, Dell) - Plus
EXPERIENCE REQUIREMENTS
Minimum: 4–5 years in a datacenter infrastructure or datacenter operations (DCO) technician role
Must Have
4+ years hands-on experience with Direct Liquid Cooling (DLC) systems in production datacenter environments, including CDU operation and coolant loop maintenance.
4+ years experience with enterprise server hardware installation and break/fix servicing at component level (CPU, DIMM, GPU, storage, NIC).
Demonstrated experience with high-density rack platforms (=20 kW/rack) including cabling, power distribution, and liquid-cooling integration.
Proven ability to follow and enforce LOTO, EHS, and physical security procedures in a critical facility environment.
Experience with out-of-band server management (BMC, IPMI, Redfish, iDRAC, iLO, or equivalent).
Familiarity with high-speed fiber optics (LC/MPO, QSFP-DD/OSFP) including cleaning, inspection, and patch panel management.
Strongly Preferred
2+ years experience with GPU compute infrastructure (Client Instinct, NVIDIA HGX, or equivalent OAM/SXM form factors).
Experience with the client EPYC server platforms (1P/2P) including BIOS, memory population, and diagnostic procedures.
Familiarity with OCP Open Rack / ORW standards and open datacenter architectures.
Exposure to AI or HPC datacenter deployments with cluster-scale interconnects (InfiniBand, Ethernet RoCE, or UALink).
Working knowledge of Client ROCm software stack for hardware validation tasks.
Nice to Have
BICSI RCDD, DCIS, or equivalent datacenter infrastructure certification.
CompTIA Server+, Linux+, or similar vendor/industry certifications.
Experience with Schneider Electric EcoStruxure, Motivair, or Vertiv CDU/DCIM platforms.
Familiarity with change management frameworks (ITIL) and CMDB tooling (ServiceNow or equivalent).
Notes:
Work Schedule: Onsite, 5 days a week Plus weekend support
Up to 10 hours a week of OT to be expected in this role
Occasional travel to other client location
VIVA is an equal opportunity employer. All qualified applicants have an equal opportunity for placement, and all employees have an equal opportunity to develop on the job. This means that VIVA will not discriminate against any employee or qualified applicant on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status