22 ago
|
SUBMER
|
Provincia di Bergamo
22 ago
SUBMER
Provincia di Bergamo
**About Radian Arc.** **What impact you will have** Mission: Install, validate, operate, and maintain regional and core GPU cloud deployments across datacenter environments**.** This role owns the practical deployment and operational readiness of the platform stack in the field. It bridges infrastructure engineering, datacenter operations, networking, host systems, storage, and platform operations. The engineer is responsible for taking a validated architecture and BOM and turning it into a working production environment, including physical deployment coordination, rack and cabling validation, switch and host bring-up, firmware and BIOS validation, operating-system installation, GPU and DPU validation, storage integration, platform stack installation, acceptance testing, and operational handover. The role is intentionally hands-on and cross-domain. It is not a pure datacenter technician role, and it is not a pure platform engineering role. It is the person who can work in a datacenter, understand cabling and optics, debug host and network issues, validate GPU servers, support platform installation, and coordinate with engineering when problems arise. The role is especially important as the platform evolves from regional deployments toward HGX-based GPU systems, east-west fabrics, and core AI infrastructure. This profile is focused on actual installation, commissioning, maintenance, and operational readiness. **What you'll do** **Datacenter deployment and commissioning**: - Coordinate physical deployment activities with datacenter providers, integrators, logistics teams, and internal engineering.
- Validate rack layouts, elevations, power feeds, airflow assumptions, cable paths, and labeling before installation.
- Support rack-and-stack activities for GPU nodes, CPU nodes, storage nodes, switches, firewalls, routers,
serial/OOB equipment, PDUs, and supporting infrastructure.
- Validate fibre and copper cabling against the deployment design, including OOB, north-south, east-west, storage, and management networks.
- Validate optics, transceivers, link speeds, breakout cables, port mappings, and redundancy assumptions.
- Maintain accurate as-built documentation, including rack elevations, cable maps, port maps, serial numbers, asset records, IP allocations, and change records. **Host and hardware bring-up**: - Bring up GPU servers, platform servers, storage nodes, and supporting infrastructure.
- Validate BIOS, BMC, firmware, NIC, DPU, GPU, NVMe, RAID/HBA, and platform firmware versions.
- Configure and validate BMC access using Redfish/IPMI and OOB management networks.
- Validate GPU visibility, PCIe topology, NUMA layout, thermals, power behavior, and hardware health.
- Run hardware acceptance tests, burn-in tests, GPU stress tests, network tests, and storage validation before handover.
- Troubleshoot hardware issues across servers, GPUs, DPUs, NICs, optics, cables, disks, memory, firmware, and BIOS. **Network deployment support**: - Support deployment and validation of OOB, north-south, storage, and east-west networking.
- Validate BGP, ECMP, VLAN/VRF segmentation, EVPN/VXLAN where applicable, OVS/OVN integration, and routing reachability.
- Validate VyOS routers, OOB firewalls, transit routers, Citrix NetScaler/WAF,
and customer connectivity.
- Support RoCE/RDMA fabric validation for distributed AI workloads where applicable.
- Troubleshoot practical network issues such as link flaps, optics issues, incorrect polarity, MTU mismatches, route leaks, VLAN errors, packet loss, PFC/ECN issues, and fabric congestion.
- Support integration with NVIDIA Cumulus / Spectrum-X environments, and assist with Cisco or SONiC-based alternatives if those become part of the roadmap. **Platform stack installation and validation**: - Support installation and validation of the platform stack across regional and core deployments.
- Install and validate host operating systems, kernel versions, NVIDIA drivers, Mellanox/NVIDIA OFED or inbox drivers, CUDA compatibility, Docker/containerd, KVM/QEMU, and platform agents.
- Support CloudStack-based deployments and Kubernetes/KubeVirt-based deployments.
- Validate GPU passthrough, SR-IOV, BlueField NIC/DPU behavior, VM networking, and container networking.
- Support Kubernetes node registration, GPU Operator validation, CSI validation, CNI validation, and node lifecycle workflows.
- Support storage integration with StorPool, Weka, local NVMe, or other supported storage platforms.
- Execute acceptance tests and produce deployment readiness reports. **Operational maintenance and Day-2 support**: - Perform controlled maintenance activities such as firmware upgrades, switch upgrades, host OS updates, GPU driver updates, BIOS changes, and hardware replacements.
- Support incident response for infrastructure issues affecting GPU nodes, hosts, networking, storage, or platform components.
- Perform root-cause analysis for deployment and operational failures.
- Maintain runbooks f
📌 Senior Technical Operations & Deployment Engineer (Gpu Cloud Infrastructure) (Provincia di Bergamo)
🏢 SUBMER
📍 Provincia di Bergamo