10 ott
|
NVIDIA
|
Perugia
pIn this role you will own Day 2 fabric operations for NVIDIA’s Cloud Partner estates, guiding the reliability and performance of InfiniBand, RoCE/Ethernet, and NVLink/NVSwitch fabrics at scale. You’ll work with cross-functional teams and customers to architect, validate, and operate complex GPU-centric networks across multi-tenant environments. This is a hands-on leadership role in a fast-moving, open-source–first fabric ecosystem. You’ll drive operational excellence, incident response, and knowledge transfer to partner teams, shaping how large AI/HPC systems run reliably. /pulliOwn Day 2 fabric operations across NVIDIA Cloud Partner fleets (NVLink/NVSwitch partition management, maintenance-partition isolation, safe partition-change workflows) /liliManage switch software and firmware lifecycle including upgrades and rollout campaigns with minimal production disruption /liliSupport Day 1 fabric validation and acceptance (InfiniBand/UFM bring-up, Spectrum-X/RoCE config, burn-in against MTBI/goodput targets) /liliMinimise handover time to production by reducing duplicated validation across hardware bring-up and partner operations /liliDrive fleet-wide fabric reliability through telemetry, fault detection, remediation, and root-cause analysis of network-induced job failures /liliProvide consultative guidance and hands-on troubleshooting across fabric stack (NICs/DPUs, switch OS, routing, Kubernetes networking)
and lead knowledge transfer with runbooks for partner teams /liliServe as technical leader for assigned accounts and present to executive stakeholders /lili5+ years in data center networking, fabric engineering, or large-scale network operations /liliDeep understanding of data center architectures and RDMA fabrics (InfiniBand, RoCE/Ethernet) including topology, routing, congestion control, and QoS /liliHands-on experience with InfiniBand/UFM, Spectrum-X Ethernet, Cumulus Linux and/or SONiC, ConnectX/BlueField NICs/DPUs, NVLink/NVSwitch on NVL72-class platforms /liliStrong Linux knowledge (RedHat/Ubuntu), switch OS internals, security, and HPC/AI traffic patterns /liliAutomation and Observability skills: Python/Bash, IaC tools (Ansible, Terraform), GitOps, Grafana/Loki/Prometheus /liliProven ability to measure and improve MTBI and job goodput, and to lead architectural reviews with executive stakeholders /liliExperience with Kubernetes networking in GPU clusters and multi-tenant network isolation concepts /liliStrong consultative mindset with demonstrated capacity to transfer knowledge and create runbooks for partner teams /liliConsultative mindset /liliStrong communication and presentation skills /liliCross-functional collaboration /liliInfiniBand and UFM /liliSpectrum-X Ethernet /li /ulpSenior Cloud Infrastructure and Network Operations Solutions Architect • perugia, umbria, it /p #J-18808-Ljbffr
📌 Senior Cloud Infrastructure And Network Operations Solutions Architect (Perugia)
🏢 NVIDIA
📍 Perugia