In this role you will own Day 2 fabric operations for NVIDIA’s Cloud Partner estates, guiding the reliability and performance of InfiniBand, RoCE/Ethernet, and NVLink/NVSwitch fabrics at scale. You’ll work with cross-functional teams and customers to architect, validate, and operate complex GPU-centric networks across multi-tenant environments. This is a hands‑on leadership role in a fast-moving, open‑source–first fabric ecosystem. You’ll drive operational excellence, incident response, and knowledge transfer to partner teams, shaping how large AI/HPC systems run reliably.
- Own Day 2 fabric operations across NVIDIA Cloud Partner fleets (NVLink/NVSwitch partition management, maintenance‑partition isolation, safe partition-change workflows)
- Manage switch software and firmware lifecycle including upgrades and rollout campaigns with minimal production disruption
- Support Day 1 fabric validation and acceptance (InfiniBand/UFM bring‑up, Spectrum‑X/RoCE config, burn‑in against MTBI/goodput targets)
- Minimise handover time to production by reducing duplicated validation across hardware bring‑up and partner operations
- Drive fleet‑wide fabric reliability through telemetry, fault detection, remediation, and root‑cause analysis of network‑induced job failures
- Provide consultative guidance and hands‑on troubleshooting across fabric stack (NICs/DPUs, switch OS, routing, Kubernetes networking)
and lead knowledge transfer with runbooks for partner teams
- Serve as technical leader for assigned accounts and present to executive stakeholders
- 5+ years in data center networking, fabric engineering, or large‑scale network operations
- Deep understanding of data center architectures and RDMA fabrics (InfiniBand, RoCE/Ethernet) including topology, routing, congestion control, and QoS
- Hands‑on experience with InfiniBand/UFM, Spectrum‑X Ethernet, Cumulus Linux and/or SONiC, ConnectX/BlueField NICs/DPUs, NVLink/NVSwitch on NVL72‑class platforms
- Strong Linux knowledge (RedHat/Ubuntu), switch OS internals, security, and HPC/AI traffic patterns
- Automation and Observability skills: Python/Bash, IaC tools (Ansible, Terraform), GitOps, Grafana/Loki/Prometheus
- Proven ability to measure and improve MTBI and job goodput, and to lead architectural reviews with executive stakeholders
- Experience with Kubernetes networking in GPU clusters and multi‑tenant network isolation concepts
- Strong consultative mindset with demonstrated capacity to transfer knowledge and create runbooks for partner teams
- Consultative mindset
- Strong communication and presentation skills
- Cross‑functional collaboration
- InfiniBand and UFM
- Spectrum‑X Ethernet
Senior Cloud Infrastructure and Network Operations Solutions Architect rovigo, veneto, it
#J-18808-Ljbffr
📌 Senior Cloud Infrastructure And Network Operations Solutions Architect (Rovigo)
🏢 NVIDIA
📍 Rovigo