06 set
|
Hermes Corporate
|
Italia
06 set
Hermes Corporate
Italia
ppWe are looking for an experienced bAI HPC Infrastructure Engineer /b to operate, optimise, and continuously evolve high-performance computing infrastructure supporting advanced engineering and scientific workloads. /p pThis is a hands-on infrastructure role for someone with strong Linux and HPC expertise who is comfortable owning production platforms, troubleshooting complex systems, and driving improvements across GPU compute, storage, networking, containers, and platform performance. /p h3What You’ll Do /h3 ul liAdminister and maintain GPU compute infrastructure, including system configuration, NVIDIA drivers and software stack, firmware, and hardware health monitoring. /li liOperate, monitor, and optimise HPC clusters to ensure reliability, performance, and efficient use of compute resources. /li liDefine and refine resource allocation policies across different workload types to ensure effective and equitable use of available capacity. /li liManage containerised environments for platform users, including user provisioning, environment maintenance, GPU access, and standards for container usage. /li liAdminister shared and high-performance storage and support the operation and troubleshooting of high-speed interconnects. /li liLead performance engineering activities, including profiling, benchmarking, bottleneck analysis, and platform optimisation. /li liSupport the integration and optimisation of engineering and scientific applications on GPU-accelerated and HPC platforms. /li liProvide incident response, troubleshooting, root-cause analysis, and change management, including planning and execution of maintenance activities. /li liMaintain platform monitoring, alerting, and operational reporting. /li liDevelop and maintain technical documentation, operational procedures, and platform standards. /li liContribute to security, architecture governance, and compliance activities.
/li /ul h3Required Qualifications /h3 ul liSubstantial professional experience in Linux systems engineering / administration at an advanced level. /li liProven experience operating production HPC environments, including cluster administration and large-scale parallel workloads. /li liHands-on experience administering GPU compute infrastructure, particularly NVIDIA-based environments and the associated software stack. /li liPractical experience with containerisation in HPC environments, including GPU device access and multi-node workloads. /li liStrong troubleshooting and performance-tuning skills across complex compute infrastructure. /li liDemonstrated ability to take end-to-end ownership of production infrastructure and act as a senior technical escalation point. /li liStrong analytical and problem-solving skills with a production-focused mindset. /li liProfessional working proficiency in English. /li /ul h3Preferred Qualifications /h3 ul liExperience supporting engineering simulation, scientific computing, or computational analysis workloadson accelerated infrastructure. /li liExperience with HPC scheduling and resource management. /li liExperience with monitoring and observability for compute infrastructure. /li liExperience with high-performance storage and networking/interconnect technologies. /li liExperience in an enterprise, industrial, or regulated environment. /li liExperience with platform automation, scripting, or infrastructure tooling. /li /ul h3Ideal Candidate /h3 pThe ideal candidate is a hands-on infrastructure engineer who combines deep Linux, HPC, and GPU expertise with strong operational ownership. /p pYou are comfortable working close to the hardware and software stack, diagnosing complex performance and reliability issues, and improving infrastructure used by demanding engineering and scientific workloads. You thrive in environments where reliability, performance, and technical depth matter, and you can operate effectively as a senior technical point of reference for the platform. /p /p #J-18808-Ljbffr
📌 HPC ENGINEER (Italia)
🏢 Hermes Corporate
📍 Italia