Questa posizione è in Domyn
Il processo di selezione sarà interamente gestito Domyn.
--
We are looking for an experienced Site Reliability Engineer to join our growing team in Milan and help shape the future of our flagship project, Colosseum, one of Europe’s most powerful AI supercomputers, currently in development. Designed to run our proprietary AI models at scale, it forms the compute backbone behind the intelligence we deliver to the world’s most demanding industries. In this role, you will design and implement observability and control mechanisms that extract operational data from infrastructure and feed it into automated systems to enable continuous optimization, including key system budgets such as power, cooling and service level, security-level objectives. You will be responsible for actively guarding and maintaining these operational budgets as part of day-to-day system reliability and performance management. You will also contribute to operational excellence through blameless post-mortem analysis and structured incident learning, ensuring continuous improvement of system behavior and resilience. As a part of the team, you will work closely with Platform Engineering in a shared cybersecurity model, where SRE focuses on detection and monitoring, while Platform Engineering ensures the secure design and operation of the underlying infrastructure. ### What You Have * Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field. * At least 6 years of experience as a Site Reliability Engineer or in similar roles. * Strong experience with observability and monitoring systems such as Prometheus, Thanos, Grafana, and OpenTelemetry * Experience with low-level system instrumentation and performance visibility using technologies such as eBPF * Experience with security monitoring and threat detection tools such as Zeek, Wazuh, or equivalent SIEM / security observability platforms * Strong experience with containerized and cloud-native environments, particularly Kubernetes * Strong software development skills, particularly in Python, with the ability to build automation, integrations,
and custom tooling * Experience integrating heterogeneous infrastructure systems across multiple vendors, APIs, and evolving tool ecosystems * Familiarity with modern infrastructure automation and emerging agent-based frameworks such as MCP / A2A (or equivalent technologies) * Exposure to digital twin technologies and simulation platforms such as NVIDIA Omniverse or equivalent * Strong ability to design, build, and maintain software-driven infrastructure solutions in complex, large-scale environments ### Who You Are * A versatile engineer, comfortable operating in complex and fast-paced environments. * Driven and fearless, you proactively tackle challenges and overcome obstacles with determination. * A systems thinker, capable of understanding the broader architecture and identifying dependencies across platforms and technologies. * A collaborative team player who is enthusiastic, curious, and passionate about problem-solving, thriving both independently and within cross-functional teams. * An effective communicator with strong interpersonal skills, able to engage with diverse stakeholders and foster collaboration. * Fluent in English and eager to contribute in a multicultural and international environment. ### What we offer * **Compensation:** We offer a competitive base salary, as well as the opportunity to receive a variable premio and company equity. The typical base salary for this role ranges between €50.000 and €70.000, based on experience. As you gain experience and make more significant contributions to the business, your compensation will be reviewed to match your impact. Employment terms are governed by the CCNL Commercio, the Italian National Collective Bargaining Agreement for the Commerce, Distribution and Services sector.
* **Learning:** When our people grow, we grow with them. That’s why we give every team member a dedicated learning budget to invest in books, courses, conferences, or anything else that fuels their curiosity and supports their role. * **Flexible working:** We offer a flexible work policy that lets employees choose to work from home or the office, whatever suits them best. * **Wellbeing:** We provide resources to support the mental wellbeing of our team members. ### Why work at Domyn? Ambition, talent and teamwork-this is how we shape your future, combining a start-up mindset with the impact of a corporation. Domyn is an equal opportunity employer. We welcome applicants from all backgrounds and are committed to building a diverse and inclusive team. We welcome applicants of all backgrounds, including people with visible and non-visible disabilities, and we're committed to an accessible, bias-free hiring process for everyone. Domyn is committed to accommodating applicants with disabilities. Please let us know at
[email protected]. if you need accommodation during the application or interview process. All requests are kept confidential. ### About Domyn Domyn is a frontier AI company building sovereign AI for regulated industries. Our full-stack AI operating system brings together proprietary models, enterprise knowledge, industry-specific agents, governance, and infrastructure, giving organizations ownership and control of the intelligence they increasingly rely on. We work with some of the world’s leading banks, enterprises, and public institutions to build AI they can adapt, govern, and trust in mission-critical environments. Founded in Milan in 2016, Domyn brings together researchers, engineers, and industry experts across Europe, the United States, and India. We are united by the ambition to advance the frontier of AI and apply it to some of the world’s most important and complex challenges. At Domyn, people are trusted to think independently, move with purpose, and take real ownership of their work. Please review our Privacy Policy here https://bit.ly/4tndszN.
--
[#J-MIN]
📌 Senior Site Reliability Engineer (Milano)
🏢 Domyn
📍 Milano