SUSE is a global leader of enterprise open source software.
By transforming community innovations into secure, sovereign and AI-ready solutions, SUSE empowers customers to escape vendor lock-in and regain control of their IT destiny.
Through industry-leading Linux, Kubernetes, Edge and AI infrastructure solutions, SUSE delivers the flexibility to innovate everywhere—from the data center to multi-cloud and out to the edge.
Only SUSE also manages many Linux and Kubernetes distributions.
At SUSE, Choice Happens because we prioritize community, interoperability and relentless innovation.
Platform Engineer, AI
SUSE Internal IT is hiring Platform Engineers to join the team building and operating our internal Agentic AI Platform.
This is an hands-on engineering role.
You will be equally responsible for building new platform capabilities, keeping the platform operationally healthy, and maintaining the infrastructure-as-code and documentation that underpins it.
You will work with one other engineer as a pair: sharing ownership of the full platform, peer-reviewing each other's work, and developing complementary depth across the stack over time.
The platform is in active delivery.
You will join at a point where the core infrastructure is running and the next phase of security hardening, automation, and observability is under way.
Implement new platform capabilities from architectural designs, translating security, governance, and infrastructure requirements into production-grade infrastructure-as-code
Design and build the platform security and secrets management layer, ensuring all workloads operate with least-privilege credentials and certificates issued through a governed PKI hierarchy
Implement and enforce security policy across the cluster using admission control, covering workload configuration, image standards, network traffic, and resource constraints
Build and establish the platform observability stack, providing consistent log aggregation, metrics, distributed tracing, and alerting across all platform components
Design and implement GitOps delivery automation, ensuring all platform changes flow through version-controlled, auditable pipelines with drift reconciliation
Build and configure workload autoscaling, ensuring AI workflow workers scale efficiently and cost-effectively in response to demand
Implement the AI model routing and gateway layer, enabling governed, auditable routing of model traffic with per-consumer rate limiting
Own the day-to-day operational health of the platform: monitor for issues, respond to incidents, conduct root-cause analysis, and implement lasting remediation
Maintain the health of platform data services — database cluster, job queue, and object storage — including backup schedules, failover testing, and capacity management
Monitor and tune autoscaling and resource configuration as workload patterns evolve, ensuring the platform scales responsively without over-provisioning
Manage secrets rotation, certificate lifecycle, policy drift detection, and identity configuration as ongoing operational responsibilities
Own and evolve the infrastructure-as-code for your areas of the platform; keep all configurations versioned, peer-reviewed, and aligned with the architectural design
Proactively identify and resolve technical debt — manual processes, undocumented configurations, legacy credential management, and gaps in observability coverage
Produce and maintain operational runbooks for all platform procedures, ensuring any team member can execute them safely and independently
Peer-review all platform infrastructure changes produced by your engineering counterpart, providing challenge and quality assurance across the full stack
Contribute to platform documentation and knowledge-sharing, supporting the wider team's understanding of the platform as it matures
Candidates will need to demonstrate hands-on production delivery experience, not just conceptual familiarity.
Kubernetes — production cluster operation (RKE2, EKS, GKE, or equivalent); Secrets management — production deployment of a secrets management platform (HashiCorp Vault or equivalent), covering PKI, dynamic credentials, and workload secrets injection
Policy-as-code — admission control policy authoring and enforcement in production Kubernetes environments (OPA/Rego, Kyverno, or equivalent)
declarative drift reconciliation, rollback strategy, multi-environment targeting
Observability stack — log aggregation, log pipeline design, distributed tracing (OpenTelemetry or equivalent), and metrics dashboards (Prometheus/Grafana or equivalent)
API gateway engineering — production deployment and operation of an API or AI gateway (Kong, Envoy, or equivalent); Linux platform engineering — networking fundamentals, TLS and PKI, CSI storage operations, container runtime
Information Technology
SUSE is a dynamic environment that is evolving rapidly, thus requiring agility, strong entrepreneurship and an open mind.
If you're a big thinker, obsessed by execution and thrive in a dynamic environment in which you can tangibly create a lasting legacy, then please apply now!
You will work in a global community of unique individuals – like you – with different backgrounds, talents, skills and perspectives.
A truly open community where everyone is welcome, has a voice and is encouraged to reach their full potential regardless of age, gender, race, nationality, disability, sexual orientation, religion, or any other characteristics.
In the meantime, stay updated on the latest SUSE news and job vacancies by joining our Talent Community .
SUSE Values
SUSE's culture is centered on four key values - Choice, Community, Trust, and Innovation - which are deeply integrated with our open source ethos.
SUSE fosters a diverse and inclusive environment where our people are encouraged to be themselves.
We are "open source first, upstream first" where collaboration benefits all
We offer trust by default, and do not wait for it to be earned
Innovation
We are committed to continuous improvement, creativity and adaptability
📌 Platform Engineer, Ai (Milano)
🏢 Suse
📍 Milano