03 ago
|
Experteer Italy
|
Bardi
03 ago
Experteer Italy
Bardi
In this role you ensure the reliability and performance of Kong's cloud services.
You will contribute to building scalable infrastructure and embedding reliability into the software lifecycle.
You'll tackle production incidents, implement robust monitoring, and drive automation to reduce toil.
You work with cross-functional teams to achieve high uptime and developer delight, shaping how we scale and secure our platform.
This opportunity offers impact through infrastructure as code, capacity planning, and incident-driven improvements.
Build and maintain core infrastructure as code using Terraform and Ansible
Design and operate monitoring, logging, and alerting to meet *****% uptime
Resolve production incidents with blameless post-mortems to prevent recurrence
Write automation to reduce operational toil and enable self-service for engineers
Collaborate with developers to embed reliability and scalability into the product lifecycle
Contribute to capacity planning, disaster recovery drills, and security hardening
Participate in a fair on-call rotation to maintain platform availability
Experience operating production workloads on major cloud providers (AWS, GCP, Azure)
Proficiency in at least one programming or scripting language (Golang, Python, Bash)
Hands-on experience with Docker and Kubernetes
Knowledge of Infrastructure as Code principles and tools (Terraform a plus)
Familiarity with CI/CD concepts and tools (GitLab CI, Jenkins)
Understanding of modern observability stacks (Prometheus, Grafana, ELK)
#J-*****-Ljbffr
📌 Site Reliability Engineer (Bardi)
🏢 Experteer Italy
📍 Bardi