Overview
As a Site Reliability Engineer at Kong, you will help build and operate the core cloud infrastructure that powers our API and AI platform. You will work with cross-functional teams to maintain high reliability, performance, and fast feature delivery. You will implement automation, observability, and capacity planning to scale our services. This role offers impact through shaping reliability at scale and contributing to blameless incident resolution and continuous improvement.
Responsabilità
- Develop and maintain infrastructure as code using tools like Terraform and Ansible
- Build robust monitoring, logging, and alerting to achieve 99.99% uptime
- Investigate and resolve production incidents with blameless post-mortems
- Automate operations to reduce toil and enable self-service for engineers
- Collaborate with developers to embed reliability and scalability into the app lifecycle
- Contribute to capacity planning, DR drills, and security hardening
- Participate in a fair, sustainable on-call rotation
Requisiti fondamentali
- Experience operating production workloads on AWS, GCP, or Azure
- Proficiency in at least one language (Golang, Python, or Bash)
- Hands-on with Docker and Kubernetes
- Knowledge of Infrastructure as Code (Terraform a plus)
- Familiarity with CI/CD concepts and tools (GitLab CI, Jenkins)
- Understanding of modern observability stacks (Prometheus, Grafana, ELK)
- Collaborative mindset
- Problem-solving under pressure
- Blameless communication and post-mortem participation
- Terraform
- Ansible
- Docker
📌 Site Reliability Engineer (Monza)
🏢 Kong
📍 Monza