Site Reliability Engineer

Evlo AISeattle, United States
Full TimeOn-siteMidLimited info disclosed
15 views0 applications

Description

About The Role The role owns the reliability, scalability, and observability of core production infrastructure supporting millions of users daily. The team works alongside backend and platform engineers to build robust distributed systems where latency, high availability, and automated recovery are critical. Key Responsibilities Design, build, and maintain production Kubernetes clusters and underlying cloud infrastructure using Terraform Define and track critical service level objectives (SLOs), service level indicators (SLIs), and comprehensive alerting pipelines Automate incident response, failover mechanisms, and routine operational tasks using Python or Go Conduct root cause analysis for production incidents and implement preventative architectural improvements Manage continuous integration and continuous deployment (CI/CD) pipelines to ensure safe, zero-downtime releases Collaborate with engineering teams to review system architectures for scalability, security, and performance bottlenecks What We Are Looking For 3–6 years of experience in Site Reliability Engineering, DevOps, or systems engineering in cloud-native environments Strong hands-on experience with Kubernetes, Docker, and container orchestration at production scale Deep proficiency with Infrastructure as Code tooling, specifically Terraform and Ansible Solid understanding of networking concepts, TCP/IP, DNS, load balancing, and secure cloud architectures Experience with observability platforms such as Prometheus, Grafana, Datadog, or OpenTelemetry Bonus: Contributions to open-source infrastructure projects or certifications such as CKA or AWS Certified DevOps Engineer