DevOps / Site Reliability Engineer (SRE)
Description
Job Description: DevOps / Site Reliability Engineer (SRE) Company Overview: NextAmp LLC is a digital modernization services company focused on transforming the insurance industry. We help insurance organizations redesign their workflows and build AI-powered, automated systems that improve efficiency and accuracy. Our solutions streamline key operational areas such as claims processing, underwriting, billing, and back-office operations. Position : DevOps / Site Reliability Engineer (SRE) Location: USA (Remote) Experience level: 4+ years About the Role We are looking for a highly skilled DevOps / Site Reliability Engineer (SRE) to build, automate, and maintain reliable, scalable, and secure cloud infrastructure. The ideal candidate should have hands-on experience with CI/CD pipelines, containerization, infrastructure as code, cloud platforms, and production operations. Key Responsibilities Design, implement, and maintain CI/CD pipelines to enable reliable and automated software deployments. Build, manage, and optimize containerized applications using Docker and Kubernetes. Provision and manage cloud infrastructure using Infrastructure as Code (IaC) tools such as Terraform or CloudFormation. Deploy, monitor, and maintain applications on AWS or Azure. Ensure high availability, scalability, security, and reliability of production environments. Monitor application and infrastructure health, respond to incidents, and perform root cause analysis (RCA). Automate operational tasks to improve efficiency and reduce manual effort. Collaborate with development teams to improve deployment processes and application reliability. Implement monitoring, logging, alerting, and observability best practices. Required Skills 4+ years of experience in DevOps or Site Reliability Engineering (SRE). Strong experience building and managing CI/CD pipelines using tools such as Jenkins, GitHub Actions, Azure DevOps, or GitLab CI. Hands-on experience with Docker and Kubernetes. Strong knowledge of Infrastructure as Code (Terraform, CloudFormation, or similar). Experience with AWS or Azure cloud platforms. Experience with monitoring and observability tools such as Prometheus, Grafana, CloudWatch, Datadog, Splunk, ELK, or Azure Monitor. Good understanding of Linux system administration, networking, and security best practices. Experience with scripting using Bash, Python, or PowerShell. Strong troubleshooting and production incident management skills. Preferred Skills Experience with Helm, ArgoCD, or FluxCD. Knowledge of container security and vulnerability management. Familiarity with service mesh technologies (Istio, Linkerd). Experience with secrets management tools such as HashiCorp Vault or AWS Secrets Manager. Knowledge of high availability, disaster recovery, and backup strategies. N ice to Have Experience supporting 24x7 production environments. Exposure to SRE practices such as SLOs, SLIs, error budgets, and reliability engineering. Experience implementing cost optimization and cloud governance initiatives. Relevant certifications such as AWS Certified DevOps Engineer, AWS Solutions Architect, Certified Kubernetes Administrator (CKA), or Microsoft Azure DevOps Engineer Expert.