Senior Site Reliability Engineer
Description
Senior Site Reliability Engineer (Hybrid – Charlotte, NC) Optomi, in partnership with a leading insurance and financial services organization, is seeking a Senior Site Reliability Engineer (SRE) to join a highly mature, AI-driven engineering organization in Charlotte, NC. This engineer will play a key role in owning observability, alerting, monitoring, and production reliability initiatives across critical enterprise platforms. The ideal candidate will bring a strong blend of software development, cloud, and SRE experience, with a passion for leveraging AI to improve engineering workflows and operational excellence. This is a highly collaborative role focused on partnering with stakeholders, driving SLO/SLI conversations, supporting production services, and helping shape the future of AI-enabled Site Reliability Engineering practices. This position is hybrid and requires onsite presence in Charlotte three days per week. What the right candidate will enjoy Hybrid environment! Owning reliability, observability, and production support initiatives within a highly mature and modern SRE organization. Leveraging AI-driven tools and workflows to improve engineering efficiency, automation, and service reliability. What type of experience does the right candidate have: 5+ years of Site Reliability Engineering, DevOps Engineering, Software Engineering, or related experience Strong hands-on coding experience with Python, Go, or similar programming languages Prior experience in software development or engineering-focused DevOps/SRE environments Experience building, managing, and optimizing observability and monitoring solutions Strong understanding of alerting frameworks, incident response, and production support processes Experience with cloud platforms, preferably AWS Familiarity with SLOs, SLIs, service ownership, and reliability engineering best practices Experience working in modern engineering organizations that emphasize automation and continuous improvement Strong communication skills with the ability to engage technical and non-technical stakeholders Proven ability to take ownership of services, systems, or technical domains Experience leveraging AI tools to improve engineering workflows and productivity Understanding of AI-assisted development practices and modern coding workflows Nice-to-have experience: Experience with Dynatrace, Datadog, or similar observability platforms Exposure to Terraform or other Infrastructure-as-Code technologies Experience working with AI tools such as Claude, Gemini, Cursor, Codex, or similar platforms Background supporting AI-powered applications, platforms, or workflows Experience mentoring junior engineers and leading technical initiatives What the responsibilities are of the right candidate: Own and continuously improve alerting, monitoring, observability, and reliability practices across enterprise services Design, implement, and maintain effective alerting strategies that reduce noise and improve operational awareness Partner with development, infrastructure, and product teams to define and manage SLOs and SLIs Support production services through troubleshooting, incident response, root cause analysis, and remediation efforts Develop automation solutions and tooling to improve reliability, scalability, and operational efficiency Contribute to cloud platform initiatives and engineering enablement efforts Leverage AI tools and emerging technologies to enhance engineering productivity and operational outcomes Drive continuous improvement initiatives across monitoring, observability, and production support processes Collaborate with stakeholders to identify reliability risks and implement proactive solutions Mentor junior engineers and provide technical guidance across the organization Participate in architectural discussions and influence reliability-focused engineering decisions Maintain documentation, operational standards, and best practices supporting enterprise reliability goals Show more Show less