Senior Site Reliability Engineer

OptomiCharlotte, United States
ContractOn-siteMidLimited info disclosed
35 views0 applications

Description

Senior Site Reliability Engineer (AI Focused) (Hybrid | Charlotte, NC | Long- Term Contract) Overview: Optomi, in partnership with a leading financial company, is seeking a Senior Site Reliability Engineer to join a highly mature, modern SRE organization focused on reliability, observability, automation, and AI-driven engineering practices. This is an opportunity for a senior engineer who can immediately take ownership of monitoring, alerting, and operational readiness initiatives while partnering closely with development teams to improve the reliability and performance of critical enterprise applications. This is not a traditional operations role. We are looking for an engineer with a strong software development, DevOps, Platform Engineering, or SRE background who can lead conversations around observability, SLOs, SLIs, incident reduction, and automation. The ideal candidate is someone who can quickly step in and own conversations around monitoring strategy, alerting, reliability, remediation, and operational excellence. They bring a strong engineering background, understand how modern software systems are built and supported, actively leverage AI in their day-to-day workflow, and have the confidence to influence technical direction across teams. Responsibilities: Own and continuously improve monitoring, alerting, observability, and remediation strategies across critical production systems. Lead discussions around Service Level Objectives (SLOs), Service Level Indicators (SLIs), operational readiness, and application reliability. Partner with software engineering teams to ensure applications are designed, deployed, and supported with observability and reliability best practices. Design and implement proactive alerting strategies that reduce noise, improve signal quality, and minimize customer impact. Develop automation and tooling using modern engineering practices to eliminate operational toil and improve efficiency. Drive incident response, root cause analysis, and production troubleshooting efforts for complex distributed systems. Create and maintain operational standards, runbooks, dashboards, and reliability best practices. Mentor junior engineers and provide technical leadership across reliability initiatives. Leverage AI-powered tools and workflows to accelerate troubleshooting, development, automation, and operational decision-making. Contribute to continuous improvement initiatives within a mature Dynatrace-based observability environment. Qualifications: 5+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps Engineering, Production Engineering, or Software Engineering with production ownership responsibilities. Strong experience building and improving monitoring, alerting, observability, and operational readiness processes rather than simply responding to incidents. Experience owning or driving SLO, SLI, reliability, incident management, and production support initiatives. Hands-on software development or automation experience using Python, Go, Java, or similar programming languages. Experience supporting cloud-based environments, preferably AWS. Strong troubleshooting and root cause analysis skills within enterprise-scale distributed systems. Experience collaborating directly with development teams to improve application reliability and operational maturity. Active use of AI tools such as Claude, ChatGPT, Gemini, Cursor, GitHub Copilot, Codex, MCPs, or similar technologies to improve engineering productivity, automation, troubleshooting, or development workflows. Proven ability to communicate effectively, lead technical discussions, challenge assumptions constructively, and influence stakeholders. Demonstrated ownership mentality with the ability to identify problems, propose solutions, and drive execution independently. Preferred Qualifications: Dynatrace experience. Experience with Datadog, Splunk, AppDynamics, New Relic, or similar observability platforms. Terraform and Infrastructure as Code experience. Experience supporting cloud-native applications and modern platform engineering environments. Experience building or supporting AI-enabled workflows, AI-assisted development practices, agent-based tooling, or operational automation solutions. Show more Show less