Site Reliability Engineer
Description
Summary Job Summary (Site Reliability Engineer Shea, AZ): Support and maintain large-scale, high-performance enterprise applications across hybrid (on-premises and cloud) environments. Develop automation scripts to optimize operational processes and improve efficiency. Build and manage Application Performance Management (APM) dashboards for comprehensive transaction and system monitoring. Implement and enhance observability solutions using OpenTelemetry (OTEL), distributed tracing, and incident management tools. Manage and support containerized applications on Kubernetes platforms, participating in cloud migration and containerization projects. Monitor application health, troubleshoot production issues, and ensure high reliability and uptime of platforms. Work with relational and NoSQL databases (e.g., Oracle, PostgreSQL, MongoDB, Redis) to support production applications. Collaborate with cross-functional teams to enhance platform performance, scalability, and operational excellence. Participate in 24x7 on-call rotations, responding to incidents and ensuring timely resolutions per SLAs. Contribute to continuous improvement, focusing on reliability, automation, and best practices for highly available, customer-facing platforms.