Site Reliability Engineer
Description
OpenArc - Empowering Your Career. As a leading IT staffing firm, we are dedicated to connecting talented professionals with your ideal opportunities. We are currently seeking a qualified Site Reliability Engineer (AI & Agentic Systems) to join our client’s organization and contribute to their ongoing success. Job summary This role demands strong hands-on ownership of production reliability and troubleshooting, coupled with advanced capabilities in AI- and agentic-driven automation and performance engineering. The Site Reliability Engineer will play a critical role in ensuring reliability, scalability, performance, and operational excellence of our platforms. The ideal candidate will leverage Azure-native AI services and agentic systems to reduce toil, improve incident response, and enable intelligent operations—while also driving performance testing practices to validate system resilience under load. Responsibilities: Own end-to-end reliability of large-scale, Azure-hosted production systems, ensuring high availability, fault tolerance, and graceful degradation Lead hands-on incident troubleshooting, root cause analysis (RCA), and post-incident reviews with actionable follow-ups Build and operate resilient, scalable services on Microsoft Azure (AKS, App Services, Functions, Event Hubs, etc.) Design and maintain comprehensive observability platforms using Prometheus for metrics, Loki for log aggregation, Tempo for distributed tracing, and Grafana for dashboarding and alerting Design, develop, and execute performance testing strategies for distributed systems and microservices, including load testing, stress testing, soak testing, and capacity planning Integrate AI agents with Azure monitoring stack, CI/CD tooling, and incident management platforms Contribute to evolving SRE standards, tooling, operational processes, and knowledge base Requirements: Experience in Site Reliability Engineering, DevOps, or Production Engineering roles Strong hands-on experience in production troubleshooting of distributed systems at scale Solid understanding of Linux internals, networking (TCP/IP, DNS, HTTP, TLS), and system performance tuning Deep hands-on experience with Microsoft Azure (compute, networking, storage, managed services, AKS) Strong knowledge of Kubernetes, container orchestration, Helm charts, and microservices architectures Proficiency in one or more programming languages: Python, Go, Java, or equivalent Experience with CI/CD pipelines (Azure DevOps, GitHub Actions) and Infrastructure as Code (Terraform, ARM Templates, Bicep) At OpenArc, we prioritize your career success and strive to build exceptional technical teams for our clients. By understanding your experience and aspirations, we ensure to present you with rewarding and fulfilling opportunities. As an employee of OpenArc and our clients, you will be eligible to participate in a comprehensive benefits package. OpenArc is an equal opportunity employer and will consider all applications without regards to race, sex, age, color, religion, national origin, veteran status, disability, sexual orientation, gender identity, genetic information or any characteristic protected by law. Show more Show less