Site Reliability Engineer (SRE)

STAFFXPERT LLCOrlando, United States
ContractOn-siteMidLimited info disclosed
34 views0 applications

Description

Location: Orlando, FL (Hybrid) Job Summary STAFFXPERT LLC is seeking a Site Reliability Engineer (SRE) on behalf of our client in Orlando, FL . This role is responsible for ensuring the reliability, availability, scalability, and performance of critical applications and platforms. The ideal candidate will have strong experience in production operations, automation, cloud technologies, observability, and incident management, with a focus on improving system stability and operational excellence in a high-availability environment. Key Responsibilities Lead incident response efforts, including troubleshooting, service restoration, and issue resolution for production environments. Conduct root cause analysis (RCA) and implement preventive measures to improve system reliability. Define, monitor, and enhance service health through observability practices, SLIs, and SLOs. Design and maintain monitoring, alerting, logging, and tracing solutions. Automate operational processes and workflows to improve efficiency and reduce manual effort. Optimize system performance, scalability, and resiliency across cloud and on-premises environments. Support capacity planning, disaster recovery, failover strategies, and resilience testing. Collaborate with engineering, infrastructure, and operations teams to implement reliability best practices. Participate in on-call rotations and major incident management activities. Required Qualifications Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent work experience. Proven experience in Site Reliability Engineering, DevOps, Platform Engineering, or Production Support. Strong knowledge of cloud platforms such as AWS and/or Azure. Hands-on experience with containerization and orchestration technologies, including Kubernetes and Docker. Experience with monitoring and observability tools such as Prometheus, Grafana, Splunk, Datadog, ELK, or similar platforms. Proficiency in scripting and automation using Python, Bash, PowerShell, or comparable languages. Experience with Infrastructure as Code (IaC) tools such as Terraform or Ansible. Strong understanding of CI/CD pipelines and deployment automation. Knowledge of distributed systems, system performance tuning, and reliability engineering concepts. Excellent analytical, troubleshooting, and problem-solving skills. Preferred Qualifications Experience supporting large-scale, mission-critical production environments. Familiarity with SLI, SLO, and error budget frameworks. Experience with security, compliance, and vulnerability remediation practices. Industry certifications related to AWS, Azure, Kubernetes, DevOps, or Site Reliability Engineering. Show more Show less