MLOps Engineer

Scale.jobsAtlanta, United States
Full TimeOn-siteMidLimited info disclosed
56 views0 applications

Description

About The Role The role is responsible for bridging the gap between machine learning development and production operations. This position designs, builds, and maintains the automated infrastructure and pipelines required to deploy, monitor, and scale machine learning models reliably at enterprise scale. Working closely with data scientists, backend developers, and data engineers, the role ensures that model training is reproducible, deployments are seamless, and running models are continuously monitored for latency, drift, and performance regressions. Key Responsibilities Design and implement automated CI/CD pipelines for machine learning models, enabling seamless transition from experimental code to production deployments. Develop and maintain robust feature stores and training data pipelines using tools like Feast, Spark, or dbt. Build and manage model serving infrastructure using Kubernetes (EKS/GKE), KServe, Triton Inference Server, or BentoML. Implement comprehensive monitoring, logging, and alerting systems for deployed models to detect data drift, concept drift, and performance degradation using Prometheus and Grafana. Optimize model inference performance, utilizing techniques such as quantization, pruning, or ONNX runtime integration. Establish model governance frameworks, including robust model registry management, versioning, and automated audit trails. What We Are Looking For 3-6 years of experience in DevOps, SRE, or Software Engineering, with at least 2 years dedicated to specialized MLOps engineering. Strong proficiency in Python and solid experience with containerization technologies, specifically Docker and Kubernetes. Hands-on experience with orchestration and ML workflow tools such as Kubeflow, Airflow, MLflow, or Prefect. Deep understanding of cloud infrastructure (AWS or GCP) and Infrastructure as Code (IaC) principles using Terraform. Familiarity with standard machine learning frameworks (PyTorch, TensorFlow, scikit-learn) and data manipulation libraries. Bonus: Experience managing LLM operations (LLMOps), implementing vector databases, or working with Triton Inference Server. Show more Show less