MLOps Engineer
Description
About The Role The role bridges the gap between machine learning and infrastructure, owning the automation, deployment, monitoring, and scaling of production machine learning models. The team focuses on creating robust CI/CD pipelines for ML models to ensure that deployment is repeatable, measurable, and highly resilient. This position requires deep expertise in both software engineering and ML systems. The engineer will collaborate closely with data scientists and data engineers to establish robust MLOps practices, reducing the time to market for new models while maintaining strict service-level agreements. Key Responsibilities Design and maintain automated CI/CD pipelines for packaging, testing, and deploying machine learning models to Kubernetes environments Implement model monitoring, logging, and alerting systems using tools like Prometheus, Grafana, and Evidently AI to detect data drift and performance degradation Manage and optimize feature stores and data pipelines using Feast or Hopsworks to ensure low-latency feature serving at scale Configure and scale orchestration platforms such as Kubeflow, Airflow, or Prefect for robust model training and evaluation workflows Collaborate with infrastructure teams to optimize compute resource allocation (CPU/GPU) for deep learning training and real-time inference Enforce version control and reproducibility standards across code, data, and model artifacts using tools like DVC or MLflow What We Are Looking For 3–6 years of experience in software engineering, DevOps, or MLOps, with a proven track record of deploying ML models to production Strong proficiency in Python and shell scripting, alongside experience writing production-grade, containerized applications using Docker and Kubernetes Hands-on experience with cloud infrastructure platforms, preferably AWS or GCP, including Terraform or CloudFormation for infrastructure as code Solid understanding of the machine learning lifecycle, including training, validation, serving, and continuous monitoring Excellent communication and collaboration skills, with the ability to define technical standards and mentor junior engineers Bonus: Experience with Triton Inference Server, Ray, Spark, or deep learning frameworks like PyTorch and TensorFlow Show more Show less