MLOps Engineer
Description
About The Role The role drives the design, implementation, and maintenance of the core MLOps platform, bridging the gap between machine learning research and robust production systems. The team focuses on building automated pipelines that allow data scientists to train, deploy, and monitor complex models at scale with minimal friction. The engineer will collaborate closely with platform infrastructure, data engineering, and data science teams to establish reliable CI/CD pipelines for ML, manage feature stores, and implement high-throughput, low-latency model serving frameworks. Key Responsibilities Design and implement scalable CI/CD pipelines for machine learning models (CT/continuous training) using tools like GitLab CI, GitHub Actions, or Jenkins Deploy and orchestrate machine learning workflows using Kubernetes, Kubeflow, or Apache Airflow to manage compute resources efficiently Build and maintain robust model serving infrastructure supporting both real-time REST/gRPC endpoints (using Triton, TorchServe, or TF Serving) and batch inference pipelines Implement centralized monitoring and observability for deployed models to track performance metrics, data drift, and system latency using Prometheus, Grafana, or Arize Manage and optimize a centralized Feature Store (such as Feast or Tecton) to ensure consistent data definitions across training and serving environments Collaborate with security and compliance teams to enforce data governance, model lineage, and access controls across the entire ML lifecycle What We Are Looking For 3-6 years of experience in DevOps, SRE, or Software Engineering, with at least 2 years dedicated specifically to MLOps and ML infrastructure Strong proficiency in Python and shell scripting, alongside solid experience with containerization using Docker and orchestration with Kubernetes Hands-on experience with cloud infrastructure (AWS, GCP, or Azure) and Infrastructure as Code (IaC) tools like Terraform Familiarity with ML tracking and registry tools such as MLflow, Weights & Biases, or cloud-native model registries Bachelor's degree in Computer Science, Engineering, or a related quantitative field Bonus: Experience managing infrastructure for Large Language Models (LLMs), deep learning optimization (TensorRT), or distributed training frameworks (Ray, Horovod) Show more Show less