MLOps Engineer
Description
About The Role The role bridges the gap between machine learning research and production systems. This engineer will design, build, and maintain the automated infrastructure that allows data scientists to train, deploy, and monitor machine learning models at scale. This position is critical for transforming experimental models into robust, resilient production services. The engineer will collaborate closely with data science and DevOps teams to establish automated CI/CD pipelines, robust monitoring, and scalable orchestration frameworks. Key Responsibilities Build and maintain automated CI/CD pipelines for ML model training, testing, and deployment using tools like GitLab CI, GitHub Actions, or Jenkins Design and scale ML orchestration workflows using platforms such as Kubeflow, Airflow, or Prefect to manage complex data and training pipelines Implement model serving infrastructure using Triton Inference Server, TorchServe, or FastAPI to support high-throughput, low-latency API endpoints Deploy automated monitoring, logging, and alerting systems to track model performance, data drift, and system health in production Manage and optimize infrastructure on cloud platforms (AWS, GCP, or Azure) using Infrastructure as Code (IaC) principles with Terraform or CloudFormation Collaborate with security and data compliance teams to enforce data governance, access controls, and secure model execution practices What We Are Looking For 3–6 years of software engineering or DevOps experience, with at least 2 years focused specifically on production MLOps or ML infrastructure Strong proficiency in Python and solid experience with containerization technologies, specifically Docker and Kubernetes Hands-on experience with ML flow-tracking and registry tools such as MLflow, Weights & Biases, or cloud-native model registries Familiarity with data pipeline tools and distributed processing frameworks like Spark, Snowflake, or BigQuery Bachelor's or Master's degree in Computer Science, Software Engineering, or a related quantitative field Bonus: Experience with feature stores (e.g., Feast, Tecton), Triton optimization, or managing GPU acceleration pools in Kubernetes cluster environments Show more Show less