MLOps Engineer
Description
About The Role The role is responsible for bridging the gap between machine learning model development and production-grade software engineering. This position builds, scales, and maintains the automated infrastructure that allows data science teams to deploy models safely, monitor their performance in real-time, and retrain them continuously. Operating at the intersection of software engineering, DevOps, and data science, this role ensures that our ML platforms are highly available, secure, and capable of processing millions of inferences per day with low latency. Key Responsibilities Design and implement scalable CI/CD pipelines for machine learning models using GitOps principles and tools like GitLab CI, GitHub Actions, or Argo Workflows Deploy and orchestrate machine learning workloads on Kubernetes clusters using Kubeflow, MLflow, or Seldon Core Build and maintain a centralized feature store (such as Feast or Tecton) to ensure consistent data definitions across training and serving environments Establish automated monitoring, logging, and alerting systems for deployed models using Prometheus, Grafana, and specialized ML observability platforms to detect data drift and concept drift Optimize model serving infrastructure for low-latency and high-throughput requirements using frameworks like Triton Inference Server, TorchServe, or TF Serving Collaborate with security and compliance teams to implement access controls, model lineage tracking, and audit trails for all deployed artifacts What We Are Looking For 3–6 years of experience in an MLOps, DevOps, or Software Engineering role with a strong focus on machine learning infrastructure Expert-level Python programming skills and deep familiarity with containerization technologies including Docker and Kubernetes Hands-on experience with cloud infrastructure (AWS, GCP, or Azure) and Infrastructure as Code tools like Terraform Proven track record of building and managing ML pipelines with orchestration tools like Apache Airflow, Prefect, or Argo Solid understanding of software engineering best practices, including unit testing, integration testing, and version control for both code and data (DVC) Bonus: Experience with large language model deployment (LLMOps), Triton inference server optimizations, or vector database administration (e.g., Pinecone, Milvus, Qdrant) Show more Show less