MLOps Engineer
Description
About The Role The role owns the infrastructure, orchestration, and deployment pipelines that power machine learning and GenAI models at scale in production environments. The team works closely with applied scientists and backend engineers to ensure reliable, secure, and low-latency model serving across distributed cloud architectures. Key Responsibilities Architect and maintain scalable MLOps pipelines for continuous integration, continuous delivery, and automated testing of machine learning models Build and manage feature stores, model registries, and artifact repositories using tools like MLflow, Feast, and DVC Containerize ML workflows using Docker and orchestrate distributed training and inference jobs with Kubernetes and Kubeflow Monitor production models for data drift, concept drift, and system performance anomalies using Prometheus, Grafana, and automated alerting Optimize model serving infrastructure for high throughput and low latency using Triton Inference Server, ONNX Runtime, or vLLM Establish security, compliance, and governance frameworks for sensitive data handling and model access control What We Are Looking For 3–6 years of experience in MLOps, DevOps, or platform engineering, with a focus on machine learning infrastructure Deep expertise in Kubernetes, Docker, and infrastructure-as-code tools such as Terraform or Ansible Hands-on experience with cloud platforms (AWS, GCP, or Azure) and managed ML services like SageMaker or Vertex AI Strong software engineering background in Python and Bash, with proficiency in CI/CD pipeline configuration using GitHub Actions or GitLab CI Solid understanding of distributed systems, networking, and observability in microservices architectures Bonus: Experience managing LLM serving infrastructure, vLLM, or vector databases in production