MLOps Engineer

Scale.jobsAustin, United States
Full TimeOn-siteMidLimited info disclosed
59 views0 applications

Description

About The Role The role drives the design, implementation, and maintenance of the core machine learning infrastructure and deployment pipelines. The team focuses on bridge building between data science exploration and highly available, production-grade serving systems that scale to millions of monthly active users. The engineer will collaborate closely with machine learning researchers and backend platform teams to establish standardized MLOps practices, automate training and inference pipelines, and implement comprehensive monitoring frameworks for model health. Key Responsibilities Design and implement automated CI/CD pipelines for packaging, testing, and deploying machine learning models to production clusters using Docker and Kubernetes Build and maintain orchestration pipelines using Apache Airflow or Prefect to manage complex data prep, training, and evaluation workflows Implement real-time and batch model monitoring solutions to detect data drift, concept drift, and performance anomalies using Prometheus and Grafana Deploy and optimize centralized feature stores and model registries to ensure consistent, low-latency access to feature vectors at training and inference time Collaborate with infrastructure engineers to optimize cloud resource utilization, container orchestration, and GPU scheduling on AWS or GCP Establish robust fallback mechanisms, shadow deployments, and canary release strategies to guarantee high-availability model serving What We Are Looking For 3-6 years of experience in software engineering, DevOps, or infrastructure roles, with at least 2 years dedicated specifically to MLOps in production environments Expert-level Python programming skills and extensive experience containerizing applications using Docker and Kubernetes Hands-on experience deploying and scaling model serving frameworks such as Triton Inference Server, TF Serving, TorchServe, or vLLM Strong proficiency with infrastructure as code (IaC) tools like Terraform and workflow orchestrators like Airflow, Kubeflow, or Argo Workflows Solid understanding of software engineering best practices, including git workflows, unit testing, and design patterns Bonus: Experience with MLflow, Weights & Biases, feature stores like Feast or Tecton, or building pipelines for large language models (LLMs)