MLOps Engineer

Scale.jobsMiami, United States
Full TimeOn-siteMidLimited info disclosed
11 views0 applications

Description

About The Role The role bridges the gap between data science and platform engineering, owning the infrastructure, pipelines, and deployment frameworks that make machine learning models operational at scale. The engineer will design and maintain highly automated CI/CD and CT (continuous training) pipelines to ensure models are reliably served and monitored in production environments. This position is critical for scaling machine learning initiatives, transforming experimental models into robust, low-latency microservices. Working closely with data scientists and platform architects, the role ensures that the infrastructure remains scalable, secure, and cost-effective. Key Responsibilities Design and implement scalable ML infrastructure and deployment pipelines using Kubernetes, Kubeflow, or MLflow to automate model serving and monitoring. Build and maintain feature stores and automated data pipelines (ETL/ELT) using technologies like dbt, Apache Spark, or Feast to ensure consistent feature engineering across training and inference. Develop real-time and batch model serving systems with Docker, FastAPI, and Triton Inference Server, optimizing for ultra-low latency and high throughput. Implement robust monitoring systems to track model performance, data drift, and system health using Prometheus, Grafana, and specialized ML monitoring tools. Establish automated testing, versioning, and CI/CD pipelines for ML models and infrastructure code to support safe, repeatable, and continuous deployments. Collaborate with security and compliance teams to enforce data governance, access controls, and secure model execution standards across all cloud environments. What We Are Looking For 3 to 6 years of software engineering or data engineering experience, with at least 2 years dedicated to MLOps and production machine learning infrastructure. Advanced proficiency in Python and robust experience with shell scripting, infrastructure as code (Terraform), and containerization (Docker, Kubernetes). Proven experience building and optimizing ML pipelines on major cloud platforms, specifically AWS (SageMaker, EKS, S3) or GCP (Vertex AI, GKE). Hands-on familiarity with CI/CD tools (GitHub Actions, GitLab CI) and MLOps platforms (MLflow, Kubeflow, Weights & Biases, or Prefect). Bachelor's or Master's degree in Computer Science, Software Engineering, Data Science, or a closely related quantitative field. Bonus: Experience with Triton Inference Server, fine-tuning large language models (LLMs) in production, or working with vector databases (Pinecone, Milvus).