MLOps Engineer

Scale.jobsChicago, United States
Full TimeOn-siteMidLimited info disclosed
59 views0 applications

Description

About The Role The role bridges the gap between machine learning research and production software engineering, owning the infrastructure, pipelines, and automation required to deploy and monitor models at scale. The engineer will design and maintain robust MLOps systems that ensure model deployments are repeatable, observable, and highly reliable. This position collaborates closely with data scientists, backend developers, and data engineers to establish standardized ML lifecycles. The focus is on automating training and deployment workflows, optimizing serving latency, and safeguarding system stability under high-throughput production loads. Key Responsibilities Build and maintain automated CI/CD pipelines for machine learning models, enabling seamless transition from development to production Design and orchestrate ML workflows using tools like Kubeflow, Airflow, or Prefect to manage complex training and evaluation jobs Develop and manage model serving infrastructure using Triton Inference Server, TorchServe, or FastAPI containerized on Kubernetes Implement comprehensive monitoring and alerting systems to track model drift, concept drift, system latency, and data quality issues Collaborate on the design of a centralized feature store and model registry to ensure reproducibility and consistency across offline and online environments Optimize inference performance through quantization, pruning, and hardware-accelerated execution on GPUs or specialized cloud instances What We Are Looking For 3-6 years of experience in MLOps, DevOps, or Software Engineering with a strong focus on production machine learning environments Proficiency with containerization and orchestration technologies, specifically Docker, Kubernetes, and Helm charts Hands-on experience with cloud infrastructure (AWS or GCP) and Infrastructure as Code using Terraform Strong programming skills in Python and familiarity with ML frameworks such as PyTorch, TensorFlow, or Scikit-Learn Experience implementing data pipelines and integrating model registries like MLflow or Weights & Biases Bonus: Experience with vector databases, deploying Large Language Models (LLMs), or working with Triton Inference Server Show more Show less