MLOps Engineer

Scale.jobsAustin, United States
Full TimeOn-siteMidLimited info disclosed
52 views0 applications

Description

About The Role The role bridges the gap between machine learning and system engineering, focusing on the deployment, scaling, and monitoring of production machine learning models. The team builds robust, automated infrastructure that enables data scientists and ML engineers to deploy models rapidly and reliably, transforming raw code into resilient, self-healing services. This position is critical for scaling machine learning capabilities across the organization. The selected candidate will own the CI/CD pipelines for ML, optimize high-throughput model inference, and ensure comprehensive monitoring and lineage tracing for all models running in production. Key Responsibilities Design, build, and maintain scalable ML platform infrastructure using Kubernetes, Docker, and Kubeflow to orchestrate end-to-end model workflows Develop and optimize CI/CD pipelines tailored for machine learning (GitOps), ensuring automated testing, packaging, and deployment of models Implement automated monitoring, logging, and alerting systems to detect model drift, concept drift, and performance degradation in real-time Build and maintain centralized feature stores and model registries (such as Feast or MLflow) to ensure reproducibility and lineage tracking Collaborate with backend engineers to integrate model endpoints into high-throughput, low-latency microservices Optimize model inference performance through quantization, pruning, and GPU acceleration (TensorRT, ONNX Runtime) What We Are Looking For 3–6 years of experience in DevOps, MLOps, or Software Engineering, with at least 2 years dedicated to MLOps and production ML systems Strong proficiency in Python and hands-on experience with containerization and orchestration using Docker and Kubernetes Proven experience with ML platform tools such as MLflow, Kubeflow, Seldon Core, or BentoML Solid understanding of cloud infrastructure (AWS or GCP) and Infrastructure as Code (IaC) tools like Terraform Familiarity with data pipeline tools (Airflow, Prefect) and database technologies (SQL, NoSQL, and Vector databases) Bonus: Experience with Triton Inference Server, ML security best practices, or large-scale distributed training frameworks (Ray, PyTorch Elastic) Show more Show less