Principal MLOps Engineer

ContractOn-siteMidLimited info disclosed
28 views0 applications

Description

Hello Everyone, Title: Principal MLOps Engineer Location: Remote Duration: Long-term Role Overview We are seeking a Principal MLOps Engin eer to architect and build our next-generation Enterprise Machine Learning Operations (MLOps) platform. In this role, you will build the core infrastructure on AWS and Amazon SageMaker to streamline the entire ML lifecycle. Your work will directly enable data science teams to deploy, monitor, and scale both high-throughput batch processing and ultra-low-latency real-time models safely and efficiently. Core Responsibilities Platform Architecture & Deployment Build scalable infrastructure on Amazon SageMaker supporting both real-time hosting and batch transform workloads. Implement multi-model endpoints and auto-scaling policies to optimize resource utilization and compute costs. Configure secure networking environments utilizing AWS VPC, PrivateLink, IAM roles, and KMS encryption. CI/CD & Pipeline Automation Build automated pipelines (CI/CD) for model training, testing, and deployment using Infrastructure as Code (IaC). Design and maintain ADO (Azure DevOps) pipelines, including multi-stage approval gate strategies for controlled model promotion. Integrate model registries to enforce version control, reproducibility, and rigorous approval workflows. Develop reusable blueprints and containers to standardize the model promotion process across environments. Diagnose and resolve pipeline failures, including training job failures, failed deployments, and registry update errors. Monitoring & Governance Establish drift detection and data quality monitoring using SageMaker Model Monitor and Amazon CloudWatch, including configuration of custom data drift detection rules. Build unified monitoring and observability dashboards using Grafana Cloud alongside CloudWatch for full-stack visibility. Implement proactive alerting for endpoint downtime, high latency, error spikes, timeouts, and general performance degradation. Configure alerts covering deployment failures, pipeline failures, increased latency, and drift-detected events to ensure rapid incident response. Track data lineage and maintain audit trails to ensure compliance with enterprise security standards. Required Qualifications Technical Experience AWS Expertise : Minimum 3 years of hands-on experience managing production workloa ds using Amazon SageMaker, S3, Lambda , and IAM. Infrastructure as Code: Proven proficie ncy with Terraform for provisioning cloud r esources. Software Engineering: Strong programming s kills in Python and deep expertise with containerization. Data Pipelines : Experience integrating orchestration tools such as AWS Glue, Apache Spark, o r Airflow with ML workflows. CI/CD Tooling: Hands-on experience building and mai ntaining ADO (Azure DevOps) pipelines, including approval gate design. Observability : Practical experie nce wit h Grafana C loud and C loudWatch for monitoring model and infrastructure health, including configuring alerts for endpoint unavailability, high latency, error spikes, and timeouts. Education & Soft Skills Education: Bachelor's or Master's degree in Computer Science, Data Science, Engineering, or a related technical field. Coll aboration: Experience partnering closely with Data Science, DevOps, and Cyber Security teams. Com munication: Ability to explain complex infrastructure choices to both technical and non-technical stakeholders.