DevOps / MLOps / AIOps Engineer

Elios AIUnited States
ContractOn-siteMidLimited info disclosed
33 views0 applications

Description

DevOps / MLOps / AIOps Engineer Location: Remote (US) | Type: Contract to Hire | Experience: 5+ years About the Role We're hiring a DevOps/MLOps/AIOps Engineer to design and maintain the cloud infrastructure, CI/CD pipelines, and observability that keep the application and its AI systems reliable in production. You'll own the full deployment lifecycle, from agent serving and scaling to monitoring, cost management, and incident response. The client is a fast-growing financial advisory and accounting firm building an in-house AI platform team. Because they handle sensitive data in a regulated domain, you'll operate agent serving across two models: self-hosted workers that keep sensitive data on the firm's own infrastructure, and managed cloud sandboxes for everything else. The services are polyglot, with Elixir/OTP releases running alongside Python agent workers in containers. This is a good fit for someone with a bias toward automation, a cost-optimization mindset, and the on-call maturity to catch regressions before users do. What You'll Do Own the cloud infrastructure: Azure Container Apps, managed-identity secrets via Key Vault references, and the release pipeline for Elixir/OTP services and agent workers Enable AI capabilities across Azure, AWS, and Google Cloud Build and maintain CI/CD and its quality gates: formatting, Credo strict, compile-with-warnings-as-errors, the full test suite, and the staging auto-migration job Operate agent serving and scaling across deployment models, with connection pooling and per-session supervision Own observability across the product core and the agent runtime: telemetry over LLM requests and cron jobs, dashboards, alerting, and the signals that catch regressions early Manage LLM and infrastructure cost, including token spend, model selection, and capacity Lead incident response and on-call What You'll Work On Azure Container Apps and Container Apps Jobs, plus managed identity, Key Vault, and runtime configuration Oban background-job orchestration, Finch and HTTP pooling, and SSE stream infrastructure CI/CD pipelines and the gate set that protects main, including branch protection and required reviews Telemetry and observability for both the product core and the agent runtime, including token, cost, and latency capture AI services running across Azure, AWS, and Google Cloud Qualifications 5+ years in DevOps, MLOps, or platform engineering, with strong cloud infrastructure and CI/CD experience and a bias toward infrastructure-as-code Production observability and monitoring skills across metrics, logs, traces, and alerting On-call and incident-response maturity, plus a genuine cost-optimization mindset Comfort operating polyglot services (BEAM releases and agentic workers) in containers Nice to Have Azure, AWS, and/or Google Cloud experience, and Elixir release operations LLM cost and observability tooling specifically, like token accounting and eval-in-prod signals Depth in container orchestration and secrets management Why Join Us You'll build the operational backbone for a new AI platform inside a firm where uptime and auditability directly affect client trust, so the reliability work you do is treated as core, not overhead. You'll have real ownership over the infrastructure, the pipelines, and the cost story from the start. This is a contract-to-hire role, so both sides get to confirm the fit before going long-term. To apply, click Apply on LinkedIn and visit eliosai.com/browse-jobs