LLM / GenAI Engineer

Evlo AIMiami, United States
Full TimeOn-siteMidLimited info disclosed
2 views0 applications

Description

About The Role The role focuses on building production-grade generative AI systems, moving beyond basic prompting to architect robust RAG pipelines, agentic workflows, and fine-tuning frameworks. The engineering team collaborates closely with applied researchers and backend developers to deliver scalable AI capabilities that directly impact enterprise applications. Key Responsibilities Design and deploy advanced RAG pipelines utilizing frameworks such as LangChain and LlamaIndex for large-scale information retrieval Optimize vector database integrations including Pinecone, Weaviate, and pgvector for low-latency semantic search Implement systematic LLM evaluation frameworks incorporating LLM-as-judge pipelines and automated regression testing Execute parameter-efficient fine-tuning workflows like LoRA and QLoRA on domain-specific datasets using PyTorch and Hugging Face Build robust backend APIs in Python to serve model inferences reliably in distributed cloud environments Monitor production LLM systems for latency, cost, hallucination rates, and performance drift What We Are Looking For 3 to 6 years of software engineering experience, with a minimum of 2 years dedicated to building and deploying LLM-based applications in production Proficiency in Python and deep familiarity with ML frameworks such as PyTorch, Transformers, and major orchestration libraries Demonstrated experience with vector databases, embedding models, and semantic similarity search at scale Solid grasp of prompt engineering strategies, context window management, and model quantization techniques Bachelor's or Master's degree in Computer Science, Artificial Intelligence, or a related technical field Bonus: Experience with distributed training, custom model fine-tuning, or contributions to open-source AI projects