Staff ML Engineer – AI Evaluation & Synthetic Data
Description
Staff ML Engineer – AI Evaluation & Synthetic Data Job Type: Full-time only ( No C2C, No C2H ) Location: Remote (United States) Work Authorization: U.S. Citizen, Green Card, or EAD valid for at least the next 1 year. No H1B. No sponsorship available. Please apply only if you have direct hands-on experience with at least 3 of the following: LLM evaluation pipelines or LLM-as-a-judge systems Synthetic data generation/augmentation/paraphrasing pipelines Multilingual NLP or cross-lingual evaluation Model calibration, agreement metrics, or statistical validation AI safety/trust & safety/responsible AI/content moderation PyTorch/TensorFlow/Hugging Face model training or fine-tuning This is a specialized ML evaluation role, not a general GenAI/chatbot, RAG app, data analyst/BI, or backend integration position. We are hiring an ML Engineer to build and scale automated evaluation systems and synthetic data generation pipelines for AI safety assessments across languages and markets. This role is focused on training automated judges , designing validation frameworks , building automated performance checks , and scaling analysis/reporting workflows used to evaluate AI systems in diverse linguistic contexts. You will work closely with language experts, multilingual annotators, and cross-functional teams to validate automated approaches for safety and policy evaluation. What You’ll Work On Automated Judge Development Train, fine-tune, and validate automated judge models that score AI outputs for safety and policy compliance Develop calibration and agreement metrics to measure alignment with human evaluation Validation Frameworks Design validation techniques to assess accuracy, reliability, and cross-linguistic consistency Detect drift, bias, and failure modes across markets and language groups Synthetic Data Generation Build and maintain synthetic data generation pipelines to expand evaluation coverage Use augmentation, paraphrasing, and controlled generation methods to support low-resource language evaluation Validate synthetic data quality against human-generated benchmarks Scalable Analysis & Reporting Build automated pipelines for reporting and analysis Improve reproducibility and reduce manual effort in cross-market safety assessments Integrate outputs into existing dashboards and reporting workflows Required Qualifications 3+ years of experience in ML engineering or applied ML research Strong Python skills Strong hands-on experience with PyTorch, TensorFlow, or Hugging Face Transformers Experience training, fine-tuning, or evaluating language models and/or classifiers Experience building automated data processing, evaluation, or monitoring pipelines Experience with experiment design and statistical validation across segmented samples Ability to work independently in a fast-moving environment Strong attention to detail and execution MS or PhD in Computer Science, Machine Learning, NLP, or a related field preferred Preferred Qualifications Experience with synthetic data generation, augmentation, paraphrasing, or controlled generation Experience with multilingual NLP, cross-lingual transfer, or low-resource language modeling Experience with evaluation-as-a-service or automated red teaming frameworks Experience with Spark, Ray, or cloud ML platforms Experience in AI safety, responsible AI, content moderation, or trust & safety Experience with CI/CD for ML validation and deployment What We Offer Opportunity to work on high-impact AI evaluation systems Collaborative and highly motivated team Competitive compensation Flexible schedule Medical, dental, and vision benefits Professional development opportunities Show more Show less