NLP Engineer

HiredBuddyNew York, United States
Full TimeOn-siteJuniorLimited info disclosed
21 views0 applications

Description

About the Role We are an AI data annotation and labeling contracting company delivering high-quality text, conversation, and multilingual datasets for Large Language Models (LLMs) and NLP applications. We are seeking a detail-oriented Junior to Intermediate NLP Engineer to build text pre-processing pipelines, design tokenization and parsing workflows, automate quality evaluation for text annotations, and optimize RLHF datasets for client ML teams. Key Responsibilities Text Data Pipelines: Build and maintain automated pipelines to clean, tokenize, format, and validate large text datasets (JSON, JSONL, Parquet) for LLM fine-tuning and evaluation. Annotation Guidelines & Automation: Develop programmatic validation scripts for named entity recognition (NER), intent classification, sentiment analysis, and multi-turn prompt-response labeling. Quality Assurance & IAA: Implement statistical metrics to evaluate Inter-Annotator Agreement (IAA) and semantic consistency across complex linguistic datasets. LLM & Prompt Evaluation: Benchmark AI model outputs using evaluation frameworks (e.g., ROUGE, BLEU, RAG metrics) and support Reinforcement Learning from Human Feedback (RLHF) workflows. Client Dataset Delivery: Convert raw annotation outputs into structured dataset formats ready for fine-tuning transformer models. Qualifications & Skills Experience: 1–3 years of hands-on experience in Natural Language Processing, Computational Linguistics, or Machine Learning Data Engineering. Core Programming: Proficiency in Python and essential NLP libraries (Hugging Face Transformers, spaCy, NLTK, LangChain/LlamaIndex). Data Handling: Deep familiarity with handling text data formats (JSONL, CSV, Parquet) and regular expressions (Regex). ML Foundations: Solid understanding of modern NLP architectures (Transformers, Tokenization, Embeddings, LLM fine-tuning concepts). DevOps Basics: Experience with Git, Docker, REST APIs, and SQL. What We Offer Direct hands-on experience building text datasets for industry-leading Large Language Models. Mentorship from senior AI engineers and clear career progression. Competitive salary package.