Principal Consultant - Technology – SAP Security & Controls
<p><strong>Role :: Machine Learning Evaluation Engineer</strong></p><p><strong>Location :: Cupertino, CA</strong></p><p><strong>Longterm_Contract Role</strong></p><p><br></p><p><strong>Job Description</strong></p><p>As a Machine Learning Evaluation Engineer, you will design and build scalable evaluation systems for LLM, Generative AI, Conversational AI, and Agentic AI product. </p><p>You will:</p><p>• Design and develop automated evaluation frameworks and pipelines for AI-powered products.</p><p>• Define evaluation methodologies and quality metrics across dimensions such as accuracy, relevance, groundedness, completeness, consistency, instruction following, and task completion.</p><p>• Build and maintain high-quality evaluation datasets, including golden datasets, benchmark sets, regression suites, adversarial scenarios, and production-derived test sets.</p><p>• Develop Auto Eval capabilities that enable teams to rapidly evaluate models, prompts, retrieval systems, agents, and end-to-end AI experiences.</p><p>• Design and implement model-based evaluation approaches, including LLM-as-a-Judge, while developing appropriate calibration and validation methodologies.</p><p>• Develop Human-in-the-Loop (HITL) evaluation approaches for complex or subjective quality dimensions where automated evaluation alone is insufficient.</p><p>• Define evaluation rubrics, annotation guidelines, grading criteria, and quality standards in partnership with product teams, domain experts, and annotation teams.</p><p>• Build mechanisms to calibrate automated evaluators against human judgment and measure evaluator consistency and reliability.</p><p>• Evaluate end-to-end AI systems, including retrieval, context construction, prompts, model responses, tool use, APIs, and downstream product experiences.</p><p>• Develop evaluation methodologies for multi-turn conversations, personalization, recommendations, tool use, reasoning, and agentic task execution.</p><p>• Perform detailed error analysis and failure-mode investigation to identify opportunities for model, prompt, retrieval, dataset, and product improvements.</p><p>• Build reusable evaluation infrastructure, APIs, dashboards, and developer tooling that can scale across multiple AI products and teams.</p><p>• Integrate evaluation into development and CI/CD workflows, enabling automated regression detection, quality gates, and release-readiness assessments</p><p>• Connect offline evaluation results with production signals to continuously improve evaluation coverage and product quality.</p><p>• Partner closely with Machine Learning, Software Engineering, Product, Quality Engineering, Human Interface, and Data Science teams throughout research, development, evaluation, launch, and continuous improvement.</p><p><strong>Minimum Qualifications: </strong></p><p>• 5+ years of experience in Machine Learning Engineering, ML Evaluation, Software Engineering, Data Science, Quality Engineering, or a related technical field.</p><p>• Strong programming skills in Python and experience developing production-quality software, ML systems, data pipelines, or evaluation infrastructure.</p><p>• Experience developing or evaluating LLMs, Generative AI, Conversational AI, NLP, recommendation systems, or other machine-learning-driven products.</p><p>• Experience designing automated ML evaluation frameworks, metrics, benchmarks, datasets, or experimentation methodologies.</p><p>• Understanding of modern LLM application architectures, including prompting, embeddings, retrieval-augmented generation (RAG), tool use, and agentic workflows.</p><p>• Experience with model-based evaluation techniques and an understanding of the strengths and limitations of approaches such as LLM-as-a-Judge.</p><p>• Experience with Human-in-the-Loop evaluation, annotation, or data-quality workflows.</p><p>• Strong understanding of statistical analysis, experimentation, sampling, and measurement methodologies.</p><p>• Experience performing model error analysis, failure analysis, and root-cause investigation.</p><p>• Ability to work effectively across Machine Learning, Engineering, Product, Quality, and Data teams.</p><p>• Excellent written and verbal communication skills, with the ability to translate complex technical findings into clear, actionable recommendations.</p><p><br></p><p><strong>Preferred Qualifications: </strong></p><p><br></p><p>• Experience building evaluation infrastructure for production-scale LLM or Generative AI applications.</p><p>• Experience evaluating RAG, conversational systems, AI agents, personalization, recommendations, or multimodal AI.</p><p>• Experience building golden datasets, regression suites, automated quality gates, and continuous evaluation pipelines.</p><p>• Experience integrating ML evaluation into CI/CD and production release processes.</p><p>• Experience with prompt evaluation, model comparison, experiment tracking, and AI observability.</p><p>• Experience evaluating multilingual AI experiences across languages, locales, and markets.</p><p>• Familiarity with responsible AI evaluation, including robustness, safety, bias, and adversarial testing.</p><p>• Experience developing internal ML platforms, developer tooling, or self-service evaluation capabilities used across multiple teams.</p><p>• Experience working with large-scale datasets and distributed ML or data-processing infrastructure.</p>