Principal Consultant - Technology – SAP Security & Controls

<p><strong>Role :: Machine Learning Evaluation Engineer</strong></p><p><strong>Location :: Cupertino, CA</strong></p><p><strong>Longterm_Contract Role</strong></p><p><br></p><p><strong>Job Description</strong></p><p>As a Machine Learning Evaluation Engineer, you will design and build scalable evaluation systems for LLM, Generative AI, Conversational AI, and Agentic AI product. </p><p>You will:</p><p>• Design and develop automated evaluation frameworks and pipelines for AI-powered products.</p><p>• Define evaluation methodologies and quality metrics across dimensions such as accuracy, relevance, groundedness, completeness, consistency, instruction following, and task completion.</p><p>• Build and maintain high-quality evaluation datasets, including golden datasets, benchmark sets, regression suites, adversarial scenarios, and production-derived test sets.</p><p>• Develop Auto Eval capabilities that enable teams to rapidly evaluate models, prompts, retrieval systems, agents, and end-to-end AI experiences.</p><p>• Design and implement model-based evaluation approaches, including LLM-as-a-Judge, while developing appropriate calibration and validation methodologies.</p><p>• Develop Human-in-the-Loop (HITL) evaluation approaches for complex or subjective quality dimensions where automated evaluation alone is insufficient.</p><p>• Define evaluation rubrics, annotation guidelines, grading criteria, and quality standards in partnership with product teams, domain experts, and annotation teams.</p><p>• Build mechanisms to calibrate automated evaluators against human judgment and measure evaluator consistency and reliability.</p><p>• Evaluate end-to-end AI systems, including retrieval, context construction, prompts, model responses, tool use, APIs, and downstream product experiences.</p><p>• Develop evaluation methodologies for multi-turn conversations, personalization, recommendations, tool use, reasoning, and agentic task execution.</p><p>• Perform detailed error analysis and failure-mode investigation to identify opportunities for model, prompt, retrieval, dataset, and product improvements.</p><p>• Build reusable evaluation infrastructure, APIs, dashboards, and developer tooling that can scale across multiple AI products and teams.</p><p>• Integrate evaluation into development and CI/CD workflows, enabling automated regression detection, quality gates, and release-readiness assessments</p><p>• Connect offline evaluation results with production signals to continuously improve evaluation coverage and product quality.</p><p>• Partner closely with Machine Learning, Software Engineering, Product, Quality Engineering, Human Interface, and Data Science teams throughout research, development, evaluation, launch, and continuous improvement.</p><p><strong>Minimum Qualifications: </strong></p><p>• 5+ years of experience in Machine Learning Engineering, ML Evaluation, Software Engineering, Data Science, Quality Engineering, or a related technical field.</p><p>• Strong programming skills in Python and experience developing production-quality software, ML systems, data pipelines, or evaluation infrastructure.</p><p>• Experience developing or evaluating LLMs, Generative AI, Conversational AI, NLP, recommendation systems, or other machine-learning-driven products.</p><p>• Experience designing automated ML evaluation frameworks, metrics, benchmarks, datasets, or experimentation methodologies.</p><p>• Understanding of modern LLM application architectures, including prompting, embeddings, retrieval-augmented generation (RAG), tool use, and agentic workflows.</p><p>• Experience with model-based evaluation techniques and an understanding of the strengths and limitations of approaches such as LLM-as-a-Judge.</p><p>• Experience with Human-in-the-Loop evaluation, annotation, or data-quality workflows.</p><p>• Strong understanding of statistical analysis, experimentation, sampling, and measurement methodologies.</p><p>• Experience performing model error analysis, failure analysis, and root-cause investigation.</p><p>• Ability to work effectively across Machine Learning, Engineering, Product, Quality, and Data teams.</p><p>• Excellent written and verbal communication skills, with the ability to translate complex technical findings into clear, actionable recommendations.</p><p><br></p><p><strong>Preferred Qualifications: </strong></p><p><br></p><p>• Experience building evaluation infrastructure for production-scale LLM or Generative AI applications.</p><p>• Experience evaluating RAG, conversational systems, AI agents, personalization, recommendations, or multimodal AI.</p><p>• Experience building golden datasets, regression suites, automated quality gates, and continuous evaluation pipelines.</p><p>• Experience integrating ML evaluation into CI/CD and production release processes.</p><p>• Experience with prompt evaluation, model comparison, experiment tracking, and AI observability.</p><p>• Experience evaluating multilingual AI experiences across languages, locales, and markets.</p><p>• Familiarity with responsible AI evaluation, including robustness, safety, bias, and adversarial testing.</p><p>• Experience developing internal ML platforms, developer tooling, or self-service evaluation capabilities used across multiple teams.</p><p>• Experience working with large-scale datasets and distributed ML or data-processing infrastructure.</p>

Back to blog

Other Jobs To Apply

No other job posts for this day.

Common Interview Questions And Answers

1. HOW DO YOU PLAN YOUR DAY?

This is what this question poses: When do you focus and start working seriously? What are the hours you work optimally? Are you a night owl? A morning bird? Remote teams can be made up of people working on different shifts and around the world, so you won't necessarily be stuck in the 9-5 schedule if it's not for you...

2. HOW DO YOU USE THE DIFFERENT COMMUNICATION TOOLS IN DIFFERENT SITUATIONS?

When you're working on a remote team, there's no way to chat in the hallway between meetings or catch up on the latest project during an office carpool. Therefore, virtual communication will be absolutely essential to get your work done...

3. WHAT IS "WORKING REMOTE" REALLY FOR YOU?

Many people want to work remotely because of the flexibility it allows. You can work anywhere and at any time of the day...

4. WHAT DO YOU NEED IN YOUR PHYSICAL WORKSPACE TO SUCCEED IN YOUR WORK?

With this question, companies are looking to see what equipment they may need to provide you with and to verify how aware you are of what remote working could mean for you physically and logistically...

5. HOW DO YOU PROCESS INFORMATION?

Several years ago, I was working in a team to plan a big event. My supervisor made us all work as a team before the big day. One of our activities has been to find out how each of us processes information...

6. HOW DO YOU MANAGE THE CALENDAR AND THE PROGRAM? WHICH APPLICATIONS / SYSTEM DO YOU USE?

Or you may receive even more specific questions, such as: What's on your calendar? Do you plan blocks of time to do certain types of work? Do you have an open calendar that everyone can see?...

7. HOW DO YOU ORGANIZE FILES, LINKS, AND TABS ON YOUR COMPUTER?

Just like your schedule, how you track files and other information is very important. After all, everything is digital!...

8. HOW TO PRIORITIZE WORK?

The day I watched Marie Forleo's film separating the important from the urgent, my life changed. Not all remote jobs start fast, but most of them are...

9. HOW DO YOU PREPARE FOR A MEETING AND PREPARE A MEETING? WHAT DO YOU SEE HAPPENING DURING THE MEETING?

Just as communication is essential when working remotely, so is organization. Because you won't have those opportunities in the elevator or a casual conversation in the lunchroom, you should take advantage of the little time you have in a video or phone conference...

10. HOW DO YOU USE TECHNOLOGY ON A DAILY BASIS, IN YOUR WORK AND FOR YOUR PLEASURE?

This is a great question because it shows your comfort level with technology, which is very important for a remote worker because you will be working with technology over time...