About this role
Elastic, the Search AI Company, enables everyone to find the answers they need in real time, using all their data, at scale — unleashing the potential of businesses and people. The Elastic Search AI Platform, used by more than 50% of the Fortune 500, brings together the precision of search and the intelligence of AI to enable everyone to accelerate the results that matter. By taking advantage of all structured and unstructured data — securing and protecting private information more effectively — Elastic’s complete, cloud-based solutions for search, security, and observability help organizations deliver on the promise of AI.
What is The Role
We are looking for a skilled QA & Evaluation Engineer to join our team. The role blends strategic QA leadership with hands-on technical validation and structured evaluation to safeguard the accuracy, reliability, compliance, and ethical use of AI models. You will partner across IT and Engineering teams to identify, design, implement and run robust testing frameworks and evaluation rubrics for a portfolio of GenAI solutions that will be used across our organization.
What You Will Be Doing
• Be a primary contributor to our AI strategy, helping validate and test AI infrastructure, custom solutions and third-party SaaS offerings.
• Test Strategy & Execution: Design and implement comprehensive test strategies for AI/ML systems, including accuracy, bias, robustness, and regression testing.
• Rubric-Based Evaluation: Design and implement self-contained evaluation tasks, including prompts, supporting files, and detailed grading rubrics to assess AI performance on functional workflows.
• Automation & CI/CD: Automate validation suites for agentic/multi-agent systems, integration testing, and CI/CD pipelines for ML models.
• Data Validation: Validate that AI/ML models are consuming accurate , authorized, and properly structured data sources; ensuring data quality across training and inference.
• Observation & Reporting: Meticulously observe and document AI agent behaviors, producing crisp, precise summaries and reports on model performance and hallucinations.
• Output Grounding: Validate prompt engineering outputs from a data accuracy standpoint, ensuring responses are grounded in verified data sources.
• Refinement & Iteration: Iterate and refine evaluation tasks and rubrics based on feedback and team collaboration to ensure robust benchmarking methodologies.
• Security & Governance: Ensure all AI data sources and structures meet governance, regulatory, and compliance standards, while implementing best practices for security and data privacy.
• Collaborate with teams from different areas. These areas include IT Engineering, IT Operations, Data & Integrations, PMO, CRM, Risk & Compliance, and business technology.
• Stay current on the latest work in AI and make technical recommendations to the organization.
What You Bring
• Proficiency in Python, T