AI Test Lead | GenAI, LLM, RAG & Test Automation
AI Quality Leadership

AI Test Lead

Lead end-to-end testing for AI/ML, Generative AI, LLM and RAG-based applications while driving automation, AI evaluation and quality strategy.

Role Overview

We are looking for an experienced AI Test Lead with 7+ years of experience in software testing and test automation, including hands-on exposure to AI/ML, Generative AI, LLM and RAG-based applications.

The role combines traditional QA leadership with modern AI quality engineering, requiring the ability to validate both application behavior and AI-generated outcomes.

This role is ideal for a QA leader who can bridge conventional automation testing with AI-specific evaluation, covering functional quality, API reliability, AI response quality, security, performance and continuous delivery.

Key Responsibilities
  • Lead end-to-end testing activities for AI/ML and Generative AI applications.
  • Define test strategies, test plans, scenarios and quality standards for AI solutions.
  • Design and implement automated testing frameworks for web, API and AI-based applications.
  • Test LLM, RAG, prompt-based applications, chatbots and AI agents.
  • Validate AI responses for accuracy, relevance, consistency, toxicity, bias, hallucination and reliability.
  • Develop test datasets and evaluation methodologies for AI/GenAI applications.
  • Perform API, functional, regression, integration and end-to-end testing.
  • Collaborate with developers, data scientists, product owners and business stakeholders.
  • Identify and troubleshoot defects across application and AI components.
  • Integrate automated tests into CI/CD pipelines.
  • Mentor QA engineers and provide technical leadership to the testing team.
  • Establish quality metrics, reporting and continuous improvement practices.
  • Ensure testing standards and best practices are consistently followed.
AI Testing & Evaluation
LLM Testing Validate language-model applications and generated responses.
RAG Testing Validate retrieval quality, grounding and response reliability.
Prompt Testing Test prompt behavior, consistency and regression scenarios.
AI Agents Validate AI-agent workflows and application behavior.
Hallucination Detect unsupported, fabricated or unreliable AI responses.
Responsible AI Evaluate toxicity, bias, safety and responsible AI considerations.
Required Technical Skills

Test Automation

Selenium, Playwright, Cypress or similar automation frameworks.

Programming

Strong programming skills in Python or Java.

API Testing

REST APIs, Postman, Rest Assured or equivalent tools.

AI Evaluation

DeepEval, Promptfoo, OpenAI Evals or equivalent platforms.

Database Testing

SQL, database validation and data-quality testing.

CI/CD

Jenkins, GitHub Actions, Azure DevOps and automated pipelines.

Technology Stack
Python Java Selenium Playwright Cypress Postman Rest Assured SQL DeepEval Promptfoo OpenAI Evals LangChain LangGraph Git Jenkins GitHub Actions Azure DevOps
Good to Have
Experience with LangChain and LangGraph applications
Knowledge of vector databases and RAG architectures
Exposure to AWS, Azure or GCP
AI security testing experience
Knowledge of responsible AI principles
Performance and security testing experience
Previous QA team leadership or mentoring experience
Experience with Agile/Scrum environments
Ideal Candidate

Who Fits This Role?

A senior QA and test automation professional with 7+ years of experience who has moved beyond traditional software testing into practical AI/GenAI quality engineering. The strongest candidate combines automation expertise with hands-on LLM, RAG, AI-agent and AI-evaluation experience, while being capable of leading teams and translating AI quality requirements into effective testing strategies.

Candidate Requirements

Experience

7+ years in Software Testing, QA and Test Automation with hands-on AI/GenAI testing.

Education

Bachelor's degree in Computer Science, IT, Engineering or a related field.

Work Location

Remote

Leadership

Technical leadership, QA mentoring, quality metrics and continuous improvement.

Apply for AI Test Lead →
AI Test Lead • GenAI • LLM • RAG • Test Automation • AI Quality