AI TEST LEAD • GENAI • LLM • RAG • TEST AUTOMATION

AI Test Lead

AI Quality Engineering & Test Automation Leadership

Lead end-to-end quality strategy for AI/ML and Generative AI applications, combining test automation, LLM evaluation, RAG testing, AI quality assessment and technical QA leadership.

Experience: 7+ Years
Location: Remote
Focus: AI / GenAI Testing
Automation: Selenium / Playwright / Cypress
Languages: Python / Java
Apply Now ↗

Who Is This Role For?

This opportunity is designed for a senior QA professional who can combine strong automation leadership with practical experience testing AI, GenAI, LLM, RAG and agent-based applications.

01

AI Test Leader

An experienced QA professional capable of owning testing strategy, quality standards and delivery across AI-powered applications.

02

GenAI Testing Specialist

Someone comfortable validating LLM responses, RAG workflows, prompt-based applications, chatbots and AI agents.

03

Automation Architect

Strong hands-on automation experience with modern web, API and AI application testing frameworks.

04

Quality Strategy Owner

A leader who can define AI quality metrics, build evaluation approaches and continuously improve testing practices.

End-to-End AI Quality Leadership

The role combines conventional software quality engineering with AI-specific evaluation. The successful candidate will establish reliable testing practices across functional, API, performance, security and AI quality dimensions.

Functional Application behavior
API & Integration Services and workflows
AI Evaluation LLM and RAG quality
Non-Functional Performance and security

What You Will Do

01

Lead AI Testing Strategy

Define end-to-end testing strategies, test plans, scenarios and quality standards for AI and Generative AI solutions.

02

Architect Automation Frameworks

Design and implement reusable automation frameworks for web, API and AI-powered applications.

03

Test LLM & RAG Applications

Validate LLM, RAG, prompt-based applications, chatbots and AI agents across functional and quality dimensions.

04

Evaluate AI Responses

Assess responses for accuracy, relevance, consistency, reliability, toxicity, bias and hallucination.

05

Build AI Test Datasets

Create structured test datasets and evaluation methodologies for repeatable AI and GenAI validation.

06

Integrate Testing Into CI/CD

Embed automated testing and AI evaluation into continuous integration and deployment pipelines.

07

Lead API & System Testing

Perform functional, regression, integration and end-to-end testing across distributed application components.

08

Mentor QA Engineers

Provide technical leadership and help QA engineers adopt modern AI testing and automation practices.

09

Drive Quality Improvement

Establish quality metrics, reporting mechanisms and continuous improvement practices across projects.

What You Will Evaluate

Accuracy

Determine whether AI responses correctly answer the user's request and satisfy the expected outcome.

Relevance

Check whether generated responses remain relevant to the requested context and business intent.

Grounding

Validate whether RAG responses remain supported by the underlying knowledge sources.

Consistency

Compare repeated responses and identify unstable or contradictory model behavior.

Hallucination

Detect fabricated information, unsupported claims and responses that exceed available knowledge.

Safety & Toxicity

Identify unsafe, biased, toxic or inappropriate AI-generated content and behavioral patterns.

Testing Modern AI Applications

The role requires understanding how modern AI systems behave beyond traditional UI testing, including retrieval, prompting, tool usage and autonomous agent behavior.

01

RAG Validation

Validate retrieval quality, grounding, context relevance and response correctness against knowledge sources.

02

Prompt Testing

Test prompts and prompt-based workflows for quality, consistency, edge cases and unexpected model behavior.

03

AI Agent Testing

Validate agent workflows, decisions, tool interactions and reliability across multi-step AI tasks.

04

Response Quality

Establish repeatable evaluation methods for AI-generated outputs and measurable quality thresholds.

AI Safety & Quality Dimensions

Hallucination Detection Bias Testing Toxicity Testing Prompt Evaluation Response Reliability AI Safety Responsible AI Guardrail Validation Adversarial Testing Quality Metrics

Quality Beyond Pass / Fail

AI testing requires evaluating probabilistic outputs across multiple dimensions. The objective is not only to determine whether a test passes, but whether the AI behaves accurately, reliably and safely within its intended context.

Technical Automation Expertise

Web Automation

Build scalable UI automation using Selenium, Playwright, Cypress or comparable modern frameworks.

Python / Java

Use Python or Java for automation, test utilities, AI evaluation workflows and integration testing.

API Testing

Validate REST services using Postman, Rest Assured or equivalent API testing technologies.

Selenium Playwright Cypress Python Java REST API Postman Rest Assured SQL Git

Evaluation & AI Testing Tools

DeepEval

Exposure to automated evaluation of LLM outputs and AI-specific quality metrics.

Promptfoo

Useful for prompt evaluation, regression testing and systematic comparison of model behavior.

OpenAI Evals

Familiarity with evaluation approaches for assessing model behavior against defined criteria.

LangChain / LangGraph

Understanding of AI application and agent workflows built using modern orchestration frameworks.

Vector Databases

Knowledge of vector stores and retrieval architectures used in RAG applications.

Cloud AI Platforms

Exposure to AWS, Azure or GCP environments supporting enterprise AI workloads.

Continuous AI Quality

Automated Pipelines Integrate AI and software testing into continuous integration and deployment workflows.
Quality Gates Establish measurable criteria to determine whether AI features meet required quality thresholds.
Regression Testing Detect changes in AI behavior and application functionality across releases.
Defect Management Identify, troubleshoot and track defects across application and AI components.
Jenkins GitHub Actions Azure DevOps Git CI/CD

Technical Leadership & Collaboration

Beyond technical execution, the role requires the ability to guide teams, establish standards and communicate AI quality risks clearly to engineering, product and business stakeholders.

01

Mentoring

Coach QA engineers and improve team capability in AI testing, automation and modern quality practices.

02

Stakeholder Collaboration

Work with developers, data scientists, product owners and business stakeholders to define AI quality requirements.

03

Quality Reporting

Establish meaningful quality metrics and communicate testing results, risks and improvement opportunities.

04

Continuous Improvement

Identify systemic quality problems and introduce better testing approaches, tools and engineering practices.

What You Should Bring

Professional Experience 7+ years of experience in software testing, QA or test automation.
AI Testing Hands-on experience testing AI, Generative AI, LLM and RAG-based applications.
Automation Strong experience with Selenium, Playwright, Cypress or equivalent automation technologies.
Programming Strong programming skills in Python or Java for automation and testing.
API & Database Experience with REST API testing and good understanding of SQL and database testing.
SDLC & Agile Strong understanding of SDLC, STLC, Agile/Scrum, defect management and test management practices.

Additional Skills That Add Value

LangChain LangGraph Vector Databases RAG Architecture AWS Azure GCP AI Security Testing Responsible AI Performance Testing Security Testing QA Team Leadership

Technology & Testing Expertise

AI Testing GenAI Testing LLM Testing RAG Testing AI Agents Prompt Engineering AI Evaluation Hallucination Testing Bias Testing Toxicity Testing Python Java Selenium Playwright Cypress API Testing SQL DeepEval Promptfoo OpenAI Evals CI/CD AI Quality

Lead the Future of AI Quality

Bring together AI testing, automation engineering and quality leadership to build reliable, scalable and production-ready Generative AI applications.

Apply Now ↗