AI Test Lead
Lead end-to-end quality strategy for AI/ML and Generative AI applications, combining test automation, LLM evaluation, RAG testing, AI quality assessment and technical QA leadership.
Who Is This Role For?
This opportunity is designed for a senior QA professional who can combine strong automation leadership with practical experience testing AI, GenAI, LLM, RAG and agent-based applications.
AI Test Leader
An experienced QA professional capable of owning testing strategy, quality standards and delivery across AI-powered applications.
GenAI Testing Specialist
Someone comfortable validating LLM responses, RAG workflows, prompt-based applications, chatbots and AI agents.
Automation Architect
Strong hands-on automation experience with modern web, API and AI application testing frameworks.
Quality Strategy Owner
A leader who can define AI quality metrics, build evaluation approaches and continuously improve testing practices.
End-to-End AI Quality Leadership
The role combines conventional software quality engineering with AI-specific evaluation. The successful candidate will establish reliable testing practices across functional, API, performance, security and AI quality dimensions.
What You Will Do
Lead AI Testing Strategy
Define end-to-end testing strategies, test plans, scenarios and quality standards for AI and Generative AI solutions.
Architect Automation Frameworks
Design and implement reusable automation frameworks for web, API and AI-powered applications.
Test LLM & RAG Applications
Validate LLM, RAG, prompt-based applications, chatbots and AI agents across functional and quality dimensions.
Evaluate AI Responses
Assess responses for accuracy, relevance, consistency, reliability, toxicity, bias and hallucination.
Build AI Test Datasets
Create structured test datasets and evaluation methodologies for repeatable AI and GenAI validation.
Integrate Testing Into CI/CD
Embed automated testing and AI evaluation into continuous integration and deployment pipelines.
Lead API & System Testing
Perform functional, regression, integration and end-to-end testing across distributed application components.
Mentor QA Engineers
Provide technical leadership and help QA engineers adopt modern AI testing and automation practices.
Drive Quality Improvement
Establish quality metrics, reporting mechanisms and continuous improvement practices across projects.
What You Will Evaluate
Accuracy
Determine whether AI responses correctly answer the user's request and satisfy the expected outcome.
Relevance
Check whether generated responses remain relevant to the requested context and business intent.
Grounding
Validate whether RAG responses remain supported by the underlying knowledge sources.
Consistency
Compare repeated responses and identify unstable or contradictory model behavior.
Hallucination
Detect fabricated information, unsupported claims and responses that exceed available knowledge.
Safety & Toxicity
Identify unsafe, biased, toxic or inappropriate AI-generated content and behavioral patterns.
Testing Modern AI Applications
The role requires understanding how modern AI systems behave beyond traditional UI testing, including retrieval, prompting, tool usage and autonomous agent behavior.
RAG Validation
Validate retrieval quality, grounding, context relevance and response correctness against knowledge sources.
Prompt Testing
Test prompts and prompt-based workflows for quality, consistency, edge cases and unexpected model behavior.
AI Agent Testing
Validate agent workflows, decisions, tool interactions and reliability across multi-step AI tasks.
Response Quality
Establish repeatable evaluation methods for AI-generated outputs and measurable quality thresholds.
AI Safety & Quality Dimensions
Quality Beyond Pass / Fail
AI testing requires evaluating probabilistic outputs across multiple dimensions. The objective is not only to determine whether a test passes, but whether the AI behaves accurately, reliably and safely within its intended context.
Technical Automation Expertise
Web Automation
Build scalable UI automation using Selenium, Playwright, Cypress or comparable modern frameworks.
Python / Java
Use Python or Java for automation, test utilities, AI evaluation workflows and integration testing.
API Testing
Validate REST services using Postman, Rest Assured or equivalent API testing technologies.
Evaluation & AI Testing Tools
DeepEval
Exposure to automated evaluation of LLM outputs and AI-specific quality metrics.
Promptfoo
Useful for prompt evaluation, regression testing and systematic comparison of model behavior.
OpenAI Evals
Familiarity with evaluation approaches for assessing model behavior against defined criteria.
LangChain / LangGraph
Understanding of AI application and agent workflows built using modern orchestration frameworks.
Vector Databases
Knowledge of vector stores and retrieval architectures used in RAG applications.
Cloud AI Platforms
Exposure to AWS, Azure or GCP environments supporting enterprise AI workloads.
Continuous AI Quality
Technical Leadership & Collaboration
Beyond technical execution, the role requires the ability to guide teams, establish standards and communicate AI quality risks clearly to engineering, product and business stakeholders.
Mentoring
Coach QA engineers and improve team capability in AI testing, automation and modern quality practices.
Stakeholder Collaboration
Work with developers, data scientists, product owners and business stakeholders to define AI quality requirements.
Quality Reporting
Establish meaningful quality metrics and communicate testing results, risks and improvement opportunities.
Continuous Improvement
Identify systemic quality problems and introduce better testing approaches, tools and engineering practices.
What You Should Bring
Additional Skills That Add Value
Technology & Testing Expertise
Lead the Future of AI Quality
Bring together AI testing, automation engineering and quality leadership to build reliable, scalable and production-ready Generative AI applications.
Apply Now ↗