Lead QA – Agentic AI
Lead the quality strategy for autonomous and agentic AI systems, covering AI testing, risk assessment, autonomous behavior validation, safety, performance and production readiness.
Who Is This Role For?
This role is suited for a senior QA leader who understands that testing agentic AI is fundamentally different from traditional deterministic software testing and can establish a risk-based quality strategy for autonomous systems.
Senior AI QA Leader
An experienced QA professional capable of owning end-to-end testing strategy for complex AI and autonomous systems.
Agentic AI Testing Expert
Someone who can evaluate reasoning chains, planning, tool orchestration, multi-agent collaboration and autonomous behavior.
AI Risk & Safety Specialist
Strong in hallucination testing, adversarial testing, red teaming, unsafe behavior detection, bias and AI guardrails.
Quality Strategy Owner
Able to define acceptance criteria, quality gates, evaluation metrics and release recommendations for probabilistic systems.
Quality Engineering for Autonomous AI
The role owns the quality strategy for Agentic AI platforms. The focus is on validating autonomous behavior, systemic risks, safety, reliability and production readiness across complex agent workflows.
What You Will Do
Define End-to-End Agentic AI Test Strategy
Establish comprehensive testing approaches for autonomous, agentic and LLM-powered systems across their lifecycle.
Design Complex Agent Scenarios
Create scenarios covering reasoning chains, tool orchestration, planning behavior and multi-agent collaboration.
Lead AI Safety Testing
Test hallucinations, bias, ethical risks, unsafe behavior, adversarial inputs and other AI-specific failure modes.
Validate Human-in-the-Loop Controls
Verify escalation mechanisms, human intervention points, auditability and traceability of agent decisions.
Review AI Testing Evidence
Review and approve AI test cases, defect reports, evaluation results and validation evidence.
Partner With Engineering & Compliance
Work with engineering, product and compliance teams to identify risks and establish appropriate mitigation strategies.
Lead Performance & Scale Testing
Define and execute performance, scalability and reliability testing for AI agents and their supporting architectures.
Mentor QA Teams
Develop QA capabilities in Agentic AI testing and train resources on modern AI quality engineering practices.
Define Release Recommendations
Provide risk-based quality assessments and recommendations regarding production readiness.
What Makes Agent Testing Different?
Traditional software testing generally expects deterministic behavior. Agentic systems introduce planning, memory, tool use, probabilistic outputs and autonomous decisions that require a broader evaluation strategy.
Reasoning Chains
Evaluate whether the agent follows an appropriate reasoning path and reaches a valid outcome.
Tool Orchestration
Validate correct tool selection, parameters, sequencing and handling of tool failures.
Multi-Agent Behavior
Test communication, coordination and responsibility boundaries between multiple autonomous agents.
Test How Agents Fail
Planning Failures
Determine whether agents create invalid, incomplete or unnecessarily complex plans.
Looping Behavior
Detect agents that repeatedly execute actions without meaningful progress toward the goal.
Dead-End Workflows
Validate how agents behave when tools, data, permissions or required information are unavailable.
Self-Correction
Test whether agents recognize errors and recover appropriately instead of propagating incorrect decisions.
Escalation Failures
Verify that autonomous systems escalate uncertain or risky situations to humans at the correct point.
Stateful Workflows
Validate long-running agents that maintain state and context across multiple interactions and execution stages.
Adversarial & Safety Evaluation
Test Beyond the Happy Path
Agentic AI QA requires intentionally challenging the system with unexpected instructions, conflicting goals, adversarial prompts, unsafe requests and abnormal workflow states to expose failures before production.
Analyze Agent Behavior Through Traces
Effective agent testing requires understanding not only the final answer but also what happened during execution.
Logs
Review execution logs to identify failures, unexpected actions and system-level anomalies.
Traces
Analyze agent traces to understand planning, tool calls, transitions and execution paths.
Telemetry
Use operational telemetry to identify reliability, performance and behavioral patterns over time.
Testing Under Uncertainty
Agentic AI does not always produce identical outputs for the same input. The QA strategy must therefore establish measurable quality boundaries rather than relying only on binary pass/fail testing.
Define Acceptance Criteria
Establish clear boundaries for probabilistic outputs and determine what constitutes acceptable behavior.
Measure Reliability
Create metrics that evaluate accuracy, consistency, reliability and safety across repeated evaluations.
Assess Business Risk
Connect agent behavior to business and compliance risks rather than evaluating technical output in isolation.
Make Risk-Based Decisions
Use evaluation evidence to determine whether an AI system is ready for release or requires additional mitigation.
Automation & Engineering Skills
Python Automation
Use Python to build reusable automation for AI testing, API validation, evaluation and data processing.
Robot Framework
Exposure to Robot Framework and automation strategy for scalable quality engineering.
API Automation
Test API-driven and distributed agent architectures across services and orchestration layers.
AI Governance & Production Readiness
Additional Experience That Adds Value
What You Should Bring
Technology & Testing Expertise
Lead the Quality of Agentic AI
Shape testing strategies for autonomous AI systems and help establish the quality, safety, reliability and governance standards required for production-ready agentic applications.
Apply Now ↗