LEAD QA • AGENTIC AI • AI QUALITY ENGINEERING

Lead QA – Agentic AI

Principal Lead Engineer

Lead the quality strategy for autonomous and agentic AI systems, covering AI testing, risk assessment, autonomous behavior validation, safety, performance and production readiness.

Experience: 9–15 Years
Role: Principal Lead Engineer
Focus: Agentic AI QA
Automation: Python / Robot Framework
AI: LLM / Agent Testing
Apply Now ↗

Who Is This Role For?

This role is suited for a senior QA leader who understands that testing agentic AI is fundamentally different from traditional deterministic software testing and can establish a risk-based quality strategy for autonomous systems.

01

Senior AI QA Leader

An experienced QA professional capable of owning end-to-end testing strategy for complex AI and autonomous systems.

02

Agentic AI Testing Expert

Someone who can evaluate reasoning chains, planning, tool orchestration, multi-agent collaboration and autonomous behavior.

03

AI Risk & Safety Specialist

Strong in hallucination testing, adversarial testing, red teaming, unsafe behavior detection, bias and AI guardrails.

04

Quality Strategy Owner

Able to define acceptance criteria, quality gates, evaluation metrics and release recommendations for probabilistic systems.

Quality Engineering for Autonomous AI

The role owns the quality strategy for Agentic AI platforms. The focus is on validating autonomous behavior, systemic risks, safety, reliability and production readiness across complex agent workflows.

Plan Define AI test strategy
Challenge Test adversarial behavior
Evaluate Measure agent quality
Analyze Study traces and failures
Release Assess production readiness

What You Will Do

01

Define End-to-End Agentic AI Test Strategy

Establish comprehensive testing approaches for autonomous, agentic and LLM-powered systems across their lifecycle.

02

Design Complex Agent Scenarios

Create scenarios covering reasoning chains, tool orchestration, planning behavior and multi-agent collaboration.

03

Lead AI Safety Testing

Test hallucinations, bias, ethical risks, unsafe behavior, adversarial inputs and other AI-specific failure modes.

04

Validate Human-in-the-Loop Controls

Verify escalation mechanisms, human intervention points, auditability and traceability of agent decisions.

05

Review AI Testing Evidence

Review and approve AI test cases, defect reports, evaluation results and validation evidence.

06

Partner With Engineering & Compliance

Work with engineering, product and compliance teams to identify risks and establish appropriate mitigation strategies.

07

Lead Performance & Scale Testing

Define and execute performance, scalability and reliability testing for AI agents and their supporting architectures.

08

Mentor QA Teams

Develop QA capabilities in Agentic AI testing and train resources on modern AI quality engineering practices.

09

Define Release Recommendations

Provide risk-based quality assessments and recommendations regarding production readiness.

What Makes Agent Testing Different?

Traditional software testing generally expects deterministic behavior. Agentic systems introduce planning, memory, tool use, probabilistic outputs and autonomous decisions that require a broader evaluation strategy.

Reasoning Chains

Evaluate whether the agent follows an appropriate reasoning path and reaches a valid outcome.

Tool Orchestration

Validate correct tool selection, parameters, sequencing and handling of tool failures.

Multi-Agent Behavior

Test communication, coordination and responsibility boundaries between multiple autonomous agents.

Test How Agents Fail

Planning Failures

Determine whether agents create invalid, incomplete or unnecessarily complex plans.

Looping Behavior

Detect agents that repeatedly execute actions without meaningful progress toward the goal.

Dead-End Workflows

Validate how agents behave when tools, data, permissions or required information are unavailable.

Self-Correction

Test whether agents recognize errors and recover appropriately instead of propagating incorrect decisions.

Escalation Failures

Verify that autonomous systems escalate uncertain or risky situations to humans at the correct point.

Stateful Workflows

Validate long-running agents that maintain state and context across multiple interactions and execution stages.

Adversarial & Safety Evaluation

Hallucination Testing Bias Testing Adversarial Testing Red Teaming Prompt Evaluation Unsafe Behavior Guardrail Testing AI Risk Responsible AI Safety Validation

Test Beyond the Happy Path

Agentic AI QA requires intentionally challenging the system with unexpected instructions, conflicting goals, adversarial prompts, unsafe requests and abnormal workflow states to expose failures before production.

Analyze Agent Behavior Through Traces

Effective agent testing requires understanding not only the final answer but also what happened during execution.

Logs

Review execution logs to identify failures, unexpected actions and system-level anomalies.

Traces

Analyze agent traces to understand planning, tool calls, transitions and execution paths.

Telemetry

Use operational telemetry to identify reliability, performance and behavioral patterns over time.

Testing Under Uncertainty

Agentic AI does not always produce identical outputs for the same input. The QA strategy must therefore establish measurable quality boundaries rather than relying only on binary pass/fail testing.

01

Define Acceptance Criteria

Establish clear boundaries for probabilistic outputs and determine what constitutes acceptable behavior.

02

Measure Reliability

Create metrics that evaluate accuracy, consistency, reliability and safety across repeated evaluations.

03

Assess Business Risk

Connect agent behavior to business and compliance risks rather than evaluating technical output in isolation.

04

Make Risk-Based Decisions

Use evaluation evidence to determine whether an AI system is ready for release or requires additional mitigation.

Automation & Engineering Skills

Python Automation

Use Python to build reusable automation for AI testing, API validation, evaluation and data processing.

Robot Framework

Exposure to Robot Framework and automation strategy for scalable quality engineering.

API Automation

Test API-driven and distributed agent architectures across services and orchestration layers.

Python Robot Framework API Testing Automation Frameworks Distributed Systems

AI Governance & Production Readiness

AI Governance Understanding of governance practices and controls for AI systems.
Model Risk Management Ability to identify and assess risks associated with AI-driven decisions and autonomous behavior.
Guardrails Validate controls that restrict unsafe, unauthorized or inappropriate agent behavior.
Auditability Ensure agent activity can be traced, reviewed and supported with appropriate validation evidence.
Release Readiness Provide evidence-based recommendations for production release based on risk and quality assessments.
Compliance Alignment Experience or exposure to regulated environments and AI validation expectations is valuable.

Additional Experience That Adds Value

CSA GxP AI Validation AI Governance Regulatory Audits Automation Strategy Claude AWS Bedrock LLM Applications Agentic AI

What You Should Bring

Experience 9–15 years of professional experience in QA, testing or quality engineering.
Education Bachelor's degree preferred in Computer Science, Electrical Engineering, Physics, Mathematics or related discipline.
Agent Frameworks Deep understanding of agent frameworks and orchestration layers.
Autonomous Workflows Experience testing autonomous workflows, decision policies, planning and self-correction.
Distributed Architecture Strong experience testing API-driven and distributed agent architectures.
AI Evaluation Ability to define evaluation metrics for accuracy, reliability, safety and probabilistic outputs.

Technology & Testing Expertise

Agentic AI AI QA LLM Testing Autonomous Agents Agent Orchestration Prompt Testing Red Teaming Adversarial Testing Hallucination Testing Bias Testing AI Safety Python Robot Framework API Testing Performance Testing AI Governance Model Risk AWS Bedrock Claude

Lead the Quality of Agentic AI

Shape testing strategies for autonomous AI systems and help establish the quality, safety, reliability and governance standards required for production-ready agentic applications.

Apply Now ↗