Role Overview
We are looking for an experienced AI Agent Evaluation Engineer to validate AI-powered applications, conversational agents, and Large Language Models (LLMs). This role combines AI testing, automation engineering, safety evaluation, adversarial testing, and Responsible AI practices to ensure secure, reliable, and production-ready AI systems.
6–8 Years
AI Agent Evaluation
70% Automation • 30% Manual
Python
Key Responsibilities
- Evaluate AI agents, conversational AI, and LLM-powered applications.
- Design automated and exploratory AI evaluation frameworks.
- Conduct AI safety assessments and Responsible AI validation.
- Perform red teaming, adversarial testing, and jailbreak analysis.
- Validate prompt injection resistance and model robustness.
- Create Python automation using PyTest.
- Integrate Google ADK and Vertex AI evaluation workflows.
- Measure AI quality, reliability, fairness, and security.
Required Skills
Primary Focus Areas
- Responsible AI and AI Governance
- AI Safety Evaluation
- Prompt Engineering Validation
- Adversarial Testing & Red Teaming
- Conversational AI Testing
- Google ADK & Vertex AI Ecosystem
- LLM Evaluation Frameworks
- Automation Testing with Python
Ideal Candidate
The ideal candidate has strong Software QA experience with deep expertise in Generative AI testing, AI agent evaluation, Responsible AI practices, Google ADK, Python automation, and modern LLM evaluation frameworks. Experience in safety validation, adversarial testing, and AI quality measurement is highly valued.
Disclaimer: This job information is shared for educational and career guidance purposes only. Company names and trademarks belong to their respective owners. Please apply through the official job posting.