Job Position : Senior AI Quality Assurance (QA) Engineer
Location : Bengaluru
Experience : 7+ Years
Required Skills:
We are looking for an experienced Senior AI Quality Assurance (QA) Engineer to ensure the quality, reliability, and production readiness of enterprise-grade AI applications.
This role extends beyond traditional software testing and focuses on the evaluation and quality engineering of LLM applications, RAG systems, and AI agents. You will own automated testing, AI evaluations, benchmark creation, observability, prompt regression testing, and end-to-end validation of AI workflows. Working closely with AI Engineers, Product Managers, and Platform teams, you will establish measurable quality standards and ensure every release meets enterprise-grade expectations for accuracy, reliability, performance, and scalability.
Key Responsibilities:
AI Evaluation & Benchmarking
Design benchmark (golden) datasets and build automated evaluation pipelines for LLM applications. Define quality gates and continuously evaluate prompts, models, retrieval pipelines, and agent behaviour using metrics such as hallucination rate, tool selection accuracy, execution accuracy, precision, recall, latency, and cost.
RAG & Agent Quality Validation
Validate Retrieval-Augmented Generation (RAG) pipelines and AI agent workflows by testing retrieval quality, context relevance, tool invocation, reasoning flow, memory, and end-to-end task completion. Design evaluation scenarios covering ambiguous queries, multi-turn conversations, retrieval failures, and edge cases.
Python Automation & API Testing
Develop and maintain scalable automation frameworks using Python and Pytest for unit testing, integration testing, API testing, regression testing, and end-to-end validation. Build reusable test utilities and integrate automated quality checks into CI/CD pipelines.
Frontend Automation
Develop automated UI test suites using Playwright to validate AI-powered user journeys, conversational interfaces, workflow execution, and end-to-end application behaviour across releases.
Observability & Root Cause Analysis
Use OpenTelemetry, tracing platforms, and AI observability tools to analyse execution traces, latency, model responses, API calls, and workflow behaviour. Perform root cause analysis to identify regressions, hallucinations, bottlenecks, and production issues.
Performance & Enterprise Readiness
Validate AI application performance by monitoring latency, throughput, reliability, and scalability. Ensure production readiness through regression testing, API validation, workflow testing, and enterprise quality standards.
Required Skills:
7+ years of experience in Software QA, Test Automation, or AI Quality Engineering.
Strong Python programming skills.
Hands-on experience with Pytest for:
Unit testing
Integration testing
API testing
Regression testing
Experience testing REST APIs and backend services.
Experience evaluating LLM-powered applications.
Hands-on experience evaluating Retrieval-Augmented Generation (RAG) systems.
Understanding of AI agent evaluation methodologies.
Experience measuring AI quality using metrics such as:
Hallucination Rate
Tool Selection Accuracy
Execution Accuracy
Precision / Recall
Latency (P50/P95/P99)
Token Usage
Experience with Playwright or similar frontend automation frameworks.
Experience with OpenTele metry, tracing, or observability platforms.
Strong debugging and root cause analysis skills.
Experience integrating automated tests into CI/CD pipelines.
Nice to Have
Experience with Lang Smith, Lang fuse, MLflow, Arize Phoenix, or similar AI observability platforms.
Experience evaluating multi-agent systems and orchestration frameworks such as LangGraph, CrewAI, Google ADK or AutoGen.
Experience with vector databases such as Pinecone, Milvus, Weaviate, pgvector, or Vertex AI Vector Search.
Exposure to OpenAI, Anthropic, Gemini, or Azure OpenAI.
Experience with performance and load testing tools.
Prior experience in Banking, Financial Services, or Insurance (BFSI).


