Enter a job title or keyword

Hiring || Gen AI Quality Assurance (BLRPune)

2coms


Job Location:

Bengaluru - India

Monthly Salary: Not provided by the employer
Posted: 22 September 2026 (8 hours ago)
Application Deadline: 20 December 2026
Vacancies: 1 Vacancy

Job Summary

Gen AI Quality Assurance
Summary
We are seeking an experienced Quality Engineer with 7 to 12 years of background to lead the validation of advanced Generative AI applications. This role focuses on ensuring the reliability and accuracy of diverse AI agents including chatbots Retrieval-Augmented Generation (RAG) systems and various summarization tools. The successful candidate will architect and deploy robust automated testing frameworks utilizing tools like pytest DeepEval and RAGAS to assess agent performance. Proficiency in the Arize platform is highly desirable. The position involves rigorous evaluation of AI interactions via APIs employing both standard automation protocols and innovative LLM-as-a-Judge methodologies to measure critical quality indicators such as hallucination rates faithfulness contextual precision and bias.

Location: Bangalore/Pune

Responsibilities
  • Execute comprehensive testing strategies for a wide range of AI agents including conversational chatbots RAG-driven systems document summarizers and call note analyzers.
  • Validate AI-generated outputs against key performance standards ensuring accuracy relevance consistency and completeness while identifying issues like hallucinations or bias.
  • Architect and maintain automated evaluation pipelines leveraging the DeepEval and RAGAS frameworks to create reusable test suites for RAG performance summarization fidelity and groundedness.
  • Develop and refine LLM-as-a-Judge evaluation protocols establishing clear rubrics and custom scoring mechanisms to objectively assess AI responses.
  • Engineer scalable automation scripts using pytest to validate API interactions test prompt-response dynamics and manage regression testing for AI models.
  • Integrate automated test suites with CI/CD pipelines ensuring seamless execution in both local and production environments.
  • Capture and analyze detailed quality metrics including toxicity levels answer correctness and overall response reliability to drive continuous improvement.
  • Parameterize test cases to accommodate diverse prompts contexts and expected outcomes facilitating batch evaluations of large test datasets.


Requirements

  • 7 to 12 years of professional experience in software quality assurance with a specialized focus on AI and Machine Learning systems.
  • Proven expertise in designing developing and executing automated test frameworks specifically using Python and pytest.
  • Hands-on experience implementing evaluation frameworks such as DeepEval and RAGAS for assessing Generative AI quality.
  • Strong understanding of LLM-as-a-Judge techniques and the ability to define custom evaluation criteria and rubrics.
  • Experience testing AI agents that communicate via APIs including validation of request and response structures.
  • Familiarity with the Arize platform is considered a significant advantage.
  • Ability to analyze and interpret complex metrics related to faithfulness contextual recall hallucination and bias.
  • Experience integrating test automation into CI/CD workflows and managing batch processing for test datasets.



Required Skills:

Gen AI Quality Assurance Summary We are seeking an experienced Quality Engineer with 7 to 12 years of background to lead the validation of advanced Generative AI applications. This role focuses on ensuring the reliability and accuracy of diverse AI agents including chatbots Retrieval-Augmented Generation (RAG) systems and various summarization tools. The successful candidate will architect and deploy robust automated testing frameworks utilizing tools like pytest DeepEval and RAGAS to assess agent performance. Proficiency in the Arize platform is highly desirable. The position involves rigorous evaluation of AI interactions via APIs employing both standard automation protocols and innovative LLM-as-a-Judge methodologies to measure critical quality indicators such as hallucination rates faithfulness contextual precision and bias. Location: Bangalore/Pune Responsibilities Execute comprehensive testing strategies for a wide range of AI agents including conversational chatbots RAG-driven systems document summarizers and call note analyzers. Validate AI-generated outputs against key performance standards ensuring accuracy relevance consistency and completeness while identifying issues like hallucinations or bias. Architect and maintain automated evaluation pipelines leveraging the DeepEval and RAGAS frameworks to create reusable test suites for RAG performance summarization fidelity and groundedness. Develop and refine LLM-as-a-Judge evaluation protocols establishing clear rubrics and custom scoring mechanisms to objectively assess AI responses. Engineer scalable automation scripts using pytest to validate API interactions test prompt-response dynamics and manage regression testing for AI models. Integrate automated test suites with CI/CD pipelines ensuring seamless execution in both local and production environments. Capture and analyze detailed quality metrics including toxicity levels answer correctness and overall response reliability to drive continuous improvement. Parameterize test cases to accommodate diverse prompts contexts and expected outcomes facilitating batch evaluations of large test datasets. Requirements 7 to 12 years of professional experience in software quality assurance with a specialized focus on AI and Machine Learning systems. Proven expertise in designing developing and executing automated test frameworks specifically using Python and pytest. Hands-on experience implementing evaluation frameworks such as DeepEval and RAGAS for assessing Generative AI quality. Strong understanding of LLM-as-a-Judge techniques and the ability to define custom evaluation criteria and rubrics. Experience testing AI agents that communicate via APIs including validation of request and response structures. Familiarity with the Arize platform is considered a significant advantage. Ability to analyze and interpret complex metrics related to faithfulness contextual recall hallucination and bias. Experience integrating test automation into CI/CD workflows and managing batch processing for test datasets.


Required Education:

Graduate