
AI Quality Engineer (Contract)
Side โข Hyderabad, Telangana, India
**Role & seniority: ** AI Quality Engineer (contract, 3 months with possible extension based on project/performance)
**Location & work type: ** Hyderabad, WFO 5 days/week
**Stack/tools (from text): **
-
LLM & RAG testing
-
Test automation: PyTest (nice-to-have)
-
AI eval/observability tools (nice-to-have): RAGAS, DeepEval, LangSmith
-
CI/CD test automation & guardrails (nice-to-have)
-
Cloud/observability: AWS/Azure/GCP (nice-to-have)
**Top 3 responsibilities: **
-
Test/validate AI outputs: LLM-generated insights, recommendations, and decision workflows.
-
Evaluate LLM/RAG quality: accuracy, relevance, consistency, factuality, grounding, and hallucinations.
-
Define evaluation criteria/datasets/metrics; run regression testing across models, prompts, RAG configs, and AI workflows.
**Must-have skills: **
-
Strong understanding of AI/ML and Generative AI testing.
-
Hands-on testing experience with LLM and RAG-based applications.
-
Knowledge of LLM evaluation (hallucination detection, relevance, response quality).
-
Understanding of AI agents and recommendation systems.
-
Analytical/problem-solving skills and ability to collaborate with AI/ML engineers.
**Nice-to-haves: **
-
PyTest and automated testing frameworks.
-
Building automated AI evaluation/regression frameworks.
-
Familiarity with RAGAS/DeepEval/LangSmith or equivalents.
-
Experience with CI/CD-based AI tes
Full Description
AI Quality Engineer
Location: - Hyderabad
Work Mode: - WFO (5 days a week)
Role Type: - Contractual (3 months and extension would depend on project requirement and performance)
Key Responsibilities Test and validate AI-generated insights, recommendations, and decision-making workflows. Evaluate LLM and RAG systems for accuracy, relevance, consistency, factuality, and hallucinations. Validate retrieval quality, context relevance, grounding, and response quality in RAG systems. Test AI agents and autonomous workflows across functional, negative, and edge-case scenarios. Define AI evaluation criteria, test datasets, quality metrics, and validation processes. Perform regression testing for models, prompts, RAG configurations, and AI workflows. Collaborate with AI/ML engineers to identify issues and improve AI system quality. Required Skills Strong understanding of AI/ML and Generative AI testing. Hands-on experience testing LLM and RAG-based applications. Knowledge of LLM evaluation, hallucination detection, relevance, and response quality. Understanding of AI agents and recommendation systems. Strong analytical and problem-solving skills. Good to Have Experience with PyTest and automated testing frameworks. Experience building automated AI evaluation and regression frameworks. Familiarity with tools such as RAGAS, DeepEval, LangSmith, or equivalent. Experience with CI/CD-based test automation, performance testing, or AI guardrails. Familiarity with cloud platforms (AWS/Azure/GCP) and observability tools. Success Metrics High accuracy, relevance, and reliability of AI outputs. Strong evaluation coverage across critical AI workflows. Early detection and reduction of hallucinations and AI regressions. Reduced production AI quality issues. Increased confidence and trust in AI-generated insights and recommendations.