Side logo

AI Quality Engineer (Contract)

Side โ€ข Hyderabad, Telangana, India

onsitecontract
Posted Aug 28, 2026

**Role & seniority: ** AI Quality Engineer (contract, 3 months with possible extension based on project/performance)

**Location & work type: ** Hyderabad, WFO 5 days/week

**Stack/tools (from text): **

  • LLM & RAG testing

  • Test automation: PyTest (nice-to-have)

  • AI eval/observability tools (nice-to-have): RAGAS, DeepEval, LangSmith

  • CI/CD test automation & guardrails (nice-to-have)

  • Cloud/observability: AWS/Azure/GCP (nice-to-have)

**Top 3 responsibilities: **

  • Test/validate AI outputs: LLM-generated insights, recommendations, and decision workflows.

  • Evaluate LLM/RAG quality: accuracy, relevance, consistency, factuality, grounding, and hallucinations.

  • Define evaluation criteria/datasets/metrics; run regression testing across models, prompts, RAG configs, and AI workflows.

**Must-have skills: **

  • Strong understanding of AI/ML and Generative AI testing.

  • Hands-on testing experience with LLM and RAG-based applications.

  • Knowledge of LLM evaluation (hallucination detection, relevance, response quality).

  • Understanding of AI agents and recommendation systems.

  • Analytical/problem-solving skills and ability to collaborate with AI/ML engineers.

**Nice-to-haves: **

  • PyTest and automated testing frameworks.

  • Building automated AI evaluation/regression frameworks.

  • Familiarity with RAGAS/DeepEval/LangSmith or equivalents.

  • Experience with CI/CD-based AI tes

Full Description

AI Quality Engineer

Location: - Hyderabad

Work Mode: - WFO (5 days a week)

Role Type: - Contractual (3 months and extension would depend on project requirement and performance)

Key Responsibilities Test and validate AI-generated insights, recommendations, and decision-making workflows. Evaluate LLM and RAG systems for accuracy, relevance, consistency, factuality, and hallucinations. Validate retrieval quality, context relevance, grounding, and response quality in RAG systems. Test AI agents and autonomous workflows across functional, negative, and edge-case scenarios. Define AI evaluation criteria, test datasets, quality metrics, and validation processes. Perform regression testing for models, prompts, RAG configurations, and AI workflows. Collaborate with AI/ML engineers to identify issues and improve AI system quality. Required Skills Strong understanding of AI/ML and Generative AI testing. Hands-on experience testing LLM and RAG-based applications. Knowledge of LLM evaluation, hallucination detection, relevance, and response quality. Understanding of AI agents and recommendation systems. Strong analytical and problem-solving skills. Good to Have Experience with PyTest and automated testing frameworks. Experience building automated AI evaluation and regression frameworks. Familiarity with tools such as RAGAS, DeepEval, LangSmith, or equivalent. Experience with CI/CD-based test automation, performance testing, or AI guardrails. Familiarity with cloud platforms (AWS/Azure/GCP) and observability tools. Success Metrics High accuracy, relevance, and reliability of AI outputs. Strong evaluation coverage across critical AI workflows. Early detection and reduction of hallucinations and AI regressions. Reduced production AI quality issues. Increased confidence and trust in AI-generated insights and recommendations.

AI/MLGenerative AILLMRAGHallucination detectionRecommendation systemsPyTestAutomated testingRAGASDeepEvalLangSmithCI/CDPerformance testingAWSAzureGCPmulti-location

Cookies & analytics consent

We serve candidates globally, so we only activate Google Tag Manager and other analytics after you opt in. This keeps us aligned with GDPR/UK DPA, ePrivacy, LGPD, and similar rules. Essential features still run without analytics cookies.

Read how we use data in our Privacy Policy and Terms of Service.