Tata Consultancy Services logo

AI Test Engineer

Tata Consultancy Services • Bengaluru, Karnataka, India

onsitefull-time
Posted Sep 24, 2026Apply by Oct 24, 2026
  • Role & seniority

    • AI Test Engineer (seniority not specified)
  • Stack/tools

    • Evaluation frameworks: RAGAS, (good to have) LangSmith

    • LLM evaluation approach: Human eval, automated eval, LLM-as-a-judge

    • Focus areas/metrics: accuracy/relevance/grounding/consistency/completeness, precision/recall, faithfulness, toxicity, bias

    • CI/CD: embed AI evaluation into pipelines (where applicable)

  • Top 3 responsibilities

    • Evaluate and validate AI/LLM responses (human + automated + LLM-as-judge) for accuracy, relevance, grounding, consistency, completeness.

    • Design test scenarios for RAG-based apps, emphasizing response faithfulness and reducing hallucinations.

    • Support Responsible AI/governance testing (fairness, bias, safety, transparency) including prompt/guardrail and jailbreak testing.

  • Must-have skills

    • Strong understanding of AI/LLM testing and AI quality assurance.

    • Experience with AI response evaluation (human/automated/LLM-as-a-judge) and issue identification (hallucinations, bias, inconsistencies).

    • Hands-on test data/dataset creation from business/functional requirements (golden, adversarial, edge cases).

    • Awareness of Responsible AI principles and governance models.

  • Nice-to-haves

    • Experience with RAG evaluation frameworks such as RAGAS (explicitly mentioned; “good to have” also suggests tool familiarity).

    • LangSmith experience.

  • Location & work type

    • Not provided in the text.

Full Description

PFB JD for AI Test Engineers Key Responsibilities AI Response Evaluation & Validation

Evaluate AI/LLM responses using multiple approaches including

  • Human Evaluation
  • Automated Evaluation
  • LLM‑as‑a‑Judge techniques
  • Validate AI outputs for accuracy, relevance, grounding, consistency, and completeness.
  • Design and execute test scenarios for RAG‑based applications, ensuring response faithfulness and reduced hallucinations.
  • AI Evaluation Frameworks & Tooling

Apply or integrate AI evaluation frameworks such as

  • RAGAS
  • LangSmith (good to have)
  • Define evaluation metrics (precision, recall, faithfulness, toxicity, bias, etc.) and analyze results.
  • Collaborate with engineering teams to embed AI evaluation into CI/CD pipelines where applicable.
  • Responsible AI & Governance

Ensure compliance with Responsible AI guidelines, including

  • Fairness
  • Bias detection
  • Safety
  • Transparency
  • Validate system instructions, prompt templates, and guardrails.
  • Support system instruction extraction, prompt testing, and jailbreak prevention strategies.
  • Test Data & Dataset Management

Understand business and functional requirements to

  • Design and curate high‑quality datasets (golden datasets, adversarial datasets, edge‑case datasets).
  • Prepare datasets for training validation, evaluation, and regression testing.
  • Maintain dataset versioning and traceability.
  • Collaboration & Quality Advocacy
  • Work closely with AI engineers, product owners, and domain SMEs to understand AI behavior and risks.
  • Provide actionable insights on AI quality issues, improvements, and risks.
  • Contribute to AI testing best practices, standards, and reusable assets.

Required Skills & Qualifications Mandatory Skills Strong understanding of AI/LLM testing concepts and AI quality assurance. Experience in AI response evaluation techniques (human, automated, LLM‑as‑judge).

Ability to analyze AI outputs and identify

  • Hallucinations
  • Bias
  • Inconsistencies
  • Hands‑on experience in test data / dataset creation based on requirements.
  • Awareness of Responsible AI principles and governance models.
  • Good to Have
  • Experience with RAG evaluation frameworks such as RAGA
AI/LLM TestingAI Response EvaluationHuman EvaluationAutomated EvaluationLLM-as-a-JudgeRAG TestingRAGASAI Quality AssuranceHallucination DetectionBias DetectionTest Data CreationDataset ManagementResponsible AIPrompt TestingJailbreak PreventionCI/CD Integrationmulti-location

Cookies & analytics consent

We serve candidates globally, so we only activate Google Tag Manager and other analytics after you opt in. This keeps us aligned with GDPR/UK DPA, ePrivacy, LGPD, and similar rules. Essential features still run without analytics cookies.

Read how we use data in our Privacy Policy and Terms of Service.