
AI Test Engineer
Tata Consultancy Services • Bengaluru, Karnataka, India
-
Role & seniority
- AI Test Engineer (seniority not specified)
-
Stack/tools
-
Evaluation frameworks: RAGAS, (good to have) LangSmith
-
LLM evaluation approach: Human eval, automated eval, LLM-as-a-judge
-
Focus areas/metrics: accuracy/relevance/grounding/consistency/completeness, precision/recall, faithfulness, toxicity, bias
-
CI/CD: embed AI evaluation into pipelines (where applicable)
-
-
Top 3 responsibilities
-
Evaluate and validate AI/LLM responses (human + automated + LLM-as-judge) for accuracy, relevance, grounding, consistency, completeness.
-
Design test scenarios for RAG-based apps, emphasizing response faithfulness and reducing hallucinations.
-
Support Responsible AI/governance testing (fairness, bias, safety, transparency) including prompt/guardrail and jailbreak testing.
-
-
Must-have skills
-
Strong understanding of AI/LLM testing and AI quality assurance.
-
Experience with AI response evaluation (human/automated/LLM-as-a-judge) and issue identification (hallucinations, bias, inconsistencies).
-
Hands-on test data/dataset creation from business/functional requirements (golden, adversarial, edge cases).
-
Awareness of Responsible AI principles and governance models.
-
-
Nice-to-haves
-
Experience with RAG evaluation frameworks such as RAGAS (explicitly mentioned; “good to have” also suggests tool familiarity).
-
LangSmith experience.
-
-
Location & work type
- Not provided in the text.
Full Description
PFB JD for AI Test Engineers Key Responsibilities AI Response Evaluation & Validation
Evaluate AI/LLM responses using multiple approaches including
- Human Evaluation
- Automated Evaluation
- LLM‑as‑a‑Judge techniques
- Validate AI outputs for accuracy, relevance, grounding, consistency, and completeness.
- Design and execute test scenarios for RAG‑based applications, ensuring response faithfulness and reduced hallucinations.
- AI Evaluation Frameworks & Tooling
Apply or integrate AI evaluation frameworks such as
- RAGAS
- LangSmith (good to have)
- Define evaluation metrics (precision, recall, faithfulness, toxicity, bias, etc.) and analyze results.
- Collaborate with engineering teams to embed AI evaluation into CI/CD pipelines where applicable.
- Responsible AI & Governance
Ensure compliance with Responsible AI guidelines, including
- Fairness
- Bias detection
- Safety
- Transparency
- Validate system instructions, prompt templates, and guardrails.
- Support system instruction extraction, prompt testing, and jailbreak prevention strategies.
- Test Data & Dataset Management
Understand business and functional requirements to
- Design and curate high‑quality datasets (golden datasets, adversarial datasets, edge‑case datasets).
- Prepare datasets for training validation, evaluation, and regression testing.
- Maintain dataset versioning and traceability.
- Collaboration & Quality Advocacy
- Work closely with AI engineers, product owners, and domain SMEs to understand AI behavior and risks.
- Provide actionable insights on AI quality issues, improvements, and risks.
- Contribute to AI testing best practices, standards, and reusable assets.
Required Skills & Qualifications Mandatory Skills Strong understanding of AI/LLM testing concepts and AI quality assurance. Experience in AI response evaluation techniques (human, automated, LLM‑as‑judge).
Ability to analyze AI outputs and identify
- Hallucinations
- Bias
- Inconsistencies
- Hands‑on experience in test data / dataset creation based on requirements.
- Awareness of Responsible AI principles and governance models.
- Good to Have
- Experience with RAG evaluation frameworks such as RAGA