Pearlsoft Solutions Inc. logo

QA Engineer - Agentic & Generative AI

Pearlsoft Solutions Inc. • United States

onsitecontract
Posted Oct 3, 2026Apply by Nov 2, 2026

**Role & seniority: ** Agentic QA Engineer (Generative AI & Agentic Systems); Senior (7+ years QA/testing, 2+ years AI/ML/LLM)

**Stack/tools: ** Python; LLM eval metrics (exact/soft match, BLEU/ROUGE, BERTScore, embedding-based semantic similarity); guardrails & prompt testing; distributed systems testing/latency profiling; resiliency patterns (circuit breakers, retries), chaos engineering, message queues; orchestration frameworks (LangChain, LangGraph, LlamaIndex, DSPy, OpenAI Assistants/Actions, Azure OpenAI orchestration or similar); CI/CD (GitHub Actions, Azure DevOps); observability (OpenTelemetry, Prometheus/Grafana, Datadog); feature flags/canaries; AI privacy/security/compliance

**Top 3 responsibilities: **

  • Build and run test harnesses/simulators/fixtures for agentic and multi-agent architectures

  • Perform LLM and prompt/guardrail evaluation, including semantic and safety-related checks

  • Validate distributed behavior (latency, resiliency, resiliency testing/chaos, message-queue flows) and ensure CI/CD coverage

  • Must-have skills:

    • 7+ years QA/testing; 2+ years hands-on AI/ML or LLM testing

    • Strong Python for automation and testing infrastructure

    • Expertise in LLM evaluation methods + guardrails/prompt testing

    • Distributed systems testing (latency profiling, resiliency, chaos engineering, queues)

    • Experience with orchestration frameworks for agentic systems

    • CI/CD + observability + feature flags; solid security/priv

Full Description

Job title: Agentic QA Engineer – Generative AI & Agentic Systems (Agent, Multi‑Agent Testing)

Location: Dallas, TX (Onsite)

Duration: Long term Contract

Mode of interview: 2 F2F rounds interviews in Dallas, TX (Must)

No Third Party

Required Qualifications 7+ years in Software QA/Testing, with 2+ years in AI/ML or LLM-based systems; hands-on experience testing agentic/multi-agent architectures. Strong programming skills in Python experience building test harnesses, simulators, and fixtures. Experience with LLM evaluation (exact/soft match, BLEU/ROUGE, BERTScore, semantic similarity via embeddings), guardrails, and prompt testing. Expertise in distributed systems testing latency profiling, resiliency patterns (circuit breakers, retries), chaos engineering, and message queues. Familiarity with orchestration frameworks (LangChain, LangGraph, LlamaIndex, DSPy, OpenAI Assistants/Actions, Azure OpenAI orchestration, or similar). Proficiency with CI/CD (GitHub Actions/Azure DevOps), observability (OpenTelemetry, Prometheus/Grafana, Datadog), and feature flags/canaries. Solid understanding of privacy/security/compliance in AI systems (PII handling, content policies, model safety). Excellent communication and leadership skills; proven ability to work cross-functionally with Ops, Data, and Engineering. Preferred Qualifications Experience with multi-agent simulators, agent graph testing, and tooling latency emulation. Knowledge of MLOps (model versioning, datasets, evaluation pipelines) and A/B experimentation for LLMs. Background in cloud (AWS), serverless, containerization, and event-driven architectures. Prior ownership of cost/latency/SLAs for AI workloads in production.

Software QAPythonAI/ML TestingLLM EvaluationAgentic Systems TestingMulti-Agent TestingPrompt TestingDistributed Systems TestingChaos EngineeringCI/CDObservabilityMLOpsCloud ComputingPrivacy and SecurityCross-Functional Leadership

Cookies & analytics consent

We serve candidates globally, so we only activate Google Tag Manager and other analytics after you opt in. This keeps us aligned with GDPR/UK DPA, ePrivacy, LGPD, and similar rules. Essential features still run without analytics cookies.

Read how we use data in our Privacy Policy and Terms of Service.