Sii Poland logo

Senior Automation Test Engineer – AI Systems (f/m/x)

Sii Poland Szczecin, West Pomeranian Voivodeship, Poland

remotefull-time
Posted Sep 19, 2026Apply by Oct 19, 2026

**Role & seniority: ** Senior QA Engineer (AI/ML quality & evaluation ownership)

**Stack/tools: **

  • Languages: Python; JavaScript/TypeScript

  • AI testing/evaluation: LangSmith, Langfuse, EvidentlyAI, Ragas, MLflow

  • Drift/observability: Evidently/ML drift concepts; Datadog, Grafana

  • Web stack: Angular, Node.js, Python AI microservices

  • Other: CI/CD integration, telemetry/alerting, load/performance testing

  • Top 3 responsibilities:

    • Build and maintain LLM/agent evaluation frameworks (metrics, scoring, assertions, test suites).

    • Own end-to-end QA strategy across full stack (Angular ↔ Node.js ↔ Python AI services) incl. integration/API/performance/load tests and CI/CD quality gates.

    • Implement production monitoring for AI behavior (data/concept drift, latency, token usage/cost, output quality) with telemetry and alerting.

  • Must-have skills:

    • 5+ years QA/quality engineering; 2+ years testing/evaluating AI/ML, LLM workflows, or agentic systems.

    • Strong Python for automated test/evaluation scripts; working knowledge of JS/TS for Angular/Node testing.

    • Experience selecting/implementing evaluation metrics for classification and generation quality.

    • Understanding of ML drift monitoring and observability practices/tools (Datadog/Grafana).

    • Production-grade system design for test/evaluation infrastructure; advanced English.

  • Nice-to-haves:

    • Agentic design patterns and decomposing agents

Full Description

We are looking for a Senior QA Engineer to join our team and contribute to building an AI-powered operating system for healthcare marketing. In this role, you will go beyond traditional end-to-end testing, owning quality, evaluation, and monitoring for complex AI systems. You will work on real-world production-grade products that create monetizable value, ensuring the stability of our full-stack architecture while designing and scaling AI evaluation pipelines and ML drift monitoring systems. If you are passionate about developing robust test and evaluation frameworks for modern systems powered by AI and ML, we want to hear from you.

Your tasks

Designing, building, and maintaining evaluation frameworks for Large Language Models (LLMs) and custom agentic harnesses Developing automated scoring mechanisms, assertions, and test suites to measure agent accuracy, reasoning reliability, hallucination rates, and tool-use correctness Simulating complex, multi-step agent workflows to uncover edge cases in autonomous decision-making Collaborating with engineering teams to set up robust monitoring frameworks for data drift, concept drift, and model performance degradation in production Implementing telemetry and alerting systems to track latency, token usage, cost efficiency, and output quality over time Owning the end-to-end testing strategy across our entire stack, ensuring seamless integration between the Angular frontend, Node.js backend services, and Python-based AI microservices Writing clean, maintainable, and scalable automated test scripts in Python and JavaScript/TypeScript Performing rigorous API testing, integration testing, and performance/load testing for high-throughput microservices Integrating AI evaluation and testing harnesses into our CI/CD pipelines to ensure continuous quality gate enforcement before deployment Defining and tracking key quality metrics (KPIs) for software reliability and AI model behavior

Requirements

Minimum 5 years of experience in software testing or quality engineering, with at least 2 years focused on testing or evaluating AI/ML applications, LLM workflows, or agentic systems Proficiency in Python for backend/AI testing scripts and familiarity with JavaScript/TypeScript for Node.js and Angular testing ecosystems Experience with evaluation frameworks such as LangSmith, Langfuse, EvidentlyAI, Ragas, or MLflow, and the ability to select appropriate metrics for classification or generation-quality problems Familiarity with ML drift monitoring concepts, metrics tracking, and observability tools like Datadog or Grafana Strong system-level design skills and the ability to build production-grade test and evaluation infrastructure from the ground up Advanced level of English

Nice to have

Understanding of agentic design patterns and their decomposition into testable units Experience working under the regulatory conditions of healthcare systems

Job no. JOB-GTXQF

Sii ensures that all hiring decisions are made solely on the basis of qualifications and competence. We are committed to equal and fair treatment of all, regardless of legally protected characteristics. At Sii, we promote a diverse and inclusive work environment, in full compliance with applicable anti-discrimination laws.

Benefits For You

Great Place to Work Solid financial situation Contracts with the biggest brands Centre of internal trainings Many experts you can learn from Open and accessible management team Profit sharing Passion Sponsorship program Regular integration events and trips Comfortable and well-equipped offices MySii app Medical care

PythonJavaScriptTypeScriptLLM evaluationAgentic systemsTest automationCI/CDAPI testingPerformance testingLangSmithLangfuseEvidentlyAIRagasMLflowDatadogGrafanamulti-location

Cookies & analytics consent

We serve candidates globally, so we only activate Google Tag Manager and other analytics after you opt in. This keeps us aligned with GDPR/UK DPA, ePrivacy, LGPD, and similar rules. Essential features still run without analytics cookies.

Read how we use data in our Privacy Policy and Terms of Service.