Kumaran Systems logo

ETL Testing

Kumaran Systems Toronto, Ontario, Canada

onsitefull-time
Posted Sep 23, 2026Apply by Oct 23, 2026
  • Role & seniority

    • ETL QA Engineer (seniority not explicitly stated; role focuses on strong hands-on experience)
  • Stack/tools

    • Data/ETL: Databricks (notebooks, jobs, clusters, Delta Lake, SQL), PySpark, ETL/ELT, data warehousing/data modeling (star/snowflake, dimension/fact)

    • QA/testing: automated testing with Python frameworks (e.g., pytest/unittest), data validation/reconciliation, anomaly detection/quality rules

    • CI/CD & tooling: GitHub, GitHub Copilot, GitHub Actions/Runners

    • AI/knowledge: prompt engineering for LLMs, agent-based workflow/agent frameworks, AI tools for code/test analysis or knowledge management, vector search/embeddings/LLM retrieval for knowledge bases

  • Top 3 responsibilities

    • Validate and QA large-scale ETL/ELT pipelines in Databricks (data quality checks, reconciliation, validation rules)

    • Build/maintain automated regression and data validation tests (test planning/design, defect management)

    • Support AI-assisted development: craft prompts for coding/testing/docs and help build/validate AI-powered knowledge bases

  • Must-have skills

    • Strong Databricks experience (Delta Lake, SQL, jobs/clusters/notebooks)

    • Expert PySpark (transformations/actions/window functions/optimization)

    • Data validation and reconciliation techniques for ETL QA

    • CI/CD practices using GitHub + GitHub Actions/Runners (with Copilot exposure)

    • LLM prompt engineering for coding/testing/documentation; familiarity

Full Description

Role Overview We are looking for an ETL QA Engineer with strong experience in Databricks and PySpark, combined with hands-on exposure to AI-assisted development and prompt engineering. The ideal candidate will validate large-scale data pipelines, ensure data quality, and leverage tools like GitHub Copilot, GitHub Actions/Runners, and AI agents to improve testing efficiency and coverage. You will also help build and validate AI-powered knowledge bases.

Required Skills & Experience Core Technical Skills Strong hands-on experience in Databricks (notebooks, jobs, clusters, Delta Lake, SQL). Expert-level PySpark skills, including transformations, actions, window functions, and optimization techniques. Solid understanding of ETL/ELT concepts, data warehousing, and data modeling (star/snowflake schemas, dimension/fact tables). Proven experience in data validation, data quality checks, and reconciliation techniques. Practical experience with GitHub, GitHub Copilot, and GitHub Actions/Runners for CI/CD pipelines. AI & Prompt Engineering Experience crafting effective prompts for large language models (LLMs) for coding, testing, and documentation. Exposure to agent-based architectures or tools (e.g., workflow/agent frameworks that orchestrate multi-step AI tasks). Familiarity with AI-based tools for code analysis, test generation, or knowledge management. Knowledge Base & Documentation Experience building AI-driven knowledge bases or documentation systems (e.g., using vector search, embeddings, or LLM-based retrieval). Strong documentation skills, with the ability to translate complex data/QA concepts into clear knowledge articles. Testing & QA Practices

Strong understanding of QA methodologies: test planning, test design, defect management, and regression testing. Experience with automated testing frameworks (Python-based testing libraries such as pytest, unittest, or similar). Familiarity with data quality tools/practices (e.g., validation rules, thresholds, anomaly detection).

Key Attributes Strong analytical and problem-solving skills with a detail-oriented mindset. Passion for data quality, automation, and continuous improvement. Ability to work in an agile, fast-paced environment and collaborate across multiple teams like Business and operations team and AD team. Curiosity and openness to adopting new AI tools and practices for QA.

DatabricksPySparkETL/ELT TestingData ValidationData QualityData WarehousingData ModelingGitHubGitHub CopilotGitHub ActionsPrompt EngineeringLarge Language ModelsAutomated TestingPythonKnowledge BasesRegression Testingmulti-location

Cookies & analytics consent

We serve candidates globally, so we only activate Google Tag Manager and other analytics after you opt in. This keeps us aligned with GDPR/UK DPA, ePrivacy, LGPD, and similar rules. Essential features still run without analytics cookies.

Read how we use data in our Privacy Policy and Terms of Service.