
ETL Testing
Kumaran Systems • Toronto, Ontario, Canada
-
Role & seniority
- ETL QA Engineer (seniority not explicitly stated; role focuses on strong hands-on experience)
-
Stack/tools
-
Data/ETL: Databricks (notebooks, jobs, clusters, Delta Lake, SQL), PySpark, ETL/ELT, data warehousing/data modeling (star/snowflake, dimension/fact)
-
QA/testing: automated testing with Python frameworks (e.g., pytest/unittest), data validation/reconciliation, anomaly detection/quality rules
-
CI/CD & tooling: GitHub, GitHub Copilot, GitHub Actions/Runners
-
AI/knowledge: prompt engineering for LLMs, agent-based workflow/agent frameworks, AI tools for code/test analysis or knowledge management, vector search/embeddings/LLM retrieval for knowledge bases
-
-
Top 3 responsibilities
-
Validate and QA large-scale ETL/ELT pipelines in Databricks (data quality checks, reconciliation, validation rules)
-
Build/maintain automated regression and data validation tests (test planning/design, defect management)
-
Support AI-assisted development: craft prompts for coding/testing/docs and help build/validate AI-powered knowledge bases
-
-
Must-have skills
-
Strong Databricks experience (Delta Lake, SQL, jobs/clusters/notebooks)
-
Expert PySpark (transformations/actions/window functions/optimization)
-
Data validation and reconciliation techniques for ETL QA
-
CI/CD practices using GitHub + GitHub Actions/Runners (with Copilot exposure)
-
LLM prompt engineering for coding/testing/documentation; familiarity
-
Full Description
Role Overview We are looking for an ETL QA Engineer with strong experience in Databricks and PySpark, combined with hands-on exposure to AI-assisted development and prompt engineering. The ideal candidate will validate large-scale data pipelines, ensure data quality, and leverage tools like GitHub Copilot, GitHub Actions/Runners, and AI agents to improve testing efficiency and coverage. You will also help build and validate AI-powered knowledge bases.
Required Skills & Experience Core Technical Skills Strong hands-on experience in Databricks (notebooks, jobs, clusters, Delta Lake, SQL). Expert-level PySpark skills, including transformations, actions, window functions, and optimization techniques. Solid understanding of ETL/ELT concepts, data warehousing, and data modeling (star/snowflake schemas, dimension/fact tables). Proven experience in data validation, data quality checks, and reconciliation techniques. Practical experience with GitHub, GitHub Copilot, and GitHub Actions/Runners for CI/CD pipelines. AI & Prompt Engineering Experience crafting effective prompts for large language models (LLMs) for coding, testing, and documentation. Exposure to agent-based architectures or tools (e.g., workflow/agent frameworks that orchestrate multi-step AI tasks). Familiarity with AI-based tools for code analysis, test generation, or knowledge management. Knowledge Base & Documentation Experience building AI-driven knowledge bases or documentation systems (e.g., using vector search, embeddings, or LLM-based retrieval). Strong documentation skills, with the ability to translate complex data/QA concepts into clear knowledge articles. Testing & QA Practices
Strong understanding of QA methodologies: test planning, test design, defect management, and regression testing. Experience with automated testing frameworks (Python-based testing libraries such as pytest, unittest, or similar). Familiarity with data quality tools/practices (e.g., validation rules, thresholds, anomaly detection).
Key Attributes Strong analytical and problem-solving skills with a detail-oriented mindset. Passion for data quality, automation, and continuous improvement. Ability to work in an agile, fast-paced environment and collaborate across multiple teams like Business and operations team and AD team. Curiosity and openness to adopting new AI tools and practices for QA.