
AI Tester at Denvor CO _Remote
VOLTO Consulting • Colorado, United States
Salary: $45 - $50 / year
**Role & seniority: ** AI Tester / Quality Engineering (mid-senior; 4–10 years experience) for AI & Generative AI quality.
**Location & work type: ** Denver, CO (80221) — Remote; Contract.
**Stack/tools: **
-
Testing/QA practices; regression + defect lifecycle
-
Likely use of API testing and test automation/scripting (preferred)
-
Works with structured/unstructured data and evaluates model metrics (e.g., accuracy/precision/recall)
-
Tooling not specified beyond automation/API testing familiarity.
-
Top 3 responsibilities:
-
Design/execute AI test scenarios for functional + nonfunctional + responsible-use expectations, including edge/ambiguous inputs.
-
Generative AI validation: assess outputs for accuracy/relevance, consistency/stability, safety/appropriateness; detect hallucinations/undesired responses; validate prompt/context/variability.
-
Regression & risk monitoring: retest after updates; identify model/output drift; document defects/risks with evidence and recommendations.
-
-
Must-have skills:
-
Strong software testing/QA fundamentals (test case/scenario design, regression testing, defect management).
-
Ability to reason about probabilistic/nondeterministic AI outputs and define acceptable vs unacceptable behavior.
-
Analytical capability to evaluate AI output patterns/anomalies using structured/unstructured data.
-
-
Nice-to-haves:
- Test automation experience (scripting) an
Full Description
AI Tester AI & Generative AI Quality Engineering
Denver, CO (Zip Code-80221)- REMOTE
Contract Role
Role Summary
We are seeking a detail oriented AI Tester to ensure the quality, reliability, safety, and correctness of AIenabled applications, including systems that use machine learning and generative AI. This role focuses on validating AI behavior, testing probabilistic outputs, and ensuring AI solutions meet functional, nonfunctional, and responsibleuse expectations before and after production release.
The AI Tester works closely with developers, architects, and product teams to identify risks, validate outcomes, and continuously improve AI system quality.
Key Responsibilities
AI & Functional Testing
Design and execute test scenarios for AIdriven application features.
Validate AI outputs against functional expectations, business rules, and acceptance criteria.
Test AI behavior across a wide range of inputs, including edge cases and ambiguous scenarios.
Generative AI Validation
Evaluate Responses Generated By AI Systems For
Accuracy and relevance
Consistency and stability
Safety and appropriateness
Identify hallucinations, incorrect reasoning, and undesired outputs.
Validate prompt behavior, context handling, and response variability.
Data & Model Testing
Verify quality and coverage of test data used for AI validation.
Test AI systems using structured and unstructured data inputs.
Support validation of model performance metrics such as accuracy, precision, recall, or similar indicators.
Regression, Monitoring & Drift
Perform regression testing to ensure AI behavior remains acceptable after updates.
Support monitoring of AI behavior in productionlike environments.
Identify model or output drift and raise risks when behavior changes beyond acceptable thresholds.
Responsible AI & Risk Validation
Validate AI systems for fairness, explainability, and bias risks based on defined guidelines.
Ensure outputs comply with security, privacy, and ethical usage expectations.
Support humanintheloop review processes where required.
Collaboration & Reporting
Collaborate with AI developers, data teams, and product owners during delivery cycles.
Clearly document defects, risks, and test findings with evidence and recommendations.
Contribute to continuous improvement of AI testing strategies and practices.
Required Technical Skills
Testing & Quality Engineering
Strong foundation in software testing and quality assurance practices.
Experience designing test cases, test scenarios, and validation criteria.
Understanding of regression testing and defect lifecycle management.
AI & Machine Learning Understanding
Working knowledge of AI and machinelearning concepts, including probabilistic behavior.
Understanding of generative AI characteristics such as nondeterministic outputs and context sensitivity.
Ability To Reason About Acceptable Vs. Unacceptable AI Responses.
Data & Analysis
Ability to work with structured and unstructured test data.
Strong analytical skills to evaluate patterns, anomalies, and inconsistencies in AI outputs.
Automation & Tooling (Preferred)
Exposure to test automation concepts and scripting.
Familiarity with API testing and basic data validation techniques.
Experience & Qualifications
4 10 years of experience in testing, quality engineering, or validation roles.
Exposure to AIenabled or datadriven systems is strongly preferred.
Bachelor s or Master s degree in Computer Science, Engineering, Data Science, or related field.