
SDET + Python + Playwright + Document Processing (Xbrl)
Improving South America • Costa Rica
**Role & seniority: ** Lead/own engineer for production readiness of AI-assisted financial/regulatory tagging; design/implement benchmarking & validation framework (seniority implied as lead/owner).
**Stack/tools: ** LLM evaluation; XBRL/IXBRL, XHTML, XML, SEC validation; Playwright for web/HTML-level test automation (no existing automation); CI/CD/DevOps; golden datasets, custom validation frameworks; scoring/quality metrics dashboards (nice-to-have).
**Top 3 responsibilities: **
-
Design and implement a benchmarking framework measuring LLM output quality across technical correctness, SEC validation compliance, and completeness of tagged facts.
-
Build automated test harnesses (starting with Playwright-based web/HTML checks) and define quality gates with repeatable scoring.
-
Track and communicate progress indicators over time; collaborate with R&D (AI tagging core) and production engineering to close gaps and improve regulatory/business reliability.
-
Must-have skills:
-
Hands-on experience building test/validation frameworks for complex outputs; automated correctness logic and golden datasets.
-
Proven ability to design benchmarking strategies, metrics, and evaluation strategies for LLM outputs.
-
Strong understanding of XBRL/IXBRL/XHTML/XML or financial regulatory reporting; familiarity with SEC validation requirements.
-
Playwright proficiency for test automation; CI/CD experience to run benchmarks consistently.
-
Full Description
In this role, we are focused on production readiness for AI-assisted financial/regulatory tagging. You will design and implement a comprehensive benchmarking framework that measures Large Language Model (LLM) put quality across multiple dimensions, then establish the test automation and scoring mechanisms needed to track improvements over time. You will work closely with our R&D team on the AI tagging core and with production engineering to identify gaps, iterate on quality gates, and ensure the system can reliably meet regulatory and business requirements.
Responsibilities
We will have you lead the design and implementation of a benchmarking and validation framework for LLM outputs, with clear, measurable quality dimensions and automated checks.
Design and implement a comprehensive benchmarking framework that measures LLM output quality across multiple dimensions, including
- Technical correctness, such as valid XHTML and IXBRL schema compliance.
- SEC validation compliance, ensuring generated XBRL instance documents pass SEC validation without errors.
- Completeness, verifying all required facts are tagged (including hidden or non-obvious data points).
- Create complexity “tranches” for benchmarking, ranging from single fund / single share class through multi-fund scenarios and 200+ share class cases, and define success metrics for each tier.
- Build test automation for web and HTML-level testing using Playwright (with no existing automation today), including the first stable automated test harness.
- Create measurable progress indicators that track improvements as the AI tagging core is enhanced and refined.
- Validate and score LLM outputs against both regulatory and business requirements, producing repeatable results suitable for ongoing quality evaluation.
- Collaborate with the R&D team (focused on AI tagging) and production engineering to identify gaps, prioritize iteration steps, and support production readiness.
Required Skills & Experience
We are looking for a strong technical engineer who can design evaluation strategies, implement validation frameworks, and turn regulatory requirements into automated, measurable quality gates.
Strong technical background validating and testing complex systems, with hands-on experience building test harnesses. Experience building automated test frameworks, including creation of golden datasets, custom validation frameworks, and automated correctness/validation logic. Proven experience designing and implementing benchmarking strategies and quality metrics for complex outputs. Familiarity with LLM evaluation, validation techniques, and quality metrics (including how to interpret results and drive iteration). Proficiency with test automation tools, with Playwright preferred for web/HTML-level validation. Solid understanding of XBRL, XHTML, XML, or financial regulatory reporting (strongly preferred). Strong problem-solving skills and the ability to work with ambiguity and incomplete requirements while still converging on reliable quality criteria. Excellent communication and collaboration skills across R&D and production engineering stakeholders. Experience with financial or compliance testing and validation. Familiarity with SEC filing requirements and investment company regulations. Background in data quality validation and metrics design. Experience with CI/CD pipelines and DevOps practices to run benchmarks and validations consistently.
Soft skills we value: we expect clear written communication, structured thinking, ownership of quality outcomes, and an iterative mindset that balances compliance precision with practical engineering constraints.
Desirable (nice to have)
Experience with scoring models and evaluation dashboards for benchmarking results. Previous work on SEC submission tooling, XBRL instance document generation, or schema/validation tooling at scale. Experience with financial or compliance testing and validation Familiarity with SEC filing requirements and investment company regulations Background in data quality validation and metrics design Experience with CI/CD pipelines and DevOps practices
Benefits
Contrato a largo plazo. 100% Remoto. Vacaciones y PTOs Posibilidad de recibir 2 bonos al año. 2 revisiones salariales al año. Clases de inglés. Equipamiento Apple. Plataforma de cursos en linea Budget para compra de libros. Budget para compra de materiales de trabajo mucho mas..