
Chaos Test Lead
Encora • Ciudad de México, Mexico
**Role & seniority: ** Chaos Test Lead (5–10+ years relevant experience)
**Stack/tools: **
-
Chaos/Resilience: Gremlin, Chaos Native, Litmus (any)
-
Platforms: AWS, GCP, Azure
-
OS/infra: Unix/Linux, networking protocols, storage/disk internals
-
Cloud components: VPCs, proxies, load balancers, availability zones
-
Observability: monitoring, alerting, logging
-
CI/CD: troubleshooting pipeline failures; automated/continuous experiments
-
Top 3 responsibilities:
-
Lead end-to-end chaos engineering lifecycle: planning, experiment design, execution, and reporting
-
Schedule and ensure recovery/resilience testing is staffed, executed, documented with remediation and closure
-
Analyze system architecture to identify weak points likely to cause outages/failure and coordinate requirements with business/tech teams
-
-
Must-have skills:
-
Practical experience in chaos engineering, resilience, high availability testing
-
Ability to work with enterprise architecture and development teams for HA/resiliency architecture
-
Design/develop/execute automated or continuous chaos experiments
-
Troubleshoot failures in CI/CD pipelines
-
Strong Linux/Unix and distributed systems debugging skills
-
Strong communication and analytical/decision-making skills
-
-
Nice-to-haves:
-
Experience across multiple public clouds (AWS/GCP/Azure)
-
Deep expertise with VPC/proxy/LB/AZ configurations; advanced observa
-
Full Description
Job Title: Chaos Test Lead
Key Skills: Chaos Test Planning, Chaos Test Designing, and Reporting
Experience: 5+ years Location Mexico City, Mexico
Mode: Hybrid We at Coforge are hiring Position (22660) with the following skill set. Relevant experience on Chaos engineering / Resilience / High availability testing atleast 5-10 years of relevant experience is must Implement and lead execution of the chaos engineering Lifecycle - Chaos Test Planning, Chaos Test Designing, and Reporting Ensure recovery and resilience testing is scheduled, staffed, executed, and documented, including remediation and closure of issues Ability to analyse the architecture & recommend weak areas that are likely to failure / outages Ability to work with Business & technology teams to identify and report on resilience / High availability requirements Ability to work with enterprise architecture and development teams to architect applications for high availability and resiliency Design, develop and execute automated / continuous Chaos Engineering experiments, Ability to troubleshoot the failures in CI/CD pipeline Automate Chaos experiments through chaos engineering tools (Gremlin / Chaos Native / Litmus etc) to run continuously Hands on experience in Unix/Linux OS environments and operating system internals, file systems, disk/storage and networking protocols. Strong knowledge on Public cloud platforms – AWS, GCP, Azure Knowledge on Monitoring, Alerting, Logging Knowledge on VPC’s, proxy’s, load balancers, availability zones Ability in diagnosing and debugging complex distributed systems Tools (any of these) - Gremlin, Chaos Native, Litmus Strong leadership skills and ability to work in a cross-functional environment Strong interpersonal, oral, and written communication skills Strong analytical and decision-making skills
Posted On: July 31, 2026 At Coforge, we hire professionals based solely on their skills and do not discriminate based on age, disability, religion, gender, sexual orientation, socioeconomic status, or nationality.