Responsibilities
- Define and own chaos engineering strategy and roadmap
- Lead design of advanced chaos experiments (multi-layer: infra, network, application)
- Mentor engineers in resilience and chaos practices
- Standardize frameworks for automated chaos testing
- Drive adoption of resilience practices across engineering teams
- Collaborate with architecture teams to embed resilience into system design
- Define SLIs, SLOs, and error budgets aligned with reliability goals
- Lead game days/failure simulations
- Drive observability maturity across systems
Requirements
- 7+ years in SRE/DevOps/Platform Engineering
- Strong expertise in Chaos Engineering (must-have)
- Proven experience in designing custom chaos platforms or frameworks
- Strong programming background (Python / Go preferred)
- Deep experience with: On-prem systems (primary)
- Distributed systems architecture
- Networking and failure injection
- Observability leadership experience: Metrics, tracing, logging strategies
- Experience with hybrid environments (on-prem + cloud integration)
Nice to Have
- Experience with large-scale enterprise systems
- Exposure to compliance, DR, and BCP strategies
- Kubernetes chaos testing experience