Responsibilities
- Lead the planning and implementation of initiatives that strengthen system resilience across the service ecosystem in collaboration with engineering groups.
- Collaborate with high-priority services to assess system design, enhance performance at scale, and minimize manual operational work.
- Develop and refine secure platforms for handling service configurations and sensitive data across large-scale environments.
- Enhance canary deployment frameworks and expand release infrastructure to safely manage thousands of services and hundreds of daily updates with reduced failure rates.
- Promote and institutionalize reliability standards and foster a strong culture of operational excellence throughout engineering teams.
Work Arrangement
Hybrid
Other
- The company operates with a remote-first policy but allows some in-office presence.
- Team members are expected to meet in person each quarter for focused, collaborative work periods known as 'surges'.