Responsibilities
- Work closely with engineering teams to build scalable, secure, and highly available systems for the platform.
- Define and oversee service level objectives and service level agreements for the cloud product.
- Ensure monitoring and alerting are implemented across all infrastructure components, including the data plane, control plane, and core database services.
- Improve incident response procedures and lead post-mortem reviews following outages, coordinating with support to inform affected customers.
- Drive ongoing improvements in system reliability and performance for cloud-hosted database services.
- Initiate and support chaos engineering efforts across engineering groups based on internal priorities.
- Oversee on-call operations for performance and reliability issues, and develop protocols for escalation and incident resolution.
Work Arrangement
Remote (Worldwide)