Responsibilities
- Design and implement reliability frameworks and self-service tools that empower teams to manage service reliability within a 'You Build It, You Run It' model.
- Lead the development and execution of AI-powered operations strategies, including automated diagnostics, issue resolution, and predictive failure prevention.
- Promote a culture of reliability by integrating SRE principles into engineering workflows through design reviews, production readiness assessments, and operational benchmarks.
- Take command during critical system incidents, demonstrating high standards of operational conduct and ensuring thorough, blame-free post-incident analysis.
- Develop and maintain comprehensive observability solutions—monitoring, distributed tracing, and profiling—using tools like Prometheus, Grafana, OpenTelemetry, and continuous profiling systems.
- Support the growth of engineers across SRE and product teams by providing mentorship, technical leadership, and collaborative knowledge transfer.
Work Arrangement
Remote (Worldwide) — U.S., U.K., Finland, India, Singapore, Canada, Ireland