Responsibilities
- Manage, mentor, and grow a global team of SRE engineers across Israel and Ukraine, running regular 1:1s, setting goals, and supporting career development.
- Own hiring, onboarding, and performance management for the team.
- Keep the team working as one unit across sites, with shared standards, consistent handoffs, and clear ownership so that reliability work is not fragmented by location or timezone.
- Set the technical roadmap for the team's reliability, automation, and observability initiatives, and stay hands on enough to guide design decisions and unblock complex problems.
- Guide the design and implementation of solutions that improve the reliability, availability, and scalability of our production platform, and ensure the team proactively identifies and eliminates operational risks.
- Prioritize and oversee the build of internal tools and automation that eliminate manual operational work, improve engineering productivity, and streamline production workflows.
- Drive the team's roadmap for monitoring, alerting, dashboards, and production visibility, reducing alert fatigue and strengthening operational insight across services.
- Oversee the team's work on deployment processes using modern release strategies such as Canary, Blue/Green, and Feature Flags, ensuring safe and reliable releases.
- Act as an escalation point for critical production incidents, guide root cause analysis, and ensure long term preventive improvements are implemented and tracked.
- Own the team's on-call rotation and coverage across sites, participate as needed, and drive continuous improvements that reduce operational toil and prevent future incidents.
- Oversee the team's work on our Kubernetes based cloud platform, CI/CD pipelines, and production infrastructure running on GCP and AWS.
- Represent the SRE team in planning and decision making with Software Engineering, DevOps, DBA, and Product leadership, and align the team's priorities with broader engineering goals.
Nice to Have
- Prior formal people management or team lead experience.
- Experience leading or coordinating engineers who are not co-located.
- Experience with Infrastructure as Code (Terraform, Ansible, etc.).
- Experience with messaging and distributed technologies such as Kafka, Pub/Sub, or Redis.
- Experience supporting modern deployment strategies such as Canary, Blue/Green, or Feature Flags.
- Familiarity with OpenTelemetry and modern observability tooling.
- Understanding of Site Reliability Engineering principles, including SLIs, SLOs, and error budgets.
- Experience working in large scale SaaS production environments.
- Relevant cloud or Kubernetes certifications (GCP, AWS, CKA, CKAD).
Work Arrangement
Remote (Worldwide) — Israel, Ukraine