Responsibilities
- Build and operate metrics, logs, traces, and alerting capabilities to help teams establish meaningful SLOs and use production telemetry for problem diagnosis and improvement verification.
- Own incident tooling and practices, coordinate cross-team response, and turn incident reviews into engineering improvements that reduce recovery time and prevent repeat failures.
- Build and maintain load and failure testing capabilities to validate critical paths under expected demand, quantify headroom, and test recovery and production readiness with service owners.
- Lead deep engagements with internal teams on SLOs and end-to-end performance, using profiling, telemetry, and load tests to identify bottlenecks and deliver improvements with service owners.
- Stay technically engaged by reviewing designs and production changes, debugging difficult failure modes, and using AI coding tools to prototype and automate, applying rigorous review and verification to AI-generated changes.
- Build and grow a high-ownership engineering team by coaching engineers, developing technical leaders, managing performance, and hiring against agreed needs, making distributed collaboration, mentoring, and backup coverage deliberate.
- Track rollout safety, recovery time, repeat incidents, critical-path latency and throughput, test coverage, and improvements from cost and capacity analysis, agreeing on success measures and continuing ownership with partner teams.
Requirements
- Demonstrated engineering management experience leading and developing engineers, making prioritization and performance decisions, hiring thoughtfully, and delivering through a team.
- Software-oriented production systems depth with experience building and operating distributed systems or reliability platforms, reasoning across deployment behavior, Kubernetes, telemetry, service dependencies, and recovery mechanisms.
- Safe-change and performance judgment from leading consequential migrations or incidents, using measurement to diagnose reliability or performance problems, distinguishing symptoms from causes, and validating fixes under realistic conditions.
- Platform-product and cross-team judgment to build capabilities other teams adopt, lead hands-on engagements without absorbing every service's operations, and make clear tradeoffs among reliability, performance, engineering effort, and cost.
Nice to Have
- Experience with GitOps or progressive-delivery platforms such as Harness, ArgoCD, or Kargo.
- Experience with observability, profiling, load-testing, and failure-testing systems, including OpenTelemetry or comparable tooling.
- Experience with cloud cost attribution, capacity planning, and provider coordination, particularly on GCP.
- Experience growing distributed teams and using AI tools to increase engineering output while preserving production safeguards.
Benefits
- Competitive Salary & Equity
- 401(k) Program with a 4% match (US Only)
- Health, Dental, Vision and Life Insurance
- Short Term and Long Term Disability
- Paid Parental, Medical, Caregiver Leave
- Flexible Time Off (FTO) + Holidays
- Commuter Benefits (In-Office & US Only)
- Monthly Wellness Stipend
- Autonomous Work Environment
- In Office Set-Up Reimbursement (In-Office Only)
- Quarterly Team Gatherings
- In Office Amenities (In-Office Only)
Compensation
Competitive Salary & Equity
Work Arrangement
Autonomous Work Environment