Responsibilities
- Lead the long-term technical vision and strategic planning for operational maturity across all stages of service lifecycle.
- Create and enforce a structured process for introducing new services, ensuring compliance with architecture, security, resilience, observability, and support requirements prior to production deployment.
- Set enterprise-wide standards for service ownership, including documentation of service records, responsible parties, dependencies, runbooks, support models, escalation procedures, recovery goals, and on-call preparedness.
- Drive the design and ongoing enhancement of a unified observability platform, defining consistent practices for logs, metrics, distributed tracing, and profiling across systems.
- Establish policies for dashboarding, alerting, synthetic monitoring, telemetry accuracy, data retention, sampling strategies, cardinality control, and cost management.
- Manage the reliability governance framework for live services, including service-level indicators, objectives, error budgeting, and incident escalation protocols.
- Correlate system health metrics with customer and business outcomes to detect anomalies early, prevent service degradation, and quickly identify user-impacting issues.
- Improve incident response practices through standardized severity grading, incident leadership, stakeholder communication, automated evidence gathering, and coordinated handling of critical events.
- Implement blameless post-mortem processes, ensure corrective actions are completed, and lead initiatives to eliminate recurring failure patterns.
- Oversee capacity and efficiency planning, including demand modeling, cloud and Kubernetes resource scaling, performance validation, headroom policies, resource optimization, and risk assessments.
- Design and manage AI-driven operations capabilities for event correlation, noise reduction in alerts, predictive issue detection, root cause identification, automated triage, guided remediation, and controlled self-healing.
- Develop secure AI-powered workflows integrated across observability tools, service registry, issue tracking, documentation, source repositories, and deployment pipelines.
- Enhance on-call performance through well-structured rotations, readiness criteria, escalation rules, diagnostic automation, alert quality oversight, and dependable global handover processes.
- Provide direct technical leadership during major outages, complex reliability analyses, system design evaluations, resilience testing, and high-stakes service rollouts.
- Leverage operational metrics, incident trends, service performance data, capacity signals, change success rates, and automation impact to guide continuous improvement initiatives.
Work Arrangement
On-site — Santa Clara, CA
Team
The team develops, secures, and maintains scalable infrastructure for cloud-hosted SaaS offerings that include on-premises components deployed at customer locations.
Team
The team builds, secures, and operates scalable infrastructure for cloud-managed SaaS products with on-premises components deployed at customer sites.