Responsibilities
- Architect, build, and manage resilient, scalable infrastructure systems that power key product features and internal service needs.
- Lead the creation and ongoing enhancement of internal platforms, shared tools, and services that streamline development, testing, and deployment at scale.
- Take full ownership of infrastructure components from initial requirements and design through implementation, deployment, live support, and iterative optimization.
- Collaborate across engineering groups to enable complex, high-priority workloads, platform features, and integrations across large systems.
- Detect and resolve systemic issues related to reliability, performance, scalability, and operations using automation, observability, and proven engineering methods.
- Offer technical leadership during on-call duties, incident response, root cause investigations, and corrective actions to maintain system uptime and robustness.
- Develop and oversee monitoring practices, capacity forecasts, and scaling plans for live systems and infrastructure layers.
- Ensure infrastructure adheres to and improves upon security, compliance, and operational benchmarks in production settings.
- Coach engineering team members, shape technical decisions, and evaluate designs to elevate system quality and long-term maintainability.
- Produce and update comprehensive technical documentation, design specifications, and operational runbooks.
Work Arrangement
Hybrid — Menlo Park, CA
Other
- Background checks required
- Telecommuting permitted