Responsibilities
- Manage and enhance the centralized telemetry platform supporting metrics, logs, traces, alerts, dashboards, and profiling tools.
- Ensure reliable collection, long-term retention, efficient querying, visualization, and alerting of metrics using Prometheus-style systems, VictoriaMetrics, Grafana, and modern alerting solutions.
- Support log ingestion and processing pipelines using Vector, Splunk, and Loki, focusing on stability, performance, and issue resolution.
- Maintain distributed tracing and profiling infrastructure with Grafana Alloy, Tempo, OpenTelemetry, and Pyroscope for observability.
- Deploy and operate telemetry components using Infrastructure-as-Code tools like Terraform and Terragrunt, alongside container orchestration platforms across multiple environments.
- Diagnose and resolve issues such as data gaps, delayed queries, alert failures, pipeline congestion, and resource constraints.
- Develop standardized configurations and automation to help teams safely implement dashboards, alerts, and telemetry integrations.
- Take part in incident response rotations, document procedures in runbooks, and apply post-incident insights to strengthen system resilience.
Other
- Applications are reviewed continuously unless a deadline is specified.
- Applicants may omit personal details such as age, birth date, or educational dates.
- Individuals with criminal records are considered for roles in compliance with the San Francisco Fair Chance Ordinance.
- The company provides equal employment opportunities and prohibits discrimination or harassment based on race, ethnicity, age, gender identity, citizenship, religion, sexual orientation, disability, pregnancy, veteran status, or other protected traits.
- Candidates might be required to complete job-related skills or behavioral assessments during the hiring process.
- Assessments measure role-relevant abilities and are evaluated together with professional background and interview performance.