Responsibilities
- Enhance current systems by integrating observability tools for improved system insights
- Develop and manage infrastructure through code using platforms like Terraform, Pulumi, or CloudFormation
- Design and oversee Kubernetes environments with full monitoring and logging capabilities
- Create continuous integration and delivery workflows with built-in observability and automated test execution
- Define and manage key performance metrics including Service Level Indicators, Objectives, and Agreements
- Apply error budgeting, reduce operational toil, and plan resource capacity effectively
- Assist in managing incident response and lead post-incident reviews to drive improvements
- Operate observability systems across major cloud providers including AWS, GCP, and Azure
- Implement policies for security, regulatory compliance, and governance of telemetry information
- Automate evaluation of monitoring agents within CI/CD systems and backend observability platforms
- Develop and deploy OpenTelemetry solutions spanning multiple technology environments
- Set up and manage OpenTelemetry Collectors to gather, process, sample, and route telemetry data efficiently
- Build telemetry pipelines to handle metrics, traces, and logs across microservices-based systems
- Tune collector settings to balance performance, dependability, and operational cost
Work Arrangement
Hybrid — London
Other
You will be expected to work for up to four days a week in person, be it from our office in London or from client sites.