Responsibilities
- Architect, implement, and oversee observability frameworks using tools like Prometheus, Grafana LGTM+ stack, OpenTelemetry, and related technologies to monitor system performance and behavior.
- Partner with development, operations, and cross-functional teams to embed observability into CI/CD workflows and automate monitoring and alerting mechanisms.
- Build custom monitoring tools and instrumentation libraries to collect metrics, logs, traces, and events across microservices.
- Optimize the configuration of telemetry collection, storage, and visualization components for scalability, reliability, and cost efficiency.
- Deploy anomaly detection models and predictive analytics to identify and resolve issues before user impact occurs.
- Perform in-depth incident analysis and diagnose performance bottlenecks using observability data to support continuous system improvements.
- Monitor advancements in observability practices, tools, and standards, and evaluate their relevance and integration potential.
- Guide and mentor junior engineers, promoting a collaborative and knowledge-driven team culture.
- Participate in a 24/7 on-call rotation to respond to critical system incidents outside regular business hours.
- Lead automation efforts using Terraform, Git-based workflows, and DevOps tooling to enhance deployment speed and operational consistency.
Work Arrangement
Remote (Worldwide) — Malta, Budapest, Stockholm, Tallinn, Kyiv, Athens
Other
- Must participate in a 24/7 on-call schedule to assist the engineering team during off-hours incidents.
- Requires enthusiasm for using observability tools to achieve operational excellence in a fast-paced, collaborative setting.
- Demands strong technical expertise in AWS, Terraform, Git workflows, and a solid grasp of DevOps principles.