We’re looking for a Site Reliability Engineer to join our DevOps team in Tallinn and take ownership of keeping the Reconeyez platform healthy and available. You’ll monitor our infrastructure, respond to incidents during on-call shifts, diagnose issues across the stack, and continuously improve our operational posture, while building and automating the tools and processes that keep our systems running reliably.
This is a hands-on role. You’ll spend your time in dashboards, terminals, and log files. When something breaks, you’re the one who finds out why and makes sure it doesn’t happen again.
What You’ll Do
Platform Reliability & Incident Response
- Keep production running, monitor system health, respond to alerts, and resolve incidents before they impact customers
- Participate in on-call rotation with the DevOps team, taking responsibility for incident response and resolution during your shifts
- Document runbooks and incident postmortems so the team learns from every outage
- Collaborate with development teams to improve reliability, flag recurring issues, and advocate for operational improvements
Infrastructure & Automation
- Build and set up new development tools and infrastructure; deploy updates and fixes
- Work on ways to automate and improve development and release processes using Git-based workflows and PR-based operations
- Manage containerized services running on Docker/Podman, deployments, restarts, resource management, and health checks
- Configure and maintain network services including firewalls, load balancers, and VPNs
- Contribute to infrastructure-as-code practices to make infrastructure changes auditable and repeatable
Observability & Monitoring
- Manage and improve monitoring using Zabbix, Grafana, Prometheus, and Alertmanager, build dashboards, tune alerts, reduce noise
- Adopt and extend OpenTelemetry instrumentation across services for unified tracing, metrics, and logging
- Analyze logs to identify root causes, spot patterns, and catch problems early
- Monitor AI/ML inference endpoints and model-serving infrastructure, track latency, throughput, and model health alongside traditional service metrics
- Monitor and flag infrastructure cost anomalies to support cloud spend awareness across the team
Databases & Security
- Install, monitor, and maintain PostgreSQL, backups, recovery, performance tuning, and query troubleshooting
- Ensure systems are safe and secure against cybersecurity threats, including container image scanning and supply chain security practices
- Ensure systems are safe and secure against cybersecurity threats
Platform Engineering
- Reduce cognitive load for development teams through tooling, automation, and self-service capabilities
- Build internal tools and processes that help developers move faster without sacrificing reliability
Apply on company website Tallinn, Estonia On-site Full-time