Responsibilities
- Lead end-to-end ownership of a core platform domain, such as observability or CI/CD, including architecture, reliability, and long-term planning.
- Manage technical projects from concept to production, including defining requirements, creating design documentation, task breakdown, implementation, deployment, and ongoing operational support.
- Clarify ambiguous situations by establishing clear requirements, assumptions, and actionable next steps.
- Design systems for high reliability and scalability, improving platform topology, integrations, scaling strategies, and fault tolerance.
- Assist development teams with deploying and monitoring applications on-premise and in Kubernetes using Helm, troubleshooting builds, and supporting metrics, logging, and alerting.
- Eliminate manual operational work by automating repetitive tasks, provisioning, and system maintenance through code.
- Serve as the primary escalation point for production incidents in your domain, leading incident resolution, post-mortem analysis, and implementing preventive fixes; participate in on-call rotations to improve response practices.
- Guide junior engineers through design reviews, code feedback, and collaborative sessions, identifying and preventing technical debt early.
- Integrate AI tools across daily workflows, including research, debugging, and development tasks.
Requirements
- Minimum of six years in a DevOps or SRE role with similar scope and responsibilities.
- Proven experience leading technical initiatives from requirements through to production, with demonstrable ownership beyond task execution.
- Strong Linux proficiency, particularly on Ubuntu systems.
- Familiarity with the Prometheus ecosystem, including metric types, exporters, and alerting mechanisms, sufficient to maintain and expand existing setups.
- Hands-on experience designing and managing CI/CD pipelines, build orchestration, and artifact distribution.
- Practical knowledge of containerization using Docker, including image creation and registry management.
- Experience with Ansible for configuration management and automation.
- Proficiency with Git for version control.
- Scripting ability in Bash or Python to automate operations and build observability tools such as custom exporters.
- Production support background, including incident diagnosis, service restoration, and post-mortem leadership.
- Experience mentoring less experienced engineers in technical design and implementation.
- Strong sense of ownership and precision; understanding that system downtime during peak periods can result in significant financial impact.
Nice to Have
- Experience with VictoriaMetrics or Prometheus at scale, including architectural design, cardinality management, exporter integration, and alerting infrastructure.
- Large-scale log pipeline expertise using Graylog, VictoriaLogs, or ELK, covering collection, retention, sharding, and performance tuning.
- Advanced Jenkins scripted pipeline knowledge, including shared libraries, pipeline frameworks, and managing build agent fleets.
- Familiarity with container registries and artifact management tools such as Harbor and Nexus, including base image policies and lifecycle controls.
- Production experience running applications on Kubernetes, including Helm usage, workload monitoring, log routing, and deployment troubleshooting.
- Expertise in Grafana, including managing dashboards as code, configuring alerts, and optimizing performance at scale.
Benefits
- 31 days of annual leave
- Fully paid telemedicine coverage
- Financial support for home office setup, including ergonomic furniture and equipment
- Access to English language training programs
- Funding for professional development and certifications
- Subsidy for gym or swimming pool memberships
- Co-working space allowance
- Fully remote work policy
Work Arrangement
Remote (Worldwide) — Turkey Remote
Team
Russian-speaking team
Other
- Language: Russian-speaking team
- Location: Turkey Remote