Requirements
- Communicating with customers during incidents: Identifying impacts and scope of issues.
- Providing timely updates via status pages and communications channels (email, teams)
- Leading and contributing to Root Cause Analysis (RCA) documentation.
- Participating in on-call rotations and ensuring high availability of cloud services.
- Designing, building, and maintaining CI/CD pipelines using tools such as Azure DevOps.
- Automating build, test, and deployment workflows to improve release efficiency and reliability.
- Implementing pipeline governance, versioning strategies, and release approvals.
- Troubleshooting pipeline failures and optimizing performance and reliability.
- Integrating security scans, code quality checks, and automated testing into pipelines.
- Provisioning and managing cloud infrastructure using Infrastructure as Code (IaC) tools (e.g., ARM templates, Terraform, Bicep).
- Managing and optimizing containerized environments using Kubernetes.
- Ensuring scalability, reliability, and performance of cloud-native applications.
- Managing environment configurations across development, staging, and production.
- Taking part in development phases: Providing insights based on production environments, metrics, scalability, and constraints.
- Challenging system design and proposing improvements as early as possible.
- Supporting production readiness reviews and deployment validations.
- Testing and validating disaster recovery and business continuity procedures.
- Designing and implementing monitoring, alerting, and logging solutions (e.g., Azure Monitor, Log Analytics).
- Creating dashboards for system health, performance, and business metrics.
- Proactively identifying reliability risks and implementing preventive measures.
- Taking part in security reviews.
- Monitoring security tools.
- Identification of improvements.
- Implementing improvements.
- Taking part in continuous improvement: Identifying improvements, including monitoring or dashboarding.
- Taking part in lessons learnt from releases and incidents.
- Automating repetitive tasks through scripting (PowerShell, Bash, Python).
- Taking part in processes optimization, documentation, audits (e.g. ISO 27001)
- Optimizing cloud costs through rightsizing, reserved instances, and usage analysis.
- Contributing documentation, knowledge sharing, and process improvements.