Responsibilities
- Ensure Availability & Reliability: Maintain highly available and resilient cloud infrastructure, meeting agreed SLO targets.
- Monitoring & Alerting: configure, and optimize monitoring solutions to detect anomalies early and maintain system health.
- Performance Optimization: Analyze and tune system performance, networking, and workloads to improve efficiency and reduce operational costs.
- Incident Response & Change Request: Respond to infrastructure incidents, perform root cause analysis, and implement permanent preventative solutions. Perform change request.
- Collaboration with DevOps & Development Teams: Partner with ACP Platform teams, contribute to product design, and provide operation feedback to ensure seamless service delivery and support operations.
- Disaster Recovery & Resilience Engineering: Lead and ensure backup, replication, and failover plans across AWS regions for business continuity are well maintain and tested.
- Postmortem & Continuous Improvement: Document incidents, update runbooks, and improve processes based on lessons learned.
- Automation: Build self-healing systems and automated remediation workflows to reduce manual intervention
- Capacity Planning: Forecast and optimize AWS resources to handle traffic spikes using Auto Scaling and Load Balancing.
- Security & Compliance: Follow security best practice and compliance with Accor security standards.
- Infrastructure-as-Code (IaC): maintain IaC using tools Terraform , CloudFormation , or AWS CDK to deploy and manage AWS resources.
Requirements
- Senior cloud engineer with software engineering and cloud operations expertise
- Experience with AWS services
- Skills in automation, monitoring, incident response, and performance optimization
- Ability to maintain highly available and resilient cloud infrastructure
- Experience with monitoring and alerting configuration and optimization
- Experience analyzing and tuning system performance, networking, and workloads
- Incident response experience including root cause analysis and implementing preventative solutions
- Experience performing change requests
- Collaboration skills to work with DevOps and development teams
- Experience with disaster recovery and resilience engineering including backup, replication, and failover plans across AWS regions
- Documentation skills for postmortems, runbooks, and process improvements
- Automation experience building self-healing systems and automated remediation workflows
- Capacity planning experience forecasting and optimizing AWS resources using Auto Scaling and Load Balancing
- Adherence to security best practices and compliance with security standards
- Experience with Infrastructure-as-Code (IaC) using Terraform, CloudFormation, or AWS CDK