Responsibilities
- Drive the transition from manual operations to fully automated infrastructure provisioning using Terraform as the primary tool.
- Create and sustain modular, reusable Terraform code for provisioning AWS services including virtual networks, compute instances, Kubernetes clusters, databases, and access controls.
- Develop and maintain Ansible automation scripts to manage configuration across mixed Windows and Linux server environments.
- Enforce a strict policy that all infrastructure changes must be made through code and tracked in version control systems.
- Design and implement GitOps-based deployment pipelines using GitHub Actions or ArgoCD for consistent infrastructure and application rollouts.
- Implement code governance practices such as protected branches, mandatory code reviews, and automated testing for infrastructure pull requests.
- Define and promote GitOps best practices across engineering teams, including onboarding and training for new workflows.
- Take end-to-end ownership of cloud infrastructure performance, ensuring high availability, reliability, and scalability across AWS environments.
- Establish service level objectives and indicators; implement monitoring, alerting, dashboards, and incident response procedures.
- Lead incident response efforts and post-incident reviews to identify root causes and implement preventive solutions.
- Develop and apply cost-saving measures for cloud usage, including resource tagging and governance enforcement.
- Coach and develop members of the operations team in infrastructure-as-code and automation techniques.
- Document architectural patterns, standards, and key decisions to ensure organizational knowledge sharing.
- Amplify team effectiveness—success is measured not only by personal output but by the team’s increased capability.