Responsibilities
- Create and deploy secure multi-account architectures and landing zones within AWS.
- Oversee and enhance multi-cloud identity and access management, single sign-on, and role-based access controls to enforce minimal privilege access.
- Implement tagging policies, resource organization, and cost-saving measures like rightsizing and removing unused resources to ensure financial responsibility.
- Lead the setup, scaling, and administration of Kubernetes clusters, including management of network interfaces, ingress controllers, and service meshes.
- Administer and optimize Linux and Windows Server systems, ensuring secure setups and automated updates.
- Handle integration between cloud services and operating system dependencies, such as Active Directory and file system performance adjustments.
- Develop and upkeep modular infrastructure templates using tools like Terraform, CloudFormation, or Pulumi.
- Construct and support CI/CD pipelines with Jenkins and Bitbucket, covering stages from builds to deployments, and maintain GitOps workflows.
- Design security measures including encryption, network controls, and logging to meet compliance standards like SOC2, HIPAA, or FedRAMP.
- Utilize AI-driven tools and large language models to speed up infrastructure code writing, automate diagnostics, and predictively optimize cloud usage.
- Install, configure, and manage shared engineering tools, handling authentication, backups, and integrations, and secure infrastructure access.
- Assist developers in onboarding applications, managing jobs, using pipeline templates, and resolving issues across build, test, and deployment phases.
- Provision and operate AWS resources such as instances, clusters, networking, and databases using infrastructure as code.
- Design and maintain access models including single sign-on, IAM, Kubernetes roles, and secrets with least-privilege and audit capabilities.
- Define operational responsibilities for the engineering platform, covering availability, maintenance, backups, disaster recovery, and documentation.
- Establish monitoring and observability for infrastructure, clusters, CI/CD, tools, and applications, progressing from basic metrics to advanced analysis.
- Track service health indicators like uptime, resource usage, errors, and security events across various components.
- Collaborate with development and security teams to set standards, policies, controls, alerts, and documentation for platform usage.