Responsibilities
- Quickly learn the entire platform landscape, including all workloads, interdependencies, and potential risks, primarily through code, documentation, and team discussions.
- Collaborate with subject matter experts to close knowledge gaps and develop onboarding resources for the team.
- Create and update runbooks, architectural documentation, and operational procedures.
- Design highly available and fault-tolerant infrastructure on Azure, including Azure Government environments.
- Establish service level indicators, objectives, and error budgets where they are currently undefined.
- Lead incident response efforts and conduct blameless post-incident reviews.
- Convert incident learnings into system improvements and preventive measures.
- Assess reliability risks in both modern and legacy systems and implement feasible fixes within compliance boundaries.
- Improve observability by defining monitoring requirements and driving their implementation.
- Collaborate with partner teams to set standards for alerting, telemetry, and monitoring.
- Develop automation to reduce manual effort and support large-scale system management.
- Take part in on-call duties for incident response and system support.
- Work with infrastructure-as-code, CI/CD pipelines, deployment automation, and configuration management, even in isolated or compliance-heavy environments.
- Design and maintain testing frameworks, canary deployments, and release validation processes.
- Integrate chaos engineering and monitoring tools while ensuring adherence to regulatory standards.
- Collaborate across product, platform, security, legal, compliance, and operations functions.
- Take full ownership of issues from identification through resolution without waiting for direction.
- Guide junior engineers and promote SRE principles throughout the organization.
Benefits
- Shape the foundation of the government cloud reliability program from the start.
- Influence the evolution of SRE practices across a global engineering organization.
- Collaborate with skilled teams in product, cloud engineering, security, and compliance.
- Access professional growth tools such as mentorship, training, and community engagement opportunities.
- Receive competitive pay and comprehensive benefits.
Compensation
Competitive compensation and benefits
Work Arrangement
Remote (Worldwide)
Team
Cross-functional collaboration with engineering, security, compliance, and operations teams in a globally distributed environment.
Other
- Access to government infrastructure is restricted due to clearance and security requirements.
- Participation in on-call rotations is required.
- Ability to collaborate across engineering, product, security, compliance, and operations teams is essential.
- Expected to mentor engineers and promote SRE practices organization-wide.