Responsibilities
- Design, deploy, and manage high-scale production environments with an emphasis on reliability, performance, security, and observability.
- Enhance monitoring systems and improve support mechanisms to boost developer experience across platform services.
- Develop automation tools and scripts to eliminate repetitive operational tasks and increase system efficiency.
- Establish and refine alerting systems with proactive monitoring and automated response workflows.
- Develop and keep updated documented procedures for incident response and operational runbooks.
- Investigate infrastructure and platform problems, assist software engineers in root cause analysis, and coordinate with external vendors when needed.
- Produce detailed post-incident reports and implement corrective measures to prevent recurrence.
- Schedule and communicate planned maintenance activities for live systems with minimal disruption.
- Collaborate with development teams to uncover infrastructure-related workflow constraints and deliver scalable solutions.
- Evaluate and integrate industry-standard practices to ensure systems are secure, resilient, and highly available.
- Analyze open-source software architecture and implementation to improve debugging and resolution speed for platform issues.
Responsibilities (11)
- Design, deploy, and manage high-scale production environments with an emphasis on reliability, performance, security, and observability.
- Enhance monitoring systems and improve support mechanisms to boost developer experience across platform services.
- Develop automation tools and scripts to eliminate repetitive operational tasks and increase system efficiency.
- Establish and refine alerting systems with proactive monitoring and automated response workflows.
- Develop and keep updated documented procedures for incident response and operational runbooks.
- Investigate infrastructure and platform problems, assist software engineers in root cause analysis, and coordinate with external vendors when needed.
- Produce detailed post-incident reports and implement corrective measures to prevent recurrence.
- Schedule and communicate planned maintenance activities for live systems with minimal disruption.
- Collaborate with development teams to uncover infrastructure-related workflow constraints and deliver scalable solutions.
- Evaluate and integrate industry-standard practices to ensure systems are secure, resilient, and highly available.
- Analyze open-source software architecture and implementation to improve debugging and resolution speed for platform issues.