Responsibilities
- Work with service teams to establish and uphold Service Level Objectives and Indicators
- Create and deploy automated systems that enhance service reliability and operational performance
- Record technical designs and assist in integrating them across internal platforms
- Ensure technical solutions follow defined quality benchmarks, engineering standards, and IT policies
- Collaborate with development, infrastructure, and platform groups to strengthen system consistency and dependability
- Support observability by managing monitoring, logging, alerting, and incident response practices
- Participate in root cause investigations and ongoing improvement efforts after incidents
- Enhance deployment pipelines and operational workflows using automation and infrastructure upgrades
- Assist in maintaining and advancing cloud-based environments and services
- Share expertise and support initiatives aimed at improving system reliability across teams
Work Arrangement
Remote — Montréal