Vancouver, British Columbia, Canada Hybrid Full-time USD 133,000 / year

IFS is hiring a Senior Lead Site Reliability Engineer

Responsibilities

  • Design and evolve Azure infrastructure for SaaS platforms, managing end-to-end lifecycle from architecture to production operations to ensure high availability.
  • Lead the development and optimization of CI/CD pipelines using Jenkins, Azure Dev游戏副本, and GitHub Actions, with full accountability for reliability, security, and toolchain evolution.
  • Maintain and enhance Ansible playbooks to automate configuration management, provisioning, and configuration drift correction.
  • Develop and manage Infrastructure as Code using Terraform, ARM, or Bicep, covering initial deployment through ongoing operational changes.
  • Collaborate continuously with engineering teams to integrate reliability, deployability, and operational excellence into development workflows and release processes.
  • Manage the full observability stack, including Azure Monitor, Log Analytics, Application Insights, Prometheus, Grafana, and Pingdom, ensuring comprehensive monitoring and alerting.
  • Own end-to-end PagerDuty configuration, including escalation policies, service integrations, routing rules, and on-call scheduling.
  • Serve as technical escalation point during incidents and actively participate in on-call rotations.
  • Design, manage, and tune AKS clusters for production workloads, including node pools, autoscaling, networking, identity, and storage.
  • Oversee cluster health, version upgrades, and capacity planning for Kubernetes environments.
  • Implement Prometheus exporters and build Grafana dashboards to provide actionable insights into service performance, latency, errors, and resource usage.
  • Take complete technical ownership of all Cloud Operations systems, including infrastructure, pipelines, tooling, observability, and security, ensuring reliability and continuous improvement.
  • Lead root cause investigations for production outages and produce post-mortems with engineering-focused remediation actions.
  • Define, track, and manage service-level objectives, indicators, and error budgets to guide reliability investments for cloud-hosted services.
  • Implement and enforce security practices for identity, access, secrets, and certificates within Azure environments, going beyond policy to hands-on execution.
  • Support compliance with SOC 2 Type II, ISO 27001, and ISO 9001 by contributing to technical controls, evidence collection, and audit readiness.
  • Assess new Azure services and features against operational needs, conduct scalable proof-of-concepts, and drive adoption when justified.
  • Create and maintain precise technical documentation, including architecture diagrams, runbooks, and playbooks, meeting operational and ISO 9001 standards.
About company
IFS
IFS is a billion-dollar revenue company with 7000+ employees on all continents that provides award-winning enterprise software solutions powered by leading AI technology, enabling customers at their Moment of Service™.
All jobs at IFS Visit website
Job Details
Category infrastructure
Posted 14 days ago