Responsibilities
- Ensure high availability, performance, and operational stability of Kafka clusters deployed across multiple cloud and on-premises environments.
- Develop and manage automated solutions for cluster management, version updates, scaling operations, and recovery from outages.
- Collaborate with development teams to support new streaming applications, provide guidance on optimal usage patterns, and strengthen data pipeline resilience.
- Diagnose and resolve critical system issues impacting data throughput, latency, or cluster stability, and conduct thorough post-incident reviews.
- Manage Kafka software upgrades and infrastructure transitions while maintaining service continuity for dependent systems.
- Take part in rotating on-call duties supported by a global team using a 'follow-the-sun' approach to minimize individual on-call burdens.
Compensation
There may be flexibility with the range included in this posting should a candidate be leveled higher or lower than the posted range.
Work Arrangement
Remote (Country) — Canada
Team
Engineers work across distributed teams with shared ownership of system reliability and operational excellence.
Other
- This position is fully remote and does not require residency in a specific region within Canada.
- Applicants from all regions of Canada are encouraged to apply.
- The compensation range may be adjusted based on the candidate's experience level.
Not specified