Responsibilities
- Design and maintain resilient systems supporting artificial intelligence training and inference operations
- Automate and manage operational processes in key domains such as networking, compute distribution, storage, hardware setup, or machine learning platforms
- Develop monitoring tools, alert systems, response procedures, and operational guides to enhance system manageability
- Identify and resolve performance, scalability, and stability issues spanning hardware, OS layers, network infrastructure, workload schedulers, and distributed computing environments
- Collaborate with machine learning, research, and infrastructure teams to align system capabilities with computational demands
- Enhance automation in infrastructure provisioning, configuration control, testing pipelines, and deployment workflows
- Support capacity planning, cluster expansion, resource distribution, system upgrades, and hardware lifecycle coordination
- Promote a culture of operational excellence through thorough documentation, incident analysis, and practical engineering practices
Work Arrangement
Remote — Toronto
Work Arrangement
Remote — Toronto