Requirements
- Strong infrastructure builder with experience operating production, research, cloud, or high-performance compute systems
- Deeply comfortable with Linux administration, including debugging networking, storage, system services, permissions, performance issues, and node-level failures
- Experienced with Kubernetes in real environments, including cluster operations, deployments, networking, observability, scaling, and troubleshooting
- Comfortable working with cloud infrastructure on AWS, GCP, Azure, or equivalent platforms
- Familiar with infrastructure automation and configuration tools such as Terraform, Ansible, Helm, ArgoCD, GitOps workflows, or similar systems
- Experienced with GPU-heavy, compute-heavy, or HPC-style workloads, especially in environments involving AI, ML, research computing, or scientific workloads
- Able to work across bare metal and cloud environments, and interested in the practical tradeoffs between the two
- Comfortable reasoning about resource scheduling, cluster utilization, autoscaling, storage, networking, and observability for distributed workloads
- Practical and ownership-oriented; you can take ambiguous infrastructure needs and turn them into working systems
- Comfortable collaborating across disciplines, especially with researchers and engineers who may not think in infrastructure terms
- Able to operate independently as a senior or strong intermediate contributor, while knowing when to bring others into important technical decisions
- Motivated by building foundational systems that make ambitious technical and scientific work possible
Nice to Have
- Hands-on experience with production-grade LLM inference and serving engines, such as vLLM, SGLang, or TensorRT
- Experience working at an AI company, ML infrastructure team, research lab, university compute environment, HPC center, or scientific computing organization
- Experience supporting model inference, model serving, distributed training, high-throughput batch workloads, or internal ML platforms
- Hands-on experience with Slurm or similar HPC schedulers, including job scheduling, resource allocation, queue management, and cluster configuration
- Experience operating GPU infrastructure, including NVIDIA drivers, CUDA, container runtimes, scheduling, utilization, and hardware failure modes
- Experience with RDMA, InfiniBand, high-performance networking, distributed filesystems (ie. Lustre, BeeGFS), object storage, or storage systems for compute-heavy workloads
- Experience with Kubernetes operators, custom controllers, CRDs, or platform tooling for AI/ML workloads
- Experience with Prometheus, Grafana, Loki, OpenTelemetry, Datadog, or similar monitoring, logging, and observability tools
- Experience with container registries, image optimization, CI/CD systems, deployment pipelines, and secure software delivery
- Experience leading engineering operations or infrastructure efforts while remaining hands-on technically
- Familiarity with security, access control, secrets management, and reliability practices in production or research environments
Benefits
- The opportunity to work on foundational problems at the intersection of AI and physics
- A high-trust, low-bureaucracy environment with real ownership
- Remote-first work with flexibility in how you structure your day
- Exposure to cutting-edge ideas across AI, scientific discovery, infrastructure, and emerging technologies
- A culture that values curiosity, depth of thinking, and first-principles reasoning
- The chance to shape the compute and inference infrastructure behind advanced AI systems for scientific discovery
Work Arrangement
Remote (Worldwide) — Canada, US, UK