Ontario, Canada - Remote Remote (Global) Full-time

FirstPrinciples is hiring an AI & HPC Infrastructure Engineer

Requirements

  • Strong infrastructure builder with experience operating production, research, cloud, or high-performance compute systems
  • Deeply comfortable with Linux administration, including debugging networking, storage, system services, permissions, performance issues, and node-level failures
  • Experienced with Kubernetes in real environments, including cluster operations, deployments, networking, observability, scaling, and troubleshooting
  • Comfortable working with cloud infrastructure on AWS, GCP, Azure, or equivalent platforms
  • Familiar with infrastructure automation and configuration tools such as Terraform, Ansible, Helm, ArgoCD, GitOps workflows, or similar systems
  • Experienced with GPU-heavy, compute-heavy, or HPC-style workloads, especially in environments involving AI, ML, research computing, or scientific workloads
  • Able to work across bare metal and cloud environments, and interested in the practical tradeoffs between the two
  • Comfortable reasoning about resource scheduling, cluster utilization, autoscaling, storage, networking, and observability for distributed workloads
  • Practical and ownership-oriented; you can take ambiguous infrastructure needs and turn them into working systems
  • Comfortable collaborating across disciplines, especially with researchers and engineers who may not think in infrastructure terms
  • Able to operate independently as a senior or strong intermediate contributor, while knowing when to bring others into important technical decisions
  • Motivated by building foundational systems that make ambitious technical and scientific work possible

Nice to Have

  • Hands-on experience with production-grade LLM inference and serving engines, such as vLLM, SGLang, or TensorRT
  • Experience working at an AI company, ML infrastructure team, research lab, university compute environment, HPC center, or scientific computing organization
  • Experience supporting model inference, model serving, distributed training, high-throughput batch workloads, or internal ML platforms
  • Hands-on experience with Slurm or similar HPC schedulers, including job scheduling, resource allocation, queue management, and cluster configuration
  • Experience operating GPU infrastructure, including NVIDIA drivers, CUDA, container runtimes, scheduling, utilization, and hardware failure modes
  • Experience with RDMA, InfiniBand, high-performance networking, distributed filesystems (ie. Lustre, BeeGFS), object storage, or storage systems for compute-heavy workloads
  • Experience with Kubernetes operators, custom controllers, CRDs, or platform tooling for AI/ML workloads
  • Experience with Prometheus, Grafana, Loki, OpenTelemetry, Datadog, or similar monitoring, logging, and observability tools
  • Experience with container registries, image optimization, CI/CD systems, deployment pipelines, and secure software delivery
  • Experience leading engineering operations or infrastructure efforts while remaining hands-on technically
  • Familiarity with security, access control, secrets management, and reliability practices in production or research environments

Benefits

  • The opportunity to work on foundational problems at the intersection of AI and physics
  • A high-trust, low-bureaucracy environment with real ownership
  • Remote-first work with flexibility in how you structure your day
  • Exposure to cutting-edge ideas across AI, scientific discovery, infrastructure, and emerging technologies
  • A culture that values curiosity, depth of thinking, and first-principles reasoning
  • The chance to shape the compute and inference infrastructure behind advanced AI systems for scientific discovery

Work Arrangement

Remote (Worldwide) — Canada, US, UK

Required Skills
Resource AllocationQueue ManagementPrometheusGrafanaOpenTelemetryDatadogMonitoringSecurity
About company
FirstPrinciples

FirstPrinciples is a research company building AI systems for discovery in fundamental science.

Our mission is to understand the nature of reality by advancing AI-driven discovery in fundamental science. We are building systems centred around the scientific method, enabling new ways to explore the world around us.

We enable new modes of discovery by building AI systems that contribute to rigorous, original scientific research. As AI becomes foundational to discovery, it must protect knowledge as a shared asset, not a proprietary advantage. FirstPrinciples is building AI solely for fundamental research, transparent by design, and committed to knowledge as a public good.

All jobs at FirstPrinciples Visit website
Job Details
Department FirstPrinciples
Category DevOps & SRE
Posted 2 months ago