Responsibilities
- Handle escalated support tickets, including GPU workload failures, container issues, networking problems, account infrastructure, and host-side configuration
- Diagnose and resolve issues across Docker, NVIDIA CUDA/GPU drivers, and virtualization environments (KVM)
- Troubleshoot network-layer issues: VLAN, DNS, DHCP, VPN, NAT, firewall rules, and connectivity failures on host machines
- Investigate performance issues on GPU utilization, container resource constraints, thermal throttling, driver conflicts, disk I/O bottlenecks
- Advise suppliers (hosts) on installation best practices — hardware setup, driver configuration, BIOS/firmware settings, and network configuration for optimal performance
- Provide managed support for supplier onboarding and ongoing machine management, acting as a technical resource through installation, configuration, and post-setup troubleshooting
- Write and maintain internal runbooks, escalation guides, and knowledge base articles to reduce repeat escalations
- Build diagnostic and automation tooling in Python and Bash to reduce manual triage overhead
- Collaborate with the engineering team and infrastructure support team to flag and document systemic or recurring platform issues
- Assist clients and infrastructure suppliers working with AI frameworks (TensorFlow, PyTorch) and GPU-accelerated workloads
- Provide coverage for L1 support team overflow during peak periods or incidents, per a defined on-call rotation
Requirements
- Solid Linux SysOps experience: Ubuntu Server, RHEL/CentOS, Debian; comfortable with systems, networking, storage, and permissions
- Proficiency with Docker: container debugging, Docker Compose, image management, cgroup resource limits, Docker storage/filesystem management
- Experience with virtualization: Proxmox VE, VMware, or similar hypervisors; provisioning and troubleshooting VMs
- Networking fundamentals: VLAN, DNS, DHCP, NAT, VPN, firewall rules, and general L2/L3 troubleshooting
- Hands-on experience with NVIDIA GPU drivers, CUDA, and GPU workload troubleshooting (essential)
- Scripting in Python and Bash for automation and diagnostic tooling
- Strong English written communication: clear, professional, and technically precise
- Experience providing technical support in a customer-facing or internal helpdesk context
- Ability to prioritize across a concurrent queue of escalated tickets, triaging by severity and customer impact, balancing reactive resolution against proactive documentation and tooling work, and making clear judgment calls on when to escalate versus own resolution end-to-end
Nice to Have
- Familiarity with AI/ML frameworks (TensorFlow, PyTorch) and running GPU-accelerated containers
- Monitoring and observability experience (Prometheus, Grafana)
- Relevant certifications: RHCSA, CompTIA Linux+, or similar
- Knowledge of the Vast.ai platform as a client or infrastructure supplier
Benefits
- Comprehensive health, dental, vision, and life insurance
- 401(k) with company match
- Meaningful early-stage equity
- Onsite meals, snacks, and close collaboration with founders/tech leaders
- Ambitious, fast-paced startup culture where initiative is rewarded
Work Arrangement
On-site — Westwood (LA)
Additional Information
- Schedule: Sunday - Thursday
- Adaptable to a defined on-call rotation which may include weekend coverage
- Full-time
- Onsite in our office in Westwood (LA)