Responsibilities
- Design and enhance neural network operators for high-performance workloads
- Create new neural network operations using CUDA and custom runtime interfaces
- Lead optimizations at the runtime level, focusing on compute, memory, and task scheduling
- Manage the interface and interaction between runtime systems and neural network layers
- Implement and refine operator fusion techniques, such as combining matrix multiplication with bias and LayerNorm, to improve hardware efficiency
- Detect and eliminate performance bottlenecks across software and hardware layers
- Work closely with teams specializing in compilers, PyTorch framework development, and low-level system software
Compensation
Competitive market rate based on experience and qualifications
Work Arrangement
Hybrid work model with flexibility for remote and on-site collaboration
Team
Cross-functional engineering environment focused on next-generation AI accelerator development
Responsibilities
- Design and enhance neural network operators for high-performance workloads
- Create new neural network operations using CUDA and custom runtime interfaces
- Lead optimizations at the runtime level, focusing on compute, memory, and task scheduling
- Manage the interface and interaction between runtime systems and neural network layers
- Implement and refine operator fusion techniques, such as combining matrix multiplication with bias and LayerNorm, to improve hardware efficiency
- Detect and eliminate performance bottlenecks across software and hardware layers
- Work closely with teams specializing in compilers, PyTorch framework development, and low-level system software
Available for qualified candidates requiring sponsorship