Responsibilities
- Create and manage scalable, high-speed systems for training, deploying, and monitoring machine learning models.
- Develop and sustain optimized data workflows, model tracking, and experiment logging infrastructure.
- Work closely with machine learning researchers and engineering teams to resolve performance issues and enhance platform usability.
- Design and implement distributed computing and storage architectures tailored for ML workloads.
- Enhance continuous integration and deployment pipelines specific to machine learning models and infrastructure.
- Maintain platform reliability through comprehensive monitoring, logging, and alerting solutions.
- Keep up with advancements in ML infrastructure and adopt relevant technologies to improve system capabilities.
- Guide and support junior engineering staff while promoting high technical standards.
- Follow established quality management processes and support ongoing improvements in compliance and efficiency.
- Ensure team adherence to quality standards, oversee quality metrics, and lead process optimization initiatives.
Compensation
Highly competitive salary and benefits package, including 401(k) plan
Work Arrangement
Hybrid — Silicon Valley, United States, Europe
Other
- Catered free lunch, unlimited snacks and beverages.
- Highly competitive salary and benefits package, including 401(k) plan.