Responsibilities
- Provide ML Engineers with infrastructure, tools, and frameworks to enable rapid, independent iteration.
- Reduce time-to-market for production ML solutions by ensuring seamless integration, proper service connectivity, and reliable access to data and computing resources.
- Lead the development and adaptation of ML-specific CI/CD pipelines in collaboration with machine learning teams, going beyond standard framework usage.
- Ensure ML Engineers maintain direct control over models in production, enabling real-time monitoring, troubleshooting, and refinement without reliance on staging environments.
- Support large-scale experimentation through resilient, reproducible, and scalable environments for both internal validation and live A/B testing.
- Implement and maintain core MLOps components such as MLflow, Kubeflow, and KubeRay, while dynamically managing GPU infrastructure to handle supply constraints like L40 shortages during training.
- Address technical debt in current systems while establishing scalable foundations for future initiatives.
- Act as a technical bridge between machine learning and core platform teams, understanding both domains to deliver lasting solutions.
- Manage operational duties including on-call rotations, incident post-mortems, and root cause analysis for level-1 failures.
Work Arrangement
Remote (Country) — France