Responsibilities
- Manage service level objectives for availability, latency, and throughput across a broad set of generative media model APIs handling large-scale production traffic
- Develop monitoring, alerting, and observability tools to detect model-specific failures, output quality issues, pipeline disruptions, and regressions ahead of user impact
- Strengthen model deployment processes using canary releases, shadow testing, automated rollbacks, and validation checkpoints to ensure safe model version rollouts
- Lead security initiatives for the model fleet, including secure serving practices, detection of abuse or misuse, rate limiting, and defenses against adversarial inputs
- Implement and maintain safety systems for generative media, including content moderation pipelines, safety classifiers, and inference-time guardrails that operate efficiently without degrading performance
- Take ownership of incident response for model API outages or performance degradations, conduct postmortems, and implement engineering solutions to prevent recurrence
- Optimize capacity planning, autoscaling strategies, and GPU fleet utilization for inference workloads experiencing fluctuating demand
- Collaborate with model and infrastructure teams to integrate reliability, security, and safety requirements into the model onboarding process
Work Arrangement
Hybrid — India, Australia, New Zealand