Apply on company website Remote Hybrid Full-time

fal is hiring a Machine Learning Engineer, Reliability

Responsibilities

  • Manage service level objectives for availability, latency, and throughput across a broad set of generative media model APIs handling large-scale production traffic
  • Develop monitoring, alerting, and observability tools to detect model-specific failures, output quality issues, pipeline disruptions, and regressions ahead of user impact
  • Strengthen model deployment processes using canary releases, shadow testing, automated rollbacks, and validation checkpoints to ensure safe model version rollouts
  • Lead security initiatives for the model fleet, including secure serving practices, detection of abuse or misuse, rate limiting, and defenses against adversarial inputs
  • Implement and maintain safety systems for generative media, including content moderation pipelines, safety classifiers, and inference-time guardrails that operate efficiently without degrading performance
  • Take ownership of incident response for model API outages or performance degradations, conduct postmortems, and implement engineering solutions to prevent recurrence
  • Optimize capacity planning, autoscaling strategies, and GPU fleet utilization for inference workloads experiencing fluctuating demand
  • Collaborate with model and infrastructure teams to integrate reliability, security, and safety requirements into the model onboarding process

Work Arrangement

Hybrid — India, Australia, New Zealand

Job Details
Location Remote
Work mode Hybrid
Employment Full-time
Department Engineering, ML
Category Data & ML
Posted 2 months ago
Application On company website
or drop your CV first
About company
fal logo

Generative media platform for developers.

The world's best generative image, video, and audio models, all in one place. Develop and fine-tune models with serverless GPUs and on-demand clusters.

Choose from 1,000+ production ready image, video, audio and 3D models. Build products using fal model apis. Scale custom AI models with fal serverless. Access 1000s of H100, H200 and B200 VMs with fal compute.

fal powers AI features in some of the world's most demanding environments — from public companies to hypergrowth startups. The platform is SOC 2 compliant and built for enterprise scale.

All jobs at fal Visit website