Responsibilities
- Design and maintain scalable infrastructure supporting the full lifecycle of machine learning and AI systems, including training, fine-tuning, and deployment of traditional models and large language models.
- Build and manage automated CI/CD pipelines, container orchestration, and workflow systems to enable smooth deployment of ML and generative AI workloads.
- Define and implement standardized practices for model reproducibility, version control, packaging, and deployment across hybrid infrastructure environments.
- Operate and optimize infrastructure for large language models, including GPU clusters, inference servers, and vector databases, with a focus on performance metrics such as latency, throughput, and token efficiency.
- Implement monitoring and governance solutions for AI systems to track model drift, hallucinations, bias, token consumption, and hardware utilization while ensuring system reliability and compliance.
- Partner with data scientists, AI developers, and IT teams to deliver robust, user-friendly, and forward-compatible AI platforms.
- Evaluate and adopt emerging tools and frameworks in MLOps and generative AI, such as MLflow, Ray, vLLM, Hugging Face TGI, Triton, and orchestration agents, to maintain platform modernization.