Responsibilities
- Own the complete lifecycle of AI models from experimentation to production deployment, monitoring, evaluation, and optimization.
- Improve existing AI systems by enhancing accuracy, reliability, latency, and scalability.
- Drive new AI initiatives from concept to production, collaborating closely with product, engineering, and business teams.
- Design, build, and optimize LLM-powered applications and agent-based workflows. Develop (RAG) pipelines and knowledge retrieval systems.
- Build robust orchestration frameworks using LangGraph and related agent frameworks.
- Design deterministic and semi-deterministic evaluation frameworks for LLMs and agent systems.
- Build automated testing frameworks for prompts, workflows, agents, and model outputs. Analyze model failures and implement improvements through structured evaluation cycles.
- Develop and maintain scalable LLMOps infrastructure and workflows. Implement monitoring, observability, experimentation, versioning, and deployment processes.
- Optimize AI systems for performance, reliability, and cost efficiency.
- Drive best practices around model governance, testing, and release management.
Requirements
- 7+ years of experience in AI/ML, Machine Learning Engineering, Applied AI, or related roles.
- Proven experience building and shipping AI products into production at scale.
- Experience owning AI systems end-to-end, from design through deployment and monitoring.
- Strong Python programming skills and solid software engineering fundamentals.
- Hands-on experience with LLM APIs, Agentic Systems, LangGraph, RAG, MCP, and LLM Evaluation frameworks.
- Experience designing deterministic and semi-deterministic testing frameworks for AI systems and workflows.
- Strong knowledge of traditional Machine Learning models, data analysis, debugging, SQL, APIs, JSON, and workflow observability.
- Experience with LLMOps/MLOps practices, including model deployment, monitoring, evaluation, and lifecycle management.