Responsibilities
- Design, build, and operate multi-agent workflows and tool-enabled agents, implementing orchestration logic, state management, safety guardrails, and fallback strategies for resilient production pipelines.
- Architect and maintain end-to-end RAG systems, covering document ingestion, chunking, embedding, vector retrieval, reranking, and answer synthesis with a focus on quality, attribution, and latency.
- Evaluate and integrate LLMs and GenAI services across cost, performance, and privacy dimensions, selecting the right mix of managed and in-house models.
- Develop, version, and optimise prompting strategies; implement automated prompt testing and regression tracking to maintain output quality and reliability.
- Define and own evaluation frameworks for generative outputs, including automated metrics, LLM-as-judge approaches, human evaluation protocols, hallucination detection, and drift monitoring.
- Apply classical NLP techniques where appropriate and maintain awareness of data distribution shifts that could impact model behaviour in production.
- Build and operate scalable, secure AI infrastructure on AWS (Bedrock, SageMaker, Lambda, OpenSearch), following well architected principles and infrastructure-as-code practices.
- Own the full deployment lifecycle: CI/CD for models and agents, testing strategies, observability, and rollback procedures.
- Ensure data quality through rigorous validation and augmentation, and proactively source datasets for training, fine-tuning, and evaluation.
Requirements
- Agent Orchestration: Production experience designing multi-agent systems with tool use, memory/state management, and fault-tolerant routing. Familiarity with LangChain, LangGraph, AutoGen, or custom orchestrators.
- RAG & Retrieval: Hands-on experience building RAG pipelines end-to-end: chunking, embedding models, vector databases, retrieval tuning, and answer synthesis at production scale.
- Evaluations: Strong experience defining and running evaluation pipelines for generative AI — automated scoring, human evaluation design, hallucination mitigation, and drift monitoring. LLM-as-judge patterns are a plus.
- LLMs & GenAI: Demonstrated experience with foundation models and GenAI providers (AWS Bedrock, OpenAI, Anthropic, Meta). Comfortable with fine-tuning, instruction tuning, and prompt engineering at scale.
- Traditional NLP: Solid grounding in classical NLP techniques (NER, text classification, intent detection, topic modelling) and good judgement on when to apply them alongside or instead of LLMs.
- AWS Bedrock: Hands-on experience with Amazon Bedrock: foundation model APIs, Bedrock Agents, Knowledge Bases, and Guardrails. Experience with Bedrock Model Evaluation is a plus.
- AWS Ecosystem: Proficiency with SageMaker, Lambda, ECS/EKS, S3, OpenSearch, IAM, CloudWatch, and VPC networking.
- MLOps & CI/CD: Familiarity with model registries, CI/CD for ML, feature stores, canary deployments, monitoring, and rollback.
- IaC: Experience with Terraform or AWS CDK for reproducible infrastructure provisioning.
- Python: Production-quality Python: packaging, testing (pytest), type hints, async programming, and clean ML pipeline abstractions.
- ML Frameworks: Familiarity with traditional ML frameworks and fine-tuning workflows.
- Experiment Tracking: Experience with Langfuse, Arize or Langsmith, or equivalent for tracking runs, metrics, and artefacts.
- Ability to translate ambiguous business goals into concrete technical solutions and communicate tradeoffs to non-technical stakeholders.
- Strong collaborative instincts — comfortable working across engineering, product, and data teams.
- A rigorous, evidence-driven mindset: you ship with confidence because you measure, test, and monitor thoroughly.
Nice to Have
- A Master’s degree in Machine Learning, Computer Science with a preference for specialization in the NLP domain.
Work Arrangement
Remote (Worldwide)