Responsibilities
- Architect, develop, and maintain a scalable, multi-tenant AI platform on AWS, leveraging LLM orchestration frameworks and tooling
- Empower engineering teams to create, deploy, and scale customer-facing AI agents and autonomous workflows
- Develop robust generative AI microservices using frameworks like FastAPI and implement stateful, multi-agent systems
- Choose optimal orchestration strategies—such as model-driven or graph-based approaches—based on requirements for traceability, speed, and system complexity
- Design and manage an automated pipeline for content ingestion, covering upload processing, text segmentation, embedding model selection, vector storage, and indexing via OpenSearch Serverless
- Deliver a self-serve, fully managed knowledge retrieval system accessible to all product development teams
- Develop and sustain platform-level APIs and developer tools, including MCP-compatible interfaces, real-time streaming protocols (SSE/WebSocket), and agent configuration SDKs
- Oversee the usability and design of internal AI SDKs used by engineers building agents and integrating streaming APIs
- Build and support the Platform Onboarding Copilot, integrated with Slack for internal onboarding workflows
- Create clear, actionable technical documentation and architectural decision records to guide future development
- Communicate platform vision and roadmap to business and non-technical stakeholders effectively
- Treat observability infrastructure as a product, ensuring comprehensive monitoring and diagnostics
- Provide built-in tools for engineers to access distributed tracing and performance metrics automatically upon agent deployment
- Implement an LLM evaluation framework (e.g., LangSmith, Ragas) that blocks CI/CD pipelines if quality thresholds are not met
- Establish per-customer cost tracking with automated spending alerts to ensure financial accountability
- Define quality benchmarks and testing suites tailored to different categories of AI agents
- Promote evaluation practices as a core engineering function across the organization
- Ensure multi-tenant security and isolation from initial design, including namespace separation, tenant identifier injection, and policy-as-code enforcement
- Maintain a centralized policy-as-code library tailored to specific product lines and business units
- Guarantee platform adherence to SOC 2 Type II, PCI DSS, GDPR, and CCPA compliance standards
- Develop and manage the data deletion workflow required for GDPR compliance
- Implement mandatory content safety controls, including grounding verification and filtering, for all externally visible AI agents
- Serve as the gatekeeper for quality and cost efficiency before agents move into pre-production environments
- Execute standardized evaluation tests, enforce best practices, and validate per-inference cost against defined limits
- Partner with DevOps teams through defined infrastructure contracts and build version promotion workflows using tools like GitHub Actions
Work Arrangement
Hybrid — Gurugram, Jaipur
Team
Reports to Director of AI and Analytics