Responsibilities
- Create multi-cloud architectural frameworks on AWS and Azure, following best practices for reliability, efficiency, security, and cost management
- Develop fully automated infrastructure provisioning using Infrastructure-as-Code tools such as Terraform, Pulumi, Bicep, and CDK, integrated with CI/CD pipelines for seamless deployments
- Implement and manage shared platform environments with strong separation between tenants, resource limits, namespace controls, and usage-based cost tracking
- Design on-premises systems for compute, networking, and storage, including GPU/CPU clusters, high-performance computing setups, and hybrid cloud connectivity via dedicated links
- Establish scalable platform strategies using dynamic auto-scaling, Kubernetes cluster scaling (EKS, AKS), and on-prem capacity forecasting to handle fluctuating AI demands
- Build and sustain a unified Knowledge Fabric layer by integrating vector databases, document stores, relational and graph databases, and metadata systems into coherent search and retrieval workflows
- Produce standardized infrastructure templates and platform offerings that support cloud-native applications, including microservices, AI inference services, model hosting platforms, event-driven pipelines, and data processing systems
Compensation
Not specified
Work Arrangement
Not specified
Team
Not specified
Responsibilities
- Design and implement multi-cloud reference architectures on AWS and Azure, applying Well-Architected principles across reliability, performance, security, and cost optimization pillars
- Architect and deploy fully automated, Infrastructure-as-Code (IaC) driven cloud provisioning using Terraform, Pulumi, Bicep, and CDK with CI/CD pipelines enabling zero-touch infrastructure deployments
- Build and operate multi-tenant platform infrastructure enforcing strict tenant isolation, resource quotas, namespace-level governance, and chargeback reporting mechanisms
- Design on-premises infrastructure for compute, networking, and storage, including bare-metal GPU/CPU clusters, HPC configurations, and hybrid-cloud connectivity solutions (VPN, Direct Connect, ExpressRoute)
- Define and implement platform scaling strategies combining horizontal auto-scaling, cluster autoscalers (EKS, AKS), and on-premises capacity planning to support elastic AI workloads
- Architect and maintain the Knowledge Fabric layer, integrating vector databases (Pinecone, Weaviate, pgvector), document stores, relational and graph databases, and structured metadata registries into unified, semantically searchable retrieval pipelines
- Develop infrastructure blueprints and platform services supporting diverse cloud-native workloads including microservices (service mesh, API gateway, event streaming), AI inference endpoints, AI model hosting (SageMaker, Azure ML, Triton), MCP servers, and data management pipelines
Not specified