Responsibilities
- Deploy and refine an AI-powered SRE solution to align with customer requirements in both production and pre-production settings.
- Monitor customer implementations actively to maximize product effectiveness and customer success.
- Identify hidden reliability problems such as misconfigurations, deployment issues, and scaling challenges using the platform.
- Advise customers on optimal strategies for integrating the AI SRE solution within their infrastructure.
- Design, develop, and maintain robust, scalable backend systems supporting the AI SRE product.
- Act as the technical liaison between customer needs and internal engineering and product teams.
- Lead post-mortem analyses to determine root causes of incidents and establish preventive actions.
- Ensure security standards are consistently applied across customer deployments.
- Educate customer SRE, operations, and platform teams on effective use of the AI SRE tooling.
- Manage large-scale enterprise transitions to the platform, including rule replication, phased rollouts, and adherence to strict timelines.
- Develop and fine-tune alert correlation logic, including condition building and field extraction from unstructured inputs.
- Connect the platform with customer tools for tickets, collaboration, observability, and documentation, meeting certification requirements.
- Improve investigation accuracy by resolving query and enrichment failures and refining root cause analysis.
- Establish proactive monitoring systems to detect deployment issues before customer impact.
- Manage communication deliverables such as executive updates, SLA documentation, and escalation protocols.
- Develop trusted technical relationships with senior engineering stakeholders at client organizations.
- Lead full lifecycle customer deployments, from discovery to stabilization, demonstrating reduced MTTR and manual effort.
- Convert ambiguous customer needs into structured technical plans, clearly communicating trade-offs and risks.
- Create and deploy AI-driven workflows with safeguards, fallbacks, and human oversight to maintain reliability.
- Develop standardized deployment assets like modules, architectures, and runbooks to improve efficiency.
- Define and track business value through KPIs, baseline measurements, and ROI reporting to leadership.
- Provide feedback from field experience to influence product roadmap and strengthen core features.
- Promote a culture of learning and iterative improvement within the organization.
Work Arrangement
Hybrid — Pleasanton, CA, India