Responsibilities
- Be the Observability SME, providing technical consultation and hands-on expertise in implementing observability solutions.
- Manage observability solutions using Dynatrace, ensuring monitoring of applications, infrastructure, services, and user experience.
- Develop automation scripts using Python and Shell scripting or similar to improve monitoring setup, alerting, reporting, dashboard creation, and operational workflows.
- Manage observability infrastructure in cloud and support migration of On-Prem Infrastructure to Cloud.
- Configure alerts to improve proactive monitoring and reduce noise, false positives, and operational inefficiencies.
- Collaborate with DevOps, Cloud Engineering, Operations, Developers, and Platform teams to embed observability into the software delivery lifecycle.
- Help develop SLIs, SLOs, and error budgets to improve service reliability and operational visibility.
- Troubleshoot complex monitoring, telemetry, and platform issues across distributed systems, microservices, and cloud-native environments.
- Support continuous improvement of incident management workflows by integrating observability insights into operational processes.
- Contribute to observability standards, documentation, best practices, and knowledge-sharing across teams.
- Work with partners to understand monitoring requirements and translate them into scalable observability solutions.
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- 7–10 years of experience in Observability Engineering, Monitoring Engineering, DevOps, Cloud Operations, or related technology roles.
- Hands-on experience with Dynatrace as a primary observability platform.
- Automation scripting experience using Python and Shell scripting.
- Hands-on experience managing or supporting infrastructure in AWS Cloud environments.
- Experience with infrastructure monitoring, application performance monitoring, log monitoring, alerting, dashboards, and event correlation.
- 5+ years of experience working in DevOps environments with exposure to CI/CD, automation, cloud platforms, and operational support practices.
- 3+ years of experience working with your partners, development teams, platform teams, and operations teams to implement monitoring and observability solutions.
Nice to Have
- Hands-on experience with open-source observability solutions like Prometheus and Grafana.
- Any cloud observability certification (e.g., Dynatrace, Datadog, and Splunk).
- Cloud certifications (e.g., AWS Solutions Architect, DevOps Engineer, Developer – AWS preferred).
- 3+ years of experience with Kubernetes and containerized environments.
- Knowledge of ITIL, ITSM, and ServiceNow Event Management.
- Experience with Terraform for infrastructure automation and provisioning.
- 3+ years of experience with Kubernetes and Docker.
- Hands-on experience with Datadog.
- Working knowledge of OpenTelemetry for standardized telemetry collection across traces, metrics, and logs.
- Familiarity with ITIL, ITSM, incident management, and ServiceNow Event Management.
Work Arrangement
Hybrid — Hyderabad
Team
Reports to: Engineering Manager
Additional Information
- Operating in a hybrid work model.