Santa Clara, Canada On-site Full-time USD 187,945 – 269,503 / year

IonQ is hiring a Senior Staff Service Reliability and Operational Intelligence Engineer

Responsibilities

  • Lead the long-term technical vision and strategic planning for operational maturity across all stages of service lifecycle.
  • Create and enforce a structured process for introducing new services, ensuring compliance with architecture, security, resilience, observability, and support requirements prior to production deployment.
  • Set enterprise-wide standards for service ownership, including documentation of service records, responsible parties, dependencies, runbooks, support models, escalation procedures, recovery goals, and on-call preparedness.
  • Drive the design and ongoing enhancement of a unified observability platform, defining consistent practices for logs, metrics, distributed tracing, and profiling across systems.
  • Establish policies for dashboarding, alerting, synthetic monitoring, telemetry accuracy, data retention, sampling strategies, cardinality control, and cost management.
  • Manage the reliability governance framework for live services, including service-level indicators, objectives, error budgeting, and incident escalation protocols.
  • Correlate system health metrics with customer and business outcomes to detect anomalies early, prevent service degradation, and quickly identify user-impacting issues.
  • Improve incident response practices through standardized severity grading, incident leadership, stakeholder communication, automated evidence gathering, and coordinated handling of critical events.
  • Implement blameless post-mortem processes, ensure corrective actions are completed, and lead initiatives to eliminate recurring failure patterns.
  • Oversee capacity and efficiency planning, including demand modeling, cloud and Kubernetes resource scaling, performance validation, headroom policies, resource optimization, and risk assessments.
  • Design and manage AI-driven operations capabilities for event correlation, noise reduction in alerts, predictive issue detection, root cause identification, automated triage, guided remediation, and controlled self-healing.
  • Develop secure AI-powered workflows integrated across observability tools, service registry, issue tracking, documentation, source repositories, and deployment pipelines.
  • Enhance on-call performance through well-structured rotations, readiness criteria, escalation rules, diagnostic automation, alert quality oversight, and dependable global handover processes.
  • Provide direct technical leadership during major outages, complex reliability analyses, system design evaluations, resilience testing, and high-stakes service rollouts.
  • Leverage operational metrics, incident trends, service performance data, capacity signals, change success rates, and automation impact to guide continuous improvement initiatives.

Work Arrangement

On-site — Santa Clara, CA

Team

The team develops, secures, and maintains scalable infrastructure for cloud-hosted SaaS offerings that include on-premises components deployed at customer locations.

Team

The team builds, secures, and operates scalable infrastructure for cloud-managed SaaS products with on-premises components deployed at customer sites.

Job Details
Location Santa Clara, Canada
Work mode On-site
Employment Full-time
Salary USD 187,945 – 269,503 / year
Department Platform Engineering
Category other
Posted 19 days ago
or drop your CV first
About company
IonQ logo
IonQ is a quantum computing company that develops quantum computers and quantum computing technologies, focusing on creating advanced quantum systems.
All jobs at IonQ Visit website