Responsibilities
- Serve as the primary technical owner for complex customer issues and escalations.
- Investigate and resolve technical problems spanning multiple systems and services.
- Provide clear, proactive communication to customers throughout the lifecycle of an issue.
- Monitor and triage production alerts impacting customers or system reliability.
- Coordinate incident response efforts across engineering and internal teams.
- Ensure incidents are properly documented, communicated, and followed through to resolution.
- Diagnose issues using logs, system metrics, and SQL queries.
- Analyze system behavior to identify root causes of production problems.
- Escalate and partner with engineering teams to drive long-term fixes.
- Develop and maintain troubleshooting documentation, runbooks, and operational processes.
- Identify recurring patterns and contribute to systemic improvements.
- Help strengthen incident response and operational best practices as the organization scales.
- Share technical insights and best practices with colleagues.
- Act as a technical resource within the Support organization.
Requirements
- Exceptional communication skills with a customer-first mindset, capable of translating complex technical issues into clear and actionable insights.
- A demonstrated ability to get on a call or have a remote session with a customer during less-than-optimal times, to troubleshoot and put them at ease.
- 2+ years of experience in Technical Support, Production Support, Technical Operations, or a similar customer-facing technical role.
- Strong troubleshooting and analytical skills, with the ability to quickly diagnose and resolve complex technical issues.
- Proficiency in SQL for data investigation, troubleshooting, and root cause analysis within production environments.
- Ability to analyze system behavior, investigate anomalies, and debug issues across distributed systems and application layers.
- Experience partnering closely with Engineering teams to escalate issues, provide technical context, and drive timely resolution.
- Comfortable operating in fast-paced production environments, including handling incidents, alerts, and time-sensitive customer-impacting issues.
Nice to Have
- Experience participating in incident response, production alerting, or on-call rotations in a live production environment.
- Familiarity with observability, monitoring, and debugging tools used to investigate system performance and reliability issues.
- Experience supporting or operating within SaaS or cloud-based platforms.
- Working knowledge of support, collaboration, and ticketing tools such as Salesforce, Jira / Atlassian, monday.com, or Slack.
Work Arrangement
Remote (Worldwide)
Additional Information
- This is a shift-based role supporting global customers.
- Shifts may include Pacific Time business hours, evenings, weekends, and some holidays.