Bangalore, India Remote (Global) Full-time

Skylo is hiring a Staff Network Reliability Engineer, Incident Management

Responsibilities

  • Core Network Operations & Health Ownership - Own 24×7 5G Core health across Skylo’s production NTN stack: AMF/SMF/UPF/AUSF pod status, NAS/NG-AP signaling success rates, session establishment and tear-down metrics, subscriber registration KPIs, IMS registration state, and Core-layer SLA compliance.
  • Monitor and triage Core NF alarms using OSS dashboards, Grafana/other inhouse telemetry, and Loki log correlation — distinguish transient anomalies from systemic degradation before escalating or acting.
  • Execute and own Core-domain runbooks for P2–P4 fault categories: pod restarts, persistent storage recovery, certificate rotation, IMSI state reconciliation, and BSS-IIS cluster incident response — without requiring engineering team involvement for covered fault classes.
  • Maintain DMP certificate management procedures and own escalation to BOSS (BSS & OSS) for DMP outages, certificate rotation failures, and EMS alarm integration issues.
  • Own IMSI lifecycle operations: activation, deactivation, KML file management, subscriber state reconciliation, and exception handling for provisioning failures through the OSS platform.
  • 5G Core Incident Diagnosis & Escalation Authority - Serve as the L3 escalation authority for all Core-domain incidents: take ownership from the Incident Manager, diagnose at the protocol level using NAS traces, NG-AP message flows, Diameter/SIP signaling captures, and NF-specific log analysis, and deliver a resolution or a decision-grade root cause.
  • Lead Core-domain troubleshooting bridges: command the technical investigation, direct vendor and engineering participants, correlate signals across AMF, SMF, UPF, AUSF, PCF, and IMS NFs, and drive the bridge to a documented resolution or a clear engineering handoff.
  • Diagnose and resolve Core failure modes: UE registration failures, PDU session establishment drops, handover interruptions, IMS registration and call setup failures, SEPP interconnect errors, subscriber provisioning mismatches, and signaling loop conditions.
  • Engage vendors with technical specificity: reproduce failures with log evidence, own the vendor ticket lifecycle, enforce SLA response commitments, and escalate vendor delays with full impact context.
  • Participate in the global 24×7 on-call rotation as the Core domain escalation tier — reachable within defined SLA windows for Sev 1 events; function as the technical decision-maker, not the first responder.
  • Root Cause Analysis & Post-Incident Ownership - Own Core-domain RCA end-to-end: lead the post-incident investigation, document the complete causal chain from triggering condition through downstream NF impact, and deliver systemic action items with owners, timelines, and measurable success criteria.
  • Deliver Initial RCA documentation within defined SLA windows post-incident closure; own the final RCA through engineering review and sign-off.
  • Identify systemic failure patterns — configuration drift, missing alarm coverage, stale thresholds, vendor software defects — and translate them into engineering requirements with clear impact, scope, and acceptance criteria.
  • Contribute to the weekly and monthly Network Performance Report: Core NF availability, session success rates, MTTR by fault category, top recurring issues, and SLA deviation analysis.
  • KPI Ownership & Performance Assurance - Define, baseline, and continuously refine Core KPIs: NAS registration success rates, PDU session establishment ratios, NG-AP failure rates, IMS registration and session success, UPF throughput and latency, and subscriber provisioning success — calibrated to Skylo’s NTN subscriber population, not vendor defaults.
  • Proactively track Core availability metrics against MNO SLA commitments — flag degradation trends before they breach thresholds and initiate preventive action before a subscriber impact occurs.
  • Establish and maintain Core NF performance baselines; detect and investigate deviations that indicate emerging failures, capacity constraints, or configuration regressions.
  • Drive proactive network issue detection through OSS alarm tuning, alert threshold calibration, and correlation rule improvements — reduce noise, increase signal fidelity, and eliminate alarm storms that mask real events.
  • Runbook Authorship & Operational Standards - Author, own, and maintain all Core-domain runbooks and SOPs — every procedure must be tested before it is relied upon in production; runbooks are living documents, not static artifacts.
  • Define the diagnostic decision tree for each known Core fault class: entry condition, triage steps, isolation method, resolution action, and escalation criteria — written at the level where a Senior NRE can execute independently.
  • Review and accept runbook contributions from Senior NREs; maintain the Core runbook library as the authoritative operational reference for the team.
  • Identify runbook gaps from incident post-mortems and operational observations; prioritize gap closure based on incident frequency and MTTR impact.
  • Cross-Functional Collaboration & Team Development - Partner with Product Engineering on Core software release readiness: define observability and operational acceptance criteria for new NF versions before they reach production; flag missing alarm coverage, changed default configurations, and new failure modes.
  • Partner with BOSS (BSS & OSS) on OSS platform stability, IMSI management interface SLAs, and EMS alarm integration — provide operational requirements as input to OSS roadmap planning.
  • Collaborate with Cloud Infrastructure NRE on Kubernetes-layer issues affecting Core NFs: PVC availability, pod scheduling failures, network policy changes, and GKE upgrade impacts.
  • Surface toil and automation opportunities to the Service Assurance & Automation team — document the procedure, frequency, and MTTR cost as structured input to the automation backlog; contribute to closed-loop automation requirements for P3/P4 Core events.
  • Mentor Senior NREs in Core domain depth: KPI interpretation, NAS/SIP/Diameter trace reading, log correlation patterns, and escalation judgment.
  • Vendor Management & Engineering Interface - Work closely with Core NF vendors for incident resolution and RCA — manage vendor ticket lifecycle from opening through closure, ensure reproducible evidence is provided, and escalate responsiveness failures.
  • Engage Skylo’s Core engineering team with full operational context when issues exceed operational resolution authority — deliver a structured problem statement, timeline, NF log bundle, and a clear question rather than a vague escalation.
  • Support MNO partner technical discussions on Core-layer SLA definitions, IMSI provisioning workflows, and subscriber management API behavior — provide operational evidence and data to support partner conversations.

Requirements

  • 8–10+ years of experience in 5G/4G Core operations or Core engineering in a production 24×7 environment — carrier or vendor side, with direct ownership of live Core NF infrastructure.
  • Deep 3GPP Core expertise: working knowledge of TS 23.501/23.502, NAS protocol, NG-AP/S1AP, IMS architecture (SIP, Diameter, CSCF/TAS), IMSI lifecycle procedures, and subscriber provisioning flows.
  • Production 5G Core NF troubleshooting: demonstrated ability to diagnose AMF registration failures, SMF session drops, UPF forwarding anomalies, AUSF authentication failures, IMS SIP/Diameter errors, and SEPP interconnect issues using NF logs and signaling traces.
  • Packet capture and call flow analysis: proficiency with Wireshark or equivalent for NAS, SIP, Diameter, and HTTP/2 (SBI) protocol-level analysis.
  • Kubernetes-native Core operations: pod health monitoring, rolling restart procedures, PVC and persistent storage management, log aggregation (Loki, ELK, or equivalent), and kubectl proficiency.
  • Production observability: Prometheus/Grafana/other inhouse tools for Core NF KPI dashboarding; ability to build queries, set alert thresholds, and interpret metric-level degradation signals.
  • Runbook authorship: ability to write diagnostic procedures at the level where a less-experienced engineer can execute them independently under incident pressure.
  • ITSM proficiency (Jira, ServiceNow, or equivalent): incident lifecycle management, RCA documentation, and vendor ticket tracking.
  • Strong written and verbal communication: capable of delivering RCA documents, engineering escalations with structured problem statements, and MNO-facing technical summaries.

Nice to Have

  • Experience with NTN or satellite connectivity Core operations: IMSI provisioning for NTN devices, DMP integration, BSS/IIS platform familiarity.
  • IMS operational experience: SIP registration and session troubleshooting, Diameter routing, CSCF and TAS log analysis.
  • Experience with OSS/BSS platforms: FCAPS alarm management, EMS integration, SNMP trap handling, and provisioning system diagnostics.
  • Familiarity with Core NF vendors: Druid, Nokia (Core), Mavenir, Samsung, or Oracle Core — including vendor CLI tools, log formats, and support escalation processes.
  • Scripting ability in Python or Bash — sufficient to automate log parsing, build diagnostic queries, or create operational data exports.
  • Experience contributing to closed-loop automation requirements or service assurance platform development.
  • ITIL certification or demonstrated application of ITIL problem management and change management processes in a production environment.

Benefits

  • Competitive compensation packages including a stock option based equity program
  • Comprehensive benefits plans
  • Monthly allowances for wellness and education reimbursement
  • A generous time off policy, holidays, and the opportunity to temporarily work abroad
  • Once-in-a-lifetime opportunity to be part of developing and running the world’s first commercial, live direct-to-device satellite network and service
  • Access to a world-class team and talent across tech domains: software, hardware, chipsets, telecom, satellite and network virtualization
  • Open, transparent, inclusive culture that blends Silicon Valley, Nordic and South Asia characteristics

Work Arrangement

Remote (Worldwide)

Additional Information

  • Participate in the global 24×7 on-call rotation
  • Must be reachable within defined SLA windows for Sev 1 events
Required Skills
Wireshark
About company
Skylo

Skylo was founded by engineers and scientists from MIT and Stanford, with a global deployment team spanning the US, Finland, and India. Together, they invented the missing software-defined network layer that turned satellites into cell towers — so coverage never drops and life doesn't stop.

Backed by some of the world's largest organizations, Skylo is trusted infrastructure for the carriers, businesses, governments, and their constituents, who can't afford to lose signal.

The future of connectivity isn't about more satellites. It's about making every network work together — and that's exactly what we're building with the Standardized Sky.

All jobs at Skylo Visit website
Job Details
Department Engineering, Network Operations
Category other
Posted 11 days ago