Responsibilities
- Core Network Operations & Health Ownership - Own 24×7 5G Core health across Skylo’s production NTN stack: AMF/SMF/UPF/AUSF pod status, NAS/NG-AP signaling success rates, session establishment and tear-down metrics, subscriber registration KPIs, IMS registration state, and Core-layer SLA compliance.
- Monitor and triage Core NF alarms using OSS dashboards, Grafana/other inhouse telemetry, and Loki log correlation — distinguish transient anomalies from systemic degradation before escalating or acting.
- Execute and own Core-domain runbooks for P2–P4 fault categories: pod restarts, persistent storage recovery, certificate rotation, IMSI state reconciliation, and BSS-IIS cluster incident response — without requiring engineering team involvement for covered fault classes.
- Maintain DMP certificate management procedures and own escalation to BOSS (BSS & OSS) for DMP outages, certificate rotation failures, and EMS alarm integration issues.
- Own IMSI lifecycle operations: activation, deactivation, KML file management, subscriber state reconciliation, and exception handling for provisioning failures through the OSS platform.
- 5G Core Incident Diagnosis & Escalation Authority - Serve as the L3 escalation authority for all Core-domain incidents: take ownership from the Incident Manager, diagnose at the protocol level using NAS traces, NG-AP message flows, Diameter/SIP signaling captures, and NF-specific log analysis, and deliver a resolution or a decision-grade root cause.
- Lead Core-domain troubleshooting bridges: command the technical investigation, direct vendor and engineering participants, correlate signals across AMF, SMF, UPF, AUSF, PCF, and IMS NFs, and drive the bridge to a documented resolution or a clear engineering handoff.
- Diagnose and resolve Core failure modes: UE registration failures, PDU session establishment drops, handover interruptions, IMS registration and call setup failures, SEPP interconnect errors, subscriber provisioning mismatches, and signaling loop conditions.
- Engage vendors with technical specificity: reproduce failures with log evidence, own the vendor ticket lifecycle, enforce SLA response commitments, and escalate vendor delays with full impact context.
- Participate in the global 24×7 on-call rotation as the Core domain escalation tier — reachable within defined SLA windows for Sev 1 events; function as the technical decision-maker, not the first responder.
- Root Cause Analysis & Post-Incident Ownership - Own Core-domain RCA end-to-end: lead the post-incident investigation, document the complete causal chain from triggering condition through downstream NF impact, and deliver systemic action items with owners, timelines, and measurable success criteria.
- Deliver Initial RCA documentation within defined SLA windows post-incident closure; own the final RCA through engineering review and sign-off.
- Identify systemic failure patterns — configuration drift, missing alarm coverage, stale thresholds, vendor software defects — and translate them into engineering requirements with clear impact, scope, and acceptance criteria.
- Contribute to the weekly and monthly Network Performance Report: Core NF availability, session success rates, MTTR by fault category, top recurring issues, and SLA deviation analysis.
- KPI Ownership & Performance Assurance - Define, baseline, and continuously refine Core KPIs: NAS registration success rates, PDU session establishment ratios, NG-AP failure rates, IMS registration and session success, UPF throughput and latency, and subscriber provisioning success — calibrated to Skylo’s NTN subscriber population, not vendor defaults.
- Proactively track Core availability metrics against MNO SLA commitments — flag degradation trends before they breach thresholds and initiate preventive action before a subscriber impact occurs.
- Establish and maintain Core NF performance baselines; detect and investigate deviations that indicate emerging failures, capacity constraints, or configuration regressions.
- Drive proactive network issue detection through OSS alarm tuning, alert threshold calibration, and correlation rule improvements — reduce noise, increase signal fidelity, and eliminate alarm storms that mask real events.
- Runbook Authorship & Operational Standards - Author, own, and maintain all Core-domain runbooks and SOPs — every procedure must be tested before it is relied upon in production; runbooks are living documents, not static artifacts.
- Define the diagnostic decision tree for each known Core fault class: entry condition, triage steps, isolation method, resolution action, and escalation criteria — written at the level where a Senior NRE can execute independently.
- Review and accept runbook contributions from Senior NREs; maintain the Core runbook library as the authoritative operational reference for the team.
- Identify runbook gaps from incident post-mortems and operational observations; prioritize gap closure based on incident frequency and MTTR impact.
- Cross-Functional Collaboration & Team Development - Partner with Product Engineering on Core software release readiness: define observability and operational acceptance criteria for new NF versions before they reach production; flag missing alarm coverage, changed default configurations, and new failure modes.
- Partner with BOSS (BSS & OSS) on OSS platform stability, IMSI management interface SLAs, and EMS alarm integration — provide operational requirements as input to OSS roadmap planning.
- Collaborate with Cloud Infrastructure NRE on Kubernetes-layer issues affecting Core NFs: PVC availability, pod scheduling failures, network policy changes, and GKE upgrade impacts.
- Surface toil and automation opportunities to the Service Assurance & Automation team — document the procedure, frequency, and MTTR cost as structured input to the automation backlog; contribute to closed-loop automation requirements for P3/P4 Core events.
- Mentor Senior NREs in Core domain depth: KPI interpretation, NAS/SIP/Diameter trace reading, log correlation patterns, and escalation judgment.
- Vendor Management & Engineering Interface - Work closely with Core NF vendors for incident resolution and RCA — manage vendor ticket lifecycle from opening through closure, ensure reproducible evidence is provided, and escalate responsiveness failures.
- Engage Skylo’s Core engineering team with full operational context when issues exceed operational resolution authority — deliver a structured problem statement, timeline, NF log bundle, and a clear question rather than a vague escalation.
- Support MNO partner technical discussions on Core-layer SLA definitions, IMSI provisioning workflows, and subscriber management API behavior — provide operational evidence and data to support partner conversations.
Requirements
- 8–10+ years of experience in 5G/4G Core operations or Core engineering in a production 24×7 environment — carrier or vendor side, with direct ownership of live Core NF infrastructure.
- Deep 3GPP Core expertise: working knowledge of TS 23.501/23.502, NAS protocol, NG-AP/S1AP, IMS architecture (SIP, Diameter, CSCF/TAS), IMSI lifecycle procedures, and subscriber provisioning flows.
- Production 5G Core NF troubleshooting: demonstrated ability to diagnose AMF registration failures, SMF session drops, UPF forwarding anomalies, AUSF authentication failures, IMS SIP/Diameter errors, and SEPP interconnect issues using NF logs and signaling traces.
- Packet capture and call flow analysis: proficiency with Wireshark or equivalent for NAS, SIP, Diameter, and HTTP/2 (SBI) protocol-level analysis.
- Kubernetes-native Core operations: pod health monitoring, rolling restart procedures, PVC and persistent storage management, log aggregation (Loki, ELK, or equivalent), and kubectl proficiency.
- Production observability: Prometheus/Grafana/other inhouse tools for Core NF KPI dashboarding; ability to build queries, set alert thresholds, and interpret metric-level degradation signals.
- Runbook authorship: ability to write diagnostic procedures at the level where a less-experienced engineer can execute them independently under incident pressure.
- ITSM proficiency (Jira, ServiceNow, or equivalent): incident lifecycle management, RCA documentation, and vendor ticket tracking.
- Strong written and verbal communication: capable of delivering RCA documents, engineering escalations with structured problem statements, and MNO-facing technical summaries.
Nice to Have
- Experience with NTN or satellite connectivity Core operations: IMSI provisioning for NTN devices, DMP integration, BSS/IIS platform familiarity.
- IMS operational experience: SIP registration and session troubleshooting, Diameter routing, CSCF and TAS log analysis.
- Experience with OSS/BSS platforms: FCAPS alarm management, EMS integration, SNMP trap handling, and provisioning system diagnostics.
- Familiarity with Core NF vendors: Druid, Nokia (Core), Mavenir, Samsung, or Oracle Core — including vendor CLI tools, log formats, and support escalation processes.
- Scripting ability in Python or Bash — sufficient to automate log parsing, build diagnostic queries, or create operational data exports.
- Experience contributing to closed-loop automation requirements or service assurance platform development.
- ITIL certification or demonstrated application of ITIL problem management and change management processes in a production environment.
Benefits
- Competitive compensation packages including a stock option based equity program
- Comprehensive benefits plans
- Monthly allowances for wellness and education reimbursement
- A generous time off policy, holidays, and the opportunity to temporarily work abroad
- Once-in-a-lifetime opportunity to be part of developing and running the world’s first commercial, live direct-to-device satellite network and service
- Access to a world-class team and talent across tech domains: software, hardware, chipsets, telecom, satellite and network virtualization
- Open, transparent, inclusive culture that blends Silicon Valley, Nordic and South Asia characteristics
Work Arrangement
Remote (Worldwide)
Additional Information
- Participate in the global 24×7 on-call rotation
- Must be reachable within defined SLA windows for Sev 1 events