Responsibilities
- Monitor critical physical infrastructure using PLC, BMS, and DCIM platforms such as Distech, Radix IoT/Mango, Schneider Electric, Siemens, or similar systems.
- Monitor and respond to alerts involving power, cooling, network connectivity, environmental conditions, and facility infrastructure.
- Provide Tier 2 support for escalations from L1 technicians and coordinate escalation to engineering, facilities, vendors, or other teams as required.
- Own incidents through their lifecycle, from initial triage and troubleshooting through resolution, stakeholder communication, and documentation.
- Support incident response and post-incident reviews, identifying opportunities to improve operational reliability and response procedures.
- Develop and refine monitoring dashboards, alerting thresholds, and operational health indicators.
- Perform and coordinate routine health checks of critical infrastructure, including UPS systems, PDUs, CRAC/CRAH units, backup generators, and environmental monitoring systems.
- Troubleshoot and coordinate resolution of mechanical and electrical infrastructure issues using a working knowledge of MEP systems.
- Read and interpret technical documentation, including electrical one-line diagrams, network diagrams, schematics, and equipment documentation.
- Coordinate scheduled and emergency maintenance activities with internal teams, vendors, and remote hands personnel.
- Participate in change management processes and assess the operational and infrastructure impact of proposed changes.
- Ensure physical and logical security policies and procedures are followed within data center and edge environments.
- Operate and support modular, containerized, micro, and distributed edge data center environments.
- Help maintain availability, continuity, and resiliency across geographically distributed and remotely operated infrastructure.
- Coordinate remote troubleshooting and hands-on support for edge deployments.
- Implement and continuously improve operational practices for remote infrastructure monitoring, fault detection, escalation, and recovery.
- Support environmental sensors and IoT-based monitoring integrations across remote infrastructure.
- Use operational platforms such as Grafana, Zoho Desk, Zenduty, ServiceNow, Jira, SolarWinds, or similar tools for monitoring, analysis, ticketing, and incident management.
- Maintain accurate documentation of configurations, incidents, changes, maintenance activities, and recurring operational tasks.
- Support automation and reporting initiatives focused on infrastructure performance, availability, capacity, and uptime.
- Assist with gathering operational evidence and infrastructure data for compliance and audit requirements.
- Identify opportunities to automate repetitive operational activities using scripting and monitoring integrations.
- Partner closely with Facilities, Infrastructure Engineering, Networking, IT, Security, and other operational teams to maintain reliable infrastructure.
- Communicate clearly during incidents, maintenance activities, escalations, and operational handoffs.
- Participate in incident reviews and root cause analysis and help drive corrective and preventive actions.