About the Role
The Senior Systems Engineer position is a member of the Engineering group within the Technical Operations department of our client. The function of the role is to provide MS Windows systems administration and support for engineering projects, infrastructure and development. The role will be focused on automation, building for scale, and fault tolerant systems. This role will also focus on technologies such as Puppet, VMWare, SAN/NAS, and other technologies. This is a highly challenging and progressive team environment with an opportunity to work alongside infrastructure engineers and architects.
Responsibilities
- Partners with the engineering, infrastructure, security and project management teams to help implement new systems that are scalable and highly resilient
- Develops automation tools and processes to help scale operations
- Responsible for disaster recovery testing and restoration of critical services during DR events
- Provides systems and services expertise and drive operational best practices
- Participates in capacity and scalability planning
- Researches, develops, and implements new technologies and practices
- Leverages or builds appropriate technical tools to perform administration tasks, root-cause analysis and service restoration (such as back up, restore, failover, log interpretation, and performance monitoring
- Develops and maintains positive and cooperative relationships, inside and outside of teams, interacting in a friendly, open, honest, and accepting manner
- Mentors junior systems administrators
- Ensures 24/7 availability of the production servers and applications within a team environment
- Resolves Tier 3 incidents and requests. Requires on-call availability on team rotation outside of normal business hours
- Works to develop internal SOPs and Service handbooks for key systems and services
- Performs other related duties as assigned
Requirements
- MS Windows systems administration and support
- Automation, building for scale, and fault tolerant systems
- Technologies such as Puppet, VMWare, SAN/NAS, and other technologies
- Partnership with engineering, infrastructure, security, and project management teams
- Development of automation tools and processes
- Disaster recovery testing and restoration of critical services
- Operational best practices
- Capacity and scalability planning
- Root-cause analysis and service restoration (including back up, restore, failover, log interpretation, and performance monitoring)
- 24/7 availability of production servers and applications
- Resolution of Tier 3 incidents and requests
- On-call availability on team rotation outside of normal business hours
- Development of internal SOPs and Service handbooks
Additional Information
- Requires on-call availability on team rotation outside of normal business hours