Responsibilities
- Help configure, maintain, and resolve issues in server, storage, and GPU-powered infrastructure environments
- Provide support for cloud platforms including AWS, Microsoft Azure, and Google Cloud Platform (GCP)
- Create and manage automation scripts using programming languages such as Python, Shell, Java, or JavaScript
- Track system performance, uptime, and operational status to maintain efficient and reliable services
- Work with internal teams to advance infrastructure modernization and improve operational processes
- Engage in diagnosing issues related to hardware, software, networking, and cloud systems
- Support documentation, reporting, and process enhancement efforts for infrastructure operations
- Use data analysis and visualization tools like Excel, Power BI, or Tableau to interpret system metrics and inform decisions
- Investigate and assist integrating Generative AI technologies into infrastructure management and workflows
- Promote a cooperative team culture by showing initiative, flexibility, and strong problem-solving skills
- Support incident management and escalation procedures to reduce system outages and disruptions
- Assist in monitoring security controls and compliance requirements to meet organizational and regulatory standards
- Take part in planning infrastructure capacity and optimizing resources to support scalable growth
- Help build and update technical documentation and runbooks to improve team knowledge sharing and efficiency
- Support backup systems, disaster recovery plans, and business continuity strategies to safeguard critical data and systems
- Assist with setting up, monitoring, and tuning network configurations for stable enterprise connectivity
Compensation
Competitive salary and benefits package commensurate with experience
Work Arrangement
Remote-friendly with potential hybrid options depending on location
Team
Collaborative engineering team focused on infrastructure reliability and innovation
Responsibilities
- Assist in the setup, configuration, maintenance, and troubleshooting of server, storage, and GPU-based infrastructure environments
- Support cloud infrastructure and services across platforms such as AWS, Microsoft Azure, or Google Cloud Platform (GCP)
- Develop and maintain automation scripts and tools using languages such as Python, Shell scripting, Java, or JavaScript
- Monitor system performance, availability, and operational health to ensure reliability and efficiency
- Collaborate with internal teams to support infrastructure modernization and operational improvements
- Participate in troubleshooting hardware, software, networking, and cloud-related issues
- Assist with infrastructure documentation, reporting, and process improvement initiatives
- Utilize data visualization and reporting tools such as Excel, Power BI, or Tableau to analyze operational metrics and support decision-making
- Explore and support the integration of Generative AI tools and concepts into infrastructure management and operational workflows
- Contribute to a collaborative team environment while demonstrating initiative, adaptability, and problem-solving capabilities
- Support incident response procedures and escalation protocols to minimize system downtime and operational disruptions
- Assist with security and compliance monitoring activities to ensure infrastructure environments meet organizational and regulatory standards
- Participate in capacity planning and resource optimization initiatives to support scalable infrastructure growth
- Contribute to the development and maintenance of technical knowledge bases and runbooks to support team efficiency and knowledge sharing
- Support backup, disaster recovery, and business continuity procedures to protect critical infrastructure and data assets
- Assist with network configuration, monitoring, and optimization activities to ensure reliable connectivity and performance across enterprise systems
Not available for this position