Responsibilities
- Lead the design and improvement of automation solutions using PowerShell, Ansible, and Python to reduce manual work, improve consistency, and strengthen configuration control across enterprise infrastructure platforms.
- Design and maintain enterprise standards for Linux and Windows Server operating systems, virtual machine templates, baseline configurations, lifecycle management, and patch readiness across production, disaster recovery, and lab environments.
- Create scalable platform architecture patterns that improve reliability, security, efficiency, and consistency; support change management by ensuring infrastructure changes are documented, assessed for risk, communicated, tested, and aligned with operational standards.
- Provide technical leadership for monitoring and observability tools including Zabbix, LogicMonitor, and PagerDuty; improve coverage, alert quality, escalation workflows, and operational response across all platform environments.
- Work with services, security, application, and operations teams to define platform requirements, document reference architectures, and guide implementation decisions; create and maintain documentation, runbooks, architecture diagrams, build standards, and operational procedures.
- Support database infrastructure for SQL Server and Oracle in production and disaster recovery environments, including availability, recovery, monitoring, capacity planning, and operational standards; promote security, compliance, resilience, and cost-conscious engineering throughout the platform lifecycle.
Requirements
- Bachelor's degree or at least seven years of equivalent work experience, with extensive experience serving as a Systems Architect, Infrastructure Architect, or senior infrastructure administrator with responsibility for architecture decisions and platform standards in complex production environments.
- Proficiency in infrastructure automation and scripting using PowerShell, Ansible, Python, or similar technologies, with strong practical experience in Linux and Windows Server — including build standards, configuration management, troubleshooting, and lifecycle management.
- Working knowledge of monitoring and observability platforms, preferably Zabbix, LogicMonitor, or PagerDuty, with demonstrated expertise in capacity planning, performance tuning, backup and recovery, and business continuity planning across enterprise environments.
- Experience supporting infrastructure dependencies for SQL Server, Oracle, or other enterprise database platforms in production and disaster recovery environments, along with experience with hybrid infrastructure, virtualization platforms, cloud-related services, or self-service infrastructure catalogs.
- Strong spoken and written communication skills with a professional and respectful approach; proven ability to create clear technical documentation, architecture standards, runbooks, and implementation guidance for both technical and non-technical audiences, with strong collaboration and problem-solving skills.
Nice to Have
- Architecture and systems thinking — demonstrated ability to evaluate dependencies, tradeoffs, and long-term impacts to guide infrastructure decisions across virtualization, cloud, compute, storage, data protection, and data center services.
- Automation mindset — experience identifying repeatable work, selecting appropriate automation opportunities, and promoting reusable solutions across teams to drive operational consistency and efficiency at scale.
- Continuous learning and responsible AI use — proven commitment to keeping technical skills current and evaluating approved AI tools and emerging technologies responsibly before applying them to infrastructure work and platform decisions.
- Security and compliance expertise — demonstrated ability to apply risk-based judgment when balancing security, compliance, resilience, usability, and operational requirements throughout the platform lifecycle.
- Process and governance alignment — experience aligning technical work with IT asset management, change management, documentation, operational readiness, and service management requirements in enterprise environments.
Work Arrangement
Remote (City/Region) — Hyderabad, India