Responsibilities
- Architect and manage resilient cloud infrastructure across AWS, GCP, and Azure platforms.
- Operate and fine-tune Kubernetes clusters in live production settings.
- Handle dynamic scaling and resource forecasting using Karpenter and built-in Kubernetes tools.
- Develop and sustain Infrastructure as Code solutions using Terraform.
- Enhance continuous integration and delivery pipelines for both applications and infrastructure.
- Deploy and maintain middleware systems including Kafka, Redis, and NGINX.
- Automate routine operations using programming in Golang and Python.
- Strengthen system visibility through monitoring, logging, alerting, and distributed tracing setups.
- Ensure infrastructure meets high standards for reliability, scalability, security, and recovery readiness.
- Diagnose and resolve live production issues with root cause analysis.
- Collaborate with development teams to refine deployment workflows and developer platform experience.
- Take part in on-call duties and respond to critical system incidents.
Requirements
- Minimum of three years in roles such as DevOps, Platform Engineering, SRE, or Infrastructure Engineering.
- Proven expertise with at least one major cloud provider: AWS, GCP, or Azure.
- Direct experience managing Kubernetes in production environments.
- Working knowledge of Karpenter for autoscaling and node lifecycle in Kubernetes.
- Extensive use of Terraform for automating infrastructure provisioning.
- Solid background in Linux system administration and networking principles.
- Operational experience with middleware like Kafka, Redis, or NGINX.
- Strong coding abilities in Golang or Python for automation tasks.
- Familiarity with Git, GitOps practices, and CI/CD pipeline tools.
- In-depth knowledge of observability stacks including Prometheus and OpenTelemetry.
- Understanding of containerization with Docker and orchestration via Kubernetes.
- Demonstrated ability to troubleshoot and resolve complex technical issues.
Nice to Have
- Experience supporting blockchain infrastructure or Web3-based platforms.
- Background deploying and managing blockchain nodes on Layer 1 or Layer 2 networks such as Ethereum, Bitcoin, or Solana.
- Knowledge of blockchain components like RPC services, validators, indexing, or consensus mechanisms.
- Experience using Argo CD, Helm, Helmfile, or Flux for deployment management.
- Track record with multi-region and multi-cloud infrastructure deployments.
- Understanding of cloud security, IAM policies, network design, and secrets handling.
- Exposure to service mesh tools like Istio or Linkerd.
- Familiarity with message queues, distributed caching, and high-throughput architectures.
- Experience optimizing performance and capacity for large-scale distributed systems.
Other
- Passion for automation and continuous improvement.
- Strong ownership and accountability.
- Ability to work independently in a fast-paced environment.
- Excellent communication and collaboration skills.
- Curiosity to learn new technologies and solve complex infrastructure problems.
- Experience supporting production systems with high availability and reliability requirements.
- Nice to Have: Experience in cryptocurrency exchanges or blockchain infrastructure providers.
- Nice to Have: Experience operating multi-region, multi-cloud production environments.
- Nice to Have: Knowledge of RPC infrastructure, distributed databases, or high-performance networking.
- Nice to Have: Familiarity with security and compliance best practices (IAM, WAF, Vault, KMS, etc.).
- Nice to Have: Experience contributing to open-source projects.
- The opportunity to be part of one of the world’s leading blockchain ecosystems with vast career growth potential.
- Work alongside a diverse, global team of experts and innovators in a fast-paced, dynamic environment.
- Participate in cutting-edge projects that drive industry change.