Responsibilities
- Lead full lifecycle validation for two Kubernetes platforms—Azure Local and FedRAMP-compliant GCP—spanning deployment, ongoing operations, upgrades, and retirement.
- Create, run, and record comprehensive test scenarios at the system level for reliability, availability, scalability, security, observability, performance, recovery, and upgrade capabilities, using clear Given/When/Then syntax accessible to engineering, product, SRE, and operations teams.
- Perform hands-on testing and debugging using kubectl, Helm, operators, CRDs, PVCs, StatefulSets, ArgoCD or Flux, network policies, and node affinity, with the ability to isolate issues from pod to node level independently.
- Validate configurations and integrations in Azure (AKS) and GCP (GKE), including monitoring tools, key management systems, and identity services, with openness to quickly adapt to Azure Local’s hybrid on-prem architecture.
- Analyze Terraform, Ansible, and Helm configurations to understand infrastructure setup and support root cause analysis without needing to write infrastructure code.
- Recreate real-world or customer-reported incidents in test environments, trace issues from infrastructure code through Kubernetes to cloud resources, and develop regression tests to prevent recurrence.
- Intentionally disrupt nodes, pods, network connections, and storage components (e.g., S2D pool degradation, unbound PVCs, dropped VPN tunnels) to confirm system resilience and adherence to SLOs during failures.
- Utilize chaos engineering tools such as LitmusChaos or Chaos Mesh when appropriate to test system robustness.
- Assess system behavior using metrics, logs, and distributed traces (via Prometheus, Grafana, Loki, OpenTelemetry, Alertmanager) rather than relying solely on API response codes.
- Evaluate TLS termination, certificate rotation, secret handling, RBAC enforcement, and container image scanning on GCP’s FedRAMP-authorized platform.
- Ensure audit logs are properly captured in GCP Cloud Logging and produce evidence aligned with FedRAMP Moderate/High control families (AC, AU, SC, SI, IA), applying equivalent standards to Azure Local where applicable.
- Collaborate with engineering teams during design reviews, feature planning, and operational playbook development to identify quality risks early in the development cycle.
Requirements
- Minimum of 5 years in QA or SDET roles, with at least 2 years of direct Kubernetes experience and 1+ year working in Azure or GCP.
- Practical experience with Kubernetes tools and concepts: kubectl, Helm, operators, CRDs, PVCs, StatefulSets, ArgoCD or Flux, network policies, and node affinity.
- Familiarity with Azure and GCP services including AKS/GKE, monitoring solutions, key vaults/KMS, and identity management features.
- Ability to read and interpret Terraform, Ansible, and Helm charts to understand infrastructure deployments and support troubleshooting.
- Experience designing and documenting tests across key system qualities—reliability, availability, scalability, security, observability, performance, recoverability, and upgradability—using Given/When/Then format.
- Proficiency with observability tools such as Prometheus, Grafana, Loki, OpenTelemetry, and Alertmanager.
- Demonstrated experience in security-focused testing, including TLS and certificate management, secret handling, RBAC, image scanning, audit logging, and generating compliance evidence for FedRAMP control families (AC, AU, SC, SI, IA).
- Acts as a collaborative peer during design and planning phases, not just during test execution.
- Clear and accountable communication skills, with the ability to document tests for both technical and non-technical stakeholders.
- Proven ability to manage and troubleshoot across two distinct platforms simultaneously.
- Security-first mindset with a focus on preventing defects and ensuring platform stability.
Nice to Have
- Experience with Zephyr or Jira’s native test management features.
- Background in chaos engineering using LitmusChaos, Chaos Mesh, or similar tools.
- Prior work in FedRAMP or other regulated environments such as SOC 2, HIPAA, PCI DSS, or DISA STIG.
- Familiarity with distributed data systems like Kafka, Redis, MariaDB/Galera, or S3-compatible storage including MinIO.
Work Arrangement
Hybrid
Work Arrangement
Hybrid