Responsibilities
- Keep SHEIN’s mission-critical production systems running 24/7/365, participating in on-call rotations and acting decisively during incidents.
- Triage and resolve production incidents, leveraging AI-assisted log analysis and anomaly detection to accelerate root cause identification; drive continuous improvements that reduce MTTR and prevent recurrence.
- Monitor and manage capacity planning and resource utilization, partnering with cross-functional teams to ensure systems scale safely while remaining cost-effective.
- Own and operate core open-source infrastructure such as APISIX, Nginx, Kubernetes, Kafka, Elasticsearch, Redis, Consul, Etcd, Zookeeper and other large-scale distributed systems.
- Design, build, and maintain observability solutions (metrics, logs, traces, alerting), incorporating AI-powered anomaly detection and intelligent alert correlation to surface actionable signals from high-volume telemetry, improving system visibility and resiliency.
- Automate operational workflows and eliminate manual toil through scripting, tooling, and process improvements, including the use of AI-assisted development tools (e.g., Claude Code) to accelerate the building and iteration of internal operational platforms.
- Develop and maintain technical documentation, including runbooks, architecture diagrams, operational procedures, and on-call playbooks.
- Work closely with global engineering teams to improve infrastructure reliability and performance through better system design and operational discipline.
- Mentor Senior and mid-level SREs, raising the overall technical bar and operational maturity of the team.
- Lead efforts to modernize the platform in alignment with industry best practices and evolving technology standards.
Requirements
- Bachelor’s degree in Computer Science, Information Systems, or a related technical discipline, or equivalent practical experience.
- 6+ years of experience owning and operating large-scale, high-traffic, 24/7 production systems, ideally in cloud or cloud-native environments.
- Experience applying AI/LLM-powered tools to reliability engineering, including designing and building automation or internal tools using AI-assisted development tools (e.g., Claude Code).
- Solid foundations in Linux, networking, and distributed systems, with the ability to debug complex production issues end to end.
- Hands-on experience with incident response, troubleshooting, and performance optimization in distributed systems.
- Strong software engineering skills with experience building automation, tooling, or platforms in languages such as Python or Go.
- Experience operating or supporting open-source infrastructure components such as APISIX, Nginx, Kubernetes, Kafka, Elasticsearch, Redis, Consul, Etcd, Zookeeper, etc.
- Experience with observability and monitoring systems (Prometheus, Grafana, Zabbix, etc.) and performance analysis.
- Familiarity with Git, CI/CD pipelines, and configuration management tools (e.g., Ansible).
- A strong sense of ownership, a systematic approach to problem-solving, and a passion for making systems more reliable.
- Strong communication skills and the ability to collaborate effectively with geographically distributed teams.
Nice to Have
- Bilingual fluency in Mandarin and English.
- Kubernetes Administrator certification or equivalent real-world experience.
- Experience operating big data platforms (Hadoop, Yarn, HBase, Hive, Spark).
- Experience applying AI/LLM-powered tools to reliability engineering, including designing and building automation or internal tools using AI-assisted development platforms (e.g., Claude Code).
Work Arrangement
Remote (Worldwide)
Additional Information
- Participating in on-call rotations
- Must be able to act decisively during incidents
- Must collaborate effectively with geographically distributed teams