Responsibilities
- Serve as a primary contact and leader for incidents and critical issues impacting our platform during NY business hours.
- Take ownership of incidents, drive effective communication, and facilitate swift resolution by collaborating with relevant engineering teams.
- Continuously monitor the health and performance of our applications and infrastructure.
- Analyze trends, identify potential risks, and proactively implement measures to prevent incidents and improve overall system reliability.
- Troubleshoot complex technical issues across various layers of the stack (application, infrastructure, network).
- Utilize analytical skills and technical expertise to identify root causes and implement effective solutions.
- Work closely with engineering, development, and operations teams to ensure seamless collaboration during incident response and in proactive reliability initiatives.
- Communicate effectively with stakeholders at all levels, providing clear and concise updates on incidents and system status.
- Identify opportunities to automate tasks, improve operational efficiency, and enhance the resilience of our systems.
- Develop tools and scripts as needed to streamline processes and reduce manual intervention.
- Contribute to the ongoing development and improvement of our SRE practices, tools, and processes.
- Share knowledge and expertise with the team to foster a culture of learning and growth.
Requirements
- Up to 5 years of experience in a Site Reliability Engineering (SRE), DevOps, or Production Engineering role, with a deep understanding of SRE principles and best practices.
- Incident management expertise, including triaging, escalation, and resolution of high-severity outages.
- Proficiency in at least one coding language (Python or Java) for automation and debugging.
- Hands-on experience in Kubernetes (K8s) for managing and orchestrating containerized applications.
- Cloud experience (AWS preferred) with exposure to key services like EC2, S3, Lambda, and CloudWatch.
- Excellent communication skills to articulate technical challenges and solutions effectively.
- Strong troubleshooting and problem-solving skills, with experience diagnosing complex production issues.
- Ability to stay calm under pressure, multitask, and prioritize effectively in fast-moving environments.
- Fluency in English (spoken and written) is required.
- Must have the legal right to work in the country.
Nice to Have
- Experience with Terraform or CloudFormation for infrastructure-as-code.
- Experience with monitoring tools (e.g., Datadog, Prometheus, Grafana).
- Familiarity with web application architectures and best practices.
- Exposure to CI/CD pipelines and DevOps workflows.
Work Arrangement
Hybrid — Lisbon
Additional Information
- Flexible work arrangements (hybrid model) and a casual dress code.
- Modern and comfortable office located at Avenida da Liberdade (Lisbon).
- Fluency in English (spoken and written) is required.
- Must have the legal right to work in the country.
- Recruiting emails from genuine Arcesium recruiters come from @arcesium.com domain.
- Independent search firms may contact candidates; their emails should come from their firm's domain.
- Arcesium will never ask for banking information or any payment as part of the recruiting process.
- Equal opportunity employer.