Responsibilities
- Designing, building, and operating VP's backend services, including the orchestration engine and the automated decisioning pipeline, with a strong focus on reliability, scalability, and maintainability
- Delivering key orchestration features such as multi-method flow execution, branching and fallback logic, and webhook integrations
- Improving the resilience of production-critical systems — backpressure and idempotency handling, backwards-compatibility testing, SPOF elimination, and end-to-end monitoring
- Participating in on-call rotations, incident response, and postmortems to ensure our most critical flows stay healthy
- Driving engineering quality through code reviews, CI practices, runbook documentation, and mentoring teammates
Requirements
- 5+ years of experience building large-scale backend applications with Python, including event-driven and async architectures with message queues (e.g. SQS or similar)
- Proven ownership of production-critical services — you understand what it means to be on-call and take reliability seriously
- Deep experience designing for resilience in distributed systems: idempotency, backpressure, retries, backwards compatibility
- Strong understanding of modular design approaches and SOLID principles
- Hands-on experience building scalable, maintainable APIs
- Proficiency in a range of testing strategies that ensure robust and reliable software
- Solid understanding of relational database fundamentals
- Familiarity with structured logging and observability practices
- Strong communication skills — you can clearly articulate technical ideas to engineers and non-engineers alike
Nice to Have
- Experience with workflow or orchestration engines, event-driven systems, Kubernetes, or identity verification domain knowledge