Responsibilities
- Design and maintain Python-based frameworks and services that coordinate engineering workflows across distributed systems and clusters.
- Develop reusable components for task scheduling, distributed execution, resource allocation, testing, workflow orchestration, and fault recovery.
- Establish well-defined APIs, modular interfaces, extensibility hooks, and data schemas to support long-term maintainability of infrastructure systems.
- Analyze and implement solutions for concurrency, asynchronous operations, multiprocessing, state handling, retry logic, idempotent processes, cancellation, and partial failure scenarios.
- Diagnose and resolve intricate issues involving Python applications, operating systems, processes, file systems, network configurations, remote nodes, and distributed services.
- Create comprehensive automated tests and detailed documentation to ensure infrastructure reliability and broad reusability.
- Collaborate with platform, continuous integration, release, quality assurance, machine learning systems, and product engineering teams to gather requirements and design scalable software solutions.
About the Team
The Core Infrastructure team develops the foundational software that enables engineering workflows across the organization. This includes building orchestration frameworks, execution engines, schedulers, test systems, developer tools, and shared platforms that manage complex workflows across machines, clusters, and hardware. The systems are mainly implemented in Python and serve as the control plane for distributed operations, handling resource coordination, execution state, concurrency, failure handling, retries, dependencies, and observability at scale.