Responsibilities
- Create and manage data infrastructure and pipelines for both batch and real-time processing, influencing architectural direction
- Develop and support robust ETL/ELT workflows using tools like Python, SQL, Spark, Flink, Beam, or similar technologies
- Implement streaming or near-instant data solutions such as change data capture and micro-batching for time-sensitive applications
- Collaborate with data scientists, machine learning engineers, analysts, and product teams to define data needs, service level agreements, and reusable data products
- Define and enforce standards for data quality, validation, monitoring, and observability including lineage tracking and anomaly detection
- Troubleshoot and resolve production issues involving performance, data accuracy, system resilience, and root cause analysis
- Maintain clear documentation of data pipelines, schemas, architectural designs, and operational procedures
Work Arrangement
Remote (Worldwide) — over 70 countries