Responsibilities
- Lead the full lifecycle development of large-scale data platforms, including ETL/ELT pipelines, data transformations, and storage infrastructure.
- Collaborate with engineering, product, and business teams to define requirements, validate assumptions, and deliver practical solutions from concept to deployment.
- Solve complex data challenges by building optimized data processing logic using Python, efficient SQL, and tools like dbt.
- Evaluate and balance technical trade-offs across scalability, reliability, speed, maintainability, and cost when making architectural decisions.
- Develop and lead technical design proposals and RFCs that clearly outline system architecture, alternatives, risks, and operational needs.
- Shape the long-term technical vision for owned data systems and contribute to broader team-level engineering strategies.
- Enforce high standards in code quality through rigorous design reviews, automated testing, schema versioning, and CI/CD implementation.
- Drive initiatives that enhance engineering excellence, such as data observability, lineage tracking, incident response protocols, and deployment safeguards.
- Build and manage high-volume batch and streaming data pipelines using technologies including Spark, Flink, and Kafka.
- Work closely with machine learning teams to provide well-structured, accurate datasets for model development and real-time inference.
- Promote responsible data management through comprehensive documentation, access controls, privacy safeguards, and governance frameworks.
- Conduct root-cause analyses after data incidents, identify systemic fixes, and ensure implementation of corrective measures.
- Break down complex initiatives into manageable tasks and coordinate execution across engineering team members.
Compensation
The base salary range reflects expected compensation, adjusted for location, skills, and experience; total package may include bonus and equity.
Work Arrangement
Remote (Country)