Responsibilities
- Develop and enhance batch data workflows using dbt for transformations and Dagster for orchestration, scheduling, and tracking data asset lineage.
- Create and refine BigQuery data models—such as dimensional, wide-table, or domain-focused designs—to enable analytics, experimentation, and reporting.
- Strengthen real-time data processing by building and maintaining streaming pipelines with Kafka or PubSub and Flink, mainly through FlinkSQL, to provide timely event data and metrics.
- Define and implement data platform standards covering software development lifecycle, naming, modeling, incremental processing, schema evolution, and best practices for both batch and streaming with CI/CD and testing.
- Enhance system dependability and visibility through monitoring, alerts, and service level objectives for data pipelines and quality checks.
- Collaborate with analytics, product, and engineering teams to integrate new data sources, establish data contracts, and deliver high-quality datasets.
- Manage data platform operations, including performance optimization, data integrity, cost efficiency, and scaling across data warehouse and streaming infrastructure.
- Architect a unified data serving layer that consistently delivers trusted data from both batch and streaming sources.
- Implement robust data governance, reliability benchmarks, and observability protocols across the data ecosystem.
Benefits
- Competitive compensation and equity through stock options
- Flexible vacation policy with an organizational emphasis on rest and work-life balance
- Remote-first workplace featuring virtual and in-person events to support team engagement
- Full health, dental, and vision coverage as part of a comprehensive benefits package
- 16 weeks of parental leave available to all new parents
Compensation
Competitive salary and stock options
Work Arrangement
Remote-first
Team
Collaborates with analytics, product, and engineering teams