Responsibilities
- Build and maintain reliable, scalable ETL/ELT pipelines using modern tools and best practices, ensuring efficient data flow for analytics and insights.
- Design and implement effective data models that support business needs, enabling high-quality reporting and downstream analytics.
- Work closely with data analysts, product managers, and other engineers to understand data requirements and deliver solutions that meet the needs of the business.
- Develop and apply data quality checks, validation frameworks, and monitoring to ensure the consistency, accuracy, and reliability of data.
- Identify and address performance issues in pipelines, queries, and data storage. Suggest and implement optimizations that enhance speed and reliability.
- Follow data security best practices and ensure pipelines are built to meet data privacy and compliance standards.
- Test new tools and approaches by building Proof of Concepts (PoCs) and conducting performance benchmarks to find the best solutions.
- Contribute to the development of robust CI/CD pipelines (GitLab CI or similar) for data workflows, supporting automated testing and deployment.
Requirements
- 4+ years of experience in data engineering or backend development, with a strong focus on building production-grade data pipelines.
- 2-3+ years of experience working with AWS services (Administration of Redshift is a must)
- Solid experience working with AWS services (Spectrum, S3, RDS, Glue, Lambda, Kinesis, SQS).
- Proficient in Python and SQL for data transformation and automation.
- Experience with dbt for data modeling and transformation.
- Good understanding of streaming architectures and micro-batching for real-time data needs.
- Experience with CI/CD pipelines for data workflows (preferably GitLab CI).
- Familiarity with event schema validation tools/ solutions (Snowplow, Schema Registry).
- Excellent communication and collaboration skills.
- Strong problem-solving skills—able to dig into data issues, propose solutions, and deliver clean, reliable outcomes.
- A growth mindset—enthusiastic about learning new tools, sharing knowledge, and improving team practices.
Nice to Have
- Experience with additional AWS services: EMR, EKS, Athena, EC2.
- Hands-on knowledge of alternative data warehouses like Snowflake or others.
- Experience with PySpark for big data processing.
- Familiarity with event data collection tools (Snowplow, Rudderstack, etc.).
- Interest in or exposure to customer data platforms (CDPs) and real-time data workflows.
Benefits
- Grow Together: Join a culture that champions both personal and professional growth.
- Lead by Example: Every team member is empowered to inspire and make an impact.
- Results-Driven: Focus on achieving meaningful outcomes.
- We Are Well-Makers: Be part of a movement creating a healthier, happier world.