Responsibilities
- Design and manage high-performance, scalable data pipelines that integrate data from diverse sources
- Utilize dbt to transform data across multiple stages including cleansing, conformance, and presentation layers
- Develop custom data connectors using Python and integrate existing data ingestion tools
- Curate and operationalize enterprise data models to support reporting, analytics, and machine learning use cases
- Implement automated, real-time data quality monitoring, validation rules, and alerting systems across pipelines
- Use Dagster as the core platform for orchestrating and monitoring data workflows
- Manage continuous integration and deployment processes using GitHub, GitHub Actions, and Kubernetes
- Leverage zero-copy cloning and container-based, on-demand environments for development
- Enforce role-based access controls in Snowflake and associated data platforms
- Adopt a DataOps mindset with a strong emphasis on code-centric development practices
- Champion data engineering best practices and suggest enhancements to current systems and workflows
- Apply AI-powered coding assistants such as Anthropic Claude and GitHub Copilot to improve development speed and code reliability
- Demonstrate strong analytical thinking and adaptability in dynamic environments
- Communicate clearly and work effectively within cross-functional teams
- Show initiative and a proactive attitude toward learning and technical exploration
Responsibilities
- Design, develop, and maintain scalable and efficient data pipelines from a wide variety of sources
- Use dbt (Data Build Tool) to transform data through the various layers (Cleansed, Conformed, Presentation etc)
- Ability to write custom connectors (Python) and leverage out-of-the-box data-loading tools
- Operationalise enterprise data model by curating appropriate data models to service reporting, analytics, and data science use cases
- Embed real-time, automated data quality checks, validations, exception handling, and alerting across all data pipelines
- Work within Dagster as the primary orchestration and observability platform for all data pipelines
- Manage CI/CD workflows using GitHub, GitHub Actions, and Kubernetes
- Use zero-copy cloning and containerised on-demand development environments
- Implement and review RBAC policies within Snowflake and related platforms
- Embrace DataOps and a code-first approach to all data engineering work
- Identify and promote best practices in data engineering; recommend improvements to existing processes and systems
- Leverage AI-assisted development tools (Anthropic Claude, GitHub Copilot) to accelerate delivery and improve code quality
- Excellent problem-solving skills and ability to embrace change
- Effective communication and collaboration skills
- Natural self-starter, with enthusiasm for learning and research