Responsibilities
- Design and manage ETL workflows using Python and PySpark in Azure Synapse Analytics Notebooks or Pipelines to efficiently process data.
- Construct and maintain data warehouse structures using star schema design, fact and dimension tables, within a Massively Parallel Processing SQL Pool.
- Retrieve data from diverse sources such as REST APIs, SQL database tables, and CSV files.
- Leverage in-depth knowledge of Azure Synapse Analytics to build scalable and high-performing data notebooks and pipelines.
- Support the adoption of Data Fabric components including data lakes, lakehouses, delta lakes, and data cataloging practices.
- Partner with data architects to develop data models and schemas that meet business needs.
- Enforce data accuracy and consistency by implementing validation rules and quality checks.
- Optimize ETL performance by identifying bottlenecks and tuning data pipelines to meet service level agreements.
- Monitor data workflows, troubleshoot failures, and apply fixes to maintain pipeline reliability.
- Keep detailed records of ETL processes, data transformations, and system workflows.
- Collaborate with cross-functional teams to define data needs and support data-driven projects.
- Ensure data handling follows security protocols and complies with governance and privacy regulations.
Work Arrangement
Remote (Worldwide) — Latin America