Responsibilities
- Develop and maintain ETL data engineering processes using Python (PySpark) within Azure Synapse Analytics Notebooks and Pipelines for efficient data extraction, transformation, and loading.
- Apply data warehousing expertise, including star schemas, facts, and dimensions, to design and build effective data storage structures in a Massively Parallel Processing (MPP) SQL Pool.
- Extract data from various sources such as REST APIs, SQL database tables, and CSV files.
- Utilize deep knowledge of Azure Synapse Analytics to design and optimize data notebooks and pipelines for scalability and performance.
- Contribute to implementing and understanding Data Fabric concepts like data lakes, lakehouses, delta lakes, and data cataloging to enhance data management.
- Collaborate with data architects to create data models and schemas aligned with business requirements.
- Implement data quality checks and validation processes to maintain data accuracy and consistency.
- Identify and resolve performance bottlenecks, optimizing ETL data notebooks and pipelines to meet service level agreements.
- Monitor ETL jobs, diagnose issues, and implement solutions to ensure data pipeline reliability.
- Maintain comprehensive documentation of ETL data engineering processes, data flows, and data transformations.
- Work closely with cross-functional teams to understand data requirements and provide support for data-related initiatives.
- Ensure data security and compliance with data governance and privacy standards.
Compensation
Not specified
Work Arrangement
Remote (Worldwide) — Latin America
Team
Not specified
Not specified