Responsibilities
- Design and manage ETL workflows using Python and PySpark in Azure Synapse Analytics Notebooks and Pipelines to support reliable data movement and transformation.
- Apply knowledge of data warehouse design, including star schemas, fact tables, and dimension tables, to develop efficient storage models in MPP SQL Pools.
- Pull data from diverse sources such as REST APIs, SQL databases, and CSV files for integration into analytics platforms.
- Leverage in-depth experience with Azure Synapse Analytics to build high-performance data notebooks and pipelines that scale effectively.
- Support adoption of modern data fabric components including data lakes, lakehouses, delta lakes, and data cataloging to improve data accessibility and governance.
- Partner with data architects to define data models and schemas that reflect business needs and reporting requirements.
- Establish data validation rules and quality checks to ensure accuracy and consistency across datasets.
- Detect and fix performance issues in ETL processes to meet service level agreements and maintain system efficiency.
- Monitor execution of ETL jobs, troubleshoot failures, and apply fixes to maintain pipeline stability.
- Keep detailed records of data engineering workflows, transformation logic, and pipeline architecture.
- Collaborate with cross-functional teams to gather data needs and support analytics projects.
- Enforce data security practices and adhere to governance and privacy regulations across all data systems.
Compensation
Not specified
Work Arrangement
Remote, Latin America
Team
Distributed team environment with collaboration across technical and business units
Responsibilities
- Develop and maintain ETL data engineering processes using Python (PySpark) within Azure Synapse Analytics Notebooks, and/or Azure Synapse Analytics Pipelines, to ensure efficient data extractions, transformation, and loading.
- Apply expertise in data warehousing, understanding star schemas, facts, and dimensions, to design and build effective data storage structures in a Massively Parallel Processing (MPP) SWL Pool.
- Extract data from various sources, including REST APIs, SWL database tables, and CSV files.
- Utilize deep knowledge of Azure Synapse Analytics to design and optimize data notebooks/pipelines for scalability and performance.
- Contribute to the implementation and understanding of other Data Fabric concepts, such as data lakes, lakehouses, delta lakes, and data cataloging, to enhance data management capabilities.
- Collaborate with data architects to create data models and schemas that align with business requirements.
- Implement data quality checks and validation processes to maintain data accuracy and consistency.
- Identify and resolve performance bottlenecks and optimize ETL data notebooks/pipelines to meet SLAs.
- Monitoring ETL jobs, diagnose issues, and implement solutions to ensure data pipeline reliability.
- Maintain comprehensive documentation of ETL data engineering processes, data flows, and data transformations.
- Work closely with cross-functional teams to understand data requirements and provide support for data-related initiatives.
- Ensure data security and compliance with data governance and privacy standards.
Not applicable