Responsibilities
- Create reliable, reusable, and configurable data ingestion and transformation workflows using tools such as Azure Data Factory, Synapse Pipelines, Databricks, or Microsoft Fabric Data Factory.
- Apply medallion architecture principles by organizing data into Bronze, Silver, and Gold layers in Azure Data Lake Storage Gen2 using Delta Lake, Parquet, and structured streaming.
- Develop efficient ELT processes that push computation to source systems like Synapse Dedicated SQL Pool, Azure SQL, or Teradata when beneficial.
- Write and optimize PySpark scripts and batch jobs using Azure Databricks or Synapse Spark for large-scale data processing.
- Construct dimensional data models using Kimball methodologies, including star and snowflake schemas, and implement data vault structures for analytical reporting.
- Apply techniques for managing evolving data, including Slowly Changing Dimensions (Types 1, 2, and 3), Change Data Capture, and handling delayed data arrivals.
- Optimize SQL-based workloads in Synapse Dedicated SQL Pool or Fabric Warehouse through proper use of distribution keys, partitioning, and columnstore indexing.
- Establish continuous integration and continuous delivery pipelines using Azure DevOps with YAML, along with infrastructure-as-code tools like ARM, Bicep, or Terraform across multiple environments.
- Integrate comprehensive logging, auditing, and monitoring into data workflows using Azure Monitor, Log Analytics, and Kusto Query Language (KQL).
- Enforce coding standards, peer review processes, source control branching models, and deployment procedures.
- Support or lead migration initiatives from legacy platforms such as Informatica PowerCenter to Azure Data Factory, or from on-premises databases like Teradata, Oracle, or SQL Server to cloud-based solutions.
- Conduct performance analysis, resource planning, and cost estimation for future-state data architectures.
- Respond to production issues affecting critical data pipelines and ensure timely resolution.
Compensation
Competitive salary and benefits package commensurate with experience
Work Arrangement
Hybrid or remote options available; based on project and team requirements
Team
Collaborative data engineering team focused on cloud transformation and modern data platform development
Responsibilities
- Design and build robust, reusable, parameter-driven ingestion and transformation pipeline using Azure Data Factory, Synapse Pipelines, Data Bricks and/or Microsoft Fabric Data Factory.
- Implement medallion architecture (Bronze / Silver / Gold) on Azure Data Lake Storage Gen2 using Delta Lake, Parquet, and structured streaming patterns.
- Build performant ELT workflows that leverage pushdown to source systems (Synapse Dedicated SQL Pool, Azure SQL, Teradata) where appropriate.
- Develop and optimize PySpark notebooks and jobs on Azure Databricks or Synapse Spark.
- Design dimensional models (Kimball star/snowflake) and data vault patterns for analytics consumption.
- Implement Slowly Changing Dimensions (Type 1/2/3), Change Data Capture, and late-arriving data patterns.
- Tune distributed SQL workloads in Synapse Dedicated SQL Pool / Fabric Warehouse, including distribution keys, partitioning, and clustered column store indexes.
- Implement CI/CD for data pipelines using Azure DevOps (YAML pipelines, ARM/Bicep/Terraform) across Dev / SIT / UAT / Prod environments.
- Instrument pipelines with robust logging, auditing, and monitoring using Azure Monitor, Log Analytics, and KQL.
- Define and enforce coding standards, code review practices, branching strategies, and release management.
- Lead or contribute to legacy-to-cloud migrations — e.g., Informatica PowerCenter to Azure Data Factory, on-premises Teradata / Oracle / SQL Server to Synapse or Fabric.
- Perform workload assessment, capacity planning, and cost modeling for target-state architectures.
- Production incident response for critical pipelines.
Available for qualified candidates requiring sponsorship