Requirements
- Minimum of five years of hands-on experience using SQL for data manipulation and querying.
- At least five years of experience in designing and deploying scalable ETL workflows, including tools for data integration and quality assurance.
- Three or more years working with modern big data analytics, including data lakes, Spark, and columnar storage formats such as Parquet.
- Two or more years building data systems hosted in cloud environments, with strong preference for Azure platforms.
Nice to Have
- Experience developing data pipelines using Azure Databricks, Fabric, or Spark.
- Familiarity with working with data stored in Delta Lake format and querying data from Azure Data Explorer or Kusto.
- Application of artificial intelligence and machine learning techniques to data engineering tasks, including feature engineering, feature stores, and model data pipelines.
- Background in preparing, managing, and securing datasets for use in modern AI applications such as large language models, retrieval-augmented generation, A-B testing, and privacy-sensitive access controls.