Responsibilities
- Design, build and maintain ETL/ELT processes using SQL, PySpark and Python, including integration of data from different source systems.
- Develop data layers in Databricks, Data Lake and Delta Lake, including tables, views and data models used by analytical, reporting and automation solutions.
- Automate and orchestrate data loading and transformation processes to reduce manual work, improve repeatability and lower the risk of errors.
- Ensure data quality, consistency and reliability through validation rules, monitoring, alerting and incident diagnosis.
- Optimize SQL queries, Spark processes and data storage structures with a focus on performance, stability, scalability and processing costs.
- Provide reliable, ready-to-use data to analysts, Product Owners and other stakeholders as a foundation for analysis, reporting and automation
- Create and maintain technical documentation in Confluence, covering data processes, models, KPI logic, dependencies, data lineage and incident-handling procedures
- Use AI tools consciously as a work accelerator, while fully verifying generated code, configurations and documentation before implementation.
Requirements
- SQL
- PySpark
- Python
- ETL/ELT processes
- data integration from different source systems
- data quality assurance
- validation rules
- monitoring and alerting
- incident diagnosis
- technical documentation
- critical approach to AI-supported work
Nice to Have
- experience with Databricks
- Data Lake
- Delta Lake
- data modeling
- orchestration of data processes
- SQL query optimization
- Spark process optimization
- data storage optimization
- Confluence