Responsibilities
- Create, scale, and refine ETL and ELT data pipelines on Databricks using PySpark, Spark SQL, and Delta Lake technologies
- Apply Medallion architecture principles with layered data zones, enforcing data quality, schema consistency, and evolution
- Develop declarative data workflows using Delta Live Tables and ingest real-time or incremental data via Auto Loader and Structured Streaming
- Manage and supervise production data jobs using Databricks Workflows, integrating with orchestration tools such as Airflow or Azure Data Factory
- Set up and manage Unity Catalog to enforce data governance, including catalogs, schemas, access policies, data lineage, and PII protection
- Work with data science teams to produce cleaned, feature-rich datasets and assist in deploying models using MLflow
- Optimize compute resources by adjusting cluster settings, job structures, and leveraging Photon or serverless options for efficiency
- Develop and maintain automated CI/CD pipelines for Databricks assets, including notebooks, jobs, and bundles using Git, Azure DevOps, or GitHub Actions
- Analyze and evaluate large, heterogeneous datasets from diverse source systems to ensure accuracy and usability
- Collaborate with architects, solution leads, and project managers on technical design and system architecture choices
- Engage directly with clients to gather requirements, review solutions, and explain technical decisions in business terms
- Produce clear documentation, including architecture diagrams, data flow maps, code annotations, and operational runbooks
Compensation
Competitive market rate
Work Arrangement
Flexible; may include remote or hybrid options
Team
Collaborative environment with data engineers, architects, and client teams
Responsibilities
- Design, build, and optimize scalable ETL/ELT pipelines on Databricks using PySpark, Spark SQL, and Delta Lake
- Implement Medallion (Bronze/Silver/Gold) architecture patterns, applying data quality checks, schema evolution, and enforcement along the way
- Build declarative pipelines with Delta Live Tables (DLT) and ingest streaming/incremental data using Auto Loader and Structured Streaming
- Orchestrate and monitor production workloads using Databricks Workflows, integrating with tools like Airflow or Azure Data Factory where needed
- Configure and maintain Unity Catalog for data governance — catalogs, schemas, access controls, lineage, and PII masking
- Partner with data scientists to prepare feature-engineered, ML-ready datasets and support model deployment workflows using MLflow
- Tune cluster configuration, job design, and Photon/serverless compute for performance and cost efficiency
- Build and maintain CI/CD pipelines for Databricks notebooks, jobs, and asset bundles (Git-based workflows, Azure DevOps, GitHub Actions, or similar)
- Query, profile, and assess the quality of large, complex datasets from a wide variety of source systems
- Collaborate with solution leads, architects, and project managers on solution design and technical architecture decisions
- Participate directly in client-facing work: requirements gathering, solution reviews, and translating technical tradeoffs into plain-language business impact
- Document solutions clearly — architecture diagrams, data flow documentation, code comments, and runbooks
Not specified