Responsibilities
- Design and deploy data architectures on large-scale Hadoop platforms using PySpark.
- Create advanced and optimized SQL queries to retrieve and manipulate data from diverse sources.
- Construct and manage robust, scalable data workflows using Apache Airflow and Kafka.
- Build and fine-tune data processing tasks on Microsoft Azure, utilizing services including Azure Data Lake, Databricks, and Azure Data Factory.
- Partner with data analysts and stakeholders to define data needs and deliver accurate data outputs.
- Maintain consistency and reliability of data across all platforms and systems.
- Oversee, diagnose, and enhance the efficiency of data infrastructure components.
- Keep current with emerging tools and advancements in big data technologies.
Work Arrangement
Remote (Country) — India