Responsibilities
- Create and sustain reliable Python applications for data processing, automation, and software development
- Construct efficient data transformation workflows using PySpark on extensive datasets
- Develop and support real-time data pipelines using Apache Kafka, including Kafka Connect and Kafka Streams
- Utilize Hadoop ecosystem components such as HDFS, YARN, and Hive for distributed data processing
- Retrieve, query, and administer data using SQL and relational database systems like PostgreSQL and SQL Server
- Implement data warehousing methodologies and dimensional modeling to enhance data structure and performance
- Oversee source code management and team collaboration through version control tools such as Git and GitHub