Job Description

Strong proficiency in Python for data manipulation, scripting, and application development.Experience in PySpark for developing high-performance data transformation jobs on large-scale datasets.Experience with Apache Kafka for building real-time data streaming pipelines, including Kafka Connect and Kafka Streams (or similar stream processing frameworks like Flink/Spark Streaming).Understanding and hands-on experience with the Hadoop ecosystem, including HDFS, YARN, Hive, and other related technologies.Experience with SQL and relational databases (e.g., PostgreSQL, SQL Server) for data extraction and querying.Knowledge of data warehousing concepts and dimensional modeling.Familiarity with version control systems (e.g., Git, GitHub)