Description
The role is responsible for designing, implementing, optimizing, and automating batch and real-time data pipelines using Cloudera Hadoop, Hive, Spark, Impala, HBase, and Informatica PowerCenter. It includes ETL and data integration, distributed data processing, data quality and governance controls, workflow orchestration, monitoring, and performance tuning. The position requires a bachelor’s degree in a relevant field and at least six years of experience in the same field, along with strong expertise in the listed data technologies.
