Description
This role involves developing and implementing scalable ETL/ELT data pipelines for batch and streaming processing, utilizing Spark/PySpark, Python, and SQL. The individual will also create data models using Kimball, Data Vault, or Lakehouse principles, transforming and normalizing data from various sources. They will orchestrate data pipelines with Airflow or equivalent tools, supporting data analytics and machine learning pipelines, and ensuring quality through rigorous testing, documentation, and CI/CD practices.
