Description
The role builds performant data processing pipelines, manages relational databases, and constructs cloud-portable data lake environments. Responsibilities include designing Python-based ETL pipelines using Medallion architecture, standardizing storage with Parquet, Apache Iceberg, and Delta Lake, integrating DuckDB with Python, architecting PostgreSQL and SQLite databases, and scaling workloads with Databricks or Apache Spark. The requirements call for advanced SQL, PostgreSQL and SQLite expertise, Python and DuckDB pipeline experience, open-storage-format mastery, Medallion architecture experience, and familiarity with Databricks or PySpark.
