Description
The Data Engineer will own datasets end to end, from sourcing and securing multimodal physical data to building petabyte-scale batch and streaming pipelines, standardizing and storing the data, and implementing quality metrics and automated QA checks. The role also involves writing technical requirements for external vendors, collaborating with researchers to validate datasets, and ensuring data quality translates into model performance. Candidates should have experience with large-scale data pipelines, QA systems, or evaluation workflows using technologies such as Spark, Ray, or Beam.
.png&sig=GmGZl0hIEC4KeSLZOLIDYB_qjM5Xym0OKo-pzrtoUqQ)
