Description
The Data Engineer will develop and maintain scalable data pipelines for large, diverse healthcare datasets used to train and fine-tune LLMs and machine learning models. The role covers data gathering, cleaning, validation, transformation, structuring, monitoring, and documentation, while collaborating with data scientists and machine learning engineers. Required qualifications include a bachelor's degree, data engineering experience, programming proficiency, data warehousing and ETL expertise, cloud experience, and Apache Spark experience. Healthcare data standards, machine learning, and data orchestration experience are preferred. The position offers remote or hybrid work and requires U.S. citizenship, a Green Card, or a valid H1B visa.

