Description
Cartesia is hiring a Research Engineer, Data Infrastructure to build and operate scalable data processing infrastructure for acquiring, ingesting, preprocessing, filtering, deduplicating, augmenting, and curating massive text datasets for pretraining. The role partners with research and infrastructure teams, runs ablation experiments, establishes data-quality standards, and sources external datasets. Candidates should have hands-on experience with ML data infrastructure, dataset versioning, large-scale data loading, and generative-model data systems; experience with Ray, Spark, Kubernetes, or pretraining language models is preferred. The role is based in San Francisco, London, or Bangalore, with visa sponsorship support and fully covered medical, dental, and vision insurance for United States employees.
