Description
Sarvam is hiring a large-scale data infrastructure engineer to build and operate petabyte-scale data pipelines for pre-training and post-training of foundational models. The role covers ingestion, parsing, normalization, filtering, deduplication, tokenization, packing, quality classification, contamination detection, mixture design, curriculum and annealing, data provenance and licensing, and tooling for data analysis and debugging. Candidates should have a BS or MS in computer science or a related field, at least three years of experience building large-scale distributed data systems, hands-on LLM data curation experience, familiarity with distributed processing frameworks, strong Python skills, and meaningful open-source contributions.
