Skip to main content

ML Engineer (Data), Foundational Models at Sarvam

Department: Models

Setup
On-site
Location
Bengaluru, Karnataka
Level
mid
Posted

Description

Sarvam is hiring a large-scale data infrastructure engineer to build and operate petabyte-scale data pipelines for pre-training and post-training of foundational models. The role covers ingestion, parsing, normalization, filtering, deduplication, tokenization, packing, quality classification, contamination detection, mixture design, curriculum and annealing, data provenance and licensing, and tooling for data analysis and debugging. Candidates should have a BS or MS in computer science or a related field, at least three years of experience building large-scale distributed data systems, hands-on LLM data curation experience, familiarity with distributed processing frameworks, strong Python skills, and meaningful open-source contributions.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation