Description
Karya is hiring a Data Curation Intern to build high-quality datasets for AI/ML model training, with a focus on Indian and multilingual data. The role begins with auditing and cleaning large open-source text datasets, applying metadata schemas, and creating quality checklists, then progresses to preparing phonetically diverse text for read-speech and voice model training. Candidates should have strong attention to detail, Python skills for data processing, familiarity with text data formats, and curiosity about AI/ML or language technology; prior NLP dataset experience, Indian-language knowledge, data versioning experience, and basic model-training knowledge are preferred.
