Skip to main content

Data Curation Intern at Karya

Department: TechnologyEducation: education_optional

Location
Bengaluru, Karnataka
Type
Internship
Level
intern
Posted

Description

Karya is hiring a Data Curation Intern to build high-quality datasets for AI/ML model training, with a focus on Indian and multilingual data. The role begins with auditing and cleaning large open-source text datasets, applying metadata schemas, and creating quality checklists, then progresses to preparing phonetically diverse text for read-speech and voice model training. Candidates should have strong attention to detail, Python skills for data processing, familiarity with text data formats, and curiosity about AI/ML or language technology; prior NLP dataset experience, Indian-language knowledge, data versioning experience, and basic model-training knowledge are preferred.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation