Description
The Research Crawling Engineer will design and operate large-scale web data acquisition systems for research and model development. Responsibilities include building distributed crawlers, handling anti-bot systems and rate limits, developing data-cleaning and normalization pipelines, constructing datasets, monitoring crawl quality, and optimizing infrastructure for cost, latency, and reliability. The role requires programming experience in Go, Rust, Python, Java, or C++, along with experience in web crawling or large-scale data pipelines, and familiarity with HTTP, networking, browser behavior, distributed systems, and parallel processing. The position is fully remote and requires a work schedule overlapping with EST business hours.
