Research Crawling Engineer
Los Angeles, CA - USA
Job Summary
This is a hands-on engineering role focused on building and operating large-scale web crawlers and data acquisition systems that power dataset creation for frontier AI model training. You will work as an extension of AI research teams helping to collect clean and curate web data at a scale that few organizations can match. The work directly shapes the pretraining and inference pipelines used by some of the most advanced AI labs in the world.
- Build and maintain large-scale web crawlers across diverse domains including social media travel and multi-language sites.
- Design high-throughput fault-tolerant data collection systems capable of handling millions to billions of URLs per day.
- Navigate anti-bot systems rate limits and JavaScript-heavy sites finding creative solutions when standard protocols fall short.
- Develop pipelines for cleaning deduplication filtering and normalization of web data at TB to PB scale.
- Construct and maintain datasets for research and model training in close collaboration with research teams.
- Monitor crawl performance coverage and data quality iterating quickly as web environments change.
- Optimize infrastructure for cost latency and reliability across cloud and bare-metal environments.
- 3 or more years building and operating web crawlers at scale (2 or more years considered for candidates with a PhD).
- Proficiency in one or more of: Go Rust Python Java or C.
- Experience running data pipelines at TB or greater scale.
- Deep knowledge of HTTP networking and browser behavior.
- Hands-on experience with distributed systems or parallel processing.
- Experience with headless browsers such as Playwright Puppeteer or Chrome DevTools Protocol.
- Familiarity with proxy systems IP rotation or request orchestration.
- Experience with data quality evaluation scoring or benchmarking at scale.
- Experience running crawling or data workloads on cloud platforms (AWS GCP) or bare-metal infrastructure.
- Background in NLP pipelines ML dataset curation or AI lab work is a strong plus.
Salary range: $160000 to $250000 USD annually. Visa sponsorship is not available for this role.
This role is fully remote. The primary location is Los Angeles CA United States though candidates based in other major US cities are welcome.