Enter a job title or keyword

Research Crawling Engineer


Job Location:

Los Angeles, CA - USA

Monthly Salary: Not provided by the employer
Posted: 24 September 2026 (20 hours ago)
Application Deadline: 22 December 2026
Vacancies: 1 Vacancy

Job Summary

About the Role

This is a hands-on engineering role focused on building and operating large-scale web crawlers and data acquisition systems that power dataset creation for frontier AI model training. You will work as an extension of AI research teams helping to collect clean and curate web data at a scale that few organizations can match. The work directly shapes the pretraining and inference pipelines used by some of the most advanced AI labs in the world.

What Youll Do
  • Build and maintain large-scale web crawlers across diverse domains including social media travel and multi-language sites.
  • Design high-throughput fault-tolerant data collection systems capable of handling millions to billions of URLs per day.
  • Navigate anti-bot systems rate limits and JavaScript-heavy sites finding creative solutions when standard protocols fall short.
  • Develop pipelines for cleaning deduplication filtering and normalization of web data at TB to PB scale.
  • Construct and maintain datasets for research and model training in close collaboration with research teams.
  • Monitor crawl performance coverage and data quality iterating quickly as web environments change.
  • Optimize infrastructure for cost latency and reliability across cloud and bare-metal environments.
What Were Looking For
  • 3 or more years building and operating web crawlers at scale (2 or more years considered for candidates with a PhD).
  • Proficiency in one or more of: Go Rust Python Java or C.
  • Experience running data pipelines at TB or greater scale.
  • Deep knowledge of HTTP networking and browser behavior.
  • Hands-on experience with distributed systems or parallel processing.
  • Experience with headless browsers such as Playwright Puppeteer or Chrome DevTools Protocol.
  • Familiarity with proxy systems IP rotation or request orchestration.
  • Experience with data quality evaluation scoring or benchmarking at scale.
  • Experience running crawling or data workloads on cloud platforms (AWS GCP) or bare-metal infrastructure.
  • Background in NLP pipelines ML dataset curation or AI lab work is a strong plus.
Compensation & Benefits

Salary range: $160000 to $250000 USD annually. Visa sponsorship is not available for this role.

Location

This role is fully remote. The primary location is Los Angeles CA United States though candidates based in other major US cities are welcome.