Enter a job title or keyword

Member of Technical Staff Training Infrastructure


Job Location:

San Francisco, CA - USA

Monthly Salary: Not provided by the employer
Posted: 29 September 2026 (Yesterday)
Application Deadline: 27 December 2026
Vacancies: 1 Vacancy

Job Summary

About the Role

This is a founding training infrastructure role at an early-stage robotics AI company pretraining a large-scale foundation model on 15PB of video and expert demonstration data. You will own the compute the research team runs on building GPU and data clusters largely from scratch and your architectural decisions will directly set the pace of model iteration as runs scale to hundreds of GPUs.

What Youll Do
  • Own distributed training end-to-end: parallelism strategy multi-node performance scaling efficiency and GPU capacity from cloud providers.
  • Build the full data path from storage to GPU including high-throughput loaders pre-encoding pipelines and sampling infrastructure.
  • Design fault tolerance for long-running training jobs: checkpointing checkpoint durability and automatic failure detection and restart.
  • Build an evaluation harness that runs automatically against every checkpoint including probes calibration checks and research-facing dashboards.
  • Own observability and reproducibility: experiment tracking alerting pinned environments and profiling across compute networking and storage.
  • Translate research requirements into production-grade infrastructure that keeps experimentation fast and large runs routine.
What Were Looking For
  • 3 years of hands-on experience building or operating distributed training infrastructure for large-scale model pretraining.
  • Demonstrated expertise across the full training infrastructure stack: distributed training storage-to-GPU data paths fault tolerance evaluation systems and observability.
  • Experience building or operating distributed training systems across multi-GPU or multi-node setups at 100 GPU scale.
  • Hands-on experience with large-scale video or world models such as vision-language-action models image-to-video models or robot action policies.
  • Production experience with PyTorch or JAX and strong Python fundamentals.
  • Prior GPU infrastructure optimization work on pretraining runs.
  • Experience in early-stage startup environments or sole ownership of training infrastructure systems is a strong plus.
  • Publications at top-tier ML conferences (ICML ICLR NeurIPS) are a plus.
  • Must be eligible to work in the United States without visa sponsorship.
Compensation & Benefits

Salary range: $200000 to $375000 USD annually. No visa sponsorship is available.

Location

On-site in San Francisco California United States.