Enter a job title or keyword

Machine Learning Engineer Infrastructure


Job Location:

San Francisco, CA - USA

Monthly Salary: Not provided by the employer
Posted: 24 September 2026 (Yesterday)
Application Deadline: 22 December 2026
Vacancies: 1 Vacancy

Job Summary

Machine Learning Engineer - Infrastructure

Company: Causal Labs
Location: San Francisco CA (South Park office in person 5 days per week; relocation provided)
Compensation: $200000 - $400000 highly competitive early-stage equity
Employment Type: Full-time
Visa Sponsorship: Visa transfers; can sponsor visas

About Causal Labs

Causal Labs is pursuing general causal intelligence: AI that can predict the future and identify the actions that change it. It is building a Large Physics foundation Model (LPM) because domains governed by physics have inherent cause-and-effect structure that visual or textual data lacks. Its starting domain is weather the most observed physical system on earth with rapid ground-truth feedback and data volumes that dwarf LLM training sets.

The founders come from Cruise Google Research and Meta. The company is about 10 people in San Francisco growing to around 35 this year and is backed by Kindred Ventures Refactor and BoxGroup.

The Role

Causal Labs is hiring infrastructure engineers to tackle the unsolved training and inference challenges of a Large Physics foundation Model. The work demands deep expertise in standing up distributed training clusters and optimizing performance for large models. If you have built large-scale ML infrastructure for language vision robotics or biology models and want to bet on a counterintuitive technical thesis this is the role.

What You Will Do
  • Design deploy and maintain large distributed ML training and inference clusters.
  • Build efficient scalable end-to-end pipelines for petabyte-scale datasets and model training across the ML lifecycle.
  • Research and test training approaches including parallelization techniques and numerical-precision trade-offs across model scales.
  • Analyze profile and debug low-level GPU operations to optimize performance.
  • Bring new ideas from current research into the stack.
What You Bring
  • 2-10 years building large-scale ML infrastructure for core foundation models (not fine-tuning)
  • Deep expertise optimizing large-scale training and inference workloads
  • Proficiency with distributed training frameworks (FSDP DeepSpeed)
  • Knowledge of cloud platforms (GCP AWS or Azure) and containers/orchestration (Kubernetes Docker)
  • Experience with distributed task management and scalable model serving architectures
  • Strong grasp of monitoring logging and observability for ML systems
  • Ability to work in person in San Francisco 5 days a week
Nice to Have
  • Experience at a science or physical AI company (self-driving robotics biology climate/weather)
  • Generalist experience across the ML lifecycle
  • Low-level GPU performance optimization and debugging (CUDA JAX)
Interview Process

Initial call (30 min) technical screen onsite day.

Tech Stack

FSDP DeepSpeed NVIDIA GPUs Python C Linux Kubernetes Docker GCP/AWS/Azure


REVENUE: 14% of first-year salary. Est. fee per hire $28K-$56K; 3 seat(s) up to $126K if all filled.

TARGET COMPANIES (suggested (physical AI / weather)): Google DeepMind NVIDIA (Earth-2) Waymo Cruise Aurora Isomorphic Labs Microsoft Research (Aurora weather).

BEST-FIT CANDIDATE: 2-10 yrs; foundation-model training infra (not fine-tuning); FSDP/DeepSpeed distributed clusters; low-level GPU optimization; visa: transfers; can sponsor; location: SF 5 days. Physical AI/science background strongly preferred; avoid multiple <1.5yr tenures and 9-to-5 seekers.