Enter a job title or keyword

ML Infrastructure Engineer


Job Location:

Redwood City, CA - USA

Monthly Salary: Not provided by the employer
Posted: 24 September 2026 (10 hours ago)
Application Deadline: 22 December 2026
Vacancies: 1 Vacancy

Job Summary

ML Infrastructure Engineer

Company: Dyna Robotics
Location: Redwood City CA (in office 5 days per week)
Compensation: $220000 - $350000 competitive equity
Employment Type: Full-time
Visa Sponsorship: Visa transfers (OPT H-1B transfer)

About Dyna Robotics

Dyna Robotics builds general-purpose robots powered by a proprietary embodied AI foundation model that generalizes and self-improves across environments with commercial-grade performance. Its affordable intelligent robotic arms are already deployed at customer sites in hospitality and restaurants automating repetitive stationary tasks.

Founded in 2024 by repeat founders who previously built and sold Kaper AI to Instacart with a team from Google DeepMind Meta and Cruise Dyna has about 130 people and has raised $143.5M from investors including NVentures Samsung NEXT Salesforce Ventures First Round Capital and CRV.

The Role

Dyna Robotics is hiring an ML Infrastructure Engineer to own training infrastructure end to end and turn a multi-cloud GPU fleet into a world-class training engine for massive multimodal models. You will be the connective tissue between researchers and compute and your work directly speeds the path from model to deployed robot.

What You Will Do
  • Architect and scale distributed training across large GPU clusters implementing sharding activation checkpointing and memory optimization (ZeRO FSDP).
  • Build researcher-friendly tooling and job scheduling (Kubernetes SLURM) with fast iteration automated retries and failure recovery.
  • Design high-throughput pipelines that ingest terabytes of multimodal robot data (video proprioception 3D signals) so GPUs never starve.
  • Build low-latency inference pipelines for real-time robot control using quantization distillation and compilation (TensorRT Triton).
  • Profile GPU utilization I/O bottlenecks and memory fragmentation to maximize fleet performance.
What You Bring
  • 5-7 years as an infrastructure engineer including leading technical projects in HPC or ML infrastructure
  • Built and maintained ML or data infrastructure on a team with a high talent bar
  • Deep PyTorch experience and hands-on large-scale distributed training (FSDP ZeRO failure recovery)
  • GPU performance optimization and profiling (CUDA NCCL Triton)
  • Genuine interest in robotics and physical AI
  • Ability to work in the Redwood City office 5 days a week
Nice to Have
  • Robotics experience at startups or enterprise teams
  • Early-stage or founding infrastructure hire
  • Multimodal systems (video audio multimedia models)
  • DeepSpeed or Accelerate; model serving optimization and monitoring
Interview Process

Recruiter screen (30 min) system design (45 min) two coding rounds (45 and 30 min) behavioral.

Tech Stack

PyTorch DeepSpeed Accelerate FSDP Kubernetes SLURM GCP AWS TensorRT Triton NCCL Docker Python CUDA


REVENUE: 21.25% of first-year salary. Est. fee per hire $47K-$74K; 4 seat(s) up to $242K if all filled.

TARGET COMPANIES (suggested): Google DeepMind Tesla (Optimus) Physical Intelligence Figure AI Cruise Waymo NVIDIA.

BEST-FIT CANDIDATE: 5-7 yrs; large-scale PyTorch distributed training infra (FSDP/ZeRO); multimodal data pipelines at TB scale; GPU perf profiling; visa: transfers only; location: Redwood City 5 days. Avoid pure ML modelers and inference/DevOps-only profiles; robotics passion strongly preferred.