Enter a job title or keyword

Member of Technical Staff, ML Systems


Job Location:

Menlo Park, CA - USA

Monthly Salary: Not provided by the employer
Posted: 24 September 2026 (21 hours ago)
Application Deadline: 22 December 2026
Vacancies: 1 Vacancy

Job Summary

Member of Technical Staff ML Systems

Company: TensorScale
Location: Menlo Park CA (in office 5 days per week)
Compensation: $180000 - $230000 competitive equity
Employment Type: Full-time
Visa Sponsorship: Visa transfers and new sponsorship (H-1B TN)

About TensorScale

TensorScale builds the training and inference stack for world models. Todays stack was built for language models; TensorScale is rebuilding it for video image and world-model workloads by co-designing low-level GPU kernels distributed systems and the models themselves. Its public benchmarks include MiniMax running roughly ten times faster at half the cost 2K image generation in about four seconds for three cents and LTX video running faster than real time.

TensorScale has five founders covering distributed systems GPU kernel optimization cloud infrastructure and research with prior work at Fireworks AI Meta Google Apple Microsoft Snowflake and Alibaba. It has raised a $10M seed.

The Role

TensorScale is hiring ML systems engineers to own speed and efficiency across its stack: low-level kernels distributed inference engines and multi-node training and serving systems. You report directly to the cofounder and CEO. The work sits below the application layer (kernels runtimes and distributed engines for video and world models) so this is not an agents or RAG role. You will feel at home if you would rather make a video model ten times faster than train one.

What You Will Do
  • Optimize GPU and system performance for training and inference on image video and world-model workloads.
  • Profile and remove bottlenecks at the kernel memory system and cluster level using Nsight and related tooling.
  • Write low-level optimizations in CUDA and Triton on code paths that run in production.
  • Build distributed inference and training engines for diffusion models across multiple GPUs and nodes.
  • Own communication performance: NCCL RDMA over InfiniBand or RoCE and disaggregated serving.
  • Build benchmarking and regression harnesses so performance gains hold in production.
What You Bring
  • Authored core features in an inference or training framework (vLLM SGLang TensorRT-LLM Megatron or similar) not only deployment or integration work
  • 1 year of full-time hands-on ML systems or GPU performance work in your current or most recent role
  • Kernel work on NVIDIA GPUs in CUDA CUTLASS Triton or PTX
  • Inference or training performance experience: GPU kernels runtimes distributed execution
  • Strong ML and CS fundamentals; degree in CS or a related quantitative field
  • Ability to work in the Menlo Park office 5 days a week
Nice to Have
  • Optimized diffusion video image or other multimodal workloads
  • Public verifiable kernel contributions to open-source ML systems
Interview Process

Hiring manager screen with the CEO (30 min) domain deep dive (60 min) system design (60 min) optional onsite.

Tech Stack

CUDA Triton PyTorch Nsight NCCL RDMA


REVENUE: 20% of first-year salary. Est. fee per hire $36K-$46K; 3 seat(s) up to $123K if all filled.

TARGET COMPANIES (suggested): NVIDIA Together AI Fireworks AI Baseten Modal Meta (PyTorch) vLLM/SGLang contributors.

BEST-FIT CANDIDATE: 1 yrs; core contributions to vLLM/SGLang/TRT-LLM/Megatron-class framework; CUDA/Triton kernels on NVIDIA; recent hands-on GPU perf work; visa: transfers new H-1B/TN; location: Menlo Park 5 days. Avoid compiler-only (MLIR/LLVM) engineers who do not write kernels and AMD/FPGA/custom-accelerator-only backgrounds.