Enter a job title or keyword

Member of Technical Staff Model Optimization and Inference


Job Location:

Seattle, OR - USA

Monthly Salary: Not provided by the employer
Posted: 5 October 2026 (13 hours ago)
Application Deadline: 2 January 2027
Vacancies: 1 Vacancy

Job Summary

Member of Technical Staff - Model Optimization and Inference

Company: Nuance Labs
Location: Seattle WA (in office 5 days per week; relocation assistance available)
Compensation: $250000 - $350000 equity
Employment Type: Full-time
Visa Sponsorship: Visa sponsorship available

About Nuance Labs

Nuance Labs builds photorealistic real-time AI avatars with emotional intelligence: a full-duplex audiovisual system that can listen speak react and respond like a real person. The research team includes PhDs from MIT UW Oxford CMU and Johns Hopkins.

Founded in 2024 Nuance Labs has raised $60M and has about 25 people.

The Role

Nuance Labs is hiring an experienced ML Infrastructure/Systems Engineer (2 years) to own end-to-end inference optimization across LLMs audio models and diffusion components with a focus on latency throughput and cost.

What You Will Do
  • Own end-to-end inference optimization across the model stack.
  • Implement and tune KV cache strategies for long-context conversations.
  • Evaluate deploy and extend serving frameworks such as vLLM SGLang and TensorRT-LLM.
  • Profile and benchmark latency and throughput and remove bottlenecks.
  • Accelerate diffusion inference and apply quantization techniques (INT8 INT4 GPTQ AWQ).
  • Build internal tooling that makes optimization work faster and more rigorous.
What You Bring
  • 2 years building and maintaining production ML systems
  • Designing scalable infrastructure from scratch
  • Track record optimizing latency throughput and cost
  • Debugging distributed systems
Nice to Have
  • Video or audio model experience
  • CUDA kernels and low-level optimization
  • Real-time video streaming (WebRTC)
Benefits

HSA with about $2000 annual company contribution 15 days PTO plus public holidays and a company-wide office closure week.

Tech Stack

Kubernetes Terraform Python Rust Go Dagster Ray Airflow WebRTC vLLM Triton Inference Server TensorRT


REVENUE: 15.74% of first-year salary. Est. fee per hire $39K-$55K; 7 seat(s) up to $331K if all filled.

TARGET COMPANIES (suggested): NVIDIA Together AI Fireworks AI Groq ElevenLabs Deepgram Meta (FAIR/GenAI).

BEST-FIT CANDIDATE: 2 (experienced) yrs; production ML inference systems; latency/throughput optimization; distributed systems debugging; visa: sponsorship available; location: Seattle 5 days. Low competition.