Enter a job title or keyword

LLM Systems Engineer


Job Location:

Bellevue, WA - USA

Salary: Not provided by the employer
Posted: 27 September 2026 (14 hours ago)
Application Deadline: 25 December 2026
Vacancies: 1 Vacancy

Job Summary

About the role

Were looking for a motivated LLM Systems Engineer willing to explore new and unconventional inference systems based on emerging hardware.

This role is part engineering part research youll be responsible for searching and prototyping various algorithms suitable for our inference hardware as well as guiding our hardware team on the product definition. The ideal candidate has a proven track of record of pursuing ML systems research and is very familiar with industry-standard LLM inference systems.

This role will be performed on-site from one of our offices in Santa Clara CA or Boston MA.

Essential Duties and Responsibilities

Prototype and optimize emerging ML inference systems.

Develop novel memory models for expandable vRAM.

Write efficient GPU kernels for data movement.

Perform design-space exploration implementation and benchmarking of inference engines both in simulations and on real hardware.

Qualifications

MS or PhD in computer systems ideally with a focus on LLM inference and/or distributed systems.

Prior experience contributing to the core LLM inference infrastructures (vLLM SGLang TensorRT etc.).

Prior experience in accelerator programming (e.g. CUDA JAX/Pallas ROCm).

Advanced computer architectures and performance engineering skills is a big plus.

Compensation & Benefits

Competitive salary commensurate with experience including base salary incentive-based bonus and early stage equity grant.

Comprehensive benefits including health dental vision and life insurance.

Well-equipped sunny offices in Santa Clara CA and Boston MA.

Relocation assistance and visa sponsorship.

Perks include a daily lunch stipend 401k match and more.

A collaborative continuous-learning work environment with smart dedicated colleagues engaged in developing the next generation of architecture for high-performance computing.

The Opportunity

Impact: We are tackling a fundamental challenge at the infrastructure layer: unlocking greater AI capability while dramatically improving efficiency. The work we do here compounds across state-of-the-art AI models systems and real-world applications.

Timing: Joining now means real ownership of the company and meaningful influence over product direction and execution. Youll work from first principles move quickly from insight to execution and see your contributions directly reflected in what we build.

Culture: Youll work alongside a group of people who care deeply about rigor clarity and impact. We value thoughtful disagreement fast learning and intellectual fearlessness. This is a place where strong ideas shine curiosity is encouraged and growth is a daily practice.