Enter a job title or keyword

Member of Technical Staff


Job Location:

San Francisco, CA - USA

Monthly Salary: Not provided by the employer
Posted: 24 September 2026 (14 hours ago)
Application Deadline: 22 December 2026
Vacancies: 1 Vacancy

Job Summary

Member of Technical Staff

Company: Wafer AI
Location: San Francisco CA (FiDi office on-site 5 days per week)
Compensation: $200000 - $.5% - 1% equity
Employment Type: Full-time
Visa Sponsorship: Visa transfers for candidates already in the US (no new H-1B or overseas sponsorship)

About Wafer AI

Wafer is an AI cloud on a mission to maximize intelligence per watt using AI to optimize AI infrastructure. It serves serverless and dedicated inference for open-source LLMs at the best performance per dollar using autonomous agents to write and tune GPU kernels across heterogeneous hardware and is deploying its own hardware in co-located data centers.

Founded in 2025 Wafer went from $0 to $4M in revenue in 8 weeks. It has about 7 people and is backed by Fifty Years Liquid 2 Ventures and prominent AI angel investors.

The Role

Wafer is hiring a Member of Technical Staff (1-6 years) to work close to the hardware on its inference stack. On a small team with massive surface area you will do everything from talking to customers (10-20% of the role) to writing custom GPU kernels for esoteric hardware with full autonomy over how you solve problems and a hand in setting the direction of its inference serving infrastructure.

What You Will Do
  • Ship day-zero support for new open-source models tuned for latency and throughput.
  • Optimize the serving stack: batching KV cache speculative decoding and quantization.
  • Write and tune kernels in CUDA HIP and Triton for NVIDIA AMD TPU Trainium and other accelerators.
  • Design deploy and operate heterogeneous clusters across vendors.
  • Run production inference across a mixed fleet with strong reliability observability and cost per token.
What You Bring
  • 1-6 years of software engineering focused on systems close to the hardware
  • Deep understanding of computer architecture memory hierarchy and OS internals
  • Backend infrastructure or systems work at a top-tier tech company hardware-adjacent company quant firm or strong AI startup
  • CS degree from a top undergraduate program
  • Evidence of exceptionalism (competitions rankings standout projects)
  • Enthusiasm for AI agents and coding tools in your own workflow
  • High EQ and clear communication with customers
  • Ability to work on-site in San Francisco 5 days a week
Nice to Have
  • New grads with strong internships are considered
  • Deploying or operating infrastructure at scale (data centers clusters GPU fleets)
  • ML inference optimization (quantization batching KV cache speculative decoding)
  • CUDA Triton or GPU programming; OS architecture or compilers coursework
Benefits

$1000/month housing stipend for anyone living within half a mile of the office.

Interview Process

Intro screen (20 min) combined founder screen and technical round co-founder screen (20 min) technical interview (45 min) one-day on-site work trial.

Tech Stack

CUDA HIP Triton Python TPU AWS Trainium


REVENUE: 20% of first-year salary. Est. fee per hire $40K-$60K; 5 seat(s) up to $250K if all filled.

TARGET COMPANIES (client (Tesla/NVIDIA/quant named) suggested): NVIDIA Tesla Jane Street Hudson River Trading Together AI Fireworks AI Cerebras.

BEST-FIT CANDIDATE: 1-6 (strong new grads OK) yrs; hardware-adjacent systems/kernel work; top undergrad CS; markers of exceptionalism; visa: transfers in US only; no new H-1B; location: SF 5 days. Avoid full-stack/web-only profiles; production inference and hands-on kernels are the key differentiators.