Senior ML Engineer TT
Job Summary
Were hiring a Senior ML Engineer to own the design training and deployment of a novel foundation model from research through production including the custom CUDA kernels that make it fast. This is a hands-on high-ownership role for someone who has already shipped a large-scale foundation model (01) at a high-growth AI/ML startup or top-tier research lab not just published about one. Youll architect and scale distributed training and inference pipelines on cloud infrastructure profile and optimize deep learning models at the systems level and build the internal tooling that lets a small fast-moving team punch above its size. The environment is early-stage high-transparency and high-urgency decisions move quickly ambiguity is the norm and the team expects people to challenge and be challenged. Youll work closely with the founders against real production SLOs and SLAs rather than research benchmarks. If you want to be the hands-on technical owner of a first-of-its-kind product rather than one contributor among many this role is built for you.
Details
- Schedule: Full-time
- Location: UK London
- Start: ASAP
- Duration: Long-term
- English: Fluent
- Type of collaboration: B2B
About the project
The client is a VC-backed AI/ML startup building a novel foundation model that enables fully automated unsupervised software delivery for embedded control systems. Its an early-stage company at a critical growth point scaling its technical team to deliver a high-impact first-of-its-kind product. The culture is direct high-transparency and no-jargon the team values honesty urgency and strong work ethic over process and hierarchy. Technical leadership is hands-on and expects the same from every hire: this is not a role for someone who wants to hand off hard problems to others. The company operates with real ambiguity and rapid change and rewards people who take ownership and move fast. Candidates should be excited by the prospect of building something genuinely novel from the ground up not maintaining an existing system.
You have
- Shipped a large-scale foundation model (01) at a high-growth AI/ML startup or top-tier research lab hands-on delivery not purely academic or research-only experience
- Designed and implemented custom CUDA kernels for model optimization with strong proficiency in both CUDA C/C and Python
- Direct experience scaling distributed training and/or inference pipelines on cloud infrastructure (AWS Azure or GCP)
- Deep knowledge of at least one major deep learning framework ideally PyTorch and hands-on experience with recent architectures (e.g. MoE state-space models)
- Hands-on ownership of ML systems with strict SLOs or production SLAs youve operated systems in production not just built models
- A track record of building internal tooling or infrastructure that measurably accelerated a teams productivity
- Demonstrated ability to deliver quickly in ambiguous fast-paced early-stage environments
- Fluent English
What to do
- Architect and implement a large-scale foundation model from research through production deployment
- Design and write custom CUDA kernels to optimize model performance where off-the-shelf libraries fall short
- Build and scale distributed training and inference pipelines on cloud infrastructure
- Profile debug and optimize deep learning models for latency throughput and reliability against production SLAs
- Build internal tooling and infrastructure to accelerate the teams iteration speed
- Work directly with the founders to make fast high-ownership technical decisions in an ambiguous high-urgency environment
- Take hands-on technical leadership as the team scales helping set technical direction for the model and its infrastructure
Interview Process
- 1-hour online cultural interview with the CEO (focus: values urgency transparency team fit)
- 2-hour technical interview with the CPO deep technical deep-dive hands-on problem-solving system design. No live or take-home coding tasks.