Enter a job title or keyword

Staff Software Engineer- Foundation Model Inference

Databricks


Job Location:

San Francisco, CA - USA

Monthly Salary: Not provided by the employer
Posted: 5 August 2026 (30+ days ago)
Application Deadline: 2 November 2026
Vacancies: 1 Vacancy

Job Summary

P-1930

At Databricks we are passionate about enabling data and AI teams to solve the worlds toughest problems from making the next mode of transportation a reality to accelerating the development of medical breakthroughs. We do this by building and running the worlds best data and AI infrastructure platform so our customers can use deep data insights to improve their business. Founded by engineers and customer-obsessed we leap at every opportunity to solve technical challenges from designing next-gen UI/UX for interfacing with data to scaling our services and infrastructure across millions of virtual machines. And were only getting started.

As part of the AI team youll build the platforms and products that power everything from data apps AI agents model training model serving and Vector Search. Youll be joining a high-agency high-visibility team operating at the frontier of AI infrastructure with deep ties to research product and real-world enterprise use cases. Databricks Mosaic AI is one of our fastest-growing businesses helping thousands of our customers democratize AI within their organizations. Were building the products and infrastructure that power the next generation of AI.

The Foundation Model Inference team is the backbone of Databricks generative AI capabilities. We build the infrastructure that enables our customers to serve scale and optimize frontier models with enterprise-grade reliability and performance. Our Foundation Model APIs provide a unified platform that gives customers access to LLMs with the governance flexibility and scalability required for enterprise production workloads.

We are looking for high-agency engineers who are excited to work on powering model inference at enterprise scale.

The impact you will have:
  • Build LLM infrastructure powering large-scale inference workloads for customers through partner models (OpenAI Anthropic Gemini) and self-hosted models (Qwen GPT-OSS Llama)
  • Improve reliability latency and efficiency of distributed AI workloads
  • Collaborate with platform infra and ML teams to deliver seamless end-to-end experiences
  • Shape how developers and data scientists build and interact with AI on Databricks
What we look for:
  • 8 years of experience in backend or infrastructure engineering
  • Experience with distributed systems scalable APIs or cloud-native infrastructure
  • Experience with real-time serving ML infrastructure or GPU orchestration
  • Familiarity with service-oriented architecture deployment pipelines and system observability
Bonus points for:
  • Exposure to platforms like SageMaker Vertex AI or Azure ML
  • Contributions to OSS projects like MLflow PyTorch Ray vLLM SGLang
  • Built developer platforms or internal tools supporting AI workflows

Required Experience:

Staff IC


About Company

Company Logo

The Databricks Platform is the world’s first data intelligence platform powered by generative AI. Infuse AI into every facet of your business.

View Profile View Profile