Staff Software Engineer- Foundation Model Inference
San Francisco, CA - USA
Job Summary
P-1930
At Databricks we are passionate about enabling data and AI teams to solve the worlds toughest problems from making the next mode of transportation a reality to accelerating the development of medical breakthroughs. We do this by building and running the worlds best data and AI infrastructure platform so our customers can use deep data insights to improve their business. Founded by engineers and customer-obsessed we leap at every opportunity to solve technical challenges from designing next-gen UI/UX for interfacing with data to scaling our services and infrastructure across millions of virtual machines. And were only getting started.
As part of the AI team youll build the platforms and products that power everything from data apps AI agents model training model serving and Vector Search. Youll be joining a high-agency high-visibility team operating at the frontier of AI infrastructure with deep ties to research product and real-world enterprise use cases. Databricks Mosaic AI is one of our fastest-growing businesses helping thousands of our customers democratize AI within their organizations. Were building the products and infrastructure that power the next generation of AI.
The Foundation Model Inference team is the backbone of Databricks generative AI capabilities. We build the infrastructure that enables our customers to serve scale and optimize frontier models with enterprise-grade reliability and performance. Our Foundation Model APIs provide a unified platform that gives customers access to LLMs with the governance flexibility and scalability required for enterprise production workloads.
We are looking for high-agency engineers who are excited to work on powering model inference at enterprise scale.
- Build LLM infrastructure powering large-scale inference workloads for customers through partner models (OpenAI Anthropic Gemini) and self-hosted models (Qwen GPT-OSS Llama)
- Improve reliability latency and efficiency of distributed AI workloads
- Collaborate with platform infra and ML teams to deliver seamless end-to-end experiences
- Shape how developers and data scientists build and interact with AI on Databricks
- 8 years of experience in backend or infrastructure engineering
- Experience with distributed systems scalable APIs or cloud-native infrastructure
- Experience with real-time serving ML infrastructure or GPU orchestration
- Familiarity with service-oriented architecture deployment pipelines and system observability
- Exposure to platforms like SageMaker Vertex AI or Azure ML
- Contributions to OSS projects like MLflow PyTorch Ray vLLM SGLang
- Built developer platforms or internal tools supporting AI workflows
Required Experience:
Staff IC
About Company
The Databricks Platform is the world’s first data intelligence platform powered by generative AI. Infuse AI into every facet of your business.