Enter a job title or keyword

Senior Production Engineer Realtime Products

Databricks


Job Location:

Mountain View, CA - USA

Monthly Salary: Not provided by the employer
Posted: 9 October 2026 (11 hours ago)
Application Deadline: 6 January 2027
Vacancies: 1 Vacancy

Job Summary

CSQ127R154

As a Senior Production Engineer working on Databricks realtime products you will be directly contributing to our customers success. You will build advanced monitoring and incident mitigation tooling and drive changes across the stack to proactively make the realtime products like Lakebase Neon and Model Serving reliable secure and scalable in production.

Our production engineers understand the Databricks platform from end to end and partner across engineering teams building durable solutions for a platform operating across AWS Azure and GCP.

The Impact Youll Have

Observability

  • Build advanced monitoring and detection capabilities that give early warning of issues and deep insights into workload performance and customer experience.
  • Build reliable observable automation for debugging and incident mitigation.

Reliability

  • Improve service reliability scalability security and operational efficiency.
  • Develop dependable safe mitigations for production issues.

On-Call & Incident Response

  • Participate in a follow-the-sun on-call rotation and lead incident response and mitigation.
  • Perform root-cause analysis identify and implement lasting corrective actions.
  • Partner with Product Engineering Security Support and other infrastructure teams on follow up actions.
What We Look For

Experience

  • 5 years of experience in Production Engineering Customer Reliability Engineering (CRE) Site Reliability Engineering (SRE) infrastructure engineering backend software engineering or a related field.
  • Experience in holistic monitoring and alerting of complex stateful systems applying multiple strategies like workload alerting anomaly detection and probing.
  • Experience with PostgreSQL or related managed databases or distributed systems.
  • Experience in incident management and participating in on-call rotations for critical infrastructure.

Skillset

  • A mindset focused on automation root-cause resolution and continuous improvement.
  • Strong programming skills in one or more languages such as Python Go Java Scala or similar.
  • Proficiency with infrastructure automation and Infrastructure as Code.
  • Ability to work across system boundaries and collaborate effectively during complex incidents and engage with customers infrastructure teams.
  • Experience in using AI to address production and operational challenges.

Bonus

  • Experience with AWS Azure or GCP.
  • Experience with Lakebase or Neon.
  • Experience with Kubernetes Terraform.
  • Experience building internal platforms operational tooling or developer productivity systems.

Education

  • BS degree (or higher) in Computer Science or a related field.

Required Experience:

Senior IC


About Company

Company Logo

The Databricks Platform is the world’s first data intelligence platform powered by generative AI. Infuse AI into every facet of your business.

View Profile View Profile