Software Engineer, Compute Infrastructure
Mountain View, CA - USA
Job Summary
- Design build and own backend/platform services that power Gleans runtime infrastructure with a focus on reliability scalability and performance for AI and search workloads.
- Develop and evolve Kubernetesbased runtime primitives (e.g. service orchestration scheduling integrations autoscaling patterns) across our multicloud foundation (GCP AWS Azure).
- Collaborate with platform data and product engineering teams to make it easy and safe to spin up new services and batch workloads with clear golden paths for deployment configuration and runtime operations.
- Drive endtoend improvements in latency resource utilization and cost for core platform services including multitenant runtime environments and experimental AI workloads.
- Implement and harden infrastructureascode patterns observability and guardrails so teams can confidently ship and run services in production (e.g. SLOs dashboards alerts safe rollout/rollback).
- Partner with the Costs and Runtime teams to build shared mechanisms for attribution guardrails and automation that keep our runtime layer efficient as we 5x customers and traffic.
- Participate in an oncall rotation for critical platform services lead incident response when needed and translate learnings into better reliability tooling and documentation.
- Contribute to technical direction for Runtime Infra: help define roadmaps around multitenancy autoscaling capacity/placement and platformized patterns that reduce perteam handholding.
- You are a backend/platform engineer who enjoys working close to the metalwhere application behavior infrastructure and cost all intersectand you are motivated by building shared systems that many teams depend on.
- You have strong distributed systems fundamentals and experience operating highthroughput lowlatency services or batch pipelines in production environments.
- You are comfortable owning systems endtoend: design implementation testing deployment observability and ongoing operations.
- You think in terms of reliability and guardrails: SLOs incident response safe deployment strategies and clear operational runbooks are part of how you build.
- You are pragmatic and executionoriented: you can balance ideal architectures with the constraints of a fastmoving startup and ship iterative improvements.
- You communicate clearly with both infra and product engineers and you like collaborating across teams to understand requirements and translate them into platform capabilities.
- You are excited to work in a multicloud multitenant environment and to help define best practices for running AI workloads efficiently at scale.
- This role is hybrid (4 days a week in our Mountain View office)
By clicking Submit Application I confirm that I have read the Global Data Privacy Notice and the Applicant Arbitration Agreement and I agree to the terms.
Required Experience:
IC
About Company
Glean was created to deliver healthy and fresh foods - made directly from fruits and vegetables - into your hands! Sweet Potato Flour & Pumpkin Flour now available. Our team is made up of a handful of passionate people who have worked in food and agriculture most of their careers - ... View more