Senior Platform Reliability Engineer
San Francisco, CA - USA
Job Summary
Grow Therapy is on a mission to serve as the trusted partner for therapists growing their practice and patients accessing high-quality care. Powered by technology we are a three-sided marketplace that empowers providers augments insurance payors and serves patients. Following the mass increase in depression and anxiety the need for accessibility is more important than ever. To make our vision for mental healthcare a reality were building a team of entrepreneurs and mission-driven go-getters. Since launching in February 2021 weve empowered more than ten thousand therapists and hundreds of thousands of clients across the country and insurance landscape. Weve raised more than $328Mm in funding including our Series D at a $3B valuation from Sequoia Capital Transformation Capital TCV SignalFire Menlo Ventures Goldman Sachs Alternatives and others.
Were hiring a Senior Platform Reliability Engineer to help define and scale reliability as a first-class this role youll operate horizontally across the organization shaping how reliability is understood measured and built into the developer experience.
Youll work closely with other members of the platform team as well as our product engineering teams to establish standards around observability SLOs/SLAs and incident responsewhile also helping translate those standards into self-service tooling and golden paths that make it easy for teams to adopt them.
This is a high-impact highly autonomous role where youll drive both cultural and technical change ultimately enabling teams to independently build and operate reliable systems at scale.
What Youll Work On
Youll help us establish and scale reliability as a discipline at Grow by:
Defining Reliability Standards Establishing frameworks for SLOs/SLAs error budgets and operational readiness; helping teams understand what to measure and why it matters.
Improving Observability & Measurement Identifying gaps in metrics logging and tracing; ensuring services are measurable debuggable and aligned with reliability goals.
Evolving Incident Response Developing and improving incident response practices from detection to post-incident learning and helping teams build sustainable on-call and escalation patterns.
Enabling Self-Service Reliability Partnering with the platform team to build tooling and abstractions (e.g. service scorecards dashboards templates golden paths) that make it easy for teams to adopt and stay compliant with reliability standards.
Driving Adoption Across Teams Working cross-functionally to educate influence and guide engineering teamsscaling reliability practices through a combination of clear standards strong communication and developer-friendly systems
Who You Are
Experienced in production systems: You have 6 years of experience operating and improving reliability of production systems at scale.
Strong foundation in cloud and infrastructure: You have hands-on experience with AWS Kubernetes (e.g. EKS) and infrastructure as code tools like Terraform.
Deep understanding of reliability principles: Youve defined or worked with SLOs/SLAs understand error budgets and have experience improving reliability through measurement and iteration.
Observability expertise: Youve worked with modern observability tooling (we use DataDog) and understand how to build actionable monitoring systems across metrics logs and traces.
Systems thinker: Youre able to zoom out identify patterns across teams and services and design solutions that scale beyond a single system.
Impact-oriented: You focus on outcomes over output and care deeply about improving real reliability outcomesnot just adding processes.
Strong communicator and influencer: You can drive change across teams without direct authority balancing pragmatism with long-term vision.
Self-directed: You thrive in ambiguous environments and are comfortable defining problems proposing solutions and executing independently.
Team player: You collaborate well communicate with empathy and enjoy mentoring and learning from others.
Bonus Points
Youve helped introduce or scale reliability practices in a growing organization.
Youve built internal tooling or platforms used by multiple teams.
You have experience designing service-level scorecards or compliance/reporting systems.
Youve worked with both SaaS (e.g. DataDog) and self-managed observability stacks.
You were previously a product engineer and bring empathy for developer experience.
You have experience with database reliability and performance (we use PostgreSQL)
Why This Role Is Exciting
This is a rare opportunity to define what reliability looks like at a growing scaling engineering organizationand to do it in a way that actually sticks.
You wont just be responding to incidents or working within a single team. Youll be shaping how reliability is measured enforced and experienced across the entire company. Youll work alongside your team mates to turn best practices into intuitive self-service systems that engineers rely on every day.
Your work will directly improve system reliability reduce incidents and enable teams to move faster with confidence ultimately making reliability a built-in property of how we build software at Grow.
Employment Type: Full Time Exempt
Base Compensation: The base compensation range for this position is $182000$250000 USD Annually.
This is a hybrid role with the expectation to work onsite from our San Francisco NYC or Seattle hub location three days per week (Tuesday Wednesday and Thursday) and travel 23 times per year (e.g. company and department offsites).
The base compensation for this role will vary depending on several factors including relevant experience qualifications and the candidates working location.
Full Time Employee Benefits:
Health Benefits: Comprehensive medical dental vision life and disability coverage.
Grow for Grow: No cost access to therapy through the Grow platform (available to US employees)
Financial Wellness: Retirement savings programs and equity opportunities to help you invest in your future.
Flexible Time Off Paid Holidays & Winter Break: Flexible time off company paid holidays (which vary by country) and a full company wide Winter Break to rest and recharge
Parental Leave: Up to 18 weeks of paid parental leave to support you and your growing family and a new child stipend.
Mental Health Mornings/Afternoons: Dedicated weekly flexible time for self care whether thats therapy exercise journaling time with family or simply taking a break.
Wellness & Development Stipend: Annual stipend to support your personal wellbeing and professional growth from fitness and education to books and learning resources.
Commuter Benefits: Pre tax commuter benefits to support teammates working from one of our hub locations.
Home Office & Meals: Support for your home workspace plus meal benefits with offerings tailored to both remote and hub based employees.
Additional Perks: A variety of wellbeing benefits including wellness memberships virtual care pet insurance discounts (where available) and global travel assistance.
Research shows that some groups hesitate to apply unless they meet every qualification. If youre excited about this role but dont check every box we encourage you to apply. At Grow we value diverse experiences transferable skills and the unique strengths each person brings.
Grow Therapy is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race color ancestry religion sex national origin sexual orientation age citizenship marital status disability gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories consistent with legal requirements.
Use of AI Tools: We use certain AI and automated tools to support recruitment. These tools help our team review applications organize interviews detect potential application integrity issues and maintain interview records; they do not make final hiring decisions.
Application Review: Ashbys AI-Assisted Application Review helps review application materials against role-specific criteria and may flag potential fraud or integrity issues for recruiter review. All advancement decisions are made by our human recruiting team after independent review. See Ashbys AI Bias Audit Report.
Interview Recording: We use BrightHire to record interviews supporting note-taking and consistency across candidates. Learn more. Interviews are confidential; you may ask your interviewer to stop recording anytime or opt out in advance. Choosing not to be recorded will not by itself affect your candidacy.
AI-Assisted Recruiter Screen: For select roles BrightHires AI features may help structure initial recruiter screens and organize interview notes. All responses are reviewed by our recruiting team; no hiring decisions are made by AI alone.
State-specific notices: Depending on your location you may have additional rights. Illinois: If AI analyzes a video interview well provide notice explain how it works and what it evaluates and obtain consent beforehand. New York City: Where required well provide at least 10 business days notice before using an automated employment decision tool and explain how to request an alternative process. Maryland: We do not use facial recognition to create facial templates during interviews unless required notice and written consent are obtained.
Questions accommodation requests requests for an alternative selection process objections to recording or concerns about AI use Contact Were happy to help or offer an alternative way to participate.
Required Experience:
Senior IC