Cloud Performance Engineering Site Reliability Engineer
Job Summary
The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability scalability and performance of production-grade services deployed across multiple cloud vendors and infrastructure platforms for Smile Digital Health its clients and partners.
This role designs and automates performance testing frameworks integrates them into CI/CD pipelines and uses observability tools to proactively detect and resolve bottlenecks. Working closely with engineering product and security teams the SRE ensures systems meet strict SLAs for performance and availability while driving continuous optimization across multiple cloud platforms.
Responsibilities:
- Collaborate with our Security Operations teams to define and implement best practices around Cloud Service Provider configuration for Azure and other cloud providers.
- Develop implement and coordinate a multi-tenant approach around service offerings for databases container platforms authentication certificates and product registries.
- Design develop and maintain cloud performance testing strategies frameworks and environments to validate application scalability reliability and resiliency.
- Develop and automate load stress spike and endurance (soak) testing as part of CI/CD pipelines.
- Analyze application and infrastructure performance to identify bottlenecks and recommend performance optimizations across cloud-native services.
- Develop and maintain cost and utilization tracking and attribution processes across Cloud Service Providers.
- Create documentation detailing Cloud Service Provider offerings implementation patterns and best practices.
- Develop and maintain technical relationships with our core Cloud Service Providers.
- Implement and maintain secure scalable infrastructure platforms for delivering cloud services.
- Ensure internal and external SLAs are consistently met or exceeded while continuously monitoring and improving system performance reliability and availability.
- Create tools for automating deployment monitoring and platform operations.
- Implement and manage observability solutions (logging metrics tracing) using OpenTelemetry Prometheus Grafana Azure Monitor and related technologies to provide actionable performance insights.
- Plan and execute chaos engineering experiments to evaluate and improve application resiliency and fault tolerance.
Requirements:
- 5 years of experience with Cloud Service Providers and best practices around implementation and configuration preferably managing Azure environments supporting SaaS products.
- Experience working across multiple cloud providers (Azure required; AWS and/or Google Cloud Platform considered an asset).
- Strong experience in Cloud Performance Engineering including performance analysis capacity planning scalability testing and optimization of distributed cloud-native applications.
- Proven experience working with microservices architecture with a strong focus on Java-based services.
- Experience applying Chaos Engineering practices to evaluate and improve system resiliency.
- Strong experience designing and executing performance testing strategies including load stress spike and endurance (soak) testing to validate application scalability and defined latency and error-rate thresholds.
- Hands-on experience with performance testing tools such as JMeter Gatling Azure Load Testing or k6.
- Experience validating application services sustaining 500 transactions per second (TPS) while meeting defined performance objectives.
- Hands-on experience deploying and managing containerized applications using Docker and Kubernetes including autoscaling and performance optimization.
- Experience using Terraform to provision and manage cloud infrastructure using Infrastructure as Code (IaC).
- Experience tuning Kafka (partitioning consumer group sizing throughput/latency trade-offs) and other messaging/queueing platforms to sustain target transaction rates.
- Hands-on experience implementing and using observability platforms including OpenTelemetry Prometheus Grafana Azure Monitor Application Insights and Log Analytics.
- Proven experience with Security and Compliance (SOC 2 HIPAA ISO 27001) best practices and implementing controls that support high-velocity software delivery teams.
Required Experience:
IC
About Company
Built around the visionary HL7 FHIR standard and powered by HAPI, Smile Digital Health is the most proven FHIR implementation in the world.