Senior Infrastructure SRE
Mississauga - Canada
Job Summary
PointClickCare builds cloud platforms that power safer more connected care for millions of patients. Youll design and operate resilient infrastructure services across Azure AWS GCP and on-prem environments driving reliability automation and operational excellence for identity compute storage messaging and shared services with an SRE-first mindset.
- Design and implement highly available infrastructure solutions for compute storage identity messaging and shared services
- Build and maintain Infrastructure as Code using Terraform or Pulumi; establish best practices and standards
- Automate operational workflows to eliminate toil: auto-remediation self-healing systems capacity planning
- Define and track SLIs and SLOs for critical services; manage error budgets and reliability targets
- Participate in the on-call rotation and lead incident response for complex infrastructure issues; conduct blameless post-mortems and drive systemic fixes
- Develop observability strategies: metrics logs distributed tracing alerting frameworks
- Apply AI-assisted tooling to reduce toil and speed up investigation: log analysis alert triage runbook and post-mortem drafting automation scaffolding
- Own reliability for one or more infrastructure domains end to end driving multi-team initiatives with product engineering from problem definition through adoption
- Mentor intermediate SREs; review infrastructure changes; establish operational best practices
- 5 years of hands-on experience operating and designing cloud infrastructure
- Expert-level understanding of the core services of either Azure or AWS
- Working proficiency in at least one additional platform (Azure AWS or GCP)
- Experience designing and supporting production infrastructure that spans multiple cloud platforms
- 3 years of production experience with Infrastructure as Code (Terraform Pulumi CloudFormation)
- Ability to design scalable reusable IaC modules and enforce GitOps workflows
- Experience managing IaC across multiple cloud providers including module design state layout and provider-specific resource differences
- Strong proficiency in at least one programming language (Python Go Bash) for production automation
- Demonstrated ability to write tested maintainable automation and tooling
- Practical application of SRE principles in production environments
- Experience defining and managing SLIs/SLOs error budgets toil metrics
- Track record of improving system reliability (e.g. MTTR reduction availability improvements)
- Demonstrated depth across key infrastructure services:
- Expert-level experience running Kubernetes and containerized workloads in production on managed Kubernetes (AKS EKS or equivalent) plus VM-based compute
- Practical experience operating a service mesh in production (Istio preferred; Linkerd or equivalent)
- Strong proficiency with enterprise identity and SSO: SAML OAuth/OIDC LDAP and cloud IAM
- Working knowledge of an enterprise federation platform such as PingFederate Entra ID Okta or ADFS
- Strong proficiency in storage solutions (object block and file storage; Kubernetes persistent volumes)
- Working knowledge of designing and operating infrastructure in a regulated environment (HIPAA SOC 2 PCI FedRAMP or equivalent)
- Familiarity with audit evidence access controls encryption in transit and at rest and data residency constraints
- Proven track record of measurably reducing operational toil through automation (e.g. ticket volume manual runbook executions hours reclaimed)
- Strong communication and documentation skills; demonstrated ability to influence engineering teams
- 2 years in healthcare technology or highly regulated SaaS environments (HIPAA SOC 2 HITRUST)
- Cloud certifications: Azure Solutions Architect Expert AWS Solutions Architect Professional GCP Professional Cloud Architect or equivalent
- Working experience operating Kubernetes at scale across multiple clusters (CKA or CKAD certification a plus)
- Working knowledge of messaging and event-streaming platforms (Kafka Azure Service Bus Event Hubs SQS Pub/Sub)
- Practical experience building CI/CD pipelines and deployment automation (GitLab GitHub Actions ArgoCD)
- Familiarity with AI-assisted engineering and operations tooling: LLM-based coding assistants AIOps agentic incident investigation
- Judgment about where AI belongs in an operational workflow including safe handling of sensitive data in prompts and reviewing generated changes before they reach production contribution to open-source SRE tools or infrastructure projects (GitHub profile PRs merged)
- Bachelors degree in Computer Science Computer Engineering Information Technology or related technical field
- OR equivalent practical experience with a proven track record in infrastructure and SRE practices evidence of continuous learning and staying current with SRE and cloud-native trends
- Transparent collaboration: We work in the open using OKRs cross-functional retrospectives and public roadmaps so everyone knows priorities and progress
- Blameless culture: We confront problems courageously through structured post-mortems and root cause analyses focusing on systems improvement not individual blame
- Data-driven decisions: We use metrics APM and observability data to make evidence-based choices about reliability and performance investments
- Continuous learning: We learn from incidents through retrospectives and RCAs sharing knowledge across teams to prevent repeat issues
- Outcome accountability: Were accountable for customer and business results not just completing tasks measuring success by impact
- Thoughtful experimentation: We hold strong opinions loosely testing concepts and running small experiments before scaling solutions
- Iterative delivery: We think big but act small using Scrum delivery plans and frequent milestones to ship incrementally and learn fast
- Inclusive environment: We create space to listen and learn actively growing our Ally Community to support equity belonging and career growth for all
Required Experience:
Senior IC
About Company
PointClickCare is the #1 cloud-based healthcare software provider helping long-term and post-acute care (LTPAC) providers navigate the new realities of value-based healthcare.