Senior Site Reliability Engineer, CloudOps
Raleigh, WV - USA
Job Summary
Position Summary
We are seeking a Senior Site Reliability Engineer CloudOps to support scale and optimize a multi-account AWS environment hosting healthcare-oriented applications and analytics this role you will bridge infrastructure engineering operational reliability production support and cloud modernization initiatives across complex microservices architectures. The ideal candidate brings strong AWS expertise solid Linux administration skills and a proven track record of managing production systems in HIPAA/HiTrust regulated environments. You will participate in incident response on-call rotations and continuous deployment workflows while helping drive our transition toward containerized and Kubernetes-based platforms. This collaborative position is built for an analytical engineer who excels at resolving production incidents partnering with developers and continuously elevating operational excellence.
Essential Duties & Responsibilities
- Manage maintain and troubleshoot a multi-account AWS Organization environment (35 accounts) and core services including EC2 ECS/Fargate Lambda S3 CloudFront API Gateway and Aurora/RDS databases.
- Support production deployments CI/CD pipelines (Jenkins AWS CodePipeline) and infrastructure automation using Python Bash and AWS CloudFormation.
- Monitor system health and performance using Datadog CloudWatch and Zabbix; investigate alerts execute root-cause analysis and refine monitoring coverage to reduce operational noise.
- Participate in a shared on-call rotation managing incident response and performing failover/recovery validation for production applications and data stores.
- Maintain HIPAA/HiTrust compliance and security posture by managing tools like Prisma/Cortex Cloud Security Hub and GuardDuty while enforcing proper IAM policies and network segmentation.
- Support Java (Spring Boot) and Python applications running in containers assisting developers during investigations and preparing for future Kubernetes (EKS) modernization initiatives.
Knowledge & Skills
- Deep hands-on expertise with AWS core services (networking compute serverless and database technologies) and CloudFormation IaC automation.
- Strong Linux administration skills (primarily Ubuntu) along with proficiency in Python and Bash scripting for operational automation.
- Experience with containerization technologies (Docker ECS/Fargate) and familiarity with modern Kubernetes ecosystems (EKS Helm ArgoCD).
- Solid understanding of observability tools (Datadog CloudWatch Zabbix) and CI/CD pipelines (Jenkins CodePipeline Git workflows).
- Knowledge of cloud security best practices access management (IAM) and compliance frameworks within regulated sectors (HIPAA/HiTrust).
- Proven diagnostic incident-management and analytical troubleshooting skills for complex microservices architectures.
Minimum Qualifications Education & Experience
- Must be at least 18 years of age.
- High School Diploma required.
- Bachelors degree from an accredited college or university is required.
- 7 years of hands-on experience in AWS Cloud Engineering DevOps Site Reliability Engineering (SRE) or Infrastructure Engineering.
- Practical background supporting production workloads in Linux/AWS environments reading application logs and making minor code fixes.
- Direct experience participating in on-call rotations and incident response protocols.
- Prior experience in the healthcare industry maintaining HIPAA/HiTrust-compliant infrastructure.
Work Environment
- This is largely a sedentary role.
- This job operates in a professional office environment and routinely uses standard office equipment.
- Typically requires travel less than 5% of the time
Required Experience:
Senior IC
About Company
ICU Medical has consistently provided you with clinical innovations that help solve real-world challenges. With the acquisition of Hospira Infusion Systems in 2017 and Smiths Medical in 2022, we are now a global market leader with a complete line of clinically-essential IV therapy and ... View more