DevOps Engineer
Job Summary
Fold Health Engineering
DevOps Engineer
Job Title | DevOps Engineer |
Experience | 35 Years |
Location | Pune (Amar Tech Park Balewadi) Full-time In-office |
Domain | Healthcare (Preferred) |
About Fold Health
Fold Health is building an AI-powered healthcare technology platform designed to support Value-Based Care (VBC) programs. By combining advanced data integration analytics and intelligent automation we help providers payers and care teams deliver better outcomes and improve patient experiences. Our mission is to simplify healthcare technology ensure interoperability and enable innovation at scale. Join us and be part of shaping the future of AI-driven healthcare.
Role Overview
We are looking for a DevOps Engineer who will be responsible for building maintaining and scaling the cloud infrastructure and delivery pipelines that power Fold Healths healthcare platform. The ideal candidate is hands-on reliability-focused and comfortable working across AWS GCP Terraform and Prometheus in a fast-paced healthcare environment.
Responsibilities
1. Cloud Infrastructure & Platform Engineering
- Design build and maintain highly available secure and scalable cloud infrastructure across AWS and GCP.
- Manage production environments including Amazon ECS Fargate RDS PostgreSQL ElastiCache ALB SQS/SNS Lambda Step Functions Cognito Route 53 and supporting GCP services.
- Optimize cloud resources for performance reliability availability and cost efficiency.
- Design and maintain networking security groups load balancing and disaster recovery capabilities.
2. Infrastructure as Code & Automation
- Develop and maintain reusable Infrastructure as Code using Terraform following industry best practices.
- Build enhance and maintain CI/CD pipelines to enable reliable secure and automated application deployments.
- Automate infrastructure provisioning operational tasks and cloud governance through scripting and tooling.
- Standardize infrastructure components deployment patterns and operational workflows across environments.
3. Observability Monitoring & Incident Management
- Build and maintain comprehensive monitoring logging and alerting solutions using CloudWatch Prometheus AlertManager Grafana and xMatters.
- Develop actionable alerting strategies to reduce alert fatigue while ensuring rapid detection of production issues.
- Create dashboards metrics and operational insights to improve platform health and service reliability.
- Participate in production incident response root cause analysis and post-incident reviews driving preventive improvements.
4. Reliability Performance & Security
- Ensure platform reliability scalability and performance through proactive capacity planning tuning and optimization.
- Troubleshoot complex production issues across infrastructure networking databases and distributed applications.
- Implement security best practices including IAM secrets management encryption vulnerability remediation and least-privilege access.
- Support compliance initiatives such as HIPAA SOC 2 and HITRUST by implementing required technical controls and maintaining audit readiness.
5. Collaboration & Continuous Improvement
- Partner with software engineering teams to improve application deployment reliability and operational excellence.
- Advocate DevOps best practices automation and Infrastructure as Code throughout the engineering organization.
- Continuously evaluate and adopt new cloud technologies tools and processes to improve platform efficiency and developer experience.
- Create and maintain technical documentation runbooks and operational procedures to support knowledge sharing and onboarding.
Requirements
Requirements
Bachelors degree in Computer Science Engineering Information Technology or a related field.
35 years of hands-on experience in DevOps Platform Engineering or Site Reliability Engineering.
Strong experience with AWS services such as ECS RDS IAM CloudWatch ALB Lambda and networking.
Hands-on experience with Infrastructure as Code using Terraform and CI/CD pipelines using GitLab CI/CD or similar tools.
Strong experience with Docker and containerized application deployments.
Working knowledge of Kubernetes including deployments services configuration scaling and troubleshooting.
Familiarity with PostgreSQL Redis or ElastiCache and microservices architectures.
Proficiency in Linux administration Python or Bash scripting and infrastructure automation.
Experience with monitoring and observability tools such as Prometheus Grafana CloudWatch and AlertManager.
Good understanding of cloud security IAM networking secrets management and production troubleshooting.
Strong problem-solving communication and collaboration skills.
Good to Have
Experience with Google Cloud Platform services.
Experience managing Kubernetes workloads in production using EKS GKE or similar platforms.
Experience with on-call management tools such as xMatters or PagerDuty.
Knowledge of healthcare compliance standards such as HIPAA HITRUST or SOC 2.
Benefits
Why Join Us
Opportunity to build and scale an AI-powered healthcare platform impacting millions of lives
Work with modern cloud technologies in a mission-driven environment
Collaborative and innovative culture with strong growth opportunities
Required Skills:
Bachelors degree in Computer Science Engineering Information Technology or a related field. 35 years of hands-on experience in DevOps Platform Engineering or Site Reliability Engineering. Strong experience with AWS services such as ECS RDS IAM CloudWatch ALB Lambda and networking. Hands-on experience with Infrastructure as Code using Terraform and CI/CD pipelines using GitLab CI/CD or similar tools. Strong experience with Docker and containerized application deployments. Working knowledge of Kubernetes including deployments services configuration scaling and troubleshooting. Familiarity with PostgreSQL Redis or ElastiCache and microservices architectures. Proficiency in Linux administration Python or Bash scripting and infrastructure automation. Experience with monitoring and observability tools such as Prometheus Grafana CloudWatch and AlertManager. Good understanding of cloud security IAM networking secrets management and production troubleshooting. Strong problem-solving communication and collaboration skills
Required Education:
Bachelors degree in Computer Science Engineering Information Technology or a related field