Senior Site Reliability Engineer Oracle Database
Job Summary
The Autonomous Recovery Service (RCV) team is responsible for delivering highly available secure and resilient cloud services that protect Oracle Cloud Infrastructure (OCI) customer data. Our mission is to ensure customers can confidently recover from planned and unplanned events through industry-leading recovery capabilities intelligent automation and operational excellence.
As a Software Development & Site Reliability Engineer you will play a critical role in operating and continuously improving mission-critical cloud services. This is an operations-first engineering role where youll own the health reliability and lifecycle of production services while developing software and automation that reduce operational complexity improve service resilience and enhance the customer experience.
Youll work across the full service lifecyclefrom deployment patching upgrades monitoring incident response and root cause analysis to designing automation improving observability and implementing engineering solutions that eliminate repetitive operational work. Success in this role requires curiosity strong analytical thinking and a passion for solving complex operational challenges through software and automation.
Our engineers embrace AI as a force multiplier using AI-assisted development and operational tools to accelerate problem solving improve productivity and build smarter more autonomous systems. We value engineers who combine technical expertise with critical thinking sound engineering judgment and a continuous improvement mindset to challenge existing processes and drive innovation.
If you enjoy owning production services building reliable cloud infrastructure automating everything possible and working on technology that protects mission-critical customer data at cloud scale wed love to have you join our team.
Responsibilities
- Support the day-to-day operation health availability and reliability of Oracles Autonomous Recovery Service (RCV) production environments.
- Perform operational activities including deployments patching upgrades security updates infrastructure maintenance and production change management following established operational procedures.
- Monitor production services investigate alerts troubleshoot issues and assist in restoring service while helping the team achieve Service Level Objectives (SLOs).
- Participate in an on-call rotation with guidance from senior engineers to support production services and ensure a high level of customer availability.
- Assist in incident response root cause analysis (RCA) and post-incident reviews implementing corrective actions to prevent recurring issues.
- Develop test and maintain automation tools scripts and software that simplify operational tasks improve reliability and reduce operational toil.
- Contribute to Infrastructure as Code (IaC) solutions using Terraform to automate provisioning configuration and management of Oracle Cloud Infrastructure (OCI) resources.
- Collaborate with Software Development and Site Reliability Engineering teams to build reliable scalable and maintainable cloud services.
- Analyze logs metrics and operational data to identify trends troubleshoot issues and recommend improvements to service performance and reliability.
- Support capacity planning performance optimization and service scalability initiatives.
- Develop and enhance monitoring alerting dashboards and observability solutions that improve operational visibility.
- Assist in managing Oracle Database environments including backup restore recovery validation and disaster recovery operations using Oracle Recovery Manager (RMAN).
- Write and maintain automation using Python Bash REST APIs and other scripting technologies to improve operational efficiency.
- Use Git and Bitbucket to manage source code collaborate on development efforts and participate in modern CI/CD workflows.
- Leverage AI-assisted engineering tools to accelerate software development operational analysis troubleshooting and documentation while applying sound engineering judgment to validate results.
- Identify opportunities to automate repetitive operational processes and implement intelligent solutions that improve productivity service reliability and customer experience.
- Document operational procedures troubleshooting guides and lessons learned to improve team knowledge and operational readiness.
- Collaborate effectively with software engineers cloud infrastructure teams database engineers and cross-functional partners to deliver reliable cloud services.
- Continuously expand technical knowledge by learning Oracle Cloud Infrastructure (OCI) Oracle Database RCV architecture Site Reliability Engineering practices Infrastructure as Code (Terraform) automation and AI-assisted engineering techniques.
- Demonstrate curiosity critical thinking and a continuous improvement mindset by challenging existing processes and contributing innovative ideas that improve service reliability operational excellence and customer outcomes.
Technologies Youll Work With
- Cloud & Infrastructure: Oracle Cloud Infrastructure (OCI) Linux Terraform (Infrastructure as Code)
- Database Technologies: Oracle Database Oracle Recovery Manager (RMAN) backup and recovery disaster recovery
- Development & Automation: Python Bash/Shell scripting REST APIs JSON YAML
- Source Control & DevOps: Git Bitbucket CI/CD pipelines
- Site Reliability Engineering: Production monitoring observability incident response root cause analysis performance optimization capacity planning
- Modern Engineering: AI-assisted development intelligent automation operational analytics Agile engineering practices
Qualifications
Career Level - IC3
Required Experience:
Senior IC
About Company
As a world leader in cloud solutions, Oracle uses tomorrow’s technology to tackle today’s challenges. We’ve partnered with industry-leaders in almost every sector—and continue to thrive after 40+ years of change by operating with integrity. We know that true innovation starts when eve ... View more