Site Reliability Engineer
Wilmington, DE - USA
Job Summary
The Job
As Site Reliability Engineer you will serve as a reliability subject matter expert who leads major incident recovery drives observability and reliability improvements mentors associate engineers reduces operational toil influences technical decisions and improves resiliency standards.
This role requires depth across production systems telemetry batch operations automation and incident response. You will be expected to guide technical direction for reliability improvements and help teams prevent recurring failures.
Employees joining Best Eggs Information Technology organization can expect a culture centered on Continuous Delivery Total Quality Management Knowledge Sharing Personal and Career Advancement Empowerment Innovation and Collective Ownership.
- Lead technical recovery efforts for major incidents coordinating triage evidence review restoration actions and validation.
- Optimize observability strategy alert quality dashboard standards and telemetry coverage across multiple services.
- Drive reliability initiatives that reduce recurring failures noisy alerts manual work and operational risk.
- Mentor associate engineers on troubleshooting methods RCA evidence runbook quality and production support judgment.
- Influence engineering decisions by identifying reliability risks missing telemetry supportability gaps and resiliency patterns.
- Improve JAMS GoAnywhere Datadog xMatters and service support practices through automation and standards.
- Partner with leaders and technical teams to prioritize remediations based on customer impact business impact and operational exposure.
- Hands-on familiarity with production support monitoring alerting and incident response practices.
- Working knowledge of Datadog dashboards monitors logs metrics and APM concepts.
- Ability to troubleshoot application infrastructure batch or file transfer issues using runbooks and telemetry.
- Exposure to AWS or cloud operations and scripting with Python PowerShell Bash or similar tools.
- Clear communication skills during incidents service requests and post-incident follow-through.
- Strong experience leading production incident recovery and cross-system reliability investigations.
- Ability to mentor engineers and influence technical decisions without direct authority.
- Datadog AWS ITIL Linux or automation certification.
- Experience with JAMS GoAnywhere xMatters ServiceNow/Jira or CI/CD environments.
- Exposure to AIOps anomaly detection operational automation or reliability engineering.
- Familiarity with financial services controls secure file transfer or regulated operations.
- Serves as a reliability SME across observability incident response batch operations and operational platforms.
- Leads major incident recovery with calm command of telemetry dependencies impact and restoration options.
- Reduces operational toil through automation standards better alerting and durable remediation.
- Mentors engineers and improves the quality of technical support practices across the team.
- Influences design and readiness decisions that improve resiliency and operational resilience.
- Lead a major incident or complex reliability investigation with clear recovery and follow-through.
- Deliver an observability or automation improvement that measurably reduces alert noise toil or repeat issues.
- Mentor associate engineers through troubleshooting reviews runbook improvements or incident debriefs.
- Identify and influence remediation of a meaningful resiliency or supportability gap.
- Improve standards or patterns for telemetry escalation batch support or operational validation.
Please note this job description is not designed to cover every activity duty or responsibility required for the job. Duties responsibilities and activities may change at any time with or without notice.
Required Experience:
IC
About Company
Get low-interest personal loans quickly with Best Egg. Apply online in minutes & receive funds fast. Start your journey to financial stability now!