SRE (Engineering & Administration Background)
Mexico City - Mexico
Job Summary
About Fulcrum Digital
Fulcrum Digital is a global AI-first enterprise transformation company with over 25 years of experience. We partner with enterprises across financial services insurance healthcare retail manufacturing higher education and logistics to move from AI experimentation to scalable business outcomes. With over 100 global clients including Fortune 500 enterprises we combine deep industry expertise with capabilities in enterprise AI digital engineering cloud modernisation platform integration and generative AI.
The Role
We are looking for a Site Reliability Engineer (SRE) based in Mexico City on a hybrid schedule to own and optimize the health of our production environment. You will be at the center of platform reliability working across the full service lifecycle from design through operation while collaborating with a global team across multiple time zones.
What Youll Do
- Plan manage and oversee all aspects of a production environment
- Define strategies for application performance monitoring and optimization in production
- Respond to incidents drive platform improvements and measure incident reduction over time
- Support code deployment across multiple lower environments with a strong focus on automation
- Design develop and standardize monitoring and alerting mechanisms
- Take a holistic cross-stack approach to problem-solving during production events to optimize mean time to recovery (MTTR)
- Own the full service lifecycle from inception and design through deployment operation and refinement
- Analyze ITSM activity and provide feedback loops to development teams on operational gaps and resiliency concerns
- Support pre-launch activities including system design consulting capacity planning and launch reviews
- Support CI/CD pipelines through validation and operational gating championing DevOps best practices
- Monitor availability latency and overall system health to keep services running smoothly
- Drive sustainable scaling through automation and reliability-focused system changes
- Perform root cause analysis and on-call support on a rotational basis
- Collaborate with a global team spread across multiple tech hubs and time zones
- Share knowledge and mentor others on processes and procedures
- Occasional off-hours work required
Must Have
- Significant strength in monitoring including SOP creation Splunk and Dynatrace experience is required
- Strong hands-on experience with Linux
- Shell scripting proficiency
- Strong application troubleshooting skills
- Jenkins / CI-CD pipeline experience
- Solid understanding of ITIL / ITSM processes
Also Required
- SQL knowledge
- Ansible automation experience
- Groovy scripting / YAML
- Working knowledge of Git / Bitbucket
- DevOps automation experience (CI/CD Jenkins GitHub)
- Root cause analysis and capacity planning experience
Preferred Qualifications
- Card payment knowledge (payment flows switching settlements authorization flows)
- Experience with Prometheus/Grafana
- Cloud experience AWS and/or Azure
- Event-driven framework architecture experience
Required Skills:
Significant strength in monitoring including SOP creation Splunk and Dynatrace experience is required Strong hands-on experience with Linux Shell scripting proficiency Strong application troubleshooting skills Jenkins / CI-CD pipeline experience Solid understanding of ITIL / ITSM processes Also Required SQL knowledge Ansible automation experience Groovy scripting / YAML Working knowledge of Git / Bitbucket DevOps automation experience (CI/CD Jenkins GitHub) Root cause analysis and capacity planning experience