Sr. Site Reliability Engineer

MX Technologies


Job Location:

Lehi, UT - USA

Monthly Salary: Not Disclosed
Posted on: 7 hours ago
Vacancies: 1 Vacancy

Job Summary

MX is a fintech company on a mission to empower the world to be financially strong. We build technology that helps banks credit unions and fintechs deliver smarter more intuitive financial experiences to millions of people.

Like many startups weve navigated real growth challenges and weve come out stronger on the other side. Today MX is in a phase of renewed momentum and scale with a solid foundation and a clear vision for whats next. This is a place where thoughtful execution matters innovation is encouraged and individuals have real ownership over their work.

Our culture values curiosity accountability and impact. We give people the space to question assumptions design better solutions and help shape how the company grows. If youre looking to do meaningful work influence outcomes and grow alongside a company thats ready to move fast youll feel at home at MX.

At MX reliability is a product. Our infrastructure powers financial applications used by millions of people and processes billions of transactions for major financial institutions and customers feel every second of downtime.

Were building a new observability function that runs the way we run incident response: the system does the heavy lifting and people handle judgment customers and the exceptions. As a Senior Observability Engineer you build and operate an observability control plane. You scaffold baselines score coverage and turn every real incident into the detection the platform should have caught. This is a multiplier role: you raise the bar for every team through standards and automation instead of building each teams dashboards by hand.

We call it the shepherd model. You shepherd Datadog and partner with our product engineering teams so they observe the right signals for their products. Service owners get real signal instead of noise and leadership gets coverage and health as a program metric.

This role shares the team pager. Observability and incident response run one on-call roster. You take shifts with the rest of the team and act as Incident Commander when an incident needs one. It is core to the role not an afterthought.

Engineering at MX runs hybrid infrastructure (AWS and bare metal) with services in Ruby Go and Java messaging over NATS and RabbitMQ and data on PostgreSQL and Redis. Datadog is our observability platform and is our incident response platform.

What youll do:

  • Build and operate an observability control plane: automate baseline monitors dashboards and tagging standards through the Datadog API and Terraform.

  • After significant incidents produce detection and dashboard gap packs grounded in Datadog and MX investigation patterns with queries ready to apply.

  • Define what good looks like for a Ruby Go or Java service on Datadog (tags golden signals alert quality dashboard contracts) then audit services against that standard and accept or reject readiness.

  • Validate dont own. Service owners keep their alerts and dashboards; you confirm they are complete and correct then move on. Escalate to engineering managers when coverage fails or an owner is missing.

  • Own the monthly observability and service-catalog health report: departed owners stale dashboards services with no monitors SLO gaps and coverage trends.

  • Run maturity assessments (baseline through SLO launch-ready self-serve) and track them over time.

  • Tune alerting toward zero false SEV1/2 pages and actionable SEV3/4 alerts and coach teams on Datadog cost and cardinality.

  • Build self-serve onboarding so new services get baseline observability on day one without a multi-week embed.

  • Share the team pager. Rotate on the shared IR & Observability on-call triage and investigate live incidents with Datadog and MX investigation patterns and take Incident Commander or supporting technical roles as the incident needs.

  • After incidents close the detection loop (gap packs new monitors dashboards) so the pager gets quieter over time.

  • Run high-value launch and production-readiness reviews as a checkpoint not a permanent staffing model.

Basic Requirements

  • BS in Computer Science or equivalent experience

  • 5 years running production observability SRE or DevOps

  • 5 years automation-first engineering in Python Bash Go and/or Terraform plus Kubernetes proficiency

  • AI- and workflow-literate. Youve used or built scripted and AI-assisted workflows to scale reviews audits and docs

  • Distributed-systems debugging across microservices: latency connection pools queues and cascading failure on Kubernetes and bare metal with NATS RabbitMQ Postgres and Redis

  • Shared on-call Incident Commander-capable

Preferred Requirements

  • Fintech experience with MX-like architectures

  • Datadog preferred; strong Grafana/Prometheus Splunk or New Relic experience counts if you can ramp on Datadog fast

  • Google SRE practices: toil elimination incident management automation for self-healing

  • Cross-functional influence without authority. Youve improved teams that dont report to you

  • Governance and reporting: you can produce a monthly health and compliance report leadership reads (orphans stale entries gaps trends)

  • OpenTelemetry instrumentation

  • Incident response platforms ( PagerDuty OpsGenie); prior formal Incident Commander experience

  • Golang and Ruby on Rails (the MX stack)

At MX we are a high-performance organization that thrives on trust and results. This role is based in Lehi Utah. We believe in empowering our team members to deliver exceptional outcomes while taking advantage of our incredible office space when it best supports their work. Our Utah office features onsite perks such as company-paid meals a sports simulator gym mothers lounge and meditation room and meaningful interactions with amazing people. We encourage team members to come together in the office to collaborate kick off key projects or strategize cross-functionally fostering connection and innovation.

MX is proudly committed to recruiting and retaining a diverse and inclusive workforce. As an Equal Opportunity Employer we never discriminate based on race religion color national origin gender (including pregnancy childbirth or related medical conditions) sexual orientation gender identity gender expression age military or veteran status status as an individual with a disability or other applicable legally protected characteristics. We particularly welcome applications from veterans and military spouses. All your information will be kept confidential according to EEO guidelines. You may request reasonable accommodations by sending an email to


Required Experience:

Senior IC

MX is a fintech company on a mission to empower the world to be financially strong. We build technology that helps banks credit unions and fintechs deliver smarter more intuitive financial experiences to millions of people.Like many startups weve navigated real growth challenges and weve come out s...