Staff Engineer, Machine Learning Systems & Reliability Moveworks
Mountain View, CA - USA
Job Summary
We are building AI-enabled product capabilities that improve through data feedback and real-world use. We need the production systems that make those capabilities dependable: repeatable delivery measurable quality controlled learning loops and reliable operation at scale.
Were looking for a hands-on Staff Engineer who can move machine-learning models agentic workflows and self-learning approaches from promising prototypes into secure observable continuously deployable production systems.
This role sits at the intersection of ML systems platform engineering and site reliability engineering. You will partner with ML data product and infrastructure teams to create a paved path from experimentation to productionand take ownership of how those systems perform and evolve once deployed.
What youll do
- Design and build the production path for the complete ML lifecycle: data and feature preparation training experiment tracking evaluation artifact and model management serving monitoring feedback collection and retraining.
- Build continuous-delivery workflows for models prompts agent workflows data dependencies and supporting services. Establish automated quality safety performance and compatibility checks.
- Implement safe rollout patterns such as shadow traffic canaries progressive delivery feature flags versioned artifacts automated rollback and operational kill switches.
- Turn self-learning approaches into controlled production feedback loops. Build systems for collecting outcomes validating feedback maintaining lineage triggering model refreshes comparing candidates and promoting changes under explicit guardrails.
- Define and operate SLIs SLOs alerts and error budgets across infrastructure data pipelines inference services model quality and product behavior.
- Connect model analytics and product telemetry with traditional operational signals so teams can understand whether a problem originates in infrastructure data model behavior or the surrounding product.
- Improve the scalability availability latency and cost efficiency of distributed training inference and data-processing workloads. Own capacity planning and resource optimization including GPU resources where applicable.
- Participate in production ownership across the service lifecycle: architecture reviews deployment on-call incident response blameless postmortems and systemic remediation.
- Build self-service platforms and automation that reduce operational toil and shorten the time required for ML engineers and data scientists to reach production.
- Apply LLMs or agentic automation to evaluation troubleshooting and operational workflows where they produce reliable measurable improvements.
- Establish practical standards for cloud infrastructure Kubernetes infrastructure as code observability security and compliance.
- Provide technical leadership across ML data product and platform teams mentoring engineers and influencing architecture without relying on formal authority.
Qualifications :
To be successful in this role you have:
- A track record of Staff-level technical ownership typically gained through 7 years of experience in software engineering platform engineering SRE production engineering or ML infrastructure.
- Strong software-engineering skills in Python and at least one production systems language such as Go Java C or Rust.
- Experience designing operating and troubleshooting distributed production systems including failure analysis capacity planning and performance optimization.
- Hands-on experience with cloud infrastructure containers and Kubernetes infrastructure as code CI/CD and modern observability.
- Practical understanding of the ML lifecycleincluding training evaluation model deployment serving monitoring versioning and retrainingand the ability to collaborate effectively with applied ML engineers or researchers.
- Experience distinguishing service-health problems from data-quality or model-quality problems.
- Familiarity with SRE practices such as SLIs/SLOs error budgets sustainable on-call incident management and blameless postmortems.
- A strong automation and internal-customer mindset: you build platforms that are reliable understandable and pleasant for other engineers to use.
- Excellent technical judgment and communication skills especially when navigating ambiguity and coordinating across teams during production incidents.
Additional Information :
Work Personas
We approach our distributed world of work with flexibility and trust. Work personas (flexible remote or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.
Equal Opportunity Employer
ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race color religion sex sexual orientation national origin age disability gender identity veteran status or any other category protected by addition all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.
Accommodations
We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process or are unable to use this online application and need an alternative method to apply please contact for assistance.
Export Control Regulations
For positions requiring access to controlled technology subject to export control regulations including the U.S. Export Administration Regulations (EAR) ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities.
From Fortune. 2026 Fortune Media IP Limited. All rights reserved. Used under license.
Remote Work :
No
Employment Type :
Full-time
About Company
Learn here. Grow here. Make a difference here. At ServiceNow, our cloud?based platform and solutions deliver digital workflows that create great experiences and unlock productivity for employees and enterprises. Were growing fast, innovating even faster, and making an impact on our c ... View more