ML Infrastructure Engineer
San Francisco, CA - USA
Department:
Job Summary
- Design build and maintain the infrastructure tooling and workflows that enable reliable scalable deployment of ML models to production.
- Develop monitoring and observability systems to track model performance data drift data quality and overall system health.
- Create and maintain end-to-end testing frameworks and simulation environments to validate models and pipelines prior to deployment.
- Work closely with Data Engineering and Platform Engineering teams to ensure ML systems integrate cleanly with broader Gridware infrastructure and operational standards.
- Improve CI/CD pipelines for ML workloads ensuring reproducibility safe rollout and automated rollback strategies.
- 5 years of experience building production ML infrastructure
- Strong software engineering skills and proficiency in Python
- Experience with cloud platforms (AWS) and container orchestration (Kubernetes)
- Familiarity with feature stores model registries or centralized metadata systems (i.e. MLFlow)
Required Experience:
IC
About Company
**At this time, Gridware is unable to provide visa sponsorship or immigration support for this role. We’re only able to consider candidates who are currently authorized to work in the country of employment without visa sponsorship now or in the future.** This describes the ideal candi ... View more