MLOps AI Operations Engineer
Job Summary
Job Description & Summary
Industrialize AI delivery through automated deployment evaluation operations observability reliability engineering and transparent consumption management.
Build CI/CD pipelines for AI services prompts agent configurations infrastructure and evaluation assets.
Automate environment provisioning testing deployment rollback and release evidence.
Implement tracing logging model and agent monitoring alerts and operational dashboards.
Operationalize evaluation thresholds incident handling and continuous-improvement loops.
Monitor latency capacity token usage infrastructure consumption and cost drivers.
Define runbooks service ownership and production support handover.
4 years in DevOps platform engineering ML engineering SRE or cloud operations.
Strong automation containers cloud services observability and Infrastructure as Code capability.
Experience deploying or operating ML generative AI or distributed application workloads.
Understanding of release controls reliability security and cost optimization.
Hands-on experience with GitHub Actions Azure DevOps GitLab CI or equivalent plus Infrastructure as Code using Terraform Bicep or comparable tooling.
Strong container and orchestration capability using Docker and Kubernetes together with experience deploying AI or agent services across cloud and hybrid environments.
Experience operating model and prompt assets agent configurations evaluation datasets and release evidence using MLflow platform-native registries or equivalent lifecycle tooling.
Practical implementation of agent tracing and observability using OpenTelemetry and tools such as LangSmith MLflow Langfuse Azure Monitor Prometheus or Grafana.
Ability to monitor model and agent quality tool failures retrieval performance latency token usage cost capacity and workflow-level service indicators.
Experience with progressive delivery rollback secrets management vulnerability scanning incident response and reliability practices for non-deterministic AI systems.
Deployment frequency and success rate
Mean time to detect and restore
Evaluation and monitoring coverage
Service reliability and latency
Cost and consumption transparency
Other members of the AI Transformation & Agentic Systems Practice
PwC sector functional cloud cyber risk Responsible AI and change specialists
Client business owners product owners technology teams and operational users
Technology alliance and implementation partners where relevant
Support proposals client workshops and market development appropriate to seniority.
Contribute reusable methods patterns code assets and lessons learned.
Coach colleagues and participate in the capabilitys continuous learning agenda.
Uphold PwC quality independence confidentiality and risk-management requirements.
#LI-BS1 #LI-Hybrid
Required Experience:
IC
About Company
At PwC, our purpose is to build trust in society and solve important problems. We’re a network of firms in 155 countries with over 284,000 people who are committed to delivering quality in assurance, advisory and tax services. Find out more and tell us what matters to you by vis ... View more