Enter a job title or keyword

MLOps AI Operations Engineer


Job Location:

Bucharest - Romania

Monthly Salary: Not provided by the employer
Posted: 25 September 2026 (21 hours ago)
Application Deadline: 23 December 2026
Vacancies: 1 Vacancy

Job Summary

Job Description & Summary

The opportunity

Industrialize AI delivery through automated deployment evaluation operations observability reliability engineering and transparent consumption management.


What you will be doing

Build CI/CD pipelines for AI services prompts agent configurations infrastructure and evaluation assets.

Automate environment provisioning testing deployment rollback and release evidence.

Implement tracing logging model and agent monitoring alerts and operational dashboards.

Operationalize evaluation thresholds incident handling and continuous-improvement loops.

Monitor latency capacity token usage infrastructure consumption and cost drivers.

Define runbooks service ownership and production support handover.


What we need from you

4 years in DevOps platform engineering ML engineering SRE or cloud operations.

Strong automation containers cloud services observability and Infrastructure as Code capability.

Experience deploying or operating ML generative AI or distributed application workloads.

Understanding of release controls reliability security and cost optimization.


Relevant AI technologies and tooling

Hands-on experience with GitHub Actions Azure DevOps GitLab CI or equivalent plus Infrastructure as Code using Terraform Bicep or comparable tooling.

Strong container and orchestration capability using Docker and Kubernetes together with experience deploying AI or agent services across cloud and hybrid environments.

Experience operating model and prompt assets agent configurations evaluation datasets and release evidence using MLflow platform-native registries or equivalent lifecycle tooling.

Practical implementation of agent tracing and observability using OpenTelemetry and tools such as LangSmith MLflow Langfuse Azure Monitor Prometheus or Grafana.

Ability to monitor model and agent quality tool failures retrieval performance latency token usage cost capacity and workflow-level service indicators.

Experience with progressive delivery rollback secrets management vulnerability scanning incident response and reliability practices for non-deterministic AI systems.


Measures of success

Deployment frequency and success rate

Mean time to detect and restore

Evaluation and monitoring coverage

Service reliability and latency

Cost and consumption transparency


Key interfaces

Other members of the AI Transformation & Agentic Systems Practice

PwC sector functional cloud cyber risk Responsible AI and change specialists

Client business owners product owners technology teams and operational users

Technology alliance and implementation partners where relevant


Contribution to the practice

Support proposals client workshops and market development appropriate to seniority.

Contribute reusable methods patterns code assets and lessons learned.

Coach colleagues and participate in the capabilitys continuous learning agenda.

Uphold PwC quality independence confidentiality and risk-management requirements.



#LI-BS1 #LI-Hybrid


Required Experience:

IC


About Company

Company Logo

At PwC, our purpose is to build trust in society and solve important problems. We’re a network of firms in 155 countries with over 284,000 people who are committed to delivering quality in assurance, advisory and tax services. Find out more and tell us what matters to you by vis ... View more

View Profile View Profile