Enter a job title or keyword

Senior AI Reliability Engineer

EPAM Systems


Job Location:

Lviv - Ukraine

Monthly Salary: Not provided by the employer
Posted: 9 September 2026 (3 days ago)
Application Deadline: 7 December 2026
Vacancies: 1 Vacancy

Job Summary

EPAMs Operational Intelligence practice is developing a new capability called AI Reliability Engineering (AIRE) which applies SRE principles and cloud-native practices throughout the lifecycle of production AI/ML and LLM systems. As clients transition their GenAI applications and agentic systems from pilot to production they find that traditional APM provides no visibility into token latency cost per request semantic drift or hallucinations. This role exists to close that gap.

Your work will involve instrumenting monitoring and hardening production AI systems establishing AI-native service level objectives and creating accelerators and reference implementations that the practice can reuse across accounts. Note that this is an engineering position rather than an L1/L2 support role and it does not require a 24/7 on-call rotation.

What Youll Get

  • A brand-new discipline within EPAM where you shape the approach rather than follow an existing playbook
  • Focus on engineering work without 24/7 on-call responsibilities
  • Sponsored certification and training programs (Anthropic/Claude Databricks AI & Data Observability learning paths)
  • Exposure to multiple clients and a direct route into presales and solution engineering
Responsibilities
  • Add AI telemetry to production LLM RAG and agentic applications using OpenTelemetry and APM-native AI monitoring tools
  • Deploy distributed tracing across multi-model chains agent workflows and retrieval-augmented generation pipelines to identify systemic latency and failure points
  • Establish and track AI-native SLIs and SLOs including time to first token (TTFT) throughput error and refusal rates cost per request semantic drift hallucination boundaries and contextual accuracy
  • Implement structured semantic logging and prompt/response monitoring to support quality analysis
  • Create and maintain evaluation loops for output quality and safety using golden sets LLM-as-a-judge methods and Ragas/DeepEval-style frameworks integrating them into CI/CD and runtime environments
  • Monitor token-based cloud spend model API rate limits and quota usage while driving AI cost optimization efforts
  • Set up AI gateways to manage API load balancing failover and fallback models across multiple LLM providers
  • Build guardrails covering prompt injection and jailbreak filtering output compliance and bias and safety constraints
  • Develop detection triage restoration and problem management workflows for AI incidents incorporating autonomous AI agents into root cause analysis to parse logs generate hypotheses and correlate state changes
  • Enable rollback canary and fail-safe patterns for model prompt and configuration releases while maintaining reproducibility through versioning of data code prompts and models
  • Develop practice accelerators reference architectures and internal training materials and contribute to presales activities and client assessments
Requirements
  • 4 years of experience in SRE DevOps platform or observability engineering with hands-on exposure to production AI/ML or LLM workloads
  • Strong grasp of SRE fundamentals including Golden Signals SLI/SLO definition error budgets and burn rate incident lifecycle and ITIL basics
  • Strong Python skills for building instrumentation automation and evaluation tools
  • Hands-on experience with OpenTelemetry and at least one APM/observability platform such as New Relic Datadog Grafana LGTM stack Splunk or Elastic
  • Production experience with at least one cloud platform (Azure preferred AWS or GCP also acceptable) and Kubernetes
  • Solid understanding of LLM application architecture including prompts embeddings and vector stores RAG and agent orchestration tools like LangChain/LangGraph or similar
  • Awareness of MLOps concepts including model lifecycle (training vs. inference) model endpoints containerization and deployment/rollback patterns
  • Experience with Infrastructure as Code using Terraform along with CI/CD tools such as Azure DevOps GitLab CI or GitHub Actions
  • B2 English proficiency since the role involves direct client interaction and requires clear technical communication in both writing and speech
Nice to have
  • Familiarity with AI-specific observability and evaluation tools such as Traceloop/OpenLLMetry Langfuse Arize Phoenix Ragas DeepEval or MLflow
  • Experience with distributed inference serving at scale including vLLM KServe Ray Serve Kubernetes-native LLM orchestration and GPU capacity planning
  • Knowledge of AI security practices including the OWASP LLM Top 10 prompt injection defense and guardrail frameworks like NeMo Guardrails or Llama Guard
  • Experience with Databricks (including Mosaic AI/MLflow) or Azure AI Foundry
  • Relevant certifications such as Anthropic Claude Azure AI Engineer AWS ML Specialty or Databricks GenAI
  • Background in FinOps for AI workloads particularly token and GPU cost modeling
  • Experience in Data Reliability Engineering covering data quality and pipeline SLOs since AI reliability depends on data reliability
  • Prior mentoring or team lead experience

Required Experience:

Senior IC