SRE
ملخص الوظيفة
Spinomenal is a dynamic force in the online casino sector bursting onto the scene in 2014. Renowned as one of the fastest-growing content creators in the iGaming industry Spinomenal thrives on fostering creativity and collaboration. By nurturing an environment where employees excel through teamwork and communication Spinomenal maintains a rapid pace of innovation and development. Our dedication to collective effort and shared vision has propelled us to deliver captivating cutting-edge gaming experiences to players worldwide.
About the position:
Spinomenal is seeking a hands-on Production Manager / SRE to own release engineering platform stability and incident response in our high-throughput low-latency iGaming environment. Serving as the gatekeeper to production across R&D DevOps QA and Support you will ensure high availability secure CI/CD pipelines and rapid anomaly diagnosis through automation.
Responsibilities:
Validate complex configurations execute automated health checks and own recovery and rollback runbooks for production deployments.
Identify manual operational tasks (toil) and automate them using scripting and Infrastructure-as-Code principles to increase reliability.
Maintain and optimize distributed tracing dashboards alerting thresholds and log aggregation using New Relic Grafana and advanced SQL.
Serve as a critical escalation point for complex cross-layer production failures (Application Network Database and Infrastructure).
Drive deep-dive post-mortems and implement structural fixes to prevent incident recurrence.
Govern infrastructure and configuration changes across environments in tight collaboration with DevOps to maintain environment parity and prevent drift.
Partner with QA Architects DevOps and Game Producers to champion SRE best practices operational readiness standards and resilient architecture.
Requirements:
At least 2 years of experience in a dedicated Site Reliability Engineering (SRE) Production Operations or Release Engineering role handling high-transaction web applications.
Deep experience engineering and maintaining Jenkins pipelines (Pipeline-as-Code / Jenkinsfiles) and familiarity with configuration tools (e.g. Ansible Helm).
Strong hands-on experience managing and troubleshooting cloud infrastructure (AWS: EC2 ECS S3 Lambda IAM policies VPC routing) and Infrastructure as Code (IaC) using Terraform.
Proficiency in Python or Bash for writing automation scripts system utilities and internal tooling.
Advanced capability with observability stacks (New Relic Prometheus/Grafana) and strong SQL skills for log parsing and database debugging.
Mastery of web debugging (interpreting JSON payloads analyzing API contracts diagnosing HTTP status anomalies) and understanding of DNS CDN/Redis caching and proxies.
Demonstrated capability in debugging distributed microservices architectures and defining Service Level Indicators/Objectives (SLIs/SLOs) and error budgets.
Exceptional technical communication skills with a proactive mindset and the ability to stay calm under pressure during critical outages and on-call rotations.