DevOps Engineer, Agentic Operation & Live Games
Department:
Job Summary
Big Viking Games is a Canadian gaming company focused on building operating and growing long-standing online game communities. Our games have entertained players for years supported by loyal audiences live operations evolving content systems product innovation and deep player-driven economies.
Our flagship titles YoWorld and FishWorld have served millions of players over their lifetime. These are enduring live-service virtual worlds with rich in-game economies virtual goods social interaction and long-term player engagement at their core.
We are entering a new phase of modernization and growth with a focus on stronger infrastructure better automation practical AI adoption improved reliability stronger security practices and scalable systems that help our games and teams perform at a higher level.
Big Viking Games is hiring a Senior DevOps Engineer to modernize the infrastructure behind our live-service games and to do it by building automation that thinks not just scripts that run.
Our games run on mature tech stacks with large data volumes and player counts real systems with real players on them where the constraints are genuine and the consequences are visible. There is meaningful room to automate how they are operated and we want someone who builds that: tooling that takes routine work off peoples hands and runs safely against production without compromising uptime or data integrity.
This is a hands-on senior role for an engineer who is fluent in modern cloud infrastructure AWS containers IaC CI/CD observability production reliability andwho has started using agentic coding tools and tool-calling systems to do infrastructure work that previously required a person. If you have wired a coding agent into your CI/CD built an MCP server so tooling could act on your infrastructure safely or replaced a manual runbook with something that diagnoses and remediates on its own this role is aimed at you.
The defining trait isself-direction and it works in two directions. Given a backlog you improve on how the work gets done bringing agentic approaches to problems that were scoped as manual ones and solving the underlying issue rather than the individual ticket. Left to your own judgement you find the repetitive work nobody has flagged decide whats worth automating and build both cases you own the guardrails that make automation safe to run against production.
This is a hybrid role based in Toronto with an expectation of working in office three days per week. Live-service games require operational awareness outside regular business hours including periodic on-call and incident response availability.
Requirements
Build agentic automation for infrastructure work
- Design build and operate agentic tooling that performs real DevOps work diagnosis remediation provisioning routine maintenance rather than only summarizing or suggesting.
- Build and maintain the integration layer that lets tooling act safely on our systems: MCP servers API integrations webhook-driven orchestration serverless functions and the permissions and audit trails around them.
- Convert manual runbooks SOPs and recurring operational chores into automation that runs unattended with sensible escalation when it shouldnt proceed alone.
- Establish the guardrails that make automated action against production defensible: least-privilege scopes dry-run and approval paths for destructive operations logging of what automation did and why and clear rollback.
- Use agentic coding tools to accelerate your own infrastructure work IaC authoring migration scripting log and incident analysis documentation and improve how the wider engineering team does the same.
Own and modernize live production infrastructure
- Monitor maintain and improve cloud infrastructure across AWS Netlify Vercel and related platforms supporting our live games data systems internal tools and operational workflows.
- Drive infrastructure modernization while maintaining uptime for live games with active player communities every improvement ships while the plane is flying.
- Implement and maintain Infrastructure as Code using Terraform CloudFormation CDK or similar.
- Improve CI/CD pipelines release workflows and deployment reliability so teams can ship safely and frequently.
- Operate containerized workloads and GitOps-based deployment.
Reliability observability and security
- Rebuild observability across the stack: logging metrics alerting dashboards and operational visibility with particular attention to early detection of silent failures. Our previous monitoring stack is no longer functional so this is a build not a handover.
- Maintain and monitor data pipelines between game source databases (MariaDB) the Snowflake data warehouse and downstream analytics and reporting systems ensuring pipeline health freshness and alerting when data stops flowing.
- Own secrets and credential lifecycle management across platforms API key rotation access controls environment variable governance and least-privilege practices.
- Support incident response root cause analysis remediation planning and post-incident improvements and automate the parts of that loop that repeat.
- Help manage cloud spend infrastructure usage resource tagging and environment efficiency.
- Create documentation and runbooks that are executable wherever possible rather than purely descriptive.
Experience
- 7 yearsin DevOps infrastructure engineering cloud engineering site reliability engineering or a similar role.
- 1 year building with agentic coding tools and tool-calling systems to automate scale or secure infrastructure and DevOps work.We mean systems that take action agents wired into pipelines MCP servers or tool integrations you built automated diagnosis or remediation you put into production. Using an AI assistant to write code faster is useful but is not what this requirement is asking for.
Core technical
- Strong hands-on experience with AWS or similar cloud platforms.
- Experience designing maintaining and improving production infrastructure including comfort with legacy systems that predate modern cloud-native patterns.
- Proficiency with Infrastructure as Code tools such as Terraform CloudFormation CDK Pulumi or similar including managing shared state across a team.
- Experience with containerized applications (Docker) and container orchestration Kubernetes ECS EKS or similar including debugging real cluster problems in production.
- Experience with GitOps and declarative deployment (ArgoCD Flux or equivalent).
- Experience with CI/CD tooling version control deployment automation and modern release workflows.
- Strong understanding of Linux systems networking cloud security monitoring logging and operational troubleshooting.
- Experience establishing observability choosing and deploying the stack deciding what warrants an alert and tuning alert noise down not only consuming dashboards someone else built.
- Experience supporting live production systems where uptime reliability and performance matter and that cannot tolerate extended downtime.
- Experience with relational databases (MariaDB MySQL Postgres) replication backup and verified restore and schema change against systems that stay online.
- Security-aware mindset with practical experience in secrets management credential rotation access control vulnerability reduction and least-privilege practices.
How you work
- Self-starting.You identify the work rather than wait for it to be assigned and you can justify why you picked one problem over another.
- Proactive about toil.You notice repetitive work and treat it as a defect to be engineered away not a cost of doing business.
- Inventive but pragmatic.You reach for novel approaches where they genuinely help and recognize when a cron job and forty lines of Python is the better answer.
- Strong problem-solving skills and the ability to investigate complex infrastructure or production issues including silent failures and data pipeline outages.
- Comfort working directly with engineers to improve build deploy and operational workflows.
- Practical ownership mindset with the ability to prioritize execute and close loops.
- Experience supporting live games virtual worlds multiplayer systems or other real-time online products.
- Experience with GitHub Actions specifically or similar CI/CD platforms.
- Experience with Datadog Grafana Prometheus CloudWatch ELK OpenTelemetry or similar observability tooling.
- Experience with Redis Memcached queues workers or event-driven systems.
- Experience with Snowflake data warehouse connectivity ETL monitoring or data pipeline reliability.
- Experience with serverless platforms (Netlify Functions Vercel AWS Lambda) and multi-platform hosting environments.
- Experience with disaster recovery backup strategies incident management load testing and performance tuning.
- Experience improving cloud cost management tagging resource optimization or infrastructure governance.
- Experience evaluating where automation shouldnotbe applied cases where you deliberately kept a human in the loop and why.
- Experience working in small high-leverage engineering teams where infrastructure ownership is broad and hands-on.
The ideal candidate is a senior infrastructure engineer who keeps live systems stable while systematically removing the manual work involved in keeping them that way.
They are not tool-collectors. They understand uptime production risk cloud cost security release quality and operational discipline and they treat automation as an engineering discipline with its own failure modes not as a way to look modern. They know that automation acting on production needs guardrails observability and a rollback path and they build those first.
They work independently. Given a mature live game and a broad remit they can identify what matters sequence it sensibly and improve systems incrementally without disrupting what is already working. They are comfortable being the person who decides what gets automated next.
This role suits someone who wants deep ownership of production infrastructure and a genuine mandate to reinvent how a live-service game is operated.
Benefits
Compensation range: $140000 to $180000 CAD determined based on experience technical depth infrastructure ownership breadth agentic automation experience and overall fit.
- Group Retirement Savings Plan matching and participation.
- Comprehensive benefits package including health dental and vision coverage.
- Health and Wellness spending account.
- Generous time off policies.
- Opportunity to support long-running live-service games with established player communities.
- Deep ownership of cloud modernization DevOps automation security improvement and agentic infrastructure workflows.
- A high-impact role with meaningful ownership over reliability performance and engineering operations.
Big Viking Games is committed to creating an inclusive and accessible environment for all candidates. We welcome applications from individuals of all abilities and will provide accommodations throughout the hiring process as needed.
If you require accommodation during the hiring process please contact so we can work with you to support your needs.
Required Experience:
Unclear Seniority
About Company
Founded in 2011, Big Viking Games is the largest independent mobile and social game studio in Canada and a pioneer in mobile HTML5 games. Headquartered in London, Ontario, the company has grown profitably to a team of over 80 Vikings across the globe. From our beginnings with hits lik ... View more